Web & SEO

Robots.txt

A file at a site's root that gives crawlers access rules for URL paths. It is not a security control and does not reliably prevent an otherwise discoverable URL from being indexed.

How it works

robots.txt controls which paths a crawler may request. A Disallow rule can reduce crawling, but it does not reliably remove a known URL from search. A page-level noindex instruction is different: the crawler must be allowed to fetch the page to read that instruction. Blocking the page and adding noindex can therefore create conflicting intentions.

A realistic example

A preview folder may need protection while the public site must remain crawlable. Accidentally copying a staging rule such as Disallow: / to production prevents crawlers from requesting every path. Blocking script or stylesheet folders can also prevent a crawler from seeing the layout and content users receive.

What to check

Fetch /robots.txt on the exact public host and check its status and content. Read the rules for the relevant user agent, test representative page and asset paths, and check the sitemap address. For a page that should be excluded from search, inspect its actual meta robots or HTTP robots header as well.

Limits and next steps

This file is public and is not access control. Do not list sensitive filenames as a substitute for authentication. Make narrow, documented changes and test public content afterward. Search engines do not all interpret optional directives identically, so avoid assuming an unsupported rule will achieve a specific indexing outcome.

Technical sources

← All glossary terms