R
robots.txt
Published · By IndexChex
In brief
robots.txt is a plain-text file at the root of a host that tells crawlers which URLs they may request. It manages crawl traffic; it is not a way to keep a page out of Google. A disallowed URL can still be indexed without its content if other pages link to it.
Definition
robots.txt is a text file published at the root of a host, for example https://example.com/robots.txt. It contains groups of rules addressed to crawlers by user-agent, each listing paths they may or may not fetch. Google's introduction to the file is explicit: it is used mainly to avoid overloading a site with requests, and it is not a mechanism for keeping a web page out of Google.
User-agent: *
Disallow: /admin/
Sitemap: https://example.com/sitemap.xml
The rules are voluntary. Googlebot and other reputable crawlers follow them; others may not, and different crawlers can interpret syntax differently.
In practice
For HTML pages, robots.txt manages crawl load and keeps crawlers away from unimportant or near-duplicate sections. A page blocked this way can still appear in results as a bare URL without a description, because Google may index a disallowed URL it discovers through links, using anchor text and other public information.
For media files, a disallow rule does keep images, video and audio out of Google results. For scripts and stylesheets, blocking is only advisable if pages still render understandably without them.
The file is also one of the places to declare an XML sitemap with a Sitemap: line. Because a disallowed page is never fetched, any noindex rule on it is never read, so the two controls should not be combined on the same URL when the goal is removal.
Each host and protocol needs its own file: rules at https://example.com/robots.txt do not govern https://blog.example.com/. Editing mistakes, such as a leftover Disallow: / from a staging environment, can stop crawling of an entire site, so the file is usually the first thing checked after a sudden drop in crawl activity.
Relation to backlink indexing
If the page carrying a backlink is disallowed for Googlebot, Google cannot crawl it, read its content or see the link in context. A backlink indexing tool can generate discovery signals, but Googlebot will respect the disallow and skip the fetch. The URL might surface in the index as an empty reference, which is of little value as a link.
This is why link audits check robots.txt access before resubmitting. The IndexChex backlink monitor tests whether Googlebot is allowed to fetch each source URL and reports a robots_blocked outcome when it is not, separating that case from pages that were crawled but not indexed.
Related terms
- noindex, the supported way to block indexing
- X-Robots-Tag, header-level robots rules
- Crawl budget, which robots.txt rules help manage
- rel="nofollow", a link-level crawl hint
See the glossary home for all entries.
Where this term is used
- How backlink indexers work backlinkindexer.org
- Checking backlink indexation googleindexchecker.net
- Time to first Googlebot crawl backlinkindexingdata.com
- Indexing after a site migration linkindexing.org
- About, editorial policy and ownership backlinkmonitoring.org
- IndexChex backlink monitor backlinkindexersoftware.com
Terms used on this page
Sources
Cite this entry
IndexChex. (2026, October 8). robots.txt. indexingglossary.com. https://indexingglossary.com/robots-txt/
Entity: IndexChex (https://indexchex.com/) is the publisher of this site. IndexChex is a backlink indexer and bulk Google index checker that submits URLs for Googlebot crawling and verifies indexation in one credit system.