SEO
robots.txt and sitemap: the two files you use to talk to crawlers
What robots.txt may block and what it must not, why a block does not keep a page out of the index, what belongs in the sitemap and how to check both files yourself in a few minutes.
10 min read
By Timo Wessels Published
robots.txt and the XML sitemap are two files meant only for machines. robots.txt tells crawlers which addresses they should not fetch. The sitemap tells them the opposite: which addresses you consider important. No visitor ever sees them — and yet they are regularly the place where an otherwise clean website loses visibility: through a block that blocks too much, or a sitemap that says something different from the pages themselves.
What robots.txt does — and what it does not
robots.txt controls which addresses a crawler may fetch. Google names its main purpose as keeping the site from being overloaded with requests. It is explicitly not a mechanism for keeping a page out of Google.
Three misunderstandings follow from that:
- Blocked does not mean invisible. A blocked page can still end up in the index if other websites link to it — with its address, but without a description.
- Blocked does not mean protected. The standard itself says the rules are not a form of access authorisation. Not every crawler obeys them. Anything confidential belongs behind a password.
- Blocking and
noindexrule each other out. Google does not fetch a blocked page and so never sees thenoindexon it. To get something out of the index, setnoindexand allow fetching. Google has ignorednoindexinsiderobots.txtitself since 1 September 2019.
What makes sense to block: the admin area, internal search results and filter combinations that generate endless addresses. In short, everything that keeps crawlers busy without their finding anything useful.
What does not belong in robots.txt: the staging site. A blocked staging site can still show up in the index. WordPress recommends the header X-Robots-Tag: noindex, nofollow on every file for development sites, and Google names password protection as a way of keeping content out of the index. Both together is safest.
CSS and JavaScript
The most expensive mistake in robots.txt is often not a blocked page but a blocked resource. Google can only do without CSS, JavaScript and image files if the page does not look significantly different without them. Whatever Google needs to render and understand the page has to stay fetchable.
In WordPress that means in practice: lines like Disallow: /wp-content/ or Disallow: /wp-includes/ have no place in robots.txt. That is where stylesheets, scripts and images live.
The rules Google reads the file by
- One file per host and protocol. The
robots.txtathttps://example.com/does not apply tohttps://www.example.com/or tohttp://example.com/. It has to sit in the root directory. - Case matters in paths:
/Contactand/contactare two rules. - The most specific rule wins. If
AllowandDisallowboth match an address, the longer, more specific one applies. If both are equally specific, the less restrictive one applies. - No matching rule means allowed. If nothing in the file addresses a crawler, it may fetch everything.
- 500 KiB at most. Google does not read anything beyond that.
- Google caches the file for up to 24 hours. Changes do not take effect at once.
- If the file is missing (status 404), there are no restrictions. If the server answers with a 5xx error, Google first stops crawling for 12 hours and then falls back on the last cached version for up to 30 days.
- Google does not support
crawl-delay. Some other search engines do. - The
Sitemap:line needs a full address with protocol and domain.
A lean robots.txt for a WordPress site looks like this — which is exactly what WordPress generates by itself:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://example.com/wp-sitemap.xml
robots.txt and AI crawlers
robots.txt now also decides which AI systems read your content. The vendors separate their crawlers by purpose, and each one is addressed by its own name:
| Crawler | Vendor | Purpose |
|---|---|---|
GPTBot |
OpenAI | material for training the models |
OAI-SearchBot |
OpenAI | the search feature in ChatGPT |
Google-Extended |
training and answers in Gemini | |
Googlebot |
Google Search, including AI Overviews |
Two places where it often goes wrong:
- OpenAI states that sites blocking
OAI-SearchBotwill not appear in ChatGPT's search answers. To prevent only training, blockGPTBotand allowOAI-SearchBot. - According to Google,
Google-Extendedhas no effect on whether a page appears in Google Search and is not a ranking signal. AI Overviews in Search depend on the regularGooglebot. BlockingGooglebotremoves a site from search and from AI Overviews at once — which is rarely the intention.
If your robots.txt says nothing about an AI crawler, it may fetch everything. That is a decision too, even if nobody made it deliberately. Which crawlers you need for AI answers is covered in the article on AI visibility.
The sitemap: when you need one
A sitemap is a list of the addresses you consider important. Google names three cases in which it helps most:
- large sites — Google speaks of more than about 500 pages, because there it is harder to make sure every page is linked,
- new sites with few links from elsewhere,
- sites with a lot of video, images or news.
A small, well-linked site can do without one. It does no harm there either, and WordPress generates one anyway. It is never a guarantee: Google states explicitly that not every address in a sitemap will be crawled and indexed.
What belongs in the sitemap
Only canonical, indexable addresses with value of their own. No redirects, no error pages, nothing with noindex, nothing with a canonical pointing to another page. For Google, the sitemap is an additional signal of which address is the authoritative one — if it contradicts the pages, it weakens exactly that signal.
The most common finding is therefore not a missing sitemap but a contradiction: the sitemap lists addresses that are set to noindex or redirect. The site then says "here is an important page" and "please do not index it" at the same time. Or the other way round: pages that really exist are missing because they belong to a content type the SEO plugin does not pick up.
The formal rules:
- at most 50,000 addresses and 50 MB uncompressed per file. Beyond that, split it and connect the files through an index file.
- full, absolute addresses, with protocol and domain.
- UTF-8 as the character encoding.
lastmodonly if it is true. Google uses the modification date only when it is consistently and verifiably accurate — meaning significant changes to the content, not the year in the footer.- Google ignores
priorityandchangefreq. They are not worth maintaining.
WordPress: robots.txt and sitemap
WordPress generates robots.txt itself as long as there is no real file in the root directory. You extend it through the robots_txt filter or the SEO plugin. If a physical file is there, it is served directly and WordPress never gets involved. Changes in the plugin then have no effect — a typical cause of "but I changed it".
WordPress has shipped its own sitemap since version 5.5, at /wp-sitemap.xml. It is an index file with sub-files of at most 2,000 entries by default, and WordPress adds it to robots.txt automatically. SEO plugins usually replace it with their own. If "Discourage search engines from indexing this site" is ticked, WordPress switches the sitemap off.
How to check it yourself
Look at robots.txt. Open yourdomain.com/robots.txt in the browser. The file should be short and readable. Look for three things:
- Is there a
Disallow: /without any further restriction? That blocks the whole site. - Do
wp-content,wp-includes,.cssor.jsappear in aDisallowline? - Is there a
Sitemap:line, and does the address in it work?
Open the sitemap. Go to the address it names. An error message or an empty file means the entry is out of date. Spot check: open five addresses from the sitemap — if one redirects or carries noindex, the sitemap has a problem.
Search Console. The Sitemaps report shows whether the file was submitted and read without errors, and how many addresses Google found in it. In the Page indexing report you can filter by that sitemap and see how many of those addresses are actually indexed. A large gap between the two numbers is the interesting finding, and the reasons are listed right next to it — for example "Blocked by robots.txt" or "Indexed, though blocked by robots.txt".
The AI crawlers. Search your robots.txt for GPTBot, OAI-SearchBot and Google-Extended, and for the crawlers of other vendors listed in the article on AI visibility. If nothing is there, they are allowed.
The service in front. If the site runs behind a CDN or a firewall, a crawler can be turned away there even though robots.txt allows it. A look at its settings is part of the check.
What to do
- Check whether a physical
robots.txtexists, and decide: either maintain the file or delete it and let WordPress or the SEO plugin take over. Not both. - Remove blocks on CSS, JavaScript and images as far as the page needs them to render. Keep the admin area, internal search and filter parameters blocked.
- Put nothing in
robots.txtthat should really leave the index. That is whatnoindexis for. - Add the sitemap to
robots.txtand submit it in Search Console as well. - Make sure sitemap and pages say the same thing: anything set to
noindex, redirecting or carrying another page's canonical does not belong in it. - Make a deliberate decision about AI crawlers. For most service businesses it means: allow the search crawlers, handle training according to your own stance. To appear in AI answers, you have to be readable.
- After every relaunch: look at
robots.txtand the sitemap first. A block left over from staging or a sitemap full of old addresses will otherwise only surface weeks later.
Sources
- Google Search Central, Introduction to robots.txt — purpose, not a way to keep pages out of Google, blocked pages without a description, CSS and JavaScript: developers.google.com
- Google Search Central, How Google interprets the robots.txt specification — host and protocol, 500 KiB, 24 hours, 4xx and 5xx, rule precedence, case sensitivity,
crawl-delay, theSitemap:line: developers.google.com - IETF, RFC 9309 Robots Exclusion Protocol — no matching rule means allowed, not access authorisation: rfc-editor.org
- Google Search Central Blog, A note on unsupported rules in robots.txt —
noindexin robots.txt ignored since 1 September 2019, password protection: developers.google.com - Google Search Central, Block search indexing with noindex — noindex only works without a robots.txt block: developers.google.com
- WordPress Core, Changes to prevent search engines indexing sites —
X-Robots-Tag: noindex, nofollowfor development sites: make.wordpress.org - OpenAI, Overview of OpenAI Crawlers — GPTBot for training, OAI-SearchBot for search in ChatGPT: developers.openai.com
- Google, Overview of Google's common crawlers — Google-Extended has no effect on Google Search: developers.google.com
- Google Search Central, AI features and your website — Googlebot and robots.txt as the control for Search: developers.google.com
- Google Search Central, What is a sitemap — about 500 pages, new sites, media, no guarantee: developers.google.com
- Google Search Central, Build and submit a sitemap — 50,000 addresses and 50 MB, absolute addresses, UTF-8,
lastmod,priorityandchangefreq: developers.google.com - Google Search Central, How to specify a canonical URL — the sitemap as a weak canonical signal: developers.google.com
- Google Search Console Help, Sitemaps report — status, discovered addresses, filtering the Page indexing report: support.google.com
- Google Search Console Help, Page indexing report — reasons for addresses not being indexed: support.google.com
- WordPress Developer Resources,
do_robots()— default content of the robots.txt WordPress generates: developer.wordpress.org - WordPress.org Plugin Directory, Virtual Robots.txt — a physical robots.txt prevents the one WordPress generates: wordpress.org
- WordPress Core, New XML Sitemaps Functionality in WordPress 5.5 —
/wp-sitemap.xml, 2,000 entries, entry in robots.txt: make.wordpress.org