SEO
Indexability: whether your page is allowed into the index at all
Which technical settings decide whether Google takes a page in, how noindex, canonicals and redirects cancel each other out, and how to check it yourself in a few minutes.
11 min read
By Timo Wessels Published
A page can only appear in search results if it is in the index — and whether it is allowed in is decided not by its content but by a handful of technical settings: the status code, redirects, the robots meta tag, the X-Robots-Tag in the HTTP header, the canonical tag and robots.txt. Each of them is set in a different place, often by a different plugin or a different person. That is why they can contradict each other. The result is a familiar finding: the page is online, it looks good, it works — and for Google it does not exist.
The signals that decide
| Signal | Where it lives | What it says |
|---|---|---|
| Status code | the server's response | 200 = here is the page, 301/308 = moved permanently, 404 = does not exist |
| Redirect | server, rarely the HTML | where the address points and whether that is final |
| Robots meta tag | the page's <head> |
noindex = keep out of search results |
X-Robots-Tag |
HTTP header | the same as the meta tag, also for PDFs and images |
| Canonical | <head> or HTTP header |
which of several addresses is the authoritative one |
robots.txt |
the root of the domain | what may be fetched at all |
Each signal looks harmless on its own. It gets interesting when two contradict each other — say an index in the meta tag and a noindex in the header. Google has a clear rule for that case: where robots rules conflict, the more restrictive one applies. The page is out.
robots.txt prevents fetching, not indexing
This is the distinction behind the most stubborn errors. Google states explicitly that robots.txt is not a mechanism for keeping a page out of Google. It controls which addresses a crawler may fetch — mainly so the server is not overloaded.
A blocked page can still end up in the index if other websites link to it. It then appears with its address and the link text, but without a description. In Search Console it is listed as "Indexed, though blocked by robots.txt".
Worse: because Google does not fetch the blocked page, it does not see the noindex on it either. Google names this as a condition: for noindex to work, the page must not be blocked by robots.txt. Anyone who blocks a page in order to get rid of it achieves the opposite.
The rule that follows:
- If a page should stay out of the index: set
noindexand allow fetching. - If a page should not be fetched at all because it creates server load — internal search results with arbitrary parameters, for example:
robots.txt. - Never both at once for the same page.
Google is cautious about CSS, JavaScript and image files: blocking them is only harmless if the page does not look significantly different without them. Whatever Google needs to understand the page has to stay fetchable.
noindex: the most expensive single error
The most common finding in this field is also the most annoying: an important page is set to noindex by accident. This is how it happens:
- After a relaunch. On the staging site, "Discourage search engines from indexing this site" was ticked, and the setting came along with the move. Since WordPress 5.3, that box puts a
noindex,nofollowmeta tag on every page. Google names exactly this point in its site-move guidance: removenoindextags androbots.txtblocks that were only needed during development. - Through a plugin that sets
noindexacross an entire content type. - Through a header set by a host or a server rule, which nobody sees in the HTML.
Where noindex belongs: internal search result pages, login and account areas, basket and checkout, thank-you pages after a form is sent, thin archive and tag pages with no value of their own.
noindex on its own is enough. Following the links on a page is the default. Adding follow changes nothing.
And a principle that is often missing: not every page belongs in the index. Google itself says not to expect every address to be indexed. The goal is for the canonical version of every important page to be in the index — not as many pages as possible.
Canonical: which address counts
The canonical tag tells Google which address is the authoritative one when the same or very similar content can be reached at several addresses. Duplicate addresses appear faster than you would think:
- through sorting and filtering functions and campaign parameters in the address
- through
http://andhttps:// - through
www.and without - through the trailing slash
- through staging sites that are accidentally public
Worth knowing: the canonical is a hint, not a rule. Google chooses the authoritative address from several signals — redirects, the canonical, the sitemap, HTTPS — and may decide differently from what you specified. The more signals point the same way, the more reliably Google follows you.
Two classes of error
The target is wrong. The canonical points to the wrong page. The classic case is multi-page content: page 2 points to page 1. Google advises explicitly against this — every page in a sequence gets its own canonical. Otherwise the content from page 2 onwards drops out of the index.
The form is broken. The canonical is written as a relative path, sits in the body instead of the <head>, or appears twice. Google only accepts it in the <head> (or in the HTTP header). If a page carries several canonicals with different targets, Google's own blog says it ignores all of them. So a formally broken canonical can point to the right page and still have no effect. A frequent cause: the theme and the SEO plugin each set one.
The rules
- Exactly one canonical per page.
- In the
<head>. - Written as an absolute address, with protocol and domain.
- Pointing to itself — unless the page really is a variant of another.
- The target must be reachable and indexable: no error page, no
noindex. noindexis no substitute for a canonical. Google advises against handling duplicate addresses withnoindexorrobots.txt, and against specifying different targets for one page through different methods.
Redirects and chains
301 and 308 are permanent. For Google, a strong signal that the target should become the authoritative address. That is the standard case for every page that has moved.
302, 303 and 307 are temporary. Google follows them but does not treat them as a signal that the target is the authoritative address. They are only right for things that really are temporary.
Chains grow almost by themselves: A redirects to B, B to C. Every time an address changes without the old rule being cleaned up, the chain gets longer. Googlebot follows up to 10 hops. Google still recommends redirecting straight to the final destination, and where that is not possible, keeping the chain short — ideally no more than 3 and fewer than 5 hops. Every hop costs visitors loading time.
The fix is always the same: collapse the chain into a single hop. A points directly to C. Loops — A to B, B back to A — make the page unreachable for everyone and show up in Search Console as a redirect error.
And the point that decides a relaunch: internal links should point directly to the final, canonical address, not to one that redirects. Google writes that linking consistently to the canonical address helps it understand your preference.
Redirects in a relaunch
- Every old address needs a target, and it should be the closest one in subject.
- Do not send everything to the home page. Google warns that redirecting many old addresses to a single irrelevant target can be treated as a "soft 404" — in other words, like an error page.
- Leave redirects in place for as long as possible, according to Google generally at least one year.
Address variants: trailing slash, www and https
/services and /services/ are two different addresses for Google. Only directly after the host name is the slash meaningless: https://example.com and https://example.com/ are the same. Everywhere else in the path it is part of the address.
The same goes for www. and without, http:// and https://. Where versions are equivalent, Google prefers HTTPS by itself — but only as long as no other signals contradict it.
The fix: decide on one variant once, enforce it on the server with a 301, and bring canonicals, sitemap and internal links in line. Which variant you choose matters less than having only one. The easiest is usually the one your system already produces.
Sitemap and canonical have to say the same thing
The sitemap is a weak but additional signal for the authoritative address. That is why only canonical, indexable addresses belong in it: no redirects, no error pages, nothing with noindex, nothing with a canonical pointing elsewhere. An address with noindex in the sitemap is a direct contradiction. How to check the sitemap yourself is covered in the article on robots.txt and sitemaps.
How to check it yourself
The URL Inspection tool in Search Console. The most direct route. Enter the address at the top and you get the answer: "URL is on Google" or "URL is not on Google". Below it you find "Crawl allowed?", "Indexing allowed?" and two canonical fields: "User-declared canonical" and "Google-selected canonical". If the two differ, Google has not followed your hint. After a fix, "Test live URL" shows the current state.
The Page indexing report. It lists, for the whole site, the reasons why addresses are not indexed — for example "Excluded by 'noindex' tag", "Alternate page with proper canonical tag", "Duplicate, Google chose different canonical than user", "Page with redirect", "Redirect error" or "Crawled - currently not indexed".
A look at the source. Ctrl+U, then Ctrl+F. Search for name="robots" and for rel="canonical". Each may appear at most once, and the canonical belongs in the <head>.
The HTTP headers. This is the case nobody finds by looking at the page: the meta tag says index, the X-Robots-Tag in the header says noindex — and the more restrictive rule wins. You can see it in the developer tools under Network: reload the page, click the first row, read the response headers.
Redirect chains. Also in the Network tab: if the first row shows 301 or 302 and another redirect follows, you have a chain.
The variant test. Open your home page and one subpage: with www. and without, with http:// and https://, with a trailing slash and without. All of them must land on the same address, in a single hop. It takes a minute.
For the full picture you need a crawling tool that goes through the whole site and reports indexability, canonicals, redirects and status codes per address. With 300 pages, nobody checks them one by one.
What to do
- Check the home page and the five most important subpages with the URL Inspection tool. If anything there is not indexed, it takes priority over everything else.
- Clean up
noindex. It belongs on a few clearly nameable pages — and nowhere else. Check the headers too. - Make sure every page has exactly one canonical, absolute, in the
<head>, pointing to an indexable address. - Merge the address variants into one, with a 301.
- Collapse redirect chains and fix internal links that run through a redirect.
- Check
robots.txtfor anything blocked that Google needs to render the page — or a page that should really getnoindexinstead. - After every relaunch, first thing: did the staging setting come along? It is the most expensive error in this field and the easiest to avoid.
Sources
- Google Search Central, Introduction to robots.txt — not a way to keep pages out of Google, blocked pages indexed without a description, CSS and JavaScript: developers.google.com
- Google Search Central, Block search indexing with noindex — meta tag and
X-Robots-Tag, no blocking in robots.txt: developers.google.com - Google Search Central, Robots meta tag and X-Robots-Tag — default values, the more restrictive rule wins in a conflict: developers.google.com
- Google Search Central, What is canonicalization — causes of duplicate addresses, the canonical as a hint: developers.google.com
- Google Search Central, How to specify a canonical URL — signal strength, absolute addresses, only in the
<head>, no noindex, HTTPS, internal links: developers.google.com - Google Search Central Blog, 5 common mistakes with rel=canonical — multiple canonicals are ignored, target without noindex, page 2 not to page 1: developers.google.com
- Google Search Central, Pagination — every page with its own canonical: developers.google.com
- Google Search Central, Redirects and Google Search — 301/308 permanent, 302/303/307 temporary: developers.google.com
- Google Search Central, HTTP status codes and network errors — up to 10 redirect hops: developers.google.com
- Google Search Central, Site moves with URL changes — keep chains short, at least one year, no mass redirects to the home page, remove noindex after the move: developers.google.com
- Google Search Central Blog, To slash or not to slash — trailing slash, the exception after the host name: developers.google.com
- Google Search Console Help, Page indexing report — reasons, not every address needs to be indexed: support.google.com
- Google Search Console Help, URL Inspection tool — crawl allowed, indexing allowed, user-declared and Google-selected canonical: support.google.com
- WordPress Core, Changes to prevent search engines indexing sites — WordPress 5.3 sets
noindex,nofollowinstead ofDisallow: /: make.wordpress.org