X-Robots-Tag Explained: The HTTP-Header Noindex Nobody Checks
The noindex directive that lives in an HTTP header instead of the page HTML — and why that makes it so easy to miss.
Most people checking whether a page is indexable look at the page's HTML — right-click, View Source, search for "noindex." That habit misses an entire category of deindexing, because X-Robots-Tag doesn't live in the HTML at all. It's sent as an HTTP response header, invisible to a normal page-source check, and it carries exactly the same weight with Google as a meta robots tag.
What it looks like
Instead of a tag inside <head>, it's a line in the server's response, something like:
X-Robots-Tag: noindex, nofollow
It supports the same directives as the meta robots tag — noindex, nofollow, noarchive, and so on — and Google treats them identically. The only real difference is where the instruction is sent from.
Why it exists at all
The meta robots tag only works on HTML pages, because it has to live inside an HTML <head>. X-Robots-Tag was built to cover everything else — PDFs, images, and other non-HTML files that Google can still index but that have no HTML head to put a meta tag in. It's also genuinely useful on HTML pages when you want to control indexing at the server or CDN level, without touching the page template at all — which is exactly what makes it dangerous.
How it causes silent, hard-to-diagnose deindexing
Because it's set at the infrastructure layer rather than in page content, it's commonly introduced by:
A CDN rule meant for one path pattern that ends up matching more URLs than intended · A reverse proxy or load balancer config change made for an unrelated reason · A server framework's default header behavior on certain response types · A staging environment's server config accidentally carried into production
In every one of these cases, someone can check the page's HTML, find nothing wrong, and still have a fully noindexed page — because the directive was never in the HTML to begin with.
How to actually check for it
You need to look at response headers, not page source. The simplest way is a terminal command:
curl -I https://example.com/the-page
and read the response headers for an X-Robots-Tag line. In a browser, the Network tab of the developer tools shows the same thing — click the page's main document request and check the response headers. A plain "View Source" will never show this, which is exactly why it's worth checking deliberately rather than assuming it would show up in a normal review.
The takeaway
Any indexability check that only looks at page HTML has a blind spot. A page can pass every visual and source-code inspection and still be fully excluded from Google's index because of a single header set somewhere in the infrastructure layer — which is also exactly why this class of issue tends to go unnoticed the longest.
DeindexAlert checks robots.txt, meta robots, X-Robots-Tag, and canonical signals on your monitored pages on a recurring schedule, and emails you the moment something changes -- instead of finding out weeks later from a traffic graph.