DeindexAlert DeindexAlert
Blog

Why Did Google Deindex My Page? 7 Causes and How to Check Each One

Seven concrete causes of deindexing, in the order worth checking them, with exactly how to confirm each one.

A page can vanish from Google's index without a single error message, a broken link, or anything visibly wrong when you load it in a browser. The traffic just stops. If that's happened to you, the cause is almost always one of a small number of things — here they are, in the order worth checking them.

1. A noindex tag, on purpose or by accident

Look at the page's HTML <head> for <meta name="robots" content="noindex">. This is the single most common cause, and the most common way it happens by accident is a staging/development noindex tag that was meant to be removed before launch but shipped to production anyway — often because it lives in a template or environment config, not the page content itself, so it's easy to miss in a normal content review.

2. The same directive, sent as an HTTP header instead

A noindex can also arrive via the X-Robots-Tag HTTP response header rather than HTML — which means it won't show up if you only "view source" in a browser, since that only shows the page markup, not the response headers. Check with your browser's network inspector, or curl -I the URL and look for X-Robots-Tag: noindex in the response. This is especially common after a CDN or reverse-proxy rule change, since header rules are configured separately from the page template and are easy to add without realizing they apply to more URLs than intended.

3. A canonical tag pointing somewhere else

Check the <link rel="canonical"> tag in the page's head. If it points at a different URL, Google will typically treat that other URL as the authoritative version and drop this one from the index — even if the two pages have genuinely different content and that clearly wasn't the intent. This happens most often after a site migration, a templating bug that hardcodes one canonical value across many pages, or a staging environment's canonical URLs leaking into production.

4. robots.txt blocking the crawl entirely

This one's subtly different from the first two: a Disallow rule in robots.txt stops Google from crawling the page at all, rather than telling it not to index a page it can already see. In practice a disallowed page can sometimes still appear in search results (usually with no snippet, just the bare URL) if Google finds links to it elsewhere — so this cause is a little less predictable than a direct noindex, but still worth ruling out by checking yourdomain.com/robots.txt for a rule that matches the page's path.

5. A manual action or security issue

Less common, but far more serious: Google Search Console's "Manual actions" and "Security issues" reports flag cases where Google has taken direct action against the site — for violating spam policies, or because the site was compromised and started serving malware or spam to visitors. Both sections are free to check and worth ruling out early, since the fix here is entirely different from the technical causes above.

6. The page was recrawled and judged low-value or duplicate

Google doesn't only remove pages because of an explicit directive — it can also quietly drop a page from the index if a later recrawl decides the content is thin, duplicated elsewhere on the site (or across the web), or no longer worth serving. This is harder to fix with a single technical change and usually means the page's content itself needs to improve, or genuinely duplicate pages need to be consolidated with canonical tags or redirects.

7. A DNS, hosting, or redirect problem

If Google's crawler can no longer reach the page at all — a DNS misconfiguration, an expired SSL certificate, a hosting outage, or a redirect chain that breaks — repeated failed crawl attempts will eventually cause the page to fall out of the index the same way a manual noindex would, just for infrastructure reasons instead of a directive. Search Console's crawl stats and the URL Inspection tool will usually show this clearly if it's happening.

How to check quickly, without guessing

The fastest single check is Search Console's URL Inspection tool — paste in the exact URL and it will tell you the live indexing status and, in most cases, the specific reason a page is excluded. A site:yourdomain.com/the-exact-path search in Google is a rougher but instant sanity check. Neither one tells you the moment something changes, though — both are a snapshot, not a monitor, which is the gap a recurring automated check is meant to close.

Catch this automatically

DeindexAlert checks robots.txt, meta robots, X-Robots-Tag, and canonical signals on your monitored pages on a recurring schedule, and emails you the moment something changes -- instead of finding out weeks later from a traffic graph.

Monitor your pages

Related reading