Noindex Tag vs. Robots.txt Disallow: What's the Difference and Why It Matters
Two directives that sound interchangeable but aren't — and why picking the wrong one is a quiet, common mistake.
These two get confused constantly, including by people who work on SEO regularly, because they both sound like they mean "keep this out of Google." They don't do the same job, and mixing them up is a genuinely common way pages end up in search results when they shouldn't be, or disappear from search results when they shouldn't.
What robots.txt disallow actually does
A Disallow rule in robots.txt tells well-behaved crawlers not to request a URL at all. It's a crawling instruction, not an indexing one. That distinction matters more than it sounds: if Google already knows a URL exists — usually because something else on the web links to it — it can still show that URL in search results, typically with no title or snippet, just the bare link, because it was never allowed to actually fetch and read the page.
What a noindex directive actually does
A noindex directive — whether it's a <meta name="robots" content="noindex"> tag in the HTML, or the equivalent X-Robots-Tag: noindex HTTP header — is the opposite kind of instruction. It explicitly tells Google: you're allowed to crawl this page, but don't put it in the index at all. This is the correct tool whenever the actual goal is "this page should never appear in search results," full stop.
The mistake this causes, in both directions
Block a URL with robots.txt when what you actually wanted was noindex, and you can end up with an ugly, snippet-less listing in search results instead of no listing at all — because Google never got to crawl the page and see the noindex tag that would have told it to drop the URL cleanly. Do the reverse — noindex a page but forget it's also disallowed in robots.txt — and Google may never even recrawl it to notice the noindex tag has been removed later, so a page you want back in the index can stay invisible far longer than expected.
A concrete example
Say a site has a /preview/ path used for internal drafts. Disallowing /preview/ in robots.txt stops Google from crawling those pages — reasonable for saving crawl budget on content nobody should see. But if a preview URL gets linked from somewhere external (a shared draft link, an old backlink), it can still surface in search with no snippet. Adding a noindex tag directly on those pages, in addition to or instead of the robots.txt rule, actually keeps them out of the index rather than just out of the crawl.
Which one should you use?
As a simple rule: use noindex when the goal is "never show this specific page in search results," and use robots.txt disallow when the goal is "don't waste crawl budget on this section of the site" (large faceted-navigation URL spaces, internal search results pages, and similar). They're not mutually exclusive, but they solve different problems, and reaching for the wrong one is one of the most common — and hardest to notice — ways a page ends up in the wrong indexing state.
DeindexAlert checks robots.txt, meta robots, X-Robots-Tag, and canonical signals on your monitored pages on a recurring schedule, and emails you the moment something changes -- instead of finding out weeks later from a traffic graph.