Getting Indexed

Robots.txt or Noindex? Which One to Use for Tag Pages, Thank-You Pages, and Staging Sites

You have a handful of pages you would rather nobody found — the thank-you page after a form, a hundred near-empty tag archives, a staging copy of the site. Two tools get recommended for the job, and they do different things.

The real difference in one line: robots.txt asks crawlers not to fetch a page, while a noindex directive tells search engines not to show it in results. If your goal is "this must not appear in Google," use noindex and let the crawler in to read it. If your goal is "don't waste crawling on this," use robots.txt. Doing both to the same page is the classic error, because a crawler that is not allowed to fetch the page can never see the noindex sitting inside it — which is exactly how a URL you blocked ends up listed anyway, with no description under it.

What robots.txt actually controls

robots.txt is a plain text file at the root of your domain — yourdomain.com/robots.txt — that lists which paths well-behaved crawlers should not request. It is an instruction about fetching, and nothing else.

Two consequences follow, and they are the ones people are surprised by:

  • A blocked URL can still end up in search results. If other pages link to it, a search engine knows the URL exists even though it was never allowed to look at it. It may list the address with no title or description because it has nothing to describe. Blocking hid the content, not the URL.
  • It is public and advisory. Anyone can read your robots.txt, and listing a secret path there advertises it. Well-behaved crawlers respect it; nothing else has to. It is a traffic-management tool, never a security control.

What it is genuinely good at is keeping crawlers away from things there is no point fetching: internal search result pages, endless filter and sort combinations, cart and checkout paths, session-parameter URLs. These are pages that generate near-infinite variations, and stopping the fetching stops the waste at the source.

What a noindex directive actually controls

noindex is a directive on the page itself — a <meta name="robots" content="noindex"> tag in the <head>, or an X-Robots-Tag header for files like PDFs where you cannot add a meta tag. It says: fetch this if you like, but keep it out of the results.

Because it lives inside the response, the page must be crawlable for the directive to work. That is the whole trap. A page that is blocked in robots.txt and also carries noindex is a page whose noindex will never be read. If a URL is already indexed and you want it gone, the sequence is: remove any robots.txt block on it, add noindex, let it be re-crawled, and wait. Removal is not instant — it happens on the crawler's schedule, and low-value pages are crawled infrequently, which is precisely why the pages you most want removed are the slowest to go.

So why is my page still in Google after I blocked it?

Because blocking and removing are different requests, and you made the first one.

Work through it in this order:

  1. Check whether the URL is blocked, indexed, or both. In Search Console, inspect the live URL. "Indexed, though blocked by robots.txt" is the fingerprint of this exact mistake.
  2. Remove the block for that path in robots.txt.
  3. Add noindex to the page — meta tag for HTML, X-Robots-Tag header for files.
  4. Confirm a crawler can see it. View the page source and find the tag. If the tag is injected by JavaScript after load, treat it as unreliable and put it in the served HTML instead.
  5. Wait for the re-crawl, and check status again rather than repeating the change. If a page needs to be gone urgently, use Search Console's removal request as a stopgap — it hides the result temporarily while the noindex does the permanent work.

The same logic runs in reverse for pages you want found: if an important page is not being indexed, the first two things to rule out are a leftover block and a leftover noindex, which is the ten-minute check at the front of what to do when your site is live and invisible.

Which one to use, page type by page type

This is where the general rule stops helping and the specifics start. For the pages small sites actually have:

Tag and category archives. noindex the ones that are thin or duplicative — a tag with two posts competes with the posts themselves for nothing. Keep the archives that are genuinely useful entry points, with a real description and enough entries to be worth landing on. Do not block them in robots.txt: they carry links to your real content, and you want those followed.

Thank-you and confirmation pages. noindex. They are useless to a searcher and occasionally embarrassing when they surface. Leave crawling open so the directive is read.

Internal search result pages. Block the search path in robots.txt, because every query creates another URL and there is no end to them. If a few are already indexed, unblock, noindex, let them drop, then block again once they are gone.

Filter, sort and session-parameter URLs. Same reasoning as internal search: the problem is volume, so stop the fetching. If particular filtered pages are worth having in search — a genuinely popular category-plus-city combination, say — treat those as real pages with their own links rather than as parameters.

Staging and development sites. Neither, on their own. Put the whole environment behind HTTP authentication or an IP restriction. A password prompt cannot be indexed, cannot be linked into search results, and cannot be forgotten in the way a stray directive can. The commonest cause of a site disappearing after launch is a Disallow: / or a site-wide noindex that came across from staging.

PDFs and other files. Use the X-Robots-Tag header where you need to exclude one, since there is no HTML head to put a meta tag in. Most site owners should be doing the opposite here — making useful documents findable rather than hiding them.

Login, account and admin areas. Block in robots.txt to save the fetching, and protect them with actual authentication. The block is housekeeping; the login is the control.

Paginated archive pages. Usually leave them alone. Blocking or noindexing page two onward cuts the crawl path to older posts, which costs you more than the tidiness gains.

When the choice genuinely does not matter

Plenty of the time, either answer is fine and the deliberation costs more than the outcome:

  • Nothing links to the page and nothing ever did. An unlinked page with no traffic is not being crawled anyway. Whichever directive you pick, the practical result is identical.
  • The page is already absent from the index. Confirm with a live URL inspection before spending an afternoon excluding something that was never included.
  • It is a single low-stakes page. One thank-you page will not change how your site is crawled. Pick noindex, move on, and save the analysis for the patterns that generate hundreds of URLs.

The decision earns its time when it applies to a template — anything that produces pages by the dozen. One-offs rarely repay the thinking.

Check what you actually did

Verify from outside your own assumptions, not from the settings screen of a plugin:

  • Fetch yourdomain.com/robots.txt in a browser and read it. This is the file crawlers see, whatever your CMS claims it has configured.
  • View the source of the page, not the rendered view, and search for noindex.
  • Inspect the live URL in Search Console, which reports coverage plainly: indexed, excluded by noindex, blocked by robots.txt, or the contradiction of being indexed despite a block.
  • Make sure your sitemap agrees with both. A URL listed in the sitemap while carrying noindex sends contradictory instructions, and the contradiction is resolved in ways you do not control — the consistency rule is covered in the XML sitemap guide.
  • Re-check after a change, because both directives take effect on a re-crawl rather than on save.

FAQ

It depends entirely on how often that URL is crawled, which for a low-value page can be slow — often days to a few weeks, sometimes longer. Nothing legitimately guarantees a timeframe. A removal request in Search Console hides the result in the meantime.

Should I use robots.txt and noindex together for extra certainty?

No. It is the one combination that reliably fails, because the block prevents the directive from being read. Choose the one that matches your goal: crawling, or results.

Does robots.txt keep private information private?

Not at all. The file is publicly readable, listing a path in it draws attention to that path, and only well-behaved crawlers honour it. Use real authentication for anything that matters.

Links on a noindexed page can still be followed, so the page can act as a crawl path even while it stays out of results. Adding nofollow on top of noindex is what stops that — and it is rarely what you want on your own site.

Will noindexing thin pages help my other pages rank?

Sometimes, indirectly: fewer near-duplicate pages competing for the same query makes it clearer which page you meant to be the answer. Treat it as tidying rather than as a ranking tactic, and expect the gains to show up as clarity, not as a jump. The wider picture of what gets a page indexed and kept is in the getting-indexed guide.


Decide per page whether you are managing crawling or managing results, use one tool for each, and verify with a live inspection instead of memory. Then put the effort where it pays: making the pages you do want found easy to discover.

Once the right pages are the ones left open, get them in front of crawlers and people — submit your URL to the directory and see what fired for your submission.

Comments are disabled for this article.