Robots.txt vs noindex: Find and Fix Indexing Blockers
You submitted a page for indexing — and it never shows up in search. Often the problem is not "slow Google": the page is simply blocked, and you are paying for bot visits that cannot achieve anything. This guide covers the four levels of blockers — robots.txt, meta robots, the X-Robots-Tag header, and canonical — how to find them, what they mean, and how to remove them.
Robots.txt: Disallow is not noindex
robots.txt is a file at the site root that search robots read. It controls crawling, not indexing directly. The classic webmaster mistake is confusing the two mechanisms:
Disallow: /page/in robots.txt forbids the robot to visit the page. Googlebot never sees its content — and therefore never sees the meta tags on it.noindexon the page itself forbids including it in the index. The robot visits, reads the directive, and skips indexing.
The key consequence: a page blocked only in robots.txt can still appear in the index — with no content, just the URL. This happens when external sites link to it: Google knows the address, cannot show the content, but the URL lingers in results "by link".
Google's rule of thumb: to remove a page from the index, use noindex and do not block it in robots.txt — otherwise the robot will never see your noindex.
What the problem looks like in robots.txt
Open https://your-site.com/robots.txt and look for:
User-agent: *
Disallow: /
That is a full site-wide block — a common leftover on staging copies that nobody reopened after launch.
Disallow: /page
Blocks every URL starting with /page — including /page-2, /pages/price, and similar. Check whether the pattern accidentally captures sections you need.
Meta robots: noindex on the page
The second level is a tag in the page's <head>:
<meta name="robots" content="noindex, nofollow">
What the values mean:
- noindex — do not include the page in the index;
- nofollow — do not follow the links placed on this page (for crawling from the page);
- none — shorthand for
noindex, nofollow; - index, follow — explicit permission (the default; no need to set it);
- noindex, follow — skip indexing the page, but follow its links.
There are also noimageindex, notranslate, and per-bot values: <meta name="googlebot" content="noindex"> blocks the page for Google only.
Where noindex hides
- CMS and SEO plugin settings: the "Discourage search engines from indexing this site" checkbox in WordPress (Settings → Reading) sets noindex site-wide.
- Maintenance mode and "coming soon" plugins.
- Pages closed "just in case" during launch — and forgotten.
- Automatic template rules: tag archives, internal search results, pagination pages.
X-Robots-Tag: noindex in the HTTP header
The third level is a response header:
X-Robots-Tag: noindex
It works like meta robots but at the server level — and fits files that have no HTML: PDFs, DOCX, images. It is often set in nginx/Apache for whole folders, e.g. /files/*.pdf — and accidentally covers documents you need.
Check headers with any HTTP sniffer or curl -I https://site.com/page/ — look for an X-Robots-Tag line.
Canonical: "this page is not the main one"
The fourth blocker is not a ban but a signal redirect:
<link rel="canonical" href="https://another-address/">
If the canonical points to a different URL, Google treats the page as a copy and indexes not it but the canonical address. The page lands in reports as "Duplicate without user-selected canonical" or "Alternate page with proper canonical tag".
Common mistakes:
- canonical to
http://instead ofhttps://; - canonical with/without trailing slash, with/without
www— must match the real address; - an SEO plugin auto-assigning the canonical to a category;
- after a migration, canonicals still pointing at the old domain.
How to check a page for blockers: step-by-step
- Inspect robots.txt: open
site.com/robots.txt, find Disallow rules matching your URL. Mind rule order and theUser-agent: *group. - View the page source (Ctrl+U) and search for
<meta name="robots"and<link rel="canonical". - Check response headers (
curl -I): the status must be 200 OK, with noX-Robots-Tag: noindexand no unexpected 301/302. - URL Inspection in Search Console: paste the address in the GSC top bar — the tool shows whether crawling is allowed, what canonical Google sees, and any errors.
- The Page indexing report in GSC: the "Excluded by noindex" and "Blocked by robots.txt" clusters list all blocked URLs with reasons.
A quick manual check from search: site:your-site.com/page — if the URL is found, the page is in the index (possibly content-less, as with robots.txt blocks).
How to close pages correctly when you need to
Sometimes blocking is the right move. Typically closed from indexing:
- admin panels and internal sections;
- internal search results;
- carts, personal accounts, checkout steps;
- drafts and staging.
The correct method for HTML pages: noindex in meta robots with robots.txt left open. Blocking service folders in robots.txt is also acceptable — but remember the difference: Disallow stops crawling, while URLs already in the index will not be removed that way.
After the fix: how to get the page reindexed
- Remove the blocker (the robots.txt rule / meta tag / header / canonical).
- Confirm the page returns 200 OK and passes URL Inspection.
- Request indexing in Search Console — one page at a time.
- For bulk fixes (hundreds of URLs after removing noindex) — resubmit the list with bot visits: AGD Index sends Googlebot/Bingbot to every URL and restarts crawling without GSC's daily quotas.
FAQ
What is the difference between noindex and Disallow in robots.txt?
Disallow forbids the robot to visit a URL (crawling); noindex forbids including the page in the index. A page blocked only in robots.txt can linger in results without content; to remove it from the index, use noindex and open crawling.
How do I view any site's robots.txt?
Open https://site.com/robots.txt in a browser. To test rules against a specific URL, use Search Console's robots report or an online robots.txt validator.
What does meta robots content noindex, nofollow mean?
The page will not be indexed (noindex), and the robot will not follow links placed on that page (nofollow). It is the standard way to close a utility page completely.
A page is blocked in robots.txt but still in search results. Why?
Google discovered the URL via external links. Since crawling is forbidden, Google cannot read the noindex on the page and shows the URL "by link". Fix: open the page in robots.txt and block it via meta noindex — after a recrawl the URL will drop out.
How do I remove a WordPress noindex?
Check: Settings → Reading → "Discourage search engines from indexing this site" (must be unchecked), then your SEO plugin settings — the Indexing / Titles & Meta sections where noindex may be enabled for page types. Request a recrawl afterwards.
Do bot visits help against noindex?
No — and that is the honest answer. A bot visit guarantees a crawl, not index inclusion. If a page is blocked by noindex or robots.txt, remove the blocker first, then submit. That is why the checklist in this article comes before any bulk submission.
What to read next
- Pages are open but Google still refuses — 11 reasons Google is not indexing your pages.
- The "Crawled — currently not indexed" status explained: Crawled, currently not indexed.
- Clean URL lists ready for resubmission — bot visit tariffs.
Indexing bot for Google and Bing: pages and backlinks
Built for new websites and bulk backlink lists. We send real Googlebot and Bingbot crawler visits straight to your URLs, so pages get discovered instead of waiting for the next natural crawl.
Open Telegram bot