Skip to content
Technical SEO16 min readUpdated 5 Oct 2026

Robots.txt Blocked but Still Indexed: Why It Happens and How to Fix It

Why a URL blocked in robots.txt can still appear in Google, why blocking plus noindex fails, and how to keep pages out of results properly.

Read time
16 min read
Sections
20
FAQs answered
15
Topic
Technical SEO

A page blocked in robots.txt can still appear in Google Search. Robots.txt tells crawlers which pages they may fetch. It does not tell Google to keep a page out of its index. If other sites link to a blocked URL, Google can index the address without ever reading the page. To keep a page out of results, let Google crawl it and add a noindex rule, or protect it with a password.

This is one of the most misunderstood rules in technical SEO, and it causes real problems: private pages that show up in results, staging sites that get indexed, and "blocked" pages that rank with no description. This guide explains exactly what robots.txt does, why blocked URLs still appear, how to fix each situation and how to write a safe robots.txt. Everything below is based on Google's own documentation, which we link to. For the wider checklist, see our technical SEO checklist.

What robots.txt does

Robots.txt is a plain text file at the root of your site, such as yoursite.com/robots.txt. It contains rules about which paths crawlers may request. Google's documentation says a robots.txt file is used primarily to manage crawler traffic to your site. For web pages, that means you can use it if your server might be overwhelmed by requests, or to avoid crawling unimportant or similar pages.

Three limits are worth remembering.

  • It is a request, not a lock. Google says Googlebot and other respectable crawlers obey it, but other crawlers might not. It cannot enforce behaviour. If information is private, use password protection, not robots.txt.
  • Crawlers interpret syntax differently. Rules written for one search engine may be read differently by another.
  • It controls crawling, not indexing. This is the one that surprises people.

Why a blocked URL can still be indexed

Google's documentation says that a page disallowed in robots.txt can still be indexed if it is linked from other sites. Google will not crawl or index the content of the blocked page, but it might still find and index the URL, and the address and potentially other public information, such as the anchor text of links pointing to it, can appear in search results.

Here is a simple example. You block /private-offer/ in robots.txt. A partner site links to it with the words "exclusive trade offer". Google sees the link, knows the page exists and cannot fetch it. It may still list the address in results, often without a description, because it was never allowed to read the page. Google's documentation says the result will not have a description when the page is blocked this way.

The same thing explains a common Search Console message: "Indexed, though blocked by robots.txt". It does not mean Google read your content. It means Google knows the URL exists and has indexed the address.

The noindex trap: why blocking and noindexing together fails

The correct tool for keeping a page out of search results is a noindex rule, delivered as a meta tag or an HTTP header. But there is a catch that catches many sites. Google's documentation says that for noindex to work, the page must not be blocked by robots.txt and must be otherwise accessible to the crawler. If the page is blocked, the crawler never sees the noindex rule, and the page can still appear in results if other pages link to it.

SetupWhat happens
Blocked in robots.txt, no noindexGoogle cannot read the page but may index the bare URL if linked
Blocked in robots.txt, with noindex on the pageGoogle cannot see the noindex, so the URL can still be indexed
Not blocked, with noindexGoogle crawls the page, reads noindex and removes it from results
Password-protectedGoogle cannot access it and it will not be indexed from its content

Google also states that specifying noindex inside robots.txt is not supported. Older advice that suggested it no longer works. See Google's documentation on blocking indexing with noindex and on robots.txt.

Which method to use for which goal

GoalUseNotes
Keep a page out of Google resultsnoindex meta tag or X-Robots-Tag header, with crawling allowedWait for Google to recrawl. Use the Removals tool for urgent cases
Keep a page privatePassword protection or authenticationrobots.txt is not security
Reduce crawling of low-value URLs, such as filter combinationsrobots.txt disallow, plus sensible internal linkingDo not expect it to deindex them
Stop a server being overloadedrobots.txt, and fix server capacityCrawl rate is mostly handled by Google automatically
Remove an old page permanently404 or 410 status, or a 301 redirect to a replacementBlocking it in robots.txt would stop Google seeing the removal
Control how much text Google showsnosnippet, max-snippet, data-nosnippetThese also limit how content can appear in AI features

How to fix a page that is "indexed, though blocked by robots.txt"

Decide first what you want.

If you want the page to rank normally

  1. Remove the robots.txt rule that blocks it.
  2. Check the page is not set to noindex.
  3. Use the URL Inspection tool to test the live URL, then request indexing.
  4. Wait for Google to crawl and process the page.

If you want the page out of Google

  1. Remove the robots.txt block, so Google can read the page.
  2. Add a noindex meta tag or an X-Robots-Tag: noindex header.
  3. Confirm in URL Inspection that Google sees the noindex.
  4. Once the page has dropped out, you can decide whether to block crawling again. Many sites leave it open, because the noindex then keeps working.
  5. For urgent removals, use the Removals tool in Search Console, which hides a URL temporarily while you apply a permanent fix.

If the page is genuinely private

Put it behind a login. Do not rely on robots.txt or noindex. A password is the only reliable barrier.

Writing a safe robots.txt

A typical, safe file for a business site is short.

User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /internal-search/

Sitemap: https://www.example.com/sitemap.xml

That example blocks three areas that should not be crawled and points to the sitemap. Every site is different, so adjust it to your own structure. Rules for good practice:

  • Be specific. A rule like Disallow: / blocks the whole site.
  • Do not block CSS or JavaScript files that pages need to render. Google says that if the absence of resources makes a page harder to understand, you should not block them.
  • Test changes before and after you publish them. Search Console shows how Google reads your file.
  • Keep a copy of the previous version, so you can restore it quickly.
  • Use one robots.txt per host. Subdomains need their own.
  • Add your sitemap address so crawlers can find it.

The most common robots.txt disasters

  • The staging block goes live. A development site uses Disallow: / to stay out of search. The file is copied to the live site at launch, and the whole site disappears from crawling.
  • A plugin or setting adds a block. A "discourage search engines" option left switched on after launch.
  • Important resources blocked. Blocking script or style folders can leave pages looking broken to Google.
  • Blocking pages you want removed. Pages stay in the index because Google cannot see the noindex or the 404.
  • Wildcard mistakes. A pattern that matches more URLs than intended, such as blocking every URL that contains a certain word.

When you launch or relaunch a site, check robots.txt on day one, then check Search Console's indexing report a few days later.

Robots.txt and AI crawlers

Robots.txt is also where you set policy for AI crawlers. OpenAI documents that OAI-SearchBot is for ChatGPT search and GPTBot is for training, as independent settings. Perplexity documents PerplexityBot for surfacing sites in results. Google controls access to its AI features in Search through normal Googlebot access and snippet controls, and offers a separate control, Google-Extended, for some other uses. Be deliberate about each, and remember that user-initiated fetchers, such as Perplexity-User and ChatGPT-User, may not follow robots.txt. We cover the details in our guides to ChatGPT search and Perplexity.

How to audit your robots.txt in ten minutes

  1. Open yoursite.com/robots.txt in a browser.
  2. Look for Disallow: / under a general user agent. If it is there on a live site, that is critical.
  3. Read each rule and ask what it blocks and why.
  4. Test three important URLs to confirm they are not blocked.
  5. In Search Console, open the Page indexing report and check for "Indexed, though blocked by robots.txt" and "Blocked by robots.txt".
  6. Check that your sitemap line points to a working sitemap.
  7. Check subdomains that serve content, such as a blog or shop.

A worked example: a staging site in search results

This is an invented example. A Manchester agency builds a new site for a client on a staging address. The developers block it with Disallow: / in robots.txt, and three months later the client notices the staging address appearing in search results with the title "Home | Staging" and no description.

What happened

Someone shared the staging link in a blog post, and a few sites linked to it. Google found those links, could not crawl the pages because of the block, but still indexed the address and used the anchor text it found on other sites. That matches Google's documentation, which says a disallowed page can still be indexed if linked from other places and that the address and anchor text may appear in results.

The wrong fix

The team adds a noindex meta tag to every staging page. Nothing changes, because Google cannot crawl the pages to see the tag.

The right fix

  1. Put the staging site behind a password, the only reliable barrier for something that must stay private.
  2. For the already-indexed URLs, remove the robots.txt block temporarily so Google can crawl the pages and see a noindex tag, or return a 404 or 410 status if the pages should disappear completely.
  3. Use the Removals tool in Search Console to hide the URLs quickly while the permanent fix takes effect.
  4. Once the pages have dropped out and the password is in place, the robots.txt rule can stay as an extra signal, but it is no longer the control.
  5. Add a launch checklist item: check robots.txt and noindex settings on both staging and the live site on launch day.

How to diagnose an indexing problem, step by step

QuestionHow to checkIf the answer is no
Does the page return a 200 status?URL Inspection, a crawler or browser developer toolsFix server or redirect problems first
Is it allowed to be crawled?Test the URL against robots.txt, and check the Page indexing reportRemove the block if you want it indexed, or accept it is not crawled
Is it set to noindex?View source and response headers, and URL InspectionRemove noindex if you want it in the index
Is there a canonical pointing elsewhere?URL Inspection shows the Google-selected canonicalFix the canonical, or accept that the other page is the preferred one
Is the content worth indexing?Read it honestlyImprove or merge thin or duplicate pages
Is it linked from somewhere crawlable?Crawl the site and check internal linksAdd contextual links

X-Robots-Tag for PDFs and other files

A meta tag only works in HTML pages. For PDFs, images and other files, the noindex rule can be sent as an HTTP response header called X-Robots-Tag. Google's documentation describes both methods and says they have the same effect. Typical uses:

  • Keeping a downloadable price list or internal PDF out of results while allowing visitors with the link to open it.
  • Excluding certain file types site-wide through server configuration.
  • Applying noindex to pages generated by a system where you cannot edit the HTML.

The same rule applies: the file must be crawlable for Google to see the header. If a PDF contains sensitive information, put it behind a login.

Faceted navigation, parameters and crawl waste

Online shops often create thousands of URLs from filters and sort orders, such as colour, size and price. These can waste crawling and create near-duplicates. Choices:

OptionWhen to use itCaution
Allow crawling and indexingThe combination has real search demand and unique value, such as "red running shoes"Needs distinct titles, content and internal links
Allow crawling, add noindexUseful to visitors, but not worth rankingGoogle still crawls them, which uses some crawl capacity
Block crawling in robots.txtHuge numbers of low-value combinations that waste server and crawl effortBlocked URLs can still be indexed if linked, and Google cannot see canonicals or noindex on them
Canonicalise to the main categoryVariations that should consolidate to one pageGoogle treats canonicals as hints, not commands

Pick one approach per type of URL and apply it consistently, then watch the indexing report.

A launch-day checklist for robots and indexing

  • robots.txt on the live domain does not contain a blanket Disallow.
  • No site-wide noindex in templates, headers or a CMS "discourage search engines" setting.
  • The staging site is password protected and not linked from the public web.
  • Redirects from old URLs are in place and tested.
  • The sitemap is live, lists only canonical, indexable URLs and is referenced in robots.txt.
  • Search Console is verified for the live domain, and the sitemap is submitted.
  • URL Inspection shows key pages as indexable.
  • A reminder is set to check the Page indexing report after a week and a month.

Common questions about robots.txt

QuestionAnswer
Do I need a robots.txt file?Not always. If you want everything crawled and have no special needs, a file is optional. Many sites use one to block low-value sections and to point to the sitemap
Can I block images?Google's documentation says robots.txt can prevent image, video and audio files from appearing in Google Search results, though it will not stop other pages linking to them
Does robots.txt affect other search engines?Reputable crawlers generally follow it, but each may interpret syntax differently, so test for the engines that matter to you
How do I block just one crawler?Use a user-agent group naming that crawler with its own disallow rules
How often does Google re-read robots.txt?Google caches it and refreshes it periodically. Changes can take time to take effect
Can I use wildcards?Google supports some pattern characters, but behaviour can differ between crawlers. Test your rules

An indexing decision guide for common page types

Page typeIndex?Method
Core service and product pagesYesAllow crawling, no noindex, include in sitemap
Blog posts and guidesYesAs above
Thank-you and confirmation pagesUsually noNoindex with crawling allowed, or a login
Internal search resultsUsually noBlock crawling or noindex, and avoid linking to them
Cart, checkout and account pagesNoBlock crawling and require login where appropriate
Staging and test sitesNoPassword protection
Printer-friendly or duplicate versionsNoCanonical to the main page or remove
Thin tag or archive pagesUsually noNoindex or remove if they add nothing

Why this confusion persists, and how to explain it to colleagues

The idea that robots.txt hides pages from Google is deeply embedded, and it is easy to see why. The file's name suggests control over robots. The syntax is simple, and the word "Disallow" sounds like "forbid". Many older guides, and some plugins, describe robots.txt as a way to keep pages out of search results. And for many pages, in many situations, the pages do not appear in results, which reinforces the belief. It is only in the less common case, when other sites link to a blocked URL, that the URL shows up, and then the surprise is unpleasant.

A useful way to explain the difference to colleagues is with a library analogy. Robots.txt is like a sign on a door that says "Do not enter". Visitors who respect signs will not go in. But the library catalogue, which is compiled from references in other books, may still list that room, with a call number and a note from other catalogue entries. The catalogue does not need to enter the room to know that it exists. A noindex tag, by contrast, is like a label inside the room that says "Do not list this room in the catalogue". For that to work, the cataloguer must be allowed in to read the label. If the door is locked to the cataloguer, they will never see the label, and the room can stay in the catalogue.

That analogy captures the key points. Robots.txt controls whether crawlers may enter. Noindex controls whether a page should be listed. They operate at different layers, and they interact in a way that surprises people: blocking entry prevents the crawler from reading the instruction to leave the page out of the listing. Google's documentation says exactly this, in more technical terms.

When you brief developers, marketers or clients, a short rule helps. Use robots.txt to manage crawling, for example to stop bots wasting time on endless filter pages. Use noindex to keep pages out of results. Use a password to keep content private. Never use robots.txt as a privacy tool. And when something is wrongly indexed, first ask: can Google crawl the page and see our instruction? If the answer is no, fix that before anything else. Teams that learn this rule tend to avoid most of the indexing mistakes that cost time and traffic.

A final practical reminder

When in doubt, test with a single, harmless page first. Create a test URL, apply the rule you are considering, use the URL Inspection tool to see what Google sees and wait for the Page indexing report to reflect it. Learning the behaviour on one page is far cheaper than discovering it across thousands.

Where we can help

We are a digital marketing agency in Manchester, UK and Mumbai, India. If you suspect robots.txt or indexing settings are hurting your site, our technical SEO services include a crawl, an indexing review and a plain list of fixes. Contact us to talk it through.

FAQ

Your questions, answered in plain English

Yes. Google says a disallowed page can still be indexed if it is linked from other sites. The URL and, potentially, anchor text from links may appear in results, usually without a description.

Google found the URL through links, could not crawl the page because robots.txt disallows it, and indexed the address anyway. It did not read your content.

Allow Google to crawl the page and add a noindex meta tag or X-Robots-Tag header, or put it behind a password. Do not block it in robots.txt, because Google then cannot see the noindex.

Google must crawl the page to read the noindex rule. If robots.txt blocks crawling, Google never sees it, so the page can still appear in results.

No. Google states that specifying the noindex rule in the robots.txt file is not supported. Use a meta tag or HTTP header instead.

No. It is a request that reputable crawlers follow, and it cannot enforce behaviour. Use password protection or authentication for information that must stay private.

Google says it is used primarily to manage crawler traffic, such as avoiding server overload or skipping unimportant or similar pages. It is not a way to hide pages from search results.

Generally no. Google says that if missing resources make a page harder for its crawler to understand, you should not block them, or it may not analyse pages that depend on them well.

Use the Removals tool in Search Console for a temporary hide, and apply a permanent fix such as noindex, a 404 or 410 status or a redirect. Do not block it in robots.txt.

It takes effect after Google next crawls the page and processes the rule, which can take from days to weeks depending on how often the page is crawled. Requesting indexing in URL Inspection can help.

Better to use password protection, since robots.txt does not guarantee exclusion. If you do use a block, make sure it is removed from the live site at launch.

Yes. A stray Disallow rule can stop important pages being crawled, and blocking resources can stop pages rendering properly. Check robots.txt at every launch and after changes.

Documented search and training crawlers from OpenAI and Perplexity do, but user-initiated fetchers such as ChatGPT-User and Perplexity-User may not, because a person requested the fetch.

At the root of the host, for example example.com/robots.txt. Each subdomain needs its own file.

Open the file in a browser, test key URLs against it with a robots.txt testing tool, and check Search Console for indexing messages about blocked pages. Test again after every change.

Still curious? Send us your question and a strategist will get back to you.

#Robots.txt#Noindex#Indexing#Technical SEO#Googlebot
KwiqRank Team

Written by KwiqRank Team

KwiqRank is a digital marketing agency in Manchester, UK and Mumbai, India, working on SEO, GEO, AEO, paid media and websites. We check facts against official sources, and tell you when something is unconfirmed. Spotted a mistake? Tell us and we will correct it.

Found this useful?

Want this done for you?

Start with an audit tied to revenue. We will tell you what is worth fixing and what is not.

See SEO services