Research · 11 October 2026
Which AI crawlers do the top websites block?
We read the robots.txt of 6,226 of the most visited websites (the world’s top 5,000, plus the top UK and India sites) on 11 October 2026, and checked each of 18 AI crawlers against it with the same rules our free checker uses.
Share of sites blocking each AI crawler
Blocked for the home page in robots.txt, as a share of the sites we reached. Firewalls were not tested, so the true share turned away is higher.
| Crawler | Kind | World top 5,000 | Top UK (.uk) sites | Top India (.in) sites |
|---|---|---|---|---|
| OAI-SearchBotOpenAI | AI search | 7% | 9.5% | 3.8% |
| Claude-SearchBotAnthropic | AI search | 7% | 6.5% | 4% |
| PerplexityBotPerplexity | AI search | 9.5% | 10.2% | 3.9% |
| GooglebotGoogle | AI search | 2.8% | 1.4% | 2.9% |
| BingbotMicrosoft | AI search | 3.1% | 1.6% | 3.2% |
| DuckAssistBotDuckDuckGo | AI search | 7.4% | 6% | 4.3% |
| ChatGPT-UserOpenAI | Fetch on request | 8.4% | 6.8% | 4.3% |
| Claude-UserAnthropic | Fetch on request | 7.1% | 6.4% | 4.3% |
| Perplexity-UserPerplexity | Fetch on request | 7.2% | 5.3% | 4.3% |
| MistralAI-UserMistral | Fetch on request | 7.1% | 6.1% | 4.2% |
| GPTBotOpenAI | Model training | 12.6% | 13.8% | 5.8% |
| ClaudeBotAnthropic | Model training | 11.9% | 13.1% | 5.5% |
| Google-ExtendedGoogle | Model training | 10.6% | 8.4% | 4.9% |
| Applebot-ExtendedApple | Model training | 10.3% | 11.9% | 5.1% |
| Meta-ExternalAgentMeta | Model training | 11% | 12.1% | 5% |
| AmazonbotAmazon | Model training | 9.7% | 9.8% | 5.5% |
| BytespiderByteDance | Model training | 12.9% | 11.6% | 6.6% |
| CCBotCommon Crawl | Model training | 13.3% | 12.5% | 5.9% |
By country
| World top 5,000 | Top UK (.uk) sites | Top India (.in) sites | |
|---|---|---|---|
| Sites reached | 3,796 of 5,000 | 1,257 of 1,500 | 1,173 of 1,500 |
| Have a robots.txt | 73.3% | 75.8% | 58.2% |
| Block at least one AI search crawler | 9.8% | 11% | 4.2% |
| Block GPTBot, ClaudeBot, Google-Extended and CCBot (all four) | 8.6% | 5.5% | 4.6% |
| Block at least one AI crawler | 17% | 18.5% | 7.8% |
How we did it
Sites: the Tranco top 1M list, taking the top 5,000 worldwide, and the highest-ranked 1,500 .uk and 1,500 .in domains. Each robots.txt was fetched on 11 October 2026 over HTTPS (trying www. if the bare domain failed), and every crawler was checked against the rules for the home page using RFC 9309: the most specific user-agent group, the longest matching path, and allow winning a tie. A robots.txt that answered with a server error counts as blocking everything, as the standard says.
Limits: robots.txt rules for the home page only; firewalls not tested. Sites that didn’t answer within a few seconds are left out of the percentages. Domain lists include CDNs and API hosts that are not websites people visit.
Data: download the CSV (one row per site, free to use with a link to this page).
Check your own site
Readable isn’t the same as recommended
The free AI check shows whether ChatGPT names your business and who it names instead. No card needed.