Vasukai

Research · 11 October 2026

Which AI crawlers do the top websites block?

We read the robots.txt of 6,226 of the most visited websites (the world’s top 5,000, plus the top UK and India sites) on 11 October 2026, and checked each of 18 AI crawlers against it with the same rules our free checker uses.

13.3%of the world’s top sites block CCBot, the most blocked AI crawler
9.8%block at least one AI search crawler (OAI-SearchBot, Claude-SearchBot or PerplexityBot)
5.7%block OpenAI’s training crawler but let its search crawler in
73.3%have a robots.txt file at all

Share of sites blocking each AI crawler

Blocked for the home page in robots.txt, as a share of the sites we reached. Firewalls were not tested, so the true share turned away is higher.

CrawlerKindWorld top 5,000Top UK (.uk) sitesTop India (.in) sites
OAI-SearchBotOpenAIAI search7%9.5%3.8%
Claude-SearchBotAnthropicAI search7%6.5%4%
PerplexityBotPerplexityAI search9.5%10.2%3.9%
GooglebotGoogleAI search2.8%1.4%2.9%
BingbotMicrosoftAI search3.1%1.6%3.2%
DuckAssistBotDuckDuckGoAI search7.4%6%4.3%
ChatGPT-UserOpenAIFetch on request8.4%6.8%4.3%
Claude-UserAnthropicFetch on request7.1%6.4%4.3%
Perplexity-UserPerplexityFetch on request7.2%5.3%4.3%
MistralAI-UserMistralFetch on request7.1%6.1%4.2%
GPTBotOpenAIModel training12.6%13.8%5.8%
ClaudeBotAnthropicModel training11.9%13.1%5.5%
Google-ExtendedGoogleModel training10.6%8.4%4.9%
Applebot-ExtendedAppleModel training10.3%11.9%5.1%
Meta-ExternalAgentMetaModel training11%12.1%5%
AmazonbotAmazonModel training9.7%9.8%5.5%
BytespiderByteDanceModel training12.9%11.6%6.6%
CCBotCommon CrawlModel training13.3%12.5%5.9%

By country

World top 5,000Top UK (.uk) sitesTop India (.in) sites
Sites reached3,796 of 5,0001,257 of 1,5001,173 of 1,500
Have a robots.txt73.3%75.8%58.2%
Block at least one AI search crawler9.8%11%4.2%
Block GPTBot, ClaudeBot, Google-Extended and CCBot (all four)8.6%5.5%4.6%
Block at least one AI crawler17%18.5%7.8%

How we did it

Sites: the Tranco top 1M list, taking the top 5,000 worldwide, and the highest-ranked 1,500 .uk and 1,500 .in domains. Each robots.txt was fetched on 11 October 2026 over HTTPS (trying www. if the bare domain failed), and every crawler was checked against the rules for the home page using RFC 9309: the most specific user-agent group, the longest matching path, and allow winning a tie. A robots.txt that answered with a server error counts as blocking everything, as the standard says.

Limits: robots.txt rules for the home page only; firewalls not tested. Sites that didn’t answer within a few seconds are left out of the percentages. Domain lists include CDNs and API hosts that are not websites people visit.

Data: download the CSV (one row per site, free to use with a link to this page).

Check your own site

Score out of 100, crawler access and fixes in about 10 seconds. No sign-up. Nothing stored.

Readable isn’t the same as recommended

The free AI check shows whether ChatGPT names your business and who it names instead. No card needed.

Run my free AI check