Vasukai

Free tools · 18 AI crawlers

robots.txt checker for AI crawlers

This checker reads a site's robots.txt file and shows, for each AI crawler, whether the rules let it read a given page. It covers 18 crawlers from OpenAI, Anthropic, Google, Perplexity, Microsoft, Apple, Meta, Amazon, ByteDance, DuckDuckGo, Mistral and Common Crawl.

Or try: · · ·

What AI crawlers do and the three kinds

Crawlers do different jobs. Training crawlers (such as GPTBot, ClaudeBot and CCBot) collect pages to train AI models. Search crawlers (such as OAI-SearchBot, Claude-SearchBot and PerplexityBot) find pages for AI answers that search the web. User crawlers (such as ChatGPT-User and Claude-User) fetch a page when a person asks the assistant about it.

How to read your result and what to change

The checker shows, for each AI crawler, whether the rules let it read a given page. Blocking a training crawler and allowing search crawlers is a common, reasonable choice. Blocking search crawlers usually means the assistant cannot cite or link to the site.

robots.txt versus firewalls

robots.txt is a request, not a lock. Reputable crawlers follow it. A firewall or bot-protection service can block AI crawlers even when robots.txt allows them. Some hosting and security services offer a setting that blocks AI bots.

Questions

Should I block GPTBot?

Blocking a training crawler and allowing search crawlers is a common, reasonable choice. Blocking search crawlers usually means the assistant cannot cite or link to the site.

Does blocking Google-Extended remove me from AI Overviews?

No. Google AI Overviews and Google AI Mode use Google's normal search index, which Googlebot builds. Google-Extended is a separate robots.txt token: it controls whether content is used for Gemini models, and does not affect Google Search or AI Overviews.

Why does a crawler show allowed but the server blocked it?

A firewall or bot-protection service can block AI crawlers even when robots.txt allows them. The checker also sends a few requests that identify themselves as AI crawlers and reports how the server answered. A real crawler can be treated differently, because some firewalls check where a request comes from.

Is my data stored?

Results are not saved to any account or database. A result may be kept in memory for up to 30 minutes so a repeat check is quick. Every result has its own link, so it can be shared with a colleague or client.

Does being allowed mean AI will recommend me?

No. The checker shows whether AI can read a site. It cannot show whether AI recommends the business. Vasukai's free AI check does that: it asks ChatGPT the questions customers ask and shows who is named instead.