Vasukai

AI crawlers · Common Crawl · Model training

CCBot: Common Crawl's crawler for AI training

CCBot is a crawler from Common Crawl that collects pages for an open web archive used to train AI models. It follows robots.txt, but blocking it is a common and reasonable choice.

Is CCBot allowed on your site?

Or try: · · ·

At a glance

robots.txt nameCCBot
CompanyCommon Crawl
KindModel training. The open web archive many AI models are trained on
Visits sitesYes
User agentCCBot/2.0 (https://commoncrawl.org/faq/)
Official pagecommoncrawl.org

CCBot gathers content for the Common Crawl archive, which many AI models use as training data. If your site is crawled, your pages may become part of that archive. For businesses, this means your content could help train AI systems without direct benefit. Blocking CCBot does not affect AI search crawlers, so you can still appear in AI search results.

You can allow or block CCBot in robots.txt. To block it for the whole site, add a group with 'User-agent: CCBot' and 'Disallow: /'. To allow it, use 'Allow: /' or no rule. Remember that robots.txt is a request, not a lock; a firewall can enforce blocks. Blocking CCBot is a common choice and does not stop AI search crawlers.

robots.txt lines

Block CCBot
User-agent: CCBot
Disallow: /
Allow CCBot
User-agent: CCBot
Allow: /

Questions

What is CCBot?

CCBot is a crawler from Common Crawl that collects pages for an open web archive used to train AI models.

Should I block CCBot?

Blocking CCBot is a common, reasonable choice and does not stop AI search crawlers. It depends on whether you want your content in the Common Crawl archive.

How do I block CCBot in robots.txt?

Add these two lines to robots.txt: User-agent: CCBot and Disallow: /

Does blocking CCBot affect AI search or Google?

No, blocking CCBot does not affect AI search crawlers or Google. It only stops CCBot from accessing your site.