Get Started

Free AI Training Data Checker

Enter a domain to see whether it appears in Common Crawl — the free, open web corpus that shows up again and again in the published training-data descriptions of large language models. The tool queries the last five monthly snapshots, shows real captured URLs, and checks whether CCBot is blocked from adding you to the next one.

Queries the last five monthly Common Crawl snapshots. This can take up to a minute — the public index is slow.

What Common Crawl actually tells you

Common Crawl is a non-profit that has been publishing a monthly open crawl of the web since 2008. Anyone can download it, and because it is the largest freely available web corpus, it has become a standard ingredient in language-model training sets and in the derived datasets built on top of it.

That makes presence in it a useful floor, and a poor ceiling. It is good evidence that your site is reachable, crawlable and worth including. It is not evidence that any assistant will mention you, because the answers users get today are mostly assembled by live retrieval from each vendor's own index — not from a training corpus captured months ago.

What this tool can and cannot tell you

  • It can tell you whether your hostname appears in each of the last five published snapshots, roughly how many URLs were captured, examples of which ones, and whether your robots.txt currently blocks CCBot.
  • It cannot tell you which models trained on your content. The major labs do not publish their training mixes, and no external tool can resolve that.
  • It will not guess. When the public index times out — which it does regularly — the result is reported as Unknown rather than as a negative.

How to use this tool

  1. Enter your domain — the bare hostname is fine. If you get nothing back, the tool automatically checks the www variant, because they are indexed separately.
  2. Read the trend, not just the latest row — dropping out of recent snapshots after appearing in older ones usually means a robots.txt change or a site migration.
  3. Check the CCBot verdict — if it is blocked, no amount of content work will get you into the next snapshot.
  4. Then check the crawlers that actually matter for citations — use the AI Crawler Access Checker for the answer and indexing bots, which is where AI visibility is really won or lost.

Common issues

  • Checking the wrong hostname: Common Crawl keys entries by host, so example.com and www.example.com give different answers.
  • A CCBot block you inherited: a number of hosting platforms and security products ship an AI-crawler block by default, CCBot included.
  • Reading absence as a verdict about your content: a small or new site can simply not have been reached yet. Common Crawl does not attempt to crawl the whole web every month.
  • Expecting the capture count to be a page count: the tool samples the index rather than paging through all of it, so a capped number is a floor, shown as 25+, not a total.

Frequently Asked Questions

What is Common Crawl?

Common Crawl is a non-profit that has published a free, open crawl of the web roughly every month since 2008. Each snapshot contains billions of pages, and the whole archive is downloadable by anyone. It is the largest openly available web corpus, which is why it turns up repeatedly in the published training-data descriptions of large language models.

Does being in Common Crawl mean ChatGPT will cite me?

No, and this is the most common misunderstanding about it. Common Crawl is a training-data corpus, not a retrieval index. Being in it means your pages could have been part of what a model learned from; it says nothing about whether an assistant will surface you when a user asks a question today. Most modern assistants answer with live retrieval, which uses their own crawlers and indexes. Treat Common Crawl presence as a floor — a sign your site is reachable and worth crawling — not as an AI visibility metric.

Which models are trained on Common Crawl?

Common Crawl appears by name in the published training-data descriptions of many models, and it is the backbone of derived datasets like C4 and The Pile that others train on. But the major commercial labs do not fully disclose their training mixes, so nobody outside them can tell you with certainty which models used which snapshot. Any tool that claims to know exactly which models your page trained is guessing.

My site is not in Common Crawl. What should I do?

Check three things in order. First, whether CCBot is blocked in your robots.txt — this tool tells you, and a block means no future snapshot will include you. Second, whether you checked the right hostname: www.example.com and example.com are indexed as separate entries, so one can be captured while the other is not. Third, whether your pages are reachable at all — server-side rendered, linked from somewhere, and returning a 200.

Should I block CCBot?

That is a content policy decision, not an SEO one, and there is a real case on both sides. Blocking CCBot keeps your content out of an openly redistributable corpus that anyone can train on. It does not stop you being cited in AI answers today, because live answers come from each assistant’s own crawler, not from Common Crawl. If your concern is showing up in AI answers, the crawlers to check are the answer and indexing ones — not CCBot.

Why does the tool sometimes say "Unknown" instead of "Not found"?

Because they are different facts. The public Common Crawl index answers a query for a missing domain with a clean 404, but under load it answers with a gateway timeout instead. A timeout tells you nothing about your site. Reporting it as "not found" would be the tool inventing a result, so those rows say Unknown and you should re-run the check.

Grow your organic traffic from chatbots

Enter your website to track its AI visibility across ChatGPT, Gemini, Claude, and Perplexity — and turn chatbot mentions into traffic.

  • Set up in minutes
  • Cancel anytime