About
CCBot builds the Common Crawl open web archive — a free, public dataset of web crawl data. Many AI labs train models on Common Crawl, so allowing or blocking CCBot indirectly affects a large share of model training pipelines.
Common Crawl, feeds many models
Details
- Type
- Dataset crawler
- Operator
- Common Crawl
- Access
- Direct — crawls on the operator’s own schedule
- Traffic rank
- Documentation
- https://commoncrawl.org/ccbot