Block a site from being crawled by Common Crawl Crawler.
https://commoncrawl.org/
Robots.txt · AI Bot
Get access to data on 4,882,308 websites that are Common Crawl Bot Disallow Customers. We know of 3,838,613 live websites using Common Crawl Bot Disallow and an additional 1,043,695 sites that used Common Crawl Bot Disallow historically and 2,622,467 websites in the United States.
Get a list of 4,882,308 websites using Common Crawl Bot Disallow which includes location information, hosting data, contact details, 3,838,613 currently live websites and an additional 1,859,681 domains that redirect to sites in this list. 1,043,695 sites that used this technology previouslyand 2,622,467 websites in the United States currently using Common Crawl Bot Disallow.
Countries
Financial
Group
Region