Cloudflare is taking a stand against AI website scrapers

Cloudflare has released a new free tool that prevents AI companies' bots from scraping its clients' websites for content to train large language models. The cloud service provider is making this tool available to its entire customer base, including those on free plans. "This feature will automatically be updated over time as we see new fingerprints of offending bots we identify as widely scraping the web for model training," the company said.

In a blog post announcing this update, Cloudflare's team also shared some data about how its clients are responding to the boom of bots that scrape content to train generative AI models. According to the company's internal data, 85.2 percent of customers have chosen to block even the AI bots that properly identify themselves from accessing their sites.

Cloudflare also identified the most active bots from the past year. The Bytedance-owned Bytespider bot attempted to access 40 percent of websites under Cloudflare's purview, and OpenAI's GPTBot tried on 35 percent. They were half of the top four AI bot crawlers by number of requests on Cloudflare's network, along with Amazonbot and ClaudeBot.

It's proving very difficult to fully and consistently block AI bots from accessing content. The arms race to build models faster has led to instances of companies skirting or outright breaking the existing rules around blocking scrapers. Perplexity AI was recently accused of scraping websites without the required permissions. But having a backend company at the scale of Cloudflare getting serious about trying to put the kibosh on this behavior could lead to some results.

"We fear that some AI companies intent on circumventing rules to access content will persistently adapt to evade bot detection," the company said. "We will continue to keep watch and add more bot blocks to our AI Scrapers and Crawlers rule and evolve our machine learning models to help keep the Internet a place where content creators can thrive and keep full control over which models their content is used to train or run inference on."

This article originally appeared on Engadget at https://www.engadget.com/cloudflare-is-taking-a-stand-against-ai-website-scrapers-220030471.html?src=rss https://www.engadget.com/cloudflare-is-taking-a-stand-against-ai-website-scrapers-220030471.html?src=rss
Created 1y | Jul 3, 2024, 10:10:09 PM


Login to add comment

Other posts in this group

The Cult of the Lamb comic is coming back with the Schism Special this fall

We're officially getting more of the Cult of the Lamb comic expansion. Following last year's miniseries, which built on the game's existing lore and injected some real emotional depth, wri

Jul 12, 2025, 9:40:13 PM | Engadget
Grok team apologizes for the chatbot's 'horrific behavior' and blames 'MechaHitler' on a bad update

The team behind Grok has issued a rare apology and explanation of what went wrong after X's chatbot began

Jul 12, 2025, 7:30:05 PM | Engadget
Nintendo reportedly bans Switch 2 user playing preowned game cards

You might have to be extra careful who you buy your used Nintendo Switch game cards from if you don't want to get mistakenly banned. A Nintendo Switch 2 owner

Jul 12, 2025, 7:30:04 PM | Engadget
Meta reportedly closes deal to buy AI voice replicator PlayAI

Meta has finalized the agreement to purchase Play AI

Jul 12, 2025, 5:10:12 PM | Engadget
This HDMI mod lets you play Nintendo Switch Lite on a big screen

If you can't get your hands on the latest

Jul 12, 2025, 5:10:11 PM | Engadget
The best Prime Day Apple deals on iPads, AirPods, MacBooks and more still available today

There’s a reason Apple gear is so in demand. After reviewing nearly every major device out there, our current favorite

Jul 12, 2025, 2:40:19 PM | Engadget
The best Amazon Prime Day deals under $50 that you can still get today

Big ticket items like TVs and iPads might get the lion’s share of the attention during Amazon Prime Day, but you can often find affordable tech on sale for even less, too. Despite the sale being ov

Jul 12, 2025, 2:40:18 PM | Engadget