Cloudflare is taking a stand against AI website scrapers

Cloudflare has released a new free tool that prevents AI companies' bots from scraping its clients' websites for content to train large language models. The cloud service provider is making this tool available to its entire customer base, including those on free plans. "This feature will automatically be updated over time as we see new fingerprints of offending bots we identify as widely scraping the web for model training," the company said.

In a blog post announcing this update, Cloudflare's team also shared some data about how its clients are responding to the boom of bots that scrape content to train generative AI models. According to the company's internal data, 85.2 percent of customers have chosen to block even the AI bots that properly identify themselves from accessing their sites.

Cloudflare also identified the most active bots from the past year. The Bytedance-owned Bytespider bot attempted to access 40 percent of websites under Cloudflare's purview, and OpenAI's GPTBot tried on 35 percent. They were half of the top four AI bot crawlers by number of requests on Cloudflare's network, along with Amazonbot and ClaudeBot.

It's proving very difficult to fully and consistently block AI bots from accessing content. The arms race to build models faster has led to instances of companies skirting or outright breaking the existing rules around blocking scrapers. Perplexity AI was recently accused of scraping websites without the required permissions. But having a backend company at the scale of Cloudflare getting serious about trying to put the kibosh on this behavior could lead to some results.

"We fear that some AI companies intent on circumventing rules to access content will persistently adapt to evade bot detection," the company said. "We will continue to keep watch and add more bot blocks to our AI Scrapers and Crawlers rule and evolve our machine learning models to help keep the Internet a place where content creators can thrive and keep full control over which models their content is used to train or run inference on."

This article originally appeared on Engadget at https://www.engadget.com/cloudflare-is-taking-a-stand-against-ai-website-scrapers-220030471.html?src=rss https://www.engadget.com/cloudflare-is-taking-a-stand-against-ai-website-scrapers-220030471.html?src=rss
Erstellt 3mo | 03.07.2024, 22:10:09


Melden Sie sich an, um einen Kommentar hinzuzufügen

Andere Beiträge in dieser Gruppe

The best robot vacuums on a budget for 2024

If vacuuming is your least favorite chore, employing a robot vacuum can save you time and stress wh

03.10.2024, 10:20:09 | Engadget
Epic will extend its free games program to its mobile store

Until now, the mobile version of the Epic Games Store has mostly been focused on the brand’s staples like Fortnite and Fall Guys. It won’t be that way for long.

Epic Games

02.10.2024, 22:40:15 | Engadget
Tesla has stopped selling its cheapest car

Tesla's least expensive car is off the market: the Model 3 Standard Range Rear-Wheel Drive is no longer available in the online configurator.

02.10.2024, 22:40:14 | Engadget
The creepy Crow Country is coming to Nintendo Switch on October 16

One of the year’s scariest and most engrossing horror games is clawing its way to a new console. SFB Games’ Crow Country will launch on the Nintendo Switch

02.10.2024, 22:40:12 | Engadget
ChatGPT added 50 million weekly users in just two months

It's little wonder that investors were clamoring to plow money into OpenAI. Alongside an announcement that the company had

02.10.2024, 20:20:33 | Engadget
More ads are coming to Amazon Prime Video

Can you hear the soft, cherubic voices of corporate executives singing in unison? That can only mean one thing. They’ve figured out a new way to squeeze money out of our eyeballs. Amazon is adding

02.10.2024, 20:20:32 | Engadget