engadget – “Cloudflare has released a new free tool that prevents AI companies’ bots from scraping its clients’ websites for content to train large language models. The cloud service provider is making this tool available to its entire customer base, including those on free plans. “This feature will automatically be updated over time as we see new fingerprints of offending bots we identify as widely scraping the web for model training,” the company said. In a blog post announcing this update, Cloudflare’s team also shared some data about how its clients are responding to the boom of bots that scrape content to train generative AI models. According to the company’s internal data, 85.2 percent of customers have chosen to block even the AI bots that properly identify themselves from accessing their sites. Cloudflare also identified the most active bots from the past year. The Bytedance-owned Bytespider bot attempted to access 40 percent of websites under Cloudflare’s purview, and OpenAI’s GPTBot tried on 35 percent. They were half of the top four AI bot crawlers by number of requests on Cloudflare’s network, along with Amazonbot and ClaudeBot…”
Sorry, comments are closed for this post.