Cloudflare Lets Sites Disallow AI Training Without Blocking Googlebot

Cloudflare Lets Sites Disallow AI Training Without Blocking Googlebot

Cloudflare has introduced a crawler control that lets site owners refuse AI model training without blocking Google, Bing, and Apple search crawlers. The Disallow AI Training setting has been available since September 15 in Bot Management and AI Crawl Control on all plans. The switch writes a training opt-out into robots.txt; search crawlers from Google, Apple and Microsoft keep their access.

The problem the setting targets is what Cloudflare calls mixed-use crawlers, a single bot identifier that feeds both a search index and model training. Under the Search, Training, and Agent categories Cloudflare introduced in July, blocking training also shut out the crawler doing search. Heise online notes the old logic simply applied every rule a bot's roles triggered, so refusing one use refused all of them.

The mechanics blend a published preference with network enforcement. Bot Preference Sync writes the applicable no-training directives into robots.txt, and Cloudflare blocks training-only crawlers, including those run by Amazon, Anthropic, Meta, and OpenAI, upstream in its infrastructure. Mixed-use crawlers are allowed through for search on the expectation that their operators honor the directive. Cloudflare's own blog concedes the limit: a robots.txt file cannot identify a crawler or stop one that ignores it, which is why the company pairs the preference with bot classification and public reporting on Radar.

Whether an operator gets the search carve-out depends on a label Cloudflare invented for the occasion. After talks with crawler operators that started in July, the company created an "Accountable" designation with four requirements: a training opt-out via robots.txt or similar, an opt-out for AI summaries (set with the operator today, through Cloudflare next year), URL-level visibility into which pages were made available for training alongside search metrics, and an assurance that refusing training will not hurt traditional search results. Apple, Google, and Microsoft qualify through a mix of shipped features and dated commitments. Heise points out the designation is Cloudflare's own judgment, not an external certification.

Support is uneven across the three. For Google, the setting maps to a Disallow rule for Google-Extended, which Google's documentation says has no effect on Search inclusion or ranking; appearing in AI Overviews and AI Mode is governed separately in Search Console, and that toggle does not touch training. Apple's path runs through Applebot-Extended. Bing lags: Microsoft is still building robots.txt support for a domain-level no-training preference, targeted for early 2027, so for now the Cloudflare setting sends Bing nothing, and the working opt-out remains the NOARCHIVE meta tag, which per Bing's documentation keeps content out of Microsoft's generative model training and out of links in Chat and Copilot. Google plans URL-level transparency tooling for Google-Extended in the coming weeks, per Cloudflare, with Apple's equivalent due next year.

The launch also redefines the existing options. Block and "Block on pages with ads" previously spared mixed-use crawlers precisely because of the search side effect; both now apply to them, so selecting Block stops Googlebot, Applebot, and Bingbot outright, search included. Existing configurations migrate automatically: prior Training blocks become Disallow AI Training, the legacy Block AI Bots toggle and Managed Robots.txt are deprecated, and new ad-monetized domains get Disallow AI Training as their default Training preset. Cloudflare says most customers need to change nothing.

Cloudflare's usage data explains the design. Fewer than 1 percent of its sites block search crawlers, while 17 percent have enabled some training block. The company's next target is AI summaries, where it wants a single setting controlling how much content summaries can use, rather than per-operator configuration; Search Engine Journal relays an early-next-year goal, and heise gives early 2027 as the availability target. There is no ads-pages-only training option, because the URL list changes too fast for robots.txt, and no Disallow for agents yet, pending standards work such as ai-prefs.

For marketing and SEO teams, Cloudflare has separated training policy from search visibility at the CDN layer for Google and Apple; Microsoft has committed to support the distinction for Bing. Sites that blocked all AI bots to protect their content can now revise those rules without giving up search access. Cloudflare will also make Disallow AI Training the default for new ad-supported domains, so more pages may remain searchable while refusing model training. Cloudflare says referrals from AI search convert at three to five times the rate of traditional search. If that figure holds, its planned 2027 control for AI summaries may have a greater effect on brand visibility than the training setting released this month.

Related articles

WhatsAppCloudflare Lets Sites Disallow AI Training Without Blocking Googlebot – Geolix.ai