Refusing AI training does not have to shut Googlebot out. Cloudflare’s new Disallow AI Training setting combines training preferences with crawler blocking while preserving search access for qualifying mixed-use bots.
The company introduced the control on September 15. Its rollout announcement also confirms an important distinction: selecting Block can now stop Googlebot, Applebot and Bingbot, including their search crawls.
For businesses restricting AI training bots, the choice of setting matters. Training permissions, crawler access and appearance in AI-generated answers are separate decisions.
Google-Extended Separates Training Permission From Googlebot Access
Google already provides a way to restrict certain AI uses without removing a website from Search.
Its Google-Extended documentation describes a product token used in robots.txt to control whether crawled content can train future Gemini models. The documented scope includes models powering Gemini Apps and the Vertex AI API for Gemini, alongside specified grounding uses.
Grounding is the retrieval of information to support an answer when a model responds. It differs from training, which changes the model itself.
Google-Extended has no separate HTTP user-agent string. Google retrieves pages through its existing crawlers, then uses the robots.txt preference to govern covered uses of that material. A website can therefore continue receiving Google crawler requests after disallowing Google-Extended.

Google explicitly says that the token does not affect inclusion or ranking in Google Search.
That makes the distinction from Googlebot consequential. Restricting Google-Extended communicates a usage preference. Blocking Googlebot prevents the search crawler from fetching content. The two controls operate at different points in the process and should not be treated as substitutes.
Cloudflare’s Block Option Now Carries Search Consequences
Cloudflare organises its AI bot policies around three behaviours: Search, Training and Agent. Search crawlers build indexes, training crawlers collect material for models, and agents visit pages while acting on a person’s behalf.
A single crawler can serve more than one purpose.
Disallow AI Training publishes applicable robots.txt preferences through Bot Preference Sync. Mixed-use crawlers with Cloudflare’s Accountable designation retain search access; other training crawlers are blocked.
The alternative settings enforce access restrictions. Under Training, Block applies across the domain, while Block on pages with ads applies to pages detected as displaying advertising. Mixed-use crawlers are now included.
An access block cannot distinguish between the downstream purposes of the same request. If a crawler is prevented from fetching a page, its search function loses that access too.
Keeping the Search control set to Allow therefore should not be read as a blanket exemption from another applicable block. Cloudflare’s controls concern crawler behaviour, and a bot performing both search and training can fall into both categories.
AI Overviews Require a Separate Publisher Decision
Disallowing training does not automatically exclude a website from Google’s AI search experiences.
Google’s Search generative AI control manages whether a site’s links and content appear in AI Overviews, AI Mode and generative AI features in Discover. Google says the control became available to all websites worldwide on August 31.
The setting sits in Search Console under Settings > Search generative AI.
Sites that choose exclusion lose eligibility for links, displayed content and grounding within the covered features. They also stop receiving impressions and traffic from those experiences. Google says the choice is not a ranking or inclusion signal for other parts of Search.
That separates the distribution question around AI Overviews from the training question. A publisher may want ordinary search access, refuse covered model-training uses and make an independent decision about AI-generated search answers.
Google’s help page explicitly says the Search Console control does not affect AI training. It directs publishers to Google-Extended for that purpose. Exclusion also does not override separate participation choices for services such as Merchant Center or Google Ads.
Apple Supports a Training Signal; Bing’s Integration Is Pending
Apple’s system follows a similar distinction between collection and use.
According to its Applebot documentation, publishers can disallow Applebot-Extended to opt out of training Apple’s generative foundation models. Applebot-Extended does not crawl pages itself. It controls how data collected by Applebot may be used.
Apple says those restrictions do not determine search rankings. Content can remain discoverable through search experiences in Spotlight, Siri and Safari while the training restriction applies.
AI-generated responses have another control. Apple documents the nosnippet directive for excluding content from use as additional context in generated output, including broad knowledge answers. That directive also affects descriptions and web answers, so its scope extends beyond training.
Bing’s Cloudflare integration remains incomplete. Cloudflare says Disallow AI Training does not yet automatically convey a no-training preference to Bing through robots.txt; Microsoft’s support is targeted for early 2027.
The setting therefore cannot currently be described as a uniform training opt-out across all three mixed-use search crawlers.
Existing Settings Migrate, but Robots.txt Remains a Preference
Cloudflare says existing Training selections of Block or Block on pages with ads will migrate to Disallow AI Training. The controls are available on all plans.
The underlying distinction between instructions and enforcement remains relevant. Cloudflare’s robots.txt documentation explains that compliance is voluntary: publishing a directive does not technically prevent a crawler from retrieving content. Network blocking supplies a separate enforcement mechanism.
A public robots.txt file communicates instructions to crawler operators. It does not authenticate visitors or make public content private.
In practice, marketing and technical teams need to assess the resulting configuration together: the published training preference, whether search crawlers can still fetch important pages, and the separate settings governing AI answer visibility. Reviewing those outcomes distinguishes a content-use policy from an accidental search-access restriction.
Cloudflare lists the new controls in each domain’s Security Settings. Bing’s planned robots.txt support remains outstanding.


