If you manage sites behind Cloudflare, the last two weeks produced one of those rare SEO events with an actual deadline attached: 15 September 2026. The warnings that circulated beforehand were blunt, and mostly correct. What almost none of the follow-up coverage caught is that Cloudflare shipped something on the same day that changes the answer.
Here is the sequence, because the order matters more than the headlines.
What was announced in July, and what actually shipped
On 1 July 2026, Cloudflare published its plan to replace the single "Block AI Bots" switch with three separate controls covering Search, Agent and Training behaviour. Buried in that post was the line that caused the panic:
"Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service)."
Read that twice, because the parenthetical carries the weight. Googlebot crawls for search and for AI training with the same bot. Under a most-restrictive-rule-wins model, a Training block catches it. And the rule applied to the legacy toggle, meaning a checkbox someone ticked in 2025 and forgot about was scheduled to start returning 403s to Google.
This was not theoretical. Cloudflare's own community forum carried reports of homepages being de-indexed after owners set the equivalent policy, with Search Console reporting 403s on crawl attempts. Losing the homepage is bad. Losing the sitemap alongside it is how a single setting turns into a site-wide indexation problem.
Then, on 15 September itself, Cloudflare announced the other half: a new Accountable designation for crawler operators, and a Disallow AI Training setting that lets a site refuse model training while staying fully in search results.
The mechanism is a compliance test, not a whitelist. Cloudflare set four requirements an operator must meet or commit to on a timeline:
- A clear way to opt out of AI training via robots.txt or a comparable standard
- A way to opt out of AI-generated search summaries
- URL-level visibility into how content is used for search and for training
- A public confirmation that opting out of training will not affect traditional search rankings
Apple, Google and Microsoft have been designated Accountable. Other mixed-use crawlers have not, and per Cloudflare's own wording they are still blocked when a site owner blocks Training.
So which is true?
Both, and the distinction is the practical part.
If you block AI training today and the crawler belongs to an Accountable operator, you keep your search crawling. If it belongs to a mixed-use operator that has not met the criteria, blocking training still blocks that crawler outright. The tradeoff was not abolished, it was made conditional on operator behaviour.
The scale figures in Cloudflare's release explain why this got attention at all. Mixed-use crawlers (one bot collecting for both a search index and AI training) now make up 36.6% of verified crawler traffic on Cloudflare's network, the single largest category. Meanwhile fewer than 1% of site owners block search crawlers, while 17% restrict AI training. That gap is precisely the population at risk: thousands of sites that meant to refuse training and never intended to touch search.
One more detail the secondary coverage keeps blurring. The new ad-page defaults (Training and Agent blocked on pages showing ads, Search allowed) apply to new domains onboarding to Cloudflare, not retroactively to existing zones. The change that reaches existing zones is the most-restrictive-rule logic and the deprecation of the legacy toggle.
What to check this week
This is a fifteen minute audit per site, and it is worth doing on every client property rather than the ones you think are affected.
1. Open Security Settings, Configure AI bot policies. You are looking at three independent controls now: Search, Agent, Training. Each offers Block on all pages, Block on pages with ads, or Allow.
2. Find any legacy "Block AI bots" state. That option is deprecating. If a site is still carrying it, decide deliberately rather than inheriting whatever the migration does.
3. Verify with Search Console, not with the dashboard. Run a Live URL Inspection on the homepage and a money page. A 403 to Google-InspectionTool is the signal that matters.
4. Check the crawlers that are not Googlebot. Firewall and bot rules routinely catch AdsBot-Google (which polices Ads landing page quality), Storebot-Google for Merchant Center, Bingbot, and Applebot, which now feeds Siri and Apple Intelligence. A broken AdsBot crawl shows up as an advertising problem, not an SEO one, which is why it stays undiagnosed.
5. Record what you set and when. When rankings move three weeks later, you want the config change in the same timeline as the traffic data.
The wider pattern worth internalising
Crawler access has quietly become something you configure rather than something you assume. The Cloudflare episode is one instance. Research by Vinicius Stanula, published on Search Engine Land, found another: pages linked only through client-side JavaScript were largely undiscoverable to AI crawlers. Every page linked from raw HTML was found by at least one crawler. JavaScript-linked content was reached mostly by GoogleOther, which does not build the search index, and even then only 42% to 67% of it depending on how deep it sat in the hierarchy.
Two different failures, same root cause: content that exists, that the owner believes is published, and that the machines reading the web cannot reach. Ranking factors are the argument everyone enjoys having. Retrievability is the one that silently decides whether you were ever in the running.
This is the part we designed TrafficForge's scoring around. A page that scores well on structure, schema and internal linking is a page a crawler can actually parse and traverse. QualityForge Lite checks 16 criteria before a page publishes, including heading hierarchy, structured data and internal links, precisely because the expensive errors are the boring structural ones rather than the stylistic ones. It cannot see your CDN settings, which is why the audit above is manual, but it does stop you shipping pages that are unreachable by construction.
FAQ
Did Cloudflare block Googlebot for everyone on 15 September 2026?
No. The change only affected zones where the owner had chosen to block Training, either through the new controls or the legacy "Block AI bots" toggle. Sites with no AI blocking active were unaffected. The separate ad-page defaults applied only to domains newly onboarding to Cloudflare, not to existing zones.
Can I block AI training and still rank in Google?
Yes, as of 15 September 2026. Cloudflare's Disallow AI Training setting lets a site refuse training while remaining in search, provided the crawler belongs to an operator designated Accountable. Apple, Google and Microsoft currently hold that designation. Mixed-use crawlers from operators that have not met Cloudflare's four criteria are still blocked when you block Training.
How do I tell if my site was affected?
Use Live URL Inspection in Google Search Console on your homepage and a key landing page. A 403 response indicates the crawler is being refused at the edge. Cross-check the crawl stats report for a drop starting mid-September, and check Cloudflare's AI Crawl Control analytics for blocked requests by crawler and path.
Which crawlers should I explicitly allow at the CDN?
At minimum: Googlebot, Google-InspectionTool, AdsBot-Google and AdsBot-Google-Mobile, Storebot-Google, Bingbot, Applebot, DuckDuckBot, YandexBot and Slurp. Blocking the Ads and Merchant crawlers produces problems that surface in advertising reporting rather than organic reporting, so they tend to go unnoticed for longer.
Does any of this affect whether AI search engines cite my pages?
Indirectly, and significantly. A crawler that cannot fetch a page cannot cite it. If you block Training across the board, you are opting out of the corpus that some generative engines draw on. The newer controls let you make that decision per behaviour (search, agents, training) rather than as one blunt on or off, which is the first time that choice has been precise enough to be worth treating as strategy.
Ready to build your own SEO page library? Visit trafficforge.app to get started — 30-day money-back guarantee. Or email [email protected] with questions.