01Home 02Work 03GEO 04Blog 05About 06Contact 07Free AI-visibility check 中文
[email protected]

Blog / Trends

Cloudflare's 15 September Change: Check Whether You're Blocking the AI Crawlers

Cloudflare's 15 September Change: Check Whether You're Blocking the AI Crawlers

On 15 September 2026 Cloudflare changes how AI crawlers are handled by default, and retires the legacy “Block AI bots” toggle. New domains get Training and Agent bots blocked on ad-bearing pages. Existing customers using the old toggle will find it starts blocking crawlers that combine search and training purposes, which Cloudflare says includes Googlebot, Applebot and Bingbot. If you have never checked what your CDN serves to AI crawlers, this is the month to look.

I have a personal reason for writing this one. Earlier this summer I discovered that this site, a site whose entire subject is getting found by AI, was serving a managed robots.txt that told GPTBot, ClaudeBot and Google-Extended to go away. I had not set it. A platform default had, and I had never thought to check. Fixing it took fifteen minutes. Not knowing had cost considerably more than that.

What exactly is changing?

Cloudflare announced on 1 July 2026 that it is splitting crawler control into three categories: Search, Agent and Training. Those controls are live now for every customer, including free accounts. The dated change is what happens to defaults.

From 15 September, in Cloudflare’s own words, “for all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.” That is narrower than much of the commentary suggests. It applies to new domains, not to every existing free site, and only to ad-bearing pages.

The part that reaches further is the retirement of the old control. Cloudflare is deprecating the legacy “Block AI bots” setting on the same date, and from then, crawlers that serve more than one purpose get caught by any training-blocking configuration. Cloudflare specifically names Googlebot, Applebot and Bingbot as multi-purpose crawlers. If your site is running that legacy toggle, the thing it does is changing underneath you.

Why does this matter for AI visibility?

Because you cannot be cited in an answer built from content the engine was never allowed to fetch.

A peer-reviewed paper accepted to ACM SIGIR 2026, based on 11,500 live queries, reports that “websites that block Google’s AI crawler are significantly less likely to be retrieved by AIOs, despite having access to the content.” It is an observational study, so it shows association rather than proven causation, and it attaches no headline percentage. But as a working assumption for a business that wants to be found, the direction is not in doubt.

The same study found AI Overviews appearing on 51.5% of queries, with very little overlap between the sources feeding those answers and the organic results below them. The sources feeding AI answers are not a subset of the results ranking beneath them. It is a different retrieval path with different rules, and access is the first of them.

Who should actually allow these crawlers?

This is a genuine business decision, not a technical default worth accepting by accident.

A publisher whose model depends on people landing on the page, seeing the ads and staying, has a real argument for restricting training crawlers. Cloudflare’s ad-page default exists precisely because that constituency asked for it. If your content is the product, protecting it is defensible.

A professional services firm, a B2B software company, a corporate services provider or a fintech has close to the opposite problem. Nobody buys a legal engagement because they read your blog post and were pleased with the layout. They buy because when they asked an assistant who handles this kind of work in Hong Kong, your name came back. Blocking the crawlers that build those answers, in order to protect page views you were not monetising anyway, is a straightforwardly bad trade. The cost of being absent from those answers does not show up as a dip in a dashboard. It shows up as enquiries that never arrive.

The awkward middle case is the business that has never made the choice at all. That was this site until July. It is more common than it should be, because the settings live in infrastructure that marketing rarely opens and IT rarely connects to lead generation.

What this changes about technical GEO

For years the technical layer of search visibility was mostly about being crawlable and being fast. The AI era adds a permissions dimension, and the permissions are increasingly being set by platform defaults rather than by the site owner.

That has two consequences. The first is that “we did our technical SEO in 2024” no longer covers it, because a CDN policy change in 2026 can undo the access your visibility depends on without touching a line of your code. The second is that anyone selling you AI visibility should be checking this before they sell you content. A restructured, beautifully cited-ready page behind a blocked crawler is an expensive way to be invisible.

There is a straightforward version of this check: ask what your live robots.txt actually serves to an AI crawler, as opposed to what the file in your repository says, and confirm which crawler categories your CDN currently permits. On this site those two answers disagreed for weeks. It is worth knowing which way yours points before September decides for you.

If you would rather someone else looked, the audit I run starts with exactly this, because there is no point optimising for citation while the door is shut.

Frequently asked questions

What is changing at Cloudflare on 15 September 2026? Two things. New domains onboarding to Cloudflare get Training and Agent bots blocked by default on ad-bearing pages, with Search still allowed. Separately, mixed-purpose crawlers become subject to any training-blocking configuration, including the legacy “Block AI bots” toggle, which is being deprecated the same day. Cloudflare names Googlebot, Applebot and Bingbot as affected multi-purpose crawlers.

Does this mean my existing site will suddenly be blocked? Not by the ad-page default, which is scoped to new domains. But if you use the legacy “Block AI bots” setting, its behaviour changes and it is going away. Move to the newer Search, Agent and Training controls deliberately.

Why would blocking AI crawlers hurt my visibility? A SIGIR 2026 paper found sites blocking Google’s AI crawler were significantly less likely to be retrieved by AI Overviews despite the content being accessible. It is an association rather than a proven cause, with no percentage attached, but the direction is consistent.

Should I block AI crawlers or allow them? It depends on whether your content is the product or the shop window. Publishers have a real case for restricting training crawlers. Firms that want to be recommended by an assistant usually do not. The genuine mistake is having the decision made for you by a default you never read.

Frequently asked

> What is changing at Cloudflare on 15 September 2026?

Two things. For new domains onboarding to Cloudflare, bots in the Training and Agent categories will be blocked by default on pages that display ads, while Search crawlers stay allowed. Separately, and this affects existing customers, crawlers that combine search and training purposes will be blocked by any training-blocking configuration, including the legacy 'Block AI bots' toggle, which Cloudflare is deprecating on the same date. Cloudflare names Googlebot, Applebot and Bingbot among the affected mixed-purpose crawlers.

> Does this mean my existing site will suddenly be blocked?

Not by the ad-page default, which applies to new domains onboarding to Cloudflare. But if you are already using the legacy 'Block AI bots' setting, the behaviour of that setting changes, and it is being retired. Anyone relying on it should move to the newer Search, Agent and Training controls deliberately rather than letting a deprecation decide for them.

> Why would blocking AI crawlers hurt my visibility?

A peer-reviewed SIGIR 2026 study found that sites blocking Google's AI crawler were significantly less likely to be retrieved by AI Overviews, even though AI Overviews had access to the content. That is an association rather than a proven cause, and the paper attaches no percentage to it, but the direction is clear enough that blocking should be a deliberate decision rather than an accident of a default setting.

> Should I block AI crawlers or allow them?

It depends on what your content is for. A publisher whose revenue depends on people arriving on the page has a real argument for restricting training crawlers. A professional services firm or B2B company that wants to be the named answer when a buyer asks an assistant who to hire is usually working against itself by blocking. The mistake is not choosing either way, it is not knowing which one your infrastructure has quietly chosen for you.