Editorial illustration: a September wall calendar tilted slightly to the right, the 15th circled in teal with a set mousetrap resting on that date — signifying that the September 15 crawler-blocking deadline is a trap for the unwary.
Life Beyond Google · Before September 15

The AI Block That Can Quietly Delete You From Google

July 16, 2026 By Liz Micik 6 min read

The short version:

  • The danger on September 15 is not blocking AI. It is blocking it with an instrument so blunt that it also tells Google to stop indexing you.
  • Reach for robots.txt instead of a one-click category switch, keep the search crawlers explicitly welcome, and do not reflexively block CCBot.

Most of this summer's AI deadlines you can watch from a comfortable distance. But there is one coming in just eight weeks that you cannot afford to ignore, because it can quietly change whether Google can see your website at all.

On September 15, Cloudflare will turn on a new default setting that will block AI training and agent crawlers on pages that show ads. It won't affect all of the 41 million sites it serves (about 23% of all websites), but it will affect new Cloudflare sites and “free tier” sites that have not expressly adopted AI allowance settings.

Why Everyone Started Blocking

To see the risk clearly, it helps to understand why blocking became popular in the first place, because the instinct behind it is completely reasonable.

For about thirty years the web ran on a simple trade. A crawler read your pages, and in return it sent you visitors. Google's search crawler still honors that deal, reading only about five of your pages for every visitor it sends back. The AI crawlers do not. Measured across Cloudflare's network, Perplexity sits closer to a hundred pages per visitor, OpenAI's crawler near thirteen hundred, and Anthropic's at nearly twenty-four thousand pages taken for every single visitor returned.

Pages an AI crawler takes for every visitor it sends back Crawl-to-referral ratio by crawler · log scale, so every bar stays visible 1 10 100 1,000 10,000 Pages crawled per visitor 5 Google search 111 Perplexity 1,276 OpenAI ChatGPT 23,951 Anthropic Claude = 1 visitor sent back to your site For every visitor Google sends you, Anthropic’s crawler reads nearly 24,000 of your pages. Source: Cloudflare Radar crawl-to-refer ratios, January–March 2026.
Crawl-to-referral ratios vary by orders of magnitude. A search crawler trades pages for visitors; most AI crawlers mostly take.

This was not a deliberate decision by the AI platforms to steal your visitors. Their crawlers were simply not all built to send people your way. There are three kinds:

Cloudflare found that in May 2026, about half of all AI-crawler requests were for training and fewer than one in ten were for search. That left website owners paying to serve up pages for training crawlers to gobble, with no visitors to offset the cost. It is reasonable to start looking for ways to stem that.

So site owners started closing the door. By the middle of this year, more than a third of AI-crawler requests on Cloudflare's network were being turned away. Blocking is now the norm, not the exception.

The Trap Nobody Is Flagging

Here is where a reasonable instinct goes wrong.

The crawlers that matter most to you are not single-purpose. Googlebot, Bingbot, and Applebot each crawl for search and for AI in the same visit, using one bot. When you tell a system like Cloudflare to block AI training, it applies the most restrictive rule that fits — and because those bots wear more than one hat, the block meant for AI can catch the search crawler too. The result is a site that set out to keep its content out of AI models and accidentally told Google to stop indexing it for search.

That is a real, self-inflicted way to disappear, and it is nearly invisible when it happens. There is no error message and no alert — just a setting a well-meaning person flipped, and a slow slide out of the search results a few weeks later. For a business whose pipeline depends on being found, it may be the most expensive checkbox on the internet.

The One Bot You Should Not Block

There is one bot you should think long and hard about before blocking, and this cuts against the popular advice.

That bot is CCBot, the crawler for Common Crawl, and when people block AI crawlers they tend to sweep it up along with everything else. Roughly one in twenty sites already blocks it. That is usually a mistake. Common Crawl is not one company's private training set; it is a shared, open archive that feeds a huge slice of the AI ecosystem, from academic research to the models and tools your future customers will use to find you. Blocking GPTBot keeps you out of one company's model. Blocking CCBot shrinks your footprint across the entire neighborhood at once.

And the effect compounds. Tomorrow's models will be trained in part on the data CCBot gathers today, so a block now quietly narrows how well the next generation of AI knows you exist.

If your only goal is to protect your content, blocking CCBot is defensible. But if your goal is to be found, which for most professional-services firms it is, then reflexively blocking CCBot means cutting off tomorrow's traffic to save a little cost today. Allow it on purpose.

What To Do Before September 15

Cloudflare's category switches are a sledgehammer. “Block AI training” is one blunt motion that sweeps up whatever happens to match, multi-purpose search bots included. There is a better way than serving up more and more pages in exchange for fewer and fewer visitors.

Your robots.txt file is a scalpel, and that makes it the right tool for this job. You name the exact bots you want to keep out, one line each, and you leave the rest alone. Keep CCBot, Googlebot, and the other search crawlers explicitly welcome, and decide case by case which of the pure training bots to disallow. Even then, two things are worth keeping in mind:

That second point is the biggest reason I tell clients to be generous with AI crawlers right now. We do not yet know enough to read the intent behind an AI visitor, so I would rather let them in and learn how they read, understand, and ultimately act on a site than lock the door and discover later that I shut out a buyer.

The future of the web is agentic; there are already more bots visiting most websites than people. You get a vote in how they treat you, but only if you cast it. Take a few minutes before September 15 to decide, deliberately, which crawlers you welcome, or you may find your future quietly cut short and yourself slipping out of Google not long after.

And if you are not a Cloudflare customer, spend the same five minutes anyway. Your host or CDN is making its own choices about how to treat these bots, and it is worth knowing what they are.

Not sure what your site is actually telling AI and search crawlers?

Run the free Agent Signal Check and see how the major AI models find, read, and describe your business today.

LM

Liz Micik

SEO & Content Strategy · Agent Readiness

Liz Micik helps complex B2B companies prepare their website, expertise, and conversion processes for AI agents. After 28 years in SEO, content strategy, and B2B demand generation, she translates the jargon of the agentic web into decisions a business can actually act on.

See what your site is telling the crawlers

Before you change a single setting, find out how the major AI models and search engines actually find, read, and describe your business today.

Run Your Free Signal Check Explore the Agent Readiness Audit