What we do Pricing FAQs Contact Get started
All articles

Is Your Website Blocking AI Crawlers? Check Before 15 September

TL;DR

On 1 July 2026 Cloudflare confirmed that from 15 September it blocks AI training and agent crawlers by default on ad-carrying pages for newly onboarded domains, and TechCrunch reports the change reaches free-plan customers too. When AI crawlers cannot read your site, ChatGPT, Claude, Perplexity and Gemini describe your business from whatever third parties have said, or they say nothing. Five minutes with your robots.txt and your Cloudflare dashboard tells you where you stand.

On 1 July 2026, Cloudflare announced its biggest change yet to how websites treat AI crawlers, with new defaults landing on 15 September. A handful of UK business sites will feel that directly. Far more are already turning AI crawlers away and don't know it, usually because of a box a developer ticked years ago and never revisited. This piece walks through what is changing, why it matters, and how to check your own site in about five minutes.

What is changing at Cloudflare on 15 September 2026

From 15 September 2026, Cloudflare flips its defaults: AI crawlers used for training and agent activity get blocked on pages that carry advertising, unless the site owner says otherwise. Search crawlers stay allowed. Cloudflare confirmed all of this on 1 July 2026.

The mechanics are worth understanding, because Cloudflare has spelled them out. Sites on Cloudflare can now sort AI traffic into three buckets: Search, which indexes content for search results; Agent, which fetches pages in real time on behalf of a user; and Training, which harvests content to train or fine-tune AI models (Cloudflare, 1 July 2026). Come 15 September, every new domain joining Cloudflare has Training and Agent blocked by default on ad-carrying pages, with Search still waved through. Crawlers that do more than one of those jobs get treated by the strictest rule a site applies, and Cloudflare points to Googlebot, Applebot and Bingbot as exactly that kind of multi-purpose crawler (Cloudflare, 1 July 2026).

TechCrunch adds a detail Cloudflare's own post is quieter about: the new defaults also apply to new sites spun up by existing customers, and to all free-plan customers, who can opt out before the deadline (TechCrunch, 1 July 2026). That one matters, because a lot of UK small-business sites run on Cloudflare's free plan. And this is not the first tightening. Cloudflare has blocked AI crawlers by default for newly signed-up domains since 1 July 2025 (Cloudflare, 1 July 2025), so if your site moved onto Cloudflare in the past year, the door may already be shut.

Why being crawlable matters for AI visibility

An AI assistant can only describe your business accurately if its crawler can read your website. Block GPTBot, ClaudeBot, PerplexityBot and Google-Extended, and ChatGPT, Claude, Perplexity and Gemini are left working from third-party sources, stale data, or nothing at all. Crawler access is the ground everything else is built on.

The reason this has teeth now is that people increasingly ask an AI assistant for a recommendation before they ever open Google. When that assistant cannot fetch your pages, it reaches for directories, review sites and whatever else has been written about you, and that secondhand picture is often out of date or plainly wrong. Access alone won't get you recommended. But without it, the rest of your AI search optimisation work has nothing to stand on.

The main AI crawlers and what each one does

A short list of crawlers does most of the work for UK businesses: GPTBot and OAI-SearchBot from OpenAI, ClaudeBot and Claude-SearchBot from Anthropic, PerplexityBot from Perplexity, and Google's Google-Extended control. They each do a different job, and each can be allowed or blocked on its own in your robots.txt file. That last point is the useful one, because it means you are never forced into all-or-nothing.

  • GPTBot gathers content that may be used to train OpenAI's models (OpenAI, accessed July 2026).
  • OAI-SearchBot indexes pages so ChatGPT can link to and cite them in its search answers. Block it and you get harder for ChatGPT search to surface (OpenAI, accessed July 2026).
  • ChatGPT-User fetches a live page the moment a user asks ChatGPT about it (OpenAI, accessed July 2026).
  • ClaudeBot gathers content that may feed the training of Anthropic's models. Claude-SearchBot crawls to sharpen Claude's search results, and Claude-User fetches pages when users ask. All three obey robots.txt and are controlled independently, so you can allow one without allowing the others (Anthropic, accessed July 2026).
  • PerplexityBot indexes content for Perplexity's search results, and Perplexity states it is not used to train foundation models. Perplexity-User visits a page when a user asks a question (Perplexity, accessed July 2026).
  • Google-Extended is a robots.txt control, not a bot in its own right. It decides whether the content Googlebot already crawls can be used for Gemini training and for grounding, and Google states it does not affect a site's inclusion or ranking in Google Search (Google Search Central, accessed July 2026).

One warning before you touch anything. Googlebot and Bingbot also run ordinary search, so blocking either can quietly cost you normal Google or Bing traffic. Leave those two alone unless you genuinely understand what happens if you don't.

The five-minute self-check

Working out whether your site blocks AI crawlers really does take about five minutes, and it needs no technical skill. Read your robots.txt for blocked bot names, glance at your Cloudflare dashboard if you use Cloudflare, and put one plain question to your developer. Those three moves catch almost every common problem.

Step 1: read your robots.txt

Type yourdomain.com/robots.txt into your browser. Hit Ctrl+F and search for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and Google-Extended. A bot name sitting above Disallow: / means that bot is shut out of the whole site. User-agent: * followed by Disallow: / blocks everything at once. A bot that isn't mentioned anywhere is allowed by default, so silence is good news.

Step 2: check your Cloudflare dashboard

Log in, pick your domain and open AI Crawl Control. The Crawlers tab lists every AI crawler next to an allow or block action (Cloudflare Docs, accessed July 2026). While you are in there, look at your AI bot policy under the security settings, where the Search, Agent and Training presets live, and see whether Cloudflare's managed robots.txt is quietly adding block rules for you. If you are on the free plan, do this before 15 September rather than after.

Step 3: ask your developer

Copy this and send it: "Can GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot fetch our pages, or are they blocked in robots.txt, a firewall or Cloudflare?" It is worth asking because a network-level block leaves no trace in robots.txt at all. Your file can read perfectly clean while Cloudflare or a security plugin turns crawlers away at the door.

When blocking AI crawlers is the right call

None of this means everyone should throw the doors open. For some businesses, blocking is the right and rational call. Publishers who sell their content, membership sites, and firms with real scraping worries may be better off blocked, or better off charging for access through something like Cloudflare's marketplace. It comes down to one question: how does your business actually make its money?

If your words are the product, handing them to AI models for free is a genuine cost, not a rounding error. That is the gap Cloudflare says it is closing with Pay Per Use, a marketplace where publishers can charge AI companies when their content feeds an answer, with Ceramic.ai and You.com signed up as early partners (Cloudflare, 1 July 2026). There is also a sensible middle path: block the training bots, but let the search and user-triggered ones through. For most UK service businesses, though, being found and described accurately is worth far more than the training value of a few pages.

How to unblock AI crawlers in robots.txt

Unblocking a crawler is simple in principle: delete the Disallow line that names it from robots.txt, or add an explicit Allow rule for that bot. The catch is that if Cloudflare is blocking the bot at network level, editing robots.txt alone does nothing, because you also have to change the setting in your dashboard. Both layers have to say yes before a crawler gets in.

These lines allow the main AI crawlers site-wide:

User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

Touch only the AI bot lines and leave the rules that protect your admin areas exactly where they are. If you want the search benefit without the training, allow OAI-SearchBot and Claude-SearchBot while keeping GPTBot and ClaudeBot disallowed. In Cloudflare, set each crawler to Allow in AI Crawl Control and confirm your AI bot policy agrees. One WordPress gotcha: robots.txt is usually controlled by an SEO plugin there, so make the change inside the plugin rather than uploading a file that will get overwritten.

What to expect after unblocking

Don't expect the lights to come on the moment you hit save. Crawlers revisit sites on their own timetable, so it can be days or weeks before AI assistants start reflecting your content. Access is the price of entry, not the finish line, and what these tools then say about you still rides on your content and your wider footprint online.

The good news is that the crawlers come back often. Cloudflare says more than half of all AI crawler traffic is re-fetching pages that have not even changed (TechCrunch, 1 July 2026). The user-triggered fetchers like ChatGPT-User and Perplexity-User read pages live, so those pick up your content the instant access opens. Whether an assistant then actually recommends you is a separate question, and it turns on clear service pages, real evidence of expertise and consistent mentions elsewhere. Our guide to getting recommended by ChatGPT in the UK picks up exactly there.

Questions people ask

These come up again and again when UK business owners ask us about AI crawler access. The short version: the check is quick, allowing AI crawlers does not touch your Google rankings, and Cloudflare has blocked them by default on new domains for the past year, so the newer your site, the more reason to look.

How do I check if my website is blocking ChatGPT?

Visit yourdomain.com/robots.txt in a browser and search the page for GPTBot and OAI-SearchBot. If either name appears above a Disallow: / line, ChatGPT's crawlers are blocked. If you use Cloudflare, also check the AI Crawl Control section of your dashboard, because Cloudflare can block bots without any sign in robots.txt.

Will allowing AI crawlers affect my Google rankings?

No. AI crawler permissions are separate from normal search crawling. Google states that its Google-Extended control does not affect a site's inclusion or ranking in Google Search, and allowing bots such as GPTBot or ClaudeBot has no bearing on how Googlebot indexes your site.

Does Cloudflare block AI crawlers by default?

For new domains, yes. Cloudflare has blocked AI crawlers by default for newly signed-up sites since July 2025. From 15 September 2026 its defaults change again: crawlers classed as training or agent traffic will be blocked on pages that display ads, while search crawlers remain allowed. Existing customers can adjust settings in the dashboard.

Sources

Not sure what your site is telling AI crawlers?

Stagg Studios does this for you, from £100 a month. One plan, no lock-in, cancel any time.

Get startedSee pricing