Whether to block AI crawlers is presented as a technical question and it is not one. Blocking them is a few lines of configuration that takes ten minutes. The hard part is deciding whether you want to, and that decision is commercial: you are choosing between protecting content from being trained on and being present in the answers people now get instead of search results.
Most of the advice online picks a side and argues it. The honest position is that the right answer differs by business model, and that a publisher funded by page views and an agency funded by enquiries should reach opposite conclusions.
The trade in one sentence: blocking AI crawlers stops your content being ingested, and it also stops you being cited. If your revenue depends on people arriving on your page, blocking protects you. If your revenue depends on people knowing you exist and then contacting you, blocking removes you from the surface where that discovery increasingly happens.
What Cloudflare Actually Changed
The default flipped. On 1 July 2025 Cloudflare became the first major infrastructure provider to block AI crawlers by default , moving from an opt-out model to an opt-in one. New domains are asked up front whether AI crawlers may scrape them, and AI companies now need explicit permission rather than an absence of refusal.
That matters beyond Cloudflare’s own customers, because a large share of web traffic passes through it. A default is not a rule, but a default applied at that scale reshapes what AI companies can assume.
The tooling is AI Crawl Control , available on all plans including Free. It lets you see which AI services are reaching your content, set allow or block rules per crawler, and check whether a given crawler respects robots.txt. Seeing the traffic before deciding anything is the useful part; most site owners have never looked.
Cloudflare also pushed crawlers to declare purpose, so you can distinguish a bot training a model from one running inference for a live answer from one indexing for search. Those are three different bargains and until recently they arrived indistinguishable.
The Money Model Changed Again in July 2026
This is where advice written last year is now wrong.
Pay Per Crawl launched in 2025 as a marketplace where publishers charged AI companies per fetch. On 1 July 2026 Cloudflare declared that model insufficient and moved to Pay Per Use , which pays when your content creates value inside an answer rather than when a bot downloads a page.
The reasoning is straightforward once you see the number: over 50% of AI crawler traffic re-fetches pages that have not changed. Charging per fetch mostly bills for waste. Charging when content appears in an answer bills for the thing that actually displaced a visit to your site.
It is early. The launch partners are Ceramic.ai and You.com, so this is not yet a broad revenue stream for a small UK business. Treat it as a direction of travel rather than a line in next year’s budget.
There is also a deadline worth diarising. From 15 September 2026, mixed-use crawlers will be blocked by default from pages carrying advertising unless the site owner changes the setting. If you run ads and rely on AI referral traffic, that default will apply to you without you doing anything.
What Blocking Actually Costs
Three things, and only the first is obvious.
You lose citations. AI answers name sources. Being one of those names is the modern equivalent of a first-page ranking, and a blocked crawler cannot cite what it cannot read. Our generative engine optimisation guide covers what earns a citation once you have decided to be visible.
You lose the measurement. Blocked traffic does not appear in your reporting as a loss. It appears as nothing, which is indistinguishable from never having had it. Sites that blocked early often cannot say what it cost them because the counterfactual was never recorded.
You do not actually stop training. Blocking the well-behaved crawlers that identify themselves and honour robots.txt removes exactly the ones willing to follow rules. Content already ingested stays ingested, and scrapers that ignore robots.txt were never in scope. You are shaping the behaviour of the polite subset.
Against that, blocking genuinely protects a business whose product is the content itself. A subscription publisher, a research firm selling reports, a stock library: for these, an AI answer that summarises the content is a direct substitute for the sale, and citation is poor compensation.
How to Decide Whether to Block AI Crawlers
Ask what a visitor is worth and how they convert.
Block if your content is the product. Paywalled journalism, paid research, reference data you license. The AI answer competes with your sale.
Allow if your content is marketing for something else. A software agency, a consultancy, a law firm, a clinic. Your posts exist to make buyers aware you can solve their problem. Being named in an answer is the awareness you were paying for, and the enquiry still has to come to you.
Allow selectively if you are somewhere in between. Per-crawler rules let you admit search indexing while refusing training, which is a defensible middle position now that purpose is declared.
For most businesses reading this, the answer is allow. The instinct to block is protective and understandable, and it is usually protecting content that has no independent commercial value while cutting off the discovery channel that content was written to feed.
Where to Start
Look at the traffic before you change anything. AI Crawl Control shows which crawlers are reaching you and how often, on any plan. A month of that data turns an ideological argument into an arithmetic one.
Then decide per crawler rather than globally, and write the decision down with a date, because this area has changed twice in fourteen months and will change again.
Mecanik runs technical SEO audits that cover AI crawler access and citation visibility alongside the conventional checks, and implements the configuration through our website development team. If you are ranking well and never appearing in AI answers, crawler access is the first thing to rule out.
Related reading: AI SEO Agency: What They Do and What to Pay , Google AI Mode: What It Means for Your Website Traffic , WordPress Hacked: Malware Removal and Recovery Guide and Cloudflare Website Speed: A 2026 Optimisation Guide .
Frequently Asked Questions
Should I block AI crawlers on my website? It depends entirely on how you make money. Block if your content is the product, as with paywalled journalism or paid research, because an AI summary substitutes for the sale. Allow if your content is marketing for a service, because being cited in an answer is the awareness the content was written to create and the enquiry still comes to you.
Does blocking AI crawlers stop my content being used for training? Only partly. It stops the well-behaved crawlers that identify themselves and honour robots.txt, which is exactly the subset willing to follow rules. Content already ingested remains ingested, and scrapers that ignore robots.txt were never affected. You are shaping the polite crawlers, not all of them.
What is Cloudflare AI Crawl Control? A Cloudflare feature, available on all plans including Free, that lets you see which AI services are accessing your site, set allow or block rules for individual crawlers, and track whether each one respects robots.txt. It also supports pay-per-crawl pricing. Looking at the traffic before deciding anything is the most useful thing it does.
What replaced Cloudflare Pay Per Crawl? Pay Per Use, announced on 1 July 2026. Instead of charging AI companies each time a page is fetched, publishers are paid when their content creates value inside an answer. Cloudflare’s reasoning was that over 50% of AI crawler traffic re-fetches unchanged pages, so per-fetch billing mostly charged for waste. Launch partners are Ceramic.ai and You.com.
What happens on 15 September 2026? Cloudflare begins blocking mixed-use AI crawlers by default on pages that carry advertising, unless the site owner changes the setting. Mixed-use means a crawler that does not separate search indexing from AI training and agent traffic. If you run ads and depend on AI referrals, that default applies to you without any action on your part.
Comments