Share
By Joel Thomson
AI can’t recommend what it can’t reach. Strictly speaking, it can. That’s the problem.
If an AI system researching suppliers for a buyer can’t retrieve your current product pages, technical documentation or case studies, it may still answer. It can fall back on old model knowledge, review sites, cached material, competitor claims and somebody else’s version of your business.
I use AI heavily. That has made me less impressed by fluent answers and more interested in what the machine could actually reach. Confidence tells you nothing about the completeness of the research.
Last year, the argument centred on AI companies taking publisher content to train their models. Cloudflare has now made the traffic itself more legible.
It has separated three activities. Search crawlers build an index. Training crawlers absorb content into a model. Agents retrieve or use information in real time for a person.
Here, Agent isn’t another name for AI search. Nor does it necessarily mean an autonomous digital employee roaming the web. Cloudflare’s category includes chat fetchers and browser-use tools retrieving pages because somebody asked them to. To the buyer, the experience may still look like an ordinary AI search.
The controls are already available. From 15 September, new domains joining Cloudflare began blocking Training and Agent traffic by default on pages displaying ads, while Search stays open. Customers can choose different settings.
Cloudflare says more than 20 percent of the web sits behind its network. That explains the attention. It does not mean a fifth of the web will suddenly disappear from AI. The September default is narrower than that.
For B2B marketing leaders, the practical issue is still substantial. The setting may be technical. The consequence is whether a buyer can retrieve your evidence.
The agent may have been sent by your buyer
B2B purchases are slow, expensive and evidence hungry. An agent can help assemble a longlist, compare technical capabilities, summarise customer proof or prepare material for an internal business case.
Account-based marketing (ABM) starts from the premise that important buying activity happens before a form fill. Agents add another layer to that hidden work.
When an agent can’t reach a supplier’s first-party content, the research may continue without it. Marketing gets no visit or identifiable contact. The account can move while the brand’s current evidence remains absent from the decision.
The machine won’t complain. It will make do.
GEO only works on what can be reached
Marketers have spent the past year discussing generative engine optimisation: structured content, clear claims, credible proof and strong third-party sources.
All useful. But a blocked page stays a blocked page.
There’s something impressively corporate about commissioning a GEO (Generative Engine Optimisation) program before checking whether the machines can read the site.
Access belongs in the content strategy. Product information and proof intended to support consideration should usually be retrievable by legitimate buyer agents. Proprietary research may deserve protection or a licence. Customer-only material stays closed.
One setting across every page and every machine purpose is unlikely to be a serious policy.
Someone already owns the setting
Cloudflare has put a useful set of controls in front of businesses. My question is: who in the business is looking at them?
Its AI Crawl Control dashboard can show crawler operators, request activity, requested paths and current allow-or-block settings. A block is enforced through Cloudflare’s web application firewall. Other firewall rules can still stop crawlers marked as allowed.
So, I’d start with an unglamorous meeting. Get marketing, web and security looking at the same screen. Check the category settings, individual crawler policies and any conflicting firewall rules.
Then run a realistic supplier-research task through the AI services your buyers are likely to use. Review the answer and its cited sources. Where the service makes an identifiable live request, compare it with Cloudflare’s crawler activity. You won’t get perfect attribution, but you’ll learn more than you will by asking ChatGPT whether your GEO is good.
I’ve watched technical defaults become commercial rules before. Marketing usually notices when the numbers move.
An AI system researching for your next buyer won’t tell you it couldn’t reach the evidence. It will simply return with an answer built without you.
Joel Thomson is the head of strategy and brand at marketing consultancy Green Hat.
Image: Supplied
Read more: Your lead generation-focused gated content works against you in AI search
