Machines read differently from people. And from Google.
Many websites block AI crawlers without knowing it. In most answers ChatGPT reads only the title and a snippet around the main heading. Anything rendered by JavaScript does not exist for it.
40%
of landing pages effectively lock out well-behaved AI crawlers.
93%
of ChatGPT answers in default mode open no page at all.
Source: Resoneo, 26,900 pages, 1,200 ChatGPT answers, July to August 2026
< 10%
of AI crawler traffic serves search. Around 80% serves training.
Source: Cloudflare, AI crawler traffic across its network, July 2025
Six checks
Each with a measurable before and after.
Crawler access
Which bots reach you and whether a block is intended. Search and training decided separately.
- One rule per bot, with a reason
- Verified in the server log
Readability without JavaScript
Is the content in the HTML, or does it only appear in the browser?
- Raw HTML against rendered page
- Page weight within limits
The first characters
The heading and what surrounds it are often all that gets read.
- Heading answers the question
- Paragraphs extractable on their own
Internal linking
Crawlers find pages through links, not the sitemap.
- Find orphaned pages
- Link from pages already read
Company data
Name, location, services: identical everywhere. Contradictions produce false statements.
- Reconcile all your own sources
- Structured data for entity clarity
Measurability
Make AI visits and fetches visible in your own tools.
- AI assistant channel in your analytics
- Crawler analysis from the log
What we implement, what we don't
We implement
- Crawler access per bot
- Content in the delivered HTML
- Clear first characters on every page
- Internal linking
- Consistent company data
- Sourced figures and citations in your own text
We don't sell
- llms.txt as a lever: Google calls it unnecessary, Ahrefs found it almost never fetched
- Schema markup as a citation lever: no effect in an Ahrefs controlled study
- Keyword stuffing "for the AI": worse than nothing in the KDD study
- Mass AI-generated pages
- Bought mentions and directory listings
The crawlers that matter
- OAI-SearchBot
The search index behind ChatGPT. Blocked: invisible in ChatGPT answers with web search.
- GPTBot
Training data for OpenAI. Blocked: no effect on search visibility.
- ChatGPT-User
Real-time fetch on a user question. OpenAI documents that robots.txt does not always apply.
- PerplexityBot, Perplexity-User
Index and real-time fetch for Perplexity. Blocked: no citation there.
- ClaudeBot, Claude-SearchBot
Training and search for Claude. Decided separately.
- Google-Extended
Training of Google's models only. No effect on AI Overviews.
- Googlebot
The search index, and with it AI Overviews and AI Mode. Blocked: invisible on Google.
Two things first
Google and the others are not the same
Google's guide says its AI features need no special files and no special markup (Google Search Central, May 2026). True for Google. ChatGPT has its own index; Perplexity fetches sources live. We say for every measure which engine it is meant for. Where GEO goes beyond SEO is set out under GEO compared with SEO.
Part of the effect lies outside your website
The strongest measured correlations concern mentions elsewhere: trade media, review platforms, forums. We do not offer that work. We show you which third-party sources the engines cite and where you are missing. More under AI visibility.
How we work
In your system, with your people. We need access to the CMS and the server logs. Every change is dated so it can be matched to a reading.
Can machines read your website?
Give us your address. We check which crawlers reach you and what an engine finds first.
A reply within one working day, from the person who measures.