Are you blocking ChatGPT and Claude by accident? A robots.txt guide to AI crawlers
Why this matters
Many websites block AI crawlers without anyone deciding to. Common causes:
- A website template or security plugin that added broad
Disallowrules. - A CDN or firewall setting that blocks or challenges “AI bots”.
- A robots.txt copied from another site years ago.
If an AI assistant’s crawler cannot read your site, the assistant has to rely on whatever others say about you — or skip you.
The main AI crawlers, by purpose
AI companies increasingly separate crawlers by purpose. That lets you choose: be found in AI search, without necessarily contributing to model training.
| Crawler | Operator | Purpose |
|---|---|---|
| OAI-SearchBot | OpenAI | Indexes pages for ChatGPT search results |
| ChatGPT-User | OpenAI | Fetches a page when a ChatGPT user asks about it |
| GPTBot | OpenAI | Collects data for model training |
| Claude-SearchBot | Anthropic | Indexes pages for Claude’s search |
| Claude-User | Anthropic | Fetches a page when a Claude user asks |
| ClaudeBot | Anthropic | Collects data for model training |
| PerplexityBot | Perplexity | Indexes pages for Perplexity answers |
| Perplexity-User | Perplexity | Fetches a page for a user request |
| Googlebot | Google Search, including AI Overviews and AI Mode | |
| Google-Extended | A control token for Gemini training and grounding (not a separate crawler) | |
| Bingbot | Microsoft | Bing and Copilot; Bing’s index is also used by other AI search products |
| Applebot-Extended | Apple | A control token for Apple’s AI training |
| CCBot | Common Crawl | Open web dataset used to train many models |
Crawler names and behaviour change over time; check each operator’s documentation for the latest list.
A recommended robots.txt for businesses that want AI visibility
User-agent: *
Allow: /
# AI search and user-requested fetching
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
Sitemap: https://www.example.com/sitemap.xml
Two notes:
- Groups combine user-agents. Several
User-agentlines followed by rules form one group. - The most specific group wins. If you have a
User-agent: GPTBotgroup withDisallow: /, it applies to GPTBot regardless of whatUser-agent: *says.
Should you allow training crawlers?
This is a business decision, not a technical one.
- Allowing GPTBot, ClaudeBot, Google-Extended and CCBot means your public content may be used to train future models. For many service businesses, that is a benefit: future models “know” who you are.
- Blocking them protects content you consider proprietary, and — for the major operators — does not remove you from their AI search results, because search uses separate crawlers.
At TropiGEO we allow training crawlers on our own site, because being well known to AI models is part of our business.
Check your CDN and firewall too
robots.txt is a request, not a lock. Your CDN can block crawlers outright, regardless of robots.txt. If you use Cloudflare, review Security → Bots → AI Crawl Control and choose what to do for Search, Agent and Training traffic separately. Aggressive bot-fighting modes can also challenge legitimate AI crawlers.
Check it in 30 seconds
Our free AI visibility check reads your robots.txt and shows, crawler by crawler, which AI systems are allowed — plus whether you have llms.txt, a sitemap and structured data.