Are you blocking ChatGPT and Claude by accident? A robots.txt guide to AI crawlers

Short answerIf you want AI assistants to recommend your business, allow their search and user-fetch crawlers — such as OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User and PerplexityBot — in robots.txt, and check that your CDN or firewall is not blocking them. Blocking training crawlers like GPTBot or ClaudeBot is a separate choice that does not remove you from AI search.

Why this matters

Many websites block AI crawlers without anyone deciding to. Common causes:

  • A website template or security plugin that added broad Disallow rules.
  • A CDN or firewall setting that blocks or challenges “AI bots”.
  • A robots.txt copied from another site years ago.

If an AI assistant’s crawler cannot read your site, the assistant has to rely on whatever others say about you — or skip you.

The main AI crawlers, by purpose

AI companies increasingly separate crawlers by purpose. That lets you choose: be found in AI search, without necessarily contributing to model training.

Crawler Operator Purpose
OAI-SearchBot OpenAI Indexes pages for ChatGPT search results
ChatGPT-User OpenAI Fetches a page when a ChatGPT user asks about it
GPTBot OpenAI Collects data for model training
Claude-SearchBot Anthropic Indexes pages for Claude’s search
Claude-User Anthropic Fetches a page when a Claude user asks
ClaudeBot Anthropic Collects data for model training
PerplexityBot Perplexity Indexes pages for Perplexity answers
Perplexity-User Perplexity Fetches a page for a user request
Googlebot Google Google Search, including AI Overviews and AI Mode
Google-Extended Google A control token for Gemini training and grounding (not a separate crawler)
Bingbot Microsoft Bing and Copilot; Bing’s index is also used by other AI search products
Applebot-Extended Apple A control token for Apple’s AI training
CCBot Common Crawl Open web dataset used to train many models

Crawler names and behaviour change over time; check each operator’s documentation for the latest list.

User-agent: *
Allow: /

# AI search and user-requested fetching
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

Sitemap: https://www.example.com/sitemap.xml

Two notes:

  • Groups combine user-agents. Several User-agent lines followed by rules form one group.
  • The most specific group wins. If you have a User-agent: GPTBot group with Disallow: /, it applies to GPTBot regardless of what User-agent: * says.

Should you allow training crawlers?

This is a business decision, not a technical one.

  • Allowing GPTBot, ClaudeBot, Google-Extended and CCBot means your public content may be used to train future models. For many service businesses, that is a benefit: future models “know” who you are.
  • Blocking them protects content you consider proprietary, and — for the major operators — does not remove you from their AI search results, because search uses separate crawlers.

At TropiGEO we allow training crawlers on our own site, because being well known to AI models is part of our business.

Check your CDN and firewall too

robots.txt is a request, not a lock. Your CDN can block crawlers outright, regardless of robots.txt. If you use Cloudflare, review Security → Bots → AI Crawl Control and choose what to do for Search, Agent and Training traffic separately. Aggressive bot-fighting modes can also challenge legitimate AI crawlers.

Check it in 30 seconds

Our free AI visibility check reads your robots.txt and shows, crawler by crawler, which AI systems are allowed — plus whether you have llms.txt, a sitemap and structured data.

See what AI says about your business

Run the free AI visibility check in under a minute, or book a full audit for Singapore or Malaysia.