Guide · Updated 2026-10-08

robots.txt and AI Crawlers: What to Allow and What to Block

How a robots.txt file works, how training crawlers differ from AI search and user-request bots, and a sample policy that blocks training but allows search.

A robots.txt file is a plain-text request to automated visitors. It tells well-behaved crawlers which parts of your site to read and which to skip. With AI products now reading the web, the file has become a way to say whether your pages may be used for training, for AI search or for answering a user’s question. The AI Crawler Policy Builder writes the file for you.

What robots.txt can and cannot do

  • It is a request, not a lock. Reputable crawlers follow it. Others may not, and the file does not hide anything, because it is public.
  • It does not remove pages from search. Use a noindex tag for that, and note that blocking a page in robots.txt can stop crawlers from seeing the tag.
  • It works per crawler name, called a user agent.

Three kinds of AI crawler

TypePurposeExamples from the builder
TrainingCollects pages that may be used to train modelsGPTBot, ClaudeBot, anthropic-ai, Google-Extended
AI searchFinds pages to show in AI search resultsOAI-SearchBot, Claude-SearchBot
User requestFetches a page when a person asks the assistant toChatGPT-User, Claude-User

Google-Extended is a control token that tells Google whether Gemini models may use your content. Blocking it does not affect Google Search, which is controlled by Googlebot.

A sample policy: block training, allow search

With the builder’s “No training” starting policy, 12 crawlers are blocked and 8 are allowed. GPTBot, ClaudeBot, anthropic-ai and Google-Extended are blocked, while OAI-SearchBot, ChatGPT-User, Claude-SearchBot and Claude-User are allowed. This keeps your pages out of training sets while still letting AI search and assistants read them to answer questions and link back.

Writing the file

  • Start with User-agent: * and a Disallow line for private paths such as /admin/.
  • Add a separate block for each AI crawler with Allow: / or Disallow: /.
  • End with a Sitemap line giving the full address of your sitemap.
  • Save it as robots.txt at the root of your domain.

The Robots.txt Generator covers the same file with a simpler set of options, and the Sitemap Generator builds the sitemap it points to.

Trade-offs

Blocking a crawler can reduce how often your site appears in the AI product that uses it. Allowing training means your content may shape future models. There is no right answer, so decide what suits your site. Crawler names and behaviour change, so check each company’s current documentation from time to time.

Frequently asked questions

Will blocking GPTBot remove my pages from ChatGPT search? No. Search uses a different crawler, OAI-SearchBot.

Do I need an llms.txt file too? It is an optional, plain-language summary of a site for language models. It is separate from robots.txt and does not control access.

Try the tools