Guide · Updated 2026-10-08
robots.txt and AI Crawlers: What to Allow and What to Block
How a robots.txt file works, how training crawlers differ from AI search and user-request bots, and a sample policy that blocks training but allows search.
A robots.txt file is a plain-text request to automated visitors. It tells well-behaved crawlers which parts of your site to read and which to skip. With AI products now reading the web, the file has become a way to say whether your pages may be used for training, for AI search or for answering a user’s question. The AI Crawler Policy Builder writes the file for you.
What robots.txt can and cannot do
- It is a request, not a lock. Reputable crawlers follow it. Others may not, and the file does not hide anything, because it is public.
- It does not remove pages from search. Use a noindex tag for that, and note that blocking a page in robots.txt can stop crawlers from seeing the tag.
- It works per crawler name, called a user agent.
Three kinds of AI crawler
| Type | Purpose | Examples from the builder |
|---|---|---|
| Training | Collects pages that may be used to train models | GPTBot, ClaudeBot, anthropic-ai, Google-Extended |
| AI search | Finds pages to show in AI search results | OAI-SearchBot, Claude-SearchBot |
| User request | Fetches a page when a person asks the assistant to | ChatGPT-User, Claude-User |
Google-Extended is a control token that tells Google whether Gemini models may use your content. Blocking it does not affect Google Search, which is controlled by Googlebot.
A sample policy: block training, allow search
With the builder’s “No training” starting policy, 12 crawlers are blocked and 8 are allowed. GPTBot, ClaudeBot, anthropic-ai and Google-Extended are blocked, while OAI-SearchBot, ChatGPT-User, Claude-SearchBot and Claude-User are allowed. This keeps your pages out of training sets while still letting AI search and assistants read them to answer questions and link back.
Writing the file
- Start with User-agent: * and a Disallow line for private paths such as /admin/.
- Add a separate block for each AI crawler with Allow: / or Disallow: /.
- End with a Sitemap line giving the full address of your sitemap.
- Save it as robots.txt at the root of your domain.
The Robots.txt Generator covers the same file with a simpler set of options, and the Sitemap Generator builds the sitemap it points to.
Trade-offs
Blocking a crawler can reduce how often your site appears in the AI product that uses it. Allowing training means your content may shape future models. There is no right answer, so decide what suits your site. Crawler names and behaviour change, so check each company’s current documentation from time to time.
Frequently asked questions
Will blocking GPTBot remove my pages from ChatGPT search? No. Search uses a different crawler, OAI-SearchBot.
Do I need an llms.txt file too? It is an optional, plain-language summary of a site for language models. It is separate from robots.txt and does not control access.
Try the tools
AI Crawler robots.txt Policy Builder
Choose which AI crawlers to allow or block and generate a robots.txt file, with a plain-English explanation of what each rule does.
botRobots.txt Generator
Build a robots.txt file with sitemap and AI crawler rules.
xmlXML Sitemap Generator
Turn a list of URLs into a valid sitemap.xml.