A few years ago, most site owners wrote their robots.txt file once and forgot about it. You blocked the admin folder, pointed Google at your sitemap, and moved on.
That changed when AI companies started sending their own crawlers across the web. Some of these bots collect pages to train language models. Others fetch pages so a chatbot can quote them in an answer. A third group only shows up when a real person asks an assistant to read a specific link.
Each of these bots has its own name, and each one checks your robots.txt file before it does anything else. So if you want a say in how AI tools use your content, this small text file is where you start.
This guide walks through how to write a robots.txt file for AI agents. You will learn which bots exist, what each one does, how to allow or block them one by one, and where robots.txt stops being useful. There are copy-and-paste examples for the most common setups, plus answers to the questions people ask most often.
- What Robots.txt Actually Does (and What It Doesn’t)
- The Three Kinds of AI Bots You Need to Know
- AI User-Agent Reference Table
- How to Create a Robots.txt File for AI Agents, Step by Step
- Ready-to-Use Robots.txt Examples
- How to Test Your File
- Where Robots.txt Falls Short, and What to Add
- Common Mistakes to Avoid
- Frequently Asked Questions
- Will blocking AI bots hurt my Google rankings?
- Does blocking an AI bot remove my content from existing AI models?
- Can I block AI training but still show up in ChatGPT and Claude answers?
- Do AI bots really follow robots.txt?
- What’s the difference between ClaudeBot, Claude-SearchBot and Claude-User?
- Can I use a wildcard to block every AI bot at once?
- How long does it take for changes to work?
- Is llms.txt a replacement for robots.txt?
- Do I need a robots.txt for each subdomain?
- Is it legal for AI companies to ignore my robots.txt?
What Robots.txt Actually Does (and What It Doesn’t)
Robots.txt is a plain text file that sits at the root of your domain, for example https://example.com/robots.txt. It tells automated clients which parts of your site they may crawl. The idea dates back to 1994, when Martijn Koster proposed it. It finally became an official internet standard, RFC 9309, in September 2022.
The file is made of groups. Each group starts with one or more User-agent lines naming a bot, followed by Allow and Disallow rules for paths on your site.
User-agent: GPTBot
Disallow: /
Those two lines ask OpenAI’s training crawler to stay off the entire site. That is really all there is to the syntax.
Here is the part many people miss. The standard itself says these rules are not a form of access authorization. Robots.txt is a request, not a lock. Well-known AI companies publish their bot names and say they follow the file. A scraper that ignores the rules faces no technical barrier at all.
So think of robots.txt as a clear, public statement of your preferences. It works well with companies that play by the rules. For everyone else you need other tools, which we cover later in this guide.
| Robots.txt can | Robots.txt can’t |
| Ask named bots to skip your whole site or certain folders | Physically stop a bot from loading a page |
| Treat training bots and search bots differently | Remove content that was already collected |
| Point crawlers to your sitemap | Hide a page from people who have the link |
| Give you a public record of your stated preferences | Force a bot to identify itself honestly |
The Three Kinds of AI Bots You Need to Know
The biggest mistake people make is treating every AI bot the same. They aren’t. Most large AI companies now run separate bots for separate jobs, and blocking one does nothing to the others.
OpenAI’s own documentation says each of its settings is independent. You can allow its search bot so your pages show up in ChatGPT search while blocking its training bot at the same time. Anthropic works the same way. Blocking ClaudeBot only stops training collection. It leaves Claude-SearchBot and Claude-User untouched.
In practice, AI bots fall into three groups.
| Bot type | What it does | When it visits | Examples |
| Training crawler | Collects pages that may be used to train future AI models | On a schedule, crawling across the web | GPTBot, ClaudeBot, CCBot, Applebot-Extended |
| Search / index crawler | Builds an index so an AI assistant can find and cite your pages in answers | On a schedule, similar to Googlebot | OAI-SearchBot, Claude-SearchBot, PerplexityBot |
| User-triggered fetcher | Loads one page because a person asked the assistant about it | Only when someone asks, in real time | ChatGPT-User, Claude-User, Perplexity-User |
There is also a fourth, odd one out: control tokens. Google-Extended is the best-known example. It never visits your site under that name. Google still crawls with its normal Googlebot. The token only tells Google whether that crawled content may be used for Gemini.
Why does this split matter? Because the trade-offs are very different. Blocking a training crawler costs you almost nothing in traffic. Blocking a search crawler can drop you out of AI answers, which is where a growing share of people now look things up. Blocking a user-triggered fetcher means that when a reader pastes your link into a chatbot, the chatbot may not be able to read it.
AI User-Agent Reference Table
This is the list you will actually copy from. The name in the first column is the exact token you put after User-agent: in your file. Matching is not case-sensitive, but spelling and hyphens have to be right.
| User-agent token | Company | Type | Follows robots.txt? |
| GPTBot | OpenAI | Training | Yes |
| OAI-SearchBot | OpenAI | Search | Yes |
| ChatGPT-User | OpenAI | User-triggered | May not apply, per OpenAI |
| ClaudeBot | Anthropic | Training | Yes |
| Claude-SearchBot | Anthropic | Search | Yes |
| Claude-User | Anthropic | User-triggered | Yes |
| Google-Extended | Control token (Gemini training and grounding) | Yes | |
| PerplexityBot | Perplexity | Search | Yes |
| Perplexity-User | Perplexity | User-triggered | Generally not |
| Applebot-Extended | Apple | Control token (Apple AI training) | Yes |
| meta-externalagent | Meta | Training | Yes, per Meta |
| CCBot | Common Crawl | Training dataset used by many AI labs | Yes |
| Bytespider | ByteDance | Training | Reports of poor compliance |
A few details are worth knowing before you write rules.
- Old Anthropic names no longer work. Before ClaudeBot, Anthropic used anthropic-ai and Claude-Web. A file written in 2023 that only blocks those two leaves today’s bots free to crawl.
- User-triggered bots are a grey area. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person started the request. Perplexity says much the same for Perplexity-User. Anthropic says Claude-User does follow the file.
- Blocking Google-Extended does not hurt your Google rankings. Google says it has no effect on inclusion or ranking in Google Search.
- Changes are not instant. OpenAI notes it can take about 24 hours for its systems to pick up a robots.txt change. Perplexity gives the same window.
New bots appear every few months. Check each company’s official page once or twice a year rather than trusting any list, this one included.
How to Create a Robots.txt File for AI Agents, Step by Step
You don’t need any special software. A plain text editor such as Notepad, TextEdit (in plain text mode) or VS Code is enough.
Step 1: Decide What You Actually Want
Before typing a single line, answer one question for each bot type. Do you want to be in AI training data? Do you want to be cited in AI search answers? Do you want chatbots to read your page when a user shares the link?
Most businesses and bloggers land on the same answer: no to training, yes to search, yes to user fetches. News publishers and paywalled sites often say no to all three. Documentation sites and open-source projects usually say yes to everything, because being quoted by AI tools helps their users.
Each bot type is a separate yes-or-no choice. Your answers tell you exactly which groups to write in the next step.
Step 2: Check What You Already Have
Open https://yourdomain.com/robots.txt in a browser. If a file loads, copy its contents somewhere safe before you change anything. Many sites already have rules for Googlebot and a sitemap line, and you want to keep those.
If you get a 404 page, you have no file yet. That means every bot is currently allowed everywhere.
Step 3: Write One Group per Bot
Each group starts with User-agent: and is followed by its rules. You can stack several user-agent lines on top of one set of rules if they should all be treated the same way.
# Block AI training crawlers
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /
To allow a bot, either leave it out of the file completely or give it an explicit rule:
User-agent: OAI-SearchBot
Allow: /
The explicit version is better. It documents your choice, and it protects you if you later add a broad block for all bots.
Step 4: Understand How Groups Are Matched
This rule trips up a lot of people. A bot reads the whole file, finds the group that names it most specifically, and follows only that group. It ignores everything under User-agent: * if it finds its own name.
So if you block /private/ for all bots and then write a separate group for GPTBot, GPTBot will not see the /private/ rule. You have to repeat it inside the GPTBot group.
Within a group, the longest matching path wins. Allow: /blog/ beats Disallow: / for any URL that starts with /blog/.
Step 5: Block Folders Instead of the Whole Site (Optional)
You don’t have to go all or nothing. Here is a setup that lets AI search bots see your blog but keeps them away from members-only content and internal search pages:
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /blog/
Disallow: /members/
Disallow: /search
Step 6: Add Your Sitemap
Put a Sitemap: line at the bottom of the file. It isn’t tied to any group, so every crawler that respects the file will see it.
Sitemap: https://yourdomain.com/sitemap.xml
Step 7: Upload It to the Root of Every Host
Save the file as robots.txt, all lowercase, in plain UTF-8 text. Upload it to the top level of your site so it loads at https://yourdomain.com/robots.txt. A file inside a subfolder is ignored.
Each subdomain needs its own copy. shop.yourdomain.com and blog.yourdomain.com do not read the file on your main domain. Anthropic’s help page calls this out directly for anyone opting out of its bots.
On WordPress, plugins such as Yoast SEO or Rank Math let you edit robots.txt from the dashboard. On Shopify, you edit the robots.txt.liquid template. On Cloudflare-hosted sites, check whether Cloudflare is already adding managed AI rules before you add your own, so the two don’t conflict.
Ready-to-Use Robots.txt Examples
Pick the one closest to your situation, swap in your own domain and folders, and upload it. Lines starting with # are comments. Bots skip them, but they help the next person who opens the file.
Example 1: Stay Visible in AI Search, Opt Out of Training
This is the setup most businesses choose. Your pages can still be found and cited by ChatGPT, Claude and Perplexity, but you ask the training crawlers to stay away.
# — AI training: blocked —
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
User-agent: Bytespider
Disallow: /
# — AI search and user fetches: allowed —
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
Disallow: /admin/
Disallow: /cart/
# — Everyone else —
User-agent: *
Disallow: /admin/
Disallow: /cart/
Sitemap: https://example.com/sitemap.xml
One catch here: Google-Extended also covers grounding in Gemini apps, not only training. If showing up in Gemini answers matters to you, think twice before blocking it.
Example 2: Block All Known AI Bots
For publishers and anyone who does not want AI tools using their work at all.
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: meta-externalagent
User-agent: CCBot
User-agent: Bytespider
Disallow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Normal search engines like Googlebot and Bingbot are still allowed by the last group. Your regular SEO stays as it was.
Example 3: Welcome All AI Bots
For documentation sites, open-source projects and brands that want maximum reach inside AI tools.
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Google-Extended
Allow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Technically, a file with no rules does the same job. Writing the names out makes your choice clear to anyone who checks.
Example 4: Share Some Sections, Protect Others
For sites with a mix of public and premium content, such as a SaaS company with a public help center and a paid course area.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
Allow: /docs/
Allow: /blog/
Disallow: /
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Disallow: /courses/
Disallow: /account/
User-agent: *
Disallow: /account/
Sitemap: https://example.com/sitemap.xml
Here the training bots may read the docs and blog but nothing else. The search bots can see everything except paid courses and account pages.
Example 5: Slow Bots Down Instead of Blocking Them
If AI crawlers are hammering your server, you may not want to shut them out. Anthropic supports the non-standard Crawl-delay directive, which asks for a pause between requests.
User-agent: ClaudeBot
Crawl-delay: 5
Not every company honors Crawl-delay, and Google ignores it. For real rate limiting, your host or CDN is the better place to set it up.
How to Test Your File
A typo in robots.txt fails silently. Nothing breaks on the page, so you only find out weeks later when a bot you meant to block shows up in your logs. Spend ten minutes testing.
- Load the live file. Visit https://yourdomain.com/robots.txt in a private browser window. Make sure you see the new version, not a cached copy from your CDN.
- Check the response code. The file should return 200 OK as plain text. A redirect loop or a 5xx error can cause some crawlers to treat your whole site as blocked, while a 404 is read as “no rules at all”.
- Run it through a parser. Google Search Console has a robots.txt report that shows the version Google fetched and flags errors. For AI bots specifically, most SEO crawlers (Screaming Frog, Sitebulb and similar) let you set a custom user agent such as GPTBot and see which URLs it would be allowed to reach.
- Watch your server logs. After a day or two, search your access logs for the bot names. Blocked bots should only be requesting /robots.txt and nothing else.
- Verify the visitor is real. Anyone can put “GPTBot” in a user-agent string. OpenAI, Anthropic and Perplexity publish the IP ranges their bots use. If a request claims to be GPTBot but comes from an IP outside OpenAI’s published list, it is an impostor.
| Test | What a good result looks like |
| Browser check | File loads at the root, shows your latest edits |
| Status code | 200, content type text/plain |
| Parser or tester | No syntax warnings; blocked bots show as disallowed on sample URLs |
| Server logs | Blocked bots fetch only /robots.txt |
| IP check | Bot requests come from the vendor’s published IP ranges |
Where Robots.txt Falls Short, and What to Add
Robots.txt answers one question: may this bot fetch this URL? It can’t say “read this, but don’t train on it”. It can’t stop a bot that lies about its name. And it does nothing about content that was collected before you changed the file.
Several newer tools try to fill those gaps. They work alongside robots.txt, not instead of it.
| Tool | What it does | Enforced? | Good for |
| Content-Usage rule (IETF draft) | Adds a line such as Content-Usage: train-ai=n to robots.txt to state how content may be used, not just fetched | No, voluntary | Saying “crawl, but don’t train” in a standard way |
| Cloudflare Content Signals | A comment-style policy in robots.txt, such as search=yes, ai-train=no | No, voluntary | Cloudflare sites that want a ready-made policy |
| llms.txt | A Markdown file at /llms.txt that points AI tools to your most useful pages | No, it only guides | Docs and product sites that want to be read well |
| Firewall or CDN bot rules | Blocks or rate-limits requests by user agent and verified IP | Yes | Stopping bots that ignore robots.txt |
| Login or paywall | Keeps content away from any visitor without an account | Yes | Premium or private content |
The IETF AI Preferences Work
The IETF has a working group, AIPREF, building a shared vocabulary for AI usage. Its attachment draft, revised in August 2026, defines a Content-Usage rule for robots.txt and a matching HTTP header. It is still a draft, so support among AI companies is limited for now. Adding the line costs nothing, though, and it puts your preference on record.
User-agent: *
Allow: /
Content-Usage: train-ai=n
llms.txt Is Not a Blocking Tool
People often confuse llms.txt with robots.txt. They do opposite jobs. Robots.txt tells bots where not to go. llms.txt is a friendly map that tells AI tools where your best content lives, usually as a short Markdown list of links. Use it if you want AI tools to understand your site. It won’t keep anyone out.
When You Need Real Enforcement
If you see a bot ignoring your file, move the rule to your server or CDN. Cloudflare, Fastly, Akamai and most managed hosts can block by user agent. Better still, block by IP, since the big AI companies publish the ranges their bots use and a fake bot can’t easily spoof those. Keep robots.txt in place too. It remains the clearest public record of what you asked for.
Common Mistakes to Avoid
Most broken robots.txt files break in the same handful of ways.
| Mistake | What happens | Fix |
| Blocking only GPTBot and calling it done | ClaudeBot, CCBot and others keep crawling | Name every bot you care about |
| Using old names like anthropic-ai or Claude-Web | Current Anthropic bots ignore the rule | Switch to ClaudeBot, Claude-SearchBot, Claude-User |
| Blocking search bots by accident | Your pages vanish from ChatGPT, Claude or Perplexity answers | Block training bots only, unless you truly want out |
| Relying on the * group for AI bots that also have their own group | The bot follows its own group and skips the * rules | Repeat shared rules inside each named group |
| Disallow: / under User-agent: * | Google and Bing drop your whole site | Only use this if you really want no search traffic |
| File saved as Robots.txt or placed in a subfolder | Bots can’t find it, so nothing is blocked | Lowercase name, root folder, every subdomain |
| Hiding secret pages with robots.txt | The file is public, so it points people straight at them | Use a login or noindex instead |
| Never updating the file | New bots crawl freely | Review it every six months |
The one about hiding pages deserves a second mention. Anyone can open your robots.txt. Listing /secret-launch-page/ in it is like putting up a sign that says “nothing to see behind this door”.
Frequently Asked Questions
Will blocking AI bots hurt my Google rankings?
No, as long as you only block AI-specific names. Googlebot is a separate crawler. Google also says that blocking Google-Extended has no effect on inclusion or ranking in Google Search. Just don’t block Googlebot itself or use Disallow: / under User-agent: *.
Does blocking an AI bot remove my content from existing AI models?
No. Robots.txt only affects future crawling. Anthropic, for example, says blocking ClaudeBot signals that your site’s future material should be left out of training. Anything collected earlier is not pulled back out.
Can I block AI training but still show up in ChatGPT and Claude answers?
Yes. Block the training crawlers (GPTBot, ClaudeBot) and allow the search crawlers (OAI-SearchBot, Claude-SearchBot). Example 1 above does exactly this.
Do AI bots really follow robots.txt?
The major ones say they do, and publish their bot names so you can check. User-triggered fetchers are the exception. OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it. Smaller or unknown scrapers may ignore the file entirely.
What’s the difference between ClaudeBot, Claude-SearchBot and Claude-User?
ClaudeBot collects pages that may be used for training. Claude-SearchBot indexes pages to improve Claude’s search results. Claude-User fetches a page when a person asks Claude about it. Each one reads its own group in your file.
Can I use a wildcard to block every AI bot at once?
Not really. There is no shared “AI” token. User-agent: * would block every bot, search engines included. You have to list each AI bot by name.
How long does it take for changes to work?
Usually about a day. OpenAI and Perplexity both mention a window of roughly 24 hours. Other crawlers cache the file for their own periods, so give it a few days before you judge.
Is llms.txt a replacement for robots.txt?
No. llms.txt helps AI tools find and understand your content. It has no blocking power. Use robots.txt for access rules and llms.txt, if you want it, as a guide.
Do I need a robots.txt for each subdomain?
Yes. Every host reads only its own file. blog.example.com will not follow rules from example.com/robots.txt.
Is it legal for AI companies to ignore my robots.txt?
That depends on your country and on the case, and courts are still working it out. Robots.txt is not a contract on its own. If this matters for your business, talk to a lawyer. Meanwhile, back up your robots.txt rules with firewall blocks and clear terms of use on your site.
Also See: