Robots.txt for AI Agents: How To Create One?

5/5 - (6 votes)

A few years ago, most site owners wrote their robots.txt file once and forgot about it. You blocked the admin folder, pointed Google at your sitemap, and moved on.

That changed when AI companies started sending their own crawlers across the web. Some of these bots collect pages to train language models. Others fetch pages so a chatbot can quote them in an answer. A third group only shows up when a real person asks an assistant to read a specific link.

Each of these bots has its own name, and each one checks your robots.txt file before it does anything else. So if you want a say in how AI tools use your content, this small text file is where you start.

This guide walks through how to write a robots.txt file for AI agents. You will learn which bots exist, what each one does, how to allow or block them one by one, and where robots.txt stops being useful. There are copy-and-paste examples for the most common setups, plus answers to the questions people ask most often.

What Robots.txt Actually Does (and What It Doesn’t)

robots file for AI agents

Robots.txt is a plain text file that sits at the root of your domain, for example https://example.com/robots.txt. It tells automated clients which parts of your site they may crawl. The idea dates back to 1994, when Martijn Koster proposed it. It finally became an official internet standard, RFC 9309, in September 2022.

The file is made of groups. Each group starts with one or more User-agent lines naming a bot, followed by Allow and Disallow rules for paths on your site.

User-agent: GPTBot

Disallow: /

Those two lines ask OpenAI’s training crawler to stay off the entire site. That is really all there is to the syntax.

Here is the part many people miss. The standard itself says these rules are not a form of access authorization. Robots.txt is a request, not a lock. Well-known AI companies publish their bot names and say they follow the file. A scraper that ignores the rules faces no technical barrier at all.

So think of robots.txt as a clear, public statement of your preferences. It works well with companies that play by the rules. For everyone else you need other tools, which we cover later in this guide.

Robots.txt canRobots.txt can’t
Ask named bots to skip your whole site or certain foldersPhysically stop a bot from loading a page
Treat training bots and search bots differentlyRemove content that was already collected
Point crawlers to your sitemapHide a page from people who have the link
Give you a public record of your stated preferencesForce a bot to identify itself honestly

The Three Kinds of AI Bots You Need to Know

The biggest mistake people make is treating every AI bot the same. They aren’t. Most large AI companies now run separate bots for separate jobs, and blocking one does nothing to the others.

OpenAI’s own documentation says each of its settings is independent. You can allow its search bot so your pages show up in ChatGPT search while blocking its training bot at the same time. Anthropic works the same way. Blocking ClaudeBot only stops training collection. It leaves Claude-SearchBot and Claude-User untouched.

In practice, AI bots fall into three groups.

Bot typeWhat it doesWhen it visitsExamples
Training crawlerCollects pages that may be used to train future AI modelsOn a schedule, crawling across the webGPTBot, ClaudeBot, CCBot, Applebot-Extended
Search / index crawlerBuilds an index so an AI assistant can find and cite your pages in answersOn a schedule, similar to GooglebotOAI-SearchBot, Claude-SearchBot, PerplexityBot
User-triggered fetcherLoads one page because a person asked the assistant about itOnly when someone asks, in real timeChatGPT-User, Claude-User, Perplexity-User

There is also a fourth, odd one out: control tokens. Google-Extended is the best-known example. It never visits your site under that name. Google still crawls with its normal Googlebot. The token only tells Google whether that crawled content may be used for Gemini.

Why does this split matter? Because the trade-offs are very different. Blocking a training crawler costs you almost nothing in traffic. Blocking a search crawler can drop you out of AI answers, which is where a growing share of people now look things up. Blocking a user-triggered fetcher means that when a reader pastes your link into a chatbot, the chatbot may not be able to read it.

AI User-Agent Reference Table

This is the list you will actually copy from. The name in the first column is the exact token you put after User-agent: in your file. Matching is not case-sensitive, but spelling and hyphens have to be right.

User-agent tokenCompanyTypeFollows robots.txt?
GPTBotOpenAITrainingYes
OAI-SearchBotOpenAISearchYes
ChatGPT-UserOpenAIUser-triggeredMay not apply, per OpenAI
ClaudeBotAnthropicTrainingYes
Claude-SearchBotAnthropicSearchYes
Claude-UserAnthropicUser-triggeredYes
Google-ExtendedGoogleControl token (Gemini training and grounding)Yes
PerplexityBotPerplexitySearchYes
Perplexity-UserPerplexityUser-triggeredGenerally not
Applebot-ExtendedAppleControl token (Apple AI training)Yes
meta-externalagentMetaTrainingYes, per Meta
CCBotCommon CrawlTraining dataset used by many AI labsYes
BytespiderByteDanceTrainingReports of poor compliance

A few details are worth knowing before you write rules.

  • Old Anthropic names no longer work. Before ClaudeBot, Anthropic used anthropic-ai and Claude-Web. A file written in 2023 that only blocks those two leaves today’s bots free to crawl.
  • User-triggered bots are a grey area. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person started the request. Perplexity says much the same for Perplexity-User. Anthropic says Claude-User does follow the file.
  • Blocking Google-Extended does not hurt your Google rankings. Google says it has no effect on inclusion or ranking in Google Search.
  • Changes are not instant. OpenAI notes it can take about 24 hours for its systems to pick up a robots.txt change. Perplexity gives the same window.

New bots appear every few months. Check each company’s official page once or twice a year rather than trusting any list, this one included.

How to Create a Robots.txt File for AI Agents, Step by Step

You don’t need any special software. A plain text editor such as Notepad, TextEdit (in plain text mode) or VS Code is enough.

Step 1: Decide What You Actually Want

Before typing a single line, answer one question for each bot type. Do you want to be in AI training data? Do you want to be cited in AI search answers? Do you want chatbots to read your page when a user shares the link?

Most businesses and bloggers land on the same answer: no to training, yes to search, yes to user fetches. News publishers and paywalled sites often say no to all three. Documentation sites and open-source projects usually say yes to everything, because being quoted by AI tools helps their users.

Each bot type is a separate yes-or-no choice. Your answers tell you exactly which groups to write in the next step.

Step 2: Check What You Already Have

Open https://yourdomain.com/robots.txt in a browser. If a file loads, copy its contents somewhere safe before you change anything. Many sites already have rules for Googlebot and a sitemap line, and you want to keep those.

If you get a 404 page, you have no file yet. That means every bot is currently allowed everywhere.

Step 3: Write One Group per Bot

Each group starts with User-agent: and is followed by its rules. You can stack several user-agent lines on top of one set of rules if they should all be treated the same way.

# Block AI training crawlers

User-agent: GPTBot

User-agent: ClaudeBot

User-agent: Google-Extended

User-agent: CCBot

Disallow: /

To allow a bot, either leave it out of the file completely or give it an explicit rule:

User-agent: OAI-SearchBot

Allow: /

The explicit version is better. It documents your choice, and it protects you if you later add a broad block for all bots.

Step 4: Understand How Groups Are Matched

This rule trips up a lot of people. A bot reads the whole file, finds the group that names it most specifically, and follows only that group. It ignores everything under User-agent: * if it finds its own name.

So if you block /private/ for all bots and then write a separate group for GPTBot, GPTBot will not see the /private/ rule. You have to repeat it inside the GPTBot group.

Within a group, the longest matching path wins. Allow: /blog/ beats Disallow: / for any URL that starts with /blog/.

Step 5: Block Folders Instead of the Whole Site (Optional)

You don’t have to go all or nothing. Here is a setup that lets AI search bots see your blog but keeps them away from members-only content and internal search pages:

User-agent: OAI-SearchBot

User-agent: Claude-SearchBot

User-agent: PerplexityBot

Allow: /blog/

Disallow: /members/

Disallow: /search

Step 6: Add Your Sitemap

Put a Sitemap: line at the bottom of the file. It isn’t tied to any group, so every crawler that respects the file will see it.

Sitemap: https://yourdomain.com/sitemap.xml

Step 7: Upload It to the Root of Every Host

Save the file as robots.txt, all lowercase, in plain UTF-8 text. Upload it to the top level of your site so it loads at https://yourdomain.com/robots.txt. A file inside a subfolder is ignored.

Each subdomain needs its own copy. shop.yourdomain.com and blog.yourdomain.com do not read the file on your main domain. Anthropic’s help page calls this out directly for anyone opting out of its bots.

On WordPress, plugins such as Yoast SEO or Rank Math let you edit robots.txt from the dashboard. On Shopify, you edit the robots.txt.liquid template. On Cloudflare-hosted sites, check whether Cloudflare is already adding managed AI rules before you add your own, so the two don’t conflict.

Ready-to-Use Robots.txt Examples

Pick the one closest to your situation, swap in your own domain and folders, and upload it. Lines starting with # are comments. Bots skip them, but they help the next person who opens the file.

Example 1: Stay Visible in AI Search, Opt Out of Training

This is the setup most businesses choose. Your pages can still be found and cited by ChatGPT, Claude and Perplexity, but you ask the training crawlers to stay away.

# — AI training: blocked —

User-agent: GPTBot

User-agent: ClaudeBot

User-agent: Google-Extended

User-agent: Applebot-Extended

User-agent: meta-externalagent

User-agent: CCBot

User-agent: Bytespider

Disallow: /

# — AI search and user fetches: allowed —

User-agent: OAI-SearchBot

User-agent: ChatGPT-User

User-agent: Claude-SearchBot

User-agent: Claude-User

User-agent: PerplexityBot

User-agent: Perplexity-User

Allow: /

Disallow: /admin/

Disallow: /cart/

# — Everyone else —

User-agent: *

Disallow: /admin/

Disallow: /cart/

Sitemap: https://example.com/sitemap.xml

One catch here: Google-Extended also covers grounding in Gemini apps, not only training. If showing up in Gemini answers matters to you, think twice before blocking it.

Example 2: Block All Known AI Bots

For publishers and anyone who does not want AI tools using their work at all.

User-agent: GPTBot

User-agent: OAI-SearchBot

User-agent: ChatGPT-User

User-agent: ClaudeBot

User-agent: Claude-SearchBot

User-agent: Claude-User

User-agent: Google-Extended

User-agent: Applebot-Extended

User-agent: PerplexityBot

User-agent: Perplexity-User

User-agent: meta-externalagent

User-agent: CCBot

User-agent: Bytespider

Disallow: /

User-agent: *

Allow: /

Sitemap: https://example.com/sitemap.xml

Normal search engines like Googlebot and Bingbot are still allowed by the last group. Your regular SEO stays as it was.

Example 3: Welcome All AI Bots

For documentation sites, open-source projects and brands that want maximum reach inside AI tools.

User-agent: GPTBot

User-agent: OAI-SearchBot

User-agent: ClaudeBot

User-agent: Claude-SearchBot

User-agent: PerplexityBot

User-agent: Google-Extended

Allow: /

User-agent: *

Allow: /

Sitemap: https://example.com/sitemap.xml

Technically, a file with no rules does the same job. Writing the names out makes your choice clear to anyone who checks.

Example 4: Share Some Sections, Protect Others

For sites with a mix of public and premium content, such as a SaaS company with a public help center and a paid course area.

User-agent: GPTBot

User-agent: ClaudeBot

User-agent: CCBot

Allow: /docs/

Allow: /blog/

Disallow: /

User-agent: OAI-SearchBot

User-agent: Claude-SearchBot

User-agent: PerplexityBot

Disallow: /courses/

Disallow: /account/

User-agent: *

Disallow: /account/

Sitemap: https://example.com/sitemap.xml

Here the training bots may read the docs and blog but nothing else. The search bots can see everything except paid courses and account pages.

Example 5: Slow Bots Down Instead of Blocking Them

If AI crawlers are hammering your server, you may not want to shut them out. Anthropic supports the non-standard Crawl-delay directive, which asks for a pause between requests.

User-agent: ClaudeBot

Crawl-delay: 5

Not every company honors Crawl-delay, and Google ignores it. For real rate limiting, your host or CDN is the better place to set it up.

How to Test Your File

A typo in robots.txt fails silently. Nothing breaks on the page, so you only find out weeks later when a bot you meant to block shows up in your logs. Spend ten minutes testing.

  1. Load the live file. Visit https://yourdomain.com/robots.txt in a private browser window. Make sure you see the new version, not a cached copy from your CDN.
  2. Check the response code. The file should return 200 OK as plain text. A redirect loop or a 5xx error can cause some crawlers to treat your whole site as blocked, while a 404 is read as “no rules at all”.
  3. Run it through a parser. Google Search Console has a robots.txt report that shows the version Google fetched and flags errors. For AI bots specifically, most SEO crawlers (Screaming Frog, Sitebulb and similar) let you set a custom user agent such as GPTBot and see which URLs it would be allowed to reach.
  4. Watch your server logs. After a day or two, search your access logs for the bot names. Blocked bots should only be requesting /robots.txt and nothing else.
  5. Verify the visitor is real. Anyone can put “GPTBot” in a user-agent string. OpenAI, Anthropic and Perplexity publish the IP ranges their bots use. If a request claims to be GPTBot but comes from an IP outside OpenAI’s published list, it is an impostor.
TestWhat a good result looks like
Browser checkFile loads at the root, shows your latest edits
Status code200, content type text/plain
Parser or testerNo syntax warnings; blocked bots show as disallowed on sample URLs
Server logsBlocked bots fetch only /robots.txt
IP checkBot requests come from the vendor’s published IP ranges

Where Robots.txt Falls Short, and What to Add

Robots.txt answers one question: may this bot fetch this URL? It can’t say “read this, but don’t train on it”. It can’t stop a bot that lies about its name. And it does nothing about content that was collected before you changed the file.

Several newer tools try to fill those gaps. They work alongside robots.txt, not instead of it.

ToolWhat it doesEnforced?Good for
Content-Usage rule (IETF draft)Adds a line such as Content-Usage: train-ai=n to robots.txt to state how content may be used, not just fetchedNo, voluntarySaying “crawl, but don’t train” in a standard way
Cloudflare Content SignalsA comment-style policy in robots.txt, such as search=yes, ai-train=noNo, voluntaryCloudflare sites that want a ready-made policy
llms.txtA Markdown file at /llms.txt that points AI tools to your most useful pagesNo, it only guidesDocs and product sites that want to be read well
Firewall or CDN bot rulesBlocks or rate-limits requests by user agent and verified IPYesStopping bots that ignore robots.txt
Login or paywallKeeps content away from any visitor without an accountYesPremium or private content

The IETF AI Preferences Work

The IETF has a working group, AIPREF, building a shared vocabulary for AI usage. Its attachment draft, revised in August 2026, defines a Content-Usage rule for robots.txt and a matching HTTP header. It is still a draft, so support among AI companies is limited for now. Adding the line costs nothing, though, and it puts your preference on record.

User-agent: *

Allow: /

Content-Usage: train-ai=n

llms.txt Is Not a Blocking Tool

People often confuse llms.txt with robots.txt. They do opposite jobs. Robots.txt tells bots where not to go. llms.txt is a friendly map that tells AI tools where your best content lives, usually as a short Markdown list of links. Use it if you want AI tools to understand your site. It won’t keep anyone out.

When You Need Real Enforcement

If you see a bot ignoring your file, move the rule to your server or CDN. Cloudflare, Fastly, Akamai and most managed hosts can block by user agent. Better still, block by IP, since the big AI companies publish the ranges their bots use and a fake bot can’t easily spoof those. Keep robots.txt in place too. It remains the clearest public record of what you asked for.

Common Mistakes to Avoid

Most broken robots.txt files break in the same handful of ways.

MistakeWhat happensFix
Blocking only GPTBot and calling it doneClaudeBot, CCBot and others keep crawlingName every bot you care about
Using old names like anthropic-ai or Claude-WebCurrent Anthropic bots ignore the ruleSwitch to ClaudeBot, Claude-SearchBot, Claude-User
Blocking search bots by accidentYour pages vanish from ChatGPT, Claude or Perplexity answersBlock training bots only, unless you truly want out
Relying on the * group for AI bots that also have their own groupThe bot follows its own group and skips the * rulesRepeat shared rules inside each named group
Disallow: / under User-agent: *Google and Bing drop your whole siteOnly use this if you really want no search traffic
File saved as Robots.txt or placed in a subfolderBots can’t find it, so nothing is blockedLowercase name, root folder, every subdomain
Hiding secret pages with robots.txtThe file is public, so it points people straight at themUse a login or noindex instead
Never updating the fileNew bots crawl freelyReview it every six months

The one about hiding pages deserves a second mention. Anyone can open your robots.txt. Listing /secret-launch-page/ in it is like putting up a sign that says “nothing to see behind this door”.

Frequently Asked Questions

Will blocking AI bots hurt my Google rankings?

No, as long as you only block AI-specific names. Googlebot is a separate crawler. Google also says that blocking Google-Extended has no effect on inclusion or ranking in Google Search. Just don’t block Googlebot itself or use Disallow: / under User-agent: *.

Does blocking an AI bot remove my content from existing AI models?

No. Robots.txt only affects future crawling. Anthropic, for example, says blocking ClaudeBot signals that your site’s future material should be left out of training. Anything collected earlier is not pulled back out.

Can I block AI training but still show up in ChatGPT and Claude answers?

Yes. Block the training crawlers (GPTBot, ClaudeBot) and allow the search crawlers (OAI-SearchBot, Claude-SearchBot). Example 1 above does exactly this.

Do AI bots really follow robots.txt?

The major ones say they do, and publish their bot names so you can check. User-triggered fetchers are the exception. OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it. Smaller or unknown scrapers may ignore the file entirely.

What’s the difference between ClaudeBot, Claude-SearchBot and Claude-User?

ClaudeBot collects pages that may be used for training. Claude-SearchBot indexes pages to improve Claude’s search results. Claude-User fetches a page when a person asks Claude about it. Each one reads its own group in your file.

Can I use a wildcard to block every AI bot at once?

Not really. There is no shared “AI” token. User-agent: * would block every bot, search engines included. You have to list each AI bot by name.

How long does it take for changes to work?

Usually about a day. OpenAI and Perplexity both mention a window of roughly 24 hours. Other crawlers cache the file for their own periods, so give it a few days before you judge.

Is llms.txt a replacement for robots.txt?

No. llms.txt helps AI tools find and understand your content. It has no blocking power. Use robots.txt for access rules and llms.txt, if you want it, as a guide.

Do I need a robots.txt for each subdomain?

Yes. Every host reads only its own file. blog.example.com will not follow rules from example.com/robots.txt.

Is it legal for AI companies to ignore my robots.txt?

That depends on your country and on the case, and courts are still working it out. Robots.txt is not a contract on its own. If this matters for your business, talk to a lawyer. Meanwhile, back up your robots.txt rules with firewall blocks and clear terms of use on your site.

Also See:

List of AI SEO MythsQuestions To Ask An AI SEO CompanyAI SEO Services For Ecommerce
AI SEO SkillsSmall Business AI SEO SoftwareFoundation of AI

Add Comment