Free robots.txt Generator for AI Crawlers
Build a robots.txt that allows or blocks GPTBot, ClaudeBot, PerplexityBot and other AI crawlers, or check which AI bots a site already blocks. Free.
| Crawler | Used for | Rule |
|---|
How to use the robots.txt generator
- Pick a starting point. Allow all keeps every AI crawler welcome, Block AI training only stops bots that collect data to train models while leaving AI search bots alone, and Block all AI bots shuts them all out.
- Change any single bot with its Allow or Block menu.
- Add any folders you want every crawler to skip, such as
/wp-admin/or/cart/, and your sitemap URL. - Download
robots.txtand upload it to the root of your site, so it loads atyoursite.com/robots.txt.
Switch to Check a site to see which AI crawlers any live robots.txt allows or blocks. It reads the file the same way Google describes: the most specific rule wins, and a bot with its own group ignores the * rules.
Training bots, search bots and user bots
AI companies now run several crawlers each, and they do different jobs. Blocking the wrong one costs you visibility without protecting anything.
| Type | What it does | Examples | Effect of blocking |
|---|---|---|---|
| AI training | Collects pages to train future models | GPTBot, ClaudeBot, CCBot, meta-externalagent | Your content is less likely to end up in training data. No direct effect on AI search answers today |
| AI search | Builds the index an assistant searches when it answers | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Your pages can stop appearing as cited sources in that assistant's answers |
| User fetches | Loads a page because a person asked the assistant to | ChatGPT-User, Claude-User, Perplexity-User | The assistant can't open your page when someone pastes the link. Some of these say they may not follow robots.txt at all |
| Training controls | Not crawlers. Tokens that tell an existing crawler how its data may be used | Google-Extended, Applebot-Extended | Opts you out of model training without touching search |
For most businesses that want to be found, the sensible default is to allow the search and user bots. Whether to block training bots is a business decision about your content, not an SEO one.
The Googlebot trap
Google's AI Overviews and AI Mode are part of Google Search, and Google says they are controlled by Googlebot. Blocking Googlebot to keep out of AI Overviews removes the site from Google Search entirely.
Google-Extended is a different control. It covers Gemini model training and grounding outside Search, and Google says it has no effect on inclusion or ranking in Search. Blocking it does not keep you out of AI Overviews. To limit what Google shows in AI features, use nosnippet, data-nosnippet or max-snippet on the page instead.
What robots.txt can't do
- It is a request, not a lock. Well-behaved crawlers follow it, but nothing forces a bot to.
- It controls crawling, not indexing. A blocked page can still appear in search results if other sites link to it, just without a description. Use
noindexto keep a page out. - Some bots that fetch pages on a user's behalf say they may ignore robots.txt, because a person asked for the page. The generator flags these.
- Google ignores
crawl-delay. Some other crawlers, including ClaudeBot, do follow it. - Google reads only the first 500 KiB of the file and ignores the rest.
- Paths are case-sensitive, so
/Blog/and/blog/are different rules.
What a check result means
| Status | Meaning |
|---|---|
| Allowed | No rule stops that bot from crawling the site |
| Partly blocked | The bot can crawl the homepage but some paths are disallowed |
| Blocked | The rules stop the bot from crawling the whole site |
"Via its own rules" means the file has a group for that bot by name. "Via the * rules" means it falls back to the catch-all group. If a site has no robots.txt at all (a 404), every crawler is allowed. If the file returns a server error, Google pauses crawling for a while, which is worth fixing quickly.
Frequently asked questions
Should I block AI crawlers? Blocking AI search bots trades away a growing source of visibility and referral traffic. Blocking training bots has no clear SEO cost, so it comes down to how you feel about your content training models. Most sites that sell something are better off allowed.
Does blocking GPTBot keep me out of ChatGPT search? No. ChatGPT's search results come from OAI-SearchBot, and pages a user asks about are fetched by ChatGPT-User. GPTBot is the training crawler.
Will blocking AI bots hurt my Google rankings? Not unless you block Googlebot. None of the other AI crawlers affect Google Search.
How often does the bot list change? Often. The list here was checked against each company's own crawler documentation in October 2026. New bots appear every few months, so recheck your file now and then.
For the bigger picture on getting cited by AI assistants, read LLM SEO and how AI Overviews choose sources. If you also want to publish an llms.txt file, use the llms.txt generator.

