AI visibility

The 6 crawlers your robots.txt should let in

They come in pairs — one builds the AI assistant’s search index, the other goes and gets your page while a customer is sat there waiting. Six names, three companies. They are not the crawlers that collect training data, and plenty of trade sites block the lot by accident.

Free check

See where you turn up when a customer searches your trade in your suburb.

Type your own domain into a browser and put /robots.txt on the end. Whatever loads is the file every well-behaved bot reads before it touches anything else on your site. If you have never looked at yours, now is a good time.

There are 3 kinds of bot and they are not interchangeable

Lumping them together is how sites end up blocking the wrong one.

1. Training crawlers

These collect pages in bulk to build a model. Nothing about your business is being looked up when one visits. The names published today include GPTBot from OpenAI, ClaudeBot from Anthropic, CCBot from Common Crawl, and Google-Extended, which is Google’s separate control for training rather than a crawler in its own right.

Whether you allow them is a business decision about your words ending up in a model. It is a reasonable thing to say no to.

2. Search crawlers

These build the AI assistant’s own index of the web, which is what it searches when a customer asks a question. That is OAI-SearchBot, Claude-SearchBot and PerplexityBot.

Block these and you are not in the pool it draws the answer from. Not ranked low. Absent.

3. Live fetchers

Somebody asks a question right now and the AI assistant goes out and reads pages to answer it. Different bot, different name, separate rule: ChatGPT-User, Claude-User and Perplexity-User.

So it is six, in three pairs, and each company publishes its own list: OpenAI, Anthropic and Perplexity. Allowing the search half and blocking the live half means you are in the index and absent from the answer, which is the worst of both. Checked against all three on 25 August 2026.

One thing to be straight about: these names come from what each company publishes, and they change. New ones appear, old ones get retired. Any list written down in August is a snapshot, this one included. The stable part is the shape of it, which is that fetching-to-answer and collecting-for-training are separate jobs done by separately named bots, and a rule about one has no effect on the other.

Reading yours in about 30 seconds

The file is a stack of blocks. Each block starts with a User-agent line naming a bot, then Disallow lines listing what that bot is not to touch.

  • Disallow followed by a single slash means the whole site. That is a full block.
  • Disallow with nothing after it means nothing is blocked. It reads like a rule and it is permission.
  • User-agent with an asterisk is the catch-all, applying to any bot with no block of its own.
  • A block naming a specific bot wins over the catch-all for that bot. So a site can allow everything in general and still have 1 named crawler shut out further down the file.
  • A Sitemap line pointing at your sitemap belongs in there too. It costs nothing and it is the one line that helps rather than restricts.

If nothing loads at all and you get a 404, that is not a disaster. No file means no restrictions. It also means no sitemap line, and it usually means nobody has thought about any of this.

The 4 ways a trade site blocks them without meaning to

  1. The staging tick nobody unticked. WordPress has a setting that discourages search engines while a site is being built. It writes a site-wide block. It is supposed to be switched off at launch and sometimes it is not.
  2. A security plugin’s bad-bot list. Plenty of them ship with a blocklist maintained by the plugin vendor, and AI crawlers went onto those lists as a default. Nobody chose it. It arrived in an update.
  3. Bot-fight mode at the edge. This one is nastier because your robots.txt can look perfect. A firewall or CDN (a service that sits in front of your site and serves it faster) refuses anything that does not look like a person in a browser, and returns an error before your file is ever served. The file is a note on the door. The firewall is a lock on it.
  4. A virtual file. Some SEO plugins generate robots.txt on the fly, so the file sitting on your server is not the file being served. Edit the wrong one and nothing changes.

Numbers 3 and 4 are why we do not accept the file at face value. What we check is what actually comes back when the page is requested, which is a different question from what the file claims.

What we set, and what we ask the client

BotWhat we doWhy
GooglebotAllowGoogle’s search crawler.
OAI-SearchBotAllowBuilds ChatGPT’s search index.
ChatGPT-UserAllowFetches your page while someone is waiting on the answer.
Claude-SearchBotAllowSame job as OAI-SearchBot, different company.
Claude-UserAllowThe live half of the same pair.
PerplexityBotAllowBuilds Perplexity’s index, not the live fetch.
Perplexity-UserAllowThe live fetch. The one that matters when a customer asks.
GPTBot, ClaudeBot, CCBotWe ask youTraining, not answering. Your call, not ours.

We write the file by hand rather than letting a plugin generate it, and it is one of the things a page is checked for before it ships. The check looks at 4 things: that the file exists, that it is not blanket-blocking, that it references your sitemap, and that all six of the crawlers above are allowed through.

Worth doing today, costs nothing. Load your own /robots.txt. Search the text for the word Disallow followed by a lone slash. Then search for each of the 6 names in the table above. If the file has no mention of them and no site-wide block, they are allowed, and you have just ruled out the cheapest failure on the list.

What allowing them does not do

Letting a crawler in does not make an AI assistant name you. It removes 1 reason it cannot. That is the whole claim. If someone tells you it does more, ask them where that is published.

The genuinely open question is the training half. Does having your pages in a model’s training data make you more likely to be recommended later? Nobody outside those companies has published anything that answers it, and no independent study we would stand behind has measured it either. We have an opinion. We do not have evidence, so we ask the client and we do what he decides instead of dressing the opinion up as a finding.

Two smaller things worth knowing. Robots.txt is voluntary. The named crawlers above publish that they obey it, and something that ignores it is not stopped by a text file. And blocking a fetch is not the same as keeping a page out of an index, which is a different instruction entirely. If you want to see what a crawler gets once you have let it through the door, read what an AI assistant reads when it lands on your site.

See where you turn up

Your business name and your suburb. We check where you come up when someone searches your trade in your patch, who is above you, and what it would take to change that.

Free. 2 minutes. No call unless you ask for one.

We never share or sell your details, and we do not add you to a mailing list.

You can check the door yourself. We’ll check the rest of the house.

The free audit requests your pages the way a crawler does and reports what actually came back, not what the file claims. The report comes back to you regardless of what you do next.

Run my free visibility audit