AI crawler index
54 named crawlers from 27 companies, and what each one is actually for. Most companies run several, and they do different jobs, so blocking by company name is rarely the decision people think it is.
Looking for the addresses they connect from? See the published IP ranges.
The four jobs
Training
Collects pages to train or improve a model. It is not answering anyone's question at the moment it visits.
If you block it: Blocking removes your content from future training. It does not remove you from answers today.
Search index
Builds the index the assistant searches when someone asks a question. Closest to a traditional search engine crawler.
If you block it: Blocking takes you out of the index, so the assistant stops finding you.
Live fetch
Fetches your page because a person just asked about it. There is a real user waiting on the other end.
If you block it: Blocking means the assistant cannot read your page for someone who asked about you by name.
In-app browser
A person browsing your site inside the assistant's own window rather than their normal browser.
If you block it: Blocking stops a real visitor from reading your site.
Crawlers by company
OpenAI
| Name | Job | What it does |
|---|---|---|
| GPTBot | Training | OpenAI training/indexing crawler |
| ChatGPT-User | Live fetch | OpenAI fetch on user behalf |
| OAI-SearchBot | Search index | ChatGPT Search index crawler |
| OAI-AdsBot | Search index | OpenAI ads discovery crawler |
| ChatGPTBrowser | Live fetch | Classic ChatGPT browse-mode fetch |
| ChatGPT-App | Live fetch | ChatGPT app-side fetch |
| ChatGPT/ | In-app browser | Real user inside ChatGPT mobile app |
| ChatGPT%20Atlas | In-app browser | Real user inside ChatGPT Atlas browser |
| Name | Job | What it does |
|---|---|---|
| Googlebot | Search index | Standard Googlebot crawler |
| GoogleOther | Training | Google generic crawler (product/R&D fetches, not AI training) |
| AdsBot-Google | Search index | Google Ads landing-page quality crawler |
| Storebot-Google | Search index | Google Shopping/store crawler |
| GoogleWv (WKWebView) | In-app browser | Google WKWebView wrapper used by Gemini and Google apps |
Anthropic
| Name | Job | What it does |
|---|---|---|
| Claude-SearchBot | Search index | Anthropic search index crawler |
| ClaudeBot | Training | Anthropic training/indexing crawler |
| Claude-User | Live fetch | Anthropic fetch on user behalf |
| Claude-Web | Live fetch | Claude web client fetch |
Gemini
| Name | Job | What it does |
|---|---|---|
| Gemini-Deep-Research | Live fetch | Gemini Deep Research fetch for a user task |
| Google-Extended | Training | Google AI training opt-in crawler |
| GeminiiOS/ | In-app browser | Real user inside Gemini iOS app |
| GeminiAndroid/ | In-app browser | Real user inside Gemini Android app |
Meta
| Name | Job | What it does |
|---|---|---|
| Meta-ExternalAgent | Training | Meta AI training crawler |
| Meta-ExternalFetcher | Live fetch | Meta AI fetch on user behalf |
| meta-externalagent | Training | Meta AI training crawler |
| meta-webindexer | Training | Meta web indexer |
Amazon
| Name | Job | What it does |
|---|---|---|
| Amazonbot | Training | Amazon AI training crawler |
| Amzn-SearchBot | Search index | Amazon search/AI crawler |
ByteDance
| Name | Job | What it does |
|---|---|---|
| Bytespider | Training | ByteDance/TikTok training crawler |
| TikTokSpider | Training | TikTok / ByteDance crawler |
DuckDuckGo
| Name | Job | What it does |
|---|---|---|
| DuckAssistBot | Live fetch | DuckDuckGo Assist AI fetch |
| DuckDuckBot | Search index | DuckDuckGo search crawler |
Huawei
| Name | Job | What it does |
|---|---|---|
| PanguBot | Training | Huawei PanGu training crawler |
| PetalBot | Search index | Huawei Petal search crawler |
Microsoft
| Name | Job | What it does |
|---|---|---|
| bingbot | Search index | Bing/Copilot search crawler |
| BingSapphire/ | In-app browser | Real user inside the Bing/Copilot mobile app WebView |
Perplexity
| Name | Job | What it does |
|---|---|---|
| PerplexityBot | Training | Perplexity training/indexing crawler |
| Perplexity-User | Live fetch | Perplexity fetch on user behalf |
xAI
| Name | Job | What it does |
|---|---|---|
| GrokApp | Live fetch | Grok iOS app fetch |
| GrokAppAndroid | Live fetch | Grok Android app fetch |
AI2
| Name | Job | What it does |
|---|---|---|
| AI2Bot | Training | Allen Institute training crawler |
Apple
| Name | Job | What it does |
|---|---|---|
| Applebot | Search index | Apple Siri/Spotlight/Apple Intelligence crawler |
Baidu
| Name | Job | What it does |
|---|---|---|
| Baiduspider | Search index | Baidu search crawler |
Brave
| Name | Job | What it does |
|---|---|---|
| Bravebot | Training | Brave Search crawler |
Cohere
| Name | Job | What it does |
|---|---|---|
| cohere-ai | Training | Cohere AI training crawler |
CommonCrawl
| Name | Job | What it does |
|---|---|---|
| CCBot | Training | Common Crawl training-data crawler |
Copilot
| Name | Job | What it does |
|---|---|---|
| BingSapphire | In-app browser | Microsoft Copilot UA marker |
Diffbot
| Name | Job | What it does |
|---|---|---|
| Diffbot | Training | Diffbot knowledge-graph crawler |
ImageSift
| Name | Job | What it does |
|---|---|---|
| ImagesiftBot | Training | ImageSift image dataset crawler |
Mistral
| Name | Job | What it does |
|---|---|---|
| MistralAI-User | Live fetch | Le Chat fetch on user behalf |
Semrush
| Name | Job | What it does |
|---|---|---|
| SemrushBot-OCOB | Training | Semrush AI content crawler |
Timpi
| Name | Job | What it does |
|---|---|---|
| Timpibot | Training | Timpi decentralized index crawler |
Webz.io
| Name | Job | What it does |
|---|---|---|
| Webzio-Extended | Training | Webz.io AI data crawler |
Yandex
| Name | Job | What it does |
|---|---|---|
| YandexBot | Search index | Yandex search crawler |
You.com
| Name | Job | What it does |
|---|---|---|
| YouBot | Training | You.com AI search crawler |
These are not crawlers
A visit can come from an AI product without any crawler being involved. When someone clicks a link in an answer, that is a person arriving in a normal browser, and it shows up through one of the signatures below. It is worth counting separately from crawler traffic, because one is your content being read and the other is a visitor showing up.
Common questions
- Does one company only have one crawler?
- No, and assuming so is the usual mistake. OpenAI runs separate crawlers for training, for its search index, and for fetching a page live when someone asks about it. Google runs different ones again, and only Google-Extended governs Gemini training. Blocking by company name treats all of those as one decision when they are not.
- Which ones should I allow?
- That is a business decision, not a technical one. The live fetch and search index crawlers are how an assistant finds you and cites you, so blocking those removes you from answers. Training crawlers do not affect whether you appear today. Plenty of sites allow the first two and block the third.
- A crawler is hitting me but is not on this list. Is it fake?
- Not necessarily. New crawlers appear regularly and this list covers the ones we identify by name. Check the user agent against the vendor's own documentation, and check the IP against their published ranges. A name nobody documents and an IP from a hosting provider is the pattern worth blocking.
- How is a referral different from a crawler?
- A crawler is software reading your page. A referral is a person clicking through to you from an answer, which shows up as a normal browser visit carrying a referer such as chatgpt.com. They are counted separately because they mean different things: one is your content being read, the other is a visitor arriving.