28 AI crawlers currently tracked by Tailwin across customer sites: the user-agent each one sends, who operates it, and what it does with the pages it reads. Every entry is verifiable in your own server logs. This list is maintained as data, so it updates as vendors ship new agents, without waiting for anyone’s release cycle.
Crawlers that build the search layer behind AI answers. A page these bots cannot reach cannot be cited in the products they feed.
| Crawler | Operator | User-agent token | robots.txt | Verify |
|---|---|---|---|---|
| Claude-SearchBot | Anthropic | Claude-SearchBot | honored | vendor IPs |
| Applebot | Apple | Applebot | honored | – |
| Googlebot | Googlebot | honored | – | |
| Bingbot | Microsoft | bingbot | honored | – |
| OAI-SearchBot | OpenAI | OAI-SearchBot | honored | vendor IPs |
| PerplexityBot | Perplexity | PerplexityBot | honored | vendor IPs |
Agents that fetch a page during a live conversation, because a user's question needed it right then. A fetcher hit is a person, mid-question, whose answer touched your site.
| Crawler | Operator | User-agent token | robots.txt | Verify |
|---|---|---|---|---|
| Claude-User | Anthropic | Claude-User | honored | vendor IPs |
| DuckAssistBot | DuckDuckGo | DuckAssistBot | honored | – |
| Google-Agent | Google-Agent | documented as ignored | – | |
| Gemini Notebook | Google-GeminiNotebook | documented as ignored | – | |
| NotebookLM (legacy) | Google-NotebookLM | documented as ignored | – | |
| Meta-ExternalFetcher | Meta | meta-externalfetcher | documented as ignored | – |
| BingPreview | Microsoft | BingPreview | undocumented | – |
| MistralAI-User | Mistral | MistralAI-User | honored | – |
| ChatGPT-User | OpenAI | ChatGPT-User | documented as ignored | vendor IPs |
| OAI-AdsBot | OpenAI | OAI-AdsBot | undocumented | vendor IPs |
| Perplexity-User | Perplexity | Perplexity-User | documented as ignored | vendor IPs |
Crawlers that collect pages for model training corpora. Blocking a trainer is a content-rights decision; it does not remove you from any live answer.
| Crawler | Operator | User-agent token | robots.txt | Verify |
|---|---|---|---|---|
| AI2Bot | Allen Institute | AI2Bot | honored | – |
| Amazonbot | Amazon | Amazonbot | honored | vendor IPs |
| ClaudeBot | Anthropic | ClaudeBot | honored | vendor IPs |
| Bytespider | ByteDance | Bytespider | documented as ignored | – |
| Cohere Crawler | Cohere | cohere-ai | undocumented | – |
| CCBot | Common Crawl | CCBot | honored | – |
| Diffbot | Diffbot | Diffbot | undocumented | – |
| Google-CloudVertexBot | Google-CloudVertexBot | honored | – | |
| FacebookBot | Meta | FacebookBot | honored | – |
| Meta-ExternalAgent | Meta | meta-externalagent | honored | – |
| GPTBot | OpenAI | GPTBot | honored | vendor IPs |
None of these visits appear in JavaScript analytics, and Search Console shows Googlebot only. If you want to know which AI systems read your site, you have to look at the request layer itself. Tailwin logs these crawlers at the edge for its customers, scores what their access means for AI visibility, and ties published content to the crawls and citations that follow.
Run a free scan to see which of these crawlers can reach your site today.