> ## Documentation Index
> Fetch the complete documentation index at: https://docs.growthmention.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Recognised crawlers

> The 116 AI crawlers Growth Mention recognises, what each one is for, and which can be verified.

Growth Mention recognises **116 AI crawlers** from 68 companies. It works out which one sent each visit from the request's user agent, then checks the visit really came from that company.

## Only verified crawlers are counted

A user agent is only a claim, and anyone can send a request calling itself *GPTBot*. So each visit is checked against the addresses or DNS names its operator publishes for its crawlers. **Only visits that pass appear in Agent Analytics.**

That's possible for **35** of the 116 crawlers, including every crawler from OpenAI, Perplexity, Apple and Meta. Some notable ones can't be verified — xAI's and Microsoft's among them.

The other 81 mostly come from companies that publish nothing to check against, so their visits can't be proven and aren't shown. Claude Code is the exception: it runs on a developer's own computer, so there are no company addresses to check it against.

<Note>
  The list is kept up to date on our side, not in the WordPress plugin or your integration. New crawlers are recognised without you updating anything.
</Note>

## What each purpose means

| Purpose        | What it's doing                                                                                                                           | If you block it                                       |
| -------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------- |
| **Search**     | Building the index an assistant answers from. Includes retrieval providers such as Exa, Tavily and Brave, whose indexes assistants query. | You drop out of that assistant's answers.             |
| **User fetch** | Fetching a page right now, because someone asked an assistant or agent about it.                                                          | The assistant can't read the page it was asked about. |
| **Training**   | Collecting pages to train future models on. Includes scraping services that resell data.                                                  | Nothing, for your visibility today.                   |

## Verified crawlers

These are the crawlers that appear in your Agent Analytics.

| Crawler                | Company      | Purpose    |
| ---------------------- | ------------ | ---------- |
| Amzn-SearchBot         | Amazon       | Search     |
| Claude-SearchBot       | Anthropic    | Search     |
| Applebot               | Apple        | Search     |
| Google-CloudVertexBot  | Google       | Search     |
| PetalBot               | Huawei       | Search     |
| meta-webindexer        | Meta         | Search     |
| MistralAI-Index        | Mistral      | Search     |
| OAI-SearchBot          | OpenAI       | Search     |
| PerplexityBot          | Perplexity   | Search     |
| YandexAdditional       | Yandex       | Search     |
| YandexAdditionalBot    | Yandex       | Search     |
| YouBot                 | You.com      | Search     |
| Amzn-User              | Amazon       | User fetch |
| Claude-User            | Anthropic    | User fetch |
| YiyanBot               | Baidu        | User fetch |
| DuckAssistBot          | DuckDuckGo   | User fetch |
| Gemini-Deep-Research   | Google       | User fetch |
| Google-Agent           | Google       | User fetch |
| Google-GeminiNotebook  | Google       | User fetch |
| GoogleAgent-Mariner    | Google       | User fetch |
| GoogleAgent-URLContext | Google       | User fetch |
| meta-externalfetcher   | Meta         | User fetch |
| MistralAI-User         | Mistral      | User fetch |
| ChatGPT-User           | OpenAI       | User fetch |
| OAI-AdsBot             | OpenAI       | User fetch |
| Perplexity-User        | Perplexity   | User fetch |
| Amazonbot              | Amazon       | Training   |
| ClaudeBot              | Anthropic    | Training   |
| ERNIEBot               | Baidu        | Training   |
| CCBot                  | Common Crawl | Training   |
| CloudVertexBot         | Google       | Training   |
| GoogleOther            | Google       | Training   |
| FacebookBot            | Meta         | Training   |
| meta-externalagent     | Meta         | Training   |
| GPTBot                 | OpenAI       | Training   |

## Recognised but not verifiable

These are recognised, but their operators publish no addresses or DNS names, so a visit can't be told apart from an impostor using the same name. They don't appear in Agent Analytics.

**Search:** AddSearchBot (AddSearch), amazon-kendra (Amazon), atlassian-bot (Atlassian), Bravebot (Brave), Channel3Bot (Channel3), Cloudflare-AutoRAG (Cloudflare), Diffbot (Diffbot), Anomura (Direqt), ExaBot (Exa), LinkupBot (Linkup), AIWebIndex (Lyrenth), AzureAI-SearchBot (Microsoft), Kimi-SearchBot (Moonshot AI), Mozilla-Tabstack (Mozilla), ShapBot (Parallel), Querit-SearchBot (Querit), QueritBot (Querit), TavilyBot (Tavily), HenkBot (Valyu), xAI-SearchBot (xAI), ZanistaBot (Zanista).

**User fetch:** AI2Bot-DeepResearchEval (Ai2), TongyiBot (Alibaba), amazon-QBusiness (Amazon), AmazonBuyForMe (Amazon), NovaAct (Amazon), Claude-Code (Anthropic), bigsur.ai (Big Sur AI), Manus-User (Butterfly Effect), Trae (ByteDance), Devin (Cognition), GeistHaus-PageFetcher (GeistHaus), Google-Gemini-CLI (Google), kagi-fetcher (Kagi), KlaviyoAIBot (Klaviyo), LinerBot (Liner), Kimi-User (Moonshot AI), opencode (opencode), Shap-User (Parallel), PhindBot (Phind), Poggio-Citations (Poggio), QualifiedBot (Qualified), TwinAgent (Twin).

**Training:** AI2Bot (Ai2), Ai2Bot-Dolma (Ai2), SalamandraVLM (Aithlas), QwenBot (Alibaba), bedrockbot (Amazon), ApifyBot (Apify), ApifyWebsiteContentCrawler (Apify), Brightbot (Bright Data), Bytespider (ByteDance), Doubaobot (ByteDance), imageSpider (ByteDance), Terra Cotta (Ceramic), TerraCotta (Ceramic), cohere-training-data-crawler (Cohere), CohereBot (Cohere), CragCrawler (CragSoftware), Crawlspace (Crawlspace), DeepSeekBot (DeepSeek), FirecrawlAgent (Firecrawl), PanguBot (Huawei), Kangaroo Bot (Kangaroo LLM), laion-huggingface-processor (LAION), MistralBot (Mistral), KimiBot (Moonshot AI), KimiCrawler (Moonshot AI), MoonshotBot (Moonshot AI), MoonSpider (Moonshot AI), Datenbank Crawler (netEstate), netEstate Imprint Crawler (netEstate), ICC-Crawler (NICT), micro-crawl (Reflection AI), Cotoyogi (ROIS), SBIntuitionsBot (SB Intuitions), Timpibot (Timpi), VelenPublicWebCrawler (Velen), omgili (Webz.io), webzio-extended (Webz.io), ChatGLM-Spider (Zhipu AI).

## Not crawlers at all

**Google-Extended** and **Applebot-Extended** are names you can use in `robots.txt` to opt out of Google's and Apple's AI training. They never visit your site — Google and Apple crawl with their usual bots and honour the setting — so they never appear here. Blocking them is still a real choice; it just isn't one you can see in traffic.
