Crawlers directory

AI crawlers, search bots, and training agents

A compact reference of the crawler identities OpenRevenue checks: what they are for, how they identify themselves, and where their verification sources live.

62 crawlers27 with verifiable IP rangesTrack them on your site

62 crawlers found

AI answers

User-triggered fetches when an assistant opens a page to answer someone.

16

Search indexes

Discovery crawlers that refresh content for search and answer indexes.

14

Training crawlers

Public-content crawlers that may improve models, datasets, or products.

19

GPTBot

IP verified
OpenAI

OpenAI crawler for collecting public content that may improve future models.

ClaudeBot

IP verified
Anthropic

Anthropic crawler for public content that may be used to improve Claude.

GoogleOther

IP verified
Google

Generic Google crawler used by product teams for public content fetches, including internal research and development.

Google-Extended

UA match
Google

Robots token controlling whether Google may use crawled content for Gemini model training and grounding; it is not a separate HTTP crawler user agent.

Google-CloudVertexBot

IP verified
Google

Google crawler used for targeted crawls requested by site owners building Vertex AI agents.

Applebot-Extended

IP verified
Apple

Robots token controlling whether Apple may use crawled content for AI training.

Applebot

IP verified
Apple

Apple crawler Cloudflare classifies across search and training use cases.

Amazonbot

UA match
Amazon

Amazon crawler used to improve Amazon products and services.

meta-externalagent

UA match
Meta

Meta crawler for indexing or improving Meta products and AI systems.

KimiBot

IP verified
Moonshot AI

Public-content crawler from Moonshot AI.

Bytespider

UA match
ByteDance

Public-content crawler from ByteDance.

ERNIEBot

UA match
Baidu

Public-content crawler from Baidu.

QwenBot

UA match
Alibaba

Public-content crawler from Alibaba.

ChatGLM-Spider

UA match
Zhipu AI

Public-content crawler from Zhipu AI.

DeepSeekBot

UA match
DeepSeek

Public-content crawler from DeepSeek.

cohere-ai

UA match
Cohere

Public-content crawler from Cohere.

cohere-training-data-crawler

UA match
Cohere

Public-content crawler from Cohere.

AI2Bot

UA match
Allen AI

Allen Institute crawler used to find documents for AI research systems.

CCBot

IP verified
Common Crawl

Common Crawl crawler that builds open web crawl datasets.

Other AI bots

Recognised AI bots that don't fit the categories above.

13

See which AI bots crawl your site

OpenRevenue tracks every crawl server-side, verifies it against the provider's official IP ranges, and groups it by AI answers, indexing, and training.