Registry reviewed September 2026

AI crawlers and search bots, from the operators' own documentation

109 documented crawlers and fetchers from 15 operators. For each one: the exact user-agent token, the robots.txt token, whether the operator says it honors robots.txt, how to verify a request is genuine, and the official IP range file where one exists. Nothing here is guessed; bots without official documentation are not listed.

109
documented bots
15
operators
46
with official IP range files
12,153
IP prefixes in those files

AI training crawlers

Collects public web content that may be used to train or improve AI models. These are the tokens most site owners decide about first.

AI search indexs

Builds an index that an AI assistant searches at answer time, so it can cite and link pages. Blocking it removes a site from those answers.

User-triggered fetchers

Fetches a page because a person asked an assistant or app to open it. Operators generally treat these like a browser acting on the user's behalf.

Search engine crawlers

Classic web search indexing. Blocking it removes a site from that search engine.

Applebot
Apple
Applebot
robots.txt: yes33 IP prefixesreverse DNS
Bytespider
ByteDance
Bytespider
robots.txt: undocumented12 IP prefixesreverse DNS
DuckDuckBot
DuckDuckGo
DuckDuckBot
robots.txt: yes486 IP prefixes
Google StoreBot
Google
Storebot-Google
robots.txt: yes317 IP prefixesreverse DNS
Googlebot
Google
Googlebot
robots.txt: yes317 IP prefixesreverse DNS
Googlebot Image
Google
Googlebot-Image
robots.txt: yes317 IP prefixesreverse DNS
Googlebot News
Google
Googlebot-News
robots.txt: yes317 IP prefixesreverse DNS
Googlebot Video
Google
Googlebot-Video
robots.txt: yes317 IP prefixesreverse DNS
BingVideoPreview
Microsoft Bing
BingVideoPreview/1.0
robots.txt: undocumented
Bingbot
Microsoft Bing
bingbot/2.0
robots.txt: yes28 IP prefixesreverse DNS
YaComBot
Yandex
YaComBot
robots.txt: yesreverse DNS
YandexBlogs
Yandex
YandexBlogs
robots.txt: yesreverse DNS
YandexBot
Yandex
YandexBot
robots.txt: yesreverse DNS
YandexComBot
Yandex
YandexComBot
robots.txt: partialreverse DNS
YandexFavicons
Yandex
YandexFavicons
robots.txt: partialreverse DNS
YandexImages
Yandex
YandexImages
robots.txt: yesreverse DNS
YandexMedia
Yandex
YandexMedia
robots.txt: yesreverse DNS
YandexMobileBot
Yandex
YandexMobileBot
robots.txt: partialreverse DNS
YandexOntoDB
Yandex
YandexOntoDB
robots.txt: yesreverse DNS
YandexRenderResourcesBot
Yandex
YandexRenderResourcesBot
robots.txt: partialreverse DNS
YandexSitelinks
Yandex
YandexSitelinks
robots.txt: yesreverse DNS
YandexVerticals
Yandex
YandexVerticals
robots.txt: yesreverse DNS
YandexVertis
Yandex
YandexVertis
robots.txt: yesreverse DNS
YandexVideo
Yandex
YandexVideo
robots.txt: yesreverse DNS
YandexVideoParser
Yandex
YandexVideoParser
robots.txt: partialreverse DNS

Social preview fetchers

Fetches a page to build the link card shown when a URL is shared on a platform.

SEO tool crawlers

Crawls to build backlink and site-audit indexes sold to marketers.

Archive crawlers

Crawls to build a public archive or open corpus of the web.

Others

Documented bots that do not fit the categories above.

iTMS
Apple
iTMS
robots.txt: noreverse DNS
APIs-Google
Google
APIs-Google
robots.txt: partial272 IP prefixesreverse DNS
AdSense
Google
Mediapartners-Google
robots.txt: partial272 IP prefixesreverse DNS
AdsBot
Google
AdsBot-Google
robots.txt: partial272 IP prefixesreverse DNS
AdsBot Mobile Web
Google
AdsBot-Google-Mobile
robots.txt: partial272 IP prefixesreverse DNS
Google-Safety
Google
Google-Safety
robots.txt: no272 IP prefixesreverse DNS
GoogleOther
Google
GoogleOther
robots.txt: yes317 IP prefixesreverse DNS
GoogleOther-Image
Google
GoogleOther-Image
robots.txt: yes317 IP prefixesreverse DNS
GoogleOther-Video
Google
GoogleOther-Video
robots.txt: yes317 IP prefixesreverse DNS
Meta-ExternalAds
Meta
meta-externalads
robots.txt: yes
AdIdxBot
Microsoft Bing
adidxbot/2.0
robots.txt: undocumented
OAI-AdsBot
OpenAI
OAI-AdsBot
robots.txt: undocumented2 IP prefixes
YaDirectFetcher
Yandex
YaDirectFetcher
robots.txt: noreverse DNS
Yandex
Yandex
Yandex
robots.txt: partial
YandexAccessibilityBot
Yandex
YandexAccessibilityBot
robots.txt: partialreverse DNS
YandexAdNet
Yandex
YandexAdNet
robots.txt: yesreverse DNS
YandexCheckBot
Yandex
YandexCheckBot
robots.txt: partialreverse DNS
YandexDirect
Yandex
YandexDirect
robots.txt: partialreverse DNS
YandexDirectDyn
Yandex
YandexDirectDyn
robots.txt: partialreverse DNS
YandexImageResizer
Yandex
YandexImageResizer
robots.txt: yesreverse DNS
YandexMarket
Yandex
YandexMarket
robots.txt: partialreverse DNS
YandexMetrika
Yandex
YandexMetrika
robots.txt: noreverse DNS
YandexMobileScreenShotBot
Yandex
YandexMobileScreenShotBot
robots.txt: partialreverse DNS
YandexOntoDBAPI
Yandex
YandexOntoDBAPI
robots.txt: partialreverse DNS
YandexPartner
Yandex
YandexPartner
robots.txt: partialreverse DNS
YandexRCA
Yandex
YandexRCA
robots.txt: partialreverse DNS
YandexScreenshotBot
Yandex
YandexScreenshotBot
robots.txt: partialreverse DNS
YandexSpravBot
Yandex
YandexSpravBot
robots.txt: yesreverse DNS
YandexTracker
Yandex
YandexTracker
robots.txt: partialreverse DNS

By operator

How this registry is built

Each entry was compiled from the operator's own documentation, then handed to a second reviewer whose only job was to re-open every source and strike any claim not found there. What survived is what you see. Fields the operator does not document are marked as such rather than filled in.

IP-range figures come from a script that downloads each operator's official range file, counts the prefixes, and records the file's stated creation time and a SHA-256 of the body. The last run was 2026-09-06. Every bot page links the live file so you can compare.

The registry is read-only reference. It documents what operators say about their own bots so site owners can identify and control them; it does not publish third-party IP lists or speculate about undocumented behaviour.

Common questions

What is the crawler and bot registry?+

A read-only reference of 109 web crawlers and fetchers from 15 operators, including the AI crawlers site owners ask about most: GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and Googlebot. Each entry gives the exact user-agent token, the robots.txt token, whether the operator says it honors robots.txt, the documented way to verify a request, and the official IP range file where one is published. Reviewed September 2026.

Where does the data come from?+

Only from each operator's own documentation and official machine-readable range files. Every field was checked against the operator's page and re-checked by an independent reviewer. IP-range counts are produced by a script that downloads the operator's file and records its creation time and a hash, last run 2026-09-06. Bots without official documentation are not listed.

How do I block AI crawlers from training on my site?+

Add a robots.txt rule per token. Training crawlers documented here include GPTBot, ClaudeBot and CCBot; Google uses the policy token Google-Extended, which has no crawler of its own. Each entry has copy-ready allow and block blocks. Note that user-triggered fetchers such as ChatGPT-User act on a person's request and some operators document that they do not consult robots.txt.

How can I tell whether a request is really from the bot it claims to be?+

Anyone can send any User-Agent string. The two documented checks are a reverse DNS lookup (the IP must resolve to the operator's domain and back again) and a match against the operator's published IP ranges. Each entry states which of those the operator supports.

Is the dataset free to use?+

Yes. The full registry is published as an open dataset under CC BY 4.0, downloadable as CSV, NDJSON or JSON, with citation blocks on the dataset page.

Related references

The anti-bot protection registry covers the other side of this exchange: which protection popular sites run against automated visitors. The glossary defines the terms both registries use.

All open datasets