We run a large fleet of real mobile devices and a signature-verified registry of the anti-bot web, which lets us measure things most people can only guess at. When we find something worth publishing, it lands here - with the raw data attached and free to cite.
Anti-bot protection observed on 94 popular sites across 15 verticals, from public HTTP signatures. 61% are hard targets. Vendor market share, difficulty by vertical, and the raw data.
PROXIES.SX. (July 2026). Anti-Bot Census: anti-bot protection observed across 94 popular websites. https://www.proxies.sx/anti-bot/dataset
@misc{proxiessx_antibot_census,
title = {Anti-Bot Census: anti-bot protection observed across 94 popular websites},
author = {{PROXIES.SX}},
year = {2026},
note = {Reviewed July 2026},
url = {https://www.proxies.sx/anti-bot/dataset}
}109 documented crawlers and fetchers from 15 operators, including GPTBot, ClaudeBot, PerplexityBot and Googlebot. Exact user-agent tokens, robots.txt tokens, robots compliance as the operator states it, and official IP range files with probe-recorded prefix counts. Compiled from operator documentation only; bots without official docs are not listed.
134 mobile operators in 44 countries. Each ASN carries its RIPEstat holder name, whether it is announced, IPv4 and IPv6 prefix counts, its PeeringDB record, and the documented role (mobile, fixed, enterprise) with a confidence grade and an evidence URL. MCC/MNC codes are listed with their source.
Every figure traces to an evidence-backed source we can re-run. If we cannot verify it, we do not publish it.
Datasets are released under CC BY 4.0. Use them in articles, decks and research - just attribute proxies.sx and link back.
The web changes, so studies carry a review date and recurring editions, with the differences published as a changelog.
Data Works is our done-for-you pipeline for the hard targets in these datasets - you send the URLs and fields, we deliver clean data on real 4G/5G carrier IPs.
20 resolver operators and 167 documented addresses, each checked word for word against the operator's pages. We queried every IPv4 address for DNSSEC validation, client-subnet forwarding and NXDOMAIN handling, and tested every DoH and DoT endpoint.
20 selected files, with 383,301 prefix strings counted across those files. Each row records its source, fetch time, publisher fields and SHA-256 hash. Overlapping ranges remain separate.
Diagnose HTTP 407 with curl. Separate proxy credentials, authentication methods and CONNECT failures before changing IPs or retrying.
Handle HTTP 429 with Retry-After, limited retries and a tested Python example. Distinguish request rate, concurrency and shared quotas.
Compare local and proxy-side DNS in curl and Python Requests. Test hostname resolution and understand what a browser DNS leak test establishes.
Match IPv4 and IPv6 addresses against an AWS range file with Python. Keep source hashes, overlapping matches and clear classification limits.
Verify crawler traffic using official IP ranges or forward-confirmed reverse DNS. Check trusted client IPs and avoid broad cloud allowlists.
Research library
Find the reference behind a network decision. Browse the records, follow their sources and download the available datasets.