AI Bot & Crawler Traffic 2026

AI Bot & Crawler Traffic Statistics 2026

How much of the web is now automated, which AI crawlers generate the most requests, how many pages they take for every visitor they send back, and what that traffic costs the sites serving it. Every figure below carries a named source and a date, and live measurements are labelled with the exact window they were read.

📊 40+ sourced statistics 🔍 Cloudflare, Thales, Fastly & Pew data 📅 Live figures read 2–8 Aug 2026 🔄 Updated August 2026
Our analysis

Why “what percent of web traffic is bots” has three different answers

The three most-quoted sources on bot traffic disagree by nearly 20 percentage points. They are not contradicting each other — they are measuring different things. Here is the reconciliation, because picking the wrong figure is the most common error in reporting on this subject.

35.7%of HTTP requests are botsCloudflare Radar, worldwide, 2–8 Aug 2026
>50%of internet traffic is non-humanCloudflare, July 2026
53%of internet traffic is botsThales, full-year 2025 data
See how the three reconcile →

Original analysis by Paid Hosting · August 2026 · methodology · full source list

Key Bot Traffic Statistics at a Glance

The headline numbers that define automated web traffic in 2026. Every figure links to a named report or a dated live measurement in the sources list.

53%of all internet traffic came from bots across 2025, with bad bots alone at 40%Thales · Apr 2026
1.2K:1pages crawled by Anthropic’s bots for every visitor referred backCloudflare Radar · 2–8 Aug 2026
52%of crawler requests were for AI training as of June 2026, up from 22% in spring 2025Cloudflare · Jul 2026
22.3%of AI bot and crawler traffic comes from Googlebot — the single largest sourceCloudflare Radar · 2–8 Aug 2026
17.2Tbad bot requests blocked across 2025 — about 47 billion per dayThales · Apr 2026
11.7%of the top 10,000 domains’ robots.txt files fully disallow GPTBot — versus 0.7% for GooglebotPaid Hosting calculation from Cloudflare Radar · 3 Aug 2026

How Much of Web Traffic Is Bots

This is the most-requested figure on the subject and the one most often misquoted, because the three authoritative sources report numbers between 35% and 53%. The difference is methodological, not a dispute about facts.

📌
Reconciling 35.7%, “more than 50%”, and 53%

Cloudflare Radar’s live dashboard reported 35.7% of worldwide HTTP requests as bot traffic for 2–8 August 2026. Cloudflare’s own July 2026 report states that more than 50% of internet traffic is now non-human. Thales, analysing full-year 2025 data across its customer base, puts bots at 53% of all internet traffic. Three reasons account for the gap. First, scope: Radar’s “bot” classification counts requests it identifies as automated, while “non-human traffic” is a broader category that includes AI agents acting on behalf of users, which Thales began counting as a separate third category in its 2026 report. Second, network: each vendor sees only its own customers’ traffic, and the mix of sites behind Cloudflare, Thales and Fastly differs. Third, period: Radar reports a rolling seven-day window that moves daily, while the vendor reports are annual and lag by months. If you need one figure, say “roughly a third to a half of web traffic is automated, depending on how automation is defined” and cite the specific measurement you mean.

  • Bots generated 35.7% of worldwide HTTP requests, with humans at 64.3%.Source: Cloudflare Radar, live worldwide reading for 2–8 August 2026
  • More than 50% of traffic on the internet is now non-human — a threshold crossed for the first time this year.Source: Cloudflare, “Content Independence Day, one year on,” 1 July 2026
  • Bots accounted for 53% of all internet traffic across 2025, with bad bots alone at 40% — a rise of 3 percentage points year on year.Source: Thales 2026 Bad Bot Report, 30 April 2026 (13th edition, full-year 2025 data)
  • Taking those two Thales figures together, roughly three-quarters of all bot traffic (75.5%) is classified as malicious rather than benign.Source: Paid Hosting calculation from Thales 2026 Bad Bot Report figures (40% ÷ 53%)
  • The United States originates 41.5% of worldwide bot traffic — more than the next nineteen countries combined. Germany follows at 6.0%, the Netherlands at 4.7%, Singapore at 4.5% and France at 3.2%.Source: Cloudflare Radar, 2–8 August 2026
  • Amazon’s network (AS16509) is the single largest source of bot traffic at 8.7%, followed by Google (AS396982) at 5.7%, Google LLC (AS15169) at 2.2%, Servers.com at 2.1% and DigitalOcean at 1.1% — a reminder that most crawling originates from cloud infrastructure rather than residential networks.Source: Cloudflare Radar, 2–8 August 2026
  • Generative AI reached 2.5 billion active users — over 30% of humanity — in three and a half years, adopted at more than twice the speed of smartphones.Source: Cloudflare, 1 July 2026

Which AI Crawlers Generate the Most Traffic

Five operators account for two-thirds of all AI bot and crawler traffic. The ranking has shifted substantially in twelve months, which is why the older figures further down this section are worth reading alongside the current ones.

Share of AI bot and crawler traffic by bot

Top five most active AI bots as a share of AI bot and crawler HTTP requests, worldwide.

Googlebot 22.3%
Meta-ExternalAgent 14.0%
ClaudeBot 13.5%
GPTBot 8.6%
Amazonbot 8.2%

Source: Cloudflare Radar, worldwide, 2–8 August 2026. Read 9 August 2026.

Chart may be reproduced with attribution to paidhosting.com
  • The five most active AI bots — Googlebot, Meta-ExternalAgent, ClaudeBot, GPTBot and Amazonbot — together account for 66.6% of AI bot and crawler traffic.Source: Paid Hosting calculation from Cloudflare Radar, 2–8 August 2026
  • Within OpenAI’s traffic, GPTBot (the training crawler) generates 84.9% of requests, ChatGPT-User 9.2% and OAI-SearchBot 5.9%.Source: Cloudflare Radar, 2–8 August 2026
  • Within Anthropic’s traffic, ClaudeBot generates 75.8%, Claude-SearchBot 23.7% and Claude-User 0.5%.Source: Cloudflare Radar, 2–8 August 2026
  • Within Perplexity’s traffic, PerplexityBot generates 78.4% and Perplexity-User 21.6%.Source: Cloudflare Radar, 2–8 August 2026
  • A year earlier the ranking looked very different: across April–July 2025, Meta’s AI crawlers alone generated 52% of AI crawler traffic on Fastly’s network, more than Google (23%) and OpenAI (20%) combined.Source: Fastly Q2 2025 Threat Insights Report, 19 August 2025
  • In that same period, crawlers made up almost 80% of AI bot traffic and real-time fetchers the remaining 20% — and OpenAI’s bots generated 98% of all fetcher requests.Source: Fastly, 19 August 2025
  • Fetcher traffic reached bursts of 39,000 requests per minute against individual sites — enough to degrade or overwhelm an unprepared origin server.Source: Fastly, 19 August 2025
  • Crawler traffic is heavily concentrated in North America: nearly 90% of North American AI bot traffic came from crawlers, versus 72% in Latin America, 58% in Asia and 41% in Europe. Meta and Google sent over 90% of their crawler traffic to North American sites.Source: Fastly, 19 August 2025
💡
Why Fastly and Cloudflare rank the bots differently

Fastly’s 2025 report put Meta far ahead of every other crawler; Cloudflare’s 2026 live data puts Googlebot first and Meta second. Both are accurate for what they measure. The two networks serve different customer bases in different regions, and Fastly’s figure covers AI crawlers specifically while Cloudflare’s covers AI bots and crawlers together, including search. The Fastly figures are also a year older. Use Cloudflare for a current reading and Fastly as a historical comparison point — not as competing claims about the same quantity.

Crawl-to-Refer Ratios

The crawl-to-refer ratio is the number of pages an operator’s bots request for every visitor it sends back to the site. It is the clearest single measure of the exchange between AI companies and the sites they read, and the spread between operators is enormous.

Pages crawled per visitor referred back

Ratio of HTML page crawl requests to HTML page referrals, by platform. Shown on a scale capped at 1,200 for legibility.

Anthropic 1.2K:1
Perplexity 746:1
OpenAI 256.8:1
Microsoft 40.3:1
Yandex 31.7:1
Baidu 13.6:1
ByteDance 9.3:1
Google 5.1:1
DuckDuckGo 1.8:1

Source: Cloudflare Radar, worldwide, 2–8 August 2026. Read 9 August 2026. Mistral registered no referrals at all in this period.

Chart may be reproduced with attribution to paidhosting.com
  • Anthropic’s bots requested roughly 1,200 pages for every visitor referred back — the highest ratio of any major operator.Source: Cloudflare Radar, 2–8 August 2026
  • Google’s ratio is 5.1:1, meaning Google returns a visitor roughly 235 times more often per page crawled than Anthropic does. DuckDuckGo is the only platform close to parity at 1.8:1.Source: Paid Hosting calculation from Cloudflare Radar, 2–8 August 2026
  • Mistral’s user agent registered no referrals whatsoever during the measurement week, making its ratio effectively infinite.Source: Cloudflare Radar, 2–8 August 2026
  • Ratios are volatile week to week. Against the previous seven days, ByteDance rose 37.9%, Perplexity rose 24.6% and Baidu rose 13.2%, while OpenAI fell 63.5% and Anthropic fell 12.7%. Any single reading should be quoted with its date.Source: Cloudflare Radar, 2–8 August 2026
  • Around July 2025, Cloudflare observed crawl-to-referral ratios ranging from 118:1 to nearly 50,000:1 across leading AI companies — a far wider spread than the current range, indicating the extreme outliers have narrowed considerably.Source: Cloudflare, “Unmasking the crawls with Attribution Business Insights,” 1 July 2026
⚠️
A caution on quoting these numbers

Crawl-to-refer ratios are among the most frequently cited statistics in coverage of AI and publishing, and they are also among the most volatile. The figures above moved by more than 60% in a single week for one operator. They also measure HTML page requests against HTML page referrals on Cloudflare’s network only, so a site with a different traffic mix will see different ratios. Quote them with the operator, the source and the exact week, and treat any figure more than a few months old as historical rather than current.

What AI Crawlers Are Collecting

Cloudflare classifies crawler requests by declared purpose, which shows how much of the crawling is for model training rather than for search indexing that returns traffic to the source.

AI crawler traffic by declared purpose

Share of crawler request volume by the purpose the operator declares.

Training 41.4%
Training & search (mixed-use) 36.8%
Search only 15.8%
User action 4.6%
Undeclared 1.5%

Source: Cloudflare Radar, worldwide, 2–8 August 2026. Read 9 August 2026.

Chart may be reproduced with attribution to paidhosting.com
  • 78.2% of AI crawler traffic involves model training, either as the sole declared purpose (41.4%) or combined with search (36.8%). Search-only crawling accounts for 15.8%.Source: Paid Hosting calculation from Cloudflare Radar, 2–8 August 2026
  • 52% of crawler requests were for AI training as of June 2026, up from 22% in spring 2025 — more than doubling in roughly a year.Source: Cloudflare, 1 July 2026
  • More than one-third of crawler activity comes from mixed-use bots that do not distinguish between crawling for search and crawling for training, leaving site owners unable to allow one and refuse the other.Source: Cloudflare, 1 July 2026
  • Content served to AI bots is overwhelmingly HTML at 71.8%, followed by plain text at 8.3%, JSON at 6.0% and images at 4.8%. Markdown and documents each account for under 0.1%.Source: Cloudflare Radar, 2–8 August 2026
  • Of the leading AI operators tracked by Cloudflare, Amazon, Anthropic, Meta and OpenAI both verify their bots and run distinct bots by purpose. Apple, Google and Microsoft verify but do not separate by purpose; ByteDance does neither.Source: Cloudflare Radar AI bot transparency tracker, read 9 August 2026
  • Google’s combined discovery-and-AI crawler gives it access to roughly twice as much information as leading AI companies that separate the two functions.Source: Cloudflare, 1 July 2026

How Site Owners Are Responding

Two datasets show what site owners actually do about AI crawlers: what they declare in robots.txt, and what their servers return when a bot arrives. The two tell noticeably different stories.

Share of top-10,000-domain robots.txt files fully disallowing each bot

Calculated from 4,236 robots.txt files found across the top 10,000 domains. “Fully disallowed” excludes partial restrictions.

GPTBot (OpenAI) 11.7%
CCBot (Common Crawl) 11.6%
Bytespider (ByteDance) 11.1%
ClaudeBot (Anthropic) 10.5%
Google-Extended 9.7%
meta-externalagent 9.4%
Amazonbot 8.9%
Applebot-Extended 8.4%
PerplexityBot 4.2%
Googlebot (search) 0.7%

Source: Paid Hosting calculation from Cloudflare Radar robots.txt data, updated 3 August 2026. Read 9 August 2026.

Chart may be reproduced with attribution to paidhosting.com
  • GPTBot is the most-blocked bot among the top 10,000 domains: 495 of 4,236 robots.txt files (11.7%) fully disallow it, and 656 (15.5%) restrict it fully or partially.Source: Paid Hosting calculation from Cloudflare Radar, updated 3 August 2026
  • Site owners block AI training crawlers roughly seventeen times more often than they block Google’s search crawler: 495 domains fully disallow GPTBot against 29 for Googlebot.Source: Paid Hosting calculation from Cloudflare Radar, updated 3 August 2026
  • Google-Extended, the opt-out token for Google’s AI training, is fully disallowed by 411 domains (9.7%) — while Googlebot itself is disallowed by 29. Site owners are drawing a clear line between being indexed and being used for training.Source: Paid Hosting calculation from Cloudflare Radar, updated 3 August 2026
  • Perplexity’s crawler is blocked far less often than the major training crawlers, at 4.2% fully disallowed against GPTBot’s 11.7%.Source: Paid Hosting calculation from Cloudflare Radar, updated 3 August 2026
  • In practice, servers refuse or throttle more traffic than robots.txt suggests: 14.6% of AI bot requests receive a 403 Forbidden and a further 2.6% receive a 429 Too Many Requests, for a combined 17.2% refused or rate-limited. Only 63.7% receive a normal 200 OK.Source: Paid Hosting calculation from Cloudflare Radar, 2–8 August 2026
  • Cloudflare changed its default in 2025 so that AI training crawlers are blocked for all new domains unless the owner opts in — a network-level enforcement mechanism rather than the voluntary robots.txt standard.Source: Cloudflare, 1 July 2026
💡
robots.txt understates how much blocking actually happens

Only 11.7% of large sites fully disallow GPTBot in robots.txt, yet 17.2% of AI bot requests across Cloudflare’s network are refused or throttled at the server. The gap exists because robots.txt is an advisory standard that a crawler can ignore, while 403 and 429 responses are enforcement — and much of that enforcement now happens at the network layer through firewall rules and bot management products rather than through anything the site owner writes in a text file. Reporting that uses robots.txt adoption as a proxy for “how many sites block AI crawlers” will therefore undercount.

Bandwidth & Server Cost

Crawler traffic is a hosting cost before it is anything else. It consumes bandwidth, occupies server resources and is billed identically to human traffic — but the load it creates is disproportionate to its share of requests.

  • Wikimedia’s bandwidth for serving multimedia content grew 50% between January 2024 and April 2025, driven principally by automated scraping rather than by growth in human readership.Source: Wikimedia Foundation, “How crawlers impact the operations of the Wikimedia projects,” 1 April 2025
  • At least 65% of Wikimedia’s most resource-intensive traffic came from bots, while bots accounted for only about 35% of total pageviews — meaning automated traffic was roughly 1.9 times more expensive to serve than its request share implies.Source: Wikimedia Foundation, 1 April 2025; ratio calculated by Paid Hosting
  • Serving markdown instead of HTML to AI bots reduces the median response to 7.2% of its original size — a 92.8% reduction in bytes transferred for the same content.Source: Cloudflare Radar, 2–8 August 2026; reduction calculated by Paid Hosting
  • 71.8% of what AI bots receive is HTML, the heaviest of the common formats, and under 0.1% is markdown — so almost none of that potential saving is currently being realised.Source: Cloudflare Radar, 2–8 August 2026
  • Bursts of AI fetcher traffic have reached 39,000 requests per minute against a single site, a rate comparable to a denial-of-service event for an unprepared origin server.Source: Fastly Q2 2025 Threat Insights Report, 19 August 2025
  • Indiscriminate crawling creates unnecessary bandwidth burden for publishers and wastes compute for AI companies, according to Cloudflare, which is investing in freshness signals specifically to reduce redundant crawling.Source: Cloudflare, 1 July 2026
📌
What the cost figures do and do not show

There is no reliable published figure for the average dollar cost of AI crawler traffic to a typical website, and any number presented as one should be treated with suspicion. What the evidence does establish is directional and well documented: crawler traffic consumes disproportionately more bandwidth than its share of requests, it concentrates on the most expensive resources to serve, and it arrives in bursts that size infrastructure requirements rather than averages. Site owners assessing the impact should measure their own logs rather than apply a general figure.

Referral Traffic & Publisher Impact

The counterpart to rising crawl volume is falling referral traffic. The exchange that funded the open web — access to content for visitors in return — is measurably weakening.

  • Google users who encountered an AI summary clicked a traditional search result in 8% of visits, against 15% for users who did not see one — close to half the click rate.Source: Pew Research Center, 22 July 2025 (browsing data from 900 U.S. adults, March 2025)
  • Users clicked a link inside the AI summary itself in just 1% of visits to pages containing one.Source: Pew Research Center, 22 July 2025
  • 26% of visits to a search page with an AI summary ended the browsing session entirely, against 16% for pages with only traditional results.Source: Pew Research Center, 22 July 2025
  • 58% of respondents encountered at least one AI-generated summary in their Google searches during March 2025.Source: Pew Research Center, 22 July 2025
  • For every hour spent searching for information online, only about 15 minutes is now spent on the open web.Source: Cloudflare, 1 July 2026
  • Some of the most heavily crawled site categories have seen human traffic fall by as much as 40% in less than a year.Source: Cloudflare, 1 July 2026
  • Google still accounts for approximately 88% of referral traffic, making it simultaneously the largest source of visitors and a growing substitute for them.Source: Cloudflare, 1 July 2026
  • More than 50 publisher–AI licensing agreements have been signed since 2023 as publishers seek to replace lost referral revenue.Source: Cloudflare, 1 July 2026

Malicious Bots & Automated Attacks

Not all automated traffic is crawling for content. The majority of bot traffic is classified as malicious, and AI has sharply accelerated the volume.

  • Bad bots accounted for 40% of all internet traffic across 2025, up 3 percentage points year on year.Source: Thales 2026 Bad Bot Report, 30 April 2026
  • Thales blocked 17.2 trillion bad bot requests during 2025 — an average of roughly 47 billion per day.Source: Thales 2026 Bad Bot Report, 30 April 2026; daily average calculated by Paid Hosting
  • Daily AI-driven bot attacks rose more than twelvefold in a single year, from about 2 million per day in 2024 to 25 million per day in 2025.Source: Thales 2026 Bad Bot Report, 30 April 2026
  • Financial services absorbed 24% of all bot attacks and the business sector 19%, while the sports sector accounted for just 0.1%.Source: Thales 2026 Bad Bot Report, 30 April 2026
  • Retail — not financial services — is the sector most targeted by AI-driven bots specifically, because dynamic pricing, limited inventory and timed promotions reward continuous automated monitoring.Source: Thales 2026 Bad Bot Report, 30 April 2026
  • The most common API threats are data leakage (26%), remote code execution or remote file inclusion (13%), business logic abuse (13%), automated attacks (8%) and path traversal or local file inclusion (7%).Source: Thales 2026 Bad Bot Report, 30 April 2026
  • Detectable AI traffic represents only a fraction of total AI-enabled activity, because attackers can deploy self-hosted models that never identify themselves — creating a permanent gap between measured and actual volume.Source: Thales 2026 Bad Bot Report, 30 April 2026

Cite this page

Paid Hosting. AI Bot & Crawler Traffic Statistics 2026. August 2026.
https://www.paidhosting.com/bot-traffic-statistics/

Figures, charts and calculations on this page may be reproduced with attribution and a link to this page. Where a statistic is attributed to a third party, please cite that source directly. Press and data enquiries: [email protected]

Sources & Methodology

Every statistic on this page has been read directly from the source listed below and carries its publisher and publication date. Live dashboard measurements are labelled with the exact window they cover and the date they were read, because those figures change continuously — a reader checking Cloudflare Radar today will see a different number than the one quoted here, and that is expected rather than an error. Figures marked as Paid Hosting calculations are arithmetic performed on the published source data, with the inputs stated so the working can be checked. Where sources disagree, we explain the methodological reason rather than choosing one and omitting the others. This page is updated as new data is published; live figures are re-read monthly. This page and our AI Agent Traffic analysis both read Cloudflare Radar, but over adjacent seven-day windows — 2–8 August here, 4–10 August there. Figures for the same metric therefore differ slightly between the two pages. Both are accurate for the window stated beside them; neither supersedes the other.

  • Cloudflare Radar — Bot Traffic, AI Insights and robots.txt datasets (radar.cloudflare.com). Live figures on this page were read on 9 August 2026 and cover 2–8 August 2026 worldwide, except robots.txt data, which Cloudflare last updated 3 August 2026.
  • Cloudflare — “Content Independence Day, one year on: building the business model for the agentic Internet,” 1 July 2026
  • Cloudflare — “Unmasking the crawls with Attribution Business Insights,” 1 July 2026
  • Thales — 2026 Bad Bot Report: Bad Bots in the Agentic Age, 30 April 2026 (13th edition; based on full-year 2025 bot activity across Thales and Imperva networks)
  • Fastly — Q2 2025 Threat Insights Report, 19 August 2025 (AI bot traffic observed mid-April to mid-July 2025)
  • Pew Research Center — “Google users are less likely to click on links when an AI summary appears in the results,” 22 July 2025 (browsing data from 900 U.S. adults who consented to share activity, March 2025)
  • Wikimedia Foundation — “How crawlers impact the operations of the Wikimedia projects,” Diff, 1 April 2025

A note on vendor data: Cloudflare, Thales and Fastly each measure only the traffic crossing their own networks, and each also sells products that mitigate bot traffic. Their figures are the best available primary measurements of automated traffic at scale, and we cite them as such, but they are not neutral third-party audits and none of them observes the whole internet.