Home / Learn / Does crawl budget still matter?
Learn

Crawl budget: settled for Google, still bleeding for the crawlers that don't render

Google has said for years that crawl budget is not something most sites need to manage. Gary Illyes put a number on it in 2020: the vast majority of sites don't have to care, and the ones that do are generally north of a million URLs. That guidance describes Googlebot, a crawler with two decades of infrastructure behind it that absorbs a 30 to 40 percent 404 rate in Search Console without any real cost, according to John Mueller. But GPTBot and Anthropic's crawler are a different kind of visitor. Vercel's measurement of close to a billion requests found ChatGPT's crawler spending 34.82% of its fetches on 404 pages and another 14.36% following redirects, against 8.22% and 1.49% for Googlebot on the same property. A separate, wider measurement across the open web in 2026 found AI crawlers as a group succeeding on only 73% of requests, against an even lower 33.3% for crawlers outside that group, in a dataset where barely 46% of all crawler requests returned a plain 200. None of that is Google's problem to fix. It is yours, if you want an AI engine to keep finding what you publish.

What Google actually said, and who it was talking to

Gary Illyes, from Google's Search Relations team, addressed crawl budget directly on the Search Off the Record podcast on 25 August 2020. His line has been repeated ever since because it is unusually direct for Google: "the vast majority of the people don't have to care about it." He didn't stop there. He also said "there is a substantial segment of the ecosystem that has to care about it," and put a rough number on the split: sites with fewer than about a million URLs generally sit outside that segment.

Illyes returned to the subject on LinkedIn on 2 April 2024, reported the same day by Barry Schwartz at Search Engine Roundtable. His stated mission, in his own words, is to "figure out how to crawl even less, and have fewer bytes on wire," through smarter caching rather than a cruder cut. "Decreasing crawling without sacrificing crawl-quality would benefit everyone," he wrote. Both statements are Google optimizing its own crawler against its own infrastructure bill. Neither one is a claim about GPTBot, ClaudeBot or PerplexityBot, which are different companies' crawlers, built on different budgets, with no obvious reason to inherit Google's tolerance for waste.

The 404 rate Google can shrug off

John Mueller made the tolerance explicit during a Google Search Central SEO hangout on 25 February 2021, when asked about a site showing 30 to 40 percent of its Search Console URLs as 404. "That's perfectly fine," he said, "that's completely natural especially for a site that has a lot of churn," as reported by Search Engine Journal two days later. The only scenario he flagged as worth worrying about was the homepage itself returning a 404.

Google can afford that calm because the scale behind it is enormous. On one property measured by Vercel in its December 2024 crawler analysis, Googlebot made 4.5 billion fetches in a single month and still only spent 8.22% of them on 404s and 1.49% on redirects. A near-40% error rate sitting on top of that kind of volume is noise. The same error rate sitting on top of a crawler making a few hundred million visits a month, as the next section shows, is not noise at all.

The same kind of waste, spent by a much smaller visitor

Vercel's analysis, published 17 December 2024, measured GPTBot making 569 million fetches on that same property in the same month, with 34.82% landing on 404 pages and another 14.36% following redirects. Anthropic's crawler, at 370 million fetches, spent 34.16% of its visits on 404s. Roughly a third to more than a third of every visit these crawlers make is spent on something that was never going to produce a citation, against under a tenth for Google on the identical site. These are one property's numbers, not a universal constant, but the gap between the two crawlers on the same URLs is what matters: a 404 costs GPTBot about four times what it costs Googlebot, because GPTBot has roughly an eighth of Googlebot's visits to spend finding anything at all. Our guide to whether AI crawlers render JavaScript covers the rest of this same dataset, including why a page can rank in Google and return nothing to GPTBot at all; this is the half of the story about what happens to the visits that do land on a real URL.

A fresher measurement across the open web says it again, differently

A separate report gives a 2026 reading on the same question at a much larger scale. SEOmator's Crawl Waste Report, published 15 July 2026 by Ben Kaiser, drew on Cloudflare Radar's Web Crawlers dataset across the 28 days ending 19 July 2026, supplemented by the firm's own crawl of 12 sites. Across that sample, only 45.9% of crawler requests on the open web returned a plain 200, with another 1.2% landing on an efficient 304. The rest split across 403 Forbidden (20.6%, the single largest non-200 bucket), 301 redirects (8.0%), 404s (7.8%) and 429 rate-limiting (6.3%).

Inside that mix, AI crawlers as a category actually came out ahead of non-AI bots: a 73.0% success rate against 33.3%. That is a real point in AI crawlers' favor, not something to bury by only citing the worse number. It still means more than a quarter of every AI-crawler visit, across the wider web, lands on something other than a clean page. The report separately breaks out AI crawler traffic by volume rather than success: Googlebot carries the largest share of that traffic at 24.6%, ahead of ClaudeBot (18.7%), Meta's crawler (11.2%) and GPTBot (9.7%). That ranking describes how much of the AI-crawler traffic each bot accounts for, not how often each one succeeds, and the report does not publish a per-bot success rate, so the two numbers should not be collapsed into one ranking.

What Google's own documentation says the waste actually is

Google's large site owner's guide to crawl budget, last updated 22 July 2026, defines crawl budget as the set of URLs Google "can and wants" to crawl, set by a capacity limit on one side and crawl demand on the other. The one lever the guide hands back to the site owner is what it calls perceived inventory, and it names the exact waste categories: long redirect chains, which it says have "a negative effect on crawling"; soft 404s, which it warns "will continue to be crawled, and waste your budget" precisely because they return a 200 status while being empty; and duplicate URL variants, which split one page's worth of value across many crawlable addresses instead of one. The guide also draws a distinction worth knowing before reaching for a quick fix: a noindex tag still gets requested and then dropped, which still costs a crawl, where a robots.txt disallow stops the request from happening at all.

Those are exactly the categories the Vercel and SEOmator numbers above are measuring from the outside. Google is describing, in its own documentation, the same redirect chains and dead ends that show up as a 34% 404 rate for GPTBot and a 7.8% 404 rate across the open web. The difference is which crawler can absorb the cost without anyone noticing.

Where this usually shows up: an old redirect map nobody retired

The most common source of a long redirect chain is not a mistake made on purpose. It is a site that has moved once, been restructured once, and never cleaned up the URLs each move left behind, so a crawler hits three hops before reaching a live page. Our guide to the migration recovery timeline covers a dataset of 1,052 tracked migrations where the median full recovery took 304 days, and a rushed or incomplete redirect map is one of the named reasons recovery drags that long. The website migration SEO guide covers the mechanics of building a one-to-one map instead of letting old redirects stack on top of newer ones. Fixing the chain matters for Google's crawl efficiency too, but it matters more for a crawler that has, by the numbers above, roughly an eighth of Googlebot's total visits to spend finding your site in the first place.

What we could not verify

SEOmator's report does not disclose a total request count for the Cloudflare Radar portion of its analysis, only the 28-day window, so the 45.9% and 73.0% figures should be read as that dataset's finding rather than a number with a known sample size attached. Its own supplementary crawl, 389 internal URLs across 12 sites, is small next to the open-web numbers and is reported separately rather than blended into them, which this piece has kept separate as well. Vercel's figures come from one property at one point in time, not a market-wide census, so the exact percentages are illustrative of the gap between crawlers rather than a number every site should expect to reproduce. Illyes' one-million-URL threshold is a rule of thumb from a podcast, not a published specification, and Google has not restated a precise number since. No published study ties crawl-budget efficiency directly to AI citation rate; the connection drawn here, that a visit wasted on a 404 is a visit that can't discover a real page, is a mechanism argument, not a measured result.

Sources

Crawl budget guidance and the million-URL threshold: Matt G. Southern, Google: Most Sites Don't Need to Worry About Crawl Budget, Search Engine Journal, 25 August 2020, reporting Gary Illyes on the Search Off the Record podcast. Illyes' 2024 statement: Barry Schwartz, Gary Illyes From Google Wants Googlebot To Crawl Less, Search Engine Roundtable, 2 April 2024. The 404-tolerance quote: Matt G. Southern, Google: Fine if 30-40% of URLs in Search Console Are 404s, Search Engine Journal, 27 February 2021. Crawler fetch-mix and error rates: Vercel, The rise of the AI crawler, 17 December 2024. Open-web crawl success rates: Ben Kaiser, Crawl Waste Report 2026, SEOmator, 15 July 2026. Google's own crawl budget definitions and recommendations: Google Search Central, Large site owner's guide to managing your crawl budget, last updated 22 July 2026.

Common questions

Does crawl budget still matter for SEO?

For Google, usually not. Gary Illyes said in August 2020 that the vast majority of sites don't have to care about crawl budget, drawing the practical line at roughly a million URLs. John Mueller separately said a 30 to 40 percent 404 rate in Search Console is perfectly fine. Neither statement was about AI crawlers, which behave differently on the same kind of site.

Do AI crawlers like GPTBot have their own version of crawl budget problems?

Yes, and the numbers are worse than Google's. Vercel measured GPTBot spending 34.82% of its fetches on 404 pages and 14.36% on redirects, against 8.22% and 1.49% for Googlebot on the same property. A 2026 open-web measurement found AI crawlers succeeding on 73% of requests, ahead of non-AI bots at 33.3%, but still losing more than a quarter of every visit.

What actually wastes crawl budget?

By Google's own documentation: long redirect chains, soft 404s that keep getting crawled instead of dying cleanly, and duplicate URL variants that split one page's value across many addresses. A noindex tag still gets requested and dropped, which still costs a visit, where a robots.txt block prevents the request entirely.

Find out what your own crawlers are wasting

The Technical SEO Audit maps redirect chains, orphan pages and crawl waste alongside indexing, Core Web Vitals and structured data, scored and ordered into a fix roadmap your developer can run Monday.

Technical SEO Audit, from $295
AI Search Ready badge by Faro for heykadima.com
AI Search ReadyOur own site is built to the exact standard we sell, so AI engines can find, read and recommend it. Verified by Faro, built by the same team.
See the proof →