Notes · Note

Googlebot Fetch Limits and the New Crawler IP Path

Googlebot still documents a 2MB HTML fetch cutoff. Crawler IP JSON files moved, and the directory URL is now a docs page. What operators should change.

Updated Sep 21, 2026 Reviewed Sep 21, 2026 en

On 2026-03-31, Google Search Central published two crawler posts on the same day. One explained how Googlebot fetches bytes. The other moved the public IP range files. Both posts were still live on 2026-09-21. The IP files themselves had moved again in practice: the directory URL is now a documentation page, and the lists live as named JSON files. This brief records the operator facts. It is not a rewrite of either announcement.

Treat the two posts as one infrastructure change: how much HTML Googlebot will take, and which files allowlist teams should fetch.

What changed about fetching

Google said Googlebot is not a single program. It is one client of a shared crawling platform. Other Google products use that infrastructure under other crawler names.

The fetch limit that matters for HTML pages is 2MB per URL, including HTTP headers. PDFs go to 64MB. Google’s crawler overview still states a 15MB default for crawlers that do not set a tighter product limit. Googlebot’s HTML cap is the tighter one.

Oversize HTML is truncated, not rejected. The first 2MB is treated as the complete file. Bytes after the cutoff are not fetched, rendered, or indexed. Referenced resources have their own per-URL byte counters. The Web Rendering Service only executes code that was actually retrieved, and it is stateless between requests.

The operator implication is placement, not a new robots token. Titles, canonicals, answers, evidence, and important internal links need to sit early in the HTML. A page can return 200, stay allowed in robots.txt, and still lose the useful text if that text starts after a large inline payload.

CheckWhy it matters now
HTML weight before the first answerGooglebot may never see the answer if it sits below the 2MB cut.
Critical tags high in the documentTitle, robots, and canonical signals that arrive late can be truncated away.
Heavy inline CSS, SVG, or JSONThose bytes count toward the same per-URL budget.
Separate resource URLsImages and scripts have their own counters, but the HTML budget is still 2MB.
PDF sourcesThe documented PDF cap is 64MB, not 2MB.

Do not invent a smaller or larger Googlebot HTML cap from a vendor blog. Recheck the 2026-03-31 post before you change the number.

What changed about IP files

The second post said Google crawler IP range JSON files were moving from /search/apis/ipranges/ to developers.google.com/crawling/ipranges/. Google said old paths would redirect within six months of 2026-03-31.

On 2026-09-21, that is only half true:

URL you might still have bookmarkedWhat it did on 2026-09-21
developers.google.com/crawling/ipranges/ (no filename)Redirected to the verify-Google-requests documentation page. Not a JSON index.
.../crawling/ipranges/common-crawlers.jsonServed JSON. This is the common-crawler list, including Googlebot. File creationTime was 2026-09-18.
.../crawling/ipranges/special-crawlers.jsonServed JSON.
.../crawling/ipranges/user-triggered-fetchers.jsonServed JSON.
Old .../search/apis/ipranges/googlebot.jsonRedirected to common-crawlers.json and still served JSON.

The live operator page is Verify Google crawler requests. It names those JSON files and still says to confirm Googlebot with reverse DNS, then a forward lookup to the same IP. Do not allowlist on user-agent alone.

Do not tell a firewall job to fetch the directory with no filename. Fetch common-crawlers.json for Googlebot. Use the other two files only if you are also allowing special-case crawlers or user-triggered fetchers.

Google still identifies crawlers by user-agent, source IP, and reverse DNS. An IP allowlist that still requests googlebot.json, or that bookmarks a directory that now loads a docs page, is a crawlability incident, not a content incident.

What these posts are not

These announcements are not a citation report. They are not a ranking change. They are not an llms.txt endorsement. They do not replace Search Console performance views or generative AI impression reports.

A page that is truncated before its answer can fail retrieval and later fail citation. That is still an access problem. Fix the HTML, then measure answers. Do not fold a fetch-limit incident into a single AI visibility score.

QuestionThese posts answerStill needs another source
How many HTML bytes will Googlebot fetch?Yes: 2MB including headers, as of the 2026-03-31 postNo
Where should IP allowlists read Google crawler ranges?Yes: named JSON files under /crawling/ipranges/, confirmed 2026-09-21Recheck the verify page if a filename 404s
Will Google cite the page?NoPrompt capture and citation review
Should the site add llms.txt?NoA separate access and maintenance decision
Did generative AI impressions change in Search Console?NoThe later GSC generative AI report brief

What operators should do

  1. Weigh a sample of important public HTML pages, including headers. If the useful answer starts late, move it up. Do not wait for a crawl error that says “truncated.”
  2. Keep public source pages allowed in robots.txt. A fetch limit does not make blocking safer.
  3. Point IP allowlists at https://developers.google.com/crawling/ipranges/common-crawlers.json for Googlebot. Add special-crawlers.json and user-triggered-fetchers.json only if those clients matter. Do not fetch the directory URL.
  4. If a job still requests googlebot.json under /search/apis/ipranges/, change the filename. The old URL redirected on 2026-09-21; do not depend on that redirect remaining.
  5. Verify suspected Googlebot hits with reverse DNS, then a forward lookup to the same IP. Do not allowlist on user-agent alone.
  6. Recheck the verify-Google-requests page and the 2026-03-31 fetch-limit post before you change the byte number or the JSON filenames again.

Example: a category guide returns 200 and is in the sitemap, but the first 1.8MB is theme CSS and a product JSON blob. The definition and comparison table sit after that. Googlebot can fetch the URL and still never store the table. The next edit is the template, not a new Markdown map.

Counterexample: treating the IP path change as proof that Googlebot stopped crawling the site, or bookmarking /crawling/ipranges/ with no filename and calling the docs page a broken IP feed. If logs show fewer hits, check the JSON URL, DNS verification, and robots rules before rewriting content.

Why this belongs in a brief

Notes exist to date a source change and name the operator move. This one is: Google documented a 2MB Googlebot HTML cutoff, moved crawler IP files, and the directory URL is now the verify page rather than a file listing.

It does not become a Guide. If you need the later measurement signal from Search Console, use Search Console generative AI performance reports.