Crawl Budget for Large Aggregators: Why Travel and Comparison Sites Keep Fixing the Wrong Problem

1. What Happened? – Crawl Budget for Large Aggregators

Nothing announced, which is exactly why this is worth writing about.

Google updated its crawl budget and faceted navigation documentation over the past few months, and buried in the revisions are two sentences that change how large aggregators should think about their architecture.

The first: while each crawler has a different crawl demand, the crawl capacity limit is shared across all crawlers – meaning high demand from one crawler can reduce the capacity available for others.

Read that again with a comparison site or an OTA in mind. Googlebot is not the only Google crawler on your infrastructure. AdsBot is there too, and Google is explicit that AdsBot generally has a higher demand when a site is running dynamic ad targets. Travel and insurance aggregators are among the heaviest Dynamic Search Ads users in existence, and that crawl logs roughly every three weeks.

Which means, on a great many aggregator sites, the paid media team is quietly consuming the capacity that the organic team is trying to win. Nobody owns that trade-off because nobody has ever seen it stated as a trade-off.

The second: blocking or hiding already-crawled pages from recrawls will not shift your crawl budget to another part of your site unless Google is already hitting your serving limits.

That single caveat invalidates the most common crawl budget work in the industry. The standard playbook – robots.txt the facets, watch the guides get crawled more – only works if you are capacity-bound. If you are demand-bound, you have just spent a quarter of engineering time to change nothing.

Almost nobody checks which one they are before starting.

We looked at this while building the Car Insurance Category Leaderboard, where the comparison sites show a consistent and revealing shape: large brand demand, comparatively weaker search visibility. Compare the Market ranks 1st on brand and category demand but 11th on Google Search Visibility. GoCompare ranks 2nd on demand and 14th on visibility. Money Supermarket, 3rd and 9th.

Brand demand that large with visibility that far behind is a gap worth interrogating. Part of it is SERP composition and part is intent mix. But on sites with URL inventories in the millions, part of it is simply that Google is not reaching or refreshing enough of the site to compete on the queries that matter.

That part is fixable. It just is not fixable with the workflow most SEO teams are running.

2. Why Does It Matter? – Crawl Budget for Large Aggregators

Aggregators have the worst possible URL shape

Google defines the problem -> faceted navigation based on URL parameters can generate infinite URL spaces, and because crawlers cannot tell whether those URLs are useful without crawling them first, they typically access a very large number of them before determining they are useless.

Now apply that to travel, where dates are a facet. Destination × origin × outbound date × return date × passengers × cabin × stops × airline × price band. A mid-sized OTA can generate more unique URLs in a week than a national newspaper has published in its history, and essentially none of them deserves to be indexed.

Insurance comparison has the same problem with different dimensions glued into the daily generated URLS: vehicle, cover type, driver age, postcode, excess, add-ons. Same infinite space, same wasted crawl.

The waste has a second-order cost

Google’s framing is that when crawling is spent on useless URLs, crawlers have less time to spend on new, useful URLs – so the damage is not only wasted server capacity. It is slower discovery of the content you actually want ranking.

For an aggregator, that content is usually the commercially valuable layer: the guides, the destination hubs, the provider comparison pages, the editorial that wins the top-of-funnel queries and feeds the AI surfaces. Those pages are a rounding error in your URL count and the majority of your organic value, and they are queued behind a million date permutations.

Empty result pages are a soft 404 farm

Google names an empty internal search result page as a cause of soft 404s, and warns that soft 404 pages continue to be crawled and waste budget.

Every travel aggregator serves empty result pages constantly. No flights on that route on that date. No hotels matching those filters. No policies for that combination. If those pages return HTTP 200 with a friendly “no results” message – and on most aggregators they do – you have built a soft 404 farm at industrial scale, and it is being recrawled indefinitely.

Google’s guidance for faceted navigation is clear on the fix: return a 404 status code when a filter combination returns no results, including for duplicate filters, nonsensical combinations and nonexistent pagination URLs. And specifically, do not redirect to a common “not found” page – serve the 404 at the URL where it was encountered.

That last clause is the one dev teams get wrong. Redirecting empty results to a generic error page feels tidier and is worse.

noindex is the wrong solution, and everyone reaches for it

Google states it directly: don’t use noindex for crawl budget purposes, because Google will still request the page and then drop it when it sees the tag, wasting crawling time.

noindex, follow on facet pages is close to universal on aggregator sites. It is a sensible instruction about indexing and it does nothing whatsoever for crawl efficiency – the request still happens, the render still happens, the capacity is still consumed. Teams have been paying the full cost of crawling those pages for years while believing they had solved the problem.

Your subdomains have separate budgets

Google’s crawling infrastructure defines a site as a unique hostname, so https://www.example.com/ and https://flights.example.com/ are treated as separate sites with separate crawl budgets.

That is an architectural lever, and aggregators are unusually well placed to use it because many already split the funnel across hosts – content on www, the quote or search engine on a separate subdomain. Where that split already exists, the volatile transactional layer is not competing with the editorial layer for the same capacity pool. Where it does not exist, moving it is a serious piece of work with real trade-offs, and it should be a deliberate architectural decision rather than an SEO deliverable.

The AI crawlers changed the arithmetic

The shared-capacity rule was written about Google’s own crawlers. The principle generalises. Your infrastructure now also serves OAI-SearchBot, GPTBot, ClaudeBot, PerplexityBot and a growing list of others, all requesting pages from the same origin, all competing for the same server headroom.

Crawl efficiency used to be an SEO concern about Google. It is now a question of whether every system that decides what to recommend can actually read your site.

3. Who Is Affected? – Crawl Budget for Large Aggregators

Google scopes its own guidance carefully, and it is worth repeating honestly because most sites do not need this article. The advice applies primarily to sites with more than a million unique pages whose content changes at least weekly; medium or larger sites with 10,000+ pages and very rapidly changing content; and sites with a large portion of their URLs sitting in Search Console as Discovered – currently not indexed.

That third criterion is the useful one. It is a diagnostic you can run today, and it is the clearest signal that Google knows about your URLs and is choosing not to fetch them.

  • Travel aggregators and OTAs.
    The most affected category, because dates make the URL space genuinely infinite and inventory changes hourly. Flights, hotels, packages, car hire, rail.
  • Insurance and finance comparison sites.
    Very large facet spaces, heavy Dynamic Search Ads investment, and – as the leaderboard shows – large brand demand that the organic estate is not fully converting into visibility.
  • Marketplaces and classifieds.
    Property, automotive, recruitment, tickets. Listings churn constantly, and expired listings that return 200 are the same soft 404 problem wearing different clothes.
  • Retail marketplaces with deep facets.
    Colour × size × brand × price × availability, multiplied across a catalogue.
  • Multi-market aggregators.
    Every facet problem multiplied by every locale. This is where crawl waste becomes genuinely large and where it is least visible, because nobody audits the Portuguese site.
  • Anyone running heavy Dynamic Search Ads alongside a large organic traffic.
    Whatever the vertical. If your DSA targets cover your whole site, AdsBot demand is drawing from the same capacity pool as Googlebot, and no one in either team has been briefed on it.
  • Who this does not apply to: if your pages are generally crawled the same day they are published, you do not have a crawl budget problem. Google says so explicitly. Keep the sitemap current, check the Page Indexing report, and go and work on something that matters.

4. What Should Businesses Do? – Crawl Budget for Large Aggregators

The order of operations matters more than any individual fix. Diagnose first.

4A. For the whole team: diagnose capacity-bound or demand-bound

  • Open the Crawl Stats report. Look at the host availability graphs for instances where Googlebot requests exceeded the red limit line, click into the graph to identify the failing URLs, and correlate them with what was happening on your infrastructure.
  • Run the URL Inspection tool on a sample. If it returns Hostload exceeded, Googlebot cannot crawl as many URLs as it has discovered. That is your answer.

Then branch:

  • If you are capacity-bound – hitting the limit line, seeing hostload warnings, returning 5xx or 429 under crawl load – then reducing waste genuinely reallocates budget, and everything in 4C will pay off. Google’s own test for this is direct: if crawling seems to cross the limit line often, increase serving resources for a month and see whether crawl requests rose in the same period.
  • If you are demand-bound – no availability problems, Googlebot comfortably under capacity, large volumes sitting in Discovered – currently not indexed – then robots.txt will not move the needle, because Google will not reallocate freed capacity it was not using. Your problem is that Google does not consider enough of your inventory worth crawling.

That is a content and architecture problem wearing a technical costume, and Google frames the it accordingly: the two ways to get more crawl budget are adding server resources, or optimising content quality for the Google product you are targeting – with popularity, overall user value, content uniqueness and serving capacity named as the relevant factors for Search.

Four million thin pages are not going to become high-quality. The honest answer for a demand-bound aggregator is to have dramatically fewer, dramatically better indexable pages. Which brings us to the content team.

4B. For the content and marketing team: decide what deserves to exist

Crawl budget work fails when it is handed to devs as a purely technical ticket. Someone has to decide which pages are supposed to rank, and that decision is editorial.

  • Build the indexable inventory deliberately. For each template, answer one question: is there real, recurring search demand for this page, and does it say something a generic filtered view cannot? “Cheap flights to Malaga” almost certainly qualifies. “Flights to Malaga departing 14 March returning 22 March for 3 passengers with 1 stop” does not, and never will.

The output is a defensible list: these facet combinations are pages we will invest in, everything else is a filtered view that exists for users and not for search.

  • Promote the survivors to real pages. A facet combination worth ranking deserves to stop being a query string. Give it a clean path, an editorial introduction, unique supporting content, internal links from the hub, and a place in the sitemap. Google’s own guidance for path-encoded filters is that the logical order of filters must always stay the same and no duplicate filters can exist – so if you encode /flights/uk/malaga/direct, that order is fixed forever and the routing must reject variants.
  • Write the hub layer that actually earns crawl demand. Destination guides, provider comparisons, “best X for Y” pages, seasonal editorial. This is what popularity and uniqueness look like in practice, and for a demand-bound site it is the lever, not an adjunct to the lever.
  • Own the empty-results decision. Returning a 404 on a zero-result search feels hostile to marketing teams. It is not: the page still renders, the user still sees helpful alternatives, the status code just tells crawlers the truth. Agree this once, in a room, with marketing present – otherwise engineering will implement it and marketing will reverse it in three months.
  • Ask paid media about DSA coverage. If Dynamic Search Ads target your entire site, AdsBot is crawling the whole estate every three weeks from the same capacity pool as Googlebot. Google’s own solution when that becomes a problem is to either limit ad targets or increase serving capacity. That is a conversation between two teams who have probably never had it.

4C. For the development team: implement

Block facet crawling at the source. Where facets are parameter-based and not on the indexable list, disallow them in robots.txt. Google’s own example pattern:

user-agent: Googlebot
disallow: /*?*products=
disallow: /*?*color=
disallow: /*?*size=
allow: /*?products=all$

Applied to travel, that looks like disallowing depart=, return=, pax=, cabin=, stops=, sort= and friends, while allowing the unfiltered route page that you do want indexed.

Or move filters to URL fragments entirely. Google generally does not support URL fragments in crawling and indexing, so a fragment-based filtering mechanism has no crawling impact at all – positive or negative:

https://example.com/flights/uk-malaga#depart=2027-03-14&pax=3&cabin=economy

This is the cleanest available answer and the hardest to retrofit. On a rebuild, it should be the default.

  • Understand why canonical and nofollow are the weaker options. Google acknowledges both as ways to signal preference, and says they are generally less effective in the long term than robots.txt or fragments. Use them as supplements, not as the strategy.
  • Fix the status codes. Zero-result filter combinations, duplicate filters, nonsensical combinations and pagination beyond the last page all return 404 at the URL requested – not a redirect to a shared error page. Permanently removed listings return 404 or 410; Google notes that a 404 is a strong signal not to crawl that URL again, while blocked URLs stay in the crawl queue much longer.
  • Audit for soft 404s. Check the Page Indexing report. Then find the causes: pages rendering blank or near-blank, prominent error messages in the rendered output, missing critical resources. Google flags blocked resources, too many resources on a page, server errors and slow or very large resources as reasons a page can be interpreted as a soft 404.
  • Implement conditional requests. Google generally supports If-Modified-Since and If-None-Match, and will reuse the previously crawled version on a 304. Better still, Google notes you can return a 304 with no response body for any Googlebot request where the content has not changed since the last visit, independently of the request headers.
  • For an aggregator this is close to free money. Your guides, destination pages and provider profiles change rarely; your pricing changes constantly. Serving 304s on the stable layer preserves capacity for the volatile one.
  • Reference shared resources from one URL. Where a common image or script is reused across pages, Google asks that it be referenced from the same URL each time so it can cache and reuse it rather than re-requesting it. Aggregators with aggressive per-page CDN transforms are re-requesting the same asset thousands of times.
  • Block heavy non-critical resources. Google’s guidance is to prevent large but unimportant resources from being loaded by Googlebot via robots.txt – blocking only what is not needed to understand the page, such as decorative images. Be careful here: block something the page needs to render and you have created a soft 404.
  • Kill redirect chains. Named explicitly as having a negative effect on crawling. On aggregators they accumulate silently through years of route and campaign changes.
  • Keep sitemaps honest. Include <lastmod> and mean it. Google’s list of things to avoid is worth pinning up: do not resubmit the same unchanged sitemap several times a day, do not expect everything in a sitemap to be crawled or crawled immediately, and do not include URLs you do not want in Search – that wastes budget on pages you do not want indexed.
  • Emergency throttling, with the safety note. If Googlebot genuinely overwhelms you, returning 503 or 429 temporarily will slow it, and Google retries those URLs for about two days. But returning those codes for more than a couple of days causes Google to permanently slow or stop crawling, and URLs will be dropped from the index. This is a fire extinguisher, not a policy.

4D. Rolling it out across the site

  1. Diagnose. Crawl Stats host availability, URL Inspection on a sample, and the proportion of URLs in Discovered – currently not indexed. Capacity-bound or demand-bound. Write the answer down before anyone opens a ticket.
  2. Inventory your URL space. Server logs, not a crawler. Group every Googlebot request by template and parameter pattern, and rank by request volume. The top ten patterns will typically account for the overwhelming majority of crawl, and on most aggregators the top pattern is something nobody wants indexed.
  3. Segment log data by user agent. Split Googlebot from AdsBot from the AI crawlers. This is the report that makes the shared-capacity problem visible to people outside SEO, and it is the one that gets budget approved.
  4. Agree the indexable inventory with content and marketing. Signed off, written down, one list.
  5. Ship status codes first. Zero-result 404s and expired-listing 410s are the highest-value, lowest-risk change, and they stop the bleeding while the bigger work is scoped.
  6. Then ship the crawl blocks. robots.txt patterns or a fragment migration. Roll out by parameter group, not all at once, so regressions are attributable.
  7. Then 304 support on the stable content layer.
  8. Segment your sitemaps by template – guides, hubs, listings – with accurate <lastmod>. Segmented sitemaps make the Page Indexing report diagnostic rather than decorative.
  9. Measure over 8 to 12 weeks. Total crawl requests, average response time, requests by response code, crawl distribution across templates, and indexed coverage of the pages you actually care about. Not rankings. Rankings move later and for other reasons too.
  10. Re-diagnose quarterly. A site that was capacity-bound in March may be demand-bound by September, and the treatment changes with it.

5. What We’re Watching Next

  • Crawl efficiency becoming an AI visibility issue. Google’s shared-capacity principle describes Google’s crawlers, but your origin serves everyone. As assistant crawlers become a material share of traffic, the question stops being “can Googlebot keep up” and becomes “can every system that decides what to recommend actually read us”. Sites that solved this for Google will find they solved it for the rest.
  • Conditional requests becoming table stakes. 304 support is currently a differentiator among large sites, which is remarkable given how long the mechanism has existed. We expect it to become an expectation, particularly as crawl volumes from non-Google agents rise.
  • More pressure on thin permutation pages. The long-standing aggregator strategy of indexing a vast combinatorial estate and letting Google sort it out gets less viable as retrieval consolidates around fewer, better answers. Demand-bound sites will keep discovering that the fix was never technical.
  • The paid and organic capacity conversation reaching the boardroom. Once the user-agent-segmented log report exists, the trade-off between DSA coverage and organic crawl becomes a budget conversation rather than an SEO complaint. We expect that report to become a standard artefact in aggregator technical audits within a year.
  • Brand demand and visibility diverging further. The leaderboard pattern – very strong brand demand, comparatively weaker search visibility – is the aggregator signature. As AI surfaces mediate more discovery, converting demand into visibility depends increasingly on being fully readable and fully current. Crawl efficiency is the foundation of that.

6. About Szymaniak Digital

Szymaniak Digital is an enterprise AI SEO consultancy working with senior marketing teams and the engineering teams who ship for them. We run log-level crawl diagnostics on large aggregators, separate the capacity problems from the demand problems, and write clean specs developers can implement without a translation.

We also publish the Category Leaderboard, which scores brands on Google Search Visibility, Brand and Category Demand, and AI Recommendation Score – because the gap between how much demand a brand has and how visible it actually is usually has a cause you can fix.

If your URL count runs to seven figures and nobody can tell you what Googlebot spent last month on, that is where we would start. Book a technical SEO consultation.

Contact Us!

Need More Enquiries from Google and ChatGPT? 

📞 0330 223 7866

Close Welcome Bar
Scroll to Top