Reducing Google’s Crawl Rate: Retry-After – Documentation Update

1. What Happened? – Reducing Google’s Crawl Rate: Retry-After – Documentation Update

Google has restructured the emergency section of its Reduce the Google crawl rate documentation and added guidance and examples for the Retry-After HTTP header.

It is worth being precise about what actually changed, because some of the coverage around this will probably make it sound more significant than it is.

Google’s own changelog makes it clear that support for Retry-After is not new. It was already documented elsewhere, in Google’s guide about temporarily pausing or disabling a website.

What has changed is where the information now appears.

Google has moved it directly into the guide people are likely to read when their server is struggling with Googlebot traffic.

So this is a documentation change, not a change to Google’s crawling behaviour.

Googlebot does not suddenly work differently because of this update.

The reason I think the change is still important is because of the page Google has put this information on.

Read the updated guidance carefully and there are several warnings that, taken together, describe a very powerful emergency control — and one that can have a much wider impact than the person switching it on might expect.

What the guide now says

The mechanism

If you need to reduce crawling quickly for a short period, Google suggests returning a 500, 503 or 429 status code instead of 200 to crawl requests.

Google gives a short period as the example: a couple of hours or one to two days.

The impact

When Google’s crawling systems see a significant number of these status codes, Google reduces the crawl rate for the whole hostname.

That includes URLs that are still working normally.

The recovery

Google says it will increase crawling again automatically once the errors reduce.

That does not mean recovery is necessarily immediate, which we will come back to later.

The Retry-After header

For 503 and 429, you can also return a Retry-After header.

Google gives two forms:

HTTP/1.1 503 Service Unavailable
Retry-After: 120

Or an absolute UTC date and time:

HTTP/1.1 503 Service Unavailable
Retry-After: Wed, 21 Oct 2026 07:28:00 GMT

The two-day warning

Google also gives a very important warning: do not keep returning these status codes for too long.

If Googlebot sees them on the same URL for multiple days, that URL may be removed from the index.

The side effects

This is probably the most important part of the page.

Google says reduced crawling can mean:

  • fewer new pages being discovered
  • existing pages being refreshed less often
  • prices and product availability taking longer to update in Search
  • removed pages staying in the index for longer

And there is another consequence that is particularly interesting for paid media.

Google says your Google Ads campaigns may be cancelled or paused, and your ads may stop serving.

That creates a connection between crawling, engineering and paid media that many businesses will not have thought about.

What causes a crawl spike?

Google lists several common causes and tells you to check your access logs before doing anything.

The examples include:

  • faceted navigation and sorting or filtering
  • calendars with lots of date-based URLs
  • Dynamic Search Ads targets

The last resort

Google also says that, if serving errors is not practical, you can submit a request about an unusually high crawl rate.

There is an important limitation here.

It may take several days to evaluate.

And you cannot use the request to ask Google to crawl your site more quickly.

So the control is largely one-way: it is easy to reduce crawling, but increasing it again is not something you can simply request.

2. Why Does It Matter?

The emergency brake is connected to your ad account

There is one detail in the updated documentation that deserves much more attention.

Google lists Dynamic Search Ad targets as one possible cause of a crawl spike.

The same page also says reducing crawl activity can cause Google Ads campaigns to be cancelled or paused, and ads may stop serving.

Put those two points together and you get a situation like this.

Paid media launches or expands a Dynamic Search Ads campaign.

The targeting covers a large part of the site, including a large number of parameterised or dynamically generated URLs.

Crawling of those URLs rises.

Server load increases.

Engineering sees the traffic, identifies Google crawlers and responds exactly as Google recommends by returning 503 or 429 responses.

Google reduces crawling across the hostname.

Google’s own documentation warns that Ads campaigns may then pause or stop serving.

Paid media sees conversions disappear and has no idea why.

Nobody has necessarily made a bad decision.

The problem is that the teams are working separately.

One team’s campaign contributes to the crawl increase.

A second team’s response to the server problem affects the first team’s campaign.

Every decision can make sense on its own.

The failure happens because there is no connection between the decisions.

This is a good example of a wider problem we see across technical SEO.

One team makes a sensible change.

That change affects a system another team depends on.

Nobody has a process that connects the two.

We have seen the same pattern with security teams blocking legitimate crawlers with 403 responses, and with paid media activity increasing crawl demand on large websites.

Here, the loop is even more obvious.

There is one important caveat.

Google documents the Ads consequence, but it does not explain exactly why it happens.

The most likely explanation is that Google’s crawling systems share infrastructure or crawl capacity across some of these processes, meaning problems with landing pages or reduced fetching can affect the checks Google Ads depends on.

That is our interpretation, not something Google explicitly confirms.

The consequence is documented.

You cannot crawl just one part of the site

The second warning that is easy to miss is that the crawl reduction applies across the whole hostname.

That includes URLs that are still returning content normally.

This matters because the obvious engineering response is often to try to isolate the problem.

For example:

“The faceted URLs are causing the problem, so we will return errors on those URLs and leave the product pages alone.”

That may help reduce the immediate server load.

But it does not mean Google will keep crawling your product pages at the same rate.

Google says the crawl-rate reduction applies to the hostname.

So your product pages can also be crawled less frequently.

This is why it is important to separate two different ideas.

You can scope the URLs that return the error.

You cannot necessarily scope Google’s resulting crawl-rate reduction to just those URLs.

Scoping is still useful because it can reduce the number of URLs exposed to the longer-term indexing risk.

But it does not stop the hostname-wide crawl reduction.

The hostname matters too

The unit here is the hostname, not the entire domain.

For example:

  • shop.example.com
  • www.example.com

are different hosts.

That means architecture can affect how isolated the problem is.

If the problematic faceted navigation sits on a separate subdomain, crawling that hostname does not necessarily affect the main website.

If the faceted navigation and the main product catalogue share the same hostname, both are affected.

A technical architecture decision made years ago can therefore determine how precisely you can deal with a crawl problem today.

“It automatically increases again” does not mean “it comes back immediately”

Google says crawling will increase again after the errors reduce.

That is true.

But it would be a mistake to read that as:

“Turn the crawl off and everything goes straight back to normal.”

It does not necessarily work that way.

The crawl timing data Gary Illyes presented at Search Central Live Deep Dive gives typical crawl-capacity update windows of around four hours to one to two weeks, with recovery potentially taking one to three weeks. Also documented by Szymaniak Digital in our crawl, index, and serve timings update.

Google can reduce crawling very quickly.

Recovery can take much longer.

So the practical version of the documentation is:

The brake can engage quickly. The recovery can take days or weeks.

That means a six-hour crawl does not necessarily mean only six hours of reduced crawling.

The lower crawl rate can continue after the errors stop.

During that recovery period, the same side effects still matter:

  • new pages can take longer to appear
  • existing pages may be refreshed less often
  • prices can become stale
  • deleted pages can remain in search for longer

This is the same basic problem we have seen in other crawl incidents: a relatively short infrastructure problem can have a much longer SEO impact.

The difference here is that the decision to slow crawling is deliberate.

500, 503 and 429 are not the same thing

Google groups 500, 503 and 429 together because all three can reduce crawling.

That does not mean you should treat them as interchangeable.

The documentation itself gives us a clue.

Google specifically discusses Retry-After with 503 and 429.

Status codeWhat it tells the clientRetry-AfterAppropriate for a deliberate reduction?
500 Internal Server Error“Something is broken.”Not in this guidanceNo
503 Service Unavailable“The service is temporarily unavailable.”YesYes, when the service is genuinely degraded
429 Too Many Requests“You are sending too many requests.”YesYes, especially when users are still being served normally

The practical point is that your status code should match what is actually happening.

If customers can browse the site normally but you are rate-limiting Google’s crawler, 429 is the more accurate message.

A 503 suggests the service is temporarily unavailable.

A 500 suggests something has gone wrong.

There is also a practical operational reason to avoid using 500 unless there really is an internal server error.

Your monitoring systems will see it as an error.

Your alerts may start firing.

Your error dashboards will light up.

During an incident, that is exactly when you want your monitoring to be telling the truth.

Google may respond to all three codes in a similar way for crawl reduction.

Your own systems will not.

Choose the status code that accurately describes what you are doing.

The 48-hour warning is per URL

The deindexing warning is worth reading carefully.

Google says that if Googlebot keeps seeing these status codes on the same URL for multiple days, that URL may be removed from the index.

There are two important consequences.

An emergency change can become an indexing problem

Emergency changes made at 2am have a habit of staying in place longer than planned.

The load drops.

Everyone goes home.

The post-mortem is scheduled for later.

The configuration stays active over the weekend.

By Monday, the affected URLs may have been returning 429 or 503 for several days.

That is no longer just an infrastructure issue.

It can become an indexing issue.

Retry-After does not reset the clock

A Retry-After header is not a promise that Google will keep the URL indexed for as long as you specify.

It tells the client when it should try again.

It does not create extra indexing protection.

So setting:

Retry-After: 1209600

and telling Google to wait two weeks does not mean you have bought two weeks of safety.

It simply tells Google that you expect the resource to remain unavailable for a long time.

That is not a good outcome for a page you care about.

Keep the value short.

This is also an important refinement to the comparison we made previously between 403, 503 and 429.

Broadly speaking:

  • 403 can cause indexed URLs to be removed but does not reduce the overall crawl rate
  • 503 and 429 can reduce crawling while allowing indexed URLs to remain available

But there is a time limit.

After a couple of days of continued errors, the pages returning those responses can start moving towards the same indexing problem you were trying to avoid.

Most crawl spikes are not caused by Google

The first thing Google’s documentation tells you to do is check your access logs and understand where the traffic is coming from.

This is probably the instruction most likely to be skipped.

It is also one of the most important.

A large amount of crawler traffic that looks like “Googlebot” is not necessarily Googlebot.

It can come from:

  • AI crawlers
  • retrieval systems
  • scrapers
  • monitoring tools
  • badly configured bots
  • bots pretending to be Googlebot

Some of these systems deliberately use a Googlebot-style user agent because they know many websites automatically trust it.

This creates a dangerous situation.

You see traffic labelled as Googlebot.

You reduce traffic matching the user agent.

The real Googlebot gets reduction.

The fake Googlebot keeps running because it was never Google in the first place.

You have then created all of the SEO side effects without solving the actual problem.

Verify the crawler first

Before anything, verify that the crawler really belongs to Google.

Use reverse DNS followed by forward DNS, or Google’s published IP ranges.

Do not rely on the user agent string alone.

Then separate the traffic by crawler type:

  • Googlebot
  • AdsBot
  • inspection tools
  • other Google crawlers

This distinction matters.

If the crawl increase is coming from AdsBot because of Dynamic Search Ads, the solution may be in the Ads account rather than on the web server.

Only a temporary fix

All three causes Google lists have something in common.

They can create far more crawlable URLs than the site has useful pages.

Faceted navigation

A site might have a few hundred products but generate combinations such as:

  • colour
  • size
  • price
  • brand
  • sort order

Combine enough of those options and the number of URLs can become enormous.

Date calendars

Events, bookings and travel sites can create a URL for every date.

Without sensible limits, the calendar can keep generating pages into the future and the past.

Dynamic Search Ads

Broad URL rules can direct Google’s systems at large groups of dynamically generated or parameterised URLs.

Returning 429 or 503 can reduce the immediate pressure.

It does not fix the reason the crawl volume increased.

Once the reduction is removed, the same URLs still exist.

The same links still exist.

The same problem comes back.

So the emergency response is the tourniquet.

The real fix is reducing the number of URLs Google is being asked to crawl.

That might mean:

  • better robots.txt rules
  • preventing unwanted filter combinations from becoming crawlable links
  • putting sensible limits on calendars
  • tightening Dynamic Search Ads targets

Google’s crawl guidance says a robots.txt change can take around 24 hours to be picked up.

That is useful because it can fit within the one-to-two-day emergency window.

You can ask Google to crawl less. You cannot ask it to crawl more.

The request form at the bottom of the documentation is another interesting detail.

You can report an unusually high crawl rate and provide an optimal rate.

But you cannot use the same process to request a higher crawl rate.

Search Console used to offer a crawl-rate setting, but that was retired.

That leaves you with a very simple model:

Reducing crawling can be immediate. Increasing it again is slower and largely controlled by Google.

That means any decision to reduce crawling needs to be treated as something with a real recovery cost.

Undoing the consequences can take much longer.

3. Who Is Affected?

Large ecommerce websites with faceted navigation

These are probably the clearest example.

Google specifically identifies faceted navigation as a source of crawl spikes, and reduced crawling can delay price and availability updates.

Advertisers running Dynamic Search Ads

This group is especially interesting because Google’s documentation mentions them twice.

They can contribute to the crawl spike.

They can also be affected by the resulting crawl reduction.

Paid media teams need to know that connection exists.

Events, booking, travel and hospitality

Date-based URLs are a named source of crawl problems.

Hotels, venues, ticketing sites and booking platforms can generate huge numbers of date-based URLs by design.

Publishers with date-based archives

Daily archive pages, date-based navigation and long chains of historical or future pages can create the same type of problem as booking calendars.

Websites on limited hosting infrastructure

Smaller servers have less spare capacity.

They can hit the point where a crawl spike becomes a genuine infrastructure problem much sooner.

They are also less likely to have sophisticated log analysis in place to determine exactly what is generating the traffic.

Businesses in the middle of a migration

Migrations are particularly risky because they can combine:

  • new infrastructure
  • cold caches
  • increased crawling
  • large numbers of URL changes
  • teams working under pressure

That is exactly the sort of environment where an emergency can be introduced quickly and then forgotten.

Websites seeing heavy AI crawler traffic

The load can be very real.

But the traffic may not be coming from Google.

Reducing Google’s crawlers in response to AI crawler traffic is one of the easiest ways to make the problem worse.

Any organisation where engineering can change server responses without notifying marketing

That probably describes most organisations.

Who is less affected?

Small websites that never come close to their server capacity are unlikely to need this.

Fully static websites served almost entirely from edge caching are also less exposed because crawler requests may have little impact on the origin server.

4. What Should Businesses Do?

The response can be broken into three parts:

1. Find out what is actually causing the problem.

2. If you have to reduce crawling, make sure this cannot stay on indefinitely.

3. Fix the underlying problem while the emergency window is still open.

4A. Everyone: agree the process before there is an incident

Do not wait until 2am to decide who should know about a crawl reduction.

At a minimum, the following people should be notified:

  • SEO lead
  • paid media lead
  • engineering or infrastructure owner

This does not mean everyone needs to approve the emergency change.

An incident may require action immediately.

The point is simply to tell the teams whose numbers are about to change.

A one-line message in the right Slack or Teams channel can prevent hours of confusion later.

Use a standard sequence

The order matters.

1. Verify the crawler

Use your logs and verify the crawler.

Do not trust the user agent string.

2. Work out which URLs are being hit

Look for:

  • parameterised URLs
  • faceted navigation
  • internal search pages
  • date URLs
  • Dynamic Search Ads targets

3. Apply carefully

Use 429 where it accurately describes the situation.

Use a short Retry-After.

Scope the response to the URLs causing the immediate problem.

Build in a hard expiry.

4. Fix the underlying cause

Work on the actual source of the crawl increase at the same time:

  • robots.txt
  • internal links
  • faceted navigation
  • calendar limits
  • DSA targeting

5. Remove the status code well before the two-day mark

Do not rely on someone remembering to turn it off.

6. Watch the recovery

Do not assume that removing the reduction means crawl activity instantly returns to normal.

Monitor it for several weeks.

The key point is simple:

“Reducing Google’s crawl rate” is not a free fix.

You are trading an immediate infrastructure problem for potential SEO and advertising problems.

Sometimes that is the right trade.

But it should be a conscious decision.

4B. For SEO, content and marketing teams: know when the brake is on

Paid media: review your Dynamic Search Ads targets now

Do not wait for an incident.

Look at broad URL rules that cover:

  • parameterised pages
  • faceted pages
  • dynamically generated pages
  • large sections of the site

Narrow your targeting to URLs you actually want advertisements generated from.

Using a page feed of specific, canonical URLs gives you much more control than an open-ended URL rule.

Paid media: be included in the incident channel

If an infrastructure change can cause campaigns to pause, paid media needs to know about it immediately.

Do not wait until the weekly performance report to find out why conversions disappeared.

Be careful with promotions during recovery

Google says that when crawling is reduced, prices and product availability can take longer to update in Search.

That means launching a price-led promotion during or immediately after a crawl reduction carries an obvious risk.

You may have changed the price on your website.

Google may still be showing the old price in search results.

That can create a conversion problem and, in some industries, potentially a compliance problem.

Do not publish time-sensitive content expecting instant discovery

Reduced crawling means new pages may take longer to be discovered.

That matters if you are launching:

  • a major product
  • an event
  • a new service
  • time-sensitive content
  • a promotion

Need a page removed urgently?

Google warns that removed pages may stay in the index longer when crawling is reduced.

If a page needs to disappear from Search urgently, the Search Console removal tool can be useful because it does not depend on waiting for Googlebot to crawl the page again.

Review faceted navigation

If every filter combination creates a crawlable URL, you are storing up a technical SEO problem.

This is not just an SEO issue.

It is a product and engineering issue because the way the filters work determines how many URLs Google can discover.

4C. For development teams: diagnose first, reduce safely, and make it expire automatically

Step 1 — Find the source

Before changing anything, split the traffic by verified crawler and URL type.

One of the most useful numbers is the number of distinct URLs being requested.

-- Who is crawling, and what?
-- Run this BEFORE reducing anything.

SELECT
  CASE
    WHEN user_agent LIKE '%AdsBot-Google%'                 THEN 'AdsBot'
    WHEN user_agent LIKE '%Google-InspectionTool%'         THEN 'InspectionTool'
    WHEN user_agent LIKE '%Googlebot%' AND verified_google THEN 'Googlebot (verified)'
    WHEN user_agent LIKE '%Googlebot%'                     THEN 'Claims Googlebot — UNVERIFIED'
    ELSE 'Other'
  END AS crawler,

  CASE
    WHEN path LIKE '/search%'                       THEN 'internal search'
    WHEN path ~ '/\d{4}-\d{2}-\d{2}'               THEN 'date URL'
    WHEN query_string <> ''                         THEN 'parameterised'
    ELSE 'clean path'
  END AS url_class,

  count(*) AS requests,
  count(DISTINCT path || '?' || query_string) AS distinct_urls,
  round(avg(origin_ms)) AS avg_origin_ms

FROM access_logs

WHERE ts >= now() - INTERVAL '24 hours'

GROUP BY 1, 2

ORDER BY requests DESC;

How to read the output

A very large distinct_urls value for parameterised or date-based URLs usually points towards a URL-space problem.

A large UNVERIFIED row means you may be dealing with traffic that is pretending to be Google.

A large AdsBot row points towards Dynamic Search Ads and should trigger a conversation with the paid media team.

The verified_google flag should come from proper crawler verification using reverse and forward DNS or Google’s published IP ranges.

Do not create that flag from the user agent string alone.

Step 2 — Build a reduction that cannot stay on forever

The most important property of an emergency crawler reduction is a hard expiry.

Not:

  • a reminder
  • a ticket
  • a note in Slack
  • someone’s memory

The configuration itself should contain the expiry.

After that time, the reduction should switch itself off.

For example:

// Emergency crawler reduction:
// scoped, self-expiring and fails open.
//
// Requires upstream middleware to set
// req.isVerifiedGoogleCrawler using reverse + forward DNS
// or Google's published IP ranges.

const MAX_REDUCTION_HOURS = 36;
const DEFAULT_RETRY_AFTER = 3600;

const config = loadReductionConfig(process.env);

function loadReductionConfig(env) {
  const enabled = env.CRAWL_REDUCTION === 'on';

  const expiresAt = Date.parse(
    env.CRAWL_REDUCTION_EXPIRES || ''
  );

  const now = Date.now();

  if (!enabled) {
    return { enabled: false };
  }

  // Refuse configurations with no expiry
  // or an expiry that has already passed.

  if (!Number.isFinite(expiresAt) || expiresAt <= now) {
    console.error(
      '[crawl-reduction] missing or past expiry — reduction DISABLED'
    );

    return { enabled: false };
  }

  // Never allow the emergency reduction
  // to run beyond the configured maximum.

  if (
    expiresAt - now >
    MAX_REDUCTION_HOURS * 3600_000
  ) {
    console.error(
      `[crawl-reduction] expiry beyond ${MAX_REDUCTION_HOURS}h — reduction DISABLED`
    );

    return { enabled: false };
  }

  return {
    enabled: true,
    expiresAt,

    retryAfter:
      Number(env.CRAWL_REDUCTION_RETRY_AFTER) ||
      DEFAULT_RETRY_AFTER,

    // Scope the response to the URLs causing the problem.
    //
    // Google can still reduce crawl rate across the hostname,
    // but this limits the URLs exposed to the longer-term
    // indexing risk.

    scope: [
      /^\/search(\/|\?|$)/,
      /[?&](sort|order|colour|color|size|price|brand|page)=/i,
      /^\/(events|calendar)\/\d{4}-\d{2}-\d{2}/
    ]
  };
}

function crawlerReduction(req, res, next) {
  if (!config.enabled) {
    return next();
  }

  const now = Date.now();

  // Expired: fail open.
  if (now >= config.expiresAt) {
    return next();
  }

  // Never trust the user agent string alone.
  if (!req.isVerifiedGoogleCrawler) {
    return next();
  }

  // Only reduce URLs that match the emergency scope.
  if (
    !config.scope.some(
      re => re.test(req.originalUrl)
    )
  ) {
    return next();
  }

  const secondsLeft = Math.ceil(
    (config.expiresAt - now) / 1000
  );

  res.set(
    'Retry-After',
    String(
      Math.min(
        config.retryAfter,
        secondsLeft
      )
    )
  );

  res.set(
    'Cache-Control',
    'no-store'
  );

  res.status(429).send(
    'Too Many Requests'
  );
}

module.exports = {
  crawlerReduction,
  loadReductionConfig
};

There are several details here that matter.

1. It fails open

If the configuration is broken, the expiry is missing or the expiry has passed, the reduction turns itself off.

That is safer than leaving the reduction running indefinitely.

A reduction that fails closed can become an outage.

2. It has a hard limit

The example refuses to run for longer than 36 hours.

That keeps the configuration comfortably inside Google’s one-to-two-day guidance.

3. It does not trust the user agent

A user agent can be spoofed.

The reduction should only act on a crawler that has actually been verified.

4. Retry-After cannot extend beyond the expiry

There is no point telling Google to wait for an hour if the reduction itself is going to disappear in ten minutes.

The retry period should respect the remaining lifetime of the reduction.

5. The error response must not get cached

This is extremely important.

If a CDN caches a 429 response and does not vary it correctly, you can end up serving the crawler-specific error page to normal users.

That would turn a crawl-control measure into a customer-facing outage.

Cache-Control: no-store helps prevent that.

Prefer seconds for Retry-After

Both Retry-After formats are valid under RFC 9110.

In practice, a simple number of seconds is easier to manage.

For example:

Retry-After: 3600

The HTTP-date version needs to use the exact GMT date format required by the standard.

If you do use an absolute time, generate it rather than typing it manually:

// Always generates a valid HTTP date in GMT.

const retryAt =
  new Date(
    Date.now() + 2 * 3600_000
  ).toUTCString();

Step 3 — Monitor the two-day warning per URL

The indexing risk is tied to individual URLs continuing to return errors.

So monitor those URLs directly.

-- URLs where the current unbroken run of
-- reduction/error responses to verified Google crawlers
-- has lasted more than 24 hours.

WITH g AS (
  SELECT
    path ||
      coalesce(
        '?' || nullif(query_string, ''),
        ''
      ) AS url,

    ts,
    status

  FROM access_logs

  WHERE verified_google

    AND ts >= now() - INTERVAL '72 hours'
),

last_ok AS (
  SELECT
    url,
    max(ts) AS last_success

  FROM g

  WHERE status NOT IN (429, 500, 503)

  GROUP BY url
)

SELECT
  g.url,
  min(g.ts) AS streak_started,
  max(g.ts) AS last_error,
  count(*) AS error_responses

FROM g

LEFT JOIN last_ok l USING (url)

WHERE g.status IN (429, 500, 503)

  AND (
    l.last_success IS NULL
    OR g.ts > l.last_success
  )

GROUP BY g.url

HAVING max(g.ts) - min(g.ts)
  > INTERVAL '24 hours'

ORDER BY streak_started;

The important part is last_ok.

A simpler query that only looks at error responses can give you the wrong answer.

Imagine a URL returned an error for six hours, then returned 200, then returned another error for 20 hours.

The total time between the first and last error is 26 hours.

But the URL has never actually been unavailable for 26 hours continuously.

The successful request in the middle broke the streak.

That is why you need to measure the current run of errors since the last successful response.

Anything appearing in this report that you actually want indexed should be back to returning 200 before you approach the multi-day warning.

Step 4 — Fix the cause while the emergency window is open

Use that time to reduce the number of unnecessary URLs being crawled.

robots.txt

Block parameter combinations and internal search URLs that should never be crawled.

Google’s own timing data suggests changes can take around a day to be picked up.

Make unwanted filter combinations less crawlable

For filter combinations you do not want indexed, do not create normal crawlable links unnecessarily.

A button or JavaScript interaction can sometimes make more sense than creating an <a href> for every possible combination.

Put sensible limits on calendars

Do not allow endless chains of dates into the future and past.

Only expose dates that make sense for the business.

Tighten Dynamic Search Ads

Review broad targeting rules with paid media.

The people managing the Ads account need to know that the targeting can affect crawling.

Step 5 — Make the site cheaper to crawl

The best emergency response is not needing the emergency response in the first place.

Make crawler requests as cheap as possible.

Return 304 Not Modified responses to conditional requests when a page has not changed.

Keep important crawlable responses fast.

The more efficiently your infrastructure handles crawling, the less likely you are to need an emergency crawl reduction.

4D. Rolling this out

This is the part I would put into an actual incident runbook.

  1. Write down the response process and name the people who must be notified.
  2. Build the application now, but leave it disabled.
  3. Make sure it has the expiry limit and fail-open behaviour already built in.
  4. Test that your CDN cannot cache crawler-specific error responses and serve them to normal users.
  5. Build the access-log query and make sure crawler verification works.
  6. Set up the 24-hour alert for URLs approaching the indexing risk.
  7. Review Dynamic Search Ads targets with paid media.
  8. Review faceted navigation and calendars for uncontrolled URL generation.
  9. Run a tabletop exercise with SEO, paid media and engineering.
  10. After a real incident, monitor crawl recovery for several weeks and record how long it actually took.
  11. Review the process after the incident rather than waiting until the next one.

4E. Governance

This does not need to become a giant approval process.

It needs a few simple controls.

Notify when the reduction is enabled

Treat “reducing Google’s crawlers” as a change that automatically alerts SEO and paid media.

It does not need to wait for approval.

The important thing is that the teams whose performance is about to change know what has happened.

That one message could prevent a week of unexplained paid media problems.

Put the expiry in code

Do not rely on somebody remembering to turn the reduction off.

Process is useful.

Code is better.

A hard expiry is a technical control that protects you when people are tired, busy or simply forget.

Require a follow-up fix

Every reduction should create a task to fix the cause.

Otherwise the same problem will come back.

Have one owner for crawler verification

The same verification process should be used whenever you need to distinguish a real Google crawler from a fake one.

Keep that logic in one place with one clear owner.

Make paid media part of the conversation

Dynamic Search Ads can influence crawling.

A crawl reduction can influence whether ads serve.

Those two teams therefore need a way to talk to each other.

Record recovery times

After each incident, record how long it took for crawl activity to return to normal.

After several incidents, you will have your own data.

That is far more useful when making internal decisions than simply quoting a generic recovery range.

5. What We’re Watching Next

Will Google explain how long Retry-After values are actually honoured?

The documentation gives examples such as two minutes and a specific date.

It also warns against keeping the reduction active for more than one to two days.

What it does not explain is exactly what happens if somebody sends a much longer value.

Until that is clearer, keep Retry-After values short.

Treat the header as a signal to the crawler, not as a way of extending the life of a page.

Will Ads and Search crawl capacity become more clearly separated?

Google warns that crawl reduction can cause ads to stop serving.

That suggests some shared infrastructure or shared dependency.

As more advertisers become aware of the issue, it would make sense for Google to explain the relationship more clearly.

Will AI crawler traffic make this problem more common?

Probably.

Many websites now see significant crawler traffic from AI systems, scraping tools and retrieval systems.

The traffic itself can create a real infrastructure problem.

But that does not mean Google is responsible for it.

The more crawler traffic websites receive, the more important proper crawler verification becomes.

Do other crawlers consistently respect Retry-After?

Retry-After is a general HTTP mechanism.

It is not just for Google.

But behaviour across Bing, AI crawlers and commercial scraping systems is not always consistent.

A proper comparison of how different crawlers respond would be useful.

Will crawl controls return to Search Console?

The old Search Console crawl-rate control was simple but visible.

The current approach is more technical.

Status codes and server configuration give engineers more control, but they are much easier for the wider business to miss.

A clearly labelled, time-limited control in Search Console would reduce some of the risk described in this article.

6. About Szymaniak Digital

Szymaniak Digital is an enterprise AI SEO consultancy.

A large part of our work sits between technical engineering, paid media and organic search.

Those teams often depend on the same Google systems without sharing the same dashboards, incident processes or communication channels.

That is where problems like this become expensive.

A server decision can affect crawling.

Crawling can affect Search.

Search can affect product visibility.

And in this case, crawling can also affect Google Ads.

If your business does not already have a written process for reducing Google crawling, it is worth creating one before you need it.

At a minimum, that process should define:

  • who verifies the crawler
  • who gets notified
  • which status code is used
  • how long the reduction can remain active
  • how it automatically expires
  • who fixes the underlying URL problem
  • how crawl recovery is monitored afterwards

The first serious crawl incident will force your business to create that process anyway.

It is better to write it before the incident happens.

Contact Us!

Scroll to Top