Fix Googlebot Getting 403 Error

5/5 - (7 votes)

A 403 for Googlebot almost always means something in front of your site (a firewall, CDN, security plugin or server rule) is refusing Google’s crawler while letting normal visitors through. The fix is to find which layer returns the 403, allow verified Googlebot through it, and then ask Google to recrawl.

What Is The Googlebot 403 Error?

HTTP 403 Forbidden means the server understood the request and refused to serve it. When the page loads fine in your browser but Google reports 403, the refusal is conditional: it depends on the user agent, the IP address, the request rate or the country the request comes from.

In Google Search Console it appears as Blocked due to access forbidden (403) in the Page indexing report, or as Page fetch: Failed: Access forbidden (403) in URL Inspection.

Is It Important To Fix The Googlebot 403 Error?

Google does not index content from URLs that return a 4xx status, which makes a Googlebot 403 one of the more damaging technical SEO problems a site can have. The practical effects build up over time:

  • New pages that return 403 are never indexed.
  • Pages already indexed are dropped if they keep returning 403 on repeat crawls.
  • Blocked CSS, JavaScript or image files stop Google from rendering pages correctly, even when the HTML itself returns 200.
  • A 403 on robots.txt is treated as if no robots.txt exists, so the crawl rules in your robots.txt file are ignored.
  • A 403 on the sitemap means Google cannot read it, which slows discovery of new URLs.

Google also advises against using 403 or 401 to slow crawling down. To reduce crawl rate, return 429, 500 or 503 for a short period instead.

Confirm the Problem

Confirm that the real Googlebot is getting the 403 before changing anything. A report from a third-party SEO tool is not proof, because those tools crawl from their own IP addresses and are often blocked for unrelated reasons.

  1. Check the Page indexing report. In Search Console, open Indexing > Pages and look for “Blocked due to access forbidden (403)”. Note whether it affects every URL, one directory or one file type.
  2. Run a live test. Paste an affected URL into URL Inspection and click Test live URL. This sends a real Google request, so a 403 here confirms the block is active right now.
  3. Check Crawl stats. Under Settings > Crawl stats, open the “By response” breakdown. A rising share of 403 responses shows when the block started, which you can match against deployments or settings changes.
  4. Read your server logs. Filter for Googlebot requests that returned 403. The log shows the exact URL, time and source IP.
  5. Check the logs of every layer. If the 403 appears in Search Console but not in your origin server logs, the request never reached your server. The CDN or firewall in front of it returned the 403.

A log search on an Nginx or Apache server looks like this:

grep -i “googlebot” /var/log/nginx/access.log | awk ‘$9 == 403’ | tail -n 50

Check That the Blocked Requests Are From Real Googlebot

Anyone can send “Googlebot” as a user agent, and blocking those impostors is correct behaviour. Verify the IP address from your logs with a reverse and forward DNS lookup, the method Google documents:

host 66.249.66.1

# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.

host crawl-66-249-66-1.googlebot.com

# crawl-66-249-66-1.googlebot.com has address 66.249.66.1

The request is genuine when the hostname ends in googlebot.com, google.com or googleusercontent.com and the forward lookup returns the original IP. If the lookup points anywhere else, the request was fake and the 403 was justified.

Reproduce the Block Yourself

A request with Googlebot’s user agent from your own machine shows whether the block is based on user agent:

curl -I -A “Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)” https://example.com/page

  • 403 with the Googlebot user agent, 200 with a browser user agent: a rule matches the user agent string.
  • 200 in both cases, but Search Console still shows 403: the block is based on IP address, country or request rate, so it only affects Google’s real IP ranges. Only the URL Inspection live test can reproduce it.

Some firewalls block a Googlebot user agent that comes from a non-Google IP as a fake bot. A 403 from the curl test on its own therefore does not prove real Googlebot is blocked.

Common Causes of Googlebot Getting 403 Error

The 403 comes from one of the layers a request passes through on its way to your page. Work from the outside in: CDN first, then firewall, web server, application, and file system last.

LayerTypical causeTelltale sign
CDN or edge (Cloudflare, Akamai, Fastly, CloudFront)Bot protection, managed challenge, custom firewall rule or country block catches Googlebot403 in Search Console, nothing in origin logs; block event visible in the CDN’s security log
Web application firewall (ModSecurity, Sucuri, Imunify360, AWS WAF)A rule set flags crawler behaviour or the user agentEntry in the WAF audit log with a rule ID at the time of the crawl
Hosting providerHost-level bot throttling or IP blocklist you cannot see in your own config403 on all sites on the account; config files look clean
Server rate limiting (fail2ban, mod_evasive, Nginx limit_req)Googlebot’s request bursts trip a threshold and its IP gets banned403s arrive in clusters after crawl spikes; Google IPs in the ban list
Web server config (.htaccess, Nginx)A deny rule matches the user agent or an IP range, often pasted from a “block bad bots” listcurl test with Googlebot user agent returns 403
Security plugin or application (Wordfence, iThemes, All In One Security)Country blocking, rate limiting or a “block fake Google crawlers” option misfiresBlock recorded in the plugin’s live traffic or firewall log
Hotlink protectionRequests without a referrer for images, CSS or JS are refusedPages return 200, only assets return 403
Geo-blockingCountries are blocked; Googlebot crawls mostly from US IP addressesSite works for you locally, fails for Google and for US visitors
Login or IP allowlistStaging rules, password protection or a maintenance mode left on in productionWhole site or directory returns 403 to all outside visitors
File permissions and ownershipFiles or directories unreadable by the web server user403 for everyone, not only Googlebot

The last two rows affect every visitor. They are listed because Search Console is often where the site owner first notices them.

Two patterns narrow the search quickly. A 403 on every URL points to an edge, firewall or hosting block. A 403 on one directory or one file type points to a server rule, hotlink protection or permissions.

How to Fix The Googlebot Getting 403 Error

Apply the fix for the layer your logs point to. One rule applies to all of them: allow Googlebot by verified IP or by the provider’s built-in “verified bot” setting, never by user agent alone. A user-agent allowlist lets any scraper in by claiming to be Googlebot.

1. CDN and Edge Bot Protection

Open the CDN’s security event log, filter for the Googlebot user agent or the 403 action, and note which feature made the decision. The fix depends on that feature.

  • Custom firewall rules: edit the rule so it excludes verified bots. On Cloudflare, add and not cf.client.bot to the rule expression, or create a higher-priority rule with the Skip action for cf.client.bot.
  • Bot protection modes: confirm that verified or “known good” bots are set to Allow. Most CDNs have this as a single toggle in their bot management settings.
  • Country or ASN blocks: remove the United States from any country block, or add an exception for verified bots ahead of it.
  • “Under attack” or challenge-everything modes: switch them off once the attack has passed. Googlebot does not solve CAPTCHAs or JavaScript challenges.

2. Web Application Firewall

Find the rule ID in the WAF audit log for a blocked Googlebot request. Then add an exclusion for that rule scoped to Google’s published crawler IP ranges, and leave the rule active for everyone else.

On managed services such as Sucuri or AWS WAF, check for an IP blocklist entry or a rate-based rule first. These are the most common reasons a verified crawler is refused.

3. Hosting Provider

If your own configuration is clean and the CDN logs show nothing, open a support ticket. Give the host an affected URL, the timestamp and Google IP from URL Inspection or your logs, and ask them to check server-level firewall and bot-throttling rules. Shared hosts sometimes throttle crawlers account-wide during load spikes.

4. Rate Limiting and Automatic Bans

Check whether Google IPs are currently banned, unban them, then exempt Google’s ranges from the limiter.

# fail2ban: list banned IPs, then unban one

fail2ban-client status <jail-name>

fail2ban-client set <jail-name> unbanip 66.249.66.1

Add Google’s crawler ranges to ignoreip in jail.local so the ban does not recur. In Nginx, give Google’s ranges an empty rate-limit key, which exempts them:

geo $is_google {

    default 0;

    66.249.64.0/27 1;   # example only: generate from Google’s common-crawlers.json

}

map $is_google $limit_key {

    0 $binary_remote_addr;

    1 “”;

}

limit_req_zone $limit_key zone=perip:10m rate=10r/s;

Google refreshes its IP range files regularly. Generate the list with a scheduled script instead of pasting it once.

If Googlebot is genuinely overloading the server, return 503 or 429 for a day or two. Google slows down in response to those codes and keeps your pages indexed in the short term. A 403 does not slow crawling.

5. Apache and Nginx Rules

Search your configuration for rules that deny by user agent or IP. These are the usual offenders, often copied from a “block bad bots” snippet with a pattern broad enough to match Googlebot:

# .htaccess: this blocks Googlebot because “bot” matches its user agent

RewriteCond %{HTTP_USER_AGENT} (bot|crawl|spider) [NC]

RewriteRule .* – [F,L]

# Nginx: same problem

if ($http_user_agent ~* (bot|crawler|spider)) {

    return 403;

}

Replace the broad pattern with the specific bot names you want to block, or delete the rule. Also look for Deny from and Require not ip lines (Apache) and deny directives (Nginx) covering ranges that start with 66.249, and remove them.

Remember that .htaccess files apply per directory. A 403 limited to one folder usually means a second .htaccess inside it.

6. Security Plugins and CMS Settings

Open the plugin’s traffic or firewall log, find the blocked Googlebot request and read the reason given. Then adjust the matching setting:

  • Rate limiting: set verified Google crawlers to unlimited. Wordfence has a dedicated option for how to treat Google’s crawlers under its rate limiting rules.
  • Country blocking: remove the United States, or turn the feature off.
  • Blocked IP lists: remove any Google addresses that were added automatically.
  • Fake crawler blocking: if your site sits behind a CDN or proxy, make sure the plugin reads the real visitor IP from the forwarded header. Otherwise it sees the proxy’s IP, fails verification and blocks real Googlebot as fake.

To confirm a plugin is responsible, deactivate it briefly and rerun the URL Inspection live test.

7. Hotlink Protection

Googlebot does not send a referrer when it fetches images, CSS and JavaScript. Change the rule to allow empty referrers:

RewriteCond %{HTTP_REFERER} !^$

RewriteCond %{HTTP_REFERER} !^https?://(www\.)?example\.com [NC]

RewriteRule \.(jpe?g|png|gif|webp)$ – [F,NC]

The first line is the one that matters. Without it, every request with no referrer is refused. Never apply hotlink protection to CSS or JavaScript files.

8. Access Restrictions Left On

Check for password protection, IP allowlists and maintenance or “coming soon” modes carried over from staging. If the content is meant to be public, remove the restriction. If it is meant to be private, the 403 is correct and you can ignore the report or remove those URLs from your sitemap.

9. File Permissions and Ownership

If everyone gets the 403, reset permissions to the standard values: 755 for directories and 644 for files, owned by the user your web server runs as.

find /var/www/example.com -type d -exec chmod 755 {} \;

find /var/www/example.com -type f -exec chmod 644 {} \;

A directory with no index file and directory listing disabled also returns 403. Add an index file or fix the link that points to the bare directory.

Validate the Fix and Request a Recrawl

The fix is confirmed only when Google itself fetches the page with a 200. Follow these steps in order:

  1. Rerun the live test. In URL Inspection, click Test live URL for several affected pages. Each should show “Page fetch: Successful”.
  2. Check the rendered page. Click View tested page and open the screenshot and the page resources list. Any CSS, JavaScript or image still returning 403 will be listed there.
  3. Test robots.txt and the sitemap. Inspect /robots.txt and your sitemap URL the same way. Both must return 200. If the file still fails, work through the checks for when Google is unable to crawl robots.txt.
  4. Request indexing for your most important pages from URL Inspection. There is a daily quota, so prioritise.
  5. Start validation. In Indexing > Pages, open the “Blocked due to access forbidden (403)” issue and click Validate fix. Google then rechecks the affected URLs and reports progress.
  6. Resubmit the sitemap under Indexing > Sitemaps to prompt rediscovery of the rest.
  7. Watch the logs. Over the next few days, Googlebot requests in your server logs and in Crawl stats should return 200.

Recovery is gradual. Validation in Search Console typically runs for days to a couple of weeks, and pages that were dropped from the index return as they are recrawled. Rankings for pages that were out of the index for a long time can take longer to settle. To shorten the wait, follow these steps to make Google index your site faster.

Prevent It From Happening Again

Most Googlebot 403s are introduced by a security change that nobody tested against crawlers. These habits catch the problem before it costs indexed pages:

  • Run a URL Inspection live test after every firewall, CDN, plugin or hosting change.
  • Allow crawlers through the provider’s verified-bot setting or Google’s published IP ranges, never by user agent.
  • Refresh any stored copy of Google’s IP ranges on a schedule.
  • Set an alert on 403 responses to verified Googlebot in your logs or CDN analytics.
  • Review Crawl stats in Search Console monthly for changes in the “By response” breakdown.
  • Keep staging restrictions out of production deployments, with a check in the release process.
  • Read “block bad bots” snippets line by line before adding them, and test with the curl command above.
  • Use 503 or 429, not 403, when you need Google to slow down temporarily.

Frequently Asked Questions

The page loads fine in my browser. Why does Google see a 403? The block is conditional. Your browser, IP address and country pass the rule that Googlebot’s user agent, IP range or request rate fails.

Is this a robots.txt problem? No. A robots.txt block appears in Search Console as “Blocked by robots.txt” and Google never requests the page. A 403 means Google requested the page and the server refused.

Will a 403 remove my pages from Google? Yes, if it persists. Google treats 403 like other 4xx errors and removes indexed URLs that keep returning it. A block lasting a few hours rarely causes lasting damage.

Only some pages show 403. What does that mean? The block is tied to a path, a file type or a rate limit. Compare blocked and working URLs for a shared directory or pattern, and check whether the 403s cluster in time, which indicates rate limiting.

Other Google crawlers are blocked too. Is the fix the same? The method is the same. AdsBot, Google-InspectionTool and user-triggered fetchers use separate IP range files, all linked from Google’s verification page, so include those ranges if you depend on Google Ads or other Google products.

Should I allow everything that says it is Googlebot? No. Fake Googlebot traffic is common. Allow only requests that pass DNS verification or come from Google’s published ranges.

Also See:

Fix Sitemap Submitted But Not Indexed ErrorGoogle Showing Old Titles? Fix It
Why Your Website Ranks On Google Not in AIFix Google Indexing Issues
Google Rankings Dropped Overnight?Common SEO Mistakes and Fixes

Add Comment