Search Console says “Blocked due to access forbidden (403)”, yet the page opens perfectly in your browser. That mismatch is the whole problem in one sentence. Your server, or something in front of it, is treating Googlebot differently from you, and is telling it “you are not allowed here”.
A 403 is not Google’s decision. It is your site’s response. Google simply records it, can’t read the page, and leaves it out of the index. The fix is to find which piece of your stack is sending that 403 to Googlebot and teach it to let real Google crawlers through, without opening the door to fake ones.
- What the 403 Status Actually Means
- First, Decide Whether the Block Is Intentional
- Where the 403 Is Coming From
- Steps To Fix Indexing Failed Due To Access Forbidden Issue in Search Console
- Step 1: Reproduce the Error
- Step 2: Find the 403 in Your Logs
- Step 3: Fix the Layer That Is Blocking
- Step 4: Ask Google to Recrawl
- Block Bad Bots Without Blocking Google
- FAQs
What the 403 Status Actually Means
Table 1: 403 Compared With Similar Indexing Errors
| Status | What Your Server Is Saying | Typical Cause |
|---|---|---|
| 403 Forbidden | “I know who you are, and you can’t come in.” | Firewall, bot protection, IP or user-agent block, file permissions |
| 401 Unauthorized | “You need to log in.” | Password-protected or staging pages |
| 404 Not Found | “This page doesn’t exist.” | Deleted or mistyped URLs |
| 429 Too Many Requests | “Slow down.” | Rate limiting |
| 5xx Server Error | “Something broke on my side.” | Overloaded or misconfigured server |
| Blocked by robots.txt | “Please don’t crawl this.” | A disallow rule, not an HTTP error |
If you want Google to crawl more slowly, don’t use 403. Google’s crawling guidance recommends 429 or 503 for temporary rate limiting. A 403 tells Google the page is off limits, and pages that keep returning it drop out of the index.
First, Decide Whether the Block Is Intentional
Some 403s are correct. Admin areas, customer accounts, members-only content and staging copies should be forbidden. If the URLs in the report are private, the fix is not to unblock them but to stop advertising them:
- Remove them from your XML sitemap.
- Remove internal links to them from public pages.
- Let the report entries stay. A 403 on a private URL is not hurting your public pages.
If the report lists blog posts, product pages, category pages or your homepage, the block is a mistake and needs fixing.
Where the 403 Is Coming From
A common trap is checking only the first layer. In one Cloudflare community case, a site owner switched off every bot setting and added a skip rule for verified bots, yet Googlebot still got a 403. Cloudflare’s logs showed nothing, because the block came from the web host’s own user-agent filtering further down the chain.
Steps To Fix Indexing Failed Due To Access Forbidden Issue in Search Console
Step 1: Reproduce the Error
Start with tests that tell you what kind of block you are dealing with.
Table 2: Three Tests and What They Reveal
| Test | How | If It Fails, the Block Is Probably |
|---|---|---|
| Search Console live test | URL Inspection, then Test Live URL | Anywhere in the stack, because this uses a real Google IP |
| Googlebot user agent from your computer | curl with Googlebot’s user agent (below) | A user-agent rule |
| Normal browser from another country or a VPN | Open the page from a different network | A country or IP reputation block |
# Request the page as a normal browser
curl -sI https://example.com/page/
# Request the page with Googlebot Smartphone's user agent
curl -sI -A "Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" https://example.com/page/
If the first returns 200 and the second returns 403, something is filtering on the user agent. If both return 200 but the live test in Search Console fails, something is filtering on Google’s IP addresses or on request patterns.
Watch for this pattern: the live test passes but normal crawling keeps failing. That usually means a rate limit or a reputation-based rule that only triggers when Googlebot crawls several pages in a row.
Step 2: Find the 403 in Your Logs
Your server access log shows exactly which requests got a 403. Filter for Googlebot and status 403:
# Apache or Nginx access log
grep "Googlebot" /var/log/nginx/access.log | grep '" 403 ' | tail -20
Then confirm the requests are really from Google, because plenty of scrapers fake the Googlebot user agent. Google’s verification guide describes a reverse and forward DNS check:
host 66.249.66.1
# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
# crawl-66-249-66-1.googlebot.com has address 66.249.66.1
A real Google crawler’s hostname ends in googlebot.com, google.com or googleusercontent.com, and the forward lookup returns the same IP. For bulk checks, Google publishes its crawler IP ranges as JSON files, including common-crawlers.json for Googlebot.
Step 3: Fix the Layer That Is Blocking
CDN and WAF (Cloudflare, Sucuri and Similar)
- Check security or firewall events for blocked requests from Google IPs or with Googlebot user agents.
- Make sure verified bots are allowed. In Cloudflare, a custom rule that skips security for
cf.client.botallows known good bots, including Googlebot. - Review country blocks. Googlebot crawls mostly from US IP addresses, so blocking the US, or allowing only your own country, will block Google.
- Loosen rate limits for verified crawlers, or apply them only to unverified traffic.
- Avoid challenge pages (JavaScript or CAPTCHA challenges) for search engine bots.
Host Firewall and Server Security
- Ask your host whether Imunify360, ModSecurity or a similar tool is blocking Googlebot by user agent or IP reputation. Many shared hosts filter bots without telling you.
- Ask them to allowlist Google’s published crawler IP ranges, not just the user agent string.
- If your host can’t or won’t fix it, that’s a good reason to move. Our list of SEO-friendly hosting services covers hosts that handle crawlers properly.
Web Server Rules (.htaccess, Nginx, IIS)
Look for deny rules aimed at bots, user agents or IP ranges. Rules copied from “block bad bots” lists often include patterns broad enough to catch Googlebot.
# Risky: blocks any user agent containing "bot", including Googlebot
RewriteCond %{HTTP_USER_AGENT} bot [NC]
RewriteRule .* - [F,L]
- Remove or narrow any rule like the one above.
- Check hotlink protection rules, which can return 403 for images and CSS that Google needs to render pages.
- Check file and folder permissions. Files usually need 644 and folders 755 on Linux servers.
CMS Security Plugins
- In Wordfence, Solid Security or similar plugins, review blocked requests, country blocking and “fake Google crawler” settings.
- Make sure fake-crawler checks verify by reverse DNS rather than blocking everything with a bot user agent.
- Temporarily disable the plugin and run the live test again. If the error disappears, you’ve found the source.
Login Walls and Staging Protection
- Check that public pages are not behind HTTP authentication left over from a staging site.
- Make sure maintenance mode or a “coming soon” plugin is switched off.
- For paywalled content you want indexed, use Google’s structured data for paywalled content rather than a hard 403.
Step 4: Ask Google to Recrawl
- Run URL Inspection’s live test until it returns 200 and the page renders.
- Request indexing for your most important URLs. The feature has a daily quota.
- In the Page indexing report, open the 403 issue and click Validate fix.
- Check the Crawl stats report (Settings, then Crawl stats) over the next few days to confirm 403 responses are falling.
For other statuses in the same report, our guide on fixing Google indexing issues walks through each one.
Block Bad Bots Without Blocking Google
Most 403s for Googlebot are side effects of good intentions. Bot traffic is a real cost, as our bot traffic statistics show, and blocking scrapers is sensible. The safe way to do it:
Table 3: Safe Bot-Blocking Practices
| Do | Don’t |
|---|---|
| Verify crawlers by reverse DNS or Google’s IP ranges | Block or allow based only on the user agent string |
| Allow verified search engine bots before other rules run | Put broad bot rules above your allowlist |
| Use 429 or 503 to slow crawlers down temporarily | Use 403 to manage crawl load |
| Block specific bad IPs and abusive patterns | Block whole countries Google crawls from |
| Test with URL Inspection after every security change | Assume a firewall update didn’t affect Googlebot |
FAQs
Why Does My Page Load in the Browser but Show 403 for Google?
Something on your server treats Googlebot differently: a firewall, CDN rule, security plugin or user-agent filter. Your browser has a different user agent, IP and location, so it gets through.
Will a 403 Error Remove My Page From Google?
Yes, if it continues. Google can’t read a page that returns 403, so it won’t index it, and already indexed pages are dropped over time if the error persists.
Does Cloudflare Block Googlebot?
Not by default for verified Googlebot traffic, but strict bot settings, country blocks, rate limits or custom rules can. Your host can also block Google behind Cloudflare.
How Long Does It Take for Google to Reindex After Fixing a 403?
Important pages can be recrawled within days after requesting indexing. Large sites may take a few weeks for every affected URL to be recrawled and validated.
Should I Allow Every Bot That Says It Is Googlebot?
No. Many scrapers fake the Googlebot name. Verify by reverse DNS or Google’s published IP ranges and allow only real Google crawlers.