Last month a client sent me a panicked email: "Google says it can't crawl half my site." I opened Search Console, looked at the coverage report, and laughed. Not because it was funny—because I'd made the exact same mistake two years earlier on my own blog.
Here's what was happening: their sitemap pointed to a version of the site that no longer existed. Googlebot kept knocking on a door that had been bricked over. Every crawl attempt failed. Nothing got indexed. Traffic fell off a cliff over about six weeks, and they didn't notice until a sale campaign flopped.
Crawl errors are boring right up until they cost you money. Then they're the only thing you care about. Let me walk you through how to actually fix them, and—more importantly—how to tell a real crawl problem apart from noise.
Key takeaways
- Crawl errors split into two families: pages Google can't reach and pages Google refuses to index. Different fixes.
- Use the URL Inspection tool and the crawl stats report together—one shows you a single URL, the other shows the pattern.
- Fixing a crawl error is often a five-minute job in robots.txt or a canonical tag. The hard part is finding it.
- Crawl budget matters on large sites, not on a 40-page brochure site. Stop worrying about it if you have 40 pages.
- If you want to force a recrawl, the URL Inspection tool does it—but on a per-URL basis, and only for pages you own in Search Console.
What a crawl error actually means (and why most advice gets it wrong)
Googlebot is a program. It requests a URL, waits for a response, reads the HTML, and follows links it finds. A crawl error is anything that stops that loop from completing cleanly.
That's it. No mysticism.
The advice you'll find everywhere lumps everything into one bucket: "fix your crawl errors." Unhelpful, because the causes are completely different depending on what stage broke.
The three stages where things break
- Reachability. Googlebot can't get a 200 response. Server down, DNS misconfigured, a firewall blocking Google's IP ranges, robots.txt disallowing the path.
- Parseability. The page loads, but the content is JavaScript-rendered and blocked, or the HTML is malformed enough that links don't resolve.
- Indexability. Googlebot crawls fine, then decides not to index—noindex tag, canonical pointing elsewhere, thin content, duplicate URL parameters.
Stage one is a technical emergency. Stage three is usually a deliberate choice someone made and forgot about. I've seen both in the same week.
What is crawl in SEO?
Crawling is the discovery step. Before Google can rank anything, it has to find the page, fetch it, and store a copy. Indexing is what happens after. Ranking happens after that. If crawling fails, nothing downstream matters—you can have the best content on the internet and it will never appear in a result page.
Think of it as a library. Crawling is the librarian walking the shelves. Indexing is the card catalog. If the librarian can't find the shelf, the book might as well not exist.
How to fix crawl errors: a diagnostic sequence that works
Random troubleshooting wastes time. Work in this order and you'll find the cause in twenty minutes instead of two days.
Step 1: confirm it's real, not a reporting lag
Search Console reports lag. Sometimes by days. I once spent an afternoon chasing a 5xx error that had already been fixed—the report was just stale. Check the live page first. If it returns 200 in your browser and in a fetch tool, the error may already be resolved.
Step 2: isolate one URL, not the whole pattern
Open the URL Inspection tool on a single failing page. It tells you:
- Whether the URL is indexed
- What Googlebot saw last time it crawled
- Whether the page is blocked by robots.txt
- The canonical Google chose (which is often not the one you set)
That last point catches more people than any other. You set a canonical, Google ignores it, and now two versions of the same page compete. Not technically a crawl error, but it produces the same symptom: pages you expected to rank that don't.
Step 3: read the crawl stats report like a data source
The crawl stats report shows how many requests Googlebot made to your site over time, and what happened to them. Two things to look for:
A sudden drop in total requests. That usually means a server-side problem—slow response times, a CDN change, a firewall rule. Googlebot backs off when your site is slow. Fix the speed, and the crawl rate recovers on its own over a few weeks.
A surge in a specific response code. A spike in 404s means something changed in your URL structure. A spike in 5xx means the server is failing intermittently. The report breaks responses down by type—that's your signal.
The common causes, ranked by how often I actually see them
| Cause | How it looks in Search Console | Fix |
|---|---|---|
| Robots.txt blocking a needed path | "Blocked by robots.txt" | Edit the rule. Test it before saving. |
| Broken internal links (404s) | Spike in "Not found (404)" | Update the links or 301 the old URLs to the new ones. |
| Server errors (5xx) | "Server error (5xx)" cluster | Check hosting logs, database timeouts, plugin conflicts. |
| Redirect chains | "Page with redirect" and long chains | Point redirects directly to the final URL. Chains hurt. |
| Accidental noindex | "Excluded by noindex tag" | Remove the tag. Often left in from a staging build. |
| Faceted navigation explosion | Thousands of near-identical URLs | Block or canonicalize the filter parameters. |
Notice the redirect chain row. It's the one people skip. A chain of three hops doesn't technically break crawling, but it wastes crawl budget and dilutes link signals. I've fixed chains on a site and seen indexed pages jump by roughly 15% within a month—no content change, just cleaner redirects.
Crawl budget: when it matters and when it's a distraction
Here's the thing. If your site has 50 pages, crawl budget is irrelevant. Google will crawl all of them. Stop reading about it.
Crawl budget becomes real somewhere above a few thousand URLs, or when a large portion of your pages are low-value—filter combinations, session-parameter URLs, endless pagination. That's when Google starts rationing time on your site and the pages you care about get skipped.
How do I force a Google crawl?
You can request a recrawl of a specific URL through the URL Inspection tool. Paste the URL, run the live test, and click the request indexing button. Google will usually fetch it within a day or two.
What you can't do is force a site-wide recrawl. There's no button for that. If you need Google to revisit hundreds of pages quickly, the honest answer is: publish fresh internal links pointing at them, resubmit your sitemap, and wait. Fetch requests are rate-limited, and Google decides how much it trusts the urgency.
Crawl rate versus crawl budget—don't confuse them
Crawl rate is how fast Googlebot hits your server. Crawl budget is how many URLs it will bother fetching in a given window. A rate drop shows up as fewer total requests in the crawl stats report; a budget problem shows up as Google skipping pages you actually want crawled while wasting effort on junk URLs.
The fix for a rate problem is usually technical: faster server responses, less JavaScript blocking, better CDN caching. The fix for a budget problem is structural: kill the junk URLs.
One thing nobody warned me about: non-Google crawlers
Most crawl-error tutorials stop at Google. That's a mistake now. AI crawlers—GPTBot, PerplexityBot, Bingbot, and others—hit sites constantly, and they have their own rules and their own robots.txt user-agent strings. If you wrote a robots.txt disallow rule for one bot, it might not apply to the others.
A tight robots.txt can quietly block an AI crawler while leaving Google untouched, and you won't see it in Search Console at all. Check your server logs for user-agent strings you don't recognize. On a client site last year, I found three separate AI crawlers hammering a faceted URL structure and generating thousands of requests a day. The fix was a targeted robots.txt rule for those user-agents—Google's crawling was unaffected.
If AI visibility matters to you—and by 2026 it increasingly does—crawler access for those bots is now part of the job.
The honest priority order
You can't fix everything at once. Here's the order I'd actually work in:
- Server errors (5xx). These kill everything else.
- Robots.txt blocks on important paths. One wrong character can deindex a section.
- Broken internal links and redirect chains. Cheap to fix, immediate payoff.
- Accidental noindex tags. Usually a staging-to-production leak.
- Faceted URL bloat. Only worth it if you're above a few thousand URLs.
- Non-Google crawler access. Depends on whether you care about AI visibility.
The pattern: fix what breaks reachability first, then what wastes crawl effort, then what's cosmetic.
What I keep coming back to is how unglamorous this work is. There's no clever strategy in fixing a broken canonical. It's plumbing. But plumbing is what makes the rest of the house function. I've watched technically immaculate sites with mediocre content outrank brilliant content stuck behind a crawl error, and I don't think that's a bug in the system—I think it's a reminder that search engines can only judge what they can reach.
Fix the plumbing. Then write the content.