HTTP Status Codes and SEO: 4xx, 5xx, Soft 404
Every time Googlebot requests a URL, the HTTP status code decides what happens next, before anyone looks at the content. According to Google's crawling documentation, content from URLs that return a 4xx code (401, 403, 404, 410 and the rest) is not used, and indexed URLs that start returning one are removed from the index. 5xx errors and 429 make Google slow down crawling, and URLs that keep failing eventually drop out too. A soft 404 is the odd one out: a 200 response that looks like an error page, which Google excludes from search even though the server reports success.
How Google Handles Each Class of Status Code
A status code is the three-digit number at the start of every server response. Browsers rarely show it unless something breaks, but a crawler reads it on every single fetch, and it is the first thing that shapes how a page is crawled and whether it can be indexed at all.
Google documents its behavior in a reference page titled How HTTP status codes affect Google's crawlers, which now lives in Google's crawling infrastructure documentation. The logic follows the four classes:
- 2xx (success). Google passes the content to the next processing step, which for Search means the indexing pipeline. The documentation is careful to add that a 2xx response doesn't guarantee indexing.
- 3xx (redirection). Google's crawlers follow up to 10 redirect hops by default and process the content of the final target. A 301 or 308 is a strong signal that the target should be processed, a 302 or 307 a weak one.
- 4xx (client errors). Google doesn't use the content. New URLs with a 4xx code are not indexed, and URLs already in the index are removed. Every 4xx code except 429 is treated the same way.
- 5xx (server errors). Google temporarily slows down crawling. Indexed URLs are preserved for a while, but eventually dropped if the errors persist.
Search Console generates error messages for the whole 4xx–5xx range and for failed redirects, which is why these codes end up as long lists in your reports, not just as an annoying screen for the occasional visitor.
A Quick Reference for the Codes That Matter
Most sites only ever need to reason about a dozen codes. The table below summarizes what each one means and how Google's crawlers respond to it, based on the same documentation, the HTTP specification (RFC 9110) and RFC 6585 for 429.
| Code | What it means | What Google does |
|---|---|---|
| 200 OK | The request succeeded | Passes the content to indexing; indexing is not guaranteed |
| 301 / 308 | Moved permanently | Follows the redirect; strong signal to process the target |
| 302 / 307 | Temporary redirect | Follows the redirect; weak signal to process the target |
| 304 Not Modified | Content unchanged since the last crawl | Reuses the previous version; no effect on indexing |
| 400 Bad Request | The server won't process a malformed request | Treated like any 4xx: content considered nonexistent |
| 401 Unauthorized | Valid authentication credentials are missing | Treated like 404; content not used |
| 403 Forbidden | The server understood the request but refuses it | Treated like 404; content not used |
| 404 Not Found | No current representation of the resource | Not indexed; indexed URLs removed; crawled less often over time |
| 410 Gone | The resource is gone, likely permanently | Same as 404 |
| 429 Too Many Requests | Rate limiting | Treated as server overload, i.e. a server error; crawling slows down |
| 500 Internal Server Error | The server failed | Crawl rate drops in proportion to failing URLs; persistent errors removed |
| 503 Service Unavailable | Temporary overload or maintenance | Crawling slows down; prolonged 503 leads to removal |
The two rows that surprise people most are 429, the only 4xx code Google treats as a server problem rather than a missing page, and 401 and 403, which look like access issues to a human but are, for Google, simply pages that don't exist.
401 and 403: Gated and Blocked Content
A 401 Unauthorized response means the request lacks valid authentication credentials, and the server must say how to authenticate in a WWW-Authenticate header. That is the definition in RFC 9110. A 403 Forbidden means the server understood the request and refuses to fulfill it, for reasons that may have nothing to do with credentials. The specification even allows a server that wants to hide a resource's existence to answer with 404 instead.
For Google the distinction disappears. Both are ordinary 4xx codes, so a page behind a login is, from the crawler's point of view, a page that isn't there. Google's JavaScript SEO basics name 401 as the meaningful status code for pages behind a login. If you want content in search, it can't sit behind either code.
The crawling documentation adds a warning worth repeating: don't use 401 or 403 to limit the crawl rate. Codes in the 4xx range other than 429 have no effect on how fast Google crawls, so instead of easing the load on your server they quietly remove pages from the index.
The robots.txt file is a special case. If it returns 401 or 403, Google's robots.txt specification says the crawlers behave as if no valid robots.txt existed and assume there are no crawl restrictions at all. Our robots.txt glossary entry covers how the file itself works.
In our audits we regularly find the reverse of what the site owner intended: the site works fine in a browser, while Googlebot gets a 403 because a web application firewall or a bot protection service decided it was an intruder. The owner always sees a 200, so the problem only surfaces in the reports. That is why we check status codes with the URL Inspection tool in Search Console, which, according to Google's documentation, shows the code returned to the crawler and the rendered page, rather than relying on a browser.
404 vs. 410: When a Page Is Gone
A 404 Not Found means the server found no current representation of the resource, or isn't willing to disclose that one exists. Per RFC 9110, a 404 says nothing about whether the condition is temporary or permanent, and a 410 Gone is preferred when the server knows the resource is gone for good. The usual causes are a mistyped link, a deleted page, or a page that moved without a redirect.
In Google's documentation, 404 and 410 behave like every other 4xx code except 429, so for Google Search the choice between them makes no practical difference, even though RFC 9110 still prefers 410 when you know a page is gone for good. Newly discovered 404 pages aren't processed, and crawling of those URLs gradually slows. Google's guide to crawl budget recommends returning 404 or 410 for permanently removed pages and explains that while Google never forgets a URL it knows about, a 404 is a strong signal not to crawl it again. A URL blocked in robots.txt, by contrast, stays in the crawl queue much longer. On sites where crawl budget is a real constraint, that difference matters.
Fixing a 404 starts with asking whether it is an error at all. If the page moved or has a clear replacement, a 301 redirect sends both people and crawlers to the new URL. If it was removed with no replacement, a 404 or 410 is the correct answer and there is nothing to fix. Only a 404 on a page that should exist is a genuine bug.
A custom 404 page should help visitors: say clearly that the page wasn't found, keep the site's look and navigation, and point to popular content and the homepage. Google stresses that such a page is for users only, so the server must still return a 404 status code with it, not a 200.
Soft 404s: Pages That Pretend to Exist
A soft 404 is the reverse problem. The server returns 200 OK, yet the page tells the visitor there is nothing here. According to Google's guide to troubleshooting crawling errors, it can also be a page with no main content or an empty page. Google lists typical causes: a missing server-side include file, a broken database connection, an empty internal search results page, or a JavaScript file that failed to load.
Such pages are excluded from search. When Google's algorithms detect from the content that a page is really an error page, Search Console flags it as a soft 404 in the Page Indexing report. The crawl budget guide adds a second cost: soft 404 pages keep getting crawled and waste budget that could go to pages you care about.
The fix depends on what is really going on. If the content no longer exists, return 404 or 410. If it moved, use a 301. If a perfectly good page was flagged, it most likely didn't load properly for Googlebot: critical resources were missing, or it showed a prominent error message during rendering. Resources blocked by robots.txt, too many resources on the page, server errors, and slow or very large resources are the usual suspects.
Single-page apps have their own version of the problem, since with client-side routing, meaningful status codes can be impossible or impractical. Google suggests either a JavaScript redirect to a URL where the server responds with 404, or adding a robots meta tag with noindex to the error page through JavaScript.
In our audits, Search Console most often flags four kinds of pages as soft 404: empty category and listing pages, internal search results with no matches, pages for discontinued products or services that still return 200 with an "unavailable" note, and client-rendered apps that display an error message without changing the status code.
5xx Errors, 429 and Planned Downtime
A 5xx code says the server failed, not the request, and Google responds by easing off rather than concluding the content is gone. Indexed URLs are kept for a while, but not indefinitely. With a 500, the slowdown is proportional to the number of URLs returning the error, and URLs that persistently fail are removed from the index. Once the server returns 2xx again, Google gradually raises the crawl rate.
The effect isn't limited to the failing URLs. Google's page on how to reduce the crawl rate explains that when its crawlers hit a significant number of 500, 503 or 429 responses, they slow down across the whole hostname, including URLs that return content normally. An outage in one section of a site can slow crawling of all of it.
That same mechanism is Google's recommended emergency brake. If Googlebot is overwhelming your server, the troubleshooting guide says to return 503 or 429 temporarily; Googlebot retries those URLs for about two days. Keep it up for more than two days and Google drops those URLs from the index, so rate-limiting rules in a firewall that also catch Googlebot deserve a careful review.
For planned maintenance, 503 Service Unavailable is the right tool. Google's guidance on how to pause or disable a website recommends serving an informational page with a 503 instead of the content for an outage of a day or two, and keeping an indexable homepage with a 200 as a placeholder for longer closures. It also recommends a Retry-After header with a best-effort date or duration, and warns that Google can't refresh titles, descriptions or structured data while a page returns 503.
The advice isn't new. A Google blog post from January 25, 2011, on handling planned site downtime, already argued that 503 beats both a 404 and an error page served with 200, and that a long-lasting 503 can be read as a sign the server is permanently unavailable, leading to URLs being removed. Google now labels that post as older and points to its current guidance, but both principles survive in today's documentation.
One file must never return 503: robots.txt. The pause-a-website guide says plainly that a 503 on robots.txt blocks all crawling. The robots.txt specification spells out the timeline for server errors: for the first 12 hours Google stops crawling the site while retrying the file, for the next 30 days it uses the last good version, and after that it either behaves as if there were no robots.txt, if the site is otherwise available, or stops crawling if the whole site has availability problems.
Where Status Code Problems Show Up in Search Console
The Page Indexing report is the first stop: it lists URLs that aren't indexed along with the reason, and it is where soft 404s are flagged. Our Google Search Console guide walks through that report and its statuses.
The Crawl Stats report is the second. Google's documentation describes it as the history of Googlebot's crawling of your site, including the moments it ran into host availability issues. In the host availability graphs you can see where crawl requests crossed the red limit line, click through to the URLs that were failing, and match them with what was happening on your server at the time.
The third tool is URL Inspection, which shows the returned status code and the rendered page for a single URL. What Search Console doesn't offer is a crawl history you can filter by URL or path. To know whether Googlebot fetched specific pages and what it got back, you need your server logs.
What We Look for in an Audit
The raw number of errors in a report tells you very little. In an SEO audit, most of the work is separating real problems from the normal state of a living website, and that is where status codes repay attention.
Status codes on the pages that matter. We start with the pages the business depends on (homepage, service and offer pages) and compare what URL Inspection reports with what a browser sees. This is where firewall and bot protection blocks come to light.
Linked and unlinked 404s. We split 404s into two groups. URLs that internal links or the sitemap still point to are errors to fix. Old URLs Google remembers but nothing links to anymore are usually fine. We don't bulk-redirect every deleted page to the homepage, because a visitor then lands somewhere they never meant to go. Dead URLs typically pile up during redesigns and site moves.
Soft 404s, one by one. For each flagged URL we check whether the page is really empty, has moved, or is a good page that failed to render for the crawler.
5xx and host availability. We read server errors alongside the Crawl Stats report and the server logs, looking for recurring windows such as backup jobs or deployments, and for firewall rules that answer the crawler with 429 or 503 for longer than they should.
Leftovers from a migration. After a redesign we check that old URLs 301 to their equivalents instead of returning 404 or a homepage with a 200.
These checks sit alongside the rest of our technical SEO checklist, and status codes are one of the core areas of technical SEO in general.
Summary
The HTTP status code decides whether Google uses a page's content at all. Every 4xx code, including 401, 403, 404 and 410, means no content and, for indexed URLs, removal from the index. 5xx and 429 slow crawling down, across the whole host when there are enough of them, and prolonged errors end in removal as well. A soft 404 returns 200 but looks like an error page, and Google excludes it regardless of the status code.
Not every error in a report needs fixing. A 404 for a page removed without a replacement is correct, and a 503 during a short planned outage is good practice. What needs fixing are errors on pages that should work, links that lead nowhere, and blocks the site owner doesn't know about.
If you want to know which status codes on your site are the normal state and which are costing you a place in the index, get a free quote for an audit.