Wayback Machine for SEO: Old URLs and Domains
The Wayback Machine is the Internet Archive's public collection of archived web pages, available at web.archive.org. For SEO, it does three jobs well: it recovers the old URLs you need for a redirect map when the old site is gone, it shows what a page used to say before a redesign, and it reveals what a domain hosted before you buy it. The archive is incomplete and it is not Google's view of your site, but it is often the only record left of a site that no longer exists.
What the Internet Archive and the Wayback Machine Are
The Internet Archive set out to preserve digital artifacts and build an internet library for researchers, historians and scholars. According to the Wayback Machine general information page in the archive's help center, it began archiving the web in 1996. The name is a nod to Mr. Peabody's WABAC (pronounced "way-back") machine from the Rocky and Bullwinkle cartoon show.
Using it is simple: you type in a URL, pick a date on the calendar and browse the copy captured that day. The dots on the calendar mark captures, and when you hover over one, the archive shows which crawl it came from. The archive gathers pages through many separate crawls, and it notes that behind each one is a story of who ran it, why, when and how.
Every capture has its own address: web.archive.org/web/, followed by a 14-digit timestamp (year, month, day, hour, minute, second) and the original URL, for example https://web.archive.org/web/20130919044612/http://example.com/. That address is a handy piece of evidence in an audit report, because it pins down exactly what a page looked like on a given day.
Why Old URLs Matter in a Site Move
When URLs change, whether you're moving to a new domain, restructuring paths or switching platforms, Google's documentation on site moves with URL changes tells you to prepare a URL mapping: each old URL gets a new destination, and the server redirects one to the other. For a simple domain change, a wildcard server-side redirect may be enough. For anything more complex, you need an actual list of old URLs, and Google suggests where to get it: your sitemaps, your server logs or analytics, the Links report in Search Console and your content management system.
Every one of those sources assumes the old site, or at least its data, still exists. In post-redesign audits we regularly run into cases where it doesn't. The hosting has lapsed, the previous agency never handed over an export or the logs, and the new site has been live for months. At that point, the URL list from the Wayback Machine is often the only place to start rebuilding a redirect map. We treat it as a supplement rather than a full inventory, since the archive only holds what its crawlers reached, and we merge it with every other source we can still recover.
A site move involves a lot more than the redirect map, from verifying both the old and new sites in Search Console to watching traffic on both. This article sticks to the part the archive helps with.
Pulling Old URLs with the Availability and CDX APIs
The archive offers two interfaces that matter here. The Wayback Availability JSON API answers a simple question: is a given URL archived? A request to https://archive.org/wayback/available?url=example.com returns the closest capture with its archived URL, timestamp and status. The optional timestamp parameter sets the date to search around, and if nothing is archived, the response contains an empty archived_snapshots object. The archive itself suggests using this API in a 404 handler that checks whether an archived copy is available to display.
The Wayback CDX Server API goes deeper. It serves the same index the Wayback Machine uses to look up captures, one row per capture, with the fields urlkey, timestamp, original, mimetype, statuscode, digest and length. One query is enough to pull the archived URLs of a domain:
https://web.archive.org/cdx/search/cdx?url=example.com/*&output=json&fl=original,timestamp,statuscode,mimetype&filter=statuscode:200&filter=mimetype:text/html&collapse=urlkey
The trailing /* turns the query into a prefix match on everything under that path, fl picks the columns, the two filters keep only HTML pages captured with a 200 response, and collapse=urlkey keeps one row per URL. The documentation also describes matchType=domain, which includes subdomains, and from and to for narrowing the date range. On large sites, you'll want result limits, pagination or the resumption key, and the documentation itself warns that a prefix query collapsed by urlkey may be slow.
Keep in mind that statuscode is the response the archive's crawler got on the capture date, not the page's status today. The same list also helps you find orphan pages: old URLs that still return 200 but no longer have a single link pointing to them from the current site.
Turning the URL List into Redirects
A list of old URLs fixes nothing on its own. Each one has to be requested today, and the result compared with what Google's site move documentation says about redirects:
- Content that moved gets a server-side permanent redirect. Google recommends permanent redirects such as 301 and 308, pointing straight at the final destination. If a chain is unavoidable, the documentation advises keeping it short: ideally no more than three hops, and fewer than five.
- Many old URLs sent to one irrelevant page, such as the new homepage, are a mistake by Google's account: they can confuse users and might be treated as a soft 404. The exception is content from several pages consolidated into one, in which case the old URLs can point to that new page.
- Content you didn't move should return a 404 or 410 on the new site, which Google's documentation lists as part of preparing the new site. What those codes mean to Googlebot is covered in our article on HTTP status codes and SEO.
Google recommends keeping redirects in place for as long as possible, generally at least a year, because that window lets Google transfer all signals to the new URLs, including recrawling and reassigning links from other sites. From the user's perspective, the documentation suggests considering keeping them indefinitely.
Checking What a Page Used to Say
The second thing you can't reconstruct from the live site is what it used to say: the exact wording, the order of sections and the links that led out of each page.
When a page loses visibility after a redesign for queries it used to answer, we compare the archived version with the current one as part of the audit. More often than not, the answer is right there: the paragraph that answered the question is gone, the title or heading changed, or the page lost its internal links from the menu and related articles. An archived copy turns a debate about "what was there before" into a comparison of two concrete versions.
Two caveats apply. According to the archive, pages that render standard HTML archive well, but forms, JavaScript and anything that needs to talk to the original server lose their functionality in the copy, so judge the text and structure, not how the page behaved. And a capture date tells you when the archive's crawler saved the page, not when the page changed: the change happened somewhere between two consecutive captures.
Vetting a Domain Before You Buy It
Google's spam policies define expired domain abuse as buying an expired domain and repurposing it primarily to manipulate search rankings by hosting content that provides little to no value to users. The policy's examples are affiliate content on a site previously used by a government agency, commercial medical products sold on a site previously used by a nonprofit medical charity, and casino-related content on a former elementary school site. In other words, the policy describes the abuse, not the purchase itself.
The Wayback Machine lets you see what a domain hosted before you own it. Before buying a domain, or when taking over someone else's site, we go through its captures across the years. We look for gaps after which the domain comes back with a completely different topic, pages in a language that has nothing to do with the target market, content from the kinds of industries in Google's examples, and stretches when the address showed nothing but a domain-for-sale page.
A missing history proves nothing: according to the archive, pages may not be archived because of robots exclusions, and some sites are excluded at the owner's direct request. Once you own the domain, the same Google site move documentation tells you to clean it up in Search Console: check for a manual action from previous spam and for URL removals the previous owner left behind, especially a site-wide one.
What the Archive Doesn't Capture, and Removal Requests
The Wayback Machine is not a copy of the whole web. According to its help center, the archive collects publicly available pages and does not archive pages that require a password or pages that are only reachable by submitting a form. Pages may also be left out because of robots exclusions, and some sites are excluded at the owner's direct request. The archive also notes that the public pages it collects may include personal information.
Site owners can ask for their archived copies to be excluded from web.archive.org. Under the archive's instructions for removal requests, you send an email to info@archive.org with the URL or URLs, the time period you want excluded, the period during which you controlled the site or account, and anything else that helps the team understand the request. That starts a review, and the archive makes no guarantee about the outcome beforehand. Copyright claims go through a separate process under the archive's copyright policy, which says the archive may, in appropriate circumstances and at its discretion, remove content or disable access to it.
We're describing the process as the archive presents it; this isn't legal advice. It's worth remembering, though, that excluded copies are gone for anyone who later tries to reconstruct the site's history, including the owner during the next migration.
The Wayback Machine and Google Search
On September 11, 2024, the Internet Archive announced a new feature in Google Search: clicking the three dots next to a result opens the "About this Result" panel, and selecting "More About This Page" there reveals a link to the site's page in the Wayback Machine. According to the announcement, the link isn't available when the rights holder has opted out of having their site archived or when the page violates content policies. We describe the feature as announced on that date; the panel may look different by now.
Around the same time, Google was updating its documentation after retiring its own cached pages. The Google Search Central documentation changelog records that on September 24, 2024, Google removed the cache: search operator documentation because the operator no longer works in Google Search, and on October 2, 2024, it moved the noarchive rule to a historical section because the cached link feature is no longer available in search results. Google added that you don't need to remove the tag, since other search engines and services may still use it.
An archived copy is still not a view of what Google sees. It shows what the archive's crawler saved, not what Googlebot fetched. The same changelog notes that cached links aren't reliable for debugging and points to the URL Inspection tool in Search Console instead, because it has the most up-to-date version of the page.
What We Look for in an Audit
We open the archive during an audit with a specific question, not to browse an old homepage. Most of the time it's one of four.
Do the old URLs land where they should? Once we have the list from the CDX API, we request every URL on the live site and check whether it redirects in a single hop to its equivalent, returns 404 or 410 where content was removed on purpose, or gets dumped on the homepage. URLs with links from other sites, which show up in the Search Console Links report, go first, because a lost redirect there costs the most.
What disappeared from a redesigned page? We compare the capture from before the change with the current version, including the copy, title, headings and internal links, before looking for the cause of a drop anywhere else.
What history does the domain carry? Before a purchase or a takeover, we go through captures year by year, looking for changes in topic and language and for periods when the domain was parked for sale.
Do other data back up what the archive suggests? A capture shows what the archive's crawler saved on a given day. It says nothing about what Googlebot saw or whether the page was indexed, so we always check conclusions from the archive against Search Console and server data.
Summary
The Internet Archive has been preserving public web pages since 1996, and the Wayback Machine lets anyone browse them. For SEO, its value is narrow but concrete: the Availability and CDX APIs rebuild the list of old URLs for a redirect map, archived copies show what a page said before it changed, and a domain's history tells you whether there's anything to worry about under Google's expired domain abuse policy before you buy.
The archive is incomplete and doesn't show what Google sees, so its findings always need to be checked against Search Console and server data. Google's guidance on the redirects themselves is clear: make them permanent, point them straight at the equivalent page, and keep them in place for generally at least a year.
If you're planning a rebuild or a domain move and want to make sure no old URL gets lost, get a free quote. We prepare the redirect map and check it after launch as part of our web development projects.