- →A broken link is a state, link rot is a rate: clearing today's 404s changes nothing about the share of your link graph that will decay next quarter on servers you don't own.
- →Pew Research Center (May 2024) found 25% of pages that existed between 2013 and 2023 gone by October 2023, including about 20% of pages published in 2021, so decay hits recent placements too.
- →The 66.5% figure from the Ahrefs link rot analysis describes a ten year horizon across 2,062,173 domains. It's a planning assumption for contracts, not a forecast for your campaign.
- →The dangerous variant returns HTTP 200: the article stays live, your link gets edited out. No status-code monitor catches it. Store anchor plus paragraph hash at placement time and diff quarterly.
- →Track net velocity, gross acquisitions minus verified losses. Teams that only count links acquired systematically overstate how fast the profile is actually growing.
- →Redirecting dead inbound URLs to the homepage is treated as a soft 404 and transfers nothing. Map to the closest surviving equivalent or accept the 404 honestly.
What link rot really is once you stop calling it broken links
A broken link is a state you can fix tonight. Link rot is a rate, and rates don't get fixed, they get managed. That distinction is the whole operational point: you can clear every 404 in your crawl report this evening and still lose the same share of your link graph next quarter, because most of the decay is happening on servers that belong to other people.
Rot comes in at least five shapes and only one of them shows up red in a standard crawler. The hard 404 is the easy case. Then there is the expired domain that gets re-registered and repointed at something unrelated, the migration that rewrites URL patterns without a redirect map, the redirect chain collapsed into a homepage rule during a replatform, and the worst one, the page that still returns 200 while your link has been quietly edited out of the paragraph. That last variant is invisible to every HTTP status check ever written, and in our audits it is the most common form of loss on paid placements older than eighteen months.
A short visual primer before we get into the numbers:
Google's own position is calmer than the industry's. Guidance summarised in a May 2026 review of Search Console documentation is that 404 responses are a normal part of the web and do not by themselves depress indexing or ranking. That's true, and it's beside the point. Rot doesn't hurt because Google penalises dead URLs. It hurts because a link that no longer exists no longer passes anything, and because the equity you bought was priced as a permanent asset when it was really a lease with no renewal clause.
How much of the web is already gone
The Pew Research Center study « When Online Content Disappears » (May 2024) is still the cleanest baseline. Working from roughly one million pages sampled out of Common Crawl, Pew found that 25% of pages that existed at some point between 2013 and 2023 were no longer accessible by October 2023. Split by cohort the picture sharpens: 38% of pages from 2013 were gone within a decade, and about one in five pages from 2021 had already disappeared two years later. Decay is not a long-tail phenomenon that only touches ancient content, it starts biting inside a normal campaign horizon.
The figure the link building world quotes is the other one. The Ahrefs link rot analysis, built on crawl data going back to January 2013 and covering 2,062,173 domains, puts the share of links pointing at those domains that are now dead at 66.5%, with rot unevenly distributed: domains carrying more than ten live links rot considerably faster than very small sites. Search Engine Journal cited that figure in January 2026, in its guide to hiring a link building agency, as the argument for demanding a written replacement policy. Use the number, but use it honestly. It describes a decade across two million domains. Repeated as « two thirds of your backlinks will die », it's a slogan, not a forecast.
The government sample is the one worth showing a sceptical client. Pew looked at around 500,000 government pages from a Common Crawl snapshot taken in March and April 2023 and found 6% of links on those pages leading somewhere inaccessible; a separate analysis put the share of government webpages carrying at least one broken link at 21%. These are institutions with budgets, archival mandates and no commercial incentive to churn URLs. If they rot at that rate, the forty page vendor blog run by one person on a shared host is not going to do better.
Where rot hits a netlinking campaign
Three surfaces, three owners, three very different remediation costs. Outbound links from your own content are yours to fix and cost nothing but attention. Inbound links pointing at URLs you deleted or moved are also yours, and they're the highest return work on this list, since the equity already exists and you're only reconnecting a pipe. The third surface, the pages that host your paid or earned links, is the one you cannot fix, and it's where the money is.
Placements rot in ways that never trigger a status code. A publisher restructures categories and your article moves from /finance/ to /guides/finance/ with nothing behind it. An editor refreshes a two year old post and rewrites the paragraph where the link lived: a link embedded in body copy is worth more than a footer mention precisely because it depends on that paragraph surviving, which cuts both ways. A site changes hands and the new owner strips outbound links wholesale. None of that produces an alert anywhere.
This is where owning the media rather than renting a slot changes the arithmetic. Across the French editorial media we operate in-house at Stringer Network, a rotted article is a restore we perform ourselves, not a support ticket sent to a stranger who stopped answering in 2024. If you're sourcing from a catalogue you can browse without an account, the question to settle before the invoice is who guarantees the URL in three years, and what happens concretely if nobody does.
Rot also quietly corrupts your own reporting. Most teams count links acquired and never subtract links lost, which means the pace at which a profile actually grows is overstated by whatever the decay rate happens to be. Net velocity, gross acquisitions minus verified losses, is the number that belongs on the dashboard. On a portfolio built steadily since 2021 it sits noticeably below the gross figure the agency reports, and that gap is the honest answer to why rankings flattened while the link count kept climbing.
Preservation is your problem, not the archive's
The comforting story is that the Internet Archive catches whatever falls. It doesn't, or nowhere near enough. The Internet Archive's own April 2026 analysis of the Pew dataset confirms Pew's figures and adds the one that matters operationally: the Wayback Machine had preserved around 15% of the dead pages Pew identified. Six in seven vanished URLs left no public copy at all.
The stakes are clearest in science, where a citation that no longer resolves is a claim that can no longer be checked:
Academic publishing solved the naming half of this decades ago with the DOI, an identifier that survives the publisher moving, merging or folding, because resolution is indirected through a registry instead of baked into a hostname. That's the transferable lesson, and it has nothing to do with storage. The current wave of blockchain and IPFS style answers stores bytes durably and leaves the hard part untouched: nobody links to a content hash from an editorial article, browsers don't resolve them natively, and Google doesn't index them. Permanent storage under an unusable name is an archive, not a link.
What actually works is unglamorous. Stable semantic URLs decided once and never renegotiated, no dates in paths, no CMS ids exposed, a redirect map maintained as a permanent asset rather than a migration artefact, and a snapshot of every page carrying a link you paid for, taken the day it goes live. Legal and compliance teams already work this way, because evidence that no longer resolves is evidence you cannot produce. SEO teams generally don't, and then discover in year three that they have no proof a placement ever existed.
Detecting rot, tool by tool
For internal and outbound rot, Screaming Frog is still the reference. The free desktop edition crawls up to 500 URLs, and SE Ranking's June 2026 comparison lists the paid licence at 279 dollars a year for unlimited crawling, scheduling and integrations. That covers your own architecture completely, and covers your backlink portfolio not at all, which is the confusion behind half the « we monitor broken links » claims you'll hear.
For inbound decay, two Ahrefs screens do the work: the broken backlinks report in Site Explorer, which lists external links now pointing at nothing on your site, and the best by links view filtered on 404, which ranks your dead URLs by the equity stranded behind them. Search Console adds what no third party crawler can see, the crawl stats trend and the indexing report where soft 404s surface. A soft 404 is the signal worth chasing there, because it means a page went thin rather than going away, and thin is the state that quietly stops passing anything.
Browser extensions have one honest use, checking the single page in front of you. Running a portfolio review through them is spending an afternoon to produce a spreadsheet a crawler would have generated in four minutes.
None of these catch the 200 that lies. The only method we've found that holds up is to record, at placement time, the destination URL, the exact anchor and a hash of the surrounding paragraph, then re-fetch on a schedule and diff. Twenty lines of Python and a cron entry surface the losses every status-code monitor on the market is blind to. Monthly cadence for links pointing at commercial pages, quarterly for the rest. Anything more frequent is theatre, anything less and you find out a year late.
Fixing rot without making it worse
Start with inbound 404s ranked by referring domains, never by error volume. One dead URL with fourteen referring domains outweighs four hundred parameter junk 404s nobody ever linked to. Map each one to its closest surviving equivalent and 301 it there, where closest is doing all the work in that sentence.
The practical mechanics of the repair, condensed:
Redirecting everything to the homepage is the mistake we see most often, and it fails twice over. Google treats an irrelevant catch-all redirect as a soft 404, so nothing arrives, and the visitor coming from a 2019 article lands on a homepage with no path back to what they were promised. If no equivalent page exists, either rebuild the content from an archived copy or serve the 404 honestly and move on.
Redirect files are the second trap. They accumulate, and in most implementations the first matching rule wins, so a broad early pattern silently swallows the precise rule somebody added six months later to recover one specific backlink. Audit the file top to bottom rather than grepping for the line you just wrote. Chains are the third: a 301 pointing at a 301 pointing at a 301 still resolves for the user and still leaks. Whenever you open the map, rewrite intermediate hops to the final destination.
On the buying side the real fix is contractual before it is technical. Ask what happens when a placement disappears, get the replacement window in writing, and put a number on it: once you compare what a link actually costs across its useful life, a slightly dearer placement with a guaranteed replacement beats a cheap one on a site likely to be resold within the year. The same arithmetic argues for treating a campaign as something maintained rather than delivered, which is the difference between a batch of links and having someone accountable for the profile three years out.
Nautilinks operates an owned network of editorial media. In-house written articles, transparency disclosures respected, anchor mix calibrated.
Frequently asked questions
If the page hosting my backlink is deleted, does anything survive in Google's index?
Briefly, then nothing. Google keeps a cached view of the link graph for a while after a page 404s, which is why rankings sometimes hold for weeks after a placement disappears. Once the hosting URL is dropped from the index, the edge is gone with it. The lag is what makes rot dangerous: by the time you see the ranking move, the loss is months old and you're diagnosing an algorithm update that never happened.
How often should a backlink portfolio be re-verified?
Split it by value. Links pointing at commercial pages get checked monthly, everything else quarterly. Re-checking weekly buys nothing, since publishers don't edit at that cadence and you'll spend the time triaging noise. What matters more than frequency is what you check: an HTTP 200 on the hosting page is not verification. You need the anchor still present in the body copy, which means storing it at placement time and comparing.
Is a 301 to the homepage an acceptable recovery for dead inbound URLs?
No, and it's worse than doing nothing. Google treats a redirect to an unrelated destination as a soft 404, so the equity you were trying to recover doesn't arrive, and you've also destroyed the diagnostic signal: the URL now returns 200 and disappears from your error reports while still transferring nothing. Redirect to the closest surviving equivalent, or leave the 404 in place so it stays visible in Search Console.
How do you detect rot when the URL still returns 200?
Status codes cannot help you, so stop looking there. Record the destination URL, the exact anchor text and a hash of the paragraph containing the link on the day the placement goes live, then re-fetch the page on a schedule and diff against the stored values. A changed hash with the anchor gone means the link was edited out. A changed hash with the anchor intact usually means a routine content refresh. Twenty lines of scripting covers a whole portfolio.
Do archiving services actually fix rot, or just document it?
They document it, which is still worth doing. The Internet Archive's April 2026 analysis of the Pew dataset found the Wayback Machine had preserved roughly 15% of the dead pages Pew identified, so treating it as a safety net is optimistic. An archived snapshot proves a link existed and lets you rebuild lost content, but it passes no equity and Google doesn't credit it. Archive for evidence and recovery material, not for SEO value.
Should a vendor be contractually liable when a placement rots?
Yes, and the specificity of the answer tells you who you're dealing with. Ask for a defined replacement window, a stated coverage period and what constitutes a loss, since a vendor who only recognises hard 404s will not replace a link edited out of a live article. Search Engine Journal made exactly this argument in January 2026, citing the Ahrefs link rot figure as grounds for requiring a written replacement policy before signing.
Test your knowledge
Quiz: Link rot
1/3The Ahrefs link rot analysis, covering 2,062,173 domains with crawl data from January 2013 onward, put the share of dead links at what level?