Skip to main content

HubSpot

Where Launch-Day Orphan Pages Come From


Orphan pages are usually a data-model mismatch, not an oversight. And you cannot redirect a slug a published page still occupies.

Julia HanneyBy Julia HanneyUpdated October 1, 20267 min read
A site map showing published pages with no inbound internal links after a launch.

Key Takeaways

  • Orphans usually come from two systems disagreeing about the same entity, not from someone forgetting a link.
  • A published page holds its slug: you cannot point a redirect at a URL that still exists, so unpublish or move first.
  • Crawl the staging build for pages with zero inbound internal links before launch, not after.
  • Absolute internal links written against the old domain pass every staging test and break at cutover.
  • Decide the canonical navigation value in the data model, then generate navigation from it rather than curating it by hand.

An orphan page is a published page with nothing linking to it. Visitors cannot reach it by navigating, internal link equity never flows to it, and it usually sits undiscovered for months because every automated check says the page is fine. It exists, it returns 200, it renders.

The instinct is to treat it as sloppiness: somebody forgot a link. In our experience that is rarely what happened.

The real cause is usually two systems disagreeing

The version we hit on a multi-location build is the clearest example.

The client's source records were filed under the physical town each location sat in. The website's navigation was organised by the marketed city, the larger place the business advertised into. Both were correct in their own system. Both had been correct for years.

Pages generated from the source data landed under the physical town. Navigation, built from the marketing structure, listed the marketed city. The two values matched for some locations and not for others, and every location where they diverged produced a page nobody could reach.

Nobody forgot anything. Two teams used two vocabularies for the same entity, and the launch was the first moment the disagreement became visible.

Once you know the shape, you see it everywhere:

  • Product pages generated from SKU categories, navigation built from marketing categories.
  • Service pages generated per service line, navigation grouped by customer problem.
  • Team pages generated per person, navigation grouped by department, where somebody sits in two departments or none.
  • Location pages where the source has a branch code and the site has a region.

The common factor: the page set is generated from one field and the navigation is curated from another.

The fix is a decision, not a crawl

Crawling finds orphans. It does not stop them coming back, because the next content addition reproduces them.

The durable fix is to name a canonical navigation value in the data model and generate both the pages and the navigation from it. If the business genuinely needs two groupings, physical town for operations and marketed city for marketing, then hold both as fields on the record and be explicit about which one drives the site.

That conversation is a data-modelling conversation and it belongs in discovery, not in launch week. Ask one question during scoping: what is the single field that decides where this page appears in navigation? If nobody can answer it, or if two people answer differently, you have found your orphans months early.

The trap that compounds it

Here is the part that turns a tidy fix into a bad afternoon.

You cannot redirect a slug that a published page still occupies.

The natural remediation for an orphan is often to fold it into another page and redirect the old URL. But if the page is still published at that URL, the redirect will not take effect. The page holds the slug.

The order that works:

  1. Unpublish the page, or change its slug to free the URL.
  2. Create the redirect.
  3. Verify it resolves, from a clean session, rather than assuming.

Getting this wrong produces the worst outcome available: a redirect that exists in the interface, is listed in the redirect table, and does nothing. Everybody assumes the problem is solved and the orphan is still live and still indexable.

The staging blind spot

One more mechanism worth knowing, because it defeats pre-launch QA specifically.

Absolute internal links written against the old production domain resolve perfectly during staging. The old site is still up and still answering, so a link to https://theirdomain.com/services works when you click it in the staging build. It passes review. It passes a link checker.

Then DNS moves, the old host stops answering that domain, and the links break, or worse, resolve to the wrong place.

The check is cheap and specific: crawl the staging build for absolute links pointing at the production domain and convert them to relative before cutover. Every time.

The pre-launch pass we run

Fifteen minutes, before cutover:

  1. Crawl the staging build and list every page with zero inbound internal links.
  2. Diff that list against the page inventory and the sitemap. A page in the sitemap with no internal links is an orphan that search engines will find and users will not.
  3. Crawl for absolute links to the production domain.
  4. Check the navigation against the generated page set, not against the design. The design shows what somebody intended; the page set shows what exists.
  5. For every orphan, decide: link it, fold and redirect it, or unpublish it. Doing nothing is a decision too and it should be a conscious one.
  6. Check internal links on 100 percent of pages. Sampling is fine for cross-browser QA, where problems live in templates. It is not fine for links, where problems live in individual pages and are cheap to check programmatically.

That last distinction is worth internalising: sample what fails by template, be exhaustive about what fails by instance.

The other three sources, in order of frequency

The data-model mismatch is the largest cause. Three more account for most of the rest, and each has a different fix.

Pages built for a navigation that changed. A section is designed, pages are written against it, the navigation is revised late in the project, and the pages survive the revision because deleting content feels risky. Nobody links to them because they belong to a structure that no longer exists. The fix is a content inventory reconciled against the final navigation, which sounds obvious and is skipped under deadline pressure roughly always.

Pages reachable only from something that did not launch. A page linked exclusively from a campaign asset, an email, a gated resource or a component that got cut. It was never orphaned in the plan; the thing that pointed at it did not ship. Worth an explicit question during QA: for every page, what links to it, and did that thing launch?

Pagination and filtered views. Listing pages that expose only the first page of results, or that require a filter to reveal certain entries. Technically linked, practically unreachable, and invisible to a naive crawler that only follows the first page. This one disproportionately affects large templated page sets, which is to say exactly the HubDB-style builds where the page count is highest.

A useful framing when you present findings to a client: an orphan is not a broken page, it is a page the site does not believe in. That usually turns the conversation from "who forgot the link" into the more useful "should this page exist at all," and a decent share of orphans should simply be unpublished.

What it actually costs

Worth being straight about, because the temptation is to treat orphans as tidiness rather than as a problem with a number attached.

Internal link equity does not reach them. A page with no inbound internal links is asking search engines to rank it on external signals alone, on a site where every other page is getting help.

They are still indexable. Orphaned does not mean hidden. Search engines find pages through sitemaps regardless of internal linking, so an orphan can rank, receive traffic, and deliver a visitor to a page with no navigation back into the site.

They accumulate silently. Nobody audits pages nobody visits, so orphans from a 2023 launch are still there in 2026, still indexed, still representing the client.

They dilute a templated page set. On a large generated set, a tail of unreachable near-duplicate pages is exactly the profile that thin-content assessments are looking for.

After launch

Two checks in the week after cutover, because some orphans only appear once the site is live:

Search Console coverage. Pages discovered but not linked from anywhere show up as indexed with no referring pages. This is the cheapest ongoing orphan detector available and it costs nothing to look at.

The redirect table, verified rather than read. Test a sample of redirects by requesting them. A redirect that was created while its target slug was still occupied will sit in the table looking correct.

The bottom line

Orphan pages are usually two systems disagreeing about how to name the same thing, not somebody forgetting a link. Fix the data model, generate navigation from the same field the pages come from, and the class of problem disappears rather than being cleaned up repeatedly.

Then remember the two mechanics that defeat ordinary QA: a published page holds its slug so redirects at that URL do nothing, and absolute internal links to the old domain pass every staging test and break at cutover.

If your agency delivers site launches and would rather have someone else own the pre-cutover pass, our white-label web design team runs these under partner brands. The DNS cutover runbook covers the launch itself.

Sources

  1. HubSpot: Connect a domain to HubSpot (opens in new tab)
  2. HubSpot: Manage system domains (opens in new tab)

Frequently Asked Questions

What causes orphan pages after a website launch?

Most often a data-model mismatch. The source data files each record under one value, such as a physical town, while the site's navigation is built around a different one, such as a marketed city. Pages generate correctly and nothing links to them, because the two values never match.

Can you redirect a URL that a published page still uses?

No. A published page occupies its slug, so a redirect pointing at that URL will not take effect. Unpublish the page or change its slug first, then create the redirect, then verify it resolves rather than assuming.

How do you find orphan pages before launch?

Crawl the staging build and list every page with zero inbound internal links, then compare that list against the sitemap and the page inventory. Do it before cutover, because after cutover the same pages are live, indexable and invisible to your own navigation.

Why do internal links break only after DNS changes?

Because absolute links written against the old production domain still resolve during staging QA, since the old site is still answering. They pass every test right up to the moment the domain moves, then break. Crawl for absolute links to the production domain before launch.

Should navigation be built by hand or generated?

Generated from the same field the pages are generated from, wherever the page set is large or repeating. Hand-curated navigation over a generated page set is the exact condition that produces orphans, because the two drift the moment anything is added.

White-Label Web Design

A Creative Team On Demand, Under Your Brand

Conversion-focused design your clients will love and your PMs can rely on: 12,000+ projects and tasks delivered.