SEO Fundamentals

How to Fix Location Page Duplicate Content at Scale

Location page duplicate content is a sameness problem, not a copying one. Here is how to measure sibling distinctness and fix the pages worth keeping.

S SparkCliks 0 21 min read
Share
How to Fix Location Page Duplicate Content at Scale

Location page duplicate content is almost never a copying problem. Nobody stole anything from you. You built one template, fed it 240 city names, and now most of those URLs sit unindexed while two of them carry every impression in the report. The fix is not more words. It's knowing exactly how similar your pages are, deciding which locations deserve a page at all, and giving the survivors something a search engine cannot get from their siblings.

The short answer

Three separate mechanisms get filed under "duplicate content", and they have different symptoms and different fixes. Most location page problems are the second one, and most advice on the internet is written about the first.

MechanismWhat the engine is doingWhere you see itThe real fix
Near-duplicate clusteringGrouping URLs it considers the same page and picking one to show`Duplicate without user-selected canonical`, `Duplicate, Google chose different canonical than user`Make the pages genuinely different, or accept the cluster and pick your own canonical
Selection declineFetching the page, judging it low value, declining to store it`Crawled - currently not indexed`Add substance the sibling pages do not have
Doorway policyTreating a set of near-identical geo pages as manipulationA manual action, or a set that quietly never gains tractionDelete the set, or rebuild it around real service differences

There is no generic penalty for having similar pages. Search engine documentation on duplicate content is explicit that duplicate content on a site is not grounds for action unless it looks intended to deceive. What you're actually fighting is a storage decision. An index slot costs money to fill and money to keep, and a page that repeats a sibling has to justify why anyone should pay twice.

Location page duplicate content is a sameness problem

Run this thought experiment. You have /plumbing/austin/ and /plumbing/round-rock/. Both say you're a family-run plumber with 20 years of experience, both list the same three services, both carry the same testimonial, both close with the same phone number. The only difference is the city name in the H1, the title tag, and four places in the body copy.

Nothing was copied. Both pages were generated the same way, from the same source, at the same time. And a search engine reading them has a genuine problem: it cannot tell what a searcher in Round Rock would get from the second page that they would not already get from the first. Neither can you.

That's the whole diagnosis. If the difference between two of your pages is a find-and-replace, the pages are one page with a variable in it. Every fix that follows is really the same fix stated at different levels of effort: put something on each page that would be false if you moved it to the page next door. If your crawl and indexing setup is shaky underneath all this, check that first, because a sameness problem sitting on top of a delivery problem looks identical in Search Console. The technical SEO fundamentals cover that layer.

Free trial

Stuck on page two?

Real human clicks that lift your CTR and move you up the rankings.

Three ways a location set goes wrong

Location sets fail in three recognizable patterns, and the repair differs for each.

Spun copy. Someone took 400 words and rewrote them 200 times, by hand or with a model, so no two pages match word for word while all 200 say the same thing. It's the worst of the three, because it costs real money and buys nothing. Distinct wording is not distinct information: swapping "our skilled engineers" for "our experienced technicians" adds zero facts, 200 times over.

Template sameness. The honest version. One template, real data fields, and the fields are thin: city name, a map embed, maybe a phone number. Every page is truthful and every page is interchangeable. Most well-intentioned programmatic sets sit here, and this one is recoverable.

Doorway intent. Pages built for the query rather than for the visitor, funneling everyone into the same conversion path regardless of which location they landed on. Search engine spam policies name this directly, and the published examples include multiple pages targeted at specific regions or cities that funnel users to one page. If your Austin page and your Round Rock page both submit to the same generic form with no location context, you're closer to this than you think.

The tell that separates pattern two from pattern three is whether you'd still want the page if it never ranked. If the answer is no, it's a doorway.

Two tests you can run in five minutes

Before any tooling, two manual tests catch the majority of bad location sets.

The swap test. Open two sibling pages side by side and swap the location names. Does any other sentence now need to change? Does any statement become false? If nothing has to change and nothing becomes false, the two pages are the same page and a search engine will treat them that way. A page that passes says something like "we run same-day callouts in Round Rock but next-day in Austin, because the depot is on the north side", and that sentence breaks the moment you move it.

The row test. List every fact unique to one page: city name, phone number, service radius, opening hours. Now try to fit all of them into a single row of a table where the location name is the row key. If they fit, the page is that row, and the honest structure is one page with a table of locations rather than 240 separate URLs. A location earns its own page when its unique facts overflow a table cell: local pricing rules, a different service mix, named local projects, market-specific questions.

Most teams have never run either test against their own set, which is why so many location directories are 200 pages that should have been one page and six pages.

Measure sibling distinctness properly

Those tests are qualitative. For a set of any size you want a number, and the number worth tracking is sibling distinctness: the share of the main content on page A that does not also appear on page B, measured between two typical siblings rather than against the whole web.

One setting decides whether that number means anything. Compare full HTML and your global nav, mega-menu, footer and cookie banner dominate the result, so every page on every site looks like a near-duplicate. Compare the main content area only.

Here's the practical route with Screaming Frog SEO Spider, which most teams already have installed:

  1. Open Configuration, then Content, then Area. Restrict the content area to your main content container, or exclude the nav, header and footer classes. This step is what makes every number after it meaningful.
  2. Open Configuration, then Content, then Duplicates. Enable Near Duplicates and set the similarity threshold. The default is 90% similarity, which flags anything under 10% distinct.
  3. Crawl the location set, restricted to the location path, so you're not diffing blog posts against city pages.
  4. Run Crawl Analysis. The Near Duplicates filter stays empty until you do, which is the most common reason people conclude their pages are fine.
  5. Read the Content tab. It reports the closest similarity match and a near-duplicate count per URL.

No budget for tooling? The crude version works: extract the visible text of the main content block from two siblings, paste both into any text diff tool, count changed lines against total lines. Rougher than a token-level comparison, and plenty to tell 1% apart from 20%.

Now the interpretation. These bands are SparkCliks' working thresholds from auditing programmatic sets, including our own. No search engine publishes a distinctness number, and anyone quoting you one as official is guessing.

Sibling distinctnessWhat it actually isWhat to do about it
Under 2%A clone with the name swappedDo not publish. Build one page with a location table
2% to 5%A template with a name variableConsolidate, or add real per-location substance before publishing
5% to 15%A template with real data fields (language, hours, coverage)Publishable for a small set. Expect slow, partial indexing
15% to 30%Template plus per-location guidance and local FAQsThe realistic target for a programmatic set
Over 30%Genuinely different pagesNormal editorial territory, no special handling

Notice where the 90% default lands: right at the bottom edge of the publishable band. Treat a flagged near-duplicate as a stop, not a warning.

Worked example: two country pages, diffed

SparkCliks runs its own programmatic location set, one page per country, on the pattern /web-traffic/buy-website-traffic-from-COUNTRY/, so none of this is hypothetical. You can read one at buy website traffic from Australia. During a rebuild, the team diffed the live Australia page against the live Zimbabwe page. Two lines differed out of 5,064, and both of them were CMS post IDs. Everything a visitor could read was identical. That is a textbook doorway build, and the internal notes call it exactly that.

The rebuild's first template did better and still not well: siblings measured about 1% distinct, because only the hero section varied by country. An earlier version of the same rebuild, carrying per-market guidance and a country-specific FAQ, measured 12 to 15% distinct. Same product, same data, same designer. The difference was two content blocks.

Three things from that audit are worth stealing.

What gets left out matters as much as what goes in. The rebuild's country data file carries no population figures, no internet-user counts and no market-share percentages. Those would have raised distinctness immediately, and they need a dated source, they go stale silently, and a wrong figure repeated across 246 pages is 246 wrong figures to correct. Padding distinctness with unsourced statistics trades an indexing problem for a credibility problem.

A locale database will hand you self-cannibalizing pages. The country list is generated from ICU and CLDR data rather than typed by hand, and CLDR still names deprecated region codes: DD for East Germany, SU for the Soviet Union, YU for Yugoslavia. Left in, the build would have produced two pages competing for /germany/ and two for /russia/, with the winner decided by whichever sorted last. If you generate locations from any reference dataset, dedupe on the human-facing name, not on the code.

Only publish a location you actually serve. The generated list covers 246 inhabited regions. The product's own copy says 160-plus countries. Publishing the gap would mean promising service in a market the engine cannot reach, which is a broken promise at checkout and a page with nothing true to say. Your coverage list is the publishing gate, not the locale database.

What genuinely varies by location

Most location pages stay thin because nobody ever sat down and listed what actually differs. Do it once, as a table, and the template writes itself.

ElementReal variation?Notes
PriceUsually noIf price is identical everywhere, say so once on a shared page and stop pretending otherwise
Service availability and coverage radiusYesThe most under-used differentiator. Which services are not offered here, and why
Delivery or response timeYesDepot distance, ferry crossings, rural drive time
Language and locale conventionsYesSpelling, currency format, date order, the language the page should be written in
Operating hours and time zone overlapYesReal hours in the local zone, plus the window when your support actually overlaps theirs
Regulation, tax and complianceYes, often heavilyFrequently the single richest source of legitimately unique copy
Local proofYes, where it existsNamed projects, local photos, a review that mentions the place. Do not invent this
Local demand and vocabularyYesWhat buyers in this market call the thing they're looking for
Competitors and alternativesYesWho a buyer here compares you against
How the product worksNoShared explanation, written once, kept short on the location page

SparkCliks' country data shows what the middle rows look like in practice. Each country carries its ISO code, the languages a campaign can target with the exact BCP 47 tag the engine's language setting takes, and real IANA time zones for the region. Three fields, true per country, verifiable, impossible to fabricate. Pricing is not among them, because it genuinely does not vary by country, and inventing per-country prices to inflate distinctness would have been a lie bought for a few percentage points. For the geographic angle on the traffic side rather than the content side, geo-targeting website traffic covers the delivery mechanics.

Keep, consolidate, noindex or delete

With a distinctness number and an honest list of what varies, every page in the set falls into one of five buckets.

SituationActionWhy
Real service difference, real local demand, unique factsKeep and indexThis is what a location page is for
Two locations serving one market with identical deliveryMerge, 301 the weaker URL into the strongerTwo pages splitting one market's signals help nobody
Real location, page not enriched yet`noindex, follow`, ship the content, then remove the tagBetter than an indexed page teaching the engine your set is thin
Location you do not actually serveDeleteA page that cannot be honest about a market has nothing to say about it
Print, sort, tracking-parameter or session variants of one page`rel=canonical` to the primary URLThis is the classic duplicate case, and it is not your location problem

The row people get wrong is the third, and the specific mistake is reaching for rel=canonical. Canonicalizing your Austin page to your Dallas page does not fix duplicate content; it withdraws Austin from search results and hands its signals to Dallas. A canonical is a statement that two URLs are the same page. If your location pages are meant to be different pages, declaring them identical is the opposite of what you want. Use noindex while a page is being built out, and be precise about which directive does what, because the crawl-versus-index distinction underneath them trips people constantly: robots.txt versus noindex covers where each one applies.

Worked example: audit a location set in Search Console

Here's the sequence to run against a live set. Every number below is an illustrative example, not SparkCliks data.

Step 1: segment the sitemap first. Put the location set in its own sitemap file, referenced from your sitemap index. This is the highest-value 10 minutes in the whole audit, because the Page indexing report can be filtered by sitemap, and without that filter you're reading your location set's health mixed in with every blog post and product page you own. If sitemap submission is new to you, submitting a sitemap and checking indexed pages walks through the setup.

Step 2: read the Page indexing report for that sitemap. In Search Console, open Indexing, then Pages, then switch the source from "All known pages" to your location sitemap. Four rows matter most:

  • Duplicate without user-selected canonical: the engine clustered these siblings and picked a representative. Your set has already collapsed.
  • Duplicate, Google chose different canonical than user: you declared a canonical and it was overruled. Usually means the pages really are the same.
  • Alternate page with proper canonical tag: expected and fine, if you deliberately canonicalized these.
  • Crawled - currently not indexed: fetched, judged, declined. A value verdict rather than a fault, and the most common outcome for a thin location set. What crawled but not indexed really means covers how to read that status without burning weeks on resubmits.

Step 3: compute the activation rate. Open Performance, then Search results. Set the date range to the last 3 months if the set has been live longer than that, otherwise last 28 days. Open the Pages tab, add a filter of Page contains /locations/ (use your real path), and export.

Divide the number of location URLs with at least one impression by the number of location URLs you published. That share is your activation rate, and for a programmatic set it beats clicks or average position as a health metric, because it tells you how much of the set exists at all. An example: 240 pages published, 31 with any impression in 90 days, gives an activation rate of 13%. An enriched set runs high. A doorway build runs in single digits, with two URLs holding nearly all the impressions.

Step 4: inspect a sample sibling. Take three URLs with zero impressions, run each through URL Inspection, and compare "User-declared canonical" against "Google-selected canonical". If the selected canonical points at a different location page, that cluster has already merged, and no amount of title-tag tweaking will separate them. Only new information will.

Step 5: sort the results into three piles. URLs at Crawled - currently not indexed have a value problem. URLs with a duplicate status have a sameness problem. URLs that are indexed and collecting impressions without clicks have neither, and that's a snippet question instead: measuring organic CTR in Search Console is the right report for those.

Publishing gates for a programmatic set

Every rule below exists because some set got published all at once and sank together.

Ship in batches, not in bulk. Publish 10 to 20 pages, wait four full weeks, read the activation rate for that batch, then decide whether to continue. A set of 240 pages published in one afternoon gives you exactly one experiment and no way to change course.

Gate on distinctness, not on word count. No page ships unless a diff against its nearest sibling clears your floor. Word count gates produce padding, which lifts the count and leaves distinctness where it was. And every unique fact has to be verifiable: if a location needs a fabricated statistic to clear the floor, that location does not get a page.

Give each page a real internal link. A location page whose only inbound link is a 240-item footer list has a discovery problem stacked on top of its sameness problem. Link from the parent service page, and link between locations that genuinely relate: neighboring markets, same region, shared depot.

Keep the shared boilerplate short. If 80% of the page is the same explanation of how your service works, move most of that explanation to one page and link to it. Shorter shared blocks raise distinctness without adding a single new word.

A before and after test design

Enriching 240 pages is expensive, so prove the enrichment works on a slice first. Numbers here are examples for shaping the design, not results.

Split the set into three groups. Rank all location URLs by baseline impressions, then assign them round-robin into groups A, B and C. Round-robin off a ranked list keeps the baseline comparable across groups, which alphabetical splitting does not.

  • Group A: full enrichment. Per-location guidance section, local FAQ block, local proof where it exists, shortened shared boilerplate.
  • Group B: data fields only. Hours, coverage, language, response time. No new prose.
  • Group C: control. Touch nothing, including the shared template, for the duration.

Windows. Baseline is the 28 days before the change. Wait 14 days after publishing for recrawl and recomputation, then read the following 28 days. Do not read the gap period; it's a mix of both states and it will mislead you.

Metrics, in priority order. Activation rate for the group. Count of URLs sitting at Crawled - currently not indexed in that sitemap. Total impressions for the group. Rankings are the wrong primary metric here, because the open question is whether the pages exist in the index at all.

What counts as signal. With 80 pages per group, three or four pages changing state is noise. Look for a shift of at least 10 percentage points in activation rate, moving in the same direction as the indexed count, with the control group flat. If group C moves too, something site-wide or seasonal happened and the test tells you nothing.

Then price it. If group A moves and group B doesn't, per-location prose is the lever, and you now know what the remaining pages cost to fix. If both move about the same, ship group B's treatment everywhere and keep the money.

What does not fix it

  • Rewriting the same content 200 different ways. Distinct wording, identical information. The clustering that grouped your pages was never based on exact string matching.
  • Word count padding. Adding 600 words of shared boilerplate lowers distinctness, because the shared portion of every page just grew.
  • Map embeds, weather widgets and population blurbs. Third-party embeds usually aren't read as your content, and a scraped population figure is an unsourced number you now maintain in 240 places.
  • Canonicalizing siblings to a hub. It removes the pages from results. That ends the problem rather than solving it.
  • Requesting indexing over and over. The status you were given is a decision, not a queue position.
  • Buying traffic or clicks. SparkCliks sells search clicks and website traffic, and this is the honest answer: neither one is an indexing lever. An engine that declined to store a page because it repeats its siblings does not reopen that decision because visits arrived. Click and traffic work operates on pages that are already indexed and already ranking, where the question is what happens on the results page. Different problem, different tools, and pointing them at an unindexed location set is spending money on a symptom. One more limit worth naming: if your location pages carry network ad code, automated visits counted as ad impressions are invalid traffic under every major ad network's rules, and the penalty lands on your account rather than the vendor's.

The one thing that reliably moves a location set is information a searcher in that location cannot get from the page next door. Everything on this list is an attempt to avoid producing it.

Frequently asked questions

FAQ

Does location page duplicate content cause a Google penalty?

Not in the way most people mean it. Search engine documentation states that duplicate content on a site is not grounds for action unless it appears intended to deceive. The normal outcome is that near-identical pages get clustered and one is chosen to represent them, so the rest quietly stop appearing. Doorway pages are a separate, explicitly named spam policy, and a large set of near-identical city pages funneling to one destination can fall under it.

How similar is too similar for two location pages?

No search engine publishes a threshold, so treat any number you see as a working heuristic rather than a rule. In practice, siblings under about 5% distinct behave as one page, 5% to 15% gets slow and partial indexing, and 15% to 30% is a realistic target for a programmatic set. Screaming Frog's near-duplicate default flags anything under 10% distinct, which is a sensible line to gate on.

Should I use canonical tags on my location pages?

Only if you genuinely want one page to represent the others in search results. Canonicalizing your Austin page to your Dallas page withdraws Austin from results and passes its signals to Dallas. If the pages are meant to be separate, use noindex, follow while you build them out, then remove it once they carry real substance.

How many location pages should I publish at once?

Publish 10 to 20, wait four full weeks, then read the activation rate for that batch before continuing. Shipping the full set at once gives you a single result with no way to diagnose it, and a thin set published in bulk teaches an engine something about your site that takes a long time to unlearn.

Do location pages work without a physical address in each city?

They can, provided every claim on them is true. A service-area page that states what you actually cover, how long you take to get there, and which services are available in that market is legitimate. A page implying an office that does not exist is a different problem, and local map results carry their own verification requirements that a service-area page does not satisfy.

Will buying traffic or clicks get my location pages indexed?

No, and SparkCliks says so despite selling both. Indexing is a value decision made after the crawl, and incoming visits do not reopen it. Traffic and click services act on pages that already appear in results, where the question is how a listing performs on the results page rather than whether the page gets stored at all.

About the Author

The SparkCliks Team builds and measures search-click and website-traffic services at sparkcliks.com, and runs its own programmatic country pages, which is where the diff numbers in this post come from. We publish what our own audits turn up, including the parts showing that a build we shipped was thin. Questions or corrections: hello@sparkcliks.com.

Keep reading

Related articles