Most advice about duplicate pages is written from the inside of one website. We wanted the view from the outside, so we took thirty websites we had never looked at, ran the site scan on each, and then looked at what it said before believing any of it.
The short version: the highest duplicate scores went to shops doing exactly what shops should do, the most serious problems were not duplicate pages at all but one page served at several addresses, and most of the sites we could not scan were refusing bots rather than missing a sitemap. We changed the scanner on all three counts, and this article includes what it still gets wrong.
No site is named. A problem on someone’s website is theirs to hear about first.
Edition 1, September 2026. We run this study every quarter: the same sites again, plus new ones, so each edition can say which problems were fixed and which were not. The next edition is due in December 2026.
The thirty websites
We picked four kinds of site on purpose, because each one tests a different mistake a duplicate-page checker can make, and added one SEO tool as a control.
| Kind of site | Sites | Why it is in the study |
|---|---|---|
| Local dental practices | 9 | The small business owner a scan like this is for: one location, a page per treatment |
| Leather-goods shops | 7 | Product pages that repeat on purpose: the same bag in several colours |
| Software companies | 6 | Pages built from templates at scale: integrations, comparisons, feature pages |
| Web design agencies and listicle sites | 7 | Heavily templated content: city pages, “best X in 2026” lists |
| A site from our own industry | 1 | A control |
The scan read 25 pages from each site, taken from across the whole site rather than from the top of the list, and compared every page with every other one on the wording they share once navigation and footers are set aside. Pages written separately share very little. Pages filled in from one template share most of it.
1. The highest duplicate scores went to shops working exactly as intended
Sorted by the scan’s own risk score, the site at the top was a leather-goods shop at 86 out of 100. Its closest pair of pages: one product in two sizes. They matched 100%, and still matched 100% once the site’s shared template text was set aside. The scan was right that the pages were identical. It was wrong that this was a problem.
That was not a one-off. Three of the five highest scores belonged to shops whose worst pair was one product in two sizes or two colours. The fifth-highest was the site with the most serious real problem in the whole study, described in the next section.
A shop that sells a bag in four colours has four pages that are supposed to read alike. Telling its owner that their pages are duplicates is accurate and useless, and it is how a report loses its reader before it gets to anything that matters. The same goes for a cart, a login page, an appointment confirmation or a set of terms and conditions.
So the scan now tells a product page, a utility page and a legal page from the rest. When two pages repeat each other because that is what those pages are for, the pair is still listed, with the reason, but it no longer counts toward the score. Re-scanned on the same sites:
| Leather-goods shop | Score before | Score after |
|---|---|---|
| Shop A — one product in two sizes | 86 | 34 |
| Shop B — one bag in several colours | 78 | 26 |
| Shop C — one bag in several colours | 71 | 19 |
| Shop D — one bag in two colours | 62 | 10 |
| Shop E | 62 | 10 |
| Shop F | 31 | 19 |
| Shop G — one product in several designs | 60 | 60 |
Six of the seven dropped. The seventh did not, and it is not a bug in the new rule: its remaining score comes from a brand paragraph repeated on some of its pages but not all of them. More on that below.
Across the 25 sites the scan could read, 10 had at least one pair set aside this way.
2. The worst problems were not similar pages. They were the same page at several addresses
A duplicate-page score tops out at 100%. So does one product in two sizes. Which means the most serious thing this study found scored no higher than a shop’s product catalogue until we gave it a finding of its own.
A web design agency listed five blog posts in its sitemap. The scan found all five had exactly the same text. We opened three of them by hand: each returned the same page, byte for byte, with an HTTP 200 status that tells a search engine the page exists. The blog posts do not exist.
A dental practice had the same fault. All five blog-post addresses the scan read returned the practice’s blog index instead of the post. The two we opened by hand were identical to the byte.
A software review site had two unrelated articles that both returned the same page.
None of these pages looks broken on its own. Each one loads, each one has content, and each one would pass a check that reads one page at a time. The fault only appears when two of them are put side by side, and it is the kind of fault an owner almost certainly does not know about: addresses listed in their own sitemap, promising pages that are not there. Google’s own name for a URL that answers with a real page when it should not exist is a soft 404, and Search Console has a report for it.
The scan now reports this as its own finding, “one page served at several addresses”, scored above ordinary near-copies. After the change, the three sites above became three of the four highest scores in the study, alongside a site whose case studies are built from one template.
3. Most sites we could not scan were refusing bots, not missing a sitemap
Eight of the thirty sites could not be read on the first run. The scanner said the same thing to all eight: it could not find a sitemap. Checked again, one at a time, that turned out to be right about one of them.
| What actually happened | Sites |
|---|---|
| The site refused the scanner outright (bot protection) | 5 |
| No sitemap at all | 1 |
| A sitemap that exists, but refused the scanner | 1 |
| A sitemap that listed no pages on the site’s own domain | 1 |
Three of the five refusals were dental practices, which matters: the kind of site this scan is most meant for is also the kind most likely to sit behind bot protection it never chose.
The scan now reads the three it can, even without a usable sitemap. It also says which of the three situations it was in, because the advice is different: a site with no sitemap should publish one, while a site whose sitemap refuses automated requests should check that search engine crawlers are let through. The five that refuse every request still cannot be read, and the report now says the site blocked the scan instead of blaming a missing sitemap. Getting round bot protection is not something a scanner should do.
4. Template text can make two unrelated pages look like copies
On one dental practice, the closest pair was two patient forms at 51%. With the site’s shared template text set aside, the same pair measured 0%. Every word they shared was the site’s own furniture.
That is why every pair the scan reports now carries a second figure, the overlap with the site’s template text set aside. When the two numbers are far apart, the pages are not copies of each other; they share a template. When both are high, the writing itself repeats.
The limit is Shop G from section 1. Its repeated paragraph sits on some of its pages rather than across the whole site, so it is not set aside, and the pair still scores. A block shared by one family of pages is the hardest thing for a whole-site comparison to separate from real duplication, and the scan does not do it yet.
5. One site in five was clean
Five of the 25 readable sites scored zero: nothing duplicated, nothing thin enough to mention, no templated titles. Three of the five were software companies. A clean result is a real answer, and a scan that always finds something is not measuring anything.
The numbers, all in one place
How many of the 25 readable sites had each finding, overall and by kind of site. The five sites that refused the scanner are not in these counts, and pairs the scan set aside as repeating by design are counted only in the row that says so.
| Readable sites with… | All (25) | Dental (6) | Shops (7) | SaaS (6) | Agency (5) |
|---|---|---|---|---|---|
| One page served at several addresses | 3 | 1 | 0 | 1 | 1 |
| At least one near-copy pair | 9 | 2 | 1 | 3 | 3 |
| Enough short pages to report | 8 | 4 | 1 | 1 | 2 |
| The same title on more than one page | 9 | 2 | 4 | 1 | 2 |
| The same description on more than one page | 10 | 1 | 5 | 2 | 2 |
| Titles or descriptions differing by one word | 6 | 3 | 2 | 0 | 1 |
| Templated titles or descriptions | 3 | 1 | 0 | 0 | 2 |
| Pairs set aside as repeating by design | 10 | 3 | 6 | 1 | 0 |
| Nothing to report at all | 5 | 0 | 0 | 3 | 1 |
Two things stand out. Four of the six dental practices had enough short pages for the scan to report it: treatment pages that say very little beyond the treatment’s name. That is the kind of site this scan is most meant for, and the finding most worth fixing.
And the shops shared titles and descriptions more than anyone, but every shared title and description we checked was on variants of one product: the same by-design repetition as section 1, in the page’s metadata rather than its text.
Check your own site for these three problems
You do not need a tool for a first look. Three checks, in order of how much they are worth:
- Open two pages from your sitemap that should be different, especially blog posts and location pages, and compare what loads. If the text is identical, one address is serving a page that is not its own. Check the status code too: a page that should not exist should answer 404, not 200.
- Pick two pages from the same template, two towns, two services, two products, and read them side by side. If everything except a name is the same, that is the pattern a site-wide review is built to notice. If they are product variants, that is fine.
- Look at the text that repeats. If what two pages share is navigation and a footer, they are not duplicates. If it is the body of the page, they are.
The free site scan runs all three across 25 of your pages and names the addresses behind every number. For why whole-site repetition matters to Google in the first place, the history of the Panda update explains where judging a site as a whole came from, and the helpful content survival guide covers what changed since.
What we got wrong along the way
Scanning thirty sites turned up four faults in the scanner, and three of them would have put a confident, false statement in front of a site owner:
- It ranked product variants and confirmation pages above real problems, as described above.
- Its first attempt at the “same page at several addresses” finding fired on two product variants sharing a description. We caught it by scanning the same thirty sites a second time.
- It told five sites they had no sitemap when they were in fact refusing the scanner.
- It told one site to publish a sitemap it already had, because the sitemap refused the scanner while the home page did not.
All four are fixed and covered by tests, and written down in a log we keep of every way a real website has proved the scanner wrong. The rule that comes out of it is the one this study was run to test: a wrong finding about someone’s own website is worse than no finding, so the scan has to be checked against real sites before anyone is asked to believe it.
Frequently asked questions
Are product variants duplicate content?
They repeat by design. A shop that sells one bag in four colours has four pages that are supposed to read alike, and nothing on them is broken. In this study the highest duplicate score of all belonged to a product sold in two sizes. The site scan now lists product-variant pairs with the reason they were left out, instead of counting them against the site.
What is a soft 404?
A URL that should not exist but answers as if it does: the server returns a normal page with a 200 status instead of a 404. Two of the clearest problems in this study looked like this — sitemap addresses for blog posts that did not exist, each returning the same page. Search Console has a report of the same name, and it is worth checking if a scan finds this.
How can I tell whether two of my pages are really the same page?
Open them in two tabs and compare what loads. If the text is identical word for word, it is one page served at two addresses, not two similar pages. The site scan reports this separately from pages that are merely alike, and names both addresses.
Why could some websites not be scanned?
Eight of the thirty could not be read at first. Five of them refused the scanner outright: bot protection turning away automated requests. One had no sitemap, one had a sitemap that refused the scanner, and one had a sitemap listing no pages on its own domain. The scan can now read the last three. The five that refuse every request still cannot be read, and the report says so rather than guessing.
Were any of the thirty websites told about their problems?
Not by this study. The sites are described here by type only, because a problem on someone’s website is theirs to hear about first.
Can I check a scan’s numbers myself?
Yes. Every number the scan reports names the pages behind it, so you can open those pages side by side and see the repeated wording for yourself. You can run the same site scan used in this study on your own site, free.