Most lists of website mistakes to avoid are written from opinion. This one is written from evidence. We took 50 websites we had never looked at — local businesses, online shops, software companies, blogs, directories, and sites in five languages — found the way a customer would find them, by searching. We ran our site scan on each, then verified every finding against the live pages before counting it.
42 of the 45 sites we could read had at least one of the common website mistakes below. Almost none of them is visible on a single page. They appear when pages are put side by side, which is how a search engine reads a site and how its owner almost never does. Most are not design choices at all but website maintenance mistakes: left behind by a migration, a plugin, a renamed page or an empty template field.
No site is named. A problem on someone’s website is theirs to hear about first.
The common website mistakes, at a glance
| Mistake | Sites (of 45 readable) |
|---|---|
| Pages built from one template with only a name swapped | 23 |
| The same title or description on different pages | 23 |
| Thin pages that say little beyond their title | 21 |
| Fill-in-the-blank titles across a set of pages | 20 |
| Pages with almost no text of their own | 18 |
| A sitemap listing redirects, broken pages or nothing usable | 13 |
| Titles with a blank or a code where text should be | 6 |
| The same page published at two or more addresses | 6 |
| Copy that only appears after JavaScript runs | 5 |
Every count is a site where the problem was verified on the live pages: each page a finding named was fetched again, independently of the scan, rendered in a real browser where it mattered, and compared before it was counted. Where the scan was wrong, it is not counted here; what it got wrong is at the end.
Mistake 1: Building pages from one template and swapping the name
Tied for the most common mistake, on 23 sites: a set of pages for each town, service, route or product line, written once and filled in many times. Two of them side by side read the same apart from the place name. We saw it on moving companies with a page for every pair of cities, on local service businesses with a page per suburb, and on directories with a page per town whose text differed only in the town’s name and its numbers.
This is the shape Google’s spam policies describe as doorway pages: many similar pages built to rank for slight variations of one search. One such page is harmless. A hundred of them is the pattern the policy names, and it is judged across the site, not page by page.
Check yours: open two pages from the same set — two towns, two services — and read them side by side. If everything but the name is the same, give each page something only it can say: local prices, photos, staff, real questions from customers in that place. If there is nothing to add, one stronger page beats twenty thin ones.
Mistake 2: Thin pages that say little beyond their title
On 21 sites a real share of the pages carried very little copy of their own: a heading, a sentence, a contact form. Category and tag archives, treatment and service pages, news posts that were never finished. Each one looks fine in a menu. Together they tell a search engine that much of the site has little to offer.
Check yours: list your shortest pages and ask of each one what a visitor learns there that they could not learn elsewhere on your site. Expand the ones that matter, merge the ones that overlap, and remove the ones that exist only because a plugin made them.
Mistake 3: Reusing one title or description across pages
On 23 sites, different pages carried the same title tag or meta description: two different articles with one title, every page in a language section sharing one description, a terms page titled exactly like the home page. When titles repeat, the pages compete for the same search and Google often rewrites the title it shows.
A related mistake turned up on 20 sites: titles from a fill-in-the-blank pattern, where a set of pages differs by one word. On a set of location pages that is the template from mistake 1 showing in the metadata. On two articles about almost the same topic, it is a sign the two pages are competing with each other.
Check yours: search your site’s titles for repeats (a crawler export or Search Console’s page list both work). Every page should have a title that says what only that page does.
Mistake 4: Pages with almost no text of their own, still in the sitemap
On 18 sites we found pages whose body was empty or close to it: parent pages that only exist to hold child pages, glossary entries with a heading and no definition, empty collection pages, and on 2 sites test pages anyone could reach. A page like this is listed in the sitemap, so a search engine visits it, finds nothing, and draws a conclusion about the site.
On 6 sites the opposite fault: one page published at two or more addresses — an article at a second web address, a pricing page at two paths, a guide chapter in two sections, three ad landing pages with one text. Each copy competes with the other for the same search.
Check yours: open a sample of addresses from your own sitemap rather than from your menu. Any page with nothing on it should be filled, redirected to the page that replaced it, or removed. An old address of a moved page should redirect to the new one, not serve a second copy.
Mistake 5: A sitemap that sends crawlers somewhere else
13 sites had a sitemap problem: addresses that redirect to another page instead of the page itself, addresses that no longer load, a sitemap named in robots.txt that returns an error, and a sitemap whose pages sit on a different domain from the site. Each one costs a search engine a visit and tells it the sitemap cannot be trusted.
A close cousin, found on 1 site: every page’s canonical tag pointed at an old hosting address that no longer exists. A canonical tag tells search engines which address is the real one, and these told them to use a page that returns an error.
Check yours: open your sitemap in a browser, follow a handful of its addresses and confirm each one loads directly, without a redirect. Then view the source of one page and check that its canonical tag names that page on your own domain.
Mistake 6: Broken title templates
On 6 sites, at least one title had a blank where text should be, or a code the site’s software was meant to replace: a separator with nothing after it, an unfilled field name. Search results show exactly what the title says.
Check yours: search your page titles for curly braces, percent signs and doubled separators. They usually come from one template field left empty, so one fix repairs every page.
Mistake 7: Copy that only appears after JavaScript runs
On 5 sites some pages had almost no text in the HTML the server sends; the content arrives later from JavaScript, or the page really is empty, and from the HTML the two look the same. Google can render JavaScript, but later than it reads the HTML, and many other crawlers never do.
Check yours: view the page source (not the inspector) and search for a sentence from the page. If it is not there, the server is not sending it, and the page should be server-rendered or pre-rendered.
Mistake 8: Review-star markup that does not describe the page
On 2 sites every page, articles included, carried structured data calling the page a product with a star rating, so the stars could appear in search results. Google’s structured data rules require markup to describe the page’s own content, and review stars that a business gives itself have not been shown for local businesses since 2019. Markup that misdescribes a page is one of the things Google issues a manual action for.
Check yours: run a page through Google’s Rich Results Test and read what it says the page is. A law firm’s article is not a product.
Mistake 9: Bot protection that turns away more than bots
5 of the 50 sites refused our scanner outright, 3 of them while letting an ordinary browser in. Turning away an unknown scanner is a reasonable choice. The risk is protection configured by user agent rather than by verified crawler, which can turn away a search engine too, and the owner would never see it.
Check yours: use Search Console’s URL Inspection on a few pages and on your sitemap. If Google’s own fetch fails, the protection is set too tight.
Things that look like website mistakes but are not
11 sites had something a careless check would call a mistake. None of these needs fixing:
- Product variants. One item in four colours is four pages that are supposed to read alike, including their titles.
- Numbered archive pages. Page 2, 3 and 4 of a blog share its title by design.
- A series with similar titles. “A beginner’s guide to” two different dishes is a series, not a template, when the articles themselves are different.
- A repeated widget. The same block of customer reviews on every service page is shared furniture, not copied writing.
A report that flags these is wasting the owner’s time, and it is the fastest way to stop anyone believing the findings that do matter.
Website maintenance mistakes: a 15-minute monthly check
- Open five addresses from your sitemap, not your menu. Each should load directly, have real content and a title of its own.
- Read two pages from the same template side by side. If only a name differs, they need more than a name.
- Search your titles for repeats and for blanks or codes.
- View the source of one page: is the text there, and does the canonical tag name this page on this domain?
- Check Search Console’s page indexing report after any redesign, migration or plugin change. That is when most of the faults above were introduced.
The free site scan does the comparison part across many pages at once and lists the addresses behind every finding, so each one can be opened and checked. For why whole-site patterns matter to Google in the first place, see the history of the Panda update, the helpful content survival guide and how Google ranks pages. For the duplicate-page side in more depth, our earlier study of 30 websites shows the cases one by one.
What our own scanner got wrong
Verifying every finding is how this list was made, and it is also how we found our own mistakes. On the first pass, 32 of 196 findings (16%) were wrong: pages called unreachable that had only redirected or been slowed down by our own crawl, short pages called JavaScript-rendered, pages matched on their menus rather than their words, and a home page whose title is just the brand counted as a near-duplicate of every other title.
We fixed each one, re-scanned the same 50 websites and checked again. Two of the fixes broke something new, which the second check caught. On the final pass, 7 of 199 (under 4%) were wrong. That is not zero, which is why every finding names the pages behind it: a claim about your own website should be one you can check in a browser in a minute.
Frequently asked questions
What are the most common website mistakes?
On the 50 websites in this check, the most common were pages built from one template with only a name swapped, thin pages that say little beyond their title, titles and descriptions reused across pages, and pages with almost no text of their own still listed in the sitemap. None of them shows up when you look at one page at a time; they appear when pages are compared with each other.
Which website maintenance mistakes hurt search traffic the most?
The ones a site owner never sees: an old address that still serves a page, a sitemap full of redirects, a canonical tag pointing at a domain that no longer exists, test pages left public. Each is invisible from the home page and obvious to a crawler that reads every address the site lists.
How do I check my website for these mistakes?
Open two pages that should be different and compare what loads, read two pages made from the same template side by side, and open a few addresses from your own sitemap to see what they return. A free site scan does the comparison across many pages at once and names the addresses behind every finding, so each one can be checked in a browser.
Are product colour variants a website mistake?
No. A shop that sells one item in four colours has four pages that are supposed to read alike, and the same goes for numbered archive pages and a blog series with similar titles. They are not what a search engine’s spam systems look for, and a report that flags them is wasting your time.
Does bot protection stop Google from crawling my website?
It can, if it is configured by user agent rather than by verified crawler. Five of the 50 websites here turned our scanner away, three of them while letting an ordinary browser in. That is their choice, but it is worth confirming in Search Console that Google can still fetch your pages and your sitemap.
How often should I check a website for these mistakes?
After every batch of new pages and after any redesign or migration, and once a month otherwise. Most of the faults in this check were left behind by a change, not built in on purpose: a renamed page, a plugin that adds archive pages, a template field left empty.