Browse documentation

Feature guide · 4 minute read

Review duplicate body content

Matching body content is a review signal, not an automatic instruction to delete a page. Use duplicate content groups to compare the URLs together, understand why each exists, and choose a consolidation decision that matches search intent and site ownership.

Best for migrations, faceted URLs, location pages, and copied templatesResult: one decision for every matching group

12-second workflow

Isolate the matches, then read the group.

Silent screen recording

Recorded with deliberate local fixtures that contain no customer data. The matching theme loads on demand; the alternate recording is not fetched.

What happens in the demo

Start with the completed crawl

The Pages table contains the complete public fixture crawl. Content shows two exact body-content matches.

Open Duplicate body content

The issue reduces the table to the two matching URLs while keeping the issue panel open as context.

Read and sort the group

Move horizontally to Canonical URL and Duplicate Content Group, then sort the group column so matching pages stay together.

Start with the group

A shared hash tells you which pages deserve comparison.

SEO Crawler compares the effective body-content hash of successful HTML pages. When two or more URLs share that value, the app assigns them the same Duplicate Content Group number. Matching titles or similar word counts alone do not create a group.

Keep the URL and duplicate group visible together.

ColumnQuestion it answers
URLWhich routes share the body, and do their paths imply different purposes?
Duplicate Content GroupWhich rows belong to the same exact-match comparison set?
IndexabilityCan each page currently enter search results?
Canonical URLDoes each page declare the same preferred indexing target?
Status CodeIs a redirect or error already changing the intended consolidation?

Review one complete group at a time.

  1. 1.

    Open Duplicate body content

    In the Issues panel, expand Content and select Duplicate body content. The table keeps only pages that share an effective content hash.

  2. 2.

    Show the decision columns

    Keep URL, Status Code, Indexability, Canonical URL, and Duplicate Content Group visible. Hide unrelated fields before moving horizontally.

  3. 3.

    Sort by Duplicate Content Group

    Sorting brings every member of a group together. Do not compare one isolated row without reading its matching URLs.

  4. 4.

    Inspect page purpose

    Open or review each route. Compare intent, internal links, sitemap membership, canonical declarations, and ownership before choosing an action.

Choose the action from the pages' intended role.

Keep both and differentiate

Each URL serves a distinct intent, audience, location, product, or stage of a journey.

Rewrite the shared main content so each page answers its own purpose completely.

Consolidate and redirect

The pages compete for the same intent and only one URL needs to remain accessible.

Choose the durable destination, move useful content, and redirect the redundant URL.

Keep both with a canonical

Both URLs must remain available but one version should be the preferred indexing target.

Confirm internal links, sitemap entries, and canonical declarations support the same URL.

Exclude an intentional duplicate

The URL is a print view, parameter variant, test route, or another deliberately non-indexable copy.

Verify the exclusion is intentional and that crawl paths do not create unnecessary duplication.

Recrawl and confirm the group changed for the intended reason.

After publishing the correction, clear the previous crawl and run it again. A page rewritten to serve a distinct purpose should leave the duplicate group. A redirected URL should report the intended destination. Canonicalized variants may remain visually identical, but their indexability, canonical target, internal links, and sitemap signals should agree.

Record the group number, affected URLs, chosen primary page, action, and post-release evidence in the handoff. Group numbers are crawl-local identifiers, so URLs remain the durable reference.