Guide

Why duplicate products keep appearing, and how to stop it

Max Pechonis · Founder, Rebridge ·

Duplicates appear when a line on a vendor document is created as a new product instead of matched to something you already carry. Three causes account for nearly all of them: color codes that do not match, style numbers that change format between seasons, and imports that never check.

Nobody decides to create a duplicate

That is the part worth sitting with. Duplicates are not produced by carelessness, and not by someone choosing wrongly. They are produced by someone who had no way to see that a decision existed.

A line arrives on a vendor document. It matches nothing the search returns. The obvious, reasonable, defensible action is to create it. Everything about that moment says the product is new: unfamiliar code, unfamiliar formatting, nothing found. The catalog does not object, the purchase order reconciles against the invoice, and the day ends looking like a normal day.

Which is why telling people to be more careful changes nothing. The failure is informational, not behavioral, and the fix has to put the missing information in front of the person receiving.

Cause one: the color code does not match

The most common by a distance. A vendor ships a colorway as a code, your catalog holds it as a word, and nothing reconciles the two. Apparel meets it as a color code, a tennis supplier as a grip size abbreviation, an electrical wholesaler as a gauge notation. Same failure, different field.

A brand sends BWH. Your catalog says Black/White. Those are the same colorway and no string comparison will ever say so. The search finds nothing, so the line reads as new.

What makes this compounding rather than one-off is that codes are stable within a vendor and arbitrary across them. That brand will send BWH every season for years. Resolve it once and it never costs you again. Leave it unresolved and it produces a duplicate every single time that colorway ships.

Cause two: the identifier changed, so you are deciding blind

Many vendors encode the season into the style number, so the string legitimately changes each year. HYPRFRK-24-BLK becomes HYPRFRK-25-BLK, and a catalog that stored the full string last year now holds something that matches nothing arriving.

The identifier is not one field. It is a style root, a season and a color code travelling together, and only the first part identifies the style. Matching on the whole string works for exactly one season and then quietly stops.

A changed style number is not, on its own, evidence of a duplicate. It means you cannot tell what you are looking at, and creating without checking is the failure. Vendors also reissue genuinely new products each season under a near-identical style root, and filing one of those onto last year's record is the same mistake pointing the other way.

What settles it is the UPC. A UPC identifies a sellable unit globally, so two records carrying different UPCs are two products and creating the second is correct. The same UPC arriving under a new style number is the same product, and a second record for it is a duplicate. Where neither side carries a UPC, the question drops to a merchandising decision about whether you need to compare seasons, which the boundary section below covers.

This cause is seasonal and therefore invisible in the moment. Whatever it produces arrives in a batch at the start of a buying cycle, which is also the point in the year when nobody has time to check.

Cause three: nothing checked

The first two causes are near misses that could have matched. The third is a process that never attempted to.

A flat import reads a document and writes what it finds. It does not ask whether the style already exists, because asking requires a catalog lookup and a judgement about what counts as the same thing. Every line becomes a record. The import succeeds, the totals are correct, and the catalog gains one new product per line whether or not you already carry nine of them.

Worth separating from the other two, because it fails differently. Color and style-number mismatches produce occasional duplicates. An unchecked import produces them at exactly the rate documents arrive.

What it costs, in numbers you can check

Take a shop carrying forty styles, where each style picks up two new colorways a season across two seasons. That is one hundred and sixty of these decisions a year.

Suppose the process gets ninety percent of them right, which would be good going. Sixteen duplicates a year. Against a two thousand product catalog that is under one percent, and it sounds survivable.

It is not, because nothing removes them. Year three you are carrying forty-eight, year five eighty. And they do not distribute evenly: they cluster on the styles you reorder most often, which are your best sellers, which are precisely the numbers you most need to be true.

YearAt 90% accuracyAt 80% accuracy
11632
23264
34896
580160
Share of a 2,000-product catalog at year 54%8%

Every figure there is 160 decisions times a miss rate, accumulated, so substitute your own style count and hit rate and the arithmetic still holds. The shape does not change. The problem is not the size of any single mistake, it is that the count only ever moves in one direction. That is catalog bloat, and it is why a catalog that felt fine three years ago no longer does.

The three things that actually reduce the rate

Write the color codes down. One short list per vendor mapping their codes to your colorway values. This is the highest-value item on the page: it converts the most common cause from a judgement into a lookup, and the knowledge stays good for years.

Search the style root, not the identifier. Strip the season and color segments and search what remains. If nothing comes back, widen once to a distinctive word from the product name before concluding the style is new. A single naming inconsistency should never be enough to create a product.

Keep the manufacturer SKU populated. It is the only identifier that appears identically on the vendor document and in your catalog, so matching on it is exact rather than fuzzy. An empty field means every future document from that vendor falls back to comparing names, which is where the near misses come from.

None of this is difficult. All of it is the sort of thing that gets done diligently for three months and then stops, which is the honest reason duplicate rates are what they are.

The check, in the order that catches the most

Run these in order and stop at the first one that answers. The sequence matters: it goes from exact to fuzzy, so the cheapest reliable check runs first.

  1. Search the manufacturer SKU exactly. If the vendor prints one and you stored it at creation, this is the only comparison that is exact rather than approximate. A hit means the style exists. A clean miss is meaningful only if you know the field was populated.
  2. Strip the identifier to its style root and search that. Remove the season and color segments. A number reading HYPRFRK-24-BLK is a style root, a season and a color code travelling together, and only the first part identifies the style. An electrical part number carrying gauge and length behaves the same way.
  3. Search one distinctive word from the product name. Not the whole name, which will not match, and not a generic word like racquet or pendant, which matches everything. Pick the word only this family uses.
  4. Check the colorway against your vendor code list before concluding the color is new. Most apparent new colors are known colors under an unfamiliar code.
  5. If two records come back that look like the same style, stop. You already have a duplicate and are about to create a third. Reconcile the existing pair first: adding to the wrong one compounds the problem instead of fixing it.
  6. Only now create. If all five came back empty, the style is genuinely new.

Steps one to four take under a minute between them. Step five is the one people skip, and it is the one that turns a single duplicate into a cluster.

Telling a duplicate from a legitimate second record

Not everything that looks like a duplicate is one, and merging the wrong pair destroys information you cannot get back. Three cases that look identical in a product list and are not.

Two colorways of one style. If they sit as separate variants under one parent, that is the structure working correctly. If they sit as separate products, that is the duplicate. The test is whether they share a parent, not whether the names look alike.

The same style bought from two vendors. If you genuinely buy it both ways at different costs, two records may be the honest representation. Merging them makes the cost history meaningless.

Prior-season carryover. Same style, same color, new style number, a year apart. Whether that is one record or two depends on whether you need to compare seasons, which is a merchandising decision rather than a data-quality one. Either answer can be right. What is never right is having it settled by accident because a string did not match.

The practical rule: detection finds candidates, a human decides. Anything that merges automatically will be right about most pairs and wrong about a few, and the few are unrecoverable once sales history has been rewritten.

Prevention and cleanup are different problems

Worth being blunt: cleaning up existing duplicates while receiving still creates new ones is wasted effort. The catalog regains what you remove, and merging is far more expensive than creating. Every merge means reconciling two stock counts, two sales histories and any open purchase orders pointing at either record. That asymmetry is the whole argument for treating this as a receiving problem rather than a data-hygiene problem.

So fix receiving first, even partially, then triage what already exists by sales volume and merge the highest-value clusters by hand. The long tail can wait, and some of it should be left alone entirely.

Doing that resolution on every line, every time, is the systematic version of everything above. It is the work Rebridge does before anything reaches your POS: resolving the color code against your catalog's own values, decomposing the style number, and deciding whether a line extends something you already carry or is genuinely new.

Worked example

A worked example: the same colorway, three seasons running

You carry a boardshort in Black/White. The vendor ships it as BWH, their style number encodes the season, and the garment itself is unchanged year to year: same product, same UPC.

Season one. The line matches nothing, because your catalog holds Black/White and the document says BWH. A new product is created. You now hold two records for one product, and stock splits across them.

Season two. The style number has changed. It matches neither existing record. A third appears. Sell-through is now spread across three products, none of which describes the thing you actually sell.

Season three. Same again. Four records, one UPC. The buyer opens the reorder report, sees four mediocre products where there is really one good one, and orders accordingly.

The UPC was identical on all four documents, so a match on that field alone would have caught every one of them in the moment. Where a vendor prints no UPC, the fallback is cause one's color-code list: this vendor's BWH is your Black/White, written down once.

Now change one fact. Suppose the vendor had reworked the garment each season and issued a new UPC each time. Then those are four different products, four records are correct, and merging them would destroy exactly the history that tells you which version sold. The documents look almost identical in both cases. The UPC is what tells them apart.

The same thing, another trade

The same three causes in a wine list

A wine merchant carries a producer's Sauvignon Blanc. The supplier codes vintage into the product identifier, so the 2024 arrives with a code that matches nothing from the 2023.

Here the right answer is obvious in a way it rarely is in apparel. A 2024 and a 2023 are different wines with different UPCs, so two records are correct and nobody would merge them. That obviousness is worth borrowing: it shows that a changed identifier proves nothing on its own. The apparel case only feels different because the garment may genuinely be unchanged, and the UPC is doing the same work in both.

Cause one shows up in the same catalog as varietal abbreviations, where SB means Sauvignon Blanc on one supplier's list and Sancerre Blanc on another. Resolved per supplier it is a lookup; left unresolved it is a duplicate every season.

Nothing here is about wine. It is the same two failures a firearms dealer meets in caliber codes and a bike shop meets in model years.

Questions

How do I know whether my catalog already has duplicates?

Sort by product name and look for near-identical neighbors, then check whether the same style appears under more than one record. Duplicates cluster by vendor, so triage by vendor before working item by item.

Is it safe to bulk merge duplicates once I have found them?

No. Merging rewrites stock counts and sales history and can orphan open purchase orders. A script will be right about most pairs and wrong about a few, and the few are unrecoverable. Merge by hand, highest value first.

The vendor changed their style number format. Do I need to update my catalog?

Not usually. Store the style root rather than the full identifier and the seasonal suffix stops mattering. Store the UPC too, because that is the field that tells you whether next season's arrival is the same product or a new one. Rewriting the catalog every time a vendor changes format is the more expensive of the two options.

What is an acceptable duplicate rate?

There is no safe rate, because nothing removes them. A rate that sounds survivable at under one percent a year compounds to four percent in five years, and duplicates cluster on the styles you reorder most, which are your best sellers. The number to watch is the trend, not the level.

Which duplicates should I fix first?

The ones on styles you reorder most often, not the largest count. A duplicate on a style you buy twice a year distorts a decision twice a year; a duplicate on a discontinued line distorts nothing. Sort by units sold or reorder frequency, not alphabetically.

Do duplicates hurt anything if my stock counts are still right?

Your stock counts are not still right. Six units on hand read as four and two across two records, which is enough to trigger a reorder you do not need or suppress one you do. Even where the total is correct, sell-through, open-to-buy and reorder points all read per record.

Can I stop duplicates by making the SKU field unique?

No, and it can make things worse. A uniqueness constraint blocks the exact-duplicate case, which is the rare one, and does nothing about the common one where the same style arrives under a different identifier. It can also push someone into inventing a SKU to get past the error, which creates a record that matches nothing on any future document.

We import from a spreadsheet. Does that change the causes?

It changes the volume, not the causes. An import applies the same three failures a hundred rows at a time, and it removes the moment where a person might have noticed. If you import, the color-code list and the manufacturer SKU field matter more, not less, because nothing downstream will catch what the mapping missed.

Does this happen outside apparel?

Identically, and the encoded field is the tell. Wine suppliers encode vintage, firearms dealers caliber, bike shops model year, electrical wholesalers gauge and length, tennis suppliers grip size. Any identifier carrying something that changes each season stops matching on a yearly cycle. Note that several of those are genuinely different products rather than the same one relabelled: a 2024 vintage and a 2025 bike model year are new products with new UPCs. The failure is the same either way, because in both cases the identifier stopped matching and you had to decide without it.

Two records look identical but have different UPCs. Are they duplicates?

No. A UPC identifies a sellable unit globally, so different UPCs mean different products however alike the names look. This is common with seasonal reissues, where a vendor reworks a product slightly and assigns a new code. Merging them would destroy the history that tells you which version sold.

Should I delete a duplicate or merge it?

Merge, and only delete a record that has never transacted. Deleting a record with sales history removes the history, which is usually the thing you were trying to protect. If your POS has no merge, the safe move is to keep the record with the transaction history and retire the other from sale rather than removing it.