Glossary

Catalog bloat

Catalog bloat is the accumulation of records that should not exist: duplicates, orphaned variants, and styles split across several products. It grows quietly at receiving, and its cost is that the numbers you buy from stop describing what you actually sell.

It compounds, which is what makes it dangerous

Bloat does not arrive at once. Each unreconciled document line adds one record, and the arithmetic runs away from you quietly. A shop carrying forty styles, where each style picks up two new colourways a season across two seasons, faces one hundred and sixty of these decisions a year. Get them wrong and a two thousand product catalog gains eight percent duplicates annually, on top of last year's.

Substitute your own style count and season cadence; the shape of the result does not change. The problem is not the size of any one mistake, it is that nothing removes them.

What it actually costs

Inventory splits, so a style with six units on hand reads as four and two and triggers a reorder it does not need. Sell-through fragments, so a strong seller looks average twice. Counting takes longer because staff meet two records for one shelf position. And every downstream report inherits the error without flagging it.

Cleanup is harder than prevention, by a lot

Creating a duplicate takes one careless import. Removing one means reconciling two stock counts, two sales histories, and any open purchase orders pointing at either record. That asymmetry is the argument for fixing receiving before attempting any cleanup: otherwise the catalog regains what you remove.

It also argues against bulk operations. A merge script that runs across the whole catalog will be right about most pairs and wrong about a few, and the few are unrecoverable once sales history has been rewritten. Triage by sales volume, merge the highest-value clusters by hand, and leave the long tail alone until receiving has stopped producing new ones.

Questions

How do I know if my catalog is bloated?

Sort by product name and look for near-identical neighbours, then check whether the same style appears under more than one record. Duplicates usually cluster in one or two vendors.