Guide
Measure bearing catalog completeness honestly
How to define denominators, applicable attributes and separate measures for completeness, source support and accuracy so catalog quality numbers mean something.
“Our catalog is 90% complete” is one of the most common and least informative statements in product data. It does not say 90% of what, which attributes count, whether the values have sources, or whether anyone checked that they are right. Decisions about enrichment budgets, supplier onboarding and storefront launches deserve better numbers.
This guide sets out a way to measure catalog quality that holds up when someone asks how the figure was calculated.
Three measures, not one
Separate three questions that a single percentage usually blurs together:
| Measure | Question | Numerator | Denominator |
|---|---|---|---|
| Completeness | Do we have a value where one should exist? | Populated applicable attributes | Applicable attributes |
| Source support | Can we show where each value came from? | Populated attributes with at least one source assertion | Populated attributes |
| Accuracy | Are the values correct? | Values confirmed by expert review in a sample | Values reviewed in the sample |
A catalog can score high on completeness and low on source support, for example after a bulk fill from undocumented spreadsheets. It can have strong source support and still contain errors if values were mapped to the wrong field. Reporting the three separately shows where effort is needed.
Define the attribute set first
Completeness is always relative to a list of attributes. Before calculating anything, write down:
- The families in scope. Measure per family; aggregates across families hide large differences.
- The attribute list per family. Which attributes the catalog intends to carry.
- Applicability rules. For each attribute and family, whether it applies, does not apply, or depends on configuration.
A common shortcut is to use “core attributes” chosen for a pilot. That is fine if the report says so. A figure measured against a selected list of core attributes is not a statement about every attribute a product could have.
Build the denominator from applicability
The denominator is the set of attributes that apply to each product. Attributes marked not_applicable leave the denominator entirely. Attributes whose applicability is unknown are reported separately, because placing them either side would bias the result.
Illustrative, synthetic values for one family:
| Count | Value |
|---|---|
| Products in family | 400 |
| Attributes defined for family | 12 |
| Product–attribute pairs | 4,800 |
| Pairs not applicable (configuration rules) | 600 |
| Pairs with applicability unknown | 200 |
| Applicable pairs (denominator) | 4,000 |
| Populated applicable pairs | 3,100 |
| Completeness (applicable) | 77.5% |
| Naive figure (populated ÷ all pairs) | 64.6% |
Both numbers come from the same data. Only one of them tells you how much work remains, and only if the 200 unknown-applicability pairs are reported next to it.
Report the status breakdown
Alongside the ratio, show how the non-populated applicable attributes break down:
not_researched: work not yet attempted.unknown: researched without a result.conflicting: values exist but disagree.
These have different costs. A backlog dominated by not_researched can often be reduced quickly. A backlog dominated by unknown may need different sources. Conflicts need review time, not research time.
Measure source support
For every populated value, check whether at least one assertion records a source, a location within it and a method. Report source support per attribute: some attributes are routinely well documented, others are often filled from memory or legacy files.
Source support does not establish correctness. It establishes that a reviewer could check the value.
Measure accuracy on a sample
Accuracy requires a reviewer comparing values against primary evidence. Practical guidance:
- Sample randomly within each family, and stratify by source if several suppliers are involved.
- Hold out the sample from rule development. Evaluating on the same records used to tune mapping rules overstates accuracy.
- Check identity and field mapping separately from value extraction. A correctly extracted value attached to the wrong product, or mapped to the wrong field, is still an error, and each kind has a different fix.
- Record the sample size and scope with the result. An accuracy figure from 50 values in one family is not a catalog-wide figure.
- Keep reviewer disagreements. If two reviewers disagree, the attribute definition may be unclear.
Avoid common distortions
- Counting placeholder text as populated. Strings such as “see datasheet” are not values.
- Counting defaults as evidence. A value filled by a default rule should be labelled as such, not counted as source-supported.
- Mixing product and listing scopes. Supplier SKUs and manufacturer products are different populations; report them separately.
- Treating confidence scores as accuracy. A method’s confidence estimate is not a measured probability of correctness until it has been calibrated against reviewed samples.
A reporting template
For each family, publish:
- Product count and attribute list version.
- Applicable completeness, with denominator and unknown-applicability count.
- Status breakdown of non-populated applicable attributes.
- Source support per attribute.
- Accuracy on a held-out sample, with sample size, date and reviewer role.
- Known limitations.
Apply it
- Get a structural input-completeness view of your file with the local catalog checker. It reports required-field completeness, not domain completeness, and says so.
- Read how MyBearings defines coverage, completeness and accuracy on the quality and methodology page.
- If you need a defensible baseline for a catalog project, discuss your catalog.
Methodology article using synthetic examples only; it states no manufacturer specifications. Independent technical review: not yet assigned. Found an error? Tell us.
