Guide
Missing versus not applicable in a bearing catalog
A field status model that separates unknown, not applicable, conflicting and not-yet-researched values, so blank cells stop distorting search, filters and completeness scores.
A blank cell in a bearing export looks like one problem. In practice it hides at least five different situations, and each one needs a different next action. Treating them all as “missing” inflates research backlogs, produces misleading completeness scores and, worst of all, invites people to fill fields that should stay empty.
This guide describes a small status model that makes the difference explicit, shows how to audit an existing catalog against it and explains how it changes the way completeness is measured.
Why a blank is ambiguous
Consider a contact angle column in a flat category export that mixes several bearing families. A blank in that column could mean any of the following:
- The attribute applies to the product, but nobody has found a value yet.
- The attribute does not apply to this product type at all, so there is nothing to find.
- Two sources disagree, and someone cleared the cell rather than pick one.
- Nobody has looked; the row was imported with only a designation and a description.
- The value exists in the source system but was dropped during export because the column did not exist for that category.
A merchandiser filtering by contact angle sees the same empty cell in every case. A data engineer computing completeness counts it the same way. A research team asked to “fill the gaps” will spend time on rows where there is no gap to fill.
A five-state field status model
Store a status next to every attribute value. The MyBearings draft record envelope uses these five states:
| Status | Meaning | Value field | Typical next action |
|---|---|---|---|
populated | A supported value exists | Required, with at least one assertion | Review if unverified |
unknown | Applies to the product, value not established | Null | Research with an appropriate source |
not_applicable | Does not apply to this product type or configuration | Null | None; exclude from completeness |
conflicting | Two or more supported values disagree | Null until a decision is recorded | Review sources and decide |
not_researched | No one has checked yet | Null | Prioritise for research or triage |
Two properties make the model useful rather than decorative.
First, the status determines what the value may contain. A populated attribute must carry a value and at least one assertion explaining where it came from. Every other status carries a null value. That rule is enforceable: the draft schema expresses it with if/then conditions, so a validator rejects a record that says “not applicable” and still holds a number.
Second, not_applicable is coupled to applicability. An attribute cannot be not_applicable while its applicability is marked applicable. Applicability is a property of the product type and configuration, decided by a documented rule, not by whoever happens to edit the row.
Unknown versus not researched
These two statuses are often merged, and the merge hides useful information. unknown means someone looked at the available evidence and could not establish the value. not_researched means nobody has looked.
The distinction matters for planning. A thousand not_researched fields might be resolved quickly from one manufacturer document. A thousand unknown fields have already absorbed research effort, and a different source or a direct manufacturer inquiry may be needed. Reporting both as “missing” makes the backlog look uniform when it is not.
Conflicting values are information
When two sources disagree, the tempting shortcut is to keep the first value or clear the cell. Both lose information. A conflicting status keeps every supported assertion, each with its source and location, until a review policy selects an accepted value. The history of that decision stays attached to the record.
Conflicts are also a quality signal. A cluster of conflicts on one attribute across one supplier’s rows usually points to a mapping problem, such as a column carrying a different dimension than its label suggests, rather than to many independent errors.
Auditing an existing catalog
You can apply the model to a catalog that never recorded status. Work through it in this order.
- Split by family. Applicability depends on bearing type. Profile each family separately rather than the whole export at once.
- Decide applicability per attribute and family. Write the rule down: “contact angle applies to angular-contact families; not applicable to the selected radial deep-groove family in this catalog”. Where you cannot decide, mark applicability
unknownrather than guessing. - Classify existing blanks. Blanks in non-applicable attributes become
not_applicable. Blanks in applicable attributes becomenot_researchedunless you have a record that someone tried, in which case they becomeunknown. - Look for placeholder values. Strings such as ”-”, “n/a”, “0” or “see datasheet” are often blanks in disguise. Classify them explicitly rather than treating them as values.
- Record conflicts. Where your inputs carry more than one value for the same product and attribute, keep both and mark the attribute
conflicting.
The example below shows what the audit produces for a few rows. Illustrative, synthetic values.
| Row | Designation | Family | Attribute | Raw cell | Status after audit | Reason |
|---|---|---|---|---|---|---|
| 14 | DEMO-101-OPEN | demo-radial | contact_angle | (blank) | not_applicable | Family rule: attribute not used for this family |
| 15 | DEMO-220-AC | demo-angular | contact_angle | (blank) | not_researched | Applies; no research recorded |
| 16 | DEMO-220-AC | demo-angular | total_width | “see datasheet” | not_researched | Placeholder text, not a value |
| 17 | DEMO-310-SEAL | demo-radial | total_width | 18 / 19 | conflicting | Two supplier files disagree |
| 18 | DEMO-412-TH | demo-thrust | ring_width | (blank) | not_applicable | Thrust family specified by height |
How the denominator changes
Completeness is a ratio, and the denominator decides whether the number means anything. A naive calculation divides populated cells by all cells in the export. That counts every non-applicable column in every row as a failure, so a catalog with many families looks far worse than it is, and a single-family catalog looks better than a mixed one for no real reason.
A defensible calculation uses applicable fields only:
- Numerator: attributes with status
populated. - Denominator: attributes whose applicability is
applicable. - Reported separately: attributes whose applicability is
unknown, because they cannot honestly be placed on either side.
Report conflicts and unknowns alongside the ratio instead of folding them in. A completeness figure without its denominator and status breakdown cannot be compared across catalogs or over time. The companion guide on measuring catalog completeness honestly goes further into sampling and accuracy.
Practical consequences for search and filters
Once status exists, the storefront can behave sensibly:
- Filters can hide products only where an attribute is
not_applicable, rather than treating every blank as a non-match. - Product pages can show “not specified by source” for
unknownvalues instead of a blank or a dash. - Comparison tables can flag
conflictingattributes rather than present one supplier’s value as fact.
Apply it
- Run your export through the local catalog checker to find blank, placeholder and invalid required fields before classifying them. It reports input completeness only; it does not decide applicability.
- Inspect the status rules in the draft bearing record schema.
- Read how MyBearings defines coverage and completeness on the quality and methodology page.
- If you want help designing family-specific applicability rules for your catalog, discuss your catalog.
Sources
- JSON Schema 2020-12 — Specification used by the MyBearings draft record envelope to express status-dependent value rules.
Methodology article using synthetic examples only; it states no manufacturer specifications. Independent technical review: not yet assigned. Found an error? Tell us.
