Guide

Missing versus not applicable in a bearing catalog

A field status model that separates unknown, not applicable, conflicting and not-yet-researched values, so blank cells stop distorting search, filters and completeness scores.

A blank cell in a bearing export looks like one problem. In practice it hides at least five different situations, and each one needs a different next action. Treating them all as “missing” inflates research backlogs, produces misleading completeness scores and, worst of all, invites people to fill fields that should stay empty.

This guide describes a small status model that makes the difference explicit, shows how to audit an existing catalog against it and explains how it changes the way completeness is measured.

Why a blank is ambiguous

Consider a contact angle column in a flat category export that mixes several bearing families. A blank in that column could mean any of the following:

  • The attribute applies to the product, but nobody has found a value yet.
  • The attribute does not apply to this product type at all, so there is nothing to find.
  • Two sources disagree, and someone cleared the cell rather than pick one.
  • Nobody has looked; the row was imported with only a designation and a description.
  • The value exists in the source system but was dropped during export because the column did not exist for that category.

A merchandiser filtering by contact angle sees the same empty cell in every case. A data engineer computing completeness counts it the same way. A research team asked to “fill the gaps” will spend time on rows where there is no gap to fill.

A five-state field status model

Store a status next to every attribute value. The MyBearings draft record envelope uses these five states:

StatusMeaningValue fieldTypical next action
populatedA supported value existsRequired, with at least one assertionReview if unverified
unknownApplies to the product, value not establishedNullResearch with an appropriate source
not_applicableDoes not apply to this product type or configurationNullNone; exclude from completeness
conflictingTwo or more supported values disagreeNull until a decision is recordedReview sources and decide
not_researchedNo one has checked yetNullPrioritise for research or triage

Two properties make the model useful rather than decorative.

First, the status determines what the value may contain. A populated attribute must carry a value and at least one assertion explaining where it came from. Every other status carries a null value. That rule is enforceable: the draft schema expresses it with if/then conditions, so a validator rejects a record that says “not applicable” and still holds a number.

Second, not_applicable is coupled to applicability. An attribute cannot be not_applicable while its applicability is marked applicable. Applicability is a property of the product type and configuration, decided by a documented rule, not by whoever happens to edit the row.

Unknown versus not researched

These two statuses are often merged, and the merge hides useful information. unknown means someone looked at the available evidence and could not establish the value. not_researched means nobody has looked.

The distinction matters for planning. A thousand not_researched fields might be resolved quickly from one manufacturer document. A thousand unknown fields have already absorbed research effort, and a different source or a direct manufacturer inquiry may be needed. Reporting both as “missing” makes the backlog look uniform when it is not.

Conflicting values are information

When two sources disagree, the tempting shortcut is to keep the first value or clear the cell. Both lose information. A conflicting status keeps every supported assertion, each with its source and location, until a review policy selects an accepted value. The history of that decision stays attached to the record.

Conflicts are also a quality signal. A cluster of conflicts on one attribute across one supplier’s rows usually points to a mapping problem, such as a column carrying a different dimension than its label suggests, rather than to many independent errors.

Auditing an existing catalog

You can apply the model to a catalog that never recorded status. Work through it in this order.

  1. Split by family. Applicability depends on bearing type. Profile each family separately rather than the whole export at once.
  2. Decide applicability per attribute and family. Write the rule down: “contact angle applies to angular-contact families; not applicable to the selected radial deep-groove family in this catalog”. Where you cannot decide, mark applicability unknown rather than guessing.
  3. Classify existing blanks. Blanks in non-applicable attributes become not_applicable. Blanks in applicable attributes become not_researched unless you have a record that someone tried, in which case they become unknown.
  4. Look for placeholder values. Strings such as ”-”, “n/a”, “0” or “see datasheet” are often blanks in disguise. Classify them explicitly rather than treating them as values.
  5. Record conflicts. Where your inputs carry more than one value for the same product and attribute, keep both and mark the attribute conflicting.

The example below shows what the audit produces for a few rows. Illustrative, synthetic values.

RowDesignationFamilyAttributeRaw cellStatus after auditReason
14DEMO-101-OPENdemo-radialcontact_angle(blank)not_applicableFamily rule: attribute not used for this family
15DEMO-220-ACdemo-angularcontact_angle(blank)not_researchedApplies; no research recorded
16DEMO-220-ACdemo-angulartotal_width“see datasheet”not_researchedPlaceholder text, not a value
17DEMO-310-SEALdemo-radialtotal_width18 / 19conflictingTwo supplier files disagree
18DEMO-412-THdemo-thrustring_width(blank)not_applicableThrust family specified by height

How the denominator changes

Completeness is a ratio, and the denominator decides whether the number means anything. A naive calculation divides populated cells by all cells in the export. That counts every non-applicable column in every row as a failure, so a catalog with many families looks far worse than it is, and a single-family catalog looks better than a mixed one for no real reason.

A defensible calculation uses applicable fields only:

  • Numerator: attributes with status populated.
  • Denominator: attributes whose applicability is applicable.
  • Reported separately: attributes whose applicability is unknown, because they cannot honestly be placed on either side.

Report conflicts and unknowns alongside the ratio instead of folding them in. A completeness figure without its denominator and status breakdown cannot be compared across catalogs or over time. The companion guide on measuring catalog completeness honestly goes further into sampling and accuracy.

Practical consequences for search and filters

Once status exists, the storefront can behave sensibly:

  • Filters can hide products only where an attribute is not_applicable, rather than treating every blank as a non-match.
  • Product pages can show “not specified by source” for unknown values instead of a blank or a dash.
  • Comparison tables can flag conflicting attributes rather than present one supplier’s value as fact.

Apply it

Sources

  1. JSON Schema 2020-12 — Specification used by the MyBearings draft record envelope to express status-dependent value rules.

Methodology article using synthetic examples only; it states no manufacturer specifications. Independent technical review: not yet assigned. Found an error? Tell us.

Apply it

Related resources

Bring one messy category.
Leave with a clear next step.

Share what your catalog looks like today. We will tell you what a scoped assessment would cover, what it would not, and what we would need from you.