Agree on Visual Defects Before Inspecting the Parts Lot
Share
A visual inspection can produce a precise defect count and still be unreliable. One inspector may call a mark a scratch, another may call it handling evidence, and the same inspector may classify it differently on a second viewing. Counting lot defects before resolving those disagreements turns ambiguity into data without fixing the inspection system.
An attribute agreement study tests whether human classifications are consistent and aligned with an authorized reference. It begins with agreed defect categories, controlled sample presentation and repeated blind ratings. Field symptoms and failure mechanisms remain with the existing track-roller failure-analysis context. This article does not invent cosmetic limits, set an acceptance-sampling plan or authorize a lot decision.
Define the defect categories and reference decisions
Start with the part, surface and inspection zone. A track roller’s machined seat, seal-adjacent area, painted body and nonfunctional casting surface may have different requirements. Name the defect type in observable terms and connect it to the controlled drawing, visual standard or customer specification. “Bad scratch” is not an operational category because it leaves type, boundary and acceptance undefined.
Choose the category structure that fits the decision. A binary rating might distinguish acceptable from unacceptable. Nominal categories can separate unrelated types such as scratch, dent, coating void and contamination without ranking them. Ordinal categories place states in an ordered sequence, such as defined severity levels. Do not analyze an ordered scale as though its categories were unrelated, or imply an order where the standard provides none.
Define the observable boundary for each category. Record size or area criteria only when they come from the approved requirement. Include location, permitted zones, viewing distance, angle, lighting, surface cleanliness and any scale or magnification. If the inspection depends on color or texture, define the controlled reference and environment. Do not enhance images or alter contrast unless that presentation has been validated as equivalent to the real inspection.
Build the catalogue with acceptable, unacceptable and borderline examples from authorized decisions. Show the feature in context as well as close detail where needed. Record part revision, surface condition, scale, lighting and the catalogue version. Generic internet photographs should not become acceptance standards for actual undercarriage parts.
Assign a reference decision to each study sample before or through an authorized expert process. Record who owns that decision and what method or evidence supports it. Majority vote among participating inspectors cannot create the reference because it would compare the system with itself. If the reference owner cannot classify a sample confidently, mark it borderline, unknown or unratable under the approved route rather than forcing a clean label.
Minitab’s attribute agreement overview describes studies of subjective nominal or ordinal ratings and separates within-appraiser, between-appraiser and known-standard comparisons. It does not establish universal defect limits or guarantee that the designated standard is technically correct.
Freeze the catalogue revision and sample references for the study. If the instruction changes during testing, stop or clearly separate the data collected under each version. Otherwise, an apparent improvement may reflect a different rule rather than better inspector agreement.
Control how samples are presented
Select samples that represent the process and decision range, not only obvious good and bad parts. Include common conditions, borderline examples near approved category boundaries, rare but consequential defect types, and normal variation in surface, geometry or finishing. Record why each sample is included and which categories it tests.
A few memorable samples repeated many times can make consistency look better than it is. Use enough distinct items to cover the actual process range, with sample quantity determined by the study method, number of categories, expected prevalence and customer procedure. There is no universal count that makes every visual system adequate.
Preserve sample identity while hiding the reference and prior ratings from the appraisers. Use neutral study IDs and retain the controlled map separately. Present samples in randomized order for each trial so inspectors cannot infer the expected answer from sequence. If physical orientation matters, define whether it is controlled or intentionally varied.
Include appraisers who represent the people, shifts and sites that will perform the live inspection. Repeated trials are needed to evaluate whether each person classifies the same samples consistently. Allow a suitable interval or washout between repeats and change the order. Immediate repetition of a distinctive sample may test memory more than inspection repeatability.
Keep training separate from the formal study. During training, appraisers can discuss the catalogue, ask questions and receive feedback. During the study, they should rate independently without seeing the reference, another appraiser’s decision or their own previous result. If coaching occurs, end the current study phase and document the retraining before collecting a fresh comparison.
Control lighting, viewing distance, time allowance, cleanliness and other presentation factors stated by the inspection method. Record deviations. If the live process uses physical parts, a study based only on photographs needs evidence that the images preserve the relevant visual information. Camera exposure, compression, screen calibration and lack of depth can change the decision.
Record sample damage or change between trials. Cleaning, handling or repeated examination can alter a surface. If a feature changes, retire or re-reference that sample rather than treating inconsistent ratings as entirely an inspector problem.
The study plan should therefore identify the catalogue revision, samples and references, process coverage, appraisers, trial count, randomization method, blind controls, presentation conditions and exception handling. Missing any of those fields can limit what the resulting percentages mean.
Measure three kinds of agreement
Within-appraiser agreement asks whether each inspector gives the same sample the same rating over repeated trials. It reveals repeatability of the human classification under the study conditions. High within-appraiser consistency means a person is stable; it does not show that the person agrees with other inspectors or the authorized standard.
Between-appraiser agreement compares inspectors with one another. It shows whether the visual system produces reproducible classifications across people. A group can agree strongly because everyone applies the same mistaken interpretation. That is why between-appraiser agreement must remain separate from agreement to the standard.
Each-appraiser-versus-standard results show how often each person matches the authorized reference. An all-appraisers-versus-standard measure asks whether all participants agree with the reference for the same items. Minitab’s key-results guidance presents consistency, correctness, between-appraiser and all-versus-standard outputs separately in that software context. A buyer should preserve those distinctions even when using another tool.
Report percent agreement with its denominator and category scope. Overall agreement can hide a weak defect category when most samples are easy acceptable parts. Break results out by category, appraiser, trial and relevant part or surface zone. Show how many reference-acceptable and reference-unacceptable cases were present.
Record misclassification direction. A false accept occurs when an appraiser classifies a reference-unacceptable sample as acceptable; a false reject is the reverse. For nominal or ordinal categories, use the exact confusion pair, such as calling a dent a scratch or assigning an adjacent severity class. The operational consequence depends on the actual requirement and should not be invented by the study author.
Agreement estimates have uncertainty. Include the applicable confidence intervals and the method used. Small or imbalanced samples can produce wide intervals even when the observed percentage looks high. Minitab’s attribute agreement graph guidance distinguishes within-appraiser and appraiser-versus-standard views and their intervals. The graphs do not design the study or decide acceptance by themselves.
Chance-adjusted statistics such as kappa can add context when category prevalence and the method make them appropriate. Record which statistic was used, its assumptions and the category structure. Do not impose one universal kappa threshold or use a single statistic to hide specific disagreements. For ordinal ratings, use methods that preserve order when the approved analysis calls for them.
AIAG’s Measurement Systems Analysis manual overview frames measurement-system assessment as a basis for improvement. Its public description does not provide a free universal method or threshold. The customer procedure and actual study plan control the evaluation.
Review disagreements and revise the instructions
Move from summary statistics to a sample-by-appraiser log. For every disagreement, preserve the sample ID, reference category, each appraiser’s rating by trial, defect zone, presentation condition and catalogue revision. Classify the pattern as within-appraiser inconsistency, between-appraiser disagreement, disagreement with the reference, or a combination.
Look for concentrated confusion. Several inspectors may disagree only on scratches near an edge, or one person may change ratings when lighting shifts. A reference may be disputed because the image lacks scale or because the catalogue uses two overlapping definitions. These patterns point to specific instruction or presentation changes.
Do not automatically treat the authorized reference as infallible. If evidence challenges it, send the sample to the named reference owner or technical authority. Record the reconsideration and new basis. Inspectors should not change the reference by majority vote, and the analyst should not conceal uncertainty to improve agreement.
Revise the smallest element that addresses the evidence: clarify a category boundary, add a controlled example, define a zone, improve lighting instructions, separate two categories, add an unknown route or retrain on a specific confusion pair. Update the catalogue revision and preserve the superseded version.
Run a fresh blinded, randomized study after material changes. Include both the problem samples and new representative items so inspectors cannot pass through memory. Use the customer’s acceptance criteria and uncertainty requirements rather than a universal threshold.
| Sample and reference | Appraiser and trial pattern | Misclassification or evidence gap | Revision and status |
|---|---|---|---|
| Representative part, approved reference category and catalogue version recorded | Each appraiser repeats consistently; appraisers agree together and with the reference | Category-specific results and confidence intervals meet the customer criteria | Normal: study evidence supports the defined visual system for authorized review |
| Sample IDs exist, but reference owner and borderline coverage are absent | Only one trial was performed | Within-appraiser consistency cannot be evaluated | Missing: establish authorized references and repeated representative presentation |
| Reference says unacceptable scratch; all inspectors rate acceptable | Inspectors agree strongly with one another but not with the standard | Consistent false-accept direction or disputed reference requires investigation | Conflict: reference owner reviews basis; agreement is not treated as correctness |
| Borderline surface condition is unratable under the current catalogue | Ratings change across trials and lighting conditions | Instruction and presentation ambiguity are both plausible | Unresolved: revise method, retrain and repeat with controlled conditions |
Failed agreement should pause reliance on the visual classification system under the responsible quality procedure. It does not automatically reject the inspected product lot. Lot sampling, acceptance and disposition remain separate decisions made by authorized owners.
The useful outcome is a better inspection method, not merely a higher summary number. Define observable categories, present representative samples without cues, keep repeatability, reproducibility and reference agreement separate, and revise instructions from specific disagreement evidence before counting defects in production.