Build a Supplier Scorecard from Parts Orders You Can Verify
Share
One late purchase-order line can produce two different delivery results. In a hypothetical period with 20 eligible lines and 200 eligible pieces, one late line containing four pieces gives 95% on time when the score counts lines, but 98% when it counts pieces. Neither calculation is useful until the scorecard states which event it measures and why.
A supplier scorecard should turn traceable order records into a repeatable review. Before calculating a result, define the purpose, numerator, denominator, time window, date basis, exclusions, sample size and action. Keep missing or conflicting inputs unresolved. A familiar metric name does not make two definitions comparable.
Decide what the scorecard is for
Write the decision the scorecard is meant to support. A buyer trying to reduce late receipts needs detail that reveals where delivery failed. A sourcing team reviewing continued supply may need a broader record of delivery, quality and issue response. Those purposes can use some of the same order data, but they do not automatically require the same measures, weights or action thresholds.
PCB's Supplier Scorecard Guidelines, PD2048 Rev B, describe a company-specific system used for joint problem solving, purchased-component management and supplier decisions. Its delivery and quality measures are combined using stated weights and a rolling 12-month period. That is a useful example of a documented system, not a universal formula for parts buyers.
Start with the few questions that will change an action. For example: Did eligible lines arrive against the agreed date? What proportion of inspected pieces produced verified defect records? Are the results supported by enough eligible observations to review? Keep each answer visible before considering a combined score. A weighted total can hide a serious quality event behind a strong delivery result, or exaggerate movement when the sample is small.
The same caution applies to targets and weights. Lapasar's guide to building and using supplier scorecards recommends agreeing definitions and setting weights around the relationship. Its example weighting is an illustration from that publisher. Your own weights and triggers require an approved business rationale, and two suppliers should only be compared where the definitions and eligible populations are genuinely comparable.
Define the delivery and quality metrics
Give every metric a definition version. The definition should identify the source records, counting unit, formula, eligible population and treatment of corrections. Save that version with each reporting period. If the definition changes, calculate a new series or clearly mark the break; do not join unlike periods into a smooth trend.
Requested and confirmed dates
The requested date records when the buyer wants the goods. The confirmed date records the date the supplier has acknowledged, subject to the buyer's order and change process. They answer different questions. Performance against the requested date shows whether the original need was met; performance against an accepted confirmed date shows whether the supplier met the recorded commitment. Label the measure accordingly.
For each purchase-order line, retain the original requested date, every proposed change, the party proposing it, the acceptance record and its timestamp. Freeze the scorecard's baseline according to a declared rule. A supplier's later promise should not silently replace an earlier commitment after a delay has already become visible. A buyer-approved schedule change may create a legitimate new baseline, but the original date and change evidence should remain recoverable.
Next define the arrival event. It might be receipt at the named facility, an accepted goods-receipt transaction, or another contract-specific milestone. Record the relevant time zone and how early, on-date and late receipts are treated. PCB's own formula uses total receipts, its promise-date field, a stated allowance and a preceding 12-month period. Those choices belong to that published system. They do not establish the date field, grace period or denominator for another buyer.
Partial deliveries need an explicit rule before scoring. If the metric is on-time-in-full by PO line, a line is successful only when the defined complete quantity reaches the defined arrival event by the applicable date. If the metric counts receipt events or pieces, it answers a different question. Keep the numerator and denominator in the same counting unit: on-time lines divided by eligible lines, or on-time pieces divided by eligible pieces.
Defect counts and inspected quantities
Define what the numerator counts. A defective piece counts units with at least one verified nonconformity. A defect occurrence counts individual nonconformities, so one piece may contribute more than one. A rejection or corrective-action record is another event again. Do not label these different numerators simply as “quality PPM” and compare them.
If the chosen measure is defect occurrences per million inspected pieces, write the formula as verified defect occurrences divided by inspected pieces, multiplied by 1,000,000. If the numerator instead counts defective pieces, name the result accordingly. The denominator should be the quantity actually covered by the defined inspection evidence, not the quantity ordered or received unless every eligible piece was inspected under that method.
Also define the cohort. A receipt-based window groups the lots received during the period and can attach later findings back to those lots. A discovery-based window counts issues opened during the period, including some from earlier receipts. Either can be useful, but combining them creates a moving denominator that is difficult to interpret. Record the reporting window and the reference event used to place each observation in it.
Always show sample size beside the rate. Zero verified defects among five inspected pieces and zero among 5,000 inspected pieces produce the same percentage but provide different amounts of evidence. A small or unknown sample should remain visible and should follow the predefined review rule; it should not be promoted to a pass because the calculated rate is zero.
Treat exceptions consistently
Write exception rules before looking at a supplier's result. Otherwise, the same event can be included when it improves one score and excluded when it improves another. Each exception needs a reason, supporting record, approver and treatment that can be applied to comparable orders.
| Metric | Numerator | Denominator | Time window | Reference date | Exclusions | Sample size | Action |
|---|---|---|---|---|---|---|---|
| On-time delivery by PO line | Eligible lines fully received by the accepted confirmed date | All eligible lines due in the window | 1 Apr–30 Jun 2026 | Confirmed date accepted before the defined freeze point | Buyer-requested postponements approved before the due date, with change record | 20 eligible lines | Open receipt-level evidence review if the buyer's approved trigger is reached |
| Verified defect occurrences per million inspected pieces | Verified defect occurrences linked to eligible inspected pieces | Pieces actually inspected under the stated method | Receipt lots dated 1 Apr–30 Jun 2026 | Receipt date assigns the lot; later verified findings return to that lot | Documented damage after buyer-controlled storage, if the review owner accepts causation | 200 inspected pieces from identified lots | Review the defect records, affected lots and containment when the approved trigger is reached |
A fillable supplier metric definition sheet provides the same required fields and three test cases. A complete, internally consistent row becomes ready for evidence review, not passed. A missing field, an explicit unknown or a conflict between the delivery numerator and denominator keeps the row on hold.
Apply these rules to common exceptions:
- Buyer-requested schedule changes: exclude or re-baseline only when the policy allows it and the request, acceptance and timing are documented. Preserve the earlier dates.
- Supplier-requested delays: record the proposal separately. Do not turn it into the comparison date merely because it appears in a later acknowledgement.
- Partial receipts: use the preselected line, receipt or piece rule. Do not switch units after seeing which produces the better result.
- Disputed defects: retain a pending state until the defined review resolves the defect record. Do not count the item as both accepted and defective in measures whose definitions make those states mutually exclusive.
- Small or missing samples: publish the eligible count and inspection count with the result. If the minimum evidence rule is absent or the inspected quantity is unknown, hold the assessment for review.
Corrections should leave an audit trail. Keep the original entry, the evidence for the change, the approver, date and revised value. If two systems disagree on a receipt date, quantity or defect status, do not choose the convenient record. Resolve which event each system captures and preserve the conflict until an authorized owner decides the scorecard input.
Use the score to start an evidence-based review
A score is a trigger for inspection of the underlying records. Zigaflow's overview of a supplier scorecard describes purchase orders and delivery notes as transaction records from which actual delivery and accuracy figures can be developed. The practical review starts by drilling from the metric into those order lines, receipts, inspection records, approved changes and open disputes.
When a result crosses your organization's approved trigger, first verify that the metric version, window and eligible population are correct. Then identify the few records driving the result. Check excluded events and unresolved conflicts, compare the current period with prior periods calculated on the same basis, and agree an owner, action, evidence requirement and review date. If the data is incomplete, the action is to repair or verify the record—not to label the supplier as good or poor.
Use combined scores cautiously. Before weighting delivery and quality, confirm that each component has a valid definition and adequate evidence. Record who approved the weights and targets, when they apply and what happens when a critical event requires review regardless of the total. PCB, Lapasar and other publishers show different systems because their purposes and relationships differ; their percentages and bands are not industry-wide requirements.
This process supports internal performance review. It does not produce a public ranking of real suppliers, and it does not establish KTSU's performance. For the broader task of choosing a source, see the undercarriage parts supplier overview. Once your definition sheet identifies the affected part, order, dates, quantities and unresolved evidence, you can send the part and order details for review. Keep the scorecard open until every exception has a supported disposition and the assigned action has an owner.