OPEXSTUDIO
← All lessons and training aids OPEX learning notes / Measurement systems

Check inspection agreement and escape risk

Establish the reference classifications and difficult items.. Follow the visual, practise a decision, then check your thinking.

Fictional teaching examples and AI-generated illustrations. Proposed changes and goals are not achieved results. Use the written instructions and check local conditions before applying a method.

Download this exact reviewed edition ↗

Teaching view 1 of 2

Check inspection agreement and escape risk

Two-by-two inspection matrix: of 20 reference-bad items,18 are rejected and 2 accepted; of 20 reference-good items,1 is rejected and 19 accepted. Overall agreement is 92.5%, false acceptance among bad items 10%, and false rejection among good items 5%. One pass does not establish repeatability.
Original OPEX teaching diagram. Follow the steps below, then try the practice question. View full size ↗

Overall agreement can conceal the error that matters most. Establish a defensible reference classification and include representative, difficult items. A full attribute study uses blinded repeats and multiple appraisers to examine consistency within people, between people and against the reference. The single matrix here illustrates denominator choices only. False acceptance asks what fraction of reference-bad items escaped; false rejection asks what fraction of reference-good items were rejected. Both differ from overall agreement. Report the counts and uncertainty rather than declaring success from one percentage. A reference itself can be uncertain, so unresolved classifications need an explicit adjudication method.

Follow the method

  1. Reference bad
  2. Reference good
  3. Classified reject
  4. Classified accept
  5. Overall agreement: 37/40 = 92.5%
  6. False acceptance among bad: 2/20 = 10%
  7. False rejection among good: 1/20 = 5%

Read the example carefully

Reference bad20:18reject,2accept; reference good20:1reject,19accept.

Agreement37/40=92.5%; false acceptance2/20=10%; false rejection1/20=5%.

A single matrix does not establish within-appraiser repeatability.

Teaching view 2 of 2

Read agreement and both error directions separately

Completed record derives 37/40 agreement,2/20 false acceptance and 1/20 false rejection while leaving repeatability unproven.
Original OPEX teaching diagram. Follow the steps below, then try the practice question. View full size ↗

Fictional case: inspector Leo assesses 40 reference items. Of 20 reference-bad items,18 are rejected and 2 accepted. Of 20 reference-good items,19 are accepted and 1 rejected. The core matrix retains these exact counts. The team calls 92.5% agreement proof that inspection is repeatable.

Follow the method

  1. Overall agreement
  2. False acceptance among bad
  3. False rejection among good
  4. Repeatability

Read the example carefully

Resolve or explicitly classify the reference uncertainty before using that item to score an appraiser’s correctness. Preserve the disagreement as useful method evidence.

A forced reference label could turn legitimate ambiguity into an apparent operator failure. This is a method-definition issue as well as a training question.

Apply the method

High agreement can hide an important error direction

Read a reference-classification matrix, distinguish overall agreement from the two error directions, and recognize the limits of a single pass.

Fictional case: inspector Leo assesses 40 reference items. Of 20 reference-bad items,18 are rejected and 2 accepted. Of 20 reference-good items,19 are accepted and 1 rejected. The core matrix retains these exact counts. The team calls 92.5% agreement proof that inspection is repeatable.

Role: Inspection-method owner and appraiser coach

Normal condition

Reference status is trustworthy for the defined characteristic; item identity and decision are retained; repeated and between-appraiser comparisons are designed separately.

The gap

One aggregate percentage hides false acceptance and false rejection, and one pass has no repeated classifications from which to judge repeatability.

  • This is an illustrative single pass, not a completed attribute measurement-system qualification.
  • The balanced 20/20 reference sample is not a production prevalence estimate.
Supplied case inputs
Reference classRejectedAccepted
Bad182
Good119
  1. Verify reference and observation identity

    Leo’s coach confirms the definition of good/bad and how the reference decisions were established. Each item keeps one identity linked to Leo’s classification.

    Why: Agreement against a questionable reference does not establish correctness. The comparison must concern the same characteristic and decision rule.

    Evidence: Forty item records reconcile to the two reference groups.

  2. Read both correct cells

    Correct classifications are 18 bad rejected plus 19 good accepted, totaling 37. Overall agreement is 37/40=92.5%.

    Why: Agreement counts matches to the reference across both classes. It says nothing yet about consistency on a repeated blind pass.

    Evidence: The reported numerator is 37, not merely the count of accepted items.

  3. Separate false acceptance

    Two of the 20 reference-bad items were accepted, giving 10% false acceptance among bad items in this study.

    Why: The denominator is the reference-bad group. Dividing by all 40 would answer a different question and conceal the conditional risk direction.

    Evidence: The record identifies the two item IDs for review without claiming a production escape rate.

  4. Separate false rejection

    One of 20 reference-good items was rejected, giving 5% false rejection among good items. The coach preserves this distinction from false acceptance.

    Why: Both errors matter but can have different consequences. A single combined error percentage can obscure what instruction or decision boundary needs attention.

    Evidence: The review retains separate error directions and their respective denominators.

  5. Design the next evidence

    The coach plans blinded repeat classifications and, where appropriate, other appraisers using representative cases including difficult boundaries. Prior answers are not shown during repeat assessment.

    Why: Within-appraiser consistency, between-appraiser agreement and reference agreement are different questions. Repetition must be designed to measure them rather than rehearse remembered labels.

    Evidence: The next study plan states the comparison and keeps its results blank until performed.

Completed classification interpretation record
MeasureCalculationWhat it establishes here
Overall agreement37/40=92.5%Single-pass reference matches
False acceptance among bad2/20=10%Error direction in selected bad group
False rejection among good1/20=5%Error direction in selected good group
RepeatabilityNo repeated pass suppliedNot established

Reference status is disputed

One item marked reference-bad has conflicting expert assessments and no resolved decision rule.

Resolve or explicitly classify the reference uncertainty before using that item to score an appraiser’s correctness. Preserve the disagreement as useful method evidence.

A forced reference label could turn legitimate ambiguity into an apparent operator failure. This is a method-definition issue as well as a training question.

The study record states reference uncertainty and the responsible adjudication process.

A new matrix with different class sizes

New fictional pass:30 reference-bad items yield 27 rejected and 3 accepted;10 reference-good items yield 8 accepted and 2 rejected.

Changed practice inputs
Reference classRejectedAccepted
Bad273
Good28

Your task

  1. Calculate overall agreement and each conditional error rate.
  2. Explain why overall agreement alone hides the poorer good-item result.
  3. State what remains unknown about repeatability.

Prepare your worksheet

  • Reference group counts
  • Correct cells
  • False-accept calculation
  • False-reject calculation
  • Next study question
Reveal the answer and reasoning

Overall agreement=(27+8)/40=87.5%. False acceptance among bad is 3/30=10%; false rejection among good is 2/10=20%.

The unequal reference groups weight the overall result differently. A single pass still cannot establish within-appraiser repeatability; blinded repeated classifications are needed for that question.

Worked answer record
MetricNumerator / denominatorResult
Agreement35/4087.5%
False acceptance3/3010%
False rejection2/1020%

Check these interpretations

  • 92.5% single-pass agreement is not repeatability.
  • The chosen reference mix is not automatically the production mix.

Check your work

  • Use the correct reference-group denominators.
  • Explain each error direction.
  • Keep reference uncertainty and repeated-study design visible.

Run a practice session

Materials

  • Reference/classification cards
  • Blank matrix
  • Separate answer sheet
  1. Define a matrix cell · 5 minutes

    Which items belong in false acceptance?

  2. Compare three percentages · 8 minutes

    What denominator changes?

  3. Work unequal groups · 10 minutes

    Which error is hidden by one total?

  4. Debrief next evidence · 5 minutes

    How would a repeat be kept blind?

Debrief

  • Ask learners to explain the consequence of each error without inventing a universal threshold.
  • Treat disputed reference examples as method evidence, not automatic appraiser blame.

Color the two matching cells, then calculate each conditional rate from its own reference row.

Transfer into the work

Owner: Inspection-method owner

Record: Reference basis, item-level classifications, agreement/error summaries and action record

Review: After instruction changes and during planned measurement review

Evidence: Representative blind comparisons and resolved decision criteria

Clarify references or method, coach targeted errors and verify with new blinded opportunities.

Build on reliable methods

Sources and further reading

  • ASQ: Attribute agreement analysis scope ↗

    Attribute agreement examines within-appraiser, between-appraiser and reference-standard agreement.

    Only the public scope is used; no ISO/TR14468 tables or examples are copied and current standard adoption is not inferred.
Free learning resources

Take the lesson into your team.

Read the lessons online or use these PDFs to prepare, practise and review with your team. No sign-in needed.

Facilitators and team leads

Facilitator guide

Case objectives, demonstration plans, debriefs, common mistakes and application checks across all 81 workplace cases and method lessons.

Download Facilitator guide PDF · 166 pages · 65.1 MB
Learners and improvement teams

Learner workbook

Printable case worksheets, blank observation records and five calculation exercises; answers are separate.

Download Learner workbook PDF · 169 pages · 10.7 MB
Learners after practice and facilitators

Answer key and coaching notes

Reasoned sample responses, worked calculations and coaching guidance; fictional examples are clearly labelled.

Download Answer key and coaching notes PDF · 105 pages · 8.5 MB
Practitioners and facilitators seeking detailed worked methods

Method and application reference

The native method mechanisms and worked applications for all 68 detailed lessons, in a separate bookmarked portrait reference.

Download Method and application reference PDF · 141 pages · 10.2 MB
Self-study learners and workshop groups

Illustrated systems atlas

Five illustrated system chapters: 15 Flare concept maps and 26 original workplace teaching cards, with links to all 81 supporting cases and method lessons.

Download Illustrated systems atlas PDF · 69 pages · 55.8 MB
Connect the methods

Use the next tool for the next question.

  • Adjudicated reference classifications
  • Representative items and blinded repeats
Explore all chapters and detailed lessons →