INSTRUMENT 06 · ATTRIBUTE AGREEMENT STUDY
How do you validate this in a language quality already accepts?
Design a study with known cases, two appraisers, two trials and an explicit ceiling on system performance. Think of it like calibrating a scale against a known weight: you can't call a reading "wrong" if you never established how consistent the reference itself is.
Example: two senior inspectors each grade the same 50 welds, twice. They agree with each other (and themselves) 80% of the time — people are inconsistent too. That 80% is the ceiling: if the AI system scores 85% against that same reference, the extra 5% is noise, not proof of superiority, because you never showed humans could do better than 80% in the first place.
ACCEPTANCE CRITERIA · SET IN ADVANCE
AGREEMENT CEILING
80%No system can score above 80% against a reference that is only 80% reproducible.
VERDICT
Restriction: welded structural packages, named suppliers, evaluated language, production approval support only
Cases: 50 Appraisers: 2 human + system Trials: 2 each Human agreement: 80% System effectiveness: 85% Achievable score ceiling: 80%