§8.2 · BUILDING · 5 MIN READ · EDITION 1.0
Building a golden reference set
A defensible reference set spans clean, difficult, borderline, incomplete-evidence and no-finding cases, each with an agreed ground truth.
A set built only from easy cases overstates performance; a set with no no-finding cases cannot reveal a tool that over-generates findings.
Ground truth should be agreed by a competent reviewer independent of the tool being evaluated, and disagreements resolved before testing begins.
PUT THIS INTO PRACTICE
- 01Include no-finding cases deliberately.
- 02Agree ground truth before running the tool.
- 03Track case composition against the target mix.
RELATED INSTRUMENTS
CITE THIS SECTION
§8.2 · Edition 1.0 · Building a golden reference set