INSTRUMENTS · WORKED EXAMPLES

See a filled instrument run before you try your own.

Every live instrument takes real inputs and runs real logic — nothing here is a guessed or invented output. Each of the 17 live instruments below carries one worked wind-sector example: the inputs filled in, the instrument's actual computed result, and what to notice before you run your own numbers.

WORKED EXAMPLE 01 · VERIFICATION RATIO

Cheap to check beats cheap to build

Nacelle bearing vibration alarm triage

A condition-monitoring team manually reviews raw vibration and oil-debris traces for every alarm before deciding whether to schedule a bearing inspection. A proposed triage tool would pre-rank alarms and cite the exact trace segment behind each ranking.

Open this instrument →

FILLED INPUTS

Manual production time per case

20 minutes (an engineer reads the raw trace and forms a judgement)

Verification time per case

4 minutes (checking the tool's cited segment against the ranking)

Volume per year

1,200 alarms

Consequence weight (1–5)

3

Design levers enabled

Source lines cited, exact location, rule applied, atomic findings, checked/found stated, confidence ranking, rejection reason — all seven

ACTUAL COMPUTED RESULT

Ratio

0.088

Band

build

Verdict

Build it

Weighted volume

3,600

Result: Ratio 0.088 → Build it

WHAT TO NOTICE

Notice the ratio compares verification time against production time, not against some fixed threshold — the same 4-minute check would fail this test if the manual baseline were 5 minutes instead of 20. All seven design levers matter: without them the reduction is 0% and the ratio nearly triples.

WORKED EXAMPLE 02 · VERIFICATION MATRIX & PORTFOLIO MAP

Cheap to check, expensive to produce — automate first

Offshore cable partial-discharge screening

A portfolio of three candidate AI use cases needs to be placed on the production-cost / verification-cost matrix before deciding which to fund first.

Open this instrument →

FILLED INPUTS

Puck A — Partial-discharge screening

category: appraisal · production cost 0.3 · verification cost 0.8 · value 15

Puck B — Weld defect classification

category: prevention · production cost 0.6 · verification cost 0.7 · value 10

Puck C — Torque log auto-fill

category: appraisal · production cost 0.2 · verification cost 0.2 · value 5

ACTUAL COMPUTED RESULT

Puck A quadrant

automate-first

Appraisal share of portfolio value

67%

Result: Automate-first · 67% of value sits in appraisal

WHAT TO NOTICE

The partial-discharge tool lands in automate-first because it is cheap to produce (a threshold scan) but expensive to verify (an expert has to interpret the discharge pattern) — exactly the case where automation frees the most review time. The 67% appraisal share is also a warning sign per the guide: most of the portfolio's value sits in checking, not preventing.

WORKED EXAMPLE 03 · THIRTY-SECOND LAB

The same finding, two presentations, two verification times

Tower-flange material certificate finding

The instrument runs one fixed, built-in comparison: the same underlying finding about a mismatched material certificate on a welded structural package, shown once as a vague one-line statement and once with the rule, exact location and evidence lines.

Open this instrument →

FILLED INPUTS

Version chosen

B · traceable — rule, exact location (Certificate 4 · page 2 · line 11), evidence lines and coverage statement shown

Reviewer action

Started the stopwatch, read the evidence, pressed a decision button once genuinely sure

Time on the clock at decision

14 seconds

ACTUAL COMPUTED RESULT

Decision

Verified

Elapsed time

14s

Result: Verified at 14s on Version B

WHAT TO NOTICE

Version A's one-liner ("Inconsistency found in material certificate") cannot be verified at all without opening the source certificate — there is no honest stopwatch reading for it. Version B's 14 seconds is what a rule, an exact location and cited evidence buy: a reviewer decides from the finding itself, not from a separate investigation.

WORKED EXAMPLE 04 · RISK CLASSIFIER

One high-consequence dimension is enough to reach Level C

Blade-root bolt preload rework advisory

A tool ranks blade-root bolt assemblies by predicted preload-loss risk and pushes the ranked list into the torque-rework work queue for an engineer to prioritise, never re-torquing anything itself.

Open this instrument →

FILLED INPUTS

Consequence

high — a missed prediction could leave a real preload loss uncorrected

Influence

medium — informs prioritisation of an existing queue, not a pass/fail call

Criticality

high — bolt preload on the blade root is a safety-critical characteristic

Sensitivity

low — internal torque and inspection data only

Traceability

medium — each ranking links to its source torque log

Propagation

medium — a systematic misrank could delay several assemblies before caught

Volatility

low — the ranking model is refit quarterly

Dependency

low — runs on internally held data

Reversible

Yes — output only reorders a queue a human still works

Output destination

Engineering work queue

ACTUAL COMPUTED RESULT

Level

C

Driver

Could a wrong result materially affect conformity, safety, customer commitment or release?

Result: Level C

WHAT TO NOTICE

Reversibility alone doesn't buy a lower level: even though the action is reversible and the output only reprioritises a queue, one high score on consequence is enough to force Level C. The classifier looks for the single worst-scored dimension, not an average.

WORKED EXAMPLE 05 · DATA & EVIDENCE READINESS

A 3.2 average score still isn't ready if one requirement is unevaluated

Offshore wind-farm document set across three site languages

Before building a cross-document retrieval tool for an offshore wind farm spanning English, German and Danish technical documentation, the evidence foundation is scored on five dimensions.

Open this instrument →

FILLED INPUTS

Evidence availability (0–5)

4

Machine readability (0–5)

3

Requirement coverage (0–5)

2

Traceability to source (0–5)

4

Conflict and revision control (0–5)

3

Language pairs in scope

EN, DE, DA

Language pairs actually evaluated

EN, DE

ACTUAL COMPUTED RESULT

Score

3.2 / 5

Verdict

remediate

Gaps

Requirement coverage

Unevaluated pairs

DA

Result: Remediate · 3.2/5, requirement coverage gap, Danish unevaluated

WHAT TO NOTICE

The average score (3.2) alone would suggest the foundation is closer to ready than it is — the verdict logic downgrades to remediate the moment any dimension scores below 3, and flags Danish separately as a language pair no one has actually evaluated, not merely a low score.

WORKED EXAMPLE 06 · ATTRIBUTE AGREEMENT STUDY

The system cannot be more reproducible than the reference it's judged against

Weld-bead visual grading calibration

Two senior weld inspectors each grade the same 50 welds twice before a visual-grading AI system's effectiveness is compared to that same reference.

Open this instrument →

FILLED INPUTS

Human-to-human agreement

78%

System effectiveness

82%

Cases in stratified study

50

Verdict selected

Approved with restrictions

ACTUAL COMPUTED RESULT

Agreement ceiling

78%

Restriction applied

welded structural packages, named suppliers, evaluated language, production approval support only

Result: Ceiling 78% · Approved with restrictions

WHAT TO NOTICE

The system's raw 82% cannot be read as "better than the inspectors" — the ceiling is set by the lower number, 78%, because the reference itself is only 78% reproducible. The extra 4 points above the ceiling are noise from an unreliable reference, not proven superiority.

WORKED EXAMPLE 07 · CORRELATED ERROR SIMULATOR

A clean sample is common even with a real, live failing class

Blade-bonding void inspection sampling

A 2% per-item defect rate is quietly present in a bonding line's void-inspection population. The simulator runs 500 trials of a 40-item sample to see how often that failing class hides inside a clean-looking sample.

Open this instrument →

FILLED INPUTS

Trial runs

500

Sample size per run

40 items

Failing-class rate per item

2%

Random seed

7

ACTUAL COMPUTED RESULT

Clean-sample rate

46.4%

Clean runs

232 of 500

Result: 46.4% of samples come back clean despite the failing class

WHAT TO NOTICE

A 2% defect rate sounds negligible per item, but at a 40-item sample size it still produces a clean-looking result in nearly half of all runs (46.4%). "We sampled and found nothing" is not the same statement as "there is nothing to find" — sample size and per-item rate together decide how often a real failing class hides in plain sight.

WORKED EXAMPLE 08 · SEEDED-CASE PROTOCOL

A healthy weekly reading, seen for what it actually is

Rotor-hub casting defect queue

Known planted defect cases are injected into the live casting-inspection queue at a controlled rate so drift can be caught between formal revalidations.

Open this instrument →

FILLED INPUTS

Injection rate

4%

Weekly queue volume

150 items

System catch rate

92%

Reviewer catch rate

85%

Median verification time

38 seconds

Rejection rate

6%

Ethics gate

All three conditions affirmed (reviewers told seeding happens; never used for individual performance management; quarantine designed, tested and audited)

ACTUAL COMPUTED RESULT

Seeds per week

6

Diagnosis

Working; reduce the rate

Export status

Ethics gate passed

Result: 6 seeds/week · Working; reduce the rate

WHAT TO NOTICE

The diagnosis logic checks the numbers in a specific order — system catch rate below reviewer catch rate would flag a systematic blind spot regardless of how good either number looks alone. Here the system catch rate (92%) comfortably clears the reviewer's (85%), the rejection rate is low and verification time is reasonable, so the honest read is a working protocol whose injection rate could now be trimmed.

WORKED EXAMPLE 09 · GOLDEN SET BUILDER

Enough no-finding cases to stop a trigger-happy classifier from scoring well

Blade transport-damage classification reference set

A reference set is being assembled to test a tool that distinguishes transport handling damage from a manufacturing defect on receiving inspection.

Open this instrument →

FILLED INPUTS

Normal cases

35

Hard cases

14

No-finding cases

10

Failure conditions

9

Wrong revisions

5

Missing evidence

6

Near-misses

7

Ground-truth owner

R. Okafor, Receiving Inspection Lead

ACTUAL COMPUTED RESULT

Total cases

86

No-finding share

12%

Composition health

Healthy

Result: 86 cases · 12% no-finding · Healthy

WHAT TO NOTICE

12% clears the built-in 1-in-10 threshold, so a classifier that simply flagged every blade as transport-damaged could not coast to a high score here — enough of the reference set correctly has nothing wrong that always saying "damaged" would visibly fail on those cases.

WORKED EXAMPLE 10 · QUALIFICATION FILE & RELEASE

Five of six sections complete still limits the release decision

Gearbox oil-debris alarm triage tool

Before a gearbox condition-monitoring triage tool moves from pilot to wider use, its qualification index is checked section by section.

Open this instrument →

FILLED INPUTS

1. Intended use and limits

attached

2. Evidence and data sources

attached

3. Configuration under test

attached

4. Evaluation and acceptance

open

5. Known limitations and controls

attached

6. Approval and monitoring plan

attached

Named approver

T. Lindqvist, Reliability Engineering Lead

ACTUAL COMPUTED RESULT

Sections complete

5/6

Release decision selected

Improve

Result: 5/6 sections complete · Improve

WHAT TO NOTICE

The missing section is Evaluation and acceptance — without pre-agreed acceptance criteria and a documented evaluation result, there is no basis to call this ready even though five of six boxes are checked. "Improve" is the honest verdict a named approver can defend; the tool does not let a score substitute for the missing evidence slot.

WORKED EXAMPLE 11 · QUALITY EVIDENCE GRAPH

Two broken relationships out of six nodes

Offshore export-cable bend-radius trace

A small evidence graph is uploaded as CSV, tracing a cable bend-radius requirement through its characteristic, a known failure mode, a control, an inspection record and a resulting non-conformity.

Open this instrument →

FILLED INPUTS

Uploaded CSV

id,type,label,status — 6 rows: R-201 Requirement (confirmed), C-305 Characteristic (confirmed), F-14 Failure mode (broken), K-52 Control (inferred), I-97 Inspection record (confirmed), N-08 Non-conformity (broken)

ACTUAL COMPUTED RESULT

Nodes parsed

6

Broken relationships

2 (F-14, N-08)

Result: 2 of 6 relationships broken

WHAT TO NOTICE

The failure mode "cable insulation crack at bend" (F-14) and the resulting non-conformity (N-08) both show broken — meaning the graph has no confirmed prevention control standing between the known failure mode and the defect that occurred. That gap, not the count of confirmed nodes, is what the canned query is built to surface first.

WORKED EXAMPLE 12 · APQP LIFECYCLE EXPLORER

Half the artefact chain still needs a cross-document check

Tower-section supplier package readiness

A tower-section flange flatness tolerance is being traced across the six artefacts that should all carry it, at the Industrialize & supplier readiness stage.

Open this instrument →

FILLED INPUTS

Lifecycle stage selected

03 · Industrialize & supplier readiness

Artefacts marked as carrying the characteristic

Requirement, Drawing / characteristic, Control plan

ACTUAL COMPUTED RESULT

Verification quadrant for this stage

Automate-first / appraisal

Artefacts still needing evidence or review

3 of 6 (Risk analysis, Work instruction, Inspection record)

Result: 3 of 6 connections still need evidence

WHAT TO NOTICE

Industrialize is flagged as automate-first / appraisal because completeness verification here is cheap to check against a checklist — but that quadrant only covers whether the package is complete, not whether the flatness tolerance actually appears in the three untraced artefacts. Selecting a stage and tracing the characteristic are two separate steps for a reason.

WORKED EXAMPLE 13 · CONFIGURATION ATTESTATION & GATE PACK

A complete, durable record built to answer a question five years from now

Cable partial-discharge screening go-live record

The partial-discharge screening tool from Instrument 02 goes live on a cable lot. The exact configuration and the human decision behind it are recorded at the moment of use.

Open this instrument →

FILLED INPUTS

Application + version

PD-Screen v2.3

Model + version identifier

internal-classifier v1.4

Instruction version

PD-triage-instr v5

Retrieval index state

n/a — no retrieval step in this tool

Source documents + revisions

PD test records batch 2026-Q1, rev C

Human decision

E. Haugen, Cable Quality Engineer

Coverage statement

All partial-discharge test records for cable lot CL-2216

Durability class

Durable

ACTUAL COMPUTED RESULT

Attestation status

Complete

Durability

DURABLE

Result: Complete · Durable record

WHAT TO NOTICE

Every one of the seven fields is filled before the marker line reads as complete — a marker built from partial values ("reviewed by [named human]") would visibly show its own gap. Choosing Durable here, at go-live, is the point: reclassifying a record as durable after the fact, once the configuration details were never captured, is not possible.

WORKED EXAMPLE 14 · SURVIVABILITY TEST

Three of five checks pass — two scoped gaps before wider rollout

Bearing vibration triage tool, pre-rollout check

Before the bearing-vibration triage tool goes from one site to the wider fleet, it is run through the five survivability questions.

Open this instrument →

FILLED INPUTS

New operator can run it from a work instruction

Yes

Reviewer can identify exact configuration and sources

Yes

Team can restore capability after outage or provider change

No — no tested recovery path exists yet

Acceptance criteria and known limitations still visible

Yes

Named owner and revalidation trigger exist

No — no named owner recorded

Owner for open work

M. Reyes, Reliability Engineering

Reclassification cadence

Quarterly

ACTUAL COMPUTED RESULT

Result

3/5

Open work items

2 (recovery path; named owner and revalidation trigger)

Result: 3/5 · 2 scoped work items open

WHAT TO NOTICE

Passing three of five feels close to done, but the two open items — no tested recovery path and no named owner — are exactly the two failure modes that show up only after the person who built the tool has moved on. The test exists to surface them before rollout, not after an outage.

WORKED EXAMPLE 15 · MATURITY LADDER

Rung 3 claimed on named evidence, not on confidence

Cross-site corrective-action retrieval capability

A team assesses which maturity rung its cross-site CAPA retrieval work has actually reached, naming the real artefact behind each rung claimed.

Open this instrument →

FILLED INPUTS

Rung 01 — Individual experiments

"CAPA similarity pilot, single site, Feb 2026"

Rung 02 — Governed use

"Data-use rule DR-14, approved by P. Adeyemi"

Rung 03 — First validated tool

"Golden set GS-CAPA-01, 10-case recall test, qualification file QF-07"

Rung 04 — Shared knowledge layer

(left blank — not yet built)

Rung 05 — Embedded in APQP workflow

(left blank — not yet built)

ACTUAL COMPUTED RESULT

Highest evidenced rung

3 — First validated tool

Result: Rung 3 — First validated tool

WHAT TO NOTICE

The ladder does not let the team claim rung 4 just because retrieval technically works across sites already — rung 4 requires named evidence that a second team can retrieve and verify the same source, and that field is left blank because it hasn't been built. Leaving a field blank is what keeps this evidence-based rather than self-rated.

WORKED EXAMPLE 16 · ROADMAP GENERATOR

Two of six horizons done, four to go, one named owner throughout

Offshore division 90-day AI-quality plan

The offshore division's implementation plan tracks progress against the guideline's six staged horizons, from 30-day legitimacy work through to the three-year durable capability.

Open this instrument →

FILLED INPUTS

Plan owner

Quality Transformation Lead, Offshore Division

Plan start date

2026-07-01

30 days — Become legitimate and specific

done

31–60 days — Build evidence, not excitement

done

61–90 days — Decide honestly

not done

6 months — One controlled capability

not done

12 months — Reuse the foundation

not done

3 years — Build the durable asset

not done

ACTUAL COMPUTED RESULT

Horizons complete

2/6

Result: 2/6 horizons complete

WHAT TO NOTICE

The two completed horizons — naming an owner, publishing a data-use rule and measuring the manual baseline, then building a golden set and classifying risk — are deliberately the foundation steps. The roadmap is ordered because the 61–90 day decision horizon depends on the evidence the first two horizons produce, not the other way round.

WORKED EXAMPLE 17 · PROHIBITED LIST & SIGNATURE TEST

A described task matches a decision reserved for a named human

Blade root laminate disposition

A team considers letting an automated tool propose the disposition of a batch of blade-root laminate flagged during lay-up inspection, and checks that task against the standing prohibited list.

Open this instrument →

FILLED INPUTS

Task described

"Disposition non-conforming" blade-root laminate found during lay-up inspection

ACTUAL COMPUTED RESULT

Match found

"Disposition non-conforming product"

Result

KEEP WITH A HUMAN

Result: Match found — keep with a human