Increase the share of verdicts traceable to a stated condition
Measure: traceable verdict share
AML's evaluation model versions three things in a fixed order: the grounds an evaluation stands on, the concerns and criteria it judges by, and the examples that calibrate it. A verdict is never a bare number — it resolves to the criterion, the condition and the ground that produced it.
Each outcome names the measure it moves and the direction it moves it — the value is stated as a measure, not as a claim.
Increase the share of verdicts traceable to a stated condition
Measure: traceable verdict share
Reduce undetected drift between evaluation runs
Measure: undetected instrument drift
Increase agreement between independent evaluators
Measure: inter-evaluator agreement
When criteria live in a prompt and evidence lives in someone's memory, two runs of the same evaluation are not comparable and drift is invisible. The evaluation aggregate freezes each part under a revision, so a re-score either uses the same instrument or declares that it did not.
An evaluation is an aggregate root over three member sets. The ground set fixes the evaluand and the sources. The concern set descends concern to dimension, subdimension, criteria, and finally the SBVR conditions each criterion is decided by, each with a modality, a polarity and a ground reference. The example set carries the golden and nuanced cases the instrument is calibrated against. Every embedded address resolves against the frozen revision it names.