A changing score does not necessarily mean the system changed. The evaluator may have changed too. Datasets evolve, parsers are corrected, thresholds move, admissibility rules shift, and reporting conventions are revised.
Once both the artifact and the evaluator are moving, a numerical difference can be completely real and still fail to tell us what actually caused it. That is where benchmark comparison stops being only a measurement problem and becomes an attribution problem.
The more reliable approach is to separate the two sources of movement. Hold the artifact fixed while the evaluator changes. Hold the evaluator fixed while the artifact changes. Then ask whether the resulting differences actually affect the decision being made. This matters far beyond AI leaderboards. Security scores, software benchmarks, safety tests, compliance systems, and readiness assessments all depend on observation systems that evolve.
The important question is not simply whether the score moved, but whether the evaluation constitution remained comparable enough to justify the conclusion.
https://www.linkedin.com/pulse/when-ruler-moves-system-rogerio-figurelli-vrvbf
