As Washington and Brussels build formal AI evaluation regimes, a repeatable but circular benchmark can convert methodological error into policy at scale ...