How we measure provider performance.
AOPX doesn't rank providers on opinion. Every result is produced by a pre-registered, prospective evaluation protocol.
How we measure provider performance.
AOPX does not rank providers on opinion. Every policy shown in production is backed by a pre-registered, prospective evaluation protocol — LAB-003.5.
1. Freeze the policy first
Before any mission runs, the routing policy (which provider is called, in which mode) is frozen and hashed. This prevents retroactive tuning on the results.
2. Fresh, paired missions
Each provider under test receives the exact same mission set, run in parallel, so that differences in outcome reflect the provider — not the task.
3. Holdout, not backtest
Missions are new at the time of the test, never previously seen by AOPX or the providers. This is what separates a prospective holdout from a backtest.
4. Pre-registered decision rule
Before results are known, AOPX defines the GO / PIVOT / STOP thresholds. A policy only ships to production if it clears the pre-registered bar.
5. Evidence ID & hash
Every run produces a verifiable evidence ID and content hash, so the result can be checked independently rather than taken on trust.
See the current benchmark
150 fresh missions, 300 paired provider calls, frozen policy & pre-registered rules.