LAB-003.5 prospective holdout
150 fresh SEARCH missions, 300 paired provider calls, frozen policy, frozen mission set and pre-registered decision rules.
AOPX
Evidence over claims
AOPX measures how search providers perform across software-agent tasks and uses that evidence to support provider selection and fallback decisions. The goal is not to declare one universal winner, but to make provider choice measurable across reliability, quality, cost and latency.
AOPX is an evidence-based provider-selection layer for software agents. It recommends a primary provider and a fallback path using measurable evidence, while reported production outcomes build an evolving record of provider performance.
AOPX is not a search engine and does not replace Brave Search, Tavily Search or other provider APIs. The search provider executes the actual search. AOPX focuses on the decision above that execution layer: which provider should the agent use, and what should happen if the first path fails?
150 fresh SEARCH missions, 300 paired provider calls, frozen policy, frozen mission set and pre-registered decision rules.
The observed result supports the value of fallback and provider redundancy more strongly than a claim that contextual routing alone caused the improvement.
| Provider path | Final successes | Success rate | Interpretation |
|---|---|---|---|
| AOPX policy | 40 / 50 | 80% | Measured policy result. |
| Tavily fixed | 37 / 50 | 74% | Fixed-provider baseline. |
| Brave → Tavily | 40 / 50 | 80% | Post-hoc cascade; supports fallback value. |
A single search API does not necessarily produce the same outcome across every agent task. Results can vary with query type, freshness requirements, source constraints, language, latency, cost and provider availability.
Provider routing gives an agent a decision layer rather than hard-coding one search provider forever. AOPX evaluates supported providers using measured evidence and can return a primary recommendation together with an optional fallback path.
Choose a supported search provider according to the active policy and available evidence.
Define a second provider path when the primary route does not satisfy the required outcome.
Record reported production outcomes so future provider decisions can be grounded in measurable history.
In an agent workflow, failure does not always mean that every provider would have failed. A fallback path gives the system another supported route before the task is considered unsuccessful. LAB-003.5 is especially relevant here because the observed CRITICAL result was reproduced by a simple Brave → Tavily cascade.
Controlled evaluation used to compare provider behaviour under a defined mission set and decision policy.
Outcomes returned after real external execution. Benchmarks, tests, replays and shadow runs are excluded from production totals.
Frozen artifacts, methodology and hashes make it possible to audit what was measured and when.
AOPX does not claim that Brave Search, Tavily Search or any future supported provider is universally best. Provider performance can depend on mission type, required freshness, source constraints, language, cost, latency and fallback policy.
AOPX also does not treat a benchmark result as a production outcome. The objective is to maintain a clear evidence trail and avoid turning a limited experiment into a broader claim than the data supports.
Use AOPX from software-agent workflows through the public API documentation.
Open API docsReview how the benchmark was frozen, measured and interpreted before using the results.
Read methodologyAOPX is an evidence-based provider-selection layer for software agents. It recommends supported providers and fallback paths using measurable provider performance and reported production outcomes.
No. AOPX does not replace search providers. The provider executes the search; AOPX adds provider selection, fallback policy and an evidence history above supported provider APIs.
The current SEARCH benchmark evaluates Brave Search and Tavily Search.
AOPX does not claim that either provider is universally superior. The benchmark measures provider behaviour under defined tasks and policies rather than declaring a permanent winner.
Fallback provides a second route when the primary provider does not satisfy the required outcome. In LAB-003.5, a Brave → Tavily cascade reached the same 80% CRITICAL final-success rate as the measured AOPX policy.
In the current consultation model, AOPX recommends the provider path and the calling agent executes the provider request externally.
It is the result returned to AOPX after a real externally executed recommendation. Benchmarks, tests, replays and shadow runs are excluded from production-outcome totals.
Review the public evidence trail behind AOPX provider-selection research and inspect the methodology before drawing conclusions from benchmark results.