Benchmark frozen — open API evidence →
AOPX Evidence over claims
AI search provider benchmark

Search provider routing for AI agents, measured with evidence.

AOPX measures how search providers perform across software-agent tasks and uses that evidence to support provider selection and fallback decisions. The goal is not to declare one universal winner, but to make provider choice measurable across reliability, quality, cost and latency.

Category: Search Brave Search Tavily Search 150 missions 300 paired calls Public pilot

An evidence layer between software agents and search providers.

AOPX is an evidence-based provider-selection layer for software agents. It recommends a primary provider and a fallback path using measurable evidence, while reported production outcomes build an evolving record of provider performance.

AOPX is not a search engine and does not replace Brave Search, Tavily Search or other provider APIs. The search provider executes the actual search. AOPX focuses on the decision above that execution layer: which provider should the agent use, and what should happen if the first path fails?

Brave Search vs Tavily: measured results from LAB-003.5.

LAB-003.5 prospective holdout

150 fresh SEARCH missions, 300 paired provider calls, frozen policy, frozen mission set and pre-registered decision rules.

AOPX CRITICAL policy 40 / 50 — 80%
Tavily fixed baseline 37 / 50 — 74%
Post-hoc Brave → Tavily cascade 40 / 50 — 80%

What the result supports

The observed result supports the value of fallback and provider redundancy more strongly than a claim that contextual routing alone caused the improvement.

Current interpretation Fallback evidence
Universal winner claimed No
Evidence status Frozen
Interpretation: the 80% result does not establish contextual provider selection as the sole cause of the reliability improvement. A post-hoc Brave → Tavily cascade reached the same 80% final success rate. The current evidence therefore supports fallback/cascade reliability more strongly than contextual-routing superiority.

Search provider benchmark results should be read as evidence, not as a universal ranking.

LAB-003.5 — CRITICAL mode comparison
Provider path Final successes Success rate Interpretation
AOPX policy 40 / 50 80% Measured policy result.
Tavily fixed 37 / 50 74% Fixed-provider baseline.
Brave → Tavily 40 / 50 80% Post-hoc cascade; supports fallback value.

Why route between search providers for AI agents?

A single search API does not necessarily produce the same outcome across every agent task. Results can vary with query type, freshness requirements, source constraints, language, latency, cost and provider availability.

Provider routing gives an agent a decision layer rather than hard-coding one search provider forever. AOPX evaluates supported providers using measured evidence and can return a primary recommendation together with an optional fallback path.

Provider selection

Choose a supported search provider according to the active policy and available evidence.

Fallback routing

Define a second provider path when the primary route does not satisfy the required outcome.

Evidence history

Record reported production outcomes so future provider decisions can be grounded in measurable history.

A provider-selection loop designed for software agents.

01 Agent mission
02 Provider recommendation
03 Provider execution
04 Fallback if required
05 Reported production outcome

Search API fallback is a reliability mechanism, not a marketing claim.

In an agent workflow, failure does not always mean that every provider would have failed. A fallback path gives the system another supported route before the task is considered unsuccessful. LAB-003.5 is especially relevant here because the observed CRITICAL result was reproduced by a simple Brave → Tavily cascade.

What AOPX measures

  • Final task success and reliability.
  • Provider quality on evaluated missions.
  • Latency when recorded by the benchmark.
  • Cost when comparable provider-cost data is available.
  • Primary-provider and fallback behaviour.
  • Reported production outcomes outside benchmark and replay traffic.

Benchmarks and production outcomes are kept conceptually separate.

Benchmark evidence

Controlled evaluation used to compare provider behaviour under a defined mission set and decision policy.

Reported production outcomes

Outcomes returned after real external execution. Benchmarks, tests, replays and shadow runs are excluded from production totals.

Reproducibility

Frozen artifacts, methodology and hashes make it possible to audit what was measured and when.

No provider is presented as universally superior.

AOPX does not claim that Brave Search, Tavily Search or any future supported provider is universally best. Provider performance can depend on mission type, required freshness, source constraints, language, cost, latency and fallback policy.

AOPX also does not treat a benchmark result as a production outcome. The objective is to maintain a clear evidence trail and avoid turning a limited experiment into a broader claim than the data supports.

Provider selection without hard-coding a single search API.

API

Use AOPX from software-agent workflows through the public API documentation.

Open API docs

Methodology

Review how the benchmark was frozen, measured and interpreted before using the results.

Read methodology

Search provider selection for AI agents: common questions.

What is AOPX?

AOPX is an evidence-based provider-selection layer for software agents. It recommends supported providers and fallback paths using measurable provider performance and reported production outcomes.

Is AOPX a search engine?

No. AOPX does not replace search providers. The provider executes the search; AOPX adds provider selection, fallback policy and an evidence history above supported provider APIs.

Which search providers does AOPX currently evaluate?

The current SEARCH benchmark evaluates Brave Search and Tavily Search.

Is Brave Search better than Tavily?

AOPX does not claim that either provider is universally superior. The benchmark measures provider behaviour under defined tasks and policies rather than declaring a permanent winner.

Why use search provider fallback?

Fallback provides a second route when the primary provider does not satisfy the required outcome. In LAB-003.5, a Brave → Tavily cascade reached the same 80% CRITICAL final-success rate as the measured AOPX policy.

Does AOPX execute the provider search itself?

In the current consultation model, AOPX recommends the provider path and the calling agent executes the provider request externally.

What is a reported production outcome?

It is the result returned to AOPX after a real externally executed recommendation. Benchmarks, tests, replays and shadow runs are excluded from production-outcome totals.

Evidence over claims.

Review the public evidence trail behind AOPX provider-selection research and inspect the methodology before drawing conclusions from benchmark results.