Methodology

A benchmark, not a black box.

The credibility of a ranking begins with what was tested, when it was tested, and how the result was calculated.

What we measure

Benchmark Bureau studies how AI systems represent, rank, and recommend companies in response to realistic buyer questions. The unit of analysis is an observed AI response, not an impression, website visit, lead, or sale.

Recommendation share is the percentage of relevant measured opportunities in which a company is actively recommended.

Consumer experience first

The core benchmark emphasizes default-like, search-enabled configurations that resemble what ordinary users are likely to encounter. Frontier models and deeper reasoning are tested on smaller stratified samples to measure whether stronger systems change who wins.

LayerPurposePlanned share of compute
Default-like systemsRepresent likely consumer recommendation experiences80%
Frontier systemsMeasure whether stronger reasoning changes recommendation outcomes15%
Controlled experimentsTest grounding, effort, and response stability5%

Raw observations remain intact

Each response retains the prompt, provider, model, configuration, timestamp, search status, recommendations, ordering, citations, URLs, explanation, geography, and raw response. This allows scoring rules to improve without rewriting history or rerunning every observation.

Rank and share answer different questions

Rank communicates relative position. Share communicates how much of the measured opportunity a company captures. A number-one rank can represent either a narrow lead or overwhelming dominance, so public results should present both.

Supporting measures

  • Mention share: how often a company appears at all.
  • Shortlist share: how often it enters the considered set.
  • Recommendation share: how often it is actively recommended.
  • First-choice rate: how often it appears first.
  • Citation share: how much of the supporting source ecosystem it occupies.
  • Consensus score: how consistently AI systems agree.

Search is a first-class variable

A company can win through model memory or current retrieval. Controlled grounded and ungrounded runs help distinguish brand recognition from evidence retrieved through reviews, directories, press, company websites, maps, and other sources.

API measurements and consumer products are distinct

Provider APIs enable controlled, repeatable measurement. Consumer applications include routing, personalization, tool selection, and product-specific orchestration. Benchmark Bureau labels API findings precisely and periodically validates them against the actual consumer surfaces.

Publication ruleNo benchmark will be presented as current without its observation window, model configuration, sample size, scoring definition, and methodology version.

Corrections and revisions

Factual errors are corrected openly. Methodological changes receive a new version and do not silently rewrite historical findings. Material corrections identify what changed, why it changed, and when the revision was made.