Benchmark Bureau
Menu

Methodology

A benchmark, not a black box.

The credibility of a ranking begins with what was tested, when it was tested, and how the result was calculated.

Methods frameworkVersion 1.1September 2026

What we measure

Benchmark Bureau studies how AI systems represent, rank, and recommend companies in response to realistic buyer questions. The unit of analysis is an observed AI response, not an impression, website visit, lead, or sale.

Recommendation rate is the percentage of eligible measured answers in which a company is actively recommended. Recommendation event share is its portion of all recommendation events; an answer can recommend several companies.

From buyer question to public record

  1. DesignFreeze realistic buyer prompts and the analysis plan.
  2. CollectRecord configuration, timing, raw answers, and sources.
  3. ExtractApply defined entity and recommendation rules.
  4. ReviewReconcile evidence, exclusions, denominators, and uncertainty.
  5. ReleasePublish findings, methods, data, limitations, and a revision path.

Current coverage and planned comparisons

The published Utah personal-injury pilot, August CRM baseline, and September CRM replication each measure one OpenAI Responses API configuration with web search available. The September release repeats 200 prompts across three timed waves and reports prompt-bootstrap intervals, rank uncertainty, within-prompt overlap, and matched change from August. These releases do not establish cross-provider consensus or what consumer ChatGPT recommends.

The framework below describes a future allocation. Additional providers and frontier-model experiments remain planned. Each study publishes its actual configuration and allocation alongside its results.

LayerPurposePlanned share of compute
Default-like systemsRepresent likely consumer recommendation experiences80%
Frontier systemsMeasure whether stronger reasoning changes recommendation outcomes15%
Controlled experimentsTest grounding, effort, and response stability5%

Raw observations remain intact

Each response retains the prompt, provider, model, configuration, timestamp, search status, recommendations, ordering, citations, URLs, explanation, geography, and raw response. This allows scoring rules to improve without rewriting history or rerunning every observation.

Rank and frequency answer different questions

Rank communicates relative position. Rates and event shares communicate frequency with different denominators. A number-one rank can represent either a narrow lead or overwhelming dominance, so public results present the underlying counts and definitions.

Uncertainty and repeated measurements

The September CRM replication resamples whole prompts within the frozen company-scale and buying-priority strata for 10,000 bootstrap draws. It reports 95% intervals, the probability of each rank across draws, pairwise wave rank correlation, and recommendation set overlap. These measures describe uncertainty inside the designed prompt universe; they are not population polling margins or guarantees about future answers.

Supporting measures

  • Mention rate: answers in which a company appears, divided by eligible answers.
  • Shortlist rate: answers presenting it as an option, divided by eligible answers.
  • Recommendation rate: answers actively recommending it, divided by eligible answers.
  • First-choice rate: explicit leader credit divided by determinate eligible answers; co-leaders split credit and each release identifies exclusions.
  • Citation measures: consulted sources, inline citations, and reviewed support, with separate denominators.
  • Provider agreement: planned for matched comparisons across additional systems.

Search is a first-class variable

The studies distinguish consulted sources from visible inline citations and reviewed recommendation support. These associations do not establish why a company was recommended. Grounded-versus-ungrounded experiments are planned to investigate retrieval effects.

API measurements and consumer products are distinct

Provider APIs enable controlled, repeatable measurement. Consumer applications include routing, personalization, tool selection, and product-specific orchestration. Benchmark Bureau labels API findings precisely. Consumer-surface validation has not yet been completed and will be reported separately.

Publication ruleNo benchmark will be presented as current without its observation window, model configuration, sample size, scoring definition, and methodology version.

Ownership, corrections, and revisions

Bob Bodily is accountable for research direction and publication standards. Questions about methods may be sent to research@benchmarkbureau.com. Factual errors are corrected openly. Methodological changes receive a new version and do not silently rewrite historical findings. Material corrections identify what changed, why it changed, and when the revision was made.

Continue with the current CRM release, the supporting datasets, or the corrections policy.