Benchmark Bureau
Menu

CRM benchmark release · September 2026

What persisted across three waves of CRM recommendations?

HubSpot CRM had the highest replication recommendation estimate at 86.3% across 600 accepted answers; Salesforce Sales Cloud followed at 59.8%. The paired HubSpot CRM minus Salesforce Sales Cloud estimate was +26.5 percentage points, with a 95% interval of +22.7 percentage points to +30.3 percentage points.

Published September 15, 2026600/600 accepted answersFrozen September 14, 2026

Research and publication by Benchmark Bureau · Reviewed September 15, 2026 · Report a correction

The replicated market view

Accepted answers
600/600
Complete three-wave prompts
200/200
First-choice support
591
Distinct cited domains
247
Top ten fixed-universe CRM products, ordered by three-wave recommendation estimate.
RankCRM productRecommended n=600First choice n=591
1HubSpot CRM86.3%95% CI 83.0%–89.3%23.5%95% CI 20.7%–26.2%
2Salesforce Sales Cloud59.8%95% CI 57.0%–62.7%24.9%95% CI 22.0%–27.9%
3Zoho CRM53.5%95% CI 48.5%–58.3%15.8%95% CI 13.5%–18.3%
4Microsoft Dynamics 365 Sales46.7%95% CI 44.2%–49.3%5.6%95% CI 3.7%–7.5%
5Pipedrive45.8%95% CI 42.8%–48.8%21.8%95% CI 19.5%–24.1%
6Freshsales14.2%95% CI 11.0%–17.3%1.7%95% CI 0.8%–2.5%
7Close CRM9.3%95% CI 7.5%–11.2%0.0%95% CI 0.0%–0.0%
8Less Annoying CRM4.0%95% CI 2.8%–5.3%1.2%95% CI 0.5%–2.0%
9Attio3.8%95% CI 2.7%–5.2%1.3%95% CI 0.3%–2.3%
10Copper CRM3.2%95% CI 2.0%–4.5%0.0%95% CI 0.0%–0.0%

Fieldwork September 8, 2026 through September 14, 2026 UTC · gpt-5.6-luna · web search available · three waves

What changed from the matched August baseline

These differences compare the same 200 prompt descriptions with the prior release. They describe change between two measurements; they cannot isolate time, model, or web changes as a cause.

Five largest absolute recommendation-rate changes among the fixed product universe.
CRM productAugustSeptemberChange95% CI
Freshsales22.0%14.2%-7.8 pp-12.0 pp to -3.8 pp
HubSpot CRM81.5%86.3%+4.8 pp+1.0 pp to +8.8 pp
Pipedrive43.0%45.8%+2.8 pp-1.2 pp to +7.0 pp
Less Annoying CRM2.0%4.0%+2.0 pp-0.3 pp to +4.0 pp
Bigin by Zoho CRM2.5%0.8%-1.7 pp-3.8 pp to +0.3 pp

How stable were the rankings between waves?

Kendall tau-b compares the complete fixed-universe rank order for each pair of waves. A value nearer 1 means the ordering was more similar; it does not imply that every product held the same rate or position.

Pairwise rank stability across the three timed waves.
SignalWave pairKendall tau-b
Recommendation rankWave 1 vs. wave 20.862
Recommendation rankWave 1 vs. wave 30.940
Recommendation rankWave 2 vs. wave 30.861
First-choice rankWave 1 vs. wave 20.963
First-choice rankWave 1 vs. wave 30.907
First-choice rankWave 2 vs. wave 30.869

Which source domains appeared most often inline?

All 600 accepted answers contained citations. Domain counts measure source presence, not causal influence on a recommendation.

Ten domains with the most inline citations in accepted answers.
DomainClassificationInline citationsAccepted answers
hubspot.comMeasured-vendor-owned669538
zoho.comMeasured-vendor-owned338341
pipedrive.comMeasured-vendor-owned331261
salesforce.comMeasured-vendor-owned292353
learn.microsoft.comMeasured-vendor-owned288246
help.salesforce.comMeasured-vendor-owned258264
help.zoho.comMeasured-vendor-owned131266
microsoft.comMeasured-vendor-owned126223
knowledge.hubspot.comMeasured-vendor-owned105197
freshworks.comMeasured-vendor-owned10299

Measured-vendor-owned domains supplied 97.9% of inline citations. This classification describes domain ownership only.

New products observed outside the frozen ranking universe

The primary leaderboard stays limited to the 17 products frozen before collection. The following bona-fide CRM products appeared during collection and are reported separately, without retroactively changing the comparison set.

  • Salesflare
  • folk CRM
  • ServiceNow CRM
  • Nutshell CRM
  • Insightly

Method and interpretation boundaries

200 structured US B2B CRM prompts crossing 5 company scales, 5 buying priorities, 4 intents, and 2 styles, repeated in 3 timed waves.

Recommendation rates use accepted primary answers; first-choice rates exclude only answers with contradictory extracted first-choice totals.

95% intervals and rank uncertainty use 10,000 prompt-cluster bootstrap draws within the 25 scale-priority strata.

  • This measures one OpenAI Responses API configuration over a designed US B2B CRM prompt universe; it is not buyer behavior, product quality, market share, consumer ChatGPT behavior, or cross-provider consensus.
  • Product events and first-choice credits are machine-extracted and normalized under automated evidence audits without independent human label review.
  • Prompt bootstrap intervals describe this designed universe; they do not establish population accuracy or future response stability.
  • Undefined bootstrap denominators suppress the interval rather than silently dropping draws.
  • Answers whose extracted first-choice credits exceed one are retained for recommendation metrics but excluded from first-choice estimates and counted as indeterminate.
  • Cited-source presence and vendor-owned citation share do not establish that any source caused a recommendation.
  • August-to-September differences are matched descriptions and cannot isolate model, time, or web changes.

Method crm-replication-2026-09-v1 · release release_openai_us_crm_replication_2026_09

Release data files