ESCI brand studyground truth for measuring AI recommendations

10,746 human relevance judgements

Which brands did shoppers judge relevant?

AEO platforms report how often a brand appears in AI answers. Whether it belonged there is a separate question and needs ground truth. Amazon's ESCI supplies it: 10,746 human relevance judgements across 2,770 shopper queries. No model involved.

One query, as the data holds it

albany park sofa

  • RivetExact
  • AcanvaSubstitute
  • FDWSubstitute
  • HIFORTSubstitute
  • HONBAYSubstitute
  • KnowlifeSubstitute
  • NolanySubstitute

Plus 15 lower-volume sellers not shown. Exact brands are the target. Substitutes are near misses. Anything absent from this list is a dataset gap or an invention.

2,770
queries sampled
998
have competing brands
602
recognisable brands
4
median brands per query

Share of voice against agreement

Share of voice measures how much of the query set a brand covers. Agreement measures how often raters judged it a match. The same brands rank differently under each.

103
brands ranked both ways
+0.08
rank correlation
21
move more than 50 places

Share of voice does not predict agreement. The correlation stays within 0.12 of zero at every threshold tested and changes sign with the floor.

Breadth floor and sensitivity sweepWithout a floor the correlation reads -0.087, which is an exposure artifact

Agreement is not comparable across brands of different breadth. A brand judged against one query cannot be wrong, so it reaches 100% mechanically. Without a query floor those brands occupy the top of the agreement ordering and drive the correlation negative.

Rank correlation and count of perfect-agreement brands by minimum query floor
Minimum queriesBrandsCorrelationAt 100% agreement
1152-0.08718
3123+0.026
5used103+0.0943
871+0.0852
1236-0.1162

Mean agreement falls as breadth rises. Perfect scores collapse with it.

  • Brands on 1 queryn=2079.4% mean agreement, 50% at 100%
  • Brands on 2 queriesn=975.9% mean agreement, 22.2% at 100%
  • Brands on 3 to 4 queriesn=2074.8% mean agreement, 15% at 100%
  • Brands on 5 to 8 queriesn=3772.3% mean agreement, 2.7% at 100%
  • Brands on 9 or more queriesn=6670.2% mean agreement, 3% at 100%
Show

Positive gap: ranks higher on share of voice than on agreement. Negative: the reverse. Restricted to brands with 10+ judgements across 5+ distinct queries.

Brands ranked by share of voice against agreement with human judgements
Zinus1.1% #450% #93+89
Canon1% #651.2% #87+81
Walsunny1.3% #257.1% #81+79
2K0.7% #2246% #96+74
EZlifego0.7% #2546.7% #94+69
ABCCANOPY0.6% #3046.2% #95+65
Amazon Essentials0.8% #1358.3% #77+64
HyperX0.9% #965% #69+60
nixplay0.6% #3845.5% #97+59
2K GAMES0.6% #2950% #88+59
Apple0.7% #2357.9% #78+55
HONBAY0.8% #1760% #72+55
bonsaii0.9% #868.2% #61+53
Retainer Brite0.5% #5043.5% #98+48
Carhartt0.5% #4350% #90+47

Competitive sets

Each query is one shopper search with every brand judged against it.

Show sets with
210
sets shown of 400
70.9%
judged Exact
20.3%
judged Substitute

Queries (showing 60)

Judged against

albany park sofa

22 brands, 12 recognisable, across 28 judged products

  • ALBANY PARKone-offExact
  • RivetExact
  • AcanvaSubstitute
  • Casa Andrea Milano llcone-offSubstitute
  • FDWSubstitute
  • HIFORTSubstitute
  • Hommooone-offSubstitute
  • HONBAYSubstitute
  • Jennifer Taylor Homeone-offSubstitute
  • JULYFOXone-offSubstitute
  • KnowlifeSubstitute
  • Lexiconone-offSubstitute
  • Modwayone-offSubstitute
  • Mr. Kateone-offSubstitute
  • NolanySubstitute
  • NouhausSubstitute
  • POLY & BARKSubstitute
  • Signature Design by AshleySubstitute
  • SLEERWAYSubstitute
  • South Shoreone-offSubstitute
  • SZLIZCCCone-offSubstitute
  • WalsunnySubstitute
How crowded are these sets?Distribution of brand counts under the current filter

Breadth against agreement

Breadth is how many queries a brand appears against. Agreement is how often raters judged it a match. The two move independently.

Which brands
433
brands shown of 3,928
71.7%
judged Exact overall
46
widest breadth (Amazon Basics)

Breadth against exact rate

One circle per brand, sized by judgement count. Dashed lines mark the medians.

narrow but usually rightbroad and usually rightbroad but rarely right

Click any column heading to sort.

Brand judgement counts and exact rates
Nike1610381.6%
Anker357876.9%
Garmin267667.1%
EXPO186088.3%
Amazon Basics465678.6%
Motor Trend405457.4%
Gillette205356.6%
2K245046%
L'Oreal Paris145068%
Moleskine74879.2%
adidas234858.3%
ASUS204582.2%
TOZO154560%
Canon124351.2%
Herman Miller54353.5%

Substitute relationships

Two brands judged Substitute on the same query were treated as interchangeable for that intent. Read across the sample, those pairs form a competitor map from human judgement rather than inferred from co-mention.

Pairs recorded
300
Brands with rivals
24
Strongest pair
3

ABCCANOPY and Eurmax

Total pairings
334
  • AIZIYUO

    24 pairings
    • Master of Muscle2
    • 5BILLION FITNESS1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Master of Muscle

    24 pairings
    • AIZIYUO2
    • 5BILLION FITNESS1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • 5BILLION FITNESS

    23 pairings
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
    • Embracing Sport1
  • APLUGTEK

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • CutS1
    • DEGOL1
    • Denvosi1
    • Embracing Sport1
  • CutS

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • DEGOL1
    • Denvosi1
    • Embracing Sport1
  • DEGOL

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • Denvosi1
    • Embracing Sport1
  • Denvosi

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Embracing Sport1
  • Embracing Sport

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Epitomie Fitness

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • FEECCO

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Fit Vikings

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Gaoykai

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • JimboGym

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • JUSDO

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Just Jump It

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • morneve

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • OXIVE

    23 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • RDX

    19 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Redipo

    18 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Sportein

    18 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Ueasy

    17 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • Wastou

    17 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • XYLsports

    17 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1
  • zengxiaoyun

    17 pairings
    • 5BILLION FITNESS1
    • AIZIYUO1
    • APLUGTEK1
    • CutS1
    • DEGOL1
    • Denvosi1

Counts come from the sample, not the full corpus, so they are directional. The relationship is the useful part: these brands were judged interchangeable by people.

Phase 2: scoring a model against the ground truth

I put each competitive query to claude-sonnet-4-5-20250929 with no tools or web search, then matched the brands it named against the human judgements. 120 queries.

25%
recall of the Exact set
40%
named the leading brand
83%
precision among judged brands
3%
substitutes offered

The model named 10.925 brands per query on average, but ESCI had judged only 1.775 of them. So the 82%off-corpus rate is mostly real brands outside ESCI's coverage, not inventions, and raw precision (14%) is low for the same reason. Precision among the brands ESCI did judge is 83%.

Read together: when the model names a brand ESCI also judged, it is usually right (83% Exact). It recovers 25% of the full Exact set and names the single leading brand 40% of the time. The model is accurate on the overlap and incomplete against the whole set.

120 of 120
Per-query model scores against ESCI judgements
reclining bed queen100%100%hit
crosscut shredder without wastebasket100%67%hit
13 inch laptop100%50%hit
boxers briefs for men100%100%hit
fertility lubricant75%100%hit
ketchup71%100%hit
push pop variety pack67%100%hit
jandy jxi 400 pool heater67%67%hit
boxes for moving medium and large67%100%hit
extra extra large moving boxes67%100%hit
organic coffeebeans60%100%miss
1ms response time monitor57%80%miss
albany park sofa50%100%miss
steel-toe boots50%100%hit
air conditioner heater units50%80%hit
Method and metricsclaude-sonnet-4-5-20250929, no tools, 120 queries
  • Recall.Share of the query's Exact-labelled brands the model named.
  • Precision (judged). Of the named brands ESCI judged, the share labelled Exact. This is the fair precision, since it ignores brands ESCI never covered.
  • Leader. Whether the model named the Exact brand with the most judgements in the dataset.
  • Off-corpus. Named brands ESCI never judged for the query. Includes products released after ESCI and brands outside its results, so it is not a hallucination rate.

Brands are matched to ESCI by normalised name. The run cost about $0.12 in API usage. Generated 2026-08-02.

Method and caveats

The four judgementsWhat Exact, Substitute, Complement and Irrelevant mean
  • ExactProduct matches the query intent71%
  • SubstituteNot exact, but a reasonable alternative20.2%
  • ComplementUsed with the product, not instead of it1.7%
  • IrrelevantDoes not match the intent7.1%

A human rated every product against the query intent. Exact is the ground truth a recommendation would be scored against.

Brand column defects23 split spellings, 36 product titles, 12 unresolved aliases

Four defects, each handled in the numbers above.

  • 23 brands carried multiple spellings and were collapsed into one record. 1MORE and 1More are one company. Left split they divide the ground-truth set, which understates recall before a model is even involved.
  • 36 entries are product titles, not brands, because some sellers fill the field with keywords. Flagged rather than deleted, and excluded by default. 1 of them also carried split spellings, so these two counts overlap.
  • 12 entries are aliases that could not be resolved. A name written in another script with a romanisation in brackets resolves cleanly, so コールマン(Coleman) folds into Coleman. A bare one does not: キヤノン is Canon, but there is nothing to key on without an alias table. They are excluded from the substitute graph, where they would otherwise read as a brand competing with itself, and left in the brand list.
  • Only 602 of 3,928 brands appear against three or more queries. The usable brand universe is 15% of the raw count. The remainder are one-off marketplace sellers.

The substitute graph is where these surface. Uncleaned, the strongest rivalries are companies paired with themselves: 1MORE / 1More, Coleman / コールマン(Coleman), Canon / キヤノン.

LimitsTemporal mismatch is the one that can invalidate a result outright
  • Temporal mismatch. The judgements predate the paper's publication in June 2022. Model responses are current, so a model naming a 2025 product this dataset never saw is a dataset gap, not a model error. The paper states no collection window, so publication date is the only firm upper bound.
  • Relevance is not market share. Exact means the product matched the query intent, not that it sells well. A brand can be relevant and obscure.
  • Near-duplicate queries. The sample holds several variants of the same search, so a print-on-demand seller can clear a three-query threshold without being recognisable in any ordinary sense. Deduplicating by query stem would fix this and has not been done.
  • Amazon-shaped intent. These are product searches, not open-ended advice questions, so findings apply to purchase-adjacent prompts.
Sampling and sourcetasksource/esci, Apache-2.0

US locale only. The corpus is ordered by locale, so evenly spaced offsets draw from the US block at the front and the final offsets return no US rows. The sample is therefore even across the US portion, not across all 2.68M rows. ESCI has no category column, so the unit of analysis is the query, not the product category.

Reddy et al., Shopping Queries Dataset (ESCI), arXiv:2206.06588. Generated 2026-08-02 from a 0.4009% sample of 2,680,364 rows. Aggregates are precomputed and committed, both because the source API rate-limits anonymous access and because frozen numbers stay citable.