Which brands did shoppers judge relevant?
AEO platforms report how often a brand appears in AI answers. Whether it belonged there is a separate question and needs ground truth. Amazon's ESCI supplies it: 10,746 human relevance judgements across 2,770 shopper queries. No model involved.
One query, as the data holds it
albany park sofa
- RivetProduct matches the query intentExact
- AcanvaNot exact, but a reasonable alternativeSubstitute
- FDWNot exact, but a reasonable alternativeSubstitute
- HIFORTNot exact, but a reasonable alternativeSubstitute
- HONBAYNot exact, but a reasonable alternativeSubstitute
- KnowlifeNot exact, but a reasonable alternativeSubstitute
- NolanyNot exact, but a reasonable alternativeSubstitute
Plus 15 lower-volume sellers not shown. Exact brands are the target. Substitutes are near misses. Anything absent from this list is a dataset gap or an invention.
- 2,770
- queries sampled
- 998
- have competing brands
- 602
- recognisable brands
- 4
- median brands per query
Share of voice against agreement
Share of voice measures how much of the query set a brand covers. Agreement measures how often raters judged it a match. The same brands rank differently under each.
- 103
- brands ranked both ways
- +0.08
- rank correlation
- 21
- move more than 50 places
Share of voice does not predict agreement. The correlation stays within 0.12 of zero at every threshold tested and changes sign with the floor.
Breadth floor and sensitivity sweepWithout a floor the correlation reads -0.087, which is an exposure artifact
Agreement is not comparable across brands of different breadth. A brand judged against one query cannot be wrong, so it reaches 100% mechanically. Without a query floor those brands occupy the top of the agreement ordering and drive the correlation negative.
| Minimum queries | Brands | Correlation | At 100% agreement |
|---|---|---|---|
| 1 | 152 | -0.087 | 18 |
| 3 | 123 | +0.02 | 6 |
| 5used | 103 | +0.094 | 3 |
| 8 | 71 | +0.085 | 2 |
| 12 | 36 | -0.116 | 2 |
Mean agreement falls as breadth rises. Perfect scores collapse with it.
- Brands on 1 queryn=2079.4% mean agreement, 50% at 100%
- Brands on 2 queriesn=975.9% mean agreement, 22.2% at 100%
- Brands on 3 to 4 queriesn=2074.8% mean agreement, 15% at 100%
- Brands on 5 to 8 queriesn=3772.3% mean agreement, 2.7% at 100%
- Brands on 9 or more queriesn=6670.2% mean agreement, 3% at 100%
Positive gap: ranks higher on share of voice than on agreement. Negative: the reverse. Restricted to brands with 10+ judgements across 5+ distinct queries.
| Zinus | 1.1% #4 | 50% #93 | +89 |
| Canon | 1% #6 | 51.2% #87 | +81 |
| Walsunny | 1.3% #2 | 57.1% #81 | +79 |
| 2K | 0.7% #22 | 46% #96 | +74 |
| EZlifego | 0.7% #25 | 46.7% #94 | +69 |
| ABCCANOPY | 0.6% #30 | 46.2% #95 | +65 |
| Amazon Essentials | 0.8% #13 | 58.3% #77 | +64 |
| HyperX | 0.9% #9 | 65% #69 | +60 |
| nixplay | 0.6% #38 | 45.5% #97 | +59 |
| 2K GAMES | 0.6% #29 | 50% #88 | +59 |
| Apple | 0.7% #23 | 57.9% #78 | +55 |
| HONBAY | 0.8% #17 | 60% #72 | +55 |
| bonsaii | 0.9% #8 | 68.2% #61 | +53 |
| Retainer Brite | 0.5% #50 | 43.5% #98 | +48 |
| Carhartt | 0.5% #43 | 50% #90 | +47 |
Competitive sets
Each query is one shopper search with every brand judged against it.
- 210
- sets shown of 400
- 70.9%
- judged Exact
- 20.3%
- judged Substitute
Queries (showing 60)
Judged against
albany park sofa
22 brands, 12 recognisable, across 28 judged products
- ALBANY PARKone-offExact
- RivetExact
- AcanvaSubstitute
- Casa Andrea Milano llcone-offSubstitute
- FDWSubstitute
- HIFORTSubstitute
- Hommooone-offSubstitute
- HONBAYSubstitute
- Jennifer Taylor Homeone-offSubstitute
- JULYFOXone-offSubstitute
- KnowlifeSubstitute
- Lexiconone-offSubstitute
- Modwayone-offSubstitute
- Mr. Kateone-offSubstitute
- NolanySubstitute
- NouhausSubstitute
- POLY & BARKSubstitute
- Signature Design by AshleySubstitute
- SLEERWAYSubstitute
- South Shoreone-offSubstitute
- SZLIZCCCone-offSubstitute
- WalsunnySubstitute
How crowded are these sets?Distribution of brand counts under the current filter
Breadth against agreement
Breadth is how many queries a brand appears against. Agreement is how often raters judged it a match. The two move independently.
- 433
- brands shown of 3,928
- 71.7%
- judged Exact overall
- 46
- widest breadth (Amazon Basics)
Breadth against exact rate
One circle per brand, sized by judgement count. Dashed lines mark the medians.
Click any column heading to sort.
| Nike | 16 | 103 | 81.6% |
| Anker | 35 | 78 | 76.9% |
| Garmin | 26 | 76 | 67.1% |
| EXPO | 18 | 60 | 88.3% |
| Amazon Basics | 46 | 56 | 78.6% |
| Motor Trend | 40 | 54 | 57.4% |
| Gillette | 20 | 53 | 56.6% |
| 2K | 24 | 50 | 46% |
| L'Oreal Paris | 14 | 50 | 68% |
| Moleskine | 7 | 48 | 79.2% |
| adidas | 23 | 48 | 58.3% |
| ASUS | 20 | 45 | 82.2% |
| TOZO | 15 | 45 | 60% |
| Canon | 12 | 43 | 51.2% |
| Herman Miller | 5 | 43 | 53.5% |
Substitute relationships
Two brands judged Substitute on the same query were treated as interchangeable for that intent. Read across the sample, those pairs form a competitor map from human judgement rather than inferred from co-mention.
- Pairs recorded
- 300
- Brands with rivals
- 24
- Strongest pair
- 3
- Total pairings
- 334
ABCCANOPY and Eurmax
AIZIYUO
24 pairings- Master of Muscle2
- 5BILLION FITNESS1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Master of Muscle
24 pairings- AIZIYUO2
- 5BILLION FITNESS1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
5BILLION FITNESS
23 pairings- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
- Embracing Sport1
APLUGTEK
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- CutS1
- DEGOL1
- Denvosi1
- Embracing Sport1
CutS
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- DEGOL1
- Denvosi1
- Embracing Sport1
DEGOL
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- Denvosi1
- Embracing Sport1
Denvosi
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Embracing Sport1
Embracing Sport
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Epitomie Fitness
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
FEECCO
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Fit Vikings
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Gaoykai
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
JimboGym
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
JUSDO
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Just Jump It
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
morneve
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
OXIVE
23 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
RDX
19 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Redipo
18 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Sportein
18 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Ueasy
17 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Wastou
17 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
XYLsports
17 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
zengxiaoyun
17 pairings- 5BILLION FITNESS1
- AIZIYUO1
- APLUGTEK1
- CutS1
- DEGOL1
- Denvosi1
Counts come from the sample, not the full corpus, so they are directional. The relationship is the useful part: these brands were judged interchangeable by people.
Phase 2: scoring a model against the ground truth
I put each competitive query to claude-sonnet-4-5-20250929 with no tools or web search, then matched the brands it named against the human judgements. 120 queries.
- 25%
- recall of the Exact set
- 40%
- named the leading brand
- 83%
- precision among judged brands
- 3%
- substitutes offered
The model named 10.925 brands per query on average, but ESCI had judged only 1.775 of them. So the 82%off-corpus rate is mostly real brands outside ESCI's coverage, not inventions, and raw precision (14%) is low for the same reason. Precision among the brands ESCI did judge is 83%.
Read together: when the model names a brand ESCI also judged, it is usually right (83% Exact). It recovers 25% of the full Exact set and names the single leading brand 40% of the time. The model is accurate on the overlap and incomplete against the whole set.
| reclining bed queen | 100% | 100% | hit |
| crosscut shredder without wastebasket | 100% | 67% | hit |
| 13 inch laptop | 100% | 50% | hit |
| boxers briefs for men | 100% | 100% | hit |
| fertility lubricant | 75% | 100% | hit |
| ketchup | 71% | 100% | hit |
| push pop variety pack | 67% | 100% | hit |
| jandy jxi 400 pool heater | 67% | 67% | hit |
| boxes for moving medium and large | 67% | 100% | hit |
| extra extra large moving boxes | 67% | 100% | hit |
| organic coffeebeans | 60% | 100% | miss |
| 1ms response time monitor | 57% | 80% | miss |
| albany park sofa | 50% | 100% | miss |
| steel-toe boots | 50% | 100% | hit |
| air conditioner heater units | 50% | 80% | hit |
Method and metricsclaude-sonnet-4-5-20250929, no tools, 120 queries
- Recall.Share of the query's Exact-labelled brands the model named.
- Precision (judged). Of the named brands ESCI judged, the share labelled Exact. This is the fair precision, since it ignores brands ESCI never covered.
- Leader. Whether the model named the Exact brand with the most judgements in the dataset.
- Off-corpus. Named brands ESCI never judged for the query. Includes products released after ESCI and brands outside its results, so it is not a hallucination rate.
Brands are matched to ESCI by normalised name. The run cost about $0.12 in API usage. Generated 2026-08-02.
Method and caveats
The four judgementsWhat Exact, Substitute, Complement and Irrelevant mean
- ExactProduct matches the query intent71%
- SubstituteNot exact, but a reasonable alternative20.2%
- ComplementUsed with the product, not instead of it1.7%
- IrrelevantDoes not match the intent7.1%
A human rated every product against the query intent. Exact is the ground truth a recommendation would be scored against.
Brand column defects23 split spellings, 36 product titles, 12 unresolved aliases
Four defects, each handled in the numbers above.
- 23 brands carried multiple spellings and were collapsed into one record. 1MORE and 1More are one company. Left split they divide the ground-truth set, which understates recall before a model is even involved.
- 36 entries are product titles, not brands, because some sellers fill the field with keywords. Flagged rather than deleted, and excluded by default. 1 of them also carried split spellings, so these two counts overlap.
- 12 entries are aliases that could not be resolved. A name written in another script with a romanisation in brackets resolves cleanly, so コールマン(Coleman) folds into Coleman. A bare one does not: キヤノン is Canon, but there is nothing to key on without an alias table. They are excluded from the substitute graph, where they would otherwise read as a brand competing with itself, and left in the brand list.
- Only 602 of 3,928 brands appear against three or more queries. The usable brand universe is 15% of the raw count. The remainder are one-off marketplace sellers.
The substitute graph is where these surface. Uncleaned, the strongest rivalries are companies paired with themselves: 1MORE / 1More, Coleman / コールマン(Coleman), Canon / キヤノン.
LimitsTemporal mismatch is the one that can invalidate a result outright
- Temporal mismatch. The judgements predate the paper's publication in June 2022. Model responses are current, so a model naming a 2025 product this dataset never saw is a dataset gap, not a model error. The paper states no collection window, so publication date is the only firm upper bound.
- Relevance is not market share. Exact means the product matched the query intent, not that it sells well. A brand can be relevant and obscure.
- Near-duplicate queries. The sample holds several variants of the same search, so a print-on-demand seller can clear a three-query threshold without being recognisable in any ordinary sense. Deduplicating by query stem would fix this and has not been done.
- Amazon-shaped intent. These are product searches, not open-ended advice questions, so findings apply to purchase-adjacent prompts.
Sampling and sourcetasksource/esci, Apache-2.0
US locale only. The corpus is ordered by locale, so evenly spaced offsets draw from the US block at the front and the final offsets return no US rows. The sample is therefore even across the US portion, not across all 2.68M rows. ESCI has no category column, so the unit of analysis is the query, not the product category.
Reddy et al., Shopping Queries Dataset (ESCI), arXiv:2206.06588. Generated 2026-08-02 from a 0.4009% sample of 2,680,364 rows. Aggregates are precomputed and committed, both because the source API rate-limits anonymous access and because frozen numbers stay citable.