Learn GEO · EP 28
ChatGPT, Claude and Gemini recommend different brands — we tested 73 to prove it
We asked three AI models the same 15 shopping questions about 73 well-known beauty brands. Each model crowned a different winner, and 34% of brands never appeared anywhere. What that means if you only ever check one AI.
⏰ The 30-second version
- We ran 73 well-known beauty brands through ChatGPT, Claude and Gemini using the same 15 customer-style questions.
- 34% never appeared in any of the three. Household names included.
- Each model had a different number one. One brand led ChatGPT with 12 mentions out of 15 — and scored zero in the other two models.
- Only one brand ranked at the top of all three. Everyone else's position collapsed when the model changed.
- So "we show up in AI" is not a fact about your brand. It's a fact about one model on one day.
Why we ran this
Most brands checking their AI visibility check ChatGPT, see themselves mentioned, and stop. We wanted to know whether that conclusion survives contact with a second model.
The setup was deliberately plain. Fifteen questions a real shopper would type — "recommend a Korean sunscreen for sensitive skin," "what's a good vegan skincare brand," "best moisturizer for dehydrated skin" — with no brand names in the prompt. Then we counted which of our 73 brands each model named.
Finding 1 — A third of famous brands are invisible everywhere
25 of 73 brands (34%) appeared zero times across all three models and all fifteen questions.
Not "ranked low." Never named. These aren't obscure labels either; the list includes brands with heavy ad budgets and strong retail presence.
This is the part that catches teams off guard: in search you can be on page two. In an AI answer there is no page two. Three or four names get spoken and the rest are absent — invisibly, with nothing showing up in analytics to tell you it happened.
Finding 2 — The winner changes with the model
Here's where the "we're fine, we checked ChatGPT" logic breaks.
| Brand | ChatGPT | Claude | Gemini |
|---|---|---|---|
| Brand A | 12 | 0 | 0 |
| Brand B | 2 | 0 | 10 |
| Brand C | 0 | 0 | 8 |
| Brand D | 4 | 9 | 0 |
| Brand E (the outlier) | 9 | 13 | 10 |
Brand A is the single most-recommended name in ChatGPT and does not exist in the other two. Brand D is a Claude favourite and a Gemini ghost. Only Brand E holds a top position everywhere — one brand out of 73.
If your customer happens to use the model you didn't test, your visibility there is a coin flip you never called.
Finding 3 — Models differ in how generous they are
The models don't just pick different names. They name different amounts.
- ChatGPT averaged about 10 brands per answer and covered 43 of our 73 brands overall.
- Claude averaged about 7 and covered 29.
- Gemini averaged about 6 and covered 27.
That gap is opportunity size. In a generous model there's room to be one of ten. In a strict one, only the established few survive — which cuts both ways: harder to break in, far more valuable once you're there.
Finding 4 — Niches are already spoken for
Ask about vegan skincare and one brand comes back repeatedly. Ask about acne and it's a different one. Ask about lip tints and the same three names appear together across models.
These associations were remarkably stable — more stable than overall rankings. Which suggests the practical move isn't "be famous in AI." It's own the question: pick the specific need you want to be the answer for, and make sure the sentences that answer it exist on your site.
What this changes about how you check
Three things follow from the data:
- Test at least three models. A single-model check produced a wrong conclusion for most brands in our set.
- Track over time, not once. These are model outputs, and models update. A reading from March tells you nothing in July.
- Optimize per question, not per brand. You don't win "AI." You win "best sunscreen for sensitive skin," one question at a time.
Where to start if you're at zero
The fix isn't more content. It's content that contains liftable sentences — a claim with a who, a why, and a specific.
"Radiance that lasts all day" cannot answer a question. "Formulated with three hyaluronic acids for dry and combination skin, fragrance-free for sensitive skin" answers several. Keep the first line if you love it; add the second underneath.
Pick one question you currently lose. Write the page that answers it in the first paragraph. Then check again in a few weeks — in all three models this time.
Running fifteen prompts across three models by hand is exactly the chore we built Bulti to remove: enter a brand, see which questions you're losing and what to publish first.
Honest footnotes
- 73 well-known Korean beauty brands, measured July 2026 with a fixed 15-question set, brand names excluded from prompts.
- Models: ChatGPT (gpt-5), Claude, Gemini 3.1. Each received identical questions.
- This measures what models recommend from trained knowledge. Live-retrieval tools such as Perplexity can differ.
- Brand letters in the table stand in for real brands from the study; the numbers are the measured ones.
- "Appeared" means the brand name was present in the answer text. Borderline cases exist.
- Zero mentions reflects content structure, not product quality or awareness, and is a snapshot that moves as models update.
So how many of these
is your website actually answering?
Find out in 1 minute across 30 customer questions — free, no signup.
Run my free scanBulti membership
Every article + unlimited scans + 10 drafts a month — $14.99/month
Next up
AI Search Engines 2026 — ChatGPT, Perplexity, Gemini, Copilot: Which and When