FIELD NOTE / 28OPEN ACCESS / BULTI

ChatGPT, Claude and Gemini recommend different brands — we tested 73 to prove it

We asked three AI models the same 15 shopping questions about 73 well-known beauty brands. Each model crowned a different winner, and 34% of brands never appeared anywhere. What that means if you only ever check one AI.

← All field notesREAD / APPLY / MEASURE

⏰ The 30-second version

  • We ran 73 well-known beauty brands through ChatGPT, Claude and Gemini using the same 15 customer-style questions.
  • 34% never appeared in any of the three. Household names included.
  • Each model had a different number one. One brand led ChatGPT with 12 mentions out of 15 — and scored zero in the other two models.
  • Only one brand ranked at the top of all three. Everyone else's position collapsed when the model changed.
  • So "we show up in AI" is not a fact about your brand. It's a fact about one model on one day.

Brands tested

73

Korean beauty brands

same fixed prompt set

Invisible everywhere

25 / 73

34% received zero mentions across all three models

Why we ran this

Most brands checking their AI visibility check ChatGPT, see themselves mentioned, and stop. We wanted to know whether that conclusion survives contact with a second model.

The setup was deliberately plain. Fifteen questions a real shopper would type — "recommend a Korean sunscreen for sensitive skin," "what's a good vegan skincare brand," "best moisturizer for dehydrated skin" — with no brand names in the prompt. Then we counted which of our 73 brands each model named.

Finding 1 — A third of famous brands are invisible everywhere

25 of 73 brands (34%) appeared zero times across all three models and all fifteen questions.

Not "ranked low." Never named. These aren't obscure labels either; the list includes brands with heavy ad budgets and strong retail presence.

This is the part that catches teams off guard: in search you can be on page two. In an AI answer there is no page two. Three or four names get spoken and the rest are absent — invisibly, with nothing showing up in analytics to tell you it happened.

Finding 2 — The winner changes with the model

Here's where the "we're fine, we checked ChatGPT" logic breaks.

Brand ChatGPT Claude Gemini
Isntree 12 0 0
Aestura 2 0 10
Abib 0 0 8
Goodal 4 9 0
Round Lab (the outlier) 9 13 10

Isntree is the single most-recommended name in ChatGPT and does not exist in the other two. Goodal is a Claude favourite and a Gemini ghost. Only Round Lab holds a top position everywhere — one brand out of 73.

If your customer happens to use the model you didn't test, your visibility there is a coin flip you never called.

Finding 3 — Models differ in how generous they are

The models don't just pick different names. They name different amounts.

  • ChatGPT averaged about 10 brands per answer and covered 43 of our 73 brands overall.
  • Claude averaged about 7 and covered 29.
  • Gemini averaged about 6 and covered 27.

ChatGPT

43 / 73

10.3 brands named per answer

Claude

29 / 73

6.9 brands named per answer

Gemini

27 / 73

5.7 brands named per answer

That gap is opportunity size. In a generous model there's room to be one of ten. In a strict one, only the established few survive — which cuts both ways: harder to break in, far more valuable once you're there.

Finding 4 — Niches are already spoken for

Ask about vegan skincare and one brand comes back repeatedly. Ask about acne and it's a different one. Ask about lip tints and the same three names appear together across models.

These associations were remarkably stable — more stable than overall rankings. Which suggests the practical move isn't "be famous in AI." It's own the question: pick the specific need you want to be the answer for, and make sure the sentences that answer it exist on your site.

What this changes about how you check

Three things follow from the data:

  1. Test at least three models. A single-model check produced a wrong conclusion for most brands in our set.
  2. Track over time, not once. These are model outputs, and models update. A reading from March tells you nothing in July.
  3. Optimize per question, not per brand. You don't win "AI." You win "best sunscreen for sensitive skin," one question at a time.

Where to start if you're at zero

The fix isn't more content. It's content that contains liftable sentences — a claim with a who, a why, and a specific.

"Radiance that lasts all day" cannot answer a question. "Formulated with three hyaluronic acids for dry and combination skin, fragrance-free for sensitive skin" answers several. Keep the first line if you love it; add the second underneath.

Pick one question you currently lose. Write the page that answers it in the first paragraph. Then check again in a few weeks — in all three models this time.

Running fifteen prompts across three models by hand is exactly the chore we built Bulti to remove: enter a brand, see which questions you're losing and what to publish first.

Honest footnotes

  • 73 well-known Korean beauty brands, measured July 2026 with a fixed 15-question set, brand names excluded from prompts.
  • Models: ChatGPT (gpt-5), Claude, Gemini 3.1. Each received identical questions.
  • This measures what models recommend from trained knowledge. Live-retrieval tools such as Perplexity can differ.
  • "Appeared" means the brand name was present in the answer text. Borderline cases exist.
  • Zero mentions reflects content structure, not product quality or awareness, and is a snapshot that moves as models update.

FAQ

Q. Do ChatGPT, Claude and Gemini recommend the same brands? No. In this July 2026 test, each model produced a materially different brand set. Isntree led ChatGPT with 12 mentions but received zero in Claude and Gemini, while Round Lab was the only brand near the top in all three.

Q. How many beauty brands were invisible across all three AI models? Twenty-five of the 73 Korean beauty brands tested, or 34%, were not named in any of the 45 answers generated from the fixed 15-question set.

Q. Is one AI visibility check enough for a brand? No. A single-model check can give a false sense of visibility because a brand that leads one model may be absent from the others. Use the same prompt set across at least three models and rerun it on a fixed schedule.

APPLY THIS NOTE / YOUR BRAND

How many buyer questions
does your website actually answer?

Build the brand report first. Inspect the initial findings before you decide whether to unlock the complete 60-day program.

Build my free report
Browse all free Learn GEO notes →