AI summary
A 2026 experiment demonstrated that ChatGPT can rank a completely fictional deodorant brand above real, clinically-tested competitors by simply creating a website and publishing targeted content, raising critical questions about AI product authority and recommendation credibility. Unlike Google Shopping, which requires verified product data and compliance checks, ChatGPT's recommendation system lacks gatekeeping mechanisms, allowing unverified products to gain visibility based on relevance alone. The experiment reveals a fundamental gap between traditional search visibility and generative AI visibility: while search engines distribute authority across multiple sources, AI models compress recommendations into singular, authoritative statements that transfer trust from the source to the narrator. As AI agents begin making purchasing decisions autonomously, this credibility gap presents significant risks for brands, consumers, and platforms regarding product authenticity and liability.
Eleven dollars. One hour. Three pages. A deodorant that does not exist. Three weeks later, ChatGPT was recommending it to strangers, sometimes above brands with clinical trials. Deana Burke’s August 2026 experiment, covered by Inc. and Cybernews, is not a growth hack to copy. It is a stress test of how models assign product authority. For any brand chasing chatgpt shopping visibility, the question it raises is blunt: who now owns the answer to what should I buy for X.
What the experiment actually measured
Before planting anything, Burke asked Claude, ChatGPT and Gemini what to buy nearly 9,000 times. The same names kept coming back across models: Native for natural deodorant, Brooks for running shoes, Similac and Enfamil for baby formula. She calls that invisible consensus, shared across models even when sources differ, the AI shelf. That is the real subject, a short list most brands never measure.
Then came Morrowen: a natural deodorant for sensitive skin, magnesium and colloidal oatmeal, baking soda free. No inventory, no fake reviews, no Reddit spam. A site, a Substack essay, a few targeted terms, then thirty days untouched. By her account, first appearance around day 21. After a month, ChatGPT with browsing on named Morrowen in four of four answers to the two long tail queries tested, listed first every time. Browsing off, nothing. Claude and Gemini, never. And on thirteen head queries like best or aluminum free, Native kept the broad shelf. Morrowen only found a dusty corner at the back of the store.
The detail that lands for an acquisition lead is one image. In one answer, Morrowen sat above MAGS Skin, a real brand, clinically tested, carrying the National Eczema Association seal. Same list, same voice, same formatting. Nothing separated the real from the invented.
This is not just SEO
Putting up a site and watching the web find it is nothing new. The shift sits elsewhere. On Google, a sensitive query returns links, ads, Reddit threads and roundups, and the reader keeps agency: the debate clearly looks unsettled. In ChatGPT, the same brand copy comes back compressed into a recommendation, delivered with third party poise. Burke puts it plainly: the model did not invent the claims, it laundered them. Trust leaves the site and settles in the narrator. That gap is exactly what now separates classic search work from visibility inside AI answers, a line we draw in the difference between classic search and generative visibility for online retailers.
And this arrives before agents take over the buying itself. While AI only recommends, a mistake costs a sale. The day an agent picks and pays, it becomes an order placed on the wrong product, a shift we track in what agentic commerce and the UCP change for brands.
Google Shopping versus AI answers: who holds authority?
Here is the authority question every acquisition group should sit with. On Google Shopping, a phantom brand built in an hour does not get through. Merchant Center requires a real product: GTIN, stock, price, merchant of record, checked compliance. On the assistants’ shelf, the logic flips. OpenAI describes its product results as organic and unsponsored, ranked purely on relevance to the request. No auction, no Merchant Center equivalent, no gatekeeper. A browsing model can therefore cite a landing page with zero inventory. The two systems are not playing the same game, which is precisely why your product feed now decides your ChatGPT visibility as much as your Shopping performance.
A few questions worth tabling, without waiting for a perfect tool. For best X for Y, what does the shopper treat as truth: the Shopping carousel, an AI Overview, or the chat they open at night. If Shopping demands a valid feed while chat can recommend a stockless site, which signal actually wins the purchase. When Claude ignores a brand ChatGPT ranks first, which answer does your customer treat as fact. And if an agent buys a nonexistent or misformulated product, who is liable: the brand, the model, or the merchant of record. None of these answers is settled yet, which is one more reason to prepare now.
Downside and upside
The downside is sharp. Confidence laundering, visual parity between clinical proof and an empty page, a fraud window on high intent long tail, model fragmentation, and amplification once agents buy without a results page. Sensitive categories, skin, light health, products for children, are not a marketing playground. A single owned site cuts both ways too: a single point of failure and an easy target for a lookalike domain copying your clinical pages.
The upside for anyone selling a real product is just as clear. The stunt failed where an incumbent already owned the answer, and only worked where no average had formed. Precise questions, pages that truly answer them, language matched to intent: that is classic content marketing with a new distribution layer, provided the product is true, sourced and machine readable. It is the same diagnosis as for structured product data, which now weighs as much as creative, and it is the point of GEO pages built to be cited by AI like Smart GEO Page, fed straight from the catalog.
What e-commerce brands can do now
The realistic posture fits in one line: keep winning Google and Shopping, and add a second scoreboard, your share of model recommendation across twenty to fifty category questions. You cannot cleanly buy the shelf like a Shopping slot yet. You can measure it, and make your real product the cleanest answer on it.
To defend the ground:
- Monthly AI answer audit, with a fixed prompt set across ChatGPT (browsing on and off), Gemini, Claude and Perplexity. Log the order, the citations, the false claims.
- Claim inventory across your site, product pages and ads. Whatever you publish will be restated with confidence, so make it accurate.
- Impersonation watch: lookalike domains and fake clinical pages planted on your long tail.
- Feed hygiene in GMC and marketplaces: GTIN, stock, shipping, returns. Agents that compare merchants still lean on these fields, which a feed enrichment layer like Feed Enrich reinforces.
To compete, legitimately:
- Map long tail intents with no clear winner: irritations, ingredient swaps, versus queries, for X skin.
- Publish question and answer, comparison and who it is for, who it is not pages, with proof, seals and contraindications in clean HTML.
- Add third party corroboration: retailer listings, serious summaries, credible reviews. A single source is fragile and easy to imitate.
- Clarify your brand entity, with a stable name and consistent links, so you do not read like a one page shell.
To avoid, finally: building disposable brands to game assistants, stuffing pages with ingredients you do not use, or assuming Claude’s silence means a safe category and a strong Shopping rank means chat presence. If you do not know where to begin, a product feed audit remains the most concrete entry point before you scale answer pages.
FAQ
Should we drop Google Shopping to focus on ChatGPT?
No. On head queries, incumbents stay hard to dislodge inside models. Shopping, PMax and the feed keep capturing demand while you work the conversational long tail. The two tracks complement each other rather than compete.
Why didn’t Claude and Gemini name Morrowen?
The experiment shows sharp fragmentation: only one model, with browsing, promoted the fake brand, on two queries only. That is both opportunity and risk, because your shopper may not use the model you test internally.
Is the lever content volume or product proof?
Both, in that order. True documented claims first, then pages that answer questions where no average has formed. Without proof, you are building the next Morrowen, with inventory attached.
How can an assistant recommend a stockless brand when Google Shopping requires stock?
Because the two surfaces follow different rules. Merchant Center verifies a real product and a merchant of record, while a browsing model ranks on the relevance of the text it finds. That asymmetry is exactly what let Morrowen through.
Where should a large catalog start?
With a feed audit and a readiness score, then twenty high intent questions on your priority categories. It is the fastest way to see where your real product can become the least ambiguous answer.
