Abstract illustration of one AI search query splitting into many different answers

AI Visibility Scores Are Lying to You

New research on 8,609 AI responses shows the brand ChatGPT recommends shifts with user history. How I measure AI visibility without fooling myself.

By Dalin de Graff

I audit AI visibility dashboards for clients, and the pitch is always the same: a clean set of test prompts, a green score, a number that looks like a grade. The number feels objective. It is measuring something. It is not measuring what the client thinks.

A study released September 23 backs that up harder than I expected. Cassie Wilson Clark, an AI search visibility consultant, and Joao da Silva, co-founder of friction AI, ran 8,609 AI responses through ChatGPT, Gemini, Claude and Perplexity. They tested six personas, three consumer categories and 30 prompts in the UK, comparing history-free API calls, temporary logged-in sessions, and accounts primed with user histories.

The result: who asks the question changes which brands get recommended. They call it the Personalization Gap.

There is no single answer to rank in

After accounting for normal response variation, ChatGPT showed a 16.0 percentage point gap between the brands it recommends with and without persistent user history. Gemini showed 33.3 percentage points, the largest effect of the four platforms. Perplexity showed 7.2 points. Claude showed no detectable change beyond normal volatility, though its search behavior did shift when user context was introduced.

Gemini's recommended brand set diverged 33.3 percentage points above the volatility floor between history-free and history-primed accounts, the largest history effect of the four. Your visibility dashboard runs the test on a clean account and reports a single number. Your actual customers ask from history-laden accounts, not the clean test environments your dashboard uses. Those are two different measurement environments, and the study says they produce two different brand lists.

Geography was the clearest personalization dimension in the UK-based experiment. In history-primed accounts, researchers measured a 6.6 percentage point increase in UK and EU brand share, a 4.9 point decrease in US-origin brands, and a 6.0 point increase in citations from .uk domains. The same prompt, asked from a British browsing history, pulls British brands forward. The study's authors are careful to note this does not prove localization automatically improves AI visibility. In this experiment, there was no single true answer, and nothing in the study establishes how far beyond these categories that holds.

A screenshot is not a measurement system

The line from the study I keep coming back to is Clark's: "A screenshot is evidence that an answer happened. It isn't a measurement system." Which is why I treat a dashboard number as one measurement from one environment, not a grade on my visibility. The study goes further. The researchers checked whether history-free API sampling could reproduce the brand recommendations seen in personalized accounts. Across 173 matched comparisons, API sampling recovered 79% of the recurring brands from history-primed sessions. Temporary logged-in sessions recovered 95% at the same sampling depth.

That 79% is the one to sit with: if your dashboard samples through the API without user context, roughly one in five of the brands that keep showing up in history-primed sessions will not show up in your tests. I still use API-based tools. They track change over time on a fixed test fine. I read the number as one measurement now, with the sampling depth, collection method and user context printed next to the score.

The traffic this measures is getting real. BrightEdge data from September showed ChatGPT referral traffic more than doubled between January and August 2026, up 101%, and ChatGPT took a 95.1% share of AI-generated referrals in August. At that growth rate, a bad measurement habit compounds.

Abstract illustration of one AI search query splitting into many different answers

How I actually measure AI visibility

I build AI visibility reporting at First Rank and on my own campaigns, so let me be specific about what I do differently after reading this.

The first habit I dropped: trusting single snapshots. When I run a GEO engagement, I run the same prompt set multiple times across days, on ChatGPT, Gemini, Perplexity and Claude, logged in and logged out, and I score brands by recurrence, not by any single appearance. A brand mentioned in 8 of 10 runs is a pattern. A brand mentioned once is noise. The Personalization Gap study gives that habit a name: recurring brand set, measured against a volatility floor. That is the unit worth tracking.

The number I trust most comes from analytics, not the dashboard. Google added UTM parameters to Gemini's outgoing links in September so site owners can attribute referrals accurately, and I wrote earlier about how little of that data flows back to publishers. But GA4's AI referral channel plus those UTMs is ground truth about actual visitors. A dashboard score going up while your AI referral traffic stays flat means the dashboard is measuring its own test environment, not your customers. I keep both numbers. The traffic is the one I trust when they disagree.

Third, I test geography as a variable now, not a constant. The study ran in the UK. The direction should hold in Canada: if a UK browsing history shifts the brand list toward .uk domains, a Canadian customer's history will shift their own list too. But Claude's null result is the warning. The effect varies by platform, and I would not assume the geography numbers port over. Here is how I run it for a Winnipeg client: the same prompts from a clean session, then from a session primed with that market's context. If the brand shows up clean and vanishes primed, the clean test is the one lying. For local and regional businesses, the test itself costs nothing.

Finally, I stopped scoring ChatGPT alone. Comscore's Q2 2026 AI Intelligence Report showed Claude's share of AI prompt volume rising from 2% to 11% between January and June while Gemini nearly doubled from 17% to 30%, with ChatGPT falling from 70% to 50%. The history effect varied sharply by platform, from 7.2pp on Perplexity to 33.3pp on Gemini, so one platform's score tells you nothing about the others.

On a recent GEO campaign I measured LLM search traffic up 98% in 30 days, and that number came from analytics, not from a dashboard. The tracking setup is what let me prove the work did something.

Your Monday-morning visibility routine

The version of this I run for clients is simpler than the dashboards make it look. Take the five prompts that map to the client's money keywords. Run each three times, logged in and clean, across ChatGPT, Gemini, Perplexity and Claude. Split the brands into the ones that repeat and the one-shot noise. Then put that list next to the dashboard's score and your GA4 AI referral traffic. When the dashboard says you are winning and your referrals disagree, the dashboard's test environment is the problem. Build the recurring-brand list from the sessions that look like your customers, not the sessions that are easiest to automate.

Dalin de Graff

About the author

Dalin de Graff is an SEO and LLM search specialist at First Rank in Winnipeg, where he built the agency's GEO practice. He writes about AI search, technical SEO and the business of marketing at dalindegraff.com. Find him on LinkedIn, X and YouTube.