Home / Learn / Measuring AI visibility
Learn

How to measure AI visibility: four instruments, none of them traffic

There is no single number for AI visibility, and any dashboard offering you one has made it up. Four separate instruments each answer a different question: server logs prove the engines can reach you, Google Search Console reports impressions inside AI features but no clicks, referral analytics catch a shrinking minority of the visits AI sends, and repeated prompt sampling is the only thing that measures whether you are actually cited. Run all four, report them separately, and never add them together.

The reason one number does not exist

Rank tracking works because a ranking is a fact. You asked Google for a query, there was a list, you were seventh. Ask an AI engine the same question twice and you can get two different answers with two different sets of sources, so the thing you want to count refuses to hold still. On top of that, the engines report almost nothing back and each one hides a different part of the picture.

What you get instead is four partial views. Each is honest about something and blind to everything else. The discipline is knowing which is which, because the expensive mistake is not measuring badly, it is measuring one thing and believing you measured another.

Instrument one: server logs, the pass or fail underneath everything

Before any of the other three mean anything, the engines have to be able to fetch your pages. Your own server or CDN logs are the only place that fact lives, and they are the one dataset in this whole discipline that nobody can spin, because it is your infrastructure recording its own behaviour.

Pull thirty days and count successful fetches by user agent: GPTBot, OAI-SearchBot and ChatGPT-User, PerplexityBot and Perplexity-User, ClaudeBot, Google-Extended, Applebot-Extended. What you are looking for is not volume, it is zeroes and 403s. A robots.txt that reads as perfectly welcoming, sitting above a bot-management product that has been challenging those agents for months, is the most common reason a site is invisible, and it is invisible to every other instrument on this list. This is the first thing the AEO Visibility Audit checks, before it measures anything at all.

Instrument two: Google now reports impressions, and stops there

On 3 June 2026 Google launched a generative AI performance report in Search Console, initially to a subset of properties, as reported by Search Engine Land on the day. It covers AI Overviews and AI Mode, and Google's Search Console documentation defines the impression it counts as how many times links to your site were shown to a user in a generative AI feature on Google Search.

Read the limits carefully, because they decide what you are allowed to conclude:

Google shipped a control alongside it that lets a site block its content from appearing in AI features, and stated that sites opting out will not receive traffic or impressions from those features while the choice will not affect core search rankings. That is a clean trade to understand and a bad one to take by accident, which is worth checking if you inherited a property. The mechanics of getting cited on that surface in the first place are in the AI Overviews guide.

What this instrument proves: you appeared. What it cannot tell you: whether anyone read you, clicked you, or what they had asked.

Instrument three: referral data is a floor, and the floor is sinking

The most widely quoted number in this field right now comes from Previsible's AI traffic study, covering 166 GA4 properties from November 2024 through May 2026 and 6.77 million sessions, reported by Search Engine Land on 6 July 2026. It puts ChatGPT at 92.4% of trackable large language model referral traffic, with Claude growing sharply after overtaking Perplexity in March 2026.

The word carrying that entire sentence is trackable. Attrifast, an attribution vendor writing about its own 200-site cohort in May 2026, reports that roughly 70% of the sessions it attributes to ChatGPT show up in GA4 as Direct or none. The mechanism is not mysterious: the iOS and Android apps send no Referer header across the native app boundary at all, and desktop web sends origin-only under a strict referrer policy. Attrifast sells the fix, so read the figure with that in mind, but the mechanism is checkable and it is real.

Put those two together and the honest reading is uncomfortable. Every AI referral market share figure in circulation is a share of the visible slice, and nobody outside the engines knows the size of the invisible one. If ChatGPT drops referrers more often than its competitors do, its dominance is understated. If less often, overstated. There is currently no way to tell from the outside, which is exactly why we treat referral counts as a floor and never as a share.

One finding from the same Previsible data is immediately actionable and mostly gets skipped: roughly a quarter of AI-referred traffic lands on internal search result pages. If that is happening to you, a real slice of your hard-won AI visibility is arriving on a page that was never built to convert anyone. Check your landing page report before you spend another hour on citations. The tactics that earn the citation are a separate job, covered in how to get cited by ChatGPT.

Instrument four: citation tracking is polling, not rank tracking

This is the only instrument that measures the thing AEO is actually about, and it is the one most often run wrong, because people run a prompt once, screenshot the result and call it a reading.

The evidence against that is now formal. A variance components decomposition published on arXiv on 14 July 2026 by Dmitrij Zatuchin analysed 12,933 responses about 20 brands across 8 languages, 3 models and 15 prompts per combination. On its resampled stability subset, within-prompt resampling alone accounted for 34.8% of the variance in a single response. Query language accounted for 26.5% across the full corpus. Brand identity, the thing you are supposedly measuring, accounted for 1.5%. A single run is not a light reading. It is mostly noise with your brand somewhere inside it.

Kevin Indig made the same argument operationally in Search Engine Land on 10 June 2026, and his framing is the right one to steal: prompt tracking should look less like rank tracking and more like polling, with repeated runs, clear sampling rules, confidence intervals and segmented panels. His working shape is 40 seed prompts, split into brand, category and problem-led questions, run five times per prompt per platform every week, recording mention rate and citation rate with confidence intervals, average position when mentioned, and sentiment.

The same arXiv paper answers the question everyone asks next, which is how many repeats to buy. Not more than five. A sixth repeat reduced relative-error variance by 0.0003, while spending that same query budget on additional languages and models reduced it several times more. So the allocation rule is blunt and slightly counterintuitive: five runs, then stop deepening and start widening. More prompts and more engines, not more repeats of the same question.

Build the scorecard, and refuse to add it up

Four rows, four separate numbers, reported side by side every month:

The temptation is to roll these into one AI visibility score out of 100, because a single number is easier to put in a board deck. Resist it. Averaging an impression count, a broken referral count and a sampled citation rate produces a number whose movement nobody can explain, and an unexplainable metric gets ignored inside two quarters.

What nobody can tell you, including us

There is no credible public benchmark for what a good citation rate looks like in your category. Every percentage circulating comes from a vendor's own cohort, measured with that vendor's own prompt set, which is a fair thing to publish and a useless thing to be compared against. Your prompt set is not their prompt set, and citation rate is almost entirely a function of which questions you chose to ask.

So the only benchmark that means anything is your own previous month, measured the same way, with the prompt set frozen. That is a less impressive claim than an industry average and it is the one that survives contact with a sceptical CFO. It is also how AI Citation Monitoring reports: your trend against yourself and against the named competitors you nominated, never against a number we invented.

A first week, in order

Day one, logs. Thirty days by user agent, looking for zeroes and 403s. If something is blocked, stop here and fix it, because nothing else you measure this week will mean anything.

Day two, Search Console. Check whether the generative AI report has reached your property yet. If it has, note the impression baseline and the date it starts. If it has not, write down that you are flying without it rather than quietly assuming zero.

Day three, referrals. Segment sessions from chatgpt.com, perplexity.ai, claude.ai, gemini.google.com and copilot.microsoft.com, then look at what they land on before you look at how many there are.

Day four and five, the panel. Write 40 questions a real buyer would ask, weighted toward problem-led phrasing rather than your brand name, because the brand-name prompts flatter you and teach you nothing. Run each five times per engine. Record cited, mentioned or absent. That is your baseline, and it is worth more than any tool you could buy in the same week.

Common questions

Can I see AI Overviews traffic in Search Console?

You can see impressions, not clicks. Google launched the generative AI performance report on 3 June 2026 and confirmed at launch that it does not include click data. AI Overviews and AI Mode are reported together rather than separately, there are no query rows and no API, and Google's documentation notes the report is still rolling out so not every property has it.

Why does ChatGPT traffic show up as Direct in GA4?

Because the referrer frequently never arrives. Native mobile apps send no Referer header across the app boundary, and desktop web sends origin-only under a strict referrer policy. Attrifast, reporting on its own 200-site cohort in May 2026, puts roughly 70% of the ChatGPT sessions it attributes into Direct or none. Treat AI referral counts as a floor.

How many times should I run each prompt?

Five per engine, then widen instead of deepening. The July 2026 arXiv variance decomposition found a sixth repeat cut relative-error variance by 0.0003, while spending the same budget on more languages and models cut it several times more. More questions and more engines beat more repeats.

Have the panel run for you, monthly

AI Citation Monitoring tracks your mention and citation rate across five AI engines on a frozen prompt set, with the trend measured against your own baseline and your named competitors.

AI Citation Monitoring, $49/month
95/100
AI Search ReadyOur own site is built to the exact standard we sell, so AI engines can find, read and recommend it. Verified by Faro, built by the same team.
See the proof →