How do you measure AI visibility honestly?
Summary
You measure AI visibility by running a fixed set of buyer questions across each engine on a schedule and recording four things per run: whether you were named, whether one of your sources carried the citation, which source earned it, and who was named instead. You cannot read this channel out of your analytics — when an engine declines to name you there is no impression to log. This covers building a prompt set that means something, why engines must never be blended into one score, what tools report natively, and where the chain to pipeline genuinely breaks.
You measure AI visibility by running a fixed set of buyer questions across each engine on a schedule, and recording four things per run: whether you were named, whether one of your sources was the attached citation, which source earned it, and who was named instead. Those four, tracked per engine over repeated runs, are the whole measurement model. Everything else is either downstream of them or decoration.
What you cannot do is read this channel out of your analytics. That is the problem this article exists to solve, and it is worth being honest about how partial the solution is.
Why does normal analytics not work for this?
Because most of what happens leaves no trace on your site. A buyer asks an assistant for suppliers, reads a paragraph naming four companies, and forms a shortlist. If you were not named, there is no impression to log and no click that failed to happen. The absence has no telemetry at all.
Even when you are named, the referral trail is unreliable. Some assistants pass a referrer and some do not, many buyers read the answer and then type your name into a browser directly, and Cloudflare's crawl-to-refer analysis published in July 2025 shows how heavily some providers crawl relative to the traffic they send back. Sessions arriving from AI surfaces are a floor on your visibility, not a measure of it.
So the measurement has to be active. You have to go and ask the questions yourself, because nothing reports them to you.
What exactly should be measured?
Four metrics, and they answer genuinely different questions. Collapsing them into one number is the most common mistake in this category.
| Metric | How it is produced | The question it answers |
|---|---|---|
| Answer share | Share of prompts in the fixed set where your company is named at all, per engine | Are you in the conversation? |
| Citation rate | Share of those answers where a source of yours carries the attached citation | Are you the source, or just a mention? |
| Cited surface | Which specific URL or profile earned the citation | Where does the next hour of work go? |
| Competitive set | Who is named when you are not, and from which source | What does the engine currently believe your category is? |
The distinction between the first two carries most of the value. Being named is a mention the model produced from what it already knows; being cited means a specific page of yours was retrieved and attributed. The second is far more durable, and it is the one you can actually influence — the mechanism is in how AI engines actually choose which sources to cite.
How do you build a prompt set that means anything?
The prompt set is the instrument. If it changes between runs, you have no measurement — you have anecdotes with dates on them.
- Write questions a buyer would type, not keywords. Retrieval systems commonly rewrite a question into several related queries before searching, so exact-match phrasing matters less than question shape.
- Cover the buying stages. Category questions ("who supplies X in Y"), comparison questions ("X versus Y for Z"), problem questions ("how do I solve Z"), and verification questions ("is company X any good").
- Include questions you expect to lose. A set built only from prompts you already win is a marketing asset, not an instrument.
- Fix it in writing before the first run. Agreed, dated, and handed over. Prompts added later cannot be compared with earlier runs.
- Keep it small enough to actually re-run by hand. Twenty questions re-run monthly beats two hundred run once and abandoned.
Two rules that sound pedantic and are not. Never edit a prompt mid-programme — start a second set instead and report them separately. And record the date and engine version alongside every result, because a change in the engine is a far more common explanation for a swing than anything you did.
What does a real prompt set look like?
Concrete beats abstract here, so this is the shape I use for an industrial supplier, with the reasoning attached to each row rather than left implicit.
| Type | Example | What it tests |
|---|---|---|
| Category | Who supplies industrial vision sensors in Germany? | Whether you exist in the engine's map of the category at all. The highest-value question, and usually the hardest to win |
| Specification | Which vision sensors suit automotive assembly line inspection? | Whether you are attached to the application rather than only to the product noun |
| Comparison | How does [your company] compare to [named competitor]? | What the engine believes about you relative to a rival, and which source it is reading that from |
| Verification | Is [your company] a reliable supplier? | What a buyer sees at the diligence step. Often the most alarming result, because it is answered largely from sources you do not own |
| Problem | How do I reduce false rejects on an inspection line? | Whether your content is retrieved at the stage before a supplier is even being considered |
Five types, three to five prompts each, is a workable set for one category. Run the verification questions first if you only have time for one pass — they are the ones most likely to surface a problem you did not know you had.
Why must each engine be measured separately?
Because presence in one predicts very little about the others. ChatGPT, Perplexity, Google's AI Overviews, Gemini and Claude differ in where they retrieve from, how densely they cite, and how much they lean on classic search rankings underneath.
The size of that last difference is measurable. Ahrefs studied Google's AI Overviews across 863,000 keywords and 4 million cited URLs and reported in March 2026 that about 38% of cited pages also ranked in the top ten organically, down from roughly 76% in July 2025. That is one engine's relationship with one ranking system over eight months. Assuming it holds for the others is an assumption, not a finding — and a blended score across engines would have hidden the change entirely.
How do you handle the fact that answers change every time?
By treating the trend as the evidence and the single run as a sample. Answer engines are non-deterministic: the same prompt can return different answers to different users on the same day.
Three practices follow. Re-run the whole set on a fixed schedule rather than checking ad hoc. Require a change to persist across at least two consecutive runs before calling it a change. And accept that this cuts both ways — a single screenshot showing you cited is no more meaningful than one showing you absent, and any vendor presenting either as a result is telling you they have not run this repeatedly.
Are there tools that report this natively yet?
Partially, and it is worth knowing what is first-party rather than inferred.
Microsoft added an AI Performance view to Bing Webmaster Tools, announced in February 2026 as a public preview, reporting how a site appears in Copilot and Bing's AI answers. That is platform-reported data rather than sampled, which makes it materially more reliable than any external estimate for the surfaces it covers.
Beyond that, third-party trackers exist and they are all doing the same thing you would do by hand: sending prompts and parsing answers. That is legitimate, and it scales. But it does not become platform truth by being automated, and a tool that will not show you its prompt set is selling you a number you cannot audit.
How do you connect any of this to pipeline?
Imperfectly, and saying so is part of doing it properly. The question clients ask agencies most often is whether marketing performance connects to revenue — 55% of agencies report hearing it regularly, according to AgencyAnalytics' 2026 Benchmarks Report of 494 agency professionals — and reporting that stops at citation counts cannot answer it.
The chain I use, with the weak link named:
- Answer share and citation rate — measured directly, reliable, the strongest link.
- Referral sessions from AI surfaces — measurable but understated for the reasons above. A floor, not a total.
- A "how did you hear about us" field on your own enquiry form — self-reported and lossy, and the only place a buyer who read an answer and typed your name directly can ever be caught.
- Enquiry quality — whether enquiries mentioning an assistant convert differently from the rest.
Step three is where the chain is weakest, and no vendor can fix that with a dashboard. Anyone showing you clean attribution for this channel has built it from assumptions and not labelled them.
What does the conversion evidence actually show?
One case study, and it needs both halves reported together or it misleads in whichever direction you stop reading.
Seer Interactive published a client analysis in June 2025 showing ChatGPT referral traffic converting at 15.9% against 1.76% for Google organic — and noted in the same study that AI traffic amounted to roughly 0.07% of organic traffic over the period.
Quote the first number alone and you have a channel beating search ninefold. Quote the second alone and you have a rounding error. The accurate reading is both: a small, unusually well-qualified stream, because someone arriving from a generated answer has typically already had their question answered and their shortlist narrowed. It is one client, one period, and self-reported — not a benchmark for your business.
What is not worth measuring?
- A single blended "AI visibility score." It averages away the per-engine differences that tell you what to do.
- Raw mention counts with no source attached. Being named 14 times is not a finding; being cited from a profile you had forgotten about is.
- Sentiment scoring of AI answers. Interesting, rarely actionable, and expensive to produce reliably.
- Anything on a prompt set you cannot see. If you cannot audit the instrument, the reading is unfalsifiable.
- Rankings as a proxy. The Ahrefs figure above is precisely the reason this no longer substitutes.
What should a report actually contain?
Per engine, per run, with the date: answer share, citation rate, the specific source behind every citation, the competitors named in your place, and a plain statement of what changed since the last run and what did not. Plus the prompt set itself, unchanged, attached to every report so the numbers can be checked.
If a report cannot be re-derived by someone else running the same prompts, it is not a measurement. That standard is unusual in marketing reporting and it is the entire reason this is worth doing carefully — the fuller argument is in what actually works in AEO, and what does not, and the questions to put to any vendor are in what to ask before you hire one.
What to do about your own situation
Take your baseline before changing anything. Write twenty buyer questions, run them across the five engines, and record the four measures above. It takes an afternoon, it costs nothing, and without it you will never be able to say whether anything you do afterwards worked.
If you would rather see one run done properly first, that is what the free check is. The service it belongs to is getting found when buyers ask AI.