How do you measure AI visibility honestly?

Summary

You measure AI visibility by running a fixed set of buyer questions across each engine on a schedule and recording four things per run: whether you were named, whether one of your sources carried the citation, which source earned it, and who was named instead. You cannot read this channel out of your analytics — when an engine declines to name you there is no impression to log. This covers building a prompt set that means something, why engines must never be blended into one score, what tools report natively, and where the chain to pipeline genuinely breaks.

Published

You measure AI visibility by running a fixed set of buyer questions across each engine on a schedule, and recording four things per run: whether you were named, whether one of your sources was the attached citation, which source earned it, and who was named instead. Those four, tracked per engine over repeated runs, are the whole measurement model. Everything else is either downstream of them or decoration.

What you cannot do is read this channel out of your analytics. That is the problem this article exists to solve, and it is worth being honest about how partial the solution is.

Why does normal analytics not work for this?

Because most of what happens leaves no trace on your site. A buyer asks an assistant for suppliers, reads a paragraph naming four companies, and forms a shortlist. If you were not named, there is no impression to log and no click that failed to happen. The absence has no telemetry at all.

Even when you are named, the referral trail is unreliable. Some assistants pass a referrer and some do not, many buyers read the answer and then type your name into a browser directly, and Cloudflare's crawl-to-refer analysis published in July 2025 shows how heavily some providers crawl relative to the traffic they send back. Sessions arriving from AI surfaces are a floor on your visibility, not a measure of it.

So the measurement has to be active. You have to go and ask the questions yourself, because nothing reports them to you.

What exactly should be measured?

Four metrics, and they answer genuinely different questions. Collapsing them into one number is the most common mistake in this category.

The four measures, what each is calculated from, and what it tells you
Metric How it is produced The question it answers
Answer share Share of prompts in the fixed set where your company is named at all, per engine Are you in the conversation?
Citation rate Share of those answers where a source of yours carries the attached citation Are you the source, or just a mention?
Cited surface Which specific URL or profile earned the citation Where does the next hour of work go?
Competitive set Who is named when you are not, and from which source What does the engine currently believe your category is?

The distinction between the first two carries most of the value. Being named is a mention the model produced from what it already knows; being cited means a specific page of yours was retrieved and attributed. The second is far more durable, and it is the one you can actually influence — the mechanism is in how AI engines actually choose which sources to cite.

How do you build a prompt set that means anything?

The prompt set is the instrument. If it changes between runs, you have no measurement — you have anecdotes with dates on them.

  1. Write questions a buyer would type, not keywords. Retrieval systems commonly rewrite a question into several related queries before searching, so exact-match phrasing matters less than question shape.
  2. Cover the buying stages. Category questions ("who supplies X in Y"), comparison questions ("X versus Y for Z"), problem questions ("how do I solve Z"), and verification questions ("is company X any good").
  3. Include questions you expect to lose. A set built only from prompts you already win is a marketing asset, not an instrument.
  4. Fix it in writing before the first run. Agreed, dated, and handed over. Prompts added later cannot be compared with earlier runs.
  5. Keep it small enough to actually re-run by hand. Twenty questions re-run monthly beats two hundred run once and abandoned.

Two rules that sound pedantic and are not. Never edit a prompt mid-programme — start a second set instead and report them separately. And record the date and engine version alongside every result, because a change in the engine is a far more common explanation for a swing than anything you did.

What does a real prompt set look like?

Concrete beats abstract here, so this is the shape I use for an industrial supplier, with the reasoning attached to each row rather than left implicit.

A worked prompt set: five question types, and why each earns its place
Type Example What it tests
Category Who supplies industrial vision sensors in Germany? Whether you exist in the engine's map of the category at all. The highest-value question, and usually the hardest to win
Specification Which vision sensors suit automotive assembly line inspection? Whether you are attached to the application rather than only to the product noun
Comparison How does [your company] compare to [named competitor]? What the engine believes about you relative to a rival, and which source it is reading that from
Verification Is [your company] a reliable supplier? What a buyer sees at the diligence step. Often the most alarming result, because it is answered largely from sources you do not own
Problem How do I reduce false rejects on an inspection line? Whether your content is retrieved at the stage before a supplier is even being considered

Five types, three to five prompts each, is a workable set for one category. Run the verification questions first if you only have time for one pass — they are the ones most likely to surface a problem you did not know you had.

Why must each engine be measured separately?

Because presence in one predicts very little about the others. ChatGPT, Perplexity, Google's AI Overviews, Gemini and Claude differ in where they retrieve from, how densely they cite, and how much they lean on classic search rankings underneath.

The size of that last difference is measurable. Ahrefs studied Google's AI Overviews across 863,000 keywords and 4 million cited URLs and reported in March 2026 that about 38% of cited pages also ranked in the top ten organically, down from roughly 76% in July 2025. That is one engine's relationship with one ranking system over eight months. Assuming it holds for the others is an assumption, not a finding — and a blended score across engines would have hidden the change entirely.

How do you handle the fact that answers change every time?

By treating the trend as the evidence and the single run as a sample. Answer engines are non-deterministic: the same prompt can return different answers to different users on the same day.

Three practices follow. Re-run the whole set on a fixed schedule rather than checking ad hoc. Require a change to persist across at least two consecutive runs before calling it a change. And accept that this cuts both ways — a single screenshot showing you cited is no more meaningful than one showing you absent, and any vendor presenting either as a result is telling you they have not run this repeatedly.

Are there tools that report this natively yet?

Partially, and it is worth knowing what is first-party rather than inferred.

Microsoft added an AI Performance view to Bing Webmaster Tools, announced in February 2026 as a public preview, reporting how a site appears in Copilot and Bing's AI answers. That is platform-reported data rather than sampled, which makes it materially more reliable than any external estimate for the surfaces it covers.

Beyond that, third-party trackers exist and they are all doing the same thing you would do by hand: sending prompts and parsing answers. That is legitimate, and it scales. But it does not become platform truth by being automated, and a tool that will not show you its prompt set is selling you a number you cannot audit.

How do you connect any of this to pipeline?

Imperfectly, and saying so is part of doing it properly. The question clients ask agencies most often is whether marketing performance connects to revenue — 55% of agencies report hearing it regularly, according to AgencyAnalytics' 2026 Benchmarks Report of 494 agency professionals — and reporting that stops at citation counts cannot answer it.

The chain I use, with the weak link named:

  1. Answer share and citation rate — measured directly, reliable, the strongest link.
  2. Referral sessions from AI surfaces — measurable but understated for the reasons above. A floor, not a total.
  3. A "how did you hear about us" field on your own enquiry form — self-reported and lossy, and the only place a buyer who read an answer and typed your name directly can ever be caught.
  4. Enquiry quality — whether enquiries mentioning an assistant convert differently from the rest.

Step three is where the chain is weakest, and no vendor can fix that with a dashboard. Anyone showing you clean attribution for this channel has built it from assumptions and not labelled them.

What does the conversion evidence actually show?

One case study, and it needs both halves reported together or it misleads in whichever direction you stop reading.

Seer Interactive published a client analysis in June 2025 showing ChatGPT referral traffic converting at 15.9% against 1.76% for Google organic — and noted in the same study that AI traffic amounted to roughly 0.07% of organic traffic over the period.

Quote the first number alone and you have a channel beating search ninefold. Quote the second alone and you have a rounding error. The accurate reading is both: a small, unusually well-qualified stream, because someone arriving from a generated answer has typically already had their question answered and their shortlist narrowed. It is one client, one period, and self-reported — not a benchmark for your business.

What is not worth measuring?

  • A single blended "AI visibility score." It averages away the per-engine differences that tell you what to do.
  • Raw mention counts with no source attached. Being named 14 times is not a finding; being cited from a profile you had forgotten about is.
  • Sentiment scoring of AI answers. Interesting, rarely actionable, and expensive to produce reliably.
  • Anything on a prompt set you cannot see. If you cannot audit the instrument, the reading is unfalsifiable.
  • Rankings as a proxy. The Ahrefs figure above is precisely the reason this no longer substitutes.

What should a report actually contain?

Per engine, per run, with the date: answer share, citation rate, the specific source behind every citation, the competitors named in your place, and a plain statement of what changed since the last run and what did not. Plus the prompt set itself, unchanged, attached to every report so the numbers can be checked.

If a report cannot be re-derived by someone else running the same prompts, it is not a measurement. That standard is unusual in marketing reporting and it is the entire reason this is worth doing carefully — the fuller argument is in what actually works in AEO, and what does not, and the questions to put to any vendor are in what to ask before you hire one.

What to do about your own situation

Take your baseline before changing anything. Write twenty buyer questions, run them across the five engines, and record the four measures above. It takes an afternoon, it costs nothing, and without it you will never be able to say whether anything you do afterwards worked.

If you would rather see one run done properly first, that is what the free check is. The service it belongs to is getting found when buyers ask AI.

Run a free AI visibility check

Frequently asked questions

How do you measure whether AI is citing your company?

By running a fixed, written prompt set across each engine on a schedule and recording four measures per run: answer share, citation rate, which specific source earned each citation, and which competitors were named in your place. The prompt set must not change between runs, or the numbers cannot be compared.

Why can't I just use Google Analytics for AI visibility?

Because most of what happens leaves no trace. If an engine does not name you there is no impression to log and no click that failed to happen. Even when you are named, some assistants pass no referrer and many buyers read the answer then navigate directly. Referral sessions from AI surfaces are a floor on your visibility, not a measure of it.

What is the difference between answer share and citation rate?

Answer share is the proportion of prompts where your company is named at all. Citation rate is the proportion of those answers where a source of yours carries the attached citation. Being named is a mention produced from what the model already knows; being cited means a specific page was retrieved and attributed. The second is more durable and is the one you can influence.

Should AI visibility be reported as a single score?

No. ChatGPT, Perplexity, Google's AI Overviews, Gemini and Claude differ in where they retrieve from and how densely they cite, so presence in one predicts little about the others. A blended average hides the only pattern worth acting on — which engine is missing you, and what it is reading instead.

Can AI visibility be connected to revenue?

Only imperfectly, and saying so is part of doing it properly. The chain runs from answer share and citation rate, which are measured directly, through referral sessions, which understate, to a self-reported field on your own enquiry form, which is lossy. That third step is the weak link and no dashboard fixes it. Anyone showing clean attribution for this channel has built it from unlabelled assumptions.

Find out what the engines currently say about you

Send me your company name and website. I run a set of buyer questions across ChatGPT, Perplexity, Gemini, Claude and Google's AI Overviews, and send back a recorded walkthrough of what came out: where you appeared, where you did not, who was named instead, and the two or three structural reasons why. No charge, no obligation, and you keep the findings whether or not you decide to hire me.

I run these myself, so there is a queue. Expect a few working days rather than an instant report.