AI visibility tracking answers a question that traditional analytics cannot: when an assistant describes your category, are you in the answer, and what does it say about you? Because these conversations happen off your property and produce no session, none of it appears in your existing reporting.
The discipline is young and the measurement is genuinely harder than rank tracking, for a reason worth understanding before you buy anything. Generated answers are probabilistic. The same question can return different sources on different days, for different users, in different regions. A measurement practice that ignores this produces confident numbers that mean very little.
Key takeaways
- Answers vary between runs, so measure rates across repeated samples rather than positions.
- Track three things at minimum: mention presence, cited link presence, and which competitors appear.
- A stable prompt set is more valuable than a large one, because comparability is what creates signal.
- Include problem-framed prompts that never name your brand, since those reveal category visibility.
- Treat commercial attribution as partial and directional, and say so when reporting it.
Why this cannot be measured like rankings
A rank tracker works because search results are relatively stable and ordered. Ask the same query twice and you generally get the same list. Generative systems do not offer that stability. Sampling, model updates, retrieval variation, personalization, and regional differences all change the output for identical input.
This has a direct methodological consequence. One observation is a sample, not a measurement. If you check a question, see a competitor, and conclude you have lost visibility, you have learned almost nothing. Repeat the same question ten times across a week and count how often each brand appears, and now you have something worth acting on.
It also means precision claims should be treated skeptically. Any tool or report presenting your AI visibility as a single exact figure without describing its sampling method is asserting more certainty than the underlying system supports.
Building a prompt set that produces signal
Start from real questions rather than keywords. The input to an assistant is usually a full question or a described situation, not a two-word phrase, so a keyword list converts poorly. Interview your sales and support teams for the way prospects actually describe the problem.
Cover four categories deliberately. Problem-framed questions that never mention any brand reveal whether you appear in the category at all. Category questions that ask for options reveal whether you are on the shortlist. Comparison questions reveal how you are positioned against named competitors. Brand questions reveal whether the assistant describes you accurately.
Then freeze the set. The value of this measurement comes from comparability across time, and a prompt set that changes every month cannot show a trend. Add prompts to a stable core rather than rewriting it, and record the date any addition was made.
- Problem-framed: how do I stop my published content from going stale
- Category: what tools handle SEO content production for a small team
- Comparison: how does one named platform compare with another
- Brand: what does this specific company do and who is it for
- Keep the core set fixed so week-over-week comparison remains valid
The metrics worth recording
Mention rate is the share of runs in which your brand appears at all. It is the primary measure of category presence and the one most worth improving first, because being absent from the answer is a different problem from being present but unlinked.
Citation rate is the share of runs in which you appear as a linked source rather than as an unlinked name. The gap between mention rate and citation rate tells you whether the system trusts your own pages or is describing you through third-party sources.
Share of voice against a fixed competitor set turns the measure into something a leadership team can interpret. Sentiment and accuracy of the description matter too, and are frequently more actionable than either rate: an assistant that consistently misstates your pricing or your supported platforms is causing measurable commercial damage that no amount of additional visibility fixes.
Tooling, and when to build it yourself
A market of dedicated tools now exists, including platforms such as Profound, Otterly, Peec AI, and brand-monitoring features inside established SEO suites. They differ in which assistants they cover, how they sample, whether they support multiple regions, and how they price. Coverage of the specific assistants your buyers use matters more than the length of the feature list.
Ask any vendor three questions before buying: how many times is each prompt run, how is regional and personalization variance handled, and is the raw answer text retained so you can audit what was actually said. A tool that reports a score without exposing the underlying answers cannot be verified.
For a narrow set of high-value questions, a manual process is a legitimate starting point. A spreadsheet, a fixed schedule, and a disciplined recording format will tell you whether the problem is worth spending on before you commit to a subscription.
Turning findings into work
Absent from problem-framed answers usually indicates a coverage gap: you have not published a credible, retrievable source on the question. The response is content, not tactics.
Mentioned but never cited usually indicates a trust or retrievability problem with your own pages. Check crawler access first, then check whether your material makes verifiable, specific claims or merely asserts quality.
Described inaccurately usually traces to third-party sources rather than to your site. Find what the assistant is likely reading, correct what you legitimately can, and publish clear public documentation of the facts being misstated.
Present but positioned poorly against competitors is a messaging problem. Look at what the answer says your competitors are good at, and ask whether your own public material makes an equally concrete claim about the thing you are genuinely better at.
Common measurement mistakes
Checking once and concluding something is the most frequent error, and it produces both false alarms and false comfort. A competitor appearing in a single answer is not a trend, and your own appearance in one answer is not a position you hold.
Changing the prompt set to chase better numbers is the second. Once you edit the questions, the comparison to previous weeks is void, and a rising score may reflect nothing more than easier prompts. Freeze the core and log every addition with a date.
Measuring only branded prompts is the third, and it is comfortable precisely because it flatters. Asking an assistant to describe your company will usually produce a mention of your company. The questions that matter commercially are the ones a buyer asks before they know you exist.
Reporting a single composite visibility score without its method is the fourth. If nobody on the receiving end can say how many runs produced the number or which assistants were sampled, the figure will eventually be challenged and will not survive the challenge.
FAQ
Questions about this guide
How often should I run AI visibility checks?
Weekly or fortnightly, with multiple runs per prompt on each occasion. Consistency of method matters more than frequency, since the goal is a comparable trend rather than a snapshot.
Why do I get different answers to the same question?
Generated responses are probabilistic and affected by sampling, retrieval variation, personalization, region, and model updates. This is expected behavior and the reason measurement must be based on rates across repeated runs.
Can I see AI assistant traffic in my analytics?
Partially. Referral traffic from assistants that link out is visible, but answers consumed without a click leave no trace. Combine referral data with branded search trend and self-reported attribution for a fuller picture.
Is a dedicated tool necessary?
Not initially. A fixed prompt set, a schedule, and a disciplined spreadsheet will establish whether the gap is material. Buy tooling when the manual process becomes the constraint rather than before.