AI Visibility KPIs Every CMO Should Track
"Are we visible in AI?" is not a reportable question. This is a set of six metrics that make it one — each defined so that a board can follow the trend and a team can act on the number.
TL;DR
| KPI | Question it answers |
|---|---|
| Share of Answer | How often are we named in category answers? |
| Unaided vs. aided visibility | Are we discovered, or only recognised once named? |
| Competitive answer set | Who are we actually being compared against? |
| Citation source mix | Which properties drive our inclusion? |
| AI referral quality | What happens when those visitors arrive? |
| Description accuracy | Do assistants describe us correctly? |
Track the trend on each. A single reading tells you almost nothing.
1. Share of Answer
What it measures: the proportion of AI responses to category-relevant questions that mention your brand.
Why it is the headline: it is the closest available analogue to share of voice for the answer era, and it moves in response to work you actually do.
How to collect it: run a fixed set of category questions across several assistants on a fixed schedule, and record whether you are named. Keep the question set stable — changing it resets your trend.
How to read it: direction over absolute value. A category with two dominant vendors produces very different numbers from a fragmented one, so cross-industry benchmarks are close to meaningless.
2. Unaided vs. aided visibility
What it measures: two distinct things that most teams wrongly combine.
- Unaided — you are never named in the prompt. "What are the best tools for X?" This is genuine discovery.
- Aided — you are named. "How does Acme compare to Beta?" This is characterisation once someone already knows you.
Why it matters: aided visibility is nearly always far higher, because naming you in the prompt guarantees you appear. Blending the two produces a flattering number that conceals a discovery problem.
How to read it: unaided is the growth metric. Aided is the reputation metric. Report them separately, always.
3. Competitive answer set
What it measures: which brands are named alongside you, and how often.
Why it matters: your AI competitive set is frequently not your sales competitive set. Teams routinely discover they are being compared against companies they have never encountered in a deal.
How to collect it: log every brand mentioned in every response. Rank by frequency.
How to read it: the brands appearing most often are your real positioning competitors in the eyes of the systems your buyers consult, whatever your internal battlecards say.
4. Citation source mix
What it measures: which properties assistants cite when answering your category's questions, and how often your own domain appears among them.
Why it matters: this is the most directly actionable metric on the list. It converts "improve our AI visibility" into a specific list of properties to pursue.
How to collect it: log every source cited in every response. Perplexity is the easiest place to observe this because it always shows sources.
How to read it: if your domain rarely appears but three review platforms dominate, your investment belongs on those platforms, not on another blog post.
5. AI referral quality
What it measures: volume and behaviour of visitors arriving from assistants.
How to collect it: a custom channel group in GA4 filtering on assistant hostnames.
How to read it: ignore volume as a headline. Volume is small and structurally undercounted, because referrers get stripped and most AI interactions end without a click. Report engagement rate and conversion rate instead, which are typically stronger than organic search.
The caveat to state explicitly: this metric counts only people who clicked. It systematically undercounts AI's influence. Never present it as the measure of AI visibility — it is the measure of AI click-through.
6. Description accuracy
What it measures: whether assistants describe your company correctly.
Why it matters: being mentioned inaccurately can be worse than not being mentioned. Wrong category, wrong pricing, wrong positioning, or conflation with a similarly named company all cause direct damage.
How to collect it: ask several assistants to describe your company, on a schedule, and score each response for factual accuracy.
How to read it: errors that persist point at entity problems — inconsistent facts across your properties, or a name collision.
What not to report
- Raw AI referral sessions as a headline. Structurally undercounted; invites the wrong conclusion.
- A single point-in-time Share of Answer. Meaningless without a trend.
- Aided visibility presented as visibility. Flattering and misleading.
- Cross-industry benchmarks. Category concentration dominates the number.
- Mention counts without sentiment or accuracy. Being named as a cautionary example still counts as a mention.
A reporting cadence that works
Monthly: Share of Answer trend (unaided), competitive answer set changes, AI referral quality.
Quarterly: citation source mix, description accuracy audit, progress against the third-party properties you targeted.
Annually: re-derive the question set. Buyer language moves, and a stale question set measures a market that no longer exists.
Frequently asked questions
What is a good Share of Answer? There is no universal figure. Judge against your own trend and your named competitors.
How many questions should the set contain? Thirty is a workable minimum; fifty is better. Consistency between runs matters more than size.
How often should we measure? Monthly for board reporting; weekly or fortnightly for the team acting on it.
Can we do this without a platform? Yes for a baseline. It becomes impractical as a continuous discipline — the run volume compounds quickly across questions, assistants, and repetitions.
Published by the SIQA Editorial Team.