SEO Analytics & Reporting
AI Visibility KPIs: Citation Rate, Share of Voice & More
Your brand shows up in a ChatGPT answer, and your monthly report still only counts Google clicks. That gap is why most teams can’t measure AI search. These are the KPIs that close it, what to track, how to calculate each, and the measurement discipline that keeps the numbers from fooling you.
By Rahul Saini, Author at Search Counsel Co. Last updated [July] 2026.
Featured answer: what are AI visibility KPIs?
They measure whether AI answers mention, cite, and recommend your brand, since a click never happens. The core set: inclusion rate (how often you’re mentioned), citation share (how often your content is the source), AI share of voice (your slice versus competitors), and AI referral traffic. Google rankings don’t predict any of them, so you measure AI directly, per platform, on a fixed set of prompts.
Your Google rankings no longer predict your AI citations. The share of AI Overview citations coming from top-ten pages fell from around three-quarters to under 40% in under a year. You can’t use where you rank as a stand-in for whether AI recommends you. You have to measure AI visibility directly, and these are the numbers that do it.
Inclusion rate
Are you named?
How often AI answers mention your brand at all.
Citation share
Are you the source?
How often AI links to your content as evidence.
AI share of voice
Your slice
Your mentions versus competitors’ in the same answers.
AI referrals
Did they come?
The visits AI platforms actually send you.
What’s in this guide
Mention vs citation
Start here, because almost every AI dashboard blurs these two, and they’re different signals. A mention is when an AI names your brand in its answer: “companies like yours offer this.” A citation is when the engine links to your content as the source, treating it as evidence for a claim. A brand can have strong share of voice and near-zero citation share at the same time, mentioned everywhere, sourced nowhere; one analysis of roughly 600,000 citation events found brands sitting at around 40% share of voice yet only about 5% citation share. Mentions tell you the models know you exist; citations tell you they trust your content enough to send buyers to it. In SimilarWeb’s Sephora case study, citation share landed near 16% across 179 beauty prompts and clustered on high-intent, transactional queries, exactly where trust converts. Track both, separately, or you’ll mistake awareness for authority.
The KPIs that matter
Six metrics cover it. Each has a plain definition and a formula, and where a benchmark exists, treat it as directional:
| Metric | What it measures | How to calculate |
|---|---|---|
| Inclusion / mention rate | How often you’re named at all | (queries mentioning you ÷ total queries) × 100 |
| Citation share | How often your content is the source | (your citations ÷ all citations in the set) × 100 |
| AI share of voice | Your mentions versus competitors | (your mentions ÷ all brand mentions) × 100 |
| Citation position | Where you appear in the answer | Order of your first mention or citation |
| Sentiment | How the AI frames you | Positive, neutral, or negative |
| AI referral traffic | Visits AI actually sends | Sessions from AI sources in GA4 |
Two of these carry outsized weight. Inclusion rate for an unoptimized brand often sits near zero, and reaching roughly a quarter to a third of your priority queries is a strong target. And position matters as much as presence: first-position citations earn on the order of four to five times the click-through of fifth-position ones, so a brand cited in 40% of prompts but always buried near the bottom is collecting impressions, not visits. Weight your citations by position, and never report a citation number without stating the denominator, “0.5% of a 48,000-citation window” is a measurement, “12% share of voice” with no denominator is marketing.
What good looks like
Benchmarks in AI search are young and vary widely by industry, query intent, and dataset, so treat these as directional reference points, not targets to promise. As a rough maturity ladder for category (non-branded) queries, several 2026 analyses land in a similar range:
| Stage | Citation rate on category queries |
|---|---|
| Unoptimized / just starting | Roughly 0 to 5% |
| Gaining traction | Roughly 8 to 15% |
| Strong / established | Roughly 20 to 35% |
| Category leader | Roughly 35 to 50% |
| Your own branded queries | Aim for 90%+ , so competitors don’t hijack your name |
Two cautions on those numbers. First, the gap between leaders and laggards is enormous: in one B2B SaaS sample the top quartile earned roughly 31 citations a month across major platforms against about 3.7 for the bottom quartile, an eight-fold spread, so a mid-table position can feel closer to zero than to the top. Second, a strong category-query rate means little if your branded queries leak; if an AI attributes your content or product to a competitor, that’s usually an entity-clarity problem to fix, not a content-volume one. Judge yourself against your own baseline and your named competitors on a frozen prompt set, never against a headline number from someone else’s industry.
Selection, credibility, outcome
The six sort cleanly into three tiers, which mirrors the wider KPI framework this hub uses:
- Selection, are you chosen? Inclusion rate, citation share, and AI share of voice. Whether you’re in the answer at all.
- Credibility, why chosen or skipped? Citation position, sentiment, and whether the model links to you or just names you. A citation in a negative context does more harm than no citation.
- Outcome, what does it earn? AI referral traffic and the conversions it drives. This is your North Star for the channel; the others are the drivers that move it.
The measurement discipline
This is what separates a real number from a vendor’s vanity chart. AI visibility is noisy, and without discipline you’ll measure noise:
- Measure per platform, never blended. Each engine is a separate game. The same page can be cited on ChatGPT and invisible on Perplexity, and the platforms differ enormously in how often they name brands at all, one large sample found Claude mentioning a brand in the high nineties percent of responses, ChatGPT in the low seventies, and Perplexity under half but linking sources in most answers. Some models hedge to “various options” while others name brands directly, so comparing a citation rate on one to another without normalizing is comparing different sports. Aggregating across models hides exactly the differences you need to act on.
- Freeze your prompt set. Citation share only compares across time if the prompts don’t change. Swap the prompts and you’ve changed the denominator, and your trend line becomes fiction.
- Choose golden prompts, weighted to solution queries. Build 50 to 100 validated prompts, and favor “what’s the best way to solve this” over “tell me about my brand.” Branded queries only confirm the model recognizes you; solution queries reveal whether it recommends you to buyers who haven’t picked a vendor yet.
- Run each prompt many times. AI answers are a distribution, not a fact. The response to the same query changes a large share of the time, and only a minority of brands stay visible across repeated runs, so a single snapshot is noise. Run each prompt ten or more times before you trust a rate, and treat small gaps between you and a competitor as possible sampling variance, not real difference.
- Report weekly or monthly, not daily, with a margin. Measure continuously, but never present a day-over-day delta from a small prompt set as a trend. Because the underlying data is probabilistic, report a range rather than a false point, “citation share is 24%, plus or minus 3” tells the truth that “24%” doesn’t.
Why rankings aren’t a proxy
You might hope your Google data can stand in for AI visibility. It can’t, and the reason is worth internalizing. The share of Google’s own AI Overview citations coming from top-ten ranked pages fell from roughly three-quarters to under 40% in under a year, and only a small fraction of pages ranking in Google’s top ten show up in ChatGPT’s citations at all. Google’s own AI is pulling away from its own rankings as a source pool, so your rank tells you little about your citation footprint. This decoupling is the same one behind tracking SERP features and rank tracking in an AI world.
There’s a second trap that pushes teams to underrate AI entirely. Most AI-driven visits arrive with no referrer and land in your analytics as direct traffic, by one estimate the majority of them, so a team glances at the trickle of visits tagged from chatgpt.com and concludes “AI doesn’t send us traffic.” That’s a false negative, and it’s the strongest argument for measuring visibility directly rather than inferring it from referral data. The method for recovering as much of that AI traffic as you can is in tracking ChatGPT and Perplexity referrals. And once you’ve measured these KPIs, the work of actually improving them, the content and entity signals that earn citations, lives in our AI visibility guide.
FAQ
What’s the difference between a mention and a citation in AI search?
A mention is when an AI names your brand in its answer; a citation is when it links to your content as the source for a claim. They’re separate signals, and a brand can be mentioned often while rarely being cited. Mentions reflect awareness; citations reflect trust and send traffic, so track them as two different metrics.
How do you calculate AI share of voice?
Divide your brand’s mentions by the total brand mentions across your tracked prompt set, then multiply by 100. It’s a competitive metric, not an absolute one, so it only means something against a fixed set of prompts and specific competitors. Always calculate it per platform, since your share on ChatGPT and Perplexity can differ dramatically.
Why do I need to run the same prompt multiple times?
Because AI answers vary from run to run for the same query, often substantially, and only a minority of brands stay visible across repeated runs. A single response is a snapshot of noise, not a reliable measurement. Running each prompt ten or more times gives you a stable rate and stops you from mistaking random variation for a real change.
Can I use my Google rankings to estimate AI visibility?
No. The link between ranking and AI citation has weakened sharply, with the share of AI Overview citations from top-ten pages falling by roughly half in under a year, and only a small fraction of top-ranked pages appearing in ChatGPT citations. Rankings and AI citations are now largely separate, so you have to measure AI visibility on its own.
What’s a good AI inclusion rate?
For a brand that hasn’t optimized for AI, near zero is common. Reaching roughly a quarter to a third of your high-priority query clusters is a strong target. Because the metric is competitive and platform-specific, judge it against your competitors on a frozen prompt set rather than a universal number, and track the trend over time.
The one thing to do next
Pick 20 to 50 solution-focused prompts your buyers would actually ask, run each ten times across ChatGPT, Perplexity, and Google’s AI, and log three things per answer: were you mentioned, were you cited, and in what position. Freeze that prompt set and repeat it monthly. That baseline is your AI scoreboard, and it plugs straight into your reporting dashboard. Improving the numbers is the next job, covered in our AI visibility guide.
Sources used
Metric definitions, formulas, and the measurement discipline draw on analysis from CitedMe, DigitalApplied, Nick Lafferty, authoritytech, and Nuwtonic, among others. The benchmarks (the citation-rate maturity ladder and top-versus-bottom-quartile gap from Data-Mania and authoritytech, the mention-versus-citation gap and Sephora figures from SimilarWeb, per-platform mention rates, citation-position effects, and the ranking-to-citation decoupling from Ahrefs and Moz) and the estimate that most AI traffic arrives unattributed are attributed directionally, and several rest on single studies, so each should be re-verified and kept flagged as an estimate before publishing. These numbers vary widely by industry and dataset and move quickly, so present them as reference ranges, not targets.
