AI Visibility Isn't One Number

By Jeff Jenkins

AI Visibility Isn't One Number

There is no single AI visibility score. Testing five platforms across ChatGPT, Claude, Gemini, and Perplexity with two measurement systems, every metric told a different story. "AI visibility" is really six different questions, and most teams are answering the wrong one without knowing it.

<p><em>Five platforms, four AI engines, two measurement systems. Every metric told a different story.</em></p><p>There is no single AI visibility score. There are several, and they are all real. The problem is not that any one of them is wrong. It is that none of them is <em>the</em> score, and treating any one that way quietly answers a question you never asked.</p><p>I found this out by accident. I set out to compare a couple of measurement tools, expecting to learn which was more accurate. Instead, I learned they were not measuring the same thing.</p><h2><strong>How I Got Here</strong></h2><p>I have been using CARL, the AI visibility platform from Xponent21, to track something most reporting still misses: who AI actually <em>recommends</em> when a buyer asks. Not who ranks. Not who gets cited. Who gets named in the answer.</p><p>I lined CARL up against the tools I already run, including Ahrefs citation data and traditional search reporting, to see how the numbers compared. They did not. One was telling me who gets recommended. Another was telling me who gets cited. A third was telling me who gets clicked. I had been treating those as one number.</p><p>They are not one number. They are not even close.</p><h2><strong>Six Metrics Wearing One Name</strong></h2><p>"AI visibility" is not a metric. It is a category of metrics, and each one answers a different question about a different moment in the buying journey.</p><table style="min-width: 75px;"><colgroup><col style="min-width: 25px;"><col style="min-width: 25px;"><col style="min-width: 25px;"></colgroup><tbody><tr><td colspan="1" rowspan="1"><p><strong>Metric</strong></p></td><td colspan="1" rowspan="1"><p><strong>The question it answers</strong></p></td><td colspan="1" rowspan="1"><p><strong>Where it can mislead you</strong></p></td></tr><tr><td colspan="1" rowspan="1"><p>Prompt-level recommendations (CARL)</p></td><td colspan="1" rowspan="1"><p>Who does AI recommend when a buyer asks?</p></td><td colspan="1" rowspan="1"><p>Swings on the exact prompt and the model</p></td></tr><tr><td colspan="1" rowspan="1"><p>Domain citations (Ahrefs)</p></td><td colspan="1" rowspan="1"><p>Whose site does AI cite as a source?</p></td><td colspan="1" rowspan="1"><p>A brand can be cited constantly and never be the recommendation</p></td></tr><tr><td colspan="1" rowspan="1"><p>Google AI Overviews</p></td><td colspan="1" rowspan="1"><p>Who influences Google's AI answer?</p></td><td colspan="1" rowspan="1"><p>Overweights informational content; presence is not endorsement</p></td></tr><tr><td colspan="1" rowspan="1"><p>Referral traffic (GA4)</p></td><td colspan="1" rowspan="1"><p>Who actually gets the click from AI?</p></td><td colspan="1" rowspan="1"><p>Often mislabeled as "direct" or "other," so it undercounts</p></td></tr><tr><td colspan="1" rowspan="1"><p>Search impressions (Search Console)</p></td><td colspan="1" rowspan="1"><p>Who gets seen in results?</p></td><td colspan="1" rowspan="1"><p>Impressions are not clicks, and clicks are not recommendations</p></td></tr><tr><td colspan="1" rowspan="1"><p>Rankings (traditional SEO)</p></td><td colspan="1" rowspan="1"><p>Who is findable on Google?</p></td><td colspan="1" rowspan="1"><p>A number one ranking can lose the click to the AI answer above it</p></td></tr></tbody></table><p>The recommendation layer, the top row, is the newest and the least understood. Most teams are still measuring citations and traffic because those are the numbers we have always had. But "who gets recommended" is a different question, and it may be the layer closest to the shortlist decision as buyers increasingly use AI to build their shortlists.&nbsp;</p><p>That is the layer CARL is built to measure, and it is the one most dashboards skip entirely.</p><p>Read the list again. "Who gets cited," "who gets recommended," "who gets clicked," and "who gets found" are four different questions about four different moments. A brand can win one and lose the rest.</p><h2><strong>What I Found</strong></h2><p>I happened to run this on five fintech platforms because that is the category I work in, but nothing about the pattern is specific to fintech. Here is the whole picture on one screen.</p><table style="min-width: 175px;"><colgroup><col style="min-width: 25px;"><col style="min-width: 25px;"><col style="min-width: 25px;"><col style="min-width: 25px;"><col style="min-width: 25px;"><col style="min-width: 25px;"><col style="min-width: 25px;"></colgroup><tbody><tr><td colspan="1" rowspan="1"><p><strong>Brand</strong></p></td><td colspan="1" rowspan="1"><p><strong>CARL (recommended)</strong></p></td><td colspan="1" rowspan="1"><p><strong>Ahrefs (cited)</strong></p></td><td colspan="1" rowspan="1"><p><strong>ChatGPT</strong></p></td><td colspan="1" rowspan="1"><p><strong>Claude</strong></p></td><td colspan="1" rowspan="1"><p><strong>Gemini</strong></p></td><td colspan="1" rowspan="1"><p><strong>Perplexity</strong></p></td></tr><tr><td colspan="1" rowspan="1"><p>Stripe</p></td><td colspan="1" rowspan="1"><p>95%</p></td><td colspan="1" rowspan="1"><p>5,151</p></td><td colspan="1" rowspan="1"><p>100%</p></td><td colspan="1" rowspan="1"><p>100%</p></td><td colspan="1" rowspan="1"><p>79%</p></td><td colspan="1" rowspan="1"><p>100%</p></td></tr><tr><td colspan="1" rowspan="1"><p>Marqeta</p></td><td colspan="1" rowspan="1"><p>66%</p></td><td colspan="1" rowspan="1"><p>37</p></td><td colspan="1" rowspan="1"><p>71%</p></td><td colspan="1" rowspan="1"><p>71%</p></td><td colspan="1" rowspan="1"><p>64%</p></td><td colspan="1" rowspan="1"><p>57%</p></td></tr><tr><td colspan="1" rowspan="1"><p>Unit</p></td><td colspan="1" rowspan="1"><p>57%</p></td><td colspan="1" rowspan="1"><p>31</p></td><td colspan="1" rowspan="1"><p>71%</p></td><td colspan="1" rowspan="1"><p>57%</p></td><td colspan="1" rowspan="1"><p>71%</p></td><td colspan="1" rowspan="1"><p>29%</p></td></tr><tr><td colspan="1" rowspan="1"><p>Galileo</p></td><td colspan="1" rowspan="1"><p>46%</p></td><td colspan="1" rowspan="1"><p>21</p></td><td colspan="1" rowspan="1"><p>71%</p></td><td colspan="1" rowspan="1"><p>29%</p></td><td colspan="1" rowspan="1"><p>57%</p></td><td colspan="1" rowspan="1"><p>29%</p></td></tr><tr><td colspan="1" rowspan="1"><p>Lithic</p></td><td colspan="1" rowspan="1"><p>39%</p></td><td colspan="1" rowspan="1"><p>2</p></td><td colspan="1" rowspan="1"><p>0%</p></td><td colspan="1" rowspan="1"><p>57%</p></td><td colspan="1" rowspan="1"><p>57%</p></td><td colspan="1" rowspan="1"><p>43%</p></td></tr></tbody></table><p><em>CARL is prompt-level recommendation share. Ahrefs is total domain citations across AI answers. The four model columns are CARL mention rates per engine. Snapshot, US, August 2026.</em></p><p>The columns do not line up, and that is the point. If every column told the same story, we would not need six metrics.</p><p>The leaders were not surprising. Stripe topped every column, exactly as you would expect. The disagreements below it are where the story is.</p><p>Lithic is the clearest case. On CARL's prompt recommendations, it scored a respectable 39 percent, close to Galileo. On Ahrefs domain citations, it registered just 2, an order of magnitude below Galileo's 21.&nbsp;</p><p>One system says Lithic is a credible presence. The other says it is nearly invisible. Both are correct. They are answering different questions, and if you only looked at citations, you would conclude Lithic barely exists in AI.&nbsp;</p><p>At the same time, the recommendation data says it is showing up in conversations more than its footprint would suggest.</p><img src="https://access.discoveraio.com/storage/v1/object/public/public-assets/email-images/ai_visibility_one_number-2.png" alt="" class="editor-image max-w-full h-auto rounded-lg cursor-pointer transition-all hover:opacity-80" draggable="false" style="max-width: 100%; height: auto;"><h2><strong>The Model Matters as Much as the Metric</strong></h2><p>The sharpest finding was not about the tools at all. It was about the models.</p><ul><li><p><strong>Lithic:</strong> 0 percent in ChatGPT, 57 percent in both Claude and Gemini.</p></li><li><p><strong>Galileo:</strong> 71 percent in ChatGPT, 29 percent in Claude.</p></li><li><p><strong>Unit:</strong> 71 percent in ChatGPT and Gemini, 29 percent in Perplexity.</p></li></ul><p>Same brand. Same day. Same prompt. The only thing that changed was which model answered, and the verdict flipped from invisible to dominant.</p><p>This is not a measurement artifact you can tool your way around. Different models retrieve, weigh, and recommend differently, so "AI visibility" is not even one number per tool. It is one number per model, per question.</p><h2><strong>Why the Tools Disagree, and Why That Is the Point</strong></h2><p>It would be easy to conclude that one of these tools is wrong. That is the wrong conclusion.</p><p>They diverge because they are observing different parts of the same buying journey. Citation data watches what AI reads. Recommendation data watches what AI says. Referral data watches what the buyer does next.</p><p>The disagreement does not necessarily mean one of them is inaccurate. Each is observing a different part of the journey, and the gap between them is information, not error. When two instruments disagree, the gap itself is telling you something about the thing you are measuring.</p><h2><strong>What This Means for the Rest of Us</strong></h2><p>You cannot manage what you have not defined. If your team reports a single "AI visibility" number, you are almost certainly optimizing for a question you did not choose on purpose. The fix is not a better tool. It is a better question.</p><p>Before you open any dashboard, decide which decision you are trying to influence. Are you trying to get cited as a source, get recommended on a shortlist, or get the click? Those three goals lead to three different bodies of work.</p><p>The next time someone on your team asks, "What is our AI visibility score?" do not answer with a number. Answer with a question. <strong>Which decision are we trying to influence?</strong></p><p>That question is where the real work starts. It is the difference between watching a dashboard and having a strategy.</p><p></p><hr><p></p><p><em>Thanks to</em><a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://www.linkedin.com/in/willmeltonconsultant/"><em> <u>Will Melton</u></em></a><em> and the team at Xponent21 for building CARL and for making Discover AIO a place where practitioners can publish research like this under their own names, and to</em><a target="_blank" rel="noopener noreferrer nofollow" class="text-primary underline" href="https://www.linkedin.com/in/garry-callis-jr-wlf-b802852ab/"><em> <u>Garry Callis Jr.</u></em></a><em> for running a community that keeps pushing on where AI search is actually heading.</em></p><p><em>Jeff Jenkins works on AI visibility for fintech and payments companies. He is a regular contributor to Discover AIO.</em></p>