AI Visibility Metrics Are Unreliable: Why We Track Reciprocal Rank Fusion Instead

For the last few years, brands and agencies have poured serious budget into tracking AI visibility. But new research suggests a lot of that spend is resting on a metric that barely holds up to scrutiny.

Rand Fishkin, writing on SparkToro’s blog in research conducted with Patrick O’Donnell of Gumshoe.ai, ran 600 volunteers through 12 prompts across ChatGPT, Claude and Google AI, a combined 2,961 times. The finding was that there’s less than a 1 in 100 chance that two runs of the same prompt return the same list of brands, and less than a 1 in 1,000 chance they return it in the same order. Interestingly, Claude was slightly more likely than ChatGPT or Google AI to repeat the same list, but even less likely to repeat it in the same order.

So it’s not that AI responses are a bit inconsistent. A single sampled response tells you almost nothing on its own.

It felt like a big revelation when it landed, and was rightly shared widely. And then, most of us went back to our AI tracking tools and carried on as before. I think that’s because AI trackers give us something clean to measure: a citation count, a ranking position, a tidy chart for a client report.

That’s not a reason to stop tracking responses. We still do, via Ebb, our new AI visibility tracking tool, but we treat them as directional samples rather than absolutes. And it’s why, alongside standard visibility metrics, we lean on one that sits upstream of the response itself: Reciprocal Rank Fusion, or RRF.

Where RRF comes in

Instead of chasing a position in one sampled answer, RRF looks at your ranking for the fan-out queries behind a prompt, the underlying, SEO-style search signal that exists before the LLM ever generates the text you’d screenshot for a client deck. It isn’t a workaround for the problem the study exposed. It’s a different layer entirely, sitting further back in the pipeline, where there’s more confidence in what’s actually being measured, because at its core, it’s a long-established SEO metric.

Put plainly, a Google or Bing ranking for a fan-out query is likely to be more stable and consistent than a single AI response.

If your AI visibility reporting is built entirely on a handful of sampled responses, unfortunately you might just be reporting on noise.

How we’re approaching it

This is why Ebb doesn’t stop at tracking citations and mentions. It measures Reciprocal Rank Fusion alongside standard visibility metrics, so reporting isn’t resting on a single, potentially unrepresentative AI response. It’s built on a more stable, SEO-grounded signal underneath it.

If you want to understand what’s actually driving your AI visibility, get in touch. You can also sign up for the Ebb waitlist here, ahead of wider roll out later this year.

Insights