MONITORING · 5 ENGINES
GEOscanAI

geo

How Often Do AI Models Change Their Answers? The Case for a Volatility Index | GEOscanAI

8 min read
How Often Do AI Models Change Their Answers? The Case for a Volatility Index | GEOscanAI
How Often Do AI Models Change Their Answers? The Case for a Volatility Index | GEOscanAI

AI answers vary run to run even when nothing about a brand changed. A volatility index measures that baseline noise so real signal is easier to trust.

Ask an AI engine the same brand question on a Tuesday and again on a Friday, and there's a real chance you'll get two different answers, not because anything about your brand changed, but because the model's own volatility did. That variability is rarely discussed, and almost never measured, which is exactly the gap a volatility index is meant to fill.

Two Different Kinds of Change, Often Confused

It's worth separating two things that both show up as "my AI visibility changed" but mean very different things. The first is real, earned change, your brand's actual visibility shifted because you published new content, fixed a technical issue, or a competitor's signals moved. The second is noise, the model gave a different answer to the exact same question, asked the exact same way, for reasons that have nothing to do with anything you did.

Most brands, when they notice a shift, assume it's the first kind. Sometimes it is. But without a sense of how much baseline variability an engine has on its own, it's genuinely hard to tell the difference between a real signal and ordinary noise, and reacting to noise as though it were signal wastes effort and produces confusing, contradictory conclusions about what's working.

Why Answers Vary Even When Nothing About You Changed

Generative models aren't deterministic in the way a search index is. The same prompt, run multiple times, can produce different phrasing, a different order of mentioned brands, or even a different set of brands entirely, purely as a function of how the model samples its response. Add to that ordinary infrastructure-level factors, which version of a model is actually serving a given request, whether a retrieval step returned slightly different results this time, and you get a system where some amount of answer-to-answer variation is simply built in, independent of anything happening in the real world.

This isn't a flaw exactly, it's closer to a known property of how these systems generate text. But it means that a single before-and-after comparison, one answer before you made a change, one answer after, is a weak way to judge whether the change actually worked. You need enough repeated samples to separate the signal from the model's own baseline noise.

What a Volatility Index Would Actually Measure

A volatility index, as we think about it, would track how much a given engine's answers to a fixed prompt set change over a fixed period, independent of any known change on the brand side. Run the same prompts daily or weekly, hold everything about the brand constant, and measure how much the resulting mentions, ordering, and framing actually move. A low-volatility engine gives you a stable baseline, where a change in your numbers is more likely to reflect something real. A high-volatility engine requires more samples, and more caution, before drawing conclusions from any single shift.

This is different from tracking your own visibility trend, which measures your position over time. A volatility index measures the engine's own instability, a property of the tool you're using to measure, not of the thing being measured.

Why This Matters More Than It Seems

Without some sense of baseline volatility, teams risk two opposite mistakes. The first is chasing noise, seeing a dip in a single check and launching a reactive scramble to fix something that was never actually broken, only for the number to bounce back on its own the following week. The second, arguably more costly, is missing a real signal because it looks similar in size to ordinary noise, dismissing an actual, earned decline as probably just normal variation when it wasn't.

Both mistakes come from the same root cause: treating every engine as equally stable, when in practice some engines appear to vary noticeably more than others, run to run, for reasons that have nothing to do with any individual brand.

How Volatility Differs by Engine

Based on repeated testing across a range of prompts, retrieval-heavy engines tend to show more run-to-run variation than trained-knowledge-heavy engines, which makes intuitive sense: a live retrieval step introduces an additional source of variability, whatever page happened to rank highest at that exact moment, on top of the model's own generation variance. Engines leaning more on stable trained knowledge tend to give more consistent answers to the same question across repeated runs, though far from perfectly consistent.

None of this is a fixed, permanent ranking of which engine is more or less volatile, since it can shift with model updates, infrastructure changes, and product changes on the engine provider's side. It's a reason to measure volatility periodically for your specific prompt set, rather than assuming it based on general reputation.

Building a Simple Volatility Check Yourself

You don't need specialized tooling to get a rough sense of this. Pick five or six prompts relevant to your category, run each one three or four times in a single sitting on a given engine, and note how much the set of mentioned brands, their order, and their framing actually varies across those runs. A prompt set that returns nearly identical answers each time suggests low baseline volatility for that engine. A prompt set that returns meaningfully different brand lists run to run suggests you should treat any single measurement with real caution, and lean more heavily on trends across many measurements rather than any one data point.

What This Means for How You Should Read Your Own Reports

Once you have some sense of an engine's baseline volatility, it changes how you should read a month-over-month visibility report. A modest dip on a known-volatile engine warrants a wait-and-see approach, confirm it persists across the next couple of measurement cycles before treating it as real. The same size dip on a known-stable engine deserves faster attention, since it's less likely to simply be noise resolving itself. Applying one universal reaction threshold across every engine, regardless of how stable that engine actually is, is a subtle but common way that reporting ends up either too jumpy or too complacent.

Why We're Proposing This as an Index, Not Just an Observation

Naming this as an index rather than a one-off observation is a deliberate choice. An index implies something tracked consistently over time, on a defined methodology, comparable across engines and across periods, not a single blog post's worth of anecdotal testing. The GEO field doesn't yet have a widely agreed-upon standard for measuring this kind of baseline instability, which means most conversations about an AI engine changing its answer stay anecdotal rather than becoming something teams can actually plan around.

That's the gap worth closing. A shared, consistent way of expressing how volatile a given engine's answers are for a given category would let brands calibrate their own reaction thresholds sensibly, rather than either overreacting to every fluctuation or under-reacting to genuine drift because it superficially resembles noise. This doesn't need to be a single universal number across every industry, since volatility genuinely differs by category, but a consistent methodology, applied to your own specific prompt set, gets you most of the practical benefit without needing an industry-wide standard to exist first.

What We're Seeing So Far

  • Across repeated same-day runs of identical prompts, we've typically seen a meaningful minority of responses vary in which brands get mentioned at all, not just in phrasing or ordering, which is a bigger effect than most teams assume going in.
  • Volatility appears to vary noticeably by category, broad, competitive categories with many plausible answers tend to show more run-to-run variation than narrow, specific categories with fewer obvious candidates.
  • GEOscanAI's approach to this is to treat every measurement as a sample, not a verdict, weighting trend direction across repeated runs more heavily than any single check, precisely because of this baseline volatility.
  • A brand that only checks its AI visibility once, on one day, is effectively taking a single sample from a genuinely noisy process and treating it as ground truth, which is a real risk for drawing the wrong conclusion.

The Takeaway

Some of what looks like a change in your AI visibility isn't about you at all, it's the model's own baseline instability showing through. A volatility index doesn't replace visibility tracking, it makes visibility tracking more honest, by giving you a sense of how much any single number should actually be trusted before you act on it.

Frequently asked questions

If AI answers vary between identical runs, does that mean visibility tracking is unreliable?

Not unreliable, but it does mean a single check shouldn't be treated as ground truth. Tracking a trend across repeated runs over time, rather than relying on any one measurement, accounts for this baseline variability and produces a far more trustworthy signal.

Is answer volatility the same thing as a real change in my brand's visibility?

No, and conflating the two is a common mistake. Volatility is the model's own baseline instability, present even when nothing about your brand changed. A real visibility shift is a change caused by something you did, or something a competitor did, that persists across repeated measurements rather than bouncing back on its own.

Which AI engines tend to be more volatile?

Based on repeated testing, engines that lean more heavily on live retrieval tend to show more run-to-run variation than engines leaning more on stable trained knowledge, though this can shift with model and infrastructure updates, so it's worth checking periodically rather than assuming a fixed ranking.

How many times should I re-run a prompt before trusting the result?

There's no single fixed number, but running the same prompt a handful of times in one sitting, and comparing that to results from a later week, gives a far more reliable picture than a single run. The wider the spread across those repeated runs, the more caution you should apply to any one measurement.

geomeasurementmodel-updatesai-answersvolatility
W

GEOscanAI monitors how AI search engines recommend brands, providing daily visibility scores across ChatGPT, Claude, Gemini, Perplexity, and Tavily.