• Home
  • Blog Post
  • Why AI Search Results Keep Changing: Understanding GEO Volatility and Run-to-Run Variation
Homepage brand Logo image for NeuralAdX Ltd showing an AI brain and digital circuitry, representing Generative Engine Optimisation specialists focused on improving visibility and citations in AI search engines

AI search results keep changing because generative search is a probabilistic retrieval-and-generation process: the same prompt can trigger different searches, retrieve and rerank different sources, allocate different evidence to the model, and generate different wording or citations from one run to the next.

AI Search Results Change Because Generative Engines Are Probabilistic Retrieval Systems, Not Fixed Ranking Lists

That distinction is now measurable. A 2026 University of St. Gallen study found that cited-source sets overlapped by only 34% to 42% between consecutive days across 4,044 comparison pairs, while repeated runs within 24 hours averaged only 32% to 43% source overlap. A peer-reviewed ACL 2026 study separately found that even with model temperature set to zero, 9% to 27% of tested queries changed answer polarity within five minutes, depending on the generative search system. Schulte et al., 2026 ACL Findings, 2026

TL;DR
One AI result is an observation, not a stable ranking.

1) Run-to-run variation can happen within minutes because retrieval and generation are non-deterministic. 2) The web, indexes, models and retrieval systems also drift over days and weeks. 3) Small prompt changes can alter what subqueries are generated and which evidence is selected. 4) Therefore, Generative Engine Optimisation should measure visibility as a distribution across repeated runs, prompt families, engines and time, rather than treating a single screenshot as a ranking result.

The strongest 2026 evidence does not imply that AI answers are pure chaos. Concepts can remain relatively stable while the specific URLs, citations, brand mentions or answer polarity change. That is why reliable GEO measurement separates source selection, brand inclusion, citation, answer absorption and measurement uncertainty.

What Is GEO Volatility and Run-to-Run Variation?

GEO volatility is the observed change in brand mentions, citations, cited domains, source order, answer wording or recommendation outcomes when a generative engine is tested repeatedly across runs, prompts or time. Run-to-run variation is the narrowest form: the same or near-identical prompt produces a materially different result on another execution, even when little external time has passed.

This differs from ordinary search-rank movement. A conventional search result is usually observed as a ranked set of links. A generative engine first has to decide whether to search, what to search for, which sources to retrieve, which passages to retain, how to combine them and which claims or links to expose in the final answer. The 2026 critical GEO survey describes the problem as a stochastic, partially observable pipeline spanning retrieval, reranking, context allocation, citation and factual absorption. Martinez, 2026 survey

The practical definition: a brand does not have one immutable “AI rank.” It has an observed probability of being retrieved, mentioned, cited, positioned and absorbed into answers under a defined prompt set and measurement protocol.

Why the Same AI Search Prompt Can Produce a Different Answer

Current research points to several interacting causes rather than one single volatility mechanism. The most useful model is to treat every AI answer as the output of a chain in which variation can enter at multiple stages.

Diagram 1: Where volatility enters the AI answer pipeline

■ 1. Prompt interpretationAmbiguity, wording, locale and context influence the information need.
■ 2. Query fan-outOne prompt may generate multiple related searches or subquestions.
■ 3. RetrievalCandidate pages can differ between calls, indexes, engines and freshness states.
■ 4. Reranking + contextDifferent evidence can be promoted, compressed or omitted before generation.
■ 5. Generation + citationThe model can synthesize, phrase, recommend and expose citations differently.

Key: ■ Prompt   ■ Fan-out   ■ Retrieval   ■ Reranking/context   ■ Generation/citation

Mobile users: scroll horizontally where needed to view the full diagram.

1. Stochastic generation can change the final answer

Language models generate tokens probabilistically. Importantly, the problem is not solved simply by setting temperature to zero. The ACL 2026 study found answer-decision flips even at zero temperature, which indicates that the wider system, including retrieval and execution behaviour, can remain variable. Kirsten et al., ACL 2026

2. Retrieval can select a different evidence set

Generative search does not simply read a fixed top ten and rewrite it. Research comparing generative and organic search finds different retrieval footprints, while repeated-run studies show that the cited domains themselves change substantially between executions. Sielinski’s 2026 study found median domain-level Jaccard overlap of about 0.29 to 0.31 for Gemini, 0.33 to 0.40 for SearchGPT and 0.50 for Perplexity across its three tested consumer topics. Exact identical citation sets were rare. Sielinski, 2026

3. Query fan-out multiplies the opportunities for variation

Google officially documents that AI Overviews and AI Mode may use query fan-out, issuing multiple related searches across subtopics and data sources. If the model decomposes the original request differently, or a subquery returns a different source pool, downstream citations can change even when the visible user prompt has not. Google also states that AI Mode and AI Overviews can use different models and techniques, so their responses and links can vary. Google Search Central

4. The web, indexes and model systems drift over time

Run-to-run stochasticity is only one layer. Over longer periods, source pages change, new pages appear, indexes refresh, crawling eligibility changes, model versions update and retrieval systems are revised. The ACL study explicitly separates temporal drift from stochasticity. Its two-month comparison found only 18% web-page overlap for Google AI Overviews, compared with 45% for organic search, even though broad concept coverage remained comparatively stable. ACL Findings, 2026

5. Prompt wording and context can move the retrieval path

Prompt sensitivity is not a minor edge case. Grossman and colleagues’ 2026 benchmark of 11,500 real-user queries found that sources retrieved by Google organic search, AI Overviews and Gemini differed substantially, with average cross-system Jaccard similarity below 0.2. The study also found AI Overviews less consistent across two runs of the same query and less robust to small query edits. This is why professional GEO testing needs prompt families and paraphrases, not one “perfect prompt.” Grossman et al., 2026

What the Latest Research Shows About AI Search Volatility

The strongest conclusion is now consistent across independent studies: a single generative-search observation is too noisy to represent stable visibility. The exact level of volatility varies by engine, query type, topic, metric and time horizon.

Research evidence on run-to-run and temporal variation
StudyScopeVolatility findingGEO implication
Schulte, Bleeker & Kaufmann, 2026
University of St. Gallen / arXiv
4 engines, 4 verticals, 8 prompts per campaign, 45 to 46 days; 4,044 consecutive-day pairs plus repeated runs.Day-to-day source Jaccard 0.34 to 0.42; within-24-hour source overlap 0.32 to 0.43; brand overlap also variable.Repeated runs and rolling windows are necessary; one run can materially overstate or understate visibility.
Kirsten et al., 2026
Findings of ACL 2026
Google organic plus five generative search systems across diverse query datasets.At temperature 0, 9% to 27% of tested ternary-answer queries changed decision within five minutes. AIO page overlap across two time-separated runs was 18%, versus 45% organic.Neither low temperature nor short time intervals guarantee reproducibility.
Sielinski, 2026
arXiv v2, June 2026
3 platforms, 3 topics, 200 generated queries per topic, repeated daily samples; 374,052 extracted daily citations.Median repeated-run domain Jaccard roughly 0.29 to 0.50 by platform; 95% confidence intervals often make small apparent citation-share differences indistinguishable from noise.Report uncertainty, not just point estimates; platform-specific sample requirements matter.
Grossman et al., 2026
Generative search benchmark
11,500 real-user queries comparing Google Search, AI Overviews and Gemini.Cross-system retrieved-source overlap averaged below 0.2; AIOs were less consistent across repeated runs and more sensitive to small query edits.Prompt robustness and engine-specific retrieval behaviour need to be tested directly.

Mobile users: scroll horizontally to view the full research table.

Diagram 2: Answer decisions can flip within five minutes

Share of tested ternary-answer queries whose overall decision changed from t0 to t+5 minutes at temperature 0.

GPT-Tool

9%

Gemini

15%

GPT-Search

16%

Google AIO

17%

Sonar

27%

Key: ■ GPT-Tool 9%   ■ Gemini 15%   ■ GPT-Search 16%   ■ Google AIO 17%   ■ Sonar 27%

Source: Kirsten et al., Findings of ACL 2026, Table 4. Mobile users: scroll horizontally to view the full chart.

Diagram 3: Most of the cited-source set can change from one day to the next

100% stacked representation of mean consecutive-day Jaccard overlap and its complement for each campaign vertical.

Consumer electronics

33.6% overlap
66.4% changed

Real estate

37.8% overlap
62.2% changed

Sporting goods

35.5% overlap
64.5% changed

Telecommunications

42.3% overlap
57.7% changed

Key: ■ Overlapping cited-source set   ■ Non-overlapping / changed portion

Source: Schulte, Bleeker & Kaufmann, 2026, 4,044 consecutive-day pairs. “Changed” is shown as 100% minus the mean Jaccard overlap. Mobile users: scroll horizontally to view the full stacked chart.

The Five Types of AI Search Volatility GEO Measurement Should Separate

Calling every change “AI randomness” is too crude. A stronger GEO measurement model identifies what changed and at which level.

Volatility taxonomy for generative search measurement
TypeWhat changesBest controlWhat to report
Run-to-run stochasticitySame prompt, same short time window, different sources or answer.Repeat executions under controlled conditions.Mention probability, citation prevalence, overlap, variance.
Temporal driftResults change across days or weeks.Fixed prompt set, fixed locale, rolling windows.Trend with confidence range, not one date.
Prompt sensitivityParaphrases or small wording changes alter retrieval and answer.Use prompt families representing the same intent.Coverage across intent variants, not one query phrase.
Engine varianceDifferent AI platforms choose different source pools and synthesis strategies.Measure engines separately before aggregating.Platform-specific citation and brand metrics.
Measurement uncertaintyFinite samples create noisy estimates even if the underlying system distribution is stable.Larger samples, repeated runs and uncertainty estimation.Confidence intervals or clearly stated sampling error.

Mobile users: scroll horizontally to view the full volatility table.

Critical distinction: system-level stochasticity and measurement uncertainty are not the same thing. Sielinski explicitly separates them. The first is variation in the answer engine itself; the second is uncertainty introduced because we observe only a finite sample of that system. A rigorous GEO report should make both visible. Statistical framework, 2026

Volatility Does Not Mean AI Search Is Random Chaos

This is an important correction. The evidence shows instability, but it also shows structure. The ACL study found that broad conceptual coverage could remain similar even while the specific retrieved pages changed. Sielinski found platform-specific citation consistency patterns, and Schulte found a highly concentrated citation landscape in which a relatively small group of domains captured a large share of citations.

In Schulte’s dataset, the mean Gini coefficient for source citation concentration was 0.715 across campaigns and engines, with Google AI Mode at 0.782 and Perplexity at 0.671. That indicates concentration and volatility can coexist: an engine may repeatedly draw from a relatively privileged source pool while still changing which specific sources appear in any individual response. Schulte et al., 2026

For GEO, the target is therefore not “make every run identical.” The realistic target is to increase the probability that a brand or source survives repeated retrieval, selection and synthesis across commercially relevant prompts.

How AI Visibility Should Be Measured When Results Keep Changing

The central measurement shift is from snapshot ranking to probabilistic visibility. A credible GEO system should ask: how often is the brand included, how often is the domain cited, in what position, for which intents, on which engines, with what uncertainty, and over what period?

Repeat the same promptSeparates one-off outcomes from repeatable inclusion.
Use prompt familiesTests the intent, not a single wording accident.
Measure over timeDistinguishes sustained visibility from short-lived noise.
Keep engines separateDifferent systems have different retrieval and citation regimes.

How many runs are enough?

There is no universal number that is proven for every engine, market and prompt set. However, Schulte’s bootstrap convergence analysis provides one concrete empirical benchmark: in that dataset, the standard error of per-brand detection fell below 0.10 at seven runs and below 0.08 at eight runs; source coverage required eight runs to get below a 0.10 standard error threshold. The authors recommend at least seven runs per prompt per day for brand visibility and eight when source coverage matters. Convergence analysis, 2026

That should be treated as study-specific evidence, not a universal law. Sielinski’s separate statistical framework finds sample needs can differ substantially by platform and metric. For citation-share confidence intervals spanning five percentage points, its tested conditions required roughly 40 to 50 queries for Gemini, about 100 for Perplexity and at least 150 for SearchGPT. The paper also explicitly leaves principled universal minimum-sample guidance as an open research question. Sielinski, 2026

How long should the observation window be?

Schulte’s rolling-window analysis found per-brand standard error below 0.10 at about 10 days and below 0.05 at 24 days. The reported 95% confidence interval half-width was approximately ±0.105 at 21 days and ±0.065 at 28 days. The paper therefore recommends rolling aggregation over roughly two to four weeks for sustained brand-level monitoring in its setting. Again, that is evidence from one empirical design, not a rule that erases the need for market-specific validation.

What a volatility-aware GEO benchmark should record
MetricWhat it answersWhy one run failsBetter reporting unit
Brand mention probabilityHow often is the brand included?A single inclusion or omission can be stochastic.Mentions / valid repeated responses.
Citation prevalenceHow often is the domain cited at least once?Citation selection varies between runs.Cited responses / valid responses.
Citation shareWhat share of observed citations belongs to the domain?Small differences may sit inside sampling noise.Share plus confidence interval.
Average brand positionWhere does the brand appear when present?Order and inclusion can both move.Mean/median over included runs plus coverage.
Prompt coverageAcross how many commercially relevant intents is the brand visible?One prompt can be unusually easy or hard.Coverage across a fixed prompt portfolio.
Answer absorptionDid the source merely receive a citation, or did its information shape the answer?Citation selection and factual use can diverge.Repeated claim-level attribution / absorption checks.

Mobile users: scroll horizontally to view the full measurement table.

Why One-Off AI Visibility Tests Can Produce False Confidence

A single test can create two opposite errors. It can produce a false positive, where a brand appears strongly once and is assumed to have durable visibility. Or it can produce a false negative, where a relevant brand is absent once and is assumed to be invisible.

Sielinski’s 2026 paper shows why confidence intervals matter. For frequently cited SearchGPT domains in its sample, 95% confidence-interval spans were often 3 to 6 percentage points. Across the studied platforms and topics, apparent citation-share differences below roughly 5 to 7 percentage points commonly had overlapping confidence intervals. In one example, a single sample would have incorrectly made nationalgeographic.com appear like a top-cited Gemini domain, while later samples contradicted that interpretation. Sielinski, 2026

“Single-run visibility metrics provide a misleadingly precise picture” of generative-search performance.

Paraphrased context and short quotation from Sielinski, 2026. The statistical point is that visibility estimates require uncertainty, not false precision.

How Generative Engine Optimisation Reduces the Business Risk of Volatile AI Results

Generative Engine Optimisation cannot force a probabilistic system to return an identical answer every time. Its practical role is to improve the conditions that make a business more consistently retrievable, interpretable, credible and citation-ready across repeated opportunities.

Entity clarityMake the organisation, people, services, locations, evidence and relationships unambiguous so the system has fewer reasons to misclassify the entity.
Citation readinessUse precise claims, attributable statistics, named evidence and answerable passages that can survive retrieval and citation selection.
Prompt coverageMap the real commercial questions customers ask and create evidence that answers the complete intent, not just one keyword variation.
Trust and source diversitySupport important statements with credible primary and independent sources, and reinforce the entity beyond its own website where appropriate.
Technical crawlabilityA page cannot be reliably retrieved if access is blocked or the information is hard to parse. Google requires normal Search eligibility for AI features, and OpenAI advises publishers not to block OAI-SearchBot if they want content discoverable in ChatGPT search.
Repeated retrieval testingMeasure citation, brand mention, position, share of voice, sentiment and answer absorption repeatedly so optimisation decisions are based on persistent patterns.

Technical sources: Google Search Central AI features guidance and OpenAI publisher guidance.

Why Longitudinal GEO Benchmarks Are More Informative Than Screenshots

A screenshot can prove that an answer occurred. It cannot, by itself, establish the probability that the same brand or source will appear again. A benchmark becomes more informative when it fixes the comparison set and prompt methodology, preserves dated reporting periods, repeats observations and publishes limitations.

That is why NeuralAdX Ltd, a specialist Generative Engine Optimisation company, publishes separate ongoing AI Citation Benchmark and AI Answer Visibility and Share of Voice Benchmark pages rather than presenting one retrieval result as permanent proof.

For example, the AI Citation Benchmark currently publishes eight monthly periods. Its latest published period, 24 June to 23 July 2026, records 1,333 citations and 13% citation share for NeuralAdX Ltd in that fixed comparison set; earlier monthly NeuralAdX Ltd citation counts on the same page range from 440 to 1,539. The responsible interpretation is not that every monthly move has a known cause. It is that longitudinal reporting exposes movement that a one-off screenshot would hide. NeuralAdX Ltd benchmark

Industry Expert Quotes

“A single AI-search result is not a ranking; it is one draw from a changing response distribution. When research finds only 32% to 43% source overlap across repeated runs within 24 hours and 34% to 42% across consecutive days, NeuralAdX Ltd treats durable GEO visibility as a probability to be measured across repeated prompts, engines and time, not as a screenshot.”

Quoted expert: Paul Rowe, Founder, Chief Generative Engine Optimisation Officer & CEO of NeuralAdX Ltd

Evidence supporting the quoted figures: Schulte et al., 2026 NeuralAdX Ltd 11-Factor GEO Methodology

The Correct Way to Interpret a Free AI Visibility Assessment

An initial assessment is useful for identifying retrieval patterns, competitor presence, citation opportunities, entity problems and prompt-level weaknesses. It should not be mistaken for a statistically complete longitudinal benchmark. If the first diagnostic reveals commercially important gaps, the next step is repeated monitoring across a defined prompt portfolio and reporting window.

FREE
AI Visibility Assessment

NeuralAdX Ltd

Request Your Free AI Visibility Assessment

Initial website check against our 11-Factor GEO Framework plus 5 Live AI Retrieval Tests.

Find out whether AI recommends your business, cites your website, prefers competitors, or leaves your business invisible in AI answers.

1111-Factor GEO Framework
Checked
5Commercial AI Prompts Tested

Start With A Free Assessment

Call NeuralAdX Ltd or send your assessment request by email.

Emailing Your Request?

For your convenience, your email is already prepared with simple placeholders. Just add your website URL, best contact number, 5 priority AI prompts and any useful information.

Initial assessment only · No obligation · Serious business enquiries answered within one UK business day · View live AI retrieval proof

What Current Platform Guidance Adds to the Volatility Picture

Platform guidance does not expose the complete ranking logic, but it confirms several moving parts. Google says AI features may use query fan-out and that AI Mode and AI Overviews may use different models and techniques. Google also states that meeting technical eligibility does not guarantee crawling, indexing or serving. OpenAI separately states that publishers who want their content discoverable and included in ChatGPT search summaries should not block OAI-SearchBot.

In June 2026, Google also began rolling out dedicated Generative AI performance reports in Search Console to a subset of websites, with visibility broken down by impressions, pages, countries, devices and dates. That improves first-party observability for Google’s own AI surfaces, but it does not remove cross-engine run-to-run variation or provide a universal measurement layer for every generative platform. Google Search Console, June 2026

Schulte and colleagues summarize the practical reality succinctly: “What is cited in one run today is not necessarily cited in the same run tomorrow.”

Frequently Asked Questions About GEO Volatility

Why does ChatGPT or another AI answer differently when I ask the same question twice?

Because the complete answer pipeline is not fixed. Retrieval can return different candidates, reranking can change which evidence reaches the model, and generation itself can vary. Repeated-run studies now show substantial source-set and answer variation even over short intervals.

What is GEO volatility?

GEO volatility is the change in generative-search visibility across repeated executions, prompts, engines or time. It can affect brand mentions, citation inclusion, source order, recommendation position, answer wording and whether a source’s facts are absorbed into the final answer.

Are AI search results deterministic at temperature zero?

No. Temperature zero can reduce one source of generation randomness, but it does not make the whole search system deterministic. ACL 2026 found 9% to 27% decision flips within five minutes across tested generative systems even at temperature zero.

How many times should AI visibility be measured?

There is no universal run count proven for every platform and use case. Schulte et al. recommend at least seven repeated runs per prompt for brand visibility and eight when source-level coverage matters in their dataset. A separate statistical study shows required sample sizes can be larger when the target is a narrow confidence interval for citation share. The correct run count depends on the metric, platform, volatility and precision required.

How often should GEO visibility be monitored?

For ongoing measurement, use repeated observations over a rolling time window rather than isolated spot checks. Schulte et al.’s 2026 dataset supports rolling aggregation over roughly two to four weeks for more stable per-brand estimates in its setting, with short windows producing materially more uncertainty.

Does a disappearing citation mean GEO has failed?

Not necessarily. A single missing citation can be normal run-to-run variation. The stronger question is whether citation prevalence or brand-mention probability is declining across a repeated, controlled sample and whether that change is larger than the expected measurement noise.

Can AI visibility be measured like a Google ranking position?

Only partially. Position can still be useful when a brand appears, but generative search adds an inclusion-exclusion problem: the brand or source may be absent entirely on another run. GEO therefore needs repeated probability, coverage, citation and answer-level metrics alongside position.

What is the best way to reduce exposure to AI search volatility?

Build more consistent retrieval conditions: clear entities, technically accessible pages, authoritative evidence, citation-ready claims, broad commercial prompt coverage and independent trust signals, then test them repeatedly. The goal is not to eliminate stochasticity; it is to increase the probability of being retrieved, cited, mentioned and accurately represented across changing runs.

Related NeuralAdX Ltd GEO Evidence and Methodology

For readers who want to inspect the measurement framework, longitudinal evidence and live retrieval examples behind NeuralAdX Ltd’s Generative Engine Optimisation work, the following resources are directly relevant to volatility-aware GEO measurement.

Primary Research and Platform Sources

This article prioritises 2026 primary research, peer-reviewed conference findings and first-party platform documentation. Preprints are identified as such and should be interpreted with their stated limitations.

Schulte, J., Bleeker, M. & Kaufmann, P. (2026). Don’t Measure Once: Measuring Visibility in AI Search (GEO). University of St. Gallen / arXiv preprint.

Kirsten, E. et al. (2026). Characterizing Web Search in the Age of Generative AI. Findings of the Association for Computational Linguistics: ACL 2026.

Sielinski, R. (2026). Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement. arXiv v2, revised June 2026.

Grossman, R. et al. (2026). How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews. Public benchmark of 11,500 real-user queries.

Martinez, O. (2026). Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026). Critical survey of 45 studies through July 2026.

Google Search Central. AI features and your website and Guide to optimizing for generative AI features.

Google Search Central (3 June 2026). Introducing Search Generative AI performance reports in Search Console.

OpenAI. Publishers and Developers FAQ, including OAI-SearchBot discoverability guidance.

Editorial status: reviewed against sources available on 15 August 2026. Research on generative-search stability is developing rapidly, so numerical thresholds should be treated as study-specific unless independently replicated across the relevant engine, language, location and prompt population.

Author and GEO methodology context

Paul Rowe

Paul Rowe, Founder, Chief Generative Engine Optimisation Officer and CEO of NeuralAdX Ltd

Paul Rowe
Founder, Chief Generative Engine Optimisation Officer and CEO.

Paul Rowe is the Founder, Chief Generative Engine Optimisation Officer and CEO of NeuralAdX Ltd, a UK-based Generative Engine Optimisation agency focused on helping brands become visible, retrievable, cited, mentioned and trusted inside AI-generated answers.

His work focuses on AI citation visibility, answer-engine retrieval, entity clarity, structured content, source trust, prompt coverage and measurable AI answer visibility across ChatGPT, Google AI Mode, Google Gemini, Microsoft Copilot, Perplexity, Grok, Claude and other major AI search and answer platforms.

Paul’s optimisation process is built around the 11-factor GEO methodology, combining citation addition, statistics, quotations, fluency, easy-to-understand content, authority signals, schema markup, recency, author bios, source diversity and technical-term clarity.

NeuralAdX Ltd publishes proof-led GEO work through live AI retrieval testing, the Proof That Generative Engine Optimisation Works evidence hub, the AI Citation Benchmark and the AI Answer Visibility and Share of Voice Benchmark. This author bio is used to connect each article with clear expertise, transparent methodology and verifiable AI visibility evidence.

Founder
CEO
11-factor GEO
AI citation visibility
Answer-engine retrieval
Entity clarity
Evidence-led GEO
Live AI retrieval
Share this post

Subscribe to our newsletter

Keep up with the latest blog posts by staying updated. No spamming: we promise.

By clicking Sign Up you’re confirming that you agree with our Terms and Conditions.

Related posts