Saturday, 10 October 2026
Your paper, your pace.
AI & Robotics

An AI Answer Is a Sample, Not a Stored Page

The same prompt can lead down different token paths, which is useful for invention but never a substitute for checking evidence.

6 min 3 sources Confidence 88/100

In short

What happened. Ask a generative AI system the same question twice and you may receive two answers that differ in wording, emphasis or even conclusion.

What it means. The answer is usually assembled token by token from a distribution of possibilities. It is not simply fetched like a stored page, so polished consistency is not proof of truth and variation is not proof of failure.

Risks and impact. People can mistake one fluent output for a settled fact, or treat several similar outputs as independent confirmation when they come from the same model and evidence gap.

What can be done. Save the prompt, relevant context, model or product version and date; require inspectable sources; and test important tasks against examples with known answers.

What to watch. Changes in tools, retrieved documents, system instructions and model versions can matter as much as visible wording. For consequential decisions, verify the claim outside the model.

Shown as a summary because of your reading settings.

What happened

Modern language models generate text sequentially. The Transformer architecture described by Google researchers in 2017 made it practical to relate each position in a sequence to the surrounding context through attention. During generation, the model assigns probabilities to possible next tokens, selects one according to a decoding rule, adds it to the context and repeats.

Hugging Face’s documentation distinguishes several decoding strategies. Greedy decoding takes the highest-probability next token. Sampling draws from the probability distribution and can yield more varied text. Beam search keeps several candidate sequences before choosing among them. Products may also add hidden instructions, search results, tools, safety rules or conversation memory. Not every interface exposes those choices.

Therefore, “same prompt” does not always mean “same generation conditions.” Even under fixed conditions, a sampling-based setup may take a different path. Under changed conditions, a deterministic decoder can also produce a different result. None of those mechanisms checks whether the resulting sentence is true.

What the evidence supports

The most useful distinction is between repeatability and validity. A greedy decoder in a fixed software pipeline may repeat its output closely because it keeps choosing the locally most probable token. That makes a run easier to reproduce; it does not turn the output into evidence. A sampling decoder may produce several reasonable formulations from the same probability landscape. That can help with brainstorming, but disagreement among samples does not tell us which claim is correct.

NIST’s generative-AI risk profile explains why factual checking remains separate. It describes “confabulation” as confidently presented but false or erroneous content and notes that large language models predict the next token from statistical patterns. A statistically plausible continuation can therefore be factually wrong. NIST also warns that generated citations can themselves be fabricated or misleading.

The practical consequence is slightly uncomfortable: asking the model again is not the same as seeking an independent source. Ten runs may share the same training blind spot, retrieval failure or unsupported premise. Conversely, two differently phrased answers may express the same well-supported fact. The evidence behind the claim—not the vote count among outputs—has to decide.

How the story is being framed

Variation is not automatically a defect. A writer may want alternative openings; a programmer may want several implementations; a teacher may want explanations at different levels. Sampling can widen the option set. The user can then compare those options against a clear brief.

Operational work has a different need. If an organisation uses generated text to classify cases, draft regulated documents or support a recurring workflow, unexplained variation can make review and auditing harder. Teams may prefer constrained formats, pinned models, lower-variation settings and test suites. Yet a perfectly stable mistake is still a mistake.

The tension is not “creative versus reliable” in the abstract. It is whether the generation strategy fits the job and whether a separate evaluation checks the property that matters: factual accuracy, completeness, tone, safety or reproducibility.

A provider may frame variability as creative range; a compliance team may frame it as loss of control; a researcher may see a measurement problem. Each view catches something real, but each can hide a different weakness. Range is useful only when someone can judge the options. Control can freeze an error. Measurement can become detached from the reader’s actual task. The shared requirement is an evaluation made for the use case, not a single score that declares the system generally “good.”

The background

A conventional web page has an address and stored content, though the publisher can update it. A search engine mainly retrieves and ranks documents. A generative assistant composes an answer, and many current products combine composition with retrieval, calculators, code tools or personal context. The visible paragraph can therefore be the last step in a longer, partly hidden pipeline.

That hybrid design is useful, but it complicates comparisons. One run may retrieve a fresh document while another does not. A provider may update the model or its instructions. A conversation may contain different earlier messages. Even the calendar date can alter an answer to “current” questions.

This is why screenshots of isolated outputs are weak records. They often omit the prompt, sources, model, date and preceding context needed to understand how the answer arose.

There is another source of confusion: the word temperature. In many systems it changes how sharply or broadly probabilities influence sampling. It is often described as a creativity dial, but that is only shorthand. A lower setting does not verify a claim, and a higher one does not guarantee useful originality. Some products do not expose the setting at all. Others may transform the user’s request before it reaches the model.

We also rarely know the full probability distribution behind a hosted answer. That limits what an outside observer can infer from two outputs. The honest conclusion is usually modest: the pipeline produced these texts under the conditions visible to us. Why one token path won may remain partly hidden.

The deeper story

The better question is not “Can I make the model say the same thing every time?” It is “What failure would matter in this task, and how will I detect it?” For ideation, diversity may be the goal. For a factual brief, traceable sources and accurate quotations matter more. For a structured workflow, format compliance and repeatability may be essential. For a high-stakes judgment, the model should not be the final authority.

That shift changes prompting too. Instead of asking only for confidence, specify the evidence standard: distinguish known facts from inference, link the original record, state uncertainty and decline when the source is missing. A good prompt cannot guarantee truth, but it can make unsupported certainty easier to notice.

This is a broader lesson about tools that speak in complete sentences. Fluency collapses the distance between calculation and conversation. We hear a voice where the system is producing a sequence, and we instinctively look for settled intention behind it. The metaphor of a “sample” restores a little useful friction. It reminds us that another path was possible, that likelihood is not knowledge and that the reader still has work to do.

That work is not suspicion for its own sake. It is choosing an appropriate burden of proof. A playful slogan can tolerate variation. A historical date needs a source. A medical, legal or financial decision needs qualified human judgment and current authoritative guidance. The more the answer can change a life, the less its surface confidence should decide.

Something to sit with

When an answer changes, which part matters: the wording, the recommendation or the underlying fact?

What independent record would let you decide between two plausible outputs?

If the answer stayed identical ten times, what would you still need to verify?

Sources

We report facts from the sources above in our own words and link to the originals. Interpretation is ours, not theirs.

QUICK UNDERSTANDING CHECK

What does a different answer to the same prompt prove by itself?

♻︎ Free to republish

Copy this HTML into your CMS. Credit line and licence are included. Republish our work — free

Every headline has a deeper story. This is ours.

What we are doing here

My Paper in your inbox

Every morning, or once on Saturday — you choose. The story worth your time, and the question underneath it. One click to leave.

One click to leave. We never sell or share your address.