How a language model turns one prompt into two different answers

Published

An engineer sends a support bot the same ticket twice, back to back, with the same settings. The ticket reads: “Deploy to staging failed. Log ends with Killed after npm run build.” The first reply says the operating system probably killed the build for running out of memory and suggests raising the container’s memory limit. The second says the build likely exceeded a resource limit and suggests checking for a timeout or a memory cap. Neither reply is wrong, and neither copies the other. If you come from backend or QA work, you expect a function given the same input to return the same output. So you start looking for a broken cache, a load balancer splitting traffic across two model versions, or a race condition.

This article argues that the model computed the same kind of thing both times: a list of probabilities for every possible next piece of text. The difference between the replies was decided afterwards, by the sampling rule that picks one piece from that list, and by the settings that shape the rule: temperature, top-k and top-p. Once you see the probabilities and the draw as two separate steps, “the model is flaky” turns into a question you can check.

The model reads tokens, not words

A language model never sees your ticket as words. A tokenizer first converts the text into tokens, the units the model reads and writes. A token can be a whole word, part of a word, a word with its leading space, or a punctuation mark. A rare word is split into several tokens because the tokenizer’s fixed vocabulary has no single entry for it.

Think of it as bytes in a network buffer. You do not size a payload in words, because the limits and costs are counted in bytes. A model’s limits are counted in tokens in the same way.

The most important of those limits is the context window. It holds everything the model can use for one request: your instructions, earlier turns, pasted logs, and the reply it is writing. Input and output share that one budget. The ticket and the reply both had to fit on this board before anything was computed.

Each step computes a distribution, not an answer

Suppose the reply so far reads The deploy failed because the, and the model is working out the next token. It does not produce one word. It produces a probability for every token in its vocabulary, and those probabilities add up to 1. This is the next-token distribution. Most entries are close to zero.

Here are five candidates from a small teaching example. Their probabilities have been rescaled to add up to 1 on their own, as if these were the only tokens. A real vocabulary has many thousands of entries.

Five candidate tokens with probabilities from 0.445 for server down to 0.022 for cat

The model’s output for one position: a probability for each candidate next token

That table is the model’s entire output for this position. The model does not choose which token appears in the reply. A separate step does that.

The chosen token is then added to the end of the sequence. The model computes a fresh distribution for the next position, and the loop repeats until something stops it. This is called autoregressive generation: each new token is chosen using all the tokens before it, and then becomes part of the input for the next one.

A loop from token sequence to distribution to sampling to appending, until a stop condition ends it

Each chosen token is appended, and the model computes a fresh distribution from the longer sequence

Think of it as phone autocomplete run in a loop, with a far better sense of what comes next. Your keyboard offers three suggestions. The model scores every possible token. Your thumb picks one suggestion; a sampling rule picks one token. Nothing in the loop plans the end of the sentence before it gets there.

One consequence matters for our two replies. Each token is chosen with only the earlier tokens in view, and the model never goes back. If cat were picked above, every later token would be computed to follow The deploy failed because the cat.

The draw, not the model, decides which reply you get

There are two basic ways to pick a token from the distribution.

Greedy decoding always takes the single most probable token. Here that is server, every time. Given the same distribution, greedy decoding always makes the same choice. Its known weakness is that it tends to produce repetitive text.

Sampling picks at random, in proportion to each token’s probability. With the table above, server comes up about 44 or 45 times in a hundred, build about 27 times, and cat about twice.

This is the full explanation of the support bot. With sampling, two runs can take different tokens at the first position where the model is uncertain. After that, each run computes its next distribution from a different sequence. The two replies then drift further apart, and both can stay reasonable.

Both runs receive the same distribution, draw different tokens, then get different follow-up distributions

Two runs share one distribution; one different draw sends each reply down its own path

So the setting that decides whether the replies can differ is the decoding rule: greedy or sampling. When sampling is on, three further settings decide how likely a difference is.

Temperature reshapes the odds before the draw

Temperature sharpens or flattens the distribution before the draw. Below 1, it makes the most likely tokens even more likely. Above 1, it spreads probability toward the long shots.

On the worked example, the arithmetic looks like this:

server rises to 0.643 at temperature 0.5 and falls to 0.325 at 2.0; cat moves from 0.002 to 0.072

Lower temperature concentrates probability on the top token; higher temperature spreads it out

Think of it as a dimmer switch for surprise. Turn it down, and the draw almost always takes the obvious token. Turn it up, and cat goes from roughly one draw in five hundred to roughly seven in a hundred.

The effect on your two replies is direct. A low temperature makes the first uncertain position more likely to resolve the same way on both runs, so the replies diverge less often. Notice what temperature does not do: it never removes a token from the draw. It only changes the odds.

Top-k and top-p decide which tokens may be drawn at all

The other two settings are filters. They decide which tokens are eligible for the draw. After filtering, the remaining probability is shared out again among the survivors.

Top-k keeps the k most probable tokens. With k = 2, only server and build remain, and their shares become about 0.62 and 0.38. The pool size stays fixed, however confident or unsure the model is.

Top-p, also called nucleus sampling, keeps the smallest set of most-probable tokens whose probabilities add up to at least p. Walk down the table and add as you go: 0.445, then 0.715, then 0.879. With p = 0.8, the running total first passes 0.8 at network, so the pool is server, build and network. The pool size adapts. When the model is confident, a few tokens reach p quickly; when it is unsure, many are needed.

These settings interact, and the order matters. Apply temperature 0.5 first, and the table becomes 0.643, 0.237, 0.087, 0.032, 0.002. Now the running total passes 0.8 at build (0.643 + 0.237 = 0.880). The same top-p of 0.8 now keeps two tokens instead of three.

Distribution passes through temperature, then a top-k or top-p filter, then redistribution, then a random draw

Temperature reshapes the odds, the filters shrink the pool, and only then is a token drawn

Removing randomness narrows the gap but does not close it

The obvious fix is to make generation deterministic. That helps, but providers’ own documentation says it is not a guarantee.

  • One provider offers a seed parameter that makes “a best effort to sample deterministically” and states plainly: “Determinism is not guaranteed.” It notes a small chance that responses differ even when the request settings and a backend fingerprint field match. It asks you to watch that field for backend changes.
  • Another provider’s API reference accepts only a temperature of 1.0 on its newer models. On those models, you cannot turn the dimmer down at all.

The settings you can control therefore differ by provider, and even by model version. The structure does not differ: a distribution is computed, and something picks from it.

In our view, the dependable way to get repeatable behaviour is to stop relying on the draw. Test the property you need, such as “the reply names a memory limit”, rather than an exact string. Record the settings and the backend identifier with every response.

Evidence, not guesswork, tells you which cause you have

Different replies to the same ticket can have more than one cause. A sampling draw is the designed-in one. A backend change is another. A reply that was cut off is a third. Each response carries the fields you need to tell these apart.

Read the request settings, the stop reason and the backend identifier, then compare both responses

Three things to read before blaming the model for two different replies

Read the request settings. If sampling was on, two different draws are the expected outcome, not a fault. Note the temperature, top-k and top-p values.

Read the stop reason. This is a field on the response that says why generation ended. The model may have finished naturally. It may have hit your length ceiling, produced one of your stop sequences, called a tool, or filled the context window. Suppose one reply is cut off by a length limit and the other is not. Then the difference is truncation, not the draw.

Read the backend identifier. If it changed between the two requests, the distribution itself may have changed.

If the settings used sampling, both replies finished naturally, and the backend did not change, then the draw is the most likely explanation. Even then, the provider’s own caveat holds: outputs can differ when every one of those fields matches.

Here is how the whole answer sounds when someone asks you after an incident. The tokenizer turned the ticket into tokens. Those tokens, plus the reply, had to fit in the context window. At each position the model computed a probability for every possible next token. A sampling rule, shaped by temperature, top-k or top-p, picked one, and the loop continued from it. The logged settings, stop reason and backend identifier confirm which of these made the difference.

Key takeaways

  • At every step, a language model computes a probability for every possible next token. It does not pick the word itself.
  • A decoding rule picks the token. Greedy decoding always takes the top token; sampling draws in proportion to the probabilities, so two runs can diverge at the first uncertain position and then drift apart.
  • Temperature reshapes the odds before the draw. Top-k and top-p limit which tokens may be drawn. Their combination decides how often replies differ, and they interact: lower temperature shrinks a top-p pool.
  • A fixed seed or low temperature narrows variation but is documented as best effort. Choose property-based tests over exact-string tests whenever a model’s output is involved.
  • Before calling a model flaky, read the settings, the stop reason and the backend identifier on both responses.

This article settles where variation between identical requests comes from and which settings govern it. It does not settle how much context is too much, since accuracy is documented to degrade as the window fills, or how to constrain a reply’s format so it can be checked automatically. Those questions need their own treatment.