Temperature 0 Does Not Make LLM Output Deterministic
Temperature controls token selection, not the full inference path. Learn where LLM output variation enters and how to test the right reliability contract.
temperature=0 removes one source of variation from an LLM request. It does not turn the entire inference system into a deterministic function.
The distinction matters in production. Temperature acts at token selection, after the model has computed its logits. Those logits can still change when the model snapshot, backend, numerical execution path, request batch, tool results, retrieval context, or conversation state changes. If two top tokens are close, a small upstream difference can change the selected token and send the rest of the generation down another path.
I treat this as an architectural contract failure, not a mysterious property of the model. A test harness assumes exact repeatability while the serving contract offers only best-effort consistency.
- Temperature controls token selection. It does not control the numerical computation or serving context that produces the logits.
- At temperature zero, many APIs use greedy selection: choose the highest-logit token. The same logits produce the same choice; slightly different logits may not.
- Floating-point execution, batching, model or backend revisions, retrieval, tools, and state can all change a result before or around token selection.
- Seeds and backend fingerprints improve reproducibility, but providers describe them as best effort rather than a bit-for-bit guarantee.
- Production systems should test behavioral invariants and accepted outcomes, not assume that exact text equality is the contract.
What Temperature Actually Controls
Temperature reshapes the probability distribution used to choose the next token. For a positive temperature \(T\), logits \(z_i\) are scaled before softmax: \(\text{softmax}(z_i / T)\). Lower values concentrate probability on the highest-logit tokens. At zero, APIs commonly switch to greedy selection rather than divide by zero.
Greedy selection is deterministic given an identical logit vector and identical decoding rules. That condition is the part production systems often skip. Temperature does not pin the computation that created the logits, the model version, or any external context supplied to the model.
Temperature zero makes token selection less variable. It does not make the complete request path deterministic.Where Variation Enters
Numerical execution. Transformer inference is dominated by parallel floating-point operations. Floating-point addition is not associative, so changing reduction order can change the last bits of a result. PyTorch's numerical-accuracy documentation states that mathematically identical floating-point computations are not guaranteed to be bitwise identical across releases, platforms, devices, or batched and unbatched execution.
Serving context. High-throughput inference systems combine requests, choose kernels, reuse computed state, and distribute work across hardware. A different batch shape or backend path can produce a slightly different logit vector. NVIDIA's cuBLAS reproducibility guidance, for example, limits its bitwise guarantee to a fixed toolkit and matching GPU conditions, and documents cases where concurrent streams break that guarantee. Recent work on deterministic LLM inference treats batch-dependent kernels and scheduling as system-level problems, not sampling-parameter problems.
Application context. Even a numerically stable model call is not repeatable if retrieved documents, tool outputs, prompt templates, conversation history, safety policies, or model aliases have changed. These are often the larger sources of behavioral drift because they alter the actual input or execution path rather than only its last bits.
None of these mechanisms implies that every repeated request will differ. Usually the highest-logit token wins by enough margin that small numerical changes leave the ordering intact. The important point is narrower: exact equality is not a property that temperature alone establishes.
What Provider Controls Do
| Control | What It Stabilizes | What It Does Not Guarantee |
|---|---|---|
| Temperature | Randomness in token selection | Identical logits or backend execution |
| Seed | Best-effort repeatability of sampling | Identical output across backend changes |
| Model snapshot | Model weights and version | Stable retrieval, tools, prompts, or hardware |
| Backend fingerprint | Detection of some serving changes | Bit-for-bit equality |
OpenAI's reproducible-output guidance recommends matching the seed, request parameters, and system_fingerprint. It still describes the result as “mostly” identical and says determinism is not guaranteed. That is a useful engineering control, not a formal contract.
Choose the Reproducibility Contract
“Reproducible” can mean at least three different things:
- Bitwise reproducibility: every token and byte is identical.
- Semantic reproducibility: wording may vary, but the same claims, classifications, or decisions survive.
- Operational reproducibility: the system stays within an accepted quality and risk envelope even when individual outputs differ.
Most production applications need the second or third contract. Exact text equality is useful for low-level debugging, but it is usually the wrong acceptance criterion for a probabilistic application. A support agent may phrase an answer differently while still citing the correct policy and taking the same permitted action. Conversely, an identical-looking answer can still be wrong after the policy or source data changes.
The contract should follow the consequence. A creative assistant may tolerate broad variation. A classifier that triggers a payment needs a narrow decision boundary, explicit abstention, and deterministic downstream enforcement. A regulated workflow may require storing the model snapshot, complete inputs, tool results, output, and review decision even when replay cannot reproduce every token.
Reproducibility is a system contract. Define what must remain invariant, then control and record the layers that can change it.A Practical Test Strategy
I would start by pinning what the platform lets you pin: model snapshot, prompts, decoding parameters, tool definitions, retrieval corpus, and seed. Record the backend fingerprint when one is exposed. For self-hosted inference, also pin framework, kernel, hardware, and batching configuration when exact replay is genuinely required.
Then test at the level the user experiences:
- Run repeated trials on a versioned evaluation set rather than a single golden completion.
- Assert hard invariants separately: required citations, allowed actions, schema validity, policy constraints, and numerical bounds.
- Measure outcome distributions for qualities that can vary: accuracy, abstention, latency, cost, and escalation rate.
- Store enough request and execution context to explain a regression, not merely the final text.
This is the same architectural shift described in The End of Determinism: reliability comes from controlling consequences and detecting drift, not from pretending variation has disappeared.
Frequently Asked Questions
Why does temperature zero look deterministic most of the time?
Because the leading token is often ahead by enough margin that small upstream differences do not change the winner. Variation becomes visible when contenders are close or when the application context changes materially. Stable observations are evidence of high consistency under those conditions, not proof of an unconditional guarantee.
Does a seed solve the problem?
It helps control sampling and makes evaluations easier to compare. It does not freeze model updates, backend execution, retrieval results, tool calls, or the rest of the application. Treat it as one coordinate in a reproducibility record.
Should tests ever use exact match?
Yes—when exact syntax is the requirement, such as a constrained enum, protocol token, or canonical transformation. Enforce that requirement with schemas and deterministic code where possible. For open-ended model output, exact match usually tests wording more than correctness.
Related: I later ran a cross-model benchmark that surfaced the same evaluation problem in how seven models judged each other: evaluator behavior is part of the system.
Related Reading
- The Validator Asymmetry Principle — why generation and verification have different cost structures
- The Probabilistic State Machine — a stronger runtime model for agentic systems
- Seven Models Judged Each Other. The Mistakes Were the Most Useful Part. — why evaluation cannot be treated as an oracle