The Supervisor Pattern: How to Use Cheap Models Without Losing Control
Cheap models can expand execution, but the architecture must preserve scarce judgment. The Supervisor Pattern separates bounded work from consequential review.
Lower-cost models do not create reliability problems simply because they cost less. The problem begins when an output crosses a consequential boundary without a contract, a test, or an accountable reviewer.
- Route work by consequence, reversibility, and uncertainty — not by a permanent ranking of model brands.
- The Supervisor Pattern has three gates: a scoped handoff, an auditable output contract, and review before promotion.
- Use direct execution for stable, reversible work; broaden search only when the decision terrain warrants it.
- Several agents using the same evidence and assumptions create false diversity, not independent judgment.
- Stop deliberating when a cheap reality probe can resolve the uncertainty more directly.
The useful question is which work can be delegated, what evidence must come back, and what may happen with the result.
Lower inference cost makes delegation attractive, but choosing an execution model and authorizing its output are separate decisions. Classification, extraction, comparison, formatting, and bounded research can move to worker-tier models when direct checks establish sufficient quality.
What the Supervisor Pattern Is
The Supervisor Pattern is a three-gate arrangement:
- Scoped handoff. Define the intended outcome, available means, local decision rights, and conditions the worker must flag. Distinguish continuing restrictions from temporary setup steps.
- Auditable output contract. Specify what the next owner must be able to decide or do. Return the supporting evidence in a checkable form: structured fields, source links, a comparison table, or a patch against a named precondition.
- Review before promotion. An accountable supervisor verifies the result before it reaches a consequential consumer or triggers a side effect. A delivered artifact can pass its task checks while the wider outcome remains unresolved.
The gates solve different problems. A worker may choose means within its assigned scope; changing the mission or crossing an external-action boundary requires the corresponding authority.
| Missing gate | Failure | What exposes it |
|---|---|---|
| Scoped handoff | The worker fills ambiguity or answers a nearby question | Canonical identifiers, exclusions, and stop conditions were never stated |
| Output contract | Fluent prose hides missing evidence and structural defects | Two returns cannot be compared or checked mechanically |
| Review gate | A plausible error reaches a consumer or side effect | Confidence is accepted as evidence of correctness |
The pattern resembles technical delegation: define the work, specify the deliverable, review before merge. A cheaper model is useful when it passes the required checks without shifting excessive repair work to the supervisor.
Make the Handoff Testable
Bounded extraction, cross-link mapping, table construction, classification, and constrained file edits have inspectable outputs. A handoff can specify which source facts must survive, which files may change, and how the result will be checked.
Asking for "concrete" analysis does not supply evidence for a number. An ambiguous model identifier leaves room for substitution. A fluent research summary can hide whether a claim came from a primary source or commentary. Each is a failure to test at the output boundary.
Require claim provenance, name exact identifiers, isolate write scope, and separate production from acceptance. Prompt instructions express the contract; direct checks establish whether the returned work met it.
A worker can be excellent at executing a task and poorly calibrated to decide whether the task was framed correctly, whether its evidence is sufficient, or whether the result should ship.Route by Decision Terrain
A supervisor should not turn every task into a committee. The amount of search and review should match the decision terrain.
First identify what is missing: an observation, an interpretation, an owner's decision, or execution. More reasoning cannot supply an unobserved fact or confer permission. Input volume is a separate variable: a large, checkable extraction may need less judgment than a short request whose assumptions determine the commitment.
| Terrain | Routing pattern | Supervisor focus |
|---|---|---|
| Stable and reversible | Direct execution or one bounded worker | Check the result with the cheapest direct test |
| Several plausible routes, reversible | Finite divergence across distinct mechanisms | Compare evidence, choose, then stop the unused branches |
| Stable but consequential | One implementation path plus direct and adversarial validation | Protect the irreversible boundary and rollback path |
| Several routes and consequential | Staged divergence, synthesis, then a gated decision | Stabilize the artifact before final review |
This is a better routing surface than a fixed split between "cheap work" and "smart work." Model capability changes. Consequence, reversibility, and uncertainty describe the decision that exists now.
Amazon's distinction between reversible and irreversible decisions is useful here: reversible choices should move with less ceremony, while hard-to-reverse choices deserve a different process. Google SRE's treatment of risk and reliability adds the other half: control effort should follow the consequence of failure rather than an abstract preference for maximum caution.
Three Failure Modes at the Boundary
| Failure mode | What it looks like | Architectural response |
|---|---|---|
| Confident specificity | Numbers, timeframes, citations, or entities appear without adequate support | Require provenance per consequential claim |
| Silent scope drift | The worker solves a nearby problem without declaring the substitution | Name canonical inputs, exclusions, and abort conditions |
| Aggregator fluency | Secondary summaries are treated as equivalent to primary evidence | Bind high-stakes claims to source hierarchy and supervisor review |
These failures create Review Debt: the cost of discovering downstream what should have been checked at the boundary. The immediate inference bill looks smaller, but the system pays later through rework, correction, and trust loss.
The answer is not to send every step to the most capable model. It is to spend scarce judgment where an error can cross into durable state. Worker models can generate breadth. Supervisors protect the promotion boundary.
False Diversity Is Not Review
Adding agents does not automatically create independent judgment. If several agents use the same sources, inherit the same assumptions, and evaluate through the same frame, their different wording is correlated elaboration.
A second route earns its cost when it introduces a distinct mechanism, evidence base, test, or failure surface. One agent might inspect source fidelity while another tests the implementation directly. One might defend the leading design while another searches for a cheap falsifier. Personas alone do not create that independence.
This matters for supervisors because superficial agreement can look like confidence. Five similar answers should not be counted as five independent observations. The supervisor must inspect why the routes agree and whether they could fail together.
Stop When Reality Is Cheaper
Supervision can fail in the other direction: more reviews, more agents, and more deliberation after the decision is already testable. When a reversible probe can resolve the uncertainty directly, run it.
A focused unit test is better than another opinion about whether the parser works. A dry run is better than another debate about the shape of the output. A small representative sample is better than exhaustive review when the sample can falsify the leading assumption.
The stopping rule should be part of the handoff: what evidence is sufficient, what result cancels remaining work, and what uncertainty is acceptable at this decision's consequence level.
The supervisor's job is not to maximize deliberation. It is to allocate judgment: broaden search when uncertainty warrants it, protect consequential boundaries, and stop when direct evidence has resolved the decision.Minimum Viable Implementation
- Choose one consequential output surface. Start where worker output can reach a client, publication, deployment, shared state, or another acting system.
- Define a reusable handoff. Include inputs, output shape, success criteria, exclusions, known failure modes, and the stop condition.
- Make evidence inspectable. Use the lightest structure that exposes missing fields, unsupported claims, and conflicting returns.
- Name the promotion owner. One person or authorized system decides whether the result advances.
- Log what review catches. Repair the handoff or test around recurring failures instead of relying on reviewer memory.
The pattern has a clear boundary. Stable, reversible tasks with a cheap direct test often do not need a separate supervisor. Consequential work may require stronger validation than one model reviewing another. And when the central uncertainty is human intent, ownership, or an exception, the right next step may be a conversation rather than another agent.
The durable asset is not a static hierarchy of models. It is the routing discipline that keeps execution cheap, evidence legible, and authority explicit as model capabilities change.
Related Reading
- Seven Models Judged Each Other — a cross-model evaluation that exposed identifier and review failures
- The Validator Asymmetry Principle — why bounded verification can be easier than generation
- Temperature 0 Doesn't Mean Deterministic — why stable settings do not guarantee stable outputs