τ³-BENCH · SOPBENCH · τ²-BENCH

Keep the model.
Enforce the policy.

A deterministic layer that checks every tool call an LLM agent proposes, before it runs. No training. One declaration file per domain.

agentpolicy layertool runsrefusal

An open 27B model, not fine-tuned, goes from 49.5 to 57.2 pass^1 on τ³-bench banking_knowledge. That is above the top public entry, 55.2. The model, its weights, the tools and the benchmark's policy documents are unchanged; only the layer is added.

FIGURE 1 · τ³-BENCH BANKINGOne layer lifts the same model past the board leader
Grey is the model without the layer, indigo the same model with it, the hollow ring the top public entry. 97 tasks × 4 trials · user simulator gpt-5.2 · same harness for all arms. pass^k is the share of tasks solved in all k of k trials. The no-layer pass^4 is undefined because one simulation hit an infrastructure error. Development disclosure: the banking rules were written on these tasks; rules traced to reference solutions were removed before this run (see How we measured).

PERFORMANCE

+7.7

pass^1 on banking from the layer alone. Tasks failed in all four trials: 35 → 25.

COST TO MOVE

1 file

A new domain is a new declaration file. The engine holds no domain words and stays the same.

RELIABILITY

0 blocks

correct actions wrongly blocked on 20 unseen airline tasks. Every decision is deterministic and logged.

The failure

Slips the agent knows better than

A customer-service agent reads the policy, calls tools, and writes to a database. Most of its failures are not missing knowledge. They are small, checkable slips: a write before the required lookup, a sum done in the model's head, a value that nobody said. One slip is enough, because a write to a record cannot be taken back.

FIGURE 2 · ONE CANCELLATIONTrust the customer, or trust the record?
CUSTOMER
MODEL
POLICY LAYER
TOOL · RECORD
1I need to cancel my reservation. I booked it about 10 hours ago.
2get_reservation_details
3record created_at 2024-05-02 · policy clock 05-15
4cancel_reservation(…) takes the customer's word
executed
reservation cancelled

DENY · NOT EXECUTED

✕ booked within 24 h?
record 05-02 · policy clock 05-15

① the unexecuted call cancel_reservation

② the fact from the record created_at 2024-05-02

③ the policy's cancellation sentence, verbatim

not executed
record unchanged
5refusal (no value, no next step)
6The record shows May 2, so it is past 24 hours and I can't cancel on that basis. Let me check the other conditions.
Task failed. The agent believed "10 hours ago". The record says May 2 and the policy's clock says May 15, far outside the 24-hour window.
The refusal holds only the unexecuted call, the fact read from the record and the policy sentence. It proposes no value and no next step. The model writes its own reply.
Adapted from τ²-bench airline task 48 (IDs elided; record and refusal abridged; the model's reply is illustrative). The real refusal writes ①②③ as one paragraph and quotes the policy sentence verbatim. In recorded GPT-4o-mini runs without this rule the cancellation went through and the task failed; replaying the rule on those runs, it fires at exactly that call. Source: VOCAB_EXT.md §5-2.

The idea

Enforce only what a program can decide

Some policy conditions can be settled by looking: did a lookup happen, what a field in a fetched record says, whether a value appears in a tool output. Others need judgment: is this request reasonable, what did the customer mean. The layer enforces the first kind and leaves the second kind to the model.

FIGURE 3 · CLOSED VS OPEN CONDITIONSIf the answer is in the record, the layer decides

CLOSEDthe layer enforces

The answer is fixed by events, record fields and strings in the transcript.

  • eventWas the reservation read before cancel?
  • recordIs created_at within 24 h of the policy's clock?
  • stringDoes the ZIP code being written appear in a tool output or customer message?
  • arithmeticDoes the refund equal the value computed from the fetched records?

OPENthe model decides

The answer needs interpretation. The layer may show the rule text, but never blocks on it.

  • intentDoes the customer really want to close the account?
  • toneIs this a complaint that should be escalated?
  • relevanceWhich of three products fits what they described?
  • meaningIs "a few days ago" inside the window?
Why this split matters. A check that needs judgment would need another model, and another model makes its own mistakes. Restricting enforcement to closed conditions means the layer cannot misread a request, and every denial is explained by a fact the agent can see.

The mechanism

One turn, step by step

As you scroll, the part of the figure for each step lights up and a dot follows the path the information takes.

1 · PROPOSE

The model reads the conversation and proposes its next turn: a tool call or a message. Nothing has run yet.

2 · READ

The engine loads the domain's declaration file and the transcript so far: which tools ran, what they returned, what the customer said.

3 · CHECK

Seven levers test closed conditions on the proposed call. Each lever either stays silent or reports one violated fact. Each lever is explained in "Seven levers" below.

4a · PASS

No violation: the call goes to the real tool unchanged. The layer never edits arguments.

4b · DENY

A violation: the call does not run. The model receives the call, the unconfirmed fact and the policy sentence, and proposes again. A small fixed budget caps retries, so the dialog never stalls.

5 · LOG

Every pass, denial, injected computation and retry is written to an audit log with the tool output it relied on. The same input gives the same decision.

Why the refusal carries no value

If the layer wrote the fix ("cancel with refund 0"), a mistake in a rule would become a mistake in an action. Returning only the fact and the policy sentence keeps the model responsible for the next move, and keeps the layer from acting as a second agent.

Seven levers

One lever per kind of slip

The levers came from reading failed trajectories and naming each recurring slip. Pick one to see what it reads, what it decides, and what happens to the score when it is switched off.

FIGURE 4 · LEVER EXPLORERReads → decides → does, and what happens without it
Examples are illustrative and written for this page; they are not the declaration's rule text. Ablation: tasks where the full layer fired, one lever switched off at a time. The rose band is at or below the no-layer score. Source: ABLATION*.md.

Moving to a new domain

The engine stays. One file changes.

The engine contains no domain words: no "card", no "flight", no "fee". Everything specific to a domain lives in one declaration file: which lookups a write needs, which values must be sourced, which computations are available, which policy sentences go with which tool.

FIGURE 5 · SAME ENGINE, SWAPPED DECLARATIONSSwap the file, and the same engine runs a new domain
Moving to τ² airline needed four operators (hours_between · count_of · prefix · eq) and one operand (now) added to the engine (+74 lines). After the change, banking decisions over 6,197 recorded turns were byte-identical. Source: VOCAB_EXT.md.

What the numbers say

Gains where models break closed rules

We swapped only the declaration and ran the same layer on two more benchmarks and ten model configurations. The gain is large where the model commits closed-rule slips and disappears where it does not.

FIGURE 6 · TRANSFER, pass^1Same layer, other domains and models

Scores: grey dots without the layer, indigo dots with it. Gain vs baseline: x is the no-layer score, y the gain; filled dots are airline, hollow dots SOPBench library, grey dots retail. Hover a dot for its values. On airline the gain shrinks as the baseline rises; the 27B model already follows that policy and the layer fired on one task. Retail shows no gain for any model: its failures are dominated by open judgment and simulator behaviour. SOPBench hotel (not shown) gains +11.0 under the official scorer and −0.4 after correcting a scorer defect; we do not claim it.
FIGURE 7 · SWITCH ONE LEVER OFFWithout grounding (LB3), all three settings fall into the rose band
Grounding (LB3) carries the stack. Without it, the other levers inject computed values and policy text that the model mixes with its own guesses, and all three settings fall to or below the no-layer baseline. LB1's weight follows how often a model breaks ordering: GPT-4o-mini hit 40 order denials and loses 9.2 without it; Qwen3-8B hit 10 and loses nothing. LB4 is slightly negative and not part of our claim.
FIGURE 8 · RELIABILITYOn unseen tasks, and next to another method

Unseen airline tasks

Rules written from 30 training tasks only, scored on 20 unseen tasks. GPT-4o-mini 31.2 → 52.5, with 0 correct actions blocked.

Against a published gate, same harness

Reason Less, Verify More (2607.07405) ships four airline-specific gates. Re-run under our simulator and tasks. Simulations lost on fired tasks: this layer 1, the gates 5.

Where it does not help

Three places with no gain

≈0

τ² retail, all 4 models

Failures there need judgment (which item, which variant) or come from the user simulator ending early. Few are closed-rule slips.

≈0

τ² airline, Qwen3.8-27B

The model already follows the policy. The layer fired on one task, net zero. A strong model leaves little to catch.

−0.4

SOPBench hotel

+11.0 under the official scorer, −0.4 after fixing an OR→AND scorer defect on 56 tasks. The official gain is an artefact.

Where the layer does not help it does not hurt: losses of 0 to 3 simulations out of hundreds, none traced to a wrong denial.

How we measured

What could be miscounted as our gain

Same engine instance for both arms

Switching vLLM instances changed the first response in 30 of 66 SOPBench tasks and flipped 8.9% of non-fired outcomes, the size of a real effect. Every A/B here ran both arms on one instance.

Gains counted on fired tasks

A gain is attributed to the layer only on tasks where it actually denied or injected something. Tasks where it stayed silent report the noise band (SOPBench library: −0.25 tasks per run).

SOPBench scorer defects

The official scorer re-reads calls with eval(), so JSON booleans fail and dates are read as arithmetic; goal actions that return a value always score as failures; on hotel an OR in the reference graph is folded to AND on 56 of 195 tasks. Disabling our fixes reproduces the published scores of all 21 official function-calling models. We report the official number first.

User-simulator early stop (τ²)

In 32 of 80 audited failures the simulator says "Yes" and ends the conversation in one message, so a compliant agent never gets the turn to act. The ceiling is shared by every entry that uses this simulator.

Development disclosure (banking)

The banking rules were developed on the 97 evaluation tasks, looking at rewards and reference solutions. We audited every rule by where it came from: (A) quotes the policy the agent receives, (B) general expert practice, (C) chosen because it matched reference solutions. All C items were removed before the submitted run. This entry is not test-blind; the held-out airline check is.

Run-time information boundary

At run time the layer reads the proposed turn, the tool outputs the agent received, the policy documents the agent receives and the declaration file. It has no access to task goals, reference actions or evaluator state.