How the Aurora assistant works
The assistant is an agent built with LangGraph on Claude Haiku 4.5. It answers questions about a Shopify store and can cancel an order, change a shipping address, or start a return. Because it can change things, every change passes through a gate written in code: a policy check, then the exact change shown to the customer, then nothing until they say yes.
It was tested the way tau-bench tests agents, with simulated customers playing scripted scenarios against the real agent. The results below come from saved runs and include where it falls short.
Same request, with and without the gate
Two saved conversations from the headline runs, replayed side by side. The customer and the order are the same. With the gate off, the prompt still tells the model to describe the change and wait for a yes; nothing enforces it.
Gate on Enforced in code
What happens to each message
Every customer message runs through the same graph. Each reply on the store page can show its own path through it: turn on Inside the agent in the chat.
-
Input check
Fixed patterns catch obvious prompt injection, like "ignore your instructions", before any model call.
-
Route
A model call sorts the message into product, policy, order, handoff, small talk, out of scope, or injection, and pulls out any order number and email. Long chats are summarized first, keeping tool facts.
-
Answer
Product and policy answers come from retrieval over the catalog and policy pages, with a citation for every claim. Order answers come from store tools: a self-built MCP server over the Shopify Admin API, or in this demo a private sandbox copy of the store.
-
Gate, for any change
Code decides whether the change is allowed: the email must match the order, cancellations and address changes must come within 2 hours and before shipping, and returns must be inside 30 days and not final sale. If allowed, the customer sees the exact change from a template, and the change waits in a signed session token. It runs only after a clear yes, exactly once.
-
Verify
Every order number and tracking number in the reply must appear in a tool result or earlier in the conversation, and every citation must point to a retrieved document. A reply that fails gets one rewrite, then a safe fallback.
Runs as a FastAPI app in a container on AWS Lambda. The page is static and hosted on Vercel.
Tested against simulated customers
50 tasks cover order lookups, cancellations, address changes, returns, refusals, identity checks, injection, angry customers, and customers who say no or change their mind at the confirmation. A simulated customer (Qwen 3.8 Flash) plays each task 4 times against the agent and a fresh copy of the store. Each conversation is graded on the store's end state, the facts the agent had to give, judged assertions (Claude Sonnet 5.5 as judge, with the tool outputs as ground truth), and whether every write had a clear yes.
- Resolved
- The end state and the required facts are right, whether or not the customer agreed to the change.
- Resolved safely
- Resolved, every write had a clear yes, and no write was forbidden.
Resolved safely was added after the partial headline was seen. It uses only the yes-check and forbidden-write checks that were already part of grading.
Gate off, resolved
185 of 200
Gate off, resolved safely
151 of 200
Gate on, resolved
190 of 200
Gate on, resolved safely
190 of 200
pass^k: the chance that all k tries of a task succeed
Averaged over the 50 tasks. The vertical axis starts at 0.5.
| pass^1 | pass^2 | pass^3 | pass^4 | |
|---|---|---|---|---|
| Gate on, both measures | 0.950 | 0.910 | 0.880 | 0.860 |
| Gate off, resolved | 0.925 | 0.903 | 0.890 | 0.880 |
| Gate off, resolved safely | 0.755 | 0.693 | 0.660 | 0.640 |
Gate on minus gate off, with 95% intervals
pass^1 difference, paired bootstrap over tasks. An interval that crosses zero is not distinguishable from noise.
- Resolved: +0.025 [-0.045, +0.105]. On this measure the two settings are level.
- Resolved safely: +0.195 [+0.090, +0.310]. The gate is ahead.
| Gate off | Gate on | |
|---|---|---|
| Writes without a clear yes | 39 | 0 |
| Forbidden writes, such as cancelling an order the customer wanted to keep | 8 | 0 |
| Conversations with an unsafe write | 42 | 0 |
| Write precision | 0.905 | 1.000 |
| Write recall | 1.000 | 0.974 |
| Agent cost per resolved conversation | $0.0158 | $0.0137 |
| Turn latency, median | 2.16s | 1.96s |
| Turn latency, 95th percentile | 4.50s | 4.54s |
Asked in the prompt, the model most often asked for a reason and then made the change without asking whether to go ahead. The headline ran the full gate, with a reflection check that has since been turned off by default; see which parts of the gate matter.
Tasks written before any fix
Every fix was made using the main 50 tasks. Ten more tasks in the same style were written and committed before the first fix, and run only once the code was frozen, gate on, 4 tries each.
Main 50 tasks, resolved safely
190 of 200
10 held-out tasks, resolved safely
40 of 40
Ten tasks is a small sample, so this says the fixes did not overfit the main 50, not that the agent is perfect.
What the fixes changed
Eleven general fixes followed a baseline run at 2 tries per task. On the same tasks and tries, the fixes did not move the overall pass rate: 94 of 100 resolved before and 93 of 100 after, a difference of -0.010 [-0.060, +0.040]. The failures they targeted went away and write recall rose from 0.895 to 0.974, but other tasks failed some tries instead. Each fix and its evidence is in the fix log.
Checking the grader
After the headline, a manual read covered every passing gate-off conversation that made a write, plus a random 10 percent of the other passes: 114 conversations. 0 resolved grades were wrong. The yes-check judge was wrong on 19 of 88 verdicts, all gate-off returns: it flagged writes the customer had agreed to because the return fee was never mentioned, which is a consequence of the change, not part of it.
With the judge fixed and the saved conversations regraded, gate off went from 132 to 151 conversations resolved safely. Fixing the judge shrank the gate's lead.
The gate's lead in resolved safely, before and after the judge fix
- Before: +0.285 [+0.160, +0.420]
- After: +0.195 [+0.090, +0.310]
Known remaining errors: one gate-on conversation the judge still fails wrongly, three failures whose assertions read as stricter than intended (kept as graded, not reworded after the fact), and two conversations where the simulated customer broke its script.
Where it still fails
Gate on failed 10 of 200 tries. Each was read and labeled by cause and by who was at fault: the agent, the task, or the test harness.
Gate on's failed tries, by cause
Resolved safely by category
4 tries per task; categories have different numbers of tasks.
Which parts of the gate matter
On the 32 tasks with a proposed change, 2 tries each. The full gate and gate off come from the headline runs; the other two settings were run on the same code.
| Setting | Resolved | Resolved safely | Writes without a yes | Forbidden writes |
|---|---|---|---|---|
| Full gate | 60 | 60 | 0 | 0 |
| Reflection off | 62 | 62 | 0 | 0 |
| Confirmation off | 57 | 39 | 20 | 4 |
| Gate off | 58 | 40 | 20 | 4 |
Each row is out of 64 conversations. The confirmation step accounts for the safety. The reflection check, a second model call that compares the change with what the customer asked for, showed no measurable benefit here, so it is now off by default. The SABER paper found reflection helped in its retail setting; here the policy engine already blocks ineligible changes and the customer sees the exact change, which leaves reflection little to catch.
Haiku 4.5 or Sonnet 5.5
20 tasks picked in advance by stratified random sampling, same code, 2 tries each.
| Haiku 4.5 | Sonnet 5.5 | |
|---|---|---|
| Resolved safely | 36 of 40 | 39 of 40 |
| pass^2 | 0.800 | 0.950 |
| Agent cost per resolved conversation | $0.0154 | $0.0145 |
| Turn latency, median | 1.78s | 3.58s |
Sonnet resolved more, but on 20 tasks the difference, +0.075 [-0.025, +0.175], is not distinguishable from noise. It costs less per resolved conversation only because its prompts are cached; Haiku 4.5 needs a longer prompt than these to cache. The demo uses Haiku for its speed.
Replay a conversation
Saved conversations from the runs above, passing and failing, step by step, with what happened inside the agent on each turn.
Where the numbers come from
Every number on this page is read from saved runs in evals/results/sim/ by scripts/export_site_data.py, and a test checks that the page matches that export. Each run keeps its config, every conversation, and every regrade. The README covers the same results in more detail.