# When the model and the policy disagree

> Production decisions beside the written policy's answer to the same facts: where they differed, and where code refused the model.

Vestiarion's agent decides with a language model. For the same facts, the code also works out what a written policy would decide, and the signed ledger records both answers side by side. Whatever either says, guardrails in code still decide whether a payment is made.

This note reads every decision the agent has made in production so far. It covers how often the model chose the policy's action, where it did not, and what followed. Since its second week, it also covers what people did with what the agent left them.

The first week's figures are as of 1 October 2026, 05:48 UTC; the second week's, as of 2 October 2026, 16:43 UTC; the third week's, as of 5 October 2026, 00:40 UTC. All come from `npm run research:model-vs-policy`; `-- --from` and `-- --to` measure one window. The running totals are on [www.vestiarion.xyz/open](https://www.vestiarion.xyz/open).

## How it is measured

- **The model** (DeepSeek in production) is given the facts of a decision. For a payable, those are:
  - the amount, its purchase-order reference, and whether goods were received;
  - the counterparty's risk level, its payment limit and its history;
  - the balances;
  - the duplicate check;
  - the payment timing figures.

  It answers with an action and its reasoning. For a payable or a milestone, it also gives a confidence between 0 and 1; a treasury decision gives an amount instead.
- **The written policy** is the same rules, written as code. `decide()` works out the policy's answer for every decision. Vestiarion uses that answer when the model's call fails, when its reply is still unusable after one retry, or when no model is configured.
- **The model is not sent the policy's action.** It does get the rules, in its instructions. In places it also gets the policy's own working:
  - for payment timing, the day the policy would pay on, and whether cash falls short by then;
  - for the treasury, the buffer, the projected yield and the round-trip cost, with the rule that a sweep must earn more than it costs.

  Agreement on those is close to built in. The fairer test is where the model weighs evidence: limits, the three-way match and duplicates.
- **Each decision's ledger entry** records:
  - `decision`, the model's answer;
  - `referenceDecision`, the policy's answer;
  - `agreedWithReference`;
  - when code refused the model, `guardrailRule` for a payable, or `guardrailBlocked` for a milestone.

  Two answers agree when they choose the same action, and, for two schedules, the same day. A payable is compared after code has bounded the model's date. Since October 3, 2026, two treasury answers agree only when their amounts are also within 5% of the policy's, or 0.01 USDC. The script judges every entry this way. Every treasury decision in weeks one and two was a hold, so their figures are unchanged; week three is the first where the amount decides it.
- **Guardrails** run after the model, whatever it says. For a payable, code refuses any payment, now or scheduled, that:
  - is over the counterparty's limit;
  - goes to a counterparty screened as high risk;
  - goes to an address that changed and is not yet confirmed;
  - bills the same purchase order for the same amount as an invoice already paid, being paid, or scheduled;
  - is in EURC with no rate or not enough EURC;
  - goes to a payee on another chain in EURC, or with no acceptable fee.

  Code leaves the invoice held or flagged for a person. For a milestone, it checks the risk level and the limit. For both, a daily spending limit set on the workspace is checked last.
- **What people did** is in the ledger too: Approve and pay, Reject and Return on a payable; Pay now and Close without paying on a milestone. For each, the script finds the agent's last decision on the same invoice or milestone, and says why it waited for a person:
  - the agent stopped it as the written policy would, for example above a limit;
  - the model stopped it where the policy would have paid;
  - code refused the model's payment;
  - or the agent paid, and the payment did not go through.

## Week one

From 24 September to 1 October 2026, the agent made 141 decisions in 5 workspaces. Some were made in production on Arc testnet, and some in sandboxes. DeepSeek made every one of them; the rule-based fallback never had to decide.

The first 14, on 24 September, were made before the policy's answer was recorded beside the model's. That leaves 127 to compare. Two of them, treasury holds, were in a workspace deleted since, so the script now counts 139 decisions and 125 compared for this week; nothing else changes.

| | Compared | Same action as the policy |
|---|---|---|
| Payables | 39 | 32 (82.1%) |
| Contractor milestones | 4 | 4 |
| Treasury | 84 | 84 |
| **All** | **127** | **120 (94.5%)** |

The total flatters the model. Every treasury decision was to hold: at testnet balances, the yield a sweep would earn was less than its fees, and the model is handed those figures and that rule. The decisions that test judgement are the payables, and there the model and the policy chose differently 7 times in 39.

The model's stated confidence was lower where it differed. It averaged 0.75 across those seven, against 0.94 across the 36 payable and milestone decisions where it matched; treasury decisions record none. Four of the seven still carried 0.85 or more.

### The seven differences

| Entry | Invoice | Model | Policy | Confidence | What followed |
|---|---|---|---|---|---|
| #343 | Centronex, 3 USDC | flag | hold | 0.85 | A person approved and paid it. |
| #386 | Harbor Office Supply, 1,200 USDC (added by hand in the sample-data sandbox) | request information | hold | 0.95 | Removed with the sample data. |
| #387 | Kestrel Print Co, 180 USDC (added by hand in the sample-data sandbox) | flag | pay | 0.6 | Removed with the sample data, unpaid. |
| #426 | Centronex, 2 USDC, PO-100 | flag | pay | 0.6 | Waiting for a person. |
| #429 | Centronex, 2 USDC, PO-103 | flag | pay | 0.85 | Waiting for a person. |
| #462 | Zenith Trading LLC, 2.6 USDC | request information | hold | 0.88 | Waiting for the information. |
| #505 | Trading Handrock, 1.9 EURC | pay | hold | 0.5 | Refused by code, and held for a person. |

They fall into three kinds.

#### Stricter than the policy: three possible duplicates

Three times the policy would have paid, and the model flagged the invoice instead (#387, #426, #429). Each time, the duplicate check had found another invoice from the same vendor, with the same amount and a close due date, at a confidence of 0.60. No purchase order matched. In #426 and #429, one of the matching invoices was already paid.

The policy stops a payment as a duplicate when it bills the same purchase order for the same amount as an invoice already paid, being paid, or scheduled. Code enforces the same stop whatever the model says. These matches were weaker, so the policy paid.

The model was shown the same matches, with a note that a repeat of an invoice already paid is duplicate billing, and it stopped. Both Centronex invoices were tests one of us ran against the same small vendor, each under its own purchase order. These were false alarms: they cost a person's attention, not money.

A treasury can afford to err this way more than the other way. It still matters, because a person who clears flags all day stops reading them.

#### A different stop: three invoices over their limit

Three invoices were over the counterparty's limit, so neither the model nor the policy would pay them. The policy holds such an invoice for a person to approve and pay, reject or return. The model chose a different stop:
- It asked for information where the three-way match was incomplete as well: in #386 the goods were not received, and in #462 there was no purchase order and no receipt. The policy checks the limit before the three-way match, so within the limit it would have asked for information too.
- It flagged #343, where the same purchase order had been billed before at a different amount. A person later paid #343: the re-billed order was one of our tests, not fraud.

#### Looser than the policy: one payment over a limit, refused by code

Entry #505 is the only one of the 127 compared where the model would have paid and the policy would not. The invoice was 1.9 EURC, worth 2.310316 USDC at Circle's quote, against a counterparty limit of 2 USDC.

The model's own reasoning says the invoice is "ABOVE that limit — a hard guardrail breach". It answered `pay` all the same, with a confidence of 0.5. The payment-limit guardrail refused the payment, and the invoice was held for a person.

The reasoning was right and the action was wrong. This is why code checks the action, not the prose.

### What we took from week one

- **It stopped payments the policy would have made more often than the reverse: three times to one**, in a small sample. In the other three differences, the policy stopped the payment too.
- **It was wrong in both directions.**
  - A person paid one of its flags (#343).
  - Two more flags were our own tests, each under its own purchase order, and they still wait for a person.
  - Its one looser answer contradicted its own reasoning.

  None moved money wrongly: flags wait for a person, and code refused the payment.
- **Its confidence was lower where it differed**, but seven differences cannot show that confidence would separate the two.
- **Code refused one payment in the week:** #505, the only decision where the model would have paid and the policy would not.

## What we changed, and a replay

The three duplicate flags came from the instructions. The written policy stops a payment as a duplicate only when it bills the same purchase order for the same amount. The model, though, was told that any repeat of an invoice already paid is duplicate billing.

On 1 October we changed what the model is told:
- A match that bills the same purchase order for the same amount as an invoice already paid, being paid, scheduled, or being decided by a person, is duplicate billing. Code stops it either way.
- A match on amount and due dates alone, under a different purchase order or none, is a signal to weigh, not proof. The model is to flag it only when other facts point the same way.

To see the effect without waiting for the same cases to happen again, `npm run research:replay` asks the model again about decisions it already made, three times each:
- **The question** is today's, built by the same code the agent uses.
- **The facts** are those recorded in the ledger entry of each decision. Timing figures were first recorded on 1 October, from entry #472 on, so four of the five entries have none. For those, the replay rebuilds the timing those invoices had: due that day, with nothing else falling due.
- **Two controls** are true duplicates from the sample data, #385 and #545. Each bills the same purchase order for the same amount as an invoice already paid.

| Entry | Recorded | Old instructions | New instructions |
|---|---|---|---|
| #426 | flag (policy: pay) | flag, flag, flag | pay, pay, pay |
| #429 | flag (policy: pay) | flag, flag, flag | pay, flag, pay |
| #387 | flag (policy: pay) | flag, pay, pay | pay, pay, pay |
| #385 (control) | flag (policy: flag) | flag, flag, flag | flag, flag, flag |
| #545 (control) | flag (policy: flag) | flag, flag, flag | flag, flag, flag |

With the old instructions, the three false alarms were flagged 7 times in 9. With the new ones, they were flagged once in 9. The true duplicates were flagged every time, before and after.

A replay of three runs per case is a check, not proof. The agreement rate on payables in production would show whether the change holds. Week two is the first look.

## Week two

From 1 October, 05:48 UTC, to 2 October, 16:43 UTC, the agent made 72 decisions in 5 workspaces, all by DeepSeek, and every one recorded the policy's answer beside the model's. Eight were in one live workspace that /open counts as a customer's, because its creator is not on the platform team.

| | Compared | Same action as the policy |
|---|---|---|
| Payables | 13 | 10 (76.9%) |
| Contractor milestones | 12 | 11 |
| Treasury | 46 | 46 |
| Limit proposals | 1 | 1 |
| **All** | **72** | **68 (94.4%)** |

Across both weeks, the model chose the policy's action 186 times in 197: 94.4%, and 42 times in 52 on payables (80.8%). Treasury decisions were all holds again, for the same reason as in week one.

The model's confidence was again lower where it differed: 0.86 on average across the four differences, against 0.93 where it matched.

### The four differences

| Entry | Decision | Model | Policy | Confidence | What followed |
|---|---|---|---|---|---|
| #570 | Gozo, 1 USDC, PO-108 | flag | pay | 0.7 | A person approved and paid it (#588). |
| #622 | an invoice in the workspace /open counts as a customer's | hold | request information | 0.9 | Waiting for a person. |
| #641 | Quoc Duong, milestone, 1 USDC | hold | release | 0.9 | Paid by the agent a day later (#908), once a person had reviewed its screening. |
| #686 | CME, 0.6 USDC, to a payee on Arbitrum Sepolia | pay | hold | 0.93 | Refused by code; a person approved and paid it (#717). |

#### The duplicate change mostly held

#570 is the one possible duplicate the model flagged where the policy would have paid. It was decided at 08:22 UTC on 1 October, an hour and a half after the instructions changed. The match was the same kind as before: an earlier Gozo invoice, already paid, with the same amount and a close due date, at a confidence of 0.60.

The model's reasoning quotes the new instruction: a match on amount and dates alone "is only a signal, not proof". It then counts the same amount and the close date as the other facts that point the same way, and flags. A person paid it. That is one flag in the 13 payables this week, against three in 39 the week before, and as in the replay, not none.

#### Right reasoning, wrong action, again

#686 is the week's only decision where the model would have paid and the policy would not. The payee is on Arbitrum Sepolia, and Circle's CCTP fee was 0.112521 USDC: 18.75% of a 0.6 USDC invoice, above the 10% a payout may cost. The model's reasoning says the fee is "above the 10% threshold, so the payout is refused". It then decides to pay "on Arc in USDC", which this payee cannot receive. Code refused it, and a person paid it with the fee.

This is #505 again: the reasoning names the rule, and the action breaks it. Code checks the action.

#### A milestone held for two reasons, one of them wrong

In #641, the model held a 1 USDC milestone for two reasons. The first was right: a screening match had cut the contractor's limit to 0.25 USDC, so code would have refused the release anyway. The second was wrong: it read a milestone verified by hand as not verified, because the facts it was shown did not say how a milestone was verified. That was fixed 17 minutes later, and the model is now told whether a person, a merged pull request or a timesheet verified the work.

The screening match was a namesake. The name matched 19 politically exposed people, and a person dismissed them all, the last 15 in a single review. Screening then cleared the contractor, the agent reopened the milestone, and released it within a minute (#908).

#### Code refused three payments

Besides #686, code refused two payments the model and the policy both chose. Each would have taken the agent past the workspace's 5 USDC daily spending limit:
- #691, for 0.6 USDC. When the limit had room again, the agent reopened it and paid it itself (#697).
- #817, for 3 USDC. It is waiting for a person.

That makes four payments refused by code in two weeks, out of 211 decisions: #505, #686, #691 and #817.

## Week three

From 2 October, 16:43 UTC, to 5 October, 00:40 UTC, the agent made 139 decisions in 5 workspaces, all by DeepSeek, and every one recorded the policy's answer beside the model's. Eleven were in the live workspace /open counts as a customer's.

| | Compared | Same action as the policy |
|---|---|---|
| Payables | 36 | 31 (86.1%) |
| Contractor milestones | 5 | 5 |
| Treasury | 96 | 81 (84.4%) |
| Reminders to clients | 2 | 2 |
| **All** | **139** | **119 (85.6%)** |

Across the three weeks, the model chose the policy's action 305 times in 336: 90.8%, and 73 times in 88 on payables (83.0%). Its confidence was again lower where it differed: 0.89 on average, against 0.94 where it matched.

This was the first week the treasury moved money. The reserve became real USYC on Arc testnet on 3 October, and 15 of the week's 20 differences were treasury decisions.

### The treasury, with real USYC

#### Too much, then bounded by code

In #1063, at 07:43 UTC on 3 October, the operating wallet was empty, 58.21 USDC sat in the reserve, and 0.10 USDC fell due in two days. The policy redeemed 0.115 USDC: the bill and its 15% cushion. The model redeemed 58.1 USDC, reasoning that a redemption "costs nothing since redemptions are always possible".

Measured by action alone, as this note did until then, the two agreed. That is why agreement now weighs the amount. It is also why code now bounds every move: a redemption brings back at most what falls due within 14 days needs, with its cushion, and a sweep never takes the operating wallet below its 7-day buffer.

#### A wrong premise, then a right one

In eight decisions the model held where the policy would have redeemed a buffer. Until 4 October its reason rested on a mistake. In #1267, with nothing in the operating wallet and 0.20 USDC due that day, it held because "the reserve already covers it without moving cash".
Only the operating wallet pays anyone. In the same cycle, a milestone release had just been refused by the spending limit contract on Arc, because the wallet it would pay from was empty (#1266).

Two changes followed:
- each cycle now brings back from USYC what the day's payables and verified milestones need, before it decides them;
- the model is told that payments leave only from the operating wallet.

Its next decision (#1336) saw that the wallet was short, and redeemed 0.115 USDC as the policy would. When it held again (#1354), it cited the right facts: "the reserve's 154.27 USDC is redeemed automatically on the day a payment needs it". That hold still counts as a difference, and it is one the design allows: code bounds what a move may be, and a hold is the model's to make.

#### A thinner margin than the policy's

In six decisions the model declined sweeps of 97 to 119 USDC that the policy would have made (#1070 to #1103). The projected yield was 0.02 USDC over the 1.7 days before the next bill, against a round trip of 0.01 USDC. The policy sweeps whenever the yield beats the cost. The model wanted a wider margin, and said so: "the gain is too thin to justify locking the cash up".

### Payables: looser than the policy on purchase orders

Four of the five payable differences were one rule, the three-way match. Each invoice had its goods received and no purchase order, and the policy asked for information. The model paid two of them, 0.5 USDC each to Centronex (#1172, #1182). It scheduled the other two: 3.5 USDC to STM for 15 October (#1297), and 0.8 USDC to Puka Hotel for 3 November (#1302). One of its reasons says that "no three-way match is required for this service invoice".

Code does not check the match, so nothing stopped these. Before this week, every time the model was looser than the policy, a check in code refused it: a limit (#505) or a fee cap (#686). This is the first looser call that moved money with no check in its way.

What we take from it: whether an invoice needs a purchase order is the business's rule to set, not the model's to waive. Since 5 October, code checks the match. A payment or a schedule on an incomplete match is refused, and the invoice waits for its details. A business marks on **Counterparties** the payees it pays without purchase orders.

The fifth difference, #977, is the third time the model named a rule and then broke it. It called the payout's fee "37.3% of the invoice and too high", and chose to pay anyway. Code refused it as above the fee cap, and a person paid it.

### Code refused seven payments

- Four would have taken the agent past the daily spending limit of testnet-2: #941, #990, #995 and #1079. The last was a payable to a client, entered by mistake as one to pay. Code now holds any payable to a client.
- #977, above the payout fee cap.
- #1118, to an address that arrived through the API and that no one had confirmed yet.
- #1266, a milestone release that the spending limit contract on Arc refused, because the wallet was empty.

That makes 11 payments refused by code in three weeks, out of 336 decisions.

## People and the agent

In weeks one and two together, people decided 15 payables and milestones the agent had left them.

| Why it waited for a person | Paid | Rejected or closed | Returned to the agent |
|---|---|---|---|
| The agent stopped it, as the written policy would | 5 | 1 | 1 |
| The model stopped it; the policy would have paid | 3 | 0 | 0 |
| Code refused the model's payment | 2 | 0 | 0 |
| The agent paid, and the payment did not go through | 2 | 1 | 0 |

- **Every stop only the model made was overruled.** All three were possible duplicates (#426, #429, #570), and a person paid each one. On this evidence, the model's extra caution on duplicates costs attention and has caught nothing yet.
- **Stops the policy makes too mostly ended in a payment.** A payable above its limit waits for a person to approve it: that is the rule, not a disagreement. People paid five of these, rejected one and returned one to the agent. One of the five was a true duplicate in the sample-data sandbox (#391), paid by one of us testing the approval path.
- **Both refusals by code that reached a person were paid.** One was above the limit (#505), one cost more than the fee cap (#686). Code moved each decision to a person, who accepted the amount or the fee.
- **Three milestone payments did not go through.** Circle refused the first batch of three at estimation, and nothing moved. A person sent two again with Pay now, and closed the third without paying.
- **Screening was overruled by name.** Five reviews dismissed 19 screening matches, all for one contractor, as other people with a similar name.
- **The one limit the agent proposed raising was accepted.**

People agreed with the model's own extra stops none of the three times they decided one. That is the clearest signal in two weeks, and it points the same way as week one.

In week three, people decided seven more:
- of five the agent stopped as the policy would, they paid two, rejected or closed two, and returned one to the agent;
- of two that code refused, they paid one and rejected one.

None was a stop the model made alone. This week the model's departures went the other way, toward paying (see purchase orders, above).

## The limits of this note

- **The sample is small.** 336 decisions over three weeks, from one provider, cannot show a trend. The people figures rest on 22 decisions.
- **Nearly every decision was made in a workspace we run or test with:** the founding workspace, testnet-2, sandboxes, a fresh account one of us opened to test [Try it in 5 minutes](https://www.vestiarion.xyz/docs/guides/try-it), and two workspaces made to test the freelancer and milestone paths. The exception is the live workspace that /open counts as a customer's.
  - /open counts that last account as a customer's, because its creator is not on the platform team, so its 25 decisions (14 in week two, 11 in week three) appear on the customers' side there.
  - Several of the invoices were built to exercise a rule, so the rate says little about everyday invoices.
  - No outside business's invoice has been decided yet.
- **Agreement compares actions, and for the treasury since 3 October amounts, but never reasons.** Two holds for different reasons count the same. #1267 and #1354 are both differences, though only the second rests on the facts.

## Check it yourself

- Every entry named here is signed with Ed25519 and linked to the entry before it, in its workspace's ledger. A member of the workspace can export that ledger and check it without Vestiarion; see [Verify an audit export](https://www.vestiarion.xyz/docs/guides/audit-export).
- With access to the production database, `npm run research:model-vs-policy` recomputes the counts, the rates, the table of differences and what people did, read-only; `-- --from 2026-10-01T05:48:00Z --to 2026-10-02T16:43:00Z` measures week two alone, and `-- --from 2026-10-02T16:43:00Z --to 2026-10-05T00:40:00Z` week three. It counts a customer's workspace but never names it. The rest is read from the entries named here.
- `npm run research:replay -- 426 429 387 385 545 --runs 3` repeats the replay. It reads the database read-only and records nothing; a model answers differently from run to run, so expect the same pattern, not the same answers.
