More work to check, and less trust in the answers
Volume
The work that needs checking is about to skyrocket.
So checking has to get faster than a person reviewing a sample.
Trust
Confidence in what an agent hands back is going the other way.
So work nobody trusts has to be re-doable on demand.
Those two pull against each other.
Re-derivation resolves both.
The answer gets recomputed from the declared inputs under a pinned rule, by code the
doubter runs on their own machine, in a third of a second. What follows is what that is
worth against a regulated wage bill.
The five numbers, banking and insurance
| $2.67bn | A year off one large bank's cost base, $258m at one
large insurer. Each figure built from that firm's own disclosed headcount and wage bill,
worked line by line on its own tab. |
| 29–32% | Of the wage bill sitting in work where the rule is
written down and somebody downstream can dispute the answer. Not a share of revenue, of
profit or of headcount. |
| $0 | For the verifier, for good. Apache-2.0, offline, on the
auditor's own machine. The only invoice is the code that packages what the firm's systems
already hold. |
| 0.335s | To re-derive one answer, median across 40 pilot bundles,
with no network on the verdict path. Checking stops being the thing that gets
rationed. |
| 2 Aug 2026 | The first dated obligation is already in force, and
the high-risk record-keeping article lands 2 December 2027. Every rule asks the same
thing: show how the machine got there. |
What goes wrong when an agent pays an invoice
Routine, valid, matched to the purchase order and the goods receipt. Nothing about it needs
a person. Then one bank-account field comes back changed, and real money goes to the wrong
place.
So the agent proposes. The verifier decides.
Any model, any code, any workflow may produce the figure. A pinned rule recomputes it from
the declared inputs, and that rule is the metric for correctness, never the agent's
method.
What the verifier is: open source, agnostic, re-derivable
This changes what an agent can be trusted to do. From answers that
need faith, to work that can be accepted: bounded, monitored, receipted, independently
re-derivable. That is the line between assistance and delegation.
The share of the wage bill that comes off
Six kinds of work, and what re-derives in each
The split is load-bearing. Deterministic work is re-run: recomputed from the declared
inputs and rules, and the committed output reproduced bit-for-bit. An agent output is
never re-run, because no model reproduces its own words. It carries a receipt instead:
the sources it used, the criteria it included and excluded, re-checkable against the declared
evidence.
-
01
Claims and underwriting settled on the filed rule
45–50%
Prove every adjudicated claim and every rated premium traces to the filed terms and the declared data, before the supervisor, the auditor or the policyholder asks.
Legacy · deterministic
the computation is re-run
- Premium rating from filed rate tables: the quote reproduces bit-for-bit
- Claims adjudication against policy terms: recomputed from the declared policy
- Deductible, co-insurance and sub-limit application: re-derived, not re-asserted
- Reinsurance treaty settlement: the ceded figure re-derives or REJECTs
AI systems
the agent shows its work
- ML pricing models: each price shows the factors it stood on
- Claims triage and fraud scoring: every referral cites its exact signals
- Underwriting risk scores: a receipt of inclusion and exclusion criteria
- Straight-through decisions: the evidence each one rested on, re-checkable
The number
63% of injury claims settled automatically
operator disclosure
The trigger
AI Act Art 12 · 2 Dec 2027
life & health pricing is high-risk
Re-derivation settles that the number matches its declared inputs and filed rules. Whether the rate table or the triage policy is fair stays an underwriting judgement.
-
02
Surveillance, screening and financial-crime alerts
40–50%
Prove each alert, each escalation and each cleared name came out of the declared transactions and the scenario version in force on the day.
Legacy · deterministic
the computation is re-run
- Threshold and scenario arithmetic: recomputed from the declared transaction set
- Sanctions and PEP list matching: re-run against the list version as it stood
- Aggregation and velocity rules: the trigger value re-derives exactly
- Case-closure arithmetic: the disposition reproduces from the declared record
AI systems
the agent shows its work
- AML transaction-monitoring alerts: each alert cites the transactions behind it
- Trade surveillance escalations: every referral shows the evidence it stood on
- Adverse-media screening: the sources read and the ones excluded
- False-positive suppression: the criteria applied, re-checkable on demand
The number
Voice surveillance cost down 65%
operator disclosure, >90% accuracy
The trigger
DORA Art 10 · in force 17 Jan 2025
EU financial entities
It settles that the alert is the deterministic result of the declared inputs under the declared scenario. It does not say the scenario catches what it should.
-
03
Commission, fee and premium calculation
40%
Prove every commission, every fee line and every premium is the arithmetic the contract or the filed schedule actually specifies. Both sides holding the same copy.
Legacy · deterministic
the computation is re-run
- Commission runs from a contracted rate card: recomputed line by line
- Fee schedules and tiered pricing: the invoice re-derives or REJECTs
- Interest accrual and day-count conventions: re-run from the declared basis
- Clawback and true-up arithmetic: reproduced from the committed period data
AI systems
the agent shows its work
- Anomaly flags on a payment run: the lines that triggered them, named
- Contract-term extraction: the clause each figure was read from
- Dispute triage: a receipt of the terms compared and the version used
The number
No operator figure published
set by analogy, marked as such
The trigger
MiFID II Art 16 · RTS 24 · in force
records reconstructable on demand
Pure arithmetic under a written rule, which is why it re-derives cleanly. The share it gives up is the weakest number on this page and the first thing a pilot should measure.
-
04
Reconciliation and settlement, matched on both sides
35%
Prove the settled figure is what both ledgers say it is, with each side re-running the other's arithmetic against its own copy of the rule.
Legacy · deterministic
the computation is re-run
- Trade settlement and reconciliation: the settled figure re-derives or REJECTs
- Nostro and intercompany matching: recomputed from the declared postings
- Corporate-action entitlement: re-run from the declared terms and holdings
- Cash-break arithmetic: the residual reproduces bit-for-bit
AI systems
the agent shows its work
- Break classification: each category bound to the evidence behind it
- Auto-matching proposals: the fields compared and the tolerance applied
- Exception routing: a receipt of why this break went to this desk
The number
Straight-through processing 77.2% to 81.8%
operator disclosure, one year
The trigger
T+1 settlement discipline
shorter windows, same evidence burden
Two-party reconciliation is the highest-value case here and the longest build, because each side has to hold its own copy of the rule and the copies have to be shown to match.
-
05
Regulatory and statutory reporting
30%
Prove every submitted figure reproduces from the inputs and the template version in force, so the return can be defended long after the people who filed it moved on.
Legacy · deterministic
the computation is re-run
- Regulatory transaction reporting: each submitted record reproduced bit-for-bit
- Capital and reserving arithmetic: deterministic runs re-derive exactly
- Statutory return assembly: every cell traced to its committed source
- Restatement arithmetic: the corrected figure re-derives from the corrected input
AI systems
the agent shows its work
- Narrative drafting from the numbers: each sentence bound to the cell it cites
- Threshold and materiality checks: the criteria applied, shown
- Cross-return consistency: the comparisons run and the ones skipped
The number
Assurance cost falls about a fifth
on better-structured evidence
The trigger
CSRD / ESRS assurance · phasing in
limited assurance already applies
It settles that the figure is the deterministic result of the committed computation. Whether the return was scoped correctly is still the firm's judgement.
-
06
Control testing, model validation and the close
30–35%
Prove the control result, the back-test and the consolidated figure all re-derive from the evidence and the model as committed, rather than from a reconstruction after the fact.
Legacy · deterministic
the computation is re-run
- Control re-performance against the written test script: recomputed, not re-asserted
- Model back-testing: re-run against the committed snapshot, seed pinned
- Consolidation and elimination arithmetic: reproduced from declared postings
- Journal support: every entry traced to the file that justified it
AI systems
the agent shows its work
- Evidence assembly for an auditor request: what was pulled and what was left out
- Deficiency roll-up: the tests each conclusion rested on
- Model-card claims: assertions traceable to the declared evidence
The number
$2.3M a year · 15,580 staff-hours
KPMG, 2025
The trigger
SR 26-2 superseded SR 11-7 · 17 Apr 2026
Fed / OCC / FDIC
A faithfully re-derived wrong-but-committed rule still re-derives. This says the arithmetic held, never that the control was the right control.
Every figure above sizes the
problem. They come from public regulation, operator disclosures and the two firms' own
annual reports. What gets added is the re-derivation, run by independent code on somebody
else's machine.
The same arithmetic, run globally
The two tabs take this arithmetic to one disclosed payroll each. Run it globally and it
lands here. Customer value, not our revenue.
Only the first two lines are published. Everything under them is a modelling
assumption, which is exactly why the two firm tabs run the same chain against real disclosed
headcount and real disclosed payroll instead of a global average.
Why the bank's number is ten times the insurer's
Not because banking suits this better. On both rates that actually matter, the insurer
is the better fit. The bank is simply bigger and pays more per head. Three multipliers,
and only the third is about the product.
Start at the insurer's $258m. More people in scope takes it to about $2.06bn,
a higher cost per head to about $2.95bn, and a slightly lower share brings it back to
$2.67bn. Rounding aside, that is the whole of the gap.
Why more of the insurer's people are in scope
Sixty-one per cent of the insurer's staff sit in these kinds of work against forty-one per
cent at the bank. An insurer is mostly claims, underwriting and policy administration, which
is rule-bound by construction: a filed rate table and a policy wording decide the answer. A
bank carries large populations in relationship management, markets, technology and branch
operations that never enter this model.
Why the bank pays more for each of them
$123,831 against $86,256, each taken from that firm's own disclosed payroll over its own
disclosed headcount. Business mix is most of it. European supervisors put retail banking at
$59,030 a head and investment banking at $176,578 in the same year, and a bank's control and
markets functions sit at the top of that spread. Geography does the rest, and that part is a
reading rather than a disclosure.
So compare the percentage, never the total. The absolute figure
scales with payroll, which says more about the firm's size than about the work. The share is
the claim: 29% at the bank, 32% at the insurer, and the smaller firm is the cleaner
fit.
What removing the wait is worth, counted separately
Everything above prices the person doing the work. It does not price the wait. A day off a
cycle that is gated by a check releases one 365th of the value waiting in it, which at a 4%
cost of funding is $110,000 a year for every $1bn in flight, per day removed. That
unit works against any flow a firm can size.
The larger effect is the ramp. Regulated work takes five years to change because each step
waits to be accepted, and acceptance today is a sampled manual review on a quarterly
cadence. A check that is free and takes a third of a second turns the sample into the
population. Pulling the curve forward one year moves 78 points of the run rate inside
the window: $2.08bn at the bank and $201m at the insurer. Same money, sooner,
so it is never added to the run rate. Each firm tab shows the arithmetic.
Where to find the full working
The tabs at the top carry this arithmetic against one large bank and one large insurer: the
ten kinds of work each one runs, the wage bill, the share that comes off, the build price,
and every assumption written out in full at the bottom.
The general method sits on its own page. The eight-step chain this study applies,
with the qualifying gates, the sourcing order, the stamp lattice and a fill-in template for
taking the same arithmetic to any other sector:
the agentic value
methodology.
The five numbers for this bank
| $2.67bn | A year off the cost base once every kind of work below is running. Middle of three cases2. Converted from euros1. |
| 29% | Of the wage bill in that work. Not of revenue, not of profit, not of headcount. |
| $5.9m | Paid once, for the whole build. Under one day of the saving. The verifier itself is free forever8. |
| Year five | When all of it is running. 22% lands in year one, 70% by year three3. |
| 74,418 | People doing that work today, out of 181,510. 26,694 roles come off across the five years7. |
The ten kinds of work, and the saving from each
Each of these is somebody applying a written rule to figures already committed, then
handing the answer to a person entitled to challenge it. The share is what an agent takes
over once it is fully running2. The money is what is left
after the agent's own running cost5.
Full bar is that work's annual wage bill. The green part is what comes off. All ten together: 74,418 people, $9.21bn of wage bill, $2.67bn a year off the cost base. Judgement is excluded from every kind of work before the share is applied6.
Why this is possible now and was not before
What changed in the technology
An agent is software that reads the inputs, applies the written rule and produces the
answer on its own, then hands it to a person. It is not a chatbot waiting to be asked a
question. Nothing could do this on a real capital return or a real credit file until roughly two years ago. That part is
no longer in doubt, and it is not what anyone is buying.
What changed in the law
Automated decisions in banking now carry written obligations, listed in the next section
with the day each one bites. Every one of them comes down to the same demand: when somebody
challenges a figure a machine produced, the firm has to be able to show how it got there.
A log says a decision happened. It does not reproduce the decision.
What was still missing
A signature proves who sent a file. A hash proves the file has not changed since. Neither
one recomputes the number. So the honest position for a banking board until now was that
an agent could do the work and nobody could answer for it afterwards, which is why almost
none of this work has moved.
The verifier closes exactly that gap. The agent ships its answer together with
everything needed to recompute it: the inputs as they stood, the rule as it was pinned, and
a hash of each. The auditor, the regulator or the counterparty runs a free open-source checker on their own laptop, offline,
with no account and no call to any supplier. The checker recomputes the answer using its
own code and returns one of three verdicts: it matches, it does not, or it could not tell.
An "it could not tell" never becomes a pass.
Signing standards, notarisation and supply-chain attestation all answer one of two
questions: did the bytes change, or who signed this. This answers a third one, which is
whether the answer still comes out. Two research sweeps across more than thirty products
found nobody else independently recomputing a declared output9.
What it does not do, said once. It settles integrity, never
correctness, and it never re-runs the agent's own words. No language model reproduces its
own output, so what the agent carries is a receipt of what it read and what it excluded.
The arithmetic underneath is what re-derives
10.
The regulations, and the date each one applies
None of this is optional and none of it is far away. Each line below already has a date.
Insurance-side rules are listed on the other tab. Dates from our own regulation tracker, sources checked 4 July
2026.
Nothing here makes anyone compliant, and that word is never used.
A re-derivable record is evidence a compliance function can hand over and stand behind. The
determination stays with the regulator and the firm, exactly where it was.
What it costs to build and run
The whole build is under one day of the saving it makes. $5.9m
against $2.67bn a year. Price was never what held this back. Agents able to do the work
arrived about two years ago, and a way for an outside party to check what they produced
arrived with the verifier.
What the saved time is worth, counted separately
Everything above prices the person. It does not price the wait. They are
different pots and they are never added together here. Money moves twice when a cycle
shortens: capital stops sitting still, and the whole programme lands sooner.
The unit: per day, per billion in flight
A day taken off a cycle that is gated by a check releases one 365th of the value waiting
in it. At a 4% cost of funding, that is $110,000 a year for every $1bn in flight, per
day removed12. Ten days off a $20bn flow is $22m a
year. Hold that against any process the firm can put a number on.
The effect on the five-year ramp
Regulated work takes five years to change, and the reason is not the software. Each step
waits to be accepted, and acceptance today means a sampled manual review on a quarterly or
annual cadence. When the check is free and runs in a third of a second, the sample becomes
the population and acceptance stops being a scheduled event. Pull the curve forward by one
year and the same money arrives earlier.
Seventy-eight points of the run rate, arriving inside the five years instead
of after them. At this bank that is $2.08bn of saving arriving inside the window rather than beyond it.
This is not new money and it is never added to the run
rate. It is the same saving landing sooner, which is worth something to a finance
director and nothing to anyone who adds the two together. What makes the mechanism
credible is measured: the check runs in 0.335 seconds
9.
What sits on the clock today is the scheduled review. One control-testing programme runs once a year on 15,580 staff-hours, and when the annual check misses something the correction loop averaged 438 days.
The full working, from headcount to yearly saving
181,510 people · $22.0bn payroll · $123,831 per head
FY2025 Universal Registration Document and results, converted from
euros
1. 74,418 of those people sit in the ten kinds of work
above.
The saving, year by year
Bars are the middle case. The whisker is low to high. The first quarter and
the first year both cost money before they pay3.
Cost per person, by pay tier
Three pay tiers rather than one average, scaled so the whole population still reproduces
the bank's own disclosed payroll per head4.
Against the number the bank has already given the market. Its
cost-to-income target implies about $2.93bn of annual cost out by 2028, across every lever
it has. This reaches $2.67bn from agents alone, and not until year five. Close enough to be
worth noticing, far enough apart in time to be a different claim.
Every assumption, and the case against each
Four of the ten shares rest on analogy rather than a published result. Six of the ten pay
mixes are our own reading. The build price has no external comparison. The phasing curve
comes from cost programmes, not AI ones. Population data on 25,000 workers finds no effect
at all. Every one of those is written out below, because a reader who finds one unaided
stops believing the rest.
- Currency. This firm discloses in euros. Every figure here is converted at
a flat $1.10 to the euro, a rate that is stated rather than taken off a fixing on any
particular day. Every dollar on this tab moves one-for-one with it, so a reader who
disagrees with the rate can rescale the whole page in one step.
- The share taken, and how it is set. Per kind of work, three cases. Low
counts only work under a rule pinned by digest and held in the same version by both sides.
Middle adds work under an internal written rule the firm controls. High adds two-party
reconciliations where each side holds its own copy. Every headline on this page is the
middle case. Eight of the shares sit on a result an operator has published: Ping An,
Allianz, Lemonade, Manulife, JPMorgan, HSBC, Deutsche Bank, ING and Aviva. Four rest on
analogy to the eight and are marked with this note in the tables: commission and fee
calculation, control testing, model validation, finance close. Those four are the first
things a pilot should measure. No catch rate or false-positive rate is stated anywhere,
because neither has been measured.
- Phasing. The chain produces a run rate, and a run rate is not a first-year
figure. 22% of it lands by the end of year one, 43% by year two, 70% by year three, 91% by
year four and all of it in year five. That curve is the mean of two published ramps: BNP
Paribas realised $550m, $660m, $1.10bn, $880m and $660m across 2022 to 2026 against a
$3.85bn programme, and UBS booked $4.0bn, then $3.4bn, then $3.2bn against its $13.5bn
Credit Suisse integration ambition and called that trajectory non-linear itself. Neither
programme was an AI deployment, and no AI programme has published a year-by-year ramp, so
the curve is an analogue. No programme publishes a first-quarter realisation either, so
quarter one is taken as an eighth of the year-one share, back-loaded inside the year. It is
the weakest cell on the curve. Two traps worth naming: HSBC actioned $1.2bn of annualised
simplification savings by end-2025 and took only $0.6bn into the P&L, so actioned is not
realised; and McKinsey's 28/57/66/74% first-year transformation curve is top-quartile
performers only, so it is a ceiling and never a median.
- Pay tiers. One blended cost per head is the weakest number a model like
this can carry, so the population is split into clerical, professional and senior in the
published ratio 1 : 1.73 : 3.58, from US Bureau of Labor Statistics occupational wage
estimates for insurance carriers. The shape is imported; the level is not. All three tiers
are scaled by a single factor so the weighted average reproduces the firm's own disclosed
payroll divided by its own disclosed headcount, exactly. Nothing enters that the firm has
not already published in aggregate. Cross-checked against the European Banking Authority's
remuneration benchmarking of 2.19 million banking staff, which finds a comparable
three-times spread on a different cut: retail banking $59,030, corporate functions $76,472,
independent control functions $95,524, investment banking $176,578. What is still assumed
is the tier mix inside each kind of work. Where an occupation maps one-to-one it is read
off published employment shares; elsewhere, in six of ten cases, it is our
classification.
- Running cost. Deducted from every figure on this page at 20% of the human
cost replaced in the low case, 10% in the middle and 5% in the high. That covers the
agent's own compute plus the people still supervising it. Anchored on a measured
manual-to-electronic step in which one administrative transaction fell from $1.02 to $0.10,
and a status inquiry from $4.38 to $0.04, across plans covering 63% of one market's insured
lives.
- Judgement is never in a share. Materiality, scope, hardship, discretion, a
call between two defensible readings, and the decision to take risk are all excluded from
every kind of work before the share is applied. They stay with people.
- Severance. 26,694 roles come off across the five years, 2.9% of
headcount a year. Whether that costs anything
depends on the firm's own turnover, so it is treated as a separate class and never enters a
cash figure.
- The build price is a quote, not a benchmark. $5.9m, staged. No comparable engagement is published
anywhere, so nothing external checks this number. It can only be carried because trebling
it barely moves the answer. Cost tracks the number of system estates the inputs live in
rather than the size of the saving, and this bank runs 64 countries of estate. The verifier
itself is Apache-2.0 and carries no licence fee, ever.
- The check, measured. A clean re-derivation runs at 0.335 seconds median
across 40 pilot bundles, range 0.214s to 0.413s, wall clock including interpreter startup,
on a laptop. Bundles are a few megabytes at most and no scaling curve has been fitted.
Tamper runs against the same 40: one byte changed with the manifest untouched was caught 40
times out of 40; one byte changed with the manifest digest recomputed to match was caught
37 out of 40, and the three that pass are declaring that file as carried for storage rather
than as an input, which is now visible instead of implied. What re-derives is the
arithmetic, never the whole cycle, so the exception queue and the human judgement stay on
the clock. On the wider claim in the body: two independent research sweeps across more than
thirty products found no vendor doing independent deterministic recomputation of a declared
output. Their verifiers recompute a digest of what was declared. This one recomputes the
declaration itself.
- The rule is yours. This settles that a figure is the
deterministic result of the committed computation over the committed inputs. It says
nothing about whether the rule behind it was the right rule: a faithfully re-derived
wrong-but-committed rule still re-derives, and that is correct behaviour as well as the
limit. The agent's own output is never re-run, because language models are not
bit-reproducible even at temperature zero. What the agent carries is a receipt of the
sources it used and the criteria it included and excluded. Three commercial legal research
tools sold as avoiding hallucination were wrong between 17% and 33% of the time in a
pre-registered study, which is why a supplier's assurance about its own output is not
treated as evidence about that output.
- Sources. Firm data from BNP Paribas's Universal Registration Document 2025
and FY2025 results. Pay tiers from US Bureau of Labor Statistics occupational wage
estimates, cross-checked against European Banking Authority remuneration benchmarking. The
ramp from BNP Paribas's and UBS's own disclosed year-by-year cost programmes.
Counter-evidence from Stanford HAI, Gartner, METR and national administrative-data
research. Full workings:
the agentic value method. These
are modelled bands on stated assumptions, not measured savings. The logo is the property of its owner and appears
here for identification only.
- The clock arithmetic. Value released a year equals the value in
flight multiplied by days removed, divided by 365, multiplied by the cost of funding. The
cost of funding is taken at 4.0%, which is stated rather than sourced, and every clock
figure moves one-for-one with it. The value in flight is the one input a firm has to supply
from its own balance sheet, so no figure on this page multiplies it out. The ramp
pull-forward is arithmetic on this page's own curve and nothing else: the modelled shares
less the same shares shifted one year left, summed, which comes to 78 points of the run
rate. It assumes acceptance is what sets the pace of the ramp. Where a workstream is held
up by a system migration or a works council rather than by a review, this line is worth
nothing and should be struck.
The five numbers for this insurer
| $258m | A year off the cost base once every kind of work below is running. Middle of three cases2. |
| 32% | Of the wage bill in that work. Not of revenue, not of profit, not of headcount. |
| $2.75m | Paid once, for the whole build. Under four days of the saving. The verifier itself is free forever8. |
| Year five | When all of it is running. 22% lands in year one, 70% by year three3. |
| 9,341 | People doing that work today, out of 15,338. 3,648 roles come off across the five years7. |
The ten kinds of work, and the saving from each
Each of these is somebody applying a written rule to figures already committed, then
handing the answer to a person entitled to challenge it. The share is what an agent takes
over once it is fully running2. The money is what is left
after the agent's own running cost5.
Full bar is that work's annual wage bill. The green part is what comes off. All ten together: 9,341 people, $806m of wage bill, $258m a year off the cost base. Judgement is excluded from every kind of work before the share is applied6.
Why this is possible now and was not before
What changed in the technology
An agent is software that reads the inputs, applies the written rule and produces the
answer on its own, then hands it to a person. It is not a chatbot waiting to be asked a
question. Nothing could do this on a real claims file or a real policy schedule until roughly two years ago. That part is
no longer in doubt, and it is not what anyone is buying.
What changed in the law
Automated decisions in insurance now carry written obligations, listed in the next section
with the day each one bites. Every one of them comes down to the same demand: when somebody
challenges a figure a machine produced, the firm has to be able to show how it got there.
A log says a decision happened. It does not reproduce the decision.
What was still missing
A signature proves who sent a file. A hash proves the file has not changed since. Neither
one recomputes the number. So the honest position for a insurance board until now was that
an agent could do the work and nobody could answer for it afterwards, which is why almost
none of this work has moved.
The verifier closes exactly that gap. The agent ships its answer together with
everything needed to recompute it: the inputs as they stood, the rule as it was pinned, and
a hash of each. The regulator, the auditor or the policyholder's lawyer runs a free open-source checker on their own laptop, offline,
with no account and no call to any supplier. The checker recomputes the answer using its
own code and returns one of three verdicts: it matches, it does not, or it could not tell.
An "it could not tell" never becomes a pass.
Signing standards, notarisation and supply-chain attestation all answer one of two
questions: did the bytes change, or who signed this. This answers a third one, which is
whether the answer still comes out. Two research sweeps across more than thirty products
found nobody else independently recomputing a declared output9.
What it does not do, said once. It settles integrity, never
correctness, and it never re-runs the agent's own words. No language model reproduces its
own output, so what the agent carries is a receipt of what it read and what it excluded.
The arithmetic underneath is what re-derives
10.
The regulations, and the date each one applies
None of this is optional and none of it is far away. Each line below already has a date.
Banking-side rules are listed on the other tab. Dates from our own regulation tracker, sources checked 4 July
2026.
Nothing here makes anyone compliant, and that word is never used.
A re-derivable record is evidence a compliance function can hand over and stand behind. The
determination stays with the regulator and the firm, exactly where it was.
What it costs to build and run
The whole build is under four days of the saving it makes. $2.75m
against $258m a year. Price was never what held this back. Agents able to do the work
arrived about two years ago, and a way for an outside party to check what they produced
arrived with the verifier.
What the saved time is worth, counted separately
Everything above prices the person. It does not price the wait. They are
different pots and they are never added together here. Money moves twice when a cycle
shortens: capital stops sitting still, and the whole programme lands sooner.
The unit: per day, per billion in flight
A day taken off a cycle that is gated by a check releases one 365th of the value waiting
in it. At a 4% cost of funding, that is $110,000 a year for every $1bn in flight, per
day removed12. Ten days off a $20bn flow is $22m a
year. Hold that against any process the firm can put a number on.
The effect on the five-year ramp
Regulated work takes five years to change, and the reason is not the software. Each step
waits to be accepted, and acceptance today means a sampled manual review on a quarterly or
annual cadence. When the check is free and runs in a third of a second, the sample becomes
the population and acceptance stops being a scheduled event. Pull the curve forward by one
year and the same money arrives earlier.
Seventy-eight points of the run rate, arriving inside the five years instead
of after them. At this insurer that is $201m of saving arriving inside the window rather than beyond it.
This is not new money and it is never added to the run
rate. It is the same saving landing sooner, which is worth something to a finance
director and nothing to anyone who adds the two together. What makes the mechanism
credible is measured: the check runs in 0.335 seconds
9.
What sits on the clock today is the scheduled review. Claims, reserving and reporting all run to a scheduled review rather than a continuous one, because a manual re-check has always had to be rationed.
The full working, from headcount to yearly saving
15,338 people · $1.32bn payroll · $86,256 per head
FY2025 Annual Report and Form 20-F. 16.6m policies and about 3.0m claims a year. 9,341 of
those people sit in the ten kinds of work above.
The saving, year by year
Bars are the middle case. The whisker is low to high. The first quarter and
the first year both cost money before they pay3.
Cost per person, by pay tier
Three pay tiers rather than one average, scaled so the whole population still reproduces
the insurer's own disclosed payroll per head4.
This insurer is already ahead on one line. It runs 70%
auto-underwriting against a market average of 11%. The underwriting row above is sized on
the remaining 30%, not on the whole book.
Leakage is not in any figure above. The firm recovered over $100m of
fraud, waste and abuse in 2025 from focused effort. Running the same check across the whole
book is worth $30m low, $75m middle and $150m high. That is a different class of money and
it is never added to a payroll figure.
Every assumption, and the case against each
Four of the ten shares rest on analogy rather than a published result. Six of the ten pay
mixes are our own reading. The build price has no external comparison. The phasing curve
comes from cost programmes, not AI ones. Population data on 25,000 workers finds no effect
at all. Every one of those is written out below, because a reader who finds one unaided
stops believing the rest.
- Currency. Every figure on this tab is disclosed in dollars by the firm
itself and is carried untouched. Nothing here is converted, so no exchange-rate assumption
sits underneath any of it.
- The share taken, and how it is set. Per kind of work, three cases. Low
counts only work under a rule pinned by digest and held in the same version by both sides.
Middle adds work under an internal written rule the firm controls. High adds two-party
reconciliations where each side holds its own copy. Every headline on this page is the
middle case. Eight of the shares sit on a result an operator has published: Ping An,
Allianz, Lemonade, Manulife, JPMorgan, HSBC, Deutsche Bank, ING and Aviva. Four rest on
analogy to the eight and are marked with this note in the tables: commission and fee
calculation, control testing, model validation, finance close. Those four are the first
things a pilot should measure. No catch rate or false-positive rate is stated anywhere,
because neither has been measured.
- Phasing. The chain produces a run rate, and a run rate is not a first-year
figure. 22% of it lands by the end of year one, 43% by year two, 70% by year three, 91% by
year four and all of it in year five. That curve is the mean of two published bank ramps. The
first realised $550m, $660m, $1.10bn, $880m and $660m across 2022 to 2026 against a $3.85bn
programme, and is named in the banking tab. UBS booked $4.0bn, then $3.4bn, then $3.2bn against its $13.5bn
Credit Suisse integration ambition and called that trajectory non-linear itself. Neither
programme was an AI deployment, and no AI programme has published a year-by-year ramp, so
the curve is an analogue. No programme publishes a first-quarter realisation either, so
quarter one is taken as an eighth of the year-one share, back-loaded inside the year. It is
the weakest cell on the curve. Two traps worth naming: HSBC actioned $1.2bn of annualised
simplification savings by end-2025 and took only $0.6bn into the P&L, so actioned is not
realised; and McKinsey's 28/57/66/74% first-year transformation curve is top-quartile
performers only, so it is a ceiling and never a median.
- Pay tiers. One blended cost per head is the weakest number a model like
this can carry, so the population is split into clerical, professional and senior in the
published ratio 1 : 1.73 : 3.58, from US Bureau of Labor Statistics occupational wage
estimates for insurance carriers. The shape is imported; the level is not. All three tiers
are scaled by a single factor so the weighted average reproduces the firm's own disclosed
payroll divided by its own disclosed headcount, exactly. Nothing enters that the firm has
not already published in aggregate. Cross-checked against the European Banking Authority's
remuneration benchmarking of 2.19 million banking staff, which finds a comparable
three-times spread on a different cut: retail banking $59,030, corporate functions $76,472,
independent control functions $95,524, investment banking $176,578. What is still assumed
is the tier mix inside each kind of work. Where an occupation maps one-to-one it is read
off published employment shares; elsewhere, in six of ten cases, it is our
classification.
- Running cost. Deducted from every figure on this page at 20% of the human
cost replaced in the low case, 10% in the middle and 5% in the high. That covers the
agent's own compute plus the people still supervising it. Anchored on a measured
manual-to-electronic step in which one administrative transaction fell from $1.02 to $0.10,
and a status inquiry from $4.38 to $0.04, across plans covering 63% of one market's insured
lives.
- Judgement is never in a share. Materiality, scope, hardship, discretion, a
call between two defensible readings, and the decision to take risk are all excluded from
every kind of work before the share is applied. They stay with people.
- Severance. 3,648 roles come off across the five years, 4.8% of
headcount a year. Whether that costs anything
depends on the firm's own turnover, so it is treated as a separate class and never enters a
cash figure.
- The build price is a quote, not a benchmark. $2.75m, staged. No comparable engagement is published
anywhere, so nothing external checks this number. It can only be carried because trebling
it barely moves the answer. Cost tracks the number of system estates the inputs live in
rather than the size of the saving, and this insurer runs 20 markets. The verifier
itself is Apache-2.0 and carries no licence fee, ever.
- The check, measured. A clean re-derivation runs at 0.335 seconds median
across 40 pilot bundles, range 0.214s to 0.413s, wall clock including interpreter startup,
on a laptop. Bundles are a few megabytes at most and no scaling curve has been fitted.
Tamper runs against the same 40: one byte changed with the manifest untouched was caught 40
times out of 40; one byte changed with the manifest digest recomputed to match was caught
37 out of 40, and the three that pass are declaring that file as carried for storage rather
than as an input, which is now visible instead of implied. What re-derives is the
arithmetic, never the whole cycle, so the exception queue and the human judgement stay on
the clock. On the wider claim in the body: two independent research sweeps across more than
thirty products found no vendor doing independent deterministic recomputation of a declared
output. Their verifiers recompute a digest of what was declared. This one recomputes the
declaration itself.
- The rule is yours. This settles that a figure is the
deterministic result of the committed computation over the committed inputs. It says
nothing about whether the rule behind it was the right rule: a faithfully re-derived
wrong-but-committed rule still re-derives, and that is correct behaviour as well as the
limit. The agent's own output is never re-run, because language models are not
bit-reproducible even at temperature zero. What the agent carries is a receipt of the
sources it used and the criteria it included and excluded. Three commercial legal research
tools sold as avoiding hallucination were wrong between 17% and 33% of the time in a
pre-registered study, which is why a supplier's assurance about its own output is not
treated as evidence about that output.
- Sources. Firm data from Prudential plc's AR2025, Form 20-F and
Sustainability Report 2025. Pay tiers from US Bureau of Labor Statistics occupational wage
estimates for insurance carriers, cross-checked against European Banking Authority
remuneration benchmarking. The ramp from two banks' own disclosed year-by-year cost
programmes.
Counter-evidence from Stanford HAI, Gartner, METR and national administrative-data
research. Full workings:
the agentic value method. These
are modelled bands on stated assumptions, not measured savings. The logo is the property of its owner and appears
here for identification only.
- The clock arithmetic. Value released a year equals the value in
flight multiplied by days removed, divided by 365, multiplied by the cost of funding. The
cost of funding is taken at 4.0%, which is stated rather than sourced, and every clock
figure moves one-for-one with it. The value in flight is the one input a firm has to supply
from its own balance sheet, so no figure on this page multiplies it out. The ramp
pull-forward is arithmetic on this page's own curve and nothing else: the modelled shares
less the same shares shifted one year left, summed, which comes to 78 points of the run
rate. It assumes acceptance is what sets the pace of the ramp. Where a workstream is held
up by a system migration or a works council rather than by a review, this line is worth
nothing and should be struck.