NEXIVERIFY

What an agent takes off a regulated cost base.

Volume is rising and trust in agent output is falling. Re-derivation answers both, and here is what that is worth against two real balance sheets.

More work to check, and less trust in the answers

Volume

The work that needs checking is about to skyrocket.

So checking has to get faster than a person reviewing a sample.

Trust

Confidence in what an agent hands back is going the other way.

So work nobody trusts has to be re-doable on demand.

Those two pull against each other. Re-derivation resolves both.

The answer gets recomputed from the declared inputs under a pinned rule, by code the doubter runs on their own machine, in a third of a second. What follows is what that is worth against a regulated wage bill.

The five numbers, banking and insurance

$2.67bnA year off one large bank's cost base, $258m at one large insurer. Each figure built from that firm's own disclosed headcount and wage bill, worked line by line on its own tab.
29–32%Of the wage bill sitting in work where the rule is written down and somebody downstream can dispute the answer. Not a share of revenue, of profit or of headcount.
$0For the verifier, for good. Apache-2.0, offline, on the auditor's own machine. The only invoice is the code that packages what the firm's systems already hold.
0.335sTo re-derive one answer, median across 40 pilot bundles, with no network on the verdict path. Checking stops being the thing that gets rationed.
2 Aug 2026The first dated obligation is already in force, and the high-risk record-keeping article lands 2 December 2027. Every rule asks the same thing: show how the machine got there.

What goes wrong when an agent pays an invoice

Routine, valid, matched to the purchase order and the goods receipt. Nothing about it needs a person. Then one bank-account field comes back changed, and real money goes to the wrong place.

So the agent proposes. The verifier decides.

Any model, any code, any workflow may produce the figure. A pinned rule recomputes it from the declared inputs, and that rule is the metric for correctness, never the agent's method.

What the verifier is: open source, agnostic, re-derivable

The propertyWhat it buys
Open source Apache-2.0. It runs on the doubter's own laptop, offline, with no account and no call to any supplier. Nobody has to trust the vendor, including the vendor that wrote it. That is what lets an auditor accept the verdict rather than negotiate about it.
Agnostic It does not care which system produced the figure, which model wrote it or whose stack it sits on. It reads a declared package and a pinned rule. The same checker settles a premium, a capital return and a commission run, which is why one build covers ten kinds of work rather than one.
Re-derivable Signing says who sent a file. Hashing says the file has not changed. This recomputes the answer from the declared inputs and says whether it still comes out. That third question is the one nobody else was answering, and the only one that settles an argument about a number.
This changes what an agent can be trusted to do. From answers that need faith, to work that can be accepted: bounded, monitored, receipted, independently re-derivable. That is the line between assistance and delegation.

The share of the wage bill that comes off

29–32% comes off Settled by the written rule The share an agent takes over once every kind of work is live. $2.67bn a year at one large bank, $258m at one large insurer. Stays with people Judgement, discretion, materiality, and the decision to take risk. Excluded from every figure before the share is applied.

Six kinds of work, and what re-derives in each

The split is load-bearing. Deterministic work is re-run: recomputed from the declared inputs and rules, and the committed output reproduced bit-for-bit. An agent output is never re-run, because no model reproduces its own words. It carries a receipt instead: the sources it used, the criteria it included and excluded, re-checkable against the declared evidence.

  • 01 Claims and underwriting settled on the filed rule 45–50%

    Prove every adjudicated claim and every rated premium traces to the filed terms and the declared data, before the supervisor, the auditor or the policyholder asks.

    Legacy · deterministic the computation is re-run

    • Premium rating from filed rate tables: the quote reproduces bit-for-bit
    • Claims adjudication against policy terms: recomputed from the declared policy
    • Deductible, co-insurance and sub-limit application: re-derived, not re-asserted
    • Reinsurance treaty settlement: the ceded figure re-derives or REJECTs

    AI systems the agent shows its work

    • ML pricing models: each price shows the factors it stood on
    • Claims triage and fraud scoring: every referral cites its exact signals
    • Underwriting risk scores: a receipt of inclusion and exclusion criteria
    • Straight-through decisions: the evidence each one rested on, re-checkable
    The number 63% of injury claims settled automatically operator disclosure
    The trigger AI Act Art 12 · 2 Dec 2027 life & health pricing is high-risk

    Re-derivation settles that the number matches its declared inputs and filed rules. Whether the rate table or the triage policy is fair stays an underwriting judgement.

  • 02 Surveillance, screening and financial-crime alerts 40–50%

    Prove each alert, each escalation and each cleared name came out of the declared transactions and the scenario version in force on the day.

    Legacy · deterministic the computation is re-run

    • Threshold and scenario arithmetic: recomputed from the declared transaction set
    • Sanctions and PEP list matching: re-run against the list version as it stood
    • Aggregation and velocity rules: the trigger value re-derives exactly
    • Case-closure arithmetic: the disposition reproduces from the declared record

    AI systems the agent shows its work

    • AML transaction-monitoring alerts: each alert cites the transactions behind it
    • Trade surveillance escalations: every referral shows the evidence it stood on
    • Adverse-media screening: the sources read and the ones excluded
    • False-positive suppression: the criteria applied, re-checkable on demand
    The number Voice surveillance cost down 65% operator disclosure, >90% accuracy
    The trigger DORA Art 10 · in force 17 Jan 2025 EU financial entities

    It settles that the alert is the deterministic result of the declared inputs under the declared scenario. It does not say the scenario catches what it should.

  • 03 Commission, fee and premium calculation 40%

    Prove every commission, every fee line and every premium is the arithmetic the contract or the filed schedule actually specifies. Both sides holding the same copy.

    Legacy · deterministic the computation is re-run

    • Commission runs from a contracted rate card: recomputed line by line
    • Fee schedules and tiered pricing: the invoice re-derives or REJECTs
    • Interest accrual and day-count conventions: re-run from the declared basis
    • Clawback and true-up arithmetic: reproduced from the committed period data

    AI systems the agent shows its work

    • Anomaly flags on a payment run: the lines that triggered them, named
    • Contract-term extraction: the clause each figure was read from
    • Dispute triage: a receipt of the terms compared and the version used
    The number No operator figure published set by analogy, marked as such
    The trigger MiFID II Art 16 · RTS 24 · in force records reconstructable on demand

    Pure arithmetic under a written rule, which is why it re-derives cleanly. The share it gives up is the weakest number on this page and the first thing a pilot should measure.

  • 04 Reconciliation and settlement, matched on both sides 35%

    Prove the settled figure is what both ledgers say it is, with each side re-running the other's arithmetic against its own copy of the rule.

    Legacy · deterministic the computation is re-run

    • Trade settlement and reconciliation: the settled figure re-derives or REJECTs
    • Nostro and intercompany matching: recomputed from the declared postings
    • Corporate-action entitlement: re-run from the declared terms and holdings
    • Cash-break arithmetic: the residual reproduces bit-for-bit

    AI systems the agent shows its work

    • Break classification: each category bound to the evidence behind it
    • Auto-matching proposals: the fields compared and the tolerance applied
    • Exception routing: a receipt of why this break went to this desk
    The number Straight-through processing 77.2% to 81.8% operator disclosure, one year
    The trigger T+1 settlement discipline shorter windows, same evidence burden

    Two-party reconciliation is the highest-value case here and the longest build, because each side has to hold its own copy of the rule and the copies have to be shown to match.

  • 05 Regulatory and statutory reporting 30%

    Prove every submitted figure reproduces from the inputs and the template version in force, so the return can be defended long after the people who filed it moved on.

    Legacy · deterministic the computation is re-run

    • Regulatory transaction reporting: each submitted record reproduced bit-for-bit
    • Capital and reserving arithmetic: deterministic runs re-derive exactly
    • Statutory return assembly: every cell traced to its committed source
    • Restatement arithmetic: the corrected figure re-derives from the corrected input

    AI systems the agent shows its work

    • Narrative drafting from the numbers: each sentence bound to the cell it cites
    • Threshold and materiality checks: the criteria applied, shown
    • Cross-return consistency: the comparisons run and the ones skipped
    The number Assurance cost falls about a fifth on better-structured evidence
    The trigger CSRD / ESRS assurance · phasing in limited assurance already applies

    It settles that the figure is the deterministic result of the committed computation. Whether the return was scoped correctly is still the firm's judgement.

  • 06 Control testing, model validation and the close 30–35%

    Prove the control result, the back-test and the consolidated figure all re-derive from the evidence and the model as committed, rather than from a reconstruction after the fact.

    Legacy · deterministic the computation is re-run

    • Control re-performance against the written test script: recomputed, not re-asserted
    • Model back-testing: re-run against the committed snapshot, seed pinned
    • Consolidation and elimination arithmetic: reproduced from declared postings
    • Journal support: every entry traced to the file that justified it

    AI systems the agent shows its work

    • Evidence assembly for an auditor request: what was pulled and what was left out
    • Deficiency roll-up: the tests each conclusion rested on
    • Model-card claims: assertions traceable to the declared evidence
    The number $2.3M a year · 15,580 staff-hours KPMG, 2025
    The trigger SR 26-2 superseded SR 11-7 · 17 Apr 2026 Fed / OCC / FDIC

    A faithfully re-derived wrong-but-committed rule still re-derives. This says the arithmetic held, never that the control was the right control.

Every figure above sizes the problem. They come from public regulation, operator disclosures and the two firms' own annual reports. What gets added is the re-derivation, run by independent code on somebody else's machine.

The same arithmetic, run globally

The two tabs take this arithmetic to one disclosed payroll each. Run it globally and it lands here. Customer value, not our revenue.

   Where it comes from
25%Of the world's jobs sit in an occupation potentially exposed to generative AI ILO working paper 140, with NASK, May 2025
× ~$62TThe global labour-income base IMF world GDP for 2025 × the 52.4% ILOSTAT labour share
= ~$15.5TWage bill a year sitting in exposed work a quarter of that base
× 25%The share of it where a written rule settles the answer, so the answer re-derives modelled, and the load-bearing assumption
= ~$3.9TContestable production spend a year modelled
≈ $1.1TWhat a customer actually keeps after execution, verification, maintenance, exceptions and adoption

Only the first two lines are published. Everything under them is a modelling assumption, which is exactly why the two firm tabs run the same chain against real disclosed headcount and real disclosed payroll instead of a global average.

Why the bank's number is ten times the insurer's

Not because banking suits this better. On both rates that actually matter, the insurer is the better fit. The bank is simply bigger and pays more per head. Three multipliers, and only the third is about the product.

$0 $1bn $2bn $3bn $258m One large insurer +$1,797m 8.0× the people +$896m 1.4× cost per head −$283m 0.9× share taken $2.67bn One large bank

Start at the insurer's $258m. More people in scope takes it to about $2.06bn, a higher cost per head to about $2.95bn, and a slightly lower share brings it back to $2.67bn. Rounding aside, that is the whole of the gap.

 One large insurer One large bankRatio
People doing this work9,34174,418 8.0×
What one of them costs a year$86,256 $123,8311.4×
Wage bill in this work$806m$9.21bn 11.4×
Share of it that comes off32%29% 0.9×
Share of the whole workforce in scope61% 41%0.7×
A year off the cost base$258m $2.67bn10.3×

Why more of the insurer's people are in scope

Sixty-one per cent of the insurer's staff sit in these kinds of work against forty-one per cent at the bank. An insurer is mostly claims, underwriting and policy administration, which is rule-bound by construction: a filed rate table and a policy wording decide the answer. A bank carries large populations in relationship management, markets, technology and branch operations that never enter this model.

Why the bank pays more for each of them

$123,831 against $86,256, each taken from that firm's own disclosed payroll over its own disclosed headcount. Business mix is most of it. European supervisors put retail banking at $59,030 a head and investment banking at $176,578 in the same year, and a bank's control and markets functions sit at the top of that spread. Geography does the rest, and that part is a reading rather than a disclosure.

So compare the percentage, never the total. The absolute figure scales with payroll, which says more about the firm's size than about the work. The share is the claim: 29% at the bank, 32% at the insurer, and the smaller firm is the cleaner fit.

What removing the wait is worth, counted separately

Everything above prices the person doing the work. It does not price the wait. A day off a cycle that is gated by a check releases one 365th of the value waiting in it, which at a 4% cost of funding is $110,000 a year for every $1bn in flight, per day removed. That unit works against any flow a firm can size.

The larger effect is the ramp. Regulated work takes five years to change because each step waits to be accepted, and acceptance today is a sampled manual review on a quarterly cadence. A check that is free and takes a third of a second turns the sample into the population. Pulling the curve forward one year moves 78 points of the run rate inside the window: $2.08bn at the bank and $201m at the insurer. Same money, sooner, so it is never added to the run rate. Each firm tab shows the arithmetic.

Where to find the full working

The tabs at the top carry this arithmetic against one large bank and one large insurer: the ten kinds of work each one runs, the wage bill, the share that comes off, the build price, and every assumption written out in full at the bottom.

The general method sits on its own page. The eight-step chain this study applies, with the qualifying gates, the sourcing order, the stamp lattice and a fill-in template for taking the same arithmetic to any other sector: the agentic value methodology.

The five numbers for this bank

$2.67bnA year off the cost base once every kind of work below is running. Middle of three cases2. Converted from euros1.
29%Of the wage bill in that work. Not of revenue, not of profit, not of headcount.
$5.9mPaid once, for the whole build. Under one day of the saving. The verifier itself is free forever8.
Year fiveWhen all of it is running. 22% lands in year one, 70% by year three3.
74,418People doing that work today, out of 181,510. 26,694 roles come off across the five years7.

The ten kinds of work, and the saving from each

Each of these is somebody applying a written rule to figures already committed, then handing the answer to a person entitled to challenge it. The share is what an agent takes over once it is fully running2. The money is what is left after the agent's own running cost5.

The wage bill in each kind of work Comes off Customer enquiry and complaints $516m 35% Surveillance and monitoring $403m 50% Reconciliation and settlement $396m 35% Regulatory and statutory reporting $348m 30% Commission, fee and payroll $265m 40% Finance close and consolidation $256m 30% Model validation and back-testing $187m 30% Financial crime screening $161m 40% Document and evidence assembly $75m 40% Control testing and internal audit $59m 35%

Full bar is that work's annual wage bill. The green part is what comes off. All ten together: 74,418 people, $9.21bn of wage bill, $2.67bn a year off the cost base. Judgement is excluded from every kind of work before the share is applied6.

Why this is possible now and was not before

What changed in the technology

An agent is software that reads the inputs, applies the written rule and produces the answer on its own, then hands it to a person. It is not a chatbot waiting to be asked a question. Nothing could do this on a real capital return or a real credit file until roughly two years ago. That part is no longer in doubt, and it is not what anyone is buying.

What changed in the law

Automated decisions in banking now carry written obligations, listed in the next section with the day each one bites. Every one of them comes down to the same demand: when somebody challenges a figure a machine produced, the firm has to be able to show how it got there. A log says a decision happened. It does not reproduce the decision.

What was still missing

A signature proves who sent a file. A hash proves the file has not changed since. Neither one recomputes the number. So the honest position for a banking board until now was that an agent could do the work and nobody could answer for it afterwards, which is why almost none of this work has moved.

The verifier closes exactly that gap. The agent ships its answer together with everything needed to recompute it: the inputs as they stood, the rule as it was pinned, and a hash of each. The auditor, the regulator or the counterparty runs a free open-source checker on their own laptop, offline, with no account and no call to any supplier. The checker recomputes the answer using its own code and returns one of three verdicts: it matches, it does not, or it could not tell. An "it could not tell" never becomes a pass.

Why that holds up against someone trying to game it What it means in practice
The producer cannot ship the code that grades it A package names what to recompute. The code that does the recomputing lives in the checker, and a name the checker does not already know fails closed.
The rule is pinned by whoever is checking Swap in a softer rule and its hash is not the one the checker holds, so nothing resolves. The firm being checked does not get to choose the standard.
An inconclusive run cannot be retried into an approval Three verdicts, and the error state is not a pass. This is the failure mode that kills most audit tooling.
A file nobody declared is a rejection Evidence cannot be added quietly after the fact.

Signing standards, notarisation and supply-chain attestation all answer one of two questions: did the bytes change, or who signed this. This answers a third one, which is whether the answer still comes out. Two research sweeps across more than thirty products found nobody else independently recomputing a declared output9.

What it does not do, said once. It settles integrity, never correctness, and it never re-runs the agent's own words. No language model reproduces its own output, so what the agent carries is a receipt of what it read and what it excluded. The arithmetic underneath is what re-derives10.

The regulations, and the date each one applies

None of this is optional and none of it is far away. Each line below already has a date.

ALREADY IN FORCE GDPR Art 22 DORA Art 10 MiFID II RTS 24 2026 2027 2028 2029 AI Act Art 50 2 Aug 2026, in force Machine-readable marking 2 Dec 2026 AI Act Art 12 2 Dec 2027, high-risk logging
The ruleIn force What it obliges
EU AI Act, Article 502 Aug 2026A system that interacts with a person or generates content has to disclose that it is a machine. Machine-readable marking follows on 2 December 2026.
GDPR, Article 22In forceAnyone on the receiving end of a decision made without a human can demand human review and contest it. The 2023 SCHUFA judgment put credit scoring itself inside that right.
DORA, Article 1017 Jan 2025EU financial firms have to detect anomalous activity in their own systems and evidence that they did.
MiFID II, Article 16 and RTS 24In forceOrder and decision records have to be reconstructable on demand, for years afterwards.
EU AI Act, Article 122 Dec 2027High-risk systems must log events automatically across their whole life. Assessing the creditworthiness of a person is on the high-risk list.

Insurance-side rules are listed on the other tab. Dates from our own regulation tracker, sources checked 4 July 2026.

Nothing here makes anyone compliant, and that word is never used. A re-derivable record is evidence a compliance function can hand over and stand behind. The determination stays with the regulator and the firm, exactly where it was.

What it costs to build and run

WhatPrice What it is
The verifier$0 Open source under Apache-2.0. Runs offline on the auditor's own machine, no account, no network call on the verdict path. No licence, no seat fee, no charge per check, and no supplier sitting in the audit path.
The build$5.9m The producer-side code that turns what the bank's systems already hold into a package the verifier can re-run. Fixed price, four stages: $275k pilot, $990k for the plumbing, $715k for the first workstream into production, then $440k each for the remaining nine8.
Running it10% Of the human cost each kind of work replaces. The agent's own compute plus the people still supervising it. Already deducted from every figure on this page5.
Moving people off it$0 26,694 roles come off across five years, 2.9% of headcount a year. A bank whose ordinary turnover runs above that pays nobody off. A bank in a hurry pays severance, which is its own choice and is not in these figures7.
The whole build is under one day of the saving it makes. $5.9m against $2.67bn a year. Price was never what held this back. Agents able to do the work arrived about two years ago, and a way for an outside party to check what they produced arrived with the verifier.

What the saved time is worth, counted separately

Everything above prices the person. It does not price the wait. They are different pots and they are never added together here. Money moves twice when a cycle shortens: capital stops sitting still, and the whole programme lands sooner.

The unit: per day, per billion in flight

A day taken off a cycle that is gated by a check releases one 365th of the value waiting in it. At a 4% cost of funding, that is $110,000 a year for every $1bn in flight, per day removed12. Ten days off a $20bn flow is $22m a year. Hold that against any process the firm can put a number on.

The effect on the five-year ramp

Regulated work takes five years to change, and the reason is not the software. Each step waits to be accepted, and acceptance today means a sampled manual review on a quarterly or annual cadence. When the check is free and runs in a third of a second, the sample becomes the population and acceptance stops being a scheduled event. Pull the curve forward by one year and the same money arrives earlier.

0% 25% 50% 75% 100% As modelled One year earlier Year 1 Year 2 Year 3 Year 4 Year 5 $2.08bn arrives inside the window

Seventy-eight points of the run rate, arriving inside the five years instead of after them. At this bank that is $2.08bn of saving arriving inside the window rather than beyond it.

This is not new money and it is never added to the run rate. It is the same saving landing sooner, which is worth something to a finance director and nothing to anyone who adds the two together. What makes the mechanism credible is measured: the check runs in 0.335 seconds9. What sits on the clock today is the scheduled review. One control-testing programme runs once a year on 15,580 staff-hours, and when the annual check misses something the correction loop averaged 438 days.

The full working, from headcount to yearly saving

BNP Paribas
181,510 people · $22.0bn payroll · $123,831 per head
FY2025 Universal Registration Document and results, converted from euros1. 74,418 of those people sit in the ten kinds of work above.

The saving, year by year

$0 $1bn $2bn $3bn $4bn $18m Next quarter $583m Year one $1.16bn Year two $1.86bn Year three $2.67bn Year five

Bars are the middle case. The whisker is low to high. The first quarter and the first year both cost money before they pay3.

Cost per person, by pay tier

Three pay tiers rather than one average, scaled so the whole population still reproduces the bank's own disclosed payroll per head4.

Clerical and processing $79,010 Professional and analyst $136,535 Senior and managerial $282,687 Weighted average, matches the disclosure $123,831
Against the number the bank has already given the market. Its cost-to-income target implies about $2.93bn of annual cost out by 2028, across every lever it has. This reaches $2.67bn from agents alone, and not until year five. Close enough to be worth noticing, far enough apart in time to be a different claim.

Every assumption, and the case against each

Four of the ten shares rest on analogy rather than a published result. Six of the ten pay mixes are our own reading. The build price has no external comparison. The phasing curve comes from cost programmes, not AI ones. Population data on 25,000 workers finds no effect at all. Every one of those is written out below, because a reader who finds one unaided stops believing the rest.

  1. Currency. This firm discloses in euros. Every figure here is converted at a flat $1.10 to the euro, a rate that is stated rather than taken off a fixing on any particular day. Every dollar on this tab moves one-for-one with it, so a reader who disagrees with the rate can rescale the whole page in one step.
  2. The share taken, and how it is set. Per kind of work, three cases. Low counts only work under a rule pinned by digest and held in the same version by both sides. Middle adds work under an internal written rule the firm controls. High adds two-party reconciliations where each side holds its own copy. Every headline on this page is the middle case. Eight of the shares sit on a result an operator has published: Ping An, Allianz, Lemonade, Manulife, JPMorgan, HSBC, Deutsche Bank, ING and Aviva. Four rest on analogy to the eight and are marked with this note in the tables: commission and fee calculation, control testing, model validation, finance close. Those four are the first things a pilot should measure. No catch rate or false-positive rate is stated anywhere, because neither has been measured.
  3. Phasing. The chain produces a run rate, and a run rate is not a first-year figure. 22% of it lands by the end of year one, 43% by year two, 70% by year three, 91% by year four and all of it in year five. That curve is the mean of two published ramps: BNP Paribas realised $550m, $660m, $1.10bn, $880m and $660m across 2022 to 2026 against a $3.85bn programme, and UBS booked $4.0bn, then $3.4bn, then $3.2bn against its $13.5bn Credit Suisse integration ambition and called that trajectory non-linear itself. Neither programme was an AI deployment, and no AI programme has published a year-by-year ramp, so the curve is an analogue. No programme publishes a first-quarter realisation either, so quarter one is taken as an eighth of the year-one share, back-loaded inside the year. It is the weakest cell on the curve. Two traps worth naming: HSBC actioned $1.2bn of annualised simplification savings by end-2025 and took only $0.6bn into the P&L, so actioned is not realised; and McKinsey's 28/57/66/74% first-year transformation curve is top-quartile performers only, so it is a ceiling and never a median.
  4. Pay tiers. One blended cost per head is the weakest number a model like this can carry, so the population is split into clerical, professional and senior in the published ratio 1 : 1.73 : 3.58, from US Bureau of Labor Statistics occupational wage estimates for insurance carriers. The shape is imported; the level is not. All three tiers are scaled by a single factor so the weighted average reproduces the firm's own disclosed payroll divided by its own disclosed headcount, exactly. Nothing enters that the firm has not already published in aggregate. Cross-checked against the European Banking Authority's remuneration benchmarking of 2.19 million banking staff, which finds a comparable three-times spread on a different cut: retail banking $59,030, corporate functions $76,472, independent control functions $95,524, investment banking $176,578. What is still assumed is the tier mix inside each kind of work. Where an occupation maps one-to-one it is read off published employment shares; elsewhere, in six of ten cases, it is our classification.
  5. Running cost. Deducted from every figure on this page at 20% of the human cost replaced in the low case, 10% in the middle and 5% in the high. That covers the agent's own compute plus the people still supervising it. Anchored on a measured manual-to-electronic step in which one administrative transaction fell from $1.02 to $0.10, and a status inquiry from $4.38 to $0.04, across plans covering 63% of one market's insured lives.
  6. Judgement is never in a share. Materiality, scope, hardship, discretion, a call between two defensible readings, and the decision to take risk are all excluded from every kind of work before the share is applied. They stay with people.
  7. Severance. 26,694 roles come off across the five years, 2.9% of headcount a year. Whether that costs anything depends on the firm's own turnover, so it is treated as a separate class and never enters a cash figure.
  8. The build price is a quote, not a benchmark. $5.9m, staged. No comparable engagement is published anywhere, so nothing external checks this number. It can only be carried because trebling it barely moves the answer. Cost tracks the number of system estates the inputs live in rather than the size of the saving, and this bank runs 64 countries of estate. The verifier itself is Apache-2.0 and carries no licence fee, ever.
  9. The check, measured. A clean re-derivation runs at 0.335 seconds median across 40 pilot bundles, range 0.214s to 0.413s, wall clock including interpreter startup, on a laptop. Bundles are a few megabytes at most and no scaling curve has been fitted. Tamper runs against the same 40: one byte changed with the manifest untouched was caught 40 times out of 40; one byte changed with the manifest digest recomputed to match was caught 37 out of 40, and the three that pass are declaring that file as carried for storage rather than as an input, which is now visible instead of implied. What re-derives is the arithmetic, never the whole cycle, so the exception queue and the human judgement stay on the clock. On the wider claim in the body: two independent research sweeps across more than thirty products found no vendor doing independent deterministic recomputation of a declared output. Their verifiers recompute a digest of what was declared. This one recomputes the declaration itself.
  10. The rule is yours. This settles that a figure is the deterministic result of the committed computation over the committed inputs. It says nothing about whether the rule behind it was the right rule: a faithfully re-derived wrong-but-committed rule still re-derives, and that is correct behaviour as well as the limit. The agent's own output is never re-run, because language models are not bit-reproducible even at temperature zero. What the agent carries is a receipt of the sources it used and the criteria it included and excluded. Three commercial legal research tools sold as avoiding hallucination were wrong between 17% and 33% of the time in a pre-registered study, which is why a supplier's assurance about its own output is not treated as evidence about that output.
  11. Sources. Firm data from BNP Paribas's Universal Registration Document 2025 and FY2025 results. Pay tiers from US Bureau of Labor Statistics occupational wage estimates, cross-checked against European Banking Authority remuneration benchmarking. The ramp from BNP Paribas's and UBS's own disclosed year-by-year cost programmes. Counter-evidence from Stanford HAI, Gartner, METR and national administrative-data research. Full workings: the agentic value method. These are modelled bands on stated assumptions, not measured savings. The logo is the property of its owner and appears here for identification only.
  12. The clock arithmetic. Value released a year equals the value in flight multiplied by days removed, divided by 365, multiplied by the cost of funding. The cost of funding is taken at 4.0%, which is stated rather than sourced, and every clock figure moves one-for-one with it. The value in flight is the one input a firm has to supply from its own balance sheet, so no figure on this page multiplies it out. The ramp pull-forward is arithmetic on this page's own curve and nothing else: the modelled shares less the same shares shifted one year left, summed, which comes to 78 points of the run rate. It assumes acceptance is what sets the pace of the ramp. Where a workstream is held up by a system migration or a works council rather than by a review, this line is worth nothing and should be struck.

The five numbers for this insurer

$258mA year off the cost base once every kind of work below is running. Middle of three cases2.
32%Of the wage bill in that work. Not of revenue, not of profit, not of headcount.
$2.75mPaid once, for the whole build. Under four days of the saving. The verifier itself is free forever8.
Year fiveWhen all of it is running. 22% lands in year one, 70% by year three3.
9,341People doing that work today, out of 15,338. 3,648 roles come off across the five years7.

The ten kinds of work, and the saving from each

Each of these is somebody applying a written rule to figures already committed, then handing the answer to a person entitled to challenge it. The share is what an agent takes over once it is fully running2. The money is what is left after the agent's own running cost5.

The wage bill in each kind of work Comes off Claims adjudication $85m 45% Reconciliation and settlement $38m 35% Customer enquiry and complaints $35m 35% Commission, fee and payroll $32m 40% Underwriting $29m 50% Regulatory and statutory reporting $12m 30% Finance close and consolidation $11m 30% Control testing and internal audit $6m 35% Financial crime screening $5m 40% Model validation and back-testing $5m 30%

Full bar is that work's annual wage bill. The green part is what comes off. All ten together: 9,341 people, $806m of wage bill, $258m a year off the cost base. Judgement is excluded from every kind of work before the share is applied6.

Why this is possible now and was not before

What changed in the technology

An agent is software that reads the inputs, applies the written rule and produces the answer on its own, then hands it to a person. It is not a chatbot waiting to be asked a question. Nothing could do this on a real claims file or a real policy schedule until roughly two years ago. That part is no longer in doubt, and it is not what anyone is buying.

What changed in the law

Automated decisions in insurance now carry written obligations, listed in the next section with the day each one bites. Every one of them comes down to the same demand: when somebody challenges a figure a machine produced, the firm has to be able to show how it got there. A log says a decision happened. It does not reproduce the decision.

What was still missing

A signature proves who sent a file. A hash proves the file has not changed since. Neither one recomputes the number. So the honest position for a insurance board until now was that an agent could do the work and nobody could answer for it afterwards, which is why almost none of this work has moved.

The verifier closes exactly that gap. The agent ships its answer together with everything needed to recompute it: the inputs as they stood, the rule as it was pinned, and a hash of each. The regulator, the auditor or the policyholder's lawyer runs a free open-source checker on their own laptop, offline, with no account and no call to any supplier. The checker recomputes the answer using its own code and returns one of three verdicts: it matches, it does not, or it could not tell. An "it could not tell" never becomes a pass.

Why that holds up against someone trying to game it What it means in practice
The producer cannot ship the code that grades it A package names what to recompute. The code that does the recomputing lives in the checker, and a name the checker does not already know fails closed.
The rule is pinned by whoever is checking Swap in a softer rule and its hash is not the one the checker holds, so nothing resolves. The firm being checked does not get to choose the standard.
An inconclusive run cannot be retried into an approval Three verdicts, and the error state is not a pass. This is the failure mode that kills most audit tooling.
A file nobody declared is a rejection Evidence cannot be added quietly after the fact.

Signing standards, notarisation and supply-chain attestation all answer one of two questions: did the bytes change, or who signed this. This answers a third one, which is whether the answer still comes out. Two research sweeps across more than thirty products found nobody else independently recomputing a declared output9.

What it does not do, said once. It settles integrity, never correctness, and it never re-runs the agent's own words. No language model reproduces its own output, so what the agent carries is a receipt of what it read and what it excluded. The arithmetic underneath is what re-derives10.

The regulations, and the date each one applies

None of this is optional and none of it is far away. Each line below already has a date.

ALREADY IN FORCE GDPR Art 22 UK DUAA 2025 US FRE 902(13)/(14) 2026 2027 2028 2029 AI Act Art 50 2 Aug 2026, in force Machine-readable marking 2 Dec 2026 AI Act Art 12 2 Dec 2027, high-risk logging
The ruleIn force What it obliges
EU AI Act, Article 502 Aug 2026A system that interacts with a person or generates content has to disclose that it is a machine. Machine-readable marking follows on 2 December 2026.
GDPR, Article 22In forceAnyone on the receiving end of a decision made without a human can demand human review and contest it. A declined claim or a priced policy sits squarely inside that.
UK DUAA 20255 Feb 2026Reformed automated-decision rules, with safeguards that include the right to contest the outcome and obtain human intervention.
US Federal Rules of Evidence 902(13) and (14)In forceA machine-generated record can authenticate itself in court on a certification, with no witness called, if the process behind it can be shown.
EU AI Act, Article 122 Dec 2027High-risk systems must log events automatically across their whole life. Risk assessment and pricing in life and health insurance is on the high-risk list.

Banking-side rules are listed on the other tab. Dates from our own regulation tracker, sources checked 4 July 2026.

Nothing here makes anyone compliant, and that word is never used. A re-derivable record is evidence a compliance function can hand over and stand behind. The determination stays with the regulator and the firm, exactly where it was.

What it costs to build and run

WhatPrice What it is
The verifier$0 Open source under Apache-2.0. Runs offline on the auditor's own machine, no account, no network call on the verdict path. No licence, no seat fee, no charge per check, and no supplier sitting in the audit path.
The build$2.75m The producer-side code that turns what the insurer's systems already hold into a package the verifier can re-run. Fixed price, four stages: $150k pilot, $450k for the plumbing, $350k for the first workstream into production, then $200k each for the remaining nine8.
Running it10% Of the human cost each kind of work replaces. The agent's own compute plus the people still supervising it. Already deducted from every figure on this page5.
Moving people off it$0 3,648 roles come off across five years, 4.8% of headcount a year. An insurer whose ordinary turnover runs above that pays nobody off. One in a hurry pays severance, which is its own choice and is not in these figures7.
The whole build is under four days of the saving it makes. $2.75m against $258m a year. Price was never what held this back. Agents able to do the work arrived about two years ago, and a way for an outside party to check what they produced arrived with the verifier.

What the saved time is worth, counted separately

Everything above prices the person. It does not price the wait. They are different pots and they are never added together here. Money moves twice when a cycle shortens: capital stops sitting still, and the whole programme lands sooner.

The unit: per day, per billion in flight

A day taken off a cycle that is gated by a check releases one 365th of the value waiting in it. At a 4% cost of funding, that is $110,000 a year for every $1bn in flight, per day removed12. Ten days off a $20bn flow is $22m a year. Hold that against any process the firm can put a number on.

The effect on the five-year ramp

Regulated work takes five years to change, and the reason is not the software. Each step waits to be accepted, and acceptance today means a sampled manual review on a quarterly or annual cadence. When the check is free and runs in a third of a second, the sample becomes the population and acceptance stops being a scheduled event. Pull the curve forward by one year and the same money arrives earlier.

0% 25% 50% 75% 100% As modelled One year earlier Year 1 Year 2 Year 3 Year 4 Year 5 $201m arrives inside the window

Seventy-eight points of the run rate, arriving inside the five years instead of after them. At this insurer that is $201m of saving arriving inside the window rather than beyond it.

This is not new money and it is never added to the run rate. It is the same saving landing sooner, which is worth something to a finance director and nothing to anyone who adds the two together. What makes the mechanism credible is measured: the check runs in 0.335 seconds9. What sits on the clock today is the scheduled review. Claims, reserving and reporting all run to a scheduled review rather than a continuous one, because a manual re-check has always had to be rationed.

The full working, from headcount to yearly saving

Prudential plc
15,338 people · $1.32bn payroll · $86,256 per head
FY2025 Annual Report and Form 20-F. 16.6m policies and about 3.0m claims a year. 9,341 of those people sit in the ten kinds of work above.

The saving, year by year

$0 $100m $200m $300m $400m $1m Next quarter $56m Year one $111m Year two $179m Year three $258m Year five

Bars are the middle case. The whisker is low to high. The first quarter and the first year both cost money before they pay3.

Cost per person, by pay tier

Three pay tiers rather than one average, scaled so the whole population still reproduces the insurer's own disclosed payroll per head4.

Clerical and processing $57,190 Professional and analyst $98,828 Senior and managerial $204,617 Weighted average, matches the disclosure $86,256
This insurer is already ahead on one line. It runs 70% auto-underwriting against a market average of 11%. The underwriting row above is sized on the remaining 30%, not on the whole book.
Leakage is not in any figure above. The firm recovered over $100m of fraud, waste and abuse in 2025 from focused effort. Running the same check across the whole book is worth $30m low, $75m middle and $150m high. That is a different class of money and it is never added to a payroll figure.

Every assumption, and the case against each

Four of the ten shares rest on analogy rather than a published result. Six of the ten pay mixes are our own reading. The build price has no external comparison. The phasing curve comes from cost programmes, not AI ones. Population data on 25,000 workers finds no effect at all. Every one of those is written out below, because a reader who finds one unaided stops believing the rest.

  1. Currency. Every figure on this tab is disclosed in dollars by the firm itself and is carried untouched. Nothing here is converted, so no exchange-rate assumption sits underneath any of it.
  2. The share taken, and how it is set. Per kind of work, three cases. Low counts only work under a rule pinned by digest and held in the same version by both sides. Middle adds work under an internal written rule the firm controls. High adds two-party reconciliations where each side holds its own copy. Every headline on this page is the middle case. Eight of the shares sit on a result an operator has published: Ping An, Allianz, Lemonade, Manulife, JPMorgan, HSBC, Deutsche Bank, ING and Aviva. Four rest on analogy to the eight and are marked with this note in the tables: commission and fee calculation, control testing, model validation, finance close. Those four are the first things a pilot should measure. No catch rate or false-positive rate is stated anywhere, because neither has been measured.
  3. Phasing. The chain produces a run rate, and a run rate is not a first-year figure. 22% of it lands by the end of year one, 43% by year two, 70% by year three, 91% by year four and all of it in year five. That curve is the mean of two published bank ramps. The first realised $550m, $660m, $1.10bn, $880m and $660m across 2022 to 2026 against a $3.85bn programme, and is named in the banking tab. UBS booked $4.0bn, then $3.4bn, then $3.2bn against its $13.5bn Credit Suisse integration ambition and called that trajectory non-linear itself. Neither programme was an AI deployment, and no AI programme has published a year-by-year ramp, so the curve is an analogue. No programme publishes a first-quarter realisation either, so quarter one is taken as an eighth of the year-one share, back-loaded inside the year. It is the weakest cell on the curve. Two traps worth naming: HSBC actioned $1.2bn of annualised simplification savings by end-2025 and took only $0.6bn into the P&L, so actioned is not realised; and McKinsey's 28/57/66/74% first-year transformation curve is top-quartile performers only, so it is a ceiling and never a median.
  4. Pay tiers. One blended cost per head is the weakest number a model like this can carry, so the population is split into clerical, professional and senior in the published ratio 1 : 1.73 : 3.58, from US Bureau of Labor Statistics occupational wage estimates for insurance carriers. The shape is imported; the level is not. All three tiers are scaled by a single factor so the weighted average reproduces the firm's own disclosed payroll divided by its own disclosed headcount, exactly. Nothing enters that the firm has not already published in aggregate. Cross-checked against the European Banking Authority's remuneration benchmarking of 2.19 million banking staff, which finds a comparable three-times spread on a different cut: retail banking $59,030, corporate functions $76,472, independent control functions $95,524, investment banking $176,578. What is still assumed is the tier mix inside each kind of work. Where an occupation maps one-to-one it is read off published employment shares; elsewhere, in six of ten cases, it is our classification.
  5. Running cost. Deducted from every figure on this page at 20% of the human cost replaced in the low case, 10% in the middle and 5% in the high. That covers the agent's own compute plus the people still supervising it. Anchored on a measured manual-to-electronic step in which one administrative transaction fell from $1.02 to $0.10, and a status inquiry from $4.38 to $0.04, across plans covering 63% of one market's insured lives.
  6. Judgement is never in a share. Materiality, scope, hardship, discretion, a call between two defensible readings, and the decision to take risk are all excluded from every kind of work before the share is applied. They stay with people.
  7. Severance. 3,648 roles come off across the five years, 4.8% of headcount a year. Whether that costs anything depends on the firm's own turnover, so it is treated as a separate class and never enters a cash figure.
  8. The build price is a quote, not a benchmark. $2.75m, staged. No comparable engagement is published anywhere, so nothing external checks this number. It can only be carried because trebling it barely moves the answer. Cost tracks the number of system estates the inputs live in rather than the size of the saving, and this insurer runs 20 markets. The verifier itself is Apache-2.0 and carries no licence fee, ever.
  9. The check, measured. A clean re-derivation runs at 0.335 seconds median across 40 pilot bundles, range 0.214s to 0.413s, wall clock including interpreter startup, on a laptop. Bundles are a few megabytes at most and no scaling curve has been fitted. Tamper runs against the same 40: one byte changed with the manifest untouched was caught 40 times out of 40; one byte changed with the manifest digest recomputed to match was caught 37 out of 40, and the three that pass are declaring that file as carried for storage rather than as an input, which is now visible instead of implied. What re-derives is the arithmetic, never the whole cycle, so the exception queue and the human judgement stay on the clock. On the wider claim in the body: two independent research sweeps across more than thirty products found no vendor doing independent deterministic recomputation of a declared output. Their verifiers recompute a digest of what was declared. This one recomputes the declaration itself.
  10. The rule is yours. This settles that a figure is the deterministic result of the committed computation over the committed inputs. It says nothing about whether the rule behind it was the right rule: a faithfully re-derived wrong-but-committed rule still re-derives, and that is correct behaviour as well as the limit. The agent's own output is never re-run, because language models are not bit-reproducible even at temperature zero. What the agent carries is a receipt of the sources it used and the criteria it included and excluded. Three commercial legal research tools sold as avoiding hallucination were wrong between 17% and 33% of the time in a pre-registered study, which is why a supplier's assurance about its own output is not treated as evidence about that output.
  11. Sources. Firm data from Prudential plc's AR2025, Form 20-F and Sustainability Report 2025. Pay tiers from US Bureau of Labor Statistics occupational wage estimates for insurance carriers, cross-checked against European Banking Authority remuneration benchmarking. The ramp from two banks' own disclosed year-by-year cost programmes. Counter-evidence from Stanford HAI, Gartner, METR and national administrative-data research. Full workings: the agentic value method. These are modelled bands on stated assumptions, not measured savings. The logo is the property of its owner and appears here for identification only.
  12. The clock arithmetic. Value released a year equals the value in flight multiplied by days removed, divided by 365, multiplied by the cost of funding. The cost of funding is taken at 4.0%, which is stated rather than sourced, and every clock figure moves one-for-one with it. The value in flight is the one input a firm has to supply from its own balance sheet, so no figure on this page multiplies it out. The ramp pull-forward is arithmetic on this page's own curve and nothing else: the modelled shares less the same shares shifted one year left, summed, which comes to 78 points of the run rate. It assumes acceptance is what sets the pace of the ramp. Where a workstream is held up by a system migration or a works council rather than by a review, this line is worth nothing and should be struck.