A receipt goes in, the fields come back filled, and next to them sits a number: 97 per cent confident. It reads like a fact about the document. It is not. It is a claim the software is making about itself, and the useful question for anyone who carries filing liability is what that claim entitles you to skip.
Here is the argument of this piece: a confidence score is not a probability handed down by the machine. It is a policy decision about when a human looks. The number itself is instrumentation; the threshold that decides whether an item waits for review is policy, and someone chose where it sits. Once you see scores this way, the questions you should ask of any tool that produces them get much sharper.
What 97 per cent does and does not tell you
When software reads a document, it produces a self-estimate alongside each extracted value. That estimate is built from internal signals: how clean the image was, how familiar the layout looked, how strongly different reading passes agreed with one another. It is generated at the moment of extraction, by the same system doing the extracting. It is not a measurement against the truth of your documents, because at that moment nobody knows the truth yet.
So a 97 tells you the system considers this reading a strong one by its own lights. It does not tell you:
- That the figure is right. High confidence and wrong is a real and routine combination, especially on look-alike layouts where a familiar supplier template contains an unfamiliar surprise.
- Anything about the bookkeeping judgement. Reading £120 correctly is not the same as knowing whether the VAT is reclaimable, whether the expense is allowable, or whether the date belongs to an open period. Confidence covers "did I read the characters into the right fields", not "should this be posted as it stands".
- What happens next. A score only acquires meaning when it is attached to a consequence. Two tools with identical reading quality but different thresholds expose you to very different risk, because one of them shows a human more of the doubtful cases than the other.
That last point is the crux. The vendor sets a default threshold, and that threshold encodes a trade-off between your review time and your error rate. It is a legitimate trade-off to make. It is not legitimate for it to be invisible.
Calibration, in plain terms
Calibration is the first thing a sceptical practitioner should probe: when the tool says 90, is it right about nine times in ten? In a well-calibrated system, yes, roughly. The scores can be read at face value, and a threshold set at 90 is an explicit decision that the weakest items passing unreviewed can be wrong about one time in ten. You may or may not like that trade, but at least you know what you are trading.
An overconfident system says 95 and is right eight times in ten. It feels wonderful to use, because everything glows with assurance, and it is quietly the more dangerous of the two. An underconfident system wastes your time instead, dragging you through items that were fine. Of the two failure modes, overconfidence is the one that ends up in a client's accounts.
You can test calibration yourself without any statistics. Pull a sample of items the tool scored in a narrow band, say the low nineties. Check each one against its source document. Count the errors. If a band that implies roughly one error in ten produces one in four, the scores flatter the system, and every threshold built on them has been set blind. Half a day with fifty documents will tell you more about a tool than any demonstration.
Two caveats. First, calibration is a property of a document mix, not of the tool alone. Scores earned on crisp digital invoices may not survive contact with phone photos of crumpled receipts, and your client base decides which of those worlds you live in. Second, calibration drifts. A supplier redesigns an invoice, a client changes till systems, and yesterday's reliable band quietly is not. A calibration check is a habit, not a one-off.
One blended percentage hides field-level risk
A document is not one reading. It is many: date, supplier, net, VAT, gross, currency, often a suggested category. Each can be right or wrong on its own. A single blended score averages them, and averages hide.
Imagine a supermarket receipt for £120. The total is printed large and clean, and the software reads it perfectly. But the basket mixes standard-rated items with zero-rated food, and the tool assumes the whole supply is standard-rated at 20 per cent, booking £20 of input VAT where the receipt supports less. A blended score of 96 can sit on top of exactly this: a 99 on the total and a much weaker read on the VAT split. The ledger looks nearly right, and the error flows straight into the VAT return, in the expensive direction: input tax claimed that the evidence does not support.
The errors that matter to a practice concentrate in precisely the fields a blended number can bury: VAT treatment, dates that fall near a period boundary, and supplier identity when it drives your coding rules. If a tool can only show one score per document, treat the document as its weakest important field, because that is what the audit quality of the file ultimately rests on.
Three ways to run review, honestly compared
Review everything
The highest apparent safety, and it has two failure modes. The first is cost: at volume, checking every field of every document is a full-time job that nobody billed for. The second is attention. A person who confirms three hundred correct items in a row stops seeing them. Approval becomes muscle memory, and you end up with the liability of having reviewed without the substance of review. Checking everything and checking nothing converge more often than either name suggests.
Review nothing
Silent automation posts whatever it reads. The errors do not announce themselves; they compound quietly and surface at the VAT quarter or the year end, when the cost of unwinding them has multiplied and the client is the one asking questions. I have written separately about review-first versus silent automation; the short version is that a system that never asks is not confident, it is unaccountable.
Review by threshold
The defensible middle: humans see the uncertain slice, and the routine slice flows. But a threshold workflow is only as good as the machinery beneath it. It needs calibrated scores, or the threshold means nothing. It needs field-level scoring, or the VAT split rides through on the coat-tails of a cleanly read total. It needs the below-threshold behaviour to be a hard stop rather than a dismissible nudge. And it needs the threshold to be yours to move, because your risk appetite for a new client with chaotic paperwork is not your risk appetite for a tidy client three years in.
One more honest note: threshold review does not remove your responsibility for the items that pass. It changes your sampling basis. Periodically pull a handful of high-confidence items and check them anyway; that spot check is your ongoing calibration audit, and it is cheap.
Questions worth asking any vendor
None of these are gotchas. They are the questions a tool built for accountable review will answer readily.
- Do you publish measured accuracy? Not a headline percentage from a brochure, but a figure with a stated methodology: what document mix, measured when, broken down by field. Headline accuracy claims are everywhere in this market; published figures with a methodology behind them are hard to find, and the gap between those two things is worth noticing.
- What happens below the threshold? Is the item blocked until a person reviews it, or flagged and posted anyway? A flag that can be ignored under deadline pressure is not a control.
- Can the threshold be changed? Per field, and ideally per client. If the answer is no, the vendor has decided your risk appetite for you, permanently.
- Are overrides logged? When someone waves through a doubtful item or overrules a reading, is there a record of who, when and why that cannot be quietly edited later? If an enquiry ever asks how a figure got into the ledger, that record is your answer.
- Does correcting a mistake teach the system, and is that learning scoped to your firm? A correction should make the same mistake less likely next month. And your clients' corrections should improve your own setup, not drain into a pool shared with every other customer.
A vendor who is comfortable with this list treats review as part of the product. A vendor who treats it as friction to be engineered away is telling you what they think your review is for.
Where AIONA fits
AIONA is built on the policy view described here. Every extracted figure carries its own GREEN, AMBER or RED confidence tier, low-confidence items are blocked from posting until a person reviews them, and overrides require a written reason that is recorded in a tamper-evident audit log. Nothing posts to the ledger, and nothing goes to HMRC, without explicit human sign-off; our application for HMRC software recognition is in progress, so live submission is not yet available. If that matches how you already think about review, you can read more at aionatech.com.