There are two ways to build AI for bookkeeping, and they disagree about one question: what should the software do when it is not sure?
The first philosophy says post anyway. Read the document, pick the most likely coding, put the entry in the ledger, and rely on someone noticing later if it was wrong. This is the philosophy behind every impressive automation headline. Software that posts everything it sees has an automation rate close to one hundred per cent by construction, and it demos beautifully.
The second philosophy says stop. If the software cannot vouch for a figure, the figure does not post. Uncertain work goes into a queue for a person, the automation rate is visibly lower, and part of every demo is a human clicking approve. It looks slower, and on the day the document arrives, it genuinely is.
I want to argue that the second philosophy is the only one an accounting practice can responsibly build on. Not because automation is oversold, and not because the people building the first kind are careless. The argument is about where liability sits when the software is wrong, and about a cost that never appears in an automation rate: the cost of finding a silently posted error before you can fix it.
The industry is moving the other way
Automated bank reconciliation is fast becoming a default expectation rather than a novelty. Xero began rolling out automatic bank reconciliation in beta in November 2025, and it is worth being fair about the design: it reconciles a transaction only when its engine is highly confident, and every reconciled line stays visible for review afterwards. Even the highest-profile implementation of automatic reconciliation is confidence-gated. The real disagreement in the industry is not about whether confidence matters. It is about what happens below the threshold, how visible that threshold is, and who reviews the work that falls under it.
The commercial pressure behind the post-first camp is easy to understand. Vendors compete on the automation number because it fits on a homepage. The error rate among auto-posted entries fits on no homepage, and in fairness it is genuinely hard to measure, because measuring it means paying qualified people to re-review work the product was sold to eliminate.
Where the liability actually sits
Start with the legal position, because it is often stated backwards. When an agent submits a VAT return, responsibility for its accuracy stays with the taxpayer. HMRC looks to the client, not the practice, and certainly not the software. But that is not how the client experiences it. They engaged you precisely so they would not have to think about codings and boxes, and when an error surfaces, the questions come to you: through the engagement letter, the complaints procedure, the professional indemnity policy, and the quiet reputational ledger that decides referrals.
The size of the damage depends on behaviour. A careless inaccuracy can attract a penalty of up to 30 per cent of the potential lost revenue; a deliberate one, 20 to 70 per cent; deliberate and concealed, up to 100 per cent. The word doing the work in that scale is careless, which turns on whether reasonable care was taken. "The software posted it and nobody looked" is a difficult account of reasonable care to give with a straight face, and it gets harder as the unreviewed entries pile up.
Notice who is absent from this chain. The software vendor's terms will tell you, in careful language, that the figures are your responsibility. That asymmetry should worry you as a matter of design: software whose mistakes cost it nothing has no structural reason to hold back when it is unsure. The restraint has to be designed in on purpose.
Review debt
Here is the cost that automation rates conceal. Every entry posted without review is a small loan taken out against a future reviewer. The debt has two components: the cost of fixing wrong entries, which is modest, and the cost of finding them, which is not. When uncertain work is held in a queue, the finding cost is zero, because the queue is the list. When everything posts silently, the wrong entries are camouflaged among the right ones. They carry the same tick, sit in the same reconciled screen, and announce nothing.
Imagine a 40-client firm where each client generates around 250 bank and purchase transactions a month, and an automated system posts all of them, getting 96 in 100 right. That is a genuinely good hit rate, and it produces roughly 400 wrong entries a month distributed invisibly through 10,000 correct-looking ones. Nobody is going to re-review 10,000 entries to find 400. So nobody does, and the debt rolls forward until something forces repayment: year-end, a VAT enquiry, a due diligence exercise, or a new bookkeeper inheriting the file and slowly losing trust in it.
Review debt is invisible in every metric that sold the software. The automation rate went up. The time per client went down. The file looks immaculate. The bill arrives later, denominated in partner hours and, in the worst case, in penalty percentages.
"The books look done" is not "the books are done"
A file with no exceptions on screen can mean one of two things. Either everything was checked, or everything was guessed and no guess has been challenged yet. The two files look identical. The difference only becomes visible under inspection, which is exactly the moment you want your file at its strongest, and exactly the moment you cannot go back and add the review that did not happen. I have written separately about the tests an inspecting eye applies to a bookkeeping file; the short version is that an inspector does not ask whether the books balance. They ask how you know.
Done, in the sense that matters, means every posted entry was either verified by a person or posted by a system that was entitled to its confidence, and the file records which was which. Anything short of that is "probably done", and probably is not a word you want in your reply to an enquiry letter.
What review-first means concretely
Review-first is not a sentiment on a vendor's about page. It is a set of mechanisms, and you can test for each one.
- Confidence that gates rather than decorates. Every extracted figure should carry a confidence score, and below a threshold, posting should be blocked, not merely flagged. A warning that can be scrolled past is decoration. I have written more about what a confidence score should actually commit a system to.
- Overrides that leave a mark. A human should be able to overrule the system; that is the entire point of review. But the overrule should require a written reason, attached to a named person, in a log that cannot be quietly edited afterwards.
- Correction by reversal, never by editing. Posted entries should be immutable. Mistakes get corrected by a reversing entry, so the error and its correction both stay on the record. A file whose history can be rewritten proves nothing about anything.
- A source behind every journal. Any figure in the ledger should trace back to the document that produced it, in one click, years later, without an archaeology project.
What these mechanisms buy, collectively, is a change in when the review happens. The human effort does not disappear; it moves to before posting, where the uncertain items are queued and the work is cheap, instead of after posting, where the same work becomes a search. Same effort, radically different price.
Questions to ask any vendor
Every brochure now says accuracy and control. The design philosophy shows up in the answers to specific questions, so ask them:
- What happens, mechanically, when your AI is not sure? Ask to see the actual screen, not an architecture diagram.
- Can an uncertain item reach the ledger without a person approving it, under any setting a customer can choose?
- Is confidence shown per figure, and does any threshold actually block posting?
- When someone overrides the system, is a written reason captured, and can that record be altered later?
- Can a posted entry be edited in place, or only corrected by reversal?
- Can I export the complete ledger, documents and audit trail today, without an exit fee, if I choose to leave?
- What is the measured error rate among entries the system posted on its own, and how was it measured?
The last question is the revealing one. If a vendor cannot answer it, then the automation rate they do publish is the numerator of a fraction whose denominator they have never measured. None of these questions is hostile, and a vendor who is confident in their design will enjoy answering them. If you are weighing up software ahead of the next round of MTD obligations, the same questions belong in your practice checklist for 2026.
Where AIONA fits
AIONA is built entirely on the second philosophy. Every figure it reads carries a GREEN, AMBER or RED confidence tier; low confidence blocks posting until a person has reviewed it, overrides require a written reason recorded in a tamper-evident audit log, posted entries are corrected only by reversal, and every journal links back to its source document. Nothing reaches the ledger without explicit human sign-off. If you want to test the idea against a real file, you can connect a client's Xero book read-only and AIONA will grade it without being able to write anything back: aionatech.com.