Why this page exists
If you evaluate AI bookkeeping tools for a living, you may have noticed that published, measured accuracy figures with a stated methodology are hard to find. Confidence is easy to project. A measurement you can interrogate, with definitions attached, is a different thing.
AIONA's position is simple. We will measure our own accuracy against the definitions on this page, and we will publish the results. We are not publishing figures yet. Our measured sample is not large enough to be statistically meaningful, and a small sample dressed up as a benchmark is exactly the kind of claim this page exists to reject. When the figures come, this page is what they will mean. Until then, treat it as a standard you can hold us to.
We score fields, not documents
When AIONA reads a purchase invoice or a receipt, it does not produce one answer. It produces several: the supplier, the document date, the net, VAT and gross amounts, the VAT treatment, and a proposed ledger coding. Each field is scored separately.
This matters because document-level accuracy flatters the result. A read that gets the total right but splits the VAT wrongly has produced an error that will surface in a VAT return. A correct amount coded to the wrong account will surface in the accounts. Scoring at field level means those failures count as failures, rather than disappearing inside a document marked correct.
What counts as an error
We count an error whenever the system proposes any of the following and a human has to change it:
- the wrong supplier;
- the wrong date;
- the wrong net, VAT or gross amount;
- the wrong VAT treatment;
- the wrong ledger account.
An entry is counted correct only if the posted journal needed no human correction.
Not nearly right. Not right after a quick edit. If a reviewer had to touch an entry before it was fit to stand in the ledger, the fields they changed are recorded as errors. This is a deliberately unforgiving definition, and it is the only one that reflects what a practice actually pays for: work that does not need doing again. And because every journal AIONA posts stays linked to its source document, whether a field was wrong is settled by the document, not by memory.
Confidence tiers are operating policy, not decoration
AIONA attaches a confidence tier to every proposed entry. The tiers are not a colour scheme; each one changes what the software is allowed to do.
- Green means every field has been verified, the supplier is one your practice has established as trusted, and the ledger coding was not a close call. Green entries are eligible to post after review.
- Amber entries cannot post until a reviewer gives explicit one-click confirmation of the account, the VAT treatment and the total.
- Red entries are blocked until they have been reviewed. Overriding a red requires a written reason, and that reason is captured in a tamper-evident audit log.
Because the tiers gate behaviour, they can be tested. A label that changes nothing can never be wrong. A label that decides what must be confirmed can be checked against what turned out to need correcting.
Calibration: a tier must keep its promise
A confidence tier is honest only if its observed error rate matches what the tier promises. If entries marked green turn out to need correction more often than green implies, the tier is miscalibrated and the label is misleading, however reassuring it looks on screen.
This is where the review workflow earns its keep. Every correction in AIONA leaves a record: which field changed, what the system proposed, what the reviewer decided, and the tier the entry carried at the time. That record is what makes an error rate measurable at all. Without it, accuracy is an anecdote. With it, calibration is arithmetic: for each tier, in each field, compare what was promised with what was corrected.
The measurement cannot be tidied up afterwards
AIONA's ledger is immutable. Corrections are made by reversal: the original entry remains, a reversing entry cancels it, and the corrected entry follows. Any accountant will recognise the discipline, and it has a consequence for measurement that we consider a feature. The record of what the system originally proposed is preserved. We cannot quietly overwrite our mistakes and then measure the cleaned-up version. The audit trail that protects your clients' books is the same mechanism that stops us marking our own homework.
What we will publish
When the measured sample is large enough to carry statistical weight, we will publish:
- the sample size;
- the period the sample covers;
- the definitions above, stated alongside the results they produced;
- field-level results: supplier, date, amounts, VAT treatment and ledger coding reported separately.
What we will not publish is one blended headline number. A single percentage across all fields and all tiers hides exactly the failures a practitioner needs to know about, and it invites comparison with figures whose definitions are unstated. If a number on this page ever needs a caveat to survive scrutiny, the caveat will be printed next to it.
What you can do today
You do not need to wait for our figures to apply this standard. We have written up the questions we believe any practice should put to any vendor of AI bookkeeping, including AIONA: what a confidence score commits the software to, what counts as an error, and how corrections are recorded. You will find them in What AI confidence scores mean.
Ask those questions of every tool you evaluate, and judge the answers. If ours ever fall short of this page, we would like to hear about it.