NUMBER7 RESEARCH / CONFIDENCE DECOMPOSITION AND DECISION MODEL

Invoice Extraction Confidence Is Not Transaction Confidence

DIRECT ANSWER Invoice extraction confidence estimates whether a system read a field or structure correctly. Transaction confidence asks whether the underlying payable is sufficiently identified, evidenced, accounted for, controlled, authorized, and executable for the next action. A high-confidence invoice total can coexist with an ambiguous vendor, missing receipt, duplicate obligation, or unauthorized approval.

  • Field confidence applies to a prediction; transaction readiness applies to an action.
  • A single averaged confidence score can hide a failed hard control.
  • Use a confidence vector plus explicit blockers before experimenting with an aggregate score.
  • Calibrate each signal against its real outcome and cost of error.
  • Thresholds should depend on the proposed action and materiality, not one universal percentage.
NUMBER7 RESEARCH MODELTRANSACTION CONFIDENCE
01DOCUMENT · .99
02IDENTITY · .61
03EVIDENCE · PASS
04ACCOUNTING · .88
05AUTHORITY · NOT EVALUATED
06EXECUTION · —

A failed mandatory state cannot be averaged away.

What an extraction confidence score means

An invoice extraction system commonly returns a value for a field and a confidence or certainty signal. The field may be vendor name, invoice number, date, subtotal, tax, total, or a line item. Microsoft’s invoice model documentation, for example, describes structured fields returned from invoices; platform implementations may expose confidence at different levels. Source: Microsoft Learn: invoice data extraction.

The useful question is not whether the UI shows “98%.” It is whether that score is calibrated for the field, document population, and error definition. Among predictions labeled 0.98, are roughly 98% correct under the chosen ground truth? A score may instead be a ranking signal that is useful for prioritization but not a literal probability.

Even a well-calibrated field score answers only an interpretation question: what does the source appear to say?

Formal distinction

Extraction confidence: a field-, region-, line-, or document-level signal about the correctness of machine interpretation relative to a labeled source.

Transaction confidence: a decision-specific assessment of whether the current economic-event state has sufficient identity, evidence, accounting treatment, control satisfaction, authority, and execution reliability for the proposed next action.

The phrase “decision-specific” matters. The evidence needed to save a draft is not the evidence needed to post a bill, and the evidence needed to post is not necessarily the evidence needed to execute a payment.

Use a confidence vector, not a magic average

Number7 proposes a transaction-confidence vector:

C = [Cᴅ, Cɪ, Cᴇ, Cᴀ, Cʀ, Cₓ]

ComponentWhat it assessesExample blocker
Cᴅ — DocumentClassification, boundaries, fields, lineageMissing page or uncertain invoice number
Cɪ — IdentityVendor/entity/master resolutionTwo active suppliers share similar names
Cᴇ — EvidenceRequired support and consistencyReceipt missing or quantity conflict
Cᴀ — AccountingLedger, entity, tax, dimensions, periodTax treatment unresolved
Cʀ — Rights/AuthorityPolicy and permitted actor/state transitionApproval threshold not satisfied
Cₓ — ExecutionIdempotency, target acceptance, and verificationAPI timeout with unknown target state

Do not average these components by default. If authority is absent, strong document and identity signals do not make the transaction authorized. If execution state is unknown after a timeout, retry risk cannot be erased by a correct total.

A safer decision rule is:

Permit an action only when all mandatory controls pass and each action-relevant confidence component meets its calibrated threshold; otherwise create a named exception.

Worked example

An illustrative system extracts an invoice with high field confidence:

LayerSignalState
Document/fields0.99Invoice number, dates, lines, and totals appear reliable
Identity0.61Two vendor-master candidates remain plausible
EvidencePass for non-PO flowRequired source invoice present
Accounting0.88GL proposal has relevant prior context
AuthorityNot evaluatedApprover route has not been established
ExecutionNot applicable yetNo target action attempted

A naïve average of the available numeric signals is 0.83. That number is operationally meaningless. Identity is unresolved and authority is unevaluated. The correct state is not “83% ready”; it is “blocked on vendor identity and approval route.”

The smallest human decision is to choose or create the correct vendor record. After that, the system can evaluate the next control.

Thresholds depend on action and error cost

The same proposal can be acceptable for one action and unacceptable for another:

Proposed actionPossible threshold posture
Pre-fill a review formLower threshold; human sees source and proposal
Auto-assign a low-risk coding suggestionCalibrated threshold with visible correction path
Post a billHigher thresholds plus identity, evidence, accounting, and authority controls
Trigger money movementSeparate, stricter authorization and execution controls; outside AIdaptIQ’s current product claim

Threshold design should incorporate false-accept cost, false-reject cost, materiality, reversibility, and downstream detection. One global “95% confidence” rule is not a control design.

How to validate confidence

  1. Define ground truth for each field or decision.
  2. Separate critical fields and high-cost actions from low-risk metadata.
  3. Measure precision and error rate inside score bands, not only overall accuracy.
  4. Plot reliability: predicted confidence band versus observed correctness.
  5. Test by vendor, document family, scan quality, language, client/entity, and time.
  6. Record abstentions and human corrections.
  7. Sample high-confidence auto-accepted items for silent errors.
  8. Recalibrate after material model, prompt, parsing, or population changes.

For transaction components such as authority, “pass/fail/not evaluated” may be more honest than a probability.

Counterexample: the confident duplicate

Two invoices from the same supplier show different document numbers but describe the same economic obligation after a resubmission. Every character can be extracted correctly. Field-level accuracy is perfect. Only evidence, identity, and duplicate/economic-event reasoning reveal the risk.

This is why “invoice accuracy” must specify whether it means transcription, field correctness, accounting correctness, or completed transaction outcome.

Limitations

The confidence-vector model does not prescribe universal components, weights, or thresholds. Some states are categorical controls rather than predictions. Components may be dependent: an identity error can alter the accounting and duplicate assessment. Calibration requires representative labeled data and can drift.

Do not publish one aggregate transaction-confidence score until its semantics, calibration, control gates, and failure costs are documented. A vector is less marketable but more truthful.

Sources

How to cite

Number7 Research. “Invoice Extraction Confidence Is Not Transaction Confidence.” Number7AI, revision 1.0, 7 September 2026. https://number7ai.com/research/transaction-confidence.

Revision history

RevisionDateChangeReviewer
1.07 September 2026Initial field-versus-transaction definition, confidence vector, and control ruleNumber7 AI editorial review

Next step

Replace one global confidence threshold with action-specific controls and calibration reports.

ASSERTION LINEAGE

Evidence travels with the accounting assertion.

Material values retain source, locator, interpretation, control result and supersession history so a later action can be reconstructed.

ASSERTION LINEAGEEvidence travels with the action.
PAGE REGION
CLAIM
INTERPRETATION
CONTROL
DECISION
ERP OUTCOME
Read the full guide: Evidence travels with the accounting assertion.

What it includes

01

Source

Treat source as explicit transaction state with evidence, a control outcome and ownership—not an implicit reviewer assumption.

02

Locator

Treat locator as explicit transaction state with evidence, a control outcome and ownership—not an implicit reviewer assumption.

03

Assertion lineage

Treat assertion lineage as explicit transaction state with evidence, a control outcome and ownership—not an implicit reviewer assumption.

04

Control result

Treat control result as explicit transaction state with evidence, a control outcome and ownership—not an implicit reviewer assumption.

05

Supersession

Treat supersession as explicit transaction state with evidence, a control outcome and ownership—not an implicit reviewer assumption.

How to evaluate it

  1. Freeze the population and workflow boundary.
  2. Define required evidence and successful exit state.
  3. Measure exceptions, human touches and silent errors.
  4. Separate modeled outcomes from production measurements.

CONTROLLED FINANCE AI

State → controls → resolution → authorized action.

A generic agent loop is not enough for finance. State must be explicit, controls must re-evaluate, human resolution must be bounded and execution must remain separately authorized.

NUMBER7 RESEARCH MODELCONTROLLED FINANCE AI
01STATE
02CONTROLS
03RESOLUTION
04AUTHORIZED ACTION
05VERIFICATION

AI proposes. Controls permit. The executor acts. Verification proves.

Read the full guide: State → controls → resolution → authorized action.

Why a generic agent loop is not enough

A common AI-agent pattern is plan, call tools, observe, and continue until the task appears complete. That is useful for many knowledge-work problems. Finance operations add requirements that the generic loop does not guarantee: economic-event identity, accounting treatment, policy, segregation of duties, materiality, action authority, duplicate-safe execution, and evidence retained for later review.

The architecture should therefore center the transaction, not the agent conversation. The system may use models and agents inside stages, but they should read and update governed state through explicit transitions.

The operating model

1. State

State is the current, versioned representation of the economic event and its evidence. It includes source lineage, proposed fields, entity candidates, supporting documents, accounting proposal, exception status, approvals, and target-system history.

State must distinguish:

  • observed source facts;
  • model/system proposals;
  • human-confirmed decisions;
  • control results;
  • authorized actions;
  • verified external outcomes.

Without those distinctions, later reviewers cannot tell what was known, inferred, decided, or executed.

2. Controls

Controls evaluate whether a proposed state transition is permitted. Examples include:

  • required fields and source lineage;
  • arithmetic consistency;
  • vendor/entity identity;
  • duplicate/economic-event checks;
  • PO/receipt/contract evidence;
  • coding and tax requirements;
  • value thresholds and approval routes;
  • segregation of duties;
  • connector and target-system constraints;
  • idempotency and retry policy.

A control returns a named result: pass, fail, not applicable, not evaluated, or insufficient evidence. A vague “confidence looks good” is not a control result.

3. Resolution

When a mandatory control does not pass, the system creates a typed exception. Resolution supplies the missing fact, decision, authorization, or evidence needed for the next evaluation.

The system should pre-assemble a decision packet and ask the smallest question possible. The person should not need to rediscover the entire transaction.

4. Action

An executor performs only the action allowed by the current state and authority. The action should carry a stable transaction ID, idempotency key where supported, actor/authority context, target object, and expected outcome.

Actions may include saving a draft, requesting evidence, sending an approval, creating a bill, or recording a payment state. These actions have different risk and authority requirements.

5. Verification

Verification proves what changed. It records the target response and, when needed, reads the target state back. A 200 response or outbound email log may not establish the business outcome.

Verification can close the transaction, create a new exception, or safely trigger a bounded retry.

The transition record

Every material transition should record:

FieldPurpose
Transaction and state versionConnects the event across time
Evidence referencesShows what the transition relied on
Proposal and producerSeparates model/system suggestion from source fact
Control and resultExplains why the action was or was not allowed
Resolution/actorRecords human or system decision and authority
Action and idempotency keyMakes execution traceable and duplicate-safe
Target response/object IDLinks to the downstream system
Verification resultProves the actual outcome
Timestamp and environmentSupports chronology and debugging

This is not a claim that every event needs the same storage or retention. It is the logical evidence required for controlled operation.

Worked example: QBO bill posting after an API timeout

  1. State: invoice is extracted, vendor resolved, coding reviewed, and external approval received.
  2. Controls: required evidence, arithmetic, duplicate candidates, and authority pass for the supported workflow.
  3. Action: executor sends a bill-creation request with a stable local transaction ID.
  4. Observation: the request times out. The local state is “target outcome unknown,” not “failed.”
  5. Resolution/control: retry policy requires checking the target system for a corresponding object before another create.
  6. Verification: target lookup finds the bill and stores its ID. Transaction moves to verified posted state.

If the system had treated timeout as absence and retried blindly, it could have created a duplicate. The verifier is not an optional logging step; it changes the safe next action.

Counterexample: human approval over an incomplete state

A reviewer receives an email containing only a PDF and “approve?” They click approve. The workflow now has a human-in-the-loop, but the person may not have seen duplicate candidates, receipt mismatch, coding, or approval limit.

Human presence did not create control. The decision packet, authority, and state completeness determine whether the approval is meaningful.

Control before autonomy

Autonomy should increase along two independent axes:

  • proposal autonomy: how much the system can interpret, assemble, and recommend;
  • action autonomy: which state transitions it can execute without a person.

A system can have high proposal autonomy and deliberately limited action autonomy. That is often the right early design for finance: let AI do the evidence-heavy preparation while controls and authority govern irreversible or material actions.

How to implement the model

  1. Define the transaction unit and state schema.
  2. Separate source facts, proposals, decisions, and outcomes.
  3. Name each permitted state transition.
  4. Attach explicit controls and authority to each transition.
  5. Type failure and uncertainty using the AP exception taxonomy.
  6. Design the smallest-decision packet for each recurring exception.
  7. Make external actions idempotent where possible.
  8. Verify the target state and handle unknown outcomes explicitly.
  9. Log events needed for touch, STP, error, and audit analysis.
  10. Version policies, models, prompts, connectors, and state changes.

Limitations

The model does not prescribe a specific database, orchestration framework, accounting policy, or legal control environment. Not every workflow requires a human resolution. Conversely, a typed state does not make an unsafe policy safe. Organizations must design controls with qualified finance, security, risk, and systems owners.

The phrase “verifier proves” means operational evidence of the defined target outcome, not mathematical proof or assurance engagement.

Sources

How to cite

Number7 Research. “Controlled AI in Finance: State, Controls, Resolution, Action.” Number7AI, revision 1.0, 7 September 2026. https://number7ai.com/research/state-controls-resolution-action.

Revision history

RevisionDateChangeReviewer
1.07 September 2026Initial state/control/resolution/action/verification model and exampleNumber7 AI editorial review

Next step

Choose one external action and document its permitted states, controls, authority, idempotency behavior, and verification proof.

Request the diagnostic