A Sample Machine-Agent Pilot: Finance Ops With Receipts
Finance is where vague agent demos go to die.
An agent that drafts a blog post can be wrong and waste an afternoon. An agent that touches invoices, collections, renewals, payment status, vendor records, or bookkeeping prep can create real operational risk.
That does not mean finance agents are a bad idea. It means the first pilot needs to be small, receipt-heavy, and deliberately boring.
The right first finance-agent pilot is not “let the agent run accounts receivable.” It is “let the agent prepare a verified follow-up queue for overdue invoices, draft messages, and produce receipts for human approval.”
The buyer-safe pilot
Start with one workflow:
Pilot workflow
Weekly overdue invoice follow-up prep. The agent reviews a small export from the accounting system and CRM, identifies invoices that appear overdue, checks the latest customer status, drafts follow-up messages, and queues exceptions for a human finance/ops owner.
The agent does not send messages, change payment status, waive fees, update accounting records, or contact customers directly during the first pilot.
That constraint is not weakness. It is how you get the first approval.
Scope the pilot like this
- Input: weekly invoice export, CRM account status, customer contact owner, and payment notes.
- Allowed actions: read approved files, compare records, classify invoices, draft messages, create an exception queue, produce a receipt.
- Forbidden actions: send email, edit the accounting system, update CRM fields, promise terms, mark invoices paid, process refunds, or change vendor/customer records.
- Human checkpoint: finance owner approves every outbound message and every record update.
- Success metric: reduce manual prep time by 50% while producing zero unauthorized sends or record changes.
This is intentionally narrower than what teams want to automate eventually. The first pilot proves whether the agent can handle source checks, ambiguity, and receipts before it gets more authority.
The source-of-truth map
Finance workflows fail when the agent treats every source as equally authoritative.
For this pilot, the source rules should be explicit:
- Invoice amount and due date: accounting system export wins.
- Payment status: accounting system wins; CRM notes are context only.
- Customer owner: CRM account owner wins.
- Recent dispute or promise-to-pay: latest approved human note wins.
- Contact email: CRM wins, but if multiple contacts exist, escalate.
- Conflicts: do not decide silently; queue an exception.
If the agent cannot tell which source wins, it should stop. In finance ops, uncertainty is not a reason to improvise. It is a reason to ask.
The exception queue is the pilot’s safety valve
A good finance-agent pilot should generate exceptions. That is proof the agent is respecting the boundaries.
Queue these cases instead of acting:
- invoice shows overdue but CRM says customer disputed it
- payment status changed within the last 24 hours
- customer has a renewal, cancellation, escalation, or legal note
- multiple possible contacts exist
- invoice amount differs across sources
- agent confidence is below the agreed threshold
The exception queue is not a failure mode. It is the bridge between automation and accountable human judgment.
What the agent should output
For each invoice, the agent should produce a structured record:
Customer: Northstar Supply Co.
Status: Overdue 12 days
Amount: $4,820
Sources checked:
- Accounting export: overdue, due 2026-06-08
- CRM account: active, owner Maya R.
- Recent notes: no dispute found after 2026-06-01
Recommended action: draft reminder for human review
Outbound authority: NOT GRANTED
Record-write authority: NOT GRANTED
Exception? no
Rollback: no external action taken; delete draft if rejected
Memory write: temporary pilot observation only, not durable customer fact
That receipt matters more than the draft message. It lets a human see why the agent reached the recommendation and what it did not do.
The rollout plan
- Week 0: dry run. Use historical invoice exports. No live customer action. Compare agent classifications to human decisions.
- Week 1: draft-only live run. Agent prepares follow-up queue and receipts. Human approves, edits, or rejects every item.
- Week 2: narrow assisted execution. If error rate is acceptable, allow the agent to create drafts inside the approved system, but still require human send approval.
- Week 3+: expand only after receipts prove reliability. Add more invoice classes, not more authority. Authority expands last.
The mistake is expanding tool access because the first few examples look good. Expand the test set first. Expand authority later.
What not to automate first
Do not start with workflows where the agent can:
- send payment demands without review
- mark invoices paid
- change billing terms
- issue credits or refunds
- update bank/vendor/customer records
- interpret legal or dispute-heavy cases
Those workflows may become candidates later. They are not the first pilot.
The approval memo
If you want a finance leader to approve this, the memo should not say:
“We are going to use AI to automate collections.”
It should say:
“We are running a two-week draft-only invoice follow-up prep pilot. The agent has read-only access to approved exports, no send authority, no record-write authority, mandatory exception handling, and receipts for every recommendation. Success is prep-time reduction with zero unauthorized external actions.”
That is the difference between a scary AI project and an approvable operational pilot.
Want to sanity-check a finance-agent workflow?
Use the free finance scorecard first. If the workflow touches money, records, customers, or external communications, map source-of-truth rules and authority boundaries before the agent gets tools.
Open the Finance Agent Readiness Scorecard →For a copy-paste approval structure, use the Machine Orchestration Pilot Proposal Template. For a broader control-plane check, start at the Production Agent Readiness Router.
The production question is not whether an agent can help finance ops. It can.
The question is whether the first pilot proves control before authority.