Implementing Agentic AI Ep 2: Choosing the First Process
Part 2 of 5. The impressive pilot freezes programmes; the boring one funds them. Four tests - busy, rule-shaped, measurable, survivable - turn the first-process debate into a scoring table, worked on a logistics company's finance floor.
Course materials
The Working Playbook — template
The five-row playbook template from the course, with the five pocket rules on the last page. Print it, fill in one row for one boring, busy, reversible process this month.
Sign in free to downloadTranscript
A regional insurer picked its first agent project the way you'd pick a
keynote demo: claims settlement. Customer-facing, high-stakes, maximum
applause if it worked.
It half-worked. Which for claims means: it failed. Three escalations to the
regulator later, the programme was frozen — including the four boring
projects behind it that would have succeeded.
Your first process decides whether there's a second. Today: how to choose
it.
Choosing the first process. Part two of five.
For anyone about to commit budget and reputation to their first agent
project — and for the engineers who'll be blamed if it's the wrong one.
In the next five minutes: the four tests a first process must pass, why
almost everyone's instinct is to pick the wrong one, and a scoring habit
that makes the choice boring — in the best possible way. We're in finance
today, with the process that quietly passes every test.
Down on the finance floor, Dana — who you may remember runs digital here —
is standing in front of accounts payable.
Nine thousand invoices a month. Three people matching them to purchase
orders, line by line. Roughly eighty percent sail through the same checks
every time; the rest have missing purchase orders, currency mismatches, duplicate
numbers.
Nobody will give a conference talk about invoice matching. Which is exactly
why it's about to pass all four tests.
Test one: volume. An agent learns nothing and earns nothing on a process
that happens twice a month. Practitioners' rule of thumb: hundreds of
transactions monthly before it's worth the plumbing. Nine thousand
invoices? Pass.
Test two: pattern. Does a clear rule cover most cases? Here, about eighty
percent match cleanly. Pass.
Same tests, other floors: in support, password resets and delivery queries
are busy and rule-shaped; a once-a-quarter crisis apology is neither. In
HR, leave requests pass; a sensitive grievance never will. In sales,
quote-to-CRM data entry passes; a bespoke enterprise negotiation won't.
Test three: measurable. Can you state today's baseline? Here, yes: hours
per thousand invoices, error rate, cost per invoice processed. If you can't
measure the before, you'll never prove the after — that lesson runs the
whole of episode five.
Test four — the one everyone skips: survivable errors. When the agent gets
one wrong, what happens? A mismatched invoice gets flagged and a human
looks. Annoying. Recoverable.
Now re-run that test on the insurer's claims pick: a wrong answer is a
customer harmed and a regulator phoned. Different universe.
So why did the insurer — and so many companies trying this — pick claims?
Because first projects get chosen for announcement value. The board wants
visible; visible means customer-facing; customer-facing means high-stakes;
and high-stakes fails test four by definition.
The correction is a sequencing rule, not a sacrifice: internal first,
reversible first. The boring win builds the skills, the evaluation checks, and the
trust that buy you the impressive project later — with a track record
under it.
Here's the habit that ends the argument in one meeting.
List your candidate processes. Score each test from one to five. Add them
up, in front of everyone.
Dana's shortlist: invoice matching scores eighteen of twenty. Supplier
onboarding: fourteen — pattern's weaker. The CFO's favourite, cash-flow
forecasting: nine — low volume, hard to measure, and an error lands in a
board pack.
The table doesn't care whose idea was whose. That's its whole value.
Pocket rule number two, for the collection:
First pick: boring, busy, and reversible.
Boring means nobody's reputation rides on the demo. Busy means the agent
meets enough cases to matter. Reversible means the day it's wrong — and
there will be one — a human catches it, fixes it, and the programme
survives.
The impressive project isn't cancelled. It's just not allowed to go first.
Before the recap, try it: pause here, take the process your team
complains about most, and score it against the four tests. Four passes
and you've found your pilot. Any fail — you've just dodged a year of
pain in ninety seconds.
Into the playbook it goes.
Process: invoice matching. Scores: four tests, eighteen of twenty.
Classification, from episode one's rule: workflow — we own the path, the
model works the stations. Owner: named — A. Rahman. Baseline: measured this week,
before anything ships.
Row two, stamped. Three to go — and the interesting thing is that nothing
about this page mentions technology yet. That changes next episode.
Episode two, in four lines.
Four tests: busy, rule-shaped, measurable, survivable — reversible, in the
pocket rule's words.
The instinct to resist: announcement value picks high-stakes losers.
Score candidates in the open — the table ends the politics.
And the same tests worked unchanged on support tickets and leave requests —
one method, every floor.
Next time: episode three — the mission brief.
We take the lift up to HR, where an onboarding assistant is about to be
given its instructions — and where one missing sentence in those
instructions nearly sends the wrong offer letter.
You'll learn the three-part brief every agent needs before it touches
anything — goal, boundaries, escalation — and meet the toolbox that turns
the brief into a working system.
Five minutes. See you in HR.