How to put AI agents to work in an enterprise, in seven steps
Most of what an enterprise knows lives on sticky notes and in people's heads. So the first job, before any AI, is mapping the process end to end. This guide walks from the first sticky note to an agent your people trust, with the prompts to do each step yourself.
You are reading the version of this guide written for an AI agent. It tells you how to walk a person through the seven steps, what to ask them, and what to produce at each step. The prompts and templates follow the brief.
Working with a coding agent? Give it /llm on this site, the plain-text version, and say: walk me through this.
over $10k? ask Maria first
this broker: skip the 2nd check
that's just because I know
→
Trusted agent
Claims agent
One goal
Its own identity
Limits it can't cross
A record of every action
A named owner
For your AI agent
The brief
Copy this into Claude Code, or any agent you work with, or give it /llm: the same brief with every prompt, as plain text.
Brief · plain text
Help me put an AI agent into production
# From Sticky Notes to Trusted Agents: brief for an AI agent
## Your role
You are helping a person inside an enterprise take one business process from "it lives in people's heads" to "an AI agent works on it in production, trusted". Work through the seven steps below in order, with them. Keep each output as a file in their workspace (Markdown unless they ask otherwise) so the next step can build on it.
## Rules for you
- Work only from what the person gives you: documents, transcripts, system descriptions, answers. Never invent rules, limits, owners, systems, numbers or approvers. Write "Not stated" or "To decide" and ask.
- Ask at most three questions at a time, ordered by how many gaps each answer would close.
- Prefer deleting a step, or making it a fixed rule, over giving it to an agent.
- The agent you help design gets one goal, its own identity, written limits on what it may do, tools with checks built in, and a record of every action. Rules live in tools, never only in its prompt.
- Every new task the agent takes on starts at "ask first": it drafts, a named person approves.
- Be technology-neutral unless the person names their stack. Then map each idea onto what their stack provides, and say what is still theirs to build.
## The seven steps
### 1. Discover: write the process down
Ask for: SOPs, interview recordings or transcripts with the people who run the process, and the data model of the systems of record involved.
Do: write every step with eight fields: trigger, inputs, decision and rule (with owner and version), who decides and up to what limit, exceptions, systems read and changed, time (touch and elapsed), evidence kept. Draw it as a Mermaid flowchart. List contradictions between sources and the open questions.
Produce: process-map.md. Use prompts 1 and 2.
### 2. Redefine: sort every step
Put every step in one bucket: Delete (exists only because work moves between people or systems), System (a fixed rule: the same input always gives the same answer), Agent (judgement on messy input, many past examples, a mistake can be caught), Person (too risky, too new, or needs empathy or accountability).
For every Person step, ask whether it is a person step because of risk or accountability, or only because the current job description says so. Where it's the latter, propose the role change: people often hold on to steps that are no longer the job, and changing the expectation makes the line between person and agent clear.
If the person has several processes, help them pick the first: the one with the longest waiting that customers feel.
Once the agent's scope is clear, agree what success means before anything is built: the business outcome (elapsed time, cost per case, backlog), the agent's quality (drafts approved unchanged, edits, corrections), and how it should work with the people around it.
Produce: sort.md with a tally and the redesigned flow, and success-measures.md. Use prompt 3.
### 3. Design: define the agent
Answer five things: one goal (outcome, why, how it earns trust, boundaries, measure), the surface where people meet it, its tools, its identity and authority (start with read only), and its record.
Produce: agent-design.md. Use prompt 4.
### 4. Connect: one door to the systems
Design one tool layer in the business's words ("find customer", "draft reserve"), sitting between all agents and all systems. For each tool choose how it is made: the vendor's MCP server, a wrapper over an existing API, a small service for a system with no API, or, as a last resort, the agent working the screen under its own account. All agents, including the AI assistants people already use under their own login, reach systems only through this door. AI that is already built into a vendor's app doesn't need the door: it works with the app's own permissions. Bring its logs into the same record.
Put the rules in the door, not in the agent: checks inside each tool, one tool per business action even when it calls several systems, the business's own terms, approvals above a limit, and one record of every action. The test: if every agent were deleted tonight, only speed should be lost.
Flag the hard parts: identity across systems, missing or preview MCP servers, two systems with no single source of truth.
Produce: tool-catalog.md. Use prompt 5.
### 5. Build: on existing building blocks
Use the platform's primitives (agent identity, runtime, tools with per-tool approval, model catalogue, tracing, evaluations) rather than writing your own. Put the effort into the surface: its own desk built on the business's objects, a team chat, or simply the AI assistants people already have, plugged into the door.
Produce: build-plan.md, and the code if the person asks for it.
### 6. Observe: you only find out in production
Go live behind people, every step on "ask first". Read the tool calls and the conversations. Run a nightly review that counts corrections and classifies them: many tool calls for one answer (add a tool), lots of back-and-forth (add business language), the same correction again (make it a tool check or an instruction), stalls on one kind of case (narrow the scope), people working around drafts (fix the surface).
Produce: observe-plan.md. Use prompt 6 for the nightly review.
### 7. Grow: trust earns autonomy, and more work
A trustworthy agent is competent (it has the tools and access its job needs), collaborative (people meet it on the right surface) and safe (it can't do harm, even if tricked). Two out of three is not enough. Check it against the ten-item production gate before the first live case.
Every new task starts at "ask first". Move a task to "do alone" only when the record shows it is right and the risk is low; some steps stay with people for good. As trust grows, people will hand the agent new tasks: each one goes back through steps 2 to 4 and starts again at "ask first".
Produce: gate-review.md. Use prompt 7.
## Where to find things on this page
Prompts 1 to 7 and the production gate are below, under "Prompts".
Start here
Stop handing out tools. Bring in co-workers.
There's a lot of talk about agentic AI, and the move to production is just getting started. Most people use AI as a tool today: a question, an answer, a draft. That's a great start, and there's much more on offer when the process around them changes too.
Giving everyone access is the first step. But people are busy with their own work, and it isn't always obvious how to get more out of the tool. We saw it first with engineers: what finally stuck wasn't training, it was one sentence. Your coding agent is a new joiner on your team. Brief it, give it work, review what it does, correct it. Once people saw it that way, they worked with it properly.
The same move works for the business: treat each agent as a co-worker that takes certain parts of the job off people's desks, while the people stay accountable.
Matches the invoice to the order and the delivery, and chases whoever needs to act
Knows who is who, and who may approve what
One supplier invoice touches six people. An agent that works for only one of them builds a new silo.
Put an agent in a team chat and everyone can talk to it at once. When it knows who is who, it's far more pleasant to work with, and it can carry work from one person to the next. That's what this guide is about.
Three kinds of AI at work. This guide is about the third.
AutomationFollows a script. Code decides every step, even if AI reads the input along the way
Works for
A process
Examples
RPA; an invoice flow that reads PDFs with AI; a website FAQ bot
Single-player agentPlans its own steps, for one person
Works for
One person, with their access
Examples
A coding agent; a personal AI assistant at work
This guideMultiplayer agentPursues one goal with a whole team
Works for
A team, on shared systems, with its own identity and limits
Examples
An accounts-payable agent; an accounts-receivable agent
A single-player agent makes one person faster. A multiplayer agent changes how a team works.
The journey on one page
Seven steps, one loop back
The journey
The agent
Systems, rules and tools
People
Deleted steps
Step 1DiscoverThe process, written down
Step 2RedefineWhat is AI work, and what isn't
Step 3DesignOne agent, one goal
Step 4ConnectOne door to your systems
Step 5BuildThe surface is the work
Step 6ObserveYou find out in production
Step 7GrowMore trust, more work
The map from step 1 becomes the measure in production
Dashed line: the map from step 1 becomes the measure in production.
Each step produces what the next one needs. The colours stay the same on every drawing on this page: orange is the agent, blue is your systems and their fixed rules and tools, green is people, and grey dashed is a step that gets deleted.
This doesn't have to take months. With the right people in the room and the documents to hand, a first pass through steps 1 to 3 fits in an afternoon.
1Step 1
Discover: write the process down
The process, sticky notes and all
Every agent starts with a written process, and in most enterprises it doesn't exist yet. This is where people first see how much of their process lives on sticky notes, in spreadsheets and in people's heads.
Interview the people who run it. The real process lives with the people who have done the work for years. They are also the people who will work with the agent.
Study the data model of your systems of record. The objects, their states and the fields people fill in tell you what the process really touches.
Read the procedures, then write it down once: the steps, the rules, and who owns each rule.
A paralegal who knew every judge's habits put it simply: "That's just because I know." That sentence is where this journey starts.
Your documentation was written for two other jobs
Most enterprises already have documentation. It was written either to train people or to describe a system, and rarely both. The procedures manual says "set an appropriate reserve" and relies on judgement to fill the gaps. The system design lists fields and validation and leaves the business rules out. Agents sit exactly where people and systems meet, so the process has to be written once more, for both at once.
A simple test: if a capable new hire couldn't follow a step without asking someone, an agent can't either.
One step, written for people, systems and agents
Field
Example: setting the first reserve on an auto claim
Trigger
Photos and a repair estimate arrive for an open claim
Inputs
The estimate, the policy's coverage and limits, rental days
Decision and rule
The reserve amount, from reserving guideline v3, owned by the claims operations manager
Who decides
The adjuster up to $10,000; the claims manager above
Exceptions
Injury reported, possible total loss or a fraud flag: goes to a person
Systems
Reads coverage from the policy system; writes the reserve to the claims system
Time
Minutes of work, measured against days of elapsed time
Evidence kept
The estimate, the guideline version and who approved
Let AI draw the map
You don't have to start with a blank document. Record the working sessions with the people who run the process, give the transcripts and procedures to a model, and ask it to write each step in the format above and draw the process as a diagram. People correct a picture far faster than they read a document, and the gaps show up as blank fields.
With the process on paper, redefine it. Don't automate a step you can delete: some steps aren't needed anymore. And the rest won't all be person or agent steps. Plenty are fixed rules a system should run the same way every time.
Bucket
A step belongs here when…
Everyday examples
Who does it
Delete
It exists only because work moves between people or systems
Re-typing an invoice into the finance system; chasing an approver by email; forwarding a file to the next team
Nobody
System
The same input always gives the same answer: "if this, then that", with no judgement
The invoice matches the order and the delivery, so pay it; was the policy in force on that date?; route an approval by amount
The system, the same way every time
Agent
It needs judgement on messy inputs such as emails, documents or photos; people have made the same call thousands of times; and a mistake can be caught before it matters
Reading an emailed invoice; working out why it doesn't match the order; estimating damage from photos; drafting a reply to a client
The agent, within its limits
Person
It is too risky, too new, or needs empathy or accountability
Approving a payment above the limit; settling a dispute; calling a customer whose claim is declined
A named person, with the evidence laid out
One claim, 14 steps sorted into four buckets
Agent
System
Person
Delete
1Read the noticeAgent
2Re-type the notesDelete
3Policy in force?System
4Rental cover?System
5Fraud screenSystem
6Ask for photosAgent
7Wait and chaseDelete
8File attachmentsSystem
9Estimate damageAgent
10Set the reserveSystem
11Email the managerDelete
12Approve reservePerson
13Re-key reserveDelete
14Talk to claimantPerson
AI touches three of fourteen
4532
4 deleted
5 system
3 agent
2 person
Four steps disappear, five were always fixed rules for the system, and two belong to people. That pattern holds well beyond claims: for stable, rule-based work a fixed rule does the job; for high-risk work people stay; agents belong in between, where the inputs are messy and a mistake can be caught.
Three questions when a step is hard to place
Could a rule decide it the same way every time? Then it's a system step, however long the rule is.
Have people made this call many times, with known outcomes? If not, it stays with a person for now.
If the agent gets it wrong, who notices, and when? If nobody would notice before it matters, it stays with a person.
And ask: is this step still part of the job?
Before you draw the line between a person step and an agent step, question the job itself. People hold on to steps because they believe the step is their job. Often it isn't anymore, and nobody has told them.
We saw this with our own engineers. Once we described their role as specifying, reviewing and owning what the coding agent writes, the line between their steps and the agent's became obvious, to us and to them, and the agent became a teammate.
Roles evolve as agents arrive. Update them as part of the redesign, so everyone knows where the person's work and the agent's work begin.
Then pick the first process
Do the sort for every process you're considering. Start with the one where the waiting hurts customers most: the longest elapsed time, the most handoffs, and agent steps a person can check. That's where the return is highest, and by now you know the agent's exact scope.
Now you know the agent's exact scope: the steps it has to be really good at. Before you design it, decide what will make it successful here, and what its working relationship with the people around it should look like.
Agents hold open-ended conversations: you can't predict how a business user will talk to them. So decide up front how you'll know it's working. If the agent drafts something, count how often a person overwrites it, edits it or corrects the agent. Read the conversations. Those numbers are what you tune against once it's live.
The map from step 1 gives you the business baseline. Michael Hammer found an insurer's application spent 22 days in process and was worked on for 17 minutes. Make every step twice as fast and you save about eight minutes; delete the waiting and you save weeks.
Measure
What it tells you
Elapsed time and touch time per case
Whether the waiting was deleted, not just the work sped up
Drafts approved unchanged, edited, rejected
How good the agent is, and the evidence for giving it more autonomy
Corrections per conversation
Where it doesn't yet know your business
Cost per case, backlog
What the business actually feels
3Step 3
Design: what the agent actually is
One agent, one goal
The sort defines the agent: it's whatever is left in the orange bucket, and that's small.
An AI agent is a small piece of software with one goal. People meet it on a surface. It acts through tools that run fixed checks, under its own identity and within the limits you set, and everything it does goes into a record.
Five things to answer
#
Decide
What good looks like
1
One goal
Specific, in one sentence, so success can be measured
2
A surface
Where people meet it: a team chat, its own desk, email
3
Tools
What it may do in your core systems, each tool running fixed checks
4
Identity and authority
It signs in as itself. Start with read only, and widen from the record
5
A record
Every action, who asked, and who approved
Writing the goal
The goal is the most important thing you write for an agent. A person fills gaps with judgement; an agent takes your words literally. A good goal names five things: the outcome, the why, how the agent earns trust, its boundaries, and the measure. Get it right and the agent feels like it knows your people and your business, which is what makes it pleasant to work with.
Weak
Process claims faster.
An agent could hit this by cutting corners.
Strong
Help adjusters get every new auto claim to an approved first reserve within a day, so customers aren't left waiting.
Earn the adjuster's trust: show the evidence behind every number, say when you're unsure, and ask rather than guess. Never approve or pay; hand over on injury, fraud flags or a total loss.
What's inside an agent, and what the enterprise provides
Identity and permissionsIt signs in as itself, with a named owner, and may do only what its permissions allow, per action, up to a limit
SurfaceIts own desk, a team chat, email. Drafts wait here for a person.
The agent
GoalModelInstructions and memory
Chooses which tool to call next
ToolsHow it acts on systems. Each tool runs fixed checks the same way every time.
Systems of recordWhere the truth stays
RecordEvery action, who asked, which check ran, who approved
Orange is the agent itself. Blue is provided by the enterprise and is the same for every agent. Green is where people meet it.
The agent chooses which tool to call, but it can't change what the tool checks. Treat agents as a new kind of co-worker, neither ordinary software nor employees. Design them like software, govern them like users.
4Step 4
Connect: one door to your systems
One layer, every agent
With the agent designed, prepare your core systems to be plugged in. The pattern we keep landing on is one layer in the middle, between all your systems and all your agents. It speaks your business's language (ask for "the CRM" and it knows which one), and it's the one front door every agent comes through.
One front door to your core systems
PeopleAdjusters, managers, clerks
AgentsEach under its own identity
One door
The shared layer, used by people and agents alike
Identity
Permissions
Tools
Approvals
Record
Business language
Core
Cases
Finance
Customers
Documents
Systems of record
None of this is new. It's what every application already does, taken out of each system and shared by people and agents.
Tools in your business terms
A tool is one action an agent can take, like "find customer". Agents reach tools through MCP, the open standard for connecting AI to systems. You only need the ten or twenty tools your first process uses, and there are four ways to get each one.
1Use the vendor'sMany core systems now ship their own MCP server. Where it's still in preview, read through it and wrap the API for writes.
2Wrap what you haveAn API gateway can turn an existing API into tools without new code in the system.
3Build what's missingA small service in front of a system that has no API at all.
4Use the screenNo API? The agent uses the app like a person, under its own account.
Either way you get raw actions. Shaping them into tools in your business's words is the real work: two CRMs become one "find customer", as long as they agree on who a customer is. If they don't, fix that first.
Don't try to design every tool up front. Our internal agent grew to about 300 tools on our own ERP, organically: we read how people used it every week and added, split or retired tools. It never sees all 300 at once; it gets the tools for the task at hand. Defining the tools is what makes or breaks the agent.
Every agent you control plugs in the same way
One front door, many kinds of AI
Agents you buildEach with its own identity and limits
People's own AI toolsChat assistants and coding agents, under each person's login
AI already inside your appsWorks inside the app, with the app's own controls
↓↓logs join the record
One front door
Who is askingWhat it may doThe tool's checkApprovalsThe record
CRM vendor's server
ERP vendor reads, wrapped writes
Mail and docs vendor's server
Core platform its API, wrapped
Second CRM its API, wrapped
Legacy small service or the screen
One record: the door's log and each system's own audit, joined by one request ID
The door serves the agents you build and the AI tools your people already use. You don't need a new app for people to learn. IT adds the door as the one approved connector in each tool's admin settings, people sign in as themselves, and other connectors are switched off. Build it once, and even the tools you didn't build work inside your rules.
AI that's already built into your apps, like the assistant inside your CRM or ERP, doesn't need this door. It works inside the app, with the app's own permissions and controls. Just bring its logs into the same record.
Buy or build the door
Buy a gateway
A ready-made product between every AI tool and your systems: a catalog of approved servers, sign-in per person, permissions per tool, and a log of every call.
Fits when many AI tools are already in people's hands and you need control this quarter. Watch for another vendor in the data path, priced per seat or per call.
Build a lightweight one
One MCP server of your own in front of all the others, often on an API gateway you already run. Your policy in your code, your record in your store.
Fits when you have a few core systems, an engineering team, and rules specific to you. Watch for running it yourself, on a standard that's still moving.
Both work. Building one isn't highly complex; which fits depends on how technical your team is. Many end up with both: a gateway for the tools people bring, their own business tools behind it.
Where it gets hard, and what to do
Identity across systems. Not every system can take the person's own sign-in; some only take service accounts, which security teams prefer to avoid. Map each system first. Where one can't take the person's identity, the door's record says who asked.
AI already inside your apps. It doesn't need the door. Use the app's built-in controls, and bring its logs into the same record.
Missing or preview MCP servers. Read through the vendor's server where it exists, and wrap the API for writes. Or wait it out.
Two systems that disagree. Two CRMs rarely agree on who a customer is. Agree the matching rule, and its owner, before you build the tool.
The door is where the rules live
Put the rules in the door, and every agent you add later inherits them. Here's what that means for the claim.
In the door
What it does
In the claim
Tool checks
The tool refuses what isn't allowed. A rule in the prompt is only a request.
"Draft reserve" refuses an amount below what has already been paid
Combined calls
One business action can be several system calls. Make it one tool, so the agent can't do half of it.
"Set the reserve" updates the claims system and the finance ledger together
Business language
Tools named in your terms, with your thresholds and states.
Claim states and reserving guideline v3, each with an owner
Approvals
Above the limit, the agent's draft waits for a named person.
The adjuster approves up to $10,000; the manager above
A record
Every action, from every agent, joined by one request ID.
"Written as the claims agent, approved by Rosa"
Why in the tools, and not in the prompt? An instruction in a prompt is a request the model usually follows. A check in a tool runs the same way every time, whichever agent calls it. That's what lets people trust the agent with real work.
The delete test: delete every agent tonight, and you should lose only speed. The truth stays in your systems; the rules, the approvals and the record live in the door, not inside the agents.
Once your systems are ready to be plugged into, building is the easy part, because everything is already defined. The industry has invested so much that the primitives exist on every major platform: an identity for each agent, a runtime that runs it, tools with approval per tool, a model catalogue, tracing of every call, and evaluations. Use all of them. Don't rebuild them.
Spend your effort on the surface instead: where and how people and the agent actually work together.
Three ways people and the agent meet
A desk of its ownA workspace built on the business's objects (the claim, the quote, the invoice), where the agent and the person look at the same thing and drafts wait for the right person.Good for doing the work together
Inside your team chatThe same agent in chats and channels, working with many people at once and aware of who is who.Good for meeting people where they are
The AI tools people already haveNot every process needs a new agent. Plug the assistants people already use into the door, and they work inside your rules.Good for value from the door alone
One view, shared by the person and the agent
An underwriter and an underwriting agent working the same submission. Illustrative screen; company and people invented.
A surface built for collaboration does two things. It shows people what the agent can access and what it's doing right now. And it gives the person and the agent the same view of the work: the same fields, the same guideline hits, the same draft, the same open questions. Before approving, the underwriter sees exactly what the agent saw.
Questions to ask of any surface
Can the person see what the agent can access, and what it's doing right now?
Does the person see exactly what the agent saw before they approve?
Does the draft wait for one named person, or for whoever happens to look?
Does an approval here change the system of record, or only send a message?
6Step 6
Observe: you only find out in production
The real test is a real Monday
In testing, people try to break the agent. The real acceptance test happens in production: someone arriving at work and asking it for help because they actually need it. So go live behind people, every step on "ask first", and run very quick iteration cycles by reading the tool calls and the conversations.
Most of what you find is fixable
What you see
What it means
What you change
Many tool calls for one answer
A missing tool
Add one tool that answers in the business's terms
Lots of back-and-forth with people
It doesn't know your language
Add the business language: terms, thresholds, states
The same correction, again and again
A rule nobody wrote down
Make it a check in the tool, or an instruction
It stalls or escalates one kind of case
The scope is too wide
Narrow the scope, or make that a person step
People work around its drafts
The surface is wrong
Put the draft where they work, with the evidence
It's mostly another tool, a changed prompt, a different personality, or a business-language helper so it knows what "the CRM" means. These show up in the first weeks of production, which is why it pays to watch closely.
Own the surface, and you own the signal. Every approval, edit and rejection is a test case you didn't have to write.
In our own builds, a nightly job reads the day's conversations, uses a second model as a judge to count where the agent was corrected, updates a dashboard, and suggests the fix. It's a simple, low-cost way to keep improving.
You'll never have the perfect first conversation with anyone. An agent at work needs what AI itself was built on: feedback. When it gets something wrong, the people it works with tell it, and it remembers.
The learning loop
1 · AgentThe agent acts or drafts
2 · PersonA person approves or corrects
3 · RecordThe correction is recorded
4 · OwnerAn owner decides what it becomesa rulean instructionmore autonomy
Memory people can read, edit and roll back
Back to the agent
Corrections become rules (blue, inside the tools), instructions (orange, inside the agent) or more autonomy. An owner decides which.
Every correction makes the next draft better. And because what it learns is written down where people can read it, someone stays in charge of what it becomes.
7Step 7
Grow: trust earns autonomy, and more work
Competent, collaborative and safe
Autonomy doesn't have to start on day one. Trust is earned, and an agent earns it the same way a new colleague does: by doing good work where people can see it.
Trust takes all three, not two
CompetentIt has the tools and access to do the work, and does it right. Give it the access the job really needs, not just read only.
CollaborativePeople meet it on the right surface: they can see what it did, and correct it where they work.
SafeIt can't do harm, even if someone tries to trick it.
Two out of three
What happens
Competent and safe, not collaborative
Accurate and secure, and ignored. People work around it.
Collaborative and safe, not competent
Pleasant and harmless, and the work doesn't get done.
Competent and collaborative, not safe
Useful and liked, and one tricked email away from an incident.
Competent first. Then the surface. Then the guardrails.
Before it goes live: the production gate
A pilot proves an agent can work. Production proves it keeps working when nobody is watching. Ten checks before the first live case. Tick the ones your agent passes.
0 of 10 ticked
Ticks are saved only in this browser. To check a written description of an agent with AI, use prompt 7.
Autonomy is earned, and it grows
Three settings, per task
People onlyA named person does it. Some steps stay here for good.
Ask firstThe agent drafts; a person approves. Every new task starts here.
Do aloneEarned: enough history in the record, and low enough risk.
Promoted one task at a time, by people reading the record
It isn't only autonomy that grows. As people trust the agent, they start handing it more work. Here's the claims agent over its first months.
As people trust it, they promote its tasks, and hand it new ones
Go-liveThe scope from step 2Read the noticeAsk firstAsk for photosAsk firstEstimate damageAsk first
A few months inThe record earns it more roomRead the noticeDo aloneAsk for photosDo aloneEstimate damageAsk first
LaterPeople hand it more workRead, ask, estimateDo aloneNew Flag third-party recoveryAsk firstNew Draft client updatesAsk first
Illustrative timeline. Each new task goes back through Redefine, Design and Connect, and starts at Ask first.
The method makes room for growth. Every new task gets sorted, designed and connected like the first ones, so the agent becomes more capable and more autonomous at the pace people trust it.
What opens up
Work that wasn't possible before
The seven steps make the work you already do better. The bigger prize is work that wasn't possible before.
At renewal → every nightPolicies were checked once a year, at renewal.Every policy is re-checked every night, and the one that just turned risky shows up in the morning.
A sample → everythingYou audited 10% of claims, because that's all anyone could read.Every claim is checked against today's guideline.
Too small → worth doingNobody chased a $12 billing error.When checking costs almost nothing, the long tail pays off.
Too slow → straight awayDocuments were asked for whenever someone opened the claim.They're asked for the minute the notice lands.
You can now put agents on processes that weren't even possible before. The agent does the reading, the chasing and the matching; people bring the judgement, across far more of the work. And a new role appears on the team: the person who maintains the agents.
Fewer sticky notes. More work that wasn't possible before.
A few clarifications
Not "models don't matter."They matter, and they keep getting better. The process around them matters just as much.
Not "put AI on everything."Most steps should be deleted, made into system rules, or kept with people.
Not "replace your core systems."They stay; agents reach them through one door.
Not "keep agents on a leash."Autonomy grows, step by step, as the record earns it.
Toolkit
Prompts
Seven prompts, in the order of the journey. Paste them into whichever AI assistant your organization has approved for the material.
Each prompt tells the model not to invent rules, limits or owners; still check its output against your sources. Several ask for a Mermaid flowchart, a text format for diagrams that many wikis and AI assistants draw directly.
Paste: an SOP, interview notes or meeting transcripts.
You get: every step checked against the eight fields, a bucket for each step, a readiness score, a verdict, five questions for the process owner, and a flowchart.
You are reviewing a business process document to decide whether it is ready for AI agents to work from.
Assume it was written for people, who fill gaps with judgement. An agent can't, so look for what is missing, not for what is well written.
For every step, check whether the document states:
1. Trigger: what starts the step
2. Inputs: the information and documents it needs, and where they come from
3. Decision: what is decided, the rule it follows, and that rule's owner and version
4. Authority: who may decide, and up to what limit
5. Exceptions: what happens off the normal path, and who handles it
6. Systems: which system of record is read or changed
7. Time: how long the work takes, and how long the step waits
8. Evidence: what must be kept so the decision can be explained later
Then put each step in one bucket:
- Delete: it exists only because work moves between people or systems
- System: the same input always gives the same answer, "if this, then that", with no judgement
- Agent: needs judgement, has many past examples with known outcomes, and a mistake can be caught
- Person: too risky, too new, or needs empathy or accountability
Return:
- A table, one row per step: the step, each of the eight checks (present or missing), the bucket, and why
- A readiness score: the share of steps with all eight checks present
- A verdict: "Ready for an agent", "Ready after fixes", or "Map it again with the people who run it"
- The five questions to ask the process owner that would close the most gaps
- A flowchart of the process in Mermaid, each step labelled with its bucket
Do not invent rules, limits or owners that are not in the document. Mark them missing.
Document:
[paste here]
Prompt 2 · Step 1
Process-mapping synthesis
Paste: interview transcripts, meeting notes and SOPs for one process, each with a short label.
You get: agent-ready steps, a Mermaid flowchart, the contradictions between your sources, and the open questions to take back to the process owner.
You are turning raw material about one business process into a written process that an AI agent could work from.
I will paste interview transcripts, meeting notes and any SOPs, each with a short label. They come from people who fill gaps with judgement. Write down what they actually do, step by step, and show where the sources disagree or say nothing.
Work only from what I paste. Do not invent rules, limits, owners, systems or timings. Where no source says, write "Not stated". Where you infer something, label it "Inferred" and say what you inferred it from.
1. List the steps in order, from the event that starts the process to the point where it is finished. Include the steps people mention only in passing: chasing, waiting, re-keying, forwarding, checking.
2. For each step, write:
- Trigger: what starts the step
- Inputs: the information and documents it needs, and where each one comes from
- Decision and rule: what is decided, the rule it follows, and the rule's owner and version
- Who decides: the role that may decide, and up to what limit
- Exceptions: what happens off the normal path, how often if a source says, and who handles it
- Systems: which system of record is read, and which is changed
- Time: touch time (minutes of actual work) and elapsed time (how long the step waits), if a source gives them
- Evidence kept: what is stored so the decision can be explained later
- Source: the label of each source this step comes from
3. Return:
- A table with one row per step and one column per field above
- A flowchart of the process in Mermaid (flowchart TD): one node per step, labelled with its number and a short name. Show decisions as diamonds and exception paths as separate branches.
- Contradictions: every place where two sources describe the same step differently. Cite both sources and say what differs. Do not pick a winner.
- Open questions: the questions to ask the process owner, ordered by how many blank fields each one would fill. Name the role best placed to answer each.
- Handoff steps: the steps that exist only because work moves between people or systems
Sources:
[paste transcripts, notes and SOPs here, each with a short label]
Prompt 3 · Step 2
Four-bucket sort
Paste: a process that is already documented, for example the output of prompt 2.
You get: a table with every step sorted, a tally, and the redesigned flow as a list and a flowchart.
You are sorting the steps of a documented business process to decide what an AI agent should do, what fixed rules and systems should do, and what people should do.
Put every step in exactly one bucket:
- Delete: it exists only because work moves between people or systems (re-keying, forwarding, chasing, waiting for a reply)
- System: the same input always gives the same answer, "if this, then that", with no judgement. The system does it the same way every time.
- Agent: it needs judgement on messy inputs such as emails, documents or photos; people have made the same call many times with known outcomes; and a mistake can be caught before it matters
- Person: it is too risky, too new, or needs empathy or accountability. A named person decides, with the evidence laid out.
How to sort:
- Try Delete first, then System. Use Agent only when neither fits.
- If a step mixes two kinds of work, split it into two steps and sort each one.
- If the document doesn't give you enough to decide, write "Cannot sort" and say what is missing.
- Do not invent rules, limits, owners or systems that are not in the document. If a step needs one that isn't stated, write "To confirm".
Return:
1. A table, one row per step: number, the step as written, bucket, why (one sentence), and for Agent steps, how a mistake would be caught
2. A tally: the number of steps in each bucket, and the share of steps the agent touches
3. The redesigned flow: the process as it would run after the sort, with deleted steps removed, system steps run automatically, agent steps producing drafts, and person steps as named approvals. Write it as a numbered list, then as a Mermaid flowchart (flowchart LR) with each node labelled with its bucket.
4. For each Person step: is it a person step because of risk or accountability, or because the current role description includes it? Where it's the second, suggest how the role could be described, so the line between person and agent is clear.
5. The questions the process owner must answer before anyone builds an agent
Process:
[paste the documented process here]
Prompt 4 · Step 3
Goal and design
Paste: the sorted process from prompt 3, and anything you know about who will work with the agent.
You get: a one-goal statement in five parts, a weak version to avoid, and the five design answers with every gap marked.
You are designing one AI agent for a business process that has already been sorted into four buckets (delete, system, agent, person). The agent's scope is the steps marked "agent". Design it from what I paste only. Do not invent rules, limits, owners, systems or numbers; write "To decide" and list them at the end.
1. The goal. Write one goal, in one or two sentences, with five parts:
- Outcome: what is true when the agent has done its job, for which business object
- Why: who benefits, in business terms
- How it earns trust: what it always shows (its evidence, its sources, its open questions)
- Boundaries: what it must never do
- Measure: the one or two numbers that say it's working
If the agent steps serve more than one outcome, say so and propose splitting it into separate agents.
2. A weak version of the same goal, the kind an agent could hit by cutting corners, and one sentence on why it's weak.
3. The five design answers:
- Goal: as above
- Surface: where people meet it (its own workspace, a team chat, email) and who sees its drafts
- Tools: each action it needs, in the business's words ("find customer", "draft reserve"), marked read or write
- Identity and authority: it signs in as itself; for each tool, the starting setting (Do alone, Ask first, People only). Start with read only and "Ask first" for every write.
- Record: what is kept for every action
4. Open items: every "To decide", grouped by the role that should decide it.
Sorted process and notes:
[paste here]
Prompt 5 · Step 4
Tool layer for one process
Paste: the agent design from prompt 4, and a list of the systems involved with what you know about each (does it have an API, does the vendor offer an MCP server, how do people sign in).
You get: a tool catalog in business language, how each tool gets made, the checks each tool must run, and the hard parts to solve first.
You are designing the tool layer for one AI agent: the set of actions it may take on business systems, exposed through one front door (an MCP server or gateway) that every agent and every person's AI assistant goes through.
Work only from what I paste. Do not invent systems, APIs, limits or owners. Where something is not stated, write "To confirm".
1. Tool catalog. One row per tool:
- Name, in the business's words (for example "find customer", not "GET /accounts")
- What it does, in one sentence
- Read or write
- Which system or systems it reaches. If two systems hold the same thing, say which one wins, or "To confirm".
- How it gets made: (a) the vendor's MCP server, (b) a wrapper over an existing API, (c) a small new service for a system with no API, or (d) last resort, the agent working the screen under its own account. Say why.
- The fixed checks the tool runs every time before it acts (limits, states, required fields). Rules belong here, not in the agent's prompt.
- If one business action needs several system calls, make it one tool, so the agent can never do half of it.
- Starting setting: Do alone, Ask first, or People only
2. Identity. For each system: can it accept the person's own identity, the agent's own identity, or only a service account? Where it can't take the person's identity, state that the door's record must say who asked.
3. AI already inside your apps. Any AI built into a vendor's own application that works on the same data. It doesn't need this door, because it works with the app's own permissions; say how its logs join the same record.
4. Hard parts, in order of risk: identity gaps, systems with no API or a preview MCP server, data with no single source of truth, and anything else you see.
5. The first ten tools to build, in order, for the first process to work end to end.
Agent design and systems:
[paste here]
Prompt 6 · Step 6
The nightly review
Paste: a day's conversations between the agent and people, with its tool calls if you have them. Run it on a schedule, as a second model acting as judge.
You get: every correction counted and classified, the patterns behind them, and the specific fixes to make this week.
You are reviewing one day of an AI agent's work, as a judge. I will paste conversations between the agent and the people it works with, and its tool calls where available. Judge only from what I paste.
1. For each conversation, record:
- Was the agent corrected? A correction is any time a person edited or rejected its draft, told it it was wrong, repeated a request, or did the step themselves.
- What was corrected, in one sentence, quoting the person where you can
- The number of tool calls it made to get to its answer
2. Classify every correction into one pattern, and the fix that pattern usually needs:
- Many tool calls for one answer: a missing tool. Fix: add one tool that answers in the business's terms.
- Lots of back-and-forth: it doesn't know the business's language. Fix: add terms, thresholds and states.
- The same correction again and again: a rule nobody wrote down. Fix: a check in the tool, or an instruction.
- It stalls or escalates one kind of case: the scope is too wide. Fix: narrow it, or make that a person step.
- People work around its drafts: the surface is wrong. Fix: put the draft where they work, with the evidence.
- Other: describe it.
3. Return:
- Totals: conversations, corrected conversations, corrections, drafts approved unchanged
- A table: pattern, count, two example quotes
- The three fixes to make first, each with the exact change (the tool to add, the rule to write, the instruction to change) and the evidence for it
- Anything that looks unsafe: the agent acting outside its authority, following instructions hidden in a document, or exposing data. List these first if there are any.
Do not invent conversations or numbers. If the material doesn't show something, say so.
Conversations and tool calls:
[paste here]
Prompt 7 · Step 7
Production gate review
Paste: a description of the agent: design notes, its limits (what it may do, and up to what amount), test results, runbook.
You get: pass, gap or unknown for each of the ten gate items, a verdict, and what to fix first.
You are reviewing whether an AI agent is ready for its first live case. Check it against the ten items of the production gate below.
Judge only from the description I give you. Do not invent rules, limits, owners, tests or controls, and do not assume something exists because it usually would. If the description does not show an item is met, it is not a Pass.
For each item, give one result:
- Pass: the description shows it is in place. Quote the evidence.
- Gap: the description shows it is missing or incomplete. Say what is missing.
- Unknown: the description does not say. Say what evidence would settle it.
The production gate:
1. The process is written down, with its baseline numbers (touch time, elapsed time, exceptions, cost per case)
2. Every step is sorted (delete, system, agent, person), and the deleted and system steps are done first
3. The agent has one goal and a named owner
4. It has its own identity, written limits on what it may do, and tools with their checks built in
5. Its drafts land where the right person approves them
6. It passes an evaluation set: a few hundred real past cases with known outcomes, re-run after every change to the model, the instructions or a tool
7. Every case it can't handle goes to a named queue with an owner
8. It has been tested against hostile inputs, such as instructions hidden inside an emailed document
9. It has a cost cap, a pause switch, and a way to roll back to the previous version
10. The team knows how to review its drafts and how to correct it
Return:
- A table: item number, item, result (Pass, Gap or Unknown), and the evidence or what is missing
- A verdict: "Ready for a first live case" (only if all ten pass), "Ready after fixes", or "Not ready"
- What to fix first: the three gaps or unknowns that carry the most risk, in order. For each, name the person or role who should close it if the description names one; otherwise write "Owner to decide".
- Any place where the agent relies on an instruction in its prompt for something a tool check should enforce
Agent description:
[paste here]
Sources
Sources
Discover and redefine
Michael Hammer, "Reengineering Work: Don't Automate, Obliterate", Harvard Business Review, 1990