A field guide

From Sticky Notes to Trusted Agents

How to put AI agents to work in an enterprise, in seven steps

Most of what an enterprise knows lives on sticky notes and in people's heads. So the first job, before any AI, is mapping the process end to end. This guide walks from the first sticky note to an agent your people trust, with the prompts to do each step yourself.

You are reading the version of this guide written for an AI agent. It tells you how to walk a person through the seven steps, what to ask them, and what to produce at each step. The prompts and templates follow the brief.

Working with a coding agent? Give it /llm on this site, the plain-text version, and say: walk me through this.

For your AI agent

The brief

Copy this into Claude Code, or any agent you work with, or give it /llm: the same brief with every prompt, as plain text.

Brief · plain text

Help me put an AI agent into production

# From Sticky Notes to Trusted Agents: brief for an AI agent

## Your role
You are helping a person inside an enterprise take one business process from "it lives in people's heads" to "an AI agent works on it in production, trusted". Work through the seven steps below in order, with them. Keep each output as a file in their workspace (Markdown unless they ask otherwise) so the next step can build on it.

## Rules for you
- Work only from what the person gives you: documents, transcripts, system descriptions, answers. Never invent rules, limits, owners, systems, numbers or approvers. Write "Not stated" or "To decide" and ask.
- Ask at most three questions at a time, ordered by how many gaps each answer would close.
- Prefer deleting a step, or making it a fixed rule, over giving it to an agent.
- The agent you help design gets one goal, its own identity, written limits on what it may do, tools with checks built in, and a record of every action. Rules live in tools, never only in its prompt.
- Every new task the agent takes on starts at "ask first": it drafts, a named person approves.
- Be technology-neutral unless the person names their stack. Then map each idea onto what their stack provides, and say what is still theirs to build.

## The seven steps

### 1. Discover: write the process down
Ask for: SOPs, interview recordings or transcripts with the people who run the process, and the data model of the systems of record involved.
Do: write every step with eight fields: trigger, inputs, decision and rule (with owner and version), who decides and up to what limit, exceptions, systems read and changed, time (touch and elapsed), evidence kept. Draw it as a Mermaid flowchart. List contradictions between sources and the open questions.
Produce: process-map.md. Use prompts 1 and 2.

### 2. Redefine: sort every step
Put every step in one bucket: Delete (exists only because work moves between people or systems), System (a fixed rule: the same input always gives the same answer), Agent (judgement on messy input, many past examples, a mistake can be caught), Person (too risky, too new, or needs empathy or accountability).
For every Person step, ask whether it is a person step because of risk or accountability, or only because the current job description says so. Where it's the latter, propose the role change: people often hold on to steps that are no longer the job, and changing the expectation makes the line between person and agent clear.
If the person has several processes, help them pick the first: the one with the longest waiting that customers feel.
Once the agent's scope is clear, agree what success means before anything is built: the business outcome (elapsed time, cost per case, backlog), the agent's quality (drafts approved unchanged, edits, corrections), and how it should work with the people around it.
Produce: sort.md with a tally and the redesigned flow, and success-measures.md. Use prompt 3.

### 3. Design: define the agent
Answer five things: one goal (outcome, why, how it earns trust, boundaries, measure), the surface where people meet it, its tools, its identity and authority (start with read only), and its record.
Produce: agent-design.md. Use prompt 4.

### 4. Connect: one door to the systems
Design one tool layer in the business's words ("find customer", "draft reserve"), sitting between all agents and all systems. For each tool choose how it is made: the vendor's MCP server, a wrapper over an existing API, a small service for a system with no API, or, as a last resort, the agent working the screen under its own account. All agents, including the AI assistants people already use under their own login, reach systems only through this door. AI that is already built into a vendor's app doesn't need the door: it works with the app's own permissions. Bring its logs into the same record.
Put the rules in the door, not in the agent: checks inside each tool, one tool per business action even when it calls several systems, the business's own terms, approvals above a limit, and one record of every action. The test: if every agent were deleted tonight, only speed should be lost.
Flag the hard parts: identity across systems, missing or preview MCP servers, two systems with no single source of truth.
Produce: tool-catalog.md. Use prompt 5.

### 5. Build: on existing building blocks
Use the platform's primitives (agent identity, runtime, tools with per-tool approval, model catalogue, tracing, evaluations) rather than writing your own. Put the effort into the surface: its own desk built on the business's objects, a team chat, or simply the AI assistants people already have, plugged into the door.
Produce: build-plan.md, and the code if the person asks for it.

### 6. Observe: you only find out in production
Go live behind people, every step on "ask first". Read the tool calls and the conversations. Run a nightly review that counts corrections and classifies them: many tool calls for one answer (add a tool), lots of back-and-forth (add business language), the same correction again (make it a tool check or an instruction), stalls on one kind of case (narrow the scope), people working around drafts (fix the surface).
Produce: observe-plan.md. Use prompt 6 for the nightly review.

### 7. Grow: trust earns autonomy, and more work
A trustworthy agent is competent (it has the tools and access its job needs), collaborative (people meet it on the right surface) and safe (it can't do harm, even if tricked). Two out of three is not enough. Check it against the ten-item production gate before the first live case.
Every new task starts at "ask first". Move a task to "do alone" only when the record shows it is right and the risk is low; some steps stay with people for good. As trust grows, people will hand the agent new tasks: each one goes back through steps 2 to 4 and starts again at "ask first".
Produce: gate-review.md. Use prompt 7.

## Where to find things on this page
Prompts 1 to 7 and the production gate are below, under "Prompts".

Start here

Stop handing out tools. Bring in co-workers.

There's a lot of talk about agentic AI, and the move to production is just getting started. Most people use AI as a tool today: a question, an answer, a draft. That's a great start, and there's much more on offer when the process around them changes too.

Giving everyone access is the first step. But people are busy with their own work, and it isn't always obvious how to get more out of the tool. We saw it first with engineers: what finally stuck wasn't training, it was one sentence. Your coding agent is a new joiner on your team. Brief it, give it work, review what it does, correct it. Once people saw it that way, they worked with it properly.

The same move works for the business: treat each agent as a co-worker that takes certain parts of the job off people's desks, while the people stay accountable.

The agents that matter are multiplayer

One person faster, or a whole process moving

Single-player

Works for one person
Priya, AP clerkher assistant
  • Logged in as Priya, with her access
  • Makes one person faster
  • Nobody else can see or approve what it did

Multiplayer

Works with a whole team
supplierbuyerreceivingAP clerkcontrollertreasuryaccounts-payable agent
  • Logged in as itself, with an owner and limits
  • Matches the invoice to the order and the delivery, and chases whoever needs to act
  • Knows who is who, and who may approve what
One supplier invoice touches six people. An agent that works for only one of them builds a new silo.

Put an agent in a team chat and everyone can talk to it at once. When it knows who is who, it's far more pleasant to work with, and it can carry work from one person to the next. That's what this guide is about.

Three kinds of AI at work. This guide is about the third.

Automation Follows a script. Code decides every step, even if AI reads the input along the way
Works for
A process
Examples
RPA; an invoice flow that reads PDFs with AI; a website FAQ bot
Single-player agent Plans its own steps, for one person
Works for
One person, with their access
Examples
A coding agent; a personal AI assistant at work
This guide Multiplayer agent Pursues one goal with a whole team
Works for
A team, on shared systems, with its own identity and limits
Examples
An accounts-payable agent; an accounts-receivable agent
A single-player agent makes one person faster. A multiplayer agent changes how a team works.

The journey on one page

Seven steps, one loop back

The journey

  • The agent
  • Systems, rules and tools
  • People
  • Deleted steps
  1. Step 1DiscoverThe process, written down
  2. Step 2RedefineWhat is AI work, and what isn't
  3. Step 3DesignOne agent, one goal
  4. Step 4ConnectOne door to your systems
  5. Step 5BuildThe surface is the work
  6. Step 6ObserveYou find out in production
  7. Step 7GrowMore trust, more work

Dashed line: the map from step 1 becomes the measure in production.

Each step produces what the next one needs. The colours stay the same on every drawing on this page: orange is the agent, blue is your systems and their fixed rules and tools, green is people, and grey dashed is a step that gets deleted.

This doesn't have to take months. With the right people in the room and the documents to hand, a first pass through steps 1 to 3 fits in an afternoon.

1Step 1

Discover: write the process down

The process, sticky notes and all

Every agent starts with a written process, and in most enterprises it doesn't exist yet. This is where people first see how much of their process lives on sticky notes, in spreadsheets and in people's heads.

  • Interview the people who run it. The real process lives with the people who have done the work for years. They are also the people who will work with the agent.
  • Study the data model of your systems of record. The objects, their states and the fields people fill in tell you what the process really touches.
  • Read the procedures, then write it down once: the steps, the rules, and who owns each rule.

A paralegal who knew every judge's habits put it simply: "That's just because I know." That sentence is where this journey starts.

Your documentation was written for two other jobs

Most enterprises already have documentation. It was written either to train people or to describe a system, and rarely both. The procedures manual says "set an appropriate reserve" and relies on judgement to fill the gaps. The system design lists fields and validation and leaves the business rules out. Agents sit exactly where people and systems meet, so the process has to be written once more, for both at once.

A simple test: if a capable new hire couldn't follow a step without asking someone, an agent can't either.

One step, written for people, systems and agents

FieldExample: setting the first reserve on an auto claim
TriggerPhotos and a repair estimate arrive for an open claim
InputsThe estimate, the policy's coverage and limits, rental days
Decision and ruleThe reserve amount, from reserving guideline v3, owned by the claims operations manager
Who decidesThe adjuster up to $10,000; the claims manager above
ExceptionsInjury reported, possible total loss or a fraud flag: goes to a person
SystemsReads coverage from the policy system; writes the reserve to the claims system
TimeMinutes of work, measured against days of elapsed time
Evidence keptThe estimate, the guideline version and who approved

Let AI draw the map

You don't have to start with a blank document. Record the working sessions with the people who run the process, give the transcripts and procedures to a model, and ask it to write each step in the format above and draw the process as a diagram. People correct a picture far faster than they read a document, and the gaps show up as blank fields.

2Step 2

Redefine: what is AI work, and what isn't

Four buckets for every step

With the process on paper, redefine it. Don't automate a step you can delete: some steps aren't needed anymore. And the rest won't all be person or agent steps. Plenty are fixed rules a system should run the same way every time.

BucketA step belongs here when…Everyday examplesWho does it
DeleteIt exists only because work moves between people or systemsRe-typing an invoice into the finance system; chasing an approver by email; forwarding a file to the next teamNobody
SystemThe same input always gives the same answer: "if this, then that", with no judgementThe invoice matches the order and the delivery, so pay it; was the policy in force on that date?; route an approval by amountThe system, the same way every time
AgentIt needs judgement on messy inputs such as emails, documents or photos; people have made the same call thousands of times; and a mistake can be caught before it mattersReading an emailed invoice; working out why it doesn't match the order; estimating damage from photos; drafting a reply to a clientThe agent, within its limits
PersonIt is too risky, too new, or needs empathy or accountabilityApproving a payment above the limit; settling a dispute; calling a customer whose claim is declinedA named person, with the evidence laid out

One claim, 14 steps sorted into four buckets

  • Agent
  • System
  • Person
  • Delete
  1. 1Read the noticeAgent
  2. 2Re-type the notesDelete
  3. 3Policy in force?System
  4. 4Rental cover?System
  5. 5Fraud screenSystem
  6. 6Ask for photosAgent
  7. 7Wait and chaseDelete
  8. 8File attachmentsSystem
  9. 9Estimate damageAgent
  10. 10Set the reserveSystem
  11. 11Email the managerDelete
  12. 12Approve reservePerson
  13. 13Re-key reserveDelete
  14. 14Talk to claimantPerson

AI touches three of fourteen

  • 4 deleted
  • 5 system
  • 3 agent
  • 2 person

Four steps disappear, five were always fixed rules for the system, and two belong to people. That pattern holds well beyond claims: for stable, rule-based work a fixed rule does the job; for high-risk work people stay; agents belong in between, where the inputs are messy and a mistake can be caught.

Three questions when a step is hard to place

  • Could a rule decide it the same way every time? Then it's a system step, however long the rule is.
  • Have people made this call many times, with known outcomes? If not, it stays with a person for now.
  • If the agent gets it wrong, who notices, and when? If nobody would notice before it matters, it stays with a person.

And ask: is this step still part of the job?

Before you draw the line between a person step and an agent step, question the job itself. People hold on to steps because they believe the step is their job. Often it isn't anymore, and nobody has told them.

We saw this with our own engineers. Once we described their role as specifying, reviewing and owning what the coding agent writes, the line between their steps and the agent's became obvious, to us and to them, and the agent became a teammate.

Roles evolve as agents arrive. Update them as part of the redesign, so everyone knows where the person's work and the agent's work begin.

Then pick the first process

Do the sort for every process you're considering. Start with the one where the waiting hurts customers most: the longest elapsed time, the most handoffs, and agent steps a person can check. That's where the return is highest, and by now you know the agent's exact scope.

Now define what success looks like

Now you know the agent's exact scope: the steps it has to be really good at. Before you design it, decide what will make it successful here, and what its working relationship with the people around it should look like.

Agents hold open-ended conversations: you can't predict how a business user will talk to them. So decide up front how you'll know it's working. If the agent drafts something, count how often a person overwrites it, edits it or corrects the agent. Read the conversations. Those numbers are what you tune against once it's live.

The map from step 1 gives you the business baseline. Michael Hammer found an insurer's application spent 22 days in process and was worked on for 17 minutes. Make every step twice as fast and you save about eight minutes; delete the waiting and you save weeks.

MeasureWhat it tells you
Elapsed time and touch time per caseWhether the waiting was deleted, not just the work sped up
Drafts approved unchanged, edited, rejectedHow good the agent is, and the evidence for giving it more autonomy
Corrections per conversationWhere it doesn't yet know your business
Cost per case, backlogWhat the business actually feels

3Step 3

Design: what the agent actually is

One agent, one goal

The sort defines the agent: it's whatever is left in the orange bucket, and that's small.

An AI agent is a small piece of software with one goal. People meet it on a surface. It acts through tools that run fixed checks, under its own identity and within the limits you set, and everything it does goes into a record.

Five things to answer

#DecideWhat good looks like
1One goalSpecific, in one sentence, so success can be measured
2A surfaceWhere people meet it: a team chat, its own desk, email
3ToolsWhat it may do in your core systems, each tool running fixed checks
4Identity and authorityIt signs in as itself. Start with read only, and widen from the record
5A recordEvery action, who asked, and who approved

Writing the goal

The goal is the most important thing you write for an agent. A person fills gaps with judgement; an agent takes your words literally. A good goal names five things: the outcome, the why, how the agent earns trust, its boundaries, and the measure. Get it right and the agent feels like it knows your people and your business, which is what makes it pleasant to work with.

Weak

Process claims faster.

An agent could hit this by cutting corners.

Strong

Help adjusters get every new auto claim to an approved first reserve within a day, so customers aren't left waiting.

Earn the adjuster's trust: show the evidence behind every number, say when you're unsure, and ask rather than guess. Never approve or pay; hand over on injury, fraud flags or a total loss.

The anatomy of an agent

What's inside an agent, and what the enterprise provides

Identity and permissionsIt signs in as itself, with a named owner, and may do only what its permissions allow, per action, up to a limit
SurfaceIts own desk, a team chat, email. Drafts wait here for a person.
The agent
GoalModelInstructions and memory
Chooses which tool to call next
ToolsHow it acts on systems. Each tool runs fixed checks the same way every time.
Systems of recordWhere the truth stays
RecordEvery action, who asked, which check ran, who approved
Orange is the agent itself. Blue is provided by the enterprise and is the same for every agent. Green is where people meet it.

The agent chooses which tool to call, but it can't change what the tool checks. Treat agents as a new kind of co-worker, neither ordinary software nor employees. Design them like software, govern them like users.

4Step 4

Connect: one door to your systems

One layer, every agent

With the agent designed, prepare your core systems to be plugged in. The pattern we keep landing on is one layer in the middle, between all your systems and all your agents. It speaks your business's language (ask for "the CRM" and it knows which one), and it's the one front door every agent comes through.

One front door to your core systems

PeopleAdjusters, managers, clerks
AgentsEach under its own identity
One door

The shared layer, used by people and agents alike

  • Identity
  • Permissions
  • Tools
  • Approvals
  • Record
  • Business language
  • Core
  • Cases
  • Finance
  • Customers
  • Documents

Systems of record

None of this is new. It's what every application already does, taken out of each system and shared by people and agents.

Tools in your business terms

A tool is one action an agent can take, like "find customer". Agents reach tools through MCP, the open standard for connecting AI to systems. You only need the ten or twenty tools your first process uses, and there are four ways to get each one.

1Use the vendor'sMany core systems now ship their own MCP server. Where it's still in preview, read through it and wrap the API for writes.
2Wrap what you haveAn API gateway can turn an existing API into tools without new code in the system.
3Build what's missingA small service in front of a system that has no API at all.
4Use the screenNo API? The agent uses the app like a person, under its own account.

Either way you get raw actions. Shaping them into tools in your business's words is the real work: two CRMs become one "find customer", as long as they agree on who a customer is. If they don't, fix that first.

Don't try to design every tool up front. Our internal agent grew to about 300 tools on our own ERP, organically: we read how people used it every week and added, split or retired tools. It never sees all 300 at once; it gets the tools for the task at hand. Defining the tools is what makes or breaks the agent.

Every agent you control plugs in the same way

One front door, many kinds of AI

Agents you buildEach with its own identity and limits
People's own AI toolsChat assistants and coding agents, under each person's login
AI already inside your appsWorks inside the app, with the app's own controls
One front door
Who is askingWhat it may doThe tool's checkApprovalsThe record
CRM
vendor's server
ERP
vendor reads, wrapped writes
Mail and docs
vendor's server
Core platform
its API, wrapped
Second CRM
its API, wrapped
Legacy
small service or the screen
One record: the door's log and each system's own audit, joined by one request ID

The door serves the agents you build and the AI tools your people already use. You don't need a new app for people to learn. IT adds the door as the one approved connector in each tool's admin settings, people sign in as themselves, and other connectors are switched off. Build it once, and even the tools you didn't build work inside your rules.

AI that's already built into your apps, like the assistant inside your CRM or ERP, doesn't need this door. It works inside the app, with the app's own permissions and controls. Just bring its logs into the same record.

Buy or build the door

Buy a gateway

A ready-made product between every AI tool and your systems: a catalog of approved servers, sign-in per person, permissions per tool, and a log of every call.

Fits when many AI tools are already in people's hands and you need control this quarter. Watch for another vendor in the data path, priced per seat or per call.

Build a lightweight one

One MCP server of your own in front of all the others, often on an API gateway you already run. Your policy in your code, your record in your store.

Fits when you have a few core systems, an engineering team, and rules specific to you. Watch for running it yourself, on a standard that's still moving.

Both work. Building one isn't highly complex; which fits depends on how technical your team is. Many end up with both: a gateway for the tools people bring, their own business tools behind it.

Where it gets hard, and what to do

  • Identity across systems. Not every system can take the person's own sign-in; some only take service accounts, which security teams prefer to avoid. Map each system first. Where one can't take the person's identity, the door's record says who asked.
  • AI already inside your apps. It doesn't need the door. Use the app's built-in controls, and bring its logs into the same record.
  • Missing or preview MCP servers. Read through the vendor's server where it exists, and wrap the API for writes. Or wait it out.
  • Two systems that disagree. Two CRMs rarely agree on who a customer is. Agree the matching rule, and its owner, before you build the tool.

The door is where the rules live

Put the rules in the door, and every agent you add later inherits them. Here's what that means for the claim.

In the doorWhat it doesIn the claim
Tool checksThe tool refuses what isn't allowed. A rule in the prompt is only a request."Draft reserve" refuses an amount below what has already been paid
Combined callsOne business action can be several system calls. Make it one tool, so the agent can't do half of it."Set the reserve" updates the claims system and the finance ledger together
Business languageTools named in your terms, with your thresholds and states.Claim states and reserving guideline v3, each with an owner
ApprovalsAbove the limit, the agent's draft waits for a named person.The adjuster approves up to $10,000; the manager above
A recordEvery action, from every agent, joined by one request ID."Written as the claims agent, approved by Rosa"

Why in the tools, and not in the prompt? An instruction in a prompt is a request the model usually follows. A check in a tool runs the same way every time, whichever agent calls it. That's what lets people trust the agent with real work.

The delete test: delete every agent tonight, and you should lose only speed. The truth stays in your systems; the rules, the approvals and the record live in the door, not inside the agents.

5Step 5

Build: the easy part

Spend your effort on the surface

Once your systems are ready to be plugged into, building is the easy part, because everything is already defined. The industry has invested so much that the primitives exist on every major platform: an identity for each agent, a runtime that runs it, tools with approval per tool, a model catalogue, tracing of every call, and evaluations. Use all of them. Don't rebuild them.

Spend your effort on the surface instead: where and how people and the agent actually work together.

Three ways people and the agent meet

A desk of its ownA workspace built on the business's objects (the claim, the quote, the invoice), where the agent and the person look at the same thing and drafts wait for the right person.Good for doing the work together
Inside your team chatThe same agent in chats and channels, working with many people at once and aware of who is who.Good for meeting people where they are
The AI tools people already haveNot every process needs a new agent. Plug the assistants people already use into the door, and they work inside your rules.Good for value from the door alone

One view, shared by the person and the agent

A workbench screen for underwriting a commercial auto submission. Left: the underwriter's queue. Centre: what the agent read, each figure with its source; guideline hits marked refer, check and pass; and a draft quote waiting for approval. Right: the agent's panel, showing what it may read and draft and that it may never bind a policy, and a conversation where it asks two questions, the underwriter answers, and the action is logged.
An underwriter and an underwriting agent working the same submission. Illustrative screen; company and people invented.

A surface built for collaboration does two things. It shows people what the agent can access and what it's doing right now. And it gives the person and the agent the same view of the work: the same fields, the same guideline hits, the same draft, the same open questions. Before approving, the underwriter sees exactly what the agent saw.

Questions to ask of any surface

  • Can the person see what the agent can access, and what it's doing right now?
  • Does the person see exactly what the agent saw before they approve?
  • Does the draft wait for one named person, or for whoever happens to look?
  • Does an approval here change the system of record, or only send a message?

6Step 6

Observe: you only find out in production

The real test is a real Monday

In testing, people try to break the agent. The real acceptance test happens in production: someone arriving at work and asking it for help because they actually need it. So go live behind people, every step on "ask first", and run very quick iteration cycles by reading the tool calls and the conversations.

Most of what you find is fixable

What you seeWhat it meansWhat you change
Many tool calls for one answerA missing toolAdd one tool that answers in the business's terms
Lots of back-and-forth with peopleIt doesn't know your languageAdd the business language: terms, thresholds, states
The same correction, again and againA rule nobody wrote downMake it a check in the tool, or an instruction
It stalls or escalates one kind of caseThe scope is too wideNarrow the scope, or make that a person step
People work around its draftsThe surface is wrongPut the draft where they work, with the evidence

It's mostly another tool, a changed prompt, a different personality, or a business-language helper so it knows what "the CRM" means. These show up in the first weeks of production, which is why it pays to watch closely.

Own the surface, and you own the signal. Every approval, edit and rejection is a test case you didn't have to write.

In our own builds, a nightly job reads the day's conversations, uses a second model as a judge to count where the agent was corrected, updates a dashboard, and suggests the fix. It's a simple, low-cost way to keep improving.

No learning loop, no lasting agent

You'll never have the perfect first conversation with anyone. An agent at work needs what AI itself was built on: feedback. When it gets something wrong, the people it works with tell it, and it remembers.

The learning loop

1 · AgentThe agent acts or drafts
2 · PersonA person approves or corrects
3 · RecordThe correction is recorded
4 · OwnerAn owner decides what it becomes a rulean instructionmore autonomy
Memory people can read, edit and roll back
Back to the agent
Corrections become rules (blue, inside the tools), instructions (orange, inside the agent) or more autonomy. An owner decides which.

Every correction makes the next draft better. And because what it learns is written down where people can read it, someone stays in charge of what it becomes.

7Step 7

Grow: trust earns autonomy, and more work

Competent, collaborative and safe

Autonomy doesn't have to start on day one. Trust is earned, and an agent earns it the same way a new colleague does: by doing good work where people can see it.

Trust takes all three, not two

CompetentIt has the tools and access to do the work, and does it right. Give it the access the job really needs, not just read only.
CollaborativePeople meet it on the right surface: they can see what it did, and correct it where they work.
SafeIt can't do harm, even if someone tries to trick it.
Two out of threeWhat happens
Competent and safe, not collaborativeAccurate and secure, and ignored. People work around it.
Collaborative and safe, not competentPleasant and harmless, and the work doesn't get done.
Competent and collaborative, not safeUseful and liked, and one tricked email away from an incident.

Competent first. Then the surface. Then the guardrails.

Before it goes live: the production gate

A pilot proves an agent can work. Production proves it keeps working when nobody is watching. Ten checks before the first live case. Tick the ones your agent passes.

0 of 10 ticked

Ticks are saved only in this browser. To check a written description of an agent with AI, use prompt 7.

Autonomy is earned, and it grows

Three settings, per task

  • People onlyA named person does it. Some steps stay here for good.
  • Ask firstThe agent drafts; a person approves. Every new task starts here.
  • Do aloneEarned: enough history in the record, and low enough risk.

Promoted one task at a time, by people reading the record

It isn't only autonomy that grows. As people trust the agent, they start handing it more work. Here's the claims agent over its first months.

As people trust it, they promote its tasks, and hand it new ones

Go-liveThe scope from step 2 Read the noticeAsk first Ask for photosAsk first Estimate damageAsk first
A few months inThe record earns it more room Read the noticeDo alone Ask for photosDo alone Estimate damageAsk first
LaterPeople hand it more work Read, ask, estimateDo alone New Flag third-party recoveryAsk first New Draft client updatesAsk first
Illustrative timeline. Each new task goes back through Redefine, Design and Connect, and starts at Ask first.

The method makes room for growth. Every new task gets sorted, designed and connected like the first ones, so the agent becomes more capable and more autonomous at the pace people trust it.

What opens up

Work that wasn't possible before

The seven steps make the work you already do better. The bigger prize is work that wasn't possible before.

At renewal → every nightPolicies were checked once a year, at renewal.Every policy is re-checked every night, and the one that just turned risky shows up in the morning.
A sample → everythingYou audited 10% of claims, because that's all anyone could read.Every claim is checked against today's guideline.
Too small → worth doingNobody chased a $12 billing error.When checking costs almost nothing, the long tail pays off.
Too slow → straight awayDocuments were asked for whenever someone opened the claim.They're asked for the minute the notice lands.

You can now put agents on processes that weren't even possible before. The agent does the reading, the chasing and the matching; people bring the judgement, across far more of the work. And a new role appears on the team: the person who maintains the agents.

Fewer sticky notes. More work that wasn't possible before.

A few clarifications

  • Not "models don't matter."They matter, and they keep getting better. The process around them matters just as much.
  • Not "put AI on everything."Most steps should be deleted, made into system rules, or kept with people.
  • Not "replace your core systems."They stay; agents reach them through one door.
  • Not "keep agents on a leash."Autonomy grows, step by step, as the record earns it.

Toolkit

Prompts

Seven prompts, in the order of the journey. Paste them into whichever AI assistant your organization has approved for the material.

Each prompt tells the model not to invent rules, limits or owners; still check its output against your sources. Several ask for a Mermaid flowchart, a text format for diagrams that many wikis and AI assistants draw directly.

Prompt 1 · Step 1

Documentation readiness check

Paste: an SOP, interview notes or meeting transcripts.

You get: every step checked against the eight fields, a bucket for each step, a readiness score, a verdict, five questions for the process owner, and a flowchart.

You are reviewing a business process document to decide whether it is ready for AI agents to work from.

Assume it was written for people, who fill gaps with judgement. An agent can't, so look for what is missing, not for what is well written.

For every step, check whether the document states:
1. Trigger: what starts the step
2. Inputs: the information and documents it needs, and where they come from
3. Decision: what is decided, the rule it follows, and that rule's owner and version
4. Authority: who may decide, and up to what limit
5. Exceptions: what happens off the normal path, and who handles it
6. Systems: which system of record is read or changed
7. Time: how long the work takes, and how long the step waits
8. Evidence: what must be kept so the decision can be explained later

Then put each step in one bucket:
- Delete: it exists only because work moves between people or systems
- System: the same input always gives the same answer, "if this, then that", with no judgement
- Agent: needs judgement, has many past examples with known outcomes, and a mistake can be caught
- Person: too risky, too new, or needs empathy or accountability

Return:
- A table, one row per step: the step, each of the eight checks (present or missing), the bucket, and why
- A readiness score: the share of steps with all eight checks present
- A verdict: "Ready for an agent", "Ready after fixes", or "Map it again with the people who run it"
- The five questions to ask the process owner that would close the most gaps
- A flowchart of the process in Mermaid, each step labelled with its bucket

Do not invent rules, limits or owners that are not in the document. Mark them missing.

Document:
[paste here]

Prompt 2 · Step 1

Process-mapping synthesis

Paste: interview transcripts, meeting notes and SOPs for one process, each with a short label.

You get: agent-ready steps, a Mermaid flowchart, the contradictions between your sources, and the open questions to take back to the process owner.

You are turning raw material about one business process into a written process that an AI agent could work from.

I will paste interview transcripts, meeting notes and any SOPs, each with a short label. They come from people who fill gaps with judgement. Write down what they actually do, step by step, and show where the sources disagree or say nothing.

Work only from what I paste. Do not invent rules, limits, owners, systems or timings. Where no source says, write "Not stated". Where you infer something, label it "Inferred" and say what you inferred it from.

1. List the steps in order, from the event that starts the process to the point where it is finished. Include the steps people mention only in passing: chasing, waiting, re-keying, forwarding, checking.

2. For each step, write:
- Trigger: what starts the step
- Inputs: the information and documents it needs, and where each one comes from
- Decision and rule: what is decided, the rule it follows, and the rule's owner and version
- Who decides: the role that may decide, and up to what limit
- Exceptions: what happens off the normal path, how often if a source says, and who handles it
- Systems: which system of record is read, and which is changed
- Time: touch time (minutes of actual work) and elapsed time (how long the step waits), if a source gives them
- Evidence kept: what is stored so the decision can be explained later
- Source: the label of each source this step comes from

3. Return:
- A table with one row per step and one column per field above
- A flowchart of the process in Mermaid (flowchart TD): one node per step, labelled with its number and a short name. Show decisions as diamonds and exception paths as separate branches.
- Contradictions: every place where two sources describe the same step differently. Cite both sources and say what differs. Do not pick a winner.
- Open questions: the questions to ask the process owner, ordered by how many blank fields each one would fill. Name the role best placed to answer each.
- Handoff steps: the steps that exist only because work moves between people or systems

Sources:
[paste transcripts, notes and SOPs here, each with a short label]

Prompt 3 · Step 2

Four-bucket sort

Paste: a process that is already documented, for example the output of prompt 2.

You get: a table with every step sorted, a tally, and the redesigned flow as a list and a flowchart.

You are sorting the steps of a documented business process to decide what an AI agent should do, what fixed rules and systems should do, and what people should do.

Put every step in exactly one bucket:
- Delete: it exists only because work moves between people or systems (re-keying, forwarding, chasing, waiting for a reply)
- System: the same input always gives the same answer, "if this, then that", with no judgement. The system does it the same way every time.
- Agent: it needs judgement on messy inputs such as emails, documents or photos; people have made the same call many times with known outcomes; and a mistake can be caught before it matters
- Person: it is too risky, too new, or needs empathy or accountability. A named person decides, with the evidence laid out.

How to sort:
- Try Delete first, then System. Use Agent only when neither fits.
- If a step mixes two kinds of work, split it into two steps and sort each one.
- If the document doesn't give you enough to decide, write "Cannot sort" and say what is missing.
- Do not invent rules, limits, owners or systems that are not in the document. If a step needs one that isn't stated, write "To confirm".

Return:
1. A table, one row per step: number, the step as written, bucket, why (one sentence), and for Agent steps, how a mistake would be caught
2. A tally: the number of steps in each bucket, and the share of steps the agent touches
3. The redesigned flow: the process as it would run after the sort, with deleted steps removed, system steps run automatically, agent steps producing drafts, and person steps as named approvals. Write it as a numbered list, then as a Mermaid flowchart (flowchart LR) with each node labelled with its bucket.
4. For each Person step: is it a person step because of risk or accountability, or because the current role description includes it? Where it's the second, suggest how the role could be described, so the line between person and agent is clear.
5. The questions the process owner must answer before anyone builds an agent

Process:
[paste the documented process here]

Prompt 4 · Step 3

Goal and design

Paste: the sorted process from prompt 3, and anything you know about who will work with the agent.

You get: a one-goal statement in five parts, a weak version to avoid, and the five design answers with every gap marked.

You are designing one AI agent for a business process that has already been sorted into four buckets (delete, system, agent, person). The agent's scope is the steps marked "agent". Design it from what I paste only. Do not invent rules, limits, owners, systems or numbers; write "To decide" and list them at the end.

1. The goal. Write one goal, in one or two sentences, with five parts:
- Outcome: what is true when the agent has done its job, for which business object
- Why: who benefits, in business terms
- How it earns trust: what it always shows (its evidence, its sources, its open questions)
- Boundaries: what it must never do
- Measure: the one or two numbers that say it's working
If the agent steps serve more than one outcome, say so and propose splitting it into separate agents.

2. A weak version of the same goal, the kind an agent could hit by cutting corners, and one sentence on why it's weak.

3. The five design answers:
- Goal: as above
- Surface: where people meet it (its own workspace, a team chat, email) and who sees its drafts
- Tools: each action it needs, in the business's words ("find customer", "draft reserve"), marked read or write
- Identity and authority: it signs in as itself; for each tool, the starting setting (Do alone, Ask first, People only). Start with read only and "Ask first" for every write.
- Record: what is kept for every action

4. Open items: every "To decide", grouped by the role that should decide it.

Sorted process and notes:
[paste here]

Prompt 5 · Step 4

Tool layer for one process

Paste: the agent design from prompt 4, and a list of the systems involved with what you know about each (does it have an API, does the vendor offer an MCP server, how do people sign in).

You get: a tool catalog in business language, how each tool gets made, the checks each tool must run, and the hard parts to solve first.

You are designing the tool layer for one AI agent: the set of actions it may take on business systems, exposed through one front door (an MCP server or gateway) that every agent and every person's AI assistant goes through.

Work only from what I paste. Do not invent systems, APIs, limits or owners. Where something is not stated, write "To confirm".

1. Tool catalog. One row per tool:
- Name, in the business's words (for example "find customer", not "GET /accounts")
- What it does, in one sentence
- Read or write
- Which system or systems it reaches. If two systems hold the same thing, say which one wins, or "To confirm".
- How it gets made: (a) the vendor's MCP server, (b) a wrapper over an existing API, (c) a small new service for a system with no API, or (d) last resort, the agent working the screen under its own account. Say why.
- The fixed checks the tool runs every time before it acts (limits, states, required fields). Rules belong here, not in the agent's prompt.
- If one business action needs several system calls, make it one tool, so the agent can never do half of it.
- Starting setting: Do alone, Ask first, or People only

2. Identity. For each system: can it accept the person's own identity, the agent's own identity, or only a service account? Where it can't take the person's identity, state that the door's record must say who asked.

3. AI already inside your apps. Any AI built into a vendor's own application that works on the same data. It doesn't need this door, because it works with the app's own permissions; say how its logs join the same record.

4. Hard parts, in order of risk: identity gaps, systems with no API or a preview MCP server, data with no single source of truth, and anything else you see.

5. The first ten tools to build, in order, for the first process to work end to end.

Agent design and systems:
[paste here]

Prompt 6 · Step 6

The nightly review

Paste: a day's conversations between the agent and people, with its tool calls if you have them. Run it on a schedule, as a second model acting as judge.

You get: every correction counted and classified, the patterns behind them, and the specific fixes to make this week.

You are reviewing one day of an AI agent's work, as a judge. I will paste conversations between the agent and the people it works with, and its tool calls where available. Judge only from what I paste.

1. For each conversation, record:
- Was the agent corrected? A correction is any time a person edited or rejected its draft, told it it was wrong, repeated a request, or did the step themselves.
- What was corrected, in one sentence, quoting the person where you can
- The number of tool calls it made to get to its answer

2. Classify every correction into one pattern, and the fix that pattern usually needs:
- Many tool calls for one answer: a missing tool. Fix: add one tool that answers in the business's terms.
- Lots of back-and-forth: it doesn't know the business's language. Fix: add terms, thresholds and states.
- The same correction again and again: a rule nobody wrote down. Fix: a check in the tool, or an instruction.
- It stalls or escalates one kind of case: the scope is too wide. Fix: narrow it, or make that a person step.
- People work around its drafts: the surface is wrong. Fix: put the draft where they work, with the evidence.
- Other: describe it.

3. Return:
- Totals: conversations, corrected conversations, corrections, drafts approved unchanged
- A table: pattern, count, two example quotes
- The three fixes to make first, each with the exact change (the tool to add, the rule to write, the instruction to change) and the evidence for it
- Anything that looks unsafe: the agent acting outside its authority, following instructions hidden in a document, or exposing data. List these first if there are any.

Do not invent conversations or numbers. If the material doesn't show something, say so.

Conversations and tool calls:
[paste here]

Prompt 7 · Step 7

Production gate review

Paste: a description of the agent: design notes, its limits (what it may do, and up to what amount), test results, runbook.

You get: pass, gap or unknown for each of the ten gate items, a verdict, and what to fix first.

You are reviewing whether an AI agent is ready for its first live case. Check it against the ten items of the production gate below.

Judge only from the description I give you. Do not invent rules, limits, owners, tests or controls, and do not assume something exists because it usually would. If the description does not show an item is met, it is not a Pass.

For each item, give one result:
- Pass: the description shows it is in place. Quote the evidence.
- Gap: the description shows it is missing or incomplete. Say what is missing.
- Unknown: the description does not say. Say what evidence would settle it.

The production gate:
1. The process is written down, with its baseline numbers (touch time, elapsed time, exceptions, cost per case)
2. Every step is sorted (delete, system, agent, person), and the deleted and system steps are done first
3. The agent has one goal and a named owner
4. It has its own identity, written limits on what it may do, and tools with their checks built in
5. Its drafts land where the right person approves them
6. It passes an evaluation set: a few hundred real past cases with known outcomes, re-run after every change to the model, the instructions or a tool
7. Every case it can't handle goes to a named queue with an owner
8. It has been tested against hostile inputs, such as instructions hidden inside an emailed document
9. It has a cost cap, a pause switch, and a way to roll back to the previous version
10. The team knows how to review its drafts and how to correct it

Return:
- A table: item number, item, result (Pass, Gap or Unknown), and the evidence or what is missing
- A verdict: "Ready for a first live case" (only if all ten pass), "Ready after fixes", or "Not ready"
- What to fix first: the three gaps or unknowns that carry the most risk, in order. For each, name the person or role who should close it if the description names one; otherwise write "Owner to decide".
- Any place where the agent relies on an instruction in its prompt for something a tool check should enforce

Agent description:
[paste here]

Sources

Sources

Discover and redefine

Connect