Automate with permission, recovery, and cost controls
Set the authority boundary, prepare failure recovery, choose a model with real cases, and calculate the full operating cost before automation reaches production.
BEFORE LESSON ONE
Bring real material so the course produces a real result.
- 01 · WHO IT IS FOR
- For teams moving an agent or tool-using workflow from prototype toward production.
- 02 · WHAT TO BRING
- Bring a diagram of the current workflow, its tools, owners, and irreversible actions.
- 03 · HOW TO FINISH
- Every automated action has an owner, approval rule, observable result, and recovery path.
WHEN THIS COURSE FITS
Start here when any of these situations is true.
You do not need a level label. Enter when the work shows one of these signals, bring real material, and use course acceptance to decide when you are done.
- 01✓
An agent is about to change external state
The workflow writes to systems, notifies people, submits payment, or triggers irreversible action.
- 02✓
Human review exists but happens too late
The approver cannot see evidence or changes and upstream errors can no longer be contained.
- 03✓
The system can succeed but cannot fail safely
Duplicate execution, partial success, timeouts, and drift do not yet have recovery owners.
COURSE MILESTONES
Every stage leaves a change someone else can review.
These milestones are not another reading list. They show whether the lessons form one piece of work: real material at the start, connected artifacts in the middle, and evidence another person can reproduce at the finish.
- 01 BEFORE↓
Inventory every action instead of drawing one agent box.
For reads, generations, writes, notifications, and payments, record the tool, owner, impact, and reversibility. Mark every step that changes external state.
Starting evidence: each irreversible action has an owner and current approval route. - 02 MIDPOINT↓
Design authority, human review, and recovery together.
Set data boundaries and the permission ladder, place people where they can still alter the outcome, and define stop and rollback actions for timeouts, duplicate execution, partial success, and rejected approval.
Midpoint evidence: every action has observation, approval, idempotency, and a safe state. - 03 FINISH✓
Rehearse failure in a controlled setting, then calculate operating cost.
Simulate a tool timeout, duplicate request, and human rejection. After recovery works, choose a model on real cases and include retries, tools, review, and incident handling in cost.
Completion evidence: the drill recovers, switching can reverse, and cost and drift alerts have owners.
- 01✓LESSON 1
Draw the data boundary before AI sees the input
Route every field through allow, transform, isolate, or exclude so a useful workflow does not quietly become an uncontrolled data transfer.
LESSON OUTPUT · data boundary mapNot completeOpen lesson → - 02✓LESSON 2
Give an agent the smallest useful permission
Place each action on a five-level permission ladder, from read-only suggestions to explicitly approved irreversible work.
LESSON OUTPUT · permission ladderNot completeOpen lesson → - 03✓LESSON 3
Put human review where it can change the outcome
Choose review gates from impact, reversibility, and uncertainty, then give reviewers the evidence and actions needed to make a real decision.
LESSON OUTPUT · review gate mapNot completeOpen lesson → - 04✓LESSON 4
Design recovery before the agent fails
Write the stop signal, saved state, owner, rollback action, and safe retry rule before automation reaches production.
LESSON OUTPUT · recovery runbookNot completeOpen lesson → - 05✓LESSON 5
Choose a model with a job-specific scorecard
Weight quality, latency, tool use, privacy, recovery, and operating constraints using cases from the workflow you will ship.
LESSON OUTPUT · selection scorecardNot completeOpen lesson → - 06✓LESSON 6
Map the full cost of an AI workflow
Count input, output, retries, tools, waiting, review, and failure recovery instead of comparing a single token price.
LESSON OUTPUT · cost mapNot completeOpen lesson → - 07✓LESSON 7
Monitor an AI workflow before drift becomes an incident
Watch task mix, inputs, dependencies, outcomes, and human overrides against a named baseline so every alert leads to an inspectable response.
LESSON OUTPUT · drift watchboardNot completeOpen lesson →
COURSE ACCEPTANCE
Every automated action has an owner, approval rule, observable result, and recovery path.
Reading is not the finish line. Assemble the lesson outputs and ask another person to review them against the standard above.