Back to course syllabus
Automate with permission, recovery, and cost controlsLesson 6 of 7

Map the full cost of an AI workflow

Count input, output, retries, tools, waiting, review, and failure recovery instead of comparing a single token price.

LESSON OUTPUT · cost mapUPDATED · Sep 8, 2026
COURSE PROGRESS0 / 7
HOW TO USE THIS LESSON

Read with one real task in mind, then complete the workbench. The lesson is done when another person can review what you made.

Token price is one line in a workflow bill. The product pays for repeated context, generated output, retries, searches, external tools, waiting, human review, and recovery after failures. Compare systems by the cost of a completed acceptable task.

Define “completed” and “acceptable” first. If a generated answer still needs a human rewrite, or a rejected tool action counts as success, every unit cost will be distorted. The cost map must use the same terminal state as the acceptance contract.

Draw seven cost lanes

  1. Input: system rules, source material, history, and repeated context.
  2. Output: visible answer plus hidden planning or intermediate generation.
  3. Retries: automatic retries, user regenerations, and repair calls.
  4. Tools: search, databases, browsers, sandboxes, and paid APIs.
  5. Latency: infrastructure held open and user time spent waiting.
  6. Review: human minutes needed to accept or correct the result.
  7. Failure: refunds, incident work, rollback, and abandoned tasks.

Multiply each lane by real task frequency and completion rate. A cheap call with low acceptance may cost more per useful result than an expensive call that finishes cleanly.

Use one relationship across candidates:

cost per accepted task = (batch model + tool + infrastructure + labor + failure cost) / tasks that finally pass acceptance

The denominator is not request count, successful model responses, or tasks without an exception. If a batch produces no acceptable tasks, mark it undeliverable instead of hiding the result behind zero or an empty value.

Example: document review

Workflow A sends an entire document on every question and produces a long answer. Workflow B retrieves a few sections, returns a concise comparison, and includes citations. B adds retrieval cost but reduces repeated input and review time.

The winner depends on document size, question frequency, retrieval accuracy, and reviewer wage—not on the model price alone.

Example: autonomous research

An agent performs several searches, opens pages, retries blocked sources, writes notes, and assembles a report. The final generation is a small portion of the run. Tool calls, long waits, duplicate exploration, and source verification dominate.

Limiting search branches and requiring an evidence plan before browsing can reduce cost more than changing the final writing model.

Optimize the largest controllable lane

Measure a sample of completed tasks. Remove duplicated context, cap unproductive loops, cache stable facts, shorten output, improve routing, or move deterministic steps out of the model. Recalculate acceptance and review time after each change.

Savings that increase failures simply move cost into another lane.

Keep the baseline so later savings remain honestly comparable.

Complete a workflow cost record

Preserve raw quantities per run, then apply the organization’s price and time assumptions.

Cost fieldRaw observation to retain
Task and sliceWork type, input band, risk, and terminal state
Model callsVersion, input and output units, call count, cache hits
Tool callsTool, quantity, billing unit, failures, and retries
WaitingEnd-to-end duration, queue time, and human delay
Human reviewRole, minutes, accepted unchanged or after edit
Failure handlingRecovery, compensation, re-contact, and incident labor
Final stateAccepted, rejected, abandoned, or still pending

Attach effective date and currency to rates, wages, and tool prices. Keep material non-monetary measures, such as user delay or severe failure, visible in their own units. Do not invent a price merely to create one tidy total.

Compare interventions, not only invoices

Establish a baseline for the same task slices and alter one component at a time. Context compression can reduce input while deleting decisive evidence. Less review can shorten queues while increasing failure handling. A cheap routing step may avoid expensive calls or damage quality through misrouting.

Compare accepted-task cost, blocking failures, completion time, and human overrides together. If task mix or acceptance rules changed, the raw before-and-after figures no longer share a baseline. Record which lane produced a saving and where any burden moved.

Turn the map into budget protection

Convert the map into operating limits: maximum model calls, tool branches, output length, elapsed time, and human takeovers for one task. At a limit, stop, narrow the request, or seek approval instead of continuing indefinitely in the background.

A budget alert should include task identifier and dominant cost lane. An operator can then distinguish a loop, larger-than-usual source, tool price change, or review backlog. A total-spend notification alone does not suggest a recovery action.

Common questions

How should human time be valued? Use an organization-approved convention and state the assumption. Preserve minutes alongside currency so a changed labor rate does not conceal actual workload.

Does caching always save money? Only when content is stable, requests genuinely repeat, and invalidation works. Rework from stale or incorrect cache belongs in the failure lane.

Should the cheapest workflow ship? Only after every candidate passes quality, safety, and authorization gates. Cost ranks acceptable choices; it does not compensate for a failed gate.

Boundary

This map supports product decisions; it is not an accounting standard. Include the costs material to your workflow and state what is excluded. Revisit the map when traffic shape, model behavior, review policy, or tool pricing changes.

SOURCES CHECKED

Which first-party sources informed this guide?

Sources anchor definitions, risk boundaries, or operational facts. The decision framework and workbench are original to AI Vista.

  1. NIST AI RMF PlaybookChecked 2026-09-09
  2. NIST AI 600-1 Generative AI ProfileChecked 2026-09-09

TAKEAWAY TOOL

AI workflow cost map

Calculate cost per completed acceptable task.

cost / map
USE THIS WHENBefore asking someone else to run the work
YOU WILL GETA fillable, handoff-ready, reviewable artifact
HOW TO USE01—03
  1. 01
    Name the real taskDescribe the result to deliver, not an abstract goal.
  2. 02
    Fill the decision fieldsMake inputs, risk, evidence, and handoff explicit.
  3. 03
    Ask a colleague to reviewThe tool is ready when someone else can restate the decision.
DONE WHENFields are complete, boundaries are clear, and the result is reviewable.
Edits save automatically

LESSON READ

Finish the workbench, then mark it read.

The read state updates the syllabus and your course progress.

  1. 01Fields filled
  2. 02Case tested
  3. 03Reviewable

ARTICLE DISCUSSION

Leave a judgment another reader can reuse.

Record what worked, which boundary failed, or one question still worth pursuing.

DISCUSSINGMap the full cost of an AI workflowOpen the community →
0 discussionsINSIGHTS · QUESTIONS · IDEAS