Model operations
Which model makes this workflow dependable and affordable in operation?
Choose and operate models with task evidence, end-to-end cost, privacy, latency, baseline, and recovery constraints—not a generic leaderboard.
DECISION BRANCH
Do not begin with the tool name.
- ASK FIRST↓
Is the missing piece a fact, a standard, or model capability?
- IF↓
The result is verifiable and the failure cost is bounded, expand carefully.
- OTHERWISE✓
Narrow the action, keep the human handoff, and preserve recovery.
WORKED CASE
The top benchmark model may not win in ticket operations.
Candidates look similar on a public benchmark, but the real workflow includes mixed-language input, peak latency, tool use, retries, human overrides, and a fixed operating budget.
Compare one token price and one generic accuracy score, then replace the production model with the apparent leader.
Weight real task slices and record end-to-end latency, retries, tool charges, human review, and recovery. Keep the current model as an explicit baseline.
Replay one week of work and observe a limited rollout. The candidate clears quality, latency, and cost thresholds without concentrating overrides, and can fall back to the baseline after a failed switch.
ENTRY & EXIT SIGNALS
Know when to enter—and when the decision is good enough to leave.
A topic is not an endless knowledge directory. Entry signals tell you whether the problem belongs at this layer. Exit signals decide whether to continue instead of substituting time spent reading for work completed.
- Selection relies on a generic benchmark or unit price
- Retries, tools, and human review are missing from cost
- Production lacks a baseline, rollback, or drift signal
- Candidates are compared on the same real cases
- Quality, latency, cost, and recovery share one scorecard
- A limited release is observable, stoppable, and reversible
3 PRACTICES
Build the judgment, make the artifact, then review it with real cases.
BOUNDARY
What this topic page will not do
The scorecard is a relative choice for one workflow, not a permanent ranking across tasks, versions, or providers.