A working guide for teams building with AI agents

Agent orchestration: who handles what, and what happens when things go wrong

How multi-agent systems route requests, hand off what they can't handle, and recover from failure — and what separates a well-designed system from a fragile one.
16 sections · roughly 60 minutes
Interactive · click through examples as we go
Roadmap

What we'll cover

01 — Why orchestration

One generalist, or a team of specialists?

A single AI agent can be asked to do everything. In practice, that's rarely the best design — for the same reason a company doesn't ask one employee to handle sales, engineering, and legal.
One generalist agent

Everything goes through one model

  • Simple to build at first
  • Struggles as scope grows
  • Hard to improve one skill without risking another
  • No natural place to add safety checks
A team of specialists

Requests routed to the right expert

  • Each agent is narrow and accurate
  • Cheap models handle easy cases; expensive ones handle hard cases
  • Sensitive requests can be forced through review
  • Easier to test and improve piece by piece
02 — The core flow

Every request follows this shape

Incoming request Router Specialist agent Escalation Fallback handler

Click a box

Each stage in this flow has a distinct job. Click any box in the diagram to see what it does — we'll go deep on each one over the next sections.

03 — Routing logic

The router's only job: decide who handles this

A router agent looks at the incoming request and outputs a decision — "send this to Agent A" — usually paired with a confidence score. That score becomes important later: it's what triggers escalation when it's too low.
Input

A raw request — a message, a ticket, an API call — with no guarantee of clean phrasing or intent.

Output

An agent assignment, plus a confidence value the rest of the system can act on.

03 — Routing logic

Four approaches, one tradeoff

Cheap and fast on one end, flexible and accurate on the other. Drag the slider.
CheapestFastest edge casesHandles nuanceMost flexible
03 — Try it

You be the router

Read the request. Pick where it should go.
Example 1 of 3
04 — Specialist agent design

Five things a good specialist agent needs

Click each one for more detail.

Narrow scope

Does one job well, not many jobs adequately.

"Handles billing disputes" beats "handles customer service." Narrow scope means you can actually enumerate and test the cases it needs to cover.
Click to expand

Clear boundaries

Knows what it doesn't do.

A good billing agent recognizes a technical issue and hands it off, rather than muddling through. Agents that don't know their limits produce confidently wrong answers.
Click to expand

Minimal access

Only the tools and data it needs.

Smaller blast radius if something goes wrong, and fewer ways for the agent to wander off task. Broad "just in case" permissions are a liability, not a convenience.
Click to expand

Calibrated confidence

Knows when it's unsure.

This is the signal that feeds escalation. An agent that sounds equally confident whether it's right or wrong is dangerous — nothing downstream knows when to double-check it.
Click to expand

Structured output

Reliable, predictable shape.

Specialist agents often hand results to other agents, not just humans. The same output shape every time means the next step doesn't have to guess how to parse it.

In practice — a flight-booking agent doesn't reply "sure, I found flight AA123 leaving at 3pm." It returns {"intent":"book_flight","flight_id":"AA123","seat":"14C","price_usd":214}. A downstream payment agent can read that directly — no need to re-parse a sentence to guess which field is the price.
Click to expand
05 — Escalation

Four things that trigger an escalation

Escalation is what happens when the agent doesn't know what to do — not when something breaks.
Most common

Low confidence

The agent's confidence score falls below threshold — it's unsure this is even the right category.

Typical thresholds — most systems commit above ~70–85% confidence and escalate below it. A middle "gray zone" (~50–70%) often routes to a more capable agent first. Exact numbers are tuned per system.

Scope mismatch

Not my job

Not "I'm unsure" — "this genuinely isn't mine." A billing agent asked to debug an API should hand off immediately.

Policy trigger

Always reviewed

Legal threats, large sums, fraud signals — some topics get human review regardless of confidence. A deliberate design choice.

Repeated failure

Retry ceiling reached

The user keeps saying "that's not it." After N attempts, stop looping and hand off instead of forcing the automated path.

06 — Fallback

Retries need backoff, or they make things worse

Fallback ≠ escalation: the agent isn't confused — something technically broke. Retrying instantly just piles onto an already-struggling service.
try 1 · 1s
try 2 · 2s
try 3 · 4s
fallback
In practice — this is exactly how Stripe and AWS SDKs handle a failed API call. Say a ride-hailing app's driver-matching request times out: instead of hammering that service again immediately, the app waits ~1s, then ~2s, then ~4s. If it's still failing after 3 tries, it stops retrying and shows "finding you a driver is taking longer than usual" while switching to a backup dispatch service — rather than freezing the screen or crashing.
06 — Fallback

When retries run out: failover options

Redundant instance

A second copy of the same agent or service running elsewhere. Same capability, different instance.

In practice — Netflix runs services across multiple AWS zones; if one has an outage, traffic shifts to an identical instance elsewhere and most viewers never notice.

Different provider

Primary model API down → route to a secondary provider. Usually a downgrade, but keeps the system running.

In practice — a support platform's primary LLM provider has an outage, so new requests route to a backup model from a different vendor — slightly different answers, but the queue keeps moving.

Simpler agent

A smaller, faster, less capable version that gives an answer rather than timing out repeatedly.

In practice — a coding assistant's deep multi-file agent keeps timing out on a huge repo, so it falls back to a lighter single-file agent that at least answers instead of leaving the user with nothing.

Cached response

For read-heavy tasks, a clearly-labeled stale answer can beat returning an error.

In practice — a weather app's live-conditions API fails to respond in time, so it shows last hour's reading labeled "as of 40 min ago" rather than a blank screen.

One more layer worth knowing: a circuit breaker stops calling a dependency entirely for a cooldown period after repeated failures.

In practice — Netflix's Hystrix pattern: if a service fails repeatedly, the app stops calling it for 30 seconds and shows a fallback (e.g. popular titles) instead. One test request after the cooldown decides whether normal traffic resumes.

Recap — escalation vs. fallback

Which one is it?

You've now seen both failure modes. Escalation: the agent doesn't know what to do. Fallback: something technically broke. Sort these six real scenarios.
Score: 0 / 6
07 — The human's role

What happens once it reaches a person

Judgment

Decision-maker

Borderline calls involving values or risk — not something the system should decide alone.

Verification

Quality gate

For high-stakes actions, a human simply approves or stops the agent's proposed action before it executes.

Novel cases

Problem solver

The request doesn't fit any category — someone has to actually solve it, script or no script.

System improvement

Feedback source

Every resolved escalation is training signal — for tightening scope, thresholds, and routing.

The best systems treat human escalation as a precision instrument, reserved for what truly needs it — not a safety blanket for cases the system wasn't designed carefully enough to handle.

Recap

The whole system, one more time

Incoming request Router Specialist agent confident match Escalation low confidence / policy Fallback handler error / timeout Human / senior agent resolves & feeds back Retry / backup agent

Two failure modes, two paths

Right side: the agent doesn't know what to do — escalate to a person. Left side: something broke technically — retry, then fail over. Keep these signals separate; a spike in one means "improve the agents," a spike in the other means "fix the infrastructure."

08 — Takeaways

A checklist for your own system

Click each item as we walk through it.
Thank you

Questions, edge cases, war stories?

The core idea to leave with: routing decides who acts, escalation handles what the system doesn't know, and fallback handles what technically broke. Keep those three separate, and the rest of the design gets much easier to reason about.