Orchestrator + Specialist Agents
A Practical Blueprint for Multi-Agent Customer Support

Introduction
A single monolithic AI agent trying to handle billing questions, technical troubleshooting, and order tracking all at once tends to break down in predictable ways: prompt bloat, inconsistent tone, and hallucinated answers when it strays outside what it actually knows well.
The fix that's becoming standard in production support systems is orchestrator + specialist agents — one routing agent that understands intent, and several narrow agents that each own a single domain. This post is a practical blueprint: what the pattern looks like, how the orchestrator makes routing decisions, and how to keep context intact as a conversation moves between specialists.
Prerequisites: comfort with basic agent/LLM concepts (system prompts, tool calling) and any backend language — examples below are in Python-style pseudocode but the pattern is framework-agnostic.
Table of Contents
Why split into orchestrator + specialists
Architecture overview
Building the orchestrator
Defining specialist agents
Passing context between agents
Handling handoff failures
Summary & edge cases
1. Why Split Into Orchestrator + Specialists
Three problems show up quickly with a single do-everything agent:
Prompt dilution — a system prompt trying to cover billing, technical support, and shipping policy simultaneously produces vague, hedged answers in each area.
No clear ownership — when the bot is wrong, it's hard to tell which "part" of its knowledge failed, and harder to fix without risking regressions elsewhere.
Poor escalation logic — a monolith either escalates everything (defeats the automation) or nothing (frustrates the user).
Splitting responsibilities mirrors how human support teams already work: a triage agent (orchestrator) listens to the request and routes it to the right specialist, each of whom has a narrow, well-tested scope.
2. Architecture Overview
User message ──▶ │ Orchestrator │
│ (intent + route)│
└────────┬─────────┘
│
┌────────────────────┼────────────────────┐
▼ ▼ ▼
┌─────────────┐ ┌──────────────┐ ┌─────────────────┐
│ Billing │ │ Technical │ │ Order Tracking │
│ Specialist │ │ Specialist │ │ Specialist │
└─────────────┘ └──────────────┘ └─────────────────┘
│ │ │
└────────────────────┼────────────────────┘
▼
┌──────────────────┐
│ Shared Context / │
│ Conversation State│
└──────────────────┘
The orchestrator never answers domain questions itself — its only job is classification and handoff. Specialists never see traffic outside their domain, which keeps each one's prompt small and its behavior predictable.
3. Building the Orchestrator
The orchestrator's core job is a single classification call: given the latest message (and recent history), which specialist should handle it?
ORCHESTRATOR_SYSTEM_PROMPT = """
You are a routing agent for a customer support system.
Classify the user's message into exactly one category:
- billing
- technical
- order_tracking
- general (only if none of the above clearly apply)
Return JSON: {"category": "...", "confidence": 0.0-1.0}
Do not answer the question yourself.
"""
def route(message, history):
result = llm_call(
system=ORCHESTRATOR_SYSTEM_PROMPT,
messages=history + [message],
response_format="json"
)
if result["confidence"] < 0.6:
return "general" # fall back to a clarifying question, not a guess
return result["category"]
Why confidence matters: a low-confidence route is worse than no route. If the orchestrator isn't sure, it should ask a clarifying question rather than silently handing the user to the wrong specialist — that's a common source of frustrating loops.
4. Defining Specialist Agents
Each specialist gets a narrow system prompt and, critically, only the tools it needs.
BILLING_SPECIALIST = Agent(
system_prompt="""
You handle billing questions only: invoices, refunds,
payment methods, subscription changes. If asked about
anything outside billing, say so and hand back to routing.
""",
tools=[get_invoice, issue_refund, update_payment_method],
)
TECHNICAL_SPECIALIST = Agent(
system_prompt="""
You handle technical troubleshooting only: login issues,
integration errors, API problems. Escalate to a human if
the issue involves account security.
""",
tools=[check_system_status, search_kb, create_ticket],
)
Keeping tool access scoped per specialist isn't just tidy — it's a safety boundary. The technical agent should never be able to call issue_refund, even by mistake.
5. Passing Context Between Agents
The failure mode most teams hit first isn't bad routing — it's context loss. The user explains their problem to the orchestrator, gets handed to a specialist, and has to repeat themselves. A shared conversation state object avoids this:
class ConversationState:
def __init__(self, user_id):
self.user_id = user_id
self.history = []
self.extracted_entities = {} # order_id, invoice_id, etc.
self.current_specialist = None
def handoff(self, new_specialist, reason):
self.history.append({
"event": "handoff",
"from": self.current_specialist,
"to": new_specialist,
"reason": reason
})
self.current_specialist = new_specialist
Passing extracted_entities forward (an order ID mentioned two messages ago, for example) means the specialist can pick up mid-conversation instead of re-asking basic questions.
6. Handling Handoff Failures
Two failure modes to design for explicitly:
Mid-conversation topic switch — a billing conversation suddenly becomes a technical one. The specialist should detect this (via its own boundary check) and re-invoke the orchestrator rather than trying to answer outside its scope.
Specialist uncertainty — even within its domain, a specialist may not have a confident answer. Route to human escalation rather than let it guess, especially for billing/refund actions with real financial consequences.
def specialist_response(agent, state, message):
response = agent.run(message, state)
if response.out_of_scope:
return orchestrator.route(message, state.history)
if response.low_confidence and agent.has_side_effects:
return escalate_to_human(state)
return response
7. Summary & Edge Cases
The orchestrator's only job is classification and handoff — keeping it "dumb" (no domain knowledge) is what keeps routing reliable.
Scope tools per specialist as a safety boundary, not just an organizational convenience — a technical agent should never be able to touch billing actions.
Shared conversation state (not raw chat history) is what prevents the "repeat yourself" problem across handoffs.
Edge case to design for early: a conversation that legitimately spans two domains at once (e.g., "my refund didn't go through and now I can't log in") — decide upfront whether that's a dual-specialist collaboration or a human escalation, because the pattern above doesn't resolve it by default.




