Skip to main content

Command Palette

Search for a command to run...

Orchestrator + Specialist Agents

A Practical Blueprint for Multi-Agent Customer Support

Updated
•5 min read•View as Markdown
Orchestrator + Specialist Agents
B
We write about multi-agent systems, agentic AI, and omnichannel automation, the architecture decisions behind them, where they break in production, and what actually changes for teams running support, sales, and ops through them. We built BotSailor, where we help businesses run WhatsApp, Instagram, Telegram, and Messenger automation from one place so most of what we write comes from things we've actually shipped, broken, and fixed, not just trend-watching. Currently publishing two series here: Multi-Agent Systems, Practically and Agentic AI in Production. If you're building with agents, orchestration, or omnichannel automation, we'd genuinely like to hear what you're running into, drop it in the comments or reply to any post.

Introduction

A single monolithic AI agent trying to handle billing questions, technical troubleshooting, and order tracking all at once tends to break down in predictable ways: prompt bloat, inconsistent tone, and hallucinated answers when it strays outside what it actually knows well.

The fix that's becoming standard in production support systems is orchestrator + specialist agents — one routing agent that understands intent, and several narrow agents that each own a single domain. This post is a practical blueprint: what the pattern looks like, how the orchestrator makes routing decisions, and how to keep context intact as a conversation moves between specialists.

Prerequisites: comfort with basic agent/LLM concepts (system prompts, tool calling) and any backend language — examples below are in Python-style pseudocode but the pattern is framework-agnostic.

Table of Contents

  1. Why split into orchestrator + specialists

  2. Architecture overview

  3. Building the orchestrator

  4. Defining specialist agents

  5. Passing context between agents

  6. Handling handoff failures

  7. Summary & edge cases


1. Why Split Into Orchestrator + Specialists

Three problems show up quickly with a single do-everything agent:

  • Prompt dilution — a system prompt trying to cover billing, technical support, and shipping policy simultaneously produces vague, hedged answers in each area.

  • No clear ownership — when the bot is wrong, it's hard to tell which "part" of its knowledge failed, and harder to fix without risking regressions elsewhere.

  • Poor escalation logic — a monolith either escalates everything (defeats the automation) or nothing (frustrates the user).

Splitting responsibilities mirrors how human support teams already work: a triage agent (orchestrator) listens to the request and routes it to the right specialist, each of whom has a narrow, well-tested scope.

2. Architecture Overview

  User message ──▶ │   Orchestrator   │

                    │  (intent + route)│

                    └────────┬─────────┘

                             │

        ┌────────────────────┼────────────────────┐

        ▼                    ▼                     ▼

 ┌─────────────┐     ┌──────────────┐     ┌─────────────────┐

 │ Billing      │     │ Technical    │     │ Order Tracking   │

 │ Specialist   │     │ Specialist   │     │ Specialist       │

 └─────────────┘     └──────────────┘     └─────────────────┘

        │                    │                     │

        └────────────────────┼────────────────────┘

                             ▼

                    ┌──────────────────┐

                    │ Shared Context /  │

                    │ Conversation State│

                    └──────────────────┘

The orchestrator never answers domain questions itself — its only job is classification and handoff. Specialists never see traffic outside their domain, which keeps each one's prompt small and its behavior predictable.

3. Building the Orchestrator

The orchestrator's core job is a single classification call: given the latest message (and recent history), which specialist should handle it?

ORCHESTRATOR_SYSTEM_PROMPT = """
You are a routing agent for a customer support system.
Classify the user's message into exactly one category:
- billing
- technical
- order_tracking
- general (only if none of the above clearly apply)

Return JSON: {"category": "...", "confidence": 0.0-1.0}
Do not answer the question yourself.
"""

def route(message, history):
    result = llm_call(
        system=ORCHESTRATOR_SYSTEM_PROMPT,
        messages=history + [message],
        response_format="json"
    )
    if result["confidence"] < 0.6:
        return "general"  # fall back to a clarifying question, not a guess
    return result["category"]

Why confidence matters: a low-confidence route is worse than no route. If the orchestrator isn't sure, it should ask a clarifying question rather than silently handing the user to the wrong specialist — that's a common source of frustrating loops.

4. Defining Specialist Agents

Each specialist gets a narrow system prompt and, critically, only the tools it needs.

BILLING_SPECIALIST = Agent(
    system_prompt="""
    You handle billing questions only: invoices, refunds,
    payment methods, subscription changes. If asked about
    anything outside billing, say so and hand back to routing.
    """,
    tools=[get_invoice, issue_refund, update_payment_method],
)

TECHNICAL_SPECIALIST = Agent(
    system_prompt="""
    You handle technical troubleshooting only: login issues,
    integration errors, API problems. Escalate to a human if
    the issue involves account security.
    """,
    tools=[check_system_status, search_kb, create_ticket],
)

Keeping tool access scoped per specialist isn't just tidy — it's a safety boundary. The technical agent should never be able to call issue_refund, even by mistake.

5. Passing Context Between Agents

The failure mode most teams hit first isn't bad routing — it's context loss. The user explains their problem to the orchestrator, gets handed to a specialist, and has to repeat themselves. A shared conversation state object avoids this:

class ConversationState:
    def __init__(self, user_id):
        self.user_id = user_id
        self.history = []
        self.extracted_entities = {}  # order_id, invoice_id, etc.
        self.current_specialist = None

    def handoff(self, new_specialist, reason):
        self.history.append({
            "event": "handoff",
            "from": self.current_specialist,
            "to": new_specialist,
            "reason": reason
        })
        self.current_specialist = new_specialist

Passing extracted_entities forward (an order ID mentioned two messages ago, for example) means the specialist can pick up mid-conversation instead of re-asking basic questions.

6. Handling Handoff Failures

Two failure modes to design for explicitly:

  • Mid-conversation topic switch — a billing conversation suddenly becomes a technical one. The specialist should detect this (via its own boundary check) and re-invoke the orchestrator rather than trying to answer outside its scope.

  • Specialist uncertainty — even within its domain, a specialist may not have a confident answer. Route to human escalation rather than let it guess, especially for billing/refund actions with real financial consequences.

def specialist_response(agent, state, message):
    response = agent.run(message, state)
    if response.out_of_scope:
        return orchestrator.route(message, state.history)
    if response.low_confidence and agent.has_side_effects:
        return escalate_to_human(state)
    return response

7. Summary & Edge Cases

  • The orchestrator's only job is classification and handoff — keeping it "dumb" (no domain knowledge) is what keeps routing reliable.

  • Scope tools per specialist as a safety boundary, not just an organizational convenience — a technical agent should never be able to touch billing actions.

  • Shared conversation state (not raw chat history) is what prevents the "repeat yourself" problem across handoffs.

  • Edge case to design for early: a conversation that legitimately spans two domains at once (e.g., "my refund didn't go through and now I can't log in") — decide upfront whether that's a dual-specialist collaboration or a human escalation, because the pattern above doesn't resolve it by default.

5 views