"Multi-agent" gets described in the abstract a lot: multiple AI agents working together on a task. That’s accurate but not specific enough to build from. Here is what the pieces actually are once a multi-agent system is running against real software instead of a slide.
Distinct roles, not identical copies
A production multi-agent system is rarely several copies of the same agent working in parallel. It’s usually a small set of agents with different jobs: one that reads a request and figures out what’s actually being asked, one that gathers the information needed to answer it, one that drafts the resulting action, and one that checks the draft against policy before anything executes. Each agent has a narrower job than "handle the whole request," which makes each one easier to get right.
A coordinator, explicit or implicit
Something has to decide which agent runs next, what it’s handed, and what happens with its output. Sometimes that’s a dedicated orchestrator agent; sometimes it’s simpler, deterministic code that routes between agents based on the state of the task. Either way, this coordination layer is where most of the actual engineering happens. It’s also where most of the failure modes live: a handoff that drops context, an agent that receives an input it wasn’t designed for, a loop that never terminates.
Shared state, not shared memory
Agents in a production system don’t share a mind. They share a record: the task’s current state, what’s been decided so far, what’s still open. Each agent reads that record, does its part, and writes back to it. Getting that record right (what belongs in it, what doesn’t, how much of it any one agent actually needs to see) is a design decision with real consequences, not an implementation detail.
Verification before action
The step most demos skip and most production systems can’t: something checking an agent’s output before it’s allowed to touch a real system, whether that’s a second agent, a deterministic rule set, or a person. An agent that’s confidently wrong is more dangerous than one that’s visibly stuck, because a visible failure gets caught and a confident wrong answer doesn’t. This is also where a human in the loop earns its place: not as a fallback for a broken system, but as a designed checkpoint at the decisions that carry real risk.