Swarm Management: Running Agents in Numbers Is an Ops Problem
Spawning a swarm of agents is the easy part. Every framework has a way to fan out now, and the demo always looks the same: one coordinator, a spray of workers, results folded back together, applause. What the demo never shows is the second run, the one where four workers search for the same thing, one of them loops until the turn cap saves it, and the bill for the afternoon is larger than the feature it shipped.
Running agents in numbers is not a prompting problem. It is a fleet problem, and the questions are the ones you would ask about any fleet: what is each worker allowed to touch, what does it cost, who killed it, and can you tell afterwards what happened.
Start with the number that decides everything
Anthropic published real figures from their multi-agent research system in June, and they are worth sitting with before you design anything. A single agent burns roughly four times the tokens of a chat interaction. A multi-agent system burns about fifteen times. In their evaluations, token usage alone explained eighty percent of the variance in performance.
Read that last one carefully, because it cuts both ways. It means fanning out genuinely buys you something, since more parallel context is most of where the gains come from. It also means the gains are bought rather than engineered, and the invoice scales with the thing that makes the system good. A swarm is a budget decision before it is an architecture decision, and if your economics do not survive fifteen times the token spend on the tasks you plan to run, no amount of orchestration cleverness rescues it.
Their sizing heuristic is a useful anchor too. Simple fact finding gets one agent and a handful of tool calls, a comparison gets two to four subagents with ten to fifteen calls each, and only genuinely complex research justifies more than ten. That is a much more conservative ladder than most swarm demos imply.
Know which shape of work actually parallelizes
The pattern that works is orchestrator and workers, where a lead agent plans, delegates to subagents that explore separate branches, and synthesizes what comes back. It works when the subtasks are independent and the bottleneck is breadth: searching many sources, checking many candidates, reading more material than fits in one context window.
It stops working when the subtasks need each other’s context. Anthropic call this out directly, and coding is their example: work where everything depends on everything else does badly when split across agents that cannot see each other’s state. That matches what you would predict from first principles. The moment two workers need shared mutable understanding, you have built a distributed system whose only bus is natural language and whose only memory is whatever the coordinator remembered to pass along. We have decades of experience with distributed systems and none of it suggests that is the easy path.
The practical test before adding a worker is whether you can write its task description without referring to what another worker is currently doing. If you cannot, the work is not parallel and the swarm will spend its tokens discovering that on your behalf.
The task spec is the interface
Their failure list is the most useful part of that writeup, because it is boring in exactly the way real operations is boring. Early versions spawned too many subagents for simple queries, searched endlessly for sources that did not exist, and distracted each other with redundant updates. Workers duplicated each other’s searches and left gaps between them.
Every one of those is a task specification failure rather than a model failure. A subagent needs an objective, an output format, guidance on which tools to use, and explicit boundaries on what it is not covering. Treat that spec the way you would treat a function signature in a codebase with many callers, because that is what it is. When two workers overlap, the fix is in the boundaries you wrote, not in the model you picked.
The layer nobody demos
Here is the part that turns a demo into something you can leave running.
Every worker needs its own identity and its own permissions, scoped to the task it was given rather than to the swarm as a whole. A fan-out multiplies whatever blast radius one agent has by the number of workers, and a shared credential means the least careful branch defines your exposure. Read-only by default, write access granted to the specific worker that needs it, and the irreversible actions gated behind something that is not a classifier.
Budgets have to be enforced outside the model, in turns, tokens, wall-clock seconds and currency, at both the individual worker level and the swarm level. The worker-level cap is the one most frameworks give you. The swarm-level cap is the one that matters, because ten workers each politely respecting a twenty turn limit is still two hundred turns you did not intend to buy. Alongside that you want a kill switch that stops a running swarm without a deploy, and that is worth building on day one rather than during the incident.
Admission control is the piece people discover the hard way. A swarm is a thundering herd pointed at your own services and at your model provider, and the rate limit you comfortably fit under with one agent is the limit you breach at eight. Queue the workers, cap concurrency to a number you chose deliberately, and remember that a retry storm inside a swarm is indistinguishable from a small denial of service attack against your own database.
Durability matters more than it does for single agents, because long runs die and restarting a fifteen-times-tokens run from zero is a real cost rather than an annoyance. Checkpoint per subtask, make the coordinator resumable, and let a failed worker be retried on its own rather than taking the run with it. Anthropic reach the same conclusion from production experience, along with a deployment pattern worth copying: shift traffic gradually between versions so a deploy does not cut the legs off swarms that are mid-run.
Idempotency stops being optional here. Agents retry, and a swarm retries in parallel, so any write tool without an idempotency key will eventually do the same thing twice at the same moment. This is ordinary distributed systems hygiene, and the agent framing tends to hide it until the duplicate records show up.
Finally, observability has to be per worker and reconstructable. Trace every agent with parent and child spans, attribute cost per run and per worker, and keep enough of the decision trail that you can answer why the system did something a week later. If you cannot reconstruct a bad outcome from your logs, you do not have a swarm, you have a slot machine with an API.
Protocols help with plumbing, not with management
The interoperability layer has matured. MCP covers how an agent talks to tools, and A2A, which Google donated to the Linux Foundation last year, covers how agents talk to each other, including discovery and long-running task semantics. Both are worth adopting if you are crossing organizational or vendor boundaries.
Neither one schedules anything, enforces a quota, attributes a cost or stops a runaway. That layer is yours, the same way Kubernetes did not come with your capacity plan. Adopting a protocol answers how the messages are shaped, not how many of them you can afford this hour.
Then argue yourself back down
Once the management layer exists, use it to check whether you needed the swarm. Instrument the marginal value of each added worker against its marginal cost, on a real evaluation set of maybe twenty representative tasks with a rubric and some human spot checks. My guess, for most teams, is that the honest answer lands at one agent with better tools, or a workflow with a single agentic step in the middle, and that the fan-out only earns its keep on genuinely broad search problems.
That is not an argument against swarms. It is an argument for being able to tell, which is the entire job.