Docs menu
Sub-agents: the request is the unit
An agent listed in another agent's agents is called as the tool agent_<name> and runs as the engine's own agent run, one frame below its parent. Once, limits and approvals count the whole request, not each sub-agent.
What a sub-agent is#
List registered agent names in an agent's agents field and each one is exposed to that agent as a tool named agent_<name>. The tool takes one argument, task. The sub-agent's description becomes the tool description.
A sub-agent is not a separate code path. It is born through the same agent run as a direct run, one frame below the run that delegated, under the run id agent:<parent runId>:<delegation call>. Its own processors, scorers, guard and sub-agents run. The parent's caller, cancellation, timeouts and approvals come down with the delegation.
Taint crosses both ways. If the parent already read untrusted content, the sub-agent starts tainted, so its side effects go through the same taintedSideEffects ladder. If the sub-agent reads untrusted content, the parent is tainted after it returns.
The request is the unit#
Everything a request starts — the parent, every sibling sub-agent, every level below — is one request. What holds for one agent run holds for the request:
- Once per request. A call with
idempotencyKeyandidempotencyWindow: 'run', or withidempotency: 'args', is deduplicated across the whole request. The record lives at the request (xreq:<runId>:args-…); each run keeps a mirror under its own call. - Limits are the request's.
maxToolCalls,maxTokens,maxCostUsdandmaxConcurrencycount the parent's calls, every sibling's and the child's. - Approvals are addressed. A sub-agent's question has its own address, so one approval answers exactly one call.
Example: 30 employees are handed to payroll sub-agents, and one employee is handed ten times. The payment tool is keyed by employee id with idempotencyWindow: 'run'. Before 0.10 each sub-agent was its own window and 40 payments ran. In 0.10 the window is the request, and 30 payments run.
A repeat with no key is not silently dropped either: with the default sideEffectDuplicates: 'suspend' it becomes a question for a person. A repeat the person refused is not asked about again in the same request.
The record names the call that executed it (ranBy). Compensating (compensateRun) or purging (purgeRun) the sub-agent's run that did the work unwinds and erases that effect, as it does for a direct run.
Example: a desk agent with a refund sub-agent#
The desk agent can delegate to refunder. The refund tool is a side effect, needs a person's confirmation before it runs, and is keyed by order id for the request.
import { tool } from 'ai';
import { z } from 'zod';
import { createGnl, gnlTool } from '@gnldev/durable';
const refund = gnlTool(
tool({
description: 'Refund an order',
inputSchema: z.object({ orderId: z.string(), amount: z.number() }),
execute: async ({ orderId, amount }) => payments.refund(orderId, amount),
}),
{
sideEffect: true,
confirm: true, // a person approves before it runs
idempotencyKey: (input) => (input as { orderId: string }).orderId,
idempotencyWindow: 'run', // once per request, across every sub-agent
},
);
const gnl = createGnl({
storage,
agents: {
desk: {
model,
system: 'You answer customer tickets. Hand refunds to the refunder.',
agents: ['refunder'], // desk sees the tool agent_refunder({ task })
},
refunder: {
model,
description: 'Refunds an order',
system: 'You refund orders.',
tools: { refund },
},
},
});Run the request. The sub-agent stops at refund and the question surfaces on the parent's result:
import { STAFF } from '@gnldev/durable';
const request = {
runId: 'ticket-881',
prompt: 'Refund order A-1001, 40 EUR.',
caller: STAFF,
limits: { maxToolCalls: 20 }, // counts desk's calls and refunder's together
};
const first = await gnl.run('desk', request);
const [q] = first.interrupts;
// q.toolCallId 'call_0/call_0' <delegation call>/<child call>
// q.toolName 'refund'
// q.args { orderId: 'A-1001', amount: 40 }
// q.reason "The 'refunder' sub-agent asks: 'refund' requires explicit confirmation before it runs."
// q.delegatedTo { runId: 'agent:ticket-881:call_0', agent: 'refunder' }
const done = await gnl.run('desk', {
...request,
approvals: { [q.toolCallId]: true },
});
// refund ran once; done.interrupts is []Answer the toolCallId the run surfaced, unchanged. On resume, completed work replays from the journal, the delegation goes back down with the approval, and refund runs once. Running ticket-881 again does not run it a second time.
What the parent's model gets back#
A delegation returns a result the parent can branch on, read from the sub-agent's tool outcomes rather than from its prose:
// what an agent_<name> call returns to the parent's model
type DelegationResult = {
status: 'completed' | 'failed' | 'suspended';
text: string;
interrupts: Interrupt[]; // addressed <delegation call>/<child call> when suspended
failed?: Array<{ tool: string; kind: 'error' | 'blocked' | 'denied'; message: string }>;
}failed lists each tool whose last call in the sub-agent failed, was blocked or was refused. A failure followed by a success of the same tool does not count.
Approvals from a sub-agent#
A sub-agent's question is addressed <delegation call>/<child call>. Two sub-agents that both number their calls call_0 produce two different questions. Each deeper level prefixes again.
The question carries delegatedTo: { runId, agent }, the sub-agent's run and name, and via for the next hop when that sub-agent delegated too. The reason a person reads names the sub-agent, not its run id: The 'refunder' sub-agent asks: …. Two levels read The 'lead' sub-agent asks: The 'medic' sub-agent asks: ….
Resume with approvals: { [interrupt.toolCallId]: true } on the same runId. A client that answers the id it was shown needs nothing else; do not build the child's id yourself.
Run limits across sub-agents#
Pass limits on the run. maxToolCalls, maxTokens, maxCostUsd and maxConcurrency count the whole request. The delegation call itself (agent_<name>) is not a tool call and takes no concurrency slot; the calls the sub-agent makes are counted.
A tool call takes a seat before it executes, with one atomic step on the request's total, so parallel calls cannot pass maxToolCalls (200 parallel calls, limit 50: 50 ran on memory, SQLite, Postgres and Redis). A new model step starts only if token and cost budget is left.
What is not covered#
- Postgres with limits does more work than 0.9, because sub-agents' calls are now counted in the request: 200 sub-agents with limits, 2299 → 2668 ms (10 runs).
- A Redis store written by an earlier version lists keys by SCAN until storage.rebuildKeyIndex() is called once, after every process sharing it runs 0.10.
- More than 2000 sub-agents in one request is not measured. The in-memory adapter is quadratic at large fan-out.
- A network router reads a step's status and failed tools as text (FAILED (tool: reason)), not the structured delegation result.
- maxTokens and maxCostUsd can still overshoot by one step per run executing at the same time: each concurrent run may start a step while budget is left.
API reference#
AgentConfig.agentsAgentConfig.agents: string[] — names of registered agents, each exposed as the tool agent_<name> with input { task }.
agent_<name> result{ status: 'completed' | 'failed' | 'suspended', text, interrupts, failed? } — failed?: Array<{ tool, kind: 'error' | 'blocked' | 'denied', message }>.
Interrupt.delegatedToInterrupt.delegatedTo?: { runId, agent?, via? } — which sub-agent asked, its run, and the next hop.
RunLimitsRunLimits — maxToolCalls, maxTokens, maxCostUsd, maxConcurrency count the whole request.
createAgentToolcreateAgentTool(config, { description? }) — a sub-agent built from parts without createGnl; it runs through runDurable and gets the same request window, addressing and result.
listRuns({ topLevel })listRuns({ topLevel: true }) / GET /runs?topLevel=true — lists requests and leaves sub-agent runs out; a child run's summary carries parentRunId.