GNL
Docs menu
Core · Free@gnldev/durable

Sub-agents: the request is the unit

An agent listed in another agent's agents is called as the tool agent_<name> and runs as the engine's own agent run, one frame below its parent. Once, limits and approvals count the whole request, not each sub-agent.

What a sub-agent is#

List registered agent names in an agent's agents field and each one is exposed to that agent as a tool named agent_&lt;name&gt;. The tool takes one argument, task. The sub-agent's description becomes the tool description.

A sub-agent is not a separate code path. It is born through the same agent run as a direct run, one frame below the run that delegated, under the run id agent:&lt;parent runId&gt;:&lt;delegation call&gt;. Its own processors, scorers, guard and sub-agents run. The parent's caller, cancellation, timeouts and approvals come down with the delegation.

Taint crosses both ways. If the parent already read untrusted content, the sub-agent starts tainted, so its side effects go through the same taintedSideEffects ladder. If the sub-agent reads untrusted content, the parent is tainted after it returns.

The request is the unit#

Everything a request starts — the parent, every sibling sub-agent, every level below — is one request. What holds for one agent run holds for the request:

  • Once per request. A call with idempotencyKey and idempotencyWindow: 'run', or with idempotency: 'args', is deduplicated across the whole request. The record lives at the request (xreq:&lt;runId&gt;:args-…); each run keeps a mirror under its own call.
  • Limits are the request's. maxToolCalls, maxTokens, maxCostUsd and maxConcurrency count the parent's calls, every sibling's and the child's.
  • Approvals are addressed. A sub-agent's question has its own address, so one approval answers exactly one call.

Example: 30 employees are handed to payroll sub-agents, and one employee is handed ten times. The payment tool is keyed by employee id with idempotencyWindow: 'run'. Before 0.10 each sub-agent was its own window and 40 payments ran. In 0.10 the window is the request, and 30 payments run.

A repeat with no key is not silently dropped either: with the default sideEffectDuplicates: 'suspend' it becomes a question for a person. A repeat the person refused is not asked about again in the same request.

The record names the call that executed it (ranBy). Compensating (compensateRun) or purging (purgeRun) the sub-agent's run that did the work unwinds and erases that effect, as it does for a direct run.

Example: a desk agent with a refund sub-agent#

The desk agent can delegate to refunder. The refund tool is a side effect, needs a person's confirmation before it runs, and is keyed by order id for the request.

createGnl — desk delegates to refunder
import { tool } from 'ai';
import { z } from 'zod';
import { createGnl, gnlTool } from '@gnldev/durable';

const refund = gnlTool(
  tool({
    description: 'Refund an order',
    inputSchema: z.object({ orderId: z.string(), amount: z.number() }),
    execute: async ({ orderId, amount }) => payments.refund(orderId, amount),
  }),
  {
    sideEffect: true,
    confirm: true,                // a person approves before it runs
    idempotencyKey: (input) => (input as { orderId: string }).orderId,
    idempotencyWindow: 'run',     // once per request, across every sub-agent
  },
);

const gnl = createGnl({
  storage,
  agents: {
    desk: {
      model,
      system: 'You answer customer tickets. Hand refunds to the refunder.',
      agents: ['refunder'],       // desk sees the tool agent_refunder({ task })
    },
    refunder: {
      model,
      description: 'Refunds an order',
      system: 'You refund orders.',
      tools: { refund },
    },
  },
});

Run the request. The sub-agent stops at refund and the question surfaces on the parent's result:

run, read the question, resume with the approval
import { STAFF } from '@gnldev/durable';

const request = {
  runId: 'ticket-881',
  prompt: 'Refund order A-1001, 40 EUR.',
  caller: STAFF,
  limits: { maxToolCalls: 20 },   // counts desk's calls and refunder's together
};

const first = await gnl.run('desk', request);
const [q] = first.interrupts;
// q.toolCallId  'call_0/call_0'   <delegation call>/<child call>
// q.toolName    'refund'
// q.args        { orderId: 'A-1001', amount: 40 }
// q.reason      "The 'refunder' sub-agent asks: 'refund' requires explicit confirmation before it runs."
// q.delegatedTo { runId: 'agent:ticket-881:call_0', agent: 'refunder' }

const done = await gnl.run('desk', {
  ...request,
  approvals: { [q.toolCallId]: true },
});
// refund ran once; done.interrupts is []

Answer the toolCallId the run surfaced, unchanged. On resume, completed work replays from the journal, the delegation goes back down with the approval, and refund runs once. Running ticket-881 again does not run it a second time.

What the parent's model gets back#

A delegation returns a result the parent can branch on, read from the sub-agent's tool outcomes rather than from its prose:

What agent_<name> returns
// what an agent_<name> call returns to the parent's model
type DelegationResult = {
  status: 'completed' | 'failed' | 'suspended';
  text: string;
  interrupts: Interrupt[];   // addressed <delegation call>/<child call> when suspended
  failed?: Array<{ tool: string; kind: 'error' | 'blocked' | 'denied'; message: string }>;
}

failed lists each tool whose last call in the sub-agent failed, was blocked or was refused. A failure followed by a success of the same tool does not count.

Approvals from a sub-agent#

A sub-agent's question is addressed &lt;delegation call&gt;/&lt;child call&gt;. Two sub-agents that both number their calls call_0 produce two different questions. Each deeper level prefixes again.

The question carries delegatedTo: { runId, agent }, the sub-agent's run and name, and via for the next hop when that sub-agent delegated too. The reason a person reads names the sub-agent, not its run id: The 'refunder' sub-agent asks: …. Two levels read The 'lead' sub-agent asks: The 'medic' sub-agent asks: ….

Resume with approvals: { [interrupt.toolCallId]: true } on the same runId. A client that answers the id it was shown needs nothing else; do not build the child's id yourself.

Run limits across sub-agents#

Pass limits on the run. maxToolCalls, maxTokens, maxCostUsd and maxConcurrency count the whole request. The delegation call itself (agent_&lt;name&gt;) is not a tool call and takes no concurrency slot; the calls the sub-agent makes are counted.

A tool call takes a seat before it executes, with one atomic step on the request's total, so parallel calls cannot pass maxToolCalls (200 parallel calls, limit 50: 50 ran on memory, SQLite, Postgres and Redis). A new model step starts only if token and cost budget is left.

Run limits

What is not covered#

Warning
  • Postgres with limits does more work than 0.9, because sub-agents' calls are now counted in the request: 200 sub-agents with limits, 2299 → 2668 ms (10 runs).
  • A Redis store written by an earlier version lists keys by SCAN until storage.rebuildKeyIndex() is called once, after every process sharing it runs 0.10.
  • More than 2000 sub-agents in one request is not measured. The in-memory adapter is quadratic at large fan-out.
  • A network router reads a step's status and failed tools as text (FAILED (tool: reason)), not the structured delegation result.
  • maxTokens and maxCostUsd can still overshoot by one step per run executing at the same time: each concurrent run may start a step while budget is left.

API reference#

typeAgentConfig.agents

AgentConfig.agents: string[] — names of registered agents, each exposed as the tool agent_<name> with input { task }.

typeagent_<name> result

{ status: 'completed' | 'failed' | 'suspended', text, interrupts, failed? } — failed?: Array<{ tool, kind: 'error' | 'blocked' | 'denied', message }>.

typeInterrupt.delegatedTo

Interrupt.delegatedTo?: { runId, agent?, via? } — which sub-agent asked, its run, and the next hop.

typeRunLimits

RunLimits — maxToolCalls, maxTokens, maxCostUsd, maxConcurrency count the whole request.

fncreateAgentTool

createAgentTool(config, { description? }) — a sub-agent built from parts without createGnl; it runs through runDurable and gets the same request window, addressing and result.

fnlistRuns({ topLevel })

listRuns({ topLevel: true }) / GET /runs?topLevel=true — lists requests and leaves sub-agent runs out; a child run's summary carries parentRunId.