Your agent crashed halfway. Resume it where it stopped.
A restart should continue from the first step that never finished — not run the model and the finished tools again. This page shows why generateText cannot, what checkpointing by hand misses, and the change that does it.
The symptom#
A support agent handles one message in three steps: it looks the customer up, opens a ticket, then emails the customer the ticket number. The process dies after the ticket was opened and before the email went out. When the job runs again, the agent starts from the first step: it looks the customer up again, and the model decides again whether to open a ticket.
import { generateText, tool, stepCountIs } from 'ai';
import { z } from 'zod';
const tools = {
lookupCustomer: tool({
description: 'Find the customer',
inputSchema: z.object({ email: z.string() }),
execute: async ({ email }) => crm.find(email),
}),
createTicket: tool({
description: 'Open a support ticket',
inputSchema: z.object({ customerId: z.string(), subject: z.string() }),
execute: async (input) => desk.open(input),
}),
sendEmail: tool({
description: 'Email the customer',
inputSchema: z.object({ to: z.string(), body: z.string() }),
execute: async (input) => mailer.send(input),
}),
};
// A queue worker runs this for each incoming message.
await generateText({ model, tools, stopWhen: stepCountIs(10), prompt: job.message });
// The process dies after createTicket, before sendEmail. The retry starts from
// lookupCustomer, and the model decides again whether to open a ticket.Why it happens#
Everything generateText knows about the turn — the model's replies, the tool results, which step it is on — lives in memory. When the process dies, all of it goes with it. The next attempt is a new call with only the prompt, so the model is asked again from the beginning and paid for again.
What happens to the ticket depends on the model: it may open a second one, or it may notice nothing and skip the email. Either way, the half that already happened is invisible to the half that did not.
Checkpointing by hand#
The usual fix is to save the conversation after every step and start the next attempt from the saved messages.
const saved = await db.loadMessages(job.id); // [] on the first attempt
await generateText({
model, tools, stopWhen: stepCountIs(10),
messages: [{ role: 'user', content: job.message }, ...saved],
onStepFinish: async (step) => {
// runs AFTER the step's tools — a crash before this line loses the step
await db.saveMessages(job.id, step.response.messages);
},
});It gets you most of the way, and it misses three things:
- The step that crashed.
onStepFinishruns after the step's tools. If the process dies aftercreateTicketran and before the step was saved, the saved messages do not show the ticket, and the retry opens it again. - Two workers, one job. A queue that redelivers a job while the first worker is still alive gives you two agents resuming the same conversation at once — and both run the tools.
- The tools are still unguarded. Restoring messages does not make a side-effecting tool safe to run twice; each one still needs its own guard.
With GNL#
Give the job an id and a journal, and call runDurable where you called generateText. The first attempt and every retry run exactly the same code.
import { runDurable } from '@gnldev/durable';
import { PostgresStorage } from '@gnldev/durable/postgres';
const journal = new PostgresStorage({ connectionString: process.env.DATABASE_URL! }).runs;
// The queue handler: the first attempt and every retry run this same call.
export async function handle(job: { id: string; message: string }) {
return runDurable({
runId: `support:${job.id}`, // the id of THIS job — never a session id
journal,
model, tools, stopWhen: stepCountIs(10),
prompt: job.message,
});
}On a retry, the model's recorded replies are replayed from the journal instead of being requested again, so the finished steps cost nothing. lookupCustomer and createTicket return their recorded results and do not run. The agent reaches the step that never finished — sending the email — and runs it live.
Two workers that pick up the same job do not both run a tool: each tool call is claimed atomically in the journal before it executes, so its body runs once. With a run lock, the second worker is refused outright with RunBusyError.
Streaming works the same way: streamDurable takes the same runId and journal and resumes the same run.
runId from the job (support: plus the job id), not from the conversation — a session id would make every later message replay the first one.When GNL does not solve it#
The crash inside a tool call. If the process died while createTicket was running, GNL cannot know whether the ticket exists. It does not guess: runDurable throws SideEffectRetryBlockedError and a human decides — or the tool's recover() asks your ticket system.
After the crash point the model runs live. Replay reproduces what was recorded; from the first unfinished step on, the model is called again and can choose differently, and nothing checks that part — there is no record to check it against. For the recorded part, replay: 'strict' turns a replayed step whose tool arguments no longer match the record into a DivergenceError instead of a silent continue.
Storage that loses an acknowledged write. Resume can only continue from what the journal kept. On Postgres with failover, run synchronous replication.
Sources#
Every claim on this page has a test or a runnable example in the public repository:
packages/durable/test/crash-window.test.ts— a crash between a side effect and its journal write, and what resume does with itpackages/durable/test/multi-worker.test.ts— two workers resuming the same run, one execution per tool callpackages/durable/test/stream-crash-window.test.ts— the same crash while streamingexamples/incident-proofs— a tool call resent from a checkpoint after a reconnect, reproduced and blocked, no API key
The mechanism in full: deterministic replay. The duplicate half of the same problem: my AI SDK tool ran twice after a retry. github.com/Karaca7/gnldev