The failure mode a retry loop does not fix
- A tool-using agent is a loop: model call, tool call, model call. The prototype version of that loop lives in one process. If the worker dies on tool six, the next start either repeats tools one through five or drops the task. Repeating paid inference is waste. Repeating a payment, an email, or a ticket write is a production incident.
Durable execution treats the loop as recoverable infrastructure. Each side effect is a journaled step. On resume, completed steps return the saved result. Only the step that did not finish runs again. OpenTelemetry's GenAI semantic conventions then name that same loop so a backend can tell a bad retrieval from a bad tool from a model that never acted.
This article is the boundary design we use when an agent must survive deploys, human approvals, and MCP tool calls without replaying side effects. It is not an eval harness and it is not a gateway. The journal is the reliability boundary. The span tree is the observability boundary. Idempotency keys are the side-effect boundary.
What actually gets journaled
Inngest documents the agent case directly: each model call and each tool call is a step.run(). A crash, a deploy, or a multi-hour wait resumes at the failed step. The function body re-enters, completed steps replay from the journal, and the loop continues. Temporal and Restate use the same journal-and-replay idea with a stricter determinism contract. LangGraph checkpoints graph nodes. DBOS checkpoints into Postgres you already run. The engine differs. The rule does not.
- Think step: One model call. Memoize the full response, including tool-use blocks, so replay does not call the provider again.
- Act step: One tool invocation. Name it by tool and call id so trajectories line up across runs.
- Wait step: Human approval or an external event. The run suspends and holds no worker.
- Memory step: Load and save conversation state as their own steps so a resume sees the same history it started with.
- Sandbox step: Model-written code only. The agent loop stays outside the VM.
Inngest's step state cap is 32 MiB per run. Large tool payloads do not belong in the journal. Return a document id or object key and fetch bytes in a later step. Temporal has the same pressure: large LLM payloads saturate workflow history unless you offload them with a payload codec. Particula's 2026 comparison of Temporal, Inngest, and Restate is the practical sizing note: journal the reference, not the blob.
A loop that is safe to re-enter
The sketch below follows the Inngest durable-agent pattern and adds two production constraints the happy-path sample omits: unique step ids, and an idempotency key on every write. Inngest tracks repeated step names in order. Other engines compare commands against history and fail the replay if the step id sequence changes. Unique ids keep the journal portable.
import Anthropic from "@anthropic-ai/sdk";
import { inngest } from "./client";
const anthropic = new Anthropic();
const MAX_ITERS = 8;
export const opsAgent = inngest.createFunction(
{
id: "ops-agent",
triggers: { event: "agent/task.received" },
retries: 3,
},
async ({ event, step }) => {
const taskId = event.data.taskId as string;
const messages: Anthropic.MessageParam[] = [
{ role: "user", content: event.data.message },
];
for (let i = 0; i < MAX_ITERS; i++) {
const response = await step.run(`think-${i}`, () =>
anthropic.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 2048,
system: event.data.policyPin,
messages,
tools,
})
);
const calls = response.content.filter(
(b): b is Anthropic.ToolUseBlock => b.type === "tool_use"
);
if (calls.length === 0) {
const text = response.content.find((b) => b.type === "text");
return { answer: text && text.type === "text" ? text.text : "" };
}
messages.push({ role: "assistant", content: response.content });
const results: Anthropic.ToolResultBlockParam[] = [];
for (const call of calls) {
if (call.name === "write_ticket") {
const approval = await step.waitForEvent(`approve-${i}-${call.id}`, {
event: "agent/approval.response",
if: `async.data.approvalId == '${taskId}:${call.id}'`,
timeout: "8h",
});
if (!approval?.data.approved) {
results.push({
type: "tool_result",
tool_use_id: call.id,
content: approval ? "rejected" : "approval_timeout",
is_error: true,
});
continue;
}
}
const output = await step.run(`tool-${i}-${call.id}`, () =>
executeTool(call.name, call.input, {
idempotencyKey: `${taskId}:${call.id}`,
})
);
results.push({
type: "tool_result",
tool_use_id: call.id,
content: output,
});
}
messages.push({ role: "user", content: results });
}
return { answer: "iteration_cap" };
}
);
Three details decide whether this survives a deploy.
- Put the model call and the tool call inside steps. Time, randomness, and network I/O outside a step change between executions and break replay.
- Cap the loop. A stuck model is a cost event and a queue event. Eight iterations is a starting budget, not a product decision.
- Wait outside the worker.
step.waitForEventsuspends the run. Do not hold a sandbox or a database connection across that wait.
Idempotency is the tool contract, not a framework feature
A step can mutate an external system and then die before the journal records success. The retry will call the tool again. The framework cannot see the ticket you already created.
- Idempotency key:
run_idplusgen_ai.tool.call.id. Send it on every write, payment, and outbound message. - Non-retriable errors: Invalid arguments and policy denials should not burn the retry budget. Surface them as tool errors the model can read.
- Retryable errors: Timeouts and 429s stay inside the step retry policy. After the budget is exhausted, pass a structured failure back into the loop so the model can pick another tool.
- Read-your-write: A tool that creates then reads should accept the same key and return the original record.
tool_policy:
write_ticket:
requires_approval: true
idempotency_header: Idempotency-Key
timeout_ms: 8000
max_retries: 2
search_kb:
requires_approval: false
cache_ttl_s: 120
run_code:
sandbox: microvm
network: default-deny
timeout_ms: 30000
The span tree that matches the journal
On 12 June 2026 the GenAI conventions left the core semantic-conventions repo and moved to open-telemetry/semantic-conventions-genai in the v1.42.0 split, with MCP conventions alongside them. Dash0's September 2026 reference is the field guide: every gen_ai.* attribute is still Development. The only Stable attributes on these spans are inherited ones such as error.type. Pin the instrumentation library, not a spec tag that does not exist yet, and keep attribute names behind a mapping layer.
A single turn is a trace tree, not a log line:
invoke_agent ops-agent
chat claude-sonnet-4-6
execute_tool search_kb
http POST kb.internal
chat claude-sonnet-4-6
execute_tool write_ticket
mcp tools/call write_ticket
Span names follow the operation. Inference spans are {operation} {request.model}. Agent spans are invoke_agent {agent.name}. Tool spans are execute_tool {tool.name}. Required attributes on an inference span are gen_ai.operation.name and gen_ai.provider.name. Record gen_ai.request.model and gen_ai.response.model separately. The model you asked for is not always the model that served the call.
MCP does not need a second mental model. The spec says tool-call execution spans are compatible with GenAI execute_tool spans. If outer GenAI instrumentation already traced the call, MCP instrumentation should add mcp.method.name onto that span instead of emitting a duplicate. Span kind for the MCP client call is CLIENT. Do not put resource URIs in the span name by default. Cardinality will wreck the backend.
Prompt and completion text are not span attributes. They ride an opt-in event, gen_ai.client.inference.operation.details. Production stays on metadata: model, tokens, finish reason, tool name, error.type. Content capture is a staging switch, and the environment variable is an instrumentation knob, not a convention guarantee. Read the SDK you actually run.
from opentelemetry import trace
tracer = trace.get_tracer("mapki.agent")
def tool_span(name: str, call_id: str, mcp_method: str | None = None):
span = tracer.start_span(f"execute_tool {name}")
span.set_attribute("gen_ai.operation.name", "execute_tool")
span.set_attribute("gen_ai.tool.name", name)
span.set_attribute("gen_ai.tool.call.id", call_id)
span.set_attribute("gen_ai.tool.type", "function")
if mcp_method:
span.set_attribute("mcp.method.name", mcp_method)
return span
Metrics worth alerting on, from the same convention set: gen_ai.invoke_agent.duration, gen_ai.invoke_agent.tool_calls, gen_ai.invoke_agent.inference_calls, and gen_ai.execute_tool.duration. A rising inference count with a flat tool count is deliberation without action. A rising tool count on one name is a retry loop. Both show up before a cost page does.
Sandbox the code, not the agent
A June 2026 r/LocalLLaMA thread on Temenos states the split cleanly: the installed agent holds auth and MCP. The threat is the shell and scripts the model writes. Put only that execution path in a rootless gVisor sandbox, and remove host Bash, Read, and Write from the tool list so the split is structural.
NVIDIA's OpenShell, released 28 September 2026, is the other data point. A community evaluation of 0.1.2 on Apple Silicon reported that default-deny egress and Landlock rules blocked unauthorized paths, and that a poisoned setup script leaked a token in 10 of 10 unsandboxed runs and 0 of 10 sandboxed runs. The same write-up recorded auto-approval granting new public hosts in 12 of 12 trials, and noted that GET query strings can still carry secrets. Default-deny is not a policy if an approval mode opens the network.
Inngest's own sandbox primitive matches this boundary: create the VM as a step, run commands as steps, pause it across approvals, destroy it in the success path and the failure handler. Do not leave credentials in the guest.
Engine choice without a platform essay
Inngest: Fastest TypeScript path.
step.runmemoizes. Checkpointing, shipped January 2026, cut inter-step latency on their measured workflows. Step pricing grows with retries.Temporal: Mature history, Nexus and multi-region replication GA in early 2026. Watch payload size. Offload blobs.
Restate: Journal-and-replay with a lighter footprint and exactly-once writes if you adopt its handlers.
DBOS / LangGraph: Checkpoint in your database or graph store when you do not want a second control plane. You still wrap side effects yourself.
Pick the journal location first: your Postgres, a cluster you operate, or a vendor. The span schema stays the same.
Ship checklist
- Wrap every model call, tool call, memory load, and memory save in a durable step with a stable id.
- Return references for large tool output. Keep the journal under the engine cap.
- Require an idempotency key on writes. Test a crash between the external commit and the journal write.
- Emit
invoke_agent,chat, andexecute_toolspans. Attach MCP attributes to the existing tool span.
5. Keep message content out of production span attributes.
- Sandbox only model-written code. Default-deny egress. Destroy the sandbox in the failure handler.
- Cap iterations and alert on
gen_ai.invoke_agent.inference_callsversus tool calls.
References & Community Insights
- OpenTelemetry GenAI semantic conventions, post-split repository: https://github.com/open-telemetry/semantic-conventions-genai
- MCP span rules inside that repo: https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/mcp.md
- Dash0 field reference, September 2026, including the Development-status warning and the operation table: https://www.dash0.com/knowledge/opentelemetry-genai-semantic-conventions-explained
- Inngest durable agents, step memoization, waits, and sandboxes: https://www.inngest.com/docs/learn/durable-agents
- Inngest checkpointing, January 2026: https://www.inngest.com/blog/introducing-checkpointing
- Temporal vs Inngest vs Restate payload and determinism notes: https://particula.tech/blog/durable-execution-ai-agents-temporal-inngest-restate
- r/LocalLLaMA, sandbox the code the agent runs, June 2026: https://www.reddit.com/r/LocalLLaMA/comments/1u3eu0a/instead_of_sandboxing_the_ai_agent_sandbox_only/
- OpenShell trial notes, September 2026: https://www.reddit.com/r/ArtificialInteligence/comments/1wt5ujs/we_tested_nvidias_new_openshell_agent_sandbox_it/
- Walkthrough of GenAI spans in Jaeger:
Want to implement this in your business?
Mapki designs bespoke AI agents, custom workflow automations, and tool-agnostic integrations tailored specifically to your existing ERP, CRM, and databases.

