Skip to content

Security·9 min read

SealedRun: Tamper-Evident Audit Trails for AI Agents

Agent traces stored in your own database are claims, not evidence: whoever runs the store can edit it. SealedRun, an early-stage open-source recorder, signs every agent step into a hash chain that an outsider can verify offline. Here is how it works, what it proves, and how we run it for our own agents.

Anatoli NavahrodskiFounder & CEO, GlanitPublished 7 October 2026

Why does an AI agent need an audit trail if you already have traces?

A trace and an audit trail answer different questions. A trace helps your engineers understand what the agent did so they can fix it. An audit trail has to convince someone who does not trust you (a customer's security team, an auditor, a regulator) that the record was not changed afterwards. Traces kept in a database you administer fail that test by construction: anyone with write access can edit a row, and the row will not show it. SealedRun is an open-source recorder that signs each agent step into a hash chain and lets an outsider verify the result offline. We route our own agents' model calls through it.

A hypothetical. An agent with CRM access emails a contract to the wrong person on a Tuesday; on Thursday the customer's lawyer asks for the logs. Your Postgres rows and tracing export are both accurate. Neither can prove it, because the team that made the mistake controls both stores.

A signed hash chain closes that gap. Each record carries the hash of the one before it and a signature from a key the agent was delegated, so deleting or editing one step breaks every hash after it, and the verifier names the first record that fails.

Does SealedRun replace LangSmith or Langfuse?

No. Tracing platforms are built for debugging, evals, cost and latency, and do that well. SealedRun records less detail and makes it verifiable by a third party. Its README lists OpenTelemetry GenAI spans among its inputs, so one instrumented agent can feed both. Langfuse is MIT-licensed except its ee folders and self-hostable with Docker Compose (35.5k GitHub stars on 7 October 2026), while LangSmith self-hosting is an add-on to its Enterprise plan.

Read the table by its fourth column: the first two rows are good records that only you can vouch for.

Traces, logs and a signed chain side by side

Three kinds of agent records, compared qualitatively. No performance benchmarks implied.
Record typeQuestion it answersWho can change it unnoticedThird party can verify?Typical use
Tracing platform (LangSmith, Langfuse)Why did the agent do that, and what did it cost?Anyone with write access to the trace storeNo, they have to trust the operatorDebugging, evals, latency and cost tracking
Application logsWhat happened on our servers?Ops staff, root on the log hostNoIncident response, alerting
Signed hash chain (SealedRun)Was this sequence of steps altered after it was recorded?Nobody: an edit breaks the chain and the signaturesYes, offline, with the principal ID from a trusted channelEvidence for customers and auditors, incident reconstruction
Three kinds of agent records, compared qualitatively. No performance benchmarks implied.

What the EU AI Act requires for logging, and when

For high-risk AI systems, Article 12 says the system must "technically allow for the automatic recording of events (logs) over the lifetime of the system". The logs must help spot risk situations and substantial modifications, and support post-market and deployer monitoring. Since Regulation (EU) 2026/1744 (the Digital Omnibus on AI, published 24 July 2026, in force 27 July) these obligations apply from 2 December 2027 for Annex III systems and 2 August 2028 for AI inside products covered by Annex I. Article 12 prescribes no format and never uses the word "tamper-proof".

Plenty of pages still say August 2026 for Article 12. That date is gone.

Deployers have a retention duty too: Article 26(6) requires keeping the automatically generated logs under their control for at least six months, unless other EU or national law says otherwise. The standards that would define good logging are still drafts: FprEN ISO/IEC 24970 entered its formal approval vote on 27 August 2026 (forecast publication 24 January 2027), and prEN 18229-1 is forecast for 30 May 2027.

Article 50 is a separate matter, and it is live. Its transparency duties (telling people they are talking to an AI, labelling deepfakes and synthetic content) have applied since 2 August 2026, and the Commission adopted its Article 50 guidelines on 20 July 2026. The Omnibus added one narrow grace period: providers of generative systems placed on the market before 2 August 2026 have until 2 December 2026 for machine-readable marking under Article 50(2). Breaches can cost up to EUR 15 million or 3% of worldwide turnover (Article 99(4)). An audit trail does not satisfy Article 50; at most it shows whether an agent disclosed itself.

Most agents are not high-risk. Classification follows the use case, and the Annex III list is specific: employment, education, access to essential services, critical infrastructure, law enforcement and a few more. A support bot that answers shipping questions is not on it; an agent that screens job applicants is.

What is SealedRun?

SealedRun describes itself as "Every step your AI agent takes, sealed in real time." It records model calls, MCP tool calls, memory access, policy decisions and human approvals as signed records. It ships as a Python package (pip install sealedrun), a TypeScript package (@sealedrun/core, whose verifier also runs in browsers) and a Docker image with the recorder and the Inspector UI. Code and schemas are Apache-2.0; the specification texts (SPEC.md, TRUST.md) are CC-BY-4.0, and contributors sign a CLA.

It is young, and says so. The README reads "Status: early development". Per the changelog, 0.1.0 appeared on 17 September 2026 and the first public release, 0.1.1, on 21 September. The live LLM proxy (OpenAI, Anthropic, Ollama and Gemini formats) came in 0.2.0, anchoring plus the MCP, A2A and OTLP channels in 0.3.0, and 0.4.1 on 6 October. Seven releases in 19 days. The security page states that no independent third party has reviewed the cryptographic design yet. One more sign of its age: the homepage still calls live recording "in development" while the changelog lists it as shipped in 0.2.0.

The site claims format compatibility with MCP SEP-3004 and IETF agent audit trail drafts. Both are proposals: SEP-3004 was closed without merging, and draft-sharif-agent-audit-trail-06 (29 September 2026) is an individual submission with no formal IETF standing.

How SealedRun works, step by step

The specification (v0.1.0-draft) defines seven moving parts:

  1. Delegation. The Principal, the organisation accountable for the agent, signs a statement authorising an agent key set for a scope and validity window.
  2. Record. Every step becomes one record with kind, actor, target, payload digests, data_labels, outcome, an optional policy decision, prev_hash and hash.
  3. Chain. hash is SHA-256 over the RFC 8785 canonical JSON of the record minus its hash and signatures. prev_hash links to the record at seq − 1; at seq 0 it is all zeros.
  4. Hybrid signature. Each record is signed with Ed25519 and ML-DSA-65 (the post-quantum FIPS 204 scheme), and a verifier requires both.
  5. Anchor. Recorders should anchor the chain head at least every 10 minutes during a run and always at run_end, to an RFC 3161 timestamp authority or Sigstore Rekor. Anchoring is off by default. A witness that is down never blocks a run (docs).
  6. Bundle. Export produces a ZIP with a signed manifest.json, the delegations, runs/<run_id>.jsonl, and optionally payloads and anchor receipts.
  7. Offline verification. The Python or TypeScript verifier, or the Inspector in a browser (the file "is checked in this browser and is not uploaded"), reports ok or the first failing run, seq and check.

The chain holds digests of request and response bodies, so you can prove a prompt was what you say without putting personal data in the evidence; a deleted body leaves a tombstone. The record schema is public at sealedrun.com/schema/0.1/record.json.

What a verified bundle proves, and what it does not

The trust model is the most useful page in the project because it lists its own limits. A valid bundle shows:

  • the records were produced in this order, with nothing inserted, removed or reordered later;
  • each step was signed by the keys the Principal delegated, inside the authorised time window;
  • payload digests match the bodies the agent saw, and anchors bound when the chain head existed.

It does not show that the agent did nothing else (only traffic that passed through the recorder or SDK is recorded), that outputs were correct, that data labels were accurate, or that the operator's clock was right.

The trap we would warn every new user about: "The keys inside a bundle are not a root of trust." A bundle carries its own public keys, so a forger can build a valid bundle with fresh ones. The verifier must compare principal_id with an identifier obtained out of band, otherwise it has shown integrity, not origin, and the report says so.

For strong claims the trust page asks for four things: the recorder is the only path from the agent to models and tools (enforced by network policy, with no upstream credentials in the agent's hands); anchoring is on; agent keys sit where the agent process cannot read them and Principal keys are offline; evidentiary payloads are kept or their deletion is tombstoned with a stated legal basis.

How we run it at Glanit

SealedRun is an open-source project we use. Every Claude call made by our internal agents (the ones that research, draft and check articles like this one) goes through a SealedRun recorder running in Docker on the same machine, bound to localhost. A session hook starts the container before the agents run. Claude Code is pointed at it with ANTHROPIC_BASE_URL=http://127.0.0.1:8090 plus a token header. The relevant part of our setup:

# docker-compose.yml
services:
  sealedrun:
    image: ghcr.io/sealedrun/sealedrun:0.4.0
    ports:
      - "127.0.0.1:8090:8080"
    environment:
      SEALEDRUN_API_TOKEN: ${SEALEDRUN_API_TOKEN}
      SEALEDRUN_ALLOWED_HOSTS: '["127.0.0.1","localhost"]'

# upstreams.yaml
upstreams:
  - name: anthropic
    url: https://api.anthropic.com
    dialect: anthropic
    client_auth: passthrough
    models: ["claude-*"]

We trust SealedRun with the full record of our agents' work. Every prompt and response behind this article, from the first search to the final check, is sealed in the chain. Any run can be exported as a bundle and verified offline, so a client can check the evidence on their own machine without taking our word for it. Next we are anchoring the chain head to Sigstore Rekor, which gives every run an external time bound. We pin 0.4.0; the current 0.4.1 only patches the MCP stdio wrapper, which we do not use. Pointing the base URL at the recorder has one side effect we like: if the container is down, the model call fails instead of quietly going around it.

Who should adopt it now?

Adopt now if an outsider will ask you to prove what your agent did: teams selling agents to finance, health, HR or public-sector buyers, integrators running agents for clients, security teams that want incident reconstruction they can hand to someone else. Start with the LLM proxy, turn on anchoring before you rely on any bundle externally, and read the trust model with your security lead.

One practical note: keep the recorder off the network unless it has the token and TLS the README requires. By default it has no authentication and binds to localhost.

Five minutes to a verified bundle

The README's demo uses Docker Compose and a LangGraph agent with an MCP tool:

git clone https://github.com/sealedrun/sealedrun && cd sealedrun
SEALEDRUN_API_TOKEN=change-me docker compose up -d
cd examples/langgraph-mcp && uv sync && SEALEDRUN_TOKEN=change-me uv run agent.py

The recorder and Inspector come up on port 8080. Export the run, open inspector.sealedrun.com, drop the ZIP and enter the principal ID you got from the recorder, not from the file (no agent handy? the Inspector has an example bundle). Then change one character in a .jsonl file inside the ZIP and verify again. Watching it fail is the fastest explanation of a hash chain we know.

Audit trails are one control among several; our notes on securing AI agents in production, how a production agent is built and the OWASP Top 10 for LLM applications cover the others. We build and secure production agents in Python, JavaScript and PHP; see our AI development services or tell us what your agent touches.

Frequently asked questions