The problem
An AI agent calls tools: it reads data, sends requests, moves money. Its log sits on the operator's server. After an incident the operator (or the agent, or an attacker with its credentials) can delete a step, edit what a tool returned, or claim a check was made. Nobody outside can prove what the agent really did.
How it works
The agentlog SDK hashes each input and output, links each step to the previous one and signs it with the agent's wallet. Each step lands on Arkiv as a readonly entity. The page trusts a step only if its on-chain $creator is the agent wallet. Optionally a custodian owns the entries from the first block, so the agent cannot delete its own trail.
Check it below
A real run of an LLM agent (a release checker, 9 steps) is read from the public Tiramisu RPC and verified here: hashes, signatures, links, seal. An attacker wallet wrote a fake "all checks passed" step and replayed a genuine one into the same run. Both are on chain; both are shown and ignored.
Agent run
Another agent, wallet or run
A step counts only if its on-chain $creator is this wallet. Who the agent wallet is comes from the operator, out of band (for example its docs or its own Arkiv entity).
Forged records claim to be steps of this run, but their creator is not the agent wallet
Verify an export, offline
Drop an agentlog-export/v1 file (the button above makes one, so does agentlog verify --out). The check runs in this tab with no network: it recomputes every hash, recovers every signature and follows the links from step 0 to the seal. It still works after the Arkiv entities expire.
Dispute replay the operator and the client, one run, one report
The operator of an agent and its client disagree about what the agent did. Neither can change the run on Arkiv, so both accept it as the common record. Each side brings its own private file of raw data: the operator its evidence file, the client the copies it received, each with an optional claim per step. The page checks every step against what the agent hashed and signed, and gives a verdict for each disputed step.
Files you choose here are read in this tab and never uploaded: the replay makes no network request (the demo button fetches two public demo files from this site). Same check from a terminal: node agentlog/cli.ts dispute --export run.json --operator op.json --client cl.json --out report.json.
Details off chain only hashes on Arkiv, the record stays with you
Prompts, model replies and tool results can hold customer data. They never go on chain. Arkiv holds what is needed to prove the record was not changed: the step, the action, the hash links, the hashes of input and output, and the agent's signature. The operator keeps the raw record and shows it only to the other party in a dispute or to an auditor, who check it against Arkiv themselves.
On Arkiv, public step 4 of the demo run
Press the button below.
With the operator, private the same step, raw
…
Strict mode: detailsOffChain: true
A plain hash of a short value can be guessed: anyone can hash 200, [] or a likely URL and compare. In strict mode every step gets a fresh random salt; input and output hashes commit to the salt and the value, the tool or model name is replaced by a commitment (h:…), and the public note is empty. The salt stays in the evidence file with the raw data. From the chain alone one learns only that step N was a tool.call at time T, linked and signed. The demo run above was written in the plain mode, so its notes and tool names are readable.
Same direction as Verifiable Agent Arbiter, different trade-offs
On 6 October 2026 Google Cloud and Mysten Labs announced Verifiable Agent Arbiter: detailed agent telemetry stays private in customer-controlled Google Cloud Storage, cryptographic proofs go to Walrus and are coordinated on Sui, and two organisations can replay a disputed workflow against the same receipts over the A2A protocol (report). SiteLog works the same way at its core and differs here:
- Open source, MIT. The writer, the verifier, the dispute replay and this page are in the repository. Read them, fork them, run them yourself.
- Any wallet, any stack. An entry is signed with a plain EVM key (EIP-191) and stored as an Arkiv entity. Anything that can sign a message can write a run, through the TypeScript SDK or the CLI; any wallet can keep a run alive.
- No cloud tied in. The private record is a JSON file the operator keeps wherever it wants. A dispute needs the public run and the two parties' files, nothing else.
- Checked without our server. This page and the CLI read only the public Arkiv RPC; a signed export verifies offline, after the entities expire.
EU AI Act, Article 12 (record-keeping) what each field covers
A map of the log's fields to the text of Regulation (EU) 2024/1689, quoted from the official text on EUR-Lex. Not legal advice and not a compliance claim: Article 12 applies to high-risk AI systems, and whether a system is high-risk and whether its logs are adequate is for its provider, its deployer and their counsel to decide.
| Text of the Act | What the Agent Action Log records | What it does not do |
|---|---|---|
| 12(1) “High-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system.” | Every call made through wrap() or record() becomes an entry as it happens: run.start, llm.call, tool.call, tool.error, run.end. Each is hash-linked to the previous one and signed. | Only calls routed through the SDK or CLI are logged. Coverage over the system's lifetime is the integrator's job. |
| 12(2)(a) logs “shall enable the recording of events relevant for: (a) identifying situations that may result in the high-risk AI system presenting a risk within the meaning of Article 79(1) or in a substantial modification” | tool.error entries; the model name of every llm.call (a model change is visible); forks, gaps and edits found by the verifier; runs that stop without a seal show as OPEN. | The log records events; it does not decide what counts as a risk. |
| 12(2)(b) “facilitating the post-market monitoring referred to in Article 72” | Agent, run, step, action, tool and time are typed Arkiv attributes, so queries across runs (“every http_get of this agent”) need no indexer. Signed exports for archiving. | No monitoring plan or dashboards. |
| 12(2)(c) “monitoring the operation of high-risk AI systems referred to in Article 26(5)” | Per-run verification in a browser or the CLI, by the deployer or anyone it trusts, from the public RPC; the dispute replay for a disagreement about one run. | No alerting. |
| 12(3) applies only to systems of Annex III point 1(a) (remote biometric identification); listed here as a checklist. “(a) recording of the period of each use of the system (start date and time and end date and time of each use)” | timestamp of run.start and run.end (the agent's clock) and the block number of each entity (the chain's clock). | The agent's clock is not trusted; the block number is. |
| 12(3)(b) “the reference database against which input data has been checked by the system” | The tool name and input_hash of the lookup call; the raw query in the off-chain evidence. | Only if the lookup goes through a logged tool. |
| 12(3)(c) “the input data for which the search has led to a match” | output_hash of that call on chain; the result itself in the off-chain evidence. | The content is only as complete as what the tool returned. |
| 12(3)(d) “the identification of the natural persons involved in the verification of the results, as referred to in Article 14(5)” | Not recorded by default. A reviewer's decision can be logged as a step (for example action: "human.review") with the identity in the off-chain record and only its hash on chain. | No built-in reviewer identity: personal data must not go on chain. |
| Art. 19(1) and 26(6): logs kept “for a period appropriate to the intended purpose of the high-risk AI system, of at least six months” | The seal moves every entry of a run to 180 days; anyone can extend further (agentlog retain --days N); a signed export stays verifiable after expiry. | 180 days can be shorter than six calendar months: set sealedDays to 186 or more, or extend with retain. |
Read it without this page
The same run, from the public Tiramisu RPC, with no SiteLog code. It is the exact query this page just ran.
CLI: node agentlog/cli.ts verify --agent release-checker --run RUN --signer 0x…. Writing from your own agent: README.
Why Arkiv
- No server to trust or to take down. Steps are written to and read from the public RPC. If this page and our repo vanish, the run is still there for anyone with the query above.
$creatoris set by the chain. Anyone may write into a run; only the agent wallet's steps count. The filter runs on the node.- Readonly entities, permissionless extension. Nobody can patch a step. An auditor can keep a run alive past its lease without being able to change it.
- Expiration per purpose. A live run's steps lease 14 days; the seal moves the whole run to 180 days in the same transaction. Logs do not pile up forever, evidence does not vanish early.
- Attributes are a query index. Agent, run, step, tool, action, time, entry hash and previous hash are typed attributes: "every
http_getof agent X last week" or "who points at this hash" (fork detection) is one query. - Batches. Create a step and hand its ownership to a custodian in one transaction, so there is no moment when the agent alone can delete it.
Details and what was hard: schema.md · friction.md.
Same rules, another use case
SiteLog for construction sites: an inspector's defect remarks that the contractor cannot rewrite. Roles from a client roster, status from separately signed records, forgeries shown and ignored.