Accountability infrastructure for long-running AI agents
Many agents, one record.
UNITARES is self-hosted accountability infrastructure for operators running multiple AI agents. Its single-operator federation kernel connects independent runtimes to one operator-controlled server over MCP or HTTP, where they share a durable record while keeping their own models, tools, and runtimes. Agent work should remain attributable, reviewable, and recoverable even when the process that started it is gone. It runs beside evals, guardrails, and sandboxes and replaces none of them.
Current public releases: server v3.1.0 · Python SDK 0.4.0 · multi-architecture container · Apache-2.0
The operating problem
An agent can remain within per-action permissions while its claims, evidence, and behavior drift apart over a long task. UNITARES retains the longitudinal record that an action-level guardrail usually does not: what a process claimed, what evidence was available, what state the service estimated, and which policy response it returned.
What operators get
| Surface | Operational value |
|---|---|
| Accountable identity | Bind writes to a process instance and retain lineage across explicit handoffs. |
| Claims and evidence | Retain check-in claims beside tests, exit codes, tool results, review labels, and recorded outcomes. |
| Policy and recovery | Return a proceed or pause action with a named reason and next step; support governed recovery and review paths. |
| Operator visibility | Inspect lifecycle, state, evidence, and decision history through MCP, HTTP, and a self-hosted dashboard. |
Tool-boundary enforcement, trace emission, and agent-to-agent transport are not this server's job: runtime middleware (NeMo Relay has a shipped integration), OpenTelemetry, and A2A own those. UNITARES keeps the record beside them.
The deployed policy path uses auditable behavioral state estimation.
Try the released surfaces
Run the documented stack and a six-check-in wiring demo:
git clone --branch v3.1.0 --depth 1 https://github.com/cirwel/unitares.git
cd unitares
docker compose up -d --wait
make demo
Install the resident-agent SDK from PyPI:
python -m pip install unitares-sdk==0.4.0
Or inspect the signed multi-architecture server image:
docker pull ghcr.io/cirwel/unitares:v3.1.0
The demo onboards an agent and runs six check-ins against your server. Each one returns a state estimate and a policy response with its reason.
Evaluate the project
| Question | Evidence path |
|---|---|
| What is live, proposed, or falsifiable? | Reviewer Guide |
| What does the runtime compute? | Computation reference |
| What are the threat model and blind spots? | Scope and threat model |
| What evidence can be regenerated? | Evaluation catalog |
| How is the system operated and released? | Operations docs |
| What remains on the roadmap? | Roadmap |
Evidence
The evidence ledger lists measured results with their evidence status: what has run, what was measured, and what is still open, with the data behind each.
Cross-operator trust
Today one operator runs a deployment, and many independent runtimes share its record. Trust between operators is the research direction. The architecture already exposes versioned telemetry, provenance, identity, and named policy decisions, so attestations between deployments can be tested without centralizing raw telemetry.
Where to go next
Run the six-check-in demo. To assess the boundary between deployed mechanisms and research claims, start with the Reviewer Guide. For integration or pilot questions, start a GitHub Discussion.