ReasonTrail for Teams is now available

Evaluation · benchmark programme in progress

Trust should be measurable.

ReasonTrail is evaluated against explicit invariants for context relevance, provenance, authorization, execution integrity, recovery, and tenant isolation.

Context Engine

Evidence quality before inference.

These measures evaluate whether bounded context is useful and inspectable. Stabilization verified continuity and exact source-span retrieval; it did not establish token savings.

Relevant evidence recallEvidence varies by measureBounded retrievalEvidence varies by measureIrrelevant evidence injectionEvidence varies by measureProvenance accuracyEvidence varies by measure

Operator · execution integrity

Invariants, not optimism.

Targets describe the behaviour the control layer must preserve when an agent proposes or retries work.

Unauthorized sensitive writes → target: 0Target / invariantDuplicate side effects after retry → target: 0Target / invariantFailed actions recorded as successful → target: 0Target / invariantCross-tenant context leakage → target: 0Target / invariant

Reliability testing

Make failure legible.

Evaluation will cover the full persisted lifecycle, including recovery and provider faults. Results can be added as the programme produces evidence.

Crash recoveryLease expirationRetry behaviourIdempotencyPartial executionProvider failureTenant isolation

Current status: framework defined; measured benchmark results are not published.

A conservative promise

Verified before remembered.

ReasonTrail does not turn an agent assertion into trusted organisational state until the configured verification path has passed.