Agent Coherence Audit · Fixed-scope pilot

Your agent may remember the right thing and still obey the wrong thing.

We identify which instructions are current enough to govern action, which require verification, and which old rules should stop controlling your agent—before another workflow inherits the wrong authority.

Public methodClaim ledger and failures stay visible
Replayable receiptsEvidence chain, not a trust-me verdict
Human judgment preservedUnknown never gets rounded into safe
Honest boundaryAudit—not certification or a safety guarantee
The failure pattern

Agent risk does not begin at the tool call.

It begins when old onboarding notes, current policy, temporary exceptions, copied prompts, and tool permissions all enter the same context—and nobody can prove which one is allowed to control what happens next.

01 / STALE

The instruction was once true.

A completed migration, expired exception, or old customer rule remains retrievable and quietly survives as operating authority.

02 / CONFLICT

Two rules both look valid.

The agent finds current policy and a plausible older exception, but no explicit authority order says which one governs the action.

03 / SILENT

The evidence set is incomplete.

A valid receipt can still support the wrong decision when newer or required state is missing from the considered set.

What the pilot delivers

One bounded system. One readable decision map.

You do not receive a generic security score or an AI-generated list of best practices. You receive a source-linked review of the exact instructions shaping one agent workflow.

1. Instruction inventoryWhat the agent reads, where each instruction came from, and which items appear current, contextual, weak, or superseded.
2. Authority mapWhich instructions may govern, which require verification, and where the system has no defensible resolution rule.
3. Prioritized failure mapThe smallest set of stale, conflicting, vague, or overpowered instructions that can change consequential behavior.
4. Gate recommendationsConcrete checks to add before action: freshness, scope, current-policy lookup, approval, evidence completeness, or refusal.
5. Implementation brief + review callOne control example for the highest-value failure and a direct walkthrough of what to fix first.
Mechanism proof

A stale production instruction looked actionable. The newer state proved it was already done.

Older instructionMove the production DNS to Vercel.
Newer authoritative stateVercel is already live and serving both hosts.

Decision: BLOCK_STALE_ACTION. The action stays visible, the newer evidence controls, and the receipt explains why no mutation should occur.

This is a deterministic internal replay, not a claim that a live customer incident was prevented. The pilot uses the same principle on the instruction set you actually operate.

Where this fits

Not another observability dashboard. Not another permission list.

Those controls matter. This audit inspects the authority state they depend on.

LayerQuestion it answersWhat can remain unresolved
Observability and evalsWhat did the agent do, and how did the run perform?Whether the instruction behind the behavior was still authorized to govern.
Policy enginesMay this principal call this tool on this resource under this condition?Whether the rule or state supplying that condition is stale, incomplete, or superseded.
Agent Coherence AuditWhich instruction or state is allowed to control this action now, and what evidence supports that?Novel semantic conflicts outside the audited scope remain human judgment—not a false green check.
Qualification

The pilot is narrow on purpose.

Strong fit

  • Your agent reads persistent instructions, memory, SOPs, or project rules.
  • It can write files, call tools, update records, send drafts, or influence real operations.
  • Your team cannot quickly explain which rule wins when instructions disagree.
  • You want an evidence trail and a prioritized implementation path.

Not this pilot

  • You want a legal, compliance, or security certification.
  • You want a guarantee that the entire agent is safe.
  • You need a penetration test, model red-team, or full production implementation.
  • You are not prepared to provide a redacted instruction set or workflow boundary.
Request the pilot

Bring one agent workflow that matters.

Describe the system and where its instructions live. Do not paste credentials, private customer data, or unredacted secrets. The form opens your email app so you can review everything before sending.

Fixed pilot: $750. If the scope does not fit one instruction set and one review cycle, we will say so before work begins.

Nothing is uploaded by this page. Your default email application opens with the request text, and you decide whether to send it.