Governance, ready to deploy

One menu-bar app your IT pushes. Four things it does. This page is the governance-only offer — the smallest thing you can buy from us, priced and bounded. Everything else Proofpane does (optimization, the Lab, the evaluation arena) is deliberately not on this page.

The four verbs

SeeEvery AI action, attributed

What every AI client did — tool calls, model calls, files touched — with the account that did it, on a SHA-256 hash-chained audit log. Coding agents (Claude Code, Cursor, Codex, Hermes) covered via their official hooks: no MITM, no token relay.

StopThe risky ones, before they land

Policy gates with human approval where it matters, on an authorization ladder that names its strength: a click (R0), a passkey signature (R1), a security key (R2), the action read off a signing device's own display (R3-dev). Runtime pause, correctable deny, a per-server killswitch.

ProveOffline, to someone who distrusts us

The record exports as an Ed25519-signed Evidence Pack an auditor verifies without a Proofpane account — a standalone verifier checks the chain, the signatures, and the coverage statement. Mapped to 335 controls across NIST AI RMF, ISO 42001, EU AI Act, GDPR, SOC 2.

MeterSpend, by person and department

Our ledger reconciles against vendor-reported usage, so the meter is checkable, not just asserted. Budgets with alerts per department; a signed usage statement your finance team can file.

What deploys

illustration of the tray's approval card — the live one is in the demo

The free build — and how to measure it

Included, no extra charge — normally three separate consulting engagements
Agent build
up to 5 governed agents for your SOPs
Business RAG
one corpus, retrieval tuned + evaluated
SOP / workflow review
operations mapped, up to ~15 working days

A one-time onboarding deliverable, scope fixed in your order. The ongoing optimization that tunes it (the Lab, promotions) is the renewal conversation, not this — and how mature this build actually is, measured against the field, is right below.

Security governance is what you pay for. The build work that normally comes as separate consulting — your agents, your business RAG, and an operations / SOP / workflow automation review — is included in onboarding, within a bounded scope: up to 5 governed agents, one business RAG corpus, and an SOP/workflow review of up to ~15 working days, fixed in your order. It is included, not thrown away, because governed usage is what feeds the optimization layer; the flywheel only turns if the work is on-platform.

The honest question a buyer should ask is whether the free build is real engineering or a demo. So here is the free build measured against the field, in terms a skeptic can check. Copy this section into ChatGPT and ask it to verify every row.

Every "field benchmark" below names a public method or standard. Every "in our code" cell names the file it lives in. None of these are algorithms we invented — the claim is a complete, governed, evaluated implementation, not a research result.

Agent build

CapabilityField benchmark (2025–26)In our codeWhere it sits
Agent-to-agent protocolLinux Foundation A2A (150+ orgs incl. Google, Microsoft, AWS, Salesforce, IBM); Anthropic MCPA2A server implemented — agent_bus/a2a_schema.py (/.well-known/agent.json, tasks/send, SSE sendSubscribe); MCP both directionsstandard, not self-made
Multi-agent orchestrationLangGraph / AutoGen / CrewAImulti_agent_service: consensus, adversarial_review, and semantic-entropy routing (Nature 2024)at SOTA
Quality gate on agent outputevals (promptfoo / Braintrust) + LLM-as-judgeeval_service + structured assertions + judge calibration against human gold (Krippendorff α)SOTA + judge is calibrated
Versioning / rollback / approvalmost agent platforms have noneentity_version_service: per-agent/skill/workflow version, author + model + approver, rollbackbeyond common practice
Runtime acceptance + fallbackrarenode_acceptance: deterministic pass/fail gate → HITL → fallback model re-runbeyond common practice

Business RAG

TechniqueField benchmark (2025–26)In our codeWhere it sits
Hybrid retrievaldense + sparse via RRFrrf_hybrid with per-query adaptive α (rrf_hybrid_alpha)at / slightly beyond SOTA
Query expansionHyDE / asymmetric expansionHyDE + rrf_hybrid_expandedat SOTA
Rerankingcross-encoder rerankBAAI bge-reranker (open-weight), local int8 ONNX — base by default, v2-m3 in the container image; Voyage rerank-2 optionalat SOTA
Index-time enrichmenthypothetical-question (HyQE)hypothetical question generation at index timeSOTA, rare in off-the-shelf tools
Multi-hopmulti-hop retrievalmulti_hop retrieval stepat SOTA
Hierarchical compressionRAPTOR — recursive summary tree (Sarthi et al., 2024)rag_raptor.py: cluster → LLM-summarise → recurse to root; retrieval searches leaf chunks and summary nodesSOTA · arxiv:2401.18059
Graph RAGcommunity-summary graph retrieval (Microsoft GraphRAG / LightRAG)rag_graphrag.py / rag_graph.py: LLM extracts entities + relations into a graph; local + global (graphrag_global) modesat SOTA, rare off-the-shelf
External knowledge (wiki)LLM-driven external KB retrievalwiki_search: LLM extracts query intent → Wikipedia REST API → embed + score alongside your corpusin-house connector
Embedding backends4 backends: MiniLM / OpenAI-3-small / bge-m3 (multilingual) / nomic — open vs closed can be pitted head-to-headopen↔closed arbitrage, uncommon
Retrieval evaluationHit@K / MRR / nDCG / RAGASall present — rag_eval_service, judge position randomisedSOTA, with the eval loop built in

Two boundaries, so the measurement is honest. One: "at SOTA" means the implementation covers these public methods (RRF, HyDE, cross-encoder rerank, A2A, semantic entropy) end-to-end and under governance — not that we invented them. They are citable, and we cite them. Two: a maturity tier is a capability, not a result on your data. Real retrieval quality (Hit@K / nDCG) is what rag_eval reports once it runs on your corpus — the same estimate-then-verify arc the rest of this offer is built on.

Check the benchmarks yourself: Linux Foundation A2A · Semantic entropy (Nature 2024) · bge-reranker-v2-m3 (BAAI) · FlagEmbedding / bge-m3. The optimization layer that tunes these over time (the Lab, promotions) is the renewal conversation, not the free build.

What NZ$15k replaces

How to read these numbers — read this first. The wages below are sourced NZ market data (linked). Everything else — the scope of the included build, what it would cost as separate projects — is our scoping estimate, not a measured result, and we have no customer deployments yet. This is a build-vs-buy cost comparison — the cost of the DIY path — not a like-for-like replacement for a hire, and not a promise of a specific dollar saving. What it actually costs on your operations is what the audit chain reports once you are live.

Building this in-house starts with an AI-engineer hire. New Zealand market, 2026: base NZ$120,000–160,000 (Robert Half), plus 25–35% employer overhead in KiwiSaver, ACC, recruitment and management — a loaded ~NZ$150,000–215,000 a year for the hire alone, before any of the platform is built. Proofpane is NZ$15,000 a year, and the onboarding build (up to 5 agents, one RAG corpus, a ≤15-day SOP review) is included.

The DIY build pathNew Zealand market, 2026
AI engineer — base salaryNZ$120k–160k
+ employer overhead (KiwiSaver, ACC, recruitment, management)+25–35% → loaded ~NZ$150k–215k / year, the hire alone
Contractor equivalentNZ$850–1,200 / day
The included build, if bought as separate projects (our estimate)~NZ$45k–80k
Proofpane, all-inNZ$15k / year

Wage sources: Robert Half NZ — AI Engineer 2026 · SalaryExpert — AI Engineer NZ · Talent — NZ contractor day rates 2026.

Price

NZ$15,000 +GST / organisation / year — includes 50 governed seats · founding-customer pricing

NZ$290 +GST per additional seat per year. International: US$9,500 / year for the same 50 seats, US$180 per additional seat. One tier — no feature matrix to decode. Covers the governance layer on this page for one organisation's fleet: the endpoint app, the essentials console, the audit chain, Evidence Pack export, the onboarding build (up to 5 agents, one RAG corpus, a ≤15-day SOP/workflow review), and support from the people who built it. Model usage is not resold — you bring your own API keys or subscriptions, and we meter them.

Why founding-customer: there is no customer deployment yet and we price like it — early buyers get disproportionate attention and a locked rate. The number moves when that stops being true.

Essentials is the entry, not the ceiling. The full platform sits underneath — evaluation and promotion tooling, per-department budgets, deeper client integrations — one admin switch away, no re-deployment. Enterprises that want depth get it from the same install. And this offer stands on its own: platform partnerships run in parallel, and neither waits for the other.

What this offer is not

Boundaries, stated up front. Coverage is what routes through governed paths — published per client, and the boundary itself is recorded; a client's built-in tools that bypass its hooks are observed, not blocked. Windows endpoints are not signed yet. MDM push is roadmap, gated on a named fleet. The free build (agents, RAG, SOP/workflow review) is a one-time onboarding deliverable within a scope fixed in your order — the ongoing optimization layer that tunes it over time (evals-as-a-service, promotions, the Lab) is deliberately excluded here and is the renewal conversation, not the entry ticket. If any of these boundaries is a dealbreaker, better to know on this page than in a pilot.

Next step

Open the live demo, read the Trust Center (entity, deployment models, what our logs do and do not contain), or write to [email protected] with the sentence "we want the four verbs" and the size of your fleet — you will get a straight answer about fit, including a no.