One menu-bar app your IT pushes. Four things it does. This page is the governance-only offer — the smallest thing you can buy from us, priced and bounded. Everything else Proofpane does (optimization, the Lab, the evaluation arena) is deliberately not on this page.
What every AI client did — tool calls, model calls, files touched — with the account that did it, on a SHA-256 hash-chained audit log. Coding agents (Claude Code, Cursor, Codex, Hermes) covered via their official hooks: no MITM, no token relay.
Policy gates with human approval where it matters, on an authorization ladder that names its strength: a click (R0), a passkey signature (R1), a security key (R2), the action read off a signing device's own display (R3-dev). Runtime pause, correctable deny, a per-server killswitch.
The record exports as an Ed25519-signed Evidence Pack an auditor verifies without a Proofpane account — a standalone verifier checks the chain, the signatures, and the coverage statement. Mapped to 335 controls across NIST AI RMF, ISO 42001, EU AI Act, GDPR, SOC 2.
Our ledger reconciles against vendor-reported usage, so the meter is checkable, not just asserted. Budgets with alerts per department; a signed usage statement your finance team can file.
airgov_daemon connect-subscription,
in the shipped binary) and one click in the app from the next release — either
way it installs that client's official hooks and is reversible..pkg and a profile, not a vendor integration. A
machine is governed the first time its user opens the app; nobody types anything.
The
IT runbook is public. What is not done: binding
devices to your IdP (Okta / Entra / Google Workspace) for verified per-user
identity — built against your actual IdP in the Deployment Sprint, not guessed at
in advance. Windows parity (Authenticode-signed) sits behind a real Windows fleet.A one-time onboarding deliverable, scope fixed in your order. The ongoing optimization that tunes it (the Lab, promotions) is the renewal conversation, not this — and how mature this build actually is, measured against the field, is right below.
Security governance is what you pay for. The build work that normally comes as separate consulting — your agents, your business RAG, and an operations / SOP / workflow automation review — is included in onboarding, within a bounded scope: up to 5 governed agents, one business RAG corpus, and an SOP/workflow review of up to ~15 working days, fixed in your order. It is included, not thrown away, because governed usage is what feeds the optimization layer; the flywheel only turns if the work is on-platform.
The honest question a buyer should ask is whether the free build is real engineering or a demo. So here is the free build measured against the field, in terms a skeptic can check. Copy this section into ChatGPT and ask it to verify every row.
| Capability | Field benchmark (2025–26) | In our code | Where it sits |
|---|---|---|---|
| Agent-to-agent protocol | Linux Foundation A2A (150+ orgs incl. Google, Microsoft, AWS, Salesforce, IBM); Anthropic MCP | A2A server implemented — agent_bus/a2a_schema.py (/.well-known/agent.json, tasks/send, SSE sendSubscribe); MCP both directions | standard, not self-made |
| Multi-agent orchestration | LangGraph / AutoGen / CrewAI | multi_agent_service: consensus, adversarial_review, and semantic-entropy routing (Nature 2024) | at SOTA |
| Quality gate on agent output | evals (promptfoo / Braintrust) + LLM-as-judge | eval_service + structured assertions + judge calibration against human gold (Krippendorff α) | SOTA + judge is calibrated |
| Versioning / rollback / approval | most agent platforms have none | entity_version_service: per-agent/skill/workflow version, author + model + approver, rollback | beyond common practice |
| Runtime acceptance + fallback | rare | node_acceptance: deterministic pass/fail gate → HITL → fallback model re-run | beyond common practice |
| Technique | Field benchmark (2025–26) | In our code | Where it sits |
|---|---|---|---|
| Hybrid retrieval | dense + sparse via RRF | rrf_hybrid with per-query adaptive α (rrf_hybrid_alpha) | at / slightly beyond SOTA |
| Query expansion | HyDE / asymmetric expansion | HyDE + rrf_hybrid_expanded | at SOTA |
| Reranking | cross-encoder rerank | BAAI bge-reranker (open-weight), local int8 ONNX — base by default, v2-m3 in the container image; Voyage rerank-2 optional | at SOTA |
| Index-time enrichment | hypothetical-question (HyQE) | hypothetical question generation at index time | SOTA, rare in off-the-shelf tools |
| Multi-hop | multi-hop retrieval | multi_hop retrieval step | at SOTA |
| Hierarchical compression | RAPTOR — recursive summary tree (Sarthi et al., 2024) | rag_raptor.py: cluster → LLM-summarise → recurse to root; retrieval searches leaf chunks and summary nodes | SOTA · arxiv:2401.18059 |
| Graph RAG | community-summary graph retrieval (Microsoft GraphRAG / LightRAG) | rag_graphrag.py / rag_graph.py: LLM extracts entities + relations into a graph; local + global (graphrag_global) modes | at SOTA, rare off-the-shelf |
| External knowledge (wiki) | LLM-driven external KB retrieval | wiki_search: LLM extracts query intent → Wikipedia REST API → embed + score alongside your corpus | in-house connector |
| Embedding backends | — | 4 backends: MiniLM / OpenAI-3-small / bge-m3 (multilingual) / nomic — open vs closed can be pitted head-to-head | open↔closed arbitrage, uncommon |
| Retrieval evaluation | Hit@K / MRR / nDCG / RAGAS | all present — rag_eval_service, judge position randomised | SOTA, with the eval loop built in |
Two boundaries, so the measurement is honest.
One: "at SOTA" means the implementation covers these public methods (RRF, HyDE,
cross-encoder rerank, A2A, semantic entropy) end-to-end and under governance — not that
we invented them. They are citable, and we cite them.
Two: a maturity tier is a capability, not a result on your data. Real
retrieval quality (Hit@K / nDCG) is what rag_eval reports once it runs on
your corpus — the same estimate-then-verify arc the rest of this offer is built on.
Check the benchmarks yourself: Linux Foundation A2A · Semantic entropy (Nature 2024) · bge-reranker-v2-m3 (BAAI) · FlagEmbedding / bge-m3. The optimization layer that tunes these over time (the Lab, promotions) is the renewal conversation, not the free build.
How to read these numbers — read this first. The wages below are sourced NZ market data (linked). Everything else — the scope of the included build, what it would cost as separate projects — is our scoping estimate, not a measured result, and we have no customer deployments yet. This is a build-vs-buy cost comparison — the cost of the DIY path — not a like-for-like replacement for a hire, and not a promise of a specific dollar saving. What it actually costs on your operations is what the audit chain reports once you are live.
Building this in-house starts with an AI-engineer hire. New Zealand market, 2026: base NZ$120,000–160,000 (Robert Half), plus 25–35% employer overhead in KiwiSaver, ACC, recruitment and management — a loaded ~NZ$150,000–215,000 a year for the hire alone, before any of the platform is built. Proofpane is NZ$15,000 a year, and the onboarding build (up to 5 agents, one RAG corpus, a ≤15-day SOP review) is included.
| The DIY build path | New Zealand market, 2026 |
|---|---|
| AI engineer — base salary | NZ$120k–160k |
| + employer overhead (KiwiSaver, ACC, recruitment, management) | +25–35% → loaded ~NZ$150k–215k / year, the hire alone |
| Contractor equivalent | NZ$850–1,200 / day |
| The included build, if bought as separate projects (our estimate) | ~NZ$45k–80k |
| Proofpane, all-in | NZ$15k / year |
Wage sources: Robert Half NZ — AI Engineer 2026 · SalaryExpert — AI Engineer NZ · Talent — NZ contractor day rates 2026.
NZ$290 +GST per additional seat per year. International: US$9,500 / year for the same 50 seats, US$180 per additional seat. One tier — no feature matrix to decode. Covers the governance layer on this page for one organisation's fleet: the endpoint app, the essentials console, the audit chain, Evidence Pack export, the onboarding build (up to 5 agents, one RAG corpus, a ≤15-day SOP/workflow review), and support from the people who built it. Model usage is not resold — you bring your own API keys or subscriptions, and we meter them.
Why founding-customer: there is no customer deployment yet and we price like it — early buyers get disproportionate attention and a locked rate. The number moves when that stops being true.
Essentials is the entry, not the ceiling. The full platform sits underneath — evaluation and promotion tooling, per-department budgets, deeper client integrations — one admin switch away, no re-deployment. Enterprises that want depth get it from the same install. And this offer stands on its own: platform partnerships run in parallel, and neither waits for the other.
Boundaries, stated up front. Coverage is what routes through governed paths — published per client, and the boundary itself is recorded; a client's built-in tools that bypass its hooks are observed, not blocked. Windows endpoints are not signed yet. MDM push is roadmap, gated on a named fleet. The free build (agents, RAG, SOP/workflow review) is a one-time onboarding deliverable within a scope fixed in your order — the ongoing optimization layer that tunes it over time (evals-as-a-service, promotions, the Lab) is deliberately excluded here and is the renewal conversation, not the entry ticket. If any of these boundaries is a dealbreaker, better to know on this page than in a pilot.
Open the live demo, read the Trust Center (entity, deployment models, what our logs do and do not contain), or write to [email protected] with the sentence "we want the four verbs" and the size of your fleet — you will get a straight answer about fit, including a no.