ResearchField notes from the operational frontier

Notes from where decisions live.

MAIA Research is the field record behind the platform — benchmarks, methodology, and the operational gaps we’re trying to close. Each paper is written by the team that ships the code.

Read the latest paper Read the research manifesto
01 · Featured PapersEditorial ledger

The work, in print.

Nine papers shipped since the start of 2026. Each one comes out of a live deployment, not a whiteboard.

  1. 01
    2026-07-08 · MAIA Research

    What a board needs before an agent acts.

    Agentic AI is arriving in the enterprise faster than the governance around it — and the procurement documents of major regulated institutions are the leading indicator of what will actually be allowed to run. We read those asks closely: agentic controls testing, AI-generated variance narratives, continuous monitoring. All of them couple autonomy to evidence. We state the evidence-bound architecture plainly, explain why the narrative layer is where trust compounds or collapses, and give boards the five commitments to require of any agentic system before it acts.

    Read
  2. 02
    2026-07-05 · MAIA Research

    When the answer must be the same twice.

    Regulated decisions — permits, filings, audit findings — cannot ride on probability alone, and buyers of compliance AI have started writing that into their solicitations. We describe the deterministic division of labour (probabilistic extraction, deterministic evaluation over an ontology), argue for a four-verdict vocabulary whose most important member is the honest 'this needs a human,' and close with the five questions that separate production-grade compliance automation from a demo.

    Read
  3. 03
    2026-06-07 · MAIA Research

    Escalation that arrives.

    Most systems can raise an alarm. Almost none can guarantee it lands somewhere a human will see it, own it, and be able to act. We argue an escalation is a transfer of custody, not a notification — and that it is only real when it survives a reload, reaches a second account, and carries enough to act on. We catalogue the three ways escalations die, describe the spine that makes one durable and owned, and give you the four-step test to run in any procurement.

    Read
  4. 04
    2026-05-30 · MAIA Research

    Sovereignty is an architecture, not a flag.

    Every vendor we audit this year calls itself sovereign. Most of them mean one of three things — their data centre is in the right country, their support staff is cleared, or the customer gets a private VPC. None of those is sovereignty. We lay out the four architectural properties a system actually needs before the word means anything: residency you can prove, keys the customer can revoke without our help, an audit chain that survives the vendor going away, and a learning loop where the customer keeps the deltas.

    Read
  5. 05
    2026-05-26 · Mohamed Yousuf

    What we got wrong about workforce planning.

    Ten years of building rosters for airlines and a film studio, and the tools were never the problem. We had every system money could buy. What we did not have was a way to act on the schedule once it was wrong. A roster gets cut at 0600; by 0700 three operators have pushed back manually in three different systems and nobody has reconciled them. This is the part I have not seen written down — the gap between the planning surface and the operating surface, and why the gap is the actual workload.

    Read
  6. 06
    2026-05-19 · MAIA Research

    Policy-by-Proof.

    Evidence-Bound Autonomy. Every agent action must produce its proof packet — typed inputs, named policy clauses, classification floor, reversibility guarantee, signature — before it executes. The architecture, the trade-offs, and why this is the only honest model for defence-grade AI.

    Read
  7. 07
    2026-05-19 · MAIA Research

    The Living Ontology.

    An operational ontology that evolves continuously from observed operator and agent behaviour, with tamper-evident governance. Phase 1 ships agent-emitted Evolution Proposals, severity-aware HITL routing, reversible patches, federated learning by delta. We describe the architecture, the heuristics, the moat — and why no frozen-ontology platform can catch up.

    Read
  8. 08
    2026-04-22 · MAIA Research

    The State of Operational Decision-Making, 2026.

    A field survey of 217 operating environments across government, defence, and enterprise. We measure the gap between when a decision-relevant signal arrives and when an operator acts on it. The median is 41 hours. The 90th percentile is over a week. We show why dashboards have made this worse, not better, and what an operating substrate has to do to close the gap.

    Read
  9. 09
    2026-03-14 · MAIA Research

    Decision Latency Benchmark, 2026.

    We benchmark seven AI-agent stacks against 1,400 real operational decisions drawn from logistics, healthcare staffing, and energy dispatch. We score on signal-to-draft latency, evidence completeness, and policy-conformance under load. MAIA's reference stack reaches a verified decision in 11 minutes at p50, with full lineage. No competitor closes the loop under one hour without losing provenance.

    Read
  10. 10
    2026-02-05 · MAIA Research

    Multi-Domain Fusion at Sovereign Scale.

    A methodology paper on fusing classified, controlled, and open signals into a single operational picture without leaking provenance across clearance boundaries. We describe a typed fusion graph, a per-edge clearance gate, and a redaction kernel that lets the same agent reason across boundaries while emitting evidence packets distinct to each cleared audience.

    Read
  11. 11
    2026-01-18 · MAIA Research

    The Operational Ontology.

    Operations are not rows in a warehouse. They are sites, crews, assets, jurisdictions, policies, costs, and cadences — with typed relationships, lifecycle states, and authority boundaries. We propose a first-class operational ontology as the lens every agent reads through. We define the type system, the editing surface for operators, and the rules for safe drift.

    Read
02 · In the LabWhat we're working on

The next chapters.

Three threads the team is pulling on right now. No release dates — papers ship when the evidence is ready.

  1. 01

    Evidence-bound autonomy.

    What changes when an agent cannot act without producing the proof packet first — and what that costs at operational tempo.

  2. 02

    Ontology under load.

    How a typed operational model behaves when sites, crews, and jurisdictions drift faster than the schema can be re-blessed.

  3. 03

    Cross-domain fusion at sovereign scale.

    Extending the fusion graph to coalitions — multiple sovereigns, multiple clearance regimes, one shared operating picture.

03 · SubscribeNotes when published

Receive notes when published.

One email per paper. No promotion, no roundups — just the abstract and the link, sent the day it ships.