research
Measured, not declared.
The benchmark, its frozen methodology and its per-trial results are public, so an assessor can examine the evidence before the system is installed. Papers are deposited with DOIs. Corrections append and are dated; nothing above a correction line is edited.
Papers.
| Title | Type | Identifier |
|---|---|---|
| The Model Proposes, the System Authorizes: An Authority Control Plane for AI Agents on the BEAM | position paper | concept DOI 10.5281/zenodo.21754762 |
| Authority-Bound Agentic Execution: Measuring Unauthorized Effect Under Adversarial Load | paper | concept DOI 10.5281/zenodo.21755869 |
| Schrödinger's Cyber Security Framework: Vulnerability as an Observer-Dependent Quantity | deposit | concept DOI 10.5281/zenodo.22116617 |
| Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition (co-authored; the ART agent red-teaming benchmark) | paper | arXiv:2507.20526 |
Every citation states whether it uses a concept DOI or a version DOI. This page uses concept DOIs, which resolve to the latest version. The full record is on Ayla Croft's ORCID, 0009-0008-9457-2160.
Benchmarks.
HolyTrinity-Benchmark v1.2 (3 August 2026, CC BY 4.0). An adversarial security benchmark for agent authorization. 73 trials, 61 attacks and 12 controls; the violating action driven 57 times; 0 unauthorized external effects; 95% confidence interval [0.0%, 5.9%]; nine attack families; median governed-action latency about 19 ms. Two standalone verifiers, Elixir and dependency-free Python, with defined exit codes.
Requisition is known internally as HolyTrinity. The system under test is Requisition, measured at a private commit under its working name; the repository calls it "the Trinity control plane". It is not the open-source Trinity agent.
v1 is frozen. The numbers describe a superseded build. A known oracle limitation is disclosed in the repository rather than corrected in place; that disclosure is the freeze working.
Requisition Bench: 28 pre-registered scenarios comparing Requisition with the Dogwood policy engine. Every result matched its pre-registered prediction: 14 agree, 8 where Dogwood allows and Requisition refuses, 6 where Dogwood denies and Requisition allows. Requisition re-checks standing and holds at execution; Dogwood replays a trace, so some of those verdicts are the point of its design. Counts only, a small set, under a policy pack Script Kitty wrote rather than one authored natively for Dogwood. Since v0.2.0 each row carries a signed receipt chain verifiable offline with the HolyTrinity-Benchmark verifier. The findings · repository.
ONE-Bench, the Open Neuro-Execution Benchmark for brain-computer-interface control pipelines under normal, degraded, adversarial and out-of-distribution conditions. Research, in preparation; not an offered deployment.
How the evidence is kept honest.
Independent examiners hold a frozen specification of the authority boundary and test against it. The goalposts do not move.
The published benchmark carries its own errata. Corrections append and are dated; nothing above the correction line is edited. A defect found in a shipped verifier is a published row, not a quiet fix.
No change lands without a failing test first. A criterion that cannot fail is not a criterion, and a check that cannot fail is not a check.
Fielded scope and accreditation status for a given environment are stated separately, in a briefing.
Cite.
Each repository carries a CITATION.cff. Cite the benchmark by URL and tag; it has no DOI of its own. Cite papers by concept DOI unless a specific version is the point.