scriptkittyos & labs

research

Measured, not declared.

The benchmark, its frozen methodology and its per-trial results are public, so an assessor can examine the evidence before the system is installed. Papers are deposited with DOIs. Corrections append and are dated; nothing above a correction line is edited.

Papers.

TitleTypeIdentifier
The Model Proposes, the System Authorizes: An Authority Control Plane for AI Agents on the BEAMposition paperconcept DOI 10.5281/zenodo.21754762
Authority-Bound Agentic Execution: Measuring Unauthorized Effect Under Adversarial Loadpaperconcept DOI 10.5281/zenodo.21755869
Schrödinger's Cyber Security Framework: Vulnerability as an Observer-Dependent Quantitydepositconcept DOI 10.5281/zenodo.22116617
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition (co-authored; the ART agent red-teaming benchmark)paperarXiv:2507.20526

Benchmarks.

HolyTrinity-Benchmark v1.2 (3 August 2026, CC BY 4.0). An adversarial security benchmark for agent authorization. 73 trials, 61 attacks and 12 controls; the violating action driven 57 times; 0 unauthorized external effects; 95% confidence interval [0.0%, 5.9%]; nine attack families; median governed-action latency about 19 ms. Two standalone verifiers, Elixir and dependency-free Python, with defined exit codes.

Requisition is known internally as HolyTrinity. The system under test is Requisition, measured at a private commit under its working name; the repository calls it "the Trinity control plane". It is not the open-source Trinity agent.

v1 is frozen. The numbers describe a superseded build. A known oracle limitation is disclosed in the repository rather than corrected in place; that disclosure is the freeze working.

Requisition Bench: 28 pre-registered scenarios comparing Requisition with the Dogwood policy engine. Every result matched its pre-registered prediction: 14 agree, 8 where Dogwood allows and Requisition refuses, 6 where Dogwood denies and Requisition allows. Requisition re-checks standing and holds at execution; Dogwood replays a trace, so some of those verdicts are the point of its design. Counts only, a small set, under a policy pack Script Kitty wrote rather than one authored natively for Dogwood. Since v0.2.0 each row carries a signed receipt chain verifiable offline with the HolyTrinity-Benchmark verifier. The findings · repository.

ONE-Bench, the Open Neuro-Execution Benchmark for brain-computer-interface control pipelines under normal, degraded, adversarial and out-of-distribution conditions. Research, in preparation; not an offered deployment.

How the evidence is kept honest.

Independent examiners hold a frozen specification of the authority boundary and test against it. The goalposts do not move.

The published benchmark carries its own errata. Corrections append and are dated; nothing above the correction line is edited. A defect found in a shipped verifier is a published row, not a quiet fix.

No change lands without a failing test first. A criterion that cannot fail is not a criterion, and a check that cannot fail is not a check.

Fielded scope and accreditation status for a given environment are stated separately, in a briefing.

Cite.

Each repository carries a CITATION.cff. Cite the benchmark by URL and tag; it has no DOI of its own. Cite papers by concept DOI unless a specific version is the point.