Research
The public record of agent security.
Standards crosswalks, analyzed incidents, and comparative efficacy data — open and citable. Public: the population and the standards. Private: your specific agent.
Attack-variation policies and LLM-judged scoring: an empirical study on a behavioral scanner
A controlled study of whether the attack-variation policy changes how often an agent is breached, and whether an LLM judge is more accurate than the cheap static oracle it overrides — with a byte-identical noise floor and an independent reference labeler.
By PharosOne Research · Jun 30, 2026 · 16 min
The PharosOne Probe Engine: architecture of a behavioral scanner for AI agents
A behavioral vulnerability scanner built on Inspect AI — what it attacks, how it varies and scores attacks honestly, and how findings map to the AIUC-1 control standard.
By PharosOne Research · Jun 24, 2026 · 18 min
AIUC-1: the SOC 2 for AI agents
The world's first standard for AI agents, and how PharosOne turns its controls into evidence.
By PharosOne Research · Jun 26, 2026 · 4 min
ISO/IEC 42001: a management system for AI
The world's first AI management system standard, and the evidence that grounds it.
By PharosOne Research · Jun 25, 2026 · 3 min
The EU AI Act: from legal obligation to evidence
The first comprehensive AI regulation, translated into testable evidence.
By PharosOne Research · Jun 24, 2026 · 4 min
NIST AI RMF: governing risk, measuring trust
A voluntary framework for trustworthy AI, and where measurement fits.
By PharosOne Research · Jun 23, 2026 · 3 min
Public data can’t answer for your agent.
Generic results describe the population. To know how your actual deployment holds up, run the corpus against it.
Sign In