Skip to content
open.securityopen.securityBeta

Research and provenance

The benchmark is backed by work you can inspect.

The ISPM papers define the first task families, so anyone can check what a task set measures against the paper that defined it. Papers carry methods and results with their author list; articles carry a position and one person’s name.

Open research record

arXiv / named authors

  1. Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management VisibilityarXiv:2601.07880
  2. Cross-Vendor Sola ISPM Benchmark: Evaluating Agentic AI for Federated Identity Security ReasoningarXiv:2606.02674
  3. AI Native Asset IntelligencearXiv:2605.09115

The published record, entry by entry.

Methods, environments, results, and positions remain attached to the people and publications that produced them.

№ 001Paper

Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility

The foundational visibility and hygiene benchmark. Powers the ISPM benchmark’s single-platform visibility and hygiene tasks.

Powers ISPM arXiv:2601.07880

№ 002Paper

Cross-Vendor Sola ISPM Benchmark: Evaluating Agentic AI for Federated Identity Security Reasoning

Multi-hop entity resolution and cross-system correlation across eight integrated enterprise platforms. Powers the ISPM benchmark’s cross-platform correlation tasks.

Powers ISPM arXiv:2606.02674

№ 003Paper

AI Native Asset Intelligence

Turning fragmented cloud and identity data into asset prioritization. Part of the wider research behind open.security.

arXiv:2605.09115

№ 004Article

Who Grades the Agents? Why AI Security Needs More Than Vendor Scorecards

Security vendors grade their own agents against benchmarks they designed. The argument for independent evaluation, and what antivirus testing and vulnerability scoring did to earn their credibility.

Michael Arenzon Cybersecurity Insiders July 2026

What we have not published yet.

Two threads are in flight, and neither has a link yet. The OSB paper is the evaluator written down: what a run does, what the judges are asked, and how a grade is computed, in a form a reviewer can check without reading the source.

The second changes how the ISPM questions are answered. Today a published score comes from an agent reading a queryable snapshot of the enterprise; the same questions posed to an agent that has to operate the systems to get there already run here, on their own board, and that board is what we have not published.

Paper
Open Security Benchmark: Towards Autonomous Enterprise Cyber DefenseThe framework itself: the environment data gap, the two investigation modalities over one frozen enterprise, and the scoring that reports each criterion on its own axis.In preparation
Benchmark
The live-environment modalityISPM questions answered by driving real vendor CLIs and APIs instead of querying a snapshot, so one run measures the reasoning and the operating work together. It runs, and it is its own benchmark with its own board rather than a revision of this one — the two modalities share the environment and the answer criteria, and their scores are never conflated. Neither the contract nor the board is published yet, so nothing on this site is scored under it.Built, unpublished

Inspect the work, then join the conversation.

The code, the published worlds, and the address to write to if a benchmark is missing or a score looks wrong.

Code
Open Security AI on GitHubThe harness, the task sets, and the issue tracker.
Worlds
Datasets on Hugging FaceThe published synthetic worlds every benchmark runs against.
Contact
research@open.securityPropose a benchmark, or raise a problem with how something is scored.

Run the benchmark the papers describe.

Run the packs these papers define, under the evaluation this site documents, and compare what you get with what they report.