Skip to content
open.securityopen.securitybySola SecurityBeta

Open benchmarks for autonomous cyber defense.

We make the capabilities and limits of security agents visible. Through open, reproducible evaluation across realistic enterprise environments, teams can see how agents connect evidence, where their reasoning breaks down, and whether they can be trusted as defensive decisions become increasingly autonomous.

Loading the standings.

cross-accounts-don-match-current

Are there any Okta accounts that don't match a current employee or contractor in BambooHR?

2 systemsmedium

Open task

cross-accounts-active-employees-terminated

Are there any Google Workspace accounts still active for employees who have been terminated?

3+ systemshard

Open task

Security agents need independent evaluation.

Security teams are handing agents real defensive work, and the scores come from the vendors who built them. Security has solved this before, with independent testing and open scoring standards. We are building the same neutral ground for agents: one frozen synthetic enterprise, the same questions for everyone, and answers anyone can check.

Everything you need to score your agent is public.

Run every task locally with your own keys, then submit a configuration for an official run.