ISPM Enterprise SQL
ISPM-Enterprise-SQL@v1
Contract publishedWhat every conforming run holds constant.
These 6 clauses are v1. A change to the task set, the world, the agent-visible surface, the repetition policy or the evidence record is a new revision with a new identifier, not an edit to this one.
Run contract
Every task runs against the bundled ispm-crossvendor-v1 SQLite snapshot. Its manifest names the dataset and publish date but does not yet carry a content digest.
The agent receives the task prompt, SQLite world, and schema documentation. Public ground truth and witness SQL stay outside the agent workspace during the run.
A response must state the conclusion and identify the people, resources, policies, or configuration facts that support it.
Official evaluation runs each task for five epochs with two declared, floating judge aliases; the provider-returned wire model IDs are recorded once at preflight.
Within an epoch the lower non-abstaining judge grade wins. Across epochs, the modal grade wins and ties resolve downward.
Official execution allows two concurrent samples, 15 minutes per sample, and 12 hours for the complete run; a run timeout produces no partial score.
6 judged dimensions, one answer.
A judge grades the response on every dimension below. Structural coverage of the agent’s own SQL is measured deterministically beside them, so a fluent answer cannot erase a weak investigation.
answer_correctnessanswer_verdictevidence_completenesscontextual_evidence_relevancyreasoning_utilitysql_semantic_appropriatenessWhat an official run publishes.
Public run record
Verified Execution Protocol- 01
candidate and resolved model metadata
- 02
aggregate and per-task consensus metrics
- 03
per-epoch judge grades and votes
- 04
token and step accounting
- 05
reliability and consistency statistics
Read the parts.
Run it against this contract.
One binary ships this task set, its pinned dataset and the grader. The score you see locally is the protocol the board runs; a submission is what starts the official run.