Complete benchmark record
Scenario descriptions, expected check states, criteria, supplied observations and explicit limitations.
Download JSON19 / Open research
Download the exact three-scenario baseline behind the PactVerity AI Agent Evidence Benchmark. The data is deliberately small, machine-readable and designed for independent reproduction rather than promotional scoring.
01 / Downloads
JSON preserves the complete criteria and evidence objects. CSV supports analysis and comparison. The citation file gives researchers and writers a stable attribution record.
Scenario descriptions, expected check states, criteria, supplied observations and explicit limitations.
Download JSONA flat version for spreadsheets, notebooks, data catalogues and reproducibility reports.
Download CSVDataset title, version, release date, repository, licence, schema and file hashes in standard machine-readable formats.
Download CFF02 / Reproduce
Read the expected outcome and check states before executing anything.
Record the source commit, runtime, command, timestamp, output and receipt digests.
Disclose deviations, timeouts, schema mismatches and any relationship to PactVerity.