19 / Open research

A citation-ready benchmark, with the limits attached.

Download the exact three-scenario baseline behind the PactVerity AI Agent Evidence Benchmark. The data is deliberately small, machine-readable and designed for independent reproduction rather than promotional scoring.

Dataset versionv0.1Known-answer scenarios3Verified external runs0
Research boundary:These are first-party fixtures. They test deterministic evaluation of supplied observations; they do not prove that the observations are independently true or that PactVerity has customers, integrations or audit assurance.

01 / Downloads

Use the format that fits your review workflow.

JSON preserves the complete criteria and evidence objects. CSV supports analysis and comparison. The citation file gives researchers and writers a stable attribution record.

JSONMachine-readable

Complete benchmark record

Scenario descriptions, expected check states, criteria, supplied observations and explicit limitations.

Download JSON
CSVAnalysis-ready

Scenario comparison table

A flat version for spreadsheets, notebooks, data catalogues and reproducibility reports.

Download CSV
CFFCitation

Stable attribution record

Dataset title, version, release date, repository and licence in Citation File Format.

Download citation file

02 / Reproduce

A useful citation starts with an inspectable result.

  1. 01

    Inspect the fixtures

    Read the expected outcome and check states before executing anything.

  2. 02

    Run the public kit

    Record the source commit, runtime, command, timestamp, output and receipt digests.

  3. 03

    Publish failures too

    Disclose deviations, timeouts, schema mismatches and any relationship to PactVerity.

Run the benchmark Read the methodology Open a reproducibility issue