Synthetic justice datasets for testing justice tech and legal-AI systems.
Curated synthetic matter files with ground truth, answer keys, and human QA — built for evaluation, not storytelling.
Realistic evaluation is hard to source.
Real case files are hard to use
Privacy, procurement, and handling rules make real matters slow or impossible to use for testing.
Generic examples don't hold up
Toy examples don't reflect the contradictions, gaps, and messiness of real legal records.
Access to justice outcomes
Technology only delivers on its promise of a simpler, more accessible justice system if it's tested properly. That's why every scenario is designed by lawyers with real experience in courts, CLCs, pro bono practice, and legal aid.
A deliverable pack, not a tool.
You receive finished outputs and documentation — never the generator, prompts, or an AI interface.
Every synthetic matter ships with a detailed ground truth — a fictional narrative built out with the same depth of supporting evidence you'd expect in a real file.
-
✓
Synthetic matter filesCoherent case file bundles with documents and a working timeline.
-
✓
Clean ingestion bundleStructured folders, consistent filenames, and metadata ready for pipeline tests.
-
✓
Ground truth packA detailed fictional narrative with chronology and a matching evidence index — built to support every fact in the file.
-
✓
Model answers packTask-specific "gold" outputs for benchmarking your system's responses.
-
✓
RegistersContradictions, edge cases, known synthetic defects, and contamination tests, all documented.
-
✓
Optional evaluation toolkitScoring rubric, benchmark tasks, demo scripts, and QA checklists.
Basic vs Premium.
Structured, semi-self-serve
For teams running ingestion tests, demos, and regression suites on a defined timeline.
- Configure a bundle: jurisdiction, category, stage, count, formats
- QA'd delivery pack — ZIP plus documentation
- Standard turnaround, standard QA level
- Ground truth and model answers included
Custom dataset + evaluation design
For high-fidelity, comparable testing across vendors or model versions.
- Custom scenarios and engineered edge cases
- Multi-stage simulations across a matter's lifecycle
- Tailored scoring rubrics and workshops
- Defensible evaluation design, documented end to end
Six steps, start to delivery.
Configure
Jurisdiction, category, stage, count, formats, options, QA level, timing.
Scope confirmation
Bundle summary covering inclusions and exclusions before generation starts.
Generate
Manifests and documents produced to agreed conventions.
Human QA
Coherence and registers reviewed and validated by a person.
Package
Ingestion bundle, truth pack, model answers, and README assembled.
Deliver
ZIP or secure link, with an optional iteration round.
Sample bundle first — one matter.
Before committing to a volume run, start with a single one-matter sample bundle designed to surface defects early: folder structure, filenames, metadata, timeline logic, cross-references, and — if selected — contradictions and edge cases.
Request a sample bundleSample bundle — Family Law Matter
Premium Package test bundles are highly bespoke and, where required, voluminous. A single synthetic family law case file runs to 156 documents and more than 1,600 pages, produced to consistent naming and metadata conventions.
Open five documents from the bundle
These are real files from the synthetic bundle, not mock-ups. Every person, business, address and event in them is fictional.
Full document index
Documents are grouped the way an evaluator would encounter them: file controls first, then orientation material, the law-firm administrative file, court documents, lay and expert evidence, disclosure, correspondence, settlement material and hearing preparation — followed by the evaluator-only ground truth and register pack.
Sample bundle — Corporate Dispute Matter
A synthetic corporate dispute case file — the Meridian BioSystems shareholder and director dispute pack, 74 documents.
- ✓Corporate, financial, governance and communications evidence
- ✓11 document formatsIncluding DOCX, XLSX, PDF, EML, CSV, JSON, HTML and PNG.
- ✓Three embedded edge casesStanding mismatch, profit retention versus related-party extraction, and disputed dilution/digital approval.
- ✓Separate privilege and without-prejudice quarantine
- ✓Excluded control/ground-truth folder
- ✓Model chronology, evidence index and issue-spotting report
- ✓Cleanly recalculated spreadsheets and validated archive
All material is prominently identified as synthetic.
Evidence index
The 49-document ordinary evidence set, grouped by folder. The privilege quarantine, control/ground-truth folder and model outputs (chronology, evidence index, issue-spotting report) bring the pack to 74 files in total and are not listed here.
What this is — and what it isn't.
Transparent, not a calculator.
How a quote is built
Public-interest pricing
Available for courts, government, and nonprofits.
Request a non-profit quoteCommon questions.
What is synthetic justice data?
Fully synthetic case files, documents, and supporting records modelled on real-world legal matter structures — built for testing, with no real litigant data involved.
What can I test with it?
Ingestion pipelines, document classification, extraction, summarisation, chronology-building, and any task where you need a known ground truth to score against.
Do you need my confidential files?
No. Bundles are generated independently of your real case data — that's the point. We don't require access to confidential or privileged materials.
Is this suitable for my jurisdiction?
Jurisdiction is a configuration option. Tell us what you need when you start a Basic order or request a Premium quote, and we'll confirm fit during scope confirmation.
Can I use this for training models?
Permitted use is set out in delivery terms and depends on the licence selected at order time. Ask us directly if model training is your intended use case.
Ready to test?
Synthetic data. No real litigant information.