jhana research · Phase I

Legal extraction benchmark

20 HCK matters · 39 values per matter · one frozen task and evidence surface

All bundles
All documents
Hidden gold
Prefill values
Benchmark constitution

What counts as a legal task—and what counts as truth

The 780 imported values are frozen as hidden gold under the benchmark-owner assumption. Models receive only case files, task definitions, and the declared read-only harness condition.

Versioned tasks

Benchmark protocol

v0.1
  1. Freeze the source.PDF, optional graph, task, and gold hashes identify the exact benchmark release.
  2. Hide the answers.Provider requests never contain prefills, annotations, adjudications, or database access.
  3. Hold each condition constant.Systems share the same documents, prompt, submission contract, and runtime policy within each declared tool condition.
  4. Count binary checks.Field, citation, completion, and full-matter passes are reported beside time, tokens, and cost.
Supported Contradicted Conflicting Insufficient

Source register

— pages
MatterFilesPagesGold
Select a field to begin.

Choose a field

The annotation form will follow the selected versioned task field.

Pinpoint evidence

No locators added.

Adjudicate current human submission

Production gold expects two independent submissions. The pilot override remains visibly marked in the rationale.

jhana Evidence Language · v0.1

Compile assertions. Query the evidence graph.

Program

JEL → canonical JSON IR

Compiled proof / query result

Ready
Run the compiler to inspect canonical IR, or execute the program to produce evidence proof objects.

Add reviewed graph fact

Queryable immediately

An absence assertion stays insufficient until its named coverage scope is certified complete. This prevents “not found” from silently becoming “false.”

HCK CIS metadata · 39 fields

Model-and-harness leaderboard

Matters
Hidden gold
Tool conditionsGraph-assisted · PDF + vision
Dataset
Highest verifiedAwaiting runs
Best field accuracyAwaiting runs
Fastest complete runAwaiting runs
Lowest API costAwaiting runs
01Frozen PDFs + hidden gold
02Same 39-field prompt
03Declared tool condition
04Hidden binary checks
05Accuracy + efficiency

Ranked systems

Verified score = passed fields + passed citations · efficiency stays separate
System / modelVerified scoreField accuracyEvidenceMatters acceptTimeHarnessTokensEst. cost

Benchmark runs are being prepared.

Accuracy and evidence

Binary obligations passed

Completed systems will appear here.

Quality–efficiency frontier

Where models still fail

Lowest mean field accuracy

Field gaps appear after scoring.

Field-by-field performance

39 extraction obligations

The field matrix appears after the first full run.

Latest score report

Waiting for an envelope
Every displayed percentage is a count of deterministic pass/fail checks across the frozen corpus.
Post-run audit · hidden gold

Trace Inspector

Inspect every matter and system. Recorded outputs, evidence, tool contract, and deterministic verifier steps are shown without exposing gold during inference.

Choose a case file

Trace record

Model and harness

Choose a field to compare the submitted value with frozen gold.