JARVI3 / ECOKURE

Inspect the check.
Keep the evidence.

Explore the public tools behind an approach to AI verification: scoped checks, explicit unsupported outcomes, evidence receipts and replay. Run a small example, inspect its limits, then evaluate it in your workflow.

Route → check → receipt

ClaimStack demonstrates dimensional checks, explicit refusals and sealed records using Python's standard library.

Run ClaimStack →

Exact arithmetic

DTL MathGate evaluates supported arithmetic with exact values. Its public corpus and regression tests expose the implemented scope.

Inspect MathGate →

Replay and inspect drift

ReplayGate replays supported evidence packs and compares results. Inspect the implementation and its execution limits.

Explore ReplayGate →

Public evidence record

This page records development evidence and submission status. It is not an external leaderboard. ProgramBench PR #27 remains open as of 1 October 2026.
EntryTrackResultStatusEvidence
DTL MathGate lite
Public arithmetic suite
Development corpus and regressions 192 exact answers
48 expected refusals
240 cases · 0 failures
18 unit tests pass
Reproduced 1 Oct 2026

Windows/Linux CI passed. Public tests with a separate in-repository oracle.

Run record
Code and regressions
Jarvi3: Themis-G
ProgramBench stored candidates
Official benchmark submission 2/200 resolved
1,037/1,037 tests passed
Scores reproduced 1 Oct 2026

Official evaluator; registry PR pending. Generation provenance remains unresolved.

Fresh reproduction
Package
Official PR #27
DTL/SuperMath
AIME-style public seed
Case-specific arithmetic replay 14/30 correct overall
0 incorrect
16 abstain
Development evidence

Known case IDs and problem-specific calculations; not an unseen-problem reasoning score.

Full report

The public lite tools implement bounded checks. Passing a check does not establish general scientific truth, production readiness, or validation of every E-stack layer.