Skip to content

Run

7 emulators 954 tests

Results as of this run. The arrow shows each target's movement since the previous run it was tested in.

When this run was published the board led with a single correctness percentage. The figures here are that run's own results restated as divergence and coverage, so they read differently from what this page showed at the time. The changelog has the change.

What changed in the suite this run

Suite on

Per-region scoring lands complete. 2.0.0-pre put the scoring logic in place, comparing each target against every region's recorded answer, but the evidence half was never wired: no test recorded what a target actually answered and the classifier never read one, so a fail could not be credited to a region the target matched and the score could only ever subtract. 2.0.0 closes that loop, and the seed split runs its whole lifecycle in the same release.

What changed:

  • The split tests now record what the target actually answered (src/observation-sink.ts), in the same shape the registry stores each region's answer, and the classifier carries it onto the verdict. An engine that matches a rejecting region on a split is now scored as passing in that region. Evidence only ever redeems a committed fail: a pass keeps the committed assertion's deliberate wording tolerance and is never held to the byte-exact recorded string. Committed results predate the capture, so published scores move on each target's next run, not in this change.
  • Headline ties now prefer a region the registry characterises. A region absent from every split row ties the top score by having nothing recorded about it, and the Region column must not answer "conformant to what?" with a region the suite knows nothing about. Only the label is affected; a strictly higher score still wins whatever its source.
  • The seeded { NULL: false } split closed its own loop. eu-west-2 and eu-central-1 were the last regions accepting it; by the 2026-07-17 sweep both reject it again, so every region agrees, the split is retired, and its test asserts the shared rejection. The detect, admit and reconcile path the pre-release introduced, exercised end to end within days.
  • A tooling test now asserts every tracked file is text, after a raw NUL byte used as a delimiter in one source file made grep classify the file as binary and silently skip it.
Scored against 33 real regions. 1 dropped, out of scoring until it returns: me-south-1
  1. live (AWS)

    Runs via: AWS service

    Grade baseline

    no divergence in all 33 regions · covers 100.0% of the suite

    ground truth
    Tier breakdown
    Tier 1 · Core 0.0% diverges
    100.0% covered
    Tier 2 · Complete 0.0% diverges
    100.0% covered
    Tier 3 · Strict 0.0% diverges
    100.0% covered
  2. Runs via: npx, Docker, Homebrew, binary, npm, cargo, embedded, GitHub Action, source

    Disclosure: maintained by this board's author

    Grade A

    low divergence (0.7%) in 6 regions · up to 1.0% in the other 27 · covers 98.5% of the suite

    diverged 0.1 percentage points more
    Tier breakdown
    Tier 1 · Core 0.0% diverges
    100.0% covered
    Tier 2 · Complete 0.6% diverges
    91.2% covered 14 unsupported
    Tier 3 · Strict 1.9% diverges
    100.0% covered
  3. 0b5dba5153fc

    Runs via: pip, Docker, source

    Grade B

    moderate divergence (14.4%) in all 33 regions · covers 100.0% of the suite

    diverged 0.1 percentage points less
    Tier breakdown
    Tier 1 · Core 10.9% diverges
    100.0% covered
    Tier 2 · Complete 12.6% diverges
    100.0% covered
    Tier 3 · Strict 20.3% diverges
    100.0% covered
  4. Runs via: Docker, pip, Homebrew, binary

    Grade C

    moderate divergence (14.8%) in 27 regions · up to 14.9% in the other 6 · covers 99.2% of the suite · coverage lowers this row to C

    diverged 0.2 percentage points less
    Tier breakdown
    Tier 1 · Core 6.5% diverges
    100.0% covered
    Tier 2 · Complete 13.2% diverges
    95.0% covered 8 unsupported
    Tier 3 · Strict 27.8% diverges
    100.0% covered
  5. Grade C

    high divergence (15.9%) in 27 regions · up to 16.0% in the other 6 · covers 98.6% of the suite

    diverged 0.2 percentage points less
    Tier breakdown
    Tier 1 · Core 7.8% diverges
    100.0% covered
    Tier 2 · Complete 15.7% diverges
    91.8% covered 13 unsupported
    Tier 3 · Strict 28.1% diverges
    100.0% covered
  6. Grade C

    high divergence (16.0%) in all 33 regions · covers 95.5% of the suite

    diverged 0.2 percentage points less
    Tier breakdown
    Tier 1 · Core 14.5% diverges
    100.0% covered
    Tier 2 · Complete 10.1% diverges
    73.0% covered 43 unsupported
    Tier 3 · Strict 21.3% diverges
    100.0% covered
  7. Runs via: Docker, Homebrew, install script, Scoop, binary, JAR

    Grade C

    high divergence (17.6%) in all 33 regions · covers 99.1% of the suite

    diverged 0.1 percentage points less
    Tier breakdown
    Tier 1 · Core 8.6% diverges
    100.0% covered
    Tier 2 · Complete 22.6% diverges
    94.3% covered 9 unsupported
    Tier 3 · Strict 28.4% diverges
    100.0% covered
  8. Grade C

    high divergence (22.1%) in 27 regions · up to 22.3% in the other 6 · covers 92.9% of the suite

    diverged 0.3 percentage points less
    Tier breakdown
    Tier 1 · Core 8.6% diverges
    100.0% covered
    Tier 2 · Complete 49.7% diverges
    57.2% covered 68 unsupported
    Tier 3 · Strict 28.4% diverges
    100.0% covered