Skip to content
Grade A
low divergence
answers 0.2% of the suite differently from real DynamoDB
up to 0.5% in the others
88.8% of the suite implemented · 140 tests it doesn't attempt

Disclosure. Dynoxide (wasm) is maintained by the author of this board. Its score is produced by the same automated tests as every other engine, derived from the conformance suite's published results, and isn't adjusted by hand.

Tier 1 · Core 0.0% diverges
100.0% covered
Tier 2 · Complete 0.0% diverges
71.3% covered 99 unsupported
Tier 3 · Strict 0.5% diverges
90.0% covered 41 unsupported
2 of 369 attempted

How to run it

Every way Dynoxide (wasm) is distributed, each linking to the project's own documentation for it. These are the project's claims, not the suite's measurements.

Needs: a browser, on a secure context

Per-region divergence

The headline is eu-west-2 + 15 regions, measured against 16 of 33 observed regions, and scored in ap-east-1. Real DynamoDB doesn't behave identically in every region, so a target can match some regions more closely than others. Each figure below is divergence in those regions, best first. eu-west-2 is marked as the historical baseline, and a region that couldn't be resolved this sweep is flagged rather than counted as a disagreement.

0.2% 16 regions
ap-east-1ap-east-2ap-northeast-2ap-northeast-3ap-south-1ap-southeast-1ap-southeast-2ca-west-1eu-central-1eu-west-2 · baselineeu-west-3mx-central-1sa-east-1us-east-1us-west-1us-west-2
0.2% 6 regions
af-south-1ap-southeast-3ap-southeast-4ap-southeast-5eu-north-1il-central-1
0.4% 10 regions
ap-northeast-1ap-south-2ap-southeast-6ap-southeast-7ca-central-1eu-central-2eu-south-1eu-south-2eu-west-1us-east-2
0.5% 1 region
me-central-1

Fully supports

Operation areas this target implements completely - every test passes, nothing skipped. Areas it only partly gets wrong are in “Where it falls short” below; the ones it only partly attempts are in “What it doesn’t attempt”.

  • batchGetItem tier 1
  • batchWriteItem tier 1
  • createTable tier 1
  • deleteItem tier 1
  • deleteTable tier 1
  • describeTable tier 1
  • getItem tier 1
  • listTables tier 1
  • putItem tier 1
  • query tier 1
  • scan tier 1
  • updateItem tier 1
  • updateTable tier 1
  • updateTable tier 2
  • vectorSearch tier 2
  • legacy-api tier 3
  • validation-ordering tier 3

Divergence and coverage over time

Both are plotted because neither reads correctly alone. Divergence falls when a target stops attempting an operation it used to get wrong, so a fall in the first plot is only an improvement if the second one holds.

Divergence

Dynoxide (wasm) divergence, 2026-07-24 to 2026-09-20 0% 1% 2% divergence - lower is better 24 Jul 12 Aug 13 Aug 15 Aug 17 Aug 18 Aug 21 Aug 23 Aug 25 Aug 30 Aug 3 Sep 13 Sep 20 Sep 0.2%
Diverging on 0.2% of the suite as of 20 Sep, at worst 0.9% on 12 Aug. Lower is better, so a falling line is a target getting closer to real DynamoDB.

Coverage

Dynoxide (wasm) coverage, 2026-07-24 to 2026-09-20 75% 88% 100% coverage - higher is better 24 Jul 12 Aug 13 Aug 15 Aug 17 Aug 18 Aug 21 Aug 23 Aug 25 Aug 30 Aug 3 Sep 13 Sep 20 Sep 88.8%
Implementing 88.8% of the suite as of 20 Sep, as little as 78.7% on 24 Jul. Higher is better. This is the share of the suite's tests a target implements at all. A fail becoming a skip leaves both numerators over the same fixed denominator, so a fall here is matched point for point by a fall in divergence: withdrawal costs exactly as much coverage as it gains divergence.

By operation

Every operation this target implements, with how much of it diverges and how much of it it covers, on the same two axes as the figures above. The counts are fails over the operation's whole size. full, partial, failing, unsupported. The matrix compares operations across targets.

Tier 1 - Core

  • batchGetItem 0/15 0.0% 100.0% batchGetItem: supported, diverges on 0.0% of it, covers 100.0% (15 pass)
  • batchWriteItem 0/18 0.0% 100.0% batchWriteItem: supported, diverges on 0.0% of it, covers 100.0% (18 pass)
  • createTable 0/30 0.0% 100.0% createTable: supported, diverges on 0.0% of it, covers 100.0% (30 pass)
  • deleteItem 0/13 0.0% 100.0% deleteItem: supported, diverges on 0.0% of it, covers 100.0% (13 pass)
  • deleteTable 0/3 0.0% 100.0% deleteTable: supported, diverges on 0.0% of it, covers 100.0% (3 pass)
  • describeTable 0/4 0.0% 100.0% describeTable: supported, diverges on 0.0% of it, covers 100.0% (4 pass)
  • getItem 0/39 0.0% 100.0% getItem: supported, diverges on 0.0% of it, covers 100.0% (39 pass)
  • listTables 0/5 0.0% 100.0% listTables: supported, diverges on 0.0% of it, covers 100.0% (5 pass)
  • putItem 0/129 0.0% 100.0% putItem: supported, diverges on 0.0% of it, covers 100.0% (129 pass)
  • query 0/95 0.0% 100.0% query: supported, diverges on 0.0% of it, covers 100.0% (95 pass)
  • scan 0/60 0.0% 100.0% scan: supported, diverges on 0.0% of it, covers 100.0% (60 pass)
  • updateItem 0/70 0.0% 100.0% updateItem: supported, diverges on 0.0% of it, covers 100.0% (70 pass)
  • updateTable 0/15 0.0% 100.0% updateTable: supported, diverges on 0.0% of it, covers 100.0% (15 pass)

Tier 2 - Complete

  • account 0/2 · 2 skip n/a 0.0% account: unsupported, implements none of it (2 skip)
  • backups 0/5 · 5 skip n/a 0.0% backups: unsupported, implements none of it (5 skip)
  • contributorInsights 0/2 · 2 skip n/a 0.0% contributorInsights: unsupported, implements none of it (2 skip)
  • export 0/2 · 2 skip n/a 0.0% export: unsupported, implements none of it (2 skip)
  • kinesis 0/1 · 1 skip n/a 0.0% kinesis: unsupported, implements none of it (1 skip)
  • partiql 0/186 · 1 skip 0.0% 99.5% partiql: partially supported, diverges on 0.0% of it, covers 99.5% (185 pass, 1 skip)
  • resourcePolicy 0/2 · 2 skip n/a 0.0% resourcePolicy: unsupported, implements none of it (2 skip)
  • streams 0/18 · 18 skip n/a 0.0% streams: unsupported, implements none of it (18 skip)
  • tags 0/8 · 8 skip n/a 0.0% tags: unsupported, implements none of it (8 skip)
  • transactions 0/62 · 51 skip 0.0% 17.7% transactions: partially supported, diverges on 0.0% of it, covers 17.7% (11 pass, 51 skip)
  • ttl 0/7 · 7 skip n/a 0.0% ttl: unsupported, implements none of it (7 skip)
  • updateTable 0/14 0.0% 100.0% updateTable: supported, diverges on 0.0% of it, covers 100.0% (14 pass)
  • vectorSearch 0/36 0.0% 100.0% vectorSearch: supported, diverges on 0.0% of it, covers 100.0% (36 pass)

Tier 3 - Strict

  • error-messages 0/184 · 31 skip 0.0% 83.2% error-messages: partially supported, diverges on 0.0% of it, covers 83.2% (153 pass, 31 skip)
  • legacy-api 0/42 0.0% 100.0% legacy-api: supported, diverges on 0.0% of it, covers 100.0% (42 pass)
  • limits 2/150 · 10 skip 1.3% 93.3% limits: partially supported, diverges on 1.3% of it, covers 93.3% (138 pass, 2 fail, 10 skip)
  • validation-ordering 0/34 0.0% 100.0% validation-ordering: supported, diverges on 0.0% of it, covers 100.0% (34 pass)

Where it falls short

Operations Dynoxide (wasm) implements and answers differently from real DynamoDB, biggest gap first. This is what its divergence figure counts. Open one for the exact tests, or see the conformance suite.

  • limits Tier 3 2 diverging 10 not attempted
    View these tests in the suite →
    • Item size limit by surface - UpdateItem charges per action does not charge for an attribute the statement leaves alone source:362
    • Item size limit by surface - UpdateItem through a document path charges a SET through a list index exactly one byte more source:627
    • Item size limit by surface - TransactWriteItems a Put accepts an item of exactly 409,600 bytes
    • Item size limit by surface - TransactWriteItems a Put one byte over raises ValidationException, not a cancellation
    • Item size limit by surface - TransactWriteItems an Update accepts a finished item of exactly 409,600 bytes
    • Item size limit by surface - TransactWriteItems an Update one byte over surfaces as a cancellation, with the update wording
    • Item size limit by surface - TransactWriteItems answers the same way for a fresh key and for an existing row
    • Item size limit by surface - a transacted Update does not inherit the exclusion writes what UpdateItem refuses at the same key
    • TransactWriteItems limits TransactWriteItems with exactly 100 Put actions succeeds
    • TransactWriteItems limits TransactWriteItems with 101 actions fails with ValidationException
    • TransactWriteItems limits TransactWriteItems total item size approaching 4MB succeeds
    • TransactWriteItems limits TransactWriteItems total item size over 4MB fails with ValidationException

What it doesn't attempt

Operations Dynoxide (wasm) declines rather than gets wrong, so none of this counts towards its divergence. It is what the gap in its 88.8% coverage is made of. Each test here skipped itself because the target's own feature probe said the operation isn't implemented, which is often a deliberate choice rather than a defect.

  • transactions Tier 2 partly 51 not attempted
    View these tests in the suite →
    • TransactWriteItems - basic functionality executes Put + Update + Delete atomically
    • TransactWriteItems - basic functionality commits a Put whose string set carries an empty member
    • TransactWriteItems - basic functionality succeeds when ConditionCheck condition is met
    • TransactWriteItems - basic functionality rolls back entire transaction when ConditionCheck fails
    • TransactWriteItems - basic functionality applies ConditionExpression on Put action
    • TransactWriteItems - basic functionality applies ConditionExpression on Update action
    • TransactWriteItems - basic functionality applies ConditionExpression on Delete action
    • TransactWriteItems - basic functionality executes cross-table transaction (hashTableDef + compositeTableDef)
    • TransactWriteItems - basic functionality supports idempotency via ClientRequestToken
    • TransactWriteItems - basic functionality rejects same ClientRequestToken with different payload
    • TransactWriteItems - basic functionality includes CancellationReasons in error when condition fails
    • TransactWriteItems - basic functionality returns ALL_OLD item via ReturnValuesOnConditionCheckFailure
    • TransactWriteItems - multiple items puts multiple items in one transaction
    • TransactWriteItems - multiple items updates multiple items in one transaction
    • TransactWriteItems - multiple items deletes multiple items in one transaction
    • TransactWriteItems - validation rejects duplicate target items in same transaction
    • TransactWriteItems - validation rejects empty TransactItems
    • TransactWriteItems - validation rejects transaction on non-existent table
    • TransactWriteItems - validation Put with a wrong-typed table key cancels with a ValidationError reason
    • TransactWriteItems - validation Put with a non-scalar table key cancels with a ValidationError reason
    • TransactWriteItems - validation Put with an empty-string table key is a top-level ValidationException
    • TransactWriteItems - validation Update with an empty-string Key is a top-level ValidationException
    • TransactWriteItems - validation Delete with an empty-string Key is a top-level ValidationException
    • TransactWriteItems - validation ConditionCheck with an empty-string Key is a top-level ValidationException
    • TransactWriteItems - validation Update with a wrong-typed Key cancels with a ValidationError reason
    • TransactWriteItems - validation Update with a non-scalar Key cancels with a ValidationError reason
    • TransactWriteItems - validation Delete with a wrong-typed Key cancels with a ValidationError reason
    • TransactWriteItems - validation Delete with a non-scalar Key cancels with a ValidationError reason
    • TransactWriteItems - validation ConditionCheck with a wrong-typed Key cancels with a ValidationError reason
    • TransactWriteItems - validation ConditionCheck with a non-scalar Key cancels with a ValidationError reason
    • TransactWriteItems - validation Update with attribute_exists rejects non-existent item
    • TransactWriteItems - validation Update with attribute_not_exists upserts on non-existent key
    • TransactWriteItems - validation Update with comparison condition cancels on non-existent key; no ghost item
    • TransactWriteItems - validation Update with combined attribute_exists + equality cancels on non-existent key; no ghost item
    • TransactWriteItems - validation mixed transaction: one passing, one failing on non-existent cancels everything
    • TransactWriteItems - ConditionExpression parens Put with per-condition parens succeeds on fresh key
    • TransactWriteItems - ConditionExpression parens Update with full-expression wrap succeeds when condition holds
    • TransactWriteItems - ConditionExpression parens Delete with non-redundant nested parens succeeds when condition holds
    • TransactWriteItems - ConditionExpression parens ConditionCheck with per-condition parens passes when condition holds
    • TransactWriteItems - ConditionExpression parens cancels transaction when any parenthesised condition fails
    • TransactWriteItems - ConsumedCapacity charges 2 write capacity units per item
    • TransactWriteItems - ConsumedCapacity: conditional, check, replay, cancel a passing conditional write costs the same 2 WCU/item as an unconditional one
    • TransactWriteItems - ConsumedCapacity: conditional, check, replay, cancel a standalone ConditionCheck action costs 2 write capacity units
    • TransactWriteItems - ConsumedCapacity: conditional, check, replay, cancel reports write capacity on the first call and read capacity on a same-token replay
    • TransactWriteItems - ConsumedCapacity: conditional, check, replay, cancel a cancelled conditional transaction reports no consumed capacity
    • TransactWriteItems - index key validation Put with a wrong-typed index key cancels with a ValidationError reason
    • TransactWriteItems - index key validation Put with a non-scalar index key cancels with a ValidationError reason
    • TransactWriteItems - index key validation Update setting a wrong-typed index key cancels with a ValidationError reason
    • TransactWriteItems - index key validation Update setting a non-scalar index key cancels with a ValidationError reason
    • TransactWriteItems - index key validation Put with an empty-string index key is a top-level ValidationException
    • TransactWriteItems - index key validation Update setting an empty-string index key is a top-level ValidationException
  • error-messages Tier 3 partly 31 not attempted
    View these tests in the suite →
    • Conditional check - exact error messages TransactionCanceledException message format
    • TransactWriteItems - index key error messages Put wrong-typed index key: cancelled with full ValidationError reason
    • TransactWriteItems - index key error messages Put non-scalar index key: cancelled with full ValidationError reason
    • TransactWriteItems - index key error messages Update wrong-typed index key: cancelled with full ValidationError reason
    • TransactWriteItems - index key error messages Update non-scalar index key: cancelled with full ValidationError reason
    • TransactWriteItems - index key error messages Put empty-string index key: top-level ValidationException
    • TransactWriteItems - index key error messages Update empty-string index key: top-level ValidationException
    • TransactWriteItems - index key error messages Put empty-binary index key: top-level ValidationException
    • TransactWriteItems - index key error messages Update empty-binary index key: top-level ValidationException
    • TransactWriteItems - exact error messages empty TransactItems: full minimum-length error
    • TransactWriteItems - exact error messages > 100 actions: anchored regex on the constraint phrase
    • TransactWriteItems - exact error messages duplicate target keys in same transaction: full multi-op error
    • TransactWriteItems - exact error messages non-existent table: full ResourceNotFoundException message
    • TransactWriteItems - exact error messages two failing actions: multi-reason TransactionCanceledException
    • TransactWriteItems - exact error messages one passing, one failing: positional reason codes (None for the survivor)
    • TransactWriteItems - exact error messages Put wrong-typed table key: cancelled with full ValidationError reason
    • TransactWriteItems - exact error messages Put non-scalar table key: cancelled with full ValidationError reason
    • TransactWriteItems - exact error messages Put empty-string table key: top-level ValidationException
    • TransactWriteItems - exact error messages Update empty-string Key: top-level empty-value message
    • TransactWriteItems - exact error messages Delete empty-string Key: top-level empty-value message
    • TransactWriteItems - exact error messages ConditionCheck empty-string Key: top-level empty-value message
    • TransactWriteItems - exact error messages Update wrong-typed Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages Update non-scalar Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages Delete wrong-typed Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages Delete non-scalar Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages ConditionCheck wrong-typed Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages ConditionCheck non-scalar Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages Put empty-binary item key: top-level empty-value message
    • TransactWriteItems - exact error messages Update empty-binary Key: top-level empty-value message
    • TransactWriteItems - exact error messages Delete empty-binary Key: top-level empty-value message
    • TransactWriteItems - exact error messages ConditionCheck empty-binary Key: top-level empty-value message
  • streams Tier 2 none of it 18 not attempted
    View these tests in the suite →
    • DynamoDB Streams - basic table with StreamSpecification has LatestStreamArn in DescribeTable
    • DynamoDB Streams - basic table StreamSpecification.StreamEnabled is true
    • DynamoDB Streams - basic table StreamSpecification.StreamViewType matches what was requested
    • DynamoDB Streams - basic ListStreams returns the stream for our table
    • DynamoDB Streams - basic ListStreams with TableName filter returns only our table's stream
    • DynamoDB Streams - basic DescribeStream returns stream status ENABLED or ENABLING
    • DynamoDB Streams - basic DescribeStream returns at least one shard
    • DynamoDB Streams - basic each shard has a ShardId
    • DynamoDB Streams - basic TRIM_HORIZON returns a valid iterator string
    • DynamoDB Streams - basic LATEST returns a valid iterator string
    • DynamoDB Streams - basic GetRecords after PutItem contains the new image (INSERT event)
    • DynamoDB Streams - basic INSERT record has eventName INSERT
    • DynamoDB Streams - basic INSERT record dynamodb.Keys contains the item key
    • DynamoDB Streams - basic INSERT record dynamodb.NewImage contains the full item
    • DynamoDB Streams - basic GetRecords after UpdateItem contains both old and new images (MODIFY event)
    • DynamoDB Streams - basic GetRecords after DeleteItem contains old image (REMOVE event)
    • DynamoDB Streams - basic NEW_IMAGE view type records have NewImage but no OldImage on MODIFY
    • DynamoDB Streams - basic KEYS_ONLY view type records have Keys but no NewImage or OldImage
  • tags Tier 2 none of it 8 not attempted
    View these tests in the suite →
    • Tags - basic adds tags to a table
    • Tags - basic lists tags and verifies they match what was added
    • Tags - basic adds additional tags and verifies all tags are present
    • Tags - basic removes specific tag keys with UntagResource
    • Tags - basic verifies removed tags are gone after untag
    • Tags - basic overwrites an existing tag with the same key but different value
    • Tags - validation rejects TagResource with an invalid ARN format
    • Tags - validation rejects ListTagsOfResource with a non-existent ARN
  • ttl Tier 2 none of it 7 not attempted
    View these tests in the suite →
    • TTL - basic enables TTL on a table
    • TTL - basic DescribeTimeToLive returns ENABLED status and correct attribute name after enabling
    • TTL - basic DescribeTimeToLive returns DISABLED on a table with no TTL configured
    • TTL - basic enables TTL with a different attribute name
    • TTL - validation rejects empty attribute name
    • TTL - validation rejects UpdateTimeToLive on non-existent table
    • TTL - validation rejects DescribeTimeToLive on non-existent table
  • backups Tier 2 none of it 5 not attempted
    View these tests in the suite →
    • Continuous backups - PITR reports PITR DISABLED by default
    • Continuous backups - PITR enabling PITR transitions PointInTimeRecoveryStatus to ENABLED
    • On-demand backups - lifecycle and restore CreateBackup → DescribeBackup → ListBackups → DeleteBackup
    • On-demand backups - lifecycle and restore RestoreTableFromBackup initiates a restore into a new table
    • On-demand backups - lifecycle and restore DescribeBackup on a deleted backup throws BackupNotFoundException
  • account Tier 2 none of it 2 not attempted
    View these tests in the suite →
    • Account reads - DescribeLimits, DescribeEndpoints DescribeLimits returns positive account and table capacity limits
    • Account reads - DescribeLimits, DescribeEndpoints DescribeEndpoints returns at least one endpoint with an address
  • contributorInsights Tier 2 none of it 2 not attempted
    View these tests in the suite →
    • Contributor insights - enable/describe/list reports DISABLED by default
    • Contributor insights - enable/describe/list enabling transitions the status and lists the table
  • export Tier 2 none of it 2 not attempted
    View these tests in the suite →
    • Export and import - S3 ExportTableToPointInTime initiates an export and reports it
    • Export and import - S3 ImportTable ingests S3 data into a new table
  • resourcePolicy Tier 2 none of it 2 not attempted
    View these tests in the suite →
    • Resource policies - Put/Get/Delete GetResourcePolicy on a table with no policy throws PolicyNotFoundException
    • Resource policies - Put/Get/Delete Put then Get round-trips the policy, and Delete removes it
  • kinesis Tier 2 none of it 1 not attempted
    View these tests in the suite →
    • Kinesis streaming destination enables a streaming destination and reports it via Describe
  • partiql Tier 2 partly 1 not attempted
    View these tests in the suite →
    • BatchExecuteStatement - PartiQL ReturnValuesOnConditionCheckFailure on a batch member does return the item on a TransactWriteItems ConditionCheck

Run history

Every percentage here is divergence, per tier and over the whole suite, so lower is better in each column. Coverage is the exception it is named as.

Run Gradecurrent criteria Divergence Movement
Grade A 0.2% unchanged
Grade A 0.2% unchanged

Suite on

Thirteen captures against eu-west-2 turned into coverage, and the suite grew from 1056 tests to 1251. All of it was ground truth nothing asserted, so an engine could discard a PartiQL index qualifier or store an item AWS would refuse and still score a clean pass.

  • PartiQL index qualifier. What a qualified SELECT returns, when an unprojected attribute is rejected and when it is served from the base table, how the reach-back is charged, that a qualified write is rejected, and the exact wording of each rejection.
  • PartiQL comparison on non-scalar types. Equality and inequality on every set, map and list, each checked against the same predicate written as a ConditionExpression.
  • BatchExecuteStatement member fields. Per-member ConsistentRead, the primary-key requirement, and which failures echo TableName.
  • The 400KB gate, to the byte. 409,600 accepted and 409,601 refused across seven write surfaces. UpdateItem is the outlier: it charges what the statement writes plus a fixed cost per clause, so the key and every untouched attribute stay out of the figure.
  • A number's byte cost. Twenty-three literals, each measured by an accept and a refusal either side of the gate. The rule the developer guide documents is wrong in two ways, and both have a case.
  • Reads are sized before the filter. Scan, Query and PartiQL all charge before the WHERE clause and the projection, so discarding rows saves nothing.
  • Vector write capacity. VectorWriteRequestBytes pinned term by term, charged on the change to the index's own stored view like a GSI. VectorSearchRequestBytes stays shape-only: five identical searches against an unchanged index reported 14214, 13903, 14214, 14214 and 14518.
  • Vector index creation has two phases. Resource allocation, where the table sits in UPDATING, Backfilling reads false and a cancel is refused, then backfilling, which takes one. The one-online-index limit binds the table rather than the call.

Two fixes came with it. A local run against real DynamoDB takes its own table namespace, so it can no longer delete a CI run's tables. And a TransactWriteItems control, plus the whole PartiQL validation-ordering file, were feeding an unsupported fault into an assertion about wording and scoring a declared gap as a disagreement; both now skip.

BatchWriteItem with an empty RequestItems map is now a recorded regional split. eu-north-1 answers the validation framework's generic constraint message where the other 32 answering regions answer the bespoke required-parameter sentence. A capture on 2026-08-17 found every answering region on the bespoke wording, which dates the crossing to that week rather than guessing at it.

The over-25-requests assertion beside it spans both cohorts instead of becoming a second row. Neither wording is byte-stable: the table name carries a per-run suffix, the old cohort echoes all 26 requests back, and seven regions echo them as a JVM object identity rather than expanded fields. A row records one verbatim answer per region and there is nothing verbatim here to record, so the anchored pattern grew a second branch instead. It still refuses a wrong limit and a wrong table, and it matches every one of the 33 answering regions.

The capture harness can take that evidence without creating tables. --probes selects named probes and --no-tables skips the fixtures, which is all these three need: DynamoDB refuses each of them before it looks at a table, so the answers do not depend on one existing, and read-only credentials are enough. The 2026-08-17 capture was taken with a scoped script that was never committed; this is the committed way to reproduce it.

Grade A 0.2% diverged 0.2 percentage points more
Grade A 0.0% unchanged
Grade A 0.0% diverged 0.9 percentage points less
Grade B 0.9% unchanged

Suite on

AWS corrected the vector index readiness documentation, prompted by a write-up of the earlier guidance that drew on the suite's measurements. The ACTIVE-plus-backfilling state the old advice was built around, and which no index ever occupies, is gone from the three pages that described it. The wait now reads "Backfilling is not true" rather than "is false", so a check written literally from it fires on both creation paths instead of neither. The tutorial no longer says a search during backfill can return incomplete results. Two things the suite had measured but nobody had written down are documented as well: that DescribeTable reporting ACTIVE leads the dedicated search endpoint, and that the readiness check depending on neither status field is a real search in a retry loop.

That contract is now pinned rather than described. The UpdateTable walk asserts that Backfilling true is only ever reported alongside CREATING, that an ACTIVE index reports no Backfilling field at all, and that the base table goes ACTIVE while the index is still building, which is what makes a table waiter the wrong gate for a search. The first search that succeeds has to carry every seeded item, since the backfill window answers with an error rather than a partial view. On the CreateTable path a new test runs the documented check the way an application would, and every rejection before the first served response has to be the retryable ValidationException rather than a not-found.

The suite's own search wait now absorbs those two rejections and rethrows every other answer, so a fixture waiting on an index that is ACTIVE but not yet served no longer fails on the lag it was waiting out. Two files asserting exact rejection messages wait for a served search rather than for ACTIVE: "does not have the specified index" is also what a freshly ACTIVE index says, and it would otherwise stand in for whichever message the case asked for.

The tutorial's other new claim, that a table cannot be deleted while a vector index is being created, has a test of its own. Only the UpdateTable path can ask it. Across three runs an index created with its table reached ACTIVE in the same 250ms poll as the table, so the table is never ACTIVE with the index still building there, and a DeleteTable during creation is refused for the table's own status rather than with the documented index wording. Adding an index to a live table opens that window about thirty seconds in. The test cancels the index afterwards instead of waiting out the backfill, which a still-creating vector index turns out to accept the way a backfilling GSI does, so it costs a minute rather than seventeen.

A release now dispatches its measurement as the results bot rather than with the workflow's own token. GitHub raises no workflow_run event when a run started by GITHUB_TOKEN finishes, so 3.2.0 measured green for three hours and its board only landed once the results table was dispatched by hand.

Grade B 0.9% unchanged

Suite on

The board now measures the most recent release tag rather than main. Merging a test still runs it against real AWS, which is what validates it, but the published figures no longer move until a release moves them, so a dated changelog entry always sits behind a change in the denominator. Expect the figures to shift on this release: the board is switching from measuring main to measuring a tag, and those are different trees today.

Every board now says what produced it. results/summary.json carries a suite block naming the ref measured, its commit, the suite version at that ref, the region it ran against and when. It is additive, so schemaVersion stays at 1. The same block reaches /data/latest.json, /data/index.json and /data/runs.json, where each historical run carries its own copy, so a denominator that moved between two runs can be attributed to the release that moved it. Branch on kind: only tag is a released board.

A board is graded against the suite manifest and split registry as they stood at the ref it measured. Region health is the exception and is read live, so a region dropped since the tag still counts against today's cohorts. That is the one input allowed to move under a board without a new measurement, and the board carries a health date beside its measurement date to say so.

Releases are cut by one workflow dispatch: it bumps the version, dates this section, installs against the bumped tree, tags, and opens a draft release, then starts the measurement. The draft publishes itself when the board carrying that version lands, which takes about three hours.

eu-west-2 has crossed to the validation framework's generic constraint message for BatchGetItem with an empty RequestItems map, so the split row recording that behaviour is re-characterised against a 34-region capture taken on 2026-08-17. Eleven of the thirty-three regions that answer now return the generic message; twenty-two still return the bespoke required-parameter sentence, and me-central-1 joins the row.

The matching validation-ordering row is retired. Both wordings refuse an empty RequestItems before the map is read, which is all that tier asserts, so the assertion now matches the parameter name case-insensitively and spans both. The BatchWriteItem case beside it, which no region has moved yet, is written the same way.

A probe absent from the baseline is no longer reported as drift. Adding a probe to the capture script leaves every older baseline without it, and a scheduled red then named the new probe as the thing that had moved.

The weekly cross-region capture now includes eu-west-2, so the drift lens reads its baseline and the candidate regions from one capture taken at one moment rather than comparing today's candidates against an older baseline file. A scheduled red also keeps the eu-west-2 capture its drift verdict was read from, which was previously discarded with the runner.

Grade B 0.9% unchanged
Grade B 0.9% unchanged

Suite on

ExtendDB's SQLite backend joins the run, built from the same release as the PostgreSQL one and held to the same TLS, SigV4 and IAM posture, so the storage engine is the only thing that differs between them.

A project's other builds now sit behind a disclosure on its row. Every build is measured in full and has a row of its own with its own figures; the disclosure starts closed only when every build under it reads the same grade, divergence and coverage as the row above, and only when each of them was measured in that run: a carried row on either side opens it, and so does a run the suite declined to score. It is read from each run, so a build can start closed on one and open on the next. The README table has no disclosure to offer, so it lists every build outright.

Every target in the data endpoints gains two fields. collapsedIntoProject is true when the board starts that build's row closed. standsForProject says which row the board treats as a project's own, which isVariant cannot answer: on a run where a project's reference build recorded nothing, a build is promoted to stand for it and every row of that project reads isVariant: true. Both are additive, so schemaVersion is unchanged, and /data/index.json gains a projects block documenting them.

/data/index.json also gains a schema block saying what the version number entitles a consumer to: a field you already read will not change type or meaning while schemaVersion stays put, and new fields can appear at any version. The grading criteria are named there as a separate axis, since a change to the bands changes what a letter means without the schema moving.

The coverage description said a partial run was carried forward on its last clean measurement. It is not: the row stays in the run, publishes null for both figures and reads carried: false, because the target did report. Carrying forward is what happens to the baseline's unobserved lanes. Corrected, along with Dynoxide's build label, which now names its storage (native SQLite) like every other build in the registry rather than its compile target.

Grade B 0.9% unchanged
Grade B 0.9% unchanged

Suite on

Read this first if you consume the JSON. The data endpoints go from schema 2 to schema 4 in one step. Schema 3 was never published on its own, so everything on 2 crosses both steps at once. Schema 3 breaks in four ways:

  • movement.state values changed from up/down to improved/regressed. This is the one that fails silently: the old names still parse, and now mean the opposite direction.
  • movement.delta keeps its shape but is computed from divergence rather than correctness, so the sign of a delta means the opposite of what it did.
  • A tier no longer carries pct and value. It carries divergence, coverage and correctness, each with a pct and a value of its own.
  • The whole-suite correctness percentage is correctness, not total. total also names the raw test count inside counts, so the same word meant a count in one place and a percentage in another.

Schema 4 is additive on top: each target carries its grade and the full criteria in metrics.grade, every envelope gains baseline.observation (how much of the suite the live-AWS row stands on, which passes reported, and what is carried), latest.json and runs.json gain divergence, project, configuration and isVariant, results/summary.json gains regionFailures, and /data/index.json lists the split registry.

A score is two figures, never one

Divergence is the share of the whole suite a target answers differently from real DynamoDB. Coverage is the share it implements at all. They are reported apart and never summed, because a declined operation is discoverable in minutes and a wrong one in production. No target was re-run for the change and no pass, fail or skip moved: what changed is how the same counts are expressed.

Everything else that was a percentage followed the headline down - tier figures, the per-region drilldown, the per-operation table, and the colour bands, which inverted with them. A target's history is two plots rather than one, because divergence falls when a target stops attempting something it used to get wrong, so a divergence line alone can render a withdrawal as an improvement.

Every target wears a letter

Divergence sets it - A under 5%, B under 15%, C under 25%, D under 35%, F beyond

  • and coverage can only lower it, never raise it: a third of whatever a target leaves unimplemented is added to its divergence before the bands are read. A+ is exactly zero divergence at full coverage. A row says when coverage is holding its letter down, so a capped row is not mistaken for one with room above it. Real DynamoDB reads baseline rather than a letter, because grading the yardstick against itself would seat it in a band an engine had to earn its way into.

These are grading criteria version 1, dated in the methodology, which carries the derivation. Where a threshold sits is a hand-picked input to a published letter, and moving one regrades targets whose results never changed, so any change to a band, the coverage weight or the A+ gate bumps the version.

The suite grows from 998 tests to 1054

Vector search (#125). 42 tests over the deterministic surface DynamoDB shipped in August 2026: the index lifecycle on both creation paths, request validation on each plane, rejection wording, write-path validation, search on a fixture where the nearest neighbour is unambiguous, the two new capacity shapes, and PartiQL's inability to reach a vector index. Every pinned value was characterised against real DynamoDB in eu-west-2 before it was asserted. Two findings worth naming: searching during a backfill is an error, which settles which side of a contradiction in AWS's own documentation is right - three developer-guide pages say the call fails, the tutorial page says results can be incomplete - and an overwrite leaving the stored vector unchanged reports no vector write capacity at all, because index replication is delta-based. Both sides are captured in captures/2026-08-12-vector-backfill-docs.json. No emulator implements the family yet, so every target skips it and every coverage figure drops with this release while divergence is untouched. Sending the new operations needs @aws-sdk/client-dynamodb 3.1103.0 or later.

Index write costs (#124). 14 tests. The suite's only per-index capacity assertion was on a Query, so the write side - the half you get billed extra for - went unmeasured. A sub-1KB write costs one unit for the table and one for each index it lands in, and LSI units fold into the total exactly as GSI units do. Moving an item to a new GSI key costs two on that index, a delete and an insert; touching a projected attribute costs one; touching a non-projected attribute costs nothing, and the response carries no arm for that index rather than a zero. An overwrite that leaves the item unchanged reports no index cost whatsoever.

The index exclusions create no indexed table (#116). !gsi and !lsi used to select the right tests and then build the tables anyway, so an engine with no secondary-index support died in setup whatever it had asked for. Shared tables are created on demand from what the running file declared, the composite table split into indexed and plain variants, and three guards keep the declarations and the tags honest.

Corrections

  • The suite counts its own tests. "The whole suite" had meant whichever target ran the most tests, so one of the measured things was setting the denominator every figure divided by. registry/suite-manifest.json lists every test by file and full name, generated from the suite rather than inferred from a run, and CI fails if it drifts. Publishing now refuses a row whose test population disagrees with it, which catches a results file carried across a rename that keeps its old total while naming tests that no longer exist.
  • Every figure on a row comes from one region. A headline came from the target's best-matching region while its tier split and raw counts stayed on the baseline region's basis. Correctness never had to reconcile, because each tier had its own denominator; divergence is additive, so it does.
  • The per-region overlay is matched to the run it describes, rather than keyed on the date real AWS was last swept, which had put a later run's figures on an earlier run's page.
  • A methodology claim was wrong. The page said that measuring divergence over the whole suite stops an engine implementing a sliver from posting a perfect score. It doesn't: zero fails is 0.0% divergence at any coverage. What does hold is the identity the page now states - a test going from failing to skipped leaves both numerators together over the same denominator, so withdrawal costs exactly as much coverage as it gains divergence, and is disclosed rather than silent.
  • The baseline row is measured rather than pinned once real AWS has been observed across the whole suite. It runs in three passes and only the main one reached the published artefact, so the row had claimed a full suite on a run that recorded less. The passes are merged before scoring, each with its own capture date, and the row stays pinned and says so until all three report.
  • The Atom feed carries no letter on any run measured before criteria version 1 took effect. It had been rewriting each entry's summary with a grade the run never had, while leaving <updated> alone so no subscriber re-notified.
  • Badges publish the grade under a parity label, having still been publishing the correctness percentage the board retired under a conformance label. The endpoint URL is unchanged; a badge whose target has no results this run reads no data rather than disappearing, because the URL sits in other people's READMEs.
  • Two more regional splits are admitted, both BatchGetItem with an empty RequestItems map. The nesting-depth row was rewritten from a full 32-region capture, having been written from four.
  • A build of an engine nests under it rather than taking a row beside it, and maintainedByAuthor is keyed on the project, so it now reads true for the WebAssembly build. That build runs in CI like every other target, where its row had been refreshed by hand.
  • The board leads with the highest-graded engine, since the baseline moved into a panel above the standings. Where that is the board author's own engine, the conflict-of-interest disclosure sits on the card.
  • The Region column is gone. It named the cohort a target matched at its best rate, which read as breadth. The count sits beside the figure instead, and the cohort listing stays on the target page.
  • Each target lists how it is actually distributed, with the project's own page for each. These are the only claims on the board the suite does not measure, so each carries the link that backs it.
Grade B 0.9% diverged 0.9 percentage points more
Grade B 0.0% first run