Skip to content

ExtendDB

v0.1.1 25 recorded runs

83.2%
correct, across 911 of 954 operations · 43 unsupported
Tier 1 85.5%
Tier 2 86.2%
Tier 3 78.8%

Per-region conformance

The headline is all regions. Real DynamoDB doesn't behave identically in every region, so a target can match some regions more closely than others. eu-west-2 is marked as the historical baseline, and a region that couldn't be resolved this sweep is flagged rather than counted as a disagreement.

83.2% 33 regions
af-south-1ap-east-1ap-east-2ap-northeast-1ap-northeast-2ap-northeast-3ap-south-1ap-south-2ap-southeast-1ap-southeast-2ap-southeast-3ap-southeast-4ap-southeast-5ap-southeast-6ap-southeast-7ca-central-1ca-west-1eu-central-1eu-central-2eu-north-1eu-south-1eu-south-2eu-west-1eu-west-2 · baselineeu-west-3il-central-1me-central-1mx-central-1sa-east-1us-east-1us-east-2us-west-1us-west-2

Fully supports

Operation areas this target implements completely - every test passes, nothing skipped. Areas it only partly implements, or gets wrong, are in “Where it falls short” below.

  • deleteItem tier 1
  • deleteTable tier 1
  • describeTable tier 1
  • listTables tier 1
  • account tier 2
  • streams tier 2
  • tags tier 2
  • ttl tier 2
  • legacy-api tier 3

Conformance over time

ExtendDB conformance history, 2026-05-23 to 2026-07-18 80% 90% 100% 23 May 25 May 29 May 16 Jun 19 Jun 22 Jun 24 Jun 29 Jun 1 Jul 6 Jul 14 Jul 16 Jul 18 Jul 83.2%
Now 83.2% on 18 Jul, lowest 83.0% on 16 Jul.

By operation

Every operation this target implements, with the share of its tests that pass. full, partial, failing, unsupported. The matrix compares operations across targets.

Tier 1 - Core

  • batchGetItem 14/15 93.3% batchGetItem: partially supported (14 pass, 1 fail)
  • batchWriteItem 15/18 83.3% batchWriteItem: partially supported (15 pass, 3 fail)
  • createTable 23/30 76.7% createTable: partially supported (23 pass, 7 fail)
  • deleteItem 13/13 100.0% deleteItem: supported (13 pass)
  • deleteTable 3/3 100.0% deleteTable: supported (3 pass)
  • describeTable 4/4 100.0% describeTable: supported (4 pass)
  • getItem 32/40 80.0% getItem: partially supported (32 pass, 8 fail)
  • listTables 5/5 100.0% listTables: supported (5 pass)
  • putItem 87/116 75.0% putItem: partially supported (87 pass, 29 fail)
  • query 81/90 90.0% query: partially supported (81 pass, 9 fail)
  • scan 48/56 85.7% scan: partially supported (48 pass, 8 fail)
  • updateItem 68/70 97.1% updateItem: partially supported (68 pass, 2 fail)
  • updateTable 13/15 86.7% updateTable: partially supported (13 pass, 2 fail)

Tier 2 - Complete

  • account 2/2 100.0% account: supported (2 pass)
  • backups 3/5 60.0% backups: partially supported (3 pass, 2 fail)
  • contributorInsights 0/0 · 2 skip n/a contributorInsights: unsupported (2 skip)
  • export 0/0 · 2 skip n/a export: unsupported (2 skip)
  • kinesis 0/0 · 1 skip n/a kinesis: unsupported (1 skip)
  • partiql 0/0 · 36 skip n/a partiql: unsupported (36 skip)
  • resourcePolicy 0/0 · 2 skip n/a resourcePolicy: unsupported (2 skip)
  • streams 18/18 100.0% streams: supported (18 pass)
  • tags 8/8 100.0% tags: supported (8 pass)
  • transactions 52/62 83.9% transactions: partially supported (52 pass, 10 fail)
  • ttl 7/7 100.0% ttl: supported (7 pass)
  • updateTable 10/14 71.4% updateTable: partially supported (10 pass, 4 fail)

Tier 3 - Strict

  • error-messages 95/154 61.7% error-messages: partially supported (95 pass, 59 fail)
  • legacy-api 42/42 100.0% legacy-api: supported (42 pass)
  • limits 85/93 91.4% limits: partially supported (85 pass, 8 fail)
  • validation-ordering 30/31 96.8% validation-ordering: partially supported (30 pass, 1 fail)

Where it falls short

Operation groups with gaps in the latest run, biggest first. Unsupported means the target doesn't implement that feature at all, often by design. Open one for the exact tests, or see the conformance suite.

  • error-messages Tier 3 59 failing
    View these tests in the suite →
    • BatchGetItem - exact error messages mixing ProjectionExpression on one table and AttributesToGet on another is rejected
    • BatchGetItem - ProjectionExpression rejection messages duplicate paths (a, a) reject even when the key matches no item
    • BatchGetItem - ProjectionExpression rejection messages two distinct aliases resolving to one attribute (#a, #b -> a): rejected on the resolved names
    • BatchGetItem - ProjectionExpression rejection messages overlapping parent and child paths (a, a.b): full overlap message
    • BatchGetItem - ProjectionExpression rejection messages a bad projection on one table entry rejects the whole batch despite a clean entry on another
    • BatchWriteItem - exact error messages wrong-typed index key: full type-mismatch message
    • BatchWriteItem - exact error messages non-scalar index key: full type-mismatch message
    • BatchWriteItem - exact error messages empty-string index key: full secondary-index-key message
    • BatchWriteItem - exact error messages empty-binary index key: full secondary-index-key message
    • CreateTable - exact error messages GSI INCLUDE projection without NonKeyAttributes: full missing-attributes message
    • CreateTable - exact error messages LSI INCLUDE projection without NonKeyAttributes: full missing-attributes message
    • CreateTable - exact error messages StreamSpecification StreamEnabled:false with a StreamViewType: full conflict message
    • GetItem - ProjectionExpression rejection messages raw duplicate paths (a, a): full overlap message with the path on both sides
    • GetItem - ProjectionExpression rejection messages same alias twice (#a, #a): full overlap message
    • GetItem - ProjectionExpression rejection messages two distinct aliases resolving to one attribute (#a, #b -> a): rejected on the resolved names
    • GetItem - ProjectionExpression rejection messages raw path plus alias for the same attribute (a, #a): full overlap message
    • GetItem - ProjectionExpression rejection messages raw parent and child paths (a, a.b): full overlap message
    • GetItem - ProjectionExpression rejection messages aliased parent and child paths (#a, #a.#b): full overlap message
    • GetItem - ProjectionExpression rejection messages child before parent (a.b, a): rejected with the paths in request order
    • GetItem - ProjectionExpression rejection messages cross-alias overlap (#x, #y.#b with #x and #y -> a): rejected on the resolved paths
    • GetItem - ProjectionExpression rejection messages deep overlap (a, a.b.c): full overlap message with the three-element path
    • GetItem - ProjectionExpression rejection messages list parent and list index (l, l[0]): rejected as an overlap
    • GetItem - ProjectionExpression rejection messages undefined attribute name (#undef): full undefined-name message
    • PutItem - exact error messages empty string member in NS: numeric conversion error
    • PutItem - exact error messages duplicate zero-length members in BS: full duplicates error
    • PutItem - exact error messages empty-binary index key value: full secondary-index-key message
    • Query - ProjectionExpression rejection messages duplicate paths (a, a) reject even when the key condition matches no partition
    • Query - ProjectionExpression rejection messages two distinct aliases resolving to one attribute (#a, #b -> a): rejected on the resolved names
    • Query - ProjectionExpression rejection messages overlapping parent and child paths (a, a.b) reject even when the key condition matches no partition
    • Query - ProjectionExpression rejection messages duplicate paths still reject when the key condition matches a partition (control)
    • Scan - exact error messages Select SPECIFIC_ATTRIBUTES without ProjectionExpression: full required-projection message
    • Scan - ProjectionExpression rejection messages duplicate paths (a, a) reject even when the filter matches nothing
    • Scan - ProjectionExpression rejection messages two distinct aliases resolving to one attribute (#a, #b -> a): rejected on the resolved names
    • Scan - ProjectionExpression rejection messages overlapping parent and child paths (a, a.b) reject even when the filter matches nothing
    • Scan - ProjectionExpression rejection messages rejects an undefined projection name even when the scan matches nothing
    • Scan - ProjectionExpression rejection messages duplicate paths still reject when the filter matches a row (control)
    • TransactWriteItems - exact error messages Put wrong-typed table key: cancelled with full ValidationError reason
    • TransactWriteItems - exact error messages Put non-scalar table key: cancelled with full ValidationError reason
    • TransactWriteItems - exact error messages Put wrong-typed index key: cancelled with full ValidationError reason
    • TransactWriteItems - exact error messages Put non-scalar index key: cancelled with full ValidationError reason
    • TransactWriteItems - exact error messages Update wrong-typed index key: cancelled with full ValidationError reason
    • TransactWriteItems - exact error messages Update non-scalar index key: cancelled with full ValidationError reason
    • TransactWriteItems - exact error messages Put empty-string index key: top-level ValidationException
    • TransactWriteItems - exact error messages Update empty-string index key: top-level ValidationException
    • TransactWriteItems - exact error messages Update empty-string Key: top-level empty-value message
    • TransactWriteItems - exact error messages Delete empty-string Key: top-level empty-value message
    • TransactWriteItems - exact error messages ConditionCheck empty-string Key: top-level empty-value message
    • TransactWriteItems - exact error messages Update wrong-typed Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages Update non-scalar Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages Delete wrong-typed Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages Delete non-scalar Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages ConditionCheck wrong-typed Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages ConditionCheck non-scalar Key: cancelled with schema-mismatch reason
    • TransactWriteItems - exact error messages Update empty-binary Key: top-level empty-value message
    • TransactWriteItems - exact error messages Delete empty-binary Key: top-level empty-value message
    • TransactWriteItems - exact error messages ConditionCheck empty-binary Key: top-level empty-value message
    • TransactWriteItems - exact error messages Put empty-binary index key: top-level ValidationException
    • TransactWriteItems - exact error messages Update empty-binary index key: top-level ValidationException
    • UpdateItem - exact error messages empty-binary index key value: full secondary-index-key message
  • partiql Tier 2 unsupported
    View these tests in the suite →
    • BatchExecuteStatement - PartiQL batch of multiple SELECT statements
    • BatchExecuteStatement - PartiQL batch of INSERT and UPDATE statements
    • BatchExecuteStatement - PartiQL partial failure - one valid and one invalid statement
    • BatchExecuteStatement - PartiQL rejects an empty Statements array
    • ExecuteStatement - PartiQL INSERTs a new item
    • ExecuteStatement - PartiQL SELECTs an item by primary key
    • ExecuteStatement - PartiQL SELECTs with WHERE clause using comparison
    • ExecuteStatement - PartiQL UPDATEs an existing item
    • ExecuteStatement - PartiQL DELETEs an item
    • ExecuteStatement - PartiQL rejects INSERT on an existing item (INSERT is not upsert)
    • ExecuteStatement - PartiQL INSERT succeeds after DELETE of same key
    • ExecuteStatement - PartiQL SELECT returns empty results for non-matching WHERE
    • ExecuteStatement - PartiQL parameterized INSERT with ? placeholders
    • ExecuteStatement - PartiQL parameterized SELECT with ? placeholder
    • ExecuteStatement - PartiQL SELECT with nested map path
    • ExecuteStatement - PartiQL SELECT with specific attributes
    • ExecuteStatement - PartiQL SELECT with begins_with in WHERE clause
    • ExecuteStatement - PartiQL PartiQL UPDATE with SET on attribute
    • ExecuteStatement - PartiQL PartiQL UPDATE with REMOVE
    • ExecuteStatement - PartiQL returns a populated ConsumedCapacity block when requested
    • ExecuteStatement - PartiQL evaluates negated predicates (NOT begins_with, IS NOT MISSING)
    • ExecuteStatement - PartiQL DELETE with a false non-key predicate fails ConditionalCheckFailed and leaves the item
    • ExecuteStatement - PartiQL DELETE with a true non-key predicate removes the item
    • ExecuteStatement - PartiQL UPDATE with a false non-key predicate fails ConditionalCheckFailed and leaves the item
    • ExecuteStatement - PartiQL UPDATE with a true non-key predicate mutates the item
    • ExecuteStatement - PartiQL DELETE with a false NOT begins_with predicate fails ConditionalCheckFailed
    • ExecuteStatement - PartiQL DELETE with a true NOT begins_with predicate removes the item
    • ExecuteStatement - PartiQL rejects a write WHERE clause that omits the primary key
    • ExecuteStatement - PartiQL DELETE on a missing key with a non-key predicate is a silent no-op
    • ExecuteStatement - PartiQL rejects a statement with syntax error
    • ExecuteStatement - PartiQL rejects a reference to a non-existent table
    • ExecuteTransaction - PartiQL transactional INSERT and UPDATE both succeed atomically
    • ExecuteTransaction - PartiQL transaction rolls back on duplicate key INSERT
    • ExecuteTransaction - PartiQL multiple INSERTs in one transaction
    • ExecuteTransaction - PartiQL idempotent replay under the same ClientRequestToken does not double-apply
    • ExecuteTransaction - PartiQL rejects empty TransactStatements
  • putItem Tier 1 29 failing
    View these tests in the suite →
    • ReturnConsumedCapacity - PutItem reports only the aggregate, no write-capacity split PutItem with TOTAL reports 1 CapacityUnit and no read/write split
    • ReturnConsumedCapacity - PutItem reports only the aggregate, no write-capacity split PutItem with INDEXES omits the split at the top level and on Table
    • PutItem - number format rejects "+e2"
    • PutItem - number format rejects "e2"
    • PutItem - number format rejects "+1+2"
    • PutItem - number format rejects "1+2"
    • PutItem - number format rejects "+1.2.3"
    • PutItem - number format rejects "1.2.3"
    • PutItem - number format rejects "++5"
    • PutItem - number format rejects "+-5"
    • PutItem - number format rejects "-+5"
    • PutItem - number format rejects "+"
    • PutItem - number format rejects "-"
    • PutItem - number format rejects "1e"
    • PutItem - number format rejects "1e+"
    • PutItem - number format rejects "."
    • PutItem - number format rejects "1.2e3.4"
    • PutItem - number format rejects "0x5"
    • PutItem - number format rejects "NaN"
    • PutItem - number format rejects "Infinity"
    • PutItem - number format rejects "1_000"
    • PutItem - number format rejects " 5"
    • PutItem - number format rejects "5 "
    • PutItem - number format rejects "1 5"
    • PutItem - validation rejects a wrong-typed index key attribute
    • PutItem - validation rejects a non-scalar index key attribute
    • PutItem - validation rejects an empty-string index key attribute
    • PutItem - validation rejects a number with a leading space
    • PutItem - validation rejects a number with a trailing space
  • transactions Tier 2 10 failing
    View these tests in the suite →
    • TransactWriteItems - validation Put with a wrong-typed index key cancels with a ValidationError reason
    • TransactWriteItems - validation Put with a non-scalar index key cancels with a ValidationError reason
    • TransactWriteItems - validation Update setting a wrong-typed index key cancels with a ValidationError reason
    • TransactWriteItems - validation Update setting a non-scalar index key cancels with a ValidationError reason
    • TransactWriteItems - validation Put with an empty-string index key is a top-level ValidationException
    • TransactWriteItems - validation Update setting an empty-string index key is a top-level ValidationException
    • TransactWriteItems - validation Update with an empty-string Key is a top-level ValidationException
    • TransactWriteItems - validation Delete with an empty-string Key is a top-level ValidationException
    • TransactWriteItems - validation ConditionCheck with an empty-string Key is a top-level ValidationException
    • TransactWriteItems - ConsumedCapacity: conditional, check, replay, cancel reports write capacity on the first call and read capacity on a same-token replay
  • query Tier 1 9 failing
    View these tests in the suite →
    • Query - GSI pagination across tied sort keys composes LastEvaluatedKey from base and index keys and walks every tied item once
    • Query - hash-only GSI pagination on a composite base table walks every item across pages when they share the GSI hash and base partition key
    • Query - KeyConditionExpression semantics rejects a nested path on a key attribute
    • Query - LSI pagination across tied sort keys composes LastEvaluatedKey from base and index keys and walks every tied item once
    • Query - legal shared-prefix projections accepts distinct list indices (l[0], l[1]) and returns both elements in order
    • Query - Select / ProjectionExpression rejections Select ALL_ATTRIBUTES with ProjectionExpression is rejected
    • Query - Select / ProjectionExpression rejections Select ALL_PROJECTED_ATTRIBUTES without an IndexName is rejected
    • Query - Select / ProjectionExpression rejections Select COUNT with ProjectionExpression is rejected
    • Query - Select / ProjectionExpression rejections Select ALL_PROJECTED_ATTRIBUTES with ProjectionExpression and no IndexName is rejected
  • getItem Tier 1 8 failing
    View these tests in the suite →
    • ConsumedCapacity - single-item ops report only the aggregate, no read/write split a strongly-consistent GetItem reports 1 CapacityUnit and no read/write split
    • ConsumedCapacity - single-item ops report only the aggregate, no read/write split an eventually-consistent GetItem reports 0.5 CapacityUnits and no read/write split
    • ConsumedCapacity - single-item ops report only the aggregate, no read/write split a Query under INDEXES omits the split at the top level and on Table
    • ConsumedCapacity - single-item ops report only the aggregate, no read/write split an UpdateItem reports 1 CapacityUnit and no read/write split
    • ConsumedCapacity - single-item ops report only the aggregate, no read/write split a DeleteItem reports 1 CapacityUnit and no read/write split
    • Nested attribute projection GetItem projects multiple list indices compacted and index-ordered
    • Same list-index projection merge GetItem keeps distinct inner list indices separate under one shared outer index
    • Same list-index projection merge GetItem merges the shared index and keeps a distinct index separate, compacted
  • scan Tier 1 8 failing
    View these tests in the suite →
    • Scan - GSI pagination across tied sort keys walks every item once across a paged GSI scan with tied sort keys
    • Scan - legal shared-prefix projections accepts distinct list indices (l[0], l[1]) and returns both elements in order
    • Scan - Select / ProjectionExpression rejections Select ALL_ATTRIBUTES with ProjectionExpression is rejected
    • Scan - Select / ProjectionExpression rejections Select ALL_PROJECTED_ATTRIBUTES without an IndexName is rejected
    • Scan - Select / ProjectionExpression rejections Select COUNT with ProjectionExpression is rejected
    • Scan - Select / ProjectionExpression rejections Select ALL_PROJECTED_ATTRIBUTES with ProjectionExpression and no IndexName is rejected
    • Scan - Select / ProjectionExpression rejections Select SPECIFIC_ATTRIBUTES without ProjectionExpression is rejected
    • Scan - validation rejects TotalSegments above the maximum
  • limits Tier 3 8 failing
    View these tests in the suite →
    • Empty values - strings, binary, and sets empty string member in NS is rejected (not a number)
    • Expression size limit (4KB) - UpdateExpression rejects an UpdateExpression over the 4096-byte limit
    • Expression size limit (4KB) - ConditionExpression rejects a ConditionExpression over the 4096-byte limit
    • Expression size limit (4KB) - Query FilterExpression rejects a Query FilterExpression over the 4096-byte limit
    • Expression size limit (4KB) - Scan FilterExpression rejects a Scan FilterExpression over the 4096-byte limit
    • Expression size limit (4KB) - ProjectionExpression rejects a ProjectionExpression over the 4096-byte limit
    • Key-value length limits rejects an over-limit partition key on the read path
    • Nesting depth - 32-level document limit rejects a 32-level ExpressionAttributeValue with ValidationException
  • createTable Tier 1 7 failing
    View these tests in the suite →
    • CreateTable - validation rejects a GSI INCLUDE projection without NonKeyAttributes
    • CreateTable - validation rejects StreamSpecification with StreamEnabled false plus a StreamViewType
    • CreateTable - configuration parameters TableClass STANDARD_INFREQUENT_ACCESS round-trips
    • CreateTable - configuration parameters SSESpecification with the AWS-managed key round-trips
    • CreateTable - configuration parameters OnDemandThroughput round-trips on a PAY_PER_REQUEST table
    • CreateTable - LSI rejects an LSI INCLUDE projection without NonKeyAttributes
    • CreateTable - index and stream spec validation rejects a KEYS_ONLY GSI projection carrying NonKeyAttributes
  • updateTable Tier 2 4 failing
    View these tests in the suite →
    • UpdateTable - add GSI adds multiple GSIs sequentially
    • UpdateTable - add GSI accepts a conflicting redeclaration of an existing key attribute and keeps the stored type
    • UpdateTable - add GSI drops an unused AttributeDefinition supplied with a GSI add
    • UpdateTable - remove GSI removes a GSI from a table
  • batchWriteItem Tier 1 3 failing
    View these tests in the suite →
    • BatchWriteItem - validation rejects a wrong-typed index key value
    • BatchWriteItem - validation rejects a non-scalar index key value
    • BatchWriteItem - validation rejects an empty-string index key value
  • updateItem Tier 1 2 failing
    View these tests in the suite →
    • UpdateItem - SET evaluation semantics a second SET clause reads the pre-update value of another attribute
    • UpdateItem - SET evaluation semantics applies parenthesised arithmetic (SET c = (c - :v))
  • updateTable Tier 1 2 failing
    View these tests in the suite →
    • UpdateTable - configuration parameters UpdateTable changes TableClass
    • UpdateTable - configuration parameters UpdateTable changes OnDemandThroughput
  • backups Tier 2 2 failing
    View these tests in the suite →
    • On-demand backups - lifecycle and restore RestoreTableFromBackup initiates a restore into a new table
    • On-demand backups - lifecycle and restore DescribeBackup on a deleted backup throws BackupNotFoundException
  • contributorInsights Tier 2 unsupported
    View these tests in the suite →
    • Contributor insights - enable/describe/list reports DISABLED by default
    • Contributor insights - enable/describe/list enabling transitions the status and lists the table
  • export Tier 2 unsupported
    View these tests in the suite →
    • Export and import - S3 ExportTableToPointInTime initiates an export and reports it
    • Export and import - S3 ImportTable ingests S3 data into a new table
  • resourcePolicy Tier 2 unsupported
    View these tests in the suite →
    • Resource policies - Put/Get/Delete GetResourcePolicy on a table with no policy throws PolicyNotFoundException
    • Resource policies - Put/Get/Delete Put then Get round-trips the policy, and Delete removes it
  • batchGetItem Tier 1 1 failing
    View these tests in the suite →
    • BatchGetItem - legal shared-prefix projections accepts distinct list indices (l[0], l[1]) and returns both elements in order
  • kinesis Tier 2 unsupported
    View these tests in the suite →
    • Kinesis streaming destination enables a streaming destination and reports it via Describe
  • validation-ordering Tier 3 1 failing
    View these tests in the suite →
    • UpdateItem - validation ordering rejects invalid ReturnValues (UpdateItem reports the first enum error)

Run history

Run Total Movement
83.2% 0.0pp unchanged

Suite on

Per-region scoring lands complete. 2.0.0-pre put the scoring logic in place, comparing each target against every region's recorded answer, but the evidence half was never wired: no test recorded what a target actually answered and the classifier never read one, so a fail could not be credited to a region the target matched and the score could only ever subtract. 2.0.0 closes that loop, and the seed split runs its whole lifecycle in the same release.

What changed:

  • The split tests now record what the target actually answered (src/observation-sink.ts), in the same shape the registry stores each region's answer, and the classifier carries it onto the verdict. An engine that matches a rejecting region on a split is now scored as passing in that region. Evidence only ever redeems a committed fail: a pass keeps the committed assertion's deliberate wording tolerance and is never held to the byte-exact recorded string. Committed results predate the capture, so published scores move on each target's next run, not in this change.
  • Headline ties now prefer a region the registry characterises. A region absent from every split row ties the top score by having nothing recorded about it, and the Region column must not answer "conformant to what?" with a region the suite knows nothing about. Only the label is affected; a strictly higher score still wins whatever its source.
  • The seeded { NULL: false } split closed its own loop. eu-west-2 and eu-central-1 were the last regions accepting it; by the 2026-07-17 sweep both reject it again, so every region agrees, the split is retired, and its test asserts the shared rejection. The detect, admit and reconcile path the pre-release introduced, exercised end to end within days.
  • A tooling test now asserts every tracked file is text, after a raw NUL byte used as a delimiter in one source file made grep classify the file as binary and silently skip it.
83.2% +0.2pp rose 0.2 percentage points
83.0% -0.1pp fell 0.1 percentage points
83.1% 0.0pp unchanged
83.1% -2.9pp fell 2.9 percentage points

Suite on

The scores barely move in this release, but what they mean has changed.

Until now the suite pinned one region, eu-west-2, as ground truth. That was quietly unfair: real DynamoDB disagrees with itself in a handful of places, and a one-region baseline takes a side without saying so. The clearest case is the { NULL: false } attribute value - accepted and normalised to { NULL: true } in eu-west-2 and eu-central-1, rejected with a ValidationException in us-east-1 and ap-southeast-2. An engine matching us-east-1 on that behaviour was marked non-conformant for doing exactly what real DynamoDB does in Virginia. From 2.0.0, ground truth is per region.

What changed:

  • A weekly sweep runs the full suite against real DynamoDB in every commercial region and publishes per-region ground truth (.github/workflows/sweep.yml, ground-truth/). It gates nothing: PR scoring stays offline and deterministic.
  • Confirmed regional splits live in a checked-in registry (registry/splits.json), each row carrying what every named region actually returned, when, and who admitted it. Detection is automatic; admission is not. The sweep files an issue with the evidence, and only a human commits a row. The registry ships seeded with the { NULL: false } split.
  • Each target is scored against every observed region's expectations, and its published number is its best-matching region - named in the results table's new Region column, with the full per-region view in results/summary.json (a versioned, additive artefact; the per-target results/<slug>.json files are unchanged in shape).
  • A third result state, indeterminate, for a failed observation: a timeout, an exhausted throttle, a transport fault. It is excluded from both sides of the score and cannot become a split, a registry row, or a fail. An absent answer is not a different answer.
  • Region health is tracked in registry/regions.json. A region that cannot complete a sweep is published as unresolved rather than silently omitted; two consecutive misses drop it from the scored set and page a maintainer in the same act.

One deliberate departure from the RFC that proposed this (#75): the RFC suggested a behaviour conforms if it matches any real region. 2.0.0 scores each target against one region at a time and headlines the best match, so a target only passes a behaviour when at least one real region does what it did, and its headline reflects one coherent region rather than a mix. Match-any scoring would have accepted an engine that combines eu-west-2's answer on one behaviour with us-east-1's on another - a deployment that exists nowhere. That is stricter than the RFC asked for, and it is deliberate.

No score moves at release: the one admitted split pins eu-west-2, which is the only region in the health record until the first sweep runs. Per-target deltas will be published once the sweep admits more regions; the expected movement is roughly a tenth of a percent for the six engines that match us-east-1 on the { NULL: false } split.

The suite also grew to 954 tests, up 81, all characterised against real DynamoDB - the control-plane pins in eu-west-2, everything else across four regions (eu-west-2, eu-central-1, us-east-1, ap-southeast-2):

  • UpdateTable AttributeDefinitions reconciliation (#77, #78, #79): delta-fed GSI adds merge into the stored union rather than replacing it, a conflicting redeclaration of an existing key keeps the stored type, and deleting a GSI prunes only its orphaned key attributes. An unused definition on an add is silently dropped where CreateTable rejects the same shape.
  • ProjectionExpression validation (#81): duplicate paths, alias collisions and parent/child overlaps rejected identically on GetItem, Query, Scan and BatchGetItem, pinned with exact messages; legal shared-prefix projections guarded as accepted; GetItem's misfiled validation tests rehomed into Tier 3.
  • Expression-size limit (#80): every expression parameter caps at 4096 bytes, measured on the raw string before ExpressionAttributeNames substitution - 4096 accepted, 4097 rejected, on all five expression surfaces. No tracked target enforces this limit today, so the Tier 3 movement it causes is new coverage, not a regression.
  • Empty set members (#82): empty strings in an SS and zero-length members in a BS are accepted and round-trip intact through every write path, and contains() can find them; an empty NS member and duplicate empty members are rejected, with messages pinned. One existing assertion got stricter: the empty-binary round-trip now asserts byte length zero.
86.0% 0.0pp unchanged
86.0% 0.0pp unchanged
86.0% 0.0pp unchanged

Suite on

Grew to 873 tests, up 49, all characterised against real DynamoDB in eu-west-2. New coverage in three areas:

  • ConsumedCapacity: the transactional read/write split on a same-token replay and on ExecuteTransaction, and a correction that single-item operations report the aggregate CapacityUnits and omit the split, which is transactional-only.
  • Empty-binary key values: rejected as a top-level ValidationException on every path, including secondary-index keys and inside transactions, mirroring the empty-string rejection.
  • Expression, limit and response-shape parity: KeyConditionExpression operand and nested-path rules, ExpressionAttributeNames/Values hygiene, projection validation and list-index fidelity, reversed-bounds BETWEEN, read-path key-length and segment caps, whitespace numbers, batch unprocessed fields and cross-table projection mixing, the bare no-op upsert, multi-subpath UPDATED_NEW, filter operand ordering, hash-only-GSI pagination, and CreateTable spec validation.
86.0% -2.6pp fell 2.6 percentage points

Suite on

Grew to 824 tests, up 7, in two parts: two sibling-parity gaps where one half of a rule was pinned and the other was not, and the capacity accounting of conditional and idempotent transactional writes. All characterised against real DynamoDB in eu-west-2.

The first covers the LSI side of the INCLUDE-projection-without-NonKeyAttributes rejection; the GSI side already had it. An LSI declared with ProjectionType INCLUDE and no NonKeyAttributes is rejected as a ValidationException in tier1 and pinned to the exact message in tier3, the same wording the GSI case returns.

The second pins the Query message for Select SPECIFIC_ATTRIBUTES with no ProjectionExpression, which Scan already had. Query and Scan enforce the same rule but word it differently: Query wraps the phrase in the "1 validation error detected:" envelope, Scan returns it bare.

The transaction cases settle what a conditional TransactWriteItems actually costs. A passing condition adds no read capacity: a conditional write bills the same 2 WCU per sub-1KB item as an unconditional one, and a standalone ConditionCheck costs 2 WCU, billed as write not read. Idempotent replay splits the accounting - the first call reports 2 write capacity units, a same-token replay within the window reports 2 read capacity units for re-reading the stored result. A failing condition cancels the transaction, and the response carries no ConsumedCapacity at all. Answers #27.

88.6% -0.3pp fell 0.3 percentage points
88.9% 0.0pp unchanged
88.9% -3.2pp fell 3.2 percentage points

Suite on

Grew to 762 tests, up 18, covering a malformed value in the lookup Key of a TransactWriteItems Update, Delete, or ConditionCheck - the path a Put item key does not take. Captured across four regions (eu-west-2, us-east-1, ap-southeast-2, eu-central-1), where every string was identical, so they pin exactly. An empty-string Key surfaces as a top-level ValidationException with the same message a Put item key gives; a wrong-typed or non-scalar Key cancels with a ValidationError reason carrying "The provided key element does not match the schema" - the key-only form, not the "Type mismatch for key" message the item-key path returns. The same run confirmed the BatchWriteItem table-key schema-mismatch message is region-invariant.

92.1% -1.5pp fell 1.5 percentage points

Suite on

Grew to 744 tests, up 8, pinning the Select / ProjectionExpression rules on Query and Scan. A ProjectionExpression is only valid with Select SPECIFIC_ATTRIBUTES, and ALL_PROJECTED_ATTRIBUTES is only valid with an IndexName; real DynamoDB rejects both with a ValidationException before reading anything. The cases span ALL_ATTRIBUTES, COUNT, and ALL_PROJECTED_ATTRIBUTES, including the request that breaks both rules at once, where AWS reports the ProjectionExpression one. They assert the contractual phrase, so they hold whether or not the engine carries the wrapper AWS adds on Query but not Scan.

93.6% -0.8pp fell 0.8 percentage points

Suite on

Grew to 736 tests, up 30, covering what TransactWriteItems and BatchWriteItem do with an item whose key value is malformed - the wrong type, non-scalar, or an empty string - across both table and index keys. PutItem already covered this; the transactional and batch paths covered none of it.

Characterising it against real AWS turned up a split worth pinning. An empty-string key value is rejected by up-front input validation, so even inside a transaction it surfaces as a top-level ValidationException. A wrong-typed or non-scalar key value is caught while the transaction runs, so it comes back as a TransactionCanceledException carrying a ValidationError reason rather than a top-level error. BatchWriteItem has no cancellation path, so every variant there is a plain ValidationException. The tests pin both halves, which catches two opposite mistakes: an engine that wraps the empty-string case as a cancellation, and one that surfaces the type-mismatch case as a top-level error.

94.4% -3.1pp fell 3.1 percentage points

Suite on

Grew to 706 tests, up 7, tightening secondary-index behaviour in Tier 1. Query and Scan on a GSI or LSI now assert sparse membership: an item that omits the index key stays off the index but remains on the base table. And PutItem now rejects an item whose GSI or LSI key value is the wrong type, non-scalar, or an empty string while the base-table keys are valid, holding an index key to its declared scalar type the same way a table key is held.

97.4% -0.4pp fell 0.4 percentage points
97.9% 0.0pp unchanged
97.9% 0.0pp unchanged
97.9% 0.0pp unchanged

Suite on

Real DynamoDB in eu-west-2 reworded a chunk of its validation errors, and the Tier 3 error-message tests moved with it. They now assert the contract the error carries - its type, the field it objects to, and the constraint - rather than the exact prose, because AWS varies the wrapper, the echoed input value and the field casing from one region to the next. A four-region capture in June found eu-west-2 and eu-central-1 on the new wording and us-east-1 and ap-southeast-2 still on the old, so the line between contract and cosmetic is drawn from what is invariant across all four.

This moves some Tier 3 numbers. Targets that were only ever marked down for wording DynamoDB itself renders inconsistently now pass those checks, so their Tier 3 scores rise: the suite has stopped counting a cosmetic difference as a behavioural one. Genuine behavioural divergences are still pinned exactly, for example PutItem with a { NULL: false } attribute, which DynamoDB now accepts in eu-west-2 and normalises to { NULL: true }.

97.9% +4.1pp rose 4.1 percentage points
93.8% +0.1pp rose 0.1 percentage points

Suite on

Grew to 684 tests with a control-plane and table-configuration sweep: the CreateTable/UpdateTable config parameters and the secondary control-plane operations (limits, backups and PITR, exports and imports, Kinesis, resource policies, contributor insights), each characterised against real AWS and probe-skipped where a target doesn't implement it.

The published percentage changed with it. It now measures correctness over implemented operations, Pass / (Pass + Fail), so skips no longer count against the score. A skip is honest scope; a fail is a bug.

93.7% -1.1pp fell 1.1 percentage points
94.7% -2.1pp fell 2.1 percentage points

Suite on

Grew to 625 tests, up 24 on the previous run: eleven more in Tier 1 and thirteen more in Tier 3, tightening coverage of core operations and the strict edge cases.

96.8% -3.2pp fell 3.2 percentage points

Showing the 24 most recent runs. The chart above covers the full history, and every run is browsable from Runs.