Skip to main content

PK Snapshot Scan Internals

This page is the implementation-level companion to the user-facing snapshot guide.

It explains how the current Go SDK performs primary-key snapshot scan today, using the exact path exercised by the canonical E2E harness.

What SnapshotScanRows(...) Actually Does​

The public entry point is:

  • TableClient.SnapshotScanRows(ctx, schema, opts)

Internally, the flow is:

  1. resolve the table info and KV encoding mode
  2. call GetLatestKvSnapshots(...) if a snapshot ID is not already supplied
  3. call GetKvSnapshotMetadata(...) for the chosen bucket and snapshot
  4. collect the remote snapshot file list
  5. fetch all snapshot files into a local temporary directory
  6. open that directory as a local read-only RocksDB-compatible snapshot using rockyardkv
  7. iterate the local key-value state and decode rows back into public Go values

The code path is in:

Real Example from the E2E Harness​

The canonical snapshot-scan check currently looks like this:

schema := client.NewSchema(
client.Int64Type(),
client.StringType(),
client.StringType(),
)

snapshotID, err := waitForLatestKVSnapshot(ctx, cli, env.kvTable, 0)
if err != nil {
log.Fatal(err)
}

rows, err := cli.Table(env.kvTable).SnapshotScanRows(ctx, schema, client.SnapshotScanOptions{
BucketID: 0,
SnapshotID: &snapshotID,
})
if err != nil {
log.Fatal(err)
}

This is not just an illustrative snippet. It is derived directly from demo/fluss-paimon/go-e2e/main.go.

How the Local Reader Works​

After remote files are fetched, the SDK opens the local directory with:

db, err := rockyardkv.OpenForReadOnly(localDir, nil, false)

That happens in internal/snapshot/reader.go.

This means the current implementation is:

  • not decoding SST files manually in custom code
  • not requiring a running external RocksDB service
  • opening the fetched snapshot directory locally in read-only mode through rockyardkv

Once open, the reader:

  • iterates every value in the local snapshot DB
  • extracts the schema ID prefix from each value payload
  • resolves the correct decode plan for that schema ID
  • decodes either indexed or compacted KV row layout
  • projects columns onto the requested target schema when needed

Schema Evolution Handling​

Snapshot rows may have been written with an older schema than the one the caller currently wants to read.

The current implementation handles that by:

  • keeping a decode plan keyed by schema ID
  • loading source schema JSON on demand through TableClient.Schema(...)
  • decoding with the source schema
  • projecting the decoded row onto the requested target columns

That is the reason snapshotDecodePlans(...) exists in client/row.go.

Indexed vs Compacted Decoding​

The snapshot reader does not assume a single KV row encoding.

It first resolves whether the table is using indexed or compacted KV encoding, then:

  • uses DecodeIndexed(...) for indexed payloads
  • uses DecodeCompacted(...) for compacted payloads

This is one of the key reasons the docs should describe snapshot scan as “implementation-aware but still scoped” rather than pretending it is already a universal abstraction across every future cluster layout.

Current Constraints​

  • The strongest real validation is the canonical RustFS/Paimon demo path.
  • Snapshot scan is real and working, but broader portability across more filesystem backends and storage layouts is still a follow-up area.
  • The current docs should present SnapshotScanRows(...) as a public supported helper with a known, concrete implementation path rather than as a vague black box.