PK Snapshot Scan Internals
This page is the implementation-level companion to the user-facing snapshot guide.
It explains how the current Go SDK performs primary-key snapshot scan today, using the exact path exercised by the canonical E2E harness.
What SnapshotScanRows(...) Actually Does
The public entry point is:
TableClient.SnapshotScanRows(ctx, schema, opts)
Internally, the flow is:
- resolve the table info and KV encoding mode
- call
GetLatestKvSnapshots(...)if a snapshot ID is not already supplied - call
GetKvSnapshotMetadata(...)for the chosen bucket and snapshot - collect the remote snapshot file list
- fetch all snapshot files into a local temporary directory
- open that directory as a local read-only RocksDB-compatible snapshot using
rockyardkv - iterate the local key-value state and decode rows back into public Go values
The code path is in:
Real Example from the E2E Harness
The canonical snapshot-scan check currently looks like this:
schema := client.NewSchema(
client.Int64Type(),
client.StringType(),
client.StringType(),
)
snapshotID, err := waitForLatestKVSnapshot(ctx, cli, env.kvTable, 0)
if err != nil {
log.Fatal(err)
}
rows, err := cli.Table(env.kvTable).SnapshotScanRows(ctx, schema, client.SnapshotScanOptions{
BucketID: 0,
SnapshotID: &snapshotID,
})
if err != nil {
log.Fatal(err)
}
This is not just an illustrative snippet. It is derived directly from demo/fluss-paimon/go-e2e/main.go.
How the Local Reader Works
After remote files are fetched, the SDK opens the local directory with:
db, err := rockyardkv.OpenForReadOnly(localDir, nil, false)
That happens in internal/snapshot/reader.go.
This means the current implementation is:
- not decoding SST files manually in custom code
- not requiring a running external RocksDB service
- opening the fetched snapshot directory locally in read-only mode through
rockyardkv
Once open, the reader:
- iterates every value in the local snapshot DB
- extracts the schema ID prefix from each value payload
- resolves the correct decode plan for that schema ID
- decodes either indexed or compacted KV row layout
- projects columns onto the requested target schema when needed
Schema Evolution Handling
Snapshot rows may have been written with an older schema than the one the caller currently wants to read.
The current implementation handles that by:
- keeping a decode plan keyed by schema ID
- loading source schema JSON on demand through
TableClient.Schema(...) - decoding with the source schema
- projecting the decoded row onto the requested target columns
That is the reason snapshotDecodePlans(...) exists in client/row.go.
Indexed vs Compacted Decoding
The snapshot reader does not assume a single KV row encoding.
It first resolves whether the table is using indexed or compacted KV encoding, then:
- uses
DecodeIndexed(...)for indexed payloads - uses
DecodeCompacted(...)for compacted payloads
This is one of the key reasons the docs should describe snapshot scan as “implementation-aware but still scoped” rather than pretending it is already a universal abstraction across every future cluster layout.
Current Constraints
- The strongest real validation is the canonical RustFS/Paimon demo path.
- Snapshot scan is real and working, but broader portability across more filesystem backends and storage layouts is still a follow-up area.
- The current docs should present
SnapshotScanRows(...)as a public supported helper with a known, concrete implementation path rather than as a vague black box.