Architecture Overview
This page brings the current architecture guide into the docs site so the high-level system shape is available alongside the user-facing docs.
Goals
The SDK is designed to:
- talk directly to Fluss over its native RPC protocol
- provide an idiomatic Go API for Fluss admin and table workflows
- keep protocol and routing details internal where practical
- grow from a raw-but-correct foundation into a production-ready client
Package Layout
Public Surface
client- the main entry point for dialing the cluster
- admin and table operations
- user-facing config and result types
Internal Support Packages
internal/auth- pluggable auth interfaces
internal/transport- TCP connection management
- framed request/response I/O
- API negotiation and auth hook integration
internal/codec/frame- Fluss frame encode/decode logic
internal/protocol- protocol-level error types and helpers
internal/metadata- metadata cache
- table and bucket routing helpers
- schema and table descriptor parsing helpers
internal/codec/row- indexed-row and KV row encode/decode
- the row/schema/type foundation re-exported through
client - writer-state headers used by long-lived write flows
internal/codec/arrow- Fluss Arrow log batch encode/decode
- compression-aware Arrow data path
internal/proto- source
.protodefinitions and generation entrypoint
- source
internal/proto/gen- generated protobuf Go code used by transport and client layers
internal/pbutil- protobuf helper utilities
internal/snapshot- local snapshot decode and fetch helpers for KV snapshot scan flows
Demo Environments
demo/fluss-dev- lightweight Fluss-only development stack
- host-run validation harness for day-to-day SDK work
- bootstraps dev tables through the SDK itself and exercises log, Arrow, and KV flows
demo/fluss-paimon- broader real-cluster support-contract environment
- validates direct Go client access against a Fluss + Paimon-style stack
- used when the project needs stronger end-to-end coverage than the lightweight dev loop
Current Request Flow
At a high level:
client.Dial(...)builds a root client using configured endpoints.- The internal transport layer negotiates supported API versions with the cluster.
- Metadata is fetched and cached for routing.
- Admin or table methods construct requests using generated protobuf Go types from
internal/proto/gen. - The client either talks to the coordinator directly or routes table and bucket work to the correct tablet leader using cached metadata.
- The transport client sends the framed request and matches the response by correlation ID.
- Results are returned as Go types or raw record-batch bytes, depending on the API.
For higher-level row flows, the current write path often looks like this:
TableClient.BootstrapRowWriter(...)pins table metadata, schema ID, KV key decisions, Arrow compression settings, and optional writer-ID state.- Public row helpers encode indexed, KV, or Arrow batches through
internal/codec/roworinternal/codec/arrow. - Table-level write methods group batches by routed leader and issue the underlying Fluss produce or put RPCs.
Current Tradeoffs
The repo is intentionally early in a few areas:
- some data-plane APIs still expose raw or low-level batch-oriented results
- higher-level scanner abstractions are not complete yet
- writer ergonomics have started, but failover and retry hardening are still in progress
- secured-cluster support is still mostly an extension-point design, not a full implementation
- real-cluster validation is improving, but production-confidence coverage is still narrower than the roadmap target
These tradeoffs are deliberate. The current priority is protocol correctness and a stable foundation before the high-level API grows much further.
Protocol Model Direction
The active implementation shape is:
- public SDK surface: Go-native types and methods
- internal wire/protocol layer: generated protobuf Go code
Application code should work with client, AdminClient, TableClient, option structs, and result types rather than raw protobuf request or response structs.
Direction of Travel
The next major steps are:
- production-grade writer hardening around schema evolution, retries, and failover
- stronger scanner abstractions and broader typed row ergonomics
- richer real-cluster validation, especially failure-path coverage
- security and observability hardening
- broader public API and documentation polish