Cluster
etcd-backed cluster state for distributed pinning: membership, pins, the caller's refs, and the GC epoch.
etcd-backed cluster state for distributed pinning: membership, pins, the caller's refs, and the GC epoch.
Service builder.Cluster, 2 rpcs.
GetClusterInfo(GetClusterInfoRequest) -> GetClusterInfoResponse
A snapshot of this node's view of the cluster. clustered is false (and
the lists empty) on a node that has not joined an etcd cluster.
Request: GetClusterInfoRequest
message GetClusterInfoRequest {
// no fields
}Response: GetClusterInfoResponse
message GetClusterInfoResponse {
optional bool clustered = 1;
optional string self_node_id = 2;
optional uint32 replication_factor = 3;
optional uint64 gc_epoch = 4;
repeated ClusterMember members = 5;
repeated ClusterPin pins = 6;
repeated ClusterRef refs = 7;
}| Field | |
|---|---|
clustered | False if this node has not joined an etcd cluster. members still carries this node's own entry — a standalone node is a cluster of one, and its profile is worth reporting — while pins and refs are empty. |
self_node_id | This node's own id. |
replication_factor | Replication factor for placed state. |
gc_epoch | Current GC epoch (the shed-grace clock). |
members | |
pins | |
refs | Refs for the request tenant only. |
ListGcPasses(ListGcPassesRequest) -> ListGcPassesResponse
Recent distributed-GC sweeps, newest first, across every node.
Each node writes its own sweeps to etcd, so this answers cluster questions a per-node log cannot: is GC keeping up, did every node sweep after that unpin, which node is failing. Empty on a node that has not joined a cluster — there is no shared place for the records to live.
Request: ListGcPassesRequest
message ListGcPassesRequest {
optional uint32 limit = 1;
}| Field | |
|---|---|
limit | Most recent passes to return, across all nodes. Zero means the server's default. |
Response: ListGcPassesResponse
message ListGcPassesResponse {
repeated GcPass passes = 1;
optional bool clustered = 2;
}| Field | |
|---|---|
passes | |
clustered | False on a node with no etcd cluster: an empty list then means "nowhere to look", not "nothing happened". |
Types used above
ClusterMember
One live cluster member (an etcd-leased node).
message ClusterMember {
optional string node_id = 1;
repeated string addrs = 2;
optional uint64 joined_unix = 3;
optional bool is_self = 4;
optional InitialStorageSync initial_storage_sync = 6;
optional ReplicationStatus replication = 7;
oneof _profile {
MemberProfile profile = 5;
}
}| Field | |
|---|---|
node_id | Stable node identity: the peer id (multihash of the node's Ed25519 key). |
addrs | Where the member can be reached: the "host:port" internal-gRPC endpoints it advertised. Members running older daemons publish multiaddrs here instead, and a mixed cluster still returns those. Empty means the member advertised no route and cannot be dialed. |
joined_unix | Unix seconds when the member joined. |
is_self | True if this member is the querying node itself. |
profile | What this member can run and how much room it has. Absent from a member whose daemon predates the profile — read it as "unknown", never as "nothing": a member with no profile may still run everything. |
initial_storage_sync | How far this member is through re-homing the blocks it held before its last restart. A member still syncing holds blocks whose owners moved while it was down, and runs no GC until it is done. |
replication | How far behind this member's block replication is running. A member whose queue is growing writes blocks faster than it hands them to their owners, and until it catches up those blocks exist only on its disk. |
MemberProfile
What a node can run and how much room it has. Published in its member record and refreshed whenever the ledger moves.
message MemberProfile {
optional string arch = 1;
optional string os = 2;
repeated NodeFunction functions = 3;
repeated NodeRuntime runtimes = 4;
repeated string capabilities = 5;
optional MemberResources resources = 6;
optional uint64 updated_unix = 7;
}| Field | |
|---|---|
arch | Host architecture in OCI spelling ("amd64") — the node's default architecture, exposed to build code as the bldr.default-architecture build option. |
os | Host operating system ("linux"). |
functions | Every function the node can execute, sorted by name. |
runtimes | Every configured container runtime, sorted by id. |
capabilities | The optional build-code features the node implements — the same list it hands its build controllers ("pod.cache-volumes", "pod.dag.volumes", …). |
resources | The admission ledger as it stands. |
updated_unix | Unix seconds of the last publish. A profile that has not changed is not republished, so this is when the node last looked different — liveness is the etcd lease's job, not this field's. |
NodeFunction
One function a node can execute.
message NodeFunction {
optional string name = 1;
optional string kind = 2;
optional NodeResources grant = 3;
optional bool sizes_own_grant = 4;
}| Field | |
|---|---|
name | The registered name, e.g. "pod.dag", "deployment/local.shell/v0.0.0". |
kind | Which registry it belongs to: "build", "deployment", "content-provider", "content-tracking", "build-controller". |
grant | What one execution is granted when it does not size itself. |
sizes_own_grant | True when the function derives its own grant from its input (pod.dag does, from the DAG's declared resources) — grant is then the node's default for it rather than a promise. |
NodeResources
CPU and memory as one quantity, in the units the node's admission ledger counts: CPU in millicores (1000 = one core), memory in bytes.
message NodeResources {
optional uint64 cpu_milli = 1;
optional uint64 memory = 2;
}NodeRuntime
One container runtime a node has configured.
message NodeRuntime {
optional string id = 1;
optional string kind = 2;
optional string arch = 3;
}| Field | |
|---|---|
id | The configured id a pod.dag names ("native", "vm", …). |
kind | The backend behind it: "native" or "cloud-hypervisor". |
arch | The architecture it executes, in OCI spelling ("amd64", "arm64"). A runtime's full name is "<id>/<arch>": the id alone does not say whether the node can run a given image. |
MemberResources
A node's admission ledger — what it may hand out to executions, what is outstanding, and how deep the queue is.
message MemberResources {
optional NodeResources capacity = 1;
optional NodeResources granted = 2;
optional NodeResources available = 3;
optional uint32 running = 4;
optional uint32 queued = 5;
optional uint32 max_concurrent = 6;
optional bool gate_enabled = 7;
}| Field | |
|---|---|
capacity | What the node may have outstanding across all admitted executions. |
granted | The sum of every grant currently held. |
available | capacity − granted: what a new execution could be granted right now. |
running | Admitted executions running. |
queued | Executions queued for admission. |
max_concurrent | The concurrency cap running is measured against. |
gate_enabled | False when the node admits everything unconditionally; capacity then describes the host but bounds nothing. |
InitialStorageSync
A member's initial-storage-sync progress. syncing false is the steady
state; the counters are meaningful only while it is true.
message InitialStorageSync {
optional bool syncing = 1;
optional uint64 blocks_checked = 2;
optional uint64 blocks_replicated = 3;
optional uint64 total_blocks = 4;
}| Field | |
|---|---|
syncing | |
blocks_checked | Blocks whose placement has been resolved so far. |
blocks_replicated | Blocks handed to an owner that did not have them. |
total_blocks | What the store estimated it held when the scan started. An estimate, so blocks_checked may pass it slightly by the end. |
ReplicationStatus
A member's block-replication backlog.
message ReplicationStatus {
optional uint64 queued = 1;
optional uint64 last_delay_ms = 2;
}| Field | |
|---|---|
queued | Blocks written but not yet offered to every owner. |
last_delay_ms | How long the most recently replicated block waited between being written and reaching its owners. The last one rather than an average: an average over a burst hides the tail, and the question is how stale a block written right now can be. |
ClusterPin
One cluster pin: a DAG root replicated to its HRW owners.
message ClusterPin {
optional string cid = 1;
optional string origin = 2;
repeated string owners = 3;
}| Field | |
|---|---|
cid | The pinned root CID (base32). |
origin | Peer id recorded as the initial pull source. |
owners | The node ids currently responsible for it under HRW placement (rf owners). |
ClusterRef
One distributed ref: a tenant-scoped mutable name → root-CID pointer.
message ClusterRef {
optional string name = 1;
optional string cid = 2;
optional int64 revision = 3;
}| Field | |
|---|---|
name | |
cid | The CID it points at (base32). |
revision | etcd mod_revision, for compare-and-swap. |
GcPass
One node's account of one distributed-GC sweep.
message GcPass {
optional string node_id = 1;
optional google.protobuf.Timestamp started_at = 2;
optional uint64 epoch = 4;
optional GcPassStatus status = 5;
optional uint64 markers_dropped = 6;
optional uint64 pins_examined = 7;
oneof _finished_at {
google.protobuf.Timestamp finished_at = 3;
}
oneof _error {
string error = 8;
}
}| Field | |
|---|---|
node_id | The node that ran it. |
started_at | When it began. Also its identity: a node's sweeps are keyed by start time. |
finished_at | When it ended. Absent while running — and also when the node died holding the record, which is what an old started_at with no finished_at means. |
epoch | The GC epoch the sweep observed: the grace clock shedding is timed against. |
status | |
markers_dropped | Pin anchors dropped — reclaimed (nobody wants it) plus shed (someone else owns it now). |
pins_examined | Cluster pins the sweep considered. |
error | Why it failed, when it did. |
GcPassStatus
enum GcPassStatus {
GC_PASS_STATUS_UNSPECIFIED = 0;
GC_PASS_STATUS_RUNNING = 1;
GC_PASS_STATUS_COMPLETED = 2;
GC_PASS_STATUS_FAILED = 3;
}