Okube
A small Go container orchestrator built by a team to explore scheduling and distributed coordination. I implemented etcd-backed task and worker state, registration and heartbeats, leader election, pending-task dispatch, CLI flows, and network-aware multi-node scheduling.
- Scope
- Team
- Status
- Prototype
- Stack
- Go · Docker · etcd · REST · YAML
01 Context
The problem
A scheduler needs a consistent picture of workers, a single active decision-maker, and a way to place work despite changing resource and network conditions. Okube explores those mechanics without pretending to be a general-purpose cluster platform.
My role
In the team build, I implemented the CLI-to-worker path, etcd task and worker stores, registration and heartbeats, manager leadership, pending-task dispatch, network-aware scheduling, and multi-node support.
- Implemented etcd-backed worker/task state, registration, and heartbeats.
- Built leader/follower behavior, election, pending-task watching, and manager CLI flows.
- Added network-aware placement and multi-node scheduling support.
02 Architecture
System at a glance
03 Decisions & trade-offs
Coordination through explicit state
Managers and workers share task and membership state through etcd. Workers register, publish heartbeats, and report lifecycle changes. Managers participate in leader election, but only the leader watches pending tasks and dispatches placement decisions; followers retain the information needed to take over manager leadership.
The design gives the educational prototype a concrete coordination model. It does not make every component highly available, and it does not secure manager-to-worker traffic for untrusted networks.
Placement is a pipeline, not a single score
Scheduling separates feasibility from preference. Resource and network filters remove unsuitable workers first. Resource and latency-aware scores then compare the remaining candidates before selection. This makes it possible to explain why a worker was rejected instead of hiding every concern in one opaque number.
I implemented the network-aware and multi-node path alongside the task and worker stores that supply the scheduler’s inputs. A simpler round-robin strategy remains useful when richer placement data is not needed.
Service order and failure behavior
Multi-service manifests are ordered from their dependency graph. Services start in dependency order, receive environment-based discovery information, and are torn down in reverse order. Health and restart behavior handles failed containers within the project’s intentionally small operating model.
The missing capabilities are as important as the implemented ones: there is no overlay network, horizontal scaling, broad security model, or evidence for production use. Okube is a focused way to study leadership, scheduling, state, and container lifecycle—not a Kubernetes replacement.
04 Evidence
Inspect the work
- Canonical team source (opens in a new tab)Team repository
- Pal's fork (opens in a new tab)
- Verification (opens in a new tab)Build passed; repository currently has no tracked tests
Current status & limitations
- The documented model supports one replica and has no overlay network or horizontal scaling.
- Manager-to-worker communication is plain HTTP intended for a trusted LAN.
- Default vet/test currently reports a non-constant log.Printf format string.