Threat Model: 5-Spot ScheduledMachine Controller¶
Version: 1.5
Date: 2026-10-05
Status: Active — living document
Covers: ADR-0001 … ADR-0015 (the 1.5 pass covers ADR-0015, the dependency
internalization policy. Every section walked. It is a process decision with a
real posture effect in both directions, so §6 T10 and the §7 supply-chain table
now state the trade explicitly: each dependency removed is one fewer upstream to
compromise, and one fewer component the SBOM-driven controls report on. A new LOW
residual records the current exposure (19 lines). No component, asset, actor or
boundary changes. The 1.4 pass covers ADR-0014: a host-governance
conflict now drives capacity to zero instead of refusing to write. Every
section was walked. Changed: §6.6 C5 from "mitigated with a known gap" to
mitigated, and §8's MEDIUM residual removed rather than reworded, because
the gap is closed rather than accepted, replaced by a LOW in the other
direction. No component, asset, actor or boundary changes: the same writer
writes the same field, only the value chosen in one branch differs.
The 1.3 pass implements ADR-0011, which was
Accepted but unimplemented at 1.2: a second controller identity and the first
boundary where 5-Spot writes an API group it does not own. Every section was
walked. Changed: §1 scope (a second binary, a second CRD, and the consumer
declared out of scope), §2 overview and flows F8/F9, §3 three assets, §4 TB7
and its diagram, §5 two actors, §6.6 with nine threats, §7 seven controls, §8
two residuals, §9 two assumptions. §10 unchanged.)
Classification: Public. This page ships in the published documentation site
(docs/mkdocs.yml → Security → Threat Model) of a public repository. It states
the posture — what is defended, from whom, with which control in which file.
Specific unremediated findings do not belong here; report those privately per
SECURITY.md.
1. Document Scope¶
This threat model covers the 5-Spot controller — a Kubernetes operator that manages the lifecycle of physical machines in k0smotron-backed CAPI clusters based on configurable time schedules.
It identifies assets, trust boundaries, threat actors, per-component STRIDE threats, current mitigations, and residual risk. It is intended to inform security reviews, deployment hardening decisions, and future development.
In scope:
- The controller process (five_spot binary) and its Kubernetes RBAC surface
- The 5spot-capacity-controller process and its separate RBAC surface (ADR 0011)
- The ScheduledMachine and ScheduledCapacity Custom Resource Definitions and their admission paths
- All Kubernetes API interactions (CAPI Machine, Bootstrap, Infrastructure, Nodes, Pods), and the single-field patch into an allowlisted foreign API group
- The metrics (/metrics) and health (/healthz, /readyz) HTTP endpoints of both controllers
Out of scope: - The underlying k0smotron / CAPI infrastructure providers - The capacity consumer: the object 5-Spot patches, and everything it does with the capacity it is granted. 5-Spot writes a number and waits; claim binding, guest isolation and drain belong to the consumer (ADR 0011, §9 assumption 7) - The physical machines being managed - The Kubernetes API server itself - Network-level threats (CNI, firewall policy)
2. System Overview¶
Key Data Flows¶
| Flow | Description | Protocol | Auth |
|---|---|---|---|
| F1 | User → API Server: create/update ScheduledMachine | HTTPS | Kubernetes RBAC |
| F2 | Controller → API Server: watch ScheduledMachine events | HTTPS/Watch | Service account JWT |
| F3 | Controller → API Server: create/delete CAPI resources | HTTPS | Service account JWT |
| F4 | Controller → API Server: cordon Node, evict Pods | HTTPS | Service account JWT |
| F5 | Prometheus → Controller: scrape metrics | HTTP (no TLS) | None (cluster-internal) |
| F6 | Kubernetes probes → Controller: health checks | HTTP (no TLS) | None (cluster-internal) |
| F7 | Controller → Node: stamp the kata-config-ref annotation naming the workload-cluster source object |
HTTPS | Service account JWT |
| F8 | Capacity Controller → API Server: watch ScheduledCapacity, its spot-schedule provider and its target object |
HTTPS/Watch | Capacity service account JWT (separate identity) |
| F9 | Capacity Controller → API Server: patch one numeric field on an object in an allowlisted foreign API group. Never create or delete |
HTTPS | Capacity service account JWT (separate identity) |
| F8 | Kata-config agent → workload API: get the named ConfigMap/Secret, then write /etc/k0s/containerd.d/kata.toml on the host and restart a systemd unit through nsenter |
HTTPS, then host namespaces | Per-pod service account JWT, then host PID 1 |
| F9 | Reclaim agent → Node: write the three reclaim annotations | HTTPS | Per-pod service account JWT |
3. Assets¶
| Asset | Sensitivity | Description |
|---|---|---|
| Physical machine availability | Critical | Machines being added/removed from cluster; unintended removal causes workload disruption |
| CAPI cluster integrity | Critical | Bootstrap and infrastructure resources that define cluster membership |
| Node workloads | High | Running pods that could be evicted during a drain operation |
| Controller service account credentials | High | JWT token granting cluster-wide RBAC privileges |
| ScheduledMachine spec data | Medium | Contains infrastructure topology (addresses, ports, SSH config) |
| Kubernetes RBAC posture | High | Overly broad permissions on the service account expand blast radius of compromise |
| Cluster-wide node state | High | Cordon/drain operations affect all workloads on targeted nodes |
ScheduledCapacity spec data |
Medium | Names a foreign object and a field path on it. The path is user-controlled input that the capacity controller writes with its own credential, which makes it a confused-deputy surface rather than inert configuration (§6.6 C1) |
| Conceded host capacity | High | The slice of a host handed to a consumer while it stays in the cluster. Over-provisioning starves the incumbent workload the host exists to run; under-provisioning or a missed handback silently wastes it |
| Capacity controller service account credentials | Medium | JWT granting patch on one allowlisted API group and its own CRD's status. Deliberately lower sensitivity than the machine controller's: no Secret access, no CAPI verbs, no create or delete anywhere (§6.6 C6) |
4. Trust Boundaries¶
TB7 is the only place 5-Spot writes an API group it does not own (ADR 0011). The capacity controller patches one numeric field on a consumer's object so a spot schedule can gate a bounded slice of a host that stays in the cluster, which is the opposite of every other flow here: TB3 creates and deletes CAPI resources, TB7 only ever scales a field on an object someone else owns.
Three properties define the boundary, and each is a deliberate narrowing rather than an accident of implementation:
- A separate identity. Its own binary, Deployment and ServiceAccount, so the
grant cannot compound the machine controller's. Compromising the capacity
controller cannot delete a CAPI
Machine; compromising the machine controller cannot write capacity. This is the whole reason ADR 0011 chose a new binary over a second reconciler in the existing one, and it is why the two HIGH residuals in §8 do not grow. patchis the only write verb, on an allowlisted API group, with nocreate, nodelete, no Secret access and no CAPI verbs (deploy/capacity-controller/clusterrole.yaml). A compromised capacity controller cannot bring a consumer's object into existence or destroy it, which is what makes the design safe against a pool holding live claims.- The field path is user input and is treated as such.
spec.capacity.pathnames a field on a foreign object and comes from whoever can create aScheduledCapacity, so it is constrained at admission and again in the reconciler. See §6.6.
A cluster with no capacity consumer installs none of TB7.
TB6 is the highest-privilege boundary in the system. The kata-config agent
is the only component that runs privileged: true, and its work — writing a
containerd drop-in to the host filesystem and restarting a host systemd unit
through nsenter -t 1 — is node-root by construction (ADR 0002, ADR 0003). It
crosses from the Kubernetes API into the host's mount, UTS, IPC, network and PID
namespaces. Everything in §6.5 follows from that.
5. Threat Actors¶
| Actor | Capability | Motivation |
|---|---|---|
| Malicious tenant | Can create/edit ScheduledMachines in their namespace | Escape namespace, disrupt other tenants, exfiltrate data |
| Compromised CI/CD | Can push images or apply manifests | Backdoor controller binary, escalate privileges |
| Compromised controller pod | Has the controller's service account | Lateral movement to CAPI resources, node disruption |
| Rogue operator | Internal user with broad kubectl access | Misuse kill switch, drain nodes during business hours |
| Supply chain attacker | Can inject into upstream crates (kube-rs, serde, etc.) | RCE inside controller, credential theft |
| Spot-schedule provider (untrusted CRD) | Owns status.active of a spotschedules.5spot.finos.org object referenced by a ScheduledMachine |
Flap the referenced machines on/off — effectively the same control surface as spec.enabled |
| Compromised capacity controller pod | Has the capacity controller's service account: patch on the allowlisted target group and on its own CRD's status, read-only elsewhere |
Starve or over-provision a consumer's capacity. Cannot delete a CAPI Machine, read a Secret, or create/delete the governed object: the separation from the machine controller is the control |
ScheduledCapacity author |
Can create a ScheduledCapacity in their namespace, naming a target object and a field path on it |
Reach a field the author should not be able to write (metadata.ownerReferences, finalizers, labels) on an object in an API group 5-Spot holds patch on, i.e. use the controller as a confused deputy |
6. STRIDE Threat Analysis¶
6.1 ScheduledMachine Custom Resource (Trust Boundary: TB1)¶
Spoofing¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| S1 | User crafts a ScheduledMachine that mimics another tenant's resource name to confuse monitoring/alerting | Low | Low | Accepted — names are unique within namespace |
| S2 | Attacker spoofs ownerReference UID to claim ownership of existing resources |
Low | Medium | Mitigated — UID is set server-side by the controller, not from user input |
Tampering¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| T1 | Cross-namespace resource creation — user sets bootstrapSpec.metadata.namespace to kube-system |
High | Critical | Mitigated (2026-04-08, hardened 2026-05-30) — controller always uses the SM's own namespace; metadata.namespace/metadata.name are now loudly rejected at admission (VAP rules 13c–13f) and at reconcile (validate_embedded_metadata()). Only metadata.labels/metadata.annotations are accepted, reserved-prefix-checked |
| T2 | Label injection — user sets machineTemplate.labels["cluster.x-k8s.io/cluster-name"] to redirect machine to attacker-controlled cluster |
High | Critical | Mitigated (2026-04-08) — validate_labels() blocks reserved prefixes |
| T3 | Annotation injection — user injects kubectl.kubernetes.io/restartedAt to trigger rolling restarts |
Medium | Medium | Mitigated (2026-04-08) — same prefix allowlist |
| T4 | apiVersion/kind injection — user sets bootstrapSpec.kind: ClusterRole to create RBAC resources |
High | High | Mitigated (2026-04-08) — validate_api_group() enforces allowlist |
| T5 | User injects malicious content into bootstrapSpec.spec or infrastructureSpec.spec targeting provider vulnerabilities |
Medium | High | Partially mitigated — spec content is passed opaquely to providers; provider-side validation is out of scope |
| T6 | Timezone log injection — user injects newlines/control chars into timezone field to poison structured logs | Low | Low | Mitigated (2026-04-08) — CRD schema enforces pattern: ^[A-Za-z][A-Za-z0-9_+\-/]*$ and maxLength: 64 |
| T7 | Duration overflow — user sets gracefulShutdownTimeout: "9999999999999h" causing integer overflow |
Medium | High | Mitigated (2026-04-08) — checked_mul + MAX_DURATION_SECS = 86400 cap |
| T8 | User updates ScheduledMachine spec after machine is active, changing clusterName mid-lifecycle |
Medium | Medium | Residual risk — spec changes trigger reconciliation; no immutability enforcement on clusterName |
| T8a | Provider-driven activation control — a compromised or buggy spot-schedule provider flips status.active, starting/stopping the machines that reference it |
Medium | Medium | Mitigated (2026-06-14) — blast radius equals editing spec.enabled and is bounded to same-namespace SMs that explicitly named the provider (cross-namespace refs are forbidden by design); controller RBAC is read-only (get/list/watch) on spotschedules.5spot.finos.org — it can never write a provider object; the provider group is CEL-pinned at the CRD field, in the VAP (rule 4), and at reconcile (validate_activation_source()). Auditable via status.spotSchedule (resolved/reason/message) + fivespot_spot_schedule_* metrics (ADR 0006) |
Repudiation¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| R1 | No audit trail when kill switch is activated | Medium | High | Residual risk — Kubernetes audit log captures the CR edit, but controller logs only emit a single line; no structured event emitted |
| R2 | No record of which schedule window caused machine removal | Low | Low | Accepted — status conditions record transition timestamps |
Information Disclosure¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| I1 | Error messages in ReconcilerError echo user-provided values (timezone, duration, API groups) verbatim |
Low | Low | Residual risk — these are operator-visible logs, not exposed to end users; impact limited |
| I2 | bootstrapSpec.spec may contain infrastructure addresses, credentials, or SSH keys in plaintext |
High | High | Residual risk — see Section 8 |
| I3 | ScheduledMachine status exposes nodeRef, machineRef — leaks infrastructure topology to namespace readers |
Low | Low | Accepted — intentional observability; readers in the same namespace are trusted |
Denial of Service¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| D1 | Attacker creates thousands of ScheduledMachines to overwhelm the controller reconciliation queue | Medium | Medium | Partially mitigated — Kubernetes resource quotas and admission webhooks can cap CR count; controller has CPU/memory limits |
| D2 | Finalizer hang — drain operation never completes, blocking namespace deletion indefinitely | Medium | High | Mitigated (2026-04-08) — tokio::time::timeout(600s) wraps finalizer cleanup |
| D3 | Attacker crafts an oversized ScheduledMachine CR to slow admission validation |
Low | Low | Mitigated — daysOfWeek/hoursOfDay arrays now live on the TimeBasedSpotSchedule CRD (validated by its own schema, ADR 0009), not on the ScheduledMachine VAP; Kubernetes limits CR size in either case |
| D4 | User triggers kill switch repeatedly causing rapid machine add/remove cycles (thrashing) | Low | Medium | Accepted — kill switch is write-once-by-design; no automatic reactivation |
| D5 | Provider flapping — a spot-schedule provider toggles status.active rapidly, churning expensive machine create/delete cycles |
Low | Medium | Partially mitigated (2026-06-14) — transitions are bounded by the controller's reconcile back-off and the machine lifecycle's own grace/drain timers; fivespot_spot_schedule_transitions_total exposes the flap rate for alerting (see monitoring). A per-SM minimumStateDuration debounce is recorded as future work in ADR 0006 |
Elevation of Privilege¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| E1 | Attacker uses apiVersion: rbac.authorization.k8s.io/v1, kind: ClusterRole to create RBAC resources via controller |
High | Critical | Mitigated (2026-04-08) — API group allowlist |
| E1a | Escalation through the controller — user who can create a ScheduledMachine but not the embedded bootstrap/infrastructure resource has the broadly-permissioned controller create it on their behalf |
High | High | Mitigated (2026-05-29) — VAP authorizer rules (13a/13b) require the requesting user to hold create on the embedded GVKs; controller's own SA independently gated at reconcile by ensure_can_create() (SelfSubjectAccessReview) |
| E2 | Compromised controller pod gains cluster-wide node cordon/drain, CAPI write access | Medium | Critical | Partially mitigated — k0smotron.io RBAC narrowed; bootstrap/infra still use wildcards (provider-agnostic requirement) |
| E3 | Controller service account token stolen from pod filesystem | Low | Critical | Mitigated — read-only root filesystem; token mounted at standard path (Kubernetes default); no projected service account with long lifetime |
6.2 Controller Process (Trust Boundary: TB2)¶
Spoofing¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| S3 | Attacker replaces controller image with backdoored binary | Low | Critical | Residual risk — mitigated by image signing (release workflow uses Cosign); IfNotPresent pull policy means in-cluster image is trusted |
Tampering¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| T9 | Memory corruption in unsafe Rust code or via malformed serde input | Very Low | Critical | Mitigated — codebase is safe Rust; no unsafe blocks; serde handles malformed input with errors |
| T10 | Supply chain attack via malicious crate version | Low | Critical | Residual risk: mitigated by Grype container scan (VEX-aware) in CI; no automated dependency pinning beyond Cargo.lock. ADR 0015 cuts both ways here: every dependency removed is one fewer upstream that can be compromised, and also one fewer component a scanner reports on, because internalized code leaves the SBOM. The 500-line bound and the test requirement are what keep the second effect small; see the §7 supply-chain note and §8 |
Information Disclosure¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| I4 | /metrics endpoint accessible without authentication |
Medium | Low | Accepted — metrics contain no secrets; exposes operational data only; restricted to cluster network |
| I5 | Service account token exposed via RUST_LOG=trace debug output |
Low | Medium | Residual risk — RUST_LOG=debug set in deployment; kube-rs does not log JWT tokens; verify before production |
Denial of Service¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| D5 | Pod eviction during drain exhausts Kubernetes API rate limits (429 responses from PDB) |
Medium | Medium | Mitigated — evict_pod handles 429 gracefully and logs a warning rather than crashing |
6.3 Node Drain Path (Trust Boundary: TB4)¶
Tampering¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| T11 | Attacker creates ScheduledMachine whose clusterName matches a production cluster, triggering drain of production nodes outside schedule |
Medium | Critical | Partially mitigated — clusterName is used only as a label on the created CAPI Machine; actual drain targets the node resolved via the Machine's nodeRef |
Denial of Service¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| D6 | Grace period expires mid-drain; pods are forcefully killed without completing shutdown hooks | Medium | Medium | Accepted — POD_EVICTION_GRACE_PERIOD_SECS = 30 is configurable via constant; PDB protection applies |
| D7 | Api::all() for Nodes/Pods fetches cluster-wide list; maliciously large cluster could cause memory spike |
Low | Low | Accepted — field selector spec.nodeName=<node> scopes pod list to one node |
6.4 Reclaim Agent (Trust Boundary: TB5)¶
The node-side 5spot-reclaim-agent DaemonSet runs on opted-in
worker nodes only and signals the controller via three Node
annotations. It is the lowest-trust component in the system: pod
runs as UID 0 with hostPID: true, host /proc mount, host
/etc/machine-id mount, and CAP_NET_ADMIN (for rung 2 netlink
proc connector). Per-pod ServiceAccount grants only
nodes: get,patch cluster-wide and configmaps: get,list,watch
in 5spot-system.
Spoofing¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| S5 | Attacker with update daemonsets overrides the agent pod's NODE_NAME env var, causing the agent to PATCH reclaim annotations on a victim Node it isn't actually running on |
Low (precondition is cluster-admin-equivalent in most clusters) | High (would trigger emergency-remove on innocent host) | Mitigated — host-identity verification (read_host_machine_id + compare_machine_ids, security-audit Phase 4, 2026-04-26): agent reads /etc/machine-id from a hostPath.type: File mount at startup and refuses to PATCH a Node whose status.nodeInfo.machineID doesn't match. --skip-host-id-check opt-out exists but defaults off. |
| S6 | Compromised pod somewhere else in the cluster spoofs the reclaim annotations directly on a Node, triggering an emergency-remove the agent never decided on | Low | High | Partially mitigated — nodes: patch is broadly held in most clusters (kubelet itself has it); the controller treats the annotation as authoritative. Mitigation: monitor 5spot.finos.org/reclaim-requested PATCH events via API audit logs, alert when the field manager is anything other than 5spot-reclaim-agent. |
Tampering¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| T12 | Compromised agent escalates to other Nodes via stolen ServiceAccount token | Low (per-pod SA, scoped credentials) | Medium | Mitigated — RBAC: nodes: get,patch cluster-wide (no list/watch — can't enumerate). Host-identity check refuses any PATCH against a Node whose machineID doesn't match the agent's own /etc/machine-id. |
| T13 | CAP_NET_ADMIN (granted for rung 2 netlink) is misused to send rogue netlink messages on other subsystems |
Low | Low–Medium | Accepted — CAP_NET_ADMIN is Linux's coarsest-grained networking cap and grants more than just netlink connector access (route table edits, iptables, etc.). Scope-bounded in two ways: (1) the cap is added only on opted-in Nodes via the existing 5spot.finos.org/reclaim-agent: enabled nodeSelector, so only the small set of Nodes with killIfCommands set ever sees it; (2) the agent binary is single-purpose with no shell, no exec of children — a compromise would need an attacker to swap the binary, at which point they could grant themselves any cap they wanted. Operators who refuse the cap can pin --detector=poll (rung 1) which needs no extra capability. |
Repudiation¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| R1 | Agent PATCH on the Node is indistinguishable from controller writes in audit logs | Low | Low | Mitigated — the agent uses field manager 5spot-reclaim-agent; the controller uses the distinct field manager 5spot-controller-reclaim-agent. Every Node PATCH carries the field manager in managedFields, so audit attribution is unambiguous (NIST AU-2 / AU-10). |
Information Disclosure¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| I3 | Reading host /proc/<pid>/cmdline for matching exposes argv strings (which can include passwords / tokens passed on the command line) to the agent's logs |
Low (logs are operator-owned) | Low | Accepted — the agent logs matched_pattern (the operator-supplied pattern) and the matched pid, NOT the full cmdline. Patterns themselves are operator-authored config, not user input. |
Denial of Service¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| D8 | User repeatedly re-enables a SM whose conflicting process is still running, generating an ejection loop | Medium | Low | Mitigated — loop-protection: ≥3 reclaims for the same SM within 10 minutes emits a RapidReReclaim Warning Event and bumps fivespot_rapid_re_reclaims_total{namespace, name}. Operator runbook in troubleshooting covers the response. |
| D9 | Heavy-exec workload (make -j32) overwhelms the netlink subscriber's recv loop with PROC_EVENT_EXEC traffic |
Low | Low | Accepted — agent CPU is bounded by container limits (50m); ENOBUFS on the netlink socket surfaces as NetlinkError::Io and is logged, not silently dropped. Operators who anticipate exec storms should pin --detector=poll. |
Elevation of Privilege¶
| ID | Threat | Likelihood | Impact | Status |
|---|---|---|---|---|
| E4 | hostPID: true lets the agent see all host PIDs — could be abused to extract data from another container's /proc/<pid>/environ if the agent is compromised |
Low | Medium | Accepted — hostPID is architecturally required (the agent's job is to read host process state). Mitigated by: opt-in nodeSelector (only on nodes with killIfCommands), single-purpose binary with no shell, readOnlyRootFilesystem: true, drop ALL caps + add only NET_ADMIN, seccompProfile: RuntimeDefault. Trivy / Semgrep suppressions in .trivyignore document the architectural-necessity reasoning. |
| E5 | Abuse of the namespace-wide pod-security exemption — the privileged agents require 5spot-system to be exempted from the cluster's PSA/Gatekeeper/Kyverno baseline; any principal with create pods in the namespace (compromised CI, typo'd Deployment) could then run privileged / mount the host root with no admission check |
Medium | Critical | Mitigated (2026-06-10, ADR 0004) — 5spot-agent-pod-security deny-by-default ValidatingAdmissionPolicy re-imposes the baseline inside 5spot-system: risky attributes pinned to the two agent ServiceAccounts at their exact documented posture (hostPath clamped per agent, caps clamped to NET_ADMIN, privileged to the kata agent only), compensating controls mandatory, hostNetwork/hostIPC and risky ephemeral containers denied outright, failurePolicy: Fail. Residual: a principal who can both create pods and use an agent SA wears the exception — clamped to the agents' documented posture; SA RBAC is the control. |
6.5 Kata-Config Agent (Trust Boundary: TB6)¶
The node-side 5spot-kata-config-agent DaemonSet delivers a Kata containerd
drop-in to opted-in workload-cluster nodes (ADR 0002) and restarts the host k0s
service so containerd reloads it (ADR 0003). It is the only component that
runs privileged: true. Its input is the 5spot.finos.org/kata-config-ref
annotation the controller stamps on the Node (F7): a compact JSON object naming
the workload-cluster namespace, kind, object name, data key, and the systemd
unit to restart. The agent gets that one object, writes
/etc/k0s/containerd.d/kata.toml, and runs
nsenter -t 1 -m -u -i -n -p -- systemctl restart <unit>.
Its per-pod ServiceAccount grants nodes: get,patch cluster-wide and
configmaps,secrets: get (no list, no watch) in the target namespace.
| ID | Threat | STRIDE | Likelihood | Impact | Status |
|---|---|---|---|---|---|
| K1 | Drop-in written outside /etc/k0s/ — a host path supplied through the CR or annotation, or a symlink planted on the host, redirects the write anywhere on the node |
T, E | Low | Critical | Mitigated — no CRD field or annotation carries a host path at all (ADR 0005 removed destPath); the destination is the fixed KATA_CONFIG_DEST_PATH constant, and confine_dest_path canonicalizes and re-checks it against the /etc/k0s/ base before every write and unlink, fail closed (src/kata_config_agent.rs) |
| K2 | Command injection through the restart path | T, E | Low | Critical | Mitigated, twice over — nsenter_restart_argv builds an argv vector, so there is no shell anywhere in the path and metacharacters in any field are inert; and since ADR 0012 the argv carries a second --, between restart and the unit, which ends systemctl's own option parsing. The first -- only ever ended nsenter's (src/kata_config_agent.rs) |
| K3 | Drop-in content is attacker-chosen: containerd configuration is executed-adjacent, and the agent restarts the runtime that reads it | T, E | Low | Critical | Accepted — architecturally required. This is the agent's entire purpose. The control is upstream: whoever may set spec.kata on a ScheduledMachine, or write the Node annotation, is choosing containerd configuration on that node. Treat create/update on scheduledmachines carrying spec.kata, and patch nodes in the workload cluster, as node-root-equivalent grants (§7 deployment controls, §8) |
| K4 | Restart of an unintended host unit | T, D | Low | High | Mitigated for the injection case (ADR 0012) — the argv's option terminator means a value beginning with - is a unit name and not a flag such as -H (remote host over SSH), -M (container) or --root=. The CRD pattern is anchored so a leading hyphen is not admissible (src/crd.rs), and parse_kata_ref re-checks the same rule on the agent's own input. Not mitigated for the policy case: a syntactically valid unit the operator did not intend — sshd.service matches and always will — which is why §7 treats spec.kata as a node-root-equivalent grant |
| K5 | The agent trusts its Node annotation, which RBAC cannot scope to the agent's own Node | S, T, E | Low | High | Mitigated for the agent case (ADR 0012 + ADR 0013). ADR 0012 made the agent distrust the annotation's shape — parse_kata_ref validates every field and refuses a bad one, so no write and no restart. ADR 0013 addresses its authority: deploy/admission/kata-config-annotation-policy.yaml rejects any Node UPDATE that changes 5spot.finos.org/kata-config-ref when the requester is a node-agent ServiceAccount, comparing object against oldObject so their own kata-config-applied writes still pass. A compromised agent can no longer retarget another node. Residual: anything else holding cluster-wide nodes: patch — a human admin, a CI identity — still can; see §8 |
| K6 | Stolen agent token enumerates or reads the workload cluster | I | Low | Medium | Mitigated — get only on configmaps/secrets, no list/watch: a stolen token reaches objects the attacker can already name, not the namespace (deploy/kata-config-agent/rbac.yaml) |
| K7 | The agent's privileged posture is inherited by an unrelated pod in 5spot-system |
E | Medium | Critical | Mitigated (ADR 0004) — 5spot-agent-pod-security VAP pins privileged to the kata agent's ServiceAccount alone, clamps hostPath per agent, denies hostNetwork/hostIPC, failurePolicy: Fail (deploy/admission/agent-pod-security-policy.yaml) |
| K8 | Restart loop — every reconcile bounces the host service | D | Low | High | Mitigated — the applied-content hash is recorded in the kata-config-applied Node annotation and needs_restart fires only on an actual content transition (src/kata_config_agent.rs) |
6.6 Capacity Controller (Trust Boundary: TB7)¶
The 5spot-capacity-controller Deployment reconciles ScheduledCapacity and
patches one numeric field on an object in an API group 5-Spot does not own,
so a spot schedule can gate a bounded slice of a host that stays in the cluster
(ADR 0011). It never creates or deletes that object and never sets an
ownerReference on it.
Its own ServiceAccount grants patch on the allowlisted target group and on
its own CRD's /status, read-only on the spot-schedule group and on
scheduledmachines, plus create on selfsubjectaccessreviews. It holds no
Secret access, no CAPI verbs, and create/delete nowhere.
The interesting attacker here is not the compromised pod but the
ScheduledCapacity author, who supplies a field path that the controller
then writes with its own, broader credential: a confused-deputy shape.
| ID | Threat | STRIDE | Likelihood | Impact | Status |
|---|---|---|---|---|---|
| C1 | spec.capacity.path names metadata.ownerReferences, finalizers or labels, so a CR author uses the controller's credential to take ownership of, block deletion of, or re-label an object in the target group |
T, E | Medium | High | Mitigated, twice over: the path must match CAPACITY_PATH_PATTERN (dot-separated camelCase, so no indices, wildcards, .., quotes or /) and a CEL rule pins it to the spec. prefix, so metadata. and status. are inadmissible (src/crd.rs). The reconciler re-validates before any API call (src/reconcilers/capacity_path.rs), which is the check that still holds when the deployed CRD is older than the binary, precisely the case a schema cannot defend. Verified against a live API server: metadata.ownerReferences, metadata.finalizers, metadata.labels and status.claimed are all rejected at admission |
| C2 | The path is read as a JSON Pointer or JSONPath expression, so ~0, /, [0] or a quoted segment escapes the intended field |
T, E | Low | High | Mitigated: the charset is an allowlist ([a-z][a-zA-Z0-9]* per segment), not a denylist of dangerous characters, so nothing outside camelCase is expressible regardless of how the value is later interpreted. Actuation builds a nested merge patch from the validated segments rather than parsing a path expression (capacity_path.rs::build_merge_patch) |
| C3 | The target is an object in an API group the controller was never meant to write | T, E | Low | High | Mitigated: ALLOWED_CAPACITY_TARGET_API_GROUPS (src/constants.rs) is the third allowlist beside the bootstrap and infrastructure ones, checked in the reconciler so the permitted set lives in one greppable place and cannot be widened by a stale CRD. The ClusterRole grants patch on exactly those groups, so RBAC is the second gate |
| C4 | A capacity write destroys in-flight work: a consumer's pool is zeroed while claims are live | T, D | Medium | High | Mitigated by construction: actuation is a field write, never a delete, and the inactive value is a warm-pool zero: it stops replenishment while claimed work finishes on its own. The consumer owns claim binding and drain, which 5-Spot has no way to reason about. On deactivation the controller writes zero and then waits for spec.handback.drainedPath to reach zero, up to handback.timeout; on expiry it holds and reports loudly rather than escalating (ADR 0011 decision 6), because a missed handover is visible and recoverable while killed work is not |
| C5 | A host is governed by both a ScheduledMachine and a ScheduledCapacity, so it is drained for handover while capacity is still being sized on it |
T, D | Medium | Medium | Mitigated (ADR 0014). spec.nodeName is compared against every ScheduledMachine.status.nodeRef.name in the namespace, and a match sets HostGovernanceConflict=True and drives capacity to zero (src/reconcilers/capacity_decision.rs). The active value is never written to a conflicted object whatever the schedule says, so the fail-closed intent holds where it matters, while zero removes a contradiction that may already exist rather than freezing it. ADR 0011 originally said "refuse to write", which left a value written before the conflict standing: the exact state the invariant forbids, observed live. The check is controller-side, not admission, because a ValidatingAdmissionPolicy evaluates one request against its own object and its paramRef and cannot look up other objects at all. Residual, now LOW: a spec.nodeName colliding with an unrelated machine zeroes capacity the operator wanted, until the reference is corrected; spec.nodeName is optional for exactly that reason. §8 |
| C6 | The capacity grant compounds the machine controller's CAPI and bootstrap/infrastructure wildcards | E | Low | High | Mitigated: this is the reason for the separate binary. A distinct ServiceAccount, Deployment and ClusterRole; the machine controller's ClusterRole is not extended (ADR 0011 decision 3). Verified against a live API server with 26 SubjectAccessReview assertions: the capacity identity cannot delete or create a CAPI Machine, read a Secret, create or delete the governed object, patch a ScheduledMachine, or evict a Pod |
| C7 | An RBAC gap surfaces as a partial write or an opaque 403 mid-actuation | D | Low | Low | Mitigated: a pre-flight SelfSubjectAccessReview for patch on the resolved target runs before the first write, so a missing grant becomes a TargetResolved=False condition naming the resource (ADR 0011 decision 5) |
| C8 | The controller re-patches its own status every reconcile, re-triggering its own watch and flooding the API server | D | Low | Medium | Mitigated: condition lastTransitionTime is preserved unless the condition's status changes, and the computed status is compared against the stored one so a no-op patch is never sent. Found by running against a real API server, which measured ~50 resourceVersion bumps/second before the fix; pinned by a test that derives the expected status key set from the type rather than a hand-written list |
| C9 | A missing or unparseable drain counter is read as "drained", completing a handback that never happened | T, D | Low | Medium | Mitigated: Some(0) and None are deliberately different answers: an unread or non-numeric counter keeps the object in HandingBack and says so in the status message, rather than being collapsed to zero (capacity_path.rs::read_i64_at) |
7. Mitigations Summary¶
Implemented (as of 2026-09-23)¶
| Control | Where | Addresses |
|---|---|---|
EmbeddedResource.namespace field removed |
src/crd.rs |
T1 — cross-namespace creation |
validate_api_group() allowlist |
src/reconcilers/helpers.rs |
T4, E1 — kind/apiVersion injection |
validate_labels() reserved prefix rejection |
src/reconcilers/helpers.rs |
T2, T3 — label/annotation injection |
checked_mul + MAX_DURATION_SECS cap |
src/reconcilers/helpers.rs |
T7 — duration overflow |
Timezone maxLength + character pattern in CRD |
src/crd.rs |
T6 — log injection |
tokio::time::timeout(600s) in finalizer cleanup |
src/reconcilers/helpers.rs |
D2 — finalizer hang |
| k0smotron.io RBAC narrowed to explicit resources | deploy/deployment/rbac/clusterrole.yaml |
E2 — over-privileged SA |
| Non-root container, read-only root filesystem, all caps dropped | deploy/deployment/deployment.yaml |
E3 — token theft |
| CPU/memory resource limits | deploy/deployment/deployment.yaml |
D1 — resource exhaustion |
PDB 429 handling in evict_pod |
src/reconcilers/helpers.rs |
D5 — API rate limit crash |
| Cosign image signing in release CI | .github/workflows/release.yaml |
S3 — image tampering |
| Host-identity verification (machine-id cross-check) | src/reclaim_agent.rs::compare_machine_ids + deploy/node-agent/daemonset.yaml (/etc/machine-id mount) |
S5 — DaemonSet-tampering NODE_NAME spoof |
Distinct field managers (5spot-reclaim-agent vs 5spot-controller-reclaim-agent) |
All Node PATCH paths | R1 — audit attribution |
Per-pod ServiceAccount with nodes: get,patch only (no list/watch) |
deploy/node-agent/rbac.yaml |
T12 — agent enumerates other Nodes |
Opt-in 5spot.finos.org/reclaim-agent: enabled nodeSelector |
deploy/node-agent/daemonset.yaml |
T13 — CAP_NET_ADMIN scope-bounding |
Loop-protection (RapidReReclaim warning + counter) |
src/loop_protection.rs + src/reconcilers/helpers.rs::handle_emergency_remove |
D8 — re-enable loop |
Agent pod-security exception boundary (5spot-agent-pod-security VAP, deny-by-default in 5spot-system) |
deploy/admission/agent-pod-security-policy.yaml + binding (ADR 0004) |
E5 — abuse of the namespace-wide PSA/OPA/Kyverno exemption the privileged agents require |
Capacity write path constrained to spec. camelCase segments (CRD pattern + CEL rule, re-checked in the reconciler) |
src/crd.rs, src/reconcilers/capacity_path.rs |
C1, C2: confused-deputy write to metadata./status., pointer-escape |
| Capacity target API group allowlist, checked in the reconciler and mirrored in RBAC | src/constants.rs (ALLOWED_CAPACITY_TARGET_API_GROUPS), deploy/capacity-controller/clusterrole.yaml |
C3: write to an unintended API group |
Separate capacity identity: own binary, Deployment, ServiceAccount and ClusterRole; patch only, no create/delete/secrets/CAPI |
deploy/capacity-controller/ (ADR 0011) |
C6: compounding the machine controller's CAPI and bootstrap wildcards |
Pre-flight SelfSubjectAccessReview for patch before the first capacity write |
src/reconcilers/scheduled_capacity.rs |
C7: opaque 403 or partial write on an RBAC gap |
Cooperative bounded handback: write zero, wait on drainedPath, hold and report on timeout; never delete or force |
src/reconcilers/capacity_decision.rs (ADR 0011 decision 6) |
C4: destroying a consumer's in-flight work |
| No-op status patch guard and preserved condition timestamps | src/reconcilers/scheduled_capacity.rs |
C8: self-triggering reconcile loop against the API server |
Supply Chain (TB0 — contributor / CI to published artifact)¶
§5 lists a supply-chain attacker; these are the controls that answer it. All of
them run in .github/workflows/.
| Control | Where | Addresses |
|---|---|---|
| SLSA build provenance on every release | .github/workflows/build.yaml (provenance: true) |
Forged or substituted release artifact |
| Cosign keyless signing, by digest not tag | .github/workflows/build.yaml |
S3 — image tampering after publication |
| SBOM generated per image and per release | .github/workflows/build.yaml (sbom: true, anchore/sbom-action) |
Unknown component inventory |
| Auto-VEX presence + reachability gate, byte-exact, fails closed (ADR 0008) | src/bin/auto_vex_*.rs, .vex/, release workflow |
An advisory shipping untriaged |
| Bounded dependency internalization: <=500 lines, excluded categories, measured on the production target (ADR 0015) | .claude/rules/dependency-internalization.md, roadmap 04 §2a |
T10 in the reducing direction, by shrinking the set of upstreams that can be compromised at all. Its limit is the point: internalized code is not in the SBOM, so the controls above stop covering it, which is why specs, crypto, reference data and proc macros are excluded however thin their surface looks |
Base images digest-pinned on the FROM line, Dependabot-tracked (ADR 0010) |
Dockerfile, Dockerfile.chainguard, .github/dependabot.yml |
A stale or substituted base image |
| Actions SHA-pinned; Dependabot groups and a 7-day cooldown per ecosystem | .github/workflows/*, .github/dependabot.yml |
A compromised action release |
Scanning: CodeQL, Trivy (image + IaC), grype, cargo audit, cargo deny, gitleaks, Semgrep |
.github/workflows/ |
Known-vulnerable dependency, leaked credential |
No pull_request_target, no issue_comment; untrusted event fields reach run: only through env: |
.github/workflows/ |
Workflow script injection from a fork |
Deployment-Layer Controls (operator responsibility)¶
| Control | Recommendation |
|---|---|
| Kubernetes ResourceQuota | Limit ScheduledMachine count per namespace (e.g., max 50) |
ValidatingAdmissionPolicy ✅ deployed 2026-04-08 |
Validates bootstrapSpec.apiVersion, infrastructureSpec.apiVersion, kind fields, duration format, and day/hour item format at admission time — see deploy/admission/ |
| NetworkPolicy | Restrict controller pod egress to the management API server (6443), child/k0smotron-hosted control planes (NodePort apiPort 30443), and DNS — all other egress denied |
| Audit logging | Enable API server audit log at RequestResponse level for scheduledmachines resources |
| RBAC for SM creation | Only grant create on scheduledmachines to trusted identities; do not grant to end users directly |
RBAC for spec.kata |
A ScheduledMachine carrying spec.kata chooses containerd configuration on a node and triggers a host service restart (§6.5). Grant it only to identities you would trust with node root |
RBAC for patch nodes (workload cluster) |
Node-root-equivalent wherever the kata-config agent runs — the agent's input is a Node annotation. See §8 |
| Secrets for bootstrap data | Move sensitive bootstrap config out of CR spec into Secrets; reference from spec |
8. Residual Risks¶
HIGH — Sensitive data in bootstrapSpec.spec¶
Threat: EmbeddedResource.spec is an arbitrary JSON object with x-kubernetes-preserve-unknown-fields: true. It may contain SSH keys, IP addresses, tokens, or other credentials stored in plaintext in etcd and visible to anyone with get scheduledmachines access.
Recommendation: Introduce a secretRef field alongside spec that references a Secret, and merge the Secret's data at runtime inside the controller. This keeps credentials out of the CR and benefits from Kubernetes secret encryption-at-rest.
Workaround (now): Restrict get/list on scheduledmachines to the owning service account and cluster admins only.
HIGH — Bootstrap/Infrastructure RBAC wildcards¶
Threat: bootstrap.cluster.x-k8s.io and infrastructure.cluster.x-k8s.io still use resources: ["*"] in the ClusterRole because the controller is designed to be provider-agnostic. A compromised controller can create any bootstrap or infrastructure resource cluster-wide.
Recommendation: For deployments targeting a single known provider, replace wildcards with explicit resource lists (e.g., k0sworkerconfigs only). Document this in the operator deployment guide.
MEDIUM — clusterName immutability¶
Threat: A user can update spec.clusterName on an active ScheduledMachine. The controller will reconcile with the new cluster name, potentially creating CAPI resources in a different cluster while leaving orphaned resources in the original cluster.
Recommendation: Use a CEL validation rule (x-kubernetes-validations) to make clusterName immutable after creation:
LOW — Kill switch audit trail¶
Threat: Activating spec.killSwitch: true immediately removes a machine. The only record is the Kubernetes API audit log (if enabled). No Kubernetes Event is emitted by the controller.
Recommendation: Emit a Kubernetes Event with reason: KillSwitchActivated and type: Warning when the kill switch fires, so it appears in kubectl describe scheduledmachine and feeds into alerting pipelines.
LOW: A mis-set spec.nodeName zeroes capacity the operator wanted¶
Closing the former MEDIUM (a conflict leaving written capacity in place) by
driving to zero (ADR 0014) trades it for a smaller risk in the other
direction: if spec.nodeName names a node an unrelated ScheduledMachine
happens to govern, the consumer loses its warm pool until the reference is
corrected.
Bounded and self-correcting. The gated field is a warm-pool target rather than
a kill, so claimed work finishes; the condition and message name the
conflicting ScheduledMachine and the node; and the next reconcile after the
reference is fixed restores the active value, with no operator action beyond
that. spec.nodeName is optional precisely so an operator who cannot name the
node confidently leaves it unset, in which case no conflict is detected and the
invariant stays documentation only.
Revisit when: a Kubernetes Event is added for the conflict (ADR 0014 option D, not adopted), which would make this page someone rather than wait to be noticed.
LOW: Internalized code is outside dependency scanning¶
ADR 0015 internalizes a dependency when its used surface is 500 lines or fewer.
Each one removes an upstream that could be compromised, and simultaneously
removes a component that cargo audit, cargo deny and the Grype scan report
on: a defect in src/stream.rs is found by review, not by a scanner.
Accepted deliberately, and bounded by the policy rather than by a control. The 500-line ceiling is low, every internalization must be testable at least as well as upstream for the surface used, and the categories where a thin surface hides a thick moving specification (wire formats, crypto, reference data, date/time arithmetic, proc macros) are excluded outright. Current exposure is 19 lines.
Revisit when: the internalized total approaches a few hundred lines, or an internalization is proposed that cannot be pinned by tests. Either is a signal the ceiling is doing less work than it appears to.
LOW: Capacity rows written under the old scheme are not reclaimed¶
Not a 5-Spot control: noted because operators will see it. Where a consumer's status aggregates per-writer rows, rows written before a writer's identity changes are not taken over by the new identity and persist until the consumer prunes them. 5-Spot neither writes nor reads those rows.
Revisit when: a consumer asks 5-Spot to participate in pruning.
MEDIUM — Node-side agents hold cluster-wide nodes: patch¶
Threat: Both DaemonSet agents act only on their own Node — the reclaim agent
writes its three reclaim annotations, the kata-config agent records the applied
hash and clears the opt-in label on tear-down — but Kubernetes RBAC cannot
express "the Node this pod runs on". Both ServiceAccounts therefore hold
nodes: patch across the cluster. list/watch are deliberately withheld
(T12, K6), so an agent cannot enumerate the cluster; it can still write to a
Node it can name.
This matters more for the kata-config agent than for the reclaim agent, because
the 5spot.finos.org/kata-config-ref annotation is not merely data: it is an
instruction consumed by a privileged component on the node that reads it
(§6.5, K5). The Node annotation — not the CRD — is that agent's real input, and
the ScheduledMachine schema does not gate it.
Narrowed twice on 2026-09-27. ADR 0012 made the agent validate every field of
the annotation before acting on it, so a malformed or option-shaped value is
refused rather than executed. ADR 0013 then took the recommendation below and
shipped it: deploy/admission/kata-config-annotation-policy.yaml denies the
node agents any change to 5spot.finos.org/kata-config-ref, which closes the
node-to-node lateral path this entry was written about.
What is left is narrower and belongs to the operator: the agents' grant is
still cluster-wide, and other holders of nodes: patch — a human admin, a CI
identity, other tooling — can still write the key. Per-node ServiceAccounts with
resourceNames would close that and were rejected as a lifecycle system in
exchange for a bound admission gives for nothing (ADR 0013).
The control that shipped, for reference: a
ValidatingAdmissionPolicy on nodes UPDATE that permits a change to the
5spot.finos.org/kata-config-* annotation keys only when
request.userInfo.username is the controller's ServiceAccount, denying it to
every other principal including the agents themselves. This bounds the
annotation to its one legitimate writer without needing per-node ServiceAccounts.
Workaround (now): Treat patch nodes in a workload cluster running the
kata-config agent as a node-root-equivalent grant, and review every binding that
carries it. Audit changes to 5spot.finos.org/kata-config-ref — the controller's
field manager is the only expected writer.
LOW — Multi-instance hash distribution weakness¶
Threat: The consistent hash function adds priority * 1000 to a 64-bit hash, which provides negligible differentiation. High-priority resources may cluster on one instance.
Recommendation: Use a proper consistent hash ring (e.g., rendezvous hashing) when HA multi-instance support is hardened.
9. Security Assumptions¶
The following conditions are assumed to be true for this threat model to hold:
- Kubernetes API server is trusted — requests to the API server are authenticated and authorized; no API server vulnerabilities are in scope.
- etcd encryption at rest is enabled — CR specs (which may contain infrastructure details) are encrypted in etcd.
- RBAC for ScheduledMachine creation is restricted — only trusted users/service accounts have
createpermission onscheduledmachines. - Container image integrity — the controller image is pulled from a trusted registry; image signing is enforced.
- Cluster network is trusted — the metrics and health endpoints are not accessible from outside the cluster.
- Node-level isolation — physical machines managed by 5-Spot do not share sensitive workloads with other tenants.
- The capacity consumer enforces its own drain semantics: 5-Spot writes a capacity field and waits; it has no way to evict, reclaim or account for work already running inside the slice. A consumer that reports drained while work continues, or that never reports at all, is outside 5-Spot's control (C4, C9). This is the deliberate division of labour in ADR 0011: the consumer owns claim binding and drain, 5-Spot owns only the calendar.
createonscheduledcapacitiesis restricted: the field path in aScheduledCapacityis written by the controller with its own credential, so the grant is a confused-deputy surface. The path is constrained (C1, C2) and the target group is allowlisted (C3), but whoever may create one still chooses which allowlisted object gets scaled and by how much. Treat it with the same care ascreateonscheduledmachines.
10. Related Documents¶
- Architecture
- Machine Lifecycle
- RBAC Configuration — a repository file, not a site page
- API Reference
- Developer Guide — the ADD cycle and the decision log