Skip to content
AC-INFO Kft · Hungary
Back to the blog

14 min readGuides

Kubernetes Deployments, DaemonSets, and StatefulSets: a Deep Dive

Which Kubernetes workload controller to choose, and how Deployments, DaemonSets and StatefulSets behave when things break — from a production K3s cluster and a three-hour database outage.

Every introduction to Kubernetes workload controllers gives the same three-line answer: Deployments for stateless apps, DaemonSets for one pod per node, StatefulSets for databases. That answer is correct, and it is also the part that never causes an outage. What causes outages is how each controller behaves when a node is drained, a volume refuses to attach, or a rollout hits a pod that never becomes healthy.

This deep dive answers the question from a production cluster rather than from the documentation. The worked example is our own K3s cluster on Hetzner Cloud: 5 nodes running 50 Deployments, 6 DaemonSets and 12 StatefulSets, with 26 persistent volumes. And the story that ties it together is the three-hour outage of 17 June 2026, when a routine node upgrade took every WordPress site on the cluster offline, and the workload controllers decided both how bad it got and how it was repaired.

One boundary first: this article is about which controller, and how it behaves when things break. Sizing — requests, limits, what the Galera pods and the DaemonSets actually consume — is covered separately in how we cut Kubernetes resource overhead by 50% using only built-in tools.

The 30-second decision tree

Ask three questions, in this order, and stop at the first yes.

  1. Does the pod work with the node itself — its logs, its metrics, its disks, its network, or traffic that arrives at that node? → DaemonSet.
  2. Does each replica need its own identity and its own data that must follow it when it is rescheduled — a database member, a message broker, a quorum peer? → StatefulSet.
  3. Everything else → Deployment. That includes an app that mounts a single persistent volume — but then choose the update strategy on purpose (see the next section).

In one sentence: a Deployment is for interchangeable pods, a StatefulSet is for pods that are not interchangeable, and a DaemonSet is for pods that belong to a node rather than to the application.

Deployments: the default, and the persistent-volume trap

A Deployment keeps a given number of identical pods running through a ReplicaSet. Pods get random names, any one can replace any other, and updates roll out gradually, with rollback built in. That is exactly right for web servers, APIs, workers and most of what a cluster runs, which is why 50 of our 68 workload controllers are Deployments.

Can a Deployment use persistent storage?

Yes. The previous version of this article said Deployments have ephemeral storage, and that was wrong: a Deployment can mount a PersistentVolumeClaim like any other pod. The trap is not the storage. It is the default update strategy combined with a ReadWriteOnce volume.

A Deployment’s default strategy is RollingUpdate with maxSurge: 25% and maxUnavailable: 25%. Kubernetes rounds the surge up and the unavailability down. With replicas: 1, that means a surge of 1 and an unavailability of 0: the new pod starts before the old one stops. If the scheduler places the new pod on a different node, the ReadWriteOnce volume is still attached to the old node, the new pod sits in ContainerCreating with a Multi-Attach error, and the rollout stalls.

On our cluster, 15 Deployments mount a PVC. All 15 are ReadWriteOnce Hetzner volumes, and all 15 run a single replica. They include the WordPress sites, a mail stack whose five components share one 100 GiB volume, Nextcloud, Grafana and a newsletter server. Every one of them is a candidate for the trap.

After the June outage we reviewed all fifteen. The rule the review settled on: Recreate where it is necessary, RollingUpdate where it is feasible.

  • Recreate stops the old pod before starting the new one. Each rollout costs a few seconds of downtime, but it cannot deadlock on a volume. Three of the fifteen use it.
  • RollingUpdate on one node. ReadWriteOnce means one node at a time, not one pod — the per-pod mode is ReadWriteOncePod. Two pods on the same node can both mount the volume. So the WordPress sites keep their zero-downtime rollouts, and are kept on one node instead.

Our web chart does it with a pod-group label and a pod-affinity rule, so the new pod is drawn to the node where its predecessor, and any other pod sharing the volume, already runs:

strategy:
  type: RollingUpdate
template:
  metadata:
    labels:
      pod-group: {{ .Values.podAffinity.group }}
  spec:
    affinity:
      podAffinity:
        preferredDuringSchedulingIgnoredDuringExecution:
          - weight: 100
            podAffinityTerm:
              labelSelector:
                matchExpressions:
                  - key: pod-group
                    operator: In
                    values:
                      - {{ .Values.podAffinity.group }}
              topologyKey: kubernetes.io/hostname

The mail stack and Nextcloud go further and are pinned to a node with a nodeSelector.

If you copy this pattern, know its limit. preferred affinity is a hint, not a guarantee: when the node is under pressure, the scheduler is allowed to place the new pod elsewhere, and then you are back to the Multi-Attach error. required affinity or a nodeSelector gives the guarantee, at the price of a pod that stays Pending if that one node is full. If you have not thought about this trade-off for a particular app, Recreate is the safe default.

Does a DaemonSet really run on every node?

No. A DaemonSet runs one pod on every node it is allowed onto, and the difference is where the interesting decisions are.

Our cluster has a dedicated database node, tainted dedicated=mariadb:NoSchedule so that only the database lands there. This is what the six DaemonSets do with it:

DaemonSet Pods Runs on the DB node? Why
fluent-bit (logs) 5/5 Yes, explicit toleration No logs from the DB node is a blind spot
loki-canary 5/5 Yes, explicit toleration Same reason
prometheus-node-exporter 5/5 Yes, tolerates every NoSchedule taint No metrics from the DB node is a blind spot
hcloud-csi-node 5/5 Yes, tolerates every taint Without the CSI node plugin, no volume attaches on that node
ingress-nginx-controller 4/5 No, by design The DB node serves no web traffic
svclb-ingress-nginx-controller 4/5 No Follows the ingress controller

Same cluster, same taint, and the opposite answer is right for different DaemonSets. The observability agents must reach the tainted node — the database node is exactly where you want logs and metrics when something goes wrong. The ingress controller must not.

The detail that catches people: the DaemonSet controller automatically adds tolerations for the built-in node-condition taints (not-ready, unreachable, memory-pressure, disk-pressure, pid-pressure, unschedulable, and network-unavailable for host-network pods). It does not add tolerations for your own taints. Add a custom taint to a node, and every DaemonSet without a matching toleration silently stops covering it. There is no error, only a DESIRED count one lower than your node count. After tainting a node, kubectl get ds -A is the check.

Why run the ingress controller as a DaemonSet?

Because of how traffic reaches the cluster. A Hetzner load balancer forwards incoming traffic to all worker nodes, so every worker node needs an ingress controller to answer it. The chart supports both modes, and ours is set to DaemonSet:

controller:
  # -- Use a `DaemonSet` or `Deployment`
  kind: DaemonSet

The old version of this article also claimed that DaemonSets have “no built-in rolling update strategy (until Kubernetes 1.7+)”. That is long obsolete: RollingUpdate is the default in the apps/v1 API, and all six of our DaemonSets use it with maxUnavailable: 1.

StatefulSets as they actually run in production

The textbook StatefulSet gives each pod a stable ordinal name (db-0, db-1, db-2), a stable DNS entry through a headless Service, and its own PersistentVolumeClaim from a volumeClaimTemplate, which follows the pod when it is rescheduled. By default it starts pods one at a time in ordinal order (podManagementPolicy: OrderedReady) and rolls updates in reverse order, with a partition to stage them.

Our production database is a three-node MariaDB Galera cluster, and it shows how differently a real one is configured. It is not a hand-written StatefulSet. It is declared as a MariaDB custom resource, and mariadb-operator generates the StatefulSet from it:

apiVersion: k8s.mariadb.com/v1alpha1
kind: MariaDB
spec:
  replicas: 3
  galera:
    enabled: true
  affinity:
    podAntiAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        - labelSelector:
            matchExpressions:
              - key: app.kubernetes.io/instance
                operator: In
                values:
                  - mariadb-galera-new
          topologyKey: kubernetes.io/hostname
  podDisruptionBudget:
    minAvailable: 2

The ordering people choose StatefulSets for is switched off

The generated StatefulSet runs with podManagementPolicy: Parallel and updateStrategy: OnDelete. The ordered, one-at-a-time behaviour that most articles list as the main reason to use a StatefulSet is turned off — and the operator does the ordering itself, rolling the replicas first and the primary last.

The reason is how Galera works. A Galera member that cannot see its peers cannot form a quorum, so “start pod-0, wait until it is Ready, then start pod-1” can stall a full-cluster restart. The operator starts the members together, and applies the order that actually matters for a database — never restart the node taking writes until the others are updated — in its own logic.

The rest of the configuration is where the operational safety lives:

  • Required anti-affinity on the hostname, so no two members share a node, plus a nodeSelector and a toleration that admit them to the database nodes only.
  • A PodDisruptionBudget of minAvailable: 2. A three-node Galera needs two nodes for quorum, so a node drain may take down at most one member at a time.
  • persistentVolumeClaimRetentionPolicy: Retain for both deletion and scale-down. Deleting the StatefulSet or scaling it down never deletes the data. In June this was what made the repair possible.
  • Two volumeClaimTemplates per pod: a 20 GiB data volume and a 100 MiB Galera state volume. The 100 MiB claims are bound at 10 GiB, because Hetzner’s minimum volume size is 10 GB — three times over, for about 300 MB of state.
  • Volumes are tied to a location. A Hetzner volume attaches only to a server in the same location. A PVC created in one datacenter cannot follow its pod to another one; moving it means copying the data into a new PVC.

17 June 2026: what the controllers did during a three-hour outage

The Galera cluster in production at the time was the previous one: a Bitnami chart StatefulSet with the default ordered pod management. This is how a routine maintenance task became a total outage.

  1. A routine worker-node upgrade drained a node.

  2. Galera lost its Primary component. All three members went non-Primary, and every WordPress site on the cluster returned “Error establishing a database connection”. Contributing factors, as the runbook records them: one member was running in a different datacenter, the pods sat at about 99% of a 5 GiB memory limit, and heavy database dumps were running at the same time.

  3. The CSI pods could not be scheduled either. No CSI, no volume attachments — and from that point everything that needed a persistent volume was down, not just the database.

  4. Recovery started by bootstrapping node-0 with an empty cluster address (gcomm://), so it formed a new Primary component on its own.

  5. The StatefulSet’s rolling update was stuck. Under OrderedReady, the controller will not roll past a pod that never becomes Ready; the Kubernetes documentation says that in this state you have to delete the broken pods by hand. The fix was to delete the StatefulSet without deleting its pods or PVCs and recreate it:

    kubectl delete statefulset <name> --cascade=orphan
  6. Three-node HA came back through a staged rollout with partition=1. Node-0 kept serving untouched, nodes 1 and 2 rejoined through a full state transfer (SST), and node-0 rolled last. The memory limit went from 5 GiB to 7 GiB.

Total downtime: about three hours. No data was lost.

What changed afterwards

The Bitnami image in use had been moved to the frozen bitnamilegacy repository, with no further security updates, so the cluster was moved to mariadb-operator. The first cutover attempt — a full dump and reload under a write freeze — took about 60 minutes on a three-node Galera, which is far too long to freeze writes. It was redone the next day as asynchronous GTID replication from the old cluster to the new one, followed by a write freeze of seconds and a switch of the Service selector.

The database’s DNS name never changed. The old mariadb-galera Service now selects the new primary, so not a single tenant configuration file had to be touched. At the same time, the database and several application volumes were moved from the old datacenter to the main one — by copying the data into new PVCs, since a volume cannot move in place.

Lessons that apply to any cluster

  • A node drain is decided by your PodDisruptionBudgets, taints and tolerations — check what each controller will do before upgrading nodes, including the CSI components.
  • Know what ordered pod management does with a pod that never becomes Ready. It waits. The rollout will not repair itself.
  • --cascade=orphan plus a Retain PVC policy lets you rebuild a controller without touching the running pods or the data.
  • Keep the Service name stable. The DNS name is the contract with every client of the database; everything behind it can be replaced.
  • Move data safely: never dump inside the database pod (it will run out of memory), never stream large data through the kube-apiserver (a small control-plane node will fall over), copy pod to pod inside the cluster with a bandwidth cap (pv -L 8m), one database at a time, with set -o pipefail.

Key differences at a glance

Deployment DaemonSet StatefulSet
Manages Interchangeable replicas One pod per eligible node Replicas with stable identity
Pod names Random hash Random hash, one per node Ordinal: name-0, name-1
Placement Scheduler decides Every node it tolerates Scheduler decides; anti-affinity usually required
Storage Any PVC; shared RWO volumes need care Usually node-local paths Own PVC per pod via volumeClaimTemplates
Default update RollingUpdate, 25% surge / 25% unavailable RollingUpdate, maxUnavailable: 1 RollingUpdate in reverse ordinal order, with partition
Ordering None None OrderedReady by default; clustered DB operators often use Parallel
Typical failure Multi-Attach error on a single-replica RWO rollout Silently skips nodes with custom taints Rollout stuck on a pod that never becomes Ready
Use for Web apps, APIs, workers Logs, metrics, CSI, ingress on every node Databases, brokers, quorum-based systems

Final thoughts

Choosing the controller takes thirty seconds with the decision tree above. The work is in the behaviour around it: the update strategy of a Deployment with a volume, the tolerations of every DaemonSet, and the pod management, disruption budget and retention policy of every StatefulSet. On our cluster, each of those was either a cause of the June outage or part of its repair.

Choosing the controller is the easy half; keeping the resulting cluster efficient is the half that shows up on the invoice. Our production K3s audit — how we cut Kubernetes resource overhead by 50% using only built-in tools — shows what over-requested StatefulSets and cluster-wide DaemonSets actually cost once they are measured. That measurement is a one-off cloud cost audit; teams that would rather not own the work in-house afterwards hand it to a Fractional DevOps retainer, where Kubernetes management is part of the Scale plan.

Common questions about Kubernetes workload controllers

What is the difference between a Deployment and a StatefulSet?

A Deployment runs interchangeable pods: they get random names, any of them can be replaced by any other, and they usually share or skip storage. A StatefulSet runs pods that are not interchangeable: each gets a stable ordinal name (db-0, db-1, db-2), a stable DNS entry through a headless Service, and its own PersistentVolumeClaim that follows it when it is rescheduled. Use a StatefulSet when a replica's identity and data matter; use a Deployment for everything else.

When should I use a DaemonSet instead of a Deployment?

When the pod belongs to the node rather than to the application: log shippers, metrics exporters, CSI node plugins, network agents, and ingress controllers that must answer on every node a load balancer sends traffic to. A DaemonSet runs one pod on every node it is allowed onto, and adds or removes pods as nodes join or leave. Note that a custom taint keeps a DaemonSet off a node unless the DaemonSet explicitly tolerates it.

Can a Kubernetes Deployment use a PersistentVolumeClaim?

Yes. The storage is not the problem; the update strategy is. With one replica and a ReadWriteOnce volume, the default RollingUpdate starts the new pod before the old one stops, and if the new pod lands on another node the volume cannot attach there (a Multi-Attach error), so the rollout stalls. Either use the Recreate strategy or keep the new pod on the same node as the old one.

Why would a StatefulSet use Parallel pod management instead of OrderedReady?

Because some clustered databases cannot become healthy one member at a time. A Galera node that cannot see its peers cannot form a quorum, so waiting for pod-0 to become Ready before starting pod-1 can stall a full restart. Operators such as mariadb-operator set podManagementPolicy to Parallel and updateStrategy to OnDelete, and then enforce the safe order themselves: replicas first, the primary last.

More articles

10 min

Fractional DevOps vs Full-Time DevOps: A Complete Comparison Guide

You need serious DevOps expertise. The real question is: how do you get it without overpaying — or understaffing? For many growing companies, this decision comes down to two options: hire a full-time DevOps engineer on payroll, or engage a Fractional DevOps partner on a monthly retainer. On the surface, both solve the same problem.

Guides

4 min

Understanding fully-managed web hosting: what it means and when you need it

In today’s digital landscape, having a reliable online presence is crucial for businesses of all sizes. However, maintaining a website involves much more than just uploading files to a server. This is where fully-managed web hosting comes into play, offering a comprehensive solution that takes the technical burden off your shoulders. What Is Fully-Managed Web […]

Guides

Want this run for you?

A free 30-minute discovery call. Engineer to engineer.

Book a free call