No restart, no disconnect: live-migrating Kubernetes pods with Paguro

Contents

What this is about

If you run Kubernetes, you know the moment: a node has to go, maybe because a spot instance gets reclaimed or a kernel patch is due. Kubernetes then kills the pod and starts it somewhere else. For a game server with fifty players, that means everyone gets kicked.

Paguro changes that. It live-migrates running pods to another node, with their memory, volumes and even open TCP and UDP connections. The application pauses for a moment and then simply carries on.

In plain words

Picture a packed restaurant that has to move to another building. Kubernetes sends every guest out the door, reopens somewhere else and hopes they come back. Paguro moves the restaurant with its guests inside: everyone keeps their table, the food is still there, the conversation goes on. The lights flicker for the blink of an eye. That’s all anyone notices.

0.6 sfreeze of a game server with its IP kept on Cilium
0disconnects: Minecraft with eight bots on five clusters
55 → 0.6 sfreeze optimization for moving the pod IP
0.1 srestore of 1 GiB of memory with lazy pages

Demos first

The video is loaded from YouTube (Google) only when you press play. Privacy · Watch on YouTube

Amazon EKS, eight bots keep playing: 0.4 s freeze, 0 disconnects, the same TCP connection as before.

The video is loaded from YouTube (Google) only when you press play. Privacy · Watch on YouTube

An arena shooter on Amazon EKS: 0.3 s freeze, and the match just keeps going because the server process moves with its memory.

That’s all it takes:

# migrate a single pod
kubectl paguro migrate <pod> -n <namespace> --to <node> --wait

# or drain the node: migratable pods are migrated instead of restarted
kubectl drain <node> --ignore-daemonsets

How Paguro works

The trick: the pod on the target node is an ordinary pod. A 2.7 MB wrapper around runc turns its runc create into runc restore. kubelet and containerd don’t notice a thing, and nothing in the node’s runtime stack changes.

How a live migration works: pre-copy while the application keeps running, then a short freeze with the final dump, commit and restore on the target node
How a live migration works. The application only notices the orange freeze.
  1. Pre-copy: the application keeps running while its memory is copied in rounds, with only the changed pages from the second round on.
  2. Freeze and commit: the container is paused and the last bit transferred. Until the commit, every error ends in a rollback.
  3. Restore: CRIU starts the process on the target, and missing pages follow via userfaultfd.

The five hardest problems

The idea is simple, the execution isn’t. These are the five problems live migration usually fails on.

1. A single packet can destroy everything

If a TCP retransmission hits the target pod before its socket is restored, the kernel answers with a RST, and a single RST ends the connection for good. Paguro’s RST shield therefore puts an nftables drop rule into every new network namespace before the CNI even runs. Even a race of a few milliseconds while the namespaces are being created is covered: 0 failures in 500 attempts.

2. Moving the IP: from 55 to 0.6 seconds

Bar chart: the freeze keeping the IP on Cilium drops over eight optimization steps from 55.2 seconds to 0.55 to 0.73 seconds
Freeze per migration keeping the IP on Cilium, step by step.

Behind every bar is a mechanism in Cilium or kubelet that I only found in the source code. One example: kubelet retries a failed sandbox only once per second. My first idea was to time the freeze to exactly that rhythm. It sounded logical, but in an A/B test it was slower: 1,731 instead of 1,510 ms. So it was dropped. What worked was a different approach: Paguro binds the replacement pod to the node at exactly the right moment, because a pod’s very first sync doesn’t wait for that rhythm. That’s how the whole project went: every idea gets measured, and whatever doesn’t improve the numbers doesn’t make it in.

Why every millisecond counts: TCP retransmits lost packets after roughly 0.2, 0.6, 1.4 and 3.0 s. A 1.12 s freeze therefore cost the client 1.45 s, a 1.60 s freeze already 3.16 s.

3. Memory that changes faster than you can copy it

Pre-copy transfers memory while the application keeps writing. For a pod with 8 GiB, the rounds shrank from 8,194 MiB to 4 MiB, and the freeze came in at 1.59 s for 8 GiB. On restore, CRIU was the bottleneck: 1.87 s for 1 GiB. With a 13-line CRIU patch and lazy pages, it’s now 104–146 ms.

4. CPUs that aren’t equal

A Python process started on a newer CPU and migrated to an older one crashed right after the restore: at startup, glibc had picked instructions the older CPU doesn’t have. Paguro therefore gives every migratable pod a CPU baseline that glibc, Go, the JVM and friends are limited to from the start. Now the same process migrates in 326 ms, in both directions.

5. Phantom: new address, same connection

Normally, a pod keeps its IP address when Paguro moves it, just like keeping your phone number when you move house. That has two catches: handing the address over takes time, right inside the freeze. And many networks, such as Flannel, Antrea or the AWS VPC CNI, can’t move addresses between nodes at all.

Phantom solves both. The pod gets a new address at its new home, and Paguro makes sure the other side never notices. It works like mail forwarding, only in real time and in both directions: whoever keeps writing to the old address reaches the pod at the new one, and its replies still carry the old sender. The old address belongs to no one anymore; it only lives on as a phantom. Hence the name.

Phantom: a client keeps sending to the old IP 10.0.1.5, which no longer belongs to anyone. eBPF rewrites every packet to the new IP 10.0.2.9 of the migrated pod, and replies come back with 10.0.1.5 as the sender.
Phantom: running connections keep talking to the old address, and eBPF invisibly redirects them to the new one.

Under the hood, small eBPF programs in the pod’s network namespace do the work. They rewrite every packet of a running connection between old and new IP, exactly per connection, while new connections go straight to the new address. Because the new pod is fully ready before the freeze, the freeze drops to 0.3–0.9 s. And live migration now works on networks that can’t move IPs at all.

Four CNIs, one result

Every CNI had its own trap, from Cilium’s eBPF load balancer to Antrea’s Open vSwitch. After the fixes, not a single connection is lost, whether the client uses the pod IP, the service or a NodePort from outside:

CNILost connectionsFreeze
Flannel0338–635 ms
Cilium0339–491 ms
Antrea0354–609 ms
Calico0356–904 ms

The ultimate test: game servers

Game servers have no failover: players send 30 inputs per second over UDP, and the server keeps their sessions in memory. UDP brought four traps of its own, and Paguro defuses all of them. In a test with five back-to-back migrations, inputs were lost only during the 0.6 s freeze, there was no session reset, and all 625 newly joining players got in.

The real-world test with Minecraft Java (TCP) and Minecraft Bedrock (UDP), eight bots each, playing and chatting. On all five clusters, not a single player lost the connection:

ClusterFreeze Java (TCP)Freeze Bedrock (UDP)
Cilium (lab)1.1 / 1.75 s0.77 / 0.85 s
Calico (lab)1.9–2.7 s1.45–1.57 s
Flannel (lab)1.3 / 1.5 s0.48 / 0.51 s
Antrea (lab)1.2 / 4.0 s0.53 / 0.53 s
Amazon EKS0.61 / 0.66 s0.44 / 0.47 s

Storage, cloud and chaos

Volumes without data loss. The pod’s data moves along too. Paguro copies it ahead of time, so the freeze only carries what changed at the very end. In a stress test, a 25 GiB volume being written to every 100 ms moved over: not a single entry was missing.

Cloud, Karpenter and spot. When a node is drained, for example by Karpenter or a spot interruption, Paguro migrates the pods automatically instead of restarting them:

DisruptionResult
drift (new node first, then drain)651 ms freeze, 0 disconnects
drift from on-demand to spot549 ms freeze, 0 disconnects
NodeClaim deleted531 ms freeze, 0 disconnects
spot interruption warning0.53–0.70 s freeze, done before AWS terminates the instance

Chaos tests. A test script kills controllers and agents in the middle of a migration. Every run ended either in a clean rollback or a successful migration, never with a duplicate or frozen pod. Even a helm upgrade in the middle of a Minecraft migration went through without a disconnect.

What Paguro can do today

  • Live migration of memory, files and volumes
  • IP preservation on Cilium and Calico, Phantom for every network that can’t move IPs (Flannel, Antrea, the AWS VPC CNI and more)
  • Automatic migration on kubectl drain, Karpenter, the Cluster Autoscaler and spot interruptions
  • Agones integration for game servers
  • Encrypted transfer between nodes, installation with a Helm chart

Key learnings

  1. Measure, don’t assume. I thought the 1 GbE cable was the bottleneck. A direct 10G link showed that even two VMs on the same host only reach 5–6 Gbit/s. The bottleneck was virtualization, not the cable.
  2. Optimize for what the user feels. What counts isn’t the freeze, but how long the client waits. That’s why Paguro always measures at the client too.
  3. Make silent failures loud. Since Go 1.24, every Go server opens MPTCP sockets that CRIU can’t migrate. Paguro now detects that before the freeze and says clearly what to do.

You run game servers, spot fleets or jobs that can’t afford a restart? Let’s talk: david@picillo.de