ZSoftly
Talk to our team
Back to blog

The Restart That Hid a Libvirt Bug

A seven-minute VM create, two wrong turns, a self-healing timer I built and then deleted, and the package update that fixed it.

Ditah KumbongFounder and Chief Technology Officer
9 min read
ZCP branded engineering cover reading The restart that hid a bug, on a deep blue grid

On a Sunday morning a test VM on our Montreal region took 6 minutes and 44 seconds to create. It was an ordinary machine: 8 vCPUs, 8 GB of RAM, a 20 GB root disk and AlmaLinux 9. Eight seconds after CloudStack reported it created, the platform marked it Stopped, even though it was running.

Creating a VM should take seconds. This post covers how I found the cause, the two wrong turns I took on the way, and why the fix was a package update we already had access to.

Splitting the seven minutes

The CloudStack management log records every command it sends to a host and every answer it gets back. Splitting the create into steps gave me this:

StepTime
Copy the virtual router’s root disk46 s
Boot the virtual router67 s
Configure the router5 s
Create the VM’s root disk142 s
Start the VM144 s

The first VM in a new network also pays for the network’s virtual router. This accounts for about two minutes. The two VM steps were the problem. The disk itself is a Ceph RBD clone, which the host log showed finishing in under a second. The agent spent the other 141 seconds waiting.

The waiting was inside the CloudStack agent on the host. Each time the agent looked up a storage pool, 10 to 50 seconds passed between “Trying to fetch storage pool” and the next line. Running the same lookup by hand with virsh took about 50 milliseconds.

Wrong turn one: a story built on one sample

The agent on this host had been up for 40 days and sat at 100% CPU. A thread dump showed one worker thread busy in a VM stats call. The first explanation was tidy: one stuck thread, a 35-second baseline on healthy hosts, restart the agent and move on.

I did not accept it. The 35-second baseline came from a handful of hand-picked deploys. Nobody had measured the median across all of them.

So I pulled five months of management logs and measured every VM start, host by host. Four hosts show the pattern:

PeriodAMD host AAMD host BIntel host AIntel host B
June to August10-24 s4-12 sn/an/a
Late September46 s72 s4 s3 s
October 9 to 1164-68 s64-126 s4-14 s4-15 s

The AMD hosts had slowed by five to ten times since early September. The Intel hosts stayed at about 4 seconds. A second thread dump also corrected the first: all five agent worker threads had used almost the same CPU, about 374,000 CPU-seconds each. An Intel host with the same 40-day uptime had used about 8,400 each. No single thread was stuck. Every call was expensive.

I ruled out four other causes with measurements:

  • NFS: the primary storage mount answered statfs in 7 ms with no retransmits.
  • Stats traffic: about one storage stats call per minute per management server, nothing unusual.
  • TLS: an Intel host on TLS libvirt was fast, so the transport was not the cause.
  • Package drift: versions matched across hosts, apart from one Ceph client.

Then I ran perf record against the agent process. About 90% of its CPU was inside libglib, in GLib main loop code around g_source_ref. The agent loads the libvirt client library, which uses a GLib event loop. The profile fits a leak. The loop collects sources and never frees them, so every call walks a longer list than the last.

A restart proved the theory. I restarted the agent on the worst host. Running VMs were not affected, and the host reported Up again 32 seconds later. Storage lookups dropped from 8.4 seconds to 0.01 seconds, and a test VM pinned to that host was created in 12 seconds.

Wrong turn two: a self-healing timer

I now knew the cost grew with agent uptime and call volume. The busy AMD hosts went from healthy to slow in two to three weeks, and the quiet hosts took much longer.

I did not want another alert waking someone up. Our rule: once we understand a fault, a job checks for it, fixes it, and tells a human only when the fix fails. We already run one of these for outbound port 25 filtering on every hypervisor, so I copied the design:

  • A daily timer on each host measures the storage lookup gap in the agent log.
  • If the average is over one second, it restarts the agent, unless the agent is in the middle of a start, copy or migration.
  • It sends a low-priority notification when it acts and a warning only if the restart does not help.

I built it, sent it through code, context and writing review, and tested it in dry-run mode on real hosts. It flagged a slow host in Ottawa and held off on a host restarted the same morning. It was ready to deploy.

One question stopped the deploy

Before deploying, I asked one question: does CloudStack already handle this, so I am not reinventing the wheel?

CloudStack 4.22 does have something close. Point the agent property agent.health.check.script.path at a script and turn on the global setting enable.kvm.host.auto.enable.disable. When the script fails, CloudStack stops placing new VMs on the host and re-enables it when the script passes again. This keeps new VMs off a sick host. It does not make the host healthy. It also removes capacity from hosts whose VMs cannot move elsewhere.

Asking whether this was a known bug got me further. It was. In June 2024 libvirt merged rpc: avoid leak of GSource in use for interrupting main loop. Its commit message describes our symptom: “an ever growing list of GSources attached to the main context, which will gradually slow down execution of the loop, as several operations are O(N) for the number of attached GSource objects.”

An earlier GSource leak in libvirt slowed OpenStack’s compute agent the same way. The patch that fixed it reports queries growing from about one second to over a minute within an hour. The agent ended up using 100% of a CPU thread.

The fix shipped in libvirt 10.5.0. Ubuntu 24.04 ships libvirt 10.0.0, with the fix backported into package 10.0.0-2ubuntu8.19. Our hosts still ran the earlier build.

My timer would have restarted agents every few weeks, forever, to work around a bug a routine package update already fixed.

The fix

I rolled the update across our KVM hosts in both regions, one host at a time:

  1. Install only the libvirt packages, keeping our existing configuration files.
  2. Restart the CloudStack agent, so it loads the fixed library.
  3. Confirm the configuration is unchanged, every VM still runs with the same domain ID, and the host reports Up and secured.

The script stops at the first host to fail a check. None failed. Each host was back in about 30 seconds. No host rebooted and no VM restarted.

I deleted the self-healing timer from the repository before it ever reached production.

The result

Agent logs from July 1 to October 11 show lookup time building up between restarts. Through July and August it climbed, then fell to near zero each time a patch window restarted an agent. Nothing restarted the agents after September 1, and the worst host climbed for six weeks to 7 seconds.

Agent storage lookup time, daily average

Seconds, from the CloudStack agent log. July 1 to October 11, 2026, UTC.

Shaded: Agents restarted (Sep 1, the start of the last climb), Libvirt updated (Oct 11, 14:38 to 15:53 UTC).

Show the data
Time (UTC)AMD host BAMD host AOttawa hostIntel host A
Jul 11.07 s1.58 s0.003 sn/a
Jul 20.33 s0.43 s0.004 sn/a
Jul 30.022 s0.019 s0.004 sn/a
Jul 40.059 s0.046 s0.007 sn/a
Jul 50.24 s0.072 s0.005 sn/a
Jul 60.72 s0.13 s0.005 sn/a
Jul 71.12 s0.30 s0.005 sn/a
Jul 81.35 s0.56 s0.005 sn/a
Jul 91.48 s0.71 s0.006 sn/a
Jul 101.42 s0.21 s0.006 sn/a
Jul 110.46 s0.012 s0.006 sn/a
Jul 120.029 s0.055 s0.007 sn/a
Jul 130.069 s0.072 s0.015 sn/a
Jul 140.23 s0.51 s0.011 sn/a
Jul 150.39 s1.13 s0.012 sn/a
Jul 160.36 s1.59 s0.013 sn/a
Jul 170.41 s2.05 s0.013 sn/a
Jul 180.48 s1.88 s0.013 sn/a
Jul 190.55 s2.42 s0.014 sn/a
Jul 200.63 s2.93 s0.014 sn/a
Jul 210.71 s3.31 s0.014 sn/a
Jul 220.85 s3.54 s0.015 sn/a
Jul 230.25 s3.67 s0.006 sn/a
Jul 240.009 s0.95 s0.005 sn/a
Jul 250.088 s0.058 s0.017 sn/a
Jul 260.11 s0.069 s0.019 sn/a
Jul 270.25 s0.12 s0.019 sn/a
Jul 280.082 s0.16 s0.007 sn/a
Jul 290.033 s0.056 s0.005 sn/a
Jul 300.12 s0.032 s0.011 sn/a
Jul 310.059 s0.037 s0.009 sn/a
Aug 10.077 s0.067 s0.019 sn/a
Aug 20.22 s0.17 s0.036 sn/a
Aug 30.36 s0.26 s0.037 sn/a
Aug 40.60 s0.34 s0.037 sn/a
Aug 50.82 s0.42 s0.039 sn/a
Aug 60.99 s0.51 s0.040 sn/a
Aug 71.24 s0.60 s0.041 sn/a
Aug 81.49 s0.73 s0.041 sn/a
Aug 91.70 s0.85 s0.043 sn/a
Aug 101.85 s1.49 s0.042 sn/a
Aug 110.48 s0.36 s0.048 sn/a
Aug 120.023 s0.023 s0.019 sn/a
Aug 130.051 s0.045 s0.010 sn/a
Aug 140.069 s0.087 s0.009 sn/a
Aug 150.026 s0.021 s0.019 s0.024 s
Aug 160.037 s0.026 s0.027 s0.026 s
Aug 170.061 s0.031 s0.037 s0.032 s
Aug 180.20 s0.065 s0.17 s0.062 s
Aug 190.38 s0.065 s0.22 s0.051 s
Aug 200.46 s0.021 s0.069 s0.015 s
Aug 210.11 s0.005 s0.004 s0.006 s
Aug 220.023 s0.012 s0.007 s0.009 s
Aug 230.008 s0.011 s0.006 s0.010 s
Aug 240.032 s0.023 s0.011 s0.022 s
Aug 250.010 s0.014 s0.010 s0.013 s
Aug 260.021 s0.010 s0.008 s0.014 s
Aug 270.014 s0.006 s0.006 s0.009 s
Aug 280.037 s0.016 s0.017 s0.019 s
Aug 290.35 s0.044 s0.015 s0.038 s
Aug 300.50 s0.044 s0.004 s0.039 s
Aug 310.62 s0.046 s0.004 s0.040 s
Sep 10.19 s0.021 s0.012 s0.023 s
Sep 20.17 s0.038 s0.057 s0.036 s
Sep 30.33 s0.035 s0.068 s0.038 s
Sep 40.51 s0.038 s0.070 s0.039 s
Sep 50.64 s0.041 s0.075 s0.040 s
Sep 60.73 s0.045 s0.079 s0.043 s
Sep 70.82 s0.059 s0.083 s0.045 s
Sep 80.92 s0.10 s0.089 s0.046 s
Sep 91.00 s0.059 s0.099 s0.047 s
Sep 101.05 s0.066 s0.11 s0.048 s
Sep 111.13 s0.068 s0.12 s0.051 s
Sep 121.22 s0.073 s0.13 s0.052 s
Sep 131.30 s0.078 s0.14 s0.053 s
Sep 141.36 s0.11 s0.15 s0.054 s
Sep 151.45 s0.095 s0.16 s0.055 s
Sep 161.51 s0.10 s0.17 s0.056 s
Sep 171.54 s0.11 s0.19 s0.057 s
Sep 181.61 s0.13 s0.20 s0.057 s
Sep 191.64 s0.14 s0.21 s0.058 s
Sep 201.96 s0.18 s0.24 s0.060 s
Sep 212.14 s0.27 s0.30 s0.061 s
Sep 222.33 s0.40 s0.35 s0.063 s
Sep 232.44 s0.54 s0.40 s0.066 s
Sep 242.73 s0.67 s0.47 s0.067 s
Sep 252.92 s0.72 s0.52 s0.068 s
Sep 263.09 s0.78 s0.58 s0.069 s
Sep 273.32 s0.88 s0.65 s0.071 s
Sep 283.79 s0.91 s0.73 s0.073 s
Sep 294.94 s0.98 s0.86 s0.075 s
Sep 305.68 s1.07 s0.93 s0.077 s
Oct 15.78 s1.18 s1.04 s0.077 s
Oct 25.78 s1.19 s1.10 s0.078 s
Oct 35.86 s1.26 s1.18 s0.079 s
Oct 46.11 s1.34 s1.23 s0.081 s
Oct 58.16 s1.40 s1.33 s0.084 s
Oct 67.99 s1.50 s1.37 s0.083 s
Oct 76.58 s1.57 s1.48 s0.084 s
Oct 86.56 s1.70 s1.47 s0.087 s
Oct 97.05 s1.89 s1.51 s0.089 s
Oct 107.35 s2.04 s1.55 s0.090 s
Oct 116.38 s1.55 s1.39 s0.080 s
MeasureBeforeAfter
Agent storage lookup, worst host8.3 to 9.2 s, hourly average0.004 to 0.017 s
Agent worker CPU, worst hostabout 1,900 CPU-seconds/hourabout 8 CPU-seconds/hour
VM start on the AMD hosts, median64 to 126 snot measured separately
Full VM create on an existing network, test6 min 44 s, incl. about 2 min router11 to 19 s per test host

The test creates ran at 16:31 UTC on October 11. I pinned each VM to one host, timed it, then deleted it. The original VM took 6 minutes 44 seconds. About two minutes of that went to the new network’s router. On the VM alone, the time dropped from about 290 seconds to under 20.

These numbers come from the first hours after the update. A freshly restarted agent is fast with or without the patch, so the real proof is the line staying flat over the next week. I am watching agent CPU on the two worst-affected hosts and will update this post then.

What I took from it

  • Measure the whole history. One deploy produced a convincing but wrong story. Five months of data showed which hosts were slow and when the slowdown started.
  • A restart fix is a clue. It tells you the problem builds up over time. Find out what is building up before you schedule more restarts.
  • Ask whether it is a known bug before writing a workaround. Building the self-healing timer took longer than finding the upstream fix did.
  • Patch visibility matters. The fix was already in the Ubuntu archive. Our patch reporting did not surface it, so nobody saw the update waiting.
  • Push back on tidy answers. Both wrong turns began as neat answers, and both ended when I asked a plain question before acting.

Related articles

Engineering

We Moved Our Own DNS onto Our Own Platform

How we registered .ca nameserver host records, recovered from an unrequested delegation change, and moved zsoftly.ca onto the nameservers we sell to customers.