On a Sunday morning a test VM on our Montreal region took 6 minutes and 44 seconds to create. It was an ordinary machine: 8 vCPUs, 8 GB of RAM, a 20 GB root disk and AlmaLinux 9. Eight seconds after CloudStack reported it created, the platform marked it Stopped, even though it was running.
Creating a VM should take seconds. This post covers how I found the cause, the two wrong turns I took on the way, and why the fix was a package update we already had access to.
Splitting the seven minutes
The CloudStack management log records every command it sends to a host and every answer it gets back. Splitting the create into steps gave me this:
| Step | Time |
|---|---|
| Copy the virtual router’s root disk | 46 s |
| Boot the virtual router | 67 s |
| Configure the router | 5 s |
| Create the VM’s root disk | 142 s |
| Start the VM | 144 s |
The first VM in a new network also pays for the network’s virtual router. This accounts for about two minutes. The two VM steps were the problem. The disk itself is a Ceph RBD clone, which the host log showed finishing in under a second. The agent spent the other 141 seconds waiting.
The waiting was inside the CloudStack agent on the host. Each time the agent looked up a storage
pool, 10 to 50 seconds passed between “Trying to fetch storage pool” and the next line. Running the
same lookup by hand with virsh took about 50 milliseconds.
Wrong turn one: a story built on one sample
The agent on this host had been up for 40 days and sat at 100% CPU. A thread dump showed one worker thread busy in a VM stats call. The first explanation was tidy: one stuck thread, a 35-second baseline on healthy hosts, restart the agent and move on.
I did not accept it. The 35-second baseline came from a handful of hand-picked deploys. Nobody had measured the median across all of them.
So I pulled five months of management logs and measured every VM start, host by host. Four hosts show the pattern:
| Period | AMD host A | AMD host B | Intel host A | Intel host B |
|---|---|---|---|---|
| June to August | 10-24 s | 4-12 s | n/a | n/a |
| Late September | 46 s | 72 s | 4 s | 3 s |
| October 9 to 11 | 64-68 s | 64-126 s | 4-14 s | 4-15 s |
The AMD hosts had slowed by five to ten times since early September. The Intel hosts stayed at about 4 seconds. A second thread dump also corrected the first: all five agent worker threads had used almost the same CPU, about 374,000 CPU-seconds each. An Intel host with the same 40-day uptime had used about 8,400 each. No single thread was stuck. Every call was expensive.
I ruled out four other causes with measurements:
- NFS: the primary storage mount answered
statfsin 7 ms with no retransmits. - Stats traffic: about one storage stats call per minute per management server, nothing unusual.
- TLS: an Intel host on TLS libvirt was fast, so the transport was not the cause.
- Package drift: versions matched across hosts, apart from one Ceph client.
Then I ran perf record against the agent process. About 90% of its CPU was inside libglib, in
GLib main loop code around g_source_ref. The agent loads the libvirt client library, which uses a
GLib event loop. The profile fits a leak. The loop collects sources and never frees them, so every
call walks a longer list than the last.
A restart proved the theory. I restarted the agent on the worst host. Running VMs were not affected, and the host reported Up again 32 seconds later. Storage lookups dropped from 8.4 seconds to 0.01 seconds, and a test VM pinned to that host was created in 12 seconds.
Wrong turn two: a self-healing timer
I now knew the cost grew with agent uptime and call volume. The busy AMD hosts went from healthy to slow in two to three weeks, and the quiet hosts took much longer.
I did not want another alert waking someone up. Our rule: once we understand a fault, a job checks for it, fixes it, and tells a human only when the fix fails. We already run one of these for outbound port 25 filtering on every hypervisor, so I copied the design:
- A daily timer on each host measures the storage lookup gap in the agent log.
- If the average is over one second, it restarts the agent, unless the agent is in the middle of a start, copy or migration.
- It sends a low-priority notification when it acts and a warning only if the restart does not help.
I built it, sent it through code, context and writing review, and tested it in dry-run mode on real hosts. It flagged a slow host in Ottawa and held off on a host restarted the same morning. It was ready to deploy.
One question stopped the deploy
Before deploying, I asked one question: does CloudStack already handle this, so I am not reinventing the wheel?
CloudStack 4.22 does have something close. Point the agent property
agent.health.check.script.path at a script and turn on the global setting
enable.kvm.host.auto.enable.disable. When the script fails, CloudStack stops placing new VMs on
the host and re-enables it when the script passes again. This keeps new VMs off a sick host. It
does not make the host healthy. It also removes capacity from hosts whose VMs cannot move elsewhere.
Asking whether this was a known bug got me further. It was. In June 2024 libvirt merged rpc: avoid leak of GSource in use for interrupting main loop. Its commit message describes our symptom: “an ever growing list of GSources attached to the main context, which will gradually slow down execution of the loop, as several operations are O(N) for the number of attached GSource objects.”
An earlier GSource leak in libvirt slowed OpenStack’s compute agent the same way. The patch that fixed it reports queries growing from about one second to over a minute within an hour. The agent ended up using 100% of a CPU thread.
The fix shipped in libvirt 10.5.0. Ubuntu 24.04 ships libvirt 10.0.0, with the fix backported into
package 10.0.0-2ubuntu8.19. Our hosts still ran the earlier build.
My timer would have restarted agents every few weeks, forever, to work around a bug a routine package update already fixed.
The fix
I rolled the update across our KVM hosts in both regions, one host at a time:
- Install only the libvirt packages, keeping our existing configuration files.
- Restart the CloudStack agent, so it loads the fixed library.
- Confirm the configuration is unchanged, every VM still runs with the same domain ID, and the host reports Up and secured.
The script stops at the first host to fail a check. None failed. Each host was back in about 30 seconds. No host rebooted and no VM restarted.
I deleted the self-healing timer from the repository before it ever reached production.
The result
Agent logs from July 1 to October 11 show lookup time building up between restarts. Through July and August it climbed, then fell to near zero each time a patch window restarted an agent. Nothing restarted the agents after September 1, and the worst host climbed for six weeks to 7 seconds.
Seconds, from the CloudStack agent log. July 1 to October 11, 2026, UTC.
Shaded: Agents restarted (Sep 1, the start of the last climb), Libvirt updated (Oct 11, 14:38 to 15:53 UTC).
Show the data
| Time (UTC) | AMD host B | AMD host A | Ottawa host | Intel host A |
|---|---|---|---|---|
| Jul 1 | 1.07 s | 1.58 s | 0.003 s | n/a |
| Jul 2 | 0.33 s | 0.43 s | 0.004 s | n/a |
| Jul 3 | 0.022 s | 0.019 s | 0.004 s | n/a |
| Jul 4 | 0.059 s | 0.046 s | 0.007 s | n/a |
| Jul 5 | 0.24 s | 0.072 s | 0.005 s | n/a |
| Jul 6 | 0.72 s | 0.13 s | 0.005 s | n/a |
| Jul 7 | 1.12 s | 0.30 s | 0.005 s | n/a |
| Jul 8 | 1.35 s | 0.56 s | 0.005 s | n/a |
| Jul 9 | 1.48 s | 0.71 s | 0.006 s | n/a |
| Jul 10 | 1.42 s | 0.21 s | 0.006 s | n/a |
| Jul 11 | 0.46 s | 0.012 s | 0.006 s | n/a |
| Jul 12 | 0.029 s | 0.055 s | 0.007 s | n/a |
| Jul 13 | 0.069 s | 0.072 s | 0.015 s | n/a |
| Jul 14 | 0.23 s | 0.51 s | 0.011 s | n/a |
| Jul 15 | 0.39 s | 1.13 s | 0.012 s | n/a |
| Jul 16 | 0.36 s | 1.59 s | 0.013 s | n/a |
| Jul 17 | 0.41 s | 2.05 s | 0.013 s | n/a |
| Jul 18 | 0.48 s | 1.88 s | 0.013 s | n/a |
| Jul 19 | 0.55 s | 2.42 s | 0.014 s | n/a |
| Jul 20 | 0.63 s | 2.93 s | 0.014 s | n/a |
| Jul 21 | 0.71 s | 3.31 s | 0.014 s | n/a |
| Jul 22 | 0.85 s | 3.54 s | 0.015 s | n/a |
| Jul 23 | 0.25 s | 3.67 s | 0.006 s | n/a |
| Jul 24 | 0.009 s | 0.95 s | 0.005 s | n/a |
| Jul 25 | 0.088 s | 0.058 s | 0.017 s | n/a |
| Jul 26 | 0.11 s | 0.069 s | 0.019 s | n/a |
| Jul 27 | 0.25 s | 0.12 s | 0.019 s | n/a |
| Jul 28 | 0.082 s | 0.16 s | 0.007 s | n/a |
| Jul 29 | 0.033 s | 0.056 s | 0.005 s | n/a |
| Jul 30 | 0.12 s | 0.032 s | 0.011 s | n/a |
| Jul 31 | 0.059 s | 0.037 s | 0.009 s | n/a |
| Aug 1 | 0.077 s | 0.067 s | 0.019 s | n/a |
| Aug 2 | 0.22 s | 0.17 s | 0.036 s | n/a |
| Aug 3 | 0.36 s | 0.26 s | 0.037 s | n/a |
| Aug 4 | 0.60 s | 0.34 s | 0.037 s | n/a |
| Aug 5 | 0.82 s | 0.42 s | 0.039 s | n/a |
| Aug 6 | 0.99 s | 0.51 s | 0.040 s | n/a |
| Aug 7 | 1.24 s | 0.60 s | 0.041 s | n/a |
| Aug 8 | 1.49 s | 0.73 s | 0.041 s | n/a |
| Aug 9 | 1.70 s | 0.85 s | 0.043 s | n/a |
| Aug 10 | 1.85 s | 1.49 s | 0.042 s | n/a |
| Aug 11 | 0.48 s | 0.36 s | 0.048 s | n/a |
| Aug 12 | 0.023 s | 0.023 s | 0.019 s | n/a |
| Aug 13 | 0.051 s | 0.045 s | 0.010 s | n/a |
| Aug 14 | 0.069 s | 0.087 s | 0.009 s | n/a |
| Aug 15 | 0.026 s | 0.021 s | 0.019 s | 0.024 s |
| Aug 16 | 0.037 s | 0.026 s | 0.027 s | 0.026 s |
| Aug 17 | 0.061 s | 0.031 s | 0.037 s | 0.032 s |
| Aug 18 | 0.20 s | 0.065 s | 0.17 s | 0.062 s |
| Aug 19 | 0.38 s | 0.065 s | 0.22 s | 0.051 s |
| Aug 20 | 0.46 s | 0.021 s | 0.069 s | 0.015 s |
| Aug 21 | 0.11 s | 0.005 s | 0.004 s | 0.006 s |
| Aug 22 | 0.023 s | 0.012 s | 0.007 s | 0.009 s |
| Aug 23 | 0.008 s | 0.011 s | 0.006 s | 0.010 s |
| Aug 24 | 0.032 s | 0.023 s | 0.011 s | 0.022 s |
| Aug 25 | 0.010 s | 0.014 s | 0.010 s | 0.013 s |
| Aug 26 | 0.021 s | 0.010 s | 0.008 s | 0.014 s |
| Aug 27 | 0.014 s | 0.006 s | 0.006 s | 0.009 s |
| Aug 28 | 0.037 s | 0.016 s | 0.017 s | 0.019 s |
| Aug 29 | 0.35 s | 0.044 s | 0.015 s | 0.038 s |
| Aug 30 | 0.50 s | 0.044 s | 0.004 s | 0.039 s |
| Aug 31 | 0.62 s | 0.046 s | 0.004 s | 0.040 s |
| Sep 1 | 0.19 s | 0.021 s | 0.012 s | 0.023 s |
| Sep 2 | 0.17 s | 0.038 s | 0.057 s | 0.036 s |
| Sep 3 | 0.33 s | 0.035 s | 0.068 s | 0.038 s |
| Sep 4 | 0.51 s | 0.038 s | 0.070 s | 0.039 s |
| Sep 5 | 0.64 s | 0.041 s | 0.075 s | 0.040 s |
| Sep 6 | 0.73 s | 0.045 s | 0.079 s | 0.043 s |
| Sep 7 | 0.82 s | 0.059 s | 0.083 s | 0.045 s |
| Sep 8 | 0.92 s | 0.10 s | 0.089 s | 0.046 s |
| Sep 9 | 1.00 s | 0.059 s | 0.099 s | 0.047 s |
| Sep 10 | 1.05 s | 0.066 s | 0.11 s | 0.048 s |
| Sep 11 | 1.13 s | 0.068 s | 0.12 s | 0.051 s |
| Sep 12 | 1.22 s | 0.073 s | 0.13 s | 0.052 s |
| Sep 13 | 1.30 s | 0.078 s | 0.14 s | 0.053 s |
| Sep 14 | 1.36 s | 0.11 s | 0.15 s | 0.054 s |
| Sep 15 | 1.45 s | 0.095 s | 0.16 s | 0.055 s |
| Sep 16 | 1.51 s | 0.10 s | 0.17 s | 0.056 s |
| Sep 17 | 1.54 s | 0.11 s | 0.19 s | 0.057 s |
| Sep 18 | 1.61 s | 0.13 s | 0.20 s | 0.057 s |
| Sep 19 | 1.64 s | 0.14 s | 0.21 s | 0.058 s |
| Sep 20 | 1.96 s | 0.18 s | 0.24 s | 0.060 s |
| Sep 21 | 2.14 s | 0.27 s | 0.30 s | 0.061 s |
| Sep 22 | 2.33 s | 0.40 s | 0.35 s | 0.063 s |
| Sep 23 | 2.44 s | 0.54 s | 0.40 s | 0.066 s |
| Sep 24 | 2.73 s | 0.67 s | 0.47 s | 0.067 s |
| Sep 25 | 2.92 s | 0.72 s | 0.52 s | 0.068 s |
| Sep 26 | 3.09 s | 0.78 s | 0.58 s | 0.069 s |
| Sep 27 | 3.32 s | 0.88 s | 0.65 s | 0.071 s |
| Sep 28 | 3.79 s | 0.91 s | 0.73 s | 0.073 s |
| Sep 29 | 4.94 s | 0.98 s | 0.86 s | 0.075 s |
| Sep 30 | 5.68 s | 1.07 s | 0.93 s | 0.077 s |
| Oct 1 | 5.78 s | 1.18 s | 1.04 s | 0.077 s |
| Oct 2 | 5.78 s | 1.19 s | 1.10 s | 0.078 s |
| Oct 3 | 5.86 s | 1.26 s | 1.18 s | 0.079 s |
| Oct 4 | 6.11 s | 1.34 s | 1.23 s | 0.081 s |
| Oct 5 | 8.16 s | 1.40 s | 1.33 s | 0.084 s |
| Oct 6 | 7.99 s | 1.50 s | 1.37 s | 0.083 s |
| Oct 7 | 6.58 s | 1.57 s | 1.48 s | 0.084 s |
| Oct 8 | 6.56 s | 1.70 s | 1.47 s | 0.087 s |
| Oct 9 | 7.05 s | 1.89 s | 1.51 s | 0.089 s |
| Oct 10 | 7.35 s | 2.04 s | 1.55 s | 0.090 s |
| Oct 11 | 6.38 s | 1.55 s | 1.39 s | 0.080 s |
| Measure | Before | After |
|---|---|---|
| Agent storage lookup, worst host | 8.3 to 9.2 s, hourly average | 0.004 to 0.017 s |
| Agent worker CPU, worst host | about 1,900 CPU-seconds/hour | about 8 CPU-seconds/hour |
| VM start on the AMD hosts, median | 64 to 126 s | not measured separately |
| Full VM create on an existing network, test | 6 min 44 s, incl. about 2 min router | 11 to 19 s per test host |
The test creates ran at 16:31 UTC on October 11. I pinned each VM to one host, timed it, then deleted it. The original VM took 6 minutes 44 seconds. About two minutes of that went to the new network’s router. On the VM alone, the time dropped from about 290 seconds to under 20.
These numbers come from the first hours after the update. A freshly restarted agent is fast with or without the patch, so the real proof is the line staying flat over the next week. I am watching agent CPU on the two worst-affected hosts and will update this post then.
What I took from it
- Measure the whole history. One deploy produced a convincing but wrong story. Five months of data showed which hosts were slow and when the slowdown started.
- A restart fix is a clue. It tells you the problem builds up over time. Find out what is building up before you schedule more restarts.
- Ask whether it is a known bug before writing a workaround. Building the self-healing timer took longer than finding the upstream fix did.
- Patch visibility matters. The fix was already in the Ubuntu archive. Our patch reporting did not surface it, so nobody saw the update waiting.
- Push back on tidy answers. Both wrong turns began as neat answers, and both ended when I asked a plain question before acting.
