Debugging Performance Issues in a VM
A cheat-sheet for narrowing down why a VM feels slow, before escalating to Zenon.
Check disk I/O accounting first
Disk I/O is the most common cause of a VM feeling slow — especially on HDD storage, which is IOPS-limited. Check what's actually generating the load before assuming it's a Zenon-side problem:
iostat -x 1— per-device throughput, IOPS, and latency over time.iotop— per-process disk I/O, sorted by usage, to find the heaviest process.- For a specific systemd service, cgroup v2 gives you per-service I/O accounting — but unlike memory and tasks accounting,
IOAccountingis off by default (io.statreads all zeroes until it's turned on). That's because it has real kernel-side overhead, so enable it only while you're actively debugging, not as a standing default:
systemctl edit <service>
# add under [Service]:
# IOAccounting=yes
systemctl restart <service>
Enabling it for one service also enables it for the slice it's in and that slice's parents, so sibling services in the same slice start getting accounted too — be aware of that blast radius before enabling it on a shared slice. Read the numbers:
systemd-cgtop # live, per-cgroup CPU/memory/IO
cat /sys/fs/cgroup/system.slice/<service>.service/io.stat
Then turn it back off once you're done:
systemctl revert <service>
systemctl restart <service>
(DefaultIOAccounting=yes in /etc/systemd/system.conf enables it everywhere instead of per-service — same overhead and cascading caveat applies, so only use it for a time-boxed investigation, not permanently. Note this isn't specific to the manual check: io.stat doesn't exist at all until the io controller is enabled, so any continuous collector reading it — cAdvisor included — pays the same IOAccounting cost per service it watches. For continuous, low-overhead disk visibility, prefer node_exporter's device-level stats below instead of leaving per-service IOAccounting on.)
If this shows you're hitting the IOPS block on HDD storage, consider whether that data belongs on SSD instead — see disk layout.
Find the slow component with tracing
If the bottleneck isn't obviously disk I/O, don't guess from system graphs — use tracing to see where time is actually spent inside your application. Platon's observability stack includes distributed tracing (Tempo): instrument your app with OpenTelemetry and inspect the trace for a slow request to see which span — a database query, an outbound call, a disk operation — is actually responsible.
Ship system metrics: node-exporter to Mimir
For system-level visibility (CPU, memory, disk, network) over time, we recommend running node_exporter on your VM and shipping its metrics to Mimir, following the "External Services" path in Platon's observability docs.
cAdvisor optionally gives finer-grained detail than node_exporter alone. It's not just for containers — it can watch any systemd cgroup directly, giving you ongoing per-service CPU and memory metrics in Grafana for free (cgroup v2 always tracks these, accounting or not), plus per-container metrics if you're running containerized workloads inside the VM. Per-service I/O is the exception: it needs that same IOAccounting enabled per service as above, with the same overhead and cascading caveat — so don't turn it on broadly just to feed cAdvisor continuously.
Summary
| Question | Tool |
|---|---|
| Which process/service is driving disk I/O? | iostat, iotop, or cgroup io.stat for a specific service (enable IOAccounting temporarily first) |
| Which part of my application is slow? | Tracing (Tempo) via OpenTelemetry |
| What's my VM's overall resource usage over time, cheaply? | node_exporter → Mimir |
| Per-service/per-container CPU/memory over time? | cAdvisor (optional) — also works on systemd cgroups, not just containers |