Skip to main content
Gå til innhold

Debugging Performance Issues in a VM

A cheat-sheet for narrowing down why a VM feels slow, before escalating to Zenon.

Check disk I/O accounting first​

Disk I/O is the most common cause of a VM feeling slow — especially on HDD storage, which is IOPS-limited. Check what's actually generating the load before assuming it's a Zenon-side problem:

  • iostat -x 1 — per-device throughput, IOPS, and latency over time.
  • iotop — per-process disk I/O, sorted by usage, to find the heaviest process.
  • For a specific systemd service, cgroup v2 gives you per-service I/O accounting — but unlike memory and tasks accounting, IOAccounting is off by default (io.stat reads all zeroes until it's turned on). That's because it has real kernel-side overhead, so enable it only while you're actively debugging, not as a standing default:
systemctl edit <service>
# add under [Service]:
# IOAccounting=yes
systemctl restart <service>

Enabling it for one service also enables it for the slice it's in and that slice's parents, so sibling services in the same slice start getting accounted too — be aware of that blast radius before enabling it on a shared slice. Read the numbers:

systemd-cgtop # live, per-cgroup CPU/memory/IO
cat /sys/fs/cgroup/system.slice/<service>.service/io.stat

Then turn it back off once you're done:

systemctl revert <service>
systemctl restart <service>

(DefaultIOAccounting=yes in /etc/systemd/system.conf enables it everywhere instead of per-service — same overhead and cascading caveat applies, so only use it for a time-boxed investigation, not permanently. Note this isn't specific to the manual check: io.stat doesn't exist at all until the io controller is enabled, so any continuous collector reading it — cAdvisor included — pays the same IOAccounting cost per service it watches. For continuous, low-overhead disk visibility, prefer node_exporter's device-level stats below instead of leaving per-service IOAccounting on.)

If this shows you're hitting the IOPS block on HDD storage, consider whether that data belongs on SSD instead — see disk layout.

Find the slow component with tracing​

If the bottleneck isn't obviously disk I/O, don't guess from system graphs — use tracing to see where time is actually spent inside your application. Platon's observability stack includes distributed tracing (Tempo): instrument your app with OpenTelemetry and inspect the trace for a slow request to see which span — a database query, an outbound call, a disk operation — is actually responsible.

Ship system metrics: node-exporter to Mimir​

For system-level visibility (CPU, memory, disk, network) over time, we recommend running node_exporter on your VM and shipping its metrics to Mimir, following the "External Services" path in Platon's observability docs.

cAdvisor optionally gives finer-grained detail than node_exporter alone. It's not just for containers — it can watch any systemd cgroup directly, giving you ongoing per-service CPU and memory metrics in Grafana for free (cgroup v2 always tracks these, accounting or not), plus per-container metrics if you're running containerized workloads inside the VM. Per-service I/O is the exception: it needs that same IOAccounting enabled per service as above, with the same overhead and cascading caveat — so don't turn it on broadly just to feed cAdvisor continuously.

Summary​

QuestionTool
Which process/service is driving disk I/O?iostat, iotop, or cgroup io.stat for a specific service (enable IOAccounting temporarily first)
Which part of my application is slow?Tracing (Tempo) via OpenTelemetry
What's my VM's overall resource usage over time, cheaply?node_exporter → Mimir
Per-service/per-container CPU/memory over time?cAdvisor (optional) — also works on systemd cgroups, not just containers