Woke Up to 100% Disk Usage on K3s? Here’s How to Fix /var/lib/rancher Overflow in 15 Seconds!
Imagine this nightmare scenario: You wake up on a Monday morning, grab your coffee, open up your dashboard, and suddenly see red alerts everywhere. Your pods are stuck in Evicted status, services are going down one by one, and df -h shows the terrifying truth:
/dev/sda1 100G 100G 0 100% /var/lib/rancher
100% Disk Usage! Your K3s cluster has completely frozen because /var/lib/rancher ran out of space—even though your actual application data is just a few megabytes.
What just happened? Did a hacker infiltrate your cluster?
Nope. You just fell right into one of the most infamous traps in lightweight Kubernetes administration: Uncapped Container Log Bloat.
🔍 The Root Cause: Why Is /var/lib/rancher Bloating?
By default, K3s writes all container stdio logs to /var/log/pods, which symlinks directly under /var/lib/rancher/k3s/agent/containerd/.
If your application emits high-frequency logs (like access logs, debug traces, or unhandled exceptions), K3s will keep appending to these log files endlessly without rotating them. Before you know it, a single pod can spawn 50GB+ of raw text logs, triggering Kubernetes DiskPressure and evicting all your running workloads!
🚨 Step 1: Emergency Cleanup (Get Your Cluster Back Alive NOW)
Don't panic! You don't need to rebuild your server. Execute this one-liner emergency command to reclaim dozens of gigabytes in under 5 seconds by purging dangling containerd images and log caches:
# Clean unused containerd images and build cache immediately
sudo k3s crictl rmi --prune
# Quick-purge massive log files larger than 500MB without breaking file descriptors
sudo find /var/log/pods -name "*.log" -size +500M -exec truncate -s 0 {} \;
💡 Pro-Tip: Using
truncate -s 0is much safer thanrm -rfbecause it empties the file instantly without crashing active container process logging handlers!
Check df -h again—boom! You just reclaimed 50% to 80% of your disk space. Your pods will automatically recover from Evicted state.
🛡️ Step 2: The Permanent Fix (Never Suffer From Disk Overflow Again)
Emergency cleanup is just a band-aid. To permanently stop K3s from eating your disk, you must enforce Log Rotation directly at the Kubelet level.
Edit your K3s systemd service configuration file at /etc/systemd/system/k3s.service (or pass these flags into /etc/rancher/k3s/config.yaml):
# Open K3s systemd unit configuration
sudo nano /etc/systemd/system/k3s.service
Add these two critical --kubelet-arg parameters to your ExecStart execution string:
ExecStart=/usr/local/bin/k3s \
server \
--kubelet-arg="container-max-size=10Mi" \
--kubelet-arg="container-max-files=3" \
What does this do?
container-max-size=10Mi: Caps individual log files at 10 Megabytes. Once a log reaches 10MB, K3s automatically rotates it.container-max-files=3: Retains a maximum of 3 rotated files per container. Your logging overhead per pod will never exceed 30MB, guaranteed!
Finally, reload systemd and restart K3s to apply the fix:
sudo systemctl daemon-reload
sudo systemctl restart k3s
🎓 Want to Master Production-Ready K3s Architecture?
Debugging disk issues is just the tip of the iceberg when running lightweight Kubernetes in production.
If you want to build High Availability (HA) K3s clusters, automate S3 etcd backups, configure Traefik v3 Ingress, and self-host your own Private AI Cloud (with GPU pass-through, Ollama, and Qdrant), check out my hands-on Masterclass:
👉
20+ Practical Modules
Production Security Checklist & Hardening
Complete Troubleshooting Script Library
Real-world Infrastructure Architecture
💬 Over to You!
Have you ever experienced a 100% disk overflow on your Linux or Kubernetes servers? How did you recover? Drop a comment below and let's discuss!
#k3s
#kubernetes