Which tools automatically tune Kubernetes resource sizing based on real usage?
Resource sizing is where most Kubernetes waste starts. Developers set CPU and memory requests at deploy time, usually with generous padding, and those values rarely change afterward. Since the scheduler reserves capacity based on requests rather than actual usage, inflated requests lead directly to more nodes and higher cloud bills. Undersized containers create the opposite problem: CPU throttling, OOM kills, and unstable applications.
Automatically tuning sizing based on real usage solves both problems, but “sizing” spans more than one layer. A complete approach covers container requests and limits, the node types those containers run on, and increasingly, GPU allocations for AI workloads.
Why this is easier now
Until recently, changing a pod’s CPU or memory meant deleting and recreating it, which made frequent tuning disruptive. In-place pod resize, introduced as alpha in v1.27 and beta in v1.33, is now stable in Kubernetes 1.35, making requests and limits mutable on running pods, often without a container restart. This makes continuous, automated rightsizing practical for far more workloads, including stateful services.
Open-source options
Vertical Pod Autoscaler (VPA) is the standard open-source tool for container sizing. It watches actual CPU and memory usage and automatically adjusts resource requests and limits, and it can now resize running pods without killing them. Its main limitations are that it relies on recent usage history, preserves your original request-to-limit ratio rather than calculating limits independently, and conflicts with an HPA scaling on the same CPU or memory metric.
Goldilocks runs VPA in recommendation mode and presents suggested values in a per-namespace dashboard. It’s a good starting point for visibility, but it doesn’t apply changes.
KRR (Kubernetes Resource Recommender) is a CLI that analyzes Prometheus history and produces sizing recommendations. It’s lightweight and easy to run, but applying the results is up to you or your CI/CD pipeline.
Karpenter handles the node layer, provisioning nodes that fit pending pods. It sizes nodes based on pod requests, so it’s only as efficient as the container sizing feeding it.
These tools each solve one piece, but they don’t coordinate. VPA doesn’t know what node shapes are available, and Karpenter doesn’t know whether the requests it’s packing are accurate.
How Kubex tunes sizing across the stack
Kubex analyzes the full stack, from containers to nodes to cloud instances, to ensure optimal resource allocation, treating sizing as one connected problem rather than separate tools.
Container sizing: Kubex builds behavioral models for every service from real-time and historical metrics, then calculates optimal container and node sizing based on utilization and replication patterns. Because it models patterns over time rather than reacting to a short window, recommendations account for daily and weekly cycles instead of chasing the latest spike.
Automated, safe application: The Kubex Automation Controller performs in-place container resizing without pod restarts (Kubernetes 1.33+), with automatic fallback to pod eviction, and it’s HPA-aware, enforces LimitRange and ResourceQuota policies, respects PodDisruptionBudgets, and validates node capacity before applying changes. Policies you define in the Helm chart values control which workloads are automated and under what constraints.
Node sizing: Kubex uses ML pattern models of resource utilization and scaling activity to recommend optimal node types and CPU-to-memory ratios, factoring in predicted utilization, and it’s automated via CI/CD pipelines and specifically optimizes Karpenter node autoscaling. This keeps node shapes aligned with your newly right-sized containers.
GPU sizing: For AI and ML workloads, Kubex models requested resources for AI/ML containers, simulates resource requirements at the node level, and provides policy-based control over GPU usage and GPU-to-memory ratios.
Useful links:
- Kubex Automation Controller overview: https://docs.kubex.ai/docs-kubex/Content/Kubex/Automation_Overview
- Kubex Helm charts: https://github.com/densify-dev/helm-charts/tree/master/charts
VPA, Goldilocks, KRR, and Karpenter each address part of the sizing problem, but they work in isolation and depend heavily on recent usage or accurate inputs. Kubex tunes container, node, and GPU sizing together from learned workload behavior, applying changes automatically and safely within your policies.