Container Resource Limits and cgroups on Bare Metal Servers
Containers are useful because they package an application and its dependencies without requiring a separate full operating system for every workload. On a bare metal server, however, the container runtime shares the host’s kernel and physical CPU, memory, storage, and network paths. A clear resource policy is therefore essential. Container resource limits turn physical capacity into enforceable boundaries that keep one workload from consuming the headroom needed by other services.
Linux control groups, commonly called cgroups, provide the kernel mechanism behind those boundaries. They organise processes into hierarchies and account for resource usage so an operator or orchestration platform can apply policies to a container, pod, service, or host slice. Understanding cgroups is especially important on dedicated hardware, where the operator controls the kernel and can tune the relationship between the host, the container runtime, and the workload.
What cgroups control on a bare metal server
A cgroup is not a virtual machine and it does not emulate hardware. It is a kernel-level control boundary around a set of processes. Depending on the controller and the Linux version, a cgroup can limit or prioritise CPU time, memory, process IDs, block I/O, device access, and CPU or NUMA placement. Container runtimes create and manage these groups when they start workloads, while systemd and Kubernetes may create additional parent groups for services and pods.
- Most current distributions use cgroup v2, which provides a unified hierarchy and consistent interfaces such as cpu.max, memory.max, memory.high, pids.max, and io.max. Older systems may still expose cgroup v1, where controllers are mounted separately. The exact file names and runtime behaviour differ, so an operations team should confirm the host’s cgroup version before documenting limits or writing automation.
- The resource path usually has several layers: the host kernel, a system or service manager, the container runtime, and—when Kubernetes is used—the kubelet and pod hierarchy. A limit at one layer can interact with a limit at another. A container that appears to have available CPU may still be constrained by a pod quota or a system slice, while a node that has free memory may be protecting reserved capacity for the operating system and critical daemons.
For a useful comparison of the isolation model, see Dataplugs’ article on Containerization vs. Virtualization: Isolation and Performance on Bare Metal. It helps clarify why containers share the kernel while virtual machines introduce a separate guest operating system, and why resource boundaries must be designed explicitly on dedicated hardware.
The practical objective is not to make every container as small as possible. It is to give each workload a predictable operating envelope, preserve enough host headroom for recovery and platform services, and make exhaustion visible before it becomes an outage.
CPU and memory limits are the first control layer
CPU limits can be expressed as a quota over a period, a relative weight, or a set of CPUs on which a workload may run. A quota caps sustained CPU consumption; a weight influences how CPU time is shared during contention; and cpuset placement can isolate latency-sensitive work to selected cores. These controls solve different problems. Applying a quota without considering burst behaviour can increase queueing, while using weights alone does not prevent a busy container from consuming all available capacity when the node is otherwise idle.
Memory limits are harder because memory cannot always be reclaimed gracefully. In cgroup v2, memory.high is a pressure threshold that causes reclaim and throttling, while memory.max is a hard ceiling. memory.min or memory.low can protect important services when the host is under pressure. The correct values depend on the application’s working set, cache behaviour, startup peak, and the memory used by sidecars or helper processes—not only the average resident set.
- When a workload reaches its hard memory boundary and reclaim cannot recover enough space, the kernel may invoke the cgroup out-of-memory path and terminate a process. This is safer than allowing the entire host to become unstable, but it is not a substitute for application-level capacity planning. Export memory working-set, reclaim, throttling, and OOM events, then connect them to deployment and restart policy so an operator can distinguish a bad limit from a real memory leak or traffic spike.
- The pids controller limits the number of processes and threads a cgroup may create. It is a useful guard against fork bombs, runaway worker creation, and misconfigured process pools. Set it with knowledge of the runtime, application workers, language threads, health checks, and short-lived jobs. A PID limit that is too low can look like an application failure even though CPU and memory metrics are normal.
Block I/O controls such as io.max and io.weight can protect latency-sensitive services from a batch job or an image-heavy deployment. Device-level limits are most useful when the storage layout is understood: an NVMe volume, RAID device, and filesystem may expose different points at which pressure appears. If a workload needs direct access to a GPU or another device, combine device permissions with the appropriate runtime and security policy instead of treating device access as a generic resource limit.
Design limits for bare metal workloads
Start with measurements rather than a round number. Record CPU utilisation, throttled time, memory working set, page faults, I/O latency, throughput, process count, and startup peaks under normal and burst traffic. Measure a new workload during warm-up and cache fill as well as steady state. These observations provide a defensible basis for requests, limits, and host sizing, and they make later tuning less dependent on guesswork.
- Reserve host headroom deliberately. The operating system, container runtime, logging, monitoring, storage agents, security controls, and orchestration components all need resources even when application containers are busy. On bare metal, the operator should also account for kernel memory, interrupt processing, filesystem cache, and recovery tasks. Treating all installed RAM or all logical CPUs as assignable application capacity can turn a single container surge into a node-level incident.
- Separate a guaranteed floor from a maximum. A request or reservation communicates the capacity a workload should normally receive; a limit defines how far it may grow. On Kubernetes, requests influence scheduling while limits are enforced by the runtime and kernel. On a standalone Docker host, the equivalent policy may be encoded in compose files, systemd units, or deployment scripts. Naming these concepts separately makes capacity plans clearer and avoids setting every value to the same number.
Hardware topology also matters. On multi-socket or NUMA systems, a container that can move freely across nodes may see different memory latency and cache locality from one run to the next. Pinning should be used only when measurement shows a benefit, because overly rigid CPU or NUMA placement can reduce schedulability. Document the relationship between cpuset settings, memory placement, interrupt affinity, and the workload’s latency target. If Kubernetes is part of the design, Dataplugs’ guide to dedicated server models recommended for running Kubernetes provides useful context for matching CPU topology, memory capacity, NVMe storage, and network behaviour to cluster roles.
Monitor and test the resource boundaries
Monitoring should show both usage and enforcement. Track CPU usage beside throttled seconds, memory usage beside memory events and OOM kills, and I/O throughput beside latency and queue depth. For Kubernetes, combine node, pod, and container metrics; for a standalone host, include systemd slices and runtime statistics. A dashboard that shows only utilisation can miss the fact that a workload is already being throttled.
- Test contention intentionally. Run a controlled CPU burner, memory allocator, process creator, or I/O job in a non-production environment and confirm that the intended workload stays within its service objective. Verify the failure signal, restart behaviour, log message, alert, and recovery time. Also test the inverse case: remove the artificial pressure and confirm that the workload can use its permitted burst capacity again instead of remaining permanently throttled.
- Keep resource policy beside the application release. Version the cgroup settings, node labels, runtime flags, and monitoring rules together, and record why a limit changed. A review should be able to answer which workload was measured, what failure mode was being prevented, and which alert would detect an incorrect value. This is particularly important when teams migrate between cgroup v1 and v2 or change the container runtime.
Use the results for capacity planning. A limit that is repeatedly reached may indicate a need to tune the application, add nodes, move a workload to a larger bare metal profile, or change its scheduling class. Increasing the limit without checking contention only moves the bottleneck. Review saturation, throttling, OOM events, queueing, and recovery time together before changing hardware or policy.
Conclusion
Container resource limits and cgroups turn bare metal capacity into a set of controlled, measurable operating envelopes. The strongest design separates CPU, memory, process, I/O, and device concerns; distinguishes floors from ceilings; reserves host headroom; validates the generated hierarchy; and tests the behaviour that should occur when a boundary is reached.
Dataplugs dedicated servers give teams the hardware isolation and administrative access needed to tune these controls for containerised applications. For the operational tooling around deployment and observability, see Dataplugs’ article on Prometheus and Grafana monitoring for server health, then keep the resource policy versioned and review it as workloads change.
For more information about dedicated server infrastructure, visit the Dataplugs website or contact sales@dataplugs.com.
