Load Balancing Setup: Architecture, Configs, and Validation

September 28, 2026 ARPHost Uncategorized

A basic round-robin load balancing setup needs only a few configuration lines, but production validation depends on more than seeing healthy backends. HAProxy active checks default to a 2-second interval, while NGINX Plus can keep clients on the same upstream with sticky cookie, including an expires=1h policy when session affinity is required.

The popular advice is to choose round robin, click Deploy, and move on. That approach leaves the most important question unanswered: how do you prove traffic is distributing correctly, health checks detect real application failure, and failover works when a node misbehaves under customer load? In multi-tenant bare metal and Proxmox environments, the configuration is usually the easy part. The verification loop is where outages are either caught early or discovered by users.

Table of Contents

Baseline Routing and Architecture Decisions

Start with a deliberately simple configuration. HAProxy provides Layer 7 HTTP routing and health checks in one place:

global
    log /dev/log local0
    log /dev/log local1 notice
    stats socket /run/haproxy/admin.sock mode 660 level admin

defaults
    log global
    mode http
    option httplog
    timeout connect 5s
    timeout client 30s
    timeout server 30s

frontend web_front
    bind :80
    default_backend web_pool

backend web_pool
    balance roundrobin
    option httpchk GET /health
    server web01 10.0.10.11:8080 check
    server web02 10.0.10.12:8080 check

The equivalent NGINX upstream is shorter:

http {
    upstream web_pool {
        server 10.0.10.11:8080;
        server 10.0.10.12:8080;
    }

    server {
        listen 80;
        server_name app.example;

        location / {
            proxy_pass 
            proxy_set_header Host $host;
            proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
            proxy_set_header X-Forwarded-Proto $scheme;
        }
    }
}

Validate syntax before reloading:

haproxy -c -f /etc/haproxy/haproxy.cfg
systemctl reload haproxy

nginx -t
systemctl reload nginx

Round robin is a reasonable baseline when backend capacity is similar and requests have comparable cost. It isn't a promise of equal work. One request may be a lightweight asset fetch, while another may trigger a database query or a report generation job. Least connections can react better to uneven request duration, but it also needs accurate connection accounting and can behave poorly when long-lived connections dominate.

Practical rule: Start with round robin, then change the algorithm only after access logs, response times, and backend saturation show that the baseline is wrong.

Layer 4 or Layer 7

Layer 4 balancing forwards TCP or UDP flows without interpreting HTTP. It suits raw database connections, game servers, DNS, and other protocols where the proxy shouldn't rewrite headers or terminate TLS. Layer 7 balancing understands HTTP and HTTPS, which enables host and path routing, header manipulation, application health checks, and TLS offload.

RequirementLayer 4Layer 7
Raw TCP or UDP serviceStrong fitUsually unnecessary
HTTP path routingNot availableAvailable
TLS terminationPass-throughAvailable
Application response checksLimitedAvailable
Protocol transparencyStrongLower
Configuration complexityLowerHigher

Use HAProxy mode tcp for a transparent TCP service:

frontend postgres_front
    mode tcp
    bind :5432
    default_backend postgres_pool

backend postgres_pool
    mode tcp
    balance leastconn
    timeout connect 5s
    timeout server 30s
    server db01 10.0.20.11:5432 check
    server db02 10.0.20.12:5432 check

The choice belongs to the workload, not the product label. A web application commonly benefits from a reverse proxy, as described in ARPHost's guide to what a reverse proxy does. A stateful database or game service may be better served by a TCP proxy that doesn't pretend to understand application semantics. Dedicated hardware is also a legitimate option for workloads that require high connection rates, specialized inspection, or a separate failure domain. The initial configuration can be simple, but the architecture must match the protocol.

A person coding on a laptop with network equipment on a desk for baseline routing configuration.

Tuning Health Checks and Failover Mechanics

A backend can be reachable at the IP layer and still be unusable to customers. The process may have exhausted worker threads, the application may reject new sessions, or storage stalls may cause requests to time out. Health checks should test the failure that matters to the service.

HAProxy active TCP checks require the check parameter on each server line. A valid probe completes when the backend responds with SYN/ACK, and the default check interval is 2 seconds, unless inter changes it (HAProxy's health check documentation).

backend app_pool
    mode tcp
    option tcp-check
    server app01 10.0.30.11:9000 check inter 2s fall 3 rise 2
    server app02 10.0.30.12:9000 check inter 2s fall 3 rise 2

An active check is predictable and easy to reason about. It can detect a closed port quickly, but it may miss a backend that accepts probes while failing real requests. Increasing probe frequency adds monitoring traffic and can create unnecessary state changes when the service is merely slow.

Passive checks observe customer connections instead. HAProxy can combine observe layer4, error-limit, and on-error so a backend is removed after repeated connection failures. One documented example removes a server after three failed connections and reinstates it after two successful connections (HAProxy's passive health check guidance).

backend app_pool
    mode tcp
    balance roundrobin
    option tcp-check
    server app01 10.0.30.11:9000 check inter 2s observe layer4 error-limit 3 on-error mark-down
    server app02 10.0.30.12:9000 check inter 2s observe layer4 error-limit 3 on-error mark-down
Check typeWhat it catchesMain risk
Active TCPClosed or unreachable service portMay miss application-level failure
Active HTTPInvalid status or unhealthy responseProbe load and endpoint design
Passive Layer 4Failures seen during real trafficNeeds sufficient live traffic
Combined checksSynthetic and customer-facing failuresMore tuning and more state

For HTTP, use a health endpoint that checks the application state without performing destructive work:

backend web_pool
    option httpchk
    http-check send meth GET uri /health ver HTTP/1.1 hdr Host app.example
    http-check expect status 200
    server web01 10.0.10.11:8080 check
    server web02 10.0.10.12:8080 check

Resilience also requires more than one endpoint. Azure recommends at least two back-end endpoints and carefully selected probe intervals and thresholds (Azure Load Balancer guidance). Keep spare capacity at each location. Google SRE warns that load balancing combined with autoscaling, without reserve capacity, can create a feedback loop where a failure overloads the remaining sites. A subset of 20 to 100 backend tasks is common in datacenter load balancing, although the correct size depends on service behavior (Google SRE load balancing guidance).

Fallback pools should be explicit, not accidental:

backend web_pool
    balance roundrobin
    option httpchk GET /health
    server web01 10.0.10.11:8080 check
    server web02 10.0.10.12:8080 check
    server maintenance 10.0.10.99:8080 backup check

In production, I see false failovers when operators set aggressive checks against overloaded tenants. A probe succeeds, the next probe times out, and traffic flaps between nodes. Tune fall and rise around observed service behavior, then verify the state transition deliberately.

A systems engineer manages server cables while a desktop monitor displays real-time health metrics and server performance data.

Implementing SSL Offload and Session Persistence

TLS termination at the load balancer simplifies certificate management and removes cryptographic work from application nodes, but it creates a security boundary that must be monitored. Keep the private key readable only by the service account or root, forward the original protocol, and make the backend trust boundary explicit.

HAProxy can terminate HTTPS with a PEM bundle containing the certificate and private key:

frontend https_front
    bind :443 ssl crt /etc/haproxy/certs/app.example.pem alpn h2,http/1.1
    http-request set-header X-Forwarded-Proto https
    default_backend web_pool

backend web_pool
    mode http
    balance roundrobin
    server web01 10.0.10.11:8080 check
    server web02 10.0.10.12:8080 check

NGINX uses separate certificate and key directives:

server {
    listen 443 ssl http2;
    server_name app.example;

    ssl_certificate /etc/nginx/tls/app.example.fullchain.pem;
    ssl_certificate_key /etc/nginx/tls/app.example.key;
    add_header Strict-Transport-Security "max-age=31536000" always;

    location / {
        proxy_pass 
        proxy_set_header Host $host;
        proxy_set_header X-Forwarded-Proto https;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    }
}

Use a current TLS policy appropriate for the installed HAProxy or NGINX release, and test the certificate chain before production traffic arrives. ARPHost's SSL certificate configuration guide is useful for the certificate handling side, but the load balancer still needs correct hostname coverage and backend forwarding headers.

Persistence is a trade-off

Stateful applications need affinity when sessions aren't stored centrally. NGINX Plus supports cookie persistence with sticky cookie; it inserts a cookie in the first upstream response and routes later requests to the same upstream. Options include expires=1h, domain, and path, and route-based affinity is available with the route parameter (NGINX Plus HTTP load balancing documentation).

upstream web_pool {
    zone web_pool 64k;
    sticky cookie srv_id expires=1h domain=app.example path=/;
    server 10.0.10.11:8080;
    server 10.0.10.12:8080;
}

HAProxy can use a cookie to preserve affinity:

backend web_pool
    balance roundrobin
    cookie SERVER insert indirect nocache
    server web01 10.0.10.11:8080 check cookie web01
    server web02 10.0.10.12:8080 check cookie web02

Persistence improves continuity but can defeat even distribution. A busy client may remain pinned to one backend while other nodes sit lightly loaded. Prefer shared session storage when the application supports it. If affinity is unavoidable, watch per-node connections and response time, then use weights or connection limits to prevent one backend from becoming the de facto primary.

A close-up shot of a technician plugging a blue Ethernet cable into a server rack networking port.

The Verification Loop and Observability

A successful reload proves only that the configuration parses. It does not prove that the hostname reaches the intended listener, the certificate matches, requests reach every healthy backend, or failover behaves correctly. Verification should begin immediately after deployment and continue under representative load, using infrastructure monitoring best practices.

Start with local checks on the load balancer:

ss -lntp | grep -E ':(80|443)s'
curl -skI 
curl -sk 
journalctl -u haproxy --since "10 minutes ago"
journalctl -u nginx --since "10 minutes ago"

For HAProxy, expose a restricted statistics socket or stats page. The admin socket provides direct state information:

echo "show stat" | socat stdio /run/haproxy/admin.sock
echo "show servers state" | socat stdio /run/haproxy/admin.sock

Check for backend servers marked UP, request counters increasing on each expected node, and error counters consistent with normal traffic. A healthy endpoint with no counter movement can indicate a routing, DNS, cookie, or listener problem. Test from the client network as well as locally, because split DNS and edge routing can produce different results.

NGINX's stub_status exposes connection and request counters:

server {
    listen 127.0.0.1:8081;
    location /nginx_status {
        stub_status;
        allow 127.0.0.1;
        deny all;
    }
}
curl 

The response shows active, accepted, and handled connections, along with requests. Pair these values with backend access logs. Add an upstream identifier to the log format where possible, then inspect request distribution:

awk '{print $NF}' /var/log/nginx/access.log | sort | uniq -c

The field varies with the configured log format, so inspect the active definition first:

nginx -T

Test failure, not just health

A controlled failure test should remove one backend from service, send sample requests, and confirm that another node handles them. For maintenance testing, stop the application on a non-production node:

systemctl stop myapp
curl -sk -o /dev/null -w '%{http_code}n' 
echo "show stat" | socat stdio /run/haproxy/admin.sock
systemctl start myapp

Confirm that the failed node leaves rotation, fallback behavior is correct, and recovery does not trigger a traffic storm. Cloudflare's load balancing quickstart recommends reviewing analytics after sample requests, then checking pool health, DNS, and SSL settings before production traffic. Its analytics can filter by pool name, which helps isolate post-deployment routing errors.

A load balancer can be enabled and still misroute traffic because the hostname, certificate coverage, pool order, or health check is wrong.

Record the expected backend sequence, failure transition, and recovery transition as an operational test. That record proves the control plane changed traffic as designed, rather than merely showing a green dashboard.

A professional analyzing data and web traffic charts on a laptop screen while working at a desk.

Navigating Modern Topologies and Edge Routing

The hardest load balancing setup decision is often not the algorithm. It's identifying which layer owns the routing decision.

DNS-based balancing returns different endpoints during name resolution. It can distribute users across locations, but resolver caching means a backend removal won't instantly affect every client. DNS is therefore useful for coarse regional steering and site-level failover, not for precise per-request balancing or application session management.

A proxy-level balancer receives the connection after DNS resolution and can inspect the request, apply health checks, terminate TLS, preserve cookies, and route by host or path. It centralizes the decision, which simplifies observability but creates a front-door dependency that needs redundancy.

Tunnel-backed origins change the setup entirely. Cloudflare's public load balancer workflow for Cloudflare Tunnel applications uses a tunnel UUID endpoint, header values for the published route, and a fallback pool, rather than a conventional pool of origin IP addresses (Cloudflare's public load balancer tunnel documentation).

TopologyBest fitWhat complicates operations
DNS levelRegional steering and site failoverResolver caching and limited request visibility
Proxy levelHTTP routing, TLS offload, and local balancingFront-door redundancy and capacity
Tunnel basedPrivate origins without direct exposureTunnel identity, headers, and fallback behavior
Service meshService-to-service routing inside clustersSidecar or dataplane policy and observability

Choose by failure domain

For a mixed environment with on-premises bare metal, Proxmox guests, and public cloud instances, define the failure domain before defining the pool. A single pool spanning every location can route users toward a distant or degraded site when the network path is technically alive. Regional DNS or edge steering can choose the site, while a local HAProxy or NGINX pair handles backend distribution.

Service meshes add another control plane. They make sense when service identity, east-west policy, and per-request telemetry already belong inside the mesh. They don't automatically replace an internet-facing load balancer. Adding a mesh solely to obtain round robin usually increases operational cost without solving the edge problem.

The current direction is adaptive traffic management across multiple layers, including edge distribution, tunnel connectivity, and mesh-integrated routing. The practical requirement remains stable: every layer needs its own health signal, logs, and rollback path. Don't call an endpoint healthy merely because a tunnel is connected or a VM responds to TCP. Test the application path users depend on.

Optimizing for Proxmox and Bare Metal Infrastructure

A load balancer cannot rescue a congested bridge, saturated interrupt queue, or host without CPU headroom for failover. On Proxmox, start by proving that the virtual and physical network paths are healthy:

ip -br link
ip -br addr
bridge link
ip -s link show vmbr0
ethtool -k eno1
ethtool -S eno1 | egrep -i 'drop|error|miss|timeout'

Give the load balancer VM a virtio NIC on the intended VLAN-aware bridge. Match offload settings between host and guest where the workload requires it. Drops in interface counters are a network problem to fix before changing HAProxy's balancing algorithm.

Under active load, inspect CPU scheduling and interrupt distribution on bare metal:

lscpu
grep -E 'eth|mlx|ixgbe|i40e|ena' /proc/interrupts
mpstat -P ALL 1
ss -s
nstat -az | egrep 'TcpExt|Tcp'

Avoid copying kernel tuning from another server. TCP buffers, conntrack limits, and queue settings depend on the kernel, NIC, traffic pattern, and available memory. Capture a baseline, change one value, and keep a tested rollback. Connection lifetime also affects backend pressure. Google guidance recommends rotating long-lived backend connections after roughly 10 to 20 minutes, or after approximately 1,000 to 2,000 requests per connection, to reduce uneven load (Radware's practical load-balancing guidance).

Separate management, storage, migration, and service traffic in Proxmox private clouds when the hardware and design allow it. A load balancer VM sharing a busy storage path can produce application latency while its virtual CPU remains mostly idle. Keep backend pools close to expected users, consistent with the operational guidance cited above.

After deployment, verify each tenant and backend pool under realistic traffic. Watch per-VM and per-interface counters, connection distribution, retransmits, and storage latency while generating requests. On multi-tenant hardware, a noisy network queue or storage path often explains imbalance faster than changing balance roundrobin to leastconn.

Dedicated hardware fits deployments whose connection rate, packet rate, inspection needs, or failure-domain requirements exceed two software instances. Managed operations fit teams that cannot own certificate rotation, health-check testing, incident response, and Proxmox networking. ARPHost, LLC provides colocation, bare metal servers, Proxmox private clouds, VPS infrastructure, and managed services. Review ARPHost's infrastructure services when selecting the physical or virtual platform behind the balancing layer.


ARPHost, LLC supports load balancing deployments on VPS, bare metal, Proxmox private clouds, and colocated infrastructure, with managed services available for certificate rotation, monitoring, failover testing, and host networking. Contact the provider to discuss the infrastructure layer behind your load balancing setup.

Tags: , , , ,

Leave a Reply