Notes guide
Liveness is not health
The worst Kubernetes outage I keep seeing is not a node dying. It is a slow database, a liveness probe that talks to that database, and a cluster that then kills every pod that was still doing useful work.
Liveness does not mean “the app is healthy.” It means “this process is so stuck that Kubernetes should shoot it.” If you point that gun at a dependency, you will shoot the fleet.
What I believed → what I do now
I used to expose one /health endpoint and hang both liveness and readiness on it, because Spring Actuator made that easy. I now split them: liveness is process-local, readiness is “safe to send traffic,” and startup is “the JVM is still booting, leave it alone.”
Three probes, three questions
Startup. “Is the process still starting?” A Spring Boot JVM can take a minute to reach main plus context refresh. If liveness runs during that window, Kubernetes restarts a pod that was never unhealthy. startupProbe exists so you can be slow once.
Liveness. “Is this process deadlocked or stuck forever?” The check must not depend on the database, Redis, a downstream HTTP API, or disk that might hitch. If those are down, you want the pod alive and failing readiness, not restarting.
Readiness. “Should this pod receive traffic?” This is where you check the database, a broker, or a warm cache. If it fails, Kubernetes removes the pod from the Service. The process keeps running. When the dependency returns, the pod takes traffic again without a cold JVM.
Using the same Actuator /health for liveness and readiness collapses those questions into one. A blip in Postgres then looks like a crash loop.
The restart storm
- Postgres is slow (load, failover, network).
- Every pod’s liveness probe waits, then fails.
- Kubernetes restarts the pods.
- The new JVMs all connect at once.
- Postgres gets slower. Probes fail again.
You did not have an application bug. You taught the orchestrator to amplify a dependency failure.
The fix is boring: liveness is /livez or Actuator’s liveness group with no downstream checks. Readiness is /readyz with the checks that gate traffic. Timeouts on probes are short. Failure thresholds are not “retry until the universe ends.”
JVM-specific traps
- CPU limits plus a liveness HTTP server: the probe thread starves, the probe fails, the pod dies, the problem looks like a memory leak. I am slow to set CPU limits on latency-sensitive Java services.
- Huge heaps and no startup probe: first GC during boot plus a tight liveness timeout looks like a failed deploy.
- Actuator on the same port as public traffic: fine for many shops; still do not let the public path be the liveness path if that path can block on I/O.
What I put in the manifest
I want to see, in the same PR as the service:
startupProbewith a longfailureThresholdfor Boot.livenessProbeagainst a handler that only answers “the process can still run a request on this port.”readinessProbeagainst a handler that checks the dependencies this instance must have before it takes a user request.- Timeouts in milliseconds you could explain in a standup, not copied from a blog.
If those four are missing, I do not care how pretty the Helm chart is.
Takeaways
- Liveness means “kill this process.” Do not aim it at the database.
- Readiness means “stop sending traffic.” That is the right place for dependency checks.
- A slow JVM boot needs a startup probe, not a faster liveness timeout.
- One
/healthfor both probes is how a dependency blip becomes a cluster-wide restart.