Why your LLM endpoint returns 503 (it’s rarely the model)

2026/09/29

For about a year I have been the person the ticket lands on when a model endpoint answers 503, 500, or “server unreachable”. Different clusters, different models, same ticket title. Here is what it turned out to be, sorted by where the request actually died.

One request walking through client, edge, gateway, engine and model; at each layer a card shows the root cause that was found there

Almost none of these were the model. Most were something in front of it.

Edge: WAF and load balancer

The request never reached the cluster. It looked exactly like an outage.

Gateway: rate limits

Engine: how the server was launched

Model: it was the model

Sometimes it is. Twice, in a year.

The lie that cost the most time

In the last case, /health returned 200 for the whole incident while every real request returned 500. The readiness probe checked that a process was listening. It did not check that the model could answer.

graph LR
    A["/health<br/>process up?"] -->|200| B["traffic routed"]
    B --> C["POST /v1/chat<br/>500"]
    D["/ready<br/>tiny real inference"] -->|fails| E["pod pulled from rotation"]
    style C fill:#f8514922,stroke:#f85149
    style E fill:#3fb95022,stroke:#3fb950

Readiness should run a real request: one token, a fixed prompt, a tight timeout. It costs almost nothing and it is the difference between “pod is up” and “pod can serve”.

How I triage now

Outside-in, and prove each layer before moving on:

  1. Reproduce from inside the cluster with curl --resolve against the ingress. If it works inside, stop looking at the model.
  2. Diff the two paths with numbers: status code, upload speed, time to first byte. Numbers are what the network team can act on.
  3. Check the gateway config against what is in git. If nothing is in git, that is the finding.
  4. Read the engine launch line — version, KV settings, max_model_len, output bounds. Compare with the model’s recipe.
  5. Only then run the long, technical prompts. If it breaks here, it is the model, and the fix is a different build.

What I would do differently