Two servers, one connection

Almost every production web request passes through more than one HTTP server: a reverse proxy, a load balancer or a CDN in front, and an application server behind. The front end reads a request from the client and forwards it to the back end, usually over a reused connection shared by many clients.

For that to be safe, both servers must agree on where one request ends and the next begins. Request smuggling is what happens when they disagree. The front end thinks it forwarded one request; the back end thinks it forwarded one and a half, and the leftover half becomes the beginning of an attacker-influenced request for whoever is next on that connection.

The two ways to disagree

A request declares its body length in one of two ways, and a desync appears when the two servers prefer different ones.

CL.TE

The front end uses Content-Length; the back end uses Transfer-Encoding. A request that carries both headers, sized so the two readings differ, causes the back end to consume a body the front end counted as a separate request.

TE.CL

The mirror image: the front end uses Transfer-Encoding and the back end uses Content-Length. The same crafted request leaves a fragment behind, but the smuggling is driven from the other side.

There is a third shape, TE.TE, where both servers support chunked encoding but one can be persuaded to ignore it — through obfuscated header names, unusual casing, or whitespace that one parser tolerates and the other does not.

HTTP/2 is not automatically safe

HTTP/2 frames messages explicitly, so the ambiguity above does not exist at that layer. But most HTTP/2 front ends still speak HTTP/1.1 to the back end, and the downgrade reintroduces the problem. If the front end builds the back-end request by concatenating headers it received, a header value that contains a newline becomes a new header — a request line the back end will honour.

The defensive rule is precise: when translating HTTP/2 to HTTP/1.1, reject or normalise anything that could break the message framing rather than passing it through.

What a successful desync buys

Smuggling is a primitive, not a payload. What an attacker does with it depends on the path and the application:

  • Bypass front-end controls — the smuggled request is processed by the back end but never inspected by the WAF, rate limiter or authentication layer in front.
  • Poison the request queue — the next legitimate user on the connection has their request modified, or receives a response meant for someone else.
  • Cache poisoning — a smuggled request causes a response to be cached against a URL that a victim will later request.
  • Session and credential theft — in the worst case the smuggled request captures the next user's request, including cookies.

Because the front end and back end each believe they are behaving correctly, the logs disagree and the bug is invisible from either side alone.

Finding it

Detection is a timing and response-diffing problem. A probe that is harmless if the servers agree produces a tell-tale delay or a malformed response when they do not. The reliable method is to send a request shaped to expose the disagreement and observe whether the connection behaves as one request or two — checking status, timing and body against a known-good baseline.

Hugin includes smuggling probes in the scanner and exposes raw framing through the Repeater, so a suspected desync can be reproduced by hand rather than only reported as a heuristic. As always with this class, a confirmed desync is the interesting part; a timing wobble on a single request is not.