fix: prevent HAProxy from reusing stale keep-alive conns to nginx/Nextcloud
buildbot/nix-eval Build done. (1 warning)
buildbot/nix-build Build done.

DAVx5 (CalDAV/CardDAV) on greg's phone was intermittently failing every
sync type (CONTACTS/EVENTS/TASKS/RefreshCollectionsWorker) against
next.thehellings.com with:

  java.io.IOException: unexpected end of stream
  Caused by: java.io.EOFException: \n not found: limit=0

This is the classic OkHttp/HTTP client signature of the far end
silently closing a pooled keep-alive connection: the client reuses a
socket it still believes is open, gets zero bytes back while reading
response headers, and throws exactly this exception.

Root cause: HAProxy's 'next' backend proxies to nginx on
127.0.0.1:8080, and HAProxy defaults to end-to-end keep-alive (both
client- and server-side) unless told otherwise. nginx's
keepalive_timeout is 65s, so any HAProxy<->nginx connection idle past
that gets closed by nginx without HAProxy's knowledge. A request that
lands on that now-dead pooled connection right after gets nothing back
- surfacing to the client as a bare socket EOF while reading headers.
The frontend's existing 'option http-server-close'/'http-keep-alive'
pair only governs the client-facing side of HAProxy and does nothing
for the HAProxy->nginx leg.

Fix:
- backend next: add 'option http-server-close' so HAProxy opens a
  fresh connection to nginx per request instead of pooling/reusing
  one. The backend is localhost, so the extra TCP handshake cost is
  negligible, and this removes the whole class of stale-connection EOF
  errors.
- defaults: add 'timeout http-keep-alive 30s' to bound how long an
  idle client-facing keep-alive connection is held open. Previously
  unset, it fell back to 'timeout client' (500s) - unnecessarily long
  given maxconn is only 80, and tightens client-side connection churn
  to be more predictable too.

Diagnosed by pulling the nginx_access journal (enabled in #37/#38) for
the failing sync window and cross-referencing nginx's
services.nginx.appendHttpConfig / generated nginx.conf keepalive
settings against HAProxy's request-level defaults. Could not run
'haproxy -c'/'nginx -t' locally (no toolchain in the agent sandbox) -
recommend confirming via CI/garnix before merge, same as #38.
This commit is contained in:
2026-08-09 22:21:20 -05:00
parent daa33daa0c
commit 2c3607f17c
+17
View File
@@ -176,6 +176,12 @@ in
timeout connect 500s
timeout client 500s
timeout server 1h
# HAProxy defaults to end-to-end keep-alive (client AND server side)
# unless a proxy overrides it. Bound how long an idle client-facing
# keep-alive connection is held: maxconn is only 80, and leaving
# this unset falls back to "timeout client" (500s), which is far
# longer than needed just to wait for a pipelined next request.
timeout http-keep-alive 30s
listen gitsshd
bind *:${toString sshPort}
@@ -265,6 +271,17 @@ in
option accept-unsafe-violations-in-http-response
retries 3
option forwardfor
# nginx (the actual listener on 127.0.0.1:8080) has
# keepalive_timeout 65s and will silently close an idle backend
# socket after that. HAProxy's default mode is end-to-end
# keep-alive, so without this it will happily try to reuse a
# backend connection nginx already closed once a mobile client's
# own (longer) keep-alive idle assumption outlives 65s - producing
# exactly the "unexpected end of stream" / EOFException the
# CalDAV/CardDAV client saw. Since the backend is localhost, the
# cost of a fresh TCP connection per request is negligible, so
# just don't try to reuse them here.
option http-server-close
#http-response replace-value Location http://localhost:${builtins.toString nextcloudPort}/(.*) https://next.thehellings.com/\2
server nextcloud 127.0.0.1:${builtins.toString nextcloudPort}
'';