feat: enable request-level logging for bandwidth/traffic incident tracing
buildbot/nix-eval Build done. (1 warning)
buildbot/nix-build Build done.

Triggered by investigating a several-hour >10Mbps traffic spike to
linode. HAProxy's own IPAccounting confirmed ~121GB moved over ~19.6h
before it crash-looped, but with 'option httplog' commented out and no
per-backend request logs, there was no way to attribute that traffic
to a specific backend, host, or client.

- linode: enable HAProxy httplog + defaults 'log global' (was
  commented out) so every proxied HTTP request is now logged with
  timing/status/bytes.
- linode: add a haproxy 'stats' listener on 127.0.0.1:8404 for live
  per-backend/per-server connection and byte counters.
- linode: route nginx (Nextcloud's local vhost) access logs to
  journald via syslog, since the read-only monitoring account has no
  access to /var/log/nginx/*.
- linode: enable vnstat for historical per-interface bandwidth
  tracking (5-min granularity) so a reported 'traffic was high for N
  hours' can be confirmed/timestamped immediately instead of
  reconstructed after the fact from journal timestamps.
- k3s manifests: enable Traefik access logging (JSON) — this is the
  ingress layer HAProxy forwards :80 traffic to (git/matrix/immich),
  and lacked any per-request visibility.
- hosts/baseline.nix (fleet-wide): add a journald rate limit
  (2000 lines / 30s per unit). Found live while investigating that
  uptime-kuma on 'kuma' was logging a Prometheus label-validation
  error on every monitor beat (~100k lines/hour), which was itself
  degrading journalctl responsiveness on that host during the
  cross-host traffic scan.

Related but not otherwise addressed here: Nebula relay/handshake
churn on kuma's tunnel and the etcd read-latency warnings seen on
isaiah/zeke around the same incident window — noted for a future
investigation, not fixed by this PR.
This commit is contained in:
2026-08-09 16:15:35 -05:00
parent 3995eee7b0
commit 10cdf9408d
3 changed files with 56 additions and 7 deletions
+11
View File
@@ -10,6 +10,16 @@ spec:
- "--api.dashboard=true"
- "--api.insecure=true"
- "--log.level=DEBUG"
# Access logging: gives per-request visibility (client IP, host,
# path, bytes, duration) for every ingress route Traefik terminates
# (git.k3s.thehellings.lan, matrix.k3s.thehellings.lan, immich, etc).
# This is the layer HAProxy on linode forwards :80 traffic to, so
# having request-level logs here is essential for tracing bandwidth
# spikes back to a specific host/path/client rather than just a
# backend-level byte count.
- "--accesslog=true"
- "--accesslog.format=json"
- "--accesslog.fields.headers.defaultmode=keep"
ports:
postgres:
expose:
@@ -20,3 +30,4 @@ spec:
traefik:
expose:
default: true