Triggered by investigating a several-hour >10Mbps traffic spike to linode. HAProxy's own IPAccounting confirmed ~121GB moved over ~19.6h before it crash-looped, but with 'option httplog' commented out and no per-backend request logs, there was no way to attribute that traffic to a specific backend, host, or client. - linode: enable HAProxy httplog + defaults 'log global' (was commented out) so every proxied HTTP request is now logged with timing/status/bytes. - linode: add a haproxy 'stats' listener on 127.0.0.1:8404 for live per-backend/per-server connection and byte counters. - linode: route nginx (Nextcloud's local vhost) access logs to journald via syslog, since the read-only monitoring account has no access to /var/log/nginx/*. - linode: enable vnstat for historical per-interface bandwidth tracking (5-min granularity) so a reported 'traffic was high for N hours' can be confirmed/timestamped immediately instead of reconstructed after the fact from journal timestamps. - k3s manifests: enable Traefik access logging (JSON) — this is the ingress layer HAProxy forwards :80 traffic to (git/matrix/immich), and lacked any per-request visibility. - hosts/baseline.nix (fleet-wide): add a journald rate limit (2000 lines / 30s per unit). Found live while investigating that uptime-kuma on 'kuma' was logging a Prometheus label-validation error on every monitor beat (~100k lines/hour), which was itself degrading journalctl responsiveness on that host during the cross-host traffic scan. Related but not otherwise addressed here: Nebula relay/handshake churn on kuma's tunnel and the etcd read-latency warnings seen on isaiah/zeke around the same incident window — noted for a future investigation, not fixed by this PR.
34 lines
1.1 KiB
YAML
34 lines
1.1 KiB
YAML
apiVersion: "helm.cattle.io/v1"
|
|
kind: "HelmChartConfig"
|
|
metadata:
|
|
name: "traefik"
|
|
namespace: "kube-system"
|
|
spec:
|
|
valuesContent: |-
|
|
additionalArguments:
|
|
- "--entryPoints.postgres.address=:5432/tcp"
|
|
- "--api.dashboard=true"
|
|
- "--api.insecure=true"
|
|
- "--log.level=DEBUG"
|
|
# Access logging: gives per-request visibility (client IP, host,
|
|
# path, bytes, duration) for every ingress route Traefik terminates
|
|
# (git.k3s.thehellings.lan, matrix.k3s.thehellings.lan, immich, etc).
|
|
# This is the layer HAProxy on linode forwards :80 traffic to, so
|
|
# having request-level logs here is essential for tracing bandwidth
|
|
# spikes back to a specific host/path/client rather than just a
|
|
# backend-level byte count.
|
|
- "--accesslog=true"
|
|
- "--accesslog.format=json"
|
|
- "--accesslog.fields.headers.defaultmode=keep"
|
|
ports:
|
|
postgres:
|
|
expose:
|
|
default: true
|
|
port: 5432
|
|
exposedPort: 5432
|
|
protocol: TCP
|
|
traefik:
|
|
expose:
|
|
default: true
|
|
|