Remote Upstreams & Failover

Connect sockguard to a remote Docker daemon over TCP+mTLS, or configure two endpoints for active/passive HA failover with automatic health probing.

By default sockguard reaches Docker through a local unix socket. The upstream.endpoints block lifts that constraint: you can point sockguard at a remote daemon over TCP+TLS, or list two endpoints so a healthy standby takes over automatically when the primary goes down.

When to use this

  • Single remote daemon — Docker runs on a different host than sockguard (a build host, a CI worker, a remote VM). You want mTLS between them so the daemon API is not exposed as plaintext on the wire.
  • HA / redundancy — you have two daemon hosts behind keepalived or a Swarm manager HA pair and want sockguard to stay healthy when one goes down.
  • docker -H tcp://… migration — you already have DOCKER_HOST / DOCKER_TLS / DOCKER_TLS_VERIFY / DOCKER_CERT_PATH set and want zero-config drop-in (see DOCKER_* environment drop-in below).

Single remote daemon (TCP + mTLS)

The simplest remote setup: one endpoint, mutual TLS.

upstream:
  endpoints:
    - address: tcp://dockerd.internal:2376
      tls:
        ca_file: /certs/ca.pem        # verifies the daemon's server cert
        cert_file: /certs/cert.pem    # client cert sockguard presents
        key_file: /certs/key.pem

ca_file is the CA that issued the daemon's TLS certificate. cert_file / key_file are the client keypair the daemon uses to authenticate sockguard. This mirrors the standard Docker mTLS setup (dockerd --tlsverify).

Sockguard's upstream TLS client floors at TLS 1.2 when dialing a daemon — not the TLS 1.3 minimum the inbound listener enforces. The looser floor is deliberate: it keeps sockguard compatible with daemons (and the OS TLS stacks behind them) that don't yet negotiate TLS 1.3, while the listener — the side you control and expose to clients — stays at 1.3. Both sides still negotiate the highest version both peers support, so a modern daemon connects over 1.3 regardless.

When endpoints is non-empty, upstream.socket is ignored. You cannot mix a local socket fallback with remote endpoints.

SNI / hostname override

By default the hostname for TLS verification is derived from the address host. If your cert uses a different name (e.g. a SAN that doesn't match the IP):

upstream:
  endpoints:
    - address: tcp://10.0.1.5:2376
      tls:
        ca_file: /certs/ca.pem
        cert_file: /certs/cert.pem
        key_file: /certs/key.pem
        server_name: dockerd.internal   # override SNI and verified hostname

HA failover with two endpoints

List endpoints in priority order. Sockguard picks the first healthy one and routes all traffic through it. If that endpoint fails a health probe or a request dial, it is demoted and the next healthy endpoint takes over.

upstream:
  endpoints:
    - address: tcp://dockerd-a:2376
      tls:
        ca_file: /certs/ca.pem
        cert_file: /certs/cert.pem
        key_file: /certs/key.pem
    - address: tcp://dockerd-b:2376
      tls:
        ca_file: /certs/ca.pem
        cert_file: /certs/cert.pem
        key_file: /certs/key.pem
  failover:
    health_interval: "5s"   # probe period; empty = 5s default; negative disables continuous probing
    health_timeout: "2s"    # per-probe deadline; empty = 2s default

How failover works

  • Active endpoint — always the first known-healthy endpoint in list order. dockerd-a wins when both are healthy.
  • Health probe — sockguard dials each endpoint on the health_interval (TCP connect + TLS handshake for TLS endpoints). A probe that times out or is refused marks that endpoint unhealthy.
  • On dial failure during a request — the active endpoint is demoted immediately. The in-flight request fails and the client sees an error. The next request routes to the next healthy endpoint.
  • No automatic retry — the failing request is not retried. Docker writes are not idempotent, so a silent retry after a connection drop could execute an operation twice. Callers are expected to retry if the operation is safe to repeat.
  • Recovery — a demoted endpoint is re-probed on the health interval. Once it passes, it resumes its position in the priority order.

Set health_interval to a negative value to disable continuous probing. Sockguard will still detect failures at request time, but will not issue background health probes. Useful when probe traffic to the daemon is undesirable (metered links, audit-heavy environments). One probe still runs at startup regardless — the resolver seeds every endpoint's health state once before the loop checks the interval — so the active endpoint is chosen from real probe results rather than from list order alone. Only the recurring probes are suppressed.

Same-daemon constraint

All endpoints in the list MUST point to the same logical Docker daemon or Swarm cluster. This is active/passive redundancy — not load balancing or fan-out across different daemons.

Container IDs, exec sessions, volume state, and sockguard owner labels are daemon-local. Failing a live session from dockerd-a to a genuinely different dockerd-b would expose the caller to dangling IDs, missing state, and exec sessions that no longer exist. The proxy has no way to detect or compensate for that split.

Correct use cases: a Swarm manager VIP with two manager IPs behind it, a keepalived HA pair sharing state, two addresses for the same daemon on different interfaces.

Incorrect use case: two independent Docker hosts running different containers. Use separate sockguard instances for that.

Insecure and deprecated opt-ins

Two flags loosen the TLS requirement. Plaintext TCP remains an explicit risk acknowledgment. Skipping server verification is deprecated in v2.1 and scheduled for removal in v3.0.0.

Plaintext TCP (no TLS)

upstream:
  endpoints:
    - address: tcp://dockerd.internal:2376
      insecure_allow_plain_tcp: true

insecure_allow_plain_tcp: true permits a tcp:// endpoint with no TLS material at all. The Docker API is sent in plaintext — any host on the path can read or inject requests. Only use this on a private, trusted network with no external exposure. The flag mirrors the same acknowledgment on the listener side (listen.insecure_allow_plain_tcp).

Deprecated: skip server certificate verification

upstream:
  endpoints:
    - address: tcp://dockerd.internal:2376
      tls:
        cert_file: /certs/cert.pem
        key_file: /certs/key.pem
      insecure_skip_tls_verify: true   # endpoint-level, a sibling of `tls`

insecure_skip_tls_verify: true skips verification of the daemon's server certificate. Traffic is still encrypted but the daemon's identity is not verified, so a man-in-the-middle can present any certificate. It is an endpoint-level field (a sibling of tls, address, and insecure_allow_plain_tcp), not a key inside the tls block.

The field remains accepted with unchanged wire behavior throughout v2.1, but it logs a startup deprecation warning and will be removed in v3.0.0. Replace it with the correct tls.ca_file, including for self-signed or private-CA daemon certificates.

DOCKER_* environment drop-in

If you have a working docker -H tcp://… setup with the standard Docker client env vars, sockguard picks them up automatically when no endpoints are configured in YAML:

Environment variableEffect
DOCKER_HOST=tcp://host:port[/base/path], host[:port][/base/path], or unix:///pathRoutes to that TCP or Unix endpoint; portless TCP defaults to 2375, and a TCP path is prepended to every Docker API request
DOCKER_TLS=<non-empty>Enables TLS without server verification; deprecated for removal in v3.0.0
DOCKER_TLS_VERIFY=<non-empty>Enables TLS verification and takes precedence over DOCKER_TLS
DOCKER_CERT_PATH=/pathLocates ca.pem, cert.pem, and key.pem when TLS is enabled; does not enable TLS by itself
DOCKER_CONFIG=/pathSupplies the certificate directory when DOCKER_CERT_PATH is empty; otherwise the Docker default is ~/.docker

Precedence: upstream.endpoints (YAML) > DOCKER_HOST (env) > upstream.socket (YAML/default). A valid custom unix:// value is used as a literal socket name; Sockguard does not URL-decode it into a different path. Absolute-path values such as unix:///tmp/docker%2Fsock?one#two preserve valid percent escapes and raw ? or # bytes. Relative host-form values preserve raw ? and # too, but follow Docker's URL host grammar, which rejects escapes such as %2F, %3F, and %23 there. Only an absent DOCKER_HOST falls through to upstream.socket; a present empty, malformed, http:///https://, or unsupported transport value fails startup instead of silently selecting another daemon.

A TCP URL path is an API prefix, matching the Docker CLI. For example, DOCKER_HOST=tcp://gateway.internal:2375/docker sends the client's /v1.52/containers/json request to /docker/v1.52/containers/json. The prefix follows the endpoint selected for that request, including failover, readiness and inspection requests, and the initial attach, exec, and mediated BuildKit upgrade. Percent-encoded separators and dot segments are kept as configured rather than cleaned or decoded while the prefix and client path are joined.

Sockguard follows the Docker CLI's presence semantics: an empty or unset TLS variable is disabled, while every non-empty value, including 0 and false, is enabled. DOCKER_TLS_VERIFY wins when both TLS variables are non-empty, so no YAML acknowledgment is needed for the env drop-in:

  • DOCKER_TLS_VERIFY non-empty → verified TLS using ca.pem from the Docker certificate directory. If both cert.pem and key.pem exist, sockguard also presents them for mutual TLS; if either is absent, it uses server-auth-only TLS.
  • DOCKER_TLS non-empty + DOCKER_TLS_VERIFY empty or unset → encrypted, but the daemon's server certificate is not verified. Docker still requires a readable, valid ca.pem; the optional client pair is resolved the same way. This mode logs a source-specific deprecation warning and will be removed in v3.0.0. Set DOCKER_TLS_VERIFY=1 before upgrading.
  • Both TLS variables empty or unset → plaintext TCP (equivalent to insecure_allow_plain_tcp), regardless of DOCKER_CERT_PATH. The acknowledgment is implicit because the Docker CLI also treats the certificate path as a locator rather than a TLS enablement signal.

When TLS is enabled, the certificate directory is DOCKER_CERT_PATH when non-empty, otherwise DOCKER_CONFIG when non-empty, otherwise ~/.docker. ca.pem is required and replaces the system root pool. cert.pem and key.pem are optional as a pair. This is the same lookup the Docker CLI applies, including when no explicit certificate path is set.

This means an existing Docker CLI setup works with zero YAML changes — just point sockguard at the same env vars your client uses. To override any of these, set upstream.endpoints in YAML, which takes precedence over the environment.

Reload immutability

upstream.endpoints, upstream.failover and upstream.flavor are reload-immutable. Adding, removing, or changing endpoints requires a process restart, and so does changing which engine sockguard thinks it is talking to. upstream.request_timeout and upstream.hijack_inactivity_timeout are reload-mutable and take effect on hot reload without a restart. Only request_timeout's default value changed in v1.5 (unlimited → 60s); the field's mutability is unchanged.

That is the same set upstream.socket belongs to, also pinned at startup. The upstream transport is bound to long-lived connection pools that cannot be swapped safely from within a running process.

Unix socket endpoints

You can also reference a unix socket explicitly in the endpoints list, which is useful when you want the health probing and failover machinery even for a local socket:

upstream:
  endpoints:
    - address: unix:///var/run/docker.sock
    - address: /var/run/docker-secondary.sock   # bare path treated as unix://

A bare path (starting with /) is treated as a unix:// address. TLS fields on a unix endpoint aren't ignored, they're rejected: tls.server_name, ca_file, cert_file or key_file on a unix:// address fails startup, because it's almost always a copy-paste from a TCP entry.

Full schema reference

upstream:
  socket: /var/run/docker.sock     # legacy; used only when endpoints is empty
  request_timeout: "60s"            # Go duration (e.g. "30s"); default "60s"; set "off" (or "") to disable; reload-mutable
  hijack_inactivity_timeout: "10m" # positive Go duration; shared idle guard for attach/exec-start; reload-mutable
  flavor: auto                      # auto | docker | podman; auto probes GET /version once at startup; reload-immutable
  endpoints:
    - address: tcp://dockerd-a:2376
      tls:
        ca_file: /certs/ca.pem
        cert_file: /certs/cert.pem
        key_file: /certs/key.pem
        server_name: ""                # SNI override; empty = derived from address host
      insecure_allow_plain_tcp: false  # permit tcp:// with no TLS (plaintext)
      insecure_skip_tls_verify: false  # deprecated; removed in v3.0.0
    - address: tcp://dockerd-b:2376
      tls: { ca_file: /certs/ca.pem, cert_file: /certs/cert.pem, key_file: /certs/key.pem }
  failover:
    health_interval: "5s"    # empty = 5s default; negative = disable continuous probing
    health_timeout: "2s"     # empty = 2s default

Per-endpoint fields inside endpoints cannot be set via environment variable — list types require YAML. The failover timing fields have env-var equivalents:

VariableYAML fieldDefaultDescription
SOCKGUARD_UPSTREAM_REQUEST_TIMEOUTupstream.request_timeout"60s"Total per-request deadline. Set off (or "") to disable — prefer off via env var since an explicitly empty value is treated as unset. Reload-mutable.
SOCKGUARD_UPSTREAM_HIJACK_INACTIVITY_TIMEOUTupstream.hijack_inactivity_timeout"10m"Connection-wide idle deadline for hijacked attach/exec-start streams. Must be a positive duration; no off spelling. Reload-mutable.
SOCKGUARD_UPSTREAM_FLAVORupstream.flavorautoWhich engine is behind the upstream: auto, docker, or podman. auto probes GET /version once at startup and fails startup if the answer is ambiguous; an explicit value never probes. Reload-immutable. See Configuration.
SOCKGUARD_UPSTREAM_FAILOVER_HEALTH_INTERVALupstream.failover.health_interval"" (resolver default: 5s)Background probe interval per endpoint. Empty uses the 5s resolver default; negative disables probing.
SOCKGUARD_UPSTREAM_FAILOVER_HEALTH_TIMEOUTupstream.failover.health_timeout"" (resolver default: 2s)Per-probe dial+TLS-handshake timeout. Empty uses the 2s resolver default.