Security Model

Sockguard's defense-in-depth model — transport admission, client admission, method/path filtering, request-body inspection, ownership isolation, visibility-controlled reads, and structured access plus audit logging.

Why Socket Proxying Matters

The Docker socket (/var/run/docker.sock) is equivalent to root access on the host. Any container with unrestricted socket access can:

  • Create a privileged container that mounts the host filesystem
  • Execute arbitrary commands via docker exec
  • Access host PID, network, and IPC namespaces
  • Pull and run malicious images
  • Manipulate Swarm clusters

Sockguard sits between consumers and the raw socket, inspecting and filtering every request before it reaches the daemon.

Defense in Depth

Sockguard implements multiple layers of filtering:

Layer 0: Policy Integrity

Before any request is evaluated, Sockguard can verify that the loaded configuration was signed by pinned out-of-band trust. Signed mode requires sockguard serve --policy-bundle-trust-config <path>. That separate bootstrap file reuses the policy_bundle schema for enabled, signing keys or keyless identities, Rekor posture, and verify_timeout. The signed candidate carries its own policy_bundle.signature_path, but any candidate copies of trust fields are ignored and replaced by the bootstrap values.

The bootstrap and candidate must resolve to different files. The same path, symlinks to the same target, and hardlinks to the same inode are rejected. Both YAML files are capped at 16 MiB, and the referenced Sigstore bundle is capped at 4 MiB, before parsing or verification. Only regular files are accepted. Inputs are opened nonblocking and FIFOs, devices, directories, and other non-regular paths are rejected. policy_bundle.enabled: true in the candidate without the bootstrap flag is also rejected because the file being authenticated cannot select or disable its own trust gate. Verification completes before the operational logger is constructed, before any rules compile, and before a listener opens. Early diagnostics use a fixed stderr logger, so unverified logging config cannot open or truncate a requested output path.

Two verification paths are supported:

  • Keyed — PEM-encoded ECDSA, RSA, or ed25519 public keys listed under policy_bundle.allowed_signing_keys. No network round-trip required.
  • Keyless (Fulcio + Rekor)policy_bundle.allowed_keyless entries constrain the Fulcio cert chain by exact OIDC issuer URL and subject SAN regex. When policy_bundle.require_rekor_inclusion: true, a Rekor transparency-log entry is additionally required. The process loads and memoizes the public Sigstore trust root through TUF initially, then refreshes it about every 24 hours, so ongoing egress is required. The loader uses no local cache, bounds each HTTP request to 15 seconds, and remains compatible with the read-only runtime filesystem. An initial load failure aborts startup; a failed background refresh is logged and retains the last valid root.

Verification also runs on every hot reload. A reload whose bundle fails verification is rejected with result=reject_signature in sockguard_config_reload_total and never touches the running policy. Bootstrap trust is pinned for the process lifetime; only the signed candidate's signature_path is reload-mutable so an operator can re-sign without a restart.

policy_bundle.verify_timeout is a cooperative deadline for local Sigstore verification. Sockguard checks cancellation before beginning and after each synchronous verification attempt, rejects late success, and stops signer fallback. Sigstore-go does not expose context-aware verification, so an individual crypto call cannot be preempted. Rekor inclusion proofs are checked locally from the bundle. TUF root downloads use their separate 15-second HTTP bound.

Signed mode also rejects every rule-generating Tecnativa variable, including section variables, POST, GRPC/SESSION, and granular ALLOW_* variables. The presence of one at startup fails before rule compilation, even if its value is false. If one appears later, reload records result=reject_compat and preserves the active policy. This prevents unsigned environment state from changing a verified rule set or opening a BuildKit transport.

The verified signer (keyed:<spki-fingerprint> or keyless:<issuer>:<san>) and the YAML's SHA-256 digest are stamped onto GET /admin/policy/version in the bundle_signer and bundle_digest fields, giving operators a tamper-evident audit trail of exactly who signed the running policy and over which bytes.

Layer 1: Transport Admission

Non-loopback TCP listeners require mutual TLS 1.3 by default via listen.tls. Plaintext remote TCP is rejected unless you set both listen.insecure_allow_plain_tcp: true and listen.insecure_allow_unauthenticated_clients: true — two deliberate acknowledgments (one without the other is rejected) for legacy compatibility on a private network. Unix socket listeners bypass this layer because they are filesystem-bounded. listen.tls.client_ca_file defines the issuing trust root, and the optional listen.tls.common_names, dns_names, ip_addresses, uri_sans, and public_key_sha256_pins fields can narrow that trust to specific verified client certificates instead of implicitly accepting every client cert issued by the configured CA.

Layer 2: Client Admission

clients.allowed_cidrs gates incoming TCP callers by source CIDR before any rule evaluation runs. When clients.container_labels.enabled is true, Sockguard resolves the calling container by source IP and enforces per-client com.sockguard.allow.<method> label allowlists in addition to the global rule set.

Named client profiles sit on top of that admission layer. Sockguard can now select a per-client ruleset and request-body policy by source IP, verified mTLS certificate selectors (common_names, dns_names, ip_addresses, uri_sans, spiffe_ids, public_key_sha256_pins), or unix peer credentials (uids, gids, pids), with a configurable default profile for unmatched callers. That turns one proxy from "one ruleset in front of Docker" into a shared control plane for multiple consumers without collapsing back to broad allowlists.

Layer 3: Method Filtering

Block entire HTTP methods. Most consumers only need GET (read-only mode).

Layer 4: Path Filtering

Allow or deny specific Docker API endpoint paths using glob patterns. Before matching, Sockguard strips Docker API version prefixes (/v1.45/), percent-decodes the path (including double-encoded separators and mixed-case escapes such as %2F, %2E, and %252F), and resolves . / .. segments via path.Clean. That means /v1.45/containers/%2e%2e/images/json canonicalizes to /images/json before the glob matcher sees it, so adversarial path shapes cannot slip past a literal allowlist or skip a request-body inspector.

Layer 5: Request Body Inspection

POST /containers/create bodies are parsed on every request and denied when they contain dangerous configuration:

  • HostConfig.Privileged: true
  • HostConfig.NetworkMode: host
  • HostConfig.PidMode: host
  • HostConfig.IpcMode: host
  • HostConfig.UsernsMode: host
  • A non-empty HostConfig.Sysctls map (kernel parameter tuning), unless request_body.container_create.allow_sysctls is set
  • A non-empty HostConfig.Runtime value not present in request_body.container_create.allowed_runtimes (defends against runtime-escape via alternative OCI runtimes; an empty/unset runtime selects the daemon default and is always permitted)
  • Any bind mount whose source is outside request_body.container_create.allowed_bind_mounts
  • Any HostConfig.Devices host path outside request_body.container_create.allowed_devices
  • HostConfig.DeviceRequests unless explicitly allowed
  • HostConfig.DeviceCgroupRules unless explicitly allowed

Five further HostConfig fields are denied unconditionally — no policy setting opts back in — because each one opens a namespace-escape or privilege-escalation path: VolumesFrom, UTSMode: host, a non-empty CgroupParent, GroupAdd, and ExtraHosts.

POST /containers/*/exec and POST /exec/*/start can also be inspected now. When request_body.exec.allowed_commands is configured, Sockguard denies argv vectors that match no allowlist entry, denies privileged exec unless explicitly allowed, denies root-user exec unless explicitly allowed, and re-inspects POST /exec/*/start against Docker's stored exec metadata before the command runs. Each allowlist entry is an argv template whose tokens are sockguard globs (* matches a run of non-slash characters, ** matches any sequence): a command matches when its token count equals an entry's and every token matches the glob at that position, so an exec whose argv carries a variable component — a run ID, timestamp, or generated path — can be allowlisted without enumerating every literal form. Keep glob tokens as tight as the use case allows; a token of ** matches anything.

When a client sends Cmd as a single JSON string instead of an array, Sockguard tokenizes it on whitespace before matching — it does not interpret shell quoting, so "sh -c 'rm -rf /tmp/x'" becomes the five tokens sh, -c, 'rm, -rf, /tmp/x', not three. Docker itself passes the string form through without re-parsing, so the tokens Sockguard matches can differ from what ultimately executes. Prefer the array form in your Docker API clients when exec allowlists are in play, and write allowlist entries against array-form argv.

The exec-start re-inspection is necessarily best effort: Docker exposes metadata inspection and exec start as separate API calls, so Sockguard cannot make the check atomic with the eventual start operation. Treat this as a narrow TOCTOU window inherent to Docker's API shape, and prefer tight exec allowlists plus conservative per-client profile assignment for clients that do not need interactive command execution.

request_body.exec.allowed_env_vars and request_body.exec.denied_env_vars optionally restrict the Env array on exec create by variable name. denied_env_vars is checked first and always wins: a name on that list is blocked even if it also appears in allowed_env_vars. If allowed_env_vars is non-empty, any Env entry whose name isn't listed is denied. Both default to empty, meaning no name restriction — unlike allowed_commands, which denies all exec once inspection is active. Name matching is exact and case-sensitive against the substring before the first =; there is no glob support and no curated built-in denylist. A common defense-in-depth denylist targets dynamic-linker and interpreter hijack vectors that let an attacker redirect what code an exec'd binary loads:

request_body:
  exec:
    denied_env_vars:
      - LD_PRELOAD
      - LD_LIBRARY_PATH
      - LD_AUDIT
      - DYLD_INSERT_LIBRARIES
      - PYTHONPATH
      - PATH

When a permitted variable also carries security-sensitive routing or behavior, allowed_env_values can pin it to one or more exact full entries:

request_body:
  exec:
    allowed_env_vars: [CALLBACK_URL]
    allowed_env_values:
      - CALLBACK_URL=http://127.0.0.1:3000/callback

Only names represented in allowed_env_values receive a value constraint; other names continue to follow the name lists. Matching is exact—no glob or regex expansion. Sockguard compares values but never includes them in logs or denial reasons, so secrets are not reflected.

This filtering only applies at exec create — see Known Limitations below for why it cannot be re-checked at exec start.

POST /images/create is inspected by default. Sockguard blocks fromSrc imports unless explicitly allowed and constrains pulls to Docker Hub official images unless the operator opts into allow_all_registries or an explicit registry allowlist.

POST /build and native Podman's POST /libpod/build are inspected by default through the same request_body.build policy. Sockguard blocks primary remote contexts, Podman URL/image additional contexts, networkmode=host, and Dockerfiles containing RUN instructions unless those behaviors are explicitly allowed. Podman host volume controls, local-path additional contexts, multipart local contexts, and resource-usage output to a daemon-host file require insecure_allow_body_blind_writes; all other build checks remain active when that acknowledgment is set. Direct and version-prefixed libpod requests follow the same body-sensitive startup validation as Docker's classic builder.

POST /services/create and POST /services/*/update are inspected by default. Sockguard blocks host-network services, bind mounts outside request_body.service.allowed_bind_mounts, and service images outside the configured official/allowlisted registry set. The same allowed_capabilities / allow_all_capabilities capability allowlist, allow_sysctls sysctl gate, and image_trust cosign verification that apply to container-create are also enforced against the service's ContainerSpec, so swarm workloads cannot bypass container-create policy by going through the service API.

Image-trust discovery stops before hostile registry material can grow without bound. Every registry GET response stops at 4 MiB, including redirect destinations; a request can process at most 32 referrer descriptors, 16 distinct signature images, 32 layers per signature manifest, 16 aggregate verification candidates, 256 KiB of aggregate annotation keys and values, 1 MiB per simple-signing payload, and 16 MiB across all payload reads for one image. Any limit breach rejects discovery even when a valid sibling signature exists. In enforce mode that denies the request; in warn mode Sockguard logs the failed discovery and forwards it. Signature references must resolve directly to image manifests; OCI indexes are rejected instead of recursively traversed, and payload layers with alternate URLs are rejected before blob resolution. Legal media-type parameters on direct manifests are accepted. Discovery deadlines retain their cancellation cause instead of being reported as an unsigned image. Signature images are processed as they are discovered instead of retained together.

POST /volumes/create, POST /secrets/create, and POST /configs/create are inspected by default. Sockguard blocks non-local volume drivers and driver options unless explicitly allowed, and blocks custom or template drivers on secrets/configs unless explicitly allowed.

POST /swarm/init, POST /swarm/join, and POST /swarm/update are inspected by default. Sockguard blocks ForceNewCluster, external CA configuration, non-allowlisted join targets, token rotations, manager unlock-key rotations, manager autolock, and signing-CA updates unless explicitly allowed.

POST /plugins/pull, POST /plugins/*/upgrade, POST /plugins/*/set, and POST /plugins/create are inspected by default. Sockguard constrains remote registries, privilege grants, plugin-set assignments, local plugin tar config.json, host mounts, device exposure, and capability requests unless explicitly allowed. POST /plugins/create is treated as multipart/form-data as well as raw tar: Sockguard spools the upload to a temporary file, parses the multipart envelope, extracts config.json from the embedded tar, and applies the same plugin policy it applies to POST /plugins/pull. Uploads without a parseable config.json, or whose config.json fails policy, are denied before the body reaches Docker.

POST /networks/create, POST /networks/*/connect, and POST /networks/*/disconnect are inspected by default. Sockguard blocks custom network drivers, swarm/ingress/attachable/config-only networks, custom IPAM drivers/config/options, driver options, and forced disconnects unless explicitly allowed. request_body.network.allow_endpoint_config additionally gates endpoint static IP, MAC address, links, and driver options — enforced identically on POST /networks/*/connect's EndpointConfig and POST /containers/create's NetworkingConfig.EndpointsConfig, since Docker connects a create-time network entry the same way a follow-up connect call would; putting the same config on create's primary network used to be an unchecked bypass of the connect-side gate. Endpoint Aliases are always allowed at both endpoints, regardless of this flag: Docker Compose sets Aliases: [serviceName] on every endpoint it creates, so gating them broke every multi-network Compose recreate for no real security benefit — aliases were never enforced on create's primary network either.

Deployments that don't want to admit the whole EndpointSettings object can narrow this to independent per-field gates instead (#186): request_body.network.endpoint_config.allow_static_addressing, .allow_link_local_ips, .allow_mac_pinning, and .allow_gw_priority each default false, while .allow_aliases defaults true to keep the Compose behavior above unconditional. This block is only consulted when allow_endpoint_config is false — setting both is a config validation error, since the legacy flag already subsumes it. Links and DriverOpts have no granular escape hatch; they stay denied unless allow_endpoint_config: true is set. See Configuration for the exact mapping and env vars.

POST /containers/*/update and PUT /containers/*/archive are inspected by default. Sockguard blocks restart-policy/resource-control changes, privileged/device/capability-like update fields, unsafe archive target paths, tar traversal, setuid/setgid entries, device nodes, and escaping symlinks/hardlinks unless explicitly allowed.

POST /images/load is inspected by default. Image archives are denied unless image-load policy allows matching registries or untagged images; Docker manifest.json repo tags are checked against the same official/registry allowlist model used for pulls.

POST /swarm/unlock and POST /nodes/*/update are inspected by default. Swarm unlock is denied unless explicitly allowed, and node updates block role, availability, name, and arbitrary label mutations unless the corresponding node policy permits them. The default owner-label key remains allowed for controlled node claims.

Bounded JSON/tar inspectors read request bodies under per-endpoint byte caps and return 413 Payload Too Large when those caps are exceeded, instead of streaming unbounded bodies into memory. A malformed or hostile client cannot tie up the filter or the Docker daemon with oversized payloads because the bounded reader short-circuits before the JSON decode or tar parse begins. The filter also applies a 30-second read deadline to the request body before an inspector runs. The access-log and metrics wrappers preserve the response-controller interface used for that deadline, so enabling either layer cannot bypass it. On the upstream side, the reverse-proxy and side-channel transports set a 30-second response-header timeout. Attach and exec-start add a 30-second deadline across client body forwarding, upstream request writes, and the wait for response headers; the deadline is cleared only after a valid 101 Switching Protocols response, then the long-lived stream uses its inactivity guard.

These inspectors intentionally decode only the Docker request fields Sockguard actually enforces. They are not full Docker-schema validators, so full payload validation still belongs to Docker once Sockguard has checked the policy-relevant subset.

The remaining blind-write guardrail covers controls Sockguard still cannot constrain safely, chiefly arbitrary exec without an allowlist, POST /swarm/join without configured allowed_join_remote_addrs, plugin setting writes without allowed assignment prefixes, Podman build host/local/multipart or resource-usage host-file controls, and the documented uninspected libpod write surface. Validation refuses to start with broad uninspected rules unless you explicitly set insecure_allow_body_blind_writes: true, to keep the enforcement boundary honest. The flag is also wired into request-time enforcement for exec and Podman build. It lifts only the specific uninspectable gate; every other configured exec or build check still applies.

Sockguard now applies the same honesty rule to every recognized data-exfiltration surface. Validation refuses to start broad rules that would expose GET /containers/*/archive, GET /containers/*/export, GET /containers/*/logs, GET /containers/*/attach/ws, POST /containers/*/attach, GET /services/*/logs, GET /tasks/*/logs, GET /images/get, GET /images/*/get, POST /images/*/push, or POST /plugins/*/push unless you explicitly set insecure_allow_read_exfiltration: true. Image and plugin pushes are Docker API writes, but they read local artifacts and transmit them to a caller-selected registry, so they carry the same exfiltration risk; the option name is retained for configuration compatibility. Because this validation layer only sees method + path, /containers/*/logs is treated conservatively whether or not the caller also sets follow=1. That keeps backup/export, raw-stream, and deliberate registry-publish use cases possible without letting a casual wildcard rule silently include them.

Compose / BuildKit Transport

Modern docker build and docker compose build no longer speak the classic POST /build API by default — Buildx routes through BuildKit's session/gRPC tunnel (POST /session for the frontend/session bridge, POST /grpc for the moby.buildkit.v1.Control gRPC service, both tunneled over a hijacked HTTP/1.1 connection). As of issue #185's BuildKit gRPC mediation epic, Sockguard terminates both streams as h2c and mediates them at the gRPC-method level instead of hijacking the bytes opaquely: internal/buildkitproxy.Mediator runs an h2c server against the client and an h2c client against the daemon, classifies every RPC either endpoint exposes (Deny by default — only a curated, committed Mediate/Passthrough set is reachable at all), and for Mediate methods decodes the message, checks it against a request_body.buildkit policy, and forwards the client's ORIGINAL frame bytes verbatim on admission — never a re-encoded message. Supported transports, by preset:

TransportEndpoint(s)InspectableDefault posture
Classic builderPOST /buildYes — remote context, host network, RUN instructionsAllowed by an explicit POST /build rule; request_body.build denies remote context, host network, and RUN by default (see the *-with-build.yaml presets)
Mediated BuildKit session/gRPCPOST /session, POST /grpcYes — per-message policy on Control/Solve/Status and session Auth/Secrets/SSH/FileSync/FileSend/UploadAllowed when request_body.buildkit is configured (see the *-with-mediated-build.yaml presets)
Opaque BuildKit tunnel (deprecated)POST /session, POST /grpcNo — whole tunnel admitted with zero inspectionDenied; requires the deprecated insecure_accept_opaque_buildkit_tunnels: true
Native gRPC-over-h2cmoby.buildkit.v1.Control/*No — same control plane, bare HTTP/1.1 transportDenied unconditionally — sockguard's listener has no h2c support outside the two hijack-capable tunnel endpoints, so nothing can dial this path end-to-end

request_body.buildkit.control.solve gates the Solve RPC: security.insecure is always denied with no enabling knob, network.host requires request_body.build.allow_host_network, and cache import/export types, cache registries, exporter types, and exporter-push registries each have their own allowlist (empty means deny, the standard request_body.* convention). The frontend check pairs with request_body.build.allow_run_instructions: a Dockerfile synced over the session's FileSync stream gets the identical RUN-instruction hold-and-inspect scan the classic /build path already applies, unless it arrives via a genuinely remote context, which requires request_body.build.allow_remote_context for the same reason classic /build does. control.allow_status gates the Status RPC, admitted only for a ref this same trusted principal, selected profile, and BuildKit session actually Solved, so one tenant or simultaneous build cannot poll another's state. The principal comes from the verified mTLS certificate, Unix peer credentials, or normalized remote host instead of a source port or caller-controlled session header. Ref ownership and upload IDs are committed atomically only after every limit check passes. Session, ref, and upload identifiers are rejected before persistence when they exceed 256 bytes; one control tunnel can retain at most 256 distinct BuildKit session IDs; and an unconsumed upload grant expires after one hour while remaining valid across a normal control-tunnel close. Session auth, secrets, ssh, file_sync, file_send, and upload calls each stay denied until their own policy allows them, and a registry method classified for mediation but missing a dispatcher fails closed rather than becoming raw passthrough. See Configuration for the full field reference.

insecure_accept_opaque_buildkit_tunnels is now deprecated: it still works — existing configs that set it keep running unchanged — but setting it to true logs a startup warning steering operators toward request_body.buildkit, and the flag will be removed in a future major release. The flag and request_body.buildkit are mutually exclusive (mediation supersedes the wholesale acknowledgment), so a config cannot set both, and the deprecation warning only ever fires for a config using the flag on its own. If you see a build fail against a preset with a /session or /grpc denial, that means neither a classic-builder rule nor a request_body.buildkit policy is configured for that path: either switch the client to the classic builder (DOCKER_BUILDKIT=0, the *-with-build.yaml presets), configure request_body.buildkit for the fully-mediated path (the *-with-mediated-build.yaml presets), or — not recommended for new deployments — fall back to the deprecated acknowledgment. Tecnativa's GRPC=1 / SESSION=1 compat env vars still auto-set the deprecated acknowledgment with their own compat-specific warning; migrating off either warning means configuring request_body.buildkit instead. See Migration for the step-by-step move from the acknowledgment to a mediated policy.

Layer 6: Owner Label Isolation

When ownership.owner is set, Sockguard stamps label-capable creates and build-produced images with an owner label, injects owner filters into list/prune/events responses, and freshly inspects target resources on individual requests to deny cross-owner access. Ownership decisions are intentionally uncached because Docker names and image tags can be rebound to different resources. The checks cover owned containers, images, networks, volumes, services, tasks, secrets, configs, nodes, and swarm state, with service writes stamping both the service and its task template so downstream tasks inherit the same owner identity, /nodes using Docker's node.label filter key, and unlabeled node/swarm resources only claimable through their update paths.

Collection-action words are reserved only for the method and exact path that perform that action. A container, network, volume, service, secret, or config named create, prune, json, or another reused keyword still receives the normal owner check on inspect and trailing-action paths.

The same boundary applies to references embedded inside container and service payloads, not only the resource named by the URL. Container creates authorize the requested image, named volumes, custom network mode, every endpoint-config network, and every container:<ref> namespace-sharing target. Service creates and updates authorize the task image, named volumes, networks, secrets, and configs. A foreign label or unresolved workload dependency is denied before Docker sees the request; only an image that successfully resolves and is genuinely unlabeled can use allow_unowned_images. Create new volumes, networks, secrets, and configs through their owner-stamping endpoints before attaching them to a workload. This turns one shared Docker socket into N isolated identity views without leaving workload dependencies as an ownership side door.

Layer 7: Visibility-Controlled Reads

Sockguard's response filter applies to known protected Docker JSON response shapes on successful body-bearing 2xx responses across request methods, not only GET 200. If a protected successful response cannot be parsed or sanitized safely, Sockguard fails closed with a generic 502 instead of forwarding unsanitized data. Non-success responses, HEAD responses, no-body statuses, non-protected paths, and streaming endpoints (logs, attach, events) pass through unmodified — those are protected by request-side rules and the read-side exfiltration guardrail, not by response rewriting.

Together with request-side visibility and exfiltration guardrails, the read-side layer narrows what callers can see:

  • Inject label visibility selectors into GET /containers/json, GET /images/json, GET /networks, GET /volumes, and GET /events
  • Inject label visibility selectors into GET /services, GET /tasks, GET /secrets, GET /configs, and GET /nodes
  • Return 404 for hidden targets on inspect/log-style reads such as GET /containers/*/json, GET /images/*/json, GET /networks/*, GET /volumes/*, GET /exec/*/json, GET /services/*, GET /services/*/logs, GET /tasks/*, GET /tasks/*/logs, GET /secrets/*, GET /configs/*, GET /nodes/*, and GET /swarm
  • Fail startup unless raw archive/export and stream-style reads are explicitly acknowledged via insecure_allow_read_exfiltration: true
  • Redact Config.Env on GET /containers/*/json
  • Redact HostConfig.Binds host paths plus Mounts[*].Source on container list/inspect responses
  • Redact volume Mountpoint on GET /volumes and GET /volumes/*
  • Redact container and network address topology on container/network list and inspect responses
  • Redact service/task env, mount, secret/config-reference, and network metadata
  • Redact config payload data, plugin env/path metadata, node/swarm TLS material, swarm join/unlock material, and /info plus /system/df topology-sensitive fields

Single-resource inspect denials honor rollout mode: under a profile in warn or audit mode, a target that visibility policy would hide is forwarded upstream with a would_deny audit verdict instead of being hard-404'd, so visibility policy can be staged like every other deny gate. When response.name_patterns or response.image_patterns filter a list response, Sockguard buffers the upstream body under an 8 MiB cap and rejects a larger response with a 502 rather than buffering it unbounded.

Write-only collection words remain valid resource identifiers for GET and HEAD, so keyword-named networks, volumes, services, secrets, and configs do not bypass visibility. A single container or image check that combines labels with name/image patterns obtains both from one bounded inspect response. If a buffered visibility rewrite must generate its own 502, Sockguard removes stale upstream representation headers before writing the replacement JSON so clients do not see a mismatched length or encoding.

These controls are on by default where they are pure redaction because runtime env vars routinely carry credentials and Docker read APIs expose raw host mount paths plus internal network layout.

Layer 8: Structured Access And Audit Logging

Every request is stamped with a proxy-generated canonical X-Request-Id and logged with method, raw path, normalized_path, decision, matched rule index, selected client profile when present, latency, request ID, trace context, and client metadata. If the caller supplied its own request ID, Sockguard preserves it separately as client_request_id in logs instead of trusting it as the canonical correlation key.

path is the client-controlled URL path exactly as received and is retained for forensic replay. Detection logic, SIEM grouping, and policy analysis should use normalized_path, which is the canonical path after Sockguard strips Docker API version prefixes, decodes escaped separators, and resolves dot segments before rule evaluation.

When log.audit.enabled is true, Sockguard also emits a dedicated JSON audit event with a stable schema: request ID, client request ID, trace ID, trace parent/span IDs, sampled flag, raw and normalized path, decision, machine-readable reason_code, human-readable reason, matched rule, selected profile, flattened actor and transport identity fields, ownership context, and final HTTP status. Upstream reverse-proxy errors overwrite the audit reason code with bounded values such as upstream_socket_unreachable or upstream_response_rejected_by_policy, so the terminal result remains explicit even after an allow decision has already been made.

The audit ownership object is emitted on every event. If ownership.owner is configured, that owner identifier is repeated in every audit record, not only resource ownership decisions, so it should be a non-secret tenant/workload label suitable for the audit sink.

Sockguard preserves valid W3C traceparent trace IDs and sampled flags, forwards a proxy-local span ID, and includes trace_id, trace_parent_id, trace_span_id, and trace_sampled in access, audit, and upstream reverse-proxy error logs. Invalid or absent trace context starts a fresh local trace without enabling any OTLP span exporter.

When health.watchdog.enabled is true, Sockguard actively probes the upstream Docker socket, logs reachable/unreachable state transitions, and lets /health reflect the latest watchdog state. When metrics.enabled is true, Sockguard serves Prometheus text metrics from /metrics by default, including a sockguard_build_info{version,commit,build_date,go_version} gauge, a sockguard_start_time_seconds gauge, and watchdog state and check counters if the watchdog is enabled. The scrape endpoint is local to Sockguard, is never forwarded to Docker, bypasses Docker API allow rules like /health, and remains behind listener security plus client ACLs.

Dangerous Docker API Endpoints

Risk LevelEndpoints
CriticalPOST /containers/create, POST /containers/{id}/exec, POST /exec/{id}/start, PUT /containers/{id}/archive
HighPOST /images/create, POST /images/load, POST /build, POST /libpod/build, POST /services/create, POST /services/{id}/update, POST /swarm/init, POST /swarm/join, POST /swarm/update, POST /swarm/unlock, POST /nodes/{id}/update, POST /plugins/pull, POST /plugins/{name}/upgrade, POST /plugins/{name}/set, POST /plugins/create
MediumPOST /containers/{id}/update, POST /volumes/create, POST /networks/create, POST /networks/{id}/connect, POST /networks/{id}/disconnect, POST /secrets/create, POST /configs/create, DELETE /containers/{id}
LowGET /containers/json, GET /events, GET /version, GET /_ping

Image Security

Sockguard's container image is built on Wolfi (Chainguard):

  • Minimal package set, which keeps the base image's CVE exposure low
  • Built-in SBOM output and build provenance when release visibility supports attestations
  • Cosign-signed for verification — see the image verification guide for the canonical cosign verify invocation
  • No shell, no package manager in production image

Runtime Hardening

Sockguard runs as UID 65532 (Chainguard nonroot) inside the container. On stock Docker hosts where /var/run/docker.sock is owned by the docker group you may need a group_add override with the socket's numeric group ID or a matching user: / supplemental group. For a Docker socket proxy, the real security frontier is what the daemon will accept through the proxy, not the UID the proxy process reports after it has already opened the upstream socket.

The runtime controls that matter are:

  • Correct policy rules and request-body inspection
  • read_only: true
  • cap_drop: [ALL]
  • security_opt: ["no-new-privileges:true"]
  • Docker's default seccomp profile or a stricter custom profile
  • AppArmor/SELinux confinement on the host
  • Rootless dockerd on the host when available

The getting-started examples use the container-level controls above by default so the drop-in path stays simple without hiding the real hardening story.

Known Limitations

These are architectural constraints inherent to Sockguard's position in the stack. They are documented here for honest operator awareness rather than as open bugs.

IP-based client identity is soft isolation. When clients.container_labels.enabled is true, or any clients.profiles[*].match rule keys on source_cidrs, Sockguard resolves the calling container by source IP through the Docker API. This is soft isolation — adequate against configuration drift and friendly-fire mistakes, but not a hard boundary against an attacker who can influence which container a given bridge IP points at:

  • A container restart can race the label lookup: if a new container acquires the same bridge IP before the lookup completes, the lookup may return the new container's labels rather than the previous container's.
  • An attacker who can create containers on the same user-defined bridge can, in principle, claim a privileged IP and inherit its policy until the next legitimate container takes it back.
  • Host-network containers (network_mode: host) all share the host IP, so IP-keyed allowlists cannot tell them apart.

For workloads where caller identity is part of the security boundary, listen on a unix socket and use clients.unix_peer_profiles with uids/gids. SO_PEERCRED is supplied by the kernel and cannot be spoofed from within the calling container — that is the hard-isolation path.

Exec TOCTOU (inspect/start split). Docker exposes exec metadata inspection and exec start as separate API calls. Sockguard re-checks POST /exec/*/start against Docker's stored exec metadata before the command runs, but the gap between the create and start calls is an unavoidable time-of-check/time-of-use window inherent to Docker's API shape. Keep exec allowlists narrow and client profile assignments conservative for clients that do not need interactive command execution.

Exec Env allowlisting is create-time only. request_body.exec.allowed_env_vars, denied_env_vars, and allowed_env_values are enforced on POST /containers/*/exec, not on the POST /exec/*/start re-check. This isn't a gap in the re-check logic — GET /exec/{id}/json (the metadata the start-time re-check reads) never exposes the original Env, and exec instances are immutable once created, so there is no later point at which the environment could change. Once an exec is created with an allowed Env, that environment is fixed for the life of the exec instance.

Hijacked-stream redaction limits. Sockguard's response filter — including the response.redact_container_env, response.redact_mount_paths, response.redact_network_topology, and response.redact_sensitive_data toggles — operates only on structured JSON responses with known shapes. Raw streaming endpoints — GET /containers/*/logs, POST /containers/*/attach, GET /services/*/logs, GET /events, exec attach, and image-build progress output — switch the connection to a raw byte stream (or a non-JSON chunked stream) after the initial HTTP response, at which point Sockguard cannot inspect or redact the byte stream. A secret an application writes to its own stdout will reach a caller that has been allowed to attach.

These paths are gated at request time via per-profile rule allowlists and the insecure_allow_read_exfiltration guardrail (which keeps the streaming read endpoints denied by default), but there is no post-admission byte-level filtering of the stream content. Restrict these paths in your rules to only the profiles and callers that genuinely need them, and treat the redaction toggles as a guarantee for Docker's structured metadata only — not for arbitrary workload output.