Homelab Weekly: Registry Cleanup, Edge Proxying, and Safer Deployments

Sep 28, 2026 min read

This week was a cleanup and hardening week across the homelab rather than a week for a single new platform service. The GitOps repository received six commits across eighty files and all four environment groupings: dev, dmz, infra, and prd. The immediate work was to correct container-image sources after the retirement of the old Nexus registry, repair upstream references that no longer resolve on Docker Hub, and make the DMZ Traefik path understand PROXY protocol v2 from the VPN tunnel. Those changes are closely related in practice. A service that cannot pull an image cannot reconcile, and a service that can reconcile but receives traffic without trustworthy client-connection metadata is harder to operate and secure correctly.

DurpDeploy had the larger application-development stream. I landed SQLite defaults for bare database paths and addressed database locking, completed another phase of interpreter verification coverage, added versioned project runbooks, moved CI jobs from GitLab CI to GitHub Actions, and added an end-to-end verification gate with a streaming scrub check. The same week included dependency maintenance and several security-focused changes around secret masking and token presentation. The theme was making ordinary deployment behavior safer and more explainable, not adding abstraction for its own sake.

The work is best understood as maintenance of contracts. GitOps declares where images come from and where workloads belong. The edge proxy defines what transport metadata it trusts. DurpDeploy defines how it stores state, executes a command, shows a new credential, and returns API data that might otherwise disclose a secret. Each contract needs to survive the uninteresting but frequent cases: an upstream image moves, a registry is retired, a database is opened through a shorthand path, a proxy sits behind another proxy, a command runs for long enough to need observability, or a caller reads a variable through a normal API route.

Completing the Registry Transition

The most broadly visible GitOps change was the migration of registry.durp.info image references to Harbor after Nexus was retired. Container image references are deceptively small pieces of configuration. They often appear as one string in a Helm values file or Kustomize patch, but that string binds a reconciler to an availability boundary, an authentication path, an image lifecycle, and a retention policy. When the registry behind those strings changes, the correct response is to update the desired state everywhere it is declared. Leaving individual workloads on a retired endpoint and repairing them later is not a migration plan, it is a sequence of avoidable outages.

I changed the repository declarations instead of relying on image caches or a manual registry redirect. That preserves the normal GitOps model. Argo CD can render the desired image reference, Kubernetes can pull from the intended registry, and a future node or replica starts from the same configuration as the currently running workload. A cached image can hide an invalid reference until a node drain, reschedule, or version update forces a fresh pull. That kind of delayed failure is exactly why registry ownership needs to be visible in the repository.

Harbor is the replacement registry in this path, so the GitOps change also makes that dependency explicit. The registry is no longer an incidental implementation detail of a currently healthy pod. It is part of the platform that must be available to build, deploy, scale, and recover services. The specific applications touched span the platform’s operating surface: .gitlab, Argo CD, Authentik, Bitwarden, cert-manager, CrowdSec, external-dns, external-secrets, GitLab Runner, InternalProxy, LittleLink, Nebula Sync, OpenSpeedTest, Redlib, Searxng, and Traefik. That breadth is why I treated the transition as a repository-wide reconciliation change rather than a one-service fix.

There are a few useful implications here. Registry credentials, when required, should flow through the existing secret-management path rather than being copied into each application value. Image tags need to be intentional and compatible with the deployment’s update policy. Registry availability needs to be considered alongside controller behavior, because a successful Argo CD sync only means manifests were applied, not necessarily that every newly scheduled pod pulled and became ready. Finally, rollback remains simple when the image reference is versioned in Git: the known-good declaration is reviewable and can be restored through the same reconciliation path.

The main operational goal was to remove an obsolete dependency without creating hidden ones. No application should depend on a retired Nexus endpoint merely because it happened to have started before the change. Harbor-backed references put the target state in the open, where the next deployment and the next incident use the same facts.

Repairing Upstream Image References

The registry migration surfaced a related class of issue: several upstreams referenced in the manifests no longer exist on Docker Hub. The corrective commit fixed those image sources rather than trying to make Docker Hub compatibility an infrastructure responsibility. That distinction matters. A mirror, proxy, or local cache can make a valid upstream more reliable, but it cannot make an abandoned or renamed upstream a sound dependency.

Container registries are mutable ecosystems. Maintainers change publishing organizations, images move from Docker Hub to a vendor registry, old repositories disappear, and tags can be rewritten or removed. A declarative deployment repository gives these failures a crisp symptom: a pod enters an image-pull failure state. The real fix, though, is to verify the source image and update the declaration to a maintained upstream or an explicit internal copy. Retrying the pod or waiting for the cache is only treating the symptom.

I corrected the references at the GitOps layer, which is the shared source for every affected reconcile path. This is smaller and safer than putting special cases around individual applications. A workload in dev and its corresponding production deployment should not have different ad hoc recovery logic simply because a public image moved. The image declaration is the root cause, so it belongs in the desired state that all environments consume.

This also reinforces why image provenance is an engineering concern, not a cosmetic naming detail. An official upstream or documented vendor registry gives a clearer maintenance relationship than an unverified replacement with a convenient name. It provides a better basis for tracking releases, scanning images, evaluating advisories, and deciding whether a breaking change is likely. When a referenced image is no longer present on Docker Hub, the useful question is where the upstream now publishes, not which arbitrary repository happens to produce a similarly named image.

The corrected manifests cover services at different levels of exposure. Traefik and CrowdSec are edge-adjacent. Argo CD, cert-manager, external-dns, external-secrets, and GitLab Runner support the control plane and delivery workflow. Authentik, Bitwarden, LittleLink, Redlib, Searxng, OpenSpeedTest, Nebula Sync, and InternalProxy serve more focused application or integration roles. Their different purposes do not change the basic requirement: Kubernetes must be able to pull the exact image its declaration names, including after a node replacement or a clean cluster rebuild.

DMZ Traefik and PROXY Protocol v2

The other major GitOps feature was configuring DMZ Traefik to trust PROXY protocol v2 from the VPN tunnel. This is a targeted networking change, but it affects how the edge interprets connection information. A proxy protocol header carries metadata about the original connection, commonly including source and destination addresses, from a trusted upstream proxy to the next hop. Without it, Traefik sees the address of the proxy that connected to it rather than the client address that originally initiated the request.

That loss of client identity has real operational consequences. Access logs become less useful because every request appears to come from the tunnel or intermediary. Rate limiting and security policy may act on the wrong source. An application or middleware that relies on network identity has less accurate information. During incident investigation, a chain of reverse proxies that only reports the previous hop makes it unnecessarily difficult to understand what happened.

PROXY protocol v2 is a binary protocol designed for this transport-level handoff. It is distinct from HTTP headers such as X-Forwarded-For. That distinction is important because a TCP-level listener may receive connections before there is an HTTP request, and because a trusted proxy protocol path has different configuration and trust requirements than forwarded HTTP headers. Configuring the appropriate Traefik entry point to accept the protocol tells it how to parse the connection metadata arriving from the known tunnel source.

The safety property is trust scoping. I did not make Traefik accept claimed client metadata from arbitrary sources. A proxy protocol listener must only trust an upstream that is actually controlled and known to emit correct headers. If arbitrary internet clients can reach such a listener and inject PROXY metadata, they can potentially spoof the addresses used in logs or policy. The VPN tunnel is the bounded upstream in this design. The configuration is therefore an assertion about a specific network path, not a blanket instruction to trust whatever a connection says about itself.

This change belongs in the DMZ configuration because that environment owns the admission boundary. The tunnel hands off traffic to the DMZ edge, and Traefik decides how the connection should be interpreted and routed. Workloads behind Traefik do not need to solve this individually. They should receive consistent request metadata through the existing proxy chain. Centralizing the trust decision at the edge avoids duplicated, divergent application settings and makes the accepted topology easy to review.

I will validate this in the same way I validate other edge changes: confirm that the entry point accepts intended traffic, inspect Traefik access logs for the expected original source information, and ensure non-tunnel traffic does not accidentally use a PROXY-protocol listener. The feature is useful because it makes client identity visible again, but only if that visibility comes from a path that is actually trustworthy.

Reconciliation Across Four Environments

This week’s GitOps work touched dev, dmz, infra, and prd. The shared image and platform adjustments should not erase the meaning of those environments. They are boundaries for lifecycle, exposure, and operational responsibility, not just duplicate directories containing slightly different values.

infra is where foundational controllers and services naturally live. Argo CD, certificate issuance, secret delivery, DNS integration, and the systems that support ingress and storage have broad effects beyond a single application. Their values require conservative review because an error can prevent other workloads from reconciling or becoming reachable. Keeping their configuration declared provides a better recovery path than relying on shell history or a manual change made during an incident.

The DMZ is the place to be explicit about incoming traffic. Traefik, CrowdSec, external DNS behavior, certificates, and the VPN tunnel integration are collectively an edge system. Publishing a service should be a deliberate route through this boundary, with exposure and transport assumptions recorded in the configuration. A workload being internally reachable is not a reason to make it generally reachable, and an ingress annotation is not a substitute for thinking about where traffic originates and what metadata it carries.

dev provides a lower-consequence place to exercise changes, while prd carries the services that need the most conservative operation. This does not require pretending a homelab has the same scale as a large commercial fleet. It does require preserving the distinction in configuration and rollout. A valid test image reference, registry credential, or proxy setting should be proven where the impact is contained before it becomes a dependency of a production service.

Argo CD is the mechanism that gives these declarations operational force. It observes the repository and reconciles desired application state into its target cluster and namespace. The controller cannot tell whether an image source is maintained, whether a proxy trust range is sensible, or whether an environment-specific override has a sound reason. Those decisions remain in the repository and code review. What Argo CD provides is consistent application of the reviewed answer and visible drift when live state no longer matches it.

The goal is consistency of model, not artificial identity of configuration. I want common conventions for image references, secret references, labels, sync behavior, and chart layout. I do not want to flatten real differences in exposure, persistence, resource needs, or risk. This week’s broad change was successful only if every environment continues to converge on its own intentional desired state.

SQLite Defaults and Locking in DurpDeploy

DurpDeploy received a practical persistence fix around bare SQLite database paths. SQLite is often chosen because it is embedded, portable, and operationally small. Its simplicity is a benefit when the application treats it honestly. A path such as durpdeploy.db is not the same thing as a fully qualified connection string with every behavior already defined. Defaults for those bare paths need to be applied consistently so the resulting database location and connection behavior are predictable.

I applied the SQLite defaults to bare database paths and documented them. The documentation is part of the change because path interpretation is operational configuration. An operator who supplies a relative or bare path needs to know where the server will create the file, how that differs from an explicit path, and what persistence volume should contain it in a container deployment. Without that contract, a service can appear healthy while writing state into an ephemeral working directory or an unexpected filesystem location.

The related database-locking work addresses the less obvious side of SQLite’s small footprint. SQLite is reliable, but it is a single-file database with a concurrency model that needs respect. Concurrent readers and writers are ordinary for a web service. Long transactions, connections held while streaming output, or code paths that retain a database handle during slow non-database work can turn normal workload contention into avoidable database is locked failures.

I addressed the root operational behavior rather than asking callers to retry indefinitely. The change applying defaults at the bare-path boundary ensures every equivalent configuration uses the same sensible setup. The earlier work preventing slow log exports from pinning database connections follows the same rule. Streaming a potentially slow export does not require holding scarce database resources for the entire client transfer. Read the required data through the appropriate database scope, then stream the response without keeping the connection unnecessarily occupied.

This matters for a deployment system because activity arrives in bursts. Runs produce events, users read status, logs are exported, schedules update, and background workers persist results. A database implementation that behaves correctly only under a single local user is not adequate simply because it is lightweight. The objective is not to make SQLite imitate a remote database server. It is to avoid needless contention and make its documented behavior a reliable part of DurpDeploy’s supported deployment modes.

Execution Contracts and Interpreter Verification

The interpreter verification work reached phase four this week. Deployment commands are often presented as strings, which makes it tempting to let the executor assume a default shell and hope every target shares the same semantics. That works only as long as the command syntax, available interpreter, environment, quoting rules, and host assumptions all happen to line up. In a system designed to run deployment steps across real targets, those assumptions are hidden configuration.

An interpreter is part of the execution contract. If a step is intended for a particular command processor or runtime, that choice must survive creation, validation, persistence, retrieval, dispatch, and execution. Testing only the UI field or only the final subprocess invocation is not enough. The meaningful coverage traces the value across the boundary where it could otherwise be lost, substituted, or interpreted differently.

The phase-four coverage advances that verification rather than merely increasing test count. The point is to prove that supported interpreter selections arrive at the execution path intact, that invalid or unavailable choices fail clearly, and that the server does not silently fall back to an unintended shell. A deployment system should not turn a configuration error into a successful run under a different interpreter. That result is harder to detect and potentially more dangerous than an explicit failure.

This is also a security boundary. An interpreter setting should be treated as constrained metadata, not as a free-form way to run arbitrary binaries. The supported values and argument behavior need to be intentional. Once the platform accepts a deployment definition, operators should be able to understand what runtime it requests and what environment will execute it. That keeps review useful and reduces the chance that a syntactically valid command does something different on a remote agent than it did in a developer’s terminal.

The same discipline applies to output. The new streaming scrub check in the verification gate is aimed at a class of failure that is easy to miss when testing only final responses. Long-running operations can emit data incrementally. If masking is applied only after an entire output buffer is assembled, a secret can escape in an early stream chunk even though the final stored result looks clean. Verification needs to exercise the actual streaming path, where time and buffering behavior matter.

Protecting Tokens and Variables

Several DurpDeploy changes focused on secret exposure. I masked secret variable values in ordinary API reads, then extended masking across variable API responses and release API coverage, including role-based access and insecure direct object reference cases. I also moved the display of a newly created token out of the redirect URL and updated nanoid to address a high-severity npm audit finding.

These changes share one simple rule: sensitive values should appear only where the product deliberately needs to reveal them, and only for as long as that disclosure is necessary. It is not enough to protect a special endpoint while ordinary list, detail, or release responses serialize the underlying value by default. The shared API serialization boundary is where masking belongs, because every caller that reads the resource should receive the safe representation unless an explicit, separately authorized reveal flow exists.

Role coverage and IDOR coverage are especially important in a multi-user deployment tool. Authentication proves that someone has a session. Authorization determines which deployment, release, or variable that person may access. An IDOR bug occurs when an otherwise authenticated caller can alter an identifier and read another resource they were not authorized to view. Secret masking reduces the impact of an accidental read, but it does not replace correct object-level authorization. Both controls are needed.

Token creation has a different lifecycle. A token’s plaintext must be shown when it is created because the user needs to save it. After that moment, the server should store only what it needs for verification and identification, not a retrievable plaintext copy. Putting the new token in a redirect URL was an unsafe presentation path. URLs have a habit of being retained in browser history, reverse-proxy logs, analytics, referrer headers, and copied support messages. Displaying the value in a dedicated response or page state limits those accidental persistence routes.

The nanoid update is ordinary dependency hygiene, but it belongs in the same release stream. A known high-severity finding is not improved by good API masking elsewhere. Keeping dependencies current is a basic part of maintaining the application’s trust boundary. The Go dependency updates for WebAuthn, OpenAPI runtime, Goose, SQLite, pgx, Microsoft SQL Server support, Chi, OIDC, and x/crypto similarly keep the dependency graph aligned with maintained upstream releases. Each update still needs normal verification, but postponing them indefinitely is not a security strategy.

CI, End-to-End Verification, and Runbooks

I migrated DurpDeploy CI jobs from GitLab CI to GitHub Actions and added a CI end-to-end job that acts as a verification gate. CI migrations can look like administrative work, but the pipeline is part of the release contract. A job that moves platforms must preserve the checks that matter, use the required credentials and service setup safely, and fail the change when the application no longer satisfies its actual integration behavior.

The end-to-end job adds value where unit tests reach their limit. Unit coverage can validate a handler, a model method, or a scrubber in isolation. It cannot always demonstrate that a request reaches the application, persistence is configured as expected, authentication and authorization paths behave coherently, streamed output is scrubbed at the real boundary, and the response the client sees matches the intended contract. The verify gate makes that broader check a required part of the delivery path instead of a best-effort follow-up.

A gate should be narrow enough to be trusted. If it flakes, depends on undocumented network state, or attempts to validate every possible deployment target, teams learn to work around it. The useful version creates a controlled application instance, exercises a representative flow, and reports a clear failure when the contract breaks. The streaming scrub check fits this model because it verifies a concrete security invariant that a final-response assertion could miss.

I also added versioned project runbooks. Runbooks are not a substitute for a well-designed system, but they are useful when they capture the supported operational procedure for a specific release line. A versioned runbook can say which migrations, configuration keys, upgrade order, validation commands, and recovery expectations apply to that version. That is more durable than an unversioned document that quietly accumulates instructions for incompatible releases.

The right operational standard is that a release should be reproducible by someone who did not write it. CI proves a bounded set of software behavior before merge. Runbooks preserve the deployment and recovery knowledge after merge. Together they reduce the amount of undocumented personal context required to operate the system.

Looking Ahead

The next GitOps follow-up is validation, not another broad migration. I will watch Argo CD reconciliation after the Harbor and upstream-image changes, especially fresh pulls on newly scheduled workloads. The check is not just whether existing pods remain healthy from cached layers. It is whether each declared image can be resolved and pulled through its intended registry path, with the correct secrets and tags, across the environments where it is deployed.

For the DMZ proxy update, I will verify client-address visibility in Traefik logs and confirm that only the VPN tunnel path is trusted to send PROXY protocol v2 metadata. The desired result is better observability and correctly scoped policy inputs, not a listener that accepts spoofed connection identity from arbitrary sources.

For DurpDeploy, I will continue exercising the persistence and execution boundaries under realistic conditions: bare SQLite paths in local and containerized deployments, concurrent activity and slow exports, valid and invalid interpreter selections, streamed output containing values that must be scrubbed, and normal API reads across roles and object identifiers. The release path should keep rejecting unsafe behavior early and explain failures well enough to fix their cause.

The common direction is deliberately boring. Desired state stays in Git, controllers reconcile it, registry references name maintained sources, the edge trusts only known intermediaries, and the deployment application does not leak credentials or depend on accidental defaults. That is the kind of maintenance that makes future changes easier to reason about, which is a better outcome than a week of features that only look impressive in a changelog.