Homelab Weekly: Safer Deployments and a Publishable Agent

Oct 5, 2026 min read

This week I concentrated most of my development time on DurpDeploy. The work moved the deployment system from a collection of capable execution paths toward a more deliberate operational control plane: deployments now have stronger request boundaries, more useful lifecycle controls, serialized access to shared environments, explicit approval points, verification after a rollout, and a confirmed route back when verification fails. Alongside that work, I finished the agent container publishing path and removed a CI service that no longer belongs in the build.

The common theme was reducing ambiguity at the points where automation has consequences. A deployment service is not useful merely because it can start a command. It has to establish which project a caller can see, decide whether a target is currently safe to change, preserve enough state to explain what happened, and avoid treating a successful process exit as proof that a service is healthy. Most of the changes this week were small in isolation, but they close gaps between those responsibilities.

I also made a number of ongoing dotfiles updates, including model and agent configuration changes and an Alpine update. Those commits are intentionally more frequent and less narratively useful than application changes, but they keep the workstation and automation environment aligned with the tools I am actually using.

Treating deployment APIs as a trust boundary

I added strict, bounded JSON decoding for DurpDeploy API request bodies. This is not glamorous feature work, but it is the right place to start when a service accepts deployment instructions from clients. Request decoding is a trust boundary. If it accepts arbitrarily large bodies, ignores unknown fields, or proceeds despite trailing data, the application has already given up useful control before its business rules run.

The updated handling limits the body size before decoding. That gives the server a defined upper bound on the memory and parsing work attached to a request, rather than allowing a client to choose it. The decoder also rejects fields that the API does not recognize. That is important for a deployment API because a misspelled option should not quietly become a defaulted option. It also prevents clients from believing an unsupported feature was applied when it was actually discarded. Finally, decoding checks that the request contains one JSON value rather than a valid object followed by extra garbage or a second object.

The immediate result is clearer client failures, but the more important outcome is contract discipline. The service can distinguish malformed syntax, an unexpected schema, an oversized request, and a valid request that fails domain validation. Those are different operational failures and should not all become a generic server error. This also makes API evolution safer: adding a field is an intentional compatibility change, not an accidental side effect of a permissive decoder.

I paired that boundary work with authorization scoping for deployment lists. Listing is easy to underestimate because it looks read-only. In a multi-project deployment system, deployment records reveal environment names, release history, timing, status, and often the shape of the delivery process. A caller must only receive entries for projects they are authorized to access. The list query is now scoped to the caller’s authorized projects rather than filtering a broad result after the fact. Applying authorization at the data access boundary reduces the chance that a future pagination, sorting, export, or alternate endpoint reintroduces an information leak.

This is the same general design rule I use elsewhere in homelab services: enforce constraints as close as practical to the resource. UI hiding is useful for usability, but it is not authorization. Handler-level filtering is better, but it is easy to bypass through another caller. Making the scoped query the normal route means most callers cannot forget the rule.

Making the deployment runner usable for real local steps

A substantial piece of this week’s work was enabling server-side deployment steps to run through local container sockets. DurpDeploy needs to support steps that operate inside containerized tooling without pretending every operation has to be hosted by an external agent. A local runner can use a container runtime socket to create and supervise the workload where the server runs, while still keeping the deployment record and lifecycle in the central system.

The point is not to make the server a general-purpose shell endpoint. The step definition remains explicit, and execution is still part of a deployment’s tracked workflow. Using the local socket adds a practical execution option for tasks such as image-oriented tooling, controlled build or packaging steps, and operations that already expect a container boundary. It avoids requiring a separate remote agent simply because the work is local to the deployment service.

I also added files passed between local deployment steps. Deployment pipelines regularly need to carry an artifact, rendered configuration, manifest, report, or generated package from one step to the next. Requiring every small handoff to go through an external store makes simple workflows awkward and adds cleanup problems. The new capability provides an explicit way for local steps in the same deployment flow to exchange files.

That does not mean treating the server filesystem as an unbounded shared drive. Files are tied to the lifecycle of the deployment rather than silently becoming durable global state. The important distinction is between a deployment artifact with an owner and a scratch file whose retention nobody understands. Keeping the handoff within the workflow gives later verification, diagnostics, and cleanup code a defined object to reason about.

I expanded the package model with generic ZIP packages for deployments and runbooks. ZIP is a deliberately boring interchange format, which is useful here. It can package a structured set of files without forcing every consumer into one language ecosystem or one container format. A deployment package can include scripts, manifests, templates, supporting files, and metadata. A runbook package can carry operational material alongside the task it documents. The system can then make a package available to the controlled execution path instead of relying on an operator to reconstruct a working directory from scattered links.

There are security implications whenever an application accepts archives. Packaging is not an excuse to trust filenames or filesystem paths supplied by users. Archive extraction must remain bounded and must prevent entries from escaping the intended destination through absolute paths or traversal sequences. The operational design should also keep package provenance visible: which release supplied the package, which deployment consumed it, and which resulting files were exposed to the step. The value of packages comes from repeatability, not merely convenience.

Controlling concurrency with a durable environment queue

The most consequential reliability change this week was serializing deployments through a durable environment queue. An environment is shared mutable state. Two separate releases targeting the same environment at the same time can each be valid in isolation and still produce an invalid combined outcome. Database migrations, shared configuration, rolling restarts, and cleanup jobs all create ordering requirements that a generic worker pool cannot infer.

The queue makes the environment the unit of serialization. A deployment for an occupied environment waits rather than racing the current one. Persisting that queue matters. An in-memory mutex would be shorter code, but it would disappear during a process restart, would not coordinate multiple server instances, and would leave no durable explanation for why a deployment is waiting. A durable record lets the system restore the ordering decision, display it, audit it, and make recovery behavior intentional.

I also made schedule writes coordinate with release deletion. Recurring schedules and release deletion are a classic race: one path removes a release while another is calculating or saving work that refers to it. Fixing only the deletion endpoint would leave the scheduler able to recreate or retain a stale reference. The write path now participates in the same serialization rule, so the system chooses a deterministic winner rather than relying on timing.

The release deletion work itself became a complete lifecycle instead of a database delete. I added safe project release deletion, deletion confirmation, history cleanup after that confirmation, and an audit ID for the deletion operation. The service returns not found when deletion wins a race with deployment creation. That response is more truthful than creating work for an object that has already crossed the deletion boundary. I also isolated deletion end-to-end coverage from recurring schedules, so a test of deletion behavior is not accidentally dependent on unrelated background timing.

A small but meaningful follow-up handles refresh of a deleted release cleanly and gives clearer deletion guidance. Stale client state is unavoidable in a web application. A user can leave a page open, another action can delete the release, and the original page can try to refresh it later. The correct response is to recognize that state, stop presenting it as recoverable data, and guide the user toward the remaining valid choices.

The same need for predictability drove a fix to variable environment scopes for lifecycle projects. Variables must resolve against the correct environment context when a project changes phase. If scope resolution is wrong, a deployment can receive configuration that was intended for another lifecycle environment, which is a quiet and dangerous form of misconfiguration. The fix keeps variable selection tied to the effective lifecycle target rather than an incidental project-level context.

Approval, verification, and rollback as one delivery contract

I added approval gates for generated deployment artifacts. A deployment artifact is where a desired change becomes a concrete set of executable inputs, and that makes it the right approval boundary. Approving a high-level request and then generating new mutable output later creates a gap between what was reviewed and what actually runs. The gate is associated with the generated artifact so the approver is considering the resolved result.

This supports a practical separation of duties without turning routine development into ceremony. Not every environment needs a human checkpoint, but protected environments and consequential changes benefit from one. The implementation can distinguish an artifact that is prepared from one that is authorized to execute. That state distinction also helps operators understand why an otherwise ready deployment is not moving: it is waiting for an explicit decision, not stuck in an ambiguous pending state.

Approval is only one side of a controlled rollout. I added post-deployment verification and confirmed rollback. Previously, the system could know that execution completed, but process completion is not the same as service correctness. A deploy command can exit zero while the new workload fails readiness, a dependency becomes unavailable, a migration makes the service unhealthy, or traffic reaches a broken route.

The new verification stage gives a deployment a defined opportunity to prove its intended outcome. The exact verification command or probe will vary by service, but the workflow concept is stable: deploy, observe a result that matters, then mark success. This turns health checking from an optional operator habit into part of the delivery contract. It also improves the record kept by DurpDeploy. A completed deployment can now say whether it was merely applied or was applied and verified.

When verification fails, the system supports a confirmed rollback rather than an automatic destructive reaction. Rollback is not universally safe. A schema migration may not be reversible, a side effect may have been emitted, or the current state may need inspection before another change is applied. Requiring confirmation acknowledges that reality. It preserves a fast route to recovery while keeping a human in the loop for the decision that can change the failure mode from outage to data loss.

I connected that workflow to agent drain mode and fleet health alerts. A worker should not accept new deployments when it is being maintained, upgraded, or removed from service. Drain mode lets it finish or safely release existing work while preventing new assignments. Fleet health alerts provide the complementary control-plane signal: if capacity is degraded, operators need to know before a queue becomes a surprise. Together, these changes make agent availability explicit rather than an assumption embedded in job dispatch.

I also hardened HTTP timeouts while preserving long-lived streams. Timeouts must fit the semantics of each connection. A short timeout is correct for a normal request that should fail quickly, but applying it blindly to streaming logs, event feeds, or other intentionally long-lived connections turns healthy activity into an arbitrary failure. The updated behavior keeps bounded connection and request handling while allowing streams to remain open under their own expected lifecycle. This is one of those transport details that becomes very visible during incident response, when the logs needed to diagnose a deployment are the first thing that should not be cut off.

Shipping the agent image from the source of truth

In durpdeploy-agent, I merged the container publishing work and removed the Snyk CI workflow. The publishing pipeline now builds and publishes agent containers from main and release tags. Publishing from the primary branch makes the current integration artifact available for environments that intentionally track ongoing development. Publishing from release tags provides the immutable versioned path needed by environments that should move only through an explicit release.

Those two publication points serve different consumers and should not be conflated. A moving main tag is useful for fast iteration and development validation, while a release tag is the reference that can be recorded in a deployment definition and reproduced later. The build configuration now makes this distinction part of CI rather than depending on a local publish workflow or a manually remembered command.

Removing the Snyk workflow was deliberate cleanup, not a claim that dependency and image security no longer matter. A CI integration that is not delivering actionable, maintained results is operational noise. It consumes build time, adds another credential and vendor dependency, and can train maintainers to ignore alerts. Security checks need ownership and a useful remediation path. If a scanner is not part of that process, removing a stale job is better than presenting a green badge that does not meaningfully protect the deployment path.

The agent’s container publishing work also fits the broader DurpDeploy changes. The server can now model queueing, approvals, verification, rollback, and worker drain state more clearly, while the agent has a predictable distribution path. Control-plane behavior and worker packaging are separate concerns, but both must be reliable before a deployment platform earns trust.

Configuration maintenance and engineering hygiene

The dotfiles repository had a steady series of updates this week. I updated model-related configuration, agent configuration, Alpine-related settings, and other environment details. These are small commits because configuration changes are easier to inspect and revert when they are narrow. A broad workstation update with unrelated model, shell, tool, and operating-system changes is hard to reason about when one component regresses.

I treat dotfiles as production configuration for my own engineering environment. Agent settings influence how tools execute and authenticate. Model settings affect repeatability and resource usage. Alpine changes affect the base assumptions of lightweight systems. Keeping those changes in version control provides a record of what changed, but the more practical benefit is that rebuilding or synchronizing a machine is not a forensic exercise.

The week’s work also reinforced a recurring principle: the smallest feature is not always the smallest safe change. A request-body limit, a scoped query, a durable queue, or a deletion audit record can look like extra machinery compared with a direct happy-path implementation. In practice, each removes an entire category of cleanup, explanation, or recovery work later. The right level of simplicity is the one that leaves the service unsurprising when requests overlap, clients are stale, a worker is drained, or a rollout succeeds technically but fails operationally.

Looking Ahead

Next I will use the new deployment lifecycle pieces together in real workflow definitions. The key validation is not whether each state transition works independently, but whether a queued deployment moves cleanly through artifact generation, approval, execution, verification, and, when necessary, a confirmed rollback while preserving an understandable audit trail. I also want to keep tightening the agent and container release path so the image selected by a deployment is explicit, reproducible, and easy to trace back to source.

On the operational side, I will watch the new queueing and drain behavior under normal maintenance conditions. The success criterion is not maximum parallelism. It is predictable throughput, no overlapping changes to the same environment, and clear visibility when capacity or an approval gate is the limiting factor. That is the direction for DurpDeploy: fewer implicit assumptions, better records of intent and outcome, and deployment automation that is useful during both routine releases and the inconvenient cases that expose design shortcuts.