The deploy plan

dippy figures out what systems should be deployed and how via a deploy plan. The plan is a simple structure defined in Nix that is then evaluated by a function provided by dippy. That function produces all the necessary information to then perform the deploy.

t = import "${sources.korora}/types.nix";

planType = t.struct "plan" {
  meta = t.struct "planMeta" {
    nixpkgs = t.attrsOf t.str;
    specialArgs = t.attrs;
  };
  defaults = nixosModuleType;
  hosts = t.attrsOf nixosModuleType;
  phases =
    t.listOf
    <| t.struct "phase" {
      name = t.str;
      includeIf = t.optionalAttr t.function;
    };
};

nixosModuleType =
  t.union [
    t.pathLike
    t.attrs
    t.function
  ]
  |> t.rename "nixosModule";

I'm using the Korora type system for Nix as a simple way to define the schema of a deploy plan and confirm it's correct. A plan has a few pieces:

Note that when I say "NixOS module" above, I mean anything that can appear within the imports when evaluating NixOS modules. So in addition to attribute sets and functions that produce attribute sets, paths are also valid in those places.

evalHost =
  name: modules:
  let
    nixpkgsKey = plan.meta.nixpkgs.${name} or plan.meta.nixpkgs.default;
    pkgsPath = if lib.isPath nixpkgsKey then nixpkgsKey else sources.${nixpkgsKey};
    evalConfig = import "${toString pkgsPath}/nixos/lib/eval-config.nix";
  in
  evalConfig {
    modules = [
      plan.defaults
    ]
    ++ modules;
    specialArgs = {
      inherit name;
      nodes = uncheckedHostsAndVMs;
    }
    // plan.meta.specialArgs;
  };

A core operation for evaluating a plan is to evaluate a NixOS config for a host. The name of the host is used to determine where to load Nixpkgs from for that host. The defaults module from the plan is combined with the ones passed in to the function. The plan's specialArgs are augmented with two additional arguments: name and nodes. The name argument lets generic modules know the name of the host. The nodes argument lets hosts be configured using information from other hosts. This is incredibly powerful but also rough on evaluation times, because evaluating a single host may require evaluating at least part of every other host. I find it far too useful to not have, in spite of this.

uncheckedHostsAndVMs = lib.concatMapAttrs (
  name: value:
  let
    host = evalHost name [
      { _module.check = false; }
      value
    ];
  in
  { ${name} = host; } // lib.mapAttrs (_: vm: vm.config) host.config.microvm.vms
) plan.hosts;

Some of the performance penalty of evaluating other hosts is mitigated by the fact that these are evaluated "unchecked". This skips checking if every option definition matches a declaration defined somewhere. I took this from Colmena's implementation. I assume it makes a meaningful impact on performance, but I've never tested it personally.

This attribute set of evaluated hosts includes any microVMs as well. It's passed along when evaluating hosts, as shown above, and to generate some information passed along to Pulumi to set up infrastructure. This unchecked set is the not the one used to produce the final toplevel systems to build, though.

hosts =
  plan.hosts
  |> lib.filterAttrs (name: _: namesToInclude == [ ] || lib.elem name namesToInclude)
  |> lib.mapAttrs (name: module: evalHost name [ module ]);

The systems to actually build are captured in the hosts attribute set. This excludes microVMs: they will be deployed as part of their host machine, so they shouldn't be represented separately here. This evaluation does not disable checks, so all options do end up getting checked at some point. This set of hosts also checks the "namesToInclude" parameter, so it can avoid evaluating extra hosts when the dippy command was given a specific list of hosts to work with.

toplevels = hosts |> lib.mapAttrs (_: v: v.config.system.build.toplevel) |> lib.recurseIntoAttrs;

What dippy will actually see and evaluate for each host is this attribute set of toplevels. This is the complete build of each system, with everything needed to install and run it as a new generation. dippy uses nix-eval-jobs to evaluate all of the systems in parallel, and without lib.recurseIntoAttrs, it won't see the derivations for each toplevel.

tests =
  (
    hosts
    |> lib.attrValues
    |> map (v: v.config.mjm.deploy.tests)
    |> lib.mergeAttrsList
  )
  // {
    recurseForDerivations = withTests;
  };

dippy can also run NixOS VM tests as a preflight check when diffing or deploying to hosts. A module can add entries to mjm.deploy.tests and those will be built by dippy. At the CLI, running these tests can be enabled or disabled with a flag, which is passed along when evaluating as the withTests argument. By setting recurseForDerivations based on that flag, we control whether nix-eval-jobs will see and evaluate those derivations.

config = {
  deployment = lib.mapAttrs (_: v: v.config.mjm.deploy.config) hosts;
  phases = phasesWithHosts;
};

In addition to the built toplevels for each host being deployed, dippy will need some additional information to perform the deploy correctly. This information will come from a JSON file that is generated when evaluating the plan.

Some of this information is configured using the mjm.deploy.* options on each host, and that information is gathered under the deployment key. The other important piece of information is which hosts belong to which phase of the deploy, so that will be assembled now.

phasesWithHosts =
  let
    groupedHosts = hosts |> lib.attrNames |> lib.groupBy phaseForHost;
  in
  map (p: {
    inherit (p) name;
    hosts = groupedHosts.${p.name} or [ ];
  }) plan.phases;

Grouping hosts into phases is a matter of asking each host which phase it belongs to and grouping the host names by that. Then we map over all of the phases and associate the name of the phase with the list of host names that had that phase. Of course, the real logic is in phaseForHost.

phaseForHost =
  name:
  let
    config = hosts.${name}.config;
    phase = lib.findFirst (p: p ? includeIf && p.includeIf name config) defaultPhase plan.phases;
  in
  phase.name;

defaultPhase = lib.findFirst (p: !(p ? includeIf)) null plan.phases;

To determine the phase for a host, each phase that has an includeIf function defined is checked in order. The first one that returns true for that host is the phase that is assigned to that host. If none match, then the default phase (the first phase with no includeIf function) is used.

infra =
  let
    allHostsWithMicroVMs = lib.attrValues uncheckedHostsAndVMs;
  in
  {
    <<vhosts>>
    <<oidcClients>>
    <<vaultServices>>
    <<vaultRoles>>
  };

dippy also needs one more JSON file. This one is for performing infrastructure updates with Pulumi. Any information that Pulumi resources need that is derived from the configuration of the host machines will be generated into this file. This allow NixOS configuration to be an authoritative source of truth for parts of the infrastructure, which can allow for some nice automation.

vhosts =
  allHostsWithMicroVMs
  |> lib.filter (n: !n.config.mjm.ingress.enable)
  |> map (n: n.config.mjm.ingress.vhosts)
  |> lib.mergeAttrsList
  |> lib.mapAttrs (_: v: v.useIPv4Proxy);

The complete list of virtual host subdomains is used to set up DNS records. vhosts is an attribute set where the key is the subdomain and the value is a boolean for whether that subdomain is accessible via IPv4 or not, as that affects which record it is CNAME'd to.

oidcClients =
  allHostsWithMicroVMs
  |> map (n: n.config.mjm.services)
  |> lib.mergeAttrsList
  |> lib.attrValues
  |> lib.filter (s: s.http.ingress.subdomain != null && s.http.ingress.authMode == "oidc")
  |> map (s: s.http.ingress.oidc.id);

The complete list of services using OpenID Connect is used to generate client IDs and client secrets automatically. For services that can do declarative OIDC configuration, the secret is then stored in Vault so that the service can read it from there.

vaultServices =
  allHostsWithMicroVMs |> lib.concatMap (n: n.config.mjm.vault.services) |> lib.uniqueStrings;

Every service that has secrets enabled will get a Vault entity and entity alias generated for it that corresponds to its SPIFFE ID. The entity will have a policy that allows it to access its own secrets. The alias will cause this entity to be used whenever the service authenticates with Vault using a SPIFFE JWT SVID, granting it the access it needs.

vaultRoles = lib.attrNames uncheckedHostsAndVMs;

Independent of any services, each host gets an entity and entity alias in Vault for signing a host SSH certificate. The metadata on the entity is used to ensure that each machine can only issue a certificate that is valid for its own hostname, so client machines that trust the CA for those certificates can be assured they are connecting to the correct machine.

evalPlan =
  uncheckedPlan:
  {
    namesToInclude ? [ ],
    withTests ? true,
  }:
  let
    plan = planType.check uncheckedPlan uncheckedPlan;

    <<evalHost>>
    <<uncheckedHostsAndVMs>>
    <<hosts>>
    <<toplevels>>
    <<tests>>
    <<config>>
    <<phasesWithHosts>>
    <<phaseForHost>>
    <<infra>>
  in
  {
    inherit
      config
      infra
      hosts
      toplevels
      tests
      ;
    configJson = json.generate "plan-config.json" config;
    infraJson = json.generate "infra.json" infra;
    vms = lib.concatMapAttrs (_: n: lib.mapAttrs (_: vm: vm.config) n.config.microvm.vms) hosts;
  };

Finally, we have everything needed to evaluate a plan. evalPlan expects to be used in a curried fashion: the plans.nix file for the project will call it with just a single argument that is the plan to evaluate, and that will produce a function that takes in some arguments and produces the evaluated plan. dippy will evaluate the plans file, calling the function and passing in arguments based on CLI arguments passed to dippy itself.

Not every key in the resulting evaluated plan is actually used by dippy itself. It only uses toplevels, tests, configJson, and infraJson. The latter two are there in addition to config and infra because nix-eval-jobs only evaluates derivations, not arbitrary Nix data structures. Generating the content into JSON files lets dippy perform a single evaluation command to get everything it needs.

All of the other keys are included in the final result for convenience outside of dippy. In particular, I can run "nix repl -f plans.nix" and have all of this config accessible from a REPL, making it easy to inspect the config of different hosts or VMs as needed.

#| file: packages/dippy/deploy.nix
let
  <<planType>>
  sources = import ../../npins;
  pkgs = import sources.nixos-small { };
  inherit (pkgs) lib;
  json = pkgs.formats.json { };
  <<evalPlan>>
in
evalPlan

My current deploy plan

#| file: plans.nix
let
  hostNames = [
    "aion"
    "apollo"
    "arges"
    "artemis"
    "athena"
    "brontes"
    "demeter"
    "hades"
    "niobe"
    "persephone"
    "steropes"
    "uranus"
  ];

  evalPlan = import ./packages/dippy/deploy.nix;
in
evalPlan {
  <<plan>>
}

To make it easier to add new hosts, I keep a list at the top of the file with the names of each host. This will be used later to set the hosts in the plan.

Besides that, the plan file just loads the evalPlan function described above and calls it with the configuration of the plan.

meta.nixpkgs = {
  default = "nixos-small";
  athena = "nixos";
  persephone = "nixos";
  uranus = "nixos";
};

All of my server machines run against the nixos-unstable-small channel, which is the "nixos-small" source in my npins. The three desktop machines run nixos-unstable from the "nixos" source, as they run GUI applications that would be expensive to compile from source as nixos-unstable-small would demand.

I wouldn't recommend most people do this: most should probably just use nixos-unstable and call it a day. I like to stay on the bleeding edge though, and my lab has caught problems in Nixpkgs before they landed in the main channels before, so I value that as a way I can contribute.

meta.specialArgs = {
  inputs = import ./npins;
  hostConfig = null;
  localModulesPath = toString ./modules;
};

I pass a handful of special args to my NixOS modules:

defaults = ./modules/nixos;

The modules/nixos/default.nix ends up importing all the modules that should be available to one of my NixOS machines. By including it here as defaults, all of my hosts have access to all of my custom options without needing to explicitly import anything.

hosts =
  hostNames
  |> map (name: {
    inherit name;
    value = ./hosts/${name};
  })
  |> builtins.listToAttrs;

The list of host names above is used to generate the attribute set of hosts for the plan. Each host corresponds to a module at hosts/${name}/default.nix.

This could likely be expressed more succinctly with Nixpkgs lib functions, particularly lib.genAttrs, but I'm choosing to avoid importing the Nixpkgs lib from this file, so I'm restricted to builtin functions.

phases = [
  {
    name = "spire";
    includeIf = _: c: c.mjm.spire.server.enable;
  }
  { name = "main"; }
];

Finally, the list of phases for my deploys is actually really simple, and all of this logic is probably overkill for it. I used to have more complexity in the setup, but the current shape of my infrastructure doesn't really need it.

There is a single special phase at the beginning of the deploy for the machine running the SPIRE server. This needs to happen first or deploys may fail. The reason is that part of deploying that machine is a systemd service that updates the registration entries for SPIRE. If those aren't updated first, deploying a machine may fail because a service will not be able to be issued an identity that it should have.

Working with the plan from dippy

type Plan struct {
	Hosts     []*Host
	Tests     []nix.EvalJobResult
	Infra     *InfraInput
	ForceGoal string
	phases    []deployPhase
	sess      *Session
}

type InfraInput struct {
	Vhosts        map[string]bool `json:"vhosts"`
	VaultServices []string        `json:"vaultServices"`
	VaultRoles    []string        `json:"vaultRoles"`
	OIDCClients   []string        `json:"oidcClients"`
}

A Plan is a structure representing a fully evaluated deploy plan. It is a hydrated version of what is described in the Nix code above, so unlike the Nix code, it can interact with the real world hosts.

type EvalPlanOpts struct {
	Plans     string
	Workers   int
	Hostnames []string
	Tests     bool
}

func (sess *Session) EvalPlan(ctx context.Context, opts EvalPlanOpts) (*Plan, error) {
	slog.InfoContext(ctx, "evaluating plans", "file", opts.Plans, "hosts", opts.Hostnames, "workers", opts.Workers)

	<<EvalPlan-build-args>>
	<<EvalPlan-eval-jobs>>
	<<EvalPlan-gather-results>>
	<<EvalPlan-create-plan>>
}

EvalPlan creates a new Plan by evaluating the Nix code from a plan file and constructing the plan based on the eval results. A handful of options can be passed to EvalPlan:

The process of evaluating a plan can be broken down into four steps which will be covered individually below.

//| id: EvalPlan-build-args
args := map[string]string{}
if len(opts.Hostnames) > 0 {
	hostnamesBytes, err := json.Marshal(opts.Hostnames)
	if err != nil {
		return nil, fmt.Errorf("serializing hostnames to json: %w", err)
	}
	args["namesToInclude"] = fmt.Sprintf("builtins.fromJSON %q", string(hostnamesBytes))
}
if !opts.Tests {
	args["withTests"] = "false"
}

The plan file evaluates to a Nix function to allow a few things to be configured at evaluation time via CLI flags to dippy that become arguments to the plan. These are options that control how much Nix code actually gets evaluated: filtering out hosts or tests. The only way I can think of to do that while using nix-eval-jobs is to have the Nix code produce fewer things via function arguments.

The list of hostnames to filter to is serialized to JSON to pass to Nix, which will then deserialize it using builtins.fromJSON. I think this is the least effort way to pass a list to Nix, as it doesn't require generating much in the way of Nix syntax. Both possible arguments are only passed to Nix if they are different than their default values.

//| id: EvalPlan-eval-jobs
results, err := sess.EvalJobs(ctx, nix.EvalJobsOptions{
	Path:    opts.Plans,
	Args:    args,
	Workers: opts.Workers,
})
if err != nil {
	return nil, fmt.Errorf("running eval: %w", err)
}

Once the args are constructed, dippy will call out to nix-eval-jobs to do the evaluation. nix-eval-jobs outputs a line of JSON with the result for each attribute it evaluates, and the EvalJobs function will parse those results as they come in and return them to the caller via an iterator.

//| id: EvalPlan-gather-results
var planConfig struct {
	Phases     []deployPhase     `json:"phases"`
	Deployment map[string]Config `json:"deployment"`
}
var infraInput InfraInput
var errorAttrs []string
var testResults []nix.EvalJobResult
hostResults := map[string]nix.EvalJobResult{}

for r := range results {
	<<EvalPlan-process-result>>
}

if len(errorAttrs) > 0 {
	return nil, fmt.Errorf("evaluation failed for one or more attributes (%s)", strings.Join(errorAttrs, ", "))
}

To gather the results, we create a handful of variables to collect the different kinds of things that will be evaluated as part of the plan. Then each result is consumed as it is produced by nix-eval-jobs. Each result will be collected into one of the variables based on its attribute path for the most part. Once every result has been processed, if any attributes failed to evaluate, an error is returned and no Plan is produced.

//| id: EvalPlan-process-result
if r.Error != "" {
	slog.ErrorContext(ctx, "error evaluating attr", "error", r.Error, "attr_path", r.AttrPath)
	errorAttrs = append(errorAttrs, r.Attr)
} else if r.Attr == "configJson" {
	slog.InfoContext(ctx, "evaluated config")

	if err := r.RealiseJSON(ctx, sess, &planConfig); err != nil {
		return nil, fmt.Errorf("realising config json: %w", err)
	}
} else if r.Attr == "infraJson" {
	slog.InfoContext(ctx, "evaluated infra input")

	if err := r.RealiseJSON(ctx, sess, &infraInput); err != nil {
		return nil, fmt.Errorf("realising infra json: %w", err)
	}

	slog.DebugContext(ctx, "infra input", "input", infraInput)
} else if r.AttrPath[0] == "toplevels" {
	slog.InfoContext(ctx, "evaluated host", "hostname", r.AttrPath[1])
	hostResults[r.AttrPath[1]] = r
} else if r.AttrPath[0] == "tests" {
	slog.InfoContext(ctx, "evaluated test", "name", r.AttrPath[1])
	testResults = append(testResults, r)
} else {
	slog.WarnContext(ctx, "unexpected attribute evaluated", "attr_path", r.AttrPath)
}

If nix-eval-jobs reports an error for an attribute, the error is logged and the attribute is appended to the list of attributes that errored, to be reported at the end of evaluation. This lets evaluation continue so that no hosts that fail to evaluate are hidden by an early failure.

Besides errors, results are handled based on their attribute path. The configJson and infraJson attributes are derivations that produce JSON data that dippy needs to consume. For each of these, the derivation gets built immediately upon being evaluated, and the resulting JSON is parsed into the appropriate structure. This is abstracted behind the RealiseJSON method on the result. The host toplevels and NixOS VM tests are simpler: the result for these are just added to the appropriate map/list. And finally, a warning is logged for any unrecognized attributes; while these won't cause fatal issues, they do signal that evaluation is doing work that won't be used by dippy, so it's wasted time and should be fixed.

//| id: EvalPlan-create-plan
plan := &Plan{
	Tests:  testResults,
	Infra:  &infraInput,
	phases: planConfig.Phases,
	sess:   sess,
}
for name, r := range hostResults {
	if !plan.ContainsHost(name) {
		continue
	}

	h := NewHost(sess, name, r.System, r.DrvPath, r.OutPath(), planConfig.Deployment[name])
	plan.Hosts = append(plan.Hosts, h)
}
return plan, nil

With the results completely gathered without any errors, the Plan can be created. Some of the results directly become fields on the plan. Hosts are special though: each host result gets hydrated into a Host struct, which supports running all the various operations dippy need to run against hosts.

=> The Host type and the operations it supports

Note that the two fields in the planConfig are split up at this point. The phases get stored as part of the plan, as they are used to orchestrate the flow of the overall deploy. The deployment config for each host gets attached to that host when it is constructed.

In the case that the plan does not include a catch-all phase, there might be extra hosts in the plan that aren't part of any phase of the deploy. Those hosts are dropped at this point. This logic can likely be dropped, as I think it's leftover from a time when I wanted the plans file to support defining multiple plans with different phases for the same overall set of hosts. This is something I didn't end up actually needing, and I can't explain why I would have a plan that included hosts that aren't part of any phase.

func (p *Plan) Build(ctx context.Context) error {
	var drvPaths []string
	for _, h := range p.Hosts {
		drvPaths = append(drvPaths, h.DrvPath)
	}

	if err := p.sess.Realise(ctx, drvPaths); err != nil {
		return fmt.Errorf("building plan hosts: %w", err)
	}

	return nil
}

Build concurrently builds every host in the plan. It does not delegate to the Build method on Host. Instead, it runs a single Nix build command for all of the derivations. This is particularly meaningful when the Nix builder is using nix-output-monitor, as it allows it to show the progress for all of the hosts being built together.

func (p *Plan) Test(ctx context.Context) error {
	var drvPaths []string
	for _, h := range p.Tests {
		drvPaths = append(drvPaths, h.DrvPath)
	}

	if err := p.sess.Realise(ctx, drvPaths); err != nil {
		return fmt.Errorf("running plan tests: %w", err)
	}

	return nil
}

Test runs all of the NixOS VM tests that were evaluated for the plan. It largely works the same way as Build, just for a different set of derivations.

func (p *Plan) Deploy(ctx context.Context) error {
	hostsByName := map[string]*Host{}
	for _, h := range p.Hosts {
		hostsByName[h.Name] = h
	}

	for _, phase := range p.phases {
		var phaseHosts []*Host
		for _, name := range phase.Hosts {
			if h, ok := hostsByName[name]; ok {
				phaseHosts = append(phaseHosts, h)
			}
		}

		if err := p.deployPhaseHosts(ctx, phase.Name, phaseHosts); err != nil {
			return fmt.Errorf("deploying phase %s: %w", phase.Name, err)
		}
	}

	return nil
}

Deploy goes through the phases of the plan and deploys each one in order. The Plan stores the hosts as a single list and the phases store a list of host names, so some massaging of the data is needed to get the list of hosts for each phase.

func (p *Plan) deployPhaseHosts(ctx context.Context, name string, hosts []*Host) error {
	if len(hosts) == 0 {
		return nil
	}

	l := slog.Default().With("phase.name", name)
	l.InfoContext(ctx, "deploying phase", "host_count", len(hosts))

	return section.Do("Deploying "+name+" hosts", func() error {
		for _, h := range hosts {
			if err := section.Do("Deploying "+h.Name, func() error {
				return h.Deploy(ctx, p.ForceGoal)
			}); err != nil {
				return fmt.Errorf("deploying %s: %w", h.Name, err)
			}
		}

		l.InfoContext(ctx, "deployed phase")
		return nil
	})
}

deployPhaseHosts deploys the hosts for a single phase of the deploy. If the phase is empty, it bails out early, behaving as though the phase didn't exist at all. The start and end of the phase are marked with both a log message as well as section groupings when running in CI. Then each host in the phase is deployed in sequence, again with logs and sections bookending the process.

The Deploy method on Host handles the actual deploy process for a single host. If the ForceGoal field as been set on the plan before running the deploy, then that will override the normal automatic detection for whether to do a switch or boot.

func (p *Plan) EachHost(ctx context.Context, limit int, f func(context.Context, *Host) error) error {
	g, childCtx := errgroup.WithContext(ctx)
	g.SetLimit(limit)

	for _, h := range p.Hosts {
		g.Go(func() error {
			return f(childCtx, h)
		})
	}
	return g.Wait()
}

EachHost enables ad-hoc operations to be done for every host in the plan. It does not consider the phases, which are only relevant to the actual deploy process. The caller controls how many hosts are acted upon concurrently. If any of the hosts return an error, the first one is returned.

//| file: packages/dippy/deploy/plan.go
package deploy

import (
	"context"
	"encoding/json/v2"
	"fmt"
	"log/slog"
	"slices"
	"strings"

	"git.midna.dev/mjm/nix-config/packages/dippy/nix"
	"git.midna.dev/mjm/nix-config/packages/dippy/section"
	"golang.org/x/sync/errgroup"
)

<<Plan>>

// deployPhase is a single phase in a deploy. It is a group of nodes that should be
// deployed at a particular point during the deploy.
type deployPhase struct {
	Name  string   `json:"name"`
	Hosts []string `json:"hosts"`
}

<<EvalPlan>>
<<Plan.Build>>
<<Plan.Test>>
<<Plan.Deploy>>
<<Plan.deployPhaseHosts>>
<<Plan.EachHost>>

// ContainsHost returns true if any of the phases in the plan contain the host with
// the given name.
func (p *Plan) ContainsHost(name string) bool {
	return slices.ContainsFunc(p.phases, func(p deployPhase) bool {
		return slices.Contains(p.Hosts, name)
	})
}
Proxy Information
Original URL
gemini://midna.dev/homelab/dippy/plan.gmi
Status Code
Success (20)
Meta
text/gemini;lang=en-US
Capsule Response Time
6.839608 milliseconds
Gemini-to-HTML Time
0.333219 milliseconds

This content has been proxied by September (UNKNO).