Automatic updates for Nix dependencies

I have some automation for my homelab configuration to help me solve two somewhat related problems:

These are conceptually two orthogonal problems, but my solutions for both end up having to myreact to changes from upstreams, so it ends up making sense to solve them as part of a single automation workflow. This workflow is powered by a handful of dippy commands which are described here.

Overview

My solution for applying upstream changes quickly is more obvious and straightforward, so it makes sense to describe it first. I use npins to manage Nix dependencies, and I rely on a Forgejo Actions workflow that runs every half hour to check for updates to those pinned dependencies.

It's important to me to be able to vet these changes as best as I can before imposing them on the machines in my lab, so this workflow does not just merge these updates to the main branch. Instead, it opens a pull request, which allows further dippy-powered automation to build all of the hosts, run tests, and report changes in package versions on the PR. It also gives me the ability to choose when I'm ready to merge and have changes be automatically deployed. I can avoid landing a disruptive update until I'm ready to handle potential fallout.

The method I use for applying early fixes to dependencies is more unconventional. One way people might do this is to maintain a set of patch files and have them applied after a dependency is fetched. I used to do this, but it can be clunky: some changes are hard to represent this way, and resolving conflicts when things change often means reverting back to using an actual VCS to generate a new patch.

My approach is to embrace the fact that a VCS is already a very solid tool for handling this. In particular, Jujutsu has some very nice affordances for this, as you'll see in the code below. What I do is maintain a proper fork of dependencies that I need to apply changes to. These forks live under https://git.midna.dev/forks/.

In those forks, I maintain branches that combine the upstream code with my own changes using a "megamerge", sometimes also known as an "octopus merge". Both describe a merge commit with more than two parents. Instead of rebasing my changes on top of the current upstream commit one-by-one, my changes are siblings of the current upstream commit, and everything is combined into a single merge commit. This ends up being quite easy to manage and keep up-to-date as upstream moves, as you'll see in the command implementations below.

Fetching changes for forks

type ForksFetchCmd struct {
	Upstream string
	Token    string `env:"BOT_TOKEN"`
}

The "dippy forks fetch" command runs once for each forked repository. It takes the upstream Git URL for the repository as a flag, and reads a Forgejo token from the environment that it can use to push to the fork.

func (c *ForksFetchCmd) Run(ctx context.Context, cli *CLI, sess *deploy.Session) error {
	<<ForksFetchCmd.Run-init-repo>>
	<<ForksFetchCmd.Run-remotes>>
	<<ForksFetchCmd.Run-fetch>>
}

The process of fetching changes for a forked repository has two preparatory steps before the actual fetches.

//| id: ForksFetchCmd.Run-init-repo
_, err := os.Stat(".jj")
if err != nil && !errors.Is(err, os.ErrNotExist) {
	return fmt.Errorf("checking for .jj directory: %w", err)
}
if err != nil {
	if err := section.Do("Initializing jj repo", func() error {
		return cmd.Run(ctx, sess, "jj", "git", "init")
	}); err != nil {
		return fmt.Errorf("initializing jj repo: %w", err)
	}
}

First, it checks for a .jj directory inside the working directory. This command expects to already be running in the directory with the clone of the repository, but if this is the first time the workflow is running for a particular fork, that directory might be empty. If so, an empty jj repository is initialized.

//| id: ForksFetchCmd.Run-remotes
if err := section.Do("Configuring git remotes", func() error {
	pwd, err := os.Getwd()
	if err != nil {
		return fmt.Errorf("getting pwd: %w", err)
	}
	forkName := filepath.Base(pwd)

	if err := setRemoteURL(ctx, sess, "origin", fmt.Sprintf("https://ci:%s@git.midna.dev/forks/%s.git", c.Token, forkName)); err != nil {
		return fmt.Errorf("setting origin remote url: %w", err)
	}
	if err := setRemoteURL(ctx, sess, "upstream", c.Upstream); err != nil {
		return fmt.Errorf("setting upstream remote url: %w", err)
	}
	return nil
}); err != nil {
	return err
}

Each clone for a fork will have two remotes configured. "origin" will point to the fork on my Forgejo instance. The name of the repo is inferred from the basename of the current directory. "upstream" will, as you might expect, point to the upstream repository.

These remotes are configured every time the repo is fetched, so that even if they were altered out-of-band or the token has been rotated, they will be corrected each time.

func setRemoteURL(ctx context.Context, sess *deploy.Session, name, url string) error {
	if err := cmd.Run(ctx, sess, "jj", "git", "remote", "set-url", name, url); err != nil {
		if err := cmd.Run(ctx, sess, "jj", "git", "remote", "add", name, url); err != nil {
			return fmt.Errorf("setting url %q for remote %q: %w", url, name, err)
		}
	}

	return nil
}

setRemoteURL tries to update the URL of an existing remote, and falls back to adding a new one if that fails.

//| id: ForksFetchCmd.Run-fetch
return section.Do("Fetching remote changes", func() error {
	if err := cmd.Run(ctx, sess, "jj", "git", "fetch", "--remote", "origin"); err != nil {
		return fmt.Errorf("fetching origin changes: %w", err)
	}
	if err := cmd.Run(ctx, sess, "jj", "git", "fetch", "--remote", "upstream"); err != nil {
		return fmt.Errorf("fetching upstream changes: %w", err)
	}
	return nil
})

Finally, each remote is actually fetched using jj. Each remote's changes will be needed in order to update the forks properly, which is the next command's job.

//| file: packages/dippy/cmd_forks_fetch.go
package main

import (
	"context"
	"errors"
	"fmt"
	"os"
	"path/filepath"

	"git.midna.dev/mjm/nix-config/packages/dippy/cmd"
	"git.midna.dev/mjm/nix-config/packages/dippy/deploy"
	"git.midna.dev/mjm/nix-config/packages/dippy/section"
)

<<ForksFetchCmd>>
<<ForksFetchCmd.Run>>
<<setRemoteURL>>

Updating branches of forks

type ForksUpdateCmd struct {
	Branch string
}

The "dippy forks update" command runs for each branch in each forked repository. It takes the upstream branch name as a flag. It does not need a token, as that should be baked into the origin remote URL already after fetching.

func (c *ForksUpdateCmd) Run(ctx context.Context, cli *CLI, sess *deploy.Session) error {
	deployBranch := "deploy/" + c.Branch
	if err := section.Do("Tracking deploy branch", func() error {
		return cmd.Run(ctx, sess, "jj", "bookmark", "track", deployBranch, "--remote=origin")
	}); err != nil {
		return fmt.Errorf("tracking remote bookmark: %w", err)
	}

	return jjTransaction(ctx, sess, func() error {
		<<ForksUpdateCmd.Run-rebase>>
		<<ForksUpdateCmd.Run-push>>
	})
}

Each upstream branch "foo" being maintained on a fork maps to a branch "deploy/foo" in the fork that is meant to point to a merge commit that includes the tip of that upstream branch plus any additional changes I've added from upstream PRs or otherwise. The first thing the update command does is ensure that this deploy branch is tracked for the origin remote, so it can be rebased locally and then force-pushed to the remote.

The other two steps happen in a "transaction" so that they succeed or fail as one.

func jjTransaction(ctx context.Context, sess *deploy.Session, tx func() error) error {
	currentOpBytes, err := cmd.Output(ctx, sess, "jj", "op", "log", "--no-graph", "-T", "id", "--limit", "1")
	if err != nil {
		return fmt.Errorf("getting current jj operation: %w", err)
	}
	currentOp := string(currentOpBytes)

	if err := tx(); err != nil {
		if rollbackErr := cmd.Run(ctx, sess, "jj", "op", "restore", "--what", "repo", currentOp); rollbackErr != nil {
			return fmt.Errorf("rolling back jj transaction: %w, after tx error: %w", err, rollbackErr)
		}

		return err
	}

	return nil
}

jjTransaction runs a function inside a sort of makeshift transaction for jj. It records the current operation ID at the start, and then runs the function it was given. If that function succeeds, then it returns no error and the effects remain. But if that function fails, it uses "jj op restore" to restore the repository state to what it was before running the function.

//| id: ForksUpdateCmd.Run-rebase
if err := section.Do("Rebasing to include upstream changes", func() error {
	upstreamChange := c.Branch + "@upstream"
	otherChanges := fmt.Sprintf("%s- ~ ::%s", deployBranch, upstreamChange)

	return cmd.Run(ctx, sess, "jj", "rebase", "--config", "user.name=Midna Bot", "--config", "user.email=bot@midna.dev", "-s", deployBranch, "-o", otherChanges, "-o", upstreamChange)
}); err != nil {
	return fmt.Errorf("running jj rebase: %w", err)
}

This rebase command attempts to update the merge commit on the deploy branch so that it includes the latest changes from the upstream branch. It updates the parents of the merge commit to include the following:

If the upstream branch hasn't moved since this was last run, this should be a no-op. If it has advanced, the most common change is that it replaces the previous upstream tip commit in the merge commit's parents with the new one. It may drop additional parents from the merge, though, if they were commits that are now included in the upstream branch. This is usually the case for PRs that were included in the fork while they were still open but have since been merged. It's very convenient that these get dropped automatically, though this behavior does rely on the commits not being rebased or squashed.

It is somewhat common for this rebase to fail because the update introduces conflicts with the local changes on the fork. This command makes no attempt to resolve those. Instead, it simply fails, rolling back the transaction. I'll get emails about this and then need to go resolve the conflicts manually to get the workflow working again. That's fine with me.

//| id: ForksUpdateCmd.Run-push
if err := section.Do("Pushing changes", func() error {
	return cmd.Run(ctx, sess, "jj", "git", "push", "--bookmark", deployBranch)
}); err != nil {
	return fmt.Errorf("pushing changes to git: %w", err)
}

return nil

If the rebase was successful, it is pushed to the fork. This will either end up not changing anything or it will be a force push. That's fine for the way I use these: I'm not trying to maintain a proper history other than that of the upstream.

//| file: packages/dippy/cmd_forks_update.go
package main

import (
	"context"
	"fmt"

	"git.midna.dev/mjm/nix-config/packages/dippy/cmd"
	"git.midna.dev/mjm/nix-config/packages/dippy/deploy"
	"git.midna.dev/mjm/nix-config/packages/dippy/section"
)

<<ForksUpdateCmd>>
<<ForksUpdateCmd.Run>>
<<jjTransaction>>

Updating pinned dependencies

type PinsUpdateCmd struct {
	URL   string `env:"FORGEJO_SERVER_URL"`
	Repo  string `env:"FORGEJO_REPOSITORY"`
	Token string `env:"BOT_TOKEN"`
}

The "dippy pins update" command takes info from the environment to know how to talk to the Forgejo API, as it needs to work with pull requests.

func (c *PinsUpdateCmd) Run(ctx context.Context, cli *CLI, sess *deploy.Session) error {
	<<PinsUpdateCmd.Run-check-existing-pr>>
	<<PinsUpdateCmd.Run-check-channels>>
	<<PinsUpdateCmd.Run-push-update>>
	<<PinsUpdateCmd.Run-create-pr>>
}

The purpose of this command is to check for updates to the Nixpkgs channels I follow and create a pull request for my Nix config to update them when they have changed. It starts with a few checks for whether an update is needed before pushing changes and creating a PR.

//| id: PinsUpdateCmd.Run-check-existing-pr
fc, err := forgejo.NewClient(c.URL,
	forgejo.SetContext(ctx),
	forgejo.SetToken(c.Token))
if err != nil {
	return fmt.Errorf("creating forgejo client: %w", err)
}

repoParts := strings.SplitN(c.Repo, "/", 2)
owner := repoParts[0]
repo := repoParts[1]

pulls, _, err := fc.ListRepoPullRequests(owner, repo, forgejo.ListPullRequestsOptions{
	State: forgejo.StateOpen,
})
if err != nil {
	return fmt.Errorf("listing open prs: %w", err)
}

if slices.ContainsFunc(pulls, func(pull *forgejo.PullRequest) bool {
	return pull.Base.Name == "main" && pull.Head.Name == "npins-update"
}) {
	slog.InfoContext(ctx, "existing update pr is already open")
	return nil
}

The first check is to see if there is already a PR open for an automated update from this workflow. It lists any open PRs on the repo and checks for one that is merging from the npins-update branch to main. If such a PR exists, then the command exits without doing anything.

//| id: PinsUpdateCmd.Run-check-channels
pins, err := getPins()
if err != nil {
	return fmt.Errorf("reading pinned sources: %w", err)
}

channels, err := getOutdatedChannels(ctx, sess, pins)
if err != nil {
	return fmt.Errorf("getting outdated channels: %w", err)
}

if len(channels) == 0 {
	slog.InfoContext(ctx, "nothing to do")
	return nil
}

The other check is whether there are updates worth opening a PR over. The first step is to load information about the current pinned dependencies in the repo. Then those will be compared with the current state of those repos. If nothing has changed, then the command exits without doing anything.

type pinDefinition struct {
	Type       string `json:"type"`
	Repository struct {
		Type   string `json:"type"`
		Server string `json:"server"`
		Owner  string `json:"owner"`
		Repo   string `json:"repo"`
	} `json:"repository"`
	Branch   string `json:"branch"`
	Revision string `json:"revision"`
	Version  string `json:"version"`
}

func getPins() (map[string]pinDefinition, error) {
	var pins struct {
		Pins map[string]pinDefinition `json:"pins"`
	}
	f, err := os.Open("npins/sources.json")
	if err != nil {
		return nil, fmt.Errorf("opening sources.json: %w", err)
	}
	defer f.Close()

	if err := json.UnmarshalRead(f, &pins); err != nil {
		return nil, fmt.Errorf("unmarshalling sources.json: %w", err)
	}
	return pins.Pins, nil
}

getPins just reads the npins/sources.json file and unmarshals it into a map of Go structures for each pin.

func getOutdatedChannels(ctx context.Context, sess *deploy.Session, pins map[string]pinDefinition) ([]string, error) {
	channels := []string{"nixos", "nixos-small"}
	channels = slices.DeleteFunc(channels, func(name string) bool {
		slog.DebugContext(ctx, "checking pin", "name", name)
		pin := pins[name]

		url := fmt.Sprintf("%s%s/%s.git", pin.Repository.Server, pin.Repository.Owner, pin.Repository.Repo)
		latestBytes, err := cmd.Output(ctx, sess, "git", "ls-remote", url, pin.Branch)
		if err != nil {
			slog.ErrorContext(ctx, "error checking current git revision", "error", err, "name", name, "url", url, "branch", pin.Branch)
			return true
		}

		splitLatest := strings.Fields(string(latestBytes))
		latest := splitLatest[0]

		slog.InfoContext(ctx, "checked pin", "name", name, "old", pin.Revision, "new", latest)
		return pin.Revision == latest
	})

	return channels, nil
}

getOutdatedChannels takes the data loaded from getPins and looks at two entries in particular: nixos and nixos-small. These are the pins for the Nixpkgs channels I follow. For each one, it reconstructs the Git repo URL and uses "git ls-remote" to check the current revision of the pinned branch. If the revision is different than what is currently pinned, then that channel will be returned in the resulting list from getOutdatedChannels. Both of these pins are pointing at a Nixpkgs fork managed by the two commands shown above, so in the automated workflow, this command runs after all of the forks have been updated, so it will see the latest changes.

These two pins are far from the only dependencies I have that could be updated: this logic could check all of the pins if I wanted. I've done this in the past and found it far too noisy. While I've since eliminated the specific dependency that motivated that change, I've kept the logic this way. The Nixpkgs channels update with enough frequency to keep everything moving along, and they are usually the changes I'm most interested in. Other things can wait.

//| id: PinsUpdateCmd.Run-push-update
slog.InfoContext(ctx, "pins update is needed", "channels", channels)
if c.Token == "" {
	if err := cmd.Run(ctx, sess, "npins", "update", "--dry-run"); err != nil {
		return fmt.Errorf("running npins update: %w", err)
	}
	return nil
}

if err := cmd.Run(ctx, sess, "npins", "update"); err != nil {
	return fmt.Errorf("running npins update: %w", err)
}

if err := cmd.Run(ctx, sess, "git", "config", "user.email", "bot@midna.dev"); err != nil {
	return fmt.Errorf("setting git config: %w", err)
}
if err := cmd.Run(ctx, sess, "git", "config", "user.name", "Midna Bot"); err != nil {
	return fmt.Errorf("setting git config: %w", err)
}
if err := cmd.Run(ctx, sess, "git", "commit", "-m", "npins update", "npins/sources.json"); err != nil {
	return fmt.Errorf("running git commit: %w", err)
}
if err := cmd.Run(ctx, sess, "git", "push", "-f", "origin", "HEAD:refs/heads/npins-update"); err != nil {
	return fmt.Errorf("running git push: %w", err)
}

If at least one Nixpkgs channel needs an update, then the command proceeds to make it happen. If no Forgejo token is available, then instead of making the change for real, it just does a dry-run of updating the pins, printing that out, and then exits. This is useful for testing the update logic in the command.

If it does have a token, then it goes about committing and pushing the necessary changes. It runs "npins update" which will update every pinned dependency to its latest version. This includes the non-Nixpkgs pins that were not checked previously: they still get updated as long as there is also a Nixpkgs change. The changes to the npins sources are then committed and pushed to the npins-update branch.

//| id: PinsUpdateCmd.Run-create-pr
newPins, err := getPins()
if err != nil {
	return fmt.Errorf("reading new pinned sources: %w", err)
}
var body strings.Builder
names := slices.Collect(maps.Keys(newPins))
slices.Sort(names)
for _, name := range names {
	pold := pins[name]
	pnew := newPins[name]

	if pnew.Revision == pold.Revision {
		continue
	}

	if pnew.Version != "" && pold.Version != "" {
		fmt.Fprintf(&body, "- **%s**: %s -> %s\n", name, pold.Version, pnew.Version)
	} else {
		fmt.Fprintf(&body, "- **%s**: `%s` -> `%s`\n", name, pold.Revision, pnew.Revision)
	}
}

pull, _, err := fc.CreatePullRequest(owner, repo, forgejo.CreatePullRequestOption{
	Head:     "npins-update",
	Base:     "main",
	Title:    "npins update: " + strings.Join(channels, ", "),
	Body:     body.String(),
	Assignee: "mjm",
})
if err != nil {
	return fmt.Errorf("creating pull request: %w", err)
}
slog.InfoContext(ctx, "created pull request", "number", pull.Index)
return nil

The last step is to open a PR to merge the npins-update branch into main. The PR title includes the names of the Nixpkgs pins that were updated. The PR body lists all of the pins that were updated, which is done by calling getPins a second time now that the sources have been updated and comparing those entries to the previous ones.

With the PR open, I will get an email for it and a CI build will start automatically to build all of the machines and report which packages have changed. When that build succeeds, I can merge the changes at my leisure, at which point the automated updates will be able to create a new PR once there are more updates.

//| file: packages/dippy/cmd_pins_update.go
package main

import (
	"context"
	"encoding/json/v2"
	"fmt"
	"log/slog"
	"maps"
	"os"
	"slices"
	"strings"

	"codeberg.org/mvdkleijn/forgejo-sdk/forgejo/v2"
	"git.midna.dev/mjm/nix-config/packages/dippy/cmd"
	"git.midna.dev/mjm/nix-config/packages/dippy/deploy"
)

<<PinsUpdateCmd>>
<<PinsUpdateCmd.Run>>
<<getPins>>
<<getOutdatedChannels>>
Proxy Information
Original URL
gemini://midna.dev/homelab/dippy/updates.gmi
Status Code
Success (20)
Meta
text/gemini;lang=en-US
Capsule Response Time
10.548826 milliseconds
Gemini-to-HTML Time
0.374918 milliseconds

This content has been proxied by September (UNKNO).