Log aggregation with Loki

{
  pkgs,
  config,
  lib,
  ...
}:
let
  cfg = config.mjm.loki;
in
{
  options.mjm.loki = {
    enable = lib.mkEnableOption "Grafana Loki";
  };

  config = lib.mkIf cfg.enable {
    <<config>>
  };

  _class = "nixos";
}

Loki is Grafana's service for collecting logs from multiple sources to able to query them from within Grafana. I use it collect the systemd journal from every machine into one place. Honestly, I'm still really in the habit of SSHing into machines and running journalctl by hand, but it would be good to do that less.

Service settings

mjm.services.loki = {
  http = {
    port = 3100;
    health.path = "/ready";

    metrics.enable = true;
  };

  s3.enable = true;
  s3.buckets = [ "loki-logs" ];
};

Loki listens locally on port 3100, and that traffic is protected by a tunnel and advertised in Consul.

Loki stores logs in a Garage bucket.

Config

services.loki.enable = true;

The Loki service of course needs to be enabled.

services.loki.configuration.auth_enabled = false;

Loki defaults to running in multi-tenant mode, where each request is expected to receive a header indicating an org ID from a trusted proxy. I have no reason to do that, as I am my only tenant, so this is disabled.

services.loki.configuration.server.http_listen_address = "::1";
services.loki.configuration.server.grpc_listen_address = "::1";

Loki normally listens on all IPs, but I want it to only listen locally where possible. External traffic should come in via mTLS through the tunnel.

services.loki.configuration.common = {
  replication_factor = 1;
  instance_addr = "::1";
};

Since I run Loki in monolithic mode, where all components run together in the same process, I need to specify a replication factor of 1 (since there's only one instance). And because I've overridden the listen addres for gRPC to be ::1, the instances of each component need to advertise themselves with that address to others.

services.loki.configuration.storage_config = {
  aws = {
    region = "home";
    bucketnames = "loki-logs";
    insecure = true;
    s3forcepathstyle = true;
    endpoint = "http://localhost:3902";
  };
  tsdb_shipper = {
    active_index_directory = "tsdb-index";
    cache_location = "tsdb-cache";
  };
};

Loki usually relies on object storage to store logs, and my setup is no different. The logs are stored in my Garage cluster in the loki-logs bucket.

The TSDB shipper is the currently recommended way of storing the index for log chunks. It uses some directories in Loki's data directory to do its work, but the actual content of the index is stored in Garage alongside the chunks.

services.loki.configuration.schema_config.configs = [
  {
    from = "2024-04-15";
    store = "tsdb";
    object_store = "s3";
    schema = "v13";
    index = {
      prefix = "index_";
      period = "24h";
    };
  }
];

Loki supports multiple schema configurations at the same time, with each covering a different range of dates. This allows upgrading where and how data is stored without the need to actually migrate anything. To change the storage schema, you just add a new schema config with a "from" date in the future, roll out that change, and once that it is that date, data will start being stored in the new format, while older data can be still be accessed using the config for the range it falls in.

I only have one schema config right now with the latest schema version using the TSDB store. I've had other configurations in the past, as I've been running Loki for several years now, but I don't retain logs long enough for any of them to be relevant anymore.

services.loki.configuration.compactor = {
  working_directory = "compactor";
  retention_enabled = true;
  delete_request_store = "s3";
};

By default, Loki doesn't delete old logs: they are retained forever. That's ridiculous for my purposes though: I have no use for logs after a while, so I've enabled retention for the S3 store.

services.loki.configuration.limits_config = {
  retention_period = "672h";
  ingestion_rate_mb = 64;
  ingestion_burst_size_mb = 96;
};

The actual retention period the compactor configured above uses is set here in the limits config. I've set it for 4 weeks, which is plenty for what I need. I don't want my log storage to grow indefinitely.

The ingestion limits are increased from their defaults. I adjusted these when I switched from Promtail to Alloy. I must have started hitting up against the default limits at that point, though I can't entirely remember the specifics.

mjm.spire.tunnels.alloy-loki.enable = false;

The tunnel Alloy uses to communicate with Loki on other machines should be disabled here to avoid a port conflict and allow Alloy to talk directly to Loki.

mjm.deploy.tests = {
  inherit (pkgs.nixosTests) loki;
};

My deploy tooling will run the Loki VM test from Nixpkgs in CI.

Proxy Information
Original URL
gemini://midna.dev/homelab/services/loki/
Status Code
Success (20)
Meta
text/gemini;lang=en-US
Capsule Response Time
418.727633 milliseconds
Gemini-to-HTML Time
0.298827 milliseconds

This content has been proxied by September (UNKNO).