pieterpel
← All writing
Nixberry · 5 of 5

Run the toy like prod

A headless console fails silently. Its logs and metrics ship to the same Grafana as the server, so a dead service becomes a Telegram message instead of a mystery.

Published 2026-07-17 7 min read #nixos #homelab #grafana #observability

The problem: a box with no screen

The console sits under the TV. It boots on its own, runs for days, and nobody ever logs into it. That is the point — and that is the problem. There is no error dialog, no red banner, no window that will not close. A service can die at 3am and the box keeps looking exactly like it looks when it is fine: a black screen with Kodi on it.

Worse, this machine has a known bad day in it. Under the right conditions it gets slower and slower and then wedges, and nothing is logged when it does. As far as systemd is concerned nothing failed. So “is it doing it again?” is not a question I can answer from the box itself: the journal is gone on the next boot and “it feels slow” is not a measurement.

So it gets watched the way the server is watched. Not because a game console needs 99.9% availability, but because a failure nobody can see is a failure I cannot reason about.

Getting the logs off the box

The console’s logs should outlive the console, and getting them out should not mean opening anything up. The VM already runs Loki for its own journal, so the console just joins that.

Alloy runs on the Pi, tails the systemd journal and pushes to the VM’s Loki over the tailnet. Loki has no auth of its own (auth_enabled = false), and for a log store that is worse than it sounds: anyone who can reach the port can read the logs and also write to them, and every alert further down this page believes what it reads there. So the push endpoint gets a gate of its own — the password rendered from sops at boot, in front of the only port the console is allowed to reach — and the tailnet is left to carry the traffic. Nothing is exposed to the internet to make the console’s logs visible.

The interesting part is the relabelling. Journald fields become Loki labels so a query can group by service — the systemd unit, the priority. And then one rule that exists because of a gap I only noticed when a service died and nothing showed up under its name: when a unit fails, the line that says foo.service: Main process exited is written by systemd itself and carries no _SYSTEMD_UNIT field, so it arrives unlabelled and lands in the pile of anonymous lines. The rule pulls the unit name out of the message body with a regex. Without it, the one message you want when something dies is the one message you cannot group by anything.

source · modules/monitoring/alloy.nix:35-65 view on GitHub ↗
          // name from the body so these events land under the right unit label.
          rule {
            source_labels = ["unit", "__journal_message"]
            separator     = ";"
            regex         = ";([a-zA-Z0-9_.@/-]+\\.(?:service|timer|socket)): .+"
            target_label  = "unit"
            replacement   = "$1"
          }
        }

        loki.source.journal "read" {
          max_age       = "12h"
          labels        = { "job" = "systemd-journal", "host" = "${config.networking.hostName}" }
          relabel_rules = loki.relabel.journal_relabel.rules
          forward_to    = [loki.write.central.receiver]
        }

        ${
          lib.optionalString withAuth ''
            // The sops-rendered password file, read once at startup; is_secret keeps
            // it out of Alloy's own UI and logs.
            local.file "loki_password" {
              filename  = "${config.sops.templates.${templateName}.path}"
              is_secret = true
            }
          ''
        }

        loki.write "central" {
          endpoint {
            url = "${cfg.url}"

Metrics, because a hang does not write a log line

Logs cover the failures that announce themselves. The other category — slowing down until it stops — has nothing to announce. Nothing crashes, nothing logs, and afterwards there is no evidence at all, which is exactly the situation I keep ending up in: something happened, and the answer is somewhere in a window I did not record.

So the box also reports numbers. node_exporter listens on the Pi and the Prometheus on the VM scrapes it over the tailnet every fifteen seconds. That port is not open to the home LAN this box is also plugged into: it is declared once, next to the exporter that listens on it, and both machines’ firewalls are generated from that declaration — and it gets the metrics into the same Grafana, the same alerts and the same 30 days of history as the server’s own metrics. cpu, memory, load, disk, network, services — the usual list. Plus hwmon, which is the one collector added on purpose: this box’s failure mode has no CPU, memory or disk signature at all, so without a temperature series there is no way to rule heat in or out after the fact. Everything else in that list is the default. That one line is the decision.

source · modules/monitoring/node-exporter.nix:18-44 view on GitHub ↗
      config = lib.mkIf cfg.enable {
        # Declared rather than opened on every interface: the scrapers reach this
        # by MagicDNS name, so the port stays off whatever other network the host
        # is on (a home LAN, for instance).
        tailnet.openPorts.${config.networking.hostName}.${toString cfg.port} =
          "node_exporter, scraped by the central Prometheus";

        # hwmon is beyond the usual default set on purpose: a box that slows down
        # and then hangs has no CPU/memory/disk signature, so without a
        # temperature time series there is no way to rule thermal in or out
        # afterwards.
        services.prometheus.exporters.node = {
          enable = true;
          inherit (cfg) port;
          enabledCollectors = [
            "cpu"
            "diskstats"
            "filesystem"
            "hwmon"
            "loadavg"
            "meminfo"
            "netdev"
            "systemd"
            "time"
            "uname"
          ];
        };

A week in the life of the console

live · telemetry real data · scraped every 5 min
21 22 23 24 25 26 27 no data · 4 days busy 0% 54° 32° 51.6°C

dashed line = a save file was written: 26 Sep 17:46 Super Mario 64 · 27 Sep 00:20 Mario Kart 64 · 27 Sep 10:48 Donkey Kong 64

15 h since boot · 39.4 °C now · load 1.81 · 1.0 of 3.5 GB memory · 11.5 of 13.5 GB on / · scraped up

Read it as a picture of a machine with two states. Most of the week it sits between 36 and 39°C, which is a Pi 400 in a keyboard doing nothing at all. Then, three times, the line climbs to about 50°C and every core is busy for hours. Those three are the only times anyone sat down to play, and the dashed lines are the moment a save file landed: Super Mario 64 on Saturday evening, Mario Kart 64 just after midnight, Donkey Kong 64 on Sunday morning.

That is what emulation costs on this hardware: 100% of four cores and fifteen degrees of extra SoC temperature for two hours, and no throttling anywhere in sight. The number that matters is not the peak, though. It is that a question like “has it been hot all week or only when I play?” now has an answer instead of an opinion — and that the answer took one query, on a box I have not logged into for a week.

The alert that makes it loud

A dashboard nobody opens is decoration. The rules that matter are the ones that find me, and they all end in the same Telegram chat — the same chat where a bot answers questions about Grafana, so alerts and follow-up questions live in one conversation instead of two tools.

There are two kinds of rule. The log-based ones are Loki queries over the shipped journals, and the endpoint ones are blackbox probes hitting URLs from outside. The one that earns its keep watches for failed units:

Open the deep-dive — the rule that watches every service on both machines ›
source · flake/vm/grafana.nix:264-272
                  name = "service-health";
                  folder = "Alerts";
                  interval = "1m";
                  rules = [
                    (lokiAlert {
                      uid = "service-crashed";
                      title = "Systemd service crashed";
                      expr = ''sum by (unit) (count_over_time({job="systemd-journal", unit!~"grafana\\.service|loki\\.service|alloy\\.service"} |= "entered failed state" [5m]))'';
                      summary = "{{ $labels.unit }} entered failed state";

There is no mention of the console in that query, and that is the point. Because the Pi’s journal arrives with the same labels as the VM’s, a service that dies on the console trips the same rule as a service that dies on the server — one rule, both machines, and it works on the Pi for exactly the reason described above: entered failed state is one of those messages systemd writes about itself with no unit field attached.

The rest of the group is a similar list of things I want to hear about immediately: the OOM killer firing, twenty failed SSH logins in five minutes, a scrape target going quiet, a probed endpoint going down, a TLS certificate inside its last week. The monitoring stack’s own units are excluded from the service rule, so a Grafana or Loki restart does not page me about itself.

What this still does not solve

The monitoring lives on the VM. If the VM is down, so is the alerting and so is the history — the console’s telemetry is exactly as available as the machine that stores it. That is an acceptable trade here, because the VM is the one box that never gets reflashed, but it is not redundancy and I should not pretend otherwise.

The gap I closed while writing this: every rule above is either a log pattern or an HTTP probe, and neither notices a console whose exporter has stopped answering altogether. A plain up{job="node"} rule now covers that — five minutes of silence from either machine and the same chat hears about it.

And none of this fixes a hang. Knowing the box was hot does not stop it being hot. All of it is about the part I can actually control: turning a silent machine into one that leaves a trail, so that the next time something goes wrong the question is “what do the numbers say” instead of “when did I last sit down in front of it”.

Which is the whole series, really. Three machines, one of them disposable. The console rebuilds itself from a flake, its library and its saves live on the VM, and it now tells me when it breaks. Nothing on it is irreplaceable, which is what makes it safe to use — and now also what makes it possible to notice when it is not happy.