Files
healthcheck-container/README.md
T
Milan PandurovandClaude Fable 5.1 62b57a7e56 Add MQTT diagnostics, broker connection probe, and switch to Debian base
The MQTT publisher now tracks connection state, publish counters, discovery
and availability timestamps, and the last error, exposed at /api/mqtt and on
the dashboard as an MQTT card and a Home Assistant section. Connection
failures, disconnects, and dropped publishes are logged; paho's
on_connect_fail callback was not registered before, so failed attempts were
silent. MQTT connect, disconnect, and unreachable events go to the timeline.

When a connection attempt fails the service runs a probe from inside the
container: system resolver, each nameserver from resolv.conf queried directly,
/etc/hosts, and a TCP connect to every address found. The result is shown as
a summary and raw JSON on the dashboard and can be rerun via
POST /api/mqtt/probe.

The image base moves from Alpine to Debian slim. musl queries all
nameservers in parallel and accepts the first reply, so a fast public
NXDOMAIN beats a slower local server that knows the name. glibc asks the
nameservers in order.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-29 12:08:48 +02:00

144 lines
6.1 KiB
Markdown

# healthcheck-container
Small Flask service for remote devices (Raspberry Pi in Docker) that answers
`/alive` for container monitoring and, on top of that, watches the network the
device sits on: internet reachability, DNS, public IP changes, Wi-Fi link
quality, Wi-Fi networks in range, and host reboots. Metrics go to SQLite and
are shown on a dashboard at `/`. Optionally the same values are published to an
MQTT broker with Home Assistant discovery, so each device shows up in Home
Assistant as a device with sensors and no manual configuration.
## Endpoints
| Path | Purpose |
| --- | --- |
| `/alive` | Unchanged health check, returns `{"alive": true}` |
| `/` | Dashboard: status, Wi-Fi, graphs, timeline |
| `/api/status` | Latest state of every monitor as JSON |
| `/api/wifi` | Interface list, link details, last scan results |
| `/api/mqtt` | MQTT connection state, counters, last error, last published payload, last probe |
| `POST /api/mqtt/probe` | Run the broker connection probe now and return its result |
| `/api/events?limit=200` | Timeline events, newest first |
| `/api/samples?prefix=reach.&range=3600` | Bucketed samples for graphs |
## Running
The compose file pulls the published multi-arch image
`hiimmilan/health-check` (amd64, arm64, arm/v7):
```sh
docker compose pull && docker compose up -d
```
Then open `http://<device>:9999/`.
Wi-Fi needs two things from Docker:
- `network_mode: host`, otherwise the container has no wireless interface at all.
- `cap_add: NET_ADMIN`, otherwise link and interface info work but scanning fails
with a permission error that is shown on the dashboard.
Without `WIFI_INTERFACE` the dashboard lists the interfaces it can see and
nothing is scanned. Set the variable to one of the listed wireless interfaces
and restart the container.
## Storage
Samples and events are stored in SQLite at `DB_PATH` (default
`/data/healthcheck.db`). The compose file mounts a named volume at `/data`, so
the database survives `docker compose down`, image rebuilds and `up -d`. Only
`docker compose down -v` or `docker volume rm` deletes it. A bind mount works
as well, for example `- /opt/healthcheck:/data`.
Samples older than `SAMPLE_RETENTION_DAYS` (default 7) are deleted hourly.
Events are kept. Wi-Fi scan results are held in memory only and never written.
## Configuration
All settings are environment variables. Intervals are in seconds.
| Variable | Default | Meaning |
| --- | --- | --- |
| `PORT` | `9999` | HTTP port |
| `DB_PATH` | `/data/healthcheck.db` | SQLite file |
| `SAMPLE_RETENTION_DAYS` | `7` | How long graph samples are kept |
| `DEVICE_NAME` | hostname | Shown on the dashboard and in Home Assistant |
| `DEVICE_ID` | slug of `DEVICE_NAME` | Used in MQTT topics and unique ids |
| `REACH_TARGETS` | `1.1.1.1:443,8.8.8.8:443,9.9.9.9:443` | TCP connect targets, `host:port` |
| `REACH_INTERVAL` | `30` | |
| `REACH_TIMEOUT` | `3` | Connect timeout per target |
| `DNS_NAME` | `cloudflare.com` | Name resolved through the system resolver |
| `DNS_INTERVAL` | `60` | |
| `PUBLIC_IP_URLS` | ipify, ifconfig.me, icanhazip | Tried in order, first answer wins |
| `PUBLIC_IP_INTERVAL` | `300` | |
| `PUBLIC_IP_TIMEOUT` | `5` | |
| `UPTIME_INTERVAL` | `60` | Reboot detection via `/proc/uptime` |
| `WIFI_INTERFACE` | empty | Wireless interface to use, empty disables Wi-Fi |
| `WIFI_LINK_INTERVAL` | `30` | Signal, bitrate, retries of the current link |
| `WIFI_SCAN_INTERVAL` | `60` | Scan for networks in range |
| `MQTT_HOST` | empty | Broker host, empty disables MQTT |
| `MQTT_PORT` | `1883` | |
| `MQTT_USERNAME` | empty | |
| `MQTT_PASSWORD` | empty | |
| `MQTT_DISCOVERY_PREFIX` | `homeassistant` | Must match the MQTT integration setting |
| `LOG_LEVEL` | `INFO` | |
## Events on the timeline
- Internet unreachable / reachable again (all targets failed, then any succeeded)
- DNS resolution failed / working again
- Public IP is X, Public IP changed from X to Y
- Wi-Fi connected, disconnected, roamed to another BSSID
- Host rebooted
- Health check service started
- MQTT connected, disconnected, cannot reach broker
## Home Assistant
Set `MQTT_HOST` (and credentials if the broker needs them). On connect the
service publishes retained discovery messages under
`homeassistant/<component>/healthcheck_<device_id>/<key>/config`, all pointing
to one device, then publishes state to `healthcheck/<device_id>/state` and
availability to `healthcheck/<device_id>/availability`. When Home Assistant
restarts it announces itself on `homeassistant/status` and the service
re-publishes discovery.
The dashboard has an MQTT card and a Home Assistant section showing the
connection state, when it last connected or failed, how many messages were
published or dropped, when discovery and availability were last sent, and the
last state payload. The same data is at `/api/mqtt`. Use it to tell an app-side
problem (not connected, publishes dropped) from a Home Assistant-side one
(connected and publishing, but the device is still unavailable).
When a connection attempt fails the service runs a probe from inside the
container, at most every two minutes: it resolves the broker name with the
system resolver, queries each nameserver from `/etc/resolv.conf` directly,
checks `/etc/hosts`, and tries a TCP connect to every address it found. The
result is shown as a summary and as raw JSON in the Home Assistant section, and
can be rerun from the dashboard. If the nameservers disagree about the broker
name, for example a local server that knows it and a public one that does not,
the probe says so. The image is Debian based, so glibc asks the nameservers in
resolv.conf order and only moves on when one does not answer.
Entities per device:
- Internet, DNS, Wi-Fi (binary sensors, `connectivity` class)
- Internet latency, DNS latency (ms)
- Public IP
- Host uptime
- Wi-Fi SSID, BSSID, channel, signal (dBm), TX bitrate (Mbit/s), networks in range
## Development
```sh
python3 -m unittest discover -s tests
DB_PATH=./data/hc.db python3 app.py
```
Publish a new image for all Raspberry Pi architectures:
```sh
docker buildx build --platform linux/amd64,linux/arm64,linux/arm/v7 \
-t hiimmilan/health-check:latest -t hiimmilan/health-check:2.3.0 --push .
```