Observability@rhythmjs/observability
Health & shutdown
Concurrent health indicators behind cached liveness and readiness probes, and a graceful-shutdown sequence that drains before it tears down.
The health module#
@rhythmjs/observability/health is a kernel module rather than a middleware. healthModule.forRoot(options) provides a healthService; the options are indicators (the checks below), timeout (per-indicator default, 5000 ms), and cacheTtl (readiness report cache, 1000 ms).
import { Rhythm } from "@rhythmjs/rhythm";
import { healthModule } from "@rhythmjs/observability/health";
const app = new Rhythm().register(
healthModule.forRoot({
indicators: [
{ name: "db", check: async () => ({ status: (await db.ping()) ? "up" : "down" }) },
{ name: "search", critical: false, timeout: 1000, check: () => search.health() },
],
}),
({ healthService }) => ({ healthService }),
);Indicators#
A HealthIndicator is { name, check, critical?, timeout? }. check() returns (or resolves to) { status: "up" | "down", details? }. All indicators run concurrently; each races its own timeout (its timeout, else the module default), and a throw or timeout becomes status: "down" with the error message under details.error; a broken check never breaks the endpoint. Every check reports its durationMs.
critical defaults to true: a down critical indicator marks the whole report down. Mark best-effort dependencies critical: false; their status still shows in the report without failing readiness.
Liveness and readiness#
The service exposes the two probe questions separately. live() is synchronous and cannot fail - { status: "up", uptime, timestamp } with uptime in whole seconds, answering "is the process running?". ready() answers "should traffic come here?": it runs the indicators and returns a HealthReport - { status, shuttingDown, checks } keyed by indicator name. Reports are cached for cacheTtl so probe storms and dashboard polling stay cheap; once shutdown() has been called, ready() short-circuits to down with shuttingDown: true and no checks run.
Serving the probes#
healthRoutes(service, { path? }) returns a RhythmRouter mounted at /health by default: GET /health/live always answers 200 with the liveness snapshot, and GET /health/ready answers the report with 200 when up and 503 when down; the status code Kubernetes-style readiness probes act on.
import { healthRoutes } from "@rhythmjs/observability/health";
router.use(healthRoutes(healthService).middleware());
// GET /health/live -> 200 { status: "up", uptime, timestamp }
// GET /health/ready -> 200 | 503 HealthReportGraceful shutdown#
gracefulShutdown(options) ties the pieces to process signals (SIGTERM and SIGINT by default). On the first signal it runs one ordered sequence: flip the healthService to shutting-down (readiness starts failing, so load balancers drain the instance), wait drainMs (default 0), run your close hook (stop servers, close sockets), then the app's teardown() (providers dispose in reverse order), then exit 0. A force timer - timeoutMs, default 10 000 ms, unref'd so it never keeps the process alive, and exits 1 if the sequence hangs. Repeated signals are ignored once shutdown has started.
import { gracefulShutdown } from "@rhythmjs/observability/health";
const shutdown = gracefulShutdown({
healthService,
app,
close: () => server.stop(),
drainMs: 5000,
});
// the returned function triggers the same sequence manually:
process.on("uncaughtException", () => void shutdown());Set exit: false to run the sequence without exiting the process, useful under test runners, where the returned trigger function replaces the signal.