Scaleway Elastic Metal nodes that hit a silent kernel lockup on this provider now self-heal instead of wedging deploys. New nodes are provisioned with a systemd hardware watchdog (RuntimeWatchdogSec=30s) plus kernel panic/on-oops sysctls, so a frozen box reboots itself in roughly 30s with its warm cache preserved. A backstop MachineHealthCheck per fleet also remediates any existing node that has been NotReady for 10 minutes, so a single dead box no longer blocks the observability rollout, the canary job, and every production deploy behind it.
Hive
Auto-recover silently locked-up bare-metal Kura nodes
Published
Jul 02, 2026 · 09:52 UTC
Repository
tuist/tuist