Posted on July 05, 2026
Category: Technology
Tags: nemoclaw, nvidia, openshell, wsl, systemd, docker, jwt, ed25519, device-pairing, llm-fallback, process-supervision, systemd-execcondition, debugging, linux, devops
Views: 182
NVIDIA's NemoClaw is an open-source stack for running always-on AI agents inside secure sandboxes. It's impressive technology — mTLS, JWT authentication, Landlock filesystem sandboxing, OPA network policies — but it's still in early preview (launched March 2026), and one very basic operational scenario was completely broken: rebooting your machine.
Every time my WSL instance rebooted, the sandbox would die with a cryptic Connection refused (os error 111). The only official fix was to rerun the full nemoclaw onboard command — a 5-minute interactive process that requires manually pasting an API key. Not exactly something you want to do every morning.
This post tells the full story of how I traced the problem through seven distinct root causes, including the dead ends and wrong turns along the way, and built a fully automatic recovery system — right down to an intermittent inference failure that turned out not to be my bug at all, and two late-breaking bugs that only surfaced after I thought I was done.
My environment:
systemd=true enabledwill (sudo) and nemouser (no sudo, dedicated NemoClaw user)mynemoThe constraint that made this interesting: nemouser has no sudo, so the solution had to work entirely within user-space tooling.
This post is a debugging narrative, not a copy-paste runbook. The diagnostic method and the root-cause map are the transferable parts; the literal commands are shaped by my specific machine. Before you follow any step, read this section — it's the difference between recovering your sandbox and locking yourself out of it.
Every identifier below is specific to my environment. Replace it with yours before running anything:
mynemo → your sandbox name (nemoclaw list)nemouser / /home/nemouser → your dedicated NemoClaw user and homeopenshell-mynemo-... → your container name (docker ps --filter name=openshell-)Throughout this post you'll see addresses like 10.xx.0.2:18789. That is deliberately masked, and it is not a stable value. Docker assigns the sandbox bridge subnet from a pool, and it can change on any reboot — mine moved from one /28 to a completely different one between two boots. If you copy a literal IP into a script or a curl, it will rot the first time the subnet shifts.
Always discover it at runtime instead:
CT=$(docker ps --filter name=openshell- --format '{{.Names}}' | head -1)
IP=$(docker inspect "$CT" \
--format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}')
curl -s -o /dev/null -w '%{http_code}\n' "http://$IP:18789/"
One more subtlety: the address the TUI shows (its in-sandbox network-namespace view) and the address the host sees for the same container are different IPs pointing at the same gateway. Don't try to reconcile them — probe whichever side you're actually on.
The process-cleanup steps (Root Cause 7) reference concrete PIDs like 497 (the healthy gateway) and 119 (its parent). Those are my PIDs from one specific run. Yours will be different every boot. Copying a literal PID means killing the wrong process — quite possibly the healthy gateway you're trying to protect.
Identify your own tree first, and preserve it:
CT=$(docker ps --filter name=openshell- --format '{{.Names}}' | head -1)
# The real gateway and its parent nemoclaw-start:
docker exec "$CT" ps -eo pid,ppid,etime,cmd | grep -E 'openclaw|nemoclaw-start'
The healthy tree is the oldest one whose nemoclaw-start parent has been up since boot; the storm's orphans are the young duplicates. Confirm elapsed time (etime) before killing anything.
The whole systemd-user approach assumes:
systemd=true, plus loginctl enable-linger for the NemoClaw userdocker-driver-gateway-config.ts:17), constant names (JWT_TTL_SECS), the [STRIPPED_BY_MIGRATION] marker, the openclaw.json schema, and the .config-hash format can all move between versions. Grep for the concept, don't trust the line number.dangerouslyDisableDeviceAuth=true. The one-time hand-pairing (Root Cause 5) and the token-match reasoning (Root Cause 7) both depend on this. If your gateway enforces real device auth, those shortcuts will not authenticate — you'd need the proper pairing flow instead.Two fixes touch state that is painful to lose:
signing.pem present on the host. Re-signing overwrites the token file in place (bind-mounts are inode-based) — verify against the public key before writing.identity/device.json and rewrites devices/paired.json by hand. Back up every touched file with a timestamp first. Get this wrong without a backup and you lose the paired identity entirely.And the cheapest lesson of all, from Root Cause 7: do not systemctl restart a healthy workload service. Restarting a working system ran the container entrypoint over a live gateway and stormed the process table to 949% CPU. If the sandbox is already serving, prove it needs restarting before you touch it.
Before diving into debugging, I needed to understand what NemoClaw actually runs. After reading source code and inspecting running processes, the picture became clear — it's a three-layer stack:
Layer 1: Host Gateway (openshell-gateway, Rust binary, port 8080, mTLS)
|
| gRPC
v
Layer 2: Docker Container (openshell-sandbox, PID 1, restart: unless-stopped)
|
| Internal process
v
Layer 3: OpenClaw Gateway (Node.js, port 18789, AI agent runtime)
The nemoclaw onboard command sets up all three layers from scratch. The critical insight: it starts everything but persists nothing. No systemd services, no cron jobs, no startup scripts. Just running processes that vanish on reboot.
The first clue was obvious — after reboot, nothing listens on port 8080:
ss -tlnp | grep 8080
# (empty)
Looking at the onboard source code (docker-driver-gateway-launch.ts), I found exactly how the gateway is started:
spawn(gatewayBin, args, { detached: true })
A detached child process. That's it. No systemd unit, no init script, no process supervisor. When WSL reboots, the process simply vanishes. And since every CLI command and every container talks to port 8080, losing the gateway breaks everything.
A systemd user service was the obvious answer. But getting it right took several iterations.
My first service file included:
After=docker.service
Requires=docker.service
This failed immediately: docker.service not found. In WSL2 with systemd enabled, Docker runs but doesn't register as a traditional systemd service.
To keep the ExecStart readable, I split it across lines with backslashes:
ExecStart=/home/nemouser/.local/bin/openshell-gateway \
--config /path/to/config.toml \
--bind-address 127.0.0.1
systemd parsed the backslash-newline as a literal, producing unrecognized subcommand '\ '. Everything had to go on one line.
Initially I set Restart=on-failure. This worked for crashes but not for signal kills. When something sent SIGHUP to the gateway (which it doesn't handle gracefully), the process would die and not restart because a signal death isn't classified as "failure."
[Unit]
Description=OpenShell Docker-driver Gateway (NemoClaw)
After=network.target
[Service]
Type=simple
ExecStartPre=/bin/bash -c 'until [ -S /run/docker.sock ]; do sleep 1; done'
ExecStart=/home/nemouser/.local/bin/openshell-gateway --config /home/nemouser/.local/state/nemoclaw/openshell-docker-gateway/openshell-gateway.toml --bind-address 127.0.0.1 --port 8080
Restart=always
RestartSec=5
Environment=DOCKER_HOST=unix:///run/docker.sock
Environment=OPENSHELL_DB_URL=sqlite:/home/nemouser/.local/state/nemoclaw/openshell-docker-gateway/openshell.db
WorkingDirectory=/home/nemouser/.local/state/nemoclaw/openshell-docker-gateway
[Install]
WantedBy=default.target
Key decisions:
ExecStartPre polls for /run/docker.sock instead of depending on a service that doesn't existRestart=always covers both crashes and signal killsloginctl enable-linger nemouser ensures the user service starts at boot without loginWith the gateway auto-starting, I expected things to work. Instead, the container kept crash-looping with a new error: sandbox has no spec. The gateway was running, the container could connect, but it couldn't find any sandbox configuration.
My first instinct was to check the database. I wrote a script to inspect it:
import sqlite3
conn = sqlite3.connect('/home/nemouser/.local/state/nemoclaw/openshell-docker-gateway/openshell.db')
# Found: sandbox_policy(67), provider(1), inference_route(1), sandbox(1)
The database had everything. So why couldn't the gateway find it?
At this point I started running various diagnostic commands, trying to reproduce the error in different ways. The user (myself) rightfully called this out: "That seems meaningless. Why don't we just look at the source code in the .nemoclaw directory?" A good reminder that reading code beats guessing.
The decisive diagnostic was simple — check what files the gateway process actually has open:
for fd in /proc/PID/fd/*; do
target="$(readlink "$fd")"
case "$target" in *.db) echo "$target" ;; esac
done
Result: the gateway was reading from ~/.local/state/openshell/gateway/openshell.db — not the NemoClaw-specific database.
Two databases existed:
~/.local/state/openshell/gateway/openshell.db — XDG default path, empty~/.local/state/nemoclaw/openshell-docker-gateway/openshell.db — 67 policies, 1 provider, 1 sandboxWhen onboard starts the gateway, it passes the correct DB path. But when systemd starts it fresh, the gateway falls back to its XDG default and finds nothing.
An environment variable in the service file:
Environment=OPENSHELL_DB_URL=sqlite:/home/nemouser/.local/state/nemoclaw/openshell-docker-gateway/openshell.db
I found the format by reading docker-driver-gateway-env.ts:206. Note: it's sqlite: with a single colon and no //. Getting this wrong would silently default back to the empty database.
After adding the DB fix, I did a full wsl --shutdown and restart:
Victory — or so I thought. The "Ready" state held because the JWT token hadn't expired yet (less than 1 hour since the last onboard). The time bomb was already ticking.
About 80 minutes after the successful reboot test, I tried nemoclaw mynemo connect and got OpenClaw gateway is not running inside the sandbox. Then the container started crash-looping. Docker logs revealed:
invalid token: ExpiredSignature
RefreshSandboxToken returned Unauthenticated;
static token sources cannot rebootstrap automatically
Policy fetch failed after 5 attempts
The container's sandbox.jwt token had expired. The container restart count hit 43 before I noticed.
Reading the source, I traced the full token lifecycle:
onboard generates an Ed25519 keypair (signing.pem + public.pem) and a kid file{sub, iss, aud, iat, exp, sandbox_id} using EdDSATTL: Hard-coded at docker-driver-gateway-config.ts:17:
export const DOCKER_DRIVER_GATEWAY_JWT_TTL_SECS = 3600;
Delivery: Token is written to a file, bind-mounted into the container
RefreshSandboxToken — but only with a valid (non-expired) tokenThis is the chicken-and-egg problem: once the token expires, the refresh endpoint also rejects it. The gateway doesn't re-mint tokens on resume. The only way to get a new token is onboard or sandbox create.
My first idea was simple — edit the gateway's openshell-gateway.toml and set ttl_secs = 315360000 (10 years). But onboard resets this value to 3600 every time it runs. And nemoclaw rebuild --yes also regenerates the TOML. Dead end.
I tried nemoclaw mynemo rebuild --yes to force token re-issuance. Two problems:
First, I forgot to stop the systemd gateway → Address already in use (os error 98) → sandbox destroyed. (Lesson: always stop systemd gateway before running onboard/rebuild.)
After stopping the gateway and retrying: rebuild needs to back up the sandbox state via docker exec, but the container is crash-looping → docker exec fails → "Aborting rebuild to prevent data loss." The --force flag only skips the confirmation prompt, not the backup.
With no running onboard session, --resume just says "No resumable onboarding session was found."
All three "official" re-minting paths were blocked. But I had the signing key (signing.pem, Ed25519) sitting on the host. The solution: re-sign the expired token myself.
I wrote resign_token.py:
sub, iss, aud, sandbox_id)iat to now, exp to now + 10 yearsmv would break it)Result:
token re-signed and verified OK
old exp : 1783143263
new exp : 2098506257 (+315360000s, ~10 years)
Within 20 seconds, the crash-looping container picked up the new token, authenticated successfully, and the sandbox returned to Ready.
For long-term prevention, I patched the source constant:
sed -i 's/JWT_TTL_SECS = 3600/JWT_TTL_SECS = 315360000/' \
/home/nemouser/.nemoclaw/source/dist/lib/onboard/docker-driver-gateway-config.js \
/home/nemouser/.nemoclaw/source/src/lib/onboard/docker-driver-gateway-config.ts
Now any future onboard or rebuild will issue 10-year tokens. The caveat: a NemoClaw update will overwrite this patch, but since existing tokens survive updates, the only risk is if you both update and recreate the sandbox, then wait more than an hour before the next reboot.
With the token fixed and the container no longer crash-looping, I checked the container health:
docker ps --format 'table {{.Names}}\t{{.Status}}'
# Up 3 minutes (unhealthy)
The healthcheck (curl 127.0.0.1:18789/health) was failing because the OpenClaw gateway (Node.js, port 18789) wasn't running. Looking at processes inside the container:
docker exec $CONTAINER ps aux
# PID 1: /opt/openshell/bin/openshell-sandbox (root)
# PID 38: sleep infinity (sandbox user)
That sleep infinity is the "session task" — and it's there by design. The container supervisor starts a safe idle process and waits for the host to trigger the real workload.
I tried the official command: nemoclaw mynemo gateway restart. It failed with SUPERVISOR_UNAVAILABLE. Tracing through the source:
gateway-restart.ts calls executeGatewaySupervisorAction()docker exec into the container: /usr/local/bin/nemoclaw-gateway-control restart <nonce>managed-gateway-control.py_discover_supervisor() which looks for a nemoclaw-start process as a child of PID 1sleep infinity → SUPERVISOR_UNAVAILABLEThe managed restart command requires nemoclaw-start to already be running — but nemoclaw-start is exactly what we need it to start. Another chicken-and-egg.
The nemoclaw-start script (5000+ lines) says "Runs as root (via ENTRYPOINT)" and expects to be the container's actual entrypoint. In the managed topology (where openshell-sandbox is PID 1), it doesn't apply.
docker exec -d -u sandbox $CONTAINER openclaw gateway run --port 18789
# Missing config. Run `openclaw setup` or set gateway.mode=local
The OpenClaw gateway needs its configuration from $HOME/.openclaw/openclaw.json. The sandbox user's home is /sandbox, and the config lives at /sandbox/.openclaw/openclaw.json.
openclaw gateway run --port 18789 &
# Starts successfully, then: [1]+ Stopped
The process got SIGTSTP because it tried to interact with the terminal while backgrounded.
docker exec -d -u sandbox $CONTAINER \
sh -c 'HOME=/sandbox openclaw gateway run --port 18789 > /tmp/openclaw.log 2>&1'
The keys: -d (detached, no terminal), HOME=/sandbox (config discovery), output redirected (no terminal I/O). Within 8 seconds: [gateway] ready.
I created nemoclaw-openclaw-boot.sh as an ExecStartPost in the systemd service:
docker ps --filter name=openshell-mynemo-)HOME=/sandbox trickRoot Cause 4's ExecStartPost boot script worked, but it was fragile: it hard-coded openclaw gateway run and reimplemented pieces of what the container already knows how to do. Reading the container topology more carefully, I found a better trigger.
The sleep infinity process is the supervisor's idle "session task." The real workload — nemoclaw-start, which brings up the OpenClaw gateway plus the auto-pairing wrapper and a readiness watchdog — is meant to be triggered by the host through the sandbox's exec channel:
nemoclaw mynemo exec --no-tty --timeout 0 -- nemoclaw-start
Two things make this the right primitive:
nemoclaw-start launches, the supervisor adopts it, so the process survives the death of the exec/SSH channel that started it. That's exactly what a boot trigger needs.env prefix required. The nemoclaw mynemo exec session already injects the proxy and OPENCLAW_* runtime variables, and nemoclaw-start self-derives HTTP_PROXY from NEMOCLAW_PROXY_HOST/NEMOCLAW_PROXY_PORT (defaults 10.xx.0.1/3128). A bare docker exec gets a much thinner environment — a good reminder to diagnose env from the real exec path, not docker exec.So I moved the workload out of the gateway unit into its own service, leaving openshell-gateway.service clean:
[Unit]
Description=NemoClaw sandbox workload (nemoclaw-start)
After=openshell-gateway.service
Wants=openshell-gateway.service
[Service]
Type=simple
ExecStart=/home/nemouser/.local/bin/nemoclaw mynemo exec --no-tty --timeout 0 -- nemoclaw-start
Restart=always
RestartSec=10
[Install]
WantedBy=default.target
Type=simple is correct here because the exec call stays resident for the life of the workload (verified). If the sandbox isn't ready yet, the exec returns non-zero and Restart=always retries every 10 seconds until it sticks.
With the workload service in place, the gateway still refused to stabilize. The container would come up, then thrash. The OpenClaw gateway log was a wall of:
too many failed authentication attempts
pairing required: device is not approved yet
And devices/pending.json was exploding — over 150 distinct device IDs, each one used exactly once. Something was minting a brand-new device identity on every single CLI call.
The answer was in the device identity file, /sandbox/.openclaw/identity/device.json:
"privateKeyPem": "[STRIPPED_BY_MIGRATION]"
OpenClaw's own identity-migration step had removed the private key — this was set off around the time of an earlier onboard --recreate. The consequence is a nasty deadlock:
device.json still exists, so OpenClaw won't overwrite it with a fresh identity either.pending.json, the gateway's rate limiter trips, its readiness probe fails, and the workload gets respawned — forever.This one root cause explained the whole storm. Several things I had earlier suspected — a VPN interfering, rate-limits from my own repeated restarts, missing operator scopes — were all red herrings downstream of the stripped key.
While chasing this I twice fooled myself. The gateway binds to the bridge IP, not loopback:
curl http://10.xx.0.2:18789/ # 200 OK
curl http://127.0.0.1:18789/ # connection refused (false negative)
So ss | grep 18789 and curl 127.0.0.1 both scream "nothing is listening" when the gateway is perfectly healthy. Always probe the bridge address.
The repair is done entirely outside onboard, and every touched file is backed up first:
Remove the broken identity: back up and delete identity/device.json and identity/device-auth.json, and reset devices/paired.json and devices/pending.json to {}.
Restart the workload. With no device.json in the way, OpenClaw mints a fresh, valid identity — real private key, current format — and persists it. The pending flood collapses to a single device immediately.
Break the self-approval deadlock, once. A fresh device can't approve itself: approval goes through the gateway, and the gateway only trusts already-approved devices. Since the gateway runs with dangerouslyDisableDeviceAuth=true (token-match auth), I bootstrapped the pairing by hand — writing the new device into paired.json with full operator scopes and placing a matching token in identity/device-auth.json. The gateway converged instantly.
Because /sandbox/.openclaw persists across restarts, this hand-pairing is a one-time cost. Every future reboot recovers with no deadlock.
After this, the gateway served 200s, paired:1 pending:0, the TUI completed operator RPCs (sessions.list), and inference returned real completions.
Once the gateway was healthy, the TUI still failed intermittently. The failing turns all carried the same error:
ResourceExhausted: Worker local total request limit reached (32/32)
This wasn't my bug. The managed inference route (inference.local) is a shared endpoint with a 32-concurrent-worker limit. My sandbox holds a single egress connection, yet the endpoint regularly showed 9–22 concurrent requests from other tenants. When the shared total hits 32, my turn bounces too. The telling detail: failoverReason: null — there was no fallback configured to catch it.
The mitigation is a fallback chain. Crucially, inference.local already serves NVIDIA's entire catalog (100+ models) behind the same API key and base URL, so adding models needs no new egress — external providers like api.openai.com are all blocked (egress 000). Each additional model is a separate worker with its own 32-slot limit, so a fallback genuinely dodges a saturated primary.
The catch is that openclaw config set is blocked inside the sandbox ("run onboard --resume"), so the change goes straight into the config file:
In /sandbox/.openclaw/openclaw.json, add each fallback model to models.providers.inference.models[], and set agents.defaults.model.fallbacks to the provider-prefixed refs (e.g. inference/meta/llama-3.3-70b-instruct). The primary stays untouched.
Recompute .config-hash — it's a raw sha256sum openclaw.json in the format <hash> openclaw.json. Skip this and OpenClaw rejects the file as a config mismatch.
No restart needed. OpenClaw watches the file and hot-reloads:
[reload] config hot reload applied (agents.defaults.model.fallbacks, models.providers.inference.models)
The chain I landed on: primary nvidia/nemotron-3-super-120b-a12b, then nvidia/llama-3.3-nemotron-super-49b-v1.5, meta/llama-3.3-70b-instruct, deepseek-ai/deepseek-v4-pro, and qwen/qwen3.5-122b-a10b. All five verified against inference.local with a real chat completion first (embedding/vision/safety models don't accept chat and would fail silently as fallbacks).
Because openclaw.json is a managed file that onboard/rebuild regenerate, hand-editing it is normally off-design. To stay durable I wrapped the delta in an idempotent re-apply script that only ever appends the fallback models and rewrites the hash — never touching managed values like tokens or API keys. Run it after any onboard and the chain comes back.
After implementing all five fixes, I ran wsl --shutdown and let the system come back cold. Without touching anything in the nemouser session:
docker ps --format 'table {{.Names}}\t{{.Status}}'
# Up 2 minutes (healthy)
curl -s -o /dev/null -w "%{http_code}" http://10.xx.0.2:18789/
# 200
# pairing: paired:1 pending:0 (no storm)
# identity: privateKeyPem len 119 (real key, survived reboot)
# fallbacks: 4 models loaded (openclaw.json persisted)
# openclaw tui: connected, chat completion returned real answer
Every layer recovered on its own:
openshell-gateway.service): auto-started, correct DB, port 8080 servingrestart: unless-stopped brought it back, writable layer intact (/sandbox/.openclaw untouched)nemoclaw-workload.service): triggered nemoclaw-start, reparented to PID 1, stayed residentopenclaw.jsonTwo minutes from cold boot to fully operational AI sandbox. Zero manual intervention. Zero API keys to paste. Zero 5-minute onboard runs.
I thought that was the end of the story. It wasn't. Two more bugs were hiding — one triggered by a stale file left behind by an ungraceful shutdown, the other triggered by me, trying to "fix" a healthy system.
A few reboots later, a nemoclaw mynemo connect died before it reached the gateway at all:
Error: Timed out after 30000ms waiting for shields transition lock
'/home/nemouser/.nemoclaw/state/shields-transition-lock-mynemo.json':
recorded owner PID 1544 is not running (repair mutable config permissions)
NemoClaw serializes "shields transitions" (permission/config mutations) behind a lock file. Every connect first checks it. This time the lock was held by PID 1544 — a process that no longer existed.
The lock file's contents told the whole story:
{"version":1,"sandboxName":"mynemo","pid":1544,
"processStartIdentity":"proc:1593",
"command":"repair mutable config permissions","acquiredAtMs":...}
During the previous wsl --shutdown, a process was midway through a shields transition when it got SIGKILLed. It never got to release the lock. On the next boot, PID 1544 was long gone, but the lock pathname remained.
NemoClaw deliberately won't remove a stale lock automatically — the reasoning is sound: between "inspect the lock" and "delete the lock," another process could legitimately acquire it, and blindly deleting would corrupt a real transition. So it does the safe thing and refuses, leaving a 30-second timeout wall on every connect until a human intervenes.
The manual repair is trivial — verify the owner PID is dead, then move the lock aside:
# PID 1544 not running? Then it's safe to clear.
mv /home/nemouser/.nemoclaw/state/shields-transition-lock-mynemo.json \
/home/nemouser/.nemoclaw/state/shields-transition-lock-mynemo.json.stale-bak
But doing it by hand every reboot defeats the point of automatic recovery. I wrote cleanup-stale-shields-locks.sh and wired it as an ExecStartPre on the workload service. It does exactly what NemoClaw won't do on its own, but safely:
pid from each shields-transition-lock-*.jsonos.kill(pid, 0) — a signal-0 probe that tests existence without touching the processrm)The script is idempotent: no stale locks, nothing happens.
This one I caused myself. Wanting to apply the new lock-cleanup ExecStartPre, I ran the reflexive thing:
systemctl --user restart nemoclaw-workload.service
Within a minute the container's CPU was pinned at 949% across 274 processes, climbing. docker stats showed memory ballooning. The sandbox was melting down — and I had done it to a system that was, moments earlier, perfectly healthy.
The workload service runs nemoclaw mynemo exec -- nemoclaw-start. But nemoclaw-start is the container's full entrypoint, and when a gateway is already running it doesn't cleanly no-op. The log showed it detecting the conflict:
Gateway failed to start: gateway already running (pid 497); lock timeout after 5000ms
Port 18789 is already in use.
- Another process is listening on this port.
Detecting it, then dying — without cleaning up the child processes it had already spawned (an openclaw gateway attempt, an openclaw-devices call, and an auto-pair watcher). With Restart=always on the service, systemd relaunched nemoclaw-start every 10 seconds, and each doomed attempt left another orphaned watcher behind. Those orphan watchers kept spawning openclaw children of their own. The process table grew without bound.
The real gateway (PID 497, running happily for five hours) was never the problem. My restart was. Cleanup meant killing every orphaned generation while carefully preserving the original healthy tree (identified by its parent PID), then never letting the storm start again.
The guard is simple in spirit — don't run nemoclaw-start if a gateway is already up:
# guard-gateway-not-running.sh
if docker exec "$CT" pgrep -x openclaw >/dev/null 2>&1; then
echo "gateway already running — skipping nemoclaw-start"
exit 1
fi
exit 0
The subtlety is how you attach it. My first instinct was ExecStartPre, but a failing ExecStartPre drives the unit into a failed state, which then trips Restart=always into a retry loop — the very churn I was trying to stop, now dressed up as monitoring noise.
The correct primitive is ExecCondition. When an ExecCondition script exits 1–254, systemd skips the unit cleanly:
Active: inactive (dead) (Result: exec-condition)
Condition: start condition unmet
nemoclaw-workload.service: Skipped due to 'exec-condition'.
No failure. No retry storm. No monitoring false alarm. And critically, this doesn't break reboot recovery: on a cold boot the container comes up idle (sleep infinity, no gateway), so the condition is met, nemoclaw-start runs, and the gateway comes up normally. The guard only ever fires when a gateway already exists — which is exactly when starting a second one is a mistake.
Cleaning up the storm surfaced one more casualty. The TUI now refused to connect:
gateway disconnected: unauthorized: gateway token mismatch
Every failed nemoclaw-start during the storm had run [token] Gateway auth token refreshed for startup, rewriting the gateway auth token on disk. But the surviving gateway (PID 497) had loaded its token five hours earlier and held it in memory. Disk and memory had drifted apart, and the TUI — reading the freshly-rewritten disk token — no longer matched the running gateway.
The tell is in the timestamps. Compare the gateway process start against the config file's mtime:
ps -o lstart -p <gateway-pid> # started 15:13
stat -c '%y' openclaw.json # modified 20:20 ← 5 hours later
In a healthy system these are seconds apart, because the same startup writes the token and launches the gateway that reads it. A wide gap means desync. The fix is just a clean restart — kill the stale-token gateway tree and let the guarded workload service start one fresh, so disk and memory are written by the same run. A reboot does this for free.
Both new guards attach to the same unit that Root Cause 4 introduced. The order matters: the ExecCondition runs first and can skip the whole start; only if a start is going to happen does the ExecStartPre clear stale locks; then the resident ExecStart triggers the workload. The final nemoclaw-workload.service:
[Unit]
Description=NemoClaw sandbox workload (nemoclaw-start)
After=openshell-gateway.service
Wants=openshell-gateway.service
[Service]
Type=simple
# Root Cause 7: skip the whole cycle if a gateway is already up.
# ExecCondition (not ExecStartPre) so a skip is "inactive (dead)",
# not "failed" — Restart=always never sees a failure to retry.
ExecCondition=/home/nemouser/.local/bin/guard-gateway-not-running.sh
# Root Cause 6: clear shields locks orphaned by an ungraceful shutdown.
ExecStartPre=/home/nemouser/.local/bin/cleanup-stale-shields-locks.sh
ExecStart=/home/nemouser/.local/bin/nemoclaw mynemo exec --no-tty --timeout 0 -- nemoclaw-start
Restart=always
RestartSec=10
[Install]
WantedBy=default.target
Both referenced scripts live in ~nemouser/.local/bin so systemd can reach them. guard-gateway-not-running.sh is the pgrep check shown above. cleanup-stale-shields-locks.sh is the liveness-checked lock sweeper — its core is small:
# For each shields-transition-lock-*.json in the state dir:
# read the top-level "pid"; if os.kill(pid, 0) says it's dead,
# move the lock aside with a timestamp. Never touch a live owner.
import glob, json, os, shutil, sys, time
for path in glob.glob(os.path.expanduser(
"~/.nemoclaw/state/shields-transition-lock-*.json")):
pid = json.load(open(path)).get("pid")
try:
os.kill(pid, 0) # signal 0: existence probe only
except ProcessLookupError:
shutil.move(path, f"{path}.stale-bak-{int(time.time())}")
With the lock cleanup and the ExecCondition guard in place, I did one more cold wsl --shutdown. This time I was specifically watching for the two new failure modes — and for whether the guard would let a legitimate boot through:
docker ps --format 'table {{.Names}}\t{{.Status}}'
# Up 2 minutes (healthy)
# openclaw processes: 1 (no storm — orphan guard held)
# pending / paired: 0 / 1 (no identity storm)
# privateKeyPem len: 119 (identity intact)
# gateway start vs config mtime: 3 seconds apart (token in sync)
# openclaw tui: connected on the first try
The workload service came up active (running) — the guard's condition was met (idle container, no gateway yet), so it correctly ran nemoclaw-start instead of skipping. Every one of the seven root causes now recovers untouched.
| # | Root Cause | Symptom | Fix |
|---|---|---|---|
| 1 | Gateway is a detached child process | Port 8080 dead after reboot | systemd user service + linger |
| 2 | Gateway reads wrong (empty) DB | "sandbox has no spec" | OPENSHELL_DB_URL env var |
| 3 | JWT token TTL is 1 hour | Container crash loop (43x), ExpiredSignature | Ed25519 re-signing + source patch |
| 4 | Container starts sleep infinity, not OpenClaw | Port 18789 dead, unhealthy | Dedicated nemoclaw-workload.service triggering nemoclaw-start |
| 5 | Device private key stripped by migration | pending.json flood, pairing storm, endless respawn | Fresh identity mint + one-time hand-pairing |
| 6 | Stale shields lock from ungraceful shutdown | connect times out after 30s, dead owner PID |
cleanup-stale-shields-locks.sh as ExecStartPre (liveness-checked) |
| 7 | nemoclaw-start re-run over a live gateway |
Orphan process storm, CPU 949%, token desync | Gateway guard via ExecCondition (skip, don't fail) |
A separate issue wasn't a reboot bug at all — the shared inference endpoint's 32-worker limit caused intermittent turn failures, mitigated with an NVIDIA fallback chain.
One durability caveat ties it all together: inside this container, /sandbox lives on the writable layer, not a volume, and the restart policy is unless-stopped. So a plain reboot preserves everything (identity, pairing, config edits), but a recreate/rebuild/onboard wipes /sandbox — which is exactly why the identity fix and the fallback chain each need a re-apply path.
openshell-gateway.service — systemd unit for the host gateway (ExecStartPre polls the Docker socket, OPENSHELL_DB_URL pins the correct DB)nemoclaw-workload.service — systemd unit that triggers nemoclaw-start inside the sandbox and keeps it resident (guarded by an ExecCondition and a stale-lock ExecStartPre)resign_token.py — Emergency Ed25519 JWT re-signing tool for token expiry recoveryapply-inference-fallback.sh — Idempotent re-apply of the NVIDIA fallback chain into openclaw.json after any onboardcleanup-stale-shields-locks.sh — Liveness-checked removal of shields-transition locks orphaned by an ungraceful shutdownguard-gateway-not-running.sh — ExecCondition guard that skips the workload start when a gateway is already running/proc/PID/fd when something reads the wrong data. It shows exactly which files are open, no guessing needed.mv a bind-mounted file.SUPERVISOR_UNAVAILABLE from error string through three layers of code (TypeScript → shell → Python) revealed exactly why the managed restart fails.sleep infinity session task isn't a bug; it's intentional safe-idle design. Rather than reimplement the workload in a boot script, I triggered the supervisor's own nemoclaw-start through the sandbox exec channel and let it reparent to PID 1.10.xx.0.2, so every loopback check reported it dead while it was serving 200s. A wrong-address probe is a false negative that can send you chasing ghosts./sandbox looked permanent because reboots preserved it — but it's a writable container layer, not a volume. Anything that recreates the container erases it, so every out-of-band fix needs a documented way back.os.kill(pid, 0)) and only then move it aside. The safety the tool refuses to assume, you can prove.ExecStartPre failures poison the unit; ExecCondition skips it. If a systemd guard should mean "not now, quietly," use ExecCondition — an exit of 1–254 skips the unit into inactive (dead) without tripping Restart=always into a retry storm. A failing ExecStartPre does the opposite.restart. Restarting a healthy service ran the container entrypoint over a live gateway, and the failed re-runs stormed the process table to 949% CPU. When a system is already working, prove it needs restarting before you type it.token mismatch with no obvious cause. When state lives in two places, compare their timestamps (lstart vs file mtime) before chasing the wrong layer.Disclaimer: This blog post was created with assistance from Claude, an AI developed by Anthropic, under my direct supervision and guidance to ensure accuracy and alignment with my vision for the content.