GPU Fleet Operations (read-only)

SSH diagnostics on the Lambda GPU fleet (lambda-gh200, lambda-micro-ai): vLLM/embeddings health, GPU utilization, service logs. Use when diagnosing inference/embeddings issues or asked about GPU boxes. STRICTLY read-only.

GPU Fleet Operations (read-only)

Bench sandboxes have a restricted SSH identity for the Lambda GPU fleet, preconfigured in ~/.ssh/config.

Hosts

  • lambda-gh200 — GH200 480GB. Runs the embeddings stack (embed-nginx, embed-vllm; casemark/embed-v1 = Nemotron).
  • lambda-micro-ai — secondary box (key pending re-install; if SSH fails, say so).

Non-negotiable rules

  • READ-ONLY. These serve production inference for legal customers. Never restart services, kill processes, edit configs, or write files unless the user explicitly directs that exact action.
  • The key is port-forward/agent-forward restricted by design.

Canonical diagnostics

ssh lambda-gh200 'nvidia-smi'                                  # GPU util/memory/procs
ssh lambda-gh200 'docker ps --format "{{.Names}} {{.Status}}"' # container health
ssh lambda-gh200 'systemctl --no-pager --failed'               # failed units
ssh lambda-gh200 'journalctl -u embed-vllm -n 100 --no-pager'  # service logs
ssh lambda-gh200 'curl -s localhost:8000/health'               # vLLM health
ssh lambda-gh200 'df -h / && free -h && uptime'                # box vitals

Diagnosis pattern

Watchtower tells you WHAT failed (events/vercel_logs, partition-filtered); the fleet tells you WHY (GPU OOM, dead container, disk full). Correlate timestamps across both, report findings with evidence, and propose the fix for the user to approve.

Danger Zone