GPU Fleet Operations (read-only)
SSH diagnostics on the Lambda GPU fleet (lambda-gh200, lambda-micro-ai): vLLM/embeddings health, GPU utilization, service logs. Use when diagnosing inference/embeddings issues or asked about GPU boxes. STRICTLY read-only.
GPU Fleet Operations (read-only)
Bench sandboxes have a restricted SSH identity for the Lambda GPU fleet, preconfigured in ~/.ssh/config.
Hosts
lambda-gh200— GH200 480GB. Runs the embeddings stack (embed-nginx, embed-vllm; casemark/embed-v1 = Nemotron).lambda-micro-ai— secondary box (key pending re-install; if SSH fails, say so).
Non-negotiable rules
- READ-ONLY. These serve production inference for legal customers. Never restart services, kill processes, edit configs, or write files unless the user explicitly directs that exact action.
- The key is port-forward/agent-forward restricted by design.
Canonical diagnostics
ssh lambda-gh200 'nvidia-smi' # GPU util/memory/procs
ssh lambda-gh200 'docker ps --format "{{.Names}} {{.Status}}"' # container health
ssh lambda-gh200 'systemctl --no-pager --failed' # failed units
ssh lambda-gh200 'journalctl -u embed-vllm -n 100 --no-pager' # service logs
ssh lambda-gh200 'curl -s localhost:8000/health' # vLLM health
ssh lambda-gh200 'df -h / && free -h && uptime' # box vitals
Diagnosis pattern
Watchtower tells you WHAT failed (events/vercel_logs, partition-filtered); the fleet tells you WHY (GPU OOM, dead container, disk full). Correlate timestamps across both, report findings with evidence, and propose the fix for the user to approve.