Live Fleet Metrics — real-time from nvidia-smi + llama-server /metrics

CONNECTING
last poll: — net ↓ — · ↑ —
Model serving :8001
connecting...
Prompt throughput
tok/s
prompt tokens: —
Decode throughput
tok/s
predicted tokens: —
Request latency
ms TTFT
waiting for first poll…
Total GPU power draw
W / W
waiting for first poll…
GPU UTILIZATION ()0%
DECODE THROUGHPUT (tok/s)0
SECONDARY SERVERS — live
REASONING TAP — live CoT (optional: streamed reasoning_content, set THOUGHT_LOG)
tap idle — no reasoning streams captured yet. Optional: run a proxy that tees each request's reasoning_content to a log and set THOUGHT_LOG; each concurrent stream gets its own panel here.
NETWORK — Σ — this session · peak ↓ —
LAN (passive: arp + conns)

Model library — loadable inventory · served from fleet-metrics /models

loading registry…