Skip to main content

What to monitor

Availability

Node heartbeats, registry uptime, and request success rate.

Latency

P50/P95/P99 inference duration by model and node.

Cost

Credits consumed per route, model, and tenant.

Capacity

Queue depth, GPU utilization, memory pressure, and saturation.

Alert policy

Alert immediately when success rate drops below 97% for 5 minutes or registry is unreachable.

Logs and trace correlation

Use a shared request_id across:
  • client request logs
  • registry routing logs
  • node execution logs
This allows end-to-end failure triage in one query.
Last modified on February 21, 2026