What to monitor
Availability
Node heartbeats, registry uptime, and request success rate.
Latency
P50/P95/P99 inference duration by model and node.
Cost
Credits consumed per route, model, and tenant.
Capacity
Queue depth, GPU utilization, memory pressure, and saturation.
Recommended metrics
Alert policy
- Critical
- Warning
- Capacity
Alert immediately when success rate drops below 97% for 5 minutes or registry is unreachable.
Logs and trace correlation
Use a sharedrequest_id across:
- client request logs
- registry routing logs
- node execution logs