Last Minute
-
Last Hour
-
Last 24 Hours
-
Last 30 Days
-
Requests Over Time
Top Services / Sub-Endpoints
| Service | Requests | Errors | Avg ms | P95 ms |
|---|
Replica Load Distribution
| Pod | Requests |
|---|
Error Rate
-
Backend Replica Load
Load per model-serving backend pod, not the azureml-fe gateway pod - reveals skew across a
deployment's own replicas. Only currently-live pods are shown; stale entries from
scaled-down/restarted replicas are filtered out automatically.
| Deployment / Backend Pod | Requests | Errors | Avg ms |
|---|
Latency Breakdown
Overhead is total duration minus time actually spent in the model backend
(UPSTREAM_SERVICE_TIME) - i.e. azureml-fe/envoy routing & queueing time,
not model inference time.
| Deploy Tag / Service | Requests | Avg Total ms | Avg Model ms | Avg Overhead ms | P95 Total ms |
|---|
Failure Classification
Envoy's
RESPONSE_FLAGS, independent of HTTP status - distinguishes "no healthy
backend", "backend timeout", "connection failure" etc. from ordinary 4xx/5xx responses.
| Flag | Meaning | Count |
|---|
Processing Share per Deployment
How much actual work each deployment did - share of total model compute time (sum of
UPSTREAM_SERVICE_TIME) relative to other tracked services. This is not
a cost figure: these deployments reserve their full peak CPU size at all times regardless of
traffic, so a rarely-used service can barely show up here while still costing a lot - see the
Cost tab for that.
| Service | Requests | Compute Time | % Share |
|---|
Estimated Cost per Deployment
There's no per-pod Azure billing API - AKS bills at the node/VM level. This allocates the
monthly cost you enter above by each service's share of the node pool's total CPU
reserved right now (replica count × CPU request per replica) - not by
how much processing it actually did. These deployments request their full peak/burst size at
all times rather than a small baseline, and Kubernetes sizes nodes off requests, not real-time
usage - so this is the number that actually tracks your bill. It's a live snapshot, not a
time-windowed average. "Idle / Other" is node pool capacity not reserved by any tracked
service. Enter the cost of just this node pool, not the whole cluster, if it hosts unrelated
workloads too.
| Service | Replicas | Reserved Cores | % Share | Est. $/mo |
|---|
Synthetic Endpoint Checks
Every 10 minutes, healthcheck sends the sample scoring request to each active
jobtype endpoint and expects a 200. A Slack alert fires on a
down/recovered transition, not on every failed cycle.
| URL | Status | Code | Latency | Consecutive Failures | Last Checked |
|---|