AzureML Request Visibility

azureml-fe · default namespace · 3 replicas
loading...
Last Minute
-
Last Hour
-
Last 24 Hours
-
Last 30 Days
-

Requests Over Time

Top Services / Sub-Endpoints

ServiceRequestsErrorsAvg msP95 ms

Replica Load Distribution

PodRequests

Error Rate

-

Backend Replica Load

Load per model-serving backend pod, not the azureml-fe gateway pod - reveals skew across a deployment's own replicas. Only currently-live pods are shown; stale entries from scaled-down/restarted replicas are filtered out automatically.
Deployment / Backend PodRequestsErrorsAvg ms

Latency Breakdown

Overhead is total duration minus time actually spent in the model backend (UPSTREAM_SERVICE_TIME) - i.e. azureml-fe/envoy routing & queueing time, not model inference time.
Deploy Tag / Service Requests Avg Total ms Avg Model ms Avg Overhead ms P95 Total ms

Failure Classification

Envoy's RESPONSE_FLAGS, independent of HTTP status - distinguishes "no healthy backend", "backend timeout", "connection failure" etc. from ordinary 4xx/5xx responses.
FlagMeaningCount

Processing Share per Deployment

How much actual work each deployment did - share of total model compute time (sum of UPSTREAM_SERVICE_TIME) relative to other tracked services. This is not a cost figure: these deployments reserve their full peak CPU size at all times regardless of traffic, so a rarely-used service can barely show up here while still costing a lot - see the Cost tab for that.
Service Requests Compute Time % Share

Estimated Cost per Deployment

There's no per-pod Azure billing API - AKS bills at the node/VM level. This allocates the monthly cost you enter above by each service's share of the node pool's total CPU reserved right now (replica count × CPU request per replica) - not by how much processing it actually did. These deployments request their full peak/burst size at all times rather than a small baseline, and Kubernetes sizes nodes off requests, not real-time usage - so this is the number that actually tracks your bill. It's a live snapshot, not a time-windowed average. "Idle / Other" is node pool capacity not reserved by any tracked service. Enter the cost of just this node pool, not the whole cluster, if it hosts unrelated workloads too.
Service Replicas Reserved Cores % Share Est. $/mo

Synthetic Endpoint Checks

Every 10 minutes, healthcheck sends the sample scoring request to each active jobtype endpoint and expects a 200. A Slack alert fires on a down/recovered transition, not on every failed cycle.
URL Status Code Latency Consecutive Failures Last Checked