For the complete documentation index, see llms.txt. Markdown versions of all docs pages are available by appending .md to any docs URL.
Outlier detection
Eject misbehaving upstream hosts from the load balancing pool when consecutive errors occur.
Configure passive health checks and remove unhealthy hosts from the load balancing pool with an outlier detection policy.
About outlier detection
Outlier detection is an important part of building resilient apps. An outlier detection policy sets up several conditions, such as retries and ejection percentages, that kgateway uses to determine if a service is unhealthy. In case an unhealthy service is detected, the outlier detection policy defines how the service is removed from the pool of healthy destinations to send traffic to. Your apps then have time to recover before they are added back to the load-balancing pool and checked again for consecutive errors.
Before you begin
-
Follow the Get started guide to install kgateway.
-
Follow the Sample app guide to create a gateway proxy with an HTTP listener and deploy the httpbin sample app.
-
Get the external address of the gateway and save it in an environment variable.
export INGRESS_GW_ADDRESS=$(kubectl get svc -n kgateway-system http -o jsonpath="{.status.loadBalancer.ingress[0]['hostname','ip']}") echo $INGRESS_GW_ADDRESS
Set up outlier detection
-
Scale the httpbin app to 2 replicas.
kubectl scale deploy/httpbin -n httpbin --replicas=2 -
Verify that you see two replicas of the httpbin app.
kubectl get pods -n httpbinExample output:
NAME READY STATUS RESTARTS AGE httpbin-577649ddb-lsgp8 2/2 Running 0 31d httpbin-577649ddb-q9b92 2/2 Running 0 3s -
Send a few requests to the httpbin app. Because both httpbin replicas are exposed under the same service, requests are automatically load balanced between all healthy replicas.
for i in {1..5}; do curl -vi http://$INGRESS_GW_ADDRESS:8080/status/200 -H "host: www.example.com:8080" ; doneExample output for one request:
* Request completely sent off < HTTP/1.1 200 OK HTTP/1.1 200 OK < access-control-allow-credentials: true access-control-allow-credentials: true < access-control-allow-origin: * access-control-allow-origin: * < content-length: 0 content-length: 0 < x-envoy-upstream-service-time: 1 x-envoy-upstream-service-time: 1 < server: envoy server: envoy < .... -
Review the logs for both replicas. Verify that you see log entries for the 5 requests spread across both replicas. For example, one replica might have 2 log entries and the other one 3.
kubectl logs -l app=httpbin -n httpbin -f --prefixExample output for one request:
time="2025-09-15T21:11:52.8514" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.03 user_agent="curl/8.7.1" client_ip=10.X.X.XX -
Create a BackendConfigPolicy with your outlier detection policy. The following example ejects an unhealthy upstream host for one hour if the host returns one 5XX HTTP response code. Note that the maximum number of ejected hosts is set to 80%. Because of that, only one replica of httpbin can be ejected at any given time. If two replicas were ejected, that would exceed the 80% maximum threshold (because the two replicas equal 100%).
kubectl apply -f- <<EOF kind: BackendConfigPolicy apiVersion: gateway.kgateway.dev/v1alpha1 metadata: name: httpbin-policy namespace: httpbin spec: targetRefs: - name: httpbin group: "" kind: Service outlierDetection: interval: 2s consecutive5xx: 1 baseEjectionTime: 1h maxEjectionPercent: 80 EOFSetting Description intervalThe time interval after which the hosts are evaluated to determine if they are healthy or not. In this example, the hosts are evaluated every 2 seconds. If not set, this field defaults to 10s.consecutive5xxThe number of consecutive server-side error responses, such as 5XX HTTP response codes for HTTP traffic and connection failures for TCP traffic, before a host is ejected from the load balancing pool. In this example, you remove the host when one 5XX HTTP response code is returned. If not set, ejection occurs after 5 consecutive errors by default. If this field is set to 0, passive health checks are disabled. baseEjectionTimeThe duration that a host is removed from the load balancing pool before a new evaluation starts. If not set, this field defaults to 30s.maxEjectionPercentThe maximum percent of hosts that can be ejected from the load balancing pool. In this example, 80% of all hosts can be ejected at a given time. If not set, this field defaults to 10percent. -
Repeat the requests to the httpbin app. In the log output for both httpbin replicas, verify that all requests are still spread across both httpbin instances.
for i in {1..5}; do curl -vi http://$INGRESS_GW_ADDRESS:8080/status/200 -H "host: www.example.com:8080" ; done -
Force one httpbin replica to return a 503 HTTP response code. This response code triggers the outlier detection policy and automatically removes this httpbin replica from the load balancing pool for 1 hour.
curl -vik http://$INGRESS_GW_ADDRESS:8080/status/503 -H "host: www.example.com:8080"Example output:
* Request completely sent off < HTTP/1.1 503 Service Unavailable HTTP/1.1 503 Service Unavailable < access-control-allow-credentials: true access-control-allow-credentials: true < access-control-allow-origin: * ... -
Send a few more requests to the httpbin app. In the logs for both replicas, verify that all requests now only go to the instance that is still considered healthy.
for i in {1..5}; do curl http://$INGRESS_GW_ADDRESS:8080/status/200 -H "host: www.example.com:8080" ; doneExample log output of the healthy instance:
time="2025-09-15T21:19:02.0808" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.03 user_agent="curl/8.7.1" client_ip=10.X.X.XX time="2025-09-15T21:19:02.2452" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.03 user_agent="curl/8.7.1" client_ip=10.X.X.XX time="2025-09-15T21:19:02.4053" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.04 user_agent="curl/8.7.1" client_ip=10.X.X.XX time="2025-09-15T21:19:02.6067" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.01 user_agent="curl/8.7.1" client_ip=10.X.X.XX time="2025-09-15T21:19:02.7604" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.01 user_agent="curl/8.7.1" client_ip=10.X.X.XXExample log output of the unhealthy instance:
time="2025-09-15T21:17:25.4035" status=503 method="GET" uri="/status/503" size_bytes=0 duration_ms=0.04 user_agent="curl/8.7.1" client_ip=10.X.X.XX -
Force the second httpbin replica to return a 503 HTTP response code. Note that the outlier detection only allows 80% of all upstream hosts to be ejected at a given time. Since both replicas would equal 100%, the outlier detection does not remove the host from the load balancing pool. The instance is still considered healthy and can receive requests. In your log output for the healthy instance, verify that you see the log entry for the 503 request.
curl -vik http://$INGRESS_GW_ADDRESS:8080/status/503 -H "host: www.example.com:8080"Example log output of the healthy instance:
time="2025-09-15T21:20:27.1117" status=503 method="GET" uri="/status/503" size_bytes=0 duration_ms=0.02 user_agent="curl/8.7.1" client_ip=10.X.X.XX -
Send a few more requests to the httpbin app. In the logs for both replicas, verify that all requests are still routed to the same instance as the instance was not removed from the load balancing pool.
for i in {1..5}; do curl http://$INGRESS_GW_ADDRESS:8080/status/200 -H "host: www.example.com:8080" ; doneExample log output for the healthy instance:
# Previous log output time="2025-09-15T21:20:27.1117" status=503 method="GET" uri="/status/503" size_bytes=0 duration_ms=0.02 user_agent="curl/8.7.1" client_ip=10.0.9.76 # New log output time="2025-09-15T21:25:11.4236" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.02 user_agent="curl/8.7.1" client_ip=10.0.15.215 time="2025-09-15T21:25:11.5833" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.03 user_agent="curl/8.7.1" client_ip=10.0.9.76 time="2025-09-15T21:25:11.7473" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.03 user_agent="curl/8.7.1" client_ip=10.0.9.76 time="2025-09-15T21:25:11.9098" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.01 user_agent="curl/8.7.1" client_ip=10.0.15.215 time="2025-09-15T21:25:12.0824" status=200 method="GET" uri="/status/200" size_bytes=0 duration_ms=0.01 user_agent="curl/8.7.1" client_ip=10.0.9.76 -
Port-forward the Gateway pod on port 19000.
kubectl port-forward deploy/http -n kgateway-system 19000 -
Open the stats Prometheus endpoint and look for the following metrics.
envoy_cluster_outlier_detection_ejections_consecutive_5xx: The number of times a host qualified for ejection. In this example, the number is 2, because both hosts qualified for ejection.envoy_cluster_outlier_detection_ejections_enforced_consecutive_5xx: The number of times an ejection was forced. In this example, the number is 1, because the second host instance could not be ejected as it did not meet the maximum percentage setting in your outlier policy.
Example output:
envoy_cluster_outlier_detection_ejections_consecutive_5xx{envoy_cluster_name="kube_httpbin_httpbin_8000"} 2 envoy_cluster_outlier_detection_ejections_enforced_consecutive_5xx{envoy_cluster_name="kube_httpbin_httpbin_8000"} 1
Separate local-origin failures from 5xx responses
By default, outlier detection counts 5xx responses from the upstream and connection failures toward the same threshold. This behavior is a problem when your upstream returns 5xx responses during normal operation. For example, a gRPC service can return a status such as UNAVAILABLE, which outlier detection treats as an HTTP 503 error.
A local-origin failure is a request for which the upstream never sends a response, such as a refused connection, a connection reset, or a timeout. An external 5xx response is a failure that the upstream itself reports. With the default settings, outlier detection counts both toward the consecutive5xx threshold. You can instruct the gateway proxy to count these types of failures separately, and eject a host only for local-origin failures.
-
Update the BackendConfigPolicy to count local-origin failures separately from external 5xx responses. The following policy ejects a host for 1 hour after 10 consecutive local-origin failures. A host is never ejected for 5xx responses from the upstream. Only 80% of all hosts can be ejected at the same time.
kubectl apply -f- <<EOF kind: BackendConfigPolicy apiVersion: gateway.kgateway.dev/v1alpha1 metadata: name: httpbin-policy namespace: httpbin spec: targetRefs: - name: httpbin group: "" kind: Service outlierDetection: interval: 2s baseEjectionTime: 1h maxEjectionPercent: 80 consecutive5xx: 1 enforcingConsecutive5xx: 0 splitExternalLocalOriginErrors: true consecutiveLocalOriginFailure: 10 enforcingConsecutiveLocalOriginFailure: 100 EOFSetting Description interval,baseEjectionTime,maxEjectionPercentThe same settings as in the previous section. consecutive5xxThe number of consecutive 5xx responses from the upstream before a host qualifies for ejection. In this example, one 5xx response is enough to qualify for ejection. Ejection is only enforced if enforcingConsecutive5xxis set to a value greater than 0.enforcingConsecutive5xxThe percentage chance that a host is ejected after it reaches the consecutive5xxthreshold. In this example, the value is0, so 5xx responses from the upstream never eject a host. If omitted, the field defaults to100. The value must be between0and100.splitExternalLocalOriginErrorsCounts local-origin failures separately from errors that the upstream sends. If omitted, the field defaults to false, and the settings for local-origin failures have no effect.consecutiveLocalOriginFailureThe number of consecutive local-origin failures before a host is ejected. In this example, a host is ejected after 10 failures. This field takes effect only when splitExternalLocalOriginErrorsistrue. If omitted, the field defaults to5. If the value is0, ejection for local-origin failures is disabled.enforcingConsecutiveLocalOriginFailureThe percentage chance that a host is ejected after it reaches the consecutiveLocalOriginFailurethreshold. This field takes effect only whensplitExternalLocalOriginErrorsistrue. If omitted, the field defaults to100. The value must be between0and100. -
Verify that the policy is accepted and attached.
kubectl get backendconfigpolicy httpbin-policy -n httpbin -o jsonpath='{.status.ancestors[0].conditions[*].message}'Example output:
Policy accepted Attached to all targets -
Force a 503 HTTP response code from the httpbin app. This response comes from the upstream, so it does not eject the host. Both httpbin replicas serve the 503 requests.
for i in {1..3}; do curl -vik http://$INGRESS_GW_ADDRESS:8080/status/503 -H "host: www.example.com:8080"; done -
Send a few more requests to the httpbin app along the
/status/200path. In the logs for both replicas, verify that requests are still spread across both instances.for i in {1..10}; do curl http://$INGRESS_GW_ADDRESS:8080/status/200 -H "host: www.example.com:8080" ; donekubectl logs -l app=httpbin -n httpbin --prefix -
Create an HTTPRoute that sends requests on the
/refusedpath to port 9000 of the httpbin service. The httpbin app does not listen on this port, so the gateway proxy cannot connect and records a local-origin failure.kubectl apply -f- <<EOF apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: httpbin-refused namespace: httpbin spec: parentRefs: - name: http namespace: kgateway-system hostnames: - www.example.com rules: - matches: - path: type: PathPrefix value: /refused backendRefs: - name: httpbin port: 9000 EOF -
Send enough requests to the
/refusedpath for each replica to reach the threshold of 10 consecutive failures. Every request returns a 503 HTTP response code. The gateway proxy generates this response itself because it cannot connect to the upstream, so the failure counts as a local-origin failure and not as a 5xx response from the upstream.for i in {1..30}; do curl -s -o /dev/null -w "%{http_code}\n" http://$INGRESS_GW_ADDRESS:8080/refused -H "host: www.example.com:8080"; done -
Port-forward the Gateway pod on port 19000.
kubectl port-forward deploy/http -n kgateway-system 19000 -
Open the stats Prometheus endpoint and look for the following metrics.
envoy_cluster_outlier_detection_ejections_enforced_consecutive_5xx: The number of ejections for 5xx responses from the upstream. In this example, the number for thekube_httpbin_httpbin_8000cluster is 1 from the ejection in the previous section. The number does not increase after the 503 responses in step 3, becauseenforcingConsecutive5xxis0.envoy_cluster_outlier_detection_ejections_enforced_consecutive_local_origin_failure: The number of ejections for local-origin failures. In this example, the number is 1 for thekube_httpbin_httpbin_9000cluster, because one replica reached the failure threshold and was ejected.envoy_cluster_outlier_detection_ejections_overflow: The number of times the maximum ejection percentage prevented an ejection. In this example, the number is greater than 0 because the second replica also reached the threshold, but ejecting it would exceed themaxEjectionPercentof 80%.
Example output:
envoy_cluster_outlier_detection_ejections_enforced_consecutive_5xx{envoy_cluster_name="kube_httpbin_httpbin_8000"} 1 envoy_cluster_outlier_detection_ejections_enforced_consecutive_local_origin_failure{envoy_cluster_name="kube_httpbin_httpbin_9000"} 1 envoy_cluster_outlier_detection_ejections_overflow{envoy_cluster_name="kube_httpbin_httpbin_9000"} 2
Cleanup
-
Scale down the httpbin app to 1 replica.
kubectl scale deploy/httpbin -n httpbin --replicas=1 -
Remove the BackendConfigPolicy.
kubectl delete backendconfigpolicy httpbin-policy -n httpbin -
Remove the HTTPRoute for the
/refusedpath that you created in the local-origin section.kubectl delete httproute httpbin-refused -n httpbin