Loading...

XML

Word

Printable

Type: Bug
Resolution: Done
Priority: Major
Fix Version/s: None
Affects Version/s: 4.10
Component/s: Networking / router
Labels:
None

Severity:
Important
Regression:
None
Story Points:
2
Sprint:
Sprint 225, Sprint 226
sprint_count:
2
Blocked:
False
Blocked Reason:

Hide

None

Show
None
Target Version:

4.10.z

SFDC Cases Counter:
SFDC Cases Open:
SFDC Cases Links:

Manually mirroring for backport from https://bugzilla.redhat.com/show_bug.cgi?id=2076297
Description of problem:
For brief window while the openshift-router binary is starting up, it ignores shutdown signals (SIGTERMs) and will never shutdown.

This becomes a larger issue when K8S sends a graceful shutdown while the router is starting up and subsequently waits the terminationGracePeriodSeconds as specified in the router deployment, which is 1 hour.

This becomes even more of an issue with
https://github.com/openshift/cluster-ingress-operator/pull/724
which makes the ingress controller wait for all pods before deleting itself. So if these pods are stuck in Terminating for an hour, then the ingress controller will be stuck in Terminating for an hour.

OpenShift release version:

Cluster Platform:

How reproducible:
You can start/stop the router pod quickly to get it to be stuck in a hour-long Terminating state.

Steps to Reproduce (in detail):
1. Create a YAML file with the following content:

apiVersion: v1
items:

apiVersion: operator.openshift.io/v1
kind: IngressController
metadata:
name: loadbalancer
namespace: openshift-ingress-operator
spec:
replicas: 1
routeSelector:
matchLabels:
type: loadbalancer
endpointPublishingStrategy:
type: LoadBalancerService
nodePlacement:
nodeSelector:
matchLabels:
node-role.kubernetes.io/worker: ""
status: {}
kind: List
metadata:
resourceVersion: ""
selfLink: ""

2. Run the following command:

oc apply -f <YAML_FILE>.yaml && while ! oc get pod -n openshift-ingress | grep -q router-loadbalancer; do echo "Waiting"; done; oc delete pod -n openshift-ingress $(oc get pod -n openshift-ingress --no-headers | grep router-loadbalancer | awk '{print $1}');

It is considered a failure if it hangs for more than 45 seconds. You can ctrl-c after it deletes the pod and run "oc get pods -n openshift-ingress" to see that it is stuck in a terminating state with a AGE longer than 45 seconds.

The pod will take 1 hour to terminate, but you can always clean up by force deleting it.

Actual results:
Pod takes 1 hour to be deleted.

Expected results:
Pod should be deleted in about 45 seconds.

Impact of the problem:
Router pods hang in terminating for 1 hour and that will affect user experience.

Additional info:

blocks

OCPBUGS-1620 Bug 2076297 - Router process ignores shutdown signal while starting up

Closed

clones

OCPBUGS-1618 Bug 2076297 - Router process ignores shutdown signal while starting up

Closed

is blocked by

OCPBUGS-1618 Bug 2076297 - Router process ignores shutdown signal while starting up

Closed

is cloned by

OCPBUGS-1620 Bug 2076297 - Router process ignores shutdown signal while starting up

Closed

links to

openshift/router#406: [release-4.10] Bug 2098230: Fix gap in router's handling of graceful shutdowns.

Assignee:: Grant Spence

Reporter:: Grant Spence

QA Contact:: Hongan Li

Votes:: 0 Vote for this issue

Watchers:: 2 Start watching this issue

Created:: 2022/09/21 9:42 PM

Updated:: 2022/10/15 9:33 PM

Resolved:: 2022/10/12 8:14 AM

Details

Description

Attachments

Issue Links

Easy Agile Planning Poker

Activity

People

Dates