Node operating system upgrading method and device

By generating an upgrade template and displaying the upgrade status in real time, the problem of lag in upgrading the node operating system in the Kubernetes cluster is solved, and a smoother and more controllable upgrade process is achieved.

CN120075048APending Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311623571.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In Kubernetes cluster, upgrade lags may occur during the upgrade of the node operating system, which poses certain risks, and it is difficult for cluster administrators to monitor the upgrade progress in real time.

Method used

Provide a node operating system upgrade method, which generates an upgrade template by obtaining configuration information, realizes elegant offline and drainage of the node operating system, and displays the upgrade status in real time, including node drainage status, reasons for upgrade failure, upgrade progress, etc.

Benefits of technology

It realizes that the cluster administrator can view the entire process of node operating system upgrade in real time, reduces the risk of service interruption, improves the fluency of upgrades, and pauses or resumes upgrades according to the computing resource situation through dynamic adjustment of the upgrade strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075048A_ABST
    Figure CN120075048A_ABST
Patent Text Reader

Abstract

The invention discloses a node operating system upgrading method and device, and relates to the technical field of terminals, a cluster administrator can check the whole upgrading process of a node operating system in real time, and monitoring is facilitated to deal with emergencies. The method comprises the following steps: a node operating system upgrade controller obtains configuration information; the configuration information comprises isolated pod failure information and / or single pod automatic amplification information and / or default pod interruption budget template information; generating a first template according to the configuration information, and upgrading the node operating system according to the first template; nodes running in the node operating system comprise isolated pods or single-copy pods or pods without pod interrupt budget protection; displaying the upgrading state of the node; the upgrading state of the node comprises at least one of a node drainage state, a node upgrading failure reason, a proportion of the upgraded node to the total node and whether the node upgrading is successful or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and in particular, to a method and device for upgrading a node operating system. Background Art

[0002] In cloud native technology, Kubernetes (k8s) plays a crucial role and is an important container management system. The software versions of the entire Kubernetes system have an important impact on the stable operation of containers within the cluster. There are corresponding hot upgrade methods for software versions directly related to the system within the Kubernetes cluster, such as the Kubernetes application programming interface server (kube-apiser) and the node agent program (kubelet), which can be upgraded without affecting the operation of the container group (pod) on the cluster nodes.

[0003] In related technologies, operation can be used to control the upgrade of the cluster node operating system. However, in this way, there may be a phenomenon of upgrade jamming during the upgrade process, posing a certain risk. Summary of the Invention

[0004] This application provides a method and device for upgrading a node operating system. The cluster administrator can view the entire process of the node operating system upgrade in real time, which is convenient for monitoring to cope with emergencies.

[0005] To achieve the above objective, the embodiments of this application provide the following technical solutions:

[0006] In a first aspect, this application provides a method for upgrading a node operating system. The method includes: obtaining configuration information; the configuration information includes orphaned pod failure information and / or single-pod auto-scaling information and / or default pod disruption budget template information; generating a first template according to the configuration information, and upgrading the node operating system according to the first template (node pool graceful drain CR); the nodes running in the node operating system include orphaned pods or single-copy pods or pods without pod disruption budget protection; displaying the upgrade status of the nodes; the upgrade status of the nodes includes at least one of the status of node draining, the reason for node upgrade failure, the proportion of upgraded nodes in the total number of nodes, and whether the node upgrade is successful.

[0007] In this way, the embodiments of this application can not only display whether the node upgrade is successful, but also display information such as the status of node draining, the reason for node upgrade failure, and the proportion of upgraded nodes in the total number of nodes (node pool upgrade progress). The cluster administrator can view the entire process of the node operating system upgrade in real time, which is convenient for monitoring to cope with emergencies.

[0008] In a possible implementation, the method further includes: when there are isolated pods in a node, stopping the draining operation of the node; when there are pods without pod disruption budget protection in a node, creating a pod disruption budget for the pods; when there are single-copy pods in a node, scaling the single-copy pods to at least two-copy pods.

[0009] In this way, special pods in the node are processed, the risk of service interruption is reduced, and the smoothness of node operating system upgrade is improved.

[0010] In a possible implementation, the configuration information further includes at least one of grace period second information, ignore all daemonset information, stop rolling information, and global timeout information.

[0011] Optionally, the configuration information may further include other information related to node operating system upgrade, which is not specifically limited in the embodiments of the present application.

[0012] In a possible implementation, upgrading the operating system of a node according to a first template includes: generating a second template (node draining CR) according to the first template; the second template is used to perform a draining operation on the node.

[0013] Wherein, one node pool corresponds to one first template, and one node corresponds to one second template. Draining the node according to the second template to complete the upgrade of the node operating system.

[0014] In a possible implementation, the method further includes: when the computing resources of the electronic device are less than a first threshold and the number of nodes being upgraded is greater than a second threshold, changing the stop rolling information in the configuration information and stopping the node operating system upgrade operation; when the computing resources of the electronic device are greater than the second threshold, restoring the stop rolling information in the configuration information and performing the node operating system upgrade operation.

[0015] In this way, the node operating system upgrade can be paused and resumed at any time according to the computing resources of the electronic device and the current node operating system upgrade situation, realizing visual and controllable cluster node OS upgrade process, and preventing the node operating system from freezing due to insufficient computing resources of the electronic device.

[0016] In a possible implementation, after the node operating system upgrade is completed, the method further includes: deleting the created pod disruption budget and the replicas of the scaled pods.

[0017] In this way, the node state of the k8s cluster can be ensured to remain unchanged.

[0018] In a possible implementation, the states of node drainage include at least one of: the state of cleaning logs, the state of draining, the state of failure in draining or cleaning logs, and the state of successful draining and cleaning of logs.

[0019] In a second aspect, the present application provides a node operating system upgrade device, including: a receiving module for obtaining configuration information; the configuration information includes orphan pod failure information and / or single-pod auto-scaling information and / or default pod disruption budget template information; a processing module for generating a first template according to the configuration information and upgrading the node operating system according to the first template; the nodes running in the node operating system include orphan pods or single-replica pods or pods without pod disruption budget protection; a display module for displaying the upgrade status of the node; the upgrade status of the node includes at least one of the state of node drainage, the reason for node upgrade failure, the proportion of upgraded nodes in the total number of nodes, and whether the node upgrade is successful.

[0020] In a third aspect, the present application provides an electronic device, including: a processor, a display, and a memory, the memory is coupled to the processor, and the memory is used to store computer program code, the computer program code includes computer instructions, when the processor reads the computer instructions from the memory, so that the electronic device executes the method according to any one of the first aspect.

[0021] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored, when the instructions run on an electronic device, so that the electronic device executes the method according to any one of the first aspect.

[0022] In a fifth aspect, the present application provides a chip system, including at least one processor and at least one interface circuit, the at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor, when the at least one processor executes the instructions, the at least one processor executes the method according to any one of the above first aspect.

[0023] In a sixth aspect, the present application further provides a computer program product, including instructions, when it runs on a computer, so that the computer executes the method according to any one of the first aspect.

[0024] In a seventh aspect, an embodiment of the present application provides a circuit system, the circuit system includes a processing circuit, and the processing circuit is configured to execute the method according to the first aspect or any one of the implementation manners of the first aspect.

[0025] For the technical effects corresponding to the second aspect to the seventh aspect and any one of the implementation manners corresponding to the second aspect to the seventh aspect, reference may be made to the technical effects corresponding to the first aspect and any one of the implementation manners of the first aspect above, which will not be elaborated here. Description of the Drawings

[0026] Figure 1 Schematic structural diagram of a node operating system upgrade system provided by an embodiment of the present application;

[0027] Figure 2 Schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0028] Figure 3 Schematic module diagram of an electronic device provided by an embodiment of the present application;

[0029] Figure 4 Schematic flowchart of a node operating system upgrade method provided by an embodiment of the present application;

[0030] Figure 5 Schematic interface diagram of an electronic device provided by an embodiment of the present application;

[0031] Figure 6 Schematic flowchart of another node operating system upgrade method provided by an embodiment of the present application;

[0032] Figure 7 Schematic flowchart of another node operating system upgrade method provided by an embodiment of the present application;

[0033] Figure 8 Schematic interface diagram of another electronic device provided by an embodiment of the present application;

[0034] Figure 9 Schematic interface diagram of another electronic device provided by an embodiment of the present application;

[0035] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0036] In the description of the embodiments of the present application, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0037] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0038] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more. The "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.

[0039] First, some technical terms related to the embodiments of the present application are introduced:

[0040] Container: Container technology can create user space instances based on the operating system-level virtualization technology. The created user space instances are called containers. Applications (APPs) can run inside the containers. The applications running in different containers usually use different software and hardware resources to isolate the resources between the applications and improve the security of the applications. The software and hardware resources include, but are not limited to, any one or more of the following: central processing unit (CPU), memory, cache, register, peripheral device, CPU runtime, input output (IO), network, network bandwidth, file system, process, user, hostname. Exemplarily, APP1 and APP2 run in container 1, and APP3 runs in container 2. The system allows container 1 to use resource 1 and allows container 2 to use resource 2. Then, APP1 and APP2 can use some of the resources in resource 1, and APP3 can use some of the resources in resource 2.

[0041] Kubernetes: Kubernetes is an open-source mechanism for managing containerized applications on multiple hosts in a cloud platform. Kubernetes provides a mechanism for application deployment, planning, updating, and maintenance.

[0042] Node: It refers to a single device in a cluster. A node can be abstracted as a set of available central processing units (CPU for short) and memory resources. A node can be a physical device or a virtualized device. A node can be located locally, such as in a data center. A node can also be located in the cloud. For example, a node can be hosted on a cloud platform. Optionally, an agent program can run on the node, which is used to manage the container instances running on the node.

[0043] A Kubernetes cluster includes a master node and at least one worker node. Among them, the master node is used to control and manage the worker nodes, and the worker nodes are used to run applications to provide services externally. Specifically, the master node creates a pod (container group) on the worker node, where the pod contains at least one container, and each container runs a process of the same application.

[0044] Node drain: Node drain is an operation to remove a node from the cluster. This removal process includes evicting the pods running on the node to other available nodes for node maintenance, upgrade, or shutdown. Specifically, mark the node as unschedulable and evict the pods on the node to other available nodes (such as the newly upgraded node) until all pods have been successfully evicted, and then perform maintenance operations on the node. Instructions issued by the kubectl command or the Kubernetes API can indicate the execution of the node drain operation.

[0045] Node offline: Node offline is an operation to directly remove a node from the cluster, usually used in emergency situations. Different from node drain, node offline does not evict the pods on the node to other available nodes, but directly deletes the node from the cluster. This may cause the pods running on the node to malfunction. Therefore, it is necessary to manually evict the pods on the node to other available nodes before node offline, or ensure that the application has sufficient redundant pods and high availability so that the application can still run normally after the node is deleted.

[0046] Pod disruption budget (PDB): PDB is a resource object in Kubernetes used to manage and control the impact of interruptible events on pods in the cluster. When performing cluster maintenance, node upgrade, or other interference operations, PDB can define the maximum number of pods that are allowed to be interrupted to ensure the stability and availability of the application.

[0047] Orphan pod: It refers to a pod in a Kubernetes cluster that has lost its associated controller (such as ReplicaSet, Replication Controller, or DaemonSet). When the controller associated with a pod is deleted or updated, and the pod is not associated with any new controller, the pod becomes an orphan pod. Optionally, it can be determined whether a pod is an orphan pod by checking whether the metadata.ownerReferences of the pod is empty. Specifically, if metadata.ownerReferences is empty, it indicates that the pod is an orphan pod; if metadata.ownerReferences is not empty, it indicates that the pod is not an orphan pod.

[0048] Single pod: It refers to a pod with a single replica, that is, the pod contains one replica. During the node OS upgrade process, the pod replica is deleted, resulting in the interruption of the business logic running in the pod and the node OS upgrade getting stuck.

[0049] In the related art, users can upgrade the node operating system (OS) manually, and query the progress of the node OS rolling upgrade by executing command lines within the cluster. When there are few cluster nodes, this method can observe the upgrade progress at all times to handle emergencies. For example, in the provided node automatic upgrade solution, the start and pause of the upgrade can be manually controlled, and the upgrade time window can be set.

[0050] Optionally, the operation resource can be used to control the upgrade of the cluster node pool. By operating on this resource, the upgrade of the node pool can be started or stopped, the node pool upgrade strategy can be selected, and it can be checked whether the node pool upgrade is completed. The k8s engine (Google Kubernetes Engine, GKE) provides command-line execution operations and operations through the interface. Specifically, the commands used can be as follows:

[0051] 1. To start the upgrade of the node pool, use the following command:

[0052] gcloud container node-pools update NODE_POOL_NAME\

[0053] --cluster CLUSTER_NAME\

[0054] --zone COMPUTE_ZONE\

[0055] --enable-autoupgrade

[0056] 2. Check the node upgrade status using the following command:

[0057] gcloud container operations list

[0058] gcloud container operations describe OPERATION_ID

[0059] 3. Check the node pool upgrade settings using the following command:

[0060] gcloud container node-pools describe NODE_POOL_NAME--cluster=CLUSTER_NAME

[0061] 4. Disable the node pool upgrade function using the following command:

[0062] gcloud container node-pools update NODE_POOL_NAME\

[0063] --cluster CLUSTER_NAME\

[0064] --zone COMPUTE_ZONE\

[0065] --no-enable-autoupgrade

[0066] However, in actual situations, the nodes of enterprise-level k8s clusters are often in the thousands or tens of thousands, and it is difficult to manually upgrade and monitor them. Moreover, by executing commands to query the progress of the node OS rolling upgrade, one can only see whether the node pool upgrade is completed or not, and cannot view the specific upgrade progress in real time, which is inconvenient for operation and maintenance. In addition, if the k8s cluster is too large and the underlying computing resources are insufficient, upgrading too many nodes may cause upgrade delays.

[0067] To solve the above problems, the embodiment of the present application proposes a method for upgrading the node operating system. In this method, the K8s cluster administrator can configure the configuration information related to the node OS upgrade and apply this configuration information to the K8s cluster. The configuration information includes the name of the node pool to be upgraded, the processing method of a single pod of the nodes in the node pool, the default PDB, and other information. Then, through the graceful draining custom resource of the node pool and the node draining custom resource, the node OS is automatically upgraded. And the node OS upgrade status is fed back in the interface, which is convenient for the cluster administrator to view and control the node OS upgrade progress.

[0068] Optionally, the method in the embodiments of the present application can also complete node replacement, node pool replacement, etc. in the node OS.

[0069] In a possible implementation manner, the embodiments of the present application can be applied to an electronic device. The electronic device can be as Figure 1 shown. The system includes an electronic device 100. The electronic device can implement the above-mentioned node operating system upgrade method. The electronic device can be a personal computer (PC), a mobile phone, a tablet computer (Pad), a laptop computer, a desktop computer, a laptop computer, a computer with transceiver functions, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, a wearable device, an in-vehicle device and other terminal devices. The embodiments of the present application do not impose special restrictions on the specific form of the electronic device.

[0070] Optionally, the cluster administrator can perform operations such as node OS upgrade, node replacement, node pool replacement, etc. on the electronic device through the above method.

[0071] Figure 2 The following shows a schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. The electronic device includes at least one processor 201, a communication line 202, a memory 203, and at least one communication interface 204. Among them, the memory 203 can also be included in the processor 201.

[0072] The processor 201 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0073] The communication line 202 may include a path for transmitting information between the above components.

[0074] The communication interface 204 is used for communicating with other devices. In the embodiments of the present application, the communication interface may be a module, a circuit, a bus, an interface, a transceiver, or other devices capable of implementing communication functions, and is used for communicating with other devices. Optionally, when the communication interface is a transceiver, the transceiver may be an independently provided transmitter, which can be used to send information to other devices, or the transceiver may also be an independently provided receiver, which is used to receive information from other devices. The transceiver may also be a component integrating the functions of sending and receiving information. The embodiments of the present application do not limit the specific implementation of the transceiver.

[0075] The memory 203 can be a volatile memory, a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM), or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited thereto. The memory can exist independently and be connected to the processor 201 through the communication line 202. The memory 203 can also be integrated with the processor 201.

[0076] Among them, the memory 203 is used to store computer execution instructions for implementing the solution of this application, and is controlled by the processor 201 to execute. The processor 201 is used to execute the computer execution instructions stored in the memory 203, so as to implement the carrier sending method provided in the following embodiments of this application.

[0077] Optionally, the computer execution instructions in the embodiments of this application can also be referred to as application code, instructions, computer programs, or other names, and the embodiments of this application do not make specific limitations thereto.

[0078] In a specific implementation, as an embodiment, the processor 201 can include one or more CPUs, such as Figure 1 CPU0 and CPU1 in

[0079] In a specific implementation, as an embodiment, the electronic device can include multiple processors, such asFigure 2 The processors 201 and 205 therein. Each of these processors can be a single-CPU processor or a multi-CPU processor. The processors here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0080] The above-mentioned electronic device can be a general-purpose device or a special-purpose device. The embodiments of the present application do not limit the type of the electronic device.

[0081] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the first electronic device. In other embodiments of the present application, the first electronic device may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0082] As Figure 3 shown, it is an example structure of the node operating system upgrade device 300 provided by the embodiments of the present application. The device 300 can automatically orchestrate the draining and offline of nodes in the node pool. The device can include a node operating system upgrade controller 305, a node pool graceful offline custom resource 301, an offline controller 302, a node draining custom resource 303, and a draining controller 304.

[0083] Among them, the node operating system upgrade controller 305 is used to upgrade the node operating system and create the node pool graceful offline custom resource 301.

[0084] The node pool graceful offline custom resource 301 is created by the cluster administrator through the node operating system upgrade controller and belongs to a custom resource definition (CRD). The node pool graceful offline custom resource 301 is used to create the offline controller 302 and formulate the node pool offline policy.

[0085] The offline controller 302 is used to schedule the node offline batches and control the node pool to be non-expandable and non-schedulable (unavailable) during the offline process; the offline controller 302 is also used to create a node drainage custom resource 303 and notify each node in the node pool to drain water; the offline controller 302 is further used to receive the notification log of the successful node drainage from the node drainage custom resource 303 and scale down the nodes that have completed drainage by specifying node deletion; the offline controller 302 is also used to delete the node pool when there are no nodes in the node pool.

[0086] The node drainage custom resource 303 is used to create a drainage controller 304, which belongs to the custom resource definition and can formulate node drainage policies.

[0087] The drainage controller 304 is used to complete node drainage and evict the pods on the node according to the drainage policy.

[0088] Among them, one node pool corresponds to one node pool graceful offline CR for offline processing of the node pool, one node pool includes multiple nodes, and one node corresponds to one node drainage CR for offline processing of the node.

[0089] Optionally, in this embodiment of the application, the node operating system in the Kubernetes cluster is upgraded by the node operating system upgrade device as an example, as Figure 4 shown, including steps S101 - S105:

[0090] S101. The cluster administrator inputs configuration information in the control interface.

[0091] Optionally, as Figure 5 shown in the control interface 401. The control interface 401 includes configuration information 402 and an upgrade control 403.

[0092] Optionally, the cluster administrator can fill in the corresponding parameters in the configuration information 402 in the control interface 401 to complete the node operating system upgrade configuration.

[0093] Specifically, the configuration information may include node pool name information, stoprolling information, draining template information, etc.

[0094] Among them, the node pool name information is the unique name (identifier) information of the node pool, or the node pool name information is the index information of the node pool. Exemplarily, as Figure 5 shown, the node pool name can be uid - 001.

[0095] The stop rolling information can indicate the suspension of the node rolling upgrade of the k8s cluster. When the stop rolling information is false, the nodes of the k8s cluster are rolling upgraded normally; when the stop rolling information is true, the node rolling upgrade of the k8s cluster is suspended. As Figure 5 shown, the stop rolling information defaults to false.

[0096] The drainage template information is used to generate the template for node drainage. Optionally, the drainage template information can include the default PDB template information, etc.

[0097] Among them, the default PDB template information is used to add a temporary PDB to the pod workload lacking a PDB. Optionally, the default PDB template information can not specify the specific PDB match label.

[0098] Optionally, the default PDB template information can also include the max Unavailable information, which is used to represent the maximum number (or proportion) of pods allowed to be interrupted. Exemplarily, as Figure 5 shown, the max Unavailable information can be 10%.

[0099] Optionally, the drainage template information can also include the grace period seconds information, ignore all daemonset information, global timeout information, orphan pod failed information, single pod auto scale up information, etc.

[0100] Among them, the grace period seconds information is used to specify the graceful termination time during the eviction process to replace the graceful termination time declared in the pod. When the grace period seconds information defaults to -1 (i.e., the grace period seconds information is not filled), the graceful termination time during the eviction process is the graceful termination time declared in the pod. Exemplarily, as Figure 5 shown, the grace period seconds information can be 30s.

[0101] The ignore all daemonset information is used to indicate that when the node goes offline, the daemonset controller is ignored. When the ignore all daemonset information is true, the daemonset controller is ignored by default. Exemplarily, as Figure 5 shown, the ignore all daemonset information can be true.

[0102] The global timeout information is used to specify the maximum time for waiting for a pod to be deleted on each node. If the pod is not deleted within this time, it indicates that the node drainage fails. Exemplarily, as Figure 5 shown, the global timeout information can be 300s.

[0103] The orphaned pod failure information is used to indicate that when there are orphaned pods on a node, the node drainage status fails and the node drainage process stops. The orphaned pod failure information defaults to true, indicating that there are orphaned pods on the node. Exemplarily, as Figure 5 shown, the orphaned pod failure information is true.

[0104] The single-pod auto-scaling information is used to indicate that a single-replica pod automatically scales up to at least 2 replicas through the / replica sub-resource. For example, the single-pod auto-scaling information defaults to true, indicating that a single-replica pod automatically scales up to 2 replicas through the / replica sub-resource. Exemplarily, as Figure 5 shown, the single-pod auto-scaling information is true.

[0105] Optionally, the configuration information can also include the maximum parallel (max paraller) information, the maximum scale-down (maxscaledown) information, etc.

[0106] Among them, the maximum parallel information is used to specify the number of nodes that can be taken offline concurrently at the same time. The maximum number of parallel nodes includes the nodes being drained (draining nodes) and the nodes waiting for logs (waiting for log). The maximum parallel information can default to 10. Exemplarily, as Figure 5 shown, the maximum parallel information is 10.

[0107] The maximum scale-down information is used to specify the maximum number of nodes that the node pool can scale down. When the number of nodes to be scaled down is greater than or equal to this value, the node taking offline stops, so that the number of nodes taking offline in the node pool can be controlled. If the maximum scale-down information is empty, all nodes in the node pool are defaulted to be taken offline. Exemplarily, as Figure 5 shown, the maximum scale-down information is 10.

[0108] Optionally, the configuration information can also include other information related to node operating system upgrade, and the embodiments of the present application do not make specific limitations on this.

[0109] Optionally, the cluster administrator can fill in or modify the above configuration information. Subsequently, the node operating system upgrade controller can perform the upgrade of the node operating system according to this configuration information.

[0110] Create an eviction pod template for each node. Kubernetes can drain pods based on this configuration information and the eviction pod template for each node to complete the upgrade of the node operating system.

[0111] Optionally, after the cluster administrator fills in the above configuration information, they can click the upgrade control 403. Subsequently, in response to the cluster administrator's operation of clicking the upgrade control 403, the node operating system is upgraded.

[0112] S102. The node operating system upgrade controller creates a node pool graceful deletion CR for each node pool according to the configuration information in the control interface.

[0113] Optionally, the node operating system upgrade controller can apply the above configuration information to the Kubernetes cluster, that is, the node operating system upgrade controller can create a node pool graceful deletion CR in the Kubernetes cluster. After that, the node operating system upgrade controller can also create a node drain CR for each node according to the configuration information. Optionally, a node pool graceful deletion CR is created for each node pool, and a node drain CR is created for each node.

[0114] S103. The node operating system upgrade controller upgrades the node operating system in Kubernetes.

[0115] Optionally, the node operating system upgrade controller can create a new node pool, turn on the scaling switch, and scale out the new node pool to a sufficient specification (the same specification as the old node pool), and transfer the labels on the old node pool to the new node pool intact. Add a NoSchedule taint to the old node pool, turn off the elastic scaling of the target old node pool, and leave other old node pools unchanged. Obtain all the nodes under the old node pool and save them in the list of nodes to be taken offline. After that, the node pods in the old node pool can be evicted to the new node pool, and the old node pool can be deleted.

[0116] Optionally, the node operating system upgrade process can be as Figure 6 shown, specifically including the following steps S201 - S205:

[0117] S201. The cluster administrator creates a new node pool in Kubernetes through the node operating system upgrade controller.

[0118] Optionally, the node operating system image versions of the new node pool and the old node pool are different. The new node pool uses a new version of the OS image to upgrade the node operating system.

[0119] Optionally, the cluster administrator can transfer the labels on the old node pool to the new node pool through the node operating system upgrade controller.

[0120] Optionally, the resource configuration (including CPU, memory, disk, etc.) and the number of nodes of the new node pool are the same as those of the old node pool.

[0121] Optionally, the node operating system upgrade controller turns on the scale-down switch of the new node pool and transfers the labels on the old node pool to the new node pool.

[0122] Optionally, taint the old node pool with unschedulable taints and turn off the elastic scaling of the old node pool. Then, obtain all the nodes under the old node pool and save them in the list of nodes to be taken offline.

[0123] Optionally, the cluster administrator can create a graceful node pool drain CR according to the configuration information.

[0124] Optionally, the cluster administrator can default assign the configuration information without specific parameters in the graceful node pool drain CR and verify the configuration information with parameters. For example, perform anti-fooling processing on the configuration parameters, etc.

[0125] Optionally, the graceful node pool drain CR can create a drain controller to drain the nodes in the list of nodes to be taken offline.

[0126] S202. The drain controller creates a node drain CR according to the configuration parameters and drains the nodes.

[0127] Optionally, the node drain CR can create a drain controller, and the drain controller drains the nodes according to the configuration information in the node drain CR.

[0128] Optionally, the node draining process can be as Figure 7 shown, specifically including the following steps S301 - S302:

[0129] S301. The drain controller filters the status of the pods in the node and processes the pods in special status.

[0130] Optionally, the drain controller can check whether there are orphan pods in the node. If there are orphan pods in the node, the drain controller feeds back the information of the orphan pods to the node drain CR, and the node drain CR stops the subsequent operations and the node draining fails.

[0131] Optionally, the cluster administrator can also configure the orphan pod failure information as false in the configuration information. Even if there are orphan pods in the node, the node drain CR continues the subsequent operations to drain the node.

[0132] Optionally, the drain controller can also check whether there are pods in the node that are not protected by PDB. If there are pods in the node that are not protected by PDB, the drain controller creates a default PDB for the pod. Among them, if a specified PDB policy exists in the node drain CR, the drain controller can create a specified PDB for the pod according to the specified PDB policy.

[0133] The PDB created by the drain controller includes labels in the format: "wisecloud / ersdrain-NodepoolGracefulOffline CR name-workload name=true". When the entire node pool is deleted, the default PDB can be uniformly deleted by the node pool graceful offline CR.

[0134] Optionally, the drain controller can also check whether there is a single pod in the node. If there is a single pod in the node and the single pod auto-scaling parameter in the node drain CR is set to false, the drain controller can directly evict the single pod; if there is a single pod in the node and the single pod auto-scaling parameter in the node drain CR is set to true, the drain controller can scale out the single pod, expanding the single-replica pod to a double-replica pod. Then evict the double-replica pod. This eviction action meets the configuration information in the PDB.

[0135] Optionally, if horizontal pod autoscaling (HPA) is configured in the single pod, the drain controller can modify the minimum number of replicas in the HPA configuration to 2 to scale out the single pod. If HPA is not configured in the single pod, the drain controller can scale the pod to 2 replicas through the / replica subresource.

[0136] Optionally, after the node operating system upgrade is completed, the drain controller needs to restore the configuration of the expanded pod, that is, scale in the expanded pod to a single pod.

[0137] Among them, after the scaling out is completed, if the number of replicas in the HPA configuration or the / replica subresource is not 2, it means that a new service has updated the pod, and the number of replicas in the pod will not be modified further.

[0138] S302. The drain controller evicts the node pods.

[0139] Optionally, if the grace period in seconds parameter is configured in the configuration information of the node drain CR, the drain controller evicts the pod according to the specified grace period in seconds parameter; if the grace period in seconds parameter is not configured in the configuration information, the drain controller uses the eviction policy declared in the workload for eviction.

[0140] Among them, if the "too many requests" error is returned during the pod eviction process, it means that the pod eviction process is rate-limited. You can sleep for 5s and then perform the pod eviction again. Repeat this operation until the pod eviction is successful or fails.

[0141] Optionally, during the pod eviction process, if the parameter for ignoring all daemonsets configured in the configuration information is set to true, then the daemonset is ignored during pod eviction. If the parameter for ignoring all daemonsets is set to false, then the pods under the daemonset are evicted together during pod eviction.

[0142] Optionally, if a global timeout parameter is configured in the configuration information, during the pod eviction process, if the eviction time exceeds the time configured in the global timeout parameter, then the node draining fails.

[0143] If the global timeout parameter is not configured in the configuration information, the draining controller keeps waiting for the pod eviction until a clear message indicating the success or failure of the pod eviction is returned, and then ends the draining of this node. After the draining of this node is completed, the draining controller can count the number of successfully evicted pods.

[0144] Optionally, if some pods in the node eviction fail, the draining controller can write the reasons for the pod eviction failure and the number of successfully evicted pods into the status in the node pool graceful shutdown CR and display it on the control interface.

[0145] Optionally, if all pods in the node are successfully evicted, the draining controller can perform daemon set log pushing.

[0146] Among them, when the "wait for log" parameter is configured in the configuration information, if the wait for log parameter is set to true, then the draining controller needs to perform daemon set log pushing. The controller-manager can add the annotation "cluster-autoscaler.kubernetes.io / ready-for-deletion: \"false\"" to the old nodes in the downstream cluster, triggering the agent component to notify the daemon set notification process. When the daemon set log pushing is completed, the ers-agent will change this annotation to true. The controller-manager receives the message that the annotation has been changed to true, indicating that the daemon set log pushing is successful. The draining controller can return a message indicating the success of the daemon set log pushing.

[0147] Optionally, if the daemon set log push fails, the drainage controller can return the reason for the failure, indicating that the failure of the daemon set log push causes the node to fail to go offline.

[0148] Optionally, if the wait for log timeout parameter is configured in the configuration parameters, the drainage controller can perform timing for the daemon set log push. When the daemon set log push time exceeds the time set by the wait for log timeout parameter, the daemon set log push fails. If the wait for log timeout parameter in the configuration parameters is empty, the node drainage CR can set the default drainage controller wait for log timeout time to 30 minutes. The embodiments of the present application do not make specific limitations on this.

[0149] Optionally, after the node drainage is completed, the drainage controller can feedback the node drainage result to the node pool graceful shutdown CR.

[0150] S203. The offline controller queries the drainage status of each node in the old node pool and performs offline processing on the nodes with successful drainage.

[0151] Optionally, the offline controller can cyclically query the drainage status of each node in the old node pool and classify the nodes according to the drainage status of the nodes.

[0152] Exemplarily, the offline controller can query the drainage status of each node in the old node pool every 30s.

[0153] Optionally, the process of node drainage can specifically include two processes: node log cleaning and node drainage. After the node log cleaning is successful, the node drainage operation is performed. If the node drainage is successful, it means that the node drainage is completed.

[0154] Optionally, the drainage status of the node can include the status of being in the process of cleaning the log, the status of being in the process of draining, the status of failure in draining or cleaning the log, and the status of success in draining and cleaning the log.

[0155] Among them, if the node drainage or log cleaning fails, it means that the node drainage fails; if the node drainage and log cleaning are successful, it means that the node drainage is successful, and then the offline controller can delete the node.

[0156] Optionally, the offline controller can display the drainage status of the node in the status of the control interface. Specifically, the status can display the list of nodes being drained, the list of nodes being in the process of cleaning the log, the list of nodes with drainage failure, and the list of nodes with successful drainage.

[0157] Optionally, after the offline controller performs the delete operation on the node, it can check whether the node is successfully deleted. If the node is successfully deleted, the node drainage CR corresponding to the node is deleted.

[0158] Optionally, if the node deletion fails, the offline controller can try to delete the node again, and then check whether the node is successfully deleted. If the node is successfully deleted, the node drainage CR corresponding to the node is deleted. If the offline controller fails to delete the node after multiple deletion operations on the node, the node is placed in the deletion failure queue and displayed in the list of nodes that have failed to go offline. For example, if the offline controller fails to delete the node after 5 deletion operations on the node, the node is placed in the deletion failure queue and displayed in the list of nodes that have failed to go offline.

[0159] Optionally, the offline controller can check the total number of nodes in the list of nodes whose logs are being cleared and the list of nodes whose drainage is in progress. If it is less than the maximum parallel parameter configured in the node pool graceful offline CR, and there are still unprocessed nodes in the list of nodes to be taken offline, new node drainage CRs are created for the nodes in the list of nodes to be taken offline, so that the sum of the newly created node drainage CRs, the list of nodes whose logs are being cleared, and the list of nodes whose logs are being cleared is equal to the maximum parallel parameter.

[0160] In this way, the number of nodes taken offline concurrently at the same time is the maximum number of nodes specified by the offline controller, ensuring the efficiency of node offline.

[0161] Among them, the naming rule for the newly created node drainage CR can be: the CR name of the node pool graceful offline + the node name. The embodiments of the present application do not make specific restrictions on this.

[0162] S204. If the number of nodes taken offline is greater than or equal to the maximum reduction parameter, the offline controller stops arranging new nodes for drainage.

[0163] Optionally, the cluster administrator can configure parameters in the node pool graceful offline CR to set the number of nodes taken offline.

[0164] Exemplarily, the cluster administrator can configure the maximum reduction parameter in the node pool graceful offline CR to specify the maximum number of nodes that the node pool can reduce (go offline). For example, the cluster administrator can configure the maximum reduction parameter to be 30.

[0165] Among them, the number of reduced nodes = the number of nodes whose drainage is in progress + the number of nodes whose logs are being cleared + the number of nodes whose drainage or log clearing fails + the number of nodes whose drainage and log clearing are successful.

[0166] When the number of reduced nodes is greater than or equal to the maximum reduction parameter, stop arranging new nodes for offline; when the number of reduced nodes is less than the maximum reduction parameter, arrange new nodes for offline.

[0167] Optionally, the offline controller can check the status of all nodes. If not all nodes are offline, the intermediate status can then be fed back and displayed in the "status" of the control interface. Exemplarily, it can be displayed as "complete+failed", indicating that the overall task is partially successful. If all nodes have been operated on, the final status can also be fed back and displayed in the "status" of the control interface. Exemplarily, it can be displayed as "complete+succeed", indicating that the overall task is completely successful.

[0168] S205. When all nodes are successfully offline, the offline controller deletes the PDB.

[0169] Optionally, when all nodes in the node pool are successfully offline or all nodes of a set number are successfully offline, the offline controller can delete all PDBs created during the offline of the node pool. Optionally, the offline controller can confirm whether the PDB is a PDB created during this offline by checking whether the PDB carries a label.

[0170] Optionally, if the deletion of the PDB fails, the offline controller can attempt to delete the node again. For example, after the offline controller performs 5 deletion operations on the PDB. Then the offline controller can check whether the PDB is successfully deleted. If the node is successfully deleted, the result of deleting the PDB is fed back in the "status" of the control interface.

[0171] Optionally, after the node operating system upgrade is completed, the node operating system controller can delete the graceful offline CR of the node pool and the node drainage CR. Among them, when deleting the graceful offline CR of the node pool and the node drainage CR, the PDB created by the corresponding node drainage CR is also deleted.

[0172] S104. The node operating system upgrade controller receives the upgrade status of multiple nodes and sends the upgrade status of the multiple nodes to the node operating system upgrade controller.

[0173] Optionally, the node operating system upgrade controller can receive the upgrade status of the nodes that are being offline, as well as the nodes that are successfully or unsuccessfully offline in the node pool.

[0174] Optionally, the upgrade status of the node can include the status of node drainage, whether the node upgrade is successful, the reason for the node upgrade failure, the proportion of upgraded nodes in the total nodes (node pool upgrade progress), etc.

[0175] S105. The node operating system upgrade controller obtains the node operating system upgrade result based on the status of multiple nodes and displays the status of the multiple nodes and the node operating system upgrade result on the control interface.

[0176] Optionally, the node operating system upgrade controller can record the status of the node in the node drain CR and display the status of the node in the control interface.

[0177] Exemplarily, as Figure 8 shown, the status parameter in the control interface can display the status of the node. status.phase can represent the status of the node. Exemplarily, the status.phase parameter can be: "draining" which means the node is draining; status.conditions.type can represent the result of node draining. For example, if status.conditions.type is fail, it means the node draining fails. status.conditions.reason can represent the reason for the node draining failure. For example, the reasons for the node draining failure can be the existence of orphan pods in the node, drain timeout, or log clean timeout, etc. Exemplarily, status.conditions.reason can be orphan pods; status.conditions.message can represent the node name where the node draining fails. Exemplarily, status.conditions.message can be: hasophan pod xxxx; lastTransitionTime can represent the time when the node starts to drain or starts to clean the log, which is convenient for calculating whether the node draining or log cleaning times out.

[0178] Optionally, the node operating system upgrade controller can summarize the statuses of multiple nodes, record them in the node pool graceful shutdown CR, and display them in the control interface.

[0179] Exemplarily, as Figure 9As shown, the status parameter in the control interface can display the result of the node pool going offline. status.phase can represent the offline status of the nodes in the node pool. status.phase being complete+failed represents that the overall task is partially successful, and status.phase being complete+succeed represents that the overall task is completely successful. status.progress can represent the progress of the nodes in the node pool going offline. Exemplarily, status.progress can be 34%, indicating that 34% of the nodes in the node pool have gone offline. status.originalCount represents the initial number of nodes in the node pool. Exemplarily, status.originalCount can be 103. status.drainingNodes represents the list of nodes that are being drained, and this node list can include the addresses of the nodes that are being drained. Exemplarily, status.drainingNodes can be 172.12.3.1; 172.13.0.4. status.waitingForLog represents the list of nodes that are cleaning logs, and this node list can include the addresses of the nodes that are cleaning logs. Exemplarily, status.waitingForLog can be 172.135.2.12; 172.124.2.13. status.faildNodes represents the list of nodes that have failed to go offline, and this node list can include the addresses of the nodes that have failed to go offline. Exemplarily, status.faildNodes can be 172.135.2.14. Among them, the offline nodes are not included in the concurrency (the number of nodes going offline concurrently at the same time).

[0180] Optionally, status.Conditions and its parameters below can be the same as the above Figure 8 each parameter in the status.conditions described above.

[0181] Among them, status.Conditions.lastTransitionTime represents the time point when the node pool offline task is partially completed or completely completed. Exemplarily, status.Conditions.lastTransitionTime can be 2023-05-10xxx".

[0182] After that, the cluster administrator can view the progress of the node operating system upgrade in real time through the control interface.

[0183] Optionally, when there are no nodes in the node pool, the node operating system upgrade controller can delete the old node pool.

[0184] In the embodiments of the present application, when evicting nodes, special processing will be performed on single pods and pods without PDB configured to reduce service interruption. The control interface will display in real time which nodes have been upgraded, which nodes are being upgraded, and which nodes have not been upgraded, and display the specific details of node upgrades. Moreover, the upgrade process is tool-based and the process is standardized. The process of handling abnormal pods and special pods is integrated into the upgrade tool, reducing manual participation and simplifying operations.

[0185] In one embodiment, if the cluster administrator checks that the computing resources at the bottom layer of the electronic device are insufficient, such as the CPU resources of the electronic device are insufficient, the computing resources are less than the first threshold, the number of nodes being upgraded is greater than the second threshold, and the computing resources are not enough to support the simultaneous upgrade of multiple nodes, then the stopRolling parameter in the graceful draining CR of the node pool can be modified to true, and then the entire node OS rolling upgrade process will pause. When the cluster administrator checks that the computing resources at the bottom layer of the electronic device are sufficient, such as the CPU resources of the electronic device are sufficient and the computing resources are greater than the second threshold, then the stopRolling parameter in the graceful draining CR of the node pool can be modified to false, and then the node OS rolling upgrade will continue. Among them, this process can be implemented manually or automatically.

[0186] In this way, the node operating system upgrade can be paused and started at any time according to the computing resources of the electronic device and the current node operating system upgrade situation, realizing visual and controllable cluster node OS upgrade process.

[0187] It should be noted that some operations in the processes of the above method embodiments are optionally combined, and / or the order of some operations is optionally changed. Moreover, the execution order between the steps of each process is only exemplary and does not constitute a limitation on the execution order between the steps. The steps can also be in other execution orders. It is not intended to indicate that the execution order is the only order in which these operations can be performed. Those of ordinary skill in the art will think of various ways to reorder the operations described herein. In addition, it should be pointed out that the process details involved in a certain embodiment herein are also applicable to other embodiments in a similar manner, or different embodiments can be combined and used.

[0188] The embodiments of the present application can divide the function modules of the node operating system upgrade device according to the above method examples. In the case of dividing each function module according to the corresponding functions, Figure 10 shows a possible structural schematic diagram of the fault log storage system involved in the above embodiments. As Figure 10As shown in the figure, the node operating system upgrade device includes a receiving module 1101, a processing module 1102, and a display module 1103. Of course, the node operating system upgrade device may further include other modules, or the node operating system upgrade device may include fewer modules. The embodiments of the present application do not limit this.

[0189] The receiving module 1101 is used to receive the configuration information input by the cluster administrator.

[0190] The processing module 1102 is used to generate a first template according to the configuration information, and upgrade the node operating system according to the first template; the nodes running in the node operating system include isolated pods, or single-copy pods, or pods without pod disruption budget protection.

[0191] The display module 1103 is used to display Figure 8 , Figure 9 the upgrade status of the node; the upgrade status of the node includes at least one of the status of node draining, the reason for node upgrade failure, the proportion of upgraded nodes in the total nodes, and whether the node upgrade is successful.

[0192] For the specific working process of the above-described system, reference may be made to the corresponding process in the above method embodiment, which will not be elaborated here.

[0193] The embodiments of the present application provide a computer-readable storage medium storing one or more programs, and the one or more programs include instructions that, when executed by a computer, cause the computer to execute the node operating system upgrade method described in steps S102 - S105 above.

[0194] The embodiments of the present application further provide a computer program product containing instructions that, when run on a computer, cause the computer to execute the node operating system upgrade method described in steps S102 - S105 of the above embodiments.

[0195] Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0196] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0197] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed to multiple different places. During the application process, some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0198] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application.

Claims

1. A method for upgrading a node operating system, characterized in that, it includes: Obtain configuration information; the configuration information includes isolated pod failure information and / or single-pod auto-scaling information and / or default pod disruption budget template information; Generate a first template according to the configuration information, and upgrade the node operating system according to the first template; the nodes running in the node operating system include isolated pods or single-replica pods or pods without pod disruption budget protection; Display the upgrade status of the node; the upgrade status of the node includes at least one of the status of node draining, the reason for node upgrade failure, the proportion of upgraded nodes in the total nodes, and whether the node upgrade is successful.

2. The method according to claim 1, characterized in that, the method further includes: When the node includes an isolated pod, stop the draining operation of the node; When the node includes a pod without pod disruption budget protection, create a pod disruption budget for the pod; When the node includes a single-replica pod, scale the single-replica pod to a pod with at least two replicas.

3. The method according to claim 1 or 2, characterized in that, the configuration information further includes at least one of grace period in seconds information, ignore all daemonset information, stop rolling information, global timeout information, maximum parallel information, maximum reduction information.

4. The method according to any one of claims 1-3, characterized in that, upgrading the operating system of the node according to the first template includes: Generate a second template according to the first template; the second template is used to perform a draining operation on the node.

5. The method according to any one of claims 1-4, characterized in that, the method further includes: When the computing resources of the electronic device are less than the first threshold and the number of nodes being upgraded is greater than the second threshold, change the stop rolling information in the configuration information and stop the node operating system upgrade operation; When the computing resources of the electronic device are greater than the third threshold, restore the stop rolling information in the configuration information and perform the node operating system upgrade operation.

6. The method according to claim 2, characterized in that, after the node operating system upgrade is completed, the method further includes: Delete the created pod disruption budget and the replicas of the scaled pods.

7. The method according to any one of claims 1-6, characterized in that, the status of node draining includes at least one of the status of cleaning logs, the status of draining, the status of failure in draining or cleaning logs, and the status of successful draining and cleaning logs.

8. A node operating system upgrade device, characterized in that, it includes: A receiving module for obtaining configuration information; the configuration information includes isolated pod failure information and / or single-pod auto-scaling information and / or default pod disruption budget template information; A processing module, configured to generate a first template according to the configuration information and upgrade the node operating system according to the first template; the nodes running in the node operating system include isolated pods or single-copy pods or pods without pod disruption budget protection. A display module, configured to display the upgrade status of the node; the upgrade status of the node includes at least one of the status of node draining, the reason for node upgrade failure, the proportion of upgraded nodes in the total number of nodes, and whether the node upgrade is successful.

9. An electronic device characterized in that it comprises a processor, a display and a memory, the memory is coupled to the processor, the memory is used to store computer program code, the computer program code includes computer instructions, and when the processor reads the computer instructions from the memory, the electronic device is caused to execute the method according to any one of claims 1-7.

10. A computer-readable storage medium storing instructions characterized in that when the instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-7.