New establishment method of pod in K8s cluster, storage medium and electronic equipment

By constructing a node topology graph in the Kubernetes cluster and using a graph neural network to extract node vectors, the problem of deploying newly created pods to nodes with the same running mode was solved, the stability of resource consumption alarms was improved, and the operational stability of the cluster was enhanced.

CN120929189APending Publication Date: 2025-11-11MOBILE TECH COMPANY CHINA TRAVELSKY HLDG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511031081.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In a Kubernetes cluster, when creating a new pod, it is easy to deploy it to a node with the same running status as a previously removed pod, which can lead to unstable resource usage alerts.

Method used

By constructing a node topology graph and extracting node vectors using a graph neural network, the similarity between the node to be mounted for a new pod and historical abnormal nodes is matched to avoid deploying the new pod to nodes with the same running mode.

Benefits of technology

This effectively avoids instability caused by node resource consumption alarms after a pod is created, thus improving the operational stability of the cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929189A_ABST
    Figure CN120929189A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of pod scheduling, in particular to a pod new establishment method in a K8s cluster, a storage medium and electronic equipment. Comprising the following steps: matching a node vector of a to-be-mounted node corresponding to a newly-built pod with each abnormal node vector of a micro-service corresponding to the newly-built pod; and if the matching fails, deploying the new pod on the node to be mounted for operation. In the invention, through similarity matching between the node vectors, the newly built pod can be prevented from being re-deployed in the node which has the same operation mode as the node when the previous pod is removed, so that the unstable operation condition of resource occupation alarm easily occurring in the node after the pod is newly built can be further avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pod scheduling technology, and in particular to a method for creating pods in a Kubernetes cluster, a storage medium, and an electronic device. Background Technology

[0002] Kubernetes (often abbreviated as K8s) is an open-source platform for automating the deployment, scaling, and management of containerized applications. It provides a framework for running distributed systems and ensures high availability, scalability, and efficient resource utilization for applications.

[0003] A Pod is the smallest deployable unit in Kubernetes (K8s), typically containing one or more closely related containers. Applications that implement business functionalities are usually deployed within the containers corresponding to a Pod, allowing the Pod to perform its business functions when running on its corresponding node. A node is the worker machine in K8s, possessing certain computing resources (CPU, memory, etc.) and is the part of the cluster that actually performs the work; it can be a virtual machine or a physical machine. K8s and microservice architecture complement each other; the former provides the ideal runtime environment and technical support for the latter, allowing enterprises to focus more on implementing business logic rather than managing the underlying infrastructure. In current complex business scenarios, for microservices, each microservice instance (i.e., a running copy of the service) is typically run as one or more Pods on different nodes in K8s, collectively providing the functionality of the microservice to offer higher service stability.

[0004] In Kubernetes, a node typically determines whether to issue an alert based on its current resource utilization (especially CPU usage) and whether it exceeds a preset resource utilization alert threshold. If an alert is issued, existing pods on the node are then removed. Current technology primarily selects nodes with sufficient resources that meet the requirements based on the pod's resource requests, limits, and other constraints, and then deploys the pod to one of these nodes. However, when the newly created pod is identical to a previously removed or evicted pod, the existing node selection mechanism can easily lead to the pod being deployed to a node with similar operational characteristics. This can result in unstable operation and frequent resource utilization alerts after the pod is created. Summary of the Invention

[0005] To address one of the aforementioned technical problems, the present invention adopts the following technical solution:

[0006] According to one aspect of the present invention, a method for creating a new pod in a Kubernetes cluster is provided, the method comprising the following steps:

[0007] The node vector of the node to be mounted corresponding to the newly created pod is matched with the abnormal node vector of each microservice corresponding to the newly created pod. The abnormal node vector is the node vector of the node to which the pod belongs after the pod belonging to the same microservice as the newly created pod is removed. The node vector includes the feature vector of the communication interaction of each pod included in the node in the preset historical period, and the hardware attribute vector of the node. The preset historical period is the period from when the alarm of the node to which the pod belongs is removed to an earlier historical moment.

[0008] If a match fails, the newly created pod will be deployed on the node to be mounted.

[0009] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the above-described method for creating a pod in a Kubernetes cluster.

[0010] According to a third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for creating a pod in a Kubernetes cluster.

[0011] This invention has at least one of the following beneficial effects:

[0012] In this invention, after determining the node to be mounted for the newly created pod, the node is further evaluated. Specifically, the similarity between the node vector of the node to be mounted and the node vector of the node to which the same pod was removed in the historical rescheduling is determined. Since the node vector includes the feature vector of the communication interaction of each pod in the node during the preset historical period corresponding to the previously removed pod, as well as the node's hardware attribute vector, the node vector can reflect the node's operating mode. In this invention, by matching the similarity between node vectors, it is possible to avoid the newly created pod being deployed on a node with the same operating mode as the previously removed pod, thereby further avoiding the unstable operation of the node that is prone to resource consumption alarms after the pod is created. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A flowchart illustrating a method for removing a pod in a K8s cluster, provided as an embodiment of the present invention;

[0015] Figure 2 A flowchart illustrating a method for creating a pod in a K8s cluster, as provided in an embodiment of the present invention;

[0016] Figure 3 A flowchart illustrating a compensatory scheduling method for pods in a K8s cluster, provided as an embodiment of the present invention;

[0017] Figure 4 This is a schematic diagram of the node topology provided in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] According to one possible embodiment of the present invention, such as Figure 1 As shown, a method for removing a pod in a Kubernetes cluster is provided, which includes the following steps:

[0020] S100: Select other nodes in the cluster that already have the same pod as the candidate pod to be removed as the comparison nodes for the target node. The target node is the node to which the candidate pod belongs. The candidate pod is any pod in the target node that ranks highest in CPU utilization according to a preset order.

[0021] In a Kubernetes (K8s) cluster, multiple nodes are configured to deploy and run different pods to implement various business functions (i.e., multiple microservices). Taking a flight ticketing scenario as an example, the business functions such as ticket information query, seat selection, ticket payment, rescheduling, and refunds are all different microservices. Taking the ticket information query business function as an example, to handle the high traffic during peak ticket query periods, multiple identical pods are typically used to jointly implement this function, and these different pods are deployed on different nodes in the cluster. According to a pre-defined load balancing strategy, different pods can distribute the corresponding level of traffic. Furthermore, by deploying pods on different nodes, if one node fails, the replicas on other nodes can still continue to provide services, thereby improving the reliability of the service.

[0022] In actual operation, as business functions develop and upgrade, not only will more new pods be created in a node, but existing pods will also be upgraded, leading to corresponding changes in resource requirements (especially CPU requirements). Typically, existing technologies set a CPU usage alarm threshold (e.g., 80%). Once the actual CPU usage of a node exceeds this threshold, the node will be in an alarm state, requiring the removal (eviction) of some pods. This is usually done through a pre-defined eviction strategy (e.g., removing pods with the highest resource usage or removing pods with resource usage exceeding a pre-defined threshold) to identify candidate pods for removal and then evict them.

[0023] In a cluster, a microservice typically has multiple identical pods, meaning there will be replica pods identical to the currently candidate pod for removal. Furthermore, since the selection strategies for each pod are largely the same during deployment and scheduling, they will be mounted on different nodes with high similarity. Therefore, in this step, after identifying the candidate pod for removal, the nodes belonging to other replica pods will be used as comparison nodes.

[0024] Specifically, prior to S100, the method also includes the following method for determining candidate pods to be removed:

[0025] S101: When the CPU resource utilization of any node in the cluster is greater than or equal to the CPU usage alarm threshold, pods on that node with a CPU resource utilization exceeding a preset eviction threshold are considered as candidate pods for removal. The preset eviction threshold is less than the CPU usage alarm threshold. Specifically, the preset eviction threshold can be 15%, and the CPU usage alarm threshold can be 80%.

[0026] S200: Using a graph neural network, features are extracted from the node topology graphs corresponding to the target node and each pair node, generating the first target node vector corresponding to the target node and the first pair node vector corresponding to each pair node.

[0027] like Figure 4 As shown, the nodes in the node topology graph include the nodes corresponding to each pod running on the target node during a preset historical time period, such as d1, d2, and d3 in the figure, and the nodes corresponding to pods on other nodes that communicate with each pod on the target node, such as a1, a2, b1, c1, and c2 in the figure. The node attribute values ​​include pod ID, job type number, and the node ID of the node to which the pod belongs. Specifically, the node attribute value can be in the form (d1, 05, D), where d1 is the pod ID, 05 is the job type number, and D is the node ID of the node to which the pod belongs. Directed edges in the node topology graph connect any two nodes that communicate with each other, with the connection direction being the same as the communication direction. The directed edges are set with the number of communications between the two nodes in the corresponding communication direction. The preset historical time period is the period from when the alarm of the node to which the pod belongs is removed to an earlier historical time. For example, it can be a historical period of 10 minutes prior to when the alarm of the node to which the pod belongs is removed. For instance, if the time corresponding to the node alarm is 10:00:00, then the corresponding preset historical period is from 9:50:00 to 10:00:00.

[0028] A rapid increase in CPU usage on a node is usually due to increased access traffic to multiple pods, leading to higher CPU utilization in those pods. Correspondingly, this increase in access traffic is typically not sudden but rather gradual over a sustained period, causing increased CPU usage in the corresponding pods until the node's CPU resource utilization exceeds or equals the CPU usage alarm threshold, triggering an alarm. Therefore, in this embodiment, when constructing the node topology diagram, the communication topology of each pod within the node during a preset historical time period is considered. Specifically, the node topology diagram reflects the communication topology structure and traffic characteristics of each pod within the node, indicating the node's operational status during periods of increased access traffic.

[0029] Graph Neural Networks (GNNs) are deep learning models specifically designed for processing graph-structured data. In a topological graph, nodes represent entities (in this embodiment, information such as pod ID, job type number, and the node ID of the node to which the pod belongs), while edges represent the relationships between these entities (in this embodiment, the communication direction and the number of communications between two nodes). GNNs extract features by directly manipulating the graph structure. Therefore, in this embodiment, after constructing the directed node topological graph for the corresponding nodes, GNNs can be used to extract the graph's feature vectors, namely the first target node vector and the first comparison node vector in this embodiment.

[0030] S300: Concatenate the first target node vector with the hardware attribute vector of the target node to generate the second target node vector.

[0031] S400: Concatenate the first comparison node vector with the corresponding node's hardware attribute vector to generate the corresponding second comparison node vector. The node's hardware attribute vector includes the node's CPU attributes, memory attributes, storage attributes, network attributes, GPU attributes, and motherboard attributes. These specific attribute values ​​can all be encoded using existing methods to form specific element values ​​in the hardware attribute vector. For example, if the number of CPU cores in the CPU attribute is 6, then the element value in the corresponding dimension of the hardware attribute vector will be 6.

[0032] Specifically, CPU attributes can include the number of cores, architecture type, and operating frequency of each CPU core; memory attributes can include the total physical memory; storage attributes can include the total hard disk storage and hard disk type, such as HDD (Hard Disk Drive), SSD (Solid State Drive), NVMe, etc.; network attributes can include the interface type, maximum transmission rate (i.e., bandwidth) of the network interface, and average communication latency; GPU attributes can include the number of GPUs, the model of each GPU, the GPU's video memory capacity, and the GPU's operating frequency; motherboard attributes can include the motherboard model and manufacturer. Of course, other hardware attributes can also be added, such as cooling system attributes and power supply attributes.

[0033] By concatenating the vectors reflecting the node's operational status and hardware attributes as described above, a more comprehensive and accurate reflection of a node's characteristics can be obtained. Therefore, in this embodiment, the degree of similarity between two nodes can be reflected by calculating the similarity (such as cosine similarity) between their corresponding node vectors.

[0034] S500: If the similarity between the second target node vector and any second comparison node vector is greater than a preset similarity threshold, and the CPU utilization rate of the node corresponding to the second comparison node vector is greater than or equal to a first preset utilization threshold in a preset historical period, then the candidate pod to be removed will be removed; the first preset utilization threshold is less than or equal to the CPU usage alarm threshold. For example, the first preset utilization threshold can be 70%.

[0035] S600: If the similarity between the second target node vector and any second comparison node vector is greater than the preset similarity threshold, and the CPU utilization rate of the node corresponding to the second comparison node vector is always less than the first preset utilization threshold in the preset historical period, then the candidate pod to be removed will be retained.

[0036] In this embodiment, S500 and S600 compare the similarity between the target node and each comparison node, thus selecting comparison nodes in the Kubernetes cluster that have a high similarity to the target node. Then, based on whether the CPU utilization of the comparison node has consistently been less than a first preset threshold during a preset historical period, it is determined whether the comparison node's operation is stable during that period. Furthermore, by comparing the node's operational stability, it is further determined whether the current candidate pod for removal can run stably on the target node, thereby deciding whether to evict the candidate pod and reducing the possibility of mistakenly removing pods that do not affect cluster stability.

[0037] In addition, to further improve the accuracy of the judgment, S500 can be replaced by: if the stability judgment value P corresponding to the candidate pod to be removed is greater than Y1, then the candidate pod to be removed is retained. Y1 is the judgment threshold, P = P1 / P2, P1 is the number of second comparison node vectors whose similarity with the second comparison node vector is greater than the preset similarity threshold, and whose corresponding node's CPU utilization rate is consistently less than the first preset utilization threshold in a preset historical period, and P2 is the number of second comparison node vectors whose similarity with the second comparison node vector is greater than the preset similarity threshold.

[0038] As another possible embodiment of the present invention, the method for removing pods in a K8s cluster further includes:

[0039] S710: If all candidate pods to be removed in the target node are retained, then the comparison node with the highest similarity corresponding to each candidate pod is compared with the target node in descending order of CPU utilization of each candidate pod to determine the pod to be removed in the target node.

[0040] The comparison process includes:

[0041] S711: Get the set A of pods corresponding to the node with the highest similarity to the candidate pod to be removed.

[0042] S712: Get the pod collection B of the target node.

[0043] S713: Pods that exist in B but not in A will be designated as pre-removed pods.

[0044] If all candidate pods to be removed are retained after the S500 and S600 checks, it means that the other replica pods corresponding to these candidate pods can run stably on their respective nodes. This indicates that the candidate pods to be removed are not the main reason for the node's CPU utilization to exceed the limit quickly.

[0045] Therefore, in this embodiment, pods that exist in B but not in A are considered as pre-removed pods. Pods that exist in the target node but do not exist in the node corresponding to the maximum similarity of the candidate removed pods are considered as the main reason for the node's CPU utilization to rapidly exceed the limit.

[0046] After identifying the pods to be removed, the pods will be removed. The present invention provides the following two parallel pod removal embodiments.

[0047] Pod Removal Example 1

[0048] S714: Remove pods to be removed in order of increasing CPU utilization.

[0049] S715: After removing a pre-removed pod, if the CPU utilization of the target node is still in an alarm state, then continue to remove the next pre-removed pod.

[0050] S716: After removing a pod for pre-removal, if the CPU utilization of the target node is in a non-alarm state, the comparison process will be stopped.

[0051] This embodiment is applicable when the total number of candidate pods to be removed is small, such as only two. In this case, all candidate pods can be compared to determine all pods to be removed, and then removed sequentially. Since the total number of candidate pods is small, even a full comparison will not take too much time.

[0052] Pod Removal Example 2

[0053] S724: If the sum of the CPU utilization rates of all currently identified pods to be removed is greater than the total excess CPU utilization rate of the target node, then stop comparing the next candidate pods to be removed. The total excess CPU utilization rate of the target node is the difference between the target node's current CPU utilization rate and a second preset utilization threshold, which is less than a first preset utilization threshold. The second preset utilization threshold can be 65%.

[0054] S725: If the total CPU utilization of all currently identified pods to be removed is less than the total excess CPU utilization of the target node, then continue to perform comparison processing on the next candidate pod to be removed.

[0055] This embodiment is applicable to situations where the total number of candidate pods to be removed is large, such as when the total number of candidate pods to be removed is greater than 10. In this case, after performing comparison processing on each candidate pod to be removed, the CPU utilization of all currently identified pods to be removed is compared with the total CPU utilization of the target node. This allows the target node to be cleared of the alarm status as early as possible and avoids comparing all candidate pods to be removed, thus reducing the time spent on comparison processing.

[0056] As another possible embodiment of the present invention, such as Figure 2 As shown, a method for creating a new pod in a Kubernetes cluster is also provided, which includes the following steps:

[0057] A100: Matches the node vector of the node to be mounted for the newly created pod with the abnormal node vector of each microservice corresponding to the newly created pod. The abnormal node vector is the node vector of the node to which the pod belongs after it has been removed from the same microservice as the newly created pod. The node vector includes the feature vector of communication interactions of each pod included in the node within a preset historical period, as well as the node's hardware attribute vector. The preset historical period is the time period from when the alarm of the node to which the pod belongs was triggered to an earlier historical moment.

[0058] Specifically, the abnormal node vector of a microservice is obtained through the following steps:

[0059] A101: After a pod is removed from its node, retrieve the microservice tag corresponding to the removed pod.

[0060] A102: Get the node vector corresponding to the node after the pod is removed.

[0061] A103: Based on the microservice label, the node vector is used as the abnormal node vector of the corresponding microservice.

[0062] Since a single microservice may contain multiple pods deployed on different nodes within the cluster, if a pod's CPU requirements do not match the CPU supply of its current node, the pod will be unable to continue running stably on that node. In other words, the node will issue an alert, and the pod will be removed from the current node.

[0063] In this embodiment, after a pod is removed, the node vector of the node that previously belonged to that pod is obtained as the abnormal node vector corresponding to the removed pod. In this embodiment, each abnormal node vector is a node vector that does not contain any pods identical to the removed pod. Specifically, when constructing the node topology graph corresponding to a node, topology information related to pods identical to the removed pod can be deleted from the topology graph before feature extraction. Figure 4 If d3 is the same pod as the pod that was removed, then d3, a2, c2, and the edges between them can all be deleted.

[0064] Therefore, with the accumulation of pod removal work, a large number of abnormal node vectors can be obtained for each microservice. These vectors indicate that the microservice cannot run stably on the nodes corresponding to these abnormal node vectors. Thus, A100 can determine whether the node to be mounted is more suitable for deploying a new pod by matching the node vector of the node to be mounted with the abnormal node vector of each corresponding microservice.

[0065] Furthermore, in this embodiment, the abnormal node vector can also be obtained by determining the pod to be removed using the pod removal method described above in the K8s cluster, and then obtaining its corresponding abnormal node vector.

[0066] The node vector is obtained through the following steps:

[0067] A112: Obtain the node topology graph of the node to be processed. The nodes in the node topology graph include the nodes corresponding to each pod running on the node to be processed during a preset historical time period, and the nodes corresponding to pods in other nodes that communicate with each pod on the node to be processed. Node attribute values ​​include podID, job type number, and the node ID of the node to which the pod belongs. Directed edges in the node topology graph connect any two nodes that communicate with each other, with the connection direction being the same as the communication direction. The directed edges are set with the number of communication interactions between the two nodes in the corresponding communication direction.

[0068] A122: Use a graph neural network to extract features from the node topology graph and generate the first target node vector corresponding to the node to be processed.

[0069] A132: Concatenate the first target node vector with the hardware attribute vector of the node to be processed to generate the node vector corresponding to the node to be processed.

[0070] The steps for obtaining node vectors in this embodiment are the same as those in S200 and S300 above, except that the only difference is the pods included in the node.

[0071] A200: If a match fails, the newly created pod will be deployed on the node to be mounted.

[0072] A300: If a match is successful, the node to be mounted will be determined as a node that is not suitable for deployment and operation of a newly created pod.

[0073] The vector matching in the above steps can be achieved by calculating the cosine similarity of the vectors and then comparing it with a set threshold to determine whether a match is successful.

[0074] In this embodiment, after determining the node to be mounted for the newly created pod, the node is further evaluated. Specifically, the similarity between the node vector of the node to be mounted and the node vector of the node to which the same pod was removed in the historical rescheduling is determined. Since the node vector includes the feature vector of the communication interaction of each pod in the node during the preset historical period corresponding to the previously removed pod, as well as the node's hardware attribute vector, the node vector can reflect the node's operating mode. In this invention, by matching the similarity between node vectors, it is possible to avoid the newly created pod being deployed again on a node with the same operating mode as the previously rescheduling node, thereby further avoiding the unstable operation of the node that is prone to resource consumption alarms after the pod is created.

[0075] A400: If the number of consecutive failed matches of any abnormal node vector corresponding to a microservice exceeds a preset matching threshold, the abnormal node vector will be deleted.

[0076] This step allows for the timely removal of historical abnormal node vectors corresponding to microservices. As the cluster continuously schedules, unstable pods from historically unstable nodes are removed and reassigned to newer, more suitable nodes. Consequently, the node vectors corresponding to these nodes change with the scheduling process, leading to a significant difference between the current cluster topology and that of nodes corresponding to older abnormal node vectors. Therefore, older abnormal node vectors are no longer valuable and should be deleted. This step determines the usefulness of an abnormal node vector based on the number of consecutive failed matches, thus deciding whether to clean it up. By promptly deleting useless abnormal nodes, the remaining abnormal node vectors are more valuable, thereby improving the effectiveness of the matching operation.

[0077] As another possible embodiment of the present invention, such as Figure 3 As shown, a compensatory scheduling method for pods in a Kubernetes cluster is also provided, which includes the following steps:

[0078] B100: Based on the number of alarm nodes corresponding to the node to which the removed pod belongs in the cluster, generate the same number of compensation pods under the microservice corresponding to the removed pod. The compensation pods are the same as the removed pods.

[0079] The node that triggers the alert is defined as one whose similarity to the node containing the removed pod exceeds a first similarity threshold, has a CPU utilization exceeding a first preset threshold within a preset historical period, and runs a node containing the same pod as the removed pod. Specifically, the preset historical period is the time from when the node containing the removed pod issued the alert to an earlier historical moment.

[0080] To ensure load balancing and high availability and fault tolerance, multiple identical pod instances of the same microservice are typically run on different nodes. Furthermore, since multiple pods use largely the same selection criteria when choosing nodes, the nodes corresponding to these pods tend to have high similarity. Therefore, in this embodiment, when a pod is removed, nodes with high similarity to the node to which the removed pod belongs are identified. The number of nodes to issue warnings is determined based on the stability performance of these highly similar nodes over a preset historical period—specifically, whether their CPU utilization exceeds a first preset threshold during that period. This results in the number of nodes identified as issuing warnings representing the number of other unstable pods within the microservice corresponding to the removed pod. These pods have a low match with their current node and are more likely to be removed due to CPU utilization warnings in subsequent runs.

[0081] Furthermore, in this embodiment, before B100, the method further includes:

[0082] B110: Get the node vector corresponding to the node to which the pod belongs.

[0083] B120: If, within a preset historical time period, any other node in the cluster running the same pod as the removed pod has a CPU utilization rate exceeding a first preset utilization threshold, then that other node will be designated as the initial node. The first preset utilization threshold is less than or equal to the CPU usage alarm threshold.

[0084] Typically, the first preset CPU usage threshold is lower than the CPU usage alarm threshold, and the two are very close. For example, if the CPU usage alarm threshold is 80%, the first preset threshold is 75%. If, during a preset time period, other nodes running the same pod as the removed pod have a CPU usage rate exceeding the first preset threshold, it indicates that the pod running on these nodes is unstable, and these nodes also pose a certain alarm risk, potentially triggering alarms quickly during subsequent use.

[0085] B130: If the similarity between the node vector corresponding to the node to which the pod belongs and the node vector corresponding to each initial node is greater than the first similarity threshold, then the initial node is determined as the early warning node corresponding to the node to which the pod belongs.

[0086] In this step, the similarity of the working conditions between two nodes is calculated by using node vectors that can reflect the communication topology characteristics and hardware attribute characteristics of the nodes. This allows for a more accurate selection of alarm-predicting nodes from the initial nodes.

[0087] In this embodiment, when determining the number of alarm nodes corresponding to the node to which the pod belongs to be removed, a rapid initial screening is performed first based on the node's operational stability over a preset historical period (step B120). Since this method only requires threshold comparison, the screening speed is improved. Then, based on the initial screening, the node vectors of the relevant nodes are obtained for vector comparison, resulting in a finer screening (step B130). Therefore, by performing a rapid initial screening, the number of nodes required for subsequent fine screening is reduced, thereby reducing the workload of vector matching.

[0088] Specifically, in B130 of this embodiment, the node vectors are obtained based on the performance of each node in a preset historical time period. Then, the similarity between node vectors is calculated to determine the similarity between any two nodes. The specific node vectors corresponding to each node in the preset time period can be obtained according to the steps of S200, S300, and S400 described above, and will not be repeated here.

[0089] B200: Determine the target mount node for the compensation pod from multiple candidate nodes.

[0090] The target node is a candidate node whose similarity to the node to which the pod belongs is less than a second similarity threshold, and whose CPU utilization has consistently been less than a second preset utilization threshold over a preset historical period. The second similarity threshold is less than the first similarity threshold. The second preset utilization threshold is less than the first preset utilization threshold. The second preset utilization threshold can be 65%.

[0091] Prior to B200, the method also included:

[0092] B210: Select multiple initial nodes from the cluster based on the resource configuration requirements and preset constraints of the compensation pod.

[0093] B220: Prioritize the initial nodes based on their resource utilization, load balancing, and node affinity.

[0094] B230: Select the initial node whose priority is within the preset range as the candidate node for the compensation pod.

[0095] B300: Deploy the compensation pod to the corresponding target mount node and run it, and remove the pod that is the same as the pod that was removed from the alarm node.

[0096] In this embodiment, the same number of compensation pods are generated based on the number of alarm nodes, and a new node is selected as the target mounting node for each compensation pod. Since the target mounting node has low similarity to the node to which the removed pod belongs, and its operation is more stable during the preset historical period, the operation of the compensation pod on it can be guaranteed to be more stable and reliable. Furthermore, after the compensation pod is deployed to the corresponding target mounting node, the pod that is identical to the removed pod in the alarm node is removed.

[0097] Therefore, when a pod is removed, the number of other unstable pods under the corresponding microservice is determined based on the similarity between the removed pod and the node it belongs to, as well as its stability performance over a preset historical period. Then, the same number of compensation pods are created in advance and deployed to other, more suitable nodes. Finally, pods on the alerting node that are identical to the removed pod are removed. This avoids the problem of some business functions corresponding to a pod becoming unusable during the process from removal to redeployment, ensuring the high availability of the corresponding microservice and thus improving the user experience for that business function.

[0098] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0099] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0100] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.

[0101] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as “circuits,” “modules,” or “systems.”

[0102] An electronic device according to this embodiment of the invention. The electronic device is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the invention.

[0103] Electronic devices are manifested in the form of general-purpose computing devices. Components of an electronic device may include, but are not limited to: at least one processor, at least one memory, and buses connecting different system components (including memory and processor).

[0104] The memory stores program code that can be executed by a processor, causing the processor to perform the steps described in the "Exemplary Methods" section above, according to various exemplary embodiments of the present invention.

[0105] The storage may include readable media in the form of volatile storage, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).

[0106] The storage may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more applications, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0107] A bus can represent one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus architectures.

[0108] The electronic device can also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be performed via input / output (I / O) interfaces. Furthermore, the electronic device can communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0109] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0110] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the present invention may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.

[0111] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0112] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0113] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0114] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0115] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0116] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0117] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for creating a new pod in a Kubernetes cluster, characterized in that, The method includes the following steps: The node vector of the node to be mounted corresponding to the newly created pod is matched with the abnormal node vector of each microservice corresponding to the newly created pod. The abnormal node vector is the node vector of the node to which the pod belongs after the pod belonging to the same microservice as the newly created pod is removed. The node vector includes the feature vector of the communication interaction of each pod included in the node in a preset historical period, and the hardware attribute vector of the node. The preset historical period is the period from when the alarm of the node to which the pod belongs is removed to an earlier historical moment. If the match fails, the newly created pod will be deployed and run on the node to be mounted.

2. The method according to claim 1, characterized in that, After matching the node vector of the node to be mounted corresponding to the newly created pod with the abnormal node vector of each microservice corresponding to the newly created pod, the method further includes: If a match is found, the node to be mounted will be determined as a node that is not suitable for deployment and operation of a newly created pod.

3. The method according to claim 1, characterized in that, The method further includes: If the number of consecutive failed matches of any abnormal node vector corresponding to a microservice exceeds a preset matching threshold, then the abnormal node vector will be deleted.

4. The method according to claim 1, characterized in that, The abnormal node vector of a microservice is obtained through the following steps: After a pod is removed from its node, obtain the microservice tag corresponding to the removed pod; Get the node vector corresponding to the node after the pod is removed; Based on the microservice label, the node vector is used as the abnormal node vector of the corresponding microservice.

5. The method according to claim 4, characterized in that, The node vector is obtained according to the following steps: Obtain the node topology graph of the node to be processed; the nodes in the node topology graph include the nodes corresponding to each pod running in the node to be processed during a preset historical time period, and the nodes corresponding to pods in other nodes that have communication interaction with each pod in the node to be processed. The attribute values ​​of the nodes include pod ID, job type number, and node ID of the node to which the pod belongs; the directed edges in the node topology graph connect any two nodes that have communication interaction, and the connection direction is the same as the communication direction. The directed edges are set with the number of communication times between the two nodes in the corresponding communication direction. Using a graph neural network, features are extracted from the node topology graph to generate the first target node vector corresponding to the node to be processed; The first target node vector is concatenated with the hardware attribute vector of the node to be processed to generate the node vector corresponding to the node to be processed.

6. The method according to claim 5, characterized in that, The preset historical time period is the period corresponding to the 10 minutes prior to when the alarm of the node to which the pod belongs was removed.

7. The method according to claim 5, characterized in that, Removing a pod from its node includes: Other nodes in the cluster that have the same pod as the candidate pod to be removed are used as comparison nodes for the target node; the target node is the node to which the candidate pod to be removed belongs, and the candidate pod to be removed is any pod in the target node that ranks first in CPU utilization according to a preset order. Use the node vector corresponding to the target node as the second target node vector; Use the node vector corresponding to each alignment node as the second alignment node vector corresponding to each alignment node. If the similarity between the second target node vector and the second comparison node vector is greater than a preset similarity threshold, and the CPU utilization rate of the node corresponding to the second comparison node vector is greater than a first preset utilization threshold in a preset historical period, then the candidate pod to be removed will be removed; the first preset utilization threshold is less than or equal to the CPU usage alarm threshold.

8. The method according to claim 7, characterized in that, Before comparing other nodes in the cluster that have the same pod as the candidate pod to be removed, as the comparison nodes for the target node, the method further includes: When the CPU resource utilization of any node in the cluster is greater than or equal to the first preset resource utilization threshold, pods in the node whose CPU resource utilization is greater than the second preset resource utilization threshold are selected as candidate pods to be removed; the second preset resource utilization threshold is less than the first preset resource utilization threshold.

9. A non-transitory computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a method for creating a pod in a K8s cluster as described in any one of claims 1 to 8.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a method for creating a pod in a K8s cluster as described in any one of claims 1 to 8.