An Adaptive Risk-Resistant Kubernetes Cluster Scheduling Method Based on Penalty Factors
By introducing penalties and adaptive scheduling methods into Kubernetes clusters, computing resources are reserved and adjusted according to the cluster state, the problem of traditional scheduling methods insufficient risk resistance when bursts of high load is solved, and the cluster's risk resistance and load balancing are improved.
Patent Information
- Application Number
- CN202211544417.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-04
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-04
AI Technical Summary
The traditional Kubernetes scheduling method lacks risk resistance when handling bursts of high loads, which may cause node failure or crash.
Adaptive risk-resistant Kubernetes cluster scheduling method based on punishment factors is adopted to resist burst high-load computing tasks by reserveing computing resources and adaptively adjusting according to cluster state.
Improve the risk resistance of Kubernetes clusters, ensuring that resources can be effectively utilized under high load conditions, avoid node failures, and achieve load balancing.
Smart Images

Figure CN115981806B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of edge computing, and particularly relates to a Kubernetes cluster scheduling method. Background Art
[0002] With the rapid development of the Internet, various high-quality services provided by cloud computing are deeply penetrating into every corner of our lives. However, the transmission, processing, and storage of big data pose huge challenges to us. If a large amount of data is transmitted to the cloud computing center for processing and then returned, it will cause a lag in data processing. At the same time, uploading some private data to the cloud computing center will greatly increase the risk of private data leakage. Therefore, edge computing emerges as the times require. By offloading latency-sensitive and data-privacy computing tasks to the edge for processing, some problems existing in cloud computing can be effectively solved.
[0003] Edge computing is the extension of cloud computing to the edge through the network and is a part of cloud computing. In the three-layer service model of cloud computing, it is divided into Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). In the IaaS layer, resource scheduling is based on virtualization technology. However, traditional virtualization has disadvantages such as slow creation speed and increased performance loss in the intermediate call chain. Container technology can well solve the above problems. A container can be regarded as a lightweight virtual machine. Applications and their related dependencies are isolated in the container and do not affect each other, while sharing the host operating system kernel, with a faster startup speed and smaller overhead. Cloud native centered on containerization is becoming a new generation of cloud computing technology.
[0004] Kubernetes (from Greek, meaning "helmsman" or "navigator") is an open-source container orchestration framework that makes it easier for us to manage a large number of computing resources. It packages, deploys, and runs applications through containers, and runs, orchestrates, and schedules containers across nodes in a master-slave management manner between the master node Master and the worker node Worker in the cluster. While facilitating operation and maintenance and continuous integration and continuous deployment for developers, it reaches the state of optimizing the utilization rate of cluster computing resources. Kubernetes has become the most mainstream container orchestration and scheduling framework for container cloud platforms.
[0005] The Kubernetes default scheduler only schedules based on two metrics, CPU and memory of the nodes, and tries to maximize the utilization of computing resources as much as possible. However, maximizing the utilization of computing resources may not be the right solution for all clusters. Because it often takes a certain amount of time to expand the cluster to handle sudden high loads, and if not handled in time, it will lead to node failures or even crashes. Therefore, how to improve the risk resistance ability of the Kubernetes cluster is an urgent problem to be solved. Summary of the Invention
[0006] The purpose of the present invention is to provide an adaptive risk-resistant Kubernetes cluster scheduling method based on penalty factors with strong risk resistance ability to overcome the deficiencies of traditional Kubernetes scheduling methods.
[0007] In the present invention, the Kubernetes cluster consists of a master node Master and multiple worker nodes Worker; users interact with the master node Master through a graphical interface or a command-line interface, and customize the deployment of load tasks through a resource manifest file. Figure 1 It is the basic way for users to interact with the Kubernetes cluster.
[0008] The adaptive risk-resistant Kubernetes cluster scheduling method based on penalty factors provided by the present invention resists sudden high-load computing tasks by reserving a certain amount of computing resources, and adaptively changes the reserved computing resource pool according to the state of the cluster, so as to resist sudden high-load computing tasks. Figure 2 It is the flowchart of the load scheduling method of the present invention, and the specific steps are as follows:
[0009] Step 1: First, prepare a yaml file containing the resource request list of the application load image. When the user issues a command to deploy the workload through a graphical interface or a command-line interface, the Kubectl component of the master node parses the configuration in the request list and constructs the HTTP request parameters of the corresponding object according to its content; after passing the corresponding checks, the request is sent to the API Server component of the master node in the form of HTTP for processing; when the request is verified, the API Server component of the master node creates a corresponding resource object in the distributed key-value data storage system etcd, and finally the API Server component constructs an HTTP response and returns it to the client;
[0010] At this time, although the resource object is saved in etcd, it has not been deployed to the real worker node Worker yet. At the same time, we also need to determine and adjust the values of hyperparameters such as penalty factors and influence weights according to the load characteristics of the workload to be deployed.
[0011] Step 2: Use the monitoring metric collector to collect the load information of each worker node in the cluster within the sliding window. The length of the sliding window is determined by the monitoring period we expect. The monitoring metric collector collects resource load information from the Kubelet component in the worker node Worker and exposes them in the cluster through the corresponding application programming interface for various other components to use. After obtaining the load information in different dimensions, first perform a priority evaluation on the to-be-scheduled load Pods, adjust the load Pod with the highest priority to the head of the to-be-scheduled queue, and adjust the load Pod with the lowest priority to the tail of the to-be-scheduled queue. Then use the preselection algorithm of the default scheduler (Reference: Burns B, Grant B, Oppenheimer D, et al. Borg, omega, and kubernetes[J].
[0012] Communications of the ACM, 2016, 59(5): 50-57.) to screen each Worker node. It should be noted that the preselection algorithm only screens out the list of worker nodes Worker that meet the basic requirements of the additional workload and cannot obtain the optimal node. Therefore, a further optimization algorithm is needed to score the list of worker nodes Worker that meet the basic requirements.
[0013] Step 3: Use the optimization algorithm to score the list of Worker nodes obtained in the previous step and find an optimal Worker node.
[0014] Step 4: When the optimal Worker node is selected, start the asynchronous binding process, that is, write the name of the current optimal Worker node into the NodeName field of the Pod of the workload to be scheduled. So far, all the work has only been targeted at the resource objects stored in the distributed key-value data storage system etcd, and the workload has not really started. The next stage involves running the specific container workload on the Worker node, and these things are all done by the Kubelet on the Worker node. In the Kubernetes cluster, each Worker node starts a Kubelet service process. It queries the Pod in the API Server in the master node Master at regular intervals, filters the list of Pods whose NodeName matches the node where it is located. When it detects that a new Pod object has not been created on the Worker node, the Kubelet performs some pre-operations and creates a pause container through the container runtime interface CRI, sets the network for the Pod through the container network interface CNI, and finally pulls the images defined in our resource list through the container runtime interface CRI, creates the real container workload and starts it, finally achieving the load balancing of the cluster.
[0015] It should be noted that if an external interruption is triggered at this time, the scheduling process will be paused and switched to execute an application with a higher priority.
[0016] In the present invention, the optimization algorithm described in step 3 is implemented by modifying the Scheduler Framework on the basis of the core code of the scheduler. Specifically, plugin extension points are set on all critical paths of the original scheduling logic. Accordingly, plugins are developed without modifying the core code of the Kubernetes scheduler, and the scheduling logic of the optimization algorithm is extended. The specific process of finding the optimal Worker node according to this optimization algorithm is as follows.
[0017] (1) First, obtain the requested resources of the Pod to be scheduled, including CPU resource r cpu and memory resource r memory . Then calculate the average value avg cpu and standard deviation std cpu of the CPU utilization rate of the current node to be scored within a period of time, the average value avg memory and standard deviation std memory of the memory utilization rate, and at the same time obtain the computing resource load limit cap of the current node.
[0018] (2) Assume that the Pod is scheduled to the current worker node, and the normalized value μ of the average value of the current working state CPU load is calculated respectively by formulas (1)-(4)cpu The normalized value μ of the average of the memory load in the current working state memory The normalized value σ of the standard deviation of the CPU load in the current working state cpu and the normalized value σ of the standard deviation of the memory load in the current working state memory ; where α and β are hyperparameters used to adjust the influence of the standard deviation;
[0019] μ cpu =(avg cpu +r cpu ) / cap (1)
[0020] μ memory =(avg memory +r memory ) / cap (2)
[0021]
[0022]
[0023] (3) Through the normalized values of the average and standard deviation of the load in the current working state of this node, the corresponding CPU load risk value risk cpu and the memory load risk value risk memory are calculated by formulas (5) and (6), and scored according to the risk values. The CPU load score score cpu and the memory load score score memory are calculated by formulas (7) and (8). If the average and standard deviation of the current load of this node are higher, the risk value will be higher and the score will be lower. Conversely, if the average and standard deviation of the current load of this node are lower, the risk value will be lower and the score will be higher;
[0024] risk cpu =(μ cpu +σ cpu ) / 2 (5)
[0025] risk memory =(μ memory +σ memory ) / 2 (6)
[0026] score cpu =(1-risk cpu )*100 (7)
[0027] score memory =(1-risk memory )*100 (8).
[0028] (4) Finally, apply the adaptively dynamically changing penalty factors γ and η to the scores of different load types, which are given by formulas (9) and (10). The penalty factors γ and η are used to adjust the respective score weights of the CPU and memory; these penalty factors adaptively change with the cluster state. When the load type is CPU, the penalty factor γ adaptively and dynamically changes as follows: when the score is relatively high, the risk of exceeding the CPU load is relatively low, but at this time, the overall utilization rate of the cluster may be low. Therefore, γ is appropriately reduced and finally maintained at a dynamic balance state between the risk resistance ability and the cluster utilization rate. When the load type is memory, the penalty factor η also adaptively and dynamically changes; the final score finalScore is obtained from formula (11). Among them, γ0 and η0 are the initial values of the hyperparameters, and ω1, ω2, ε1, and ε2 are the redundancy correction terms. At the same time, an influence weight λ is introduced as a supplement to the scoring rule. The final score finalScore will be affected by the historical score pastScore, and the degree of influence is determined according to its own needs (formula (12)), and the obtained final score will be stored in the cache for use in the next scheduling logic scoring (formula (13));
[0029]
[0030]
[0031] finalScore = γ * score cpu + η * score memory , (11)
[0032] finalScore = λ * finalScore + (1 - λ) * pastScore, (12)
[0033] pastScore = finalScore (13).
[0034] Through the above steps, the optimal worker node Worker can be selected. Finally, on the premise of maximizing the resource utilization rate as much as possible, the risk resistance ability of the Kubernetes cluster can be effectively improved.
[0035] Among them, the penalty factors γ and η are used to adjust the respective score weights of the CPU and memory. At the same time, the value of the influence weight λ is proportional to the degree to which the current score is affected by the historical score.
[0036] Compared with the default scheduling algorithm of Kubernetes, the innovation points of the present invention are mainly in the following aspects:
[0037] (1) The default scheduling algorithm of Kubernetes is relatively simple and cannot meet the needs of the cluster to resist the risk of load breakdown. The present invention improves the preference algorithm in the default scheduling algorithm of Kubernetes, introduces the mean and standard deviation to evaluate the cluster resource status, and adds corresponding penalty factors to balance the resource status in different dimensions, optimizing the scoring rules.
[0038] (2) Introduce the influence weight to the scoring rule as a supplement. When the resource status and scoring scores of two Worker nodes are the same, incorporate the historical scores obtained by this node into the scoring rule, which can preferably select a more suitable Worker node for this load task and effectively ensure the load balance of the cluster. Description of the Drawings
[0039] Figure 1 is the basic way for users to interact with the Kubernetes cluster.
[0040] Figure 2 is the flowchart of the specific steps implementation of the load scheduling algorithm of the present invention. Detailed Embodiment
[0041] The present invention will be further described below in conjunction with the embodiments and the drawings.
[0042] Figure 1 is the basic way for users to interact with the Kubernetes cluster. A Kubernetes cluster consists of a master node Master and multiple worker nodes Worker. Users can interact with the master node Master through a graphical interface or a command-line interface, and customize the deployment of load tasks through a resource manifest file.
[0043] The embodiment comparison includes the following steps.
[0044] Step 1. Build a distributed cluster based on the Kubernetes framework of version 1.20. In the present invention, only one Intel NUC is used as the master node Master, and two Nvidia Jetson AGX Xavier are used as worker nodes Worker to form a reliable distributed cluster. At the same time, different stress testing containers are used to create a high-load environment for the distributed cluster. The CPU requirements in the resource declaration list yaml files of these two different stress testing containers are both set to 80%. Although the CPU declaration requirements in the resource declaration list yaml files are the same, the actual load of worker node Worker1 in cluster 1 is smaller than that of the other worker node Worker2 in cluster 1, which is for comparing the differences in schedulers in the two clusters. At the same time, another cluster 2 is built with the same configuration for comparison. Among them, cluster 1 uses the default scheduler of the native Kubernetes framework, and cluster 2 uses the optimized scheduler of the present invention. It should be noted that the scheduling method of the present invention is based on the default scheduler Kube-Scheduler component of the Kubernetes framework. According to the main formula in the above invention content, the scheduling logic of the optimized algorithm is coded and then deployed as a daemon process component of the cluster. Secondly, create a resource manifest file to deploy additional workloads. The resource manifest file includes the container images in the workload Pod, the CPU requirements and memory requirements of computing resources, and the corresponding label information, etc. In this embodiment, a target detection neural network YOLO is used to perform real-time inference on the video stream as the workload, and the corresponding development environment is packaged into a Docker image, and the image is declared in the resource manifest file. Then, the local Kubectl is operated through the command line interface to parse the configuration information in the resource manifest and construct the corresponding HTTP request according to its content and send it to the API Server. After a series of admission control and other processes, the API Server will create our workload Pod object in the distributed key-value data storage system etcd of the cluster. At the same time, the controller in the master node will asynchronously create the resource topology on which the Pod depends. At this time, although the deployed Pod object is now stored in etcd, it has not been deployed to the actual worker node Worker and requires the next operation.
[0045] Step 2. In this embodiment, the monitoring metric collector located in the master node Master of the Kubernets framework and the Kubelet component located in the worker node Worker are used to collect the load information in the cluster within a certain period of time. The monitoring metric collector grabs the load data on the worker node every 5 minutes, and at the same time stores the load data within 15 minutes in the cache for other components to use. Since only one additional workload is deployed in this embodiment, the process of priority sorting is omitted in the process of this embodiment, and it directly reaches the pre-selection stage of the scheduler. The pre-selection algorithms include functions such as checking whether the node is normal, checking whether there is excessive pressure on the node memory, checking whether the resource requirements of the Pod can be met by the node, and checking whether the node can meet the affinity or anti-affinity of the Pod. Finally, a list of nodes that meet the basic requirements of the load is obtained. There is no difference between Cluster 1 and Cluster 2 in all the previous processes.
[0046] Step 3. Next comes the key optimization stage of the present invention. The default scheduler of Cluster 1 does not need to be modified at all. The scheduler of Cluster 2 is modified based on the Scheduler Framework on the basis of the core code of the scheduler, and the scheduling logic of the optimization algorithm is developed and extended under the premise of the original scheduling logic. The relevant parameter settings for this embodiment are as follows: It is determined to be a CPU-intensive load according to the load characteristics of the object detection inference task. Therefore, the initial value of the weight γ that affects the CPU weight ratio is set to 0.8, which is greater than the weight η that affects the memory weight ratio of 0.2. At the same time, λ is set to 0.9, so that the score of the past evaluation in the cache can affect the score of the current evaluation to a certain extent. At the same time, after a series of parameter tuning and comparison, α and β are set to 0.5 and 2 respectively, which are most suitable for quantifying the risk value of the load change degree in this embodiment. All the above parameters can be deployed by passing parameters through the resource manifest file.
[0047] For two high-load clusters, when additional load arrives, the default scheduler Kube-Scheduler component of cluster 1 can only judge according to the CPU demand in the resource declaration list yaml file. When the stress test load CPU demand of the working node in cluster 1 is the same as 80%, the scheduling of the additional load task cannot be judged according to the actual load of the working node. In this way, the additional load task may be scheduled to the working node Worker2 with a high actual load, which may break through the bearing capacity of the working node Worker2 and cause the working node to crash. The optimized scheduler of cluster 2, after collecting the load information in the cluster within a certain period of time through the monitoring indicator collector and the Kubelet component, further quantifies the acquired information into a risk value through the algorithm of the present invention, and then converts the risk value into a percentage score. Finally, according to the score result, it is determined that the actual load in Worker1 is smaller than the actual load in Worker2, and the disaster tolerance capacity of Worker1 is larger than the disaster tolerance capacity in Worker2. The optimized scheduler is more inclined to schedule the additional load to Worker1. Similarly, if more additional loads arrive later, the scheduler will also adaptively put the entire cluster in a dynamic load balancing state based on the actual load of the cluster. The above results show that the method of the present invention has more advantages than the default scheduler. By quantifying the risk value of the actual load of the working node and accurately scheduling the additional load, the anti-risk ability of the Kubernetes cluster is improved, which reflects the effectiveness and superiority of the present invention.
[0048] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. An adaptive risk-resistant Kubernetes cluster scheduling method based on a penalty factor, wherein, The Kubernetes cluster consists of a master node (Master) and multiple worker nodes (Worker); users interact with the master node (Master) through a graphical interface or a command-line interface, and customize the deployment of load tasks through a resource manifest file; it is characterized in that a certain amount of computing resources are reserved and the reserved computing resource pool is adaptively changed according to the state of the cluster, so as to resist sudden high-load computing tasks; the specific steps are as follows: Step 1. First, prepare a yaml file containing the resource request manifest of the application load image; when the user issues a command to deploy the workload through a graphical interface or a command-line interface, the Kubectl component of the master node parses the configuration in the request manifest and constructs the HTTP request parameters of the corresponding object according to its content; after passing the corresponding checks, the request is sent to the API Server component of the master node in the form of HTTP for processing; when the request is verified, the API Server component of the master node creates a corresponding resource object in the distributed key-value data storage system etcd; finally, the APIServer component constructs an HTTP response and returns it to the client; Step 2. Use a monitoring metric collector to collect the load information within the sliding window of each worker node in the cluster, and the length of the sliding window is determined by the desired monitoring period; the monitoring metric collector collects resource load information from the Kubelet component in the worker node (Worker) and exposes them in the cluster through the corresponding application interfaces for other various components to use; after obtaining the load information in different dimensions, first evaluate the priority of the pods to be scheduled, adjust the pod with the highest priority to the head of the scheduling queue, and adjust the pod with the lowest priority to the end of the scheduling queue, and then use the preselection algorithm of the default scheduler to screen each worker node; Step 3. Use an optimization algorithm to score the list of worker nodes obtained in the previous step and find an optimal worker node; The optimization algorithm is implemented by modifying the Scheduler Framework based on the core code of the scheduler. Specifically, plugin extension points are set on all critical paths of the original scheduling logic. Accordingly, plugins are developed without modifying the core code of the Kubernetes scheduler to extend the scheduling logic of the optimization algorithm. The specific process of finding the optimal worker node according to this optimization algorithm is as follows: (1) First, obtain the requested resources of the Pod to be scheduled, including CPU resource r cpu and memory resource r memory . Then, calculate the average value avg cpu and standard deviation std cpu of the CPU utilization rate of the current node to be scored over a period of time, the average value avg memory and standard deviation std memory of the memory utilization rate, and at the same time obtain the computing resource load limit cap of the current node; (2) Assume that the Pod is scheduled to the current working node, and the normalized value μ of the average CPU load in the current working state is calculated by formulas (1)-(4) respectively cpu , the normalized value μ of the average memory load in the current working state memory , the normalized value σ of the standard deviation of the CPU load in the current working state cpu and the normalized value σ of the standard deviation of the memory load in the current working state memory ; where α and β are hyperparameters used to adjust the influence of the standard deviation μ cpu = (avg cpu + r cpu ) / cap, (1) μ memory = (avg memory + r memory ) / cap, (2) (3) Through the normalized values of the average and standard deviation of the current working state load of this node, the corresponding CPU load risk value risk is calculated by formulas (5) and (6). cpu and the memory load risk value risk memory , and a score is given according to the risk value. The CPU load score score is calculated by formulas (7) and (8). cpu and the memory load score score memory ; if the average and standard deviation of the current load of this node are higher, the risk value will be higher and the score will be lower. On the contrary, if the average and standard deviation of the current load of this node are lower, the risk value will be lower and the score will be higher. risk cpu =(μ cpu +σ cpu ) / 2, (5) risk memory =(μ memory +σ memory ) / 2, (6) score cpu =(1 - risk cpu ) * 100, (7) score memory = (1 - risk memory ) * 100; (8) (4)Finally, apply the adaptively dynamically varying penalty factors γ and η to the scores of different load types, which are given by formulas (9) and (10). The penalty factors γ and η are used to adjust the respective score weights of the CPU and memory; these penalty factors vary adaptively with the cluster state; when the load type is CPU, the penalty factor γ varies adaptively and dynamically as follows: when the score is relatively high, the risk of exceeding the CPU load is relatively low, but at this time the overall utilization rate of the cluster may be low. Therefore, γ is appropriately reduced and finally maintained in a state of dynamic balance between the risk resistance and the cluster utilization rate; when the load type is memory, the penalty factor η also varies adaptively and dynamically; the final score finalScore is obtained from formula (11); where γ0 and η0 are the initial values of the hyperparameters, and ω1, ω2, ε1, and ε2 are redundancy correction terms; at the same time, an influence weight λ is introduced to supplement the scoring rule. The final score finalScore will be affected by the historical score pastScore, and the degree of influence is determined according to its own needs (formula (12)), and the obtained final score will be stored in the cache for use in the next scheduling logic scoring (formula (13)); finalScore = γ * score cpu + η * score memory , (11) finalScore = λ * finalScore + (1 - λ) * pastScore, (12) pastScore = finalScore (13) Through the above steps, the optimal worker node is selected.
2. The adaptive anti-risk Kubernetes cluster scheduling method based on a penalty factor according to claim 1, wherein, After the optimal Worker node is selected, start the asynchronous binding process, that is, write the name of the current optimal Worker node into the NodeName field of the Pod to be scheduled. Next, run the specific container load on the worker node Worker, which is specifically completed by the Kubelet service process on the worker node Worker; in the Kubernetes cluster, each worker node Worker starts a Kubelet service process. It queries the Pod in the API Server in the master node Master at regular intervals, filters the Pod list whose NodeName matches its own node, and when it detects that a new Pod object has not been created on the worker node Worker, the Kubelet performs some pre-operations and creates a pause container through the container runtime interface CRI, sets the network for the Pod through the container network interface CNI, and finally pulls the image defined in the resource manifest through the container runtime interface CRI, creates the real container load and starts it, finally achieving the load balancing of the cluster.
Citation Information
Patent Citations
Dynamic load balancing resource scheduling method based on Kubernetes
CN110780998A
Cluster resource load balancing method and device, electronic equipment and medium
CN114443284A