Multi-factor combination virtual container cluster intelligent scheduling method
By combining the load prediction of LSTM and LightGBM models with the TOPSIS resource scheduling algorithm, a multi-factor combined intelligent scheduling method for virtual container clusters was designed. This method solves the problems of resource waste and imbalance in Kubernetes when the load changes, and achieves efficient and stable operation of the cluster.
Patent Information
- Application Number
- CN202511526900.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-10-24
AI Technical Summary
Kubernetes' autoscaling strategy lags behind load changes in cloud environments, failing to respond promptly to sudden surges in traffic and dynamic load changes, leading to resource waste and load imbalance, and lacking dynamic migration strategies.
A combined load prediction model based on LSTM and LightGBM, combined with the TOPSIS multidimensional resource scheduling algorithm, was designed to create an intelligent scheduling method for virtual container clusters with multiple factors. The method achieves automatic scaling and resource reallocation through load prediction and uses a custom Kubernetes extended scheduler for dynamic deployment and migration of Pods.
It improves the utilization rate of cloud platform resources, enhances cluster performance and load balancing, reduces resource waste, and ensures the stability and efficiency of the cluster when the load changes.
Smart Images

Figure CN121029319A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the fields of cloud computing and container orchestration technology, and in particular to an intelligent scheduling method for a multi-factor combined virtual container cluster. Background Technology
[0002] With the continuous development of computer technology, cloud computing has become a popular computing model to support the processing of large amounts of data using commercial computer clusters. It mainly has three service layers: IaaS (Infrastructure as a Service), PaaS (Platform as a Service), and SaaS (Software as a Service). As the scale and complexity of applications increase, the disadvantages of traditional monolithic applications have gradually become apparent, such as low stability and low reliability. Therefore, microservice architecture has become a better choice when building applications.
[0003] Microservice architecture is a divide-and-conquer approach based on "decomposition" proposed to address the increasing complexity of software in cloud environments. It advocates breaking down applications into a series of small services, each focusing on a single business function and running in an independent, isolated environment. In large-scale microservice-based application systems, the sheer number of microservices and their continuous online evolution lead to complex dependencies during runtime. These characteristics present new challenges to resource management during microservice operation. Microservices experience varying loads and resource demands at different times. To prevent service quality degradation and resource waste, it's necessary to continuously modify the resources allocated to microservices, ensuring appropriate resource supply under varying load pressures to meet customer demands for responsiveness.
[0004] Before the advent of container technology, microservice architectures primarily employed virtualization for system deployment. In this approach, clusters allocated, scaled, or reduced application resources by adding, removing, or migrating virtual machines. Creating a virtual machine instance required running a complete operating system for resource isolation, resulting in long startup times and high costs. Container technology offered a new solution for improving cloud resource utilization efficiency. Containers on the same host share the same host kernel, reducing the overhead of CPU and memory caused by virtual machines. Containers are more lightweight and have extremely fast startup times. Docker, as a container creation tool, has developed rapidly in recent years, bringing new opportunities and challenges to the cloud computing field. Due to Docker's simplicity and convenience in application development and deployment, as well as its rapid start-up and shutdown and low resource consumption, many domestic and international vendors have invested heavily in Docker-based cloud platforms. However, Docker does not provide advanced container management functions; it only offers image creation, downloading, uploading, and container lifecycle management. Furthermore, with the continuous growth of business, a container cluster management system is increasingly needed for the overall management of containers.
[0005] To address the challenges of container management, several open-source orchestration tools have emerged in the industry, such as Apache Mesos, Google Kubernetes, and Docker Swarm. Among these, Kubernetes is one of the most renowned container orchestration frameworks in both academia and industry, and has become the most popular open-source container cluster scheduling system within the Docker ecosystem. Kubernetes boasts powerful container orchestration capabilities, adheres to microservice architecture principles, and can provide containerized applications with functions such as resource scheduling, deployment and execution, service discovery, elastic scaling, and upgrades.
[0006] In real-world environments, cluster servers often experience unpredictable load demands. Therefore, when applications experience massive traffic, it's necessary to rapidly increase the number of nodes, while reducing the number of nodes to lower costs when traffic decreases. Currently, Kubernetes' autoscaling strategy has limitations. Adjusting the number of instances based on the current load is lagging, and it lacks timely and effective measures to handle sudden surges in traffic and dynamic load changes. Kubernetes' default scheduling strategy also has several shortcomings. For example, it cannot guarantee a balanced cluster load after elastic scaling, potentially leading to single resource bottlenecks. Furthermore, it can only allocate resources to initially deployed resources, lacking dynamic migration strategies. Therefore, with the creation and destruction of containers, a large amount of resource fragmentation occurs, resulting in resource waste. Summary of the Invention
[0007] To address the aforementioned technical issues, embodiments of this application propose a multi-factor combined intelligent scheduling method for virtual container clusters. This method optimizes and improves upon Kubernetes scheduling strategies by incorporating a cluster auto-scaling strategy based on load prediction, thereby enhancing the resource utilization of the cloud platform and improving cluster performance.
[0008] To achieve the above objectives, embodiments of this application propose a multi-factor combined intelligent scheduling method for virtual container clusters, applicable to containers orchestrated using Kubernetes. The method includes the following steps: performing feature engineering on the historical load dataset of the target cluster to obtain a training set; establishing a combined load prediction model based on an LSTM (Long Short Term Memory) model and a LightGBM (Light Gradient Boosting Machine) model; iteratively training the combined load prediction model using the training set until convergence to obtain a trained combined load prediction model; using the trained combined load prediction model to predict the future load rate of the target cluster; based on the predicted values output by the trained combined load prediction model, and combining an auto-scaling strategy based on load prediction, automatically scaling the target cluster, including predictive scaling up and predictive scaling down; obtaining the resource usage of the target cluster; and based on the resource usage of the target cluster, combining a TOPSIS (Technique for Order Preference by Similarity to Display) strategy... The solution uses a multi-dimensional resource scheduling algorithm to determine the score of each candidate node, reallocate resources in the target cluster based on the scores of each candidate node, and select a suitable Pod (the smallest schedulable and manageable basic unit in a Kubernetes cluster) for redeployment. Data collection, load prediction, autoscaling, and resource scheduling are all implemented by a self-designed Kubernetes extended scheduler.
[0009] To achieve the above objectives, embodiments of this application also propose a computer-readable storage medium storing a computer program that, when executed by a processor, enables a multi-factor combined intelligent scheduling method for virtual container clusters as described above.
[0010] Optionally, feature engineering is performed on the historical load dataset of the target cluster to obtain a training set, including: The historical load dataset of the target cluster is preprocessed, including missing anomaly handling, duplicate anomaly handling, splitting and dimensionality reduction, to obtain the preprocessed historical load dataset. Feature engineering is performed on the preprocessed historical load dataset to obtain the training set; the feature engineering constructs the training set from three aspects, namely time-series features, cross features, and aggregation features. A combined load prediction model is established based on the LSTM and LightGBM models. The combined load prediction model is iteratively trained using the training set until convergence, resulting in the trained combined load prediction model, which includes: The LSTM model and the LightGBM model are combined to establish a combined load prediction model. The LSTM model uses temporal features as input for training, while the LightGBM model uses network parameters, temporal features, cross features, and aggregation features from the LSTM model as input for training. After both models are trained to meet the preset convergence conditions, the trained combined load prediction model is obtained. The LSTM model's network structure consists of one LSTM layer and multiple fully connected layers. Each layer is configured with BatchNorm normalization and DropOut regularization.
[0011] Optionally, predictive scaling includes: The trained combined load prediction model is used to predict the load rate of the target cluster for the next five time points. If the predicted values for at least three time points are greater than the preset upper limit of the load rate, it is determined that the load rate of the target cluster is too high and expansion is required. Let the number of nodes required in the future be... The number of nodes that need to be added is , and Represented as: ; ; in, This indicates the current number of nodes. This indicates the current load rate. This represents the maximum value among the predicted values corresponding to the next five time points; After determining the number of nodes that need to be added, it is necessary to further determine the migration queue of Pods. For Pods with service status, the principle of not migrating should be adopted, and for Pods without service status, the principle of partial migration should be adopted. Based on CPU (Central Processing Unit) utilization, memory utilization, disk utilization, and bandwidth utilization, the node load rate of each node is calculated, and nodes with a load rate greater than the cluster average load rate are identified as high-load nodes. Select from each high-load node A Pod that is in a state of no service. The sum of the load rates of the non-service Pods is no greater than the absolute value of the difference between the load rate of the high-load node and the average load rate of the cluster. Sort all selected Pods in a non-service state according to their load rate from highest to lowest, and then select the Pods with the highest load rates. Pods in a high state of no service form a Pod migration queue.
[0012] Optionally, predictive scaling down includes: The trained combined load prediction model is used to predict the load rate of the target cluster at five future time points. If the predicted values at all five time points are less than the preset lower limit of the load rate, and the load rate at the current time point is also less than the preset lower limit of the load rate, then it is determined that the load rate of the target cluster is too low and a scaling-down operation is required. Iterate through each node of the target cluster, calculate the load rate of the target cluster after deleting the current node. If the load rate of the target cluster after deleting the current node is less than the preset upper limit threshold of the load rate, then migrate all pods running on the current node. After the migration is completed, reclaim the resources of the current node and complete the scaling down of the current node. When migrating a Pod with a service status across nodes, first unload the Pod with a service status from the current node, then detach its disk directory Volumes that can be shared by multiple containers from the current node, then reattach its Volumes to the new node, and finally mount the Pod with a service status to the new node.
[0013] Optionally, the resource usage of the target cluster obtained includes five metrics: CPU utilization, memory utilization, disk utilization, bandwidth utilization, and the number of devices in each node. Based on the resource usage of the target cluster, and combined with a multi-dimensional resource scheduling algorithm based on TOPSIS, the score of each candidate node is determined, including: Based on the number of network I / O requests made by the Pod and the network I / O usage of the cluster nodes, nodes with remaining bandwidth less than the preset remaining bandwidth threshold are filtered out, and the remaining nodes are selected as candidate nodes. Construction by The decision matrix consists of the number of candidate nodes and 5 standard numbers. The five standard numbers correspond to the five indicators mentioned above, and the decision matrix is... This can be expressed by the formula: ; in, Indicates the first CPU utilization of each candidate node Indicates the first Memory utilization of each candidate node. Indicates the first Disk utilization of each candidate node, Indicates the first Bandwidth utilization of each candidate node, Indicates the first The number of containers within each candidate node; Using the reciprocal method, the decision matrix is... The value of each item in the matrix is positiveized to obtain the positiveized decision matrix. ; The decision matrix after positive transformation The value of each item in the matrix is normalized to obtain the normalized decision matrix. ; The normalized decision matrix The scores are input into the TOPSIS algorithm to obtain the score matrix of each candidate node. Rating matrix This can be expressed by the formula: ; in, Indicates the first The scores of each candidate node, indicated in the upper right corner. This indicates the transpose operation.
[0014] Optionally, resources in the target cluster are reallocated based on the scores of each candidate node, and suitable Pods are selected for redeployment, including: When a new Pod is created, define a variable representing the node with the highest score. ; Based on the rating matrix Iterate through each candidate node; if the score of the current candidate node is higher than... Then Replace it with the current candidate node; otherwise, directly compare the next candidate node. After traversing all candidate nodes, the optimal candidate node with the highest score is obtained, and the new Pod is deployed to the optimal candidate node.
[0015] Optionally, the extended scheduler is implemented using a scheduling framework. The scheduling framework defines a set of extension points, which support custom scheduling logic by implementing the interfaces defined by the extension points and registering the extensions to the extension points. Each Pod scheduling is divided into two phases: the scheduling period and the binding period. The scheduling period is when the Pod selects a node, and the binding period is when the decision is applied to the target cluster. The scheduling period and the binding period together are called the scheduling context. Extension points include at least Filter extension points and Score extension points. Filter extension points are used to exclude nodes that cannot run Pods, which is equivalent to the pre-selection stage. Score extension points are used to score all candidate nodes. The score result is an integer within a range, which is equivalent to the selection stage. The same plugin can be registered on multiple extension points to perform complex or stateful tasks. When initializing the scheduler, a profile is automatically created. The profile defines the scheduler's configuration, and its implementation is KubeSchedulerProfile, which is the configuration passed in when generating the YAML (structured text format for data serialization) file. To implement a custom scheduling plugin, you need to register your custom algorithm with the target cluster and recompile the Scheduler. Finally, you can insert the scheduling plugin to be used by configuring the KubeSchedulerConfiguration object in Kubernetes.
[0016] Optionally, the extended scheduler is independent of the default scheduler in Kubernetes and consists of four parts: a data acquisition module, a load prediction module, an autoscaling module, and a resource scheduling module. The data acquisition module is implemented by running Prometheus in the target cluster. It is responsible for monitoring the load of the target cluster, including monitoring node resource usage and Pod container resource usage, obtaining the resource usage of the target cluster, and collecting, analyzing and storing the historical load data of the target cluster. The load prediction module is packaged and runs in the target cluster as a container. It is responsible for performing feature engineering on the historical load dataset of the target cluster to obtain a training set. The training set is used to iteratively train the combined load prediction model until convergence, resulting in a trained combined load prediction model. The trained combined load prediction model is then used to predict the future load rate of the target cluster. The autoscaling module is responsible for automatically scaling the target cluster based on the predicted values output by the trained combined load prediction model and the autoscaling strategy based on load prediction. The autoscaling process includes predictive scaling up and predictive scaling down. The resource scheduling module is responsible for determining the score of each candidate node based on the resource usage of the target cluster and the multi-dimensional resource scheduling algorithm based on TOPSIS. Based on the scores of each candidate node, the resources of the target cluster are reallocated, and suitable Pods are selected for redeployment.
[0017] Optionally, the data acquisition module builds a resource monitoring component based on Prometheus, and uses a combination of Node-Exporter, cAdvisor, Prometheus and Grafana to monitor resources; The Node-Exporter component is responsible for obtaining resource usage information of worker nodes in the target cluster. The selected monitoring points include cpu, meminfo, diskstats, and netstat, which monitor CPU utilization, memory utilization, disk utilization, and bandwidth utilization, respectively. The cAdvisor component is responsible for monitoring and collecting data on container resources in the Pod. The InfluxDB time-series database is used to implement persistent storage of load time-series data. The Prometheus component is responsible for collecting load metric data of the nodes. By modifying the Prometheus configuration file to set the collection period, it periodically collects and stores the data in InfluxDB. The Grafana component is responsible for visualizing the monitoring data. The Node-Exporter component is deployed as a DaemonSet on worker nodes, and cAdvisor is integrated into Kubelet as the default startup item in Kubernetes. The data collected by both is analyzed by Prometheus and stored in the time-series database InfluxDB. The data obtained after analysis is displayed on the web interface through Grafana. During monitoring, data requests are made through the HTTP API provided by Prometheus. After receiving the data request, Prometheus processes the request and returns the requested data in JSON format. The resource monitoring component then decodes, analyzes, and stores the data for use by other components.
[0018] The embodiments of this application propose an intelligent scheduling method for virtual container clusters with multiple factors, which has at least the following advantages compared to the default scheduling strategy of Kubernetes.
[0019] First, due to the complexity of real-world scenarios in cloud environments, clusters often face sudden surges in traffic. Without proper control, this can easily lead to insufficient cluster resources. The fundamental solution is to scale up or down at the node level to provide more server resources and maintain the stability of the entire cluster. Therefore, the load forecasting proposed in this application aims to provide a basis for the automatic scaling of the target cluster, scaling up before heavy traffic arrives to ensure service quality and scaling down to save resources when traffic is low. This application combines the LightGBM and LSTM models, fully leveraging the advantages of both models to effectively improve the prediction accuracy of the combined load forecasting model while ensuring predictive performance.
[0020] Second, TOPSIS is a multi-criteria decision analysis algorithm that can optimize the optimization phase. This application comprehensively considers the CPU utilization, memory utilization, disk utilization, bandwidth utilization, and number of internal devices in the target cluster. Based on the TOPSIS multi-dimensional resource scheduling algorithm, it determines the score of each candidate node, reallocates resources in the target cluster based on the score, and selects suitable Pods for redeployment. This effectively improves the overall resource balance of the target cluster and avoids resource skew among cluster nodes, thus preventing single resource bottlenecks.
[0021] Third, this application designs and implements a Kubernetes extended scheduler. The implementation of the extended scheduler provides runtime support for scheduling methods and can run effectively in the cluster. After designing the scheduler's architecture and modules, it was developed using the Go language and successfully ran on the built cluster. Because the extended scheduler can perform automatic scaling in advance, the cluster is more stable and can effectively improve the cluster's load balancing compared to Kubernetes' default scheduler. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0023] Figure 1 This is a flowchart of an intelligent scheduling method for a multi-factor combined virtual container cluster provided in one embodiment of this application; Figure 2 This is a structural diagram of a combined load prediction model provided in one embodiment of this application; Figure 3 This is a general design diagram of a resource scheduling strategy provided in one embodiment of this application; Figure 4 This is a flowchart of cluster autoscaling processing provided in one embodiment of this application; Figure 5 This is a schematic diagram of the scheduling context of a Pod and the extension points of the scheduling framework disclosed in one embodiment of this application; Figure 6 This is an overall architecture diagram of the extended scheduler provided in one embodiment of this application; Figure 7This is a functional module partitioning diagram of an extended scheduler provided in one embodiment of this application; Figure 8 This is an architecture diagram of a data acquisition module provided in one embodiment of this application; Figure 9 This is a schematic diagram of the working principle of the load prediction module provided in one embodiment of this application; Figure 10 This is a schematic diagram illustrating the working principle of an automatic telescopic module provided in one embodiment of this application; Figure 11 This is a schematic diagram of the working principle of the resource scheduling module provided in one embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0025] One embodiment of this application proposes an intelligent scheduling method for multi-factor combined virtual container clusters, applicable to containers orchestrated using Kubernetes. The implementation details of the intelligent scheduling method for multi-factor combined virtual container clusters proposed in this embodiment are described below. The following implementation details are provided for ease of understanding and are not necessary for implementing this solution.
[0026] The specific process of the intelligent scheduling method for multi-factor combined virtual container clusters proposed in this embodiment can be described as follows: Figure 1 As shown, it includes: Step 11: Perform feature engineering on the historical load dataset of the target cluster to obtain a training set. Build a combined load prediction model based on the LSTM model and the LightGBM model. Iteratively train the combined load prediction model using the training set until convergence to obtain the trained combined load prediction model.
[0027] In practical implementation, the foundation of resource scheduling is accurate load forecasting, which relies on a high-performance combined load forecasting model. Since the numerical variables in the dataset have been rounded, the originally dense variables have become sparse. The LightGBM model has a natural advantage for sparse variables, therefore this embodiment selects the LightGBM model to model the data. Furthermore, because the LSTM model has a strong ability to capture advanced time-series data patterns, this embodiment selects advanced network parameters as features, and after training, these parameters are fed into the LightGBM model. First, historical load data of the target cluster is collected to form the historical load set. Then, feature engineering is performed on the historical load dataset to obtain the training set. Next, a combined load forecasting model is built based on the LSTM and LightGBM models. The combined load forecasting model is iteratively trained using the training set until convergence, resulting in the trained combined load forecasting model. This combined model approach has higher accuracy compared to mainstream load forecasting models.
[0028] In one example, the structure of the combined load forecasting model is as follows: Figure 2 As shown, the process is mainly divided into two parts: data processing and model combination. Data processing includes preprocessing the data and mining effective feature combinations (feature engineering). First, the historical load dataset of the target cluster is preprocessed, including handling missing and duplicate anomalies, splitting, and dimensionality reduction, to obtain a preprocessed historical load dataset. Next, feature engineering is performed on the preprocessed historical load dataset to obtain a training set. Feature engineering constructs the training set from three aspects: temporal features, cross features, and aggregated features. In terms of model combination, the LSTM model and the LightGBM model are combined to establish a combined load prediction model. The LSTM model uses temporal features as input for training, while the LightGBM model uses network parameters from the LSTM model training, temporal features, cross features, and aggregated features as input for training. After both models are trained to meet the preset convergence conditions, the trained combined load prediction model is obtained.
[0029] The LSTM model's network structure consists of one LSTM layer and multiple fully connected layers. This embodiment adds seven fully connected layers to the LSTM model. While this increases the LSTM model's ability to fit nonlinear models, it also slows down the convergence speed during training as the network depth increases. This is because the distribution of activation input values before nonlinear transformations in deep neural networks gradually shifts or changes during training as the network depth increases. The slow convergence is generally due to the overall distribution gradually approaching the upper and lower limits of the nonlinear function's range, causing gradient vanishing in lower-level neural networks during backpropagation. This is the fundamental reason why training deep neural networks becomes increasingly slower. Therefore, this embodiment incorporates BatchNorm normalization into each layer of the network. BatchNorm normalization, through normalization, pulls the distribution of input values for any neuron in each layer back to a standard normal distribution with a mean of 0 and a variance of 1. This ensures that the activation input values fall within the region where the nonlinear function is highly sensitive to input. Thus, small changes in the input lead to large changes in the loss function, significantly improving training speed, accelerating the convergence process, simplifying parameter adjustment, and reducing the training difficulty of the LSTM model. Furthermore, to prevent overfitting of the LSTM model, this embodiment adds DropOut regularization to each layer. DropOut regularization is a neural network regularization technique that temporarily discards neural network units with a specified probability during model training. During training, the PReLu activation function can avoid the gradient vanishing problem caused by excessive network depth. The initial learning rate is set to 0.001, and the Adam optimizer is used to adjust the learning rate to prevent the model from getting trapped in local minima.
[0030] In the process of controlling and optimizing the algorithm, the LightGBM model controls parameters such as the number of leaves in each tree (num_leaves), learning rate, minimum amount of data in a leaf, metric function, number of parallel threads (num_threads), bagging_fraction (for sample sampling), and feature_fraction (for feature subsampling). Feature_fraction and bagging_fraction can speed up training and control overfitting.
[0031] Step 12: Use the trained combined load prediction model to predict the load rate of the target cluster. Based on the predicted value output by the trained combined load prediction model, combine the automatic scaling strategy based on load prediction to perform automatic scaling on the target cluster.
[0032] In practical implementation, after predictive analysis of the historical load data of the target cluster, the predicted load value needs to be used to automatically scale the target cluster. After the cluster is expanded or scaled down, in order to make more reasonable use of cluster resources, it is inevitable to redeploy the Pod resource objects within the cluster, that is, migrate them to more suitable nodes. Therefore, this embodiment designs appropriate resource scheduling strategies (such as...) to address the two major issues of automatic scaling and Pod migration. Figure 3 (As shown). Step 12 will first introduce the autoscaling part. After obtaining the trained combined load prediction model, the trained combined load prediction model can be used to predict the load rate of the target cluster. Based on the predicted values output by the trained combined load prediction model, combined with the load prediction-based autoscaling strategy, the target cluster will be automatically scaled. The autoscaling process includes two types: predictive scaling up and predictive scaling down.
[0033] When faced with traffic surges, resource load increases accordingly, easily leading to access blockages or system crashes, resulting in significant losses. Conversely, when system access decreases, resources may become idle, causing waste. Therefore, the cluster needs to dynamically respond to load changes. Elastic scaling is often divided into adding or removing resources at the Pod and node levels. When elastic scaling is required, the application-level capacity planning changes first. After new container replicas are deployed to nodes, node resources may fluctuate drastically. In such cases, if container replicas are not restricted, CPU, memory, and other resources may be exhausted, leading to server downtime. Therefore, when the cluster faces a sudden increase in traffic, problems such as insufficient cluster resources may arise. The fundamental solution to address the inability of cluster resources to cope with the resource changes caused by sudden traffic surges is to scale the cluster at the node level, providing more server resources to maintain the stability of the entire cluster.
[0034] To cope with the upcoming surge in traffic, the resource scheduling strategy proposed in this embodiment is based on the automatic scaling technology of the cluster. Starting from the scaling of nodes, it provides more resources to the container platform, fundamentally solving the problem of insufficient resources caused by the surge in traffic.
[0035] Load prediction-based autoscaling strategies are fundamental to dynamic resource scheduling. By analyzing the future cluster load based on the output of a combined load prediction model, it determines whether scaling up or down is necessary. For scaling up, the number of nodes to be added is calculated based on the load prediction results; for scaling down, only one node is removed at a time. Before adding or removing nodes, Pods need to be migrated to maintain a balanced use of cluster resources.
[0036] Typically, before a cluster can auto-scale, a trigger signal needs to be given to the CA component to indicate when cluster scaling will be triggered. After the cluster auto-scales, the number of nodes may increase or decrease. When nodes are added or removed from the cluster, it is inevitable to consider issues such as Pod migration and the selection of target nodes for Pods.
[0037] Cluster scaling requires specific scaling conditions to be met. In this embodiment, the scaling trigger condition is that the load metric of a certain worker node reaches a set threshold. The cluster load threshold can be manually adjusted. When the predicted load metric reaches the upper threshold, cluster expansion is required; when it reaches the lower threshold, cluster reduction is required. The principle and process of automatic cluster scaling can be described as follows: Figure 4 As shown.
[0038] In one example, when the predicted cluster load rate triggers the upper limit threshold, it indicates that the resources in the cluster are about to be exhausted, and the cluster needs to be expanded. However, there is often a certain difference between the predicted cluster load and the actual cluster load, and triggering the threshold at a certain point in time may result in a sudden peak in cluster load at some point in the future, but it will quickly drop back to the normal range. In this case, there is no need to expand the cluster.
[0039] Based on this, this embodiment uses the trained combined load prediction model to predict the load rate of the target cluster at five future time points. If the predicted values at at least three time points are greater than the preset upper limit threshold of the load rate, it is determined that the load rate of the target cluster is too high and expansion is required.
[0040] Let the number of nodes required in the future be... The number of nodes that need to be added is , and Represented as: ; ; in, This indicates the current number of nodes. This indicates the current load rate. This represents the maximum value among the predicted values corresponding to the next five time points.
[0041] After determining the number of nodes to be added, it is necessary to further determine the Pod migration queue, i.e., which Pods need to be migrated. Kubernetes contains stateful and stateless services. Stateless services do not store persistent data locally, and multiple service instances respond identically to the same user request. Therefore, dynamically starting and stopping a stateless service Pod will not affect other Pods. Stateful services, however, require local storage of persistent data, and node instances have dependent topological relationships. Stopping any instance Pod in the cluster may result in data loss. Therefore, two points need to be considered when migrating Pods: First, it is necessary to distinguish between stateful and stateless service Pods. Since the migration process for stateful service Pods is time-consuming, requiring migration to be performed on every affected container, to save migration time, a no-migration policy is adopted for stateful service Pods, primarily migrating stateless Pods. Second, migrating Pods consumes cluster resources, so it is necessary to select as few Pods as possible for migration. Therefore, a partial migration policy is adopted for stateless service Pods.
[0042] When determining the migration queue for Pods, the first step is to calculate the node load rate for each node based on CPU utilization, memory utilization, disk utilization, and bandwidth utilization. Nodes with a load rate higher than the cluster average load rate are identified as high-load nodes. Next, select from each high-load node... A Pod that is in a state of no service. The sum of the load rates of the non-service Pods should not exceed the absolute value of the difference between the load rate of the high-load node and the cluster average load rate. Finally, all the selected non-service Pods are sorted from highest to lowest load rate, and then the Pods with the highest load rates are selected. Pods in a high state of no service form a Pod migration queue.
[0043] In one example, if the predicted values for all five time points are less than the preset lower limit of the load rate, and the load rate at the current time point is also less than the preset lower limit of the load rate, then it is determined that the future load rate of the target cluster is too low, and a scaling-down operation is required. At this point, each node in the target cluster is traversed, and the load rate of the target cluster after deleting the current node is calculated. If the load rate of the target cluster after deleting the current node is less than the preset upper limit of the load rate, then all pods running on the current node are migrated. After the migration is complete, the resources of the current node are reclaimed, completing the scaling-down of the current node.
[0044] When migrating a Pod with a service status across nodes, first unload the Pod with a service status from the current node, then detach its disk directory Volumes that can be shared by multiple containers from the current node, then reattach its Volumes to the new node, and finally mount the Pod with a service status to the new node.
[0045] In general, during the automatic scaling of the cluster, the first step is to analyze historical load data and determine whether the conditions for scaling up or down have been met based on the prediction results. If the conditions for scaling up are met, the number of nodes to be added is calculated, and nodes from the node pool are added to the cluster as worker nodes. Next, target nodes for Pods are selected, Pods are selected for migration, and Pod migration is performed. If the conditions for scaling down are met, the load rate of each worker node in the cluster is calculated, and all Pods on nodes with a load rate below a lower threshold are migrated. After migration, the resources of those nodes are reclaimed.
[0046] Step 13: Obtain the resource usage of the target cluster. Based on the resource usage of the target cluster, and combined with the multi-dimensional resource scheduling algorithm based on TOPSIS, determine the score of each candidate node. Based on the score of each candidate node, reallocate the resources of the target cluster and select suitable Pods for redeployment.
[0047] In practical implementation, cluster resources are diverse, including not only CPU and memory, but also disk, network interface bandwidth, and other resources. Kubernetes' default scheduling strategy only considers CPU and memory, making it unsuitable for more complex scenarios. In real-world use, a comprehensive consideration of multi-dimensional resource usage and a balanced approach to resource utilization across all dimensions are crucial for making informed scheduling decisions. This embodiment proposes a cluster scaling solution at the node level and introduces a multi-dimensional resource scheduling algorithm based on TOPSIS. By collecting and storing cluster load data, load prediction is first performed. Then, the cluster is automatically scaled using the proposed scaling solution. Finally, the TOPSIS-based resource scheduling algorithm comprehensively considers multi-dimensional resource usage, including candidate node CPU utilization, memory utilization, disk utilization, bandwidth utilization, and the number of containers within each node. Based on the multi-dimensional resource allocation and actual usage of each node, a score is determined for each candidate node. Resources in the target cluster are then reallocated based on these scores, and suitable Pods are selected for redeployment. This process selects the optimal node for each newly created or migrated Pod, improving cluster load balancing and resource utilization.
[0048] In one example, node filtering is the first step. Based on the Pod's network I / O requests and the network I / O usage of cluster nodes, nodes with remaining bandwidth less than a preset remaining bandwidth threshold are filtered out. To prevent excessively high usage of any metric on a node, CPU utilization, memory utilization, disk utilization, and bandwidth utilization are also filtered. An upper limit of 80% is set for the utilization of each resource. If the utilization of any resource on a node exceeds the upper limit, that node is filtered out, and the remaining nodes are the candidate nodes.
[0049] In applying the TOPSIS algorithm, the core step is to construct an input matrix for the five indices and then perform standardization calculations. First, a matrix is constructed from the five indices... The decision matrix consists of the number of candidate nodes and 5 standard numbers. The five standard numbers correspond to the five indicators mentioned above, and the decision matrix is... This can be expressed by the formula: ; in, Indicates the first CPU utilization of each candidate node Indicates the first Memory utilization of each candidate node. Indicates the first Disk utilization of each candidate node, Indicates the first Bandwidth utilization of each candidate node, Indicates the first The number of containers within each candidate node.
[0050] Next, we will use the reciprocal method to process the decision matrix. The value of each item in the matrix is positiveized to obtain the positiveized decision matrix. .
[0051] Then, the decision matrix after positive transformation... The value of each item in the matrix is normalized to obtain the normalized decision matrix. The decision matrix after normalization. This can be expressed by the formula: ; in, Indicates the first CPU utilization after normalization of candidate nodes Indicates the first Memory utilization rate after normalization of candidate nodes Indicates the first Disk utilization rate after normalization of candidate nodes Indicates the first Bandwidth utilization after normalization of candidate nodes Indicates the first The number of nodes in a candidate node after normalization.
[0052] Finally, the normalized decision matrix is... The scores are input into the TOPSIS algorithm to obtain the score matrix of each candidate node. Rating matrix This can be expressed by the formula: ; in, Indicates the first The scores of each candidate node, indicated in the upper right corner. This indicates the transpose operation.
[0053] When a new Pod is created, define a variable representing the node with the highest score. Subsequently, based on the rating matrix Iterate through each candidate node; if the score of the current candidate node is higher than... Then Replace it with the current candidate node; otherwise, directly compare the next candidate node. After traversing all candidate nodes, the optimal candidate node with the highest score is obtained, and the new Pod is deployed to the optimal candidate node.
[0054] To support the resource scheduling strategy proposed in this embodiment, this embodiment also designs and develops a Kubernetes extended scheduler based on the implementation of the Kubernetes extended scheduler, in order to realize data collection, load prediction, automatic scaling and resource scheduling.
[0055] This embodiment uses a scheduling framework to implement an extended scheduler (custom scheduler). The scheduling framework defines a set of extension points, which allow users to customize scheduling logic by implementing the interfaces defined by the extension points, and register the extensions with the extension points. Figure 5 It demonstrates the Pod scheduling context and extension points in the scheduling framework.
[0056] Each Pod scheduling is divided into two phases: the scheduling period and the binding period. The scheduling period is when the Pod selects a node, and the binding period is when the decision is applied to the target cluster. The scheduling period and the binding period together are called the scheduling context.
[0057] Extension points include at least Filter extension points and Score extension points. Filter extension points are used to exclude nodes that cannot run Pods, which is equivalent to the pre-selection stage. Score extension points are used to score all candidate nodes. The score result is an integer within a range, which is equivalent to the optimization stage. The same plugin can be registered on multiple extension points to perform complex or stateful tasks.
[0058] During scheduler initialization, a profile is automatically created. This profile defines the scheduler's configuration, and its implementation is `KubeSchedulerProfile`, which is the configuration passed in when generating the YAML file. To implement a custom scheduling plugin, the custom algorithm needs to be registered in the target cluster, the scheduler needs to be recompiled, and finally, the desired scheduling plugin is inserted by configuring a `KubeSchedulerConfiguration` object in Kubernetes. This process can be described as follows: Figure 6 As shown.
[0059] Extended schedulers are independent of Kubernetes' default scheduler, such as... Figure 7 As shown, it consists of four parts: a data acquisition module, a load prediction module, an automatic scaling module, and a resource scheduling module.
[0060] The data acquisition module is implemented by running Prometheus in the target cluster. It is responsible for monitoring the load of the target cluster, including monitoring node resource usage and Pod container resource usage, obtaining the resource usage of the target cluster, and collecting, analyzing, and storing historical load data of the target cluster.
[0061] To achieve load prediction for container clusters and enable autoscaling based on these predictions, the first step is to monitor the container cluster. By monitoring cluster load metrics, historical load data can be stored, providing data support for subsequent load prediction and autoscaling.
[0062] Since this embodiment studies the scaling of node granularity, it is necessary to monitor the resource usage information of cluster nodes. There are already many mature solutions for node monitoring, such as Nagios and Zabbix. Zabbix is more suitable for monitoring physical machine environments, while Prometheus is more suitable for monitoring cloud environments. Therefore, this embodiment chooses to build a resource monitoring system based on Prometheus.
[0063] To make the monitoring system as complete as possible, this embodiment uses a combination of components such as Node-Exporter, cAdvisor, Prometheus, and Grafana to monitor resources.
[0064] The Node-Exporter component is responsible for obtaining resource usage information of worker nodes in the target cluster. The selected monitoring points include cpu, meminfo, diskstats, and netstat, which monitor CPU utilization, memory utilization, disk utilization, and bandwidth utilization, respectively. The cAdvisor component is responsible for monitoring and collecting data on container resources in Pods. The InfluxDB time-series database is used to implement persistent storage of load time-series data. The Prometheus component is responsible for collecting node load metric data. By modifying the Prometheus configuration file to set the collection period, it periodically collects and stores the data in InfluxDB. The Grafana component is responsible for visualizing the monitoring data.
[0065] The overall architecture of the data acquisition module is as follows: Figure 8 As shown, the Node-Exporter component is deployed as a DaemonSet on worker nodes, while cAdvisor is integrated into the Kubelet as the default startup item in Kubernetes. The data collected by both is analyzed by Prometheus and stored in the time-series database InfluxDB. The analyzed data is then displayed on a web interface via Grafana. During monitoring, data requests are made through the HTTP API provided by Prometheus. Upon receiving the request, Prometheus processes it and returns the data in JSON format. The resource monitoring component then decodes, analyzes, and stores the data for use by other components.
[0066] The load prediction module is packaged and runs in the target cluster as a container. It is responsible for performing feature engineering on the historical load dataset of the target cluster to obtain a training set. The training set is used to iteratively train the combined load prediction model until convergence, resulting in a trained combined load prediction model. The trained combined load prediction model is then used to predict the future load rate of the target cluster.
[0067] The load prediction module reads historical load data from InfluxDB, a language written in Go, and operates on the database using SQL-like statements. After obtaining the historical load data, it performs predictions based on the load prediction algorithm proposed in this embodiment. The prediction results are then used by the autoscaling module to calculate the number of nodes to scale. The prediction process is as follows: Figure 9As shown, the load prediction module is packaged and runs in the cluster as a container. This module preprocesses the historical load data it acquires, trains the model, predicts the cluster resource load at five future time points, and finally stores the prediction results in a time series database for the auto-scaling module to retrieve and call.
[0068] The autoscaling module is responsible for automatically scaling the target cluster based on the predicted values output by the trained combined load prediction model and the autoscaling strategy based on load prediction. The autoscaling process includes predictive scaling up and predictive scaling down.
[0069] The auto-scaling module reads the results from the load prediction module in the database, performs calculations using the auto-scaling strategy proposed in this embodiment, and uses the CA component to automatically scale the cluster based on the calculation results. After scaling up, the Pods running on the nodes need to be migrated. Similarly, before removing a node (i.e., scaling down), all Pods on that node also need to be migrated. The specific process is as follows: Figure 10 As shown.
[0070] The resource scheduling module is responsible for determining the score of each candidate node based on the resource usage of the target cluster and the multi-dimensional resource scheduling algorithm based on TOPSIS. Based on the scores of each candidate node, the resources of the target cluster are reallocated, and suitable Pods are selected for redeployment.
[0071] The resource scheduling module is the core of the extended scheduler developed in this embodiment. Its function is to find the optimal target node for the Pod queue. It can run custom scheduling algorithms and return the results to the scheduler. The extended scheduler designed in this embodiment is mainly used to verify the scheduling algorithm proposed in this embodiment, and therefore can run multi-dimensional resource scheduling algorithms based on TOPSIS. Furthermore, after the pre-selection and optimization algorithms of the resource scheduling module run, the most suitable node can be allocated to the Pod queue to be scheduled, and the allocation result is written to ETCD (an open-source distributed key-value storage system). The pre-selection stage filters out nodes that do not meet the conditions at the beginning, and the optimization stage scores the pre-selected nodes. The node with the highest score is the most suitable node for deploying Pods, and therefore, the Pod will be bound to that node. The workflow of the resource scheduling module is as follows: Figure 11 As shown.
[0072] This embodiment proposes a multi-factor combined intelligent scheduling method for virtual container clusters, which has at least the following advantages compared to the traditional default scheduling strategy of Kubernetes.
[0073] First, due to the complexity of real-world scenarios in cloud environments, clusters often face sudden surges in traffic. Without proper control, this can easily lead to insufficient cluster resources. The fundamental solution is to scale up or down at the node level to provide more server resources and maintain the stability of the entire cluster. Therefore, the load forecasting proposed in this application aims to provide a basis for the automatic scaling of the target cluster, scaling up before heavy traffic arrives to ensure service quality and scaling down to save resources when traffic is low. This application combines the LightGBM and LSTM models, fully leveraging the advantages of both models to effectively improve the prediction accuracy of the combined load forecasting model while ensuring predictive performance.
[0074] Second, TOPSIS is a multi-criteria decision analysis algorithm that can optimize the optimization phase. This application comprehensively considers the CPU utilization, memory utilization, disk utilization, bandwidth utilization, and number of internal devices in the target cluster. Based on the TOPSIS multi-dimensional resource scheduling algorithm, it determines the score of each candidate node, reallocates resources in the target cluster based on the score, and selects suitable Pods for redeployment. This effectively improves the overall resource balance of the target cluster and avoids resource skew among cluster nodes, thus preventing single resource bottlenecks.
[0075] Third, this application designs and implements a Kubernetes extended scheduler. The implementation of the extended scheduler provides runtime support for scheduling methods and can run effectively in the cluster. After designing the scheduler's architecture and modules, it was developed using the Go language and successfully ran on the built cluster. Because the extended scheduler can perform automatic scaling in advance, the cluster is more stable and can effectively improve the cluster's load balancing compared to Kubernetes' default scheduler.
[0076] The steps described above are for clarity only. In implementation, they can be combined into one step, or some steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0077] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement a multi-factor combined intelligent scheduling method for virtual container clusters as described in the above method embodiments.
[0078] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method described in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0079] Those skilled in the art will understand that the above embodiments are specific implementations of this application, and in practical applications, various changes can be made in form and detail without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A multi-factor combined intelligent scheduling method for virtual container clusters, applicable to containers orchestrated using Kubernetes, characterized in that, The method includes: Feature engineering is performed on the historical load dataset of the target cluster to obtain a training set. A combined load prediction model is built based on the LSTM model and the LightGBM model. The combined load prediction model is iteratively trained using the training set until convergence to obtain the trained combined load prediction model. The trained combined load prediction model is used to predict the future load rate of the target cluster. Based on the predicted values output by the trained combined load prediction model, combined with the load prediction-based autoscaling strategy, the target cluster is automatically scaled. The autoscaling process includes predictive scaling up and predictive scaling down. Obtain the resource usage of the target cluster. Based on the resource usage of the target cluster, and combined with the multi-dimensional resource scheduling algorithm based on TOPSIS, determine the score of each candidate node. Based on the score of each candidate node, reallocate the resources of the target cluster and select suitable Pods for redeployment. Data collection, load prediction, autoscaling, and resource scheduling are all implemented by a self-designed Kubernetes extended scheduler.
2. The intelligent scheduling method for a multi-factor combined virtual container cluster according to claim 1, characterized in that, Feature engineering is performed on the historical load dataset of the target cluster to obtain the training set, which includes: The historical load dataset of the target cluster is preprocessed, including missing anomaly handling, duplicate anomaly handling, splitting and dimensionality reduction, to obtain the preprocessed historical load dataset. Feature engineering is performed on the preprocessed historical load dataset to obtain the training set; the feature engineering constructs the training set from three aspects, namely time-series features, cross features, and aggregation features. A combined load prediction model is established based on the LSTM and LightGBM models. The combined load prediction model is iteratively trained using the training set until convergence, resulting in the trained combined load prediction model, which includes: The LSTM model and the LightGBM model are combined to establish a combined load prediction model. The LSTM model uses temporal features as input for training, while the LightGBM model uses network parameters, temporal features, cross features, and aggregation features from the LSTM model as input for training. After both models are trained to meet the preset convergence conditions, the trained combined load prediction model is obtained. The LSTM model's network structure consists of one LSTM layer and multiple fully connected layers, with each layer having BatchNorm normalization and DropOut regularization.
3. The intelligent scheduling method for a multi-factor combined virtual container cluster according to claim 1, characterized in that, Predictive scaling includes: The trained combined load prediction model is used to predict the load rate of the target cluster for the next five time points. If the predicted values for at least three time points are greater than the preset upper limit of the load rate, it is determined that the load rate of the target cluster is too high and expansion is required. Let the number of nodes required in the future be... The number of nodes that need to be added is , and Represented as: ; ; in, This indicates the current number of nodes. This indicates the current load rate. This represents the maximum value among the predicted values corresponding to the next five time points; After determining the number of nodes that need to be added, it is necessary to further determine the migration queue of Pods. For Pods with service status, the principle of not migrating should be adopted, and for Pods without service status, the principle of partial migration should be adopted. Based on CPU utilization, memory utilization, disk utilization, and bandwidth utilization, the node load rate of each node is calculated, and nodes with a load rate greater than the cluster average load rate are identified as high-load nodes. Select from each high-load node A Pod that is in a state of no service. The sum of the load rates of the non-service Pods is no greater than the absolute value of the difference between the load rate of the high-load node and the average load rate of the cluster. Sort all selected Pods in a non-service state according to their load rate from highest to lowest, and then select the Pods with the highest load rates. Pods in a high state of no service form a Pod migration queue.
4. The intelligent scheduling method for a multi-factor combined virtual container cluster according to claim 3, characterized in that, Predictive scaling down includes: The trained combined load prediction model is used to predict the load rate of the target cluster at five future time points. If the predicted values at all five time points are less than the preset lower limit of the load rate, and the load rate at the current time point is also less than the preset lower limit of the load rate, then it is determined that the load rate of the target cluster is too low and a scaling-down operation is required. Iterate through each node of the target cluster, calculate the load rate of the target cluster after deleting the current node. If the load rate of the target cluster after deleting the current node is less than the preset upper limit threshold of the load rate, then migrate all pods running on the current node. After the migration is completed, reclaim the resources of the current node and complete the scaling down of the current node. When migrating a Pod with a service status across nodes, first unload the Pod with a service status from the current node, then detach its disk directory Volumes that can be shared by multiple containers from the current node, then reattach its Volumes to the new node, and finally mount the Pod with a service status to the new node.
5. The intelligent scheduling method for a multi-factor combined virtual container cluster according to claim 1, characterized in that, The resource usage of the target cluster obtained includes five metrics: CPU utilization, memory utilization, disk utilization, bandwidth utilization, and the number of devices in each node. Based on the resource usage of the target cluster, and combined with a multi-dimensional resource scheduling algorithm based on TOPSIS, the score of each candidate node is determined, including: Based on the number of network I / O requests made by the Pod and the network I / O usage of the cluster nodes, nodes with remaining bandwidth less than the preset remaining bandwidth threshold are filtered out, and the remaining nodes are selected as candidate nodes. Construction by The decision matrix consists of the number of candidate nodes and 5 standard numbers. The five standard numbers correspond to the five indicators mentioned above, and the decision matrix is... This can be expressed by the formula: ; in, Indicates the first CPU utilization of each candidate node Indicates the first Memory utilization of each candidate node. Indicates the first Disk utilization of each candidate node, Indicates the first Bandwidth utilization of each candidate node, Indicates the first The number of containers within each candidate node; Using the reciprocal method, the decision matrix is... The value of each item in the matrix is positiveized to obtain the positiveized decision matrix. ; The decision matrix after positive transformation The value of each item in the matrix is normalized to obtain the normalized decision matrix. ; The normalized decision matrix The scores are input into the TOPSIS algorithm to obtain the score matrix of each candidate node. Rating matrix This can be expressed by the formula: ; in, Indicates the first The scores of each candidate node, indicated in the upper right corner. This indicates the transpose operation.
6. The intelligent scheduling method for a multi-factor combined virtual container cluster according to claim 5, characterized in that, Based on the scores of each candidate node, resources in the target cluster are reallocated, and suitable Pods are selected for redeployment, including: When a new Pod is created, define a variable representing the node with the highest score. ; Based on the rating matrix Iterate through each candidate node; if the score of the current candidate node is higher than... Then Replace it with the current candidate node; otherwise, directly compare the next candidate node. After traversing all candidate nodes, the optimal candidate node with the highest score is obtained, and the new Pod is deployed to the optimal candidate node.
7. The intelligent scheduling method for a multi-factor combined virtual container cluster according to claim 1, characterized in that, The extended scheduler is implemented using a scheduling framework. The scheduling framework defines a set of extension points, which support custom scheduling logic by implementing the interfaces defined by the extension points and registering the extensions to the extension points. Each Pod scheduling is divided into two phases: the scheduling period and the binding period. The scheduling period is when the Pod selects a node, and the binding period is when the decision is applied to the target cluster. The scheduling period and the binding period together are called the scheduling context. Extension points include at least Filter extension points and Score extension points. Filter extension points are used to exclude nodes that cannot run Pods, which is equivalent to the pre-selection stage. Score extension points are used to score all candidate nodes. The score result is an integer within a range, which is equivalent to the optimization stage. The same plugin can be registered on multiple extension points to perform complex or stateful tasks. When initializing the scheduler, a profile is automatically created. The profile defines the scheduler's scheduling configuration, and its implementation is KubeSchedulerProfile, which is the configuration passed in when generating the YAML file. To implement a custom scheduling plugin, you need to register the custom algorithm you have written into the target cluster and recompile the Scheduler. Finally, you can insert the scheduling plugin to be used by configuring the KubeSchedulerConfiguration object in Kubernetes.
8. The intelligent scheduling method for a multi-factor combined virtual container cluster according to claim 1, characterized in that, The extended scheduler is independent of the default scheduler in Kubernetes and consists of four parts: a data acquisition module, a load prediction module, an autoscaling module, and a resource scheduling module. The data acquisition module is implemented by running Prometheus in the target cluster. It is responsible for monitoring the load of the target cluster, including monitoring node resource usage and Pod container resource usage, obtaining the resource usage of the target cluster, and collecting, analyzing, and storing historical load data of the target cluster. The load prediction module is packaged and runs in the target cluster as a container. It is responsible for performing feature engineering on the historical load dataset of the target cluster to obtain a training set. The training set is used to iteratively train the combined load prediction model until convergence, resulting in a trained combined load prediction model. The trained combined load prediction model is then used to predict the future load rate of the target cluster. The autoscaling module is responsible for automatically scaling the target cluster based on the predicted values output by the trained combined load prediction model and the autoscaling strategy based on load prediction. The autoscaling process includes predictive scaling up and predictive scaling down. The resource scheduling module is responsible for determining the score of each candidate node based on the resource usage of the target cluster and the multi-dimensional resource scheduling algorithm based on TOPSIS. Based on the scores of each candidate node, the resources of the target cluster are reallocated, and suitable Pods are selected for redeployment.
9. The intelligent scheduling method for a multi-factor combined virtual container cluster according to claim 8, characterized in that, The data acquisition module is built on Prometheus to create a resource monitoring component, and uses a combination of Node-Exporter, cAdvisor, Prometheus and Grafana to monitor resources; The Node-Exporter component is responsible for obtaining resource usage information of worker nodes in the target cluster. The selected monitoring points include cpu, meminfo, diskstats, and netstat, which monitor CPU utilization, memory utilization, disk utilization, and bandwidth utilization, respectively. The cAdvisor component is responsible for monitoring and collecting data on container resources in the Pod. The InfluxDB time-series database is used to implement persistent storage of load time-series data. The Prometheus component is responsible for collecting load metric data of the nodes. By modifying the Prometheus configuration file to set the collection period, it periodically collects and stores the data in InfluxDB. The Grafana component is responsible for visualizing the monitoring data. The Node-Exporter component is deployed as a DaemonSet on worker nodes, and cAdvisor is integrated into Kubelet as the default startup item in Kubernetes. The data collected by both is analyzed by Prometheus and stored in the time-series database InfluxDB. The data obtained after analysis is displayed on the web interface through Grafana. During monitoring, data requests are made through the HTTP API provided by Prometheus. After receiving the data request, Prometheus processes the request and returns the requested data in JSON format. The resource monitoring component then decodes, analyzes, and stores the data for use by other components.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement a multi-factor combined intelligent scheduling method for virtual container clusters as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-factor strategy-based computing power resource optimal scheduling distribution method
CN115550370A
Enterprise computing power layout and intelligent decision mobile application system and implementation method thereof
CN120596264A
Predictive resource allocation and scheduling for a distributed workload
US20240378079A1
Cited By
Intelligent scheduling method and device for multi-hole gate and server
CN121580361A