Container cloud elastic expansion and contraction method based on load prediction
By using PatchMixer prediction model in container cloud to predict load changes and automatically perform scaling operations based on this, the problem that traditional methods cannot accurately deal with dynamic load fluctuations is solved, and more efficient resource management and system stability are achieved.
Patent Information
- Application Number
- CN202510220864.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional cloud resource elastic scaling methods cannot accurately predict and deal with dynamic load fluctuations, resulting in excessive or insufficient resource allocation, affecting system stability and performance.
By collecting container cloud load data, using historical load data to train the PatchMixer prediction model, predict future workloads, and automatically make decisions and perform scaling operations of container cloud applications based on the predicted load data and preset elastic scaling rules.
It realizes more accurate load prediction and elastic scaling of container clouds, improves the timeliness and prediction accuracy of resource adjustments, reduces the risk of overscaling or resource idleness, and improves the utilization rate of container cloud resources and the high availability and stability of application systems.
Smart Images

Figure CN120216096A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of cloud resource prediction and container elastic scaling, and particularly to a method for elastic scaling of container clouds based on load prediction. Background Art
[0002] With the continuous expansion of Internet businesses, enterprises increasingly tend to deploy their business architectures in cloud environments to flexibly respond to the ever-rising computing demands. The cloud environment, with its ability to allocate computing resources on demand, enables enterprises to dynamically adjust resource deployment strategies according to real-time changes in business loads. However, with the rapid popularization of containerization technology, the effective management of cloud computing resources has encountered unprecedented complexity challenges. Container technology, with its advantages of lightweight, high efficiency, and flexibility, has become the preferred solution for building cloud-native applications.
[0003] Traditional cloud resource elastic scaling methods usually rely on static policies or simple threshold-triggered mechanisms to automatically increase or decrease containers by monitoring the load situation. However, these methods often fail to accurately predict and respond to dynamic load fluctuations, resulting in over-allocation or under-allocation of resources, which affects the stability and performance of the system. Especially in scenarios where the load changes frequently and has a time delay, traditional elastic scaling strategies are prone to response lags or over-scaling phenomena, thus increasing the usage cost of cloud resources. In view of this, the academic and industrial communities are actively exploring how to use advanced load prediction technologies to improve elastic scaling strategies. Load prediction technology can integrate historical data and real-time monitoring information to predict future load trends, thereby providing a decision-making basis for dynamic resource adjustment. By introducing a load prediction model, the system can perform resource expansion or reduction operations in advance, effectively adapt to load changes, avoid system overload and resource waste, improve the utilization efficiency of cloud resources, reduce costs, and ensure the high availability of applications and the stability of system performance.
[0004] The present invention collects container cloud load data and stores it in a time series database, trains a PatchMixer prediction model using historical load data to predict future workloads. Based on the predicted load data and preset load elastic scaling rules, the real-time load monitoring data is filtered to obtain the data after rule filtering. Then, the filtered data is input into the load prediction model to generate an average load change trend index for the container cloud application, and based on this index, scaling decisions are made to obtain the scaling method and corresponding scaling parameters, and automatically make decisions and execute the scaling operations of the container cloud application in advance. Summary of the Invention
[0005] In order to solve the problems existing in the above-mentioned prior art, the present invention provides a method for elastic scaling of container clouds based on load prediction to solve the current technical problems.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows: The present invention provides a method for elastic scaling of container cloud based on load prediction, including the following steps: Step S1: Data collection and storage; Step S2: Data preprocessing and rule filtering; Step S3: Training of load prediction model; Step S4: Load prediction and trend analysis; Step S5: Scaling decision and parameter calculation; Step S6: Execution of scaling policy and real-time adjustment.
[0007] Preferably, the implementation process of the step S1: Data collection and storage includes: Regularly collect the real-time load data of each container instance through the monitoring system Prometheus of the container cloud platform, specifically including CPU usage rate, memory usage rate, disk I / O, network traffic and the number of requests; Store the collected multi-dimensional load data in the time series database according to the time stamp.
[0008] Preferably, the implementation process of the step S2: Data preprocessing and rule filtering includes: Clean the original data collected in step S1. The cleaning steps include denoising, missing value filling and outlier detection processing, and use the mean value of the same features in the historical data to fill the missing values; Filter the cleaned data, and only retain the load data when the CPU load exceeds the scaling threshold and in the high load state; Extract the useful features from the preprocessed data, including the average load within the time window, the standard deviation of the load and the peak-to-valley value of the load.
[0009] Preferably, the implementation process of the step S3: Training of load prediction model includes: Divide the load data set processed in step S2 into a training set and a test set, and the division ratio is 80% training set and 20% test set; Build PatchMixer to capture the long-term and short-term dependencies of the data, and initialize the relevant hyperparameters; Perform model training on the training set.
[0010] Preferably, the process of the step S4: Load prediction and trend analysis includes: Use the trained load prediction model to predict the real-time collected load data, and obtain the load change trend within a certain period in the future.
[0011] Preferably, the process of the step S5: Scaling decision and parameter calculation includes: According to the load prediction result and the preset expansion and contraction thresholds, determine whether an expansion or contraction operation is required. The determination conditions include load threshold exceeding, load recovery, and latency factors. The beneficial effects of the present invention are: Compared with traditional load prediction methods, the present invention adopts a deep learning model based on PatchMixer, which can more accurately capture the time series characteristics of cloud resource loads, especially having obvious advantages when dealing with non-linear and non-stationary data. Through high-precision load prediction, it can better provide reliable input data for the elastic scaling of container clouds, significantly improving the timeliness and prediction accuracy of resource adjustment.
[0012] The present invention ensures the high quality and continuity of load data by using the Prometheus time series database and the mean method to fill in missing values. In the data preprocessing stage, through denoising, outlier processing, and feature extraction, the stability and usability of the data are further improved. This efficient data processing method lays a solid foundation for subsequent load prediction and scaling decisions, effectively enhancing the system's response ability to real-time changes.
[0013] By integrating the load prediction result and the intelligent scaling decision-making mechanism, the present invention can dynamically adjust the configuration of cloud resources under the condition of large load fluctuations, achieving precise elastic scaling. Compared with traditional elastic scaling strategies, the present invention can predict load changes in advance and intelligently optimize the scaling parameters, thereby reducing the risk of over-expansion or resource idleness, improving the utilization rate of container cloud resources, and ensuring the high availability and stability of the application system. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above aspects and advantages of the present invention will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, where: Figure 1 is a schematic diagram of the elastic scaling method of the container cloud based on load prediction according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0016] It includes the following steps: Step S1: Data collection and storage: Through the monitoring system of the container cloud platform, collect the real-time load data of container cloud applications, including multi-dimensional metrics such as CPU usage, memory usage, disk I / O, network traffic, etc. Store these data in the form of time series. Each data point consists of a metric name, a set of labels, data, and a timestamp. For example, after collecting the CPU utilization rate, the format is: cpu_usage{container=”nginx”,instance=”9090”} 0.76 1605597776000, where cpu_usage is the metric name, {container=”nginx”,instance=”9090”} is the set of labels, 0.76 is the value, and 1605597776000 is the timestamp. And store these data in the time series database Prometheus in text form according to the timestamp; Step S2: Data preprocessing and filtering: Preprocess the original load data collected in Step S1, including denoising, filling missing values, etc. According to the preset load elastic scaling rules, filter the data to remove irrelevant or abnormal data and obtain a data set that conforms to the rules; Step S3: Train the load prediction model: The historical load data filtered in Step S2 is a multi-dimensional time series matrix. Each sample contains multi-dimensional load metrics within a continuous time window. Divide the input-output pairs according to the sliding window mechanism. Use these data to adjust the model parameters of the PatchMixer model, including the length of the time window, the number of model layers, the size of the hidden layer, the learning rate, etc. Perform training steps such as validation set evaluation and error minimization, analyze the time series characteristics of historical data, and predict the future load change trend; Step S4: Load prediction analysis: Use the trained prediction model to predict the real-time load data and obtain the future load change trend metrics; Step S5: Scaling decision and parameter calculation: Based on the predicted load trend metrics, combined with the threshold rules for scaling out and scaling in, for example, if the predicted CPU usage is greater than 80% and the duration is greater than 10 minutes, or the predicted CPU usage is greater than 85% and the duration is greater than 5 minutes, then perform a scaling out operation. If the predicted CPU usage is less than 30% and the duration is greater than 10 minutes, or the predicted CPU usage is less than 40% and the duration is greater than 5 minutes, then perform a scaling in operation. Determine the resource requirements of the container cloud application, and determine whether scaling out or scaling in is required according to the prediction results and calculate the specific parameters for scaling; Step S6: Execute the scaling policy and adjust in real time: According to the results of the scaling decision and parameter calculation, automatically execute the scaling operation of the container cloud, and continuously monitor the load changes and system response of the container cloud.
[0017] The steps of data collection and storage in step S1 are as follows: Step S11: Regularly collect the real-time load data of each container instance through the monitoring system Prometheus of the container cloud platform, specifically including CPU usage, memory usage, disk I / O, network traffic, and the number of requests; Step S12: Store the collected multi-dimensional load data in the time series database according to the time stamp; Step S13: According to the importance and storage requirements of the load data, retain the high-frequency data with a generation frequency greater than once per second or an access frequency greater than five times per minute and the data of the past seven days for short-term data retention, retain the historical data for long-term data retention and reduce the query frequency, store the data in layers according to time and importance, retain the recent data in high-performance storage, and migrate the old data to low-cost storage media; The steps of data preprocessing and rule filtering in step S2 are as follows: Step S21: Clean the original data collected in step S1. The cleaning steps include denoising, missing value filling, and outlier detection processing. Use the mean value of the same features in the historical data to fill the missing values to ensure the continuity and integrity of the data. The calculation formula for the filled missing values is: (1), In the formula, is the load value at the time point where the missing value is located, is the total number of data points used to fill the missing values, is the non-missing load value in the historical data ( represents the historical data points); Step S22: Filter the cleaned data, and only retain the load data with CPU load exceeding the scaling threshold and in the high-load state; Step S23: Extract the useful features from the preprocessed data, including the average load within the time window, the standard deviation of the load, and the peak and valley values of the load; The steps of training the load prediction model in step S3 are as follows: Step S31: Divide the load data set processed in step S2 into a training set and a test set, with a division ratio of 80% for the training set and 20% for the test set; Step S32: Build PatchMixer to capture the long-term and short-term dependencies of the data, and initialize the relevant hyperparameters; Step S33: Perform model training on the training set. The model analyzes the time series characteristics of the historical load data, adjusts the parameters and performs cross-validation, and learns the rules of load changes; Step S34: Evaluate the trained model using the validation set. The evaluation metrics include MAE (Mean Absolute Error), RMSE (Root Mean Square Error), and MAPE (Mean Absolute Percentage Error). Evaluate the prediction accuracy of the model according to the metrics and adjust the model to reduce the error. The calculation formula of MAE is: (2), where, is the mean absolute error, which is the average of the absolute errors between the predicted value and the actual value, is the sum of the data points, is the actual load value, is the predicted load value; The calculation formula of is: where, is the root mean square error, which measures the average of the squared errors between the predicted value and the actual value, , and are the same as formula (2); The calculation formula of is: where, is the mean absolute percentage error, which represents the relative error of the predicted value deviating from the actual value, , and are the same as formula (2); The steps of the above-mentioned step S4 for load prediction and trend analysis are as follows: Step S41: Use the trained load prediction model to predict the real-time collected load data to obtain the load change trend within a certain period in the future; Step S42: Analyze the prediction results to determine the load change trend, and obtain the growth rate, decline rate, and stable period of the load according to the predicted trend; Step S43: For the predicted load data, combine the real-time load situation to perform error correction. If there is a large difference between the actual load and the predicted value, update the parameters of the prediction model in time or recalculate the prediction results; The steps of the above-mentioned step S5 for scaling determination and parameter calculation are as follows: Step S51: According to the load prediction results and the preset scaling-up and scaling-down thresholds, determine whether scaling-up or scaling-down operations are required. The determination conditions include load threshold exceeding, load recovery, and delay factors; Step S52: According to the results of the load prediction and threshold judgment, calculate the specific number of container instances, resource configuration, and scaling time for scaling; The determination of scaling and the calculation of parameters in step S6 are as follows: Step S61: Automatically perform scaling operations according to the determination result of step S5, adjust the number of container instances or resource configurations, start new containers, stop redundant containers, and adjust container resource configurations; Step S62: After the scaling operation, continuously monitor the load changes and system performance of the container cloud platform. Timely feedback the effect of the scaling policy through the monitored data. If problems are found or the prediction error is large, adjust the policy in a timely manner; Step S63: Dynamically adjust the scaling policy based on real-time data and monitoring feedback to optimize resource allocation and system performance. If it is found that the accuracy of the prediction model is not high, consider retraining the model or adjusting the prediction policy.
[0018] Compared with traditional load prediction methods, the present invention adopts a deep learning model based on PatchMixer, which can more accurately capture the time series characteristics of cloud resource loads, especially having obvious advantages when dealing with non-linear and non-stationary data. Through high-precision load prediction, it can better provide reliable input data for the elastic scaling of container clouds, significantly improving the timeliness and prediction accuracy of resource adjustment.
[0019] The present invention ensures the high quality and continuity of load data by adopting the Prometheus time series database and the mean method to fill in missing values. In the data preprocessing link, data stability and availability are further improved through denoising, outlier processing, and feature extraction. This efficient data processing method lays a solid foundation for subsequent load prediction and scaling decisions, effectively enhancing the system's response ability to real-time changes.
[0020] The present invention can dynamically adjust the configuration of cloud resources through comprehensive load prediction results and an intelligent scaling decision-making mechanism to achieve precise elastic scaling in the case of large load fluctuations. Compared with traditional elastic scaling strategies, the present invention can predict load changes in advance and intelligently optimize scaling parameters, thereby reducing the risk of over-provisioning or resource idleness, improving the utilization rate of container cloud resources, and ensuring the high availability and stability of the application system.
[0021] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A container cloud elastic scaling method based on load prediction, characterized in that: The following steps are involved: Step S1: data collection and storage; Step S2: data preprocessing and rule filtering; Step S3: load prediction model training; Step S4: load forecasting and trend analysis; Step S5: capacity expansion and contraction determination and parameter calculation; Step S6: Execution and real-time adjustment of expansion and contraction strategies.
2. The method for elastically scaling container cloud based on load prediction according to claim 1 is characterized in that: The implementation process of step S1: data collection and storage includes: The container cloud platform's monitoring system Prometheus regularly collects real-time load data for each container instance, including CPU usage, memory usage, disk IO, network traffic, and number of requests. The collected multi-dimensional load data is stored in the time series database according to timestamps.
3. The method for elastically scaling container cloud based on load prediction according to claim 1 is characterized in that: The implementation process of step S2: data preprocessing and rule filtering includes: Clean the raw data collected in step S1. The cleaning steps include denoising, missing value filling and outlier detection. The missing values are filled using the mean of the same features in the historical data. Filter the cleaned data to retain only the load data when the CPU load exceeds the scaling threshold and is in a high load state; Useful features are extracted from the preprocessed data, including the average load within the time window, the standard deviation of the load, and the peak and valley values of the load.
4. The method for elastically scaling container cloud based on load prediction according to claim 1 is characterized in that: The implementation process of step S3: load prediction model training includes: Divide the load data set processed in step S2 into a training set and a test set, with a division ratio of 80% training set and 20% test set; Build PatchMixer to capture the long-term and short-term dependencies of data and initialize related hyperparameters; Train the model on the training set.
5. The method for elastically scaling container cloud based on load prediction according to claim 1 is characterized in that: The process of step S4: load forecasting and trend analysis includes: Use the trained load prediction model to predict the load data collected in real time to obtain the load change trend within a certain period of time in the future.
6. The method for elastically scaling container cloud based on load prediction according to claim 1 is characterized in that: The step S5: expansion / contraction determination and parameter calculation process includes: Based on the load prediction results and the preset expansion and reduction thresholds, it is determined whether expansion or reduction operations are required. The determination conditions include load threshold exceeding, load recovery and delay factors.
Citation Information
Cited By
Resource allocation method and device of server
CN120723484A
Server resource allocation methods and devices
CN120723484B
Elastic resource beforehand early warning and scheduling method and system based on load prediction
CN120909742A