Container cloud elastic scaling method based on GRU-attention mechanism
By adopting the load prediction model of GRU-attention mechanism and an improved resource scheduling strategy in the container cloud environment, the resource prediction accuracy and real-time problems under dynamic load and nonlinear resource requirements are solved, and efficient resource utilization and system stability are achieved.
Patent Information
- Application Number
- CN202510201669.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to efficiently handle dynamic load and nonlinear resource requirements in container cloud environments, resulting in insufficient resource prediction accuracy and poor real-time performance, and prone to insufficient or over-supply of resource supply.
The container cloud elastic scaling method based on the GRU-attention mechanism is adopted to construct a GRU-attention hybrid model to perform load prediction, and the resource allocation strategy is dynamically adjusted by combining the improved Horizontal Pod Autoscaler (HPA) mechanism and resource supply evaluation module.
It significantly improves the accuracy of load prediction and resource utilization efficiency, reduces resource waste, and ensures the stability and response speed of the system in the case of burst load and load fluctuations.
Smart Images

Figure CN120196400A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing, and particularly relates to a container cloud elastic scaling method based on GRU-attention mechanism. Background Art
[0002] In modern cloud computing environments, container technology has gradually become the first choice for enterprises and developers due to its efficient resource utilization and flexible deployment capabilities. However, resource scheduling and elastic scaling issues in container cloud environments still pose a major challenge. Traditional resource scheduling methods usually rely on static rules or simple prediction models, making it difficult to handle complex workload patterns and sudden resource demands. In the prior art, models such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have been widely used for resource prediction, but there are still problems such as high computational overhead and insufficient prediction accuracy. Especially when facing sudden loads and non-linear resource demands, the prediction accuracy and real-time performance of existing methods still need to be improved. In addition, the existing Horizontal Pod Autoscaler (HPA) mechanism usually makes scaling decisions based on real-time load data, lacking foresight and prone to causing insufficient or excessive resource supply. Therefore, there is an urgent need for an intelligent elastic scaling method that can efficiently handle dynamic loads and optimize resource allocation. Summary of the Invention
[0003] The object of the present invention is to provide a container cloud elastic scaling method based on gated recurrent unit (GRU) and attention mechanism, aiming to solve the core problems of resource scheduling and load management in container cloud environments. Through intelligent prediction and dynamic scaling strategies, optimize resource utilization efficiency, improve system performance and service quality. Especially when facing sudden loads and non-linear resource demands, it can significantly reduce resource waste and ensure system stability and response speed.
[0004] The technical solution adopted by the present invention to achieve the above object is: a container cloud elastic scaling method based on GRU-attention mechanism, comprising the following steps:
[0005] S1: Construct a GRU-attention hybrid model, and generate a load prediction result through the GRU-attention hybrid model;
[0006] S2: Based on the load prediction result generated in step S1, calculate the number of Pods required in a future period of time through an improved HPA mechanism, and execute the corresponding scaling strategy;
[0007] S3: Continuously analyze resource supply indicators through a resource supply evaluation module, and optimize the scaling decision parameters according to the quantification results of the insufficient supply rate and the over-supply rate to improve resource utilization efficiency.
[0008] The construction of the GRU-attention hybrid model is specifically as follows:
[0009] a. Construct a multi-layer neural network architecture containing bidirectional gated recurrent units;
[0010] Among them, the network architecture contains at least three temporal processing layers;
[0011] b. Introduce an attention mechanism before the output layer of the neural network. The attention mechanism dynamically adjusts the contribution degrees of feature vectors at each time step through a trainable weight matrix;
[0012] c. Use the sliding time window method to sample historical load data to generate a training data set with time correlation;
[0013] d. Optimize the model parameters through an adaptive matrix estimation algorithm to minimize the prediction error function.
[0014] In step S1, the steps of generating a load prediction result through the GRU-attention hybrid model include:
[0015] S11: Use the gated recurrent unit network, i.e., GRU, to process time series data to capture the long-term dependencies of load changes;
[0016] S12: Introduce an attention mechanism to dynamically allocate attention weights to different time steps and features, thereby identifying key load information;
[0017] S13: Generate a load prediction result for a future period of time through multi-layer non-linear mapping and multi-step prediction output.
[0018] The improved HPA mechanism includes:
[0019] (1) When it is predicted that the load will increase, expand the number of Pods in advance;
[0020] (2) When it is predicted that the load will decrease, reduce the number of Pods in advance.
[0021] The resource supply evaluation module realizes the comprehensive optimization of the under-supply rate, over-supply rate, and resource reconfiguration cost by establishing a multi-objective optimization function and using a reinforcement learning algorithm to automatically adjust the scaling policy parameters, thereby ensuring the stability of the system under load fluctuations and efficient resource utilization.
[0022] The resource supply evaluation module performs the following operations:
[0023] S21: Define the under-supply rate index as the proportion of the time when resource demand exceeds supply;
[0024] S22: Define the over-supply rate index as the proportion of the time when the resource utilization rate is lower than the set threshold;
[0025] S23: Establish a multi-objective optimization function to synchronously optimize the under-supply rate, over-supply rate, and resource reallocation cost;
[0026] S24: Use a reinforcement learning algorithm to automatically adjust the scaling policy parameters to optimize the optimization function.
[0027] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements a container cloud elastic scaling method based on the GRU-attention mechanism.
[0028] The present invention has the following beneficial effects and advantages:
[0029] 1) High prediction accuracy: By combining the GRU network and the attention mechanism, the present invention can effectively capture the long-term dependence relationships and key features in time series data, significantly improving the accuracy of load prediction. Experimental results show that compared with traditional methods, the present invention performs excellently in metrics such as RMSE, MAE, and MAPE, especially having higher prediction accuracy when dealing with bursty loads.
[0030] 2) Intelligent resource scheduling: The improved Horizontal Pod Autoscaler (HPA) mechanism adjusts resource allocation in advance based on the prediction results, reducing the situations of under-supply and over-supply of resources and improving the resource utilization efficiency. Experimental data shows that the under-supply index (θ U ) and the under-supply time (T U ) are significantly reduced, and the elastic speed improvement (ε n ) index also shows the improvement of resource scheduling efficiency.
[0031] 3) Efficiently coping with bursty loads: The present invention performs excellently when dealing with bursty loads, being able to quickly respond to the dynamically changing resource demands and ensuring the stability and response speed of the system under extreme load fluctuations. For example, when simulating the bursty traffic during the World Cup, the present invention significantly reduces resource waste while ensuring service quality.
[0032] 4) Reducing operation costs: The present invention effectively reduces the operation costs of the container cloud environment by optimizing resource allocation and reducing resource waste, especially having significant economic benefits in large-scale and highly dynamic load scenarios.
[0033] 5) Wide applicability: The present invention is not only applicable to the container cloud environment but can also be extended and applied to other cloud computing and edge computing scenarios, providing a general solution for intelligent resource scheduling.
[0034] In summary, through the innovative GRU-attention hybrid model and the improved HPA mechanism, the present invention provides an efficient and intelligent solution for elastic scaling in the container cloud environment, significantly improving the resource utilization efficiency and service quality, and having broad application prospects and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 The flowchart of the method for elastic scaling of the container cloud of the present invention;
[0036] Figure 2 The GRU-attention hybrid model diagram of the present invention;
[0037] Figure 3 The schematic diagram of the improved HPA mechanism of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0039] As Figure 2 shown, the present invention proposes a method for elastic scaling of a container cloud based on a gated recurrent unit (GRU) and an attention mechanism. By combining a GRU-attention hybrid model and an improved HPA mechanism, the present invention proposes an innovative solution aimed at overcoming the limitations of the prior art and improving the resource management efficiency of the container cloud environment.
[0040] As Figure 1 shown, it is the flowchart of the method for elastic scaling of the container cloud of the present invention. Combining with the GRU-attention hybrid model in Figure 2 , the present invention provides a method for elastic scaling of a container cloud based on a GRU-attention mechanism, including the following steps:
[0041] S1: (1) Construct a GRU-attention hybrid model, specifically:
[0042] a. Construct a multi-layer neural network architecture including bidirectional gated recurrent units;
[0043] Among them, the network architecture includes at least three time series processing layers;
[0044] b. Introduce an attention mechanism before the output layer of the neural network, and the attention mechanism dynamically adjusts the contribution degree of the feature vectors of each time step through a trainable weight matrix;
[0045] c. Use the sliding time window method to sample the historical load data to generate a training data set with time correlation;
[0046] d. Optimize the model parameters through an adaptive matrix estimation algorithm to minimize the prediction error function.
[0047] (2)Generate the load prediction result through the GRU-attention hybrid model;
[0048] S11: Use the gated recurrent unit network, i.e., GRU, to process time series data to capture the long-term dependencies of load changes;
[0049] S12: Introduce the attention mechanism to dynamically allocate attention weights to different time steps and features, thereby identifying key load information;
[0050] S13: Generate the load prediction result for a future period of time through multi-layer non-linear mapping and multi-step prediction output.
[0051] S2: Based on the load prediction result generated in step S1, calculate the number of Pods required for a future period of time through the improved HPA mechanism, and execute the corresponding scaling policy;
[0052] Among them, the improved HPA mechanism includes:
[0053] (1) When it is predicted that the load will increase, expand the number of Pods in advance;
[0054] (2) When it is predicted that the load will decrease, reduce the number of Pods in advance.
[0055] S3: Continuously analyze the resource supply indicators through the resource supply evaluation module, and optimize the scaling decision parameters according to the quantification results of the under-supply rate and over-supply rate to improve the resource utilization efficiency.
[0056] Among them, the resource supply evaluation module realizes the comprehensive optimization of the under-supply rate, over-supply rate, and resource reconfiguration cost by establishing a multi-objective optimization function and using a reinforcement learning algorithm to automatically adjust the scaling policy parameters, thereby ensuring the stability of the system under load fluctuations and efficient resource utilization.
[0057] The resource supply evaluation module performs the following operations:
[0058] S31: Define the under-supply rate indicator as the proportion of the time when the resource demand exceeds the supply;
[0059] S32: Define the over-supply rate indicator as the proportion of the time when the resource utilization rate is lower than the set threshold;
[0060] S33: Establish a multi-objective optimization function to synchronously optimize the under-supply rate, over-supply rate, and resource reconfiguration cost;
[0061] S34: Use the reinforcement learning algorithm to automatically adjust the scaling policy parameters to make the optimization function reach the optimal.
[0062] Example 1:
[0063] In this embodiment, a GRU-attention hybrid model is first constructed. This model combines the multi-layer neural network architecture of bidirectional gated recurrent units and the attention mechanism. By using the sliding time window method to sample historical load data, a training data set with time correlation is generated. The model optimizes the parameters through the adaptive matrix estimation algorithm to minimize the prediction error function, thereby generating the load prediction results for a future period of time.
[0064] As Figure 3 shown, after the model is constructed, based on the prediction results, the required number of Pods for a future period of time is calculated through the improved Horizontal PodAutoscaler (HPA) mechanism, and the corresponding scaling policy is executed. When it is predicted that the load will increase, the system expands the number of Pods in advance to ensure that user requests can be met during the peak period; when it is predicted that the load will decrease, the system reduces the number of Pods in advance to avoid waste of resources. This prediction-driven method enables the system to intelligently respond to dynamic request loads and improve resource utilization efficiency.
[0065] In addition, the system continuously analyzes resource supply metrics through the resource supply evaluation module, including the under-supply rate and the over-supply rate. The under-supply rate is defined as the proportion of time when resource demand exceeds supply, and the over-supply rate is defined as the proportion of time when resource utilization is lower than the set threshold. The resource supply evaluation module establishes a multi-objective optimization function to synchronously optimize the under-supply rate, the over-supply rate, and the resource reconfiguration cost, and uses a reinforcement learning algorithm to automatically adjust the parameters of the scaling policy to make the optimization function reach the optimum. This process realizes the efficient scheduling of resources in an automated manner, ensuring the stability and efficient resource utilization of the system under load fluctuations.
[0066] In practical applications, the system continuously monitors the operating status and load conditions, and makes dynamic adjustments by combining real-time data with prediction results. Through the GRU-attention hybrid model and the improved HPA mechanism, the system can intelligently respond to dynamic request loads, improve resource utilization efficiency, and ensure that the system has sufficient capacity to handle upcoming requests. At the same time, the resource supply evaluation module continuously adjusts and optimizes the scaling policy through the multi-objective optimization function and the reinforcement learning algorithm to achieve more efficient resource management.
[0067] This embodiment may also include an exception handling mechanism and a performance optimization module to provide a more comprehensive mechanism; among them, the exception handling mechanism is as follows:
[0068] S31: Real-time monitor the confidence interval of the model prediction output;
[0069] S32: When the width of the confidence interval exceeds the safety threshold, automatically switch to the reactive scaling mode;
[0070] S33: After the abnormal event is lifted, restore the predictive scaling mechanism through the progressive weight migration method.
[0071] The performance optimization module then performs the following steps:
[0072] S41: Implement model lightweight processing, including parameter pruning and quantization operations;
[0073] S42: Establish a distributed model inference framework to support multi-node parallel computing;
[0074] S43: Adopt a caching mechanism to store periodic load patterns and accelerate the prediction calculation process.
[0075] Combined with the above embodiments, this embodiment aims to solve the core problems of resource scheduling and load management in the container cloud environment. Through intelligent prediction and dynamic scaling strategies, it optimizes the resource utilization efficiency and improves the system performance and service quality. The present invention is particularly suitable for processing dynamically changing workloads, especially in the scenario of sudden traffic, which can significantly reduce resource waste and ensure the stability and response speed of the system.
[0076] In summary, the present invention realizes intelligent resource scheduling in the container cloud environment through the GRU-attention hybrid model and the improved HPA mechanism. This method not only improves the accuracy of load prediction, but also optimizes the resource allocation strategy, significantly enhancing the resource utilization efficiency and service quality, and providing a novel and effective solution for the elastic scaling of the container cloud environment.
[0077] Those skilled in the art can understand that the above are only the preferred embodiments of the present invention. The features described in various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. It is not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0078] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A container cloud elastic scaling method based on GRU-attention mechanism, characterized in that: The following steps are involved: S1: Build a GRU-attention hybrid model and generate load prediction results through the GRU-attention hybrid model; S2: Based on the load prediction results generated in step S1, the number of Pods required in the future is calculated through the improved HPA mechanism, and the corresponding expansion and contraction strategies are implemented; S3: The resource supply evaluation module continuously analyzes resource supply indicators and optimizes scaling decision parameters based on the quantitative results of undersupply rate and oversupply rate to improve resource utilization efficiency.
2. According to claim 1, a container cloud elastic scaling method based on GRU-attention mechanism is characterized in that: The construction of the GRU-attention hybrid model is specifically as follows: a. Build a multi-layer neural network architecture with bidirectional gated recurrent units; Wherein, the network architecture comprises at least three timing processing layers; b. Introducing an attention mechanism before the output layer of the neural network, the attention mechanism dynamically adjusts the contribution of the feature vector at each time step through a trainable weight matrix; c. Use the sliding time window method to sample historical load data and generate a training data set with time correlation; d. Optimize the model parameters through an adaptive matrix estimation algorithm to minimize the prediction error function.
3. According to the container cloud elastic scaling method based on GRU-attention mechanism according to claim 1, it is characterized in that: In step S1, the step of generating a load prediction result by using the GRU-attention hybrid model includes: S11: Gated recurrent unit network, i.e. GRU, is used to process time series data to capture long-term dependencies of load changes; S12: Introducing the attention mechanism to dynamically assign attention weights to different time steps and features to identify key load information; S13: Generate load forecast results for a period of time in the future through multi-layer nonlinear mapping and multi-step prediction output.
4. According to claim 1, a container cloud elastic scaling method based on GRU-attention mechanism is characterized in that: The improved HPA mechanism comprises: (1) When the load is predicted to increase, expand the number of Pods in advance; (2) When it is predicted that the load will decrease, reduce the number of Pods in advance.
5. According to claim 1, a container cloud elastic scaling method based on GRU-attention mechanism is characterized in that: The resource supply evaluation module establishes a multi-objective optimization function and uses a reinforcement learning algorithm to automatically adjust the expansion and contraction strategy parameters to achieve comprehensive optimization of the undersupply rate, oversupply rate and resource reconfiguration cost, thereby ensuring the stability of the system and efficient resource utilization under load fluctuations.
6. A container cloud elastic scaling method based on GRU-attention mechanism according to claim 1 or 5, characterized in that: The resource supply assessment module performs the following operations: S21: Define the undersupply rate indicator as the percentage of time that resource demand exceeds supply; S22: Define the over-provisioning rate indicator as the percentage of time when the resource utilization is below a set threshold; S23: Establish a multi-objective optimization function to simultaneously optimize the undersupply rate, oversupply rate and resource reallocation cost; S24: Use reinforcement learning algorithm to automatically adjust the expansion and contraction strategy parameters to optimize the optimization function.
7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program. When the computer program is executed by a processor, a container cloud elastic scaling method based on a GRU-attention mechanism as described in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Container resource allocation method and system based on AI drive and medium
CN120448135A
AI-driven container resource configuration method, system, and medium
CN120448135B