Real-time CPU load balancing method and system for cloud computing platform based on reinforcement learning
By applying a load balancing method based on reinforcement learning on the cloud computing platform, monitoring and analyzing CPU load in real time, and optimizing the load balancing strategy through the deep Q network model, the problem of insufficient load balancing real-time and prediction capabilities in the existing technology is solved, and efficient resource utilization and system performance improvement is achieved.
Patent Information
- Application Number
- CN202410904651.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-07-08
AI Technical Summary
The load balancing method of existing cloud computing platforms is difficult to respond to load changes in real time, and lacks effective load prediction and analysis capabilities, resulting in low resource utilization and degradation of system performance.
Using a reinforcement learning-based method, the CPU load is monitored in real time by deploying the data collection module, and timing analysis is performed using sliding window technology to build a deep Q network model to generate and optimize load balancing strategies.
It realizes rapid response and adjustment to load changes, improves resource utilization and system performance, and ensures that the system can achieve real-time CPU load balancing under different load conditions.
Smart Images

Figure CN118885291B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing platforms, and in particular to a real-time CPU load balancing method and system for a cloud computing platform based on reinforcement learning. Background Art
[0002] As a new computing model, cloud computing technology has been widely used in many fields. Cloud computing platforms provide large-scale computing resources and storage capabilities, allowing users to use computing resources on demand without having to pay attention to the underlying hardware facilities. An important feature of cloud computing platforms is the dynamic allocation and management of resources, especially with the support of virtualization technology, which can flexibly allocate computing resources between physical nodes and virtual machines. However, as the scale of cloud computing platforms continues to expand, how to efficiently manage resources and load balancing has become a problem that needs to be solved.
[0003] Existing load balancing methods mainly include static load balancing and dynamic load balancing. Static load balancing methods usually rely on preset rules and fixed algorithms, such as polling method, least connection method, etc. These methods are slow to respond to load changes and are difficult to adapt to dynamically changing load conditions in cloud computing environments. Dynamic load balancing methods achieve load balancing by real-time monitoring and dynamic adjustment of resource allocation, but most existing methods rely on simple heuristic algorithms or preset strategies, and it is difficult to find the global optimal load balancing strategy in a complex cloud computing environment. When dealing with complex and changing load conditions, these methods often cannot effectively optimize resource utilization, resulting in system performance degradation and resource waste.
[0004] In addition, traditional load balancing strategies usually have defects in the following aspects:
[0005] Existing load balancing methods are difficult to respond to load changes in real time. Due to the lack of rapid adaptability to load changes, the system may experience insufficient resources during peak load periods, and may cause idle resources during low load periods. This lack of real-time and adaptability directly affects the system's resource utilization and overall performance.
[0006] Most traditional load balancing methods lack effective load prediction and analysis capabilities and cannot accurately predict future load conditions. This causes the system to rely more on the current load status when allocating resources, and cannot make predictions based on historical data and trends, making it difficult to develop an optimized load balancing strategy.
[0007] In large-scale cloud computing platforms, the optimization and execution efficiency of load balancing strategies are crucial. Traditional methods often have difficulty finding the global optimal load balancing strategy in complex environments, and the strategy execution efficiency is low, which may lead to system overload or resource waste. For example, rule-based load balancing strategies are slow to adjust when faced with sudden load changes and are difficult to cope with instantaneous high load pressure. Summary of the invention
[0008] One object of the present invention is to propose a real-time CPU load balancing method and system for a cloud computing platform based on reinforcement learning, which can adaptively learn and optimize load balancing strategies, thereby achieving real-time and efficient CPU load balancing. This method can not only improve resource utilization and system performance, but also quickly adjust resource allocation strategies in the face of dynamically changing load conditions to ensure stable operation of the system.
[0009] A real-time CPU load balancing method for a cloud computing platform based on reinforcement learning according to an embodiment of the present invention comprises the following steps:
[0010] S1. Deploy a data collection module on the cloud computing platform to monitor the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time, and build a real-time CPU load data set;
[0011] S2, inputting the collected real-time CPU load data set into the load analysis module through a high-speed data transmission channel, parsing and preprocessing the real-time CPU load data set;
[0012] S3. In the load analysis module, the sliding window technique is used to perform time series analysis on the preprocessed real-time CPU load data set to identify the current resource utilization pattern and potential load peaks;
[0013] S4. Based on the results of timing analysis, a load balancing strategy decision model is constructed using the deep Q network model to define the state space, action space and reward function;
[0014] S5. During the training of the deep Q network model, the historical CPU load data set and the simulation environment are used for multiple iterations to optimize the strategy to improve its adaptability and efficiency.
[0015] S6. Deploy the trained and optimized load balancing strategy to the decision module, and generate specific resource allocation and task scheduling solutions through the strategy;
[0016] S7. In the decision module, the load balancing strategy of each virtual machine and physical node is calculated in real time, and CPU resource allocation is adjusted according to the current load status, and task migration and load balancing operations are performed;
[0017] S8. During the dynamic adjustment process, use the fast migration algorithm and incremental resource allocation strategy to ensure that resource adjustment is completed without affecting system performance;
[0018] S9, continuously monitor the adjusted load status, and feed back the real-time load status and strategy execution effect to the reinforcement learning module through the feedback mechanism to update and improve the strategy online;
[0019] S10. Repeat steps S1 to S9 to gradually improve the real-time CPU load balancing effect of the cloud computing platform, thereby maximizing resource utilization and optimizing system performance.
[0020] Optionally, the S1 includes the following steps:
[0021] S11. Multiple data collection modules are set up on the cloud computing platform and deployed on each virtual machine and physical node to monitor CPU usage, task queue length and other key performance indicators in real time;
[0022] S12. Each data collection module records the CPU usage U of each virtual machine and physical node through regular sampling. i (t), task queue length Q i (t) and other key performance indicators K i (t);
[0023] S13. Mark the recorded data according to the timestamp t and construct a real-time CPU load data set D(t). The data set contains the following contents:
[0024] D(t)={(U i (t),Q i (t),K i (t))|i=1,2,…,N};
[0025] Among them, U i (t) represents the CPU usage of the i-th virtual machine or physical node at time t, Q i (t) represents the task queue length of the i-th virtual machine or physical node at time t, K i (t) represents other key performance indicators of the i-th virtual machine or physical node at time t, and N is the total number of virtual machines or physical nodes.
[0026] Optionally, S2 includes the following steps:
[0027] S21, transmitting the real-time CPU load data set D(t) recorded by each data collection module to the central load analysis module through a high-speed data transmission channel;
[0028] S22, in the central load analysis module, receiving the transmitted real-time CPU load data set D(t), performing preliminary analysis on the data, and arranging the data with different timestamps t in order;
[0029] S23, performing data cleaning on the parsed real-time CPU load data set D(t) to remove possible duplicate data, outliers and noise data;
[0030] S24, standardize the cleaned data and convert the CPU usage U of each virtual machine and physical node into i (t), task queue length Q i (t) and other key performance indicators K i (t) and other key performance indicators K i (t) conversion to a uniform scale;
[0031] S25, the standardized real-time CPU load data set Stored in a central database.
[0032] Optionally, S3 includes the following steps:
[0033] S31. In the load analysis module, the sliding window technology is used to analyze the preprocessed real-time CPU load data set. For time series analysis, set the length of the sliding window to W and the number of data points in the window to n;
[0034] S32. At time t, define a sliding window W t Contains data points from time t-W+1 to t:
[0035]
[0036] Among them, W t Represents a sliding window at time t, containing data points from time t-W+1 to t;
[0037] S33: CPU usage of each virtual machine and physical node in each sliding window Perform time series analysis and calculate its mean and variance
[0038]
[0039]
[0040] in, Indicates the normalized CPU usage. Indicates the average CPU usage in the sliding window. Indicates the variance of CPU usage within the sliding window;
[0041] S34: Task queue length for each virtual machine and physical node and other key performance indicators Perform similar time series analysis and calculate its mean and variance;
[0042] S35. Based on the mean and variance in the sliding window, identify the current resource utilization pattern and potential load peak, and define the resource utilization pattern as the mean vector of each indicator
[0043]
[0044] in, Represents the mean vector within the sliding window, including the mean of CPU usage The average length of the task queue and other key performance indicators
[0045] S36: When the CPU usage of a virtual machine or physical node is detected or task queue length Exceeding the preset threshold θ U or θ Q , marks the node as a potential load peak.
[0046] Optionally, S4 includes the following steps:
[0047] S41. Based on the timing analysis results, define the state space S of the load balancing strategy decision model, where the state s∈S represents the resource utilization of each virtual machine and physical node at a specific time t, including the CPU utilization rate Task queue length and other key performance indicators
[0048]
[0049] S42, define an action space A, where action a∈A represents a specific operation of the load balancing strategy, including task migration and resource allocation adjustment;
[0050] S43, constructing a deep Q network model Q(s,a;θ), using a multi-layer neural network structure, where the parameter θ is the weight of the neural network;
[0051] S44. Define a reward function R(s,a) to evaluate the immediate reward of each state-action pair. The reward function comprehensively considers resource utilization, response time, and system performance:
[0052] R(s,a)=α·U eff (s,a)-β·T resp (s,a)+γ·P sys (s,a);
[0053] Among them, U eff (s,a) represents resource utilization, T resp (s,a) represents the response time, P sys (s,a) represents the system performance, α, β, γ are weight coefficients;
[0054] S45, using the experience replay technology, store the past state transition samples (s, a, r, s') in the replay buffer, and randomly extract small batch samples (s j ,a j ,r j ,s' j ) to train the model and update the deep Q network model parameters θ:
[0055]
[0056] Among them, η is the learning rate, m is the number of small batch samples, and y j is the target Q value, and the calculation formula is:
[0057] y j =r j +γ1max a' Q(s' j ,a′;θ - );
[0058] Among them, γ1 is the discount factor, θ - are the parameters of the target network;
[0059] S46, every fixed number of steps, synchronize the parameters θ of the current Q network to the parameters θ of the target network - , to ensure the stability of training;
[0060] S47, after the deep Q network model is trained, it is used for real-time load balancing decisions, based on the current state s t Choose the best action a t , adjust CPU resource allocation and task scheduling in real time, and optimize the resource utilization and system performance of the cloud computing platform.
[0061] Optionally, S5 includes the following steps:
[0062] S51, collect historical CPU load data set H(t), which contains the CPU usage U of each virtual machine and physical node in different time periods i (t), task queue length Q i(t) and other key performance indicators K i (t);
[0063] S52, in the simulation environment, initializing the resource configuration and task distribution state of the cloud computing platform, and building a simulation scenario based on the historical CPU load data set H(t);
[0064] S53. In each simulation scenario, the deep Q network model is used Perform strategy optimization, define state s as the resource utilization in the current simulation scenario, and action a as the specific operation of the load balancing strategy;
[0065] S54. For each state-action pair (s, a), calculate its immediate reward And update the Q value:
[0066]
[0067] Among them, α' is the learning rate, γ1 ' is the discount factor, s' is the next state after executing action a;
[0068] S55. Use experience replay technology to store state transition samples in the simulation environment To the playback buffer, randomly extract small batches of samples from it for training, and further optimize the deep Q network model parameters θ':
[0069]
[0070] Among them, η' is the learning rate, m' is the number of mini-batch samples, y' j is the target Q value, and the calculation formula is:
[0071]
[0072] S56. In multiple iterations of training, the parameters θ of the current Q network are adjusted every fixed number of steps. ′ Synchronize to the target network's parameters θ' - , ensuring the stability of the training process;
[0073] S57. During the training process, evaluate the policy adaptability and efficiency of the deep Q network model, adjust the model parameters and reward function according to the evaluation results, and further optimize the load balancing strategy;
[0074] S58. After the training is completed, the optimized deep Q network model is used for real-time load balancing decisions in the actual environment to improve the resource utilization and system performance of the cloud computing platform.
[0075] Optionally, the S6 comprises the following steps:
[0076] S61. The trained and optimized deep Q network model Deploy to the decision module and initialize the current system status as the real-time resource utilization of each virtual machine and physical node;
[0077] S62, at each decision time t, input the current state into the deep Q network model and calculate the Q value Q(s t ,a;θ'), and select the action a with the largest Q value t As the current load balancing policy:
[0078]
[0079] S63. Based on the selected action a t , generate specific resource allocation and task scheduling plans, including tasks migration, CPU resource reallocation and other operations;
[0080] S64, executing resource allocation and task scheduling solutions, migrating selected tasks from overloaded virtual machines or physical nodes to nodes with sufficient resources, and reallocating CPU resources to adjust the load of each node;
[0081] S65. During the execution process, monitor the execution effect of resource allocation and task scheduling schemes, and record the system status after execution. t+1 and its corresponding resource utilization, task queue length, and other key performance indicators;
[0082] S66, the execution result and the new system status s t+1 Feedback is given to the deep Q network model for state update and strategy optimization in the next decision;
[0083] S67. Regularly evaluate and adjust the parameters θ of the deep Q network model based on the actual execution results ′ , ensuring that the model maintains efficient and accurate load balancing decisions in dynamically changing cloud computing environments;
[0084] S68. Repeat steps S62 to S67 to continuously optimize the resource utilization and system performance of the cloud computing platform to ensure that the system can achieve real-time CPU load balancing under different load conditions.
[0085] A real-time CPU load balancing system for cloud computing platform based on reinforcement learning, including the following modules:
[0086] Data collection module: Each virtual machine and physical node deployed on the cloud computing platform monitors the CPU usage, task queue length and other key performance indicators in real time, and records these data through regular sampling to generate a real-time CPU load data set. The data collection module monitors and records the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time;
[0087] High-speed data transmission channel: used to transmit the real-time CPU load data set recorded by the data collection module of each virtual machine and physical node to the central load analysis module, ensuring low latency and high reliability of data transmission. The high-speed data transmission channel transmits the recorded data to the central load analysis module in real time;
[0088] Load analysis module: Receives and parses real-time CPU load data sets transmitted through high-speed data transmission channels, uses sliding window technology to perform time series analysis on pre-processed data, calculates resource utilization patterns and potential load peaks of each virtual machine and physical node, and identifies resource utilization patterns and potential load peaks.
[0089] Reinforcement learning module: Based on the timing analysis results of the load analysis module, the deep Q network model is used to build a load balancing strategy decision model, including defining the state space, action space and reward function. During the training process, the historical CPU load data set and the simulation environment are used for multiple iterations to optimize the strategy and improve the adaptability and efficiency of the strategy. Based on the timing analysis results, the reinforcement learning module builds and optimizes the load balancing strategy through the deep Q network model;
[0090] Decision module: Deploy the trained and optimized deep Q network model to this module, calculate the load balancing strategy of each virtual machine and physical node in real time, generate specific resource allocation and task scheduling schemes according to the current load status, and perform task migration and resource reallocation operations. The decision module generates resource allocation and task scheduling schemes according to the optimized strategy, and performs task migration and resource reallocation operations;
[0091] Feedback mechanism module: After executing the resource allocation and task scheduling plan, it continuously monitors the adjusted load situation and feeds back the real-time load status and strategy execution effect to the reinforcement learning module for online update and improvement of the strategy. The feedback mechanism module monitors and records the execution effect and feeds back the results to the reinforcement learning module for strategy update and improvement.
[0092] The beneficial effects of the present invention are:
[0093] (1) The present invention deploys a data collection module to monitor the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time, build a real-time CPU load data set to ensure the timeliness and accuracy of the data, and use sliding window technology to perform time series analysis on the pre-processed real-time CPU load data set to identify the current resource utilization pattern and potential load peaks, so as to achieve rapid response and adjustment to load changes. The deep Q network model is used to generate and execute the optimal load balancing strategy in real time to ensure that the system can achieve real-time CPU load balancing under different load conditions.
[0094] (2) The present invention uses historical CPU load data sets and simulation environments to perform multiple iterative training to optimize the deep Q network model so that it can accurately predict and respond to future load changes, formulate an optimized load balancing strategy, define a scientific and reasonable reward function, and comprehensively consider multiple indicators such as resource utilization, response time and system performance, so that the generated load balancing strategy can achieve the best global performance.
[0095] (3) The present invention deploys the trained and optimized load balancing strategy to the decision-making module, calculates the load balancing strategy of each virtual machine and physical node in real time, adjusts the CPU resource allocation according to the current load status, performs task migration and load balancing operations, and ensures the efficiency of strategy execution. Through the feedback mechanism, the real-time load status and strategy execution effect are fed back to the reinforcement learning module, and the strategy is updated and improved online to ensure that the system is continuously optimized and adapts to dynamically changing load conditions.
[0096] (4) The present invention proposes a complete system architecture, including various parts such as a data collection module, a load analysis module, a decision module and an execution module, to ensure the stability and efficiency of the system. The architecture design takes into account various challenges and requirements in practical applications, has high practicality and scalability, and can adapt to cloud computing environments of different scales and complexities. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0098] Figure 1 This is a flow chart of a real-time CPU load balancing method and system for a cloud computing platform based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION
[0099] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0100] refer to Figure 1 , a real-time CPU load balancing method for a cloud computing platform based on reinforcement learning, comprising the following steps:
[0101] S1. Deploy a data collection module on the cloud computing platform to monitor the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time, and build a real-time CPU load data set;
[0102] S2, inputting the collected real-time CPU load data set into the load analysis module through a high-speed data transmission channel, parsing and preprocessing the real-time CPU load data set;
[0103] S3. In the load analysis module, the sliding window technique is used to perform time series analysis on the preprocessed real-time CPU load data set to identify the current resource utilization pattern and potential load peaks;
[0104] S4. Based on the results of timing analysis, a load balancing strategy decision model is constructed using the deep Q network model to define the state space, action space and reward function;
[0105] S5. During the training of the deep Q network model, the historical CPU load data set and the simulation environment are used for multiple iterations to optimize the strategy to improve its adaptability and efficiency.
[0106] S6. Deploy the trained and optimized load balancing strategy to the decision module, and generate specific resource allocation and task scheduling solutions through the strategy;
[0107] S7. In the decision module, the load balancing strategy of each virtual machine and physical node is calculated in real time, and CPU resource allocation is adjusted according to the current load status, and task migration and load balancing operations are performed;
[0108] S8. During the dynamic adjustment process, use the fast migration algorithm and incremental resource allocation strategy to ensure that resource adjustment is completed without affecting system performance;
[0109] S9, continuously monitor the adjusted load status, and feed back the real-time load status and strategy execution effect to the reinforcement learning module through the feedback mechanism to update and improve the strategy online;
[0110] S10. Repeat steps S1 to S9 to gradually improve the real-time CPU load balancing effect of the cloud computing platform, thereby maximizing resource utilization and optimizing system performance.
[0111] In this implementation, S1 includes the following steps:
[0112] S11. Multiple data collection modules are set up on the cloud computing platform and deployed on each virtual machine and physical node to monitor CPU usage, task queue length and other key performance indicators in real time;
[0113] S12. Each data collection module records the CPU usage U of each virtual machine and physical node through regular sampling. i (t), task queue length Q i (t) and other key performance indicators K i (t);
[0114] S13. Mark the recorded data according to the timestamp t and construct a real-time CPU load data set D(t). The data set contains the following contents:
[0115] D(t)={(U i (t),Q i (t),K i (t))|i=1,2,…,N};
[0116] Among them, U i (t) represents the CPU usage of the i-th virtual machine or physical node at time t, Q i (t) represents the task queue length of the i-th virtual machine or physical node at time t, K i (t) represents other key performance indicators of the i-th virtual machine or physical node at time t, and N is the total number of virtual machines or physical nodes.
[0117] In this implementation, S2 includes the following steps:
[0118] S21, transmitting the real-time CPU load data set D(t) recorded by each data collection module to the central load analysis module through a high-speed data transmission channel;
[0119] S22, in the central load analysis module, receiving the transmitted real-time CPU load data set D(t), performing preliminary analysis on the data, and arranging the data with different timestamps t in order;
[0120] S23, performing data cleaning on the parsed real-time CPU load data set D(t) to remove possible duplicate data, outliers and noise data;
[0121] S24, standardize the cleaned data and convert the CPU usage U of each virtual machine and physical node into i (t), task queue length Q i (t) and other key performance indicators K i (t) and other key performance indicators K i (t) conversion to a uniform scale;
[0122] S25, the standardized real-time CPU load data set Stored in a central database.
[0123] In this implementation, S3 includes the following steps:
[0124] S31. In the load analysis module, the sliding window technology is used to analyze the preprocessed real-time CPU load data set. For time series analysis, set the length of the sliding window to W and the number of data points in the window to n;
[0125] S32. At time t, define a sliding window W t Contains data points from time t-W+1 to t:
[0126]
[0127] Among them, W t Represents a sliding window at time t, containing data points from time t-W+1 to t;
[0128] S33: CPU usage of each virtual machine and physical node in each sliding window Perform time series analysis and calculate its mean and variance
[0129]
[0130] in, Indicates the normalized CPU usage. Indicates the average CPU usage in the sliding window. Indicates the variance of CPU usage within the sliding window;
[0131] S34: Task queue length for each virtual machine and physical node and other key performance indicators Perform similar time series analysis and calculate its mean and variance;
[0132] S35. Based on the mean and variance in the sliding window, identify the current resource utilization pattern and potential load peak, and define the resource utilization pattern as the mean vector of each indicator
[0133]
[0134] in, Represents the mean vector within the sliding window, including the mean of CPU usage The average length of the task queue and other key performance indicators
[0135] S36: When the CPU usage of a virtual machine or physical node is detected or task queue length Exceeding the preset threshold θ U or θ Q , marks the node as a potential load peak.
[0136] In this implementation, S4 includes the following steps:
[0137] S41. Based on the timing analysis results, define the state space S of the load balancing strategy decision model, where the state s∈S represents the resource utilization of each virtual machine and physical node at a specific time t, including the CPU utilization rate Task queue length and other key performance indicators
[0138]
[0139] S42, define an action space A, where action a∈A represents a specific operation of the load balancing strategy, including task migration and resource allocation adjustment;
[0140] S43, constructing a deep Q network model Q(s,a;θ), using a multi-layer neural network structure, where the parameter θ is the weight of the neural network;
[0141] S44. Define a reward function R(s,a) to evaluate the immediate reward of each state-action pair. The reward function comprehensively considers resource utilization, response time, and system performance:
[0142] R(s,a)=α·U eff (s,a)-β·T resp (s,a)+γ·P sys (s,a);
[0143] Among them, U eff (s,a) represents resource utilization, T resp (s,a) represents the response time, P sys (s,a) represents the system performance, α, β, γ are weight coefficients;
[0144] S45, using the experience replay technology, store the past state transition samples (s, a, r, s') in the replay buffer, and randomly extract small batch samples (s j ,a j ,r j ,s' j) to train the model and update the deep Q network model parameters θ:
[0145]
[0146] Among them, η is the learning rate, m is the number of small batch samples, and y j is the target Q value, and the calculation formula is:
[0147] y j =r j +γ1max a 'Q(s' j ,a′;θ - );
[0148] Among them, γ1 is the discount factor, θ - are the parameters of the target network;
[0149] S46, every fixed number of steps, synchronize the parameters θ of the current Q network to the parameters θ of the target network - , to ensure the stability of training;
[0150] S47, after the deep Q network model is trained, it is used for real-time load balancing decisions, based on the current state s t Choose the best action a t , adjust CPU resource allocation and task scheduling in real time, and optimize the resource utilization and system performance of the cloud computing platform.
[0151] In this implementation, S5 includes the following steps:
[0152] S51, collect historical CPU load data set H(t), which contains the CPU usage U of each virtual machine and physical node in different time periods i (t), task queue length Q i (t) and other key performance indicators K i (t);
[0153] S52, in the simulation environment, initializing the resource configuration and task distribution state of the cloud computing platform, and building a simulation scenario based on the historical CPU load data set H(t);
[0154] S53. In each simulation scenario, the deep Q network model is used Perform strategy optimization, define state s as the resource utilization in the current simulation scenario, and action a as the specific operation of the load balancing strategy;
[0155] S54. For each state-action pair (s, a), calculate its immediate reward And update the Q value:
[0156]
[0157] Among them, α' is the learning rate, γ1 ' is the discount factor, s' is the next state after executing action a;
[0158] S55. Use experience replay technology to store state transition samples in the simulation environment To the playback buffer, randomly extract small batches of samples from it for training, and further optimize the deep Q network model parameters θ':
[0159]
[0160] Among them, η' is the learning rate, m' is the number of mini-batch samples, and y' j is the target Q value, and the calculation formula is:
[0161]
[0162] S56. In multiple iterations of training, the parameters θ of the current Q network are adjusted every fixed number of steps. ′ Synchronize to the target network's parameters θ' - , ensuring the stability of the training process;
[0163] S57. During the training process, evaluate the policy adaptability and efficiency of the deep Q network model, adjust the model parameters and reward function according to the evaluation results, and further optimize the load balancing strategy;
[0164] S58. After the training is completed, the optimized deep Q network model is used for real-time load balancing decisions in the actual environment to improve the resource utilization and system performance of the cloud computing platform.
[0165] In this implementation, S6 includes the following steps:
[0166] S61. The trained and optimized deep Q network model Deploy to the decision module and initialize the current system status as the real-time resource utilization of each virtual machine and physical node;
[0167] S62, at each decision time t, input the current state into the deep Q network model and calculate the Q value Q(s t ,a;θ'), and select the action a with the largest Q value t As the current load balancing policy:
[0168]
[0169] S63. Based on the selected action a t , generate specific resource allocation and task scheduling plans, including tasks migration, CPU resource reallocation and other operations;
[0170] S64, executing resource allocation and task scheduling solutions, migrating selected tasks from overloaded virtual machines or physical nodes to nodes with sufficient resources, and reallocating CPU resources to adjust the load of each node;
[0171] S65. During the execution process, monitor the execution effect of resource allocation and task scheduling schemes, and record the system status after execution. t+1 and its corresponding resource utilization, task queue length, and other key performance indicators;
[0172] S66, the execution result and the new system status s t+1 Feedback is given to the deep Q network model for state update and strategy optimization in the next decision;
[0173] S67. Regularly evaluate and adjust the parameters θ of the deep Q network model based on the actual execution results ′ , ensuring that the model maintains efficient and accurate load balancing decisions in dynamically changing cloud computing environments;
[0174] S68. Repeat steps S62 to S67 to continuously optimize the resource utilization and system performance of the cloud computing platform to ensure that the system can achieve real-time CPU load balancing under different load conditions.
[0175] A real-time CPU load balancing system for cloud computing platform based on reinforcement learning, including the following modules:
[0176] Data collection module: Each virtual machine and physical node deployed on the cloud computing platform monitors the CPU usage, task queue length and other key performance indicators in real time, and records these data through regular sampling to generate a real-time CPU load data set. The data collection module monitors and records the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time;
[0177] High-speed data transmission channel: used to transmit the real-time CPU load data set recorded by the data collection module of each virtual machine and physical node to the central load analysis module, ensuring low latency and high reliability of data transmission. The high-speed data transmission channel transmits the recorded data to the central load analysis module in real time;
[0178] Load analysis module: Receives and parses real-time CPU load data sets transmitted through high-speed data transmission channels, uses sliding window technology to perform time series analysis on pre-processed data, calculates resource utilization patterns and potential load peaks of each virtual machine and physical node, and identifies resource utilization patterns and potential load peaks.
[0179] Reinforcement learning module: Based on the timing analysis results of the load analysis module, the deep Q network model is used to build a load balancing strategy decision model, including defining the state space, action space and reward function. During the training process, the historical CPU load data set and the simulation environment are used for multiple iterations to optimize the strategy and improve the adaptability and efficiency of the strategy. Based on the timing analysis results, the reinforcement learning module builds and optimizes the load balancing strategy through the deep Q network model;
[0180] Decision module: Deploy the trained and optimized deep Q network model to this module, calculate the load balancing strategy of each virtual machine and physical node in real time, generate specific resource allocation and task scheduling schemes according to the current load status, and perform task migration and resource reallocation operations. The decision module generates resource allocation and task scheduling schemes according to the optimized strategy, and performs task migration and resource reallocation operations;
[0181] Feedback mechanism module: After executing the resource allocation and task scheduling plan, it continuously monitors the adjusted load situation and feeds back the real-time load status and strategy execution effect to the reinforcement learning module for online update and improvement of the strategy. The feedback mechanism module monitors and records the execution effect and feeds back the results to the reinforcement learning module for strategy update and improvement.
[0182] By introducing the deep Q network model and reinforcement learning algorithm, this implementation enables the system to adaptively learn and optimize the load balancing strategy, thereby achieving real-time and efficient CPU load balancing. This method can not only improve resource utilization and system performance, but also quickly adjust the resource allocation strategy in the face of dynamically changing load conditions to ensure the stable operation of the system.
[0183] Embodiment 1:
[0184] In the cloud computing environment of a large e-commerce platform, there are high concurrent accesses and frequent load fluctuations. The server cluster of the platform includes 100 physical nodes and 500 virtual machines. As the shopping festival approaches, the number of visits increases significantly, resulting in frequent fluctuations in the CPU load. Traditional load balancing methods are difficult to cope with such dynamic changes, resulting in some nodes being overloaded while other node resources are idle. The present invention aims to optimize resource allocation and improve the response speed and stability of the system through a real-time CPU load balancing method based on reinforcement learning.
[0185] In this embodiment, a data collection module is first deployed on the cloud computing platform to monitor the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time, and build a real-time CPU load data set. The pre-processed real-time CPU load data set is analyzed in time series by sliding window technology to identify the current resource utilization mode and potential load peak.
[0186] Based on the results of timing analysis, a load balancing strategy decision model is constructed using a deep Q network model to define the state space, action space, and reward function. During the training of the deep Q network model, multiple iterations are performed using historical CPU load data sets and simulation environments to optimize the strategy and improve its adaptability and efficiency. The trained and optimized load balancing strategy is deployed to the decision module, and a specific resource allocation and task scheduling plan is generated through the strategy.
[0187] In actual operation, the method of the present invention uses a deep Q network model to calculate the load balancing strategy of each virtual machine and physical node in real time, and adjusts the CPU resource allocation according to the current load status, and performs task migration and load balancing operations. During the execution process, the system continuously monitors the adjusted load situation, and feeds back the real-time load status and strategy execution effect to the reinforcement learning module through the feedback mechanism to perform online update and improvement of the strategy.
[0188] The test was conducted during the peak shopping season of the e-commerce platform from May 1, 2024 to May 7, 2024. The test location was the core data center of the e-commerce platform, and the data included key indicators such as CPU usage, task response time, and system performance. The method of the present invention was compared with the traditional load balancing method (such as the polling method), and the specific data is shown in Table 1 below:
[0189] Table 1 Load balancing performance comparison data
[0190]
[0191]
[0192] It can be seen from the data in Table 1 that the method of the present invention is superior to the traditional method in terms of CPU utilization, response time and system performance score. Specifically, the method of the present invention improves the average CPU utilization by about 10% through an optimized load balancing strategy, and stabilizes the peak CPU utilization at about 90%, avoiding the problem of resource overload and idleness. At the same time, the average response time is reduced by about 30ms, and the peak response time is significantly reduced, ensuring the rapid response of the system under high load conditions. In addition, the system performance score generally increased by about 10 points, proving that the method of the present invention has significant advantages in improving system stability and performance.
[0193] In summary, this embodiment verifies the application effect of the method of the present invention in a large e-commerce platform through actual data, and effectively solves the problems existing in the prior art such as insufficient real-time and adaptability, limited load prediction and analysis capabilities, low efficiency of strategy optimization and execution, and lack of self-learning and continuous optimization capabilities.
[0194] The present invention deploys a data collection module to monitor the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time, builds a real-time CPU load data set to ensure the timeliness and accuracy of the data, uses a sliding window technology to perform time series analysis on the pre-processed real-time CPU load data set, identifies the current resource utilization pattern and potential load peaks, and achieves rapid response and adjustment to load changes. The deep Q network model is used to generate and execute the optimal load balancing strategy in real time to ensure that the system can achieve real-time CPU load balancing under different load conditions.
[0195] The present invention utilizes historical CPU load data sets and simulation environments for multiple iterative training to optimize the deep Q network model, so that it can accurately predict and respond to future load changes, formulate an optimized load balancing strategy, define a scientific and reasonable reward function, and comprehensively consider multiple indicators such as resource utilization, response time and system performance, so that the generated load balancing strategy can achieve the best on a global scale.
[0196] The present invention deploys the trained and optimized load balancing strategy to the decision-making module, calculates the load balancing strategy of each virtual machine and physical node in real time, adjusts the CPU resource allocation according to the current load status, performs task migration and load balancing operations, and ensures the efficiency of strategy execution. Through the feedback mechanism, the real-time load status and strategy execution effect are fed back to the reinforcement learning module, and the strategy is updated and improved online to ensure that the system is continuously optimized and adapts to dynamically changing load conditions.
[0197] The present invention proposes a complete system architecture, including various parts such as data collection module, load analysis module, decision module and execution module, to ensure the stability and efficiency of the system. The architecture design takes into account various challenges and requirements in practical applications, has high practicality and scalability, and can adapt to cloud computing environments of different scales and complexities.
[0198] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A real-time CPU load balancing method for cloud computing platform based on reinforcement learning, characterized in that: The steps include: S1. Deploy a data collection module on the cloud computing platform to monitor the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time, and build a real-time CPU load data set; S2, inputting the collected real-time CPU load data set into the load analysis module through a high-speed data transmission channel, parsing and preprocessing the real-time CPU load data set; S3. In the load analysis module, the sliding window technique is used to perform time series analysis on the preprocessed real-time CPU load data set to identify the current resource utilization pattern and potential load peaks; S4. Based on the results of timing analysis, a load balancing strategy decision model is constructed using the deep Q network model to define the state space, action space and reward function; S5. During the training of the deep Q network model, the historical CPU load dataset and simulation environment are used to perform multiple iterations to optimize the strategy. S6. Deploy the trained and optimized load balancing strategy to the decision module, and generate specific resource allocation and task scheduling solutions through the strategy; S7. In the decision module, the load balancing strategy of each virtual machine and physical node is calculated in real time, and CPU resource allocation is adjusted according to the current load status, and task migration and load balancing operations are performed; S8. During the dynamic adjustment process, use the fast migration algorithm and incremental resource allocation strategy to ensure that resource adjustment is completed without affecting system performance; S9, continuously monitor the adjusted load status, and feed back the real-time load status and strategy execution effect to the reinforcement learning module through the feedback mechanism to update and improve the strategy online; S10, repeating steps S1 to S9 to gradually improve the real-time CPU load balancing effect of the cloud computing platform; The S4 comprises the following steps: S41. Based on the timing analysis results, define the state space S of the load balancing strategy decision model, where the state s∈S represents the resource utilization of each virtual machine and physical node at a specific time t, including the CPU utilization rate Task queue length and other key performance indicators S42, define an action space A, where action a∈A represents a specific operation of the load balancing strategy, including task migration and resource allocation adjustment; S43, constructing a deep Q network model Q(s,a;θ), using a multi-layer neural network structure, where the parameter θ is the weight of the neural network; S44. Define a reward function R(s,a) to evaluate the immediate reward of each state-action pair. The reward function comprehensively considers resource utilization, response time, and system performance: R(s,a)=α·U eff (s,a)-β·T resp (s,a)+γ·P sys (s,a); Among them, U eff (s,a) represents resource utilization, T resp (s,a) represents the response time, P sys (s,a) represents the system performance, α, β, γ are weight coefficients; S45, using the experience replay technology, store the past state transition samples (s, a, r, s') in the replay buffer, and randomly extract small batch samples (s j ,a j ,r j ,s' j ) to train the model and update the deep Q network model parameters θ: Among them, η is the learning rate, m is the number of small batch samples, and y j is the target Q value, and the calculation formula is: y j =r j +γ1max a' Q(s' j ,a′;θ - ); Among them, γ1 is the discount factor, θ - are the parameters of the target network; S46, every fixed number of steps, synchronize the parameters θ of the current Q network to the parameters θ of the target network - , to ensure the stability of training; S47, after the deep Q network model is trained, it is used for real-time load balancing decisions, based on the current state s t Choose the best action a t , adjust CPU resource allocation and task scheduling in real time to optimize resource utilization and system performance of the cloud computing platform; The S5 comprises the following steps: S51, collect historical CPU load data set H(t), which contains the CPU usage U of each virtual machine and physical node in different time periods i (t), task queue length Q i (t) and other key performance indicators K i (t); S52, in the simulation environment, initializing the resource configuration and task distribution state of the cloud computing platform, and building a simulation scenario based on the historical CPU load data set H(t); S53. In each simulation scenario, the deep Q network model is used Perform strategy optimization, define state s as the resource utilization in the current simulation scenario, and action a as the specific operation of the load balancing strategy; S54. For each state-action pair (s, a), calculate its immediate reward And update the Q value: Among them, α' is the learning rate, γ1 ' is the discount factor, s' is the next state after executing action a; S55. Use experience replay technology to store state transition samples in the simulation environment To the playback buffer, randomly extract small batches of samples from it for training and optimize the deep Q network model parameters θ': Among them, η' is the learning rate, m' is the number of mini-batch samples, y' j is the target Q value, and the calculation formula is: S56. In multiple iterations of training, the parameters θ' of the current Q network are synchronized to the parameters θ' of the target network at fixed steps. - ; S57. During the training process, evaluate the policy adaptability and efficiency of the deep Q network model, adjust the model parameters and reward function according to the evaluation results, and optimize the load balancing strategy; S58, after the training is completed, the optimized deep Q network model is used for real-time load balancing decision in the actual environment; The S6 comprises the following steps: S61. The trained and optimized deep Q network model Deploy to the decision module and initialize the current system status as the real-time resource utilization of each virtual machine and physical node; S62, at each decision time t, input the current state into the deep Q network model and calculate the Q value Q(s t ,a;θ'), and select the action a with the largest Q value t As the current load balancing policy: S63. Based on the selected action a t , generate specific resource allocation and task scheduling plans, including tasks migration, CPU resource reallocation and other operations; S64, executing resource allocation and task scheduling solutions, migrating selected tasks from overloaded virtual machines or physical nodes to nodes with sufficient resources, and reallocating CPU resources to adjust the load of each node; S65. During the execution process, monitor the execution effect of resource allocation and task scheduling schemes, and record the system status after execution. t+1 and its corresponding resource utilization, task queue length, and other key performance indicators; S66, the execution result and the new system status s t+1 Feedback is given to the deep Q network model for state update and strategy optimization in the next decision; S67. Regularly evaluate and adjust the parameters θ of the deep Q network model based on the actual execution results ′ ,Ensure that the model maintains load balancing decisions in dynamically changing cloud computing environments; S68. Repeat steps S62 to S67 to continuously optimize the resource utilization and system performance of the cloud computing platform to ensure that the system can achieve real-time CPU load balancing under different load conditions.
2. According to claim 1, a real-time CPU load balancing method for a cloud computing platform based on reinforcement learning is characterized in that: The S1 comprises the following steps: S11. Multiple data collection modules are set up on the cloud computing platform and deployed on each virtual machine and physical node to monitor CPU usage, task queue length and other key performance indicators in real time; S12. Each data collection module records the CPU usage U of each virtual machine and physical node through regular sampling. i (t), task queue length Q i (t) and other key performance indicators K i (t); S13. Mark the recorded data according to the timestamp t and construct a real-time CPU load data set D(t). The data set contains the following contents: D(t)={(U i (t),Q i (t),K i (t))∣i=1,2,…,N}; Among them, U i (t) represents the CPU usage of the i-th virtual machine or physical node at time t, Q i (t) represents the task queue length of the i-th virtual machine or physical node at time t, K i (t) represents other key performance indicators of the i-th virtual machine or physical node at time t, and N is the total number of virtual machines or physical nodes.
3. The method for real-time CPU load balancing of a cloud computing platform based on reinforcement learning according to claim 2, characterized in that: The S2 comprises the following steps: S21, transmitting the real-time CPU load data set D(t) recorded by each data collection module to the central load analysis module through a high-speed data transmission channel; S22, in the central load analysis module, receiving the transmitted real-time CPU load data set D(t), performing preliminary analysis on the data, and arranging the data with different timestamps t in order; S23, performing data cleaning on the parsed real-time CPU load data set D(t) to remove possible duplicate data, outliers and noise data; S24, standardize the cleaned data and convert the CPU usage U of each virtual machine and physical node into i (t), task queue length Q i (t) and other key performance indicators K i (t) and other key performance indicators K i (t) conversion to a uniform scale; S25, the standardized real-time CPU load data set Stored in a central database.
4. The method for real-time CPU load balancing of a cloud computing platform based on reinforcement learning according to claim 3, characterized in that: The S3 comprises the following steps: S31. In the load analysis module, the sliding window technology is used to analyze the preprocessed real-time CPU load data set. For time series analysis, set the length of the sliding window to W and the number of data points in the window to n; S32. At time t, define a sliding window W t Contains data points from time t-W+1 to t: Among them, W t Represents a sliding window at time t, containing data points from time t-W+1 to t; S33: CPU usage of each virtual machine and physical node in each sliding window Perform time series analysis and calculate its mean and variance in, Indicates the normalized CPU usage. Indicates the average CPU usage in the sliding window. Indicates the variance of CPU usage within the sliding window; S34: Task queue length for each virtual machine and physical node and other key performance indicators Perform similar time series analysis and calculate its mean and variance; S35. Based on the mean and variance in the sliding window, identify the current resource utilization pattern and potential load peak, and define the resource utilization pattern as the mean vector of each indicator in, Represents the mean vector within the sliding window, including the mean of CPU usage The average length of the task queue and other key performance indicators S36: When the CPU usage of a virtual machine or physical node is detected or task queue length Exceeding the preset threshold θ U or θ Q , marks the node as a potential load peak.
5. A cloud computing platform real-time CPU load balancing system based on reinforcement learning, used to execute a cloud computing platform real-time CPU load balancing method based on reinforcement learning according to any one of claims 1 to 4, characterized in that: Includes the following modules: Data collection module: Each virtual machine and physical node deployed on the cloud computing platform monitors the CPU usage, task queue length and other key performance indicators in real time, and generates a real-time CPU load data set by regularly sampling and recording data. The data collection module monitors and records the CPU usage, task queue length and other key performance indicators of each virtual machine and physical node in real time; High-speed data transmission channel: used to transmit the real-time CPU load data set recorded by the data collection module of each virtual machine and physical node to the central load analysis module. The high-speed data transmission channel transmits the recorded data to the central load analysis module in real time; Load analysis module: Receives and parses real-time CPU load data sets transmitted through high-speed data transmission channels, uses sliding window technology to perform time series analysis on pre-processed data, calculates resource utilization patterns and potential load peaks of each virtual machine and physical node, and identifies resource utilization patterns and potential load peaks. Reinforcement learning module: Based on the timing analysis results of the load analysis module, the deep Q network model is used to build a load balancing strategy decision model, including defining the state space, action space and reward function. During the training process, the historical CPU load data set and the simulation environment are used for multiple iterations. Based on the timing analysis results, the reinforcement learning module uses the deep Q network model to build and optimize the load balancing strategy; Decision module: Deploy the trained and optimized deep Q network model to this module, calculate the load balancing strategy of each virtual machine and physical node in real time, generate specific resource allocation and task scheduling schemes according to the current load status, and perform task migration and resource reallocation operations. The decision module generates resource allocation and task scheduling schemes according to the optimized strategy, and performs task migration and resource reallocation operations; Feedback mechanism module: After executing the resource allocation and task scheduling plan, it continuously monitors the adjusted load situation and feeds back the real-time load status and strategy execution effect to the reinforcement learning module for online update and improvement of the strategy. The feedback mechanism module monitors and records the execution effect and feeds back the results to the reinforcement learning module for strategy update and improvement.