Dynamic resource scheduling optimization method based on LSTM and Q-learning

By adopting LSTM and Q-learning methods in computing resource scheduling, the staticity, insufficient accuracy and complexity of computing resource scheduling in the prior art are solved, and efficient and real-time resource scheduling and optimization are achieved.

CN120104313APending Publication Date: 2025-06-06ANHUI GUOKE ANGHUI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510144883.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing computing resource scheduling methods have problems such as static scheduling strategies, insufficient scheduling accuracy, resource scheduling complexity, high cost and overhead, and slow response speed, making it difficult to effectively schedule computing resources in high concurrency and high load scenarios.

Method used

The dynamic resource scheduling optimization method based on LSTM and Q-learning is adopted to optimize and predict historical load data through the LSTM load prediction model, and dynamically schedule resources to optimize resource allocation strategies based on Q-learning learning unit.

Benefits of technology

It improves the efficiency, speed and cost of computing resources, realizes the combination of accurate prediction and intelligent scheduling, avoids resource overload or idleness, and has real-time adaptability and good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104313A_ABST
    Figure CN120104313A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic resource scheduling optimization method based on LSTM (Long Short Term Memory) and Q-learning. The dynamic resource scheduling optimization method comprises the following steps: step 1, acquiring historical load data X of equipment; 2, inputting the historical load data X into the LSTM load prediction model, and optimizing the LSTM load prediction model to obtain an optimal LSTM load prediction model; 3, predicting predicted load data Xw at a certain time point or within a certain time period in the future through the optimal LSTM load prediction model; 4, inputting the predicted load data Xw into a resource scheduling optimization module; and 5, calculating a Q value of a Q-learning unit of the resource scheduling optimization module, and updating the Q value. The dynamic resource scheduling optimization method based on the LSTM and the Q-learning has the advantages that the computing resources can be effectively optimized, the use efficiency of the computing resources is improved, the speed of the computing resources is increased, the use cost of the computing resources is reduced, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a computing resource scheduling optimization technology, in particular to a dynamic resource scheduling optimization method based on LSTM and Q-learning. Background Art

[0002] With the continuous development of computer hardware and cloud computing platforms, in the process of computing resource management and task scheduling, more and more tasks need to be executed simultaneously in a complex computing environment, such as data processing tasks, computing-intensive tasks, and network requests. The load of these tasks fluctuates greatly in different time periods, so the system is required to have a strong resource scheduling capability in order to efficiently allocate and utilize limited computing resources.

[0003] At present, the existing computing resource scheduling methods all have the following shortcomings. (1) Static scheduling strategy: Most of the existing resource scheduling solutions are based on fixed scheduling rules and strategies and cannot be dynamically adjusted according to the changes in real-time task load. The traditional static scheduling method is not suitable for task scenarios with large load fluctuations. It is easy to cause uneven resource allocation, resulting in some resources being idle and others being overloaded, thus affecting system performance. (2) Insufficient scheduling accuracy: The existing scheduling systems mostly rely on simple scheduling rules based on historical statistical data and lack the ability to predict task loads. Therefore, they cannot accurately predict future load changes of tasks, which can easily lead to time lag or over-scheduling of resource allocation and reduce resource utilization efficiency. (3) Resource scheduling complexity: In the scheduling process of multiple tasks and multiple resources, how to reasonably allocate CPU, memory, disk and other resources within a limited time and avoid overload and resource waste is still a huge challenge. The scheduling strategies of existing solutions are difficult to respond to changes quickly in a high-concurrency environment, and are prone to unstable system performance. (4) High cost and high overhead: In some efficient scheduling systems, existing technologies may require expensive hardware support or a more complex system architecture. Especially for high-concurrency and high-load scenarios, traditional solutions are often not suitable for small and medium-sized systems due to their complexity and high cost. (5) Slow response: In the process of task load prediction and resource scheduling, due to the complexity of the processing process or the lag in data collection, the response speed of existing technologies is slow, and it is difficult to adapt to the rapid changes in system load in real time, affecting the timely scheduling and execution of tasks.

[0004] Therefore, it is urgent to design a computing resource scheduling method to solve the above problems. Summary of the invention

[0005] The present invention aims to avoid the deficiencies in the above-mentioned prior art and provide a dynamic resource scheduling optimization method based on LSTM and Q-learning to effectively optimize computing resources and improve the utilization efficiency, speed and cost of computing resources.

[0006] The present invention adopts the following technical solutions to solve the technical problems.

[0007] The present invention provides a dynamic resource scheduling optimization method based on LSTM and Q-learning, comprising the following steps:

[0008] Step 1: Obtain the historical load data X of the equipment;

[0009] Step 2: Input the historical load data X into the LSTM load forecasting model, optimize the LSTM load forecasting model, and obtain the optimal LSTM load forecasting model;

[0010] Step 3: Use the optimal LSTM load prediction model to predict the predicted load data Xw at a certain time point or time period in the future;

[0011] Step 4: Input the predicted load data Xw into the resource scheduling optimization module;

[0012] Step 5: Calculate the Q value of the Q-learning learning unit of the resource scheduling optimization module and update the Q value.

[0013] The structural characteristics of the dynamic resource scheduling optimization method based on LSTM and Q-learning of the present invention are also:

[0014] Furthermore, in step 1, the historical load data includes load parameters: CPU usage, memory usage and I / O occupancy.

[0015] Furthermore, in step 2, the LSTM load prediction model is expressed by the following formula (1):

[0016]

[0017] In formula (1), represents the load data forecast value at the future time point t+k; f LSTM is the prediction function; t represents a certain moment t, which is the starting point of the prediction time; t+k represents the time point k time periods T after moment t.

[0018] Furthermore, in step 2, the mean square error function is used as the loss function L to optimize the parameters of the LSTM load prediction model. The mean square error optimization formula is shown in the following formula (2);

[0019]

[0020] In formula (2), y i is the i-th true value of a load parameter, is the ith predicted value of a load parameter, and N is the total number of load parameters collected within a certain time period.

[0021] Furthermore, the resource scheduling optimization module includes a Q-learning learning unit.

[0022] Furthermore, the resource scheduling optimization module includes a state space S t , state space A t and the reward function R t .

[0023] Furthermore, the value function of the Q-learning learning unit is shown in the following formula (3):

[0024]

[0025] In formula (3), α is the learning rate, which indicates the learning speed of the Q-learning unit for new information; γ is the discount factor, which indicates the importance of future rewards.

[0026] Furthermore, the Q-value calculation process of the Q-learning learning unit includes the following steps:

[0027] Step 01: Get the current state space S t ;

[0028] Step 02: Select the scheduling action space A t ;

[0029] Step 03: Calculate the instant reward R t ;

[0030] Step 04: Calculate the next state space S t+1 ;

[0031] Step 05: Calculate the next state space S t+1 The maximum Q value of

[0032] Step 06: Update the Q value.

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0034] The present invention discloses a dynamic resource scheduling optimization method based on LSTM and Q-learning, comprising the following steps:

[0035] Step 1: Obtain the historical load data X of the device; Step 2: Input the historical load data X into the LSTM load prediction model, optimize the LSTM load prediction model, and obtain the optimal LSTM load prediction model; Step 3: Use the optimal LSTM load prediction model to predict the predicted load data Xw at a certain time point or time period in the future; Step 4: Input the predicted load data Xw into the resource scheduling optimization module; Step 5: Calculate the Q value of the Q-learning learning unit of the resource scheduling optimization module, and update the Q value.

[0036] The present invention discloses a dynamic resource scheduling optimization method based on LSTM and Q-learning, which has the following characteristics.

[0037] 1. Combining accurate prediction with intelligent scheduling: LSTM is used to achieve high-precision prediction and support early decision-making. Q-learning dynamically schedules based on prediction results to improve scheduling accuracy and efficiency.

[0038] 2. Efficient resource utilization; avoiding resource overload or idleness and improving the overall resource utilization of the system.

[0039] 3. Real-time adaptability: The combination of LSTM and Q-learning modules enables the system to respond to load changes in milliseconds.

[0040] 4. Scalability: This method is suitable for complex production environments with multiple tasks and multiple resources.

[0041] The dynamic resource scheduling optimization method based on LSTM and Q-learning of the present invention has the advantages of being able to effectively optimize computing resources, and improve the utilization efficiency, speed and cost of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flow chart of a dynamic resource scheduling optimization method based on LSTM and Q-learning of the present invention.

[0043] The present invention will be further described below through specific implementation modes in conjunction with the accompanying drawings. DETAILED DESCRIPTION

[0044] See also Figure 1 The present invention provides a dynamic resource scheduling optimization method based on LSTM and Q-learning, comprising the following steps:

[0045] Step 1: Obtain the historical load data X of the equipment;

[0046] Step 2: Input the historical load data X into the LSTM load forecasting model, optimize the LSTM load forecasting model, and obtain the optimal LSTM load forecasting model;

[0047] Step 3: Use the optimal LSTM load prediction model to predict the predicted load data Xw at a certain time point or time period in the future;

[0048] Step 4: Input the predicted load data Xw into the resource scheduling optimization module;

[0049] Step 5: Calculate the Q value of the Q-learning learning unit of the resource scheduling optimization module and update the Q value.

[0050] like Figure 1 As shown, a dynamic resource scheduling optimization method based on LSTM and Q-learning of the present invention adopts an LSTM load prediction model to predict the future predicted load data Xw of the equipment, and adopts a Q-learning learning unit to process the predicted load data Xw, so as to balance the load demand and resource allocation.

[0051] During specific implementation, in step 1, the historical load data includes load parameters: CPU usage, memory usage, and I / O occupancy.

[0052] The historical load data X can be represented by an n-dimensional vector (x 1 , x 2 , ...x n ), the historical load data X includes n load parameters. In one embodiment of the present invention, assuming that x 1 is the CPU usage, x 2 is the memory usage, x 3 is the I / O occupancy rate. The three load parameters are used as an example to illustrate that X = (x 1 , x 2 , x 3 ).

[0053] In the specific implementation, in step 2, the LSTM load prediction model is expressed by the following formula (1):

[0054]

[0055] In formula (1), represents the load data forecast value at the future time point t+k; f LSTM is the prediction function; t represents a certain moment t, which is the starting point of the prediction time; t+k represents the time point k time periods T after moment t.

[0056] In specific implementation, for convenience of calculation, the time period T can be set to 1 second, 3 seconds or 5 seconds. t+k represents a time point k seconds after time t.

[0057] In the specific implementation, in step 2, the mean square error function is used as the loss function L to optimize the parameters of the LSTM load prediction model. The mean square error optimization formula is shown in the following formula (2);

[0058]

[0059] In formula (2), y i is the i-th true value of a load parameter, is the ith predicted value of a load parameter, and N is the total number of load parameters collected within a certain time period.

[0060] For the load parameter x 1 、x 2 and x 3 , you can use a collection frequency of once every 1 second. Rolling forecast the load trend for the next 5 seconds every 10 seconds. For example, at the current time t = 0, LSTM predicts that the CPU usage will reach 95% at t + k = 5 seconds, and the memory usage will reach 75% at t + k = 5 seconds.

[0061] During specific implementation, the resource scheduling optimization module includes a Q-learning learning unit.

[0062] In specific implementation, the resource scheduling optimization module includes the state space S t , state space A t and the reward function R t .

[0063] State space S t Includes the current load Xt and the predicted load data Xw predicted by LSTM. The predicted load data Xw includes multiple predicted values Right now: For example, X t ={80%, 65%} means the current CPU usage is 80% and the memory usage is 65%. It indicates that at time point t+k, the predicted CPU usage is 95% and the memory usage is 75%.

[0064] Action Space A t Includes multiple scheduling policy adjustment operations, such as increasing or decreasing resources and task migration.

[0065] A t ={Increase CPU resources, reduce memory tasks, and adjust load balancing}.

[0066] Reward function R t Combine predicted load and resource utilization to quantify scheduling effects.

[0067] R t =-(︱actual load-target value︱+resource adjustment overhead).

[0068] In specific implementation, the value function of the Q-learning learning unit is shown in the following formula (3);

[0069]

[0070] In formula (3), α is the learning rate, which indicates the learning speed of the Q-learning unit for new information; γ is the discount factor, which indicates the importance of future rewards.

[0071] In the next state S t+1 The Q value of the optimal action selected under

[0072] In specific implementation, the Q-value calculation process of the Q-learning learning unit includes the following steps:

[0073] Step 01: Get the current state space S t ;

[0074] Assume that the current time step t = 0, the state space S of the system 0 It consists of the following information: Current CPU usage: 80%; Current memory usage: 65%; LSTM predicted future load: CPU prediction 95%, memory prediction 75%.

[0075] State space S 0 ={CPU current occupancy: 80%, memory current occupancy: 65%, LSTM predicted CPU load: 95%, LSTM predicted memory load: 75%}.

[0076] Step 02: Select the scheduling action space A t ;

[0077] Assume that the system selects the following scheduling actions: A 0 =Increase CPU resources;

[0078] Step 03: Calculate the instant reward R t ;

[0079] Calculate the difference between the actual load and the target load; Actual CPU usage: 93%; Actual memory usage: 70%; Target CPU load: 80%-90%; Target memory load: 70%-80%; Resource adjustment overhead: Increase CPU resource overhead by 2%.

[0080] Instant Rewards 0 The calculation formula is shown in the following formula (4).

[0081] R 0 = -(|93% - 85%| + |70% - 75%| + 2%) = -(8% + 5% + 2%) = -15% (3);

[0082] Step 04: Calculate the next state space S t+1 ;

[0083] Assume that in the state space S 0 After scheduling, the state space S of the system 0 It becomes the state space S 1 ;

[0084] State space S 1 ={CPU current occupancy: 90%, memory current occupancy: 72%, LSTM predicted CPU load: 92%, LSTM predicted memory load: 74%}.

[0085] Step 05: Calculate the next state space S t+1 The maximum Q value of

[0086] Assume that in state S 1 Next, the optimal action A 1 The Q value is 0.5, that is:

[0087] Step 06: Update the Q value.

[0088] According to the Q value update formula (3), update Q(S 0 ,A 0 ) can be expressed by the following formula (4).

[0089]

[0090] Select learning rate α = 0.1, discount factor γ = 0.9, and substitute them into formula (4), then we can get the following formula (5).

[0091]

[0092] From formula (5), we can see that the current state S 0 Next select Action A 0 The Q value of (increase CPU resources) is updated to 0.03.

[0093] In the present invention, LSTM is a special recursive neural network that can effectively process and predict time series data, and is particularly suitable for processing long-term dependency problems. In task load prediction, LSTM can learn the law of load changes from historical load data and provide accurate load predictions for subsequent resource scheduling. Q-learning is a value-based reinforcement learning algorithm that learns how to choose the optimal action by exploring and updating Q values. In resource scheduling, Q-learning can intelligently adjust resource allocation strategies based on predicted loads and current system status to optimize the overall performance of the system. Multi-task resource scheduling: Traditional task scheduling mainly focuses on the allocation of a single resource, while in multi-resource scheduling, tasks may involve multiple resources such as CPU, memory, network bandwidth, and disk. How to coordinate the allocation of these resources to avoid conflicts and bottlenecks is an important technical challenge.

[0094] The present invention discloses a dynamic resource scheduling optimization method based on LSTM and Q-learning, which has the following advantages.

[0095] 1. Improved load prediction accuracy and solved the problem of inaccurate prediction in existing technologies.

[0096] 2. Optimized the resource scheduling strategy to avoid resource overload or idleness and improve the resource utilization of the system.

[0097] 3. Real-time scheduling and adaptive adjustment are realized to ensure efficient operation of the system under dynamic load environment.

[0098] 4. It has low computing overhead and high response speed, which meets the real-time requirements in the production environment.

[0099] 5. It has good scalability and can adapt to production environments of different scales and task complexities.

[0100] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0101] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A dynamic resource scheduling optimization method based on LSTM and Q-learning, characterized in that: The steps include: Step 1: Obtain the historical load data X of the equipment; Step 2: Input the historical load data X into the LSTM load forecasting model, optimize the LSTM load forecasting model, and obtain the optimal LSTM load forecasting model; Step 3: Use the optimal LSTM load prediction model to predict the predicted load data Xw at a certain time point or time period in the future; Step 4: Input the predicted load data Xw into the resource scheduling optimization module; Step 5: Calculate the Q value of the Q-learning learning unit of the resource scheduling optimization module and update the Q value.

2. According to the method for dynamic resource scheduling optimization based on LSTM and Q-learning according to claim 1, it is characterized in that: In step 1, the historical load data includes load parameters: CPU usage, memory usage, and I / O occupancy.

3. The dynamic resource scheduling optimization method based on LSTM and Q-learning according to claim 1 is characterized in that: In step 2, the LSTM load prediction model is expressed by the following formula (1): In formula (1), represents the load data forecast value at the future time point t+k; f LSTM is the prediction function; t represents a certain moment t, which is the starting point of the prediction time; t+k represents the time point k time periods T after moment t.

4. The dynamic resource scheduling optimization method based on LSTM and Q-learning according to claim 1 is characterized in that: In step 2, the mean square error function is used as the loss function L to optimize the parameters of the LSTM load prediction model. The mean square error optimization formula is shown in the following formula (2); In formula (2), y i is the i-th true value of a load parameter, is the ith predicted value of a load parameter, and N is the total number of load parameters collected within a certain time period.

5. The dynamic resource scheduling optimization method based on LSTM and Q-learning according to claim 1 is characterized in that: The resource scheduling optimization module includes a Q-learning learning unit.

6. The method for dynamic resource scheduling optimization based on LSTM and Q-learning according to claim 5, characterized in that: The resource scheduling optimization module includes a state space S t , state space A t and the reward function R t .

7. The method for dynamic resource scheduling optimization based on LSTM and Q-learning according to claim 5, characterized in that: The value function of the Q-learning learning unit is shown in the following formula (3); In formula (3), α is the learning rate, which indicates the learning speed of the Q-learning unit for new information; γ is the discount factor, which indicates the importance of future rewards.

8. The method for dynamic resource scheduling optimization based on LSTM and Q-learning according to claim 5, characterized in that: The Q-value calculation process of the Q-learning learning unit includes the following steps: Step 01: Get the current state space S t ; Step 02: Select the scheduling action space A t ; Step 03: Calculate the instant reward R t ; R0=-(|93%-85%|+|70%-75%|+2%)=-(8%+5%+2%)=-15% (3); Step 04: Calculate the next state space S t+1 ; Step 05: Calculate the next state space S t+1 The maximum Q value of Step 06: Update the Q value.

Citation Information

Cited By

  • Distributed collaborative management system and method based on federal learning architecture

    CN120980085A