Double-layer resource scheduling method and system based on deep reinforcement learning
By dividing microservices into two categories and building deep neural network models, using dual-deep Q network algorithm training, optimizing resource allocation strategy, the balance problem of deep neural network accuracy and computing overhead in microservice resource demand prediction is solved, and efficient resource utilization is achieved.
Patent Information
- Application Number
- CN202510354135.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
Existing deep neural networks are difficult to balance the accuracy and computational overhead of prediction in the microservice resource demand forecast, resulting in resource allocation errors and high operating costs.
Using a two-layer resource scheduling method based on deep reinforcement learning, microservices are divided into two categories, deep neural network models are built, and training through dual-deep Q network algorithms, using throttling rates instead of resource utilization as an action, and resource allocation strategies are optimized.
It realizes that without violating the service-level delay requirements, reduce resource allocation errors, improve resource utilization efficiency, and balance prediction accuracy and calculation overhead.
Smart Images

Figure CN120276847A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of resource scheduling, and in particular, to a two-layer resource scheduling method and system based on deep reinforcement learning. Background Art
[0002] To ensure a seamless end-user experience, many user-facing latency-sensitive applications set service level objectives (SLOs) for end-to-end latency. To meet these requirements, traditionally, operators of cloud applications usually avoid SLO violations by over-provisioning resources. However, this strategy not only wastes resources but also may lead to high operating costs.
[0003] The microservices architecture has become the mainstream architecture in modern distributed systems. However, the distributed nature of microservices brings complexity to resource management. The dependencies between microservices cause user requests to form execution chains, which amplify the impact of resource allocation on end-to-end performance. Especially in a dynamic environment, incorrect resource allocation may lead to service response delays and even trigger a cascade effect, further increasing the difficulty of resource management.
[0004] Cloud computing environments have gradually shifted from a single architecture (such as private clouds or hybrid clouds) to more diverse distributed systems. In these environments, resource allocation requires more efficient decision-making to meet the needs of applications and optimize system performance. Different from traditional methods, in recent years, some studies have attempted to improve the intelligence level of resource allocation through deep learning techniques to adapt to changing workload requirements. Using deep neural networks to predict the resource requirements of microservices provides a new approach to solve this problem. Compared with traditional machine learning methods, deep neural networks can better capture the complex dependencies between microservices and make resource predictions according to real-time workloads in a dynamic environment. However, the system needs to design reasonable training and inference mechanisms and it is difficult to balance the prediction accuracy and computational overhead. Summary of the Invention
[0005] In order to at least partially solve the problem that the existing method of using deep neural networks to predict the resource requirements of microservices is difficult to balance the prediction accuracy and computational overhead, the present invention provides a two-layer resource scheduling method and system based on deep reinforcement learning. The present invention divides microservices into two categories and constructs a deep neural network model to ensure that the model outputs more accurate resource control strategies for different types of microservices. Then, the double deep Q-network algorithm is used to train the model. When training, the throttling rate is used instead of the resource utilization rate as the action to achieve a stronger correlation between resource allocation and latency. The finally obtained optimal deep neural network model can effectively reduce the latency caused by incorrect resource allocation, significantly improve the resource utilization efficiency, and balance the problem of prediction accuracy and computational overhead.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] The first aspect of the present invention proposes a two-layer resource scheduling method based on deep reinforcement learning, including:
[0008] Step 1: Divide the deployed microservices into two categories according to the resource utilization status, so as to implement more precise resource control policies for different types of microservices;
[0009] Step 2: Construct a deep neural network model according to the two types of microservices;
[0010] Step 3: Train the deep neural network model through the double deep Q-network algorithm to obtain the optimal deep neural network model, which is convenient for improving the prediction accuracy;
[0011] Step 4: Input the target request rate into the optimal deep neural network model to obtain the throttling rate of the corresponding target microservice, and send the throttling rate of the corresponding target microservice to the target microservice to complete the resource scheduling of the target microservice.
[0012] Further, the specific content of Step 1 includes:
[0013] Collect the resource utilization status of the microservices; wherein, the resource utilization status includes the utilization rate of the CPU;
[0014] Divide the microservices into high-utilization microservices and low-utilization microservices according to the resource utilization status, so as to implement more precise resource control policies for different types of microservices.
[0015] Further, the state in the double deep Q-network algorithm includes the number of requests arriving per second, the action includes the throttling rate, and the reward includes the negative cost; wherein, the throttling rate includes the number of depletion times of the resources allocated to the microservice per average cycle; if the output of the deep neural network model meets the service-level objective, the cost is only related to the resource allocation, and if it does not meet the service-level objective, the cost includes the sum of the delay and the resource allocation size.
[0016] Further, the loss function when training the deep neural network model through the double deep Q-network algorithm is expressed by the following formula:
[0017] L(θ) = E (s,a,r,s′) [(y - Q(s, a; θ)) 2
[0018] y = r + γQ(s′, argmax a′ Q(s′, a′; θ); θ - )
[0019] Among them, \(L(\theta)\) is the loss function, \(E\) is the expectation calculation, \(y\) is the target Q value, \(s\) is the state, \(a\) is the action, \(\theta\) is the current Q-network parameter, \(r\) is the immediate reward, \(\gamma(0 \lt \gamma \leq 1)\) is the discount factor, and \(\theta\) - is the parameter of the target network, and \(s'\) and \(a'\) are the state and action at the next time step respectively.
[0020] Furthermore, when training the deep neural network model through the double deep Q-network algorithm, a cost filtering mechanism is added, which specifically includes:
[0021] Combining the request rate and the corresponding throttling rate into a piecewise value;
[0022] Storing the observed cost in the experience pool set by the double deep Q-network algorithm, calculating the median cost according to the historical data of the same request rate and the corresponding throttling rate combination in the experience pool, and using the median cost to replace the observed cost for training the deep neural network model;
[0023] Adjusting the throttling rate output by the deep neural network model according to the \(\xi\)-greedy policy, and sending the adjusted throttling rate to the microservice.
[0024] Furthermore, the step of adjusting the throttling rate output by the deep neural network model according to the \(\xi\)-greedy policy and sending the adjusted throttling rate to the microservice specifically includes:
[0025] Generating a random value with multiple decimal places between the output throttling rate and the throttling rates less than it;
[0026] Selecting the generated random value with a probability of \(\xi\), or querying the historical data matching the current request rate and throttling rate from the experience pool with a probability of \(1 - \xi\), and randomly selecting one of the throttling rates as the adjusted throttling rate to send to the microservice.
[0027] Furthermore, the step of sending the throttling rate of the corresponding target microservice to the target microservice to complete the resource scheduling of the target microservice specifically includes:
[0028] A resource allocation program is set on the target microservice, and the resource allocation program receives the throttling rate at a preset period;
[0029] Comparing the received throttling rate with the actual throttling rate. If the actual throttling rate is greater than the received throttling rate, the resources allocated to the microservice are increased. If the actual throttling rate is less than the received throttling rate, resource scaling is performed; among them, the resource quota before scaling is recorded. If the actual throttling rate after scaling is greater than \(\alpha(\alpha \gt 1)\) times the target throttling rate, the resource allocation amount is rolled back to before scaling.
[0030] The second aspect of the present invention proposes a two - layer resource scheduling system based on deep reinforcement learning, including:
[0031] A classification module, configured to classify the deployed microservices into two categories according to the resource utilization status;
[0032] A model construction module, configured to construct a deep neural network model according to the two categories of microservices;
[0033] A training module, configured to train the deep neural network model through the double - deep Q - network algorithm to obtain an optimal deep neural network model;
[0034] A scheduling module, configured to input the target request rate into the optimal deep neural network model to obtain the throttling rate of the corresponding target microservice, and send the throttling rate of the corresponding target microservice to the target microservice to complete the resource scheduling of the target microservice.
[0035] The third aspect of the present invention proposes an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a two - layer resource scheduling method based on deep reinforcement learning as described in the first aspect above.
[0036] The fourth aspect of the present invention proposes a computer - readable storage medium. The storage medium includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute a two - layer resource scheduling method based on deep reinforcement learning as described in the first aspect above.
[0037] Advantages of the present invention:
[0038] When training with the DDQN algorithm, the present invention proposes to use the throttling rate instead of the traditional resource utilization rate as the action, realizing a stronger correlation between resource allocation and latency. Using the DDQN algorithm for training and convergence, it realizes outputting the optimal action in the action space for two categories of microservices according to the number of requests arriving per second. Using an experience pool to reduce the noise of cost data and realizing more action selections without expanding the action space. Running a resource allocation program for each microservice, periodically receiving the throttling rate, and adjusting the resources obtained by each microservice according to the throttling rate. Through the method of deep reinforcement learning, the present invention uses the throttling rate to adjust resources in real - time, thereby realizing minimizing resource allocation as much as possible without violating the service - level latency requirements, and achieving a balance between prediction accuracy and computational overhead. Description of the Drawings
[0039] Figure 1 It is a flowchart of a two - layer resource scheduling method based on deep reinforcement learning provided by an embodiment of the present invention.
[0040] Figure 2 Schematic diagram of a two - layer resource scheduling method based on deep reinforcement learning provided by an embodiment of the present invention.
[0041] Figure 3 Schematic diagram of the dependency relationship of microservices in the microservice architecture provided by an embodiment of the present invention.
[0042] Figure 4 Architecture diagram of a two - layer resource scheduling system based on deep reinforcement learning provided by an embodiment of the present invention. Detailed implementation manners
[0043] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment 1
[0045] As Figure 1 and Figure 2 shown, a two - layer resource scheduling method based on deep reinforcement learning includes:
[0046] S101: Classify the microservices that have been deployed into two categories according to the resource utilization status.
[0047] Specifically, by pre - training the microservices deployed in the cloud or cluster, evaluate their resource utilization status (CPU utilization rate). Divide the microservices into two categories: high utilization rate and low utilization rate. This classification not only reduces the scale of the action space during training (reducing the action space scale from x 9 to 2 9 , where x is the type of microservices, and 9 is the number of values set for the throttling rate of the present invention), but also improves the efficiency and accuracy of subsequent optimization.
[0048] Preferably, for ease of understanding, as Figure 3 shown, the present invention shows the dependency relationship of microservices in a microservice architecture. Figure 3 Each box in it corresponds to a microservice, and the arrow represents the direct relationship between microservices.
[0049] S102: Construct a deep neural network model according to the two categories of microservices.
[0050] Specifically, a deep neural network is used as the allocation target generation program, and the pre-trained model is used to learn to generate the optimal CPU throttling rate for each type of microservice at different request rates (RPS).
[0051] S103: Train the deep neural network model through the double deep Q-network algorithm to obtain the optimal deep neural network model.
[0052] Specifically, the model is trained through the double deep Q-network (DDQN) algorithm to gradually converge under the long-term optimization goal.
[0053] Regard the number of requests arriving per second as the state, the throttling rate as the action. The meaning of the throttling rate is the number of times the resources allocated to the microservice are exhausted in an average cycle (how many times the resources allocated to the microservice are exhausted in an average cycle). Set the throttling rate to nine values, which are {0.00, 0.02, 0.04, 0.06, 0.08, 0.10, 0.15, 0.20, 0.30} respectively.
[0054] The reward includes the negative of the cost, where the cost includes the sum of the delay and the resource allocation size: If the service level objective (SLO) is met, the cost is only related to the resource allocation. If the SLO is not met, the cost includes the sum of the delay and the resource allocation size. By increasing the delay penalty value to constrain the timeout behavior, resource allocation can be achieved within the requirements of the service level delay.
[0055] The training of the neural network uses the DDQN algorithm. Although it is also two neural networks like DQN, DDQN prevents the Q-value estimation from being too large through the separation of selecting actions and evaluating actions. Define the loss function with the error between the target Q-value and the current Q-value. The specific formula is as follows:
[0056] L(θ) = E (s,a,r,s′) [(y - Q(s, a; θ)) 2
[0057] Achieve training optimization through the square difference between the predicted value and the actual value. Among them, L(θ) is the loss function, E is the expected calculation, y is the target Q-value, s is the state, a is the action, θ is the current Q-network parameter used to generate actions. The specific formula of y is as follows:
[0058] y = r + γQ(s′, argmax a′ Q(s′, a′; θ); θ - )
[0059] Among them, r is the immediate reward, γ(0 < γ ≤ 1) is the discount factor, θ - are the parameters of the target network, and s′ and a′ are the state and action at the next time step respectively. The two Q networks implement action selection and action evaluation respectively. By updating the current Q network parameter θ to the parameter θ of the target network. - Every certain number of training steps, synchronize the parameters of the training network to the target network.
[0060] First, select the action with the highest value it believes in the training network, and then output this action to the target network. The target network evaluates the value of this action. By optimizing the loss function L(θ), the convergence of the neural network is achieved. Facing the number of requests per second arriving, the deep neural network can generate a more accurate throttling rate probability distribution and output the optimal throttling rate in the action space for microservices with high utilization and low utilization respectively.
[0061] To reduce the noise in the cost data, use the median cost of the historical data in the experience pool to replace the real cost for model update. Specifically, combine the request rate and the corresponding throttling rate into piecewise values. Store the observed cost in the experience pool set by the double deep Q network algorithm. Calculate the median cost according to the historical data of the same request rate and the corresponding throttling rate combination in the experience pool, and use the median cost to replace the observed cost to train the deep neural network model.
[0062] During the action selection process, after selecting the throttling rate, make a downward random adjustment once. Subsequently, use the ξ-greedy strategy to select this random value or query the experience pool to obtain the throttling rate corresponding to the median cost. Without expanding the original action space, a more fine-grained action selection is achieved.
[0063] To reduce the influence of cost noise, perform a cost filtering to prevent inaccurate costs caused by unexpected situations. Add the real cost of a request rate per second and the selected throttling rate to the experience area and select the median cost corresponding to the request rate per second and the throttling rate as the cost for this time to continue training.
[0064] The cost consists of latency and resource allocation size. Latency may fluctuate due to external factors. To reduce the influence of such external factors, after obtaining the real cost, it is not directly used for the training of the neural network. Instead, query the experience area to obtain the median cost corresponding to the state and action, and use the median to replace the real cost as the reward feedback to the neural network.
[0065] After the action is selected, it is not sent directly to the resource allocation program. Instead, the action is randomized downward to obtain a multi-decimal random value between this throttling rate and the throttling rate that is less than it. The experience pool is queried once and the ξ-greedy strategy is used for selection. There is a probability of ξ to select the random value of the action at this time, and there is a probability of 1-ξ to select the random throttling rate value corresponding to the RPS and throttling rate after querying the experience pool. The selected throttling rate is periodically sent to the resource allocation program running on each microservice.
[0066] S104: Input the target request rate into the optimal deep neural network model to obtain the corresponding throttling rate of the target microservice, send the corresponding throttling rate of the target microservice to the target microservice, and complete the resource scheduling of the target microservice.
[0067] Specifically, the target request rate is input into the optimal deep neural network model to obtain the corresponding throttling rate of the target microservice.
[0068] A resource allocation program is set on the target microservice, and the resource allocation program receives the throttling rate according to a preset period. The resource allocation program is an existing program, specifically a scheduler that comes with the Linux system, which can detect resource consumption and use this interface to allocate CPU time.
[0069] The received throttling rate is compared with the actual throttling rate. If the actual throttling rate is greater than the received throttling rate, the resources allocated to the microservice are increased. If the actual throttling rate is less than the received throttling rate, resource scaling is performed. The resource quota before scaling is recorded. If the actual throttling rate after scaling is greater than α (α>1) times the target throttling rate, the resource allocation quota is returned to the amount before scaling.
[0070] The present invention divides microservices into two categories according to resource utilization and learns the probability distribution of actions. By proposing the throttling rate instead of the traditional resource utilization rate as an action, a stronger correlation between resource allocation and latency is achieved. The DDQN algorithm is used for training and convergence to output the optimal action in the action space for two types of microservices according to the number of requests arriving per second. The experience pool is used to reduce the noise of cost data and achieve more action selection without expanding the action space. The resource allocation program is run in each microservice, the throttling rate is periodically received, and the resources obtained by each microservice are adjusted according to the throttling rate. The present invention uses the throttling rate to adjust resources in real time through deep reinforcement learning, thereby reducing resource allocation as much as possible without violating the service-level latency requirements.
[0071] Example 2
[0072] Based on the above embodiments, Figure 4As shown in the figure, an embodiment of the present invention provides a two - layer resource scheduling system based on deep reinforcement learning, including:
[0073] A classification module, configured to classify the deployed microservices into two categories according to the resource utilization status.
[0074] A model construction module, configured to construct a deep neural network model according to the two categories of microservices.
[0075] A training module, configured to train the deep neural network model through the double - deep Q - network algorithm to obtain an optimal deep neural network model.
[0076] A scheduling module, configured to input the target request rate into the optimal deep neural network model to obtain the throttling rate of the corresponding target microservice, and send the throttling rate of the corresponding target microservice to the target microservice to complete the resource scheduling of the target microservice.
[0077] It should be noted that the two - layer resource scheduling system based on deep reinforcement learning provided by the embodiment of the present invention is to implement the above - mentioned two - layer resource scheduling method based on deep reinforcement learning. Its functions can be specifically referred to the above - mentioned method embodiments, and will not be elaborated here.
[0078] Embodiment 3
[0079] Based on the above - mentioned embodiment, an embodiment of the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the above - mentioned two - layer resource scheduling method based on deep reinforcement learning.
[0080] The present invention also provides a computer - readable storage medium. The storage medium includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the above - mentioned two - layer resource scheduling method based on deep reinforcement learning.
[0081] In summary, when training with the DDQN algorithm, the present invention replaces the traditional resource utilization rate with the throttling rate as an action, achieving a stronger correlation between resource allocation and latency. The DDQN algorithm is used for training and convergence to output the optimal action in the action space for two types of microservices according to the number of requests arriving per second. The experience pool is used to reduce the noise of cost data and enable more action selections without expanding the action space. A resource allocation program runs for each microservice, periodically receiving the throttling rate and adjusting the resources obtained by each microservice according to the throttling rate. Through deep reinforcement learning, the present invention uses the throttling rate to adjust resources in real time, thereby minimizing resource allocation as much as possible without violating the service-level latency requirements, achieving a balance between prediction accuracy and computational overhead.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A two - layer resource scheduling method based on deep reinforcement learning, characterized in that, Including: Step 1: Classify the deployed microservices into two categories according to the resource utilization status; Step 2: Construct a deep neural network model based on the two categories of microservices; Step 3: Train the deep neural network model through the double deep Q-network algorithm to obtain the optimal deep neural network model; Step 4: Input the target request rate into the optimal deep neural network model to obtain the throttling rate of the corresponding target microservice, and send the throttling rate of the corresponding target microservice to the target microservice to complete the resource scheduling of the target microservice.
2. The double-layer resource scheduling method based on deep reinforcement learning according to claim 1, wherein The specific content of Step 1 includes: Collect the resource utilization status of the microservices; among them, the resource utilization status includes the utilization rate of the CPU; Classify the microservices into high-utilization microservices and low-utilization microservices according to the resource utilization status.
3. A two-layer resource scheduling method based on deep reinforcement learning according to claim 1, characterized in that The state in the double deep Q-network algorithm includes the number of requests arriving per second, the action includes the throttling rate, and the reward includes the negative cost; among them, the throttling rate includes the depletion times of the resources allocated to the microservice on average per period; if the output of the deep neural network model meets the service-level objective, the cost is only related to the resource allocation, and if it does not meet the service-level objective, the cost includes the sum of the delay and the resource allocation size.
4. A two-layer resource scheduling method based on deep reinforcement learning according to claim 3, characterized in that, When training the deep neural network model through the double deep Q-network algorithm, the loss function is expressed by the following formula: L(θ) = E (s,a,r,s′) [(y - Q(s, a; θ)) 2 y = r + γQ(s′, argmax a′ Q(s′, a′; θ); θ - ) Among them, \(L(\theta)\) is the loss function, \(E\) is the expectation calculation, \(y\) is the target Q value, \(s\) is the state, \(a\) is the action, \(\theta\) is the current Q network parameter, \(r\) is the immediate reward, \(\gamma(0 \lt \gamma \leq 1)\) is the discount factor, and \(\theta\) - is the parameter of the target network, and \(s'\) and \(a'\) are the state and action at the next time step respectively.
5. A two-layer resource scheduling method based on deep reinforcement learning according to claim 1, characterized in that When training the deep neural network model through the double deep Q-network algorithm, a cost filtering mechanism is added, which specifically includes: Combining the request rate and the corresponding throttling rate into a piecewise value; Storing the observed cost into the experience pool set by the double deep Q-network algorithm, calculating the median cost according to the historical data of the same request rate and the corresponding throttling rate combination in the experience pool, and using the median cost to replace the observed cost to train the deep neural network model; Adjust the throttling rate output by the deep neural network model according to the ε-greedy strategy, and send the adjusted throttling rate to the microservice.
6. A two-layer resource scheduling method based on deep reinforcement learning according to claim 5, characterized in that, The specific content of adjusting the throttling rate output by the deep neural network model according to the ε-greedy strategy and sending the adjusted throttling rate to the microservice includes: Generate a random value with multiple decimal places between the output throttling rate and the throttling rates smaller than this throttling rate; With a probability of ε, select the generated random value, or with a probability of 1 - ε, query the historical data matching the current request rate and throttling rate from the experience pool, and randomly select one of the throttling rates as the adjusted throttling rate to send to the microservice.
7. A two-layer resource scheduling method based on deep reinforcement learning according to claim 1, characterized in that The specific content of sending the throttling rate of the corresponding target microservice to the target microservice to complete the resource scheduling of the target microservice includes: Set a resource allocation program on the target microservice, and the resource allocation program receives the throttling rate at a preset period; Compare the received throttling rate with the actual throttling rate. If the actual throttling rate is greater than the received throttling rate, increase the resources allocated to the microservice. If the actual throttling rate is less than the received throttling rate, perform resource scaling; among them, record the resource quota before scaling. If the actual throttling rate after scaling is greater than α (α > 1) times the target throttling rate, the resource allocation amount is rolled back to before scaling.
8. A two-layer resource scheduling system based on deep reinforcement learning, characterized in that, Including: A classification module, configured to classify the deployed microservices into two categories according to the resource utilization status; A model construction module, configured to construct a deep neural network model according to the two categories of microservices; A training module, configured to train the deep neural network model through a double deep Q-network algorithm to obtain an optimal deep neural network model; A scheduling module, configured to input a target request rate into the optimal deep neural network model to obtain a throttling rate of the corresponding target microservice, and send the throttling rate of the corresponding target microservice to the target microservice to complete the resource scheduling of the target microservice.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a two-layer resource scheduling method based on deep reinforcement learning according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the storage medium is located to execute a two-layer resource scheduling method based on deep reinforcement learning according to any one of claims 1 to 7.
Citation Information
Cited By
Model reasoning container deployment method and device
CN120875033A