A microservice management method based on load prediction and reinforcement learning

CN117453419BActive Publication Date: 2026-09-25BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311552801.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2026-09-25
Estimated Expiration
2043-11-20

AI Technical Summary

Technical Problem

[0011]本发明针对上述问题,提出了一种基于负载预测和强化学习的微服务管理方法,使用时序预测方法对微服务的未来负载情况进行预测,避免微服务扩缩动作的滞后性;通过强化学习算法训练、施加扩缩策略,不完全依赖于人工根据经验设计的阈值,高效地满足在保障微服务响应时间的前提下,节约资源的需求

Benefits of technology

[0048]1)本发明中添加了强化学习预测模块,可以帮助智能体更好地决策。如当预测到某个微服务访问量将大幅上涨,有很大概率影响用户请求的响应时间的情况下,智能体可以主动式、提前地执行扩容动作,进而避免恶化发生时,用户请求响应时间短时间内恶化的抖动现象。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117453419B_ABST
    Figure CN117453419B_ABST
Patent Text Reader

Abstract

The application discloses a micro-service management method based on load prediction and reinforcement learning, and belongs to the field of micro-service resource management. First, the data collection system collects the calling frequency and average response time of each micro-service in each cycle platform as an index, and stores the index in a time series database in a structured form. Then, the load prediction module uses the stored time series data to obtain the calling rate data of each micro-service, and predicts the calling rate of each micro-service in the next time cycle. The prediction result is input into a reinforcement learning agent, the reinforcement learning agent is trained, and the expansion and contraction decision is output through the trained policy network. Finally, the horizontal expansion and contraction controller adjusts the number of corresponding micro-service replicas according to the expansion and contraction decision output by the reinforcement learning agent, and increases, reduces or maintains the number of replicas. The application can help the agent to make better decisions and avoid the jitter phenomenon of shortening the response time of user requests in a short time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of microservice resource management, specifically relating to a microservice management method based on load prediction and reinforcement learning. Background Technology

[0002] In a microservice support platform, for complex scenarios involving massive numbers of microservices, the management objectives mainly consider the following two directions:

[0003] The first direction is optimizing time metrics: In the microservices domain, time metrics primarily refer to the response time to user requests, which also represents the performance of the microservice application. When a microservice application is overloaded, user requests may not be processed efficiently, leading to prolonged response times. Typically, users and cloud platform providers agree on an SLA (Service Level Agreement), such as a 99% service response time of less than 1ms. The goal of optimizing response time is to minimize violations of the SLA.

[0004] The second direction is to optimize resource utilization: In the microservices field, resource utilization mainly refers to the CPU and memory resources in the entire microservices system platform.

[0005] How to save resources as much as possible while ensuring microservice response time is a long-term research question.

[0006] Horizontal scaling of microservices is a reliable research direction: when microservice resources are scarce, leading to deteriorating response time, increasing the number of microservice replicas can improve their available resources and effectively guarantee the microservice's time metrics; when user requests are scarce, reducing the number of microservice replicas can save cluster resources.

[0007] On the Kubernetes platform, most existing horizontal scaling methods are rule-based scaling, such as adjusting the number of replicas by detecting the current resource usage of microservices. This type of method has the following problems:

[0008] 1. There is a certain lag; taking resource utilization-based rules as an example: when the load increases and reaches the set threshold, the expansion action is triggered. At this time, it takes tens of seconds or even minutes to take effect in the real environment, which is a certain lag; when the load fluctuates, it may make incorrect reduction actions, resulting in a deterioration of response time.

[0009] 2. Setting rules requires experience; for different microservices, operations personnel need to design scaling thresholds based on experience. If the threshold is too low, the number of microservice replicas will exceed the actual number needed, resulting in wasted resources; if the threshold is too high, the expansion of microservice replicas will not be timely when the load increases, leading to a deterioration of time metrics.

[0010] Therefore, using load prediction and reinforcement learning methods to build a microservice management system is a better choice to avoid lag and improve the robustness of scaling strategies. Summary of the Invention

[0011] To address the aforementioned problems, this invention proposes a microservice management method based on load prediction and reinforcement learning. It uses a time-series prediction method to predict the future load of microservices, avoiding the lag in microservice scaling. By training and applying scaling strategies through reinforcement learning algorithms, it does not rely entirely on manually designed thresholds based on experience, efficiently meeting the need to save resources while ensuring microservice response time.

[0012] The microservice management method based on load prediction and reinforcement learning has the following specific steps:

[0013] Step 1: For each time period, the number of calls and average response time of each microservice in the platform are collected as indicators by the data collection system and stored in the time series database in a structured form.

[0014] The structured form of an indicator includes a name, a label, a timestamp, and a corresponding indicator value. The label is used to classify the same indicator under different conditions, specifically represented as name{label=`label value`}value timestamp.

[0015] Step 2: The load prediction module obtains the call rate data of each microservice in the form of a time series by using the stored time series data, and predicts the call rate of each microservice in the next time period.

[0016] The load prediction steps are as follows:

[0017] Step 201: Perform feature engineering on the call data of microservices in the current system, and select the model used by the load prediction module according to the actual situation.

[0018] The feature engineering includes the complexity of the data being called and the amount of historical data. The selectable models include the ARIMA mathematical model and a deep learning model based on an attention mechanism.

[0019] Step 202: Divide the historical data into a test set and a training set, train the model, and determine its optimal parameters; at the same time, select the optimal value of the number of historical call rate data to use when performing each prediction, i.e., the historical data window length.

[0020] Step 203: Use the trained model for load prediction: During the decision cycle, obtain the call rate data of the historical data window through the relevant Prometheus interface, and combine it with the trained model to obtain the predicted value of the call rate for the next cycle.

[0021] Step 3: Input the prediction results into the reinforcement learning agent to train the reinforcement learning agent, and output the scaling decision through the trained policy network.

[0022] The reinforcement learning agent is trained using M episodes (iteration cycles), with each episode executing N steps. The specific steps are as follows:

[0023] Step 301: Initially, obtain the state space representation and action space representation of the agent, and randomly initialize the policy network θ0 and the state estimation network φ0;

[0024] The state space is represented by S, which includes the next cycle call rate prediction and the current number of microservice replicas;

[0025] Action space representation as Action χ0 means reducing the number of microservice replicas by one, action χ1 means not changing the number of microservice replicas, and action χ2 means increasing the number of microservice replicas by one.

[0026] Step 302: Starting from episode k=0, perform a cyclic iteration for each time period starting from t=0 in that episode;

[0027] Step 303: Starting from the period t=0, obtain the call rate prediction value for the next period from the load prediction module, and combine it with the current number of microservice replicas as the current state s. t Through strategy Obtain the agent's action a t ;

[0028] Step 304, when the agent performs action a on the microservice t Then, the average response time of the microservice is obtained, and the reward function r is designed based on the user-specified expected response time lim.

[0029] The reward function r includes the action reward function r. base and the legality reward function r legal The formula is:

[0030] r = r base +r legal

[0031] Wherein, the action reward function r base Represented as:

[0032]

[0033] Among them, rt t+1 This indicates the average response time of the microservice in the next cycle after the current action is applied to the system environment; K represents the basic value of the reward, which appears as a constant.

[0034] Action reward function r base The specific meaning is: when the agent performs action a t Afterwards, when the microservice response time rt t+1 If the threshold is exceeded, if action a is executed... t If ∈{χ0,χ1}, then a penalty of -4K is received; if action χ2 is performed, a reward of 4K is received; when the response time rt t+1 When the threshold is met, a reward of 2K is given when action χ0 is performed; no reward or penalty is given when action χ1 is performed; and a penalty of -3K is given when action χ2 is performed.

[0035] When the agent makes a decision using action χ0 or χ2, it determines whether the number of microservice replicas is legitimate, and obtains the legitimacy reward function r. legal Represented as:

[0036]

[0037] Step 305: Determine if the current time period satisfies t≤N. If yes, return to step 303 to continue iterating the current episode; otherwise, calculate the advantage function of the current policy network. And proceed to step 306;

[0038] For the current policy network θ k The driving strategy π, i.e. Its dominant function A π Represented as

[0039] A π (s,a)=Q π (s,a)-V π (s)

[0040] Among them, Q π (s,a) is the Q-function in the Bellman equation, expressing the expected reward obtained by acting according to policy π after performing action a in state s; Q π (s,a) is the V function in the Bellman equation, which expresses the expected reward obtained by acting according to policy π starting from state s.

[0041] Step 306: Based on the advantage function, optimize the following objective function using gradient ascent:

[0042]

[0043] in, The constant ∈ represents the cutoff interval of the policy optimization function. θ k This represents the current parameters of the driving policy network, and θ represents the direction of network optimization when the gradient of the objective function is ascending.

[0044] Step 307: Determine whether k≤M is satisfied. If yes, return to step 302 to continue iterating; otherwise, the final weights θ of the final policy network are obtained, and proceed to step 308.

[0045] Step 308: The policy network uses policy π θ At that time, the predicted call rate for the next cycle and the current number of microservice replicas are used as inputs to output scaling decisions.

[0046] Step 4: The horizontal scaling controller adjusts the number of replicas of the corresponding microservices according to the scaling decisions output by the reinforcement learning agent, increasing, decreasing, or maintaining the number of replicas.

[0047] The advantages of this invention are:

[0048] 1) This invention incorporates a reinforcement learning prediction module, which helps the agent make better decisions. For example, if it is predicted that the access volume of a certain microservice will increase significantly, which is likely to affect the response time of user requests, the agent can proactively and in advance perform scaling actions, thereby avoiding the jitter phenomenon of the user request response time deteriorating in a short period of time when the deterioration occurs.

[0049] 2) This invention uses reinforcement learning algorithms to control the number of microservice replicas. The transparency, dynamism and adaptability of the agent's strategy have inherent advantages in the complex and ever-changing cloud environment. Attached Figure Description

[0050] Figure 1 This is a framework diagram of the microservice management method based on temporal prediction and reinforcement learning in this invention;

[0051] Figure 2 This is a flowchart of the microservice management method based on load prediction and reinforcement learning of the present invention;

[0052] Figure 3 This is a flowchart of the training process for the reinforcement learning agent in this invention. Detailed Implementation

[0053] The present invention will be further described below with reference to the accompanying drawings and embodiments. The specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0054] This invention proposes a microservice management framework based on load prediction and reinforcement learning, such as... Figure 1 As shown, the microservice platform manages multiple microservices and receives external user requests; Istio and Jaeger are responsible for collecting the load of each microservice, i.e., the call status, and summarizing the metrics into the Prometheus time-series database; the load prediction module extracts data through Prometheus and predicts the call rate; after the reinforcement learning agent summarizes the response time and other data, it executes scaling decisions.

[0055] Based on the above framework, this invention proposes a microservice management method based on load prediction and reinforcement learning, such as... Figure 2 As shown, the specific steps are as follows:

[0056] Step 1: For each time period, collect the number of calls and average response time of each microservice in the platform using Istio and Jaeger as metrics, and store them in a structured form in the time-series database Prometheus.

[0057] The structured form of an indicator includes a name, a label, a timestamp, and a corresponding indicator value. The label is used to classify the same indicator under different conditions, specifically represented as name{label=`label value`}value timestamp.

[0058] Step 2: The load prediction module uses Prometheus to obtain call rate data for each microservice in time series format and predicts the call rate for the next time period.

[0059] The steps for load forecasting are as follows:

[0060] Step 201: Perform feature engineering on the call data of microservices in the current system, and select the model used by the load prediction module according to the actual situation.

[0061] For application scenarios with clear data call cycles and limited historical data, mathematical models such as ARIMA can be selected; for application scenarios with abundant historical data and complex data call patterns, deep learning models based on attention mechanisms can be selected.

[0062] Step 202: Train the model using historical data and determine the optimal parameters of the model using an independently partitioned test set. This includes selecting the optimal number of historical call rate data points to use for each prediction, i.e., the historical data window length.

[0063] Step 203: When using the model to perform predictions, within the decision-making cycle, obtain the call rate data of the historical data window through the relevant Prometheus interface, and combine it with the model to obtain the predicted value of the call rate for the next cycle.

[0064] Step 3: Input the prediction results into the reinforcement learning agent to train the reinforcement learning agent, and output the scaling decision of the number of microservice replicas through the trained policy.

[0065] The reinforcement learning agent is trained using M episodes (iteration cycles), with each episode executing N steps. Figure 3 As shown, the specific steps are as follows:

[0066] Step 301: Initially, obtain the state space representation and action space representation of the agent, and randomly initialize the policy network θ0 and the state estimation network φ0;

[0067] The state space is represented by S, which includes the next cycle call rate prediction value output by the load prediction module and the current number of microservice replicas;

[0068] Action space representation as Action χ0 means reducing the number of microservice replicas by one, action χ1 means not changing the number of microservice replicas, and action χ2 means increasing the number of microservice replicas by one.

[0069] Step 302: Starting from episode k=0, perform a cyclic iteration for each time period starting from t=0 in that episode;

[0070] Step 303: Starting from the period t=0, obtain the call rate prediction value for the next period from the load prediction module, and combine it with the current number of microservice replicas as the current state s. t Through strategy Obtain the agent's action a t ;

[0071] Step 304, when the agent performs action a on the microservice t Then, the average response time of the microservice is obtained, and the reward function r is designed based on the user-specified expected response time lim.

[0072] The reward function r includes the action reward function r. base and the legality reward function r legal The formula is:

[0073] r = r base +r legal

[0074] When designing a reward function, two main aspects need to be considered:

[0075] The first aspect is to prioritize ensuring that the average response time is less than the user's expected response time, and secondly, to conserve resources as much as possible. When the current number of replicas is insufficient, causing the response time to exceed the threshold, the agent is encouraged to increase the number of replicas, selecting action χ2. When the current number of replicas just meets the threshold, the reward function is designed to avoid oscillations caused by the agent continuously adjusting the number of replicas. When the number of replicas exceeds the threshold for meeting the response time, the agent is encouraged to reduce the number of replicas, selecting action χ0.

[0076] Action reward function r base Represented as:

[0077]

[0078] Among them, rt t+1 This indicates the average response time of the microservice in the next cycle after the current action is applied to the system environment; K represents the basic value of the reward, which appears as a constant.

[0079] Perform action a t After that, when the response time rt t+1 If the threshold is exceeded, if action a is executed... t If the action ∈{χ0,χ1}, then a reward of -4K is obtained, because this action should not be selected for execution in the current state; therefore, a negative reward is used to punish the agent. If action χ2 is executed, although the response time rt t+1 The threshold was exceeded, but the operation perfectly matched the expected operation in the current state, therefore a reward of 4K was awarded. When the response time rt... t+1 When the threshold is met, the reward design is slightly more complex: when action χ0 is executed, the number of replicas is reduced, saving cluster resources, and a reward of 2K is given; when action χ1 is executed, the number of replicas is not changed, and no reward or penalty is given; when action χ2 is executed, the number of replicas is increased, and a penalty of -3K is given.

[0080] The core part of the reward function, r base The design primarily expresses the following trends and expectations:

[0081] 1. If the current number of replicas is insufficient, causing the response time to exceed the threshold, guide the agent to select the only action χ2 that can provide a positive reward.

[0082] 2. When the current number of replicas increases to just enough to guarantee the threshold, avoid replica count fluctuations caused by multiple consecutive erroneous actions by the agent. When the replica count threshold is reached, if the agent performs the operation of first decreasing (χ0) and then increasing (χ2), it will receive an expected reward of -4K - 3K = -7K; if the agent performs the operation of first increasing (χ2) and then decreasing (χ0), it will receive an expected reward of -3K + 2K = -K. Therefore, in this state, the agent is more encouraged to choose operation χ1 to keep the number of replicas constant, thus ensuring the effectiveness of the agent's actions.

[0083] 3. When the current number of replicas exceeds the threshold for meeting the response time, guide the agent to select the only action χ0 that can provide a positive reward.

[0084] The second aspect is the legality of the actions. When an agent makes a decision using action χ0 or χ2, it may result in an invalid number of microservice replicas, which is penalized. In this case, the second part r of the reward function is defined. legal Represented as:

[0085]

[0086] Step 305: Determine if the current time period satisfies t≤N. If yes, return to step 303 to continue iterating the current episode; otherwise, calculate the advantage function of the current policy network. And proceed to step 306;

[0087] For the current policy network θ k The driving strategy π, i.e. Its advantage function is expressed as

[0088]

[0089] Among them, Q π (s,a) is the Q-function in the Bellman equation, expressing the expected reward obtained by acting according to policy π after performing action a in state s; Q π (s,a) is the V function in the Bellman equation, which expresses the expected reward obtained by acting according to policy π starting from state s.

[0090] Step 306: Based on the advantage function, optimize the following objective function using gradient ascent:

[0091]

[0092] in, The constant ∈ represents the cutoff interval of the policy optimization function. Using this constant, the stability of the policy optimization can be guaranteed. θ kThis represents the current parameters of the driving policy network, and θ represents the direction of network optimization when the gradient of the objective function is ascending.

[0093] Step 307: Determine whether k≤M is satisfied. If yes, return to step 302 to continue iterating; otherwise, the final weights θ of the final policy network are obtained, and proceed to step 308.

[0094] Step 308: The policy network uses policy π θ At that time, the predicted call rate for the next cycle and the current number of microservice replicas are used as inputs, and the output is a scaling decision to adjust the number of microservice replicas.

[0095] Step 4: The horizontal scaling controller adjusts the number of replicas of the corresponding microservices according to the scaling decisions output by the reinforcement learning agent, increasing, decreasing, or maintaining the number of replicas.

[0096] The above description is of a specific embodiment of the present invention. The above embodiment is exemplary and is not intended to limit the present invention.

Claims

1. A microservice management method based on load prediction and reinforcement learning, characterized in that, The specific steps are as follows: Step 1: For each time period, the number of calls and average response time of each microservice in the platform are collected as indicators through the data collection system and stored in the time series database in a structured form. Step 2: The load prediction module obtains the call rate data of each microservice in the form of a time series by using the stored time series data, and predicts the call rate of each microservice in the next time period. Step 3: Input the prediction results into the reinforcement learning agent to train the reinforcement learning agent, and output the scaling decision through the trained policy network. The reinforcement learning agent is trained using M episodes, with each episode performing N steps. The specific steps are as follows: Step 301: Initially, obtain the state space representation and action space representation of the agent, and randomly initialize the policy network θ0 and the state estimation network φ0; The state space is represented by S, which includes the next cycle call rate prediction and the current number of microservice replicas; Action space representation as Action χ0 means reducing the number of microservice replicas by one, action χ1 means not changing the number of microservice replicas, and action χ2 means increasing the number of microservice replicas by one. Step 302: Starting from episode k=0, perform a cyclic iteration for each time period starting from t=0 in that episode; Step 303: Starting from the period t=0, obtain the current state s of the agent based on the state-space representation of the agent at this time. t Through strategy Obtain the agent's action a t ; Step 304, when the agent performs action a on the microservice t Then, the average response time of the microservice is obtained, and the reward function r is designed based on the user-specified expected response time lim. The reward function r includes the action reward function r. base and the legality reward function r legal The formula is: r=r base +r legal Wherein, the action reward function r base Represented as: Among them, rt t+1 This represents the average response time of the microservice in the next cycle after the current action is applied to the system environment; K represents the basic value of the reward, which appears as a constant. When the agent makes a decision using action χ0 or χ2, it determines whether the number of microservice replicas is legitimate, and obtains the legitimacy reward function r. legal Represented as: Step 305: Determine whether the current time period satisfies t≤N. If yes, return to step 303 to continue the iteration of the current episode; otherwise, calculate the advantage function of the current policy network and execute step 306. For the current policy network θ k The driving strategy π, i.e. Its advantage function is expressed as Among them, Q π (s,a) is the Q-function in the Bellman equation, expressing the expected reward obtained by acting according to policy π after performing action a in state s; Q π (s,a) is the V function in the Bellman equation, which expresses the expected reward obtained by acting according to policy π starting from state s. Step 306: Based on the advantage function, optimize the following objective function using gradient ascent: in, The constant ∈ represents the cutoff interval of the policy optimization function; θ k This represents the current parameters of the driving policy network, and θ represents the direction of network optimization when the gradient of the objective function is ascended. Step 307: Determine whether k≤M is satisfied. If yes, return to step 302 to continue iterating; otherwise, the final weights θ of the final policy network are obtained, and proceed to step 308. Step 308: The policy network uses policy π θ At that time, the predicted call rate for the next cycle and the current number of microservice replicas are used as inputs to output scaling decisions; Step 4: The horizontal scaling controller adjusts the number of replicas of the corresponding microservices according to the scaling decisions output by the reinforcement learning agent, increasing, decreasing, or maintaining the number of replicas.

2. The microservice management method based on load prediction and reinforcement learning according to claim 1, characterized in that, The structured form of the indicator includes name, label, timestamp, and corresponding indicator value. The label is used to classify the same indicator under different conditions, specifically represented as name{label=`label value`}value timestamp.

3. The microservice management method based on load prediction and reinforcement learning according to claim 1, characterized in that, The load prediction module performs microservice call rate prediction in the following steps: Step 201: Perform feature engineering on the call data of microservices in the current system, and select the model used by the load prediction module according to the actual situation; Step 202: Divide the historical data into a test set and a training set, train the model, and determine its optimal parameters; at the same time, select the optimal value of the number of historical call rate data to use when performing each prediction, i.e., the historical data window length; Step 203: Use the trained model for load prediction: During the decision cycle, obtain the call rate data of the historical data window through the relevant Prometheus interface, and combine it with the trained model to obtain the predicted value of the call rate for the next cycle.

4. The microservice management method based on load prediction and reinforcement learning according to claim 3, characterized in that, The feature engineering includes the complexity of the data being called and the amount of historical data, and the models include the ARIMA mathematical model and a deep learning model based on the attention mechanism.

5. A microservice management method based on load prediction and reinforcement learning according to claim 1, characterized in that, The reward function guides the agent's actions, specifically as follows: (1) When the current number of replicas is insufficient, which will cause the response time to exceed the threshold, guide the agent to select the only action χ2 that can provide positive rewards; (2) When the current number of replicas increases to just enough to guarantee the threshold, guide the agent to select operation χ1 to keep the number of replicas unchanged, so as to ensure the effectiveness of the agent's action; (3) When the current number of copies exceeds the threshold for meeting the response time, guide the agent to select the only action χ0 that can provide a positive reward.

Citation Information

Patent Citations

  • Novel training method and device for intelligent perception diagnosis model of power distribution network

    CN115221776A

  • Pod horizontal scaling algorithm based on reinforcement learning

    CN115827161A