Micro-service chain level load balancing method based on deep reinforcement learning

By building a chain-level optimization model of response time-energy consumption and a multi-objective Markov decision-making process, combined with deep reinforcement learning algorithms, the load balancing problem of microservice chain is solved, and the coordinated optimization of response time and energy consumption is achieved, improving the dynamic adaptability and global optimization capabilities of the system.

CN120336025APending Publication Date: 2025-07-18SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510488743.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The load balancing method of existing microservice chains is difficult to take into account dynamic adaptability, global optimization and multi-objective collaboration, resulting in slow system response time, high energy consumption and unbalanced load distribution.

Method used

The microservice chain-level load balancing method based on deep reinforcement learning is used to build a chain-level optimization model of response time-energy consumption, and use the multi-objective Markov decision model and deep reinforcement learning algorithm to realize the global response time-energy consumption collaborative optimization of the microservice chain, and select the appropriate instance route for load balancing.

Benefits of technology

It realizes the balance of response time and energy consumption under different load states, ensures that optimization decisions meet global load balancing requirements, avoid local optimal traps, and improves the dynamic adaptability and global optimization capabilities of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336025A_ABST
    Figure CN120336025A_ABST
Patent Text Reader

Abstract

The invention provides a micro-service chain-level load balancing method based on deep reinforcement learning. The method comprises the following steps: establishing a response time-energy consumption model in a chain-level load balancing scene; determining a multi-objective optimization problem of chain-level load balancing; establishing a multi-target Markov decision model in an instance selection process in load balancing; constructing a chain-level load balancing model through a deep reinforcement learning method; training the chain-level load balancing model until rewards converge; and inputting load data into the trained model to obtain a balancing result. According to the method, load fluctuation is sensed in real time through the chain-level load balancing model, a proper load balancing strategy is adaptively selected, collaborative optimization of global response time and energy consumption of the micro-service chain is realized, and load balancing optimization requirements of high real-time performance and high request heterogeneity are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of load balancing under a microservices architecture, and more particularly, to a microservices chain-level load balancing method based on deep reinforcement learning. Background Art

[0002] With the continuous expansion of the scale of software systems, the tightly coupled code and modules in traditional monolithic architectures make the systems complex and difficult to maintain and expand. Therefore, the microservices architecture has emerged. The microservices architecture splits an application into multiple small autonomous services focusing on specific businesses, and forms a microservices chain through lightweight protocol communication to achieve complex business logic processing. However, in the context of the continuous growth of user requests and the continuous improvement of business complexity, the number of microservices and the dependencies between service chains have increased significantly, making load balancing a key issue affecting system performance. Among them, problems such as slow system response, high energy consumption, and uneven load distribution are particularly prominent, and there is an urgent need to effectively optimize with the entire microservices chain as the balancing granularity.

[0003] The response time of a microservices chain is the main quality of service concerned on the user side. During the process of optimizing the response time of a microservices chain, although the frequent decision-making of the load balancer and the real-time monitoring of instance status can improve the response speed, they also bring additional computational overhead to the system (such as increasing CPU load), resulting in an increase in system energy consumption. In order to reduce energy consumption, load balancing algorithms often affect the response time of the microservices chain and the load status of instances by reducing service instances or performing more scheduling and status checks on requests. Therefore, how to make a trade-off between response time and energy consumption without sacrificing service quality has become a key issue in chain-level load balancing under the microservices architecture.

[0004] Existing load balancing research for microservices chains mainly balances the allocation of instances or services. Some address the inter-chain resource competition problem and instance selection problem in the service chain, and use Nash bargaining and deep reinforcement learning to achieve load balancing at the instance level in the microservices chain. Some address the resource allocation and service interdependency problems in the microservices chain and conduct load balancing research with microservices as the unit. Although these methods optimize the performance of the microservices chain from different angles and provide diverse solutions, their balancing granularity mainly focuses on service instances or independent microservices units, lacking dynamic perception and collaborative regulation of the overall topology of the microservices chain, and it is difficult to take into account the three optimization dimensions of dynamic adaptability, global optimization, and multi-objective collaboration. Summary of the Invention

[0005] Aiming at the problem that the current chain-oriented load balancing method is difficult to balance the three optimization dimensions of dynamic adaptability, global optimization, and multi-objective coordination, the present invention proposes a microservice chain-level load balancing method based on deep reinforcement learning. By constructing a chain-level optimization model of response time - energy consumption and based on the deep reinforcement learning algorithm, the collaborative optimization of response time - energy consumption at the global level of the microservice chain is realized, and a suitable instance route is selected for heterogeneous load requests to achieve chain-level load balancing.

[0006] The technical solution adopted by the present invention to achieve the above object is as follows:

[0007] A microservice chain-level load balancing method based on deep reinforcement learning, comprising the following steps:

[0008] 1) Establish a response time - energy consumption model under the chain-level load balancing scenario;

[0009] 2) Based on the response time - energy consumption model, establish a multi-objective optimization problem for chain-level load balancing;

[0010] 3) Establish a multi-objective Markov decision model for the instance selection process in load balancing;

[0011] 4) Construct and train a chain-level load balancing model;

[0012] 5) Use the trained model to process the load data to obtain the balancing result.

[0013] The said step 1) includes the following steps:

[0014] 1.1) Construct a response time model:

[0015] The average response time T of microservice m in microservice chain c c,m is expressed as:

[0016]

[0017] Wherein, represents the service intensity of the queue, λ c,m , I c,m and μ c,m respectively represent the rate at which requests of chain c arrive at microservice m, the number of instances of microservice m in microservice chain c, and the service rate, F(X c,m ) and E[X c,m respectively represent the probability distribution function and the expectation of service time X c,m ;

[0018] When the system is in a steady state, the request arrival process of each microservice is a Poisson process, and the requests of chain c starting from the current microservice will arrive at the next microservice at the same rate λ c,m Therefore, the average response time T of microservice chain cc Expressed as:

[0019]

[0020] where s c,m indicates whether the microservice chain c passes through the microservice m, and M represents the number of microservices;

[0021] Due to the heterogeneity of user requests, a predefined response time upper limit deadline is set for each user request c , and to limit the average response time of each chain, the objective function of the response time is formulated as:

[0022]

[0023] 1.2) Construct an energy consumption model:

[0024] The average power of the server node n passed through by the chain c at the timestamp t is expressed using a linear formula Expressed as:

[0025]

[0026] where represents the power consumed in the idle state of the server, represents the maximum power consumed when the server is fully utilized, and u(t) represents the CPU utilization rate;

[0027] At the timestamp t, the normal energy consumption is normalized and expressed as:

[0028]

[0029] where represents the active state of the server node n at the timestamp t, 0 for sleep and 1 for wake-up;

[0030] The power in the overload state is expressed as:

[0031]

[0032] where CA represents the capacitance switched per clock cycle, V max represents the maximum voltage during overload, f max represents the maximum frequency of the CPU during overload, β represents the power consumption increase coefficient when the CPU overclocks, and u threshold represents the utilization threshold of the CPU when it enters the overload state;

[0033] The overload energy consumption is normalized and expressed as:

[0034]

[0035] Among them, N represents the number of server nodes, and C represents the number of microservice chains;

[0036] The total energy consumption of timestamp t is expressed as:

[0037]

[0038] The objective function ε of energy consumption c is formulated as:

[0039]

[0040] Among them, T is the time period for processing the current request.

[0041] In the said step 2), the multi-objective optimization problem of chain-level load balancing is:

[0042]

[0043] H c,m,i ≥H min

[0044] ρ c,m <1

[0045]

[0046] Among them, I c,m represents the number of instances of microservice m in microservice chain c, and I m represents the upper limit of the number of instances of microservice m, represents the CPU cycles allocated to the i-th instance at timestamp t, represents the upper limit of CPU cycles, H c,m,i represents the instance health level, and H min represents the lower limit of the health level, ρ c,m represents the service intensity, Θ represents the current load balancing degree of the system, represents the resource utilization rate of the system per unit time, P represents the set of chains serving user requests per unit time, ∑c represents the number of isomorphic chains, and Ψ represents the maximum acceptable threshold of the system.

[0047] The said step 3) is specifically:

[0048] The multi-objective Markov decision model is expressed as In the scenario of chain-level load balancing based on cloud computing under the microservices architecture, the cloud infrastructure assumes the role of the environment, and the load balancer plays the role of the Agent. According to the user request, the ETISA algorithm in the load balancer selects service instances and interacts with the environment over time. At each moment the load balancer observes the state space with the state s (t) in it, and selects an action a that takes into account both energy consumption and response time from the action space (t) . After the environment receives the action of the Agent, according to the reward function r(s (t) , a (t) ), it feeds back the reward r (t) to the Agent, and enters the next state s according to the state transition probability (t+1) . The Agent follows the policy function π(a (t) |s (t) ), ω), which represents the probability of selecting action a (t) in state s (t) and preference ω. The preference function f Ω maps the multi-dimensional reward vector to a single scalar. Using the preference ω ∈ preference space Ω, a scalar utility f ω (r(s, a)) = ω T r(s, a) is generated. The goal of the Agent is to find a policy π that maximizes the cumulative reward at all time stamps t. According to the Pareto non-dominated relationship, service instances that balance energy consumption and response time are selected, and the selected instances are sequentially selected to form a service link, and finally the optimal service link is obtained.

[0049] The state space, action space, and reward in the multi-objective Markov decision model are as follows:

[0050] State space: The state of instance selection at each decision period t is expressed as:

[0051]

[0052] where λ c represents the arrival rate of requests for chain c, E n represents the energy consumption of server node n, Θ (t) represents the system load fluctuation at period t, msc c represents the service link for processing requests of chain c, and m′ represents the next microservice to which the load balancer distributes requests;

[0053] Action space: User request After reaching the Agent, the Agent will correspondingly take the instance selection action a (t) = X(t), at each moment The Agent's perception of the energy consumption and response time of service instances will result in different selection strategies X(t). The discrete instance selection strategy is represented continuously, using e i,m to represent the probability that service instance i in microservice m is selected. Therefore, each instance selection decision is determined by the index of the maximum probability value in {e 1,m , e 2,m ,... e s,m}:

[0054] {e 1,m , e 2,m ,... e s,m},

[0055]

[0056] Reward: The purpose of instance selection in the chain-level load balancing environment is to obtain a service instance that strikes a balance between energy consumption and response time. Therefore, the main reward is the vector reward r(s, a) = [r T , r E :

[0057]

[0058] Among them, r T represents the time reward, r E represents the energy consumption reward, T(s (t) , a (t) ) and E(s (t) , a (t) ) represent the response time and energy consumption of the corresponding instance when the Agent executes action a (t) in state s (t) .

[0059] Step 4) includes the following steps:

[0060] 4.1) Construct a chain-level load balancing model through the deep reinforcement learning method;

[0061] 4.2) Train the chain-level load balancing model until the reward converges.

[0062] Specifically, step 4.1) is as follows:

[0063] Set the Agent's goal to maximize the expectation of the long-term return. At each moment t, the Agent learns the policy π(a|s, ω) and maps the state and preference to the action, constructing the state-value function V π (s, ω):

[0064]

[0065] Among them, represents the expected value under policy π, γ represents the reward decay factor, which is between [0, 1]. If it is 0, it is the greedy method, and the value is determined by the current immediate reward; if it is 1, the subsequent state reward is the same as the current reward;

[0066] When choosing action a in state s, the Q - function Q of the given policy π π (s, a, ω) is expressed as:

[0067]

[0068] Considering the trade - off between response time and energy consumption, assume there is a value space that contains the estimated total expected return of all bounded functions Q(s, a, ω) under the m - dimensional preference ω vector. At this time, Q(s, a, ω) is used as the multi - objective trade - off function of response time and energy consumption under different preferences;

[0069] A total of six networks are designed in the model, namely the Actor network, the Critic1 network, the Critic2 network, and their corresponding Target networks. By learning the entire preference space within the model, the Pareto - optimal policy set Π * is obtained. When calculating the objective value, a multi - objective envelope - optimal operator is introduced to add the action neighborhood and the smaller value between two Critic networks with the same network architecture to estimate the Q - value y of the next - state action pair:

[0070]

[0071] Among them, represents the optimal filter for the Q - value of the next state s′. By taking the convex hull of the current solution front, the utility - optimal Q is generated; arg Q takes the multi - objective value corresponding to the supremum; ∈ represents adding a small amount of normally - distributed random noise to the target action, and the noise is clipped, and the clipping range is [-c, c];

[0072] The optimal value Q in the value space Q * is expressed as:

[0073]

[0074] The parameters η of the Critic network are updated using the loss function L1(η q ) q :

[0075]

[0076] Introduce an auxiliary flat loss function \(L_2(\eta\) q ):

[0077]

[0078] Using the homotopy optimization method, increase the weight \(\lambda\) between \(L_1(\eta\) q ) and \(L_2(\eta\) q ) from 0 to 1, so that the loss function \(L(\eta\) q ) switches from \(L_1(\eta\) q ) to \(L_2(\eta\) q ) to find the global optimal parameters. The final loss function \(L(\eta\) q ) is expressed as:

[0079] \(L(\eta\) q )=(1 - \lambda)L_1(\eta\) q )+\lambda L_2(\eta\) q )

[0080] Update the parameters \(\theta\) of the Actor network by taking the negative of the Q value π . Therefore, its loss function \(J\) π is expressed as:

[0081]

[0082] The step 4.2) includes the following steps:

[0083] 4.2.1) Input the Actor network, the Criticl network, the Critic2 network, and their corresponding target networks. The network parameters are \(\theta\) π , \(\theta\) π ', user request \(r\) c \(\in R\), decay factor \(\gamma\), soft update coefficient \(\tau\), fixed update frequency \(n\) of the Critic network, maximum number of iterations random noise preference sampling distribution minimum number of samples \(N\) for preference ω , path \(p\) for the balance weight increasing from 0 to 1 λ ;

[0084] 4.2.2) Initialize the Actor network, the Critic1 network, the Critic2 network, randomize the parameters \(\theta\) π , and assign them to the corresponding parameters \(\theta\) in the target network π ' Initialize the experience pool \(\lambda = 0\)

[0085] 4.2.3) Sample a linear preference Initialize s as the first state of the current state sequence, and select an action a with exploration noise based on state s, a ∼ π(s) + ∈, Obtain the new state s′, reward r, and whether it is a terminal state is_end, and store the tuple <s, a, s′, r, is_end> into s = s′, and sample N samples from τ <s j , a j , s′ j , r j , is_end j , j = 1, 2,..., N ω , and assign to the parameter a′ i , Sample N preferences from ω Calculate the current Q value y , use the loss function L(η jk ), and update the Critic network parameters through neural network gradient backpropagation. Use the loss function J q , and update the Actor network parameters through neural network gradient backpropagation. Use the parameter τθ π +(1 - τ)θ π ′, π Update the corresponding θ in the target network ′, π ′, Increase the weight λ along the path p λ ;

[0086] 4.2.4) Repeat step 4.2.3) until the reward converges to obtain a trained chain-level load balancing model.

[0087] The said step 5) includes the following steps:

[0088] 5.1) Preprocess the load data, sort it in ascending order according to its predefined response time upper limit, and determine whether the health status of the service instances in the current microservice and the physical resources they contain meet the computing resources required for each request. If not, reject it;

[0089] 5.2) Call the deep reinforcement learning model of chain-level load balancing to obtain the microservice instance set;

[0090] 5.3) Sort the instance set in ascending order according to the load level of each instance, and take the instance with the minimum load as the balanced result routing.

[0091] The present invention has the following beneficial effects and advantages:

[0092] 1. Aiming at the overall performance of the microservice chain, based on the normal-overload energy consumption model, the present invention designs a response time model considering energy consumption through queuing theory. Further, through the Pareto optimal theory, the non-dominated relationship between response time and energy consumption is characterized, and a chain-level load balancing model considering both response time and energy consumption is obtained to ensure that the optimization decision not only meets the global load balancing requirements but also maintains the optimal delay and energy consumption balance under different load states. The chain-level load balancing problem is transformed into a multi-objective optimization problem.

[0093] 2. Based on the chain-level load balancing model, the present invention uses a multi-objective Markov decision process to represent the instance selection process in the model. A differentiable preference function is designed to map the multi-dimensional reward vector into an adaptive scalar reward, so that the load balancing strategy can dynamically coordinate different optimization objectives on a global scale, ensuring that the model can adapt to load changes and maintain global optimization.

[0094] 3. The present invention proposes a chain-level load balancing algorithm based on deep reinforcement learning. It can efficiently learn the Pareto optimal decision set in a single model and search for the chain-level optimal balancing strategy globally, ensuring that it will not fall into local optimality during the multi-objective optimization process. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 is a flowchart of the implementation method of the present invention;

[0096] Figure 2 is a model diagram of the deep reinforcement learning considered by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0097] The following further describes the present invention in detail with reference to the drawings and embodiments.

[0098] The microservice chain-level load balancing method based on deep reinforcement learning proposed by the present invention constructs a cascaded processing model of the microservice chain, constructs a multi-objective optimization problem of response time - energy consumption for dynamic scenarios, and establishes a chain-level load balancing model for adaptive instance selection of load balancing strategies to achieve the collaborative optimization of global delay and energy consumption of the microservice chain.

[0099] As Figure 1 shown, the present invention proposes a microservice chain-level load balancing method based on deep reinforcement learning. By constructing a chain-level optimization model of response time - energy consumption and based on the deep reinforcement learning algorithm, the collaborative optimization of response time - energy consumption globally for the microservice chain is achieved, and a suitable instance route is selected for heterogeneous load requests to achieve chain-level load balancing.

[0100] The present invention includes the following steps:

[0101] Step 1: Establish a response time - energy consumption model under the chain - level load balancing scenario;

[0102] Step 2: Define the multi - objective optimization problem of chain - level load balancing;

[0103] Step 3: Establish a multi - objective Markov decision model for the instance selection process in load balancing;

[0104] Step 4: Construct a chain - level load balancing model through the deep reinforcement learning method;

[0105] Step 5: Train the chain - level load balancing model until the reward converges;

[0106] Step 6: Input the load data into the trained model to obtain the balancing result.

[0107] This embodiment is implemented according to the Figure 1 flow shown below, and the specific steps are as follows:

[0108] 1. The response time - energy consumption model under the chain - level load balancing scenario is as follows.

[0109] The response time model is:

[0110] The average response time \(T\) of microservice \(m\) in microservice chain \(c\) c,m can be expressed as:

[0111]

[0112] where represents the service intensity of the queue, \(F(X c,m )\) and \(E[X c,m \) represent the probability distribution function and the expectation of service time \(X c,m \) respectively.

[0113] When the system is in a steady state, the request arrival process of each microservice is a Poisson process, and the requests of chain \(c\) starting from the current microservice will reach the next microservice at the same rate \(\lambda c,m . Therefore, the average response time of microservice chain \(c\) can be expressed as:

[0114]

[0115] Considering the heterogeneity of user requests, a predefined response time upper limit \(deadline c \) is set for each user request. To limit the average response time of each chain, the objective function of response time is formulated as:

[0116]

[0117] The energy consumption model is as follows:

[0118] The energy consumption of a server node n ∈ N in the cloud infrastructure during the time period T can be roughly divided into normal energy consumption and overload energy consumption. The normal energy consumption is calculated by the operating power consumption when the server processes requests normally without being overloaded. The overload energy consumption is the additional energy consumption of the server to handle excessive concurrent requests under overload conditions.

[0119] Use a linear formula to represent the average power of the server node n through which the chain c passes at the timestamp t

[0120]

[0121] Among them, represents the power consumed in the idle state of the server, represents the maximum power consumed when the server is fully utilized, and u(t) represents the CPU utilization rate, which changes over time.

[0122] At the timestamp t, the normal energy consumption can be normalized as:

[0123]

[0124] Among them, represents the active state of the server node n at the timestamp t, 0 for sleep and 1 for wake-up.

[0125] Power in the overload state can be expressed as:

[0126]

[0127] Among them, CA represents the capacitance switched per clock cycle, V max represents the maximum voltage during overload, f max represents the maximum frequency of the CPU during overload, β represents the power consumption increase coefficient when the CPU overclocks, represents the impact degree of overload on power, u threshold represents the utilization threshold when the CPU enters the overload state.

[0128] Overload energy consumption can be normalized as:

[0129]

[0130] Then, the total energy consumption at the timestamp t is obtained:

[0131]

[0132] Finally, the objective function of the energy consumption is formulated as:

[0133]

[0134] 2. The multi-objective optimization problem for establishing chain-level load balancing is as follows:

[0135]

[0136] H c,m,i ≥H min (3)

[0137] ρ c,m <1 (4)

[0138]

[0139] (1) is the instance quantity constraint;

[0140] (2) is to limit the CPU cycles in server node n;

[0141] (3) is the service availability constraint for microservice instances;

[0142] (4) is the service intensity constraint. Constraining ρ c,m below 1 is to prevent system crashes caused by system overload;

[0143] (5) is the load balancing constraint. The variance of the resource utilization rate of the microservice chains in the system is used to calculate the current load balancing degree Θ of the system. represents the resource utilization rate of the system per unit time, P represents the set of chains that serve user requests per unit time, ∑c represents the number of homogeneous chains, and Ψ represents the maximum threshold acceptable by the system, reflecting the maximum fluctuation of the load distribution in the system.

[0144] 3. The multi-objective Markov decision model for the instance selection process in load balancing is as follows:

[0145] The multi-objective Markov decision model is expressed as In the chain-level load balancing scenario based on cloud computing under the microservice architecture, the cloud infrastructure assumes the role of the environment, and the load balancer plays the role of the Agent. According to user requests, the ETISA algorithm in the load balancer selects service instances and interacts with the environment over time. At each moment the load balancer observes the state s in the state space (t) , and selects an action a that takes into account both energy consumption and response time from the action space (t) . After the environment receives the action of the Agent, it feedbacks the reward r (t) to the Agent according to the reward function r(s (t) ,a (t), and enter the next state s according to the state transition probability The Agent follows the policy function π(a (t+1) |s (t) , ω), which represents the probability of selecting action a (t) in state s (t) and preference ω. Define the preference function f (t) that maps the multi-dimensional reward vector to a single scalar, using the preference ω ∈ Ω, to produce a scalar utility f Ω (r(s, a)) = ω ω r(s, a). The goal of the Agent is to find a policy π that maximizes the cumulative reward at all timestamps t, selects service instances that balance energy consumption and response time, and finally obtains the optimal service link. T

[0146] State space. Express the state of instance selection at each decision period t as:

[0147]

[0148] where λ c represents the arrival rate of requests for chain c, E n represents the energy consumption of server node n, Θ (t) represents the system load fluctuation at time t, msc c represents the service link for processing requests of chain c, and m' represents the next microservice to which the load balancer distributes requests.

[0149] Action space. After the user request arrives at the Agent (load balancer), the Agent will correspondingly take the instance selection action a (t) = X(t). At each moment the Agent will have different selection strategies X(t) according to the perceived energy consumption and response time of the service instance. To more precisely capture the advantages and disadvantages between different instances, this paper represents the originally discrete instance selection strategy in a continuous form. e i,m represents the probability that service instance i in microservice m is selected. Therefore, each instance selection decision is determined by the index of the maximum probability value in {e 1,m , e 2,m ,... e s,m}.

[0150] {e 1,m , e 2,m ,... e s,m},

[0151]

[0152] ​Reward. The purpose of instance selection in the chain-level load balancing environment is to obtain service instances that achieve a trade-off between energy consumption and response time. Therefore, the main reward is the vector reward r(s,a) = [r T ,r E .

[0153]

[0154] 4. By using the deep reinforcement learning method, the chain-level load balancing model is constructed as follows:

[0155] Set the goal of the Agent to maximize the expectation of the long-term return. At each time step t, the Agent learns the policy π(a|s,ω), which maps the state and preference to an action, so that the action can achieve the maximum expectation. Therefore, the state-value function is:

[0156]

[0157] where γ represents the reward decay factor, which is between [0,1]. If it is 0, it is the greedy method, and the value is determined by the current delay reward; if it is 1, the subsequent state reward is the same as the current reward. Similarly, when choosing an action a in state s, the Q function of the given policy π is defined as:

[0158]

[0159] Considering the trade-off between response time and energy consumption, assume that there is a value space that contains the estimated total expected return of all bounded functions Q(s,a,ω) under the m-dimensional preference ω vector. At this time, Q(s,a,ω) can be regarded as a multi-objective trade-off function of response time and energy consumption under different preferences.

[0160] As Figure 2 shown, a total of six networks are designed, namely the Actor network, the Critic1 network, the Critic2 network, and the corresponding Target networks respectively. The algorithm learns the entire preference space within a single model, so as to obtain the Pareto optimal policy set Π * , when calculating the objective value, introduce a multi-objective envelope optimal operator Add the action neighborhood and the smaller value between two Critic networks with the same network architecture to estimate the Q value of the next state-action pair:

[0161]

[0162] where represents the optimal filter for the Q value of the next state s′, and generates the utility-optimal Q by taking the convex hull of the current solution front; arg QObtain the multi-objective value corresponding to the supremum; ∈ represents adding a small amount of random noise with a normal distribution to the target action, and the noise is clipped within the range [-c, c].

[0163] The optimal value Q in the value space Q * can be expressed as follows:

[0164]

[0165] Use the loss function L1(η q ) to update the parameters η of the Critic network q :

[0166]

[0167] Introduce an auxiliary flat loss function L2(η q ):

[0168]

[0169] Use the homotopy optimization method to gradually increase the weight λ between L1(η q ) and L2(η q ) from 0 to 1, so that the loss function L(η q ) switches from L1(η q ) to L2(η q ) to find the global optimal parameters. The final loss function is expressed as follows:

[0170] L(η q ) = (1 - λ)L1(η q ) + λL2(η q )

[0171] Since the loss of the Actor network is smaller when the Q value is larger and larger when the Q value is smaller, the parameters θ of the Actor network are updated by taking the negative of the Q value π , so its loss function J π is expressed as:

[0172]

[0173] 5. Train the chain-level load balancing model until the reward converges:

[0174] The above steps include the following steps:

[0175] (1) Input the Actor network, the Criticl network, the Critic2 network, and the corresponding target networks, with parameters θ π , θ π ′, user request rc ∈R, decay factor γ, soft update coefficient τ, fixed update frequency n of the Critic network, maximum number of iterations Random noise Preference sampling distribution Minimum number of samples N for preference ω , path p for the balance weight increasing from 0 to 1 λ

[0176] (2) Initialize the Actor network, Critic1 network, Critic2 network, and randomize the parameter θ π , Target network θ π ′←θ π , Initialize the experience pool λ = 0

[0177] (3) Sample a linear preference Initialize s as the first state of the current state sequence, select an action a ~ π(s) + ∈ with exploration noise based on state s, Obtain the new state s′, reward r, and whether it is a terminal state is_end, and store the tuple <s, a, s′, r, is_end> into s = s′, sample N samples from τ <s j , a j , s′ j , r j , is_end j , j = 1, 2,..., N ω , Sample N preferences from ω Calculate the current Q value y , use the loss function L(η jk ), update the Critic network parameters through neural network gradient backpropagation, use the loss function J q , update the Actor network parameters through neural network gradient backpropagation, and update the target network θ π ′←τθ π +(1 - τ)θ π ′ π ′ Increase the weight λ along the path p λ

[0178] (4) Repeat step (3) until the reward converges to obtain a trained deep reinforcement learning model.

[0179] ​6. Input the load data into the trained model to obtain the balancing result, including the following steps:

[0180] (1) Preprocess the load data, sort it in ascending order according to its predefined response time upper limit, and determine whether the health status of the service instances in the current microservice and the physical resources they contain meet the computing resources required for each request. If not, reject it;

[0181] (2) Call the deep reinforcement learning model for chain-level load balancing to obtain the microservice instance set;

[0182] (3) Sort the instance set in ascending order according to the load level of each instance, and take the instance with the minimum load as the balancing result route.

Claims

1. A microservice chain-level load balancing method based on deep reinforcement learning, characterized in that Including the following steps: 1) Establish a response time - energy consumption model under the chain - level load balancing scenario; 2) Based on the response time - energy consumption model, establish the multi - objective optimization problem of chain - level load balancing; 3) Establish a multi - objective Markov decision model for the instance selection process in load balancing; 4) Construct and train a chain - level load balancing model; 5) Use the trained model to process the load data and obtain the balancing result.

2. The microservice chain-level load balancing method based on deep reinforcement learning according to claim 1, wherein, The step 1) includes the following steps: 1.1) Construct a response time model: The average response time T of microservice m in microservice chain c c,m It is expressed as: Among them, represents the service intensity of the queue, λ c,m , I c,m and μ c,m respectively represent the rate at which requests of chain c arrive at microservice m, the number of instances of microservice m in microservice chain c, and the service rate. F(X c,m ) and E[X c,m respectively represent the probability distribution function and the expectation of the service time X c,m ; When the system is in a steady state, the request arrival process of each microservice is a Poisson process, and the requests of chain c starting from the current microservice will reach the next microservice at the same rate λ c,m Therefore, the average response time T of microservice chain c c is expressed as: where s c,m indicates whether the microservice chain c passes through the microservice m, and M represents the number of microservices; Due to the heterogeneity of user requests, a predefined upper bound on the response time, called deadline, is set for each user request c , and to limit the average response time of each chain, the objective function of the response time is formulated as: 1.2) Construct an energy consumption model: Use a linear formula to represent the average power of server node n that chain c passes through at timestamp t Expressed as: Among them, represents the power consumption of the server in the idle state, represents the maximum power consumption of the server when it is fully utilized, and u(t) represents the CPU utilization rate; At timestamp t, the normal energy consumption The normalized representation is: Among them, represents the activity status of server node n at timestamp t, where 0 indicates sleep and 1 indicates wake-up; Power under overload condition It is expressed as: Among them, CA represents the capacitance switched per clock cycle, and V max represents the maximum voltage during overload, f max represents the maximum frequency of the CPU during overload, β represents the power consumption increase coefficient when the CPU overclocks, and u threshold represents the utilization threshold for the CPU to enter the overload state; Overload energy consumption The normalized representation is as follows: Where N represents the number of server nodes, and C represents the number of microservice chains; Total energy consumption at timestamp t Expressed as: The objective function ε of energy consumption c is formulated as: Where T is the time period for processing the current request.

3. The microservice chain-level load balancing method based on deep reinforcement learning according to claim 1, characterized in that In the step 2), the multi - objective optimization problem of chain - level load balancing is: H c,m,i ≥ H min ρ c,m <1 Among them, I c,m represents the number of instances of microservice m in microservice chain c, I m represents the upper limit of the number of instances of microservice m, represents the CPU cycles allocated to the i-th instance at timestamp t, represents the upper limit of CPU cycles, H c,m,i represents the instance health, H min represents the lower limit of health, ρ c,m represents the service intensity, Θ represents the current load balancing degree of the system, represents the resource utilization rate of the system per unit time, P represents the set of chains that serve user requests per unit time, ∑c represents the number of isomorphic chains, and Ψ represents the maximum acceptable threshold of the system.

4. The microservice chain-level load balancing method based on deep reinforcement learning according to claim 1, characterized in that The step 3) is specifically: Formulate the multi-objective Markov decision model as In the scenario of chain-level load balancing based on cloud computing under the microservices architecture, the cloud infrastructure assumes the role of the environment, and the load balancer acts as the Agent. According to the user request, the ETISA algorithm in the load balancer selects service instances and interacts with the environment over time. At each moment the load balancer observes the state s in the state space (t) and selects an action a that simultaneously considers energy consumption and response time from the action space (t) . After the environment receives the action of the Agent, it feeds back the reward r (t) to the Agent according to the reward function r(s (t) , a (t) ) and enters the next state s according to the state transition probability (t+1) . The Agent follows the policy function π(a (t) |s (t) , ω), which represents the probability of selecting action a (t) in state s (t) and preference ω. Define the preference function f Ω to map the multi-dimensional reward vector to a single scalar. Using the preference ω ∈ preference space Ω, a scalar utility f ω (r(s, a)) = ω T r(s, a) is generated. The goal of the Agent is to find a policy π that maximizes the cumulative reward at all time stamps t. According to the Pareto non-dominance relationship, service instances that simultaneously balance energy consumption and response time are selected, and the selected instances are sequentially selected to form a service link, and finally the optimal service link is obtained.

5. The microservice chain-level load balancing method based on deep reinforcement learning according to claim 4, wherein The state space, action space, and reward in the multi - objective Markov decision model are respectively: State space: The state of the instance selection at each decision period t is expressed as: Among them, λ c represents the arrival rate of requests for chain c, and E n represents the energy consumption of server node n, and Θ (t) represents the system load fluctuation in period t, and msc c represents the service link for processing requests of chain c, and m' represents the next microservice to which the load balancer distributes requests; Action space: User request After reaching the Agent, the Agent will correspondingly take the instance selection action a (t) = X(t), at each moment The Agent's perception of the energy consumption and response time of service instances will result in different selection strategies X(t). Represent the discrete instance selection strategy continuously. Use e i,m to represent the probability that service instance i in microservice m is selected. Therefore, each instance selection decision is determined by the index of the maximum probability value in {e 1,m , e 2,m ,... e s,m}: {e 1,m ,e 2,m ,...e s,m}, Reward: The purpose of instance selection in the chain-level load balancing environment is to obtain service instances that achieve a trade-off between energy consumption and response time. Therefore, the main reward is the vector reward r(s,a) = [r T ,r E : Among them, r T represents the time reward, and r E represents the energy consumption reward. T(s (t) , a (t) ) and E(s (t) , a (t) ) represent the response time and energy consumption of the instance corresponding to the Agent executing the action a (t) in the state s (t) .

6. The microservice chain-level load balancing method based on deep reinforcement learning according to claim 1, characterized in that The step 4) includes the following steps: 4.1) Through the deep reinforcement learning method, construct a chain - level load balancing model; 4.2) Train the chain - level load balancing model until the reward converges.

7. The microservice chain-level load balancing method based on deep reinforcement learning according to claim 6, wherein The step 4.1) is specifically: Set the goal of the Agent to maximize the expectation of the long-term return. At each time step t, the Agent learns the policy π(a|s,ω) and maps the state and preference to an action, constructing the state-value function V π (s,ω): Among them, represents the expected value under the policy π, γ represents the reward decay factor, which is between [0, 1]. If it is 0, it is the greedy method, and the value is determined by the current immediate reward; if it is 1, the rewards of subsequent states are the same as the current reward. When selecting an action \(a\) in state \(s\), the Q - function \(Q_{\pi}\) of the given policy \(\pi\) π (s, a, \(\omega\)) is expressed as: Considering the trade-off between response time and energy consumption, assume that there is a value space containing the estimated expected total return of all bounded functions Q(s, a, ω) under the m-dimensional preference ω vector. At this time, Q(s, a, ω) is used as the multi-objective trade-off function of response time and energy consumption under different preferences; A total of six networks are designed in the model, namely the Actor network, the Critic1 network, the Critic2 network, and their corresponding Target networks. By learning the entire preference space within the model, the Pareto optimal policy set Π is obtained. - , when calculating the objective value, a multi-objective envelope optimal operator is introduced. The action neighborhood and the smaller value between two Critic networks with the same network architecture are added to estimate the Q value y of the next state-action pair: Among them, The optimal filter for the Q-value representing the next state s′ generates the utility-optimal Q by taking the convex hull of the current solution front; arg Q takes the multi-objective value corresponding to the supremum; ∈ means adding a small amount of normally distributed random noise to the target action, and the noise is clipped, and the clipping range is [-c, c]; Optimal value Q in value space Q * Is expressed as: Update the parameters η of the Critic network using the loss function L1(η q ) q : Introduce an auxiliary flat loss function \(L_2(\eta q ): Using the homotopy optimization method, increase the weight λ between L1(η q ) and L2(η q ) from 0 to 1, so that the loss function L(η q ) switches from L1(η q ) to L2(η q ) to find the global optimal parameters. The final loss function L(η q ) is expressed as: L(η q ) = (1 - λ)L1(η q ) + λL2(η q ) Update the parameters θ of the Actor network by taking the negative value of Q π , so its loss function J π is expressed as:

8. The microservice chain-level load balancing method based on deep reinforcement learning according to claim 6, characterized in that The step 4.2) includes the following steps: 4.2.1) Input the Actor network, the Critic network, the Critic2 network, and their corresponding target networks, with network parameters θ π , θ π ′, user request r c ∈R, decay factor γ, soft update coefficient τ, fixed update frequency n of the Critic network, maximum number of iterations random noise preference sampling distribution minimum number of samples N for preference ω , path p for the balance weight increasing from 0 to 1 λ ; 4.2.2) Initialize the Actor network, Critic1 network, and Critic2 network, and randomize the parameter θ π , and assign them to the corresponding parameters θ' in the target network π ′, Initialize the experience pool λ = 0 4.2.3) Sample a linear preference Initialize s as the first state of the current state sequence, and select an action a with exploration noise based on state s, a ∼ π(s)+∈, Obtain the new state s′, reward r, and whether it is a terminal state is_end, and store the tuple <s, a, s′, r, is_end> into s = s′, from Sample N τ samples, <s j , a j , s′ j , r j , is_end j , j = 1, 2,..., N ω , and assign to the parameter a′ i , Sample N preferences from ω Calculate the current Q value y , use the loss function L9η jk ), update the Critic network parameters through neural network gradient backpropagation, use the loss function J q ), update the Actor network parameters through neural network gradient backpropagation, use the parameter τθ π +(1 - τ)θ π ′, π Update the corresponding θ in the target network ′, π ′, Increase the weight λ along the path p λ ; 4.2.4) Repeat step 4.2.3) until the reward converges to obtain the trained chain - level load balancing model.

9. The microservice chain-level load balancing method based on deep reinforcement learning according to claim 1, wherein The step 5) includes the following steps: 5.1) Pre - process the load data, sort it in ascending order according to its predefined response time upper limit, and judge whether the health status of the service instances in the current microservice and the physical resources they contain meet the computing resources required for each request. If not, reject; 5.2) Call the deep reinforcement learning model of chain - level load balancing to obtain the microservice instance set; 5.3) Sort the instance set in ascending order according to the load level of each instance, and take the instance with the minimum load as the balancing result routing.

Citation Information

Cited By

  • Kubernetes micro-service optimal deployment method and device based on award accumulation deep reinforcement learning, and medium

    CN120803743A