A Method for Elastic Scaling of Multiple Microservice Replicas with Load Awareness in a Computing Power Network
Through graph attention network and deep context multi-arm gambling machine algorithm, a load-aware multi-microservice replica elastic scaling model is built, which solves the problem of inaccurate resource allocation in multi-microservice scenarios and realizes efficient resource utilization and response capabilities under dynamic load.
Patent Information
- Application Number
- CN202411685487.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-11-23
AI Technical Summary
When the existing microservice automatic scaling method is used to handle multi-microservice scenarios, it is difficult to effectively capture the coupling relationship between microservices, resulting in inaccurate resource allocation, affecting system performance and resource utilization, and online exploration is expensive, making it difficult to deal with dynamic load changes in real time.
Using an algorithm based on graph attention network and deep context multi-arm gambling machine, the microservice historical data is collected through non-invasive service mesh technology, a simulation environment is built, and the load-aware multi-service replica elastic scaling model is trained, and the number of microservice replicas is dynamically adjusted to meet the requirements of the Service Level Agreement (SLA) and improve resource utilization.
It realizes that the end-to-end response delay meets SLA in a dynamic load environment, while improving resource utilization and system response capabilities, reducing system performance fluctuations and online exploration costs.
Smart Images

Figure CN119512759B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, especially an automatic horizontal scaling algorithm for multi-microservice applications with complex business logic interweaving in the scenario of a computing power network. In this environment, the workload usually arrives in an online manner and is unpredictable and time-varying. This algorithm aims to handle such complex and dynamically changing workloads, ensuring that the end-to-end request response time meets the requirements of the service level agreement (SLA), while improving resource utilization and optimizing the system's response ability and scalability. Background Art
[0002] With the rapid development of computer network and cloud computing technologies, cloud applications are gradually evolving from traditional monolithic architectures to loosely coupled microservice architectures. The microservice architecture realizes higher scalability, flexibility, as well as fast deployment and iteration capabilities by splitting a large single application into multiple independent and clearly defined microservice modules. Each microservice can be developed, deployed, and scaled independently, greatly improving the overall efficiency and flexibility of the system. However, the advantages of this architecture also bring complexities in resource management and performance optimization.
[0003] Currently, the online workload of microservice applications is highly dynamic and unpredictable, making the reasonable allocation of computing resources a huge challenge. As the workload changes, the application system needs to dynamically adjust the computing resource allocation of each microservice on the premise of ensuring that the end-to-end response delay meets the requirements of the service quality (Service Level Agreement, SLA), so as to reduce costs and improve resource utilization. At present, the automatic scaling mechanism of microservices in the Kubernetes (K8S) cluster mainly relies on the CPU or memory utilization threshold strategy based on a single microservice. When the utilization exceeds the set threshold, it is scaled by increasing the number of container replicas. However, this independent scaling strategy fails to fully consider the interrelationships between multiple microservices, and the setting of the utilization threshold also depends on experience, making it difficult to accurately optimize according to the set SLA level.
[0004] Existing reinforcement learning (RL) methods have been widely applied to dynamic decision-making problems, especially in scenarios where the strategy needs to be continuously adjusted according to environmental changes, such as the dynamic scaling of a single microservice. However, for the scenario of multiple microservices with complex dependencies, existing methods have significant deficiencies in capturing the coupling relationship between microservices and training efficiency. Existing RL methods usually rely on frequent interactions with the real environment to explore the optimal strategy, which is particularly difficult in multi-microservice applications. Since the horizontal scaling of multiple microservices involves a huge state and action space, when exploring actions in the cluster, it is necessary to repeatedly explore the combinations of the scaling of multiple microservice replicas and continuously create and destroy replicas in the system, resulting in unacceptable time overhead. In addition, the frequent adjustment of microservice replicas will lead to significant response time fluctuations, thereby affecting the performance of the system. Therefore, in such a scenario with a huge state space and high operation costs, existing online exploration reinforcement learning methods are not only costly but also difficult to effectively respond to dynamic load changes in real time.
[0005] To address the changes in dynamic workloads in the multi-microservice architecture and improve resource utilization on the premise that the end-to-end response latency meets the SLA requirements, there is an urgent need for an efficient RL-based microservice auto-scaling method. This method should be able to effectively capture the coupling relationship between microservices and construct an accurate simulated interaction environment, thereby improving the training efficiency without directly interacting with the real environment. In addition, this method should also have the ability to handle large state and action spaces and quickly adapt to online workload changes. Summary of the Invention
[0006] The present invention addresses the problems existing in the above-mentioned background technology and discloses a method for elastic scaling of multiple microservice replicas with load awareness in a computing power network. According to the continuous change of the online request load of users, it improves the resource utilization rate of microservices on the premise of maximizing the guarantee that the end-to-end response latency meets the SLA requirements.
[0007] Step 1: Based on non-intrusive service mesh technology, design a historical metric collector in the K8S cluster based on plugins such as Prometheus and Istio, sample the cluster historical data and use it as the training dataset for the subsequent graph attention network.
[0008] 1) Deploy the Prometheus component and integrate cAdvisor and Istio. Deploy Prometheus in the Kubernetes cluster to monitor and collect the performance metrics of each microservice, and integrate cAdvisor and Istio into Prometheus. cAdvisor is responsible for monitoring the CPU consumption of each microservice instance, while Istio provides traffic information and call chain data in the service mesh to capture the call relationship and latency distribution between microservices.
[0009] 2) Cluster microservice metric data collection. Use the Locust tool to simulate time-varying workload requests with different numbers of users and multiple modes, and based on the random change of the number of replicas, comprehensive performance data can be regularly collected through the Prometheus query language. Denote the set of microservices that satisfy the DAG dependency relationship as M = {m1, m2, m3...m N}, where N is the number of microservices, and use the integrated Prometheus component to collect and classify metric data for each microservice. Since there is a certain time overhead for the creation and destruction of microservice replicas, and at the same time the sampling method of Prometheus is based on a specific time window, the sampling time interval is set to 30 seconds. At each timestamp t, the data sequence sampled for each microservice can be expressed as where represent the request arrival rate (QPS), CPU utilization, P90 latency, and the number of replicas of microservice m at time t respectively. We use the P90 latency as the response latency of a single microservice, and all the collected metric data will be stored in the time series database.
[0010] Step 2: Based on the time series database obtained in Step 1, simulate the impact of the changes in microservice workload and the number of replicas in the real environment on the CPU utilization and P90 latency of microservices, train the microservice CPU utilization and P90 latency predictors based on the graph attention network, and construct a simulation environment for interacting with the deep context multi-armed bandit.
[0011] 1) Construct a directed graph G = (M, E) according to the complex call relationships among multiple microservices, where the node set M corresponds to all microservices, and the edge set E represents the call relationships among microservices. Based on the microservice metric data at each timestamp t - 1 extracted from the time series database collected in Step 1, it is used to form the node feature vector at time t - 1 The clear task is to train the model to obtain the remaining features of each graph node at time t through the currently known features of each node at time t - 1 and the partial features of each node changed at time t Combine the features of each node to obtain the complete feature vector of each node at adjacent timestamps:
[0012]
[0013] 2) Use a two-layer graph attention network model to capture the complex dependencies between microservices. Utilize the call relationships between microservices to dynamically adjust the influence of neighbor nodes on the target node, and update the representation of each node through the feature aggregation of neighbor nodes, thereby accurately capturing the interaction between microservices. In the first layer of GAT, perform a linear transformation on the node input features and map them to a high-dimensional hidden space, calculate the attention weights between nodes to measure the influence of neighbor nodes, and after normalization, use it to weighted sum the neighbor node features and apply a non-linear activation function to obtain the updated node representation. At the same time, use the multi-head attention mechanism to enhance the model's learning ability. The second layer of GAT further extracts high-order feature representations, aggregates neighbor node information by recalculating the attention weights, and introduces a residual connection between the two layers to alleviate the problem of gradient disappearance. Use batch normalization and Dropout operations after each layer to enhance the model's stability.
[0014] 3) Use the updated node representations for prediction. Design two independent but identical-structured networks, where the output layer consists of a fully connected layer and a Sigmoid activation function, and predict the CPU utilization and P90 latency of microservices respectively.
[0015] Step 3: Based on the model trained in Step 2, design a deep contextual multi-armed bandit algorithm to interact with the model rather than the online environment, perform efficient offline training for the specific contextual information at each timestamp t, and obtain a multi-microservice auto-scaling model that adapts to different load changes.
[0016] 1) Define the context features, action space, and reward function of the multi-microservice application.
[0017] First, the context feature S t is defined as:
[0018] S t = (W t , W t-1 , R t-1 )
[0019] where W t and W t-1 represent the QPS of each microservice at time t and time t-1 respectively, and R t-1 represents the number of replicas of each microservice at time t-1. All three are N-dimensional vectors, and the value of their i-th (i ∈ [1, N]) dimension represents the value of the relevant performance index in the i-th microservice.
[0020] Then, the action space A t is defined as:
[0021]
[0022] where represents for microservice mi The scaling operation at the current moment, whose value can be the scalar 1, -1, or 0, corresponding to increasing the number of replicas by 1, decreasing by 1, or remaining unchanged, respectively. Each action selection means selecting a scaling policy for each microservice.
[0023] Next, the reward function is defined as:
[0024]
[0025] where the specific value of SLA is the upper limit of the response delay value that satisfies the service level agreement, α and β respectively represent custom weights, measuring the importance of improving the average CPU utilization rate of microservices and reducing the SLA violation rate, and C (C>0) represents the penalty coefficient when violating the SLO. rt t represents the end-to-end P90 response delay of the microservice system, which is the P90 delay of the front-end microservice in the multi-application microservice.
[0026] The optimization goal is to select appropriate replica scaling policies for all microservices given the current context state, so as to maximize the average resource utilization rate of microservices while ensuring the SLA violation rate.
[0027] 2) Design the deep context multi-armed bandit model architecture. The microservice system consists of multiple services. Based on the context features designed previously, due to its high spatial dimension, the action space also expands rapidly with the increase in the number of microservices and the available actions. To adapt to the changes in different load patterns and enhance the generalization ability of the model, a deep neural network is introduced in the CMAB to approximate the state-action value function Q(s, a; θ) to solve the problem of large state and action spaces.
[0028] 3) Use the microservice CPU utilization rate and P90 delay prediction models trained based on the graph attention mechanism in step 2 as the model environment for interacting with the deep context multi-armed bandit, and quickly train the model to optimize the resource management of the microservice system. In step 2, the prediction of the key performance indicators of the microservice system is realized through the graph attention network (GAT). The model can calculate according to the request arrival rate of each microservice at time t and the changed number of replicas and the known information at time t-1 to predict the changes in the CPU utilization rate and P90 response delay of each microservice, so as to realize the interaction process between DCMAB and the simulation environment. DCMAB continuously repeats the training process, observes the current state, selects actions, predicts performance, calculates rewards, and optimizes the strategy until the reward converges.
[0029] Step 4: Inject online time-varying workloads into the system, put the trained DCMAB into the actual environment, use the Prometheus plugin to detect the metric changes of microservices every time interval t and continuously update the context information, input the real-time detected context information into the trained DCMAB, prompt the model to make replica level scaling decisions for each microservice at each time stamp t, and obtain the final application SLA violation situation and microservice resource utilization situation after a set time T. Description of the Drawings
[0030] Figure 1 is the flowchart of the present invention;
[0031] Figure 2 is the system architecture diagram of the present invention. Detailed Embodiments
[0032] As Figure 2 shown, the specific steps of the technical solution of the present invention are as follows:
[0033] Step 1: Based on non-intrusive service mesh technology, design a historical metric collector in the K8S cluster based on plugins such as Prometheus and Istio, sample the cluster historical data and use it as the subsequent training dataset for the graph attention network. The specific steps are as follows:
[0034] 1) Deploy the Prometheus component and integrate cAdvisor and Istio. Deploy Prometheus in the Kubernetes cluster to monitor and collect the performance metrics of each microservice, and integrate cAdvisor and Istio into Prometheus. cAdvisor is responsible for monitoring the CPU consumption of each microservice instance, while Istio provides traffic information and call chain data in the service mesh, capturing the call relationships and latency distributions between microservices.
[0035] 2) Collect cluster microservice metric data. Use the Locust tool to simulate time-varying workload requests with different user numbers and multiple modes, and based on the random changes in the number of replicas, be able to regularly collect comprehensive performance data through the Prometheus query language. Denote the set of microservices that satisfy the DAG dependency relationship as M = {m1, m2, m3...m N}, where N is the number of microservices, and use the integrated Prometheus component to collect and classify the metric data of each microservice. Since there is a certain time overhead for the creation and destruction of microservice replicas, and at the same time the sampling method of Prometheus is based on a specific time window, set the sampling time interval to 30 seconds. At each time stamp t, the data sequence sampled for each microservice can be expressed as Among them respectively represent the request arrival rate (QPS), CPU utilization rate, P90 latency, and the number of replicas of microservice m at time t. We use the P90 latency as the latency of a single microservice, and all the collected metric data will be stored in a time series database.
[0036] Step 2: Based on the time series database obtained in Step 1, simulate the impact of changes in microservice workload and the number of replicas on the CPU utilization rate and P90 latency of microservices in a real environment, train predictors for the CPU utilization rate and P90 latency of microservices based on graph attention networks, and construct a simulation environment for interacting with a deep contextual multi-armed bandit. The specific steps are as follows:
[0037] 1) Construct a directed graph G=(M, E) according to the complex call relationships among multiple microservices, where the node set M corresponds to all microservices, and the edge set E represents the call relationships among microservices. If microservice m i calls m j , then there is a directed edge (m i , m j ). Based on the microservice metric data at each timestamp t - 1 extracted from the time series database collected in Step 1, form the node feature vector of the previous moment Since the online load is constantly changing in a real environment, that is constantly changing. To construct a simulation environment for interacting with a multi-armed bandit, it is necessary to simulate the contribution of changing the number of replicas to the immediate reward based on the known at time t. And the reward function is related to the CPU utilization rate and P90 latency . Therefore, the problem is transformed into: how to train a model to obtain the remaining features of each graph node at time t through the known node features at time t - 1 of each node and the known partial features of each node at time t . All the above features can be obtained from the time series database in Step 1. Combine the features of each node to obtain the complete feature vector of adjacent timestamps of each node: All the above features can be obtained from the time series database in Step 1. Combine the features of each node to obtain the complete feature vector of adjacent timestamps of each node:
[0038]
[0039] Form a complete graph data structure for model training by aggregating the complete node features and call relationships.
[0040] 2) Use a two-layer graph attention network model to capture the complex dependencies between microservices. Utilize the call relationships between microservices to dynamically adjust the influence of neighbor nodes on the target node. Update the representation of each node through the feature aggregation of neighbor nodes, thereby accurately capturing the interaction between microservices.
[0041] In the first layer of GAT, linearly transform the input features of each node to map them to a higher-dimensional hidden space:
[0042]
[0043] where is a trainable weight matrix used to map the input feature dimension d to the hidden dimension d'. Then, for each node m i and its neighbor node m j ∈N(m i ), calculate the attention weight e ij to measure the influence of neighbor node m j on node m i :
[0044]
[0045] where is a trainable attention vector, || represents the vector concatenation operation, and LeakyReLU is a non-linear activation function used to introduce non-linear features. Then, normalize the attention weights through the softmax function:
[0046]
[0047] The normalized attention weight α ij is used to represent the influence degree of neighbor node m j on the target node m i . The sum of all weights is 1, so that each node can dynamically adjust the influence degree of its neighbor nodes on itself. Next, use the normalized attention weights to perform a weighted sum of the features of neighbor nodes and apply the non-linear activation function σ to obtain the updated node representation:
[0048]
[0049] To improve the expressive power of the model, use the multi-head attention mechanism (Multi-Head Attention), that is, simultaneously use K independent attention heads to calculate the feature representation. This can capture different types of dependency relationships between microservices, enhance the learning ability of the model, and make the feature representation more robust. Concatenate these features together:
[0050]
[0051] Node representations generated by the first - layer GAT Are further input into the second - layer GAT to extract higher - order feature representations. The second - layer GAT only uses one attention head and no longer performs the concatenation operation:
[0052]
[0053] Among them, Is the weight matrix of the second layer, and α' ij Is the recalculated attention weight. In this way, the second - layer GAT can further aggregate the information of neighbor nodes to form a more comprehensive node representation.
[0054] To alleviate the problem of gradient vanishing, a residual connection is introduced between the two layers of GAT. Specifically, the input feature After linear transformation, is added to the output of the second - layer GAT:
[0055]
[0056] Among them, Is the linear transformation matrix. At the same time, to further enhance the stability of the model, batch normalization and Dropout (randomly discard a part of neurons) are used after each layer to prevent overfitting and improve the generalization ability.
[0057] 3) Use the updated node representations for prediction. Two independent but structurally identical networks are designed respectively for predicting the CPU utilization rate and P90 latency of each microservice. The output layer of each network consists of a fully - connected layer and a Sigmoid activation function. The P90 latency is normalized, and the CPU utilization rate and P90 latency predictions are respectively expressed as:
[0058]
[0059] Among them And Are both trainable weight matrices. The independent design of the two networks allows them to focus on different goals while sharing the same GAT structure, thus optimizing the training efficiency while improving the model's prediction ability.
[0060] Step 3: Based on the model trained in Step 2, design a deep - context multi - armed bandit algorithm to interact with the model rather than the online environment, and perform efficient offline training for the specific context information at each timestamp t to obtain a multi - microservice auto - scaling model that adapts to different load changes. The specific steps are as follows:
[0061] 1) Define the context features, action space, and reward function of the multi-microservice application.
[0062] First, the context feature S t is defined as:
[0063] S t = (W t , W t-1 , R t-1 )
[0064] where W t , W t-1 represent the QPS of each microservice at time t and t - 1 respectively, and R t-1 represents the number of replicas of each microservice at time t - 1. All three are N-dimensional vectors, and the value of their i-th (i ∈ [1, N]) dimension represents the value of the relevant performance index in the i-th microservice. The significance of designing the context features like this is that it is necessary to decide how to scale the replicas of each microservice at the current time t based on the microservice load situation and microservice replica settings at time t - 1 and the dynamically changed load at the current time t to maximize the immediate reward.
[0065] Then, the action space A t is defined as:
[0066]
[0067] where represents the scaling operation of microservice m i at the current time, and its value can be the scalar 1, -1, or 0, corresponding to the number of replicas +1, -1, or remaining unchanged respectively. Each action selection is to select a scaling strategy for each microservice.
[0068] Next, the reward function is defined as:
[0069]
[0070] where the specific value of SLA is the upper limit of the response delay value that satisfies the service level agreement, α and β respectively represent custom weights that measure the importance of improving the average CPU utilization rate of microservices and reducing the SLA violation rate, and C (C > 0) represents the penalty coefficient when violating the SLO. rt t represents the end-to-end P90 response delay of the microservice system, which is the P90 delay of the front-end microservice in the multi-application microservice.
[0071] The optimization goal is to select appropriate replica scaling strategies for all microservices given the current context state, so as to maximize the average resource utilization rate of microservices while ensuring the SLA violation rate.
[0072] 2) Design a deep context multi-armed bandit model architecture. The microservice system consists of multiple services. Based on the context features designed previously, due to its high spatial dimension, the action space expands rapidly as the number of microservices and the number of available actions increase. In addition, the relationship between system performance and the state and action spaces is usually non-linear, and traditional context multi-armed bandits (CMAB) are difficult to effectively handle these complexities. To adapt to changes in different load patterns and enhance the generalization ability of the model, a deep neural network is introduced into CMAB to approximate the state-action value function Q(s, a; θ) to solve the problems of large state and action spaces.
[0073] 3) Use the microservice CPU utilization and P90 latency prediction model trained based on the graph attention mechanism in step 2 as the model environment for interacting with the deep context multi-armed bandit, and quickly train the model to optimize the resource management of the microservice system. In step 2, the prediction of the key performance indicators of the microservice system is achieved through the graph attention network (GAT). The model can predict the CPU utilization of each microservice based on the request arrival rate at the current moment of each microservice and the changed number of replicas and predict the P90 response latency of each microservice, thereby realizing the interaction process between DCMAB and the simulation environment. The specific interaction process is as follows:
[0074] Action selection strategy: At time t, by observing the current context feature S t , obtain W at the previous moment t-1 , R t-1 and W at the current moment t , and then use the UCB strategy to select the action A t . The basic idea of the UCB strategy is to assign a confidence upper bound to each action to encourage the selection of actions that have not been fully explored. Each time an action is selected, the value with the largest UCB is chosen. The specific formula is:
[0075]
[0076] where Q(S t , A t ; θ) is the reward estimate of the action A t in the current state S t , θ represents the parameters of the neural network, c is the exploration parameter, which controls the balance between exploration and exploitation, and N(A t ) represents the number of times the action A t has been selected.
[0077] Interaction with the simulation environment (the model in step 2): After taking the action A t , obtain the new number of replicas R of each microservice at time tt , so by aggregating S t =(W t , W t-1 , R t-1 ), R t , the microservice P90 latency R at time t - 1 t-1 , and the microservice CPU utilization C at time t - 1 t-1 , and inputting into the trained GAT model in Step 2, the execution action A t can be obtained, and then the subsequent L t and R t , that is, the response latency and CPU utilization of each microservice at time t. Thus, the reward R t after the current execution action can be calculated.
[0078] Model Update: According to the calculated reward R t and the current state S t , use a deep neural network to update the DCMAB model parameters and minimize the loss function:
[0079]
[0080] By continuously optimizing the neural network parameters θ, the model gradually learns to select the best replica adjustment strategy under different load conditions.
[0081] DCMAB continuously repeats the above process, observes the current state, selects actions, predicts performance, calculates rewards, and optimizes the strategy until the reward converges.
[0082] Step 4: Inject online time-varying workloads into the system, put the trained DCMAB into the actual environment, use the Prometheus plugin to detect the metric changes of microservices every time t and continuously update the context information, input the real-time detected context information into the trained DCMAB, prompt the model to make replica level scaling decisions for each microservice at each timestamp t, and obtain the final application SLA violation situation and microservice resource utilization situation after a set time T.
Claims
1. A method for elastic scaling of multiple microservice replicas with load awareness in a computing power network, characterized in that The method at least includes the following steps: Step 1: Based on non-intrusive service mesh technology, design a historical metric collector in the Kubernetes cluster based on Prometheus, Istio, and cAdvisor plugins, sample the cluster historical data to obtain a time series data set for training subsequent environment models; Step 2: Based on the time series data set obtained in Step 1, train a microservice CPU utilization and response latency predictor based on a graph attention network, and construct an environment model that interacts with the deep contextual multi-armed bandit algorithm; Step 3: Based on the environment model trained in Step 2, design a deep contextual multi-armed bandit algorithm to interact with the model, perform efficient offline training for the specific context information at each timestamp t, and obtain a multi-microservice auto-scaling model that adapts to different load changes; Step 4: Based on the multi-microservice auto-scaling model obtained in Step 3, inject time-varying online loads in the Kubernetes system and deploy the auto-scaling model to verify the SLA violation rate and CPU resource utilization of microservice applications; The steps of training the microservice CPU utilization and response latency predictor based on the graph attention network in Step 2 further include the following steps: 1) Construct a directed graph G=(M, E) according to the complex call relationships between multi-microservices, where the node set M corresponds to all microservices, and the edge set E represents the call relationships between microservices. Combine the respective node features to obtain a complete feature vector for each node: Among them respectively represent the workload, CPU utilization, response latency, and number of replicas of microservice m at time t-1, and respectively represent the workload and number of replicas of microservice m at time t. By aggregating the complete node features and call relationships, a complete graph data structure is formed for model training; 2) Use a two-layer graph attention network model to capture the complex dependencies between microservices, utilize the call relationships between microservices, dynamically adjust the influence of neighbor nodes on the target node, and update the representation of each node through the feature aggregation of neighbor nodes, so as to accurately model the interaction between microservices; 3) Use the updated node representation for prediction. Design two independent but identically-architected networks, which are respectively used to predict the CPU utilization of each microservice at time t and the response latency Thereby, construct an environment model that interacts with the deep contextual multi-armed bandit algorithm.
2. The method for elastic scaling of multiple microservice replicas with load awareness in a computing power network according to claim 1, wherein The deep contextual multi-armed bandit algorithm in Step 3 at least further includes the following steps: 1) Define the context features of the multi-microservice application, construct the state space, dynamic space, and reward function, and the state space S t is defined as: S t = (W t , W t-1 , E t-1 ) where and represent the workload vectors of each microservice at time t and t - 1 respectively, represents the replica number vector of each microservice at time t - 1. All three are N - dimensional vectors, where N is the number of microservices, and the value of the i - th dimension of the vector represents the value of the relevant performance index in the i - th microservice; Action space A t is defined as: Among them represents the scaling operation of microservice m at the current moment. Its value is a scalar 1, -1, or 0, corresponding to an increase of 1 in the number of replicas, a decrease of 1, or remaining unchanged, respectively. Each action selection means selecting a scaling policy for each microservice; The reward function is defined as: where α and β represent custom weights that measure the importance of improving the average CPU utilization of microservices and reducing the SLA violation rate, C represents the penalty coefficient when the SLA is violated, C > 0, and rt t represents the end-to-end response latency of the microservice system, specifically manifested as the response latency of the front-end microservice; 2) Use the microservice CPU utilization and response latency prediction model trained based on the graph attention mechanism in step 2 as the environmental model for interacting with the deep contextual multi-armed bandit for efficient policy training, thereby optimizing the resource management ability of the microservice system; To adapt to different load pattern changes, the model needs to have good generalization ability. Introduce a deep neural network with strong representation ability into the traditional contextual multi-armed bandit to approximate the state-action value function to solve the problem of large state and action spaces; At time t, by observing the current context feature S t , including W at time t-1 t -1 , E t-1 and W at time t t , use the UCB strategy to select action A t . After taking the action, get the new Then aggregate S t , E t and the remaining known to obtain the complete input feature {W t-1 , C t-1 , L t -1 , E t-1 , W t , E t}. Split this input feature into the features of each microservice By inputting the of all microservices into the environmental model, the predicted L t and C t can be obtained, thereby calculating the reward R t of the current action and updating the context feature S t , thus realizing the interaction process between the deep contextual multi-armed bandit algorithm and the environmental model; This process continues iteratively, by observing context information, selecting actions, predicting performance, calculating rewards, and continuously updating network parameters to optimize the policy until the reward converges, so that the model gradually learns to select the optimal replica adjustment scheme under different load conditions.
Citation Information
Patent Citations
Container resource dynamic scheduling method and system based on usage amount prediction
CN115118602A
Heterogeneous graph anomaly detection method based on adaptive selection of multi-channel neighborhood nodes
CN118540239A