Micro-service automatic scaling method based on PPO-DU algorithm

The automatic microservice scaling method built through the PPO-DU algorithm uses the Actor-Critic neural network and the policy delay update mechanism to solve the resource allocation problem of automatic microservice scaling in complex load scenarios, and realizes efficient and stable resource management and cost control.

CN120429112APending Publication Date: 2025-08-05CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510519007.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing microservice automatic scaling algorithm is difficult to accurately adjust resource allocation in the face of complex dynamic load scenarios, and the response is lagging, and it is difficult to take into account resource utilization, service quality and cost control.

Method used

Using the automatic microservice scaling method based on the PPO-DU algorithm, the Markov decision-making process and the Actor-Critic neural network are constructed, combined with the strategy delay update mechanism, and the reward function is designed to optimize the number of containers to achieve automatic scaling of microservices.

Benefits of technology

It realizes accurate prediction of load demand in complex dynamic load scenarios, dynamic adjustment of resource allocation, improves system stability and resource utilization efficiency, reduces operation and maintenance costs, is highly adaptable, and can operate stably in high-load environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429112A_ABST
    Figure CN120429112A_ABST
Patent Text Reader

Abstract

The invention relates to the field of micro-service automatic scaling, in particular to a micro-service automatic scaling method based on a PPO-DU algorithm. The method comprises the following steps: defining a micro-service cluster, and converting a micro-service automatic scaling target into a target optimization problem; in order to solve a target optimization problem, a Markov decision process is constructed, and a reward function is designed by using a resource utilization rate and service quality; according to the Markov decision process, an Actor-Critic neural network is designed, and the network comprises state input, an Actor network, a Critic network and output; according to the method, an Actor-Critic neural network is combined, a delay updating strategy is introduced, and a PPO-DU algorithm is designed and used for outputting the adjustment proportion of the number of containers according to the state of the micro-service, so that automatic scaling of the micro-service is realized. According to the invention, the optimal scaling strategy can be dynamically generated according to the real-time change of the system operation state, the load demand can be accurately predicted, the resource allocation can be dynamically adjusted, the stable operation of the micro-service system in a high-load environment can be ensured, and the resource waste and the operation and maintenance cost can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of microservice automatic scaling, and in particular to a microservice automatic scaling method based on the PPO-DU algorithm. Background Art

[0002] With the rapid development of cloud computing technology, the Software as a Service (SaaS) market has experienced rapid expansion. Simultaneously, the size and complexity of systems within microservices architectures continue to rise. From a technical implementation perspective, research into autoscaling technology is crucial to ensure the efficient and stable operation of microservices systems within cloud platforms. Microservices autoscaling technology accurately matches business load demands, dynamically adjusting service instance size and resource allocation, achieving system elasticity and high availability. This technology can maximize resource utilization and significantly reduce enterprise cloud service costs.

[0003] Microservices architecture, a core cloud-native architecture, is widely adopted, characterized by agility, scalability, and ease of maintenance. However, its dynamic nature presents challenges for autoscaling. Microservice instances are numerous and widely distributed, and their workload patterns are in a state of dynamic change and instability. Simple rule-based approaches struggle to accurately grasp resource demand fluctuations, making them prone to deviations. Dynamically scheduling a large number of microservice instances requires a global resource perspective for overall planning and efficient coordination for flexible deployment, but existing methods struggle to adapt to complex environments. Finally, balancing diverse metrics across multiple dimensions, such as resource utilization, ensuring application performance, and controlling costs, presents a daunting challenge. Traditional autoscaling algorithms struggle to maximize overall system efficiency.

[0004] Currently, there are four main types of commonly used autoscaling algorithms: rule-based algorithms; machine learning-based algorithms; time series prediction-based algorithms; and deep reinforcement learning-based algorithms. Rule-based algorithms perform scaling operations based on preset rules, such as thresholds for metrics like CPU utilization, memory usage, and number of requests. These rules can be static or dynamically adjusted based on historical data. While intuitive and easy to use, these algorithms require manual setting of reasonable rules, limiting their adaptability to complex and changing scenarios. For example, Kubernetes' Horizontal Pod Autoscaler (HPA) automatically adjusts the number of application instances to meet system load requirements based on preset monitoring metrics and thresholds. Users can set different metrics and thresholds as needed to implement flexible autoscaling policies. Machine learning-based algorithms use machine learning models, such as linear regression, decision trees, and neural networks, to automatically learn resource demand patterns from historical data to achieve automated resource scheduling. These algorithms have strong data-driven and generalizable capabilities, but rely on large amounts of labeled training data. Time series prediction-based algorithms estimate future workloads and then allocate resources based on the predicted results. The algorithm based on deep reinforcement learning transforms automatic scaling into a deep reinforcement learning problem. The intelligent agent continuously optimizes its resource management strategy by interacting with the environment and obtaining feedback.

[0005] By analyzing and studying the horizontal scaler, existing reinforcement learning automatic scaling algorithms, and time series prediction algorithms, we found that they have the following shortcomings:

[0006] (1) HPA has a slow response speed and is difficult to handle complex scenarios. Because HPA triggers scaling based on periodic monitoring indicators, its response delay is large. When faced with complex workloads, service quality requirements, and multiple resource constraints, a single threshold rule makes it difficult to make reasonable scaling decisions.

[0007] (2) Existing reinforcement learning automatic scaling methods have poor decision-making stability when facing different loads. In real-world scenarios, load patterns and characteristic parameters are diverse, making it difficult to accurately select services for scaling decisions. They often rely on specific environments or architectures and are unable to automatically trigger in complex environments, resulting in unstable scaling results.

[0008] (3) Time series prediction algorithms have high requirements for data patterns. When faced with sudden load changes and complex situations, their scalability and generalization capabilities are often insufficient. Summary of the Invention

[0009] To address the issues of inaccurate resource allocation and delayed response in traditional auto-scaling technologies under dynamic load scenarios, this paper provides a microservice auto-scaling method based on the PPO-DU algorithm, which mainly includes:

[0010] S1: Define a microservice cluster and convert the microservice autoscaling goal into a target optimization problem.

[0011] S2: To solve the target optimization problem in S1, a Markov decision process is constructed and the reward function is designed using resource utilization and service quality.

[0012] S3: Based on the Markov decision process, design an Actor-Critic neural network, which includes state input, Actor network, Critic network, and output;

[0013] S4: Combined with the Actor-Critic neural network, a delayed update strategy is introduced to design the PPO-DU algorithm, which is used to adjust the number of containers according to the status of the microservice output, thereby achieving automatic scaling of the microservice.

[0014] A computer device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0015] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above method.

[0016] A computer program product includes a computer program or instructions, which implement the steps of the above method when the program or instructions are executed by a processor.

[0017] The beneficial effects of the technical solution provided by the present invention are as follows: by introducing a policy delay update mechanism and a comprehensive analysis of multi-dimensional indicators, the present invention can dynamically generate the optimal scaling strategy according to the real-time changes in the system's operating status, accurately predict load demand and dynamically adjust resource allocation, significantly improving the intelligence level of resource scheduling and system stability. This method performs well in dealing with sudden, periodic, and highly volatile load scenarios, and can ensure the stable operation of microservice systems in high-load environments, while significantly reducing resource waste and operation and maintenance costs. Through the comprehensive analysis of multi-dimensional indicators and the policy delay update mechanism, the present invention demonstrates excellent adaptability in complex dynamic environments and fully meets the actual needs of the industry. It provides enterprises with a more efficient and economical solution, and has extremely high application value in the face of complex load scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0019] Figure 1 This is a flowchart of a microservice automatic scaling method based on the PPO-DU algorithm in an embodiment of the present invention;

[0020] Figure 2Schematic diagram comparing the training process of PPO-DU and PPO algorithms in an embodiment of the present invention;

[0021] Figure 3 is a schematic diagram of a burst load in Sock-Shop according to an embodiment of the present invention;

[0022] Figure 4 is a schematic diagram of a periodic load in a Sock-Shop according to an embodiment of the present invention;

[0023] Figure 5 is a schematic diagram of fluctuating load in a Sock-Shop according to an embodiment of the present invention;

[0024] Figure 6 is an Actor-Critic neural network diagram in an embodiment of the present invention;

[0025] Figure 7 Schematic diagram of pseudo code of the PPO-DU training process in an embodiment of the present invention;

[0026] Figure 8 This is a flow chart of policy delay update in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to have a clearer understanding of the technical features, purposes and effects of the present invention, specific embodiments of the present invention are now described in detail with reference to the accompanying drawings.

[0028] Example 1

[0029] Please refer to Figure 1 , Figure 1 This is a flowchart of a microservice automatic scaling method based on the PPO-DU (Proximal Policy Optimization with Delayed Updates) algorithm in an embodiment of the present invention, which specifically includes the following steps:

[0030] S1: First, define the microservice cluster, as shown in formula (3.1), the microservice cluster MS consists of multiple microservice ms i Composition, n is the total number of microservices. Each microservice ms i It is composed of multiple microservice instances, which are deployed on containers, as shown in formula (3.2), where m is the microservice ms i The number of containers.

[0031] MS={ms1,ms2,…,ms n} (3.1)

[0032] ms i={container1,container2,…,container m},i∈[1,n] (3.2)

[0033] Each container consists of multiple states, as shown in formula (3.3). The state of a general container is It is composed of resources such as CPU, memory (RAM), storage (ROM), etc. These states reflect the resource usage of the container. These containers are deployed on multiple physical hosts, as shown in formula (3.4), where l is the total number of physical hosts. Multiple physical hosts constitute the entire host cluster, and the states of these physical hosts are also composed of resources such as CPU, memory (RAM), storage (ROM), etc. As shown in formula (3.5), the resource state of the host It also determines the maximum number of containers that can be allocated to the host.

[0034]

[0035] Host={host1,host2,…,host l} (3.4)

[0036]

[0037] After formalizing the microservice autoscaling problem and modeling the microservice autoscaling environment, we need to model the microservice autoscaling problem. Autoscaling adjusts the number of containers to meet load requirements, prioritizing the service level objective (SLO). Typically, an SLO consists of multiple objectives, as shown in Equation (3.6), where RT is the request response time, AVA is the target availability percentage, and HC is the target throughput value.

[0038] SLO = {RT, AVA, HC} (3.6)

[0039] Typically, the SLO is determined primarily by the request response time (RT), while minimizing the total resource usage of the microservice. Therefore, the microservice autoscaling goal is transformed into a target optimization problem as shown in Formula (3.7).

[0040]

[0041] in, Indicates that when the container request response time is lower than the target response time SLO RT It is 1 when the service is running, otherwise it is 0. This method maximizes the total number of containers that meet the target response time and tries to ensure that the service meets the SLO target. Indicates the total amount of container resources. Minimizes container resource usage while ensuring that the resource usage of all containers in the cluster is less than the resource usage of all physical hosts.

[0042] Step 2: Before designing the algorithm, it is necessary to construct a Markov Decision Process (MDP). The decision mainly includes the state space, action space, reward function and state transition probability. The reward function is designed according to resource utilization and service quality. Resource usage is determined by the utilization of CPU and memory, and the quality of service is determined by the response delay of the service. Therefore, the state space is composed of CPU utilization, memory utilization and service response delay. In order to meet the user service quality and ensure SLO, the target of service response delay is set to 500ms. The service response delay is set to 500ms based on many factors. First, from the perspective of user experience, it is set to 500ms to ensure that users feel smooth and less stuck during the use of the service, and it does not require too much system resources. Secondly, 500ms is a common industry reference standard. The default SLA of cloud service providers usually promises that 99.9% of request delays are ≤500ms. However, the service response delay can also be flexibly adjusted according to the specific business scenario and resource costs. SLO is defined as:

[0043]

[0044] The state space at time t is thus defined as:

[0045] s t =(U CPU ,U mem ,SLO) (3.9)

[0046] The action space is the probability distribution of container scaling ratios in autoscaling. In horizontal autoscaling, actions mainly involve reducing or doubling the number of microservice instances. Therefore, the action space is defined as:

[0047] a t ∈(0.5,2) (3.10)

[0048] The reward function is usually defined as the reward or penalty obtained by the agent when it reaches a certain state. It is mainly composed of CPU utilization, memory utilization and service response delay. At the same time, in order to reduce the resource consumption of the container scaling process and prevent the agent from frequently scaling automatically, a t It is added to the reward function. Therefore, the reward function can be defined as:

[0049] r t =2×SLO+U CPU +U mem +(1-at ) (3.11)

[0050] In practice, SLO∈[0,1](99%), SLO∈(-∞,0)(1%), U CPU ∈(0,1],U mem ∈(0,1], where a t ∈[0.5,2]. At the same time, since the main goal is to prioritize SLO delay and reduce resource utilization on this basis, the weight of SLO is set to 2. In order to reduce the impact of instance scaling on resource consumption, it is necessary to give a hyperparameter ξ to (1-a t ) to adjust the parameters. The reward function is:

[0051]

[0052] Step 3: If Figure 6 As shown in the figure, the Actor-Critic neural network mainly consists of four parts: state input, Actor network, Critic network, and output.

[0053] vector s t Represents the state of the microservice as the input of the Actor-Critic network. The Actor network consists of three fully connected layers, each connected to a Tanh activation function. The first fully connected layer maps the state input to a 64-dimensional hidden representation, the second fully connected layer further extracts high-level features, and the third fully connected layer maps it to a scalar value, representing the action a. t The mean of . Finally, the output is scaled to the range of [0,1] through the Sigmoid activation function, and linearly transformed to the range of [0.5,2], indicating the scaling ratio of the number of containers. Similar to the Actor network, the Critic network is also composed of three fully connected layers, each of which is followed by a Tanh activation function. The first two fully connected layers have the same function as the Actor network, which is used to extract high-level features of the state. The third fully connected layer maps the 64-dimensional hidden representation to a scalar value, which represents the estimated value of the state. Finally, the Actor network outputs the action, which is the scaling ratio of the number of containers, in the range of [0.5,2]. Less than 1 means reducing the container, and greater than 1 means increasing the container. The Critic network outputs the state value, which evaluates the quality of the current state. Through cyclic training of the loss function L(θ,φ), the best action a is obtained. t .

[0054] Step 4: The PPO-DU algorithm is an optimization algorithm based on PPO. It mainly uses delayed update to ensure the stability and efficient strategy optimization during the training process of the PPO algorithm. The pseudo code of the PPO-DU training process is as follows: Figure 7As shown in the figure, the process of policy delay update is as follows Figure 8 As shown in Figure 2, the specific training process of the PPO-DU algorithm is divided into three steps:

[0055] (1) Initialization phase: Initialize the Actor network π θ and Critic Network V φ , used for estimating the policy function and value function respectively. Set the hyperparameters required for training, including learning rate α, discount factor γ, clip range ∈, update frequency K, maximum training episodes M and maximum steps N per episode.

[0056] (2) Training phase: First, for each episode, the environment needs to be reset to obtain the initial state s0. Second, for each stept, the current strategy π θ (a t |s t ) to select action a t , and calculate the logarithmic probability logπ of the corresponding action θ (a t |s t ). Then, perform action a t , get reward r t , next state s t+1 And whether it is done, and (s t ,a t ,r t ,s t+1 ,done,logπ θ (a t |s t )) is stored in the buffer for subsequent training. If the current stept is a multiple of the update frequency K, the following update steps are required: For each sample i in the buffer, calculate the advantage function This represents the advantage of the current policy relative to the value function, which is calculated by calculating δ starting from the current step i j The weighted sum of δ is used to estimate j Represents the TD residual, i.e. r j +γV φ (s j+1 )-V φ (s j ).

[0057] After completing the above steps, to calculate the target value It is state s i The true value of is used to update the value function network. The calculation formula is the advantage function plus the current value function estimate. Then calculate the ratio r i(θ), which represents the probability ratio of the new and old strategies, and is used to measure the magnitude of strategy update. CLIP (θ) is the objective function of PPO, which achieves stable strategy update by limiting the change range of ratio. The specific calculation is to take the minimum value of the product of ratio and clipped ratio and advantage function, and average it over all samples. Finally, it is necessary to calculate the value function loss L VF (φ), which represents the mean squared error loss of the value function network, is used to update the critic network. The total loss L(θ,φ) must also be calculated, which is the weighted sum of the policy objective function and the value function loss, serving as the final optimization objective. Finally, the Adam optimizer is used to update the parameters θ and φ of the actor and critic networks to minimize the total loss L(θ,φ). The data in the buffer is cleared to prepare for the next update.

[0058] (4) Training ends: When all episodes are finished, the training is completed. The trained Actor network π will be obtained. θ , the network can adjust the ratio of the number of output containers according to the status of microservices, thereby achieving automatic scaling of microservices.

[0059] The PPO-DU algorithm in the present invention introduces a delayed update strategy. The PPO-DU algorithm ensures the stability and efficient policy optimization of the PPO algorithm during the training process through delayed policy updates. The strategy mainly includes six core steps. First, the policy parameters need to be initialized, then the trajectory data is sampled, and the corresponding policy loss and value function loss are calculated, and they are temporarily saved in the waiting area. If the update period K is not reached, a new data area needs to be sampled and added to the waiting area; if the update period K is reached, the policy parameters are updated once using the optimization algorithm based on the accumulated waiting area data, and then the data in the waiting area is cleared. Repeat the above operations until the termination condition is reached, such as model convergence or the number of training rounds is exhausted. Finally, the final optimized policy parameters θ are returned.

[0060] Frequent policy updates can cause the data used in two consecutive updates to be highly correlated, resulting in large variance and instability. Delayed updates reduce this correlation by mixing new and old data, improving stability and reducing policy oscillation.

[0061] In addition to selecting a fixed K value based on experience, we also consider dynamically adjusting the K value: setting a smaller K value in the early stage of training to accelerate the strategy to adapt to environmental changes; gradually increasing the K value in the later stage to enhance the stability of the strategy. θ and old strategies When the difference is large, increase the K value to stabilize updates; when the difference is small, decrease the K value to speed up responses. Properly setting and adjusting the K value is key to achieving stability and responsiveness in the PPO-DU algorithm for microservice autoscaling. Specific parameters require extensive experimentation and tuning based on actual scenarios and algorithm performance.

[0062] The PPO-DU algorithm also uses a learning rate decay mechanism. This mechanism accelerates the convergence of the PPO-DU algorithm. The learning rate gradually decreases with the number of iterations. A larger learning rate in the early stages of training accelerates the model's approach to the optimal solution, but an excessively large learning rate may cause the model to oscillate near the optimal solution and fail to converge. Reducing the learning rate in the later stages allows the model to converge smoothly to the optimal solution, improves its generalization ability, and avoids local optima.

[0063] In order to verify the effectiveness of the method disclosed in the present invention, the following experiments were performed:

[0064] 1. Experimental setup

[0065] (1) Experimental environment and experimental parameter settings. The cloud computing simulation platform used in the experiment is a simulation experimental environment built based on PyCloudSim, and PyTorch 2.1.1 is used as a related component for building neural networks. Each experiment uses 3 fixed random seeds to take the average value for verification, thereby reducing the error of the algorithm in training and testing. The main parameters of the machine used in the experiment are: CPU is 16-core AMD Ryzen76800H@4.7GHz; memory is 16GB, GPU is NVIDIAGeForce RTX 3060Laptop GPU; operating system is Windows11; software configuration includes PyCharm2022.1, Python3.10, PyCloudSim0.0.1, PyTorch 2.1.1, etc.

[0066] The experiment used PyCloudSim to simulate various entities in a microservice system. The microservice cluster was modeled after a real-world Kubernetes cluster. The cluster was configured with four master nodes: master, node 1, node 2, and node 3. Each node had identical parameters. Specific parameters are shown in Table 1 (Node Parameters). Furthermore, to meet the resource requirements of service requests, parameters were set for each microservice's corresponding container. Container parameter configuration is shown in Table 2. Specific parameters for the PPO and PPO-DU algorithms are shown in Table 3. Each experiment combined a normal load with an abnormal load and issued random requests based on the microservice call relationships to generate realistic loads.

[0067] Table 1 Node parameter settings

[0068]

[0069] Table 2 Container parameters

[0070]

[0071] Table 3 PPO and PPO-DU parameters

[0072]

[0073] (2) Comparison of algorithms. In order to comprehensively evaluate the effectiveness of the proposed PPO-based autoscaling algorithm PPO-DU, this experiment selected the DQN-based ADRL autoscaling comparison algorithm. The main microservice autoscaling methods compared include the Kubernetes native Horizontal PodAutoscaler autoscaling algorithm (HPA), the DQN-based ADRL autoscaling algorithm (ADRL), the PPO-based autoscaling algorithm (PPO), and the PPO-DU-based autoscaling algorithm (PPO-DU).

[0074] (3) Evaluation indicators. In order to comprehensively evaluate the performance of each algorithm, nine key indicators were selected as evaluation criteria, including: average CPU utilization, average memory utilization, average request completion time, service level objective (SLO) violation rate, service workflow completion rate, request completion rate, average number of Pods in algorithm scaling, time required for a single algorithm scaling, and algorithm memory consumption. These indicators can comprehensively reflect the performance of the automatic scaling algorithm in terms of resource utilization, service quality, response delay, and resource overhead, providing an objective basis for algorithm evaluation. Through comparative analysis, we can deeply understand the advantages and disadvantages of each algorithm and provide support for designing more efficient and reliable automatic scaling solutions.

[0075] 2. Results Analysis

[0076] In order to verify the stability of the proposed PPO-DU algorithm during training, an ablation experiment was conducted on the PPO algorithm for delayed policy update. The changes of Episode and Rewards during training are as follows: Figure 2 As shown, Figure 2The following figure compares the training process of the PPO-DU and PPO algorithms. The PPO-DU algorithm converges around Episode 3500, while the standard PPO algorithm converges around Episode 4000, demonstrating a clear advantage in convergence speed for the PPO-DU algorithm. The PPO algorithm, which does not employ a delayed update mechanism, exhibits significant fluctuations in its reward curve. This instability affects the learning efficiency of the PPO algorithm's strategy, resulting in slower convergence than the PPO-DU algorithm. In contrast, the PPO-DU algorithm effectively reduces this fluctuation through its delayed update mechanism, making the strategy learning process smoother and more stable, significantly accelerating convergence and demonstrating the positive role of this mechanism in improving training efficiency.

[0077] Under the same burst load intensity and the same random seed conditions, the experiment selected a representative microservice Frontend for comparative analysis. The specific results are shown in Figure 3 As shown in Table 4, Figure 3 is a schematic diagram of the Sock-Shop burst load, and Table 4 is the Sock-Shop burst load.

[0078] Table 4. Burst load in Sock-Shop

[0079]

[0080] In bursty load scenarios, the reinforcement learning-based PPO and PPO-DU algorithms excelled in request completion time, achieving shorter average response times. This was achieved by including service-level objective (SLO) latency as the optimization target in the reward function, enabling the algorithms to learn strategies that minimize request latency. Furthermore, these two algorithms exhibited minimal fluctuations in CPU and memory utilization, ensuring rapid request processing and response, improving workflow completion rates, request success rates, and SLO completion rates. The performance of the ADRL-based algorithm fell between that of the HPA and PPO / PPO-DU algorithms, demonstrating the advantages of deep reinforcement learning in the field of automatic scaling decisions, but still lags behind the PPO-DU algorithm.

[0081] In contrast, Kubernetes' native HPA algorithm has a fixed scaling interval and lags under sudden loads, resulting in large fluctuations in CPU and memory utilization, increased request completion time, decreased request success rate, and an increase in SLO-violating requests, ultimately affecting workflow completion rate and SLO completion rate.

[0082] In the case of cyclic loads of the same intensity, such as Figure 4The periodic load in Sock-Shop shown in Figure 5 and the periodic load in Sock-Shop shown in Table 5, the PPO and PPO-DU algorithms based on reinforcement learning perform well in terms of request completion time, and service requests can be responded to and processed more quickly.

[0083] Table 5 Periodic load in Sock-Shop

[0084]

[0085] This advantage stems from their ability to learn and capture cyclical load patterns, pre-allocating resources appropriately and avoiding request delays. Because cyclical loads cause periodic fluctuations in request frequency, the CPU and memory utilization of the four algorithms also exhibit cyclical variations. PPO and PPO-DU exhibit minimal utilization fluctuations, maintaining high levels and demonstrating excellent autoscaling capabilities. However, cyclical loads have longer peaks and require frequent request scheduling and resource allocation. The algorithms lag slightly behind in request success rates, workflow completion rates, and SLO completion rates in bursty load scenarios, though PPO and PPO-DU maintain their lead. In contrast, the Kubernetes-native HPA algorithm performs sluggishly under cyclical loads and fails to effectively capture cyclical patterns, resulting in large fluctuations in CPU and memory utilization and inefficient resource utilization. This directly leads to lower request success rates, workflow completion rates, and SLO completion rates, putting it behind reinforcement learning algorithms. Algorithms based on ADRL perform moderately, with metrics falling between PPO, PPO-DU, and HPA, but their performance is significantly lower than in bursty load scenarios. This instability stems from the fact that ADRL triggers are based on a fixed number of exceptions, L.

[0086] In the case of high-volatility loads of the same intensity, the volatile load in Sock-Shop is Figure 5 and Table 6. Because the generation time and amount of fluctuating loads are random, the experimental results show that the reinforcement learning-based PPO and PPO-DU algorithms perform best in terms of request completion time and SLO completion rate, with minimal fluctuations in CPU and memory utilization, indicating a more balanced container load distribution. In contrast, Kubernetes' HPA algorithm, due to its fixed scaling decision interval, experiences response delays under bursty loads, leading to large fluctuations in resource utilization and request completion time. The ADRL algorithm performed in the middle, outperforming HPA but not PPO and PPO-DU. In practical applications, a unified threshold L is difficult to account for all load patterns. Determining the optimal L requires extensive experience, increasing management complexity.

[0087] Table 6 Fluctuating load in Sock-Shop

[0088]

[0089] Finally, to measure cost and efficiency, we need to fully compare the performance of these four algorithms. Table 7 shows the algorithm cost of Sock-Shop, which represents the algorithm cost consumption in the fluctuating load of the Sock-Shop microservice system. It mainly includes three indicators: the average number of Pods, the time required for a single scaling, and memory consumption.

[0090] Table 7. Algorithm cost in Bookinfo

[0091]

[0092] The data in the table shows that the PPO-DU and PPO algorithms use more Pods, ensuring lower SLOs while maintaining stable CPU and memory utilization. HPA has advantages in single-time scaling time and memory consumption, requiring less time and memory. Reinforcement learning algorithms consume more time and memory during single-time scaling operations due to the need for online inference of complex neural network models. However, PPO and PPO-DU optimize operations through parallel inference and model compression, maintaining manageable time and memory consumption.

[0093] The above results show that PPO and PPO-DU achieve the best performance in terms of request completion time, SLO completion rate, workflow completion rate, and request success rate under bursty, periodic, and highly volatile load scenarios. Their advantage lies in their ability to learn and capture load patterns (such as periodic loads), pre-allocate resources, avoid request delays, and maintain stable CPU and memory utilization. ADRL, on the other hand, performs moderately well, outperforming HPA but not as well as PPO and PPO-DU. Its performance degrades significantly under periodic loads, particularly due to its trigger being based on a fixed number of exceptions, L, resulting in poor adaptability. HPA performs poorly under bursty and periodic loads, with response lags caused by fixed scaling intervals, large fluctuations in resource utilization, and low request success rates, workflow completion rates, and SLO completion rates.

[0094] From an efficiency perspective, the PPO and PPO-DU algorithms exhibited minimal fluctuations in CPU and memory utilization, maintaining high levels and demonstrating excellent auto-scaling capabilities. Despite high single-shot scaling time and memory consumption, these were kept manageable through parallelized inference and model compression optimization. The ADRL algorithm performed reasonably well under bursty loads, but its efficiency declined under periodic and highly volatile loads, resulting in higher time and memory consumption. While the HPA algorithm achieved the best single-shot scaling time and memory consumption, its overall efficiency was inferior to that of reinforcement learning algorithms due to its low resource utilization.

[0095] From an adaptability perspective, the PPO and PPO-DU algorithms adapt to various load patterns and perform well in complex microservices scenarios, meeting the needs of real-world scenarios. The ADRL algorithm performs adequately only under bursty loads, but exhibits poor adaptability and unstable performance under periodic and highly volatile loads. The HPA algorithm also exhibits poor adaptability, performing particularly sluggishly under periodic and highly volatile loads and failing to effectively cope with load fluctuations.

[0096] In summary, the PPO and PPO-DU algorithms perform best in terms of resource utilization efficiency and service quality. Despite their higher time and memory costs, their overall performance is superior to HPA and ADRL, making them suitable for a variety of complex load scenarios. The results show that the proposed method can achieve better microservice automatic scaling.

[0097] Example 2

[0098] A computer device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0099] Example 3

[0100] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above method.

[0101] Example 4

[0102] A computer program product includes a computer program or instructions, which implement the steps of the above method when the program or instructions are executed by a processor.

[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A microservice automatic scaling method based on the PPO-DU algorithm, characterized in that: include: S1: Define a microservice cluster and convert the microservice autoscaling goal into a target optimization problem. S2: To solve the target optimization problem in S1, a Markov decision process is constructed and the reward function is designed using resource utilization and service quality. S3: Based on the Markov decision process, design an Actor-Critic neural network, which includes state input, Actor network, Critic network, and output; S4: Combined with the Actor-Critic neural network, a delayed update strategy is introduced to design the PPO-DU algorithm, which is used to adjust the number of containers according to the status of the microservice output, thereby achieving automatic scaling of the microservice.

2. The microservice automatic scaling method based on the PPO-DU algorithm according to claim 1, characterized in that: In S1, the microservice cluster MS and microservice msi are: MS=[ms1,ms2,…,ms n } msi={cntainer1,container2,…,container m },i∈[1,n] Among them, the microservice cluster MS consists of multiple microservice ms i Composition, n is the total number of microservices, each microservice ms i It is composed of multiple microservice instances, m is the microservice ms i Number of containers; Each container consists of multiple states. Including resources CPU, memory RAM, storage ROM, these states reflect the resource usage of the container. The formula is: These containers are deployed on multiple physical hosts using the following formula: Host={host1,host2,…,host l } Where l is the total number of physical hosts; Multiple physical hosts constitute the entire host cluster. The resource status of these physical hosts includes CPU, memory RAM, and storage ROM. The formula is: Host resource status Determines the maximum number of containers that can be allocated to the host; Autoscaling adjusts the number of containers to meet load demands, prioritizing service-level objectives (SLOs): SLO={RT,AVA,HC} Where RT is the request response time, AVA is the target percentage of availability, and HC is the target value of throughput; The SLO is determined by the request response time (RT), while minimizing the total resource usage of the microservice. Therefore, the microservice autoscaling goal is transformed into the following target optimization problem: in, Indicates that when the container request response time is lower than the target response time SLO RT 1 when the service is in the state of interest, otherwise 0. This method maximizes the total number of containers that meet the target response time and tries to ensure that the service meets the SLO target. Indicates the total amount of container resources. Minimizes container resource usage while ensuring that the resource usage of all containers in the cluster is less than the resource usage of all physical hosts.

3. The microservice automatic scaling method based on the PPO-DU algorithm according to claim 2, characterized in that: In S2, resource usage is determined by CPU and memory utilization, and service quality is determined by service response latency. Therefore, the state space is composed of CPU utilization, memory utilization, and service response latency. To meet user service quality and guarantee the SLO, the service response latency target is set to 500ms. The SLO is defined as: latency SLO =500ms Among them, latency SLO Indicates service response delay, latency real Indicates the actual service response delay; So at time t the state space s t Defined as: s t =(You CPU ,YOU mem (SLO) Among them, U CPU Indicates CPU utilization, U mem Indicates memory utilization; The action space is the probability distribution of container scaling ratios in autoscaling. In horizontal autoscaling, actions include reducing or doubling the number of microservice instances. Therefore, the action space a t Defined as: a t ∈(0.5,2) The reward function is defined as the reward or penalty obtained by the agent when it reaches a certain state, which is composed of CPU utilization, memory utilization and service response delay. At the same time, in order to reduce the resource consumption of the container scaling process and prevent the agent from frequently scaling automatically, a t It is added to the reward function, so the reward function r t for: r t =2×SLO+U CPU +U mem +(1-a t ) Since the goal is to prioritize SLO delay and reduce resource utilization on this basis, in order to reduce the impact of instance scaling on resource consumption, it is necessary to give a hyperparameter ξ to (1-a t ) to adjust the parameters, the reward function obtained is:

4. The microservice automatic scaling method based on the PPO-DU algorithm according to claim 1, characterized in that: In S3, vector s t Represents the state of the microservice as the input of the Actor-Critic network; the Actor network consists of three fully connected layers, each of which is connected to a Tanh activation function; the first fully connected layer maps the state input to a 64-dimensional hidden representation, the second fully connected layer further extracts high-level features, and the third fully connected layer maps it to a scalar value, representing the action a t Finally, the output is scaled to the range of [0, 1] through the Sigmoid activation function and linearly transformed to the range of [0.5, 2], indicating the scaling ratio of the number of containers. Similar to the Actor network, the Critic network is also composed of three fully connected layers, each of which is followed by a Tanh activation function. The first two fully connected layers have the same function as the Actor network, which is used to extract high-level features of the state. The third fully connected layer maps the 64-dimensional hidden representation to a scalar value, which represents the estimated value of the state. Finally, the Actor network outputs the action, which is the scaling ratio of the number of containers, in the range of [0.5, 2]. A value less than 1 means reducing the number of containers, and a value greater than 1 means increasing the number of containers. The Critic network is the state value, which is used to evaluate the quality of the current state. The optimal action a is obtained through cyclic training using the loss function L(θ, φ). t .

5. The microservice automatic scaling method based on the PPO-DU algorithm according to claim 4, characterized in that: In S4, the specific training process of the PPO-DU algorithm is divided into three steps: (1) Initialization phase: Initialize the Actor network π θ and Critic Network V φ , used for estimating the policy function and value function respectively; set the hyperparameters required for training, including learning rate α, discount factor γ, clip range ∈, update frequency K, maximum training episodes M and maximum steps N of each episode; (2) Training phase: A: For each episode, the environment needs to be reset to obtain the initial state s0; B: For each step t, according to the current strategy π θ (a t |s t ) to select action a t , and calculate the logarithmic probability logπ of the corresponding action θ (a t |s t ); C: Execute action a t , get reward r t , next state s t+1 And whether it is done, and (s t ,a t ,r t ,s t+1 ,done,logπ θ (a t |s t )) is stored in the buffer for subsequent training; D: If the current stept is a multiple of the update frequency K, then the following update steps are required: For each sample i in the buffer, calculate the advantage function This represents the advantage of the current policy relative to the value function, which is calculated by calculating δ starting from the current step i j The weighted sum of δ is used to estimate j Represents the TD residual, i.e. r j +γV φ (s j+1 )-V φ (s j ); E: Calculate target value It is state s i The true value of is used to update the value function network. The calculation formula is the advantage function plus the current value function estimate; then calculate the ratio r i (θ), which represents the probability ratio of the new and old strategies, used to measure the magnitude of strategy update; clipped surrogate objectiveL CLIP (θ) is the objective function of PPO, which achieves stable strategy update by limiting the change range of ratio. The specific calculation is to take the minimum value of the product of ratio and clipped ratio and advantage function, and average it over all samples; F: Calculate the value function loss L VF (φ), which represents the mean squared error loss of the value function network and is used to update the critic network; calculate the total loss L(θ,φ), which is the weighted sum of the policy objective function and the value function loss as the final optimization goal; G: Use the Adam optimizer to update the parameters θ and φ of the Actor network and the Critic network to minimize the total loss L(θ,φ), and clear the data in the buffer to prepare for the next update; (3) Training ends: When all episodes are finished, the training is completed and the trained Actor network π is obtained. θ , the network can adjust the ratio of the number of output containers according to the status of microservices, thereby achieving automatic scaling of microservices.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the microservice automatic scaling method described in any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that A computer program is stored, and when the program is executed by a processor, the steps of the microservice automatic scaling method according to any one of claims 1 to 5 are implemented.

8. A computer program product, characterized in that The method comprises a computer program or an instruction, which, when executed by a processor, implements the steps of the microservice automatic scaling method according to any one of claims 1 to 5.

Citation Information

Cited By

  • Micro-service elastic telescopic arrangement method and arrangement device

    CN121078133A