Server-free computing load automatic scaling method and system based on reinforcement learning

By adopting a reinforcement learning method with mixed reward values and dynamic window gradient mechanism in serverless computing system, the elastic scaling parameters are dynamically adjusted, which solves the flexibility and stability problems of traditional methods in the face of sudden load changes, and achieves rapid and robust automatic scaling of the load, enhancing the robustness and efficiency of the system.

CN120276867AActive Publication Date: 2025-07-08NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510759039.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In the face of sudden workload changes, the traditional static threshold-driven elastic scaling algorithm is difficult to respond flexibly, resulting in waste of resources or degradation of service quality, and the existing reinforcement learning methods are not effective when adapting to new workload modes.

Method used

Using reinforcement learning-based method, the hybrid reward value of the system's health utilization rate and CPU usage rate is calculated, combined with anomaly detection and dynamic window gradient mechanism, elastic scaling parameters are dynamically adjusted to achieve automatic scaling of serverless computing load.

Benefits of technology

The rapid and robust online parameter adjustment of serverless computing load is realized, which enhances the robustness and adaptability of the system, reduces resource waste and delays, and improves the flexibility and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276867A_ABST
    Figure CN120276867A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for automatically scaling a server-free computing load based on reinforcement learning, and the method comprises the steps: employing a reinforcement learning algorithm to modify a specified elastic scaling parameter as an action, and calculating a reward value after the action is executed to achieve the automatic scaling of the server-free computing load. And calculating an average mixed reward value as a reward value by adopting the following steps: acquiring a system health utilization rate and a CPU utilization rate after an action is executed; calculating a mixed reward value formed by mixing the two; detecting an abnormal value of the mixed reward value by adopting a specified abnormal detection algorithm; removing abnormal values from the time sequence of the mixed reward values to obtain a filtered time sequence of the mixed reward values; and calculating an average value of the mixed reward value time sequence to obtain an average mixed reward value used as a reward value. The invention aims to realize quick and steady online parameter adjustment of server-free computing workload, and further realize automatic scaling of the load scale.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and particularly to a method and system for automatically scaling serverless computing loads based on reinforcement learning. Background Art

[0002] With the rapid development of cloud computing technology, the serverless computing architecture, as a new computing paradigm, has attracted extensive attention in the academic and industrial fields. Serverless computing exhibits unique advantages in fields such as artificial intelligence, data processing, and the Internet of Things, especially showing remarkable flexibility and efficiency in dynamic resource scheduling and automatic scaling. Currently, the function as a service (FaaS) is the main application mode of serverless computing. Compared with traditional architectures based on virtual machines or containers, the workload fluctuations of serverless computing applications are more obvious and unpredictable, especially in high-concurrency request and uncertain task execution time patterns. For example, a sudden surge in user requests may cause the execution frequency of serverless computing functions to increase or decrease significantly within a short period of time, resulting in drastic fluctuations in resource requirements. In addition, as the most mature container orchestration engine currently deployed in data centers, Kubernetes adopts a static threshold-driven elastic scaling algorithm, which can automatically increase or decrease the number of Pods hosting applications according to the current resource utilization rate and predefined thresholds. Load scheduling ensures the balanced distribution of newly created Pods within the cluster, and these two processes work together to improve the ability to quickly respond to sudden workload demands. However, traditional static threshold-driven elastic scaling algorithms often have difficulty flexibly coping with sudden resource changes, resulting in potential resource waste or a decline in service quality. In addition, adjusting system parameters under different workloads based on experience alone easily leads to system instability, increased latency, and higher costs. In recent years, applying reinforcement learning (RL) to perform online tuning for serverless computing systems has gradually become a new research trend. However, due to the need for long-term training, a large amount of training data, and poor performance when adapting to new workload patterns, its applicability is limited. Therefore, designing and implementing a lightweight online parameter threshold adjustment algorithm that can handle highly dynamic workloads in serverless scenarios has become the core focus of current research. Recently, a similar method has been proposed. This method relies on single-point feedback in reinforcement learning and only uses reward information at the parameter points being adjusted, without the need to model complex unknown reward functions. Compared with global modeling techniques, the lightweight nature of this algorithm stems from eliminating the additional calculations for estimating rewards in unexplored parameter states, which reduces the model complexity and makes it particularly suitable for online adjustment during deployment. However, the single reward value it uses provides insufficient information to guide the system for rapid exploration, and the gradient ascent method is sensitive to noise, which makes it prone to falling into local extrema. Summary of the Invention

[0003] Technical problem to be solved by the present invention: In view of the above problems of the prior art, a serverless computing load automatic scaling method and system based on reinforcement learning are provided. The present invention aims to achieve fast and robust online parameter adjustment of serverless computing workloads, and thus achieve automatic scaling of the load scale.

[0004] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A serverless computing load automatic scaling method based on reinforcement learning, which includes using a reinforcement learning algorithm to modify specified elastic scaling parameters as actions, calculating the reward value after executing the actions to achieve automatic scaling of the serverless computing load, and using the following steps to calculate the average mixed reward value as the reward value: S101, obtain the system health utilization rate and CPU utilization rate after executing the action; S102, calculate the mixed reward value composed of the system health utilization rate and the CPU utilization rate; S103, use a specified anomaly detection algorithm to detect the anomaly value of the mixed reward value; S104, remove the anomaly value from the time series of the mixed reward value to obtain the filtered time series of the mixed reward value; S105, calculate the average value of the time series of the mixed reward value to obtain the average mixed reward value used as the reward value.

[0005] Optionally, the specified elastic scaling parameters include some or all of the lower limit of CPU usage cpuLow, the upper limit of CPU usage cpuHigh, the lower limit of memory usage memLow, the upper limit of memory usage memHigh, the expected number of application replicas replica to run, and the CPU threshold containerCpuThreshold of the current container; if the CPU usage is less than the lower limit of CPU usage cpuLow, the number of CPUs is scaled down, if the CPU usage is greater than the upper limit of CPU usage cpuHigh, the number of CPUs is scaled up, if the memory usage is less than the lower limit of memory usage memLow, the amount of memory is scaled down, if the memory usage is greater than the upper limit of memory usage memHigh, the amount of memory is scaled up, if the number of running application replicas is less than the expected number of application replicas replica, the number of running application replicas is increased, if the number of running application replicas is greater than the expected number of application replicas replica, the number of running application replicas is decreased, if the CPU threshold of the container is less than the CPU threshold containerCpuThreshold of the current container, the number of CPUs is scaled down, and if the CPU threshold of the container is greater than the CPU threshold containerCpuThreshold of the current container, the number of CPUs is scaled up.

[0006] Optionally, the calculation function expression of the system health utilization rate in step S101 is: , where, is the system health utilization rate, is the time of the system health state within a given time period , and the time of the system health state is the time when the CPU usage is healthy and the memory usage is healthy, is the logical AND operation, and the healthy CPU usage means that the CPU usage is greater than the lower limit of CPU usage cpuLow and less than the upper limit of CPU usage cpuHigh, and the healthy memory usage means that the memory usage is greater than the lower limit of memory usage memLow and less than the upper limit of memory usage memHigh.

[0007] Optionally, the calculation function expression of the hybrid reward value in step S102 is: , where, is the hybrid reward value, is the weight coefficient, is the system health utilization rate, is the CPU usage.

[0008] Optionally, detecting outliers of the hybrid reward value by using a specified anomaly detection algorithm in step S103 includes: predicting the current hybrid reward value by using an autoregressive moving average model ARMA for the time series of the hybrid reward value, then taking the difference between the predicted hybrid reward value and the calculated hybrid reward value, and determining that the currently calculated hybrid reward value is an outlier if the difference exceeds a preset threshold; otherwise, determining that the currently calculated hybrid reward value is a normal value; or, determining whether the currently calculated hybrid reward value is an outlier according to the historical data of the hybrid reward value in combination with the 3σ principle.

[0009] Optionally, implementing the automatic scaling of the serverless computing load by using a reinforcement learning algorithm to modify specified elastic scaling parameters as actions and calculating the reward value after executing the actions includes: S201, initializing and exploring the radii of each elastic scaling parameter and the values of each elastic scaling parameter; S202, generating a random perturbation vector within the radius and calculating the values of each elastic scaling parameter after random perturbation; S203, executing the action of modifying the elastic scaling parameter to implement the automatic scaling of the serverless computing load; S204, calculating the reward value after executing the action; S205, updating the values of each elastic scaling parameter by using the gradient descent algorithm in combination with the reward value. If iteration is needed, jump to step S202; otherwise, end and exit.

[0010] Optionally, when updating the values of each elastic scaling parameter by using the gradient descent algorithm in combination with the reward value in step S205, it includes calculating the dynamic window gradient as the adjusted gradient for the current gradient of each elastic scaling parameter by using the following formula: , where is the dynamic window gradient, is a dynamically adjustable coefficient, is to take the average value, is the average value of the historical gradients at each time point within the most recent historical window, is the current gradient, and the calculation function expression for the size of the most recent historical window is: , , where is the size of the most recent historical window, , Is a constant used to control the initial value and maximum value of the history window, Is the current reward value, Is the average value of the reward values within the most recent history window, Is a dynamically adjustable coefficient Of the change amount, Is the median of the reward values within the most recent history window, Is a constant, Is to take the maximum value.

[0011] In addition, the present invention also provides a serverless computing load automatic scaling system based on reinforcement learning, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the serverless computing load automatic scaling method based on reinforcement learning.

[0012] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the serverless computing load automatic scaling method based on reinforcement learning through a processor.

[0013] In addition, the present invention also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the serverless computing load automatic scaling method based on reinforcement learning through a processor.

[0014] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: 1. When the present invention adjusts the elastic scaling related parameters through online reinforcement learning, it integrates the average mixed reward (AMR) mechanism. Calculating the average mixed reward value includes the intrinsic learning reward structure of the policy gradient (obtaining the system health utilization rate and CPU usage rate after executing the action and calculating the mixed reward value composed of the two), and the robustness reward structure (using a specified anomaly detection algorithm to detect the outliers of the mixed reward value; removing the outliers from the time series of the mixed reward value to obtain the filtered time series of the mixed reward value; calculating the average value of the time series of the mixed reward value to obtain the average mixed reward value used as the reward value). By using two different reward structures to accelerate convergence, it can achieve fast and robust online parameter adjustment of the serverless computing workload, and then realize the automatic scaling of the load scale, thereby enhancing the robustness of the system.

[0015] 2. The present invention can further integrate the dynamic window gradient (DG) mechanism when adjusting the elastic scaling related parameters through online reinforcement learning. The dynamic window gradient (DG) mechanism is used to smooth the gradient calculation by including historical values, thereby further enhancing the robustness of the system. Description of the Drawings

[0016] Figure 1 This is a schematic flow chart for calculating the average mixed reward value in an embodiment of the present invention.

[0017] Figure 2 This is a schematic flow chart of the reinforcement learning algorithm in an embodiment of the present invention.

[0018] Figure 3 This is a schematic diagram of the system structure in an embodiment of the present invention. Detailed implementation manners

[0019] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0020] The serverless computing load automatic scaling method based on reinforcement learning in this embodiment includes using a reinforcement learning algorithm to modify specified elastic scaling parameters as actions, and calculating the reward value after executing the actions to achieve the automatic scaling of the serverless computing load. As Figure 1 shown, the following steps are used to calculate the average mixed reward value as the reward value: S101, obtain the system health utilization rate and CPU utilization rate after executing the action; S102, calculate the mixed reward value composed of the system health utilization rate and the CPU utilization rate; S103, use a specified anomaly detection algorithm to detect the anomaly value (X abnormal ) of the mixed reward value; S104, remove the anomaly value X abnormal from the time series of the mixed reward value to obtain the filtered time series X of the mixed reward value; S105, calculate the average value of the time series of the mixed reward value to obtain the average mixed reward value used as the reward value.

[0021] Traditional reinforcement learning algorithms often use CPU utilization, memory occupancy rate, or SLO (Service Level Objective) as reward values because these data are easy to obtain in the system. In this embodiment, a reward value that considers both internal rewards and external rewards is designed. The internal reward value is the CPU utilization rate, and the external reward value is the Healthy Utilization Percent (HUP) of the system resources. The external reward value HUP is an indicator for measuring the overall state of the system used by the current existing work, and is defined as the ratio of the time when all nodes are in a healthy state to the total duration of the time period within a given time period. Both the CPU and memory have upper and lower bounds, and only when the utilization rates of both the CPU and memory fall within these bounds, the current time period is considered to be in a healthy state. HUP faces two limitations. One is that it will introduce noise, which will hinder the exploration of the optimal parameters and may lead to suboptimal solutions, especially in dynamic or multi-objective tasks. The other is that it may lead to instability, slow convergence, or failure to converge to the optimal policy. For this reason, this embodiment proposes an Average Mixed Reward (AMR) as shown in Figure 1 which integrates the intrinsic learning reward structure with policy gradients (LIRPG) and a robustness reward structure, thus overcoming these two limitations. The reason for selecting the CPU utilization rate as the internal reward is that compared with increasing memory, increasing the CPU can shorten the scheduling response time more. Using the CPU utilization rate as the internal reward aims to encourage optimizing the health of the system resources while optimizing resource utilization as much as possible.

[0022] To evaluate the health of the system across time dimensions, the elastic scaling parameters specified in this embodiment include the lower limit of CPU usage cpuLow, the upper limit of CPU usage cpuHigh, the lower limit of memory usage memLow, the upper limit of memory usage memHigh, the number of application replicas replica expected to run, and the CPU threshold containerCpuThreshold of the current container (some of them can also be selected according to the scenario needs); the above 6 elastic scaling parameters can establish the elastic boundary of the system, and by adjusting the above 6 parameters (i.e., k = 6) to implement a dynamic threshold-driven elastic scaling policy. In each iteration, the cluster manager receives a k dimensional parameter vector a(t) constituted by the above 6 elastic scaling parameters, and then uses it to set the system parameters.

[0023] The above six elastic scaling parameters can establish the elastic boundary of the system. If this threshold is exceeded, the number of CPUs and the memory will scale automatically, specifically including: (1) If the CPU usage rate is less than the lower limit of CPU usage rate cpuLow, the number of CPUs will be shrunk; if the CPU usage rate is greater than the upper limit of CPU usage rate cpuHigh, the number of CPUs will be expanded. (2) If the memory usage rate is less than the lower limit of memory usage rate memLow, the amount of memory will be shrunk; if the memory usage rate is greater than the upper limit of memory usage rate memHigh, the amount of memory will be expanded. (3) If the number of running application replicas is less than the expected number of running application replicas replica, the number of running application replicas will be increased; if the number of running application replicas is greater than the expected number of running application replicas replica, the number of running application replicas will be decreased. (4) If the CPU threshold of the container is less than the CPU threshold of the current container containerCpuThreshold, the number of CPUs will be shrunk; if the CPU threshold of the container is greater than the CPU threshold of the current container containerCpuThreshold, the number of CPUs will be expanded.

[0024] The system health utilization rate is defined as the ratio of the time when all nodes are in a healthy state within a given time period to the total duration of this time period. Both CPU and memory have upper and lower bound values, and only when the utilization rates of both CPU and memory fall within these bounds, the current time period is considered to be in a healthy state. The calculation function expression of the system health utilization rate in this embodiment is: , where, is the system health utilization rate, is the given time period the time of the system in a healthy state within which is the time when the CPU usage rate is healthy and the memory usage rate is healthy, is a logical AND operation, and the healthy CPU usage rate means that the CPU usage rate is greater than the lower limit of CPU usage rate cpuLow and less than the upper limit of CPU usage rate cpuHigh, and the healthy memory usage rate means that the memory usage rate is greater than the lower limit of memory usage rate memLow and less than the upper limit of memory usage rate memHigh.

[0025] The combination of internal rewards and external rewards forms a mixed reward value (MixReward), which is defined as the weighted sum of CPU and HUP, where CPU represents the utilization rate of CPU in the cluster. The above combination of the two accelerates the optimization of system parameters. The calculation function expression of the mixed reward value in step S102 of this embodiment is: , Among them, is the mixed reward value, is the weight coefficient, is the system health utilization rate, is the CPU usage rate.

[0026] To minimize the impact of noise on the reward value as much as possible, this embodiment introduces a robust reward structure to avoid overly rapid exploration of the parameter space. For most mixed rewards, historical average window backpropagation is used, which can smooth historical data and reduce the impact of noise. In addition, to avoid missing key data caused by over-smoothing, anomaly detection is employed in this embodiment to filter out mixed reward values important to the system. As an alternative implementation, the anomaly detection algorithm adopted in step S103 to detect outliers of the mixed reward value includes: using the autoregressive moving average model ARMA to predict the current mixed reward value from the time series of the mixed reward value, and then taking the difference between the predicted mixed reward value and the calculated mixed reward value. If the difference exceeds a preset threshold, the currently calculated mixed reward value is determined to be an outlier; otherwise, the currently calculated mixed reward value is determined to be a normal value. The anomaly detection algorithm used in this embodiment is the autoregressive moving average (ARMA) model, which analyzes the time series data of the mixed rewards generated by the workload and predicts the current value. The ARMA model consists of two main parts: autoregressive (AR) and moving average (MA). AR represents the linear relationship between the current value and the observed values at several previous time points. For example, if the current value is predicted from the observed values at the past 4 time points, a fourth-order autoregressive model is formed; while MA represents the linear combination of the current value and the previous error terms, assuming that the current value is affected not only by past observed values but also by previous error terms. MA modifies the prediction result through the residual terms at several previous moments to model the random volatility. As another alternative implementation, it is possible to determine whether the currently calculated mixed reward value is an outlier according to the historical data of the mixed reward value combined with the 3σ principle.

[0027] As Figure 2 shown, in this embodiment, a reinforcement learning algorithm is adopted to modify the specified elastic scaling parameters as actions and calculate the reward value after executing the actions to achieve the automatic scaling of the serverless computing load, including: S201, initialize the radius for exploring each elastic scaling parameter and the value of each elastic scaling parameter; S202, generate a random perturbation vector within the radius and calculate the values of each elastic scaling parameter after the random perturbation; S203, execute the action of modifying the elastic scaling parameters to achieve the automatic scaling of the serverless computing load; S204. Calculate the reward value after the execution of the action; S205. Combine the reward value and use the gradient descent algorithm to update the values of each elastic scaling parameter. If iteration is needed, jump to step S202; otherwise, end and exit.

[0028] In the process of parameter optimization of the reinforcement learning algorithm, traditional methods often directly rely on gradients and reward values to update parameters. This approach is prone to causing the algorithm to fall into local extrema, thus affecting the overall performance of the model. To solve this problem, this embodiment adopts the dynamic window gradient method (Dynamic Window Gradient, DG) based on historical gradient information, comprehensively considers historical reward values and changes, and dynamically adjusts the window size to improve the scheduling efficiency. As Figure 2 shown, When updating the values of each elastic scaling parameter by using the gradient descent algorithm in combination with the reward value in step S205 of this embodiment, it includes the current gradient of each elastic scaling parameter Calculate the dynamic window gradient as the adjusted gradient by using the following formula: , where, is the dynamic window gradient, is the dynamically adjustable coefficient, is to take the average value, is the average value of the historical gradients at each time point within the most recent historical window, is the current gradient. The calculation function expression for the size of the most recent historical window is: , , where, is the size of the most recent historical window, , are constants used to control the initial value and maximum value of the historical window, is the current reward value, is the average value of the reward values within the most recent historical window, is the dynamically adjustable coefficient change amount of, is the median of the reward values within the most recent historical window, is a constant, It is to take the maximum value. This dynamic adjustment mechanism enables the algorithm to flexibly adapt to different environmental conditions. Especially in rapidly changing tasks, it can effectively collect more relevant data, thereby improving the learning effect. After receiving the changed values of six elastic scaling parameters, namely cpuLow, cpuHigh, memLow, memHigh, replica, and containerCpuThreshold, the cluster manager updates the corresponding parameter configurations in the system and calls the existing functions of the cloud platform to achieve dynamic adjustment of the load scale. At the same time, it calculates the reward value based on the new system state after scheduling execution and sends it to the server.

[0029] Figure 3 It is the system structure diagram in this embodiment. The system consists of three parts: a serverless cluster manager, a server, and a client API. The serverless cluster manager receives the adjustment parameters selected by the server and completes the work of elastic scaling for the serverless workload. The cluster manager passes the elastic scaling parameters that currently affect the load scale to the server by calling the client API (as shown in ① in Figure 3 ). The server creates a service instance, and an online reinforcement learning algorithm integrating average mixed reward (AMR) and dynamic window gradient (DG) runs in the service instance. Then the cluster manager calls this service instance to conduct a round of exploration in the parameter space to predict the next set of optimal parameters (as shown in ② in Figure 3 ). Next, the cluster manager receives the predicted parameters through the client and updates the current system parameters (as shown in ③ in Figure 3 ). Subsequently, the cluster manager uses the received parameters to perform elastic expansion and workload scheduling, and calculates a new reward value based on the newly collected cluster state and sends it to the service instance (as shown in ④ in Figure 3 ).

[0030] The advantages of the serverless computing load automatic scaling method based on reinforcement learning in this embodiment are as follows in three aspects: (1) In terms of system architecture, a lightweight and online reinforcement learning method is adopted to achieve the challenge of serverless computing automatic scaling parameters. Compared with the deep reinforcement learning method that needs to explore a huge state space and thus requires high training costs, this embodiment can reduce the cost requirements brought by a large amount of training and enhance the adaptability to various complex business scenarios at the same time. (2) In the reward value calculation stage, the concept of reward is extended, and a mixed reward value is adopted, enabling the system to adjust its behavior based on internal attributes and receive external attributes from the environment at the same time. The external attributes are defined as rewards representing the overall system goal, while the internal attributes are defined as rewards that guide the system to optimize in a specific direction and are also related to the overall system goal, thus accelerating the convergence speed of reinforcement learning. (3) In the parameter update stage, the reward value, the current gradient, and the historical gradient are comprehensively considered to determine the amplitude of parameter update, thereby generating new parameter values. By integrating historical gradient information, this method can smooth the current gradient value and enhance the robustness of parameter update, effectively reducing the convergence difficulties caused by local extrema. To determine the degree of dependence on the historical gradient, a historical dynamic window is constructed to calculate the coefficient. In addition, the window size is dynamically adjusted by calculating the current change values of the reward value and the coefficient, and the window size in turn affects the size of the coefficient. The above process enhances the robustness of the method and enables efficient parameter update even under complex workload patterns. To sum up, the method in this embodiment adjusts the relevant parameters of elastic scaling through online reinforcement learning and integrates the average mixed reward (AMR) and dynamic window gradient (DG) mechanisms. The average mixed reward value (AMR) mechanism accelerates convergence by using two different reward structures, while the dynamic window gradient (DG) mechanism smooths the gradient calculation by including historical values, thereby enhancing the robustness of the system.

[0031] In addition, this embodiment also provides a serverless computing load automatic scaling system based on reinforcement learning, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the serverless computing load automatic scaling method based on reinforcement learning.

[0032] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the serverless computing load automatic scaling method based on reinforcement learning through a processor.

[0033] In addition, this embodiment also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the serverless computing load automatic scaling method based on reinforcement learning through a processor.

[0034] Those skilled in the art should understand that the technical solutions provided by the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of processes and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0035] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A serverless computing load automatic scaling method based on reinforcement learning, characterized in that, It includes using a reinforcement learning algorithm to modify specified elastic scaling parameters as actions, calculating the reward value after executing the actions to achieve automatic scaling of serverless computing loads, and using the following steps to calculate the average mixed reward value as the reward value: S101. Obtain the system health utilization rate and CPU utilization rate after executing the actions; S102. Calculate the mixed reward value composed of the system health utilization rate and the CPU utilization rate; S103. Use a specified anomaly detection algorithm to detect the outliers of the mixed reward value; S104. Remove the outliers from the time series of the mixed reward value to obtain the filtered time series of the mixed reward value; S105. Calculate the average value of the time series of the mixed reward value to obtain the average mixed reward value used as the reward value.

2. The serverless computing load automatic scaling method based on reinforcement learning according to claim 1, wherein The specified elastic scaling parameters include some or all of the lower limit of CPU utilization rate cpuLow, the upper limit of CPU utilization rate cpuHigh, the lower limit of memory utilization rate memLow, the upper limit of memory utilization rate memHigh, the expected number of application replicas to run replica, and the CPU threshold of the current container containerCpuThreshold; if the CPU utilization rate is less than the lower limit of CPU utilization rate cpuLow, then perform the contraction of the number of CPUs; if the CPU utilization rate is greater than the upper limit of CPU utilization rate cpuHigh, then perform the expansion of the number of CPUs; if the memory utilization rate is less than the lower limit of memory utilization rate memLow, then perform the contraction of the amount of memory; if the memory utilization rate is greater than the upper limit of memory utilization rate memHigh, then perform the expansion of the amount of memory; if the number of running application replicas is less than the expected number of application replicas to run replica, then increase the number of running application replicas; if the number of running application replicas is greater than the expected number of application replicas to run replica, then reduce the number of running application replicas; if the CPU threshold of the container is less than the CPU threshold of the current container containerCpuThreshold, then perform the contraction of the number of CPUs; if the CPU threshold of the container is greater than the CPU threshold of the current container containerCpuThreshold, then perform the expansion of the number of CPUs.

3. The method for automatically scaling the serverless computing load based on reinforcement learning according to claim 2, wherein The calculation function expression of the system health utilization rate in step S101 is: , wherein, is the system health utilization rate, is the given time period for the time of the system health state within which, the time of the system health state is the time when the CPU utilization rate is healthy and the memory utilization rate is healthy, is the AND logical operation, and the CPU utilization rate being healthy means that the CPU utilization rate is greater than the CPU utilization rate lower limit cpuLow and less than the CPU utilization rate upper limit cpuHigh, and the memory utilization rate being healthy means that the memory utilization rate is greater than the memory utilization rate lower limit memLow and less than the memory utilization rate upper limit memHigh.

4. The method for automatically scaling the serverless computing load based on reinforcement learning according to claim 2, wherein The calculation function expression of the mixed reward value in step S102 is: , Among them, is the mixed reward value, is the weight coefficient, is the system health utilization rate, is the CPU usage rate.

5. The method for automatically scaling the serverless computing load based on reinforcement learning according to claim 2, wherein The detection of the outliers of the mixed reward value by using the specified anomaly detection algorithm in step S103 includes: using the autoregressive moving average model ARMA to predict the current mixed reward value from the time series of the mixed reward value, and then taking the difference between the predicted mixed reward value and the calculated mixed reward value. If the difference exceeds the preset threshold, it is determined that the currently calculated mixed reward value is an outlier; otherwise, it is determined that the currently calculated mixed reward value is a normal value; or, determining whether the currently calculated mixed reward value is an outlier according to the historical data of the mixed reward value combined with the 3σ principle.

6. The method for automatically scaling the serverless computing load based on reinforcement learning according to claim 1, wherein The implementation of the automatic scaling of the serverless computing load by using a reinforcement learning algorithm to modify specified elastic scaling parameters as actions and calculating the reward value after executing the actions includes: S201, Initialize and explore the radii of each elastic scaling parameter and the values of each elastic scaling parameter; S202, generate a random perturbation vector within the radius and calculate the values of each elastic scaling parameter after the random perturbation; S203, implementing the automatic scaling of the serverless computing load by executing the action of modifying the elastic scaling parameters; S204, calculating the reward value after executing the action; S205, updating the values of the respective elastic scaling parameters by using a gradient descent algorithm in combination with the reward value. If further iteration is required, jump to step S202; otherwise, end and exit.

7. The method for automatically scaling the serverless computing load based on reinforcement learning according to claim 6, wherein When updating the values of the respective elastic scaling parameters using the gradient descent algorithm in combination with the reward value in step S205, it includes the current gradient for each elastic scaling parameter Use the following formula to calculate the dynamic window gradient as the adjusted gradient: , Among them, is the dynamic window gradient, is a dynamically adjustable coefficient, is to take the average value, is the average value of the historical gradients at each time point within the most recent historical window, is the current gradient, and the calculation function expression for the size of the most recent historical window is: , , Among them, is the size of the most recent historical window, , are constants used to control the initial value and the maximum value of the historical window, is the current reward value, is the average value of the reward values within the most recent historical window, is a dynamically adjustable coefficient is the change amount of is the median of the reward values within the most recent historical window, is a constant, is to take the maximum value.

8. A serverless computing load automatic scaling system based on reinforcement learning, including a microprocessor and a memory connected to each other, characterized in that, The microprocessor is programmed or configured to execute the method for automatically scaling a serverless computing load based on reinforcement learning according to any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the method for automatically scaling a serverless computing load based on reinforcement learning according to any one of claims 1 to 7 through a processor.

10. A computer program product comprising a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the method for automatically scaling a serverless computing load based on reinforcement learning according to any one of claims 1 to 7 through a processor.

Citation Information

Patent Citations

  • Elastic scaling method and device oriented to container cloud environment and electronic equipment

    CN118860572A

  • Resource aware scheduling in a distributed computing environment

    US20130104140A1

  • Enhanced Reinforcement Learning Algorithms Using Future State Prediction

    US20230177117A1

  • Cloud network abnormality detection model training method based on reinforcement learning, and storage medium

    WO2024104401A1

Cited By

  • Pulse power supply system control method and system based on safety deep reinforcement learning

    CN120896465A