Efficient Resource Allocation Method and System for Serverless Based on Reinforcement Learning
The reinforcement learning-based resource allocation method enhances serverless computing by optimizing resource utilization and latency through a state, policy, and reward module, along with a warm-up container management system, achieving improved efficiency and predictable delays.
Patent Information
- Application Number
- JP2024515078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-03-23
- Filing Date
- 2023-08-10
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2043-08-10
AI Technical Summary
Current serverless computing platforms face challenges in predicting and managing resource allocation efficiently, leading to high latency and low resource utilization due to the reliance on machine learning models that struggle with balancing resource efficiency and tail latency, and the overhead increases with frequent control adjustments.
An efficient resource allocation method using reinforcement learning to determine resource allocation for each request, incorporating a state module, policy module, and reward module, along with a warm-up type container management system to manage containers and reduce overhead.
The method improves resource utilization by 30% and ensures a 99% request delay target, reducing end-to-end delay variance by 1.9 times while avoiding resource and time overhead.
Smart Images

Figure 0007698797000004 
Figure 0007698797000005 
Figure 0007698797000006
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cloud computing, and particularly relates to an efficient resource allocation method and system for serverless based on reinforcement learning.
Background Art
[0002] Serverless computing has the characteristics of high scalability, easy development, small granularity and low cost, so it has become the main method of the current microservices architecture. It is supported by major cloud providers such as Amazon and has spread rapidly. It is widely used in multiple user-oriented application scenarios such as web applications, video processing, and inference by machine learning. In order to meet the high scalability and high flexibility needs of these applications, these complex application services are disassembled into a set of serverless functions to form a directed acyclic graph.
[0003] In many cases, these user-oriented applications all require strict latency requirements. However, due to the influence of various factors, relatively large tail latency occurs in the applications, making it extremely difficult to predict the performance of these applications. Current commercial platforms such as Amazon's serverless platform Lambda, or open source platforms such as Open Whisk, do not provide any guarantee for the response latency of applications and rely on developers to determine the resource allocation of functions. Therefore, in these platforms, developers have to select relatively large resource allocations (such as internal memory, CPU, etc.) to ensure the SLO (Service Level Objective) of the application, which leads to the problem of low resource utilization.
[0004] To reduce the uncertainty of application latency, the platform needs to be managed at one control layer. Conventional operations are mainly divided into active operations triggered by timing or pre-determined placement and passive operations triggered by thresholds. The main solution methods for both types of operations rely on machine learning models to allocate resources. However, based on conventional operations, in 1), it is difficult to balance resource efficiency and tail latency for passive control. Efficient resource utilization sacrifices high tail latency, and vice versa. In 2), active control depends on the accuracy of the machine learning model. When the prediction error is relatively low, the resource efficiency of active operations is higher than that of passive operations. In 3), the resource efficiency of the serverless system can be effectively improved by high-frequency control. In 4), it is further known that resource and time overhead increase as the control frequency increases, which may offset the advantages of high frequency. Moreover, these phenomena become more obvious and serious in function workflow tasks.
[0005] Also, it is discovered that the graph itself of the directed acyclic graph (DAG) consisting of function workflows is a Markov process. At each step of state transition in different function stages, the function may select one of multiple types of placements, and the transition probability is determined by the selected placement. Different from conventional supervised learning or unsupervised learning, there is no so-called right or wrong in this process, and there is no need to accurately correct non-optimal solutions. This is a reinforcement learning process that emphasizes how to obtain the maximized expected benefit by acting according to the environment. The focus is on balancing exploration and exploitation and emphasizing the "exploration-exploitation" trade-off during learning, which is consistent with the goal of exploring efficient resource allocation technologies for serverless.
Summary of the Invention
Problems to be Solved by the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides an efficient resource allocation method and system for serverless based on reinforcement learning. The method makes a clear and targeted determination and control of resource allocation for each request through reinforcement learning, effectively improving the resource utilization rate and ensuring a 99% request time delay target. At the same time, the present invention effectively avoids the impact of resource and time overhead on the system through pipeline determination and container management mechanisms.
Means for Solving the Problems
[0007] The present invention is realized by the following technical solutions.
[0008] An efficient resource allocation method for serverless based on reinforcement learning, including the following steps (1) to (4): In step (1), a reinforcement learning decision-making device is constructed. The reinforcement learning decision-making device makes predictions through a reinforcement learning model, and the reinforcement learning model includes a state module, a policy module, an action module, and a reward module. In step (2), the decision-making is pipelined. The reinforcement learning decision-making device constructed in step (1) determines the resource allocation of functions at each stage in a pipeline, and in the process of each decision-making, the reinforcement learning decision-making device always uses the remaining time until the target time as an input to derive the resource allocation of the next function and makes predictions using the recorded maximum execution time of the currently executing function. In step (3), the containers are managed. When a function resource allocation that does not match through reinforcement learning is determined in step (2), the function execution instance is transferred to the target container in the scheduling process and executed by a warm-up type container management system. The warm-up type container management system includes a prediction module responsible for predicting the future request arrival rate, a proxy module responsible for container management of nodes, and a transfer module responsible for executing high-speed scheduling requests. In the step (4), when each request arrives, the resource allocation for each stage is sequentially performed using the method in the step (2). After the resource allocation for the corresponding stage is obtained each time, the transfer module in the step (3) is used to schedule the request to the corresponding allocation container to execute the calculation.
[0009] Furthermore, in the step (1), the state module is mainly composed of three types: application state, request state, and cluster state. The dimension compression is performed on the application DAG information by using a graph neural network. The application state is used to describe the situation of the workflow application, including the structure of the workflow directed acyclic graph DAG, the average execution time of each function, and the average resource utilization rate of each function, that is, the average resource utilization rate of the CPU and internal memory obtained by the offline analysis of the function. The request state is used to describe the access situation to the load, including the number of requests per second QPS (Query Per Second) obtained by the request monitor, that is, the number of requests per second, the remaining time before reaching the target delay SLO, and the number of unexecuted functions in the workflow. The cluster state is used to describe the situation of the physical resources, including the available CPU and internal memory, obtained by the cluster monitor.
[0010] Furthermore, the state module performing dimensionality reduction on the application DAG information by using a graph neural network specifically uses the GraphSAGE method in the GCN graph convolutional neural network. The method adopts a node embedding method to extract high-dimensional information in the vicinity of the node graph into a high-density vector embedding, quickly generates the inductive ability of a new figure embedding, and conforms to the resource allocation needs for various workflow application programs. After the application state information is described and obtained, the system transmits the application state information from the tail node of its DAG graph and transmits the information of each node to the parent node and the root node in a recursive manner.
[0011] Furthermore, in step (1), the policy module calculates state information based on the Actor-Critic algorithm and the advantage function. The Actor is a policy network. The Actor selects a fully connected neural network and outputs one value for each action, converts the value into a corresponding probability by the SoftMax function, and selects an action according to the probability, undertakes the selected action and interacts with the environment. Critic is an evaluation network, that is, the advantage function A(s t ,a t )=Q(s t ,a t )-V(s t ) is used as the evaluation network Critic, and its Q(s t ,a t ) and V(s t ) respectively represent the trajectory of cumulative experience and the simulator approximation of average experience, and are used to score the actions of the Actor. The Actor further adjusts its own parameters according to the scoring situation of the Critic. The action module executes a separate policy network for various types of function resources and determines the corresponding resource allocation amount by each network.
[0012] Specifically, in the step (1), the reward module uses the resource allocation amount and the request end-to-end execution time as the reward value to train the accuracy of the decision module, constructs the following function to give rewards, and its expression is shown in Equation 1 below:
Number
[0013] Specifically, the pipeline decision in the step (2) is, when executing the first stage, the decision device simultaneously determines the resource allocation of the second stage, and determines the resource allocation of the third stage when executing the second stage, and by analogy from this.
[0014] Furthermore, the warm-up type container management system in the step (3) is specifically the prediction module reads the historical QPS data and predicts the future maximum arrival rate by means of an exponentially weighted average model, and adopts the action probability table divided by QPS in the training stage to multiply the maximum arrival rate and the action probability one by one to obtain the number of container instances in a future predetermined time window; after obtaining the number of container instances with different resource allocation numbers, the prediction module allocates the resource allocation to different cluster machines, and receives the corresponding information in the proxy module, that is, receives the command. After the proxy module receives the command, it first judges the total number of cluster containers. If the number is insufficient, it cold-starts the container, and when the container becomes idle, it modifies the allocation by Cgroup, and sends the IP and port of the instance of the final allocation to the transfer module; The transfer module maintains queues of a plurality of different resource allocation container IPs. After the function and corresponding allocation executed by the reinforcement learning decision-making device are determined, the transfer module obtains one executable idle container IP by taking it out from the queue, and transfers scheduling information to schedule the instance. After the execution of the container is completed, the relevant information is put into the queue again, thereby reusing the container to save the overhead of cold start.
[0015] Specifically, the action process of the warm-up type container management system in step (3) is to perform the allocation of the container by using the reinforcement learning action probability table and the predicted load.
[0016] Furthermore, in step (4), the system is responsible for managing the complete process of each request including decision-making, scheduling, and triggering operation of the function in the next stage.
[0017] A serverless efficient resource allocation system based on reinforcement learning, including a gateway module, a controller module, a proxy module, and a simulator module. One set of interfaces is added to the original system gateway for the gateway module, which is used to transfer requests to the transfer module. And the system gateway continuously monitors the request arrival rate by Prometheus and periodically sends statistical data to the request volume prediction module for decision-making. The controller module is the main module, including a pipeline prediction, a reinforcement learning decision-making device, a transfer module, and a request volume prediction module. The pipeline prediction manages the execution of an application program in a workflow manner by using the signal volume and coroutines in a programming language. The reinforcement learning decision-making device integrates a trained model into the system and observes the application program, requests, and cluster state through a container monitoring program, and transmits the state to the reinforcement learning decision-making device. The transfer module maintains one hash queue in the internal memory, and the hash queue stores IP information corresponding to different arrangements for quickly searching and transmitting requests. The request volume prediction module periodically determines the number of containers in the next stage based on the historical data transmitted through the gateway. The proxy module is a module executed in the daemon set of the container management platform. The proxy module receives a call by publishing a hypertext transfer protocol service port in the cluster and then executes a resource adjustment operation. The resource in the state is adjusted by the group control method of the operating system. Compared with modifying the relevant arrangement of the placement module of the container management platform, the group control method does not require restarting the container. The resource in the state is the number of processor allocations and the internal memory capacity. The simulator module is a serverless simulator, which is used to train a reinforcement learning policy network. It is an engine driven by one continuous time and discrete events. It simulates the complete process of a workflow request, accelerates the training process by maintaining one logical clock, constructs a network structure with 32 neurons and 16 neurons respectively in the model, and is executed by a machine learning platform.
Advantages of the Invention
[0018] The beneficial effects of the present invention are as follows.
[0019] To improve the resource utilization efficiency of serverless, passive and active dynamic adjustment policies have been proposed in conventional operations. However, the passive method only responds when a violation problem occurs, and it is extremely difficult to make a decision when taking the balance between tail latency and resource waste. The active prediction-based dynamic decision depends on the prediction accuracy of the machine learning model and is less effective for scenes with significant changes. To compensate for the deficiencies of the above two methods, the present invention provides a request-level granularity function workflow resource management system based on reinforcement learning, and makes resource allocation decisions for each request, thereby reducing the problems of resource utilization rate being low due to sudden situations and coarse-grained management.
[0020] Compared with conventional operations, the present invention provides a resource allocation policy for request-level granularity management and designs a reinforcement learning algorithm to determine the resource allocation for each step of the serverless function workflow according to the different arrival states of each request. At the same time, to avoid the problem of overly high time overhead caused by the decision, the present invention provides a pipeline-type prediction method, which simultaneously executes the prediction process and the function execution process, and reduces the linearly increasing prediction time overhead to a single decision time overhead. To solve the problem of cold start due to inconsistent resource allocation, the present invention provides a warm-up type container management plan, and prepares the necessary container instances in advance according to the action probability table of the decision-making device.
[0021] Compared with the prior art, the present invention has remarkable effects. Compared with the conventional state-of-the-art serverless resource management system, the present invention reduces the resource usage by 30%, provides a 99% request delay SLO guarantee, and reduces the end-to-end delay variance by 1.9 times.
Brief Description of the Drawings
[0022]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Embodiments for Carrying Out the Invention
[0023] To make it easier to understand the method technology, achieved objectives, and effects described in the present invention, the present invention will be described below by way of specific embodiments with reference to the drawings.
[0024] The present invention uses machine learning technology, namely reinforcement learning, to enable a proxy to learn through repeated trial and error in an interactive environment and obtain feedback on its actions and experiences, and aims to make the proxy learn a policy that can obtain the maximum reward. In serverless resource allocation, the proxy refers to the controller and the environment is the serverless cluster. The purpose is to find a resource allocation policy for determining the resource allocation of the next instance in the controller based on the current application program, requests, and cluster state. To realize the application in the request-level granularity function workflow scenario of reinforcement learning, the present invention first constructs a reinforcement learning model including four modules: state, policy, action, and reward.
[0025] A method for efficient serverless resource allocation based on reinforcement learning, the method comprising: Constructing a reinforcement learning decision device, wherein the reinforcement learning decision device makes predictions by means of a reinforcement learning model, and the reinforcement learning model includes a state module, a policy module, an action module, and a reward module in step (1); To pipeline the decisions, the reinforcement learning decision-making device constructed in step (1) determines the resource allocation of functions at each stage through a pipeline, and in the process of each decision, the reinforcement learning decision-making device uses the remaining time until the target time as input to derive the resource allocation of the next function, and makes a prediction using the recorded maximum execution time of the currently executing function in step (2). To manage the containers, when a function resource allocation that does not match through reinforcement learning is determined in step (2), a warm-up type container management system transfers the function execution instance to the target container in the scheduling process for execution. The warm-up type container management system includes a prediction module responsible for predicting the future request arrival rate, a proxy module responsible for container management of nodes, and a transfer module responsible for executing high-speed scheduling requests in step (3). When each request arrives, the resource allocation at each stage is sequentially performed using the method in step (2). After the resource allocation at the corresponding stage is obtained each time, the request is quickly scheduled to the corresponding allocated container using the transfer module in step (3) to execute the calculation in step (4).
[0026] Specifically, it is as follows.
[0027] 1. To build a reinforcement learning model. (1.1) State module The state module describes the specific situations of the environment and the proxy. The present invention defines the state in serverless resource allocation by three types of information. The first type of application state is for describing the situation of the workflow application, including the structure of the workflow directed acyclic graph DAG obtained by offline analysis of functions, the average execution time of each function, and the average resource utilization rate of each function, that is, the average resource utilization rate of CPU and internal memory. The second type of request state is for describing the access situation to the load, including the number of requests per second QPS (Query Per Second) obtained by the request monitor, that is, the number of requests per second, the remaining time before reaching the target delay SLO, and the number of unexecuted functions in the workflow. The third type of cluster state is for describing the situation of physical resources including available CPU and internal memory obtained by the cluster monitor.
[0028] The DAG structure of the serverless workflow may change significantly, and the size of the state space may be different. Therefore, how to match it with the fixed size of the input layer of the RL network becomes a challenge. The present invention uses the GraphSAGE (Graph SAmple and aggreGatE) method in the GCN (Graph Convolutional Network) graph convolutional neural network. This method adopts a node embedding method to extract high-dimensional information in the node graph neighborhood into a high-density vector embedding, quickly generates the inductive ability of a new graph embedding, and can adapt to the resource allocation needs of various workflow application programs.
[0029] After the application state information is described and obtained, the system transmits the application state information from the tail node of its DAG graph and transmits the information of each node to the parent node and the root node in a recursive manner. In this process, the present invention strengthens the expression effect of application features by a non-linear function.
[0030] (1.2) Policy The Actor-Critic algorithm is used as the policy network algorithm of the present invention. The Actor is a policy network, which is responsible for the selected actions and interacts with the environment. The Critic is an evaluation network, which is used to score the actions of the Actor. Then, the Actor further adjusts its parameters within a certain range according to the scoring situation of the Critic.
[0031] The present invention selects a fully connected neural network as the Actor of the present invention, outputs one value for each action, converts the value into a corresponding probability by the SoftMax function, and selects an action according to the probability, so as to effectively avoid the local optimal dilemma caused by taking the highest value. At the same time, the present invention uses the advantage function A(s t ,a t )=Q(s t ,a t )-V(s t ) as the evaluation network Critic of the present invention. The Q(s t ,a t ) and V(s t ) respectively represent the trajectory of cumulative experience and the simulator approximation of average experience. Compared with the method of directly calculating by numerical values, such a method has a smaller variance and a faster convergence.
[0032] (1.3) Action Train an independent decision network for each type of resource, and then determine the corresponding resource allocation amount by each network. While ensuring the decision accuracy, the training speed is significantly improved.
[0033] (1.4) Reward Since the controller adjusts the resource allocation based on the request and each function, an independent request is regarded as one complete training stage. The present invention constructs the following function to give rewards, and its expression is shown in the following formula 1.
Equation
[0034] 2. Pipeline Module In the case of one application workflow, when a request arrives, the reinforcement learning decision-making device determines the resource allocation of the first-stage function. When executing the first stage, the decision-making device simultaneously determines the resource allocation of the second stage. Similarly, when executing two stages, it determines the resource allocation of the third stage, and so on by analogy. In the case of a non-linear and complex DAG graph structure, the decision-making device identifies the last-started parent function in each function, predicts simultaneously when executing this parent function, and reduces the linearly increasing time overhead to a single inference time by means of pipeline decision-making.
[0035] In each decision-making process, the reinforcement learning decision-making device uses the remaining time until the target time as input to derive the resource allocation of the next function. However, when making decisions simultaneously, it cannot grasp the remaining time when executing the next stage. In contrast, the present invention makes predictions using the recorded maximum execution time of the currently executed function.
[0036] 3. Container Management It is mainly divided into a prediction module responsible for predicting the future request arrival rate, a proxy module responsible for container management of nodes, and a relay module responsible for executing high-speed scheduling requests.
[0037] (3.1) Prediction Module The prediction module reads historical QPS data and makes predictions using the exponentially weighted moving average model (EWMA), and simultaneously improves the obtained results by an offset amount of 20%. After obtaining the future maximum arrival rate, the present invention adopts an action probability table divided by QPS in the training stage, multiplies the maximum arrival rate and the action probability one by one to obtain the number of container instances in a future predetermined time window. At the same time, the present invention continuously observes the container execution waiting situation, and when the waiting time of a certain deployed container exceeds the cold start time, cold starts one new instance.
[0038] (3.2) Proxy Module After obtaining the number of container instances with different resource allocation numbers, the prediction module assigns the resource allocation greedily to different cluster machines and receives corresponding information in the proxy module. After receiving a command, the proxy module first determines the total number of cluster containers. If the number is insufficient, it cold starts the container, and when the container becomes idle, modifies the allocation by Cgroup. Furthermore, it sends the IP and port of the instance of the final allocation to the transfer module.
[0039] (3.3) Transfer Module The transfer module maintains a queue of multiple different resource allocation container IPs. After the function and corresponding allocation executed by the reinforcement learning decision device are determined, the transfer module obtains one executable idle container IP by taking it out from the queue, and transfers the scheduling information to schedule the instance. After the execution of the container is completed, the relevant information is put into the queue again, thereby reusing the container to save the overhead of cold start.
[0040] Regarding the system architecture of the present invention, the system structure diagram of the present invention is shown in FIG. 1. The gist of the present invention is that both resource efficiency and predictable delay can be simultaneously achieved by a request granularity resource allocation policy based on reinforcement learning. The transfer, expansion, and resource allocation of requests are managed by the cooperation of controllers arranged in physical servers. In the controller of the present invention, there are two core components. The reinforcement learning predictor determines the resource allocation of each instance, and pipelining hides the time overhead caused by each decision-making process, thus avoiding the SLO violation due to the cumulative effect. Since serverless computing generally uses "mapping from one request to one state" in the processing process, the resource allocation of each request is realized by calling the reinforcement learning decision-making device for each request and each function. Even if multiple requests share the same instance (i.e., reuse the existing instance for a new request), the resource allocation of the instance is adjusted by each request repeatedly calling the decision-making device.
[0041] The method steps of the present invention are specifically as follows: when a user makes a request, it first reaches the gateway. The gateway receives the request and sends it to the transfer module to search for available instances. The workload (i.e., QPS) is continuously monitored by the prediction module, and the number of requests within the next time interval is predicted by the exponentially weighted moving average (EWMA) algorithm. If the number of instances for processing requests within the next time interval is insufficient, the controller notifies its proxy module to split the instance into one or more servers by the greedy method to expand more instances and adjust the resource allocation of the instances. For each request of each workflow application program, the controller organizes resource allocation decisions and function calculations in the pipelining module, and sends the currently monitored state information based on the workflow structure to the reinforcement learning decision-making device, thereby obtaining the allocation information that can satisfy the SLO and has the least resource waste. For the determined function, the transfer module is responsible for searching for a container instance that meets the conditions and scheduling the request to the corresponding instance. After the execution of the function is completed, the controller observes whether the execution of the current request is completed. If there are remaining unexecuted requests, the above steps are repeated until the execution of the final function of the function workflow is completed.
[0042] Regarding the execution of the reinforcement learning model of the present invention, Figure 2 shows the execution architecture of the reinforcement learning model. After the system collects the corresponding states of requests, applications, and clusters in the serverless cluster environment, it sends them to the GCN model for state compression to convert two-dimensional information into the content of one-dimensional information. After processing the state data, it is sent to the policy network for calculation to obtain the corresponding actions. The actions are returned to the serverless function cluster, and the management of the request is realized by using the value as the resource allocation at the request stage. The above is a complete process of resource allocation decision-making.
[0043] Regarding the execution of the reinforcement learning model of the present invention, Figure 3 shows the execution architecture of the reinforcement learning model. The system monitors the serverless cluster for its status, collects the corresponding status of requests, applications, and the cluster in the cluster environment, and then sends it to a graph neural network for state compression. The graph neural network model converts the application pipeline from two-dimensional information to one-dimensional information content by means of a recursive transmission method. After processing the state data, it is sent to a policy network for calculation to obtain the corresponding actions. The actions are returned to the serverless function cluster, and the controller realizes the management of requests by using this value as the resource allocation at the request stage. The above is a complete process of resource allocation decision-making once.
[0044] Regarding the execution of the system of the present invention, the present invention is realized in an open-source serverless platform OpenFaaS, which is an open-source serverless platform based on Go and Kubernetes. As shown in Figure 4, the present invention reuses the request method provided by OpenFaaS, mainly modifies the modules related to scheduling requests and resource allocation, and the modules include the OpenFaaS gateway and the OpenFaaS processor. The present invention further adds two new modules, namely a proxy and a simulator, thereby arranging and quickly strengthening the learning training.
[0045] Regarding the gateway module, the present invention adds a set of interfaces for transferring requests from the original OpenFaaS gateway to the transfer module. Moreover, the OpenFaaS gateway continuously monitors the request arrival rate (QPS, query per second) by a monitoring program (Prometheus) and periodically sends the statistical data to the QPS prediction module for decision-making.
[0046] Regarding the controller module, it includes a pipeline prediction, a reinforcement learning decision-making device, a transfer module, and a QPS prediction module. The pipeline prediction manages the execution of application programs in a workflow manner by using semaphores and Goroutines in the Go language. The reinforcement learning decision-making device integrates a trained model by TensorFlow Go into OpenFaaS. At the same time, the present invention observes the application program, requests, and cluster status through two types of container monitoring programs, Prometheus and Cadvisor, and transmits the status to the reinforcement learning decision-making device. The decision-making device converts and combines the data by a graph neural network and inputs it into the corresponding reinforcement learning model for decision-making. The transfer module maintains a hash queue in the internal memory, and the hash queue stores IP information corresponding to different arrangements for quickly searching and transmitting requests. The QPS prediction module periodically determines the number of containers in the next stage based on the historical data transmitted through the OpenFaaS gateway.
[0047] Regarding the proxy module, the present invention executes this module in the DaemonSet of Kubernetes. After receiving a call by exposing the HTTP service port in the cluster, it executes a resource adjustment operation. The resources of the status (such as the number of CPU allocations, internal memory, etc.) are adjusted by Cgroup. Compared with modifying the related configuration of the Deployment module of Kubernetes, Cgroup does not require restarting the container.
[0048] Regarding the simulator, the present invention designs a serverless simulator, thereby training the reinforcement learning policy network more efficiently. By using an engine driven by one continuous time and discrete events to simulate the complete process of the workflow request and maintaining one logical clock, the training process is accelerated. It takes 4 hours to train one model. The present invention constructs the network structure with 32 neurons and 16 neurons respectively in the model and is realized by tensorflow. In the execution process, the present invention initializes the parameters by the glorot method.
[0049] As shown in Figure 5, the present invention further provides an efficient resource allocation system for serverless based on reinforcement learning. The system A set of interfaces is added to the original system gateway for transferring requests to the transfer module. And the system gateway continuously monitors the request arrival rate by a set of methods and periodically sends statistical data to the request volume prediction module for decision-making. The gateway module The main module includes a pipeline prediction, a reinforcement learning decision-making device, a transfer module and a request volume prediction module. The pipeline prediction manages the execution of the application program in a workflow manner by using semaphores and coroutines in the programming language. The reinforcement learning decision-making device integrates the trained model into the system and observes the application program, requests and cluster state by the container monitoring program and transmits the state to the reinforcement learning decision-making device. The transfer module maintains a hash queue in the internal memory, and the hash queue stores IP information corresponding to different arrangements for quickly searching and sending requests. The request prediction module is a controller module that periodically determines the number of containers in the next stage according to the historical data transmitted through the gateway. A module executed in the daemon set of a container management platform, the module receives a call by publishing a Hypertext Transfer Protocol service port in a cluster and then executes a resource adjustment operation. Since the resources in the state are adjusted by the group control method of the operating system, compared with modifying the related deployment of the deployment module of the container management platform, the group control method does not require restarting the container, and the resources in the said state are the number of processor allocations, the proxy module which is the internal memory, and A serverless simulator for training a reinforcement learning policy network, an engine driven by one continuous time and discrete events, simulating the complete process of a workflow request, accelerating the training process by maintaining one logical clock, and in the model, the network structure is composed of 32 neurons and 16 neurons respectively, and a simulator module realized by the most commonly used machine learning platform.
[0050] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art can make various changes and modifications to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should all be included within the protection scope of the present invention.
Claims
1. An efficient resource allocation method for serverless based on reinforcement learning, comprising the following steps (1) to (4), In step (1), it is to construct a reinforcement learning decision-making device, wherein the reinforcement learning decision-making device makes predictions by a reinforcement learning model, and the reinforcement learning model includes a state module, a policy module, an action module, and a reward module, In step (2), it is to pipeline the decision-making, wherein the resource allocation of functions at each stage is determined in a pipeline by the reinforcement learning decision-making device constructed in step (1), and in the process of each decision-making, the reinforcement learning decision-making device always uses the remaining time until the target time as an input to derive the resource allocation of the next function, and makes a prediction using the recorded maximum execution time of the currently executing function, In step (3), it is to manage containers. When a function resource allocation that does not match by reinforcement learning is determined in step (2), a function execution instance is transferred to a target container in the scheduling process by a warm-up type container management system for execution. The warm-up type container management system includes a prediction module responsible for predicting the future request arrival rate, a proxy module responsible for container management of nodes, and a transfer module responsible for executing high-speed scheduling requests, In step (4), when each request arrives, the resource allocation at each stage is sequentially performed using the method in step (2). After the resource allocation at the corresponding stage is obtained each time, the transfer module in step (3) is used to schedule the request to the corresponding allocated container to execute the calculation An efficient resource allocation method for serverless based on reinforcement learning, characterized by the above.
2. In step (1), the state module is mainly composed of three types: application state, request state, and cluster state, and uses a graph neural network to perform dimensionality compression on application DAG information, The application state is used to describe the situation of a workflow application, including the structure of a workflow directed acyclic graph (DAG), the average execution time of each function, and the average resource utilization rate of each function, i.e., the average resource utilization rates of the CPU and internal memory obtained by offline analysis of the function. The request state is used to describe the access situation to the load, including the number of requests per second QPS (Query Per Second) obtained by a request monitor, i.e., the number of requests per second, the remaining time before reaching the target delay SLO, and the number of unexecuted functions in the workflow. The cluster state is used to describe the situation of physical resources including available CPU and internal memory obtained by a cluster monitor. The method for efficient resource allocation of a serverless based on reinforcement learning according to claim 1, characterized in that.
3. When the state module performs dimensionality compression on application DAG information using a graph neural network, specifically, the GraphSAGE method in the GCN graph convolutional neural network is used. The method adopts a node embedding method to extract high-dimensional information near the node graph into a high-density vector embedding, quickly generating the inductive ability of a new figure embedding to conform to the resource allocation needs of various workflow application programs. After the application state information is described and obtained, the system transmits the application state information from the tail node of its DAG graph and transmits the information of each node to the parent node and root node in a recursive manner. The method for efficient resource allocation of a serverless based on reinforcement learning according to claim 2, characterized in that.
4. In step (1), the policy module calculates state information based on the Actor-Critic algorithm and the advantage function. The Actor is a policy network. The Actor selects a fully connected neural network to output one value for each action, converts the value into a corresponding probability by the SoftMax function, and selects an action according to the probability, undertakes the selected action and interacts with the environment. Critic is an evaluation network, that is, the advantage function A(s t , a t ) = Q(s t , a t ) - V(s t ) is used as the evaluation network Critic, and its Q(s t , a t ) and V(s t ) respectively represent the trajectory of cumulative experience and the simulator approximation of average experience, and are used to score the actions of Actor. Actor further adjusts its own parameters according to the scoring situation of Critic, The action module executes a separate policy network for various types of function resources and determines the corresponding resource allocation amount by each network. The method for efficiently allocating serverless resources based on reinforcement learning according to claim 1, characterized in that.
5. In step (1), the reward module uses the resource allocation amount and the request end-to-end execution time as reward values to train the accuracy of the decision module, constructs the following function to give rewards, and its expression is shown in Equation 1 below. 【Number 1】 However, R represents the number of resources allocated to the function, W represents the actual number of wasted resources, and t elapsed represents the time elapsed since the request arrived, and t slo represents the target delay set for the application, n represents the number of remaining sub-functions of the function, and θ and μ are constant parameters used to control their relationship The method for efficiently allocating serverless resources based on reinforcement learning according to claim 1, characterized in that.
6. The pipeline determination in step (2) is, when executing the first stage, the decision device simultaneously determines the resource allocation in the second stage, and determines the resource allocation in the third stage when executing the second stage, and by analogy. The method for efficiently allocating serverless resources based on reinforcement learning according to claim 1, characterized in that.
7. The warm-up type container management system in step (3) specifically. The prediction module reads the historical QPS data and predicts the future maximum reach rate by an exponentially weighted average model, and in the training stage, adopts an action probability table divided by QPS to multiply the maximum reach rate and the action probability one by one to obtain the number of container instances in a future predetermined time window. After obtaining the number of container instances with different resource allocation numbers, the prediction module allocates the resource allocation to different cluster machines and receives the corresponding information by the proxy module, that is, receives a command. After the proxy module receives the command, it first determines the total number of cluster containers. If the number is insufficient, it cold-starts the container, and when the container becomes idle, it modifies the allocation by Cgroup and sends the IP and port of the instance of the final allocation to the transfer module. The transfer module maintains queues of multiple different resource allocation container IPs. After the function and corresponding allocation executed by the reinforcement learning decision-making device are determined, the transfer module obtains one executable idle container IP by taking it out from the queue, and transfers the scheduling information to schedule the instance. After the execution of the container is completed, the relevant information is put back into the queue again, thereby reusing the container to save the overhead of cold start. The serverless efficient resource allocation method based on reinforcement learning according to claim 1, characterized in that.
8. The action process of the warm-up type container management system in the step (3) is to arrange the containers by using the reinforcement learning action probability table and the predicted load. The serverless efficient resource allocation method based on reinforcement learning according to claim 1, characterized in that.
9. In the step (4), the system is responsible for the management of the complete process of each request including decision-making, scheduling, and triggering operation of the function in the next stage. The serverless efficient resource allocation method based on reinforcement learning according to claim 1, characterized in that.
10. A serverless efficient resource allocation system based on reinforcement learning, including a gateway module, a controller module, a proxy module, and a simulator module. One set of interfaces is added to the original system gateway for the gateway module, which is used to transfer requests to the transfer module, and the system gateway continuously monitors the request arrival rate by Prometheus and periodically sends statistical data to the request volume prediction module for decision-making. The controller module is the main module, which includes a pipeline prediction, a reinforcement learning decision-making device, a transfer module, and a request volume prediction module. The pipeline prediction manages the execution of an application program in a workflow manner by using the signal volume and coroutines in a programming language. The reinforcement learning decision-making device integrates a trained model into the system and observes the application program, requests, and cluster state through a container monitoring program, and transmits the state to the reinforcement learning decision-making device. The transfer module maintains one hash queue in the internal memory, and the hash queue stores IP information corresponding to different arrangements for quickly searching and transmitting requests. The request volume prediction module periodically determines the number of containers in the next stage based on the historical data transmitted through the gateway. The proxy module is a module executed in the daemon set of the container management platform. The proxy module receives a call by publishing a hypertext transfer protocol service port in the cluster and then executes a resource adjustment operation. The resource in the state is adjusted by the group control method of the operating system. Compared with modifying the relevant arrangement of the placement module of the container management platform, the group control method does not require restarting the container. The resource in the state is the number of processor allocations and the internal memory capacity. The simulator module is a serverless simulator, which is used to train a reinforcement learning policy network. It is an engine driven by one continuous time and discrete events. It simulates the complete process of a workflow request and accelerates the training process by maintaining one logical clock. The network structure in the model is composed of 32 neurons and 16 neurons respectively, and is executed by a machine learning platform. A serverless efficient resource allocation system based on reinforcement learning, characterized by the above.
Citation Information
Patent Citations
Dynamic task placement method based on delay and cost balance in server-free computing
CN113176947A
Scheduling method and device based on micro-service link analysis and reinforcement learning
CN114780233A
Container cluster online deployment method fusing graph neural network and reinforcement learning in edge computing
CN115686846A
Jitter-less distributed function-as-a-service using flavor clustering
US20210021485A1
Network Node and Methods in a Communications Network
US20220255814A1