A Containerized Microservice Orchestration System and Method Based on Deep Reinforcement Learning
By introducing delay reward strategy and DRM-DQL algorithm in containerized microservice orchestration, the problem of delay reward in traditional reinforcement learning is solved, and efficient microservice orchestration and faster convergence speed are achieved.
Patent Information
- Application Number
- CN202310387279.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-04-11
AI Technical Summary
Traditional reinforcement learning has delay reward problems in containerized microservice orchestration, which leads to the inability to update and converge in time.
A containerized microservice orchestration system based on deep reinforcement learning is proposed, using delay reward strategy and DRM-DQL algorithm. By introducing temporary experience cache and global experience cache, the problem of delay reward matching is solved, and the epsilon-greedy strategy is used for action selection.
It effectively solves the limitations of traditional reinforcement learning in highly dynamic microservice orchestration, improves the orchestration efficiency and effect of containerized microservices, and achieves faster convergence speed and lower application completion time.
Smart Images

Figure CN116582407B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud native, and particularly to a containerized microservice orchestration system and method based on deep reinforcement learning. Background Art
[0002] With the rapid advent of the era of Internet of Everything and the rapid development of wireless networks, intelligent facilities have been widely popularized. Under the new Internet era, the amount of data has shown an explosive growth and reached the ZB level. The existing centralized system architecture can no longer support the transmission, storage, and calculation of the rapidly generated massive data, and sensitive data is not convenient to be transmitted to the central node due to security issues. Thus, edge computing has emerged as the times require. Edge computing refers to a computing model that computes the downlink data from cloud services and the uplink data from each edge device at the network edge. Its basic idea is to process computing tasks on computing resources close to the data source. Edge computing, with its unique geographical advantages, that is, it can provide services with high bandwidth and low latency and can protect the data security and privacy of users, is regarded as one of the key enabling technologies in fields such as artificial intelligence and 5G.
[0003] Edge computing is to sink some of the computing power of cloud computing to the edge of the network, which is an expansion of cloud computing. Cloud native realizes the rapid on-demand application orchestration and construction based on services through containerization, microservices, and loose coupling, meeting the demand-driven application development mode of rapid iteration, and gradually becoming the mainstream of cloud computing development. And microservice technology is the top priority of cloud native technology. In edge computing, containers have become one of the necessary components of edge computing due to their low overhead, convenient and fast deployment, high isolation degree, etc. At the same time, it is also the best partner of microservices under cloud native. Microservices often exist in the form of containers in edge nodes, that is, containerized microservices. The deployment of containerized microservices is the foundation of cloud native, and efficient deployment of microservices is also important for edge computing. Compared with the traditional monolithic centralized architecture, microservices deploy a large monolithic application by decomposing it into multiple small modules. The advantage of this method is that individual services can be built, tested, and deployed without having an adverse impact on the entire product or other services. Just because of this, currently, the vast majority of applications are split into many microservices and deployed in edge nodes. However, due to the complex dynamic changes of microservices, including the changing application requirements such as time-varying service request quantities and service types, and the changing complex background, as well as the complex dependencies between services, the orchestration of containerized microservices is not easy. In addition, there are call relationships between different microservices, which makes there be a load correlation between microservices, that is, the performance of one microservice has a great relationship with the performance of the microservices it calls, which further increases the weight of scheduling decisions and makes the difficulty of containerized microservice orchestration even greater.
[0004] Currently, all the major existing container orchestration platforms cannot well complete the deployment of containerized microservices in edge nodes. For example, Kubernetes (k8s) is mainly used to manage large-scale container applications in cloud computing, and it has significant limitations for edge nodes with limited computing resources. For example, platforms such as Kubernetes, k3s, and KubeEdge all have some problems in the orchestration of microservices with high dynamism.
[0005] An important reason for the inability of traditional container orchestration platforms to solve the management of containerized microservices in edge computing is the high dynamism of microservices in edge nodes. The high dynamism of service requirements and background loads greatly interferes with the serviceability of containerized microservice applications, resulting in problems such as low resource utilization and cascading failures. To cope with these dynamic events, multiple factors need to be considered comprehensively, such as the available resources in edge nodes, etc., which makes traditional container orchestration technologies unable to well handle the orchestration of microservices in edge computing. Thus, solutions for orchestrating containerized microservices based on reinforcement learning have been proposed. In these solutions, due to the dependency relationships between microservices, the microservices in edge computing are described as a set of Directed Acyclic Graphs (DAGs), and by scheduling these microservices to different devices for execution, the dynamic requirements in the network are thus met. Many pioneering studies have started using reinforcement learning models to complete the orchestration of containerized microservices, but the present invention has found that existing reinforcement learning solutions still have deficiencies.
[0006] When using traditional reinforcement learning for scheduling decisions, due to the execution time of microservices, after the agent in the reinforcement learning model executes an action in the environment, the expected reward cannot be immediately obtained. The present invention refers to this as "delayed reward". It is precisely because of the "delayed reward" problem that the reinforcement learning model cannot be updated in a timely manner because the learning experience of each episode is not completed.
[0007] Therefore, how to find a strategy to solve the phenomenon of delayed reward existing in traditional reinforcement learning training is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0008] To solve the above problems, the present invention provides a containerized microservice orchestration system and method based on deep reinforcement learning. First, a delayed reward strategy is designed to address the problem of delayed rewards in the system model and traditional reinforcement learning, which is used to assist the update of the DeepQ-Network in DQL. Then, the DRM-DQL algorithm is designed in combination with the delayed reward strategy. Finally, a containerized microservice orchestration system is established according to the DRM-DQL algorithm. The containerized microservice orchestration system is constructed based on the proposed "delayed reward strategy" and "DRM-DQL algorithm". Compared with traditional reinforcement learning algorithms, the DRM-DQL algorithm proposed by the present invention has better orchestration effects. First, in order to avoid falling into local optima, the epsilon-greedy strategy is adopted to make corresponding actions. Secondly, the biggest improvement is the combination of the delayed reward strategy. In this strategy, the present invention first represents the basic elements required for reinforcement learning as a tuple <S, A, R, S'>. However, due to the existence of a certain execution time for microservices, the reward R cannot be obtained immediately, which makes traditional reinforcement learning unable to work well.
[0009] The proposed delayed reward strategy well solves the problem of delayed rewards by introducing a "temporary experience cache" and a "global experience cache". First, the triple <S, A, S'> immediately obtained after the agent makes an action is stored in the "temporary experience cache". Then, after obtaining the corresponding delayed reward R, let the delayed reward R match the <S, A, S'> stored in the temporary experience cache, and store the successfully matched quadruple <S, A, R, S'> in the "global experience cache". Then, the complete experience in the global experience cache, that is, the quadruple, is used to train the DeepQ-Network. The quadruple <S, A, R, S'> is the key to training the reinforcement learning model, where S represents the current state, A represents the action made by the current agent, R represents the delayed reward of the current action, and S' represents the next state to be updated after the agent makes the action. In traditional reinforcement learning, there is a problem of delayed rewards, that is, the delayed reward R in the quadruple cannot be obtained in time, which leads to the inability of traditional reinforcement learning methods to converge. The DRM-DQL algorithm is constructed based on the delayed reward strategy and is the core component of the algorithm. The delayed reward strategy ensures the correctness of DQL training.
[0010] The containerized microservice orchestration method system based on the experience matching mechanism includes five modules: a system information acquisition module, a reward generation module, a delayed reward matching module, a reinforcement learning training module, a decision-making module, and two caches: a global experience cache and a temporary experience cache;
[0011] The system information acquisition module is used to obtain the state information before and after the edge node environment after the decision-making module makes corresponding actions, that is, to obtain the current state, the current action, and the next state to be updated. The state information before and after the environment is the key of the reinforcement learning training module. The obtained state information includes the state S of the environment before making the action, the state S' of the environment after making the action, and the current action A transmitted from the decision-making module, which is represented as a triple <S, A, S'> and stored in the temporary experience cache.
[0012] The reward generation module is used to calculate the delayed reward R obtained after the agent makes an action.
[0013] The delayed reward matching module is used to solve the matching problem between the delayed reward R and the triple <S, A, S'> already existing in the temporary experience cache. Here, S represents the current state, A represents the action made by the current agent, and an action means orchestrating a certain microservice to a certain edge node. S' represents the next state to be updated after the agent makes the action. Since the microservice always exists in the edge node during the entire training process, in order to ensure that the delayed reward R can be successfully matched with <S, A, S'>, the information of the currently orchestrated microservice, that is, the unique identifier of the microservice, is used to mark all the results after the agent makes the action. Specifically, when the agent makes an action, it will immediately obtain a triple <S, A, S'>. Then this module will read the information of the currently orchestrated microservice and find its unique identifier, and then use this identifier to mark the immediately obtained triple. After marking, it is stored in the temporary experience cache. After obtaining the delayed reward R, the delayed reward matching module will perform the same operation as above to obtain the unique identifier of the current microservice and use it to mark the obtained delayed reward. Finally, it traverses the content in the temporary experience cache to find the corresponding triple <S, A, S'>. After successful matching, the delayed reward matching module will store the successfully matched quadruple <S, A, R, S'> in the global experience cache for subsequent training.
[0014] The reinforcement learning training module is used to train and update the DeepQ-Network through the complete experience in the global experience cache.
[0015] The decision-making module is used to control the agent to make corresponding actions. After making the action, it describes the action information as a vector and sends this information to the system information acquisition module. In order to avoid falling into local optima and explore more and better actions, the present invention uses the epsilon-greedy strategy to select the actions of the agent.
[0016] The temporary experience cache is used to store the triple <S, A, S'>.
[0017] A global experience cache for storing the quadruple <S, A, R, S'> successfully matched by the delayed reward matching module.
[0018] Further, the system information acquisition module describes the obtained edge node environment state as a vector S = {q1, q2,..., q n ,..., q |N|}, where q n represents the number of microservices already running on the nth node. This can well represent the workload of each edge node and describe the obtained information as a triple <S, A, S'>.
[0019] Further, considering the complex dependencies between microservices, that is, a higher reward should be assigned to the completion time of the microservices on which a microservice depends. Therefore, in the reward generation module, the delayed reward R = -T f (·), where T f (·) represents the completion time of the microservice. This definition method can well achieve the efficient orchestration of microservices. After calculating the corresponding delayed reward R, the delayed reward R is transmitted to the delayed reward matching module for the next calculation.
[0020] Further, in the delayed reward matching module, by assigning the unique identifier of the marked microservice of the microservice being orchestrated to both the triple <S, A, S'> and the delayed reward R, it is ensured that the triple in the temporary experience cache is traversed after obtaining the delayed reward to complete the matching.
[0021] Further, in the decision-making module, the action information taken is represented as a vector A = {a1, a2,..., a n ,..., a |N|}, where a n represents scheduling the microservice to node n and sending the corresponding action information to the system information acquisition module. N represents the total number of nodes.
[0022] A containerized microservice orchestration method based on deep reinforcement learning, comprising:
[0023] S1: Randomly initialize the DeepQ-Network, and then initialize the hyperparameters and the global experience cache. The hyperparameters include α, γ, ε;
[0024] S2: Determine whether there is still an episode not trained. If so, reset the "global experience cache" and the "temporary experience cache", and reset the environment to obtain the initial state. If not, directly end;
[0025] S3: Determine whether there are microservices that have not been orchestrated. If so, use the epsilon-greedy strategy to select the actions of the agent. If not, directly end.
[0026] S4: After performing the action in step S3 in the environment, obtain the next state, and then obtain the triple <S, A, S'>, and store this triple in the temporary experience cache. Then, obtain the completion time and latency reward R of the current microservice.
[0027] S5: Determine whether the local buffer has been traversed completely. If so, directly end. If not, complete the matching of the triple <S, A, S'> and the latency reward R through the consistency of their identifiers. The consistent identifier refers to the identifier of the microservice currently being orchestrated. Then, store the successfully matched quadruple <S, A, R, S'> in the global experience cache, where S represents the current state, A represents the action taken by the current agent, and S' represents the next state to be updated after the agent takes the action.
[0028] S6: Finally, use the quadruples stored in the global experience cache to train the DeepQ-NetWork.
[0029] Furthermore, if the triple <S, A, S'> and the latency reward R do not match successfully, continue to determine whether the local buffer has been traversed completely.
[0030] The beneficial effects brought by the technical solution provided by the present invention are as follows: The DRM-DQL algorithm and the latency reward matching strategy proposed by the present invention well solve the limitations of traditional microservice orchestration platforms for highly dynamic microservice orchestration and can efficiently orchestrate containerized microservices in edge computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0032] Figure 1 is the structural diagram of a containerized microservice orchestration system based on an experience matching mechanism in an embodiment of the present invention.
[0033] Figure 2 is the flowchart of the execution of the DRM-DQL algorithm in an embodiment of the present invention.
[0034] Figure 3 is the comparison graph of the average reward obtained by using the DRM-DQL algorithm and the traditional DQL algorithm in an embodiment of the present invention as the episode increases.
[0035] Figure 4In the embodiments of the present invention, affected by network bandwidth, it is a comparison chart of the completion time finally applied by using the DRM-DQL algorithm and using the traditional DQL algorithm or greedy algorithm under different network bandwidths. Detailed implementation manners
[0036] For a clearer understanding of the technical features, objectives, and effects of the present invention, the detailed implementation manners of the present invention will now be described in detail with reference to the accompanying drawings.
[0037] In view of the problem of delayed rewards existing in the traditional DQL training process, in this embodiment, a new "delayed reward strategy" is designed to enhance DeepQ-Learning (DQL) by obtaining complete experiences from the environment. Then, considering the complex dependencies between microservices, combined with the "delayed reward strategy", an improved delayed reward matching deep Q-learning algorithm based on reinforcement learning (Delayed Reward Matched Deep Q-Learning Algorithm, DRM-DQL) is proposed to achieve efficient orchestration of containerized microservices. That is, the embodiments of the present invention provide a containerized microservice orchestration system and method based on deep reinforcement learning, which can provide efficient orchestration services in high-dynamic containerized microservice scenarios.
[0038] Please refer to Figure 1 , Figure 1 It is a structural diagram of a containerized microservice orchestration system based on deep reinforcement learning with an experience matching mechanism in the embodiments of the present invention, specifically including: an information acquisition module, a reward generation module, a delayed reward matching module, a reinforcement learning training module, a decision-making module, a temporary experience cache and a global experience cache, and two caches, namely the global experience cache and the temporary experience cache.
[0039] The system information acquisition module is used to obtain the state information before and after the edge node environment after the decision-making module makes a corresponding action.
[0040] The reward generation module is used to calculate the delayed reward R obtained after the agent makes an action, and transmit the delayed reward R to the delayed reward matching module.
[0041] The delayed reward matching module is used to solve the matching problem between the delayed reward R and <S, A, S'> existing in the temporary experience cache, where S represents the current state, A represents the action made by the current agent, and S' represents the next state to be updated after the agent makes the action.
[0042] The reinforcement learning training module is used to train the DeepQ-Network (DQN) through the complete experiences in the global experience cache.
[0043] A decision-making module for controlling the agent to perform corresponding actions;
[0044] A temporary experience cache for storing the triple <S, A, S'> immediately obtained after the agent performs an action;
[0045] A global experience cache for storing the quadruple <S, A, R, S'> successfully matched by the delayed reward matching module.
[0046] Compared with the traditional DQL algorithm, the present invention has taken optimization measures. To avoid falling into local optima, the present invention adopts the epsilon-greedy strategy to perform corresponding actions, and adopts the Delayed Reward Matched (DRM) strategy to solve the latency problem of microservices. The basic elements required for reinforcement learning are represented as a quadruple <S, A, R, S'>. However, due to the existence of a certain execution time for microservices, it is impossible to immediately obtain the reward R, which makes the traditional reinforcement learning unable to work well. The biggest improvement is the combination of the delayed reward strategy. Two caches are introduced to solve this problem. First, the tuple <S, A, S'> immediately obtained after performing an action is stored in the temporary experience cache A. Then, after obtaining the corresponding delayed reward R, the <S, A, S'> existing in the temporary experience cache A is matched through the "delayed reward matching module". The core idea of the "delayed reward matching module" is to ensure the completion of the matching by adding <S, A, S'>, the delayed reward R, and the identifier of the microservice being choreographed, and then store the obtained quadruple <S, A, R, S'> in the global experience cache B. Then, the quadruples in the global experience cache B are used to train the DeepQ-Network.
[0047] The flowchart of the execution of the DRM-DQL algorithm is as Figure 2 shown, and specifically includes the following steps:
[0048] S1: First, randomly initialize the DeepQ-Network, and then initialize hyperparameters such as α, γ, ε and the global experience cache B. By reasonably setting hyperparameters such as α, γ, ε, the neural network can converge more efficiently;
[0049] S2: After the initialization phase is completed, start the training process. As Figure 3 shown, in each training set, judge whether there is still an episode that has not been trained. If so, reset the global experience cache B (i.e., Figure 2 the experience buffer B in
[0050] S3: Determine whether there are microservices that have not been orchestrated. If so, use the epsilon-greedy strategy to select the actions of the agent, which can avoid getting stuck in local optima and explore more and better actions. If not, directly end.
[0051] S4: After performing the actions in step S3 in the environment, obtain the next state, and then obtain the triple <S, A, S'>, and store this triple in the temporary experience cache. Then, obtain the completion time and latency reward R of the current microservice. Considering the latency of the reward R, this triple and the reward R should be used to immediately train the DeepQ-Network. Therefore, the present invention will use the delayed reward strategy to generate correct experiences.
[0052] S5: Determine whether the local buffer A has been traversed completely. If so, directly end. If not, complete the matching of the triple <S, A, S'> and the identifier of the delayed reward R. To match the reward R with the corresponding <S, A, S'>, in this embodiment, the "delayed reward matching module" is used to complete the matching of the tuple <S, A, S'> and the reward R. The core idea of the "delayed reward matching module" is to complete the matching through the consistency of <S, A, S'> and the identifier of the delayed reward, and this identifier that ensures consistency is the identifier of the microservice currently being orchestrated. Then, store the successfully matched quadruple <S, A, R, S'> in the global experience cache B, where S represents the current state, A represents the action taken by the current agent, and S' represents the next state to be updated after the agent takes the action.
[0053] S6: Finally, use the quadruples stored in the global experience cache B to train the DeepQ-NetWork.
[0054] As Figure 3 and Figure 4 shown, in this embodiment, the system is deployed in an environment with 20 edge nodes and high-dynamic microservices that need to be orchestrated. From Figure 3 the results, it can be clearly seen that by examining the rewards in more than 6000 training rounds during the training process, the DRM-DQL algorithm has converged after 2100 training rounds, while the traditional DQL has not converged even after 4000 training rounds. The DRM-DQL proposed by the present invention has a faster convergence speed compared with the traditional DQL. From Figure 4 it can be known that under the influence of network bandwidth, the method proposed by the present invention can achieve the minimum application completion time, and the application completion time is reduced by 9.7% and 32.5% respectively compared with using the traditional DQL and the greedy (Greedy) method. Figure 3-4 The DQN in
[0055] The beneficial effects of the present invention are as follows: The DRM-DQL algorithm and the delayed reward matching strategy proposed by the present invention well solve the limitations of traditional microservice orchestration platforms for highly dynamic microservice orchestration and can efficiently orchestrate containerized microservices in edge computing.
[0056] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A containerized microservice orchestration system based on deep reinforcement learning, characterized in that: It includes: A system information acquisition module, a reward generation module, a delayed reward matching module, a reinforcement learning training module and a decision-making module, as well as two caches, namely a global experience cache and a temporary experience cache; The system information acquisition module is used to acquire the state information before and after the edge node environment after the decision-making module makes a corresponding action; Describe the edge node environment state as a vector S = {q1, q2,..., q n ,..., q |N|}, where q n represents the number of microservices already running on the nth node, N represents the total number of nodes, and describe the information obtained as a triple <S, A, S'>, where S represents the current state, A represents the action taken by the current agent, and S' represents the next state to be updated after the agent takes the action; The reward generation module is used to calculate the delayed reward R obtained after the agent makes an action, and transmit the delayed reward R to the delayed reward matching module; The delayed reward matching module is used to solve the matching problem between the delayed reward R and <S, A, S'> that already exists in the temporary experience cache; The delayed reward R = -T f (·), where T f (·) represents the completion time of the microservice; By assigning the unique identifier of the marked microservice of the microservice being orchestrated to both the triple <S, A, S'> and the delayed reward R, it is ensured that the triples in the temporary experience cache are traversed after the delayed reward is obtained to complete the matching; The reinforcement learning training module is used to train the Deep Q-Network with the complete experience in the global experience cache; The decision-making module is used to control the agent to make corresponding actions; The action information to be taken is represented as a vector A = {a1, a2,..., a n ,..., a |N|}, where a n represents scheduling the microservice to node n and sending the corresponding action information to the system information acquisition module, and N represents the total number of nodes; The temporary experience cache is used to store the triple <S, A, S'> obtained immediately after the agent makes an action; the global experience cache is used to store the quadruple <S, A, R, S'> successfully matched by the delayed reward matching module.
2. A containerized microservice orchestration method based on deep reinforcement learning, characterized in that: It includes: S1: Randomly initialize the Deep Q-Network, and then initialize the hyperparameters and the global experience cache. The hyperparameters include α, γ, ε; Describe the edge node environment state as a vector S = {q1, q2,..., q n ,..., q |N|}, where q n represents the number of microservices already running on the nth node, N represents the total number of nodes, and describe the information obtained as a triple <S, A, S'>, where S represents the current state, A represents the action taken by the current agent, and S' represents the next state to be updated after the agent takes the action; S2: Determine whether there is still an episode not trained. If so, reset the "global experience cache" and the "temporary experience cache", and reset the environment to obtain the initial state. If not, directly end; S3: Determine whether there is a microservice not yet orchestrated. If so, use the epsilon-greedy strategy to select the action of the agent. If not, directly end; S4: After performing the action in step S3 in the environment, obtain the next state, and then obtain the triple <S, A, S'>, and store the triple in the temporary experience cache, and then obtain the completion time of the current microservice and the delayed reward R; The delayed reward R = -T f (·), where T f (·) represents the completion time of the microservice; By assigning the unique identifier of the marked microservice of the microservice being orchestrated to both the triple <S, A, S'> and the delayed reward R, it is ensured that the triples in the temporary experience cache are traversed after the delayed reward is obtained to complete the matching; S5: Determine whether the local buffer has been traversed. If so, directly end. If not, complete the matching between the triple <S, A, S'> and the delayed reward R through the consistency of the identifiers. The consistent identifier refers to the identifier of the microservice currently being orchestrated, and then store the successfully matched quadruple <S, A, R, S'> in the global experience cache; The action information to be taken is represented as a vector A = {a1, a2,..., a n ,..., a |N|}, where a n represents scheduling the microservice to node n and sending the corresponding action information to the system information acquisition module, and N represents the total number of nodes; S6: Finally, use the quadruples stored in the global experience cache to train the Deep Q-NetWork.
3. The containerized microservice orchestration method based on deep reinforcement learning with an empirical matching mechanism according to claim 2, characterized in that: In step S5, if the triple <S, A, S'> and the delayed reward R do not match successfully, continue to determine whether the local buffer has been traversed.
Citation Information
Patent Citations
Cloud-edge-terminal equipment platform control architecture and method based on SuperEdge and EdgeXFoundry
CN113612820A
Training method and training device of music generation model, storage medium and equipment
CN115206269A