Kubernetes micro-service optimal deployment method and device based on award accumulation deep reinforcement learning, and medium
By building a multi-objective microservice deployment optimization model and reward accumulation deep reinforcement learning, we solved the problems of resource load balancing and low latency in Kubernetes deployment, and achieved stability and resource efficiency improvements in high-concurrency scenarios.
Patent Information
- Application Number
- CN202511255357.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing Kubernetes deployment strategies struggle to balance resource load balancing and low service response latency in high-concurrency scenarios, leading to resource waste or the risk of service interruption. Traditional methods have high computational complexity and state space explosion, making optimization difficult.
A reward accumulation-based deep reinforcement learning method is adopted to construct a multi-objective microservice deployment optimization model. Resource preference functions and reward functions are designed. Through offline training and online optimization, microservice deployment decisions are made step by step, reducing computational complexity and avoiding state space explosion.
It achieves the coordinated optimization of low latency and high resource utilization on Kubernetes, improves service stability and resource utilization, reduces computing complexity and training difficulty, and meets the deployment requirements of high-concurrency scenarios.
Smart Images

Figure CN120803743A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric digital data processing, in particular to cloud computing resource optimization, and specifically to a Kubernetes microservice optimal deployment method based on reward accumulation deep reinforcement learning, a device and a medium. BACKGROUND
[0002] With the rapid development of cloud computing, traditional service development methods are gradually being replaced by microservice development methods due to poor scalability and maintenance difficulties. Microservice architecture splits a single application into multiple independent microservices, which call each other through interfaces without affecting each other's running results. This method not only improves the flexibility and maintainability of the system, but also reduces resource redundancy caused by resource occupation of a single application. At the same time, with the rapid development of virtualization technology and edge computing, container-based microservice development is becoming the mainstream of microservice design. Containers on different edge nodes can provide virtualized isolation space for a single microservice to ensure that they do not interfere with each other.
[0003] However, the computing resources required by each microservice are irregular, and the total amount of resources of the node is limited. Therefore, how to efficiently utilize computing resources and reasonably allocate them within limited resources has become an important problem in microservice deployment. Traditional microservice deployment mostly involves coarse-grained resource allocation, which is easy to cause resource waste. In addition, while reasonably allocating resources, microservice deployment should also consider quality of service (QOS). In particular, when multiple user requests occur simultaneously, it is necessary to ensure fast service response. Therefore, an efficient microservice deployment method should ensure faster user request response and fine resource utilization.
[0004] In recent years, there have been many studies focusing on microservice deployment. Some methods aim to effectively shorten service response time, which includes processing time and data transmission time of each service, but ignore the limitation of resources. Some methods consider optimizing resource load balancing and service response delay, but ignore the elastic expansion and user access under high concurrency. Some methods achieve microservice traffic routing under high concurrency and high traffic, but do not deploy microservices. Kubernetes is a comprehensive tool for microservice orchestration and deployment. It makes each microservice establish a container, and containers interact with each other through network communication. Its flexibility, open source and stability make it gradually become one of the mainstream tools to realize microservice architecture. However, although deploying microservices on Kubernetes has strong convenience, existing Kubernetes deployment strategies mostly rely on static rules or single target optimization (such as only focusing on resource utilization or response delay), which is difficult to balance the low delay demand under high concurrency scenario and resource load balancing, leading to service interruption risk or resource waste.
[0005] Therefore, there is an urgent need for a Kubernetes optimal deployment method that balances resource load and low service response delay to solve the above problems. SUMMARY
[0006] The present application overcomes the deficiencies of the prior art and provides a Kubernetes microservice optimal deployment method based on reward accumulation deep reinforcement learning, equipment and medium.
[0007] To achieve the above purpose, the technical solution adopted by the present application is as follows: First aspect: provide a Kubernetes microservice optimal deployment method based on reward accumulation deep reinforcement learning, comprising the following steps: S1, construct a multi-objective microservice deployment optimization problem model; S2, construct a microservice resource preference function, and combine the optimization target to reduce the computational complexity in the optimization problem model training; S3, construct a reinforcement learning model; S4, construct a reward function, and use an offline deep reinforcement learning method to train the reinforcement learning model; S5, obtain the optimal deployment strategy according to the real state online through the trained reinforcement learning model.
[0008] In a preferred embodiment of the present application, in the step of S1, the multi-objective microservice deployment optimization model includes a user request response model, a resource load balancing model, a constraint condition, and two optimization targets of minimum user request service response delay and minimum cluster resource variance.
[0009] In a preferred embodiment of the present application, the obtaining of the user request response model includes the following steps: S11, split the user request response delay into service running delay, request queuing delay and data transmission delay, and obtain the corresponding calculation method according to the corresponding calculation strategy respectively; S12, the total user request response delay is obtained by adding the corresponding three delays.
[0010] In a preferred embodiment of the present application, the obtaining of the resource load balancing model includes the following steps: S111, calculate the resource usage rate of CPU, memory, IO and network transmission rate for one user request respectively, and obtain the variance of four types of calculation resources; S112, design resource weight, then the total resource load balancing can be calculated according to the variance and weight.
[0011] In a preferred embodiment of the present application, the obtaining of the multi-objective microservice deployment optimization problem includes the following steps: S1111. To meet target requirements, multiple optimization goals are set, including minimizing user request service response delay and maximizing resource load balance. S1112. To meet the target requirements, the constraints in the microservice deployment process must be met, including that the sum of all computing resources running on each node must not exceed the amount of computing resources available on the node, and that each microservice must ensure that a corresponding container is deployed on at least one node.
[0012] In a preferred embodiment of the present invention, in step S2, various types of computing resources are normalized according to different computing resource requirements of different microservices, and the microservice resource preference function is constructed. In combination with the optimization goal, the following steps are included to obtain: S21. Design Indicates the l The first microservice requirement j Class Resource Comparison Node i Total available resources on j The usage of each resource is compared with the CPU usage to obtain the usage of microservices. l Resource demand preference weight , which can be calculated by the following microservice resource preference function: ; in, R Indicates the available resources on the node; Representing microservices l preference for memory; Representing microservices l CPU preference; Representing microservices l Preference for IO; Representing microservices l Preference for network transmission speed; Indicates the l The memory resources required by each microservice compared to the node i The degree of utilization of the total available memory resources on Indicates the l CPU resources required by each microservice compared to the node i The degree of utilization of the total available CPU resources on the Indicates the l The IO resources required by each microservice are compared with the node i The degree of utilization of the total available IO resources on Indicates the l The network transmission rate resources required by each microservice are compared with the node i the degree of utilization of the total available network transmission rate on S22, the microservice resource preference function and the optimization target are combined, and a load balancing calculation method of different microservices before each training step is obtained: ; Among them, represents the total resource variance; represents the weight of the first j class resource; V CPU represents the variance of the CPU resource in the cluster; similarly, the smaller the minimum total resource variance of the cluster, the more balanced the cluster resource consumption is. l
[0013] In a preferred embodiment of the present application, in the step of S4, the process of training the model by the offline deep reinforcement learning method includes the following steps: S41, at the beginning of each training step, the agent calculates the deployment operation to be performed in the current state according to the current input network state using the model, and at each deployment, in order to reduce the number of action space, the microservices in the microservice chain are deployed one by one, so that each action can be represented as ; wherein, n is an edge node; S42, according to the current state, after the reward function is calculated, a reward value is obtained to evaluate the performance of the action; at the same time, the reward is regarded as the feedback of the model to constrain the model to find the deployment action that maximizes the objective function; S43, the cumulative reward estimation function is calculated using the deployment trajectory path that completes the entire user request to evaluate the advantages and disadvantages of the state; and the entire model is updated after the agent completes the entire user request deployment scheme.
[0014] In a preferred embodiment of the present application, the calculation method of the reward function in each step of reinforcement learning is: ; Among them, in the function , it indicates that there is only one working node used in the training process; otherwise , it indicates that the number of deployed working nodes exceeds the number of microservices; and indicate the weights of the reward function; , it indicates the positive bias reward when the agent successfully deploys the microservice instance; , it indicates the request queuing delay; , it indicates the data transmission delay; , it indicates the total service response delay; , it indicates the traffic on the working node i ; , it indicates the total resource variance;UM Represents a microservice.
[0015] In a second aspect, the present invention provides a Kubernetes deployment system device cluster, comprising: a master node for Kubernetes orchestration, and worker nodes for running each microservice; wherein any one of the above-described Kubernetes microservice optimal deployment methods is executed in the master node, and after the optimal deployment strategy is obtained, the microservice is deployed and run on the worker node; The master node includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute any one of the above-mentioned optimal deployment methods for Kubernetes microservices.
[0016] A third aspect: The present invention provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a processor to implement any of the above-mentioned optimal deployment methods for Kubernetes microservices when executed.
[0017] The present invention solves the defects existing in the background technology and has the following beneficial effects: (1) The present invention provides a Kubernetes microservice optimal deployment method, device and medium based on reward accumulation deep reinforcement learning. By constructing a multi-objective microservice deployment optimization model, service response delay and resource load balancing are converted into mathematically solvable problems, and constraints such as the total node resource limit and the number of container deployments are set. The reward accumulation deep reinforcement learning method dynamically adjusts the reward function weight, comprehensively evaluates the delay and load balancing indicators in each step of deployment, and optimizes the model parameters using the cumulative reward mechanism to achieve multi-objective collaborative optimization, thereby avoiding resource waste or excessive service delay caused by single-objective optimization. Compared with the traditional static scheduling strategy, the present invention solves the optimal microservice deployment requirements on Kubernetes in the existing technology, and can simultaneously meet the low latency and load balancing requirements, thereby further improving the service stability and resource utilization of Kubernetes in high concurrency scenarios.
[0018] (2) In the present application, in view of the complexity of multi-resource dimension variance calculation, a multi-objective microservice deployment optimization model combining microservice resource preference function is designed, the usage degree of memory, IO and CPU usage are compared, the resource demand preference weight is calculated, and the load balancing calculation is simplified by combining the resource weight, the multi-resource dimension variance calculation is converted into single-objective optimization based on preference weight, and the function reduces the simultaneous calculation amount of four types of resources in each step of training through normalization processing, reduces the high-dimensional matrix operation, and further reduces the calculation complexity of model training, compared with the traditional multi-objective optimization method which needs to traverse all resource combinations, thereby the calculation resource consumption can be significantly reduced, the problem that microservice needs low service response and high resource utilization rate when Kubernetes is deployed under high concurrency is effectively solved, the real-time requirement of Kubernetes deployment is further met, and the stability and efficiency of microservice architecture development on Kubernetes can be improved.
[0019] (3) In the present application, in the offline deep reinforcement learning method, a reward feedback strategy based on reward accumulation is designed, and a microservice chain step-by-step deployment strategy is adopted, the deployment of the whole user request chain is decomposed into step-by-step decision of single microservice instance, the exponential state space is decoupled into linear decision sequence, each action only selects a single microservice deployment node, and at the same time, the reward accumulation mechanism accumulates multi-step rewards (instead of single-step feedback) through trajectory path, drives the agent to learn long-term optimal strategy, and suppresses the invalidity of training caused by action combination explosion, thereby effectively avoiding the convergence difficulty of reinforcement learning algorithm caused by state space explosion, and further guaranteeing the real-time performance and reliability of the deployment strategy. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below, and obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings; Figure 1 is the flow chart of the Kubernetes microservice optimal deployment method based on reward accumulation deep reinforcement learning of embodiment 1 of the present application; Figure 2 is the system parameter correlation relationship diagram of the optimization problem model of embodiment 1 of the present application; Figure 3 is the deep reinforcement learning method implementation architecture diagram of embodiment 1 of the present application; Figure 4 is the Kubernetes implementation architecture diagram of embodiment 1 of the present application; Figure 5is a cluster load balancing comparison chart of embodiment 1 of the present application; Figure 6 is a service response comparison chart of embodiment 1 of the present application; Figure 7 A main node structure schematic diagram that can be used to implement embodiment 1 of the present application is shown. DETAILED DESCRIPTION
[0021] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0022] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, however, the present application can be practiced in other manners different from those described herein, therefore, the scope of protection of the present application is not limited by the specific embodiments disclosed below.
[0023] SUMMARY The core features of Kubernetes include declarative API, self-healing capability, service discovery and load balancing, and plug-in-based extensible scheduling mechanism, but the default scheduling algorithm has limitations in complex microservice scenarios: on the one hand, the traditional scheduling strategy uses coarse-grained indicators (such as the average CPU usage of the node) for resource load balancing evaluation, ignoring the coordinated optimization of multiple dimensions of resources such as memory and IO; on the other hand, the calculation of service response delay does not fully decompose the superimposed effect of running delay, queuing delay and transmission delay, resulting in insufficient accuracy of QoS guarantee.
[0024] To achieve low-delay response and high resource utilization at the same time, multi-objective optimization modeling is a feasible path. However, the applicant found that if multi-objective optimization (such as minimizing response delay and resource variance at the same time) is directly applied to Kubernetes scheduling, there may be potential problems, specifically: 1. High computational complexity: due to its high-dimensional calculation, the optimization target needs to calculate the variance of four types of resources (CPU / memory / IO / network) and three layers of delay (running, queuing, and transmission) at the same time. Assuming that the cluster contains N nodes, the computational complexity of single-class resource variance is O(N), and the superposition of four types of resources requires O(4N) node state scans. In a 100-node, 1000-container scenario, a single load balancing calculation requires 400 resource state traversals, and the convergence speed of real-time calculation of resource variance and delay may not meet the response requirements of Kubernetes.
[0025] 2. State space explosion: The state space needs to cover multi-dimensional variables such as node resources, container status, and user requests. Taking 10 nodes, 100 containers, and 5 microservice chains as an example (each variable is discretized into 10 states), the theoretical state space exceeds 10 15 The combination of these two problems makes it difficult for traditional optimization methods to converge or even fail to solve the problem.
[0026] To address the above problems, the present invention proposes a Kubernetes microservice optimal deployment method based on reward accumulation deep reinforcement learning. By constructing a resource preference function, the four-dimensional resource variance is simplified into a single-dimensional indicator, reducing the computational complexity. A dynamic reward mechanism is designed to decompose the deployment action into a step-by-step decision-making of the microservice chain, avoiding the state space explosion problem and optimizing the state space exploration efficiency, thereby achieving the coordinated optimization of low latency and load balancing, and improving the deployment timeliness and resource utilization of Kubernetes in high-concurrency scenarios.
[0027] Example 1: Figure 1 As shown in Figure 1, an optimal deployment method for Kubernetes microservices based on reward accumulation deep reinforcement learning includes the following steps: S1. Build a multi-objective microservice deployment optimization problem model; In this implementation, considering the limited resources of each edge node in the microservice architecture and the need for low user request response latency under high concurrency conditions, deploying microservices on Kubernetes needs to achieve the goals of load balancing and low service response latency. Therefore, a multi-objective microservice deployment optimization problem is constructed, and the optimal deployment strategy is obtained by solving the optimization problem.
[0028] Specifically, based on the Kubernetes deployment process, a system model is built, and the parameters of the deployment process are designed, such as: User Request ; Microservices ; node ; Available resources on the node ; container ; And microservice demand resources ; in, i Indicates the cluster i nodes, s Indicates the s User requests, j Indicates the node j Class resources, l Indicates the corresponding l microservices,k Indicates the k The relationship between the parameters is as follows: Figure 2 shown.
[0029] Furthermore, considering the above parameters, the multi-objective optimization problem should satisfy the constraints: a : , that is, the sum of the corresponding computing resources of all containers running on each node does not exceed the total amount of corresponding computing resources available on the node; in, i Indicates the i Worker nodes; j Indicates the first j Class computing resources; Represents a container k Runtime requirements j The amount of computing resources; Indicates in i The worker nodes target l The corresponding container running each microservice k The number of Indicates in i On the working node j The maximum available capacity of the class resource; b : , that is, each microservice ensures that a corresponding container is deployed on at least one node; in, , indicating that i The corresponding user request is deployed on the nodes s at least one corresponding container; , indicating the s Microservice corresponding to each user request l Deployed on the node i superior.
[0030] Furthermore, optimizing the target load balancing can be transformed into minimizing the variance of various resources in the cluster. Therefore, we first calculate the i Previous j Utilization of computing resources (CPU, memory, IO, network transmission rate): , get the first j Average resource usage of class resources ; Among them, the resource utilization rate is for the s User requests, then Indicates calculating a request in the user request chain at the same time; then we can get jVariance of each resource: ; Since each resource runs in microservices with different focuses, the resource weight is designed as where j is the j th resource on the th node, and ; where is the j th resource on the j th node, and V useage The smaller the , the smaller the overall variance of the cluster, and the more balanced the resources.
[0031] At the same time, the optimization goal of low service response delay can be divided into service running delay , request queuing delay and data transmission delay .
[0032] Specifically, for the s th user request, its corresponding service running delay is: ; where is the running time of a single container corresponding to the microservice; is the corresponding container deployed on the i th node for the s th user request; is the number of corresponding containers running on the i th worker node for the l th microservice.
[0033] For the s th user request, its corresponding data transmission delay is: ; where is the data transmission time between different containers within the same node; is the data transmission time between nodes.
[0034] For the th user request, its corresponding request queuing delay is calculated according to the M / M / 1 queuing theory. First, is the service rate of the microservice instance running on different edge nodes, ; where i is the physical edge nodei The calculation rate of Representing microservices l computational intensity.
[0035] For user requests s A microservice in the microservice chain l , which is on the working node i The flow on , then microservices l Total traffic in the cluster ,but , then microservices l Deployed on the node i The delay is expressed as , microservices in the microservice request chain l Total queuing delay across multiple instances at multiple edge nodes express , then the overall microservice chain queuing delay is: .
[0036] In summary, the calculation method of the total service response delay can be obtained as follows: ,but T The smaller it is, the shorter the delay in responding to user request services.
[0037] Furthermore, the multi-objective microservice deployment optimization problem includes the above two optimization objectives and two constraints, so the optimization problem can be expressed as: ; in, min V useage Indicates the minimum cluster resource variance, then the load is balanced; min T Indicates the minimum user request service response delay, a and b represents the above constraints.
[0038] S2. Build a microservice resource preference function and combine it with the optimization goal to reduce the computational complexity of optimization problem model training; In this implementation, for a microservice chain requested by a user, different microservices have different requirements for various types of computing resources, and different resources have different impacts on services. Based on the total amount of computing resources available at the edge node, different resource types are normalized and designed. Indicates the l The first microservice requirement j Class Resource Comparison Node i Total available resources on j The degree of use is: ; wherein, represents the container k the first j class of computing resources; represents the maximum available capacity of the first i class of resources on the first j work node.
[0039] Optionally, the normalization is one of linear normalization or min-max normalization. Linear normalization refers to mapping data to a specific interval through linear transformation, and the calculation formula is: ; wherein, X represents the original data value (such as the CPU usage, memory occupation, etc. of a certain node, which is the index to be normalized); X max represents the maximum value of the index in all samples; represents the normalized value; The min-max normalization formula is ; wherein, X represents the original data value; X max represents the maximum value of the index in all samples; X min represents the minimum value of the index in all samples; represents the normalized value; Further, when the CPU usage on the edge node is too high, the service call delay in the system will significantly increase, which may cause service interruption or complete unavailability, therefore, the usage degree of each resource is compared with the CPU usage to obtain the resource demand preference weight of the microservice l . , j = cpu, mem, IO, web The resource demand preference weight of the microservice can be calculated by the following microservice resource preference function: ; wherein, represents the preference degree of the microservice l to memory; represents the preference degree of the microservice l to CPU; represents the preference degree of the microservice l to IO; represents the preference degree of the microservice l to network transmission rate; represents the memory resource demand of the first l microservice compared with the node iThe degree of utilization of the total available memory resources on Indicates the l The CPU resources required by each microservice compared to the node i The degree of utilization of the total available CPU resources on the Indicates the l The IO resources required by each microservice are compared with the node i The degree of utilization of the total available IO resources on Indicates the l The network transmission rate resources required by each microservice are compared with the node i the degree of utilization of the total available network transmission rate on Furthermore, by combining the microservice resource preference function with the optimization goal, the new total resource variance of different microservices before each step of training can be obtained. Calculation method: ; in, V CPU Indicates the variance of CPU resources in the cluster.
[0040] It can be seen that before each step of training, the microservice resource preference is calculated first, which reduces the computational complexity of each step of training. Similarly, the minimum cluster total resource variance The smaller the size, the better for microservices l , the more balanced the cluster resource consumption.
[0041] S3. Build a reinforcement learning model; In this implementation, a deep reinforcement learning method based on reward accumulation is proposed to solve the aforementioned multi-objective microservice optimization problem. This method consists of two main parts: offline model training using deep reinforcement learning methods, and online optimization and solution based on the trained model. The key elements of reinforcement learning theory include state, agent, action, and reward. To build a reinforcement learning model, these elements must be set first, which can be obtained as follows: According to the previous system model, the state variables mainly include node status, container status and user request status. S Represents the state of each step of training, then ; in, , respectively represented by the number of nodes, available CPU resources on the node, available memory resources, available IO resources, and available network transmission rate resources NTR; , respectively representing the number of containers on the node and the computing resources required by the container; , respectively represent user requests, microservices corresponding to user requests, and maximum request response delay.
[0042] Furthermore, actions are performed by agents to satisfy constraints a and b Under certain conditions, containers are deployed on edge nodes. In addition, user requests may require multiple microservices to provide service responses, and microservices may deploy multiple corresponding containers on multiple edge nodes.
[0043] S4. Construct a reward function and use offline deep reinforcement learning methods to train the reinforcement learning model. The architecture diagram of the algorithm is as follows: Figure 3 As shown; In this embodiment, at the beginning of each training step, the agent uses the model to calculate the deployment operation to be performed in the current state based on the current input network state, and obtains a reward value to evaluate the performance of the action after calculating the reward function based on the current state; at the same time, the reward is regarded as feedback from the model to constrain the model to find the deployment action that maximizes the objective function.
[0044] Furthermore, in order to reduce the number of action spaces, each time a deployment is performed, the microservices in the microservice chain corresponding to the user request are deployed one by one, that is, for each user request UR s , and its corresponding microservice chain is , then each step of training starts from Start deploying one by one ; At each deployment step, the model only meets part of the user's request requirements; at the same time, each step targets All deployments need to be deployed on the corresponding edge nodes. Therefore, the action space of microservice deployment is the number of edge network nodes that can be selected when deploying microservice instances. Each action can be expressed as: ; By designing each action step, the excessive size of the action space caused by a large number of user request chains and network nodes in reinforcement learning is greatly reduced, thus avoiding the ineffectiveness of the reinforcement learning algorithm.
[0045] Furthermore, the reward function is designed as follows: ; Among them, the function Indicates that only one working node is used during training; otherwise Indicates that the number of deployed working nodes exceeds the number of microservices; and represents the weight of the reward function; represents the positive deviation reward when the agent successfully deploys a microservice instance; Indicates request queuing delay; denotes the data transmission delay; denotes the total service response delay; denotes the traffic on the worker node i ; denotes the total resource variance.
[0046] In each training step, the reward is calculated according to the reward function and the current state, and the obtained reward value is taken as feedback to constrain the model to obtain the optimal target function of the deployment action, and when the constraint condition is violated, a penalty is executed Loss .
[0047] Wherein, the cumulative reward is calculated using the deployment trajectory path completing the entire user request to evaluate the advantages and disadvantages of the state, and then the entire model is updated after the agent completes the entire user request deployment scheme.
[0048] S5, obtaining the optimal deployment strategy online according to the real state through the trained reinforcement learning model.
[0049] In this embodiment, through continuous iterative training, the goal of the agent is to obtain the optimal action and deployment strategy with the maximum reward in the current environment, and after obtaining the real data and state, the trained agent obtains the optimal deployment strategy for the current state in the real environment; The entire deep reinforcement learning method implementation architecture is as shown in Figure 3 .
[0050] Specifically, the hardware environment of the Kubernetes optimal microservice deployment method based on reward accumulation deep reinforcement learning is shown in Table 1, and the required computing resources of each microservice in Example 1 are shown in Table 2.
[0051] Table 1: Hardware environment of Kubernetes cluster in Example 1
[0052] Table 2: Required resources of each microservice in Example 1
[0053] Since the Kubernetes scheduling algorithm becomes more flexible with the development of microservices, the scheduling method can be replaced by a plug-in. In the experiment, a model and method were developed as a scheduling tool to replace the default scheduling method as a plug-in to realize real-time container scheduling, and combined with Prometheus, real-time scheduling and management of the platform were realized.
[0054] Prometheus is an open-source system monitoring tool, and its open-source characteristics and plug-in deployment make it easy to implement in Kubernetes. As Figure 4As shown, Master represents the master node, and the above algorithms are implemented on the master node; Node represents the worker node; Pod is the smallest running unit in Kubernetes, in which more than one container can be run; in the microservice deployment process of Kubernetes, the result of the optimal deployment strategy executed by Kubelet, i.e., deploying containers to nodes, contains cAdvisor, which stores the resource usage of Pods in the cluster, and the Node-Exporter component is responsible for collecting monitoring data on the node and transmitting it to Prometheus for display; Controller Manager is the cluster manager of Kubernetes, and Kube-proxy records the traffic of each container; Etcd saves the running results; the method of the optimal deployment strategy is recorded in the Scheduler, API Server serves as the communication interface of Kubernetes, and the above components communicate through the API server. According to the same user request, the optimal deployment method is implemented on the master node and deployed to the node. The results are compared in terms of response time and load balancing.
[0055] Figure 5 and Figure 6 It is shown in and that, compared with the Kubernetes default scheduling method (Default), the deep Q learning method (DQL) and the first fitting algorithm (FFD), the microservice optimal deployment method based on reward accumulation deep reinforcement learning (DRFLM) has the comparison of total resource variance (Variances of resources) and average service response time (Response time) under different number of user requests (Number of user requests).
[0056] From Figure 5 In order to clearly show the difference of variance value, the corresponding order of magnitude is increased, and the results after deployment by the above four methods can obtain the corresponding variance value through the monitoring and calculation of cluster resources; it can be seen that the variances of the above four methods increase with the increase of the number of user requests, which indicates that the more the number of user requests, the more difficult to control the resource usage balance, but when the number of user requests is the same, the variance of the microservice optimal deployment method based on reward accumulation deep reinforcement learning is the lowest, which indicates that this method can effectively ensure the more balanced use of cluster resources.
[0057] From Figure 6In the above four methods, the average service response time under different user request quantities is collected after the deployment; it can be seen that with the increase of the number of user requests, the average service response time also increases, because the queuing time and algorithm running time are increasing, but the average service response time of the micro-service optimal deployment method based on reward accumulation deep reinforcement learning is the lowest when the number of user requests is the same, which shows that the method can effectively reduce the service response delay.
[0058] Through the above comparison, it is shown that the micro-service deployment on Kubernetes implemented by the method has high efficiency and stability, and can meet the requirements of low service response, high resource utilization, stability and real-time in the development of systems based on micro-service architecture and Kubernetes.
[0059] Specifically, the present application considers the service response delay and the limited computing resources of the nodes in the Kubernetes deployment process, effectively converts the Kubernetes deployment requirements into a computable and solvable optimal problem by constructing a multi-objective micro-service deployment optimization problem, taking into account the two objectives of the lowest service response delay and resource load balancing; at the same time, a micro-service resource preference function is designed, which is combined with the optimization objective to reduce the computational complexity during model training, ensuring the real-time performance of Kubernetes deployment, thereby meeting the development requirements of micro-service architecture.
[0060] And the multi-objective micro-service deployment optimization model is a step-by-step serialization of actions taken under dynamic constraints and states to maximize expected returns, so the learning ability of the reinforcement learning method is suitable for solving the above-mentioned micro-service deployment optimization problem; by designing the action space of each training step, combining the feedback mechanism based on reward accumulation, and obtaining the best agent through offline deep reinforcement learning training, and then solving it online according to the real state, not only avoids the too large action space of reinforcement learning, but also meets the real-time requirement of deployment strategy, and also obtains the optimal deployment strategy that meets the requirements.
[0061] Embodiment 2: The Kubernetes deployment system device representing the cluster edge node in embodiment 1 of the present application is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers, which can also be cloud servers, also known as cloud computing servers or cloud hosts, which are a host product in the cloud computing service system. The components shown in this paper, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0062] The Kubernetes deployment system device cluster comprises: a master node for Kubernetes orchestration, and a worker node for running microservices; wherein the method provided by the application is executed in the master node, and after the optimal deployment strategy is obtained, the microservices are deployed and run in the worker node; As shown in Figure 7 The master node 10 comprises at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program which can be executed by the at least one processor 11, and the computer program is executed by the at least one processor 11 to enable the at least one processor 11 to execute the method provided by the application.
[0063] Further, the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the master node 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0064] The plurality of components in the master node 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the master node 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0065] Further, the processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various specialized Artificial Intelligence (AI) computing chips, various processors running machine learning model algorithms, a Digital Signal Process (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, e.g., the method of resource management for a database.
[0066] In some specific embodiments, the method of resource management for a database can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the master node 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the method of resource management for a database described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the method of resource management for a database by other any appropriate means, e.g., by means of firmware.
[0067] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0068] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0069] In the context of the present application, the computer readable storage medium stores computer instructions for causing a processor to implement the method of resource management of a database provided by the present application when executed. The computer readable storage medium can be a tangible medium which can contain or store the computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, the computer readable storage medium can be a machine readable signal medium. More specific examples of the machine readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0070] To provide for interaction with a user, the systems and techniques described here can be implemented on a host node having a display device (e.g., a Cathode Ray Tube (CRT) or a Liquid Crystal Display (LCD) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the host node. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0071] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0072] Optionally, the computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and virtual private server (VPS) services.
[0073] The above is according to the ideal embodiment of the present application, through the above description, for those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and can be realized in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, the scope of the present application is defined by the appended claims rather than the above description, therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0074] In addition, it should be understood that, although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the description of the specification is only for the sake of clarity, those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be properly combined to form other embodiments that those skilled in the art can understand.
Claims
1. A Kubernetes microservice optimal deployment method based on reward accumulation deep reinforcement learning, characterized by: The following steps are involved: S1. Build a multi-objective microservice deployment optimization problem model; S2. Build a microservice resource preference function and combine it with the optimization goal to reduce the computational complexity of optimization problem model training; S3. Build a reinforcement learning model; S4. Construct a reward function and use offline deep reinforcement learning methods to train the reinforcement learning model; S5. Obtain the optimal deployment strategy online based on the real state through the trained reinforcement learning model.
2. The optimal deployment method for Kubernetes microservices based on reward accumulation deep reinforcement learning according to claim 1 is characterized in that: In step S1, the multi-objective microservice deployment optimization problem model includes a user request response model, a resource load balancing model, constraints, and two optimization objectives: minimum user request service response delay and minimum cluster resource variance.
3. The optimal deployment method for Kubernetes microservices based on reward accumulation deep reinforcement learning according to claim 2 is characterized in that: The acquisition of the user request response model comprises the following steps: S11. Split the user request response delay into service operation delay, request queuing delay, and data transmission delay, and obtain corresponding calculation methods according to corresponding calculation strategies; S12. The total user request response delay is calculated by adding the corresponding three delays.
4. The optimal deployment method for Kubernetes microservices based on reward accumulation deep reinforcement learning according to claim 2 is characterized in that: Obtaining the resource load balancing model includes the following steps: S111, respectively calculating resource utilization rates of CPU, memory, I / O, and network transmission rate for a user request, and obtaining variances of the four types of computing resources; S112. Design resource weights, and the total resource load balance can be calculated based on the variance and weights.
5. The optimal deployment method for Kubernetes microservices based on reward accumulation deep reinforcement learning according to claim 2 is characterized in that: The multi-objective microservice deployment optimization problem is obtained by: S1111. To meet target requirements, multiple optimization goals are set, including minimizing user request service response delay and maximizing resource load balance. S1112. To meet the target requirements, the constraints in the microservice deployment process must be met, including that the sum of all computing resources running on each node must not exceed the amount of computing resources available on the node, and that each microservice must ensure that a corresponding container is deployed on at least one node.
6. The method for optimal deployment of Kubernetes microservices based on reward accumulation deep reinforcement learning according to claim 1, characterized in that: In step S2, according to the different computing resource requirements of different microservices, various types of computing resources are normalized, the microservice resource preference function is constructed, and combined with the optimization goal, The steps include: S21. Design Indicates the l The first microservice requirement j Class Resource Comparison Node i Total available resources on j The usage of each resource is compared with the CPU usage to obtain the usage of microservices. l Resource demand preference weight , which can be calculated by the following microservice resource preference function: ; in, R Indicates the available resources on the node; Representing microservices l Preference for memory; Representing microservices l CPU preference; Representing microservices l Preference for IO; Representing microservices l Preference for network transmission speed; Indicates the l The memory resources required by each microservice compared to the node i The degree of utilization of the total available memory resources on Indicates the l CPU resources required by each microservice compared to the node i The degree of utilization of the total available CPU resources on the Indicates the l The IO resources required by each microservice are compared with the node i The degree of utilization of the total available IO resources on Indicates the l The network transmission rate resources required by each microservice are compared with the node i the degree of utilization of the total available network transmission rate on S22. Combining the microservice resource preference function with the optimization goal, we can obtain the load balancing calculation method for different microservices before each training step: ; in, represents the total resource variance; Indicates the j The weight of the class resource; V CPU Indicates the variance of CPU resources in the cluster; similarly, the smaller the variance of the minimum cluster total resources, the better for microservices l , the more balanced the cluster resource consumption.
7. The method for optimal deployment of Kubernetes microservices based on reward accumulation deep reinforcement learning according to claim 1, characterized in that: In step S4, the process of training the model using the offline deep reinforcement learning method includes the following steps: S41. At the beginning of each training step, the agent uses the model to calculate the deployment operation to be performed in the current state according to the current input network state. At the same time, in order to reduce the number of action spaces, the microservices in the microservice chain are deployed one by one at each deployment. Therefore, each action can be expressed as ;in, n It is an edge node; S42. Based on the current state, after the reward function is calculated, a reward value is obtained to evaluate the performance of the action. At the same time, the reward is regarded as feedback to the model to constrain the model to find the deployment action that maximizes the objective function. S43. Use the deployment trajectory path that completes the entire user request to calculate the cumulative reward estimation function to evaluate the advantages and disadvantages of the state; and update the entire model after the agent completes the deployment plan for the entire user request.
8. The method for optimal deployment of Kubernetes microservices based on reward accumulation deep reinforcement learning according to claim 7, characterized in that: The calculation method of the reward function designed in each step of reinforcement learning is: ; Among them, the function Indicates that only one working node is used during training; otherwise Indicates that the number of deployed working nodes exceeds the number of microservices; and represents the weight of the reward function; represents the positive deviation reward when the agent successfully deploys a microservice instance; Indicates request queuing delay; Indicates data transmission delay; Indicates the total service response delay; Represents a working node i Traffic on represents the total resource variance; UM Represents a microservice.
9. A Kubernetes deployment system device cluster, comprising: a master node for Kubernetes orchestration, and worker nodes for running each microservice; wherein the Kubernetes microservice optimal deployment method according to any one of claims 1 to 8 is executed in the master node, and after the optimal deployment strategy is obtained, the microservice is deployed and run on the worker node; The master node includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so as to enable the at least one processor to execute the Kubernetes microservice optimal deployment method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a processor to implement the Kubernetes microservice optimal deployment method according to any one of claims 1 to 8 when executed.
Citation Information
Patent Citations
Micro-service combination deployment and scheduling method under multi-objective optimization
CN111027736A
Heterogeneous multi-edge cloud collaborative micro-service deployment and routing joint optimization method and system
CN116915686A
Resource scheduling method for optimizing edge energy consumption and load based on reinforcement learning
CN117194057A
Data intensive task edge service combination method based on multi-target reinforcement learning
CN117255126A
Performance-aware micro-service adaptive deployment and resource allocation method and system in cloud edge environment
CN117640378A
Cited By
Kubernetes micro-service redeployment method and device based on service communication resource sharing and medium
CN121050734A