A Distributed Microservice Scheduling Optimization Method

By introducing service dependency, data dependency and user dependency constraints in microservice scheduling, and using the DDPG model of reinforcement learning for continuous optimization, the problem of dependency impact in microservice scheduling in distributed scenarios is solved, and efficient and low-latency microservice scheduling is achieved.

CN115714820BActive Publication Date: 2025-05-27NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211424666.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-05-27
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

In distributed scenarios, microservice scheduling needs to consider the dependency constraints between services and services, services and data, and users and services, but the existing technology has not effectively solved the impact of these dependencies on scheduling and lacks continuous optimization capabilities.

Method used

A distributed microservice scheduling optimization method is adopted. By introducing service dependence, data dependence and user dependence as scheduling constraints, and using the continuous optimization capabilities of reinforcement learning, a microservice scheduling optimization solution based on the DDPG model is built to optimize the deployment location of microservice instances to achieve low service latency and high resource balance.

Benefits of technology

It effectively improves the efficiency and quality of microservice scheduling, can better meet the timeliness requirements of microservices and services obtained by microservice combinations, and ensures that the cloud environment has high resource balance and low service delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115714820B_ABST
    Figure CN115714820B_ABST
Patent Text Reader

Abstract

A distributed microservice scheduling optimization method includes the following steps: Step 1: Model the environment where the t-th scheduling task is located as a quadruple to describe the deployment of microservice instances in the distributed environment instance at this moment; Step 2: Model the action of the t-th scheduling task as a binary tuple to describe the microservice being deployed to a node at this moment; Step 3: Use the sum of the differences in service latency and the differences in resource balance of all services after and before the action is completed for reward description; Step 4: Perform training to obtain a prediction model for the microservice instance scheduling optimization scheme; Step 5: Use the prediction model for prediction to obtain the corresponding scheduling optimization scheme. The present invention introduces the dependencies between services, between services and data, and between services and users into the microservice scheduling constraints and utilizes the continuous optimization ability of reinforcement learning, which is a microservice scheduling optimization method adapted to distributed scenarios and having the continuous optimization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of microservice scheduling optimization, and particularly relates to a distributed microservice scheduling optimization method. Background Art

[0002] In a distributed service scenario, microservices are not only deployed on the cloud, but various services will also be deployed on the edge and at the end according to requirements. This distributed deployment method of microservices will bring certain difficulties to scheduling.

[0003] First, since microservices usually have the characteristic of single function, microservices need to cooperate with each other to provide services externally, so the calls between microservices are common. If the data calls between microservices occur frequently, it will increase the service response time and affect the service performance.

[0004] Second, data is the core of services, and the dependence between services and data is common. If data cannot be moved due to factors such as large data scale and data security, the deployment location of services that depend on data cannot be deployed arbitrarily. If the deployment location is inappropriate, it is very likely to increase service latency and even cause the service to be unable to execute.

[0005] Third, if users are restricted by the management domain, the distance between services and users will increase the data transmission time, so the deployment location of services should also be restricted.

[0006] In summary, in a distributed scenario, the deployment locations of microservices, data, and users are not arbitrary. The scheduling of microservices needs to consider the dependency constraints between services, between services and data, and between users and services. However, most current microservice scheduling only considers the scheduling of services on the cloud. Usually, the resource center, microservices, and users are regarded as entities, and the constraints include resource constraints, balance constraints, etc. The dependencies between services, between services and data, and between services and users are not taken as factors affecting scheduling. In addition, the current microservice scheduling methods basically operate in the way of one-time training and multiple applications, and do not have the ability of continuous optimization. Summary of the Invention

[0007] In order to overcome the above technical problems, the purpose of the present invention is to provide a distributed microservice scheduling optimization method. This method introduces the dependencies between services, between services and data, and between services and users into the microservice scheduling constraints, and utilizes the continuous optimization ability of reinforcement learning. It is a microservice scheduling method suitable for distributed scenarios and having the ability of continuous optimization.

[0008] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0009] A distributed microservice scheduling optimization method, comprising the following steps;

[0010] Step 1: Based on the microservice instance, the distributed environment instance, and the deployment location of the microservice in the distributed environment, model the environment where the t-th scheduling task is located as a quadruple to describe the deployment situation of the microservice instance in the distributed environment instance at this moment;

[0011] Step 2: Model the action of the t-th scheduling task as a binary tuple to describe that the microservice s i is deployed to the node n j ;

[0012] Step 3: With the goal of low service latency and high resource balance, use the sum of the differences in service latency and the differences in resource balance of all services after and before the action is completed for reward description;

[0013] Step 4: Let the environment E t =(t, S, N, SN), where t is a natural number greater than 0, representing the t-th microservice scheduling optimization task, S is a vector composed of all microservice instance description information, N is a vector composed of all node description information in the cloud environment instance, and SN is a vector composed of ordered pairs with the microservice instance number as the first element and the node number as the second element; the action a t =(s i , n j ), where s i is the microservice instance number, and n j is the number of the node to which s i is deployed; the reward r t =(g t+1 -g t )-(avgT t+1 -avgT t ), where g t and g t+1 are the resource balances before and after the execution of the action a t respectively, and avgT t and avgT t+1 are the service latencies before and after the execution of the action a t respectively. Input into the DDPG model for training to obtain a prediction model for the microservice instance scheduling optimization scheme;

[0014] Step 5: Use the trained prediction model for the microservice instance scheduling optimization scheme to predict the scheduling optimization scheme of the microservice instance in the cloud environment instance, and obtain the corresponding scheduling optimization scheme to solve the technical problem of deploying the microservice instance to the nodes of the cloud environment instance while ensuring high resource balance and low service latency in the cloud environment instance.

[0015] The specific content of step 1 is as follows:

[0016] Environmental description: Define the environmental state E of the t-th microservice scheduling optimization task t as a quadruple, E t =(t, S, N, SN), where t is the serial number of the scheduling optimization task; S is a vector representing the current service, i.e., S = {s 1 , s 2 , …, s i , …, s n}; N is a vector representing the current node, i.e., N = {n 1 , n 2 , … n j , … n m}; SN is a vector representing the current service deployment plan, i.e., SN = (<s s1 , n n1 >, <s s2 , n n2 >, … <s sk , n nk >, … <s sNum , n nNum >), where each element is an ordered pair <s sk , n nk >, indicating that the service s sk is deployed to the node n nk .

[0017] The specific content of step 2 is as follows:

[0018] Action description: Define the action of the t-th microservice scheduling optimization task as the operation a of deploying a service to a node in the current optimization scheduling task t =(s i , n j ), that is, the service s i is deployed to the node n j , where the selection of n j should meet the restrictions of service dependency, data dependency, and user dependency constraints, facing a distributed service environment.

[0019] The scheduling constraints of the service are designed as follows:

[0020] Service Dependency: The functions of microservices are relatively independent. In a distributed environment, user requirements are complex, and microservices need to cooperate with each other to meet user needs. By analyzing the call relationships between microservices, service dependencies of various strengths are extracted from the link data of microservices to constrain the action set. Specifically, the number of call relationships between services is counted from the microservice call link data, and then the number of calls between services is divided by the total number of calls per unit time of the service as the strength of the service dependency. Thereafter, when the service to be deployed has a strong service dependency relationship with other services, the node where the strongly dependent service is located or an adjacent node is preferentially selected;

[0021] Data Dependency: Data is the core element of services, so data dependency is used to limit the action set. The location of the data source can be obtained through prior knowledge, and the dependency relationship between the service and the data source is obtained through the parsing of the business process description file. It is described by a matrix. The rows of the matrix represent the node numbers where the services are located, and the columns of the matrix represent the node numbers where the data is located. If there is a dependency between the service and the data, the element with the service node number and the data node number as subscripts is 1, otherwise it is 0. Thereafter, when the service to be deployed has a dependency relationship with the data, the node where the data is located or an adjacent node is preferentially selected;

[0022] User Dependency: Users are the main body that invokes services, so user dependency is used to limit the action set. The dependency relationship between users and services is obtained through prior knowledge and is described by a matrix. The rows of the matrix represent the node numbers where the services are located, and the columns of the matrix represent the node numbers where users are connected. If there is a dependency between the service and the user, the element with the service node number and the user connection node number as subscripts is 1, otherwise it is 0. Thereafter, when the service to be deployed has a dependency relationship with the user, the node where the user is located or an adjacent node is preferentially selected.

[0023] The specific steps of step 3 are as follows:

[0024] Reward Description: Take the sum of the difference in service latency and the difference in resource balance of all services after and before the action a t is completed as the basis for setting the reward value. The calculation method is shown in formula (1). When the service latency decreases and the resource balance improves, the reward is positive, otherwise the reward is 0 or negative;

[0025] r t =(g t+1 -g t )-(avgT t+1 -avgT t ) Formula (1)

[0026] As shown in formula (1), r t is the reward value after the execution of the t-th microservice scheduling optimization task, g t and g t+1They are respectively the resource balance before and after the execution of action a t avgT t and avgT t+1 They are respectively the resource balance before and after the execution of action a t service latency before and after the execution of action a

[0027] The calculation methods of the resource balance and service latency are as follows:

[0028] 1) Resource balance: Define the resource balance based on CPU utilization and memory utilization;

[0029] CPU utilization: The CPU utilization of a node is defined as the ratio of the allocated CPU resources in the node to the total CPU resources. The allocated CPU resources are calculated based on the CPU usage of the containers on the node, as shown in formula (2), where n is the total number of containers on the node;

[0030] Formula (2)

[0031] The calculation method of CPU utilization is as shown in formula (3), where Capacity cpu is the total CPU resources of the node;

[0032]

[0033] Memory utilization: The memory utilization of a node is defined as the ratio of the allocated memory resources in the node to the total memory resources. The allocated memory resources are calculated based on the memory usage of the containers on the node, as shown in formula (4), where n is the total number of containers on the node;

[0034] Formula (4)

[0035] The calculation method of memory utilization is as shown in formula (5), where Capacity mem is the total memory resources of the node;

[0036]

[0037] The node resource balance is defined as the absolute value of the difference between the node CPU utilization and memory utilization, as shown in formula (6);

[0038] g i = |Ratio_Node cpu i –Ratio_Node mem i | Formula (6)

[0039] where i is the node number, Ratio_Node cpu iIndicates the CPU utilization rate of the i-th node, Ratio_Node mem i Indicates the memory utilization rate of the i-th node;

[0040] The cloud environment contains a large number of nodes. The resource balance needs to comprehensively consider all nodes. The balance of a single node does not represent good performance of the container cloud. The variance of the node resource balance of all nodes is used to evaluate the resource balance of the cloud environment, which is used to assist in reward scoring, as shown in formula (7);

[0041]

[0042] Among them, is the mean value of the resource balance of each node, g i is the resource balance of the i-th node;

[0043] 2) Service latency: The sum of the communication latency and the execution latency is used to represent the service latency, as shown in formula (8);

[0044] T i,j = comT i,j + exeT i,j Formula (8)

[0045] Among them, T i,j represents the service latency of service i on node j, comT i,j represents the communication latency of service i on node j, which is generated by the network transmission of the services and data on which service i depends, and exeT i,j represents the execution latency of service i on node j;

[0046] The communication latency of the service can be defined as the total latency generated by service dependencies and data dependencies, as shown in formula (9);

[0047] comT i,j = sevT i,j + datT i,j Formula (9)

[0048] Among them, sevT i,j represents the latency generated by service dependencies of service i on node j, which is represented by the sum of the quotient of the dependency strength and the bandwidth of each hop. The dependency strength between services is defined as the number of calls between services divided by the total number of calls per unit time of the service; datT i,j represents the latency generated by data dependencies of service i on node j, which can be represented by the sum of the quotient of the data volume and the bandwidth of each hop;

[0049] Assume that a node can deploy all the services scheduled for it, and the execution delay of a service is defined as the ratio of the service instruction length to the processing capacity of the node's CPU, as shown in formula (10);

[0050]

[0051] Among them, mips j represents the CPU instruction execution speed of node j, and cpuU i represents the CPU utilization rate of service i, and mi i represents the instruction length of service i. The first two metrics can be obtained through a monitoring system, and the third metric is prior knowledge;

[0052] In a distributed node, there are a large number of services. The average value of all service response times is used to measure the delay of services in the cloud environment, as shown in formula (11).

[0053]

[0054] Among them, i represents the service number, the total number of services is N, j represents the node number, and the total number of nodes is M.

[0055] The specific steps of step 4 are as follows:

[0056] Describe the microservice scheduling optimization problem based on deep reinforcement learning as a 5-tuple O t =(t, E t , a t , r t , E t+1 ), where t is a value between 1 and n, which is the task number of scheduling optimization; E t represents the state of the distributed environment at the t-th task moment; a t represents the action taken in state E t , that is, the service scheduling scheme; r t represents the reward obtained by taking action a t in state E t ; E t+1 represents the new state obtained by taking action a t in state E t . The 5-tuples depicting each scheduling optimization task will be stored in the experience replay pool to provide a basis for model training;

[0057] DDPG uses a dual-network structure composed of a policy network and a value network. The input of the policy network is the current state E of the environment t , and the output is the corresponding action a t . The input of the value network is the current state E of the environment t and the action a t, the output is E t Execute action a in the state t The scoring is as follows. In addition, both the policy network and the value network are divided into an online network and a target network. The online network and the policy network have the same corresponding structure but different initial parameters.

[0058] Specifically, step 5 is as follows:

[0059] Based on the trained model, input the environment of the current microservice scheduling optimization task and the scheduling constraints composed of the service dependency matrix, data dependency matrix, and user dependency matrix into the trained model, execute the model, generate a scheduling optimization plan, and feedback the reward value obtained from executing the current microservice scheduling optimization task to the environment for continuous optimization.

[0060] The beneficial effects of the present invention.

[0061] The present invention transforms the scheduling optimization problem of microservices into a prediction problem based on reinforcement learning, quantifies factors affecting microservice performance such as service dependency, data dependency, user dependency, and distributed resources into the scheduling model, considers more optimization factors, and also helps to improve the scheduling efficiency of services composed of microservices.

[0062] The present invention analyzes the characteristics of microservices, including service dependency, data dependency, and user dependency. Based on the above characteristics, three types of scheduling constraints are designed, and a model for the distributed microservice scheduling optimization problem is given. Thereafter, the scheduling problem of distributed microservices is converted into a model-based prediction problem, and the microservices are continuously scheduled and optimized based on the trained model. This application can effectively improve the efficiency of generating a scheduling plan on the premise of ensuring the effectiveness of the scheduling plan, and can better meet the timeliness requirements of microservices and services composed of microservices. Brief Description of the Drawings

[0063] Figure 1 Schematic diagram of the basic environment of the distributed service.

[0064] Figure 2 Basic process of the microservice scheduling optimization problem based on deep reinforcement learning.

[0065] Figure 3 Structural diagram of the scheduling optimization method based on deep reinforcement learning. Detailed Embodiment

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0067] This application designs a distributed microservice scheduling optimization method. In a distributed technology service environment, there are many factors affecting service performance. To ensure service quality, various scheduling optimization algorithms with different optimization goals have emerged, but the ultimate goal is to ensure the normal operation of the service. However, the service scheduling optimization algorithms proposed in different periods have corresponding limitations. In addition, using the same scheduling scheme for different types of microservices may result in unsatisfactory scheduling effects. Therefore, it is necessary to study the scheduling optimization method for microservices in a distributed environment.

[0068] Figure 1 As shown, the basic environment of a distributed service mainly includes three types of objects: distributed nodes, services, and users. Services are the core, referring to microservices that provide business functions. Nodes are the carriers of services, including cloud nodes and edge nodes, which are used to provide computing and storage capabilities. Nodes are connected through a network, and the transmission speed is affected by the bandwidth. Users are the consumers of services, and the distribution of users is restricted by the management domain of the industry or unit where they are located. In summary, in a distributed technology service environment, microservice scheduling refers to the process of deploying microservices to appropriate nodes under certain constraints with the optimization goal.

[0069] The present invention comprehensively considers the factors affecting service scheduling and scheduling efficiency. In the process of service scheduling optimization, it mainly considers the constraints of 3 control actions, the optimization goals of 2 auxiliary rewards, and uses an improved Deep Deterministic Policy Gradient (DDPG) algorithm to achieve service scheduling optimization.

[0070] (1) Scheduling constraints

[0071] For a distributed service environment, the scheduling constraints of services are designed as follows:

[0072] 1) Service Dependency: Generally, the functions of microservices are relatively independent. In a distributed environment, users' technology service requirements are complex, and microservices need to cooperate with each other to meet user needs. Therefore, dependencies between microservices are widespread. This study defines the call relationship between microservices as a dependency and divides the intensity of service dependencies into strong and weak dependencies according to the service call frequency. Generally speaking, services with strong dependencies call each other frequently. If services with strong dependencies are combined and deployed on adjacent or even the same node, it can effectively avoid the network overhead caused by data transmission, thereby optimizing the performance of technology services. Service dependencies can be used to constrain the action set in the scheduling optimization method based on deep reinforcement learning. Specifically, the number of call relationships between services can be statistically obtained from the service call link data acquired by Istio monitoring, and then the number of calls between services is divided by the total number of calls within the service unit time as the intensity of service dependencies. After that, when the service to be deployed has a strong service dependency relationship with other services, preferentially select the node where the strongly dependent service is located or adjacent to it.

[0073] 2) Data Dependency: Data is the core of services, and almost all services need to process data. Therefore, the dependency between services and data is also widespread. For a service to process data, it must access the data. If the data and the service are not located on the same node, data transmission is required. Therefore, the data transmission time also becomes a factor affecting service latency. Making the service and data as close as possible is also beneficial to ensuring service performance. Especially for data of a certain scale, the impact of transmission time on performance will be more obvious. Therefore, this application uses data dependency as a behavior constraint to limit the action set. The location of the data source can be obtained through prior knowledge, and the dependency relationship between the service and the data source can be obtained through the parsing of the business process description file.

[0074] 3) User Dependency: Users are the main body of services, and the deployment of services will also be affected by users. For services with security and management requirements, the services need to be deployed within the scope defined by users. Therefore, this application also uses data dependency as a behavior constraint to limit the action set. The dependency relationship between users and services is obtained through prior knowledge.

[0075] (2) Optimization Objectives

[0076] For a distributed service environment, the optimization objectives of services are designed as follows:

[0077] Generally speaking, controlling costs and improving the user experience are the most concerned evaluation criteria for service scheduling. In terms of cost control, reasonable resource allocation will help control costs. For the nodes carrying services, CPU and memory are the main resources. If the utilization rates of CPU and memory are unbalanced, especially when the utilization rate of CPU or memory is too high, it will lead to waste of another resource, resulting in low overall resource utilization rate and increased costs. Accuracy and timeliness are the basic quality requirements for all information, and the same applies to the services providing information. Accuracy is determined by the internal logic of the service and is not the research content of this application. Timeliness is reflected in the latency of the service. If the timeliness of the user's needs cannot be met, it will directly affect the user experience. Therefore, this application takes resource balance and service latency as the goals of microservice scheduling optimization.

[0078] 1) Resource balance: Nodes are the basic units for carrying services. The balanced utilization of resources can not only avoid the decline in service performance caused by high-load nodes, but also achieve cost control to a certain extent, that is, avoid resource waste and cost increase caused by low-load or even zero-load nodes. This application defines resource balance based on CPU utilization rate and memory utilization rate.

[0079] CPU utilization rate: The CPU utilization rate of a node is defined as the ratio of the allocated CPU resources in the node to the total CPU resources. The allocated CPU resources are calculated based on the CPU usage of the containers on the node, as shown in formula (1), where n is the total number of containers on the node.

[0080] Formula (2)

[0081] The calculation method of CPU utilization rate is shown in formula (3), where Capacity cpu is the total CPU resources of the node.

[0082]

[0083] Memory utilization rate: The memory utilization rate of a node is defined as the ratio of the allocated memory resources in the node to the total memory resources. The allocated memory resources are calculated based on the memory usage of the containers on the node, as shown in formula (4), where n is the total number of containers on the node.

[0084] Formula (4)

[0085] The calculation method of memory utilization rate is shown in formula (5), where Capacity mem is the total memory resources of the node.

[0086]

[0087] The node resource balance is defined as the absolute value of the difference between the node CPU utilization rate and the memory utilization rate, as shown in formula (6).

[0088] g i = |Ratio_Node cpu i –Ratio_Node mem i | Equation (6)

[0089] Where i is the node number, and Ratio_Node cpu i represents the CPU utilization rate of the i-th node, and Ratio_Node mem i represents the memory utilization rate of the i-th node.

[0090] The cloud environment contains a large number of nodes. The resource balance needs to comprehensively consider all nodes. The balance of a single node does not represent good performance of the container cloud. Therefore, this application uses the variance of the node resource balance of all nodes to evaluate the resource balance of the cloud environment and assist in reward scoring, as shown in Equation (7).

[0091]

[0092] Where is the mean value of the resource balance of each node, and g i is the resource balance of the i-th node.

[0093] 2) Service latency: The ultimate object of the service is the user, and latency is one of the most important indicators affecting the user experience. To balance the response performance of all distributed technology services. The main factors affecting service latency are data transmission and service execution. Therefore, the sum of communication latency and execution latency is used to represent service latency, as shown in Equation (8).

[0094] T i,j = comT i,j +exeT i,j Equation (8)

[0095] Where T i,j represents the service latency of service i on node j, comT i,j represents the communication latency of service i on node j, which is generated by the network transmission of the services and data on which service i depends. exeT i,j represents the execution latency of service i on node j.

[0096] Services are dependent on services, data, and users. That is to say, there is data transmission between services, between data sources and services, and in user requests for services. Generally, the amount of data requested by users can be ignored. Therefore, the communication delay of a service can be defined as the sum of the delays caused by service dependencies and data dependencies, as shown in Equation (9).

[0097] comT i,j =sevT i,j +datT i,j Equation (9)

[0098] Among them, sevT i,j represents the delay caused by service dependencies of service i on node j, which can be expressed as the sum of the quotients of the dependency strength and the bandwidth of each hop. The dependency strength between services is defined as the number of calls between services divided by the total number of calls within the unit time of the service; datT i,j represents the delay caused by data dependencies of service i on node j, which can be expressed as the sum of the quotients of the data volume and the bandwidth of each hop.

[0099] Assume that a node can deploy all the services scheduled for it. That is to say, the memory capacity of the node can meet the task requirements. Also, due to the small difference in memory speeds, the execution time of a service is mainly affected by the CPU capacity of the node and is also related to the complexity of the service itself. Therefore, the execution delay of a service is defined as the ratio of the service instruction length to the processing capacity of the node CPU, as shown in Equation (10).

[0100]

[0101] Among them, mips j represents the CPU instruction execution speed of node j, cpuU i represents the CPU utilization rate of service i, mi i represents the instruction length of service i. The first two indicators can be obtained through a monitoring system, and the third indicator is prior knowledge.

[0102] There are a large number of services in distributed nodes. This application uses the mean value of the response times of all services to measure the delay of services in the cloud environment, as shown in Equation (11).

[0103]

[0104] Among them, i represents the service number, the total number of services is N, j represents the node number, and the total number of nodes is M.

[0105] (3) Scheduling Optimization Method Based on Deep Reinforcement Learning

[0106] Considering the factors affecting service scheduling and scheduling efficiency comprehensively, three constraints of control actions and two optimization objectives of auxiliary rewards are mainly considered in the service scheduling optimization process of this application, and the improved Deep Deterministic Policy Gradient (DDPG) algorithm is used to achieve the scheduling optimization of services.

[0107] 1) Model

[0108] Environmental description: The environmental state E t at the current service scheduling optimization task t is defined as a quadruple, E t =(t, S, N, SN), where t is the scheduling optimization task number; S is a vector representing the current service; N is a vector representing the current node; SN is a vector representing the current service deployment plan, and each element in SN is an ordered pair <s i , n j >, indicating that service s i is deployed to node n j .

[0109] Action description: The action of service scheduling is defined as the operation of deploying a service to a node in the current optimization scheduling task, denoted as a t =(s i , n j ), that is, service s i is deployed to node n j . Specifically, the selection of actions should conform to the restrictions of service dependency, data dependency, and user dependency constraints.

[0110] Reward description: According to the goal of service optimization scheduling, a service deployment plan with low service latency and high resource balance is a better plan. This application uses the sum of the differences in service latency and resource balance of all services before and after the action a t is completed as the basis for setting the reward value, and the calculation method is shown in formula (12). When the service latency decreases and the resource balance improves, the reward is positive, otherwise the reward is 0 or negative.

[0111] r t =(g t+1 -g t )-(avgT t+1 -avgT t ) Formula (12)

[0112] 2) Method implementation

[0113] The scheduling requirements of the service continue to arise, and the scheduling algorithm should have the ability of continuous control. In addition, the services on the cloud have a certain scale, the influencing factors are complex, and the scheduling is continuous, so the scheduling algorithm should have high efficiency. The Deep Deterministic Policy Gradient (DDPG) algorithm combines deep learning and reinforcement learning, and is proposed to solve the problem of continuous action control. It has the advantages of adapting to the solution of complex nonlinear problems, high solution efficiency, and supporting parallel computing, and can adapt to the solution of service scheduling problems in the cloud environment. Therefore, based on the deep deterministic policy gradient algorithm, a scheduling algorithm based on deep reinforcement learning is designed.

[0114] As Figure 2 shown, this application describes the service scheduling optimization problem based on deep reinforcement learning as a 5-tuple O t =(t, E t , a t , r t , E t+1 ), where t is a value between 1 and n, which is the task number of scheduling optimization; E t represents the state of the distributed environment at the t-th task moment; a t represents the action taken in the state E t , that is, the service scheduling scheme; r t represents the reward obtained by taking the action a t in the state E t ; E t+1 represents the new state obtained by taking the action a t in the state E t . The 5-tuples depicting each scheduling optimization task will be stored in the experience replay pool to provide a basis for model training.

[0115] To solve the problem of slow convergence, DDPG uses a dual-network structure composed of a policy network and a value network. The structure of the service scheduling optimization method based on DDPG is as Figure 2 shown. The input of the policy network is the state E t of the current environment, and the output is the corresponding action a t . The input of the value network is the state E t of the current environment and the action a t , and the output is the score of executing the action a t in the state E t . In addition, both the policy network and the value network are divided into an online network and a target network. The online network and the policy network have the same corresponding structure, but different initial parameters.

[0116] The specific process of the scheduling optimization method based on deep reinforcement learning is asFigure 3 As shown, first initialize the system parameters, construct a neural network, initialize the network weights and hyperparameters, and initialize the environment. Then, the state E obtained from the environment t is passed to the policy network, and the policy network outputs the corresponding action a t . Use a t as the input to the environment to obtain the corresponding reward r t and the state E at the next moment t+1 . Store (E t , a t , r t , E t+1 ) as a state transition data in the experience replay pool. Obtain a small batch of data from the experience pool and train the neural network. Continuously repeat the above process until the model converges. After that, based on this model, perform service scheduling tasks and continuously optimize.

Claims

1. A distributed microservice scheduling optimization method, characterized in that, it includes the following steps; Step 1: Based on the microservice instance, the distributed environment instance, and the deployment location of the microservice in the distributed environment, model the environment where the t-th scheduling task is located as a quadruple to describe the deployment of the microservice instance in the distributed environment instance; Step 2: Model the action of the t-th scheduling task as a binary tuple to describe the microservice s at this moment i deployed to node n j ; Step 3: With the goal of low service latency and high resource balance, use the sum of the differences in service latency and the differences in resource balance of all services after and before the action is completed for reward description; Step 4: Input the environment, action, and reward into the DDPG model for training to obtain a prediction model for the microservice instance scheduling optimization scheme; Step 5: Use the trained prediction model for the microservice instance scheduling optimization scheme to predict the scheduling optimization scheme of the microservice instance in the cloud environment instance to obtain the corresponding scheduling optimization scheme; The specific content of step 2 is as follows: Action description: Define the action of the t-th microservice scheduling optimization task as the operation a of deploying the service to a node in the current optimization scheduling task t =(s i ,n j ), that is, the service s i is deployed to the node n j . Among them, the selection of n j should meet the restrictions of service dependency, data dependency and user dependency constraints, facing the distributed service environment; The scheduling constraints of the service are designed as follows: Service dependency: The functions of microservices are relatively independent. In a distributed environment, user requirements are complex, and microservices need to cooperate with each other to meet user requirements. By analyzing the call relationships between microservices, various intensities of service dependencies are extracted from the link data of microservices to constrain the action set. Specifically, the number of call relationships between services is statistically obtained from the microservice call link data, and then the number of calls between services is divided by the total number of calls per unit time of the service as the intensity of service dependency. After that, when the service to be deployed has a strong service dependency relationship with other services, preferentially select the node where the strongly dependent service is located or an adjacent node; Data dependency: Data is the core element of the service, so data dependency is used to limit the action set. The location of the data source is obtained through prior knowledge, and the dependency relationship between the service and the data source is obtained through the parsing of the business process description file. It is described by a matrix. The rows of the matrix represent the node numbers where the services are located, and the columns of the matrix represent the node numbers where the data is located. If there is a dependency between the service and the data, the element with the service location node number and the data location node number as subscripts is 1, otherwise it is 0. After that, when the service to be deployed has a dependency relationship with the data, preferentially select the node where the data is located or an adjacent node; User dependency: Users are the main body that calls the service, so user dependency is used to limit the action set. The dependency relationship between the user and the service is obtained through prior knowledge and is described by a matrix. The rows of the matrix represent the node numbers where the services are located, and the columns of the matrix represent the node numbers where the users access. If there is a dependency between the service and the user, the element with the service location node number and the user access node number as subscripts is 1, otherwise it is 0. After that, when the service to be deployed has a dependency relationship with the user, preferentially select the node where the user is located or an adjacent node.

2. The distributed microservice scheduling optimization method according to claim 1, characterized in that, the specific content of step 1 is as follows: Environmental description: Define the environmental state E of the t-th microservice scheduling optimization task t as a quadruple, E t =(t, S, N, SN), where t is the serial number of the scheduling optimization task; S is a vector representing the current service, i.e., S = {s 1 , s 2 , …, s i , …, s n}; N is a vector representing the current node, i.e., N = {n 1 , n 2 , … n j , … n m}; SN is a vector representing the current service deployment plan, i.e., SN = (<s s1 , n n1 >, <s s2 , n n2 >, … <s sk , n nk >, … <s sNum , n nNum >), where each element is an ordered pair <s sk , n nk >, indicating that the service s sk is deployed to the node n nk .

3. The distributed microservice scheduling optimization method according to claim 1, characterized in that, the specific content of step 3 is as follows: Reward description: Based on action a t After completion and before completion, the sum of the differences in service latency and resource balance of all services is used as the basis for setting the reward value. The calculation method is shown in formula (1). When the service latency decreases and the resource balance improves, the reward is positive; otherwise, the reward is 0 or negative. r t = (g t+1 - g t ) - (avgT t+1 - avgT t ) Equation (1) As shown in formula (1), r t is the reward value after the execution of the t-th microservice scheduling optimization task, g t and g t+1 are the resource balance before and after the execution of action a t respectively, avgT t and avgT t+1 are the service delays before and after the execution of action a t respectively.

4. The distributed microservice scheduling optimization method according to claim 3, characterized in that, The calculation methods for resource balance and service latency are as follows: 1) Resource balance: Define resource balance based on CPU utilization rate and memory utilization rate; CPU utilization rate: The CPU utilization rate of a node is defined as the ratio of the allocated CPU resources in the node to the total CPU resources. The allocated CPU resources are calculated based on the CPU usage of the containers on the node, where n is the total number of containers on the node; Memory utilization rate: The memory utilization rate of a node is defined as the ratio of the allocated memory resources in the node to the total memory resources. The allocated memory resources are calculated based on the memory usage of the containers on the node, where n is the total number of containers on the node; The node resource balance is defined as the absolute value of the difference between the node CPU utilization rate and the memory utilization rate, as shown in formula (6); g i = |Ratio_Node cpu i –Ratio_Node mem i | Equation (6) Among them, i is the node number, and Ratio_Node cpu i represents the CPU utilization rate of the i-th node, and Ratio_Node mem i represents the memory utilization rate of the i-th node; The cloud environment contains a large number of nodes. The resource balance needs to comprehensively consider all nodes. The balance of a single node does not represent good performance of the container cloud. The variance of the node resource balances of all nodes is used to evaluate the resource balance of the cloud environment and assist in reward scoring, as shown in formula (7); Among them, is the mean value of the resource balance of each node, and g i is the resource balance of the i-th node; 2) Service latency: The service latency is represented by the sum of the communication latency and the execution latency, as shown in formula (8); T i,j = comT i,j + exeT i,j Formula (8) Among them, T i,j represents the service delay of service i on node j, and comT i,j represents the communication delay of service i on node j, which is generated by the network transmission of the services and data that service i depends on. exeT i,j represents the execution delay of service i on node j; The communication latency of a service is defined as the total latency generated by service dependencies and data dependencies, as shown in formula (9); comT i,j = sevT i,j + datT i,j Formula (9) Among them, sevT i,j represents the latency generated by service dependency of service i on node j, which is expressed by the sum of the quotient of the dependency strength and the bandwidth of each hop. The dependency strength between services is defined as the number of calls between services divided by the total number of calls within the unit time of the service; datT i,j represents the latency generated by data dependency of service i on node j, which is expressed by the sum of the quotient of the data volume and the bandwidth of each hop; The node deploys all the services scheduled to it. The execution latency of a service is defined as the ratio of the service instruction length to the processing capacity of the node CPU; There are a large number of services in the distributed nodes. The average value of all service response times is used to measure the latency of services in the cloud environment.

5. A distributed microservice scheduling optimization method according to claim 1, characterized in that the specific steps of step 4 are as follows: The microservice scheduling optimization problem based on deep reinforcement learning is described as a 5-tuple O t =(t, E t , a t , r t , E t+1 ), where t is a value between 1 and n, which is the task number for scheduling optimization; E t represents the state of the distributed environment at the t-th task moment; a t represents the action taken in state E t , that is, the service scheduling scheme; r t represents the reward obtained by taking action a t in state E t ; E t+1 represents the new state obtained by taking action a t in state E t . The 5-tuples depicting each scheduling optimization task will be stored in the experience replay pool to provide a basis for model training; DDPG uses a dual-network structure composed of a policy network and a value network. The input of the policy network is the state E of the current environment t , and the corresponding action a is output t ; the input of the value network is the state E of the current environment t and the action a t , and the output is the score of executing the action a in the state E t . In addition, both the policy network and the value network are divided into an online network and a target network. The online network has the same structure as the corresponding policy network, but different initialization parameters t .

6. A distributed microservice scheduling optimization method according to claim 1, characterized in that the specific steps of step 5 are as follows: Based on the trained model, input the environment of the current microservice scheduling optimization task and the scheduling constraints composed of the service dependency matrix, data dependency matrix, and user dependency matrix into the trained model, execute the model to generate a scheduling optimization plan, and feedback the reward value obtained from executing the current microservice scheduling optimization task to the environment for continuous optimization.