AI service and micro service mixed arrangement method and device, equipment and storage medium
By combining reinforcement learning agents and greedy strategies, the tight coupling problem of microservices and AI services in edge computing is solved, efficient service deployment and request routing is achieved, end-to-end delay is reduced, resource utilization is optimized, and user service quality is improved.
Patent Information
- Application Number
- CN202510852521.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-24
AI Technical Summary
In an edge computing environment, the hybrid orchestration strategy of microservices and AI services faces the problem of tight coupling of instance deployment and request routing, which leads to high complexity in decision making and it is difficult to achieve low-end to end latency and efficient resource utilization.
The reinforcement learning agent is used to jointly optimize service deployment and request routing. By determining the number of to be deployed, deploying microservice instances one by one and adapting to routing, and asynchronously deploying AI service instances in combination with greedy strategies, using Markov decision-making process and dual time scale concepts for orchestration, collecting decision trajectories for intelligent training.
It realizes efficient orchestration of AI services and microservices in edge computing environments, reduces end-to-end delays, optimizes resource utilization, and improves user service quality.
Smart Images

Figure CN120386606A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of service computing, and particularly to a method, apparatus, device, and computer-readable storage medium for hybrid orchestration of AI services and microservices. Background Art
[0002] In recent years, Generative Artificial Intelligence (GAI) has been developing rapidly. It provides Artificial Intelligence Generated Content (AIGC) for users through Pre-trained Foundation Models (PFMs), giving rise to many new types of Internet applications, such as ChatGPT and DALL-E 2. However, the complex architectures of these intelligent applications pose great challenges to maintenance. To ensure the Quality of Service (QoS) of users, the microservices architecture and the Model as a Service (MaaS) architecture provide an efficient solution. The microservices architecture can effectively handle different types of user requests by decomposing complex applications into multiple independent microservice instances, and at the same time reuse common service modules to improve the operation efficiency of the system. The MaaS architecture allows developers to provide pre-trained AI models as services to users, and users can call these trained models through API interfaces. These service computing architectures bring more flexible solutions to complex new network applications.
[0003] To further improve QoS, many current service providers are sinking services from the cloud to the edge to provide users with more low-latency response services. Different from cloud data centers, these edge servers are usually geographically distributed and resource-constrained, and cannot load all services on a single edge node. Therefore, these edge servers need to communicate and cooperate with each other to jointly provide services for various intelligent application programs. However, these intelligent application programs usually rely on the collaborative processing of multiple services, forming a complete service chain through sequential interactions between different microservices and calling AI services. Among them, key microservices and AI services will be called by multiple request flows, thereby forming a service call graph with complex dependencies. A large number of mutual calls and communications between microservices and AI services form a complex intelligent system. These service instances are deployed on edge servers and run for a long time to provide services for a large number of users. Furthermore, different user requests can call multiple instances of a single service to handle high-concurrency scenarios. The user request flow is dispatched to the request queues of multiple corresponding instances for queuing processing, and multiple instances exacerbate the complex dependencies between different services. To capture the complex service dependencies under a large number of users and multiple instances, we establish a queuing network model for steady-state delay analysis.
[0004] In a scenario with a large number of users and distributed systems, an efficient hybrid orchestration strategy for microservices and AI services can reduce the end-to-end latency of user requests and significantly improve the quality of user services. However, due to the existence of the above problems, we will face some challenges in actual orchestration. In fact, the hybrid orchestration strategy involves joint decision-making on service deployment and request routing, and these two decisions are tightly coupled. On the one hand, the end-to-end response latency of the request flow depends on both the deployment and routing strategies; on the other hand, the service deployment and request routing strategies are interdependent. Without determining the routing after deployment, it is impossible to judge the quality of the deployment strategy, and without determining the deployment strategy, routing cannot be carried out. The tight coupling relationship between instance deployment and request routing increases the decision-making complexity of the orchestration strategy. For instance deployment, in addition to deciding the deployment location, the number of required instances also needs to be determined to balance the interests of service providers and users and achieve the maximum utility. For users, the lower the end-to-end latency, the better the QoS. For service providers, overly demanding the minimum response time of requests will lead to the deployment of a large number of instances, resulting in huge resource overhead. In addition, in edge computing, user requests and the network computing environment are distributed, and multiple call paths are generated during the process of user requests accessing and invoking service instances. Specifically, each service may deploy multiple instances on different servers for load balancing. The request flow is dispatched to the request buffer queues of the corresponding service instances in different server clusters through request routing. The existence of request routing makes there be multiple call paths for the request flow in the edge network, further exacerbating the complexity of service calls. Considering the above complexities comprehensively, to improve the quality of user services, it is necessary to study an orchestration method for joint optimization of service deployment and request routing in a multi-instance queuing network of microservices and AI services in the edge environment. Summary of the Invention
[0005] To solve the above technical problems, the present application provides a method, apparatus, device, and computer-readable storage medium for hybrid orchestration of AI services and microservices.
[0006] In a first aspect, an embodiment of the present application provides a method for hybrid orchestration of AI services and microservices, where the method for hybrid orchestration of AI services and microservices includes: For each sub-service graph , determine the corresponding several microservice instances and one AI service instance , where is composed of a request service chain that invokes the same AI service instance ; Based on the request arrival rate , processing capacity and deployment cost corresponding to the sub-service graph, determine The number of units to be deployed; Request arrival rate based on the corresponding sub-service graph Each microservice instance in the several microservice instances Processing power and deployment cost determination The number of units to be deployed; based on The number of deployments to be made, using reinforcement learning agents Each The microservice instances are deployed one by one in sequence, and the temporary request routing is determined adaptively in a probabilistic routing manner to obtain All Temporary deployment and request results, To remove of ; based on The number of waiting deployments, using greedy strategy for asynchronous deployment , and update according to the deployment results All Request routing, get the sub-service graph The arrangement results; The decision trajectories generated by the reinforcement learning agent during the entire orchestration process are collected, and the agent training and neural network update are performed to obtain a fully trained agent.
[0007] In a second aspect, an embodiment of the present application provides an AI service and microservice hybrid orchestration device, the AI service and microservice hybrid orchestration device comprising: Determine the module for each sub-service graph ,Sure Corresponding microservice instances and an AI service instance ,in, By calling the same AI service instance Request service chain constitute; Quantity determination module, used to determine the request arrival rate based on the sub-service graph 、 Processing power and deployment cost determination The number of pending deployments; based on the request arrival rate corresponding to the sub-service graph Each microservice instance in the several microservice instances Processing power and deployment cost determination The number of units to be deployed; Orchestration module for The number to be deployed, using a reinforcement learning agent for each of the microservice instances are deployed one by one in sequence, and adaptively determine the temporary request routing in a probabilistic routing manner to obtain all of the temporary deployment and request results, is the removed ; Based on the number to be deployed, adopt a greedy strategy to asynchronously deploy and update all of the request routing in to obtain the orchestration result of the sub-service graph; An iteration module, configured to collect the decision-making trajectories generated by the reinforcement learning agent during the entire orchestration process, perform agent training and neural network update, and obtain a fully trained agent.
[0008] In a third aspect, an embodiment of the present application provides an AI service and microservice hybrid orchestration device, where the AI service and microservice hybrid orchestration device includes a processor, a memory, and an AI service and microservice hybrid orchestration program stored on the memory and executable by the processor. When the AI service and microservice hybrid orchestration program is executed by the processor, it implements the steps of the AI service and microservice hybrid orchestration method as described in the first aspect.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which an AI service and microservice hybrid orchestration program is stored. When the AI service and microservice hybrid orchestration program is executed by a processor, it implements the steps of the AI service and microservice hybrid orchestration method as described in the first aspect.
[0010] The beneficial effects brought by the technical solutions provided in the embodiments of the present application include: Through comprehensive analysis of user requests and edge node states, the present application realizes efficient orchestration of AI services and microservices. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a schematic flowchart of an embodiment of the AI service and microservice hybrid orchestration method of the present application; Figure 2 is a schematic diagram of the functional modules of an embodiment of the AI service and microservice hybrid orchestration device of the present application; Figure 3 is a schematic hardware structure diagram of the AI service and microservice hybrid orchestration device involved in the solution of the embodiment of the present application. DETAILED DESCRIPTION
[0012] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts belong to the scope of protection of this application.
[0013] To make the purpose, technical solution and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the drawings.
[0014] In a first aspect, an embodiment of this application provides a method for hybrid orchestration of AI services and microservices.
[0015] In the intelligent edge computing scenario, there are various intelligent application requests from users. Service orchestration is to jointly make decisions on the fine-grained deployment of various service instances required by users and request routing. For example, in the intelligent edge computing scenario, there may be multiple types of intelligent application programs, such as intelligent customer service, chatbots, AI image generation, etc. These programs need to communicate and interact with multiple different AI models and microservices to jointly complete the service. Therefore, we need to perform service instantiation deployment according to user requests and determine the request routing based on the call relationship. The goal of our service orchestration is to balance the user service quality and the operator cost, that is, to minimize the weighted value of the user end-to-end response latency and the deployment cost.
[0016] Specifically, the intelligent edge computing scenario usually includes multiple distributed server nodes. Each node contains a certain amount of computing resources for various services to perform instantiation deployment, and the edge nodes cooperate through communication links. In addition, there are edge user requests. Considering the geographical distribution characteristics of user requests, users access the network through the nearest edge node, which is jointly completed by multiple microservice instances and one AI service instance in the chain call order. And a batch of requests for calling the same AI service constitutes a sub-service call graph, and we model a batch of user request traffic as a request flow that follows a Poisson distribution. The system uploads each request within the current service area range and the status data of each edge server to the central control system through the implementation of monitoring technology. The central system performs service orchestration based on the edge node status and user requests. And the orchestration process needs to solve the following problems: 1. Determine the number of instances required for each service in the current system based on the processing capacity and deployment cost of each service.
[0017] 2. When orchestrating all request chains in each sub-service graph, it is necessary to jointly optimize service deployment and request routing involved in the orchestration process.
[0018] 3. Ensure that when arranging, the service instances deployed on any edge node do not violate the resource constraints, and each service instance meets the basic service rate.
[0019] 4. On the basis of meeting the above constraints, minimize the weighted sum of the total service time and deployment cost of all requests.
[0020] The embodiment of the present application provides a service orchestration algorithm based on reinforcement learning. Through comprehensive analysis of user requests and edge node states by the central controller, efficient service orchestration is achieved. During the implementation process of the present application, the central controller first collects relevant attribute information of each edge server in the current system, including the remaining computing resources and the deployed instance information; collects user request information, which includes the user access node, service chain composition, request arrival rate, etc. of the request. Through comprehensive collection of this information, the central controller can formulate a more comprehensive global optimization orchestration strategy.
[0021] As Figure 1 shown, at the beginning stage of orchestration, the central controller determines the number of instances required for each type of service according to the service instance number calculation method, that is, the number of AI service instances required for each sub-service graph, and the number of various microservices called in all request sub-chains in the sub-service graph. During the execution of the orchestration strategy, the central controller deploys the microservice instances of each request sub-chain in the order of service chain calls for the sub-service graph to be orchestrated currently based on the reinforcement learning agent. After all the microservice instances of the current sub-chain are deployed, the central controller will adaptively determine the request routing based on the deployment result. Specifically, to balance the load of each deployed microservice instance, the request flow will evenly distribute the arrival traffic according to the currently deployed instance number to achieve load balancing of the microservice instances. After each request sub-chain is orchestrated, the edge node state will be updated. The central controller will continuously monitor and read the states of each node and perform the orchestration of the next request sub-chain until all request sub-chains in the current sub-service graph are orchestrated, that is, determine the microservice deployment locations and their respective request routing methods for all current request sub-chains.
[0022] Next, orchestrate the AI services that have not been deployed in the current sub-service graph. Specifically, we count the microservice instances located at the adjacent call positions of the current AI service, and then use the greedy strategy to iteratively select edge nodes. Under the condition of meeting the resource constraints of the edge nodes, select the edge node with the minimum additional communication time with the adjacent microservice for the actual deployment of the AI service instance. After the deployment of the AI service instance is completed, update the request routing of each request in the sub-service graph according to the deployment location.
[0023] The above process will continue to loop until all sub-service graph orchestrations in the system are completed. After the entire service orchestration operation is completed, the central controller will collect the decision trace samples generated during the entire service orchestration process to obtain training experiences, and the reinforcement learning agent in the central controller will perform model training based on the obtained experiences to enhance the decision-making ability of the agent.
[0024] In one embodiment, the AI service and microservice hybrid orchestration method includes: Step S10, for each sub-service graph , determine the corresponding several microservice instances and one AI service instance , where is composed of the request service chains that call the same AI service instance ; In this embodiment, the edge server contains a variety of computing resources, is geographically distributed and can communicate with each other: For a geographically distributed multi-edge server system, it is represented by an undirected graph . This system contains multiple edge nodes, denoted as , where represents the th edge node. And the edge node contains a large amount of computing and storage resources, including CPU, memory, and GPU resources, denoted as , , . These resources are available for AI models and microservices to be instantiated and deployed in the server. Let represent the coordinates of each edge node distributed in the planar geography, and the attribute set of each edge node is .
[0025] Edge nodes transmit data through the network to collaborate on computing tasks. We use to represent the set of links between any two edge nodes. The network environment between edge nodes may be heterogeneous due to geographical environment, infrastructure, or different communication protocols between different operators, resulting in different communication qualities of the links between edge nodes. We represent the transmission bandwidth between two edge nodes and as , and the transmission bandwidth of the entire system is:
[0026] We represent the physical distance between two edge nodes as , obtained from the Euclidean distance calculation formula , and the physical distance of the entire system is:
[0027] Specifically, the communication between service instances within the same edge node will be completed via high-speed optical fiber communication inside the node. Therefore, , and since the position coordinates of each node are fixed, therefore . The set of attributes for each link is .
[0028] In a microservices architecture, due to the reusability of service instances, the entire intelligent application can be modeled as a directed acyclic graph (DAG), which we denote as to represent this service graph. The intelligent application can support different functionalities, and a functionality can be decoupled into several service instances with data dependencies. These service instances call adjacent service instances through data communication to form a service function chain. We use the symbol to represent the set of all microservices, and the symbol to represent the set of all AI services. For the convenience of later narration, we uniformly use the service to represent the set of microservices and AI services, that is .
[0029] Each user request flow can be served by a corresponding functionality, which we represent using a chain structure as: , and we define that each request is completed by a combination of multiple microservice instances and one AI service instance. The AI service instances are usually completed in batch form. One AI service instance may be called by multiple requests, and the requests that call the same AI service instance are similar. There are several key reusable microservice instances, so the service graph forms several sub-service call graphs. Each sub-service graph is composed of the request service chains that call the same AI service instance, represented as:
[0030] Furthermore, the service graph can be decomposed into multiple sub-service graphs that call different AI service instances, that is:
[0031] To be closer to the real network environment, we model the traffic of a batch of user requests in each subgraph as a request flow with a Poisson distribution for the arrival rate, denoted as , and all the user arrival rates . At the same time, considering the geographical distribution characteristics of user requests, the geographical location of the user is represented by , and the user accesses the network through the edge node closest to the user. The access node of this request flow Can be expressed as . In addition, since the user request needs to go through the inference calculation of the AI service instance to output content, the user input, also known as the prompt, is expressed as . The AI model performs iterative calculations to generate content, and the expected number of iterations is expressed as . Therefore, each user request flow can be represented by a five-tuple .
[0032] Step S20: Determine the number of instances to be deployed based on the request arrival rate corresponding to the sub-service graph , processing capacity and the deployment cost of the instances to be deployed; Step S30: Determine the number of instances to be deployed based on the request arrival rate corresponding to the sub-service graph , the processing capacity of each of the several types of microservice instances and the deployment cost of the instances to be deployed; In this embodiment, step S20 includes: constructing a function based on the request arrival rate corresponding to the sub-service graph , processing capacity and the deployment cost. The function takes the number of instances to be deployed as the independent variable, and takes the value of the independent variable corresponding to the minimum function value of as the number of instances to be deployed of .
[0033] Step S30 includes: constructing a function based on the request arrival rate corresponding to the sub-service graph , the processing capacity of each of the several types of microservice instances and the deployment cost. The function takes the number of instances to be deployed as the independent variable, and takes the value of the independent variable corresponding to the minimum function value of as the number of instances to be deployed of .
[0034] Specifically: The processing capacity of the microservice instance is expressed as , which can be obtained by dividing the required processing computation by the CPU computing power (i.e., the number of calculations that the CPU can perform per unit time, which can be obtained by monitoring the server performance), that is .
[0035] The processing capacity of the AI service instance is expressed as the number of batch requests that can be processed per unit time, which is mainly determined by the model architecture , the input of user request prompts within a batch , the expected number of iterations and the computing power of the model after instantiation , that is, the FLOPs that can be calculated per unit time. Generally, we use a function to represent the processing capacity of the model for a certain request:
[0036] Taking the large language model with a typical decoder-only architecture as an example, a complete inference process can be divided into two stages: the initialization stage and the autoregressive stage. In the initialization stage, the user's prompt is input into the model and passes through all the transformer layers in sequence to complete the calculation and generate the first output word. In the autoregressive stage, the model uses the newly generated token as the input and re-enters it into the model for iterative calculation. Each iteration, the model generates a new token until the generated sequence reaches the expected number of iterations. For the common large model GPT-3 architecture, it is stacked by multiple transformer blocks, and each module consists of a self-attention module and a feed-forward neural network. Its model architecture can be abstractly represented as , and represent the hidden layer dimensions of the self-attention module and the feed-forward neural network respectively, represents the number of transformer blocks. In the initialization stage, for the prompt token with a length of , the FLOPs required for it to complete the calculation on all layers of the model can be expressed as:
[0037] In the autoregressive stage, the newly generated output word each time will be re-entered into the large model until the complete output is generated after iterating to the expected number of times. Therefore, the computational complexity of the autoregressive stage can be expressed as:
[0038] For a single request, the required computational complexity is the sum of the computational complexities of the initialization stage and the autoregressive stage, that is , so the total computational complexity within a batch is , and the processing capacity of the AI service instance is expressed as the number of requests that can be processed per unit time, that is, the ratio of the computing power of the model after instantiation to the computational complexity required for a batch, which can be expressed as:
[0039] The deployment of each service instance incurs certain costs, including the deployment costs of microservice instances and AI service instances in all requested service chains. Their instantiation and deployment consume a certain amount of computing resources and generate corresponding expenses. Since the edge nodes are managed by the same cloud provider, the costs of using resources from different edge nodes are the same. Let and represent the total numbers of AI service instances and microservice instances deployed in all edge nodes respectively, that is, , . The deployment cost of microservices is determined by the consumption of computing and storage resources. The total deployment cost of microservices can be expressed as:
[0040] Among them, the unit cost of microservice instance is determined by the CPU and memory resources it consumes, specifically , and are the weight factors of resource costs, and represent the CPU and memory resources consumed by microservice respectively. is the number of microservice instances deployed in all edge nodes.
[0041] Similarly, the unit cost of AI service instance depends on the usage of CPU, memory, and GPU resources, specifically , , and are the weight coefficients of CPU, memory, and GPU resources respectively. The total deployment cost of AI service instances can be expressed as:
[0042] is the number of deployed in all edge nodes.
[0043] Each edge node has available resource capacities, including CPU resources , storage resources and GPU resources . The microservice instances and AI service instances instantiated and deployed on the edge nodes must ensure that they do not exceed the resource capacities. Edge node The resource limit on , as well as Deployed on edge nodes Microservice instance and AI service instances The number of and , the number of instances on each edge node must meet the following resource limits:
[0044]
[0045]
[0046] Total cost of the system The total deployment cost of the microservice instance and the AI service instance is as follows:
[0047] For the service delay of microservices, due to the reusability of microservices, edge nodes Multiple microservice instances on the same node can collaborate to call multiple requests, and the corresponding nodes Microservices on The workload can be represented as all the microservices The total request rate of all call paths of the request flow deployed on the edge node is expressed as ,Since the user request flow is continuous and an instance can be regarded as a service node to complete the user's request, we use the multi-instance queuing theory M\M\C to analyze the service delay of microservices in the queuing network. The processing capacity of the microservice instance is expressed as ,node Microservices The number of instances is , first, if the request arrival rate Beyond edge nodes Microservice instance Total processing capacity , will lead to infinitely long queues, thus seriously reducing the quality of network service. Therefore, service intensity Should be less than 1. Edge node Microservice instance on The service intensity is defined as:
[0048] According to the little formula, the edge node Microservice instance on The queuing delay can be expressed as:
[0049] In the formula is expressed as:
[0050] And the total service time is equal to the queuing time plus the processing time, which is expressed by the following formula:
[0051] For each microservice , we need to balance the service delay of users on the microservice and the service cost brought by deploying the microservice . For each microservice instance in the service chain with an arrival rate of , the expression we need to balance is as follows
[0052] is the weight coefficient for balancing user delay and service cost.
[0053] For the service performance of the AI service for a batch of requests, this application adopts the M / M / 1 queuing model and the arrival rate shunting mechanism, that is, according to the arrival rate of each batch of requests, it is evenly distributed to all AI service instances on the edge nodes. Let represent the number of all AI service instances in all current edge nodes, then the arrival rate allocated to each instance can be expressed by the following formula:
[0054] Then the average service delay of the AI service can be expressed as:
[0055] For each sub-service graph corresponding to the called AI service instance , the formula we need to balance is as follows:
[0056] Next, based on the idea of gradient descent, we iteratively determine the set of all microservice instance numbers required for each request sub-chain after removing the AI service , and the number of AI service instances required for each . The specific steps are: set a minimum initial solution that satisfies the basic service rate. For and , are respectively set as and , and then continuously increase the number of instances until, when increasing to a certain value, further increasing the number of instances causes the optimization objective to increase rather than decrease, that is, the following formula is satisfied:
[0057]
[0058] The number of instances at this time is the local minimum and also the desired instance value.
[0059] Step S40, based on the number of deployments to be made, use the reinforcement learning agent to each microservice instance in all is deployed one by one in sequence, and adaptively determines the temporary request routing in a probabilistic routing manner to obtain All temporary deployments and request results in ; In this embodiment, the decision-making process of the reinforcement learning agent can be modeled by a Markov decision process MDP. Specifically, the agent continuously interacts with the environment. The agent observes the environmental state at the current moment, and based on the state makes a decision through the policy network and outputs an action . The environment performs a state transition according to the agent's decision, changing from to . The agent will obtain an immediate reward after executing the action to evaluate the quality of the policy. The continuous interaction between the agent and the environment forms a trajectory sequence, that is, . The purpose of the agent is to make decisions at each moment to maximize the long-term reward . Among them, the design of the MDP quadruple is as follows: : At time , the state observed by the agent is expressed as:
[0060] Among them, represents the remaining resources of the edge node , represents the information on the number of microservice instances already deployed on each edge node at the current moment, It includes the resource requirements of the microservice instances to be deployed currently, the micro-instance numbers, and the information of the adjacent called micro-instances in the request service chain, including the instance numbers and deployment locations; : At time , the action of the agent is to select the edge node for deploying the microservice instance at the current deployment time, and the action space is represented as a one-dimensional discrete vector: , where is a binary variable indicating whether to deploy the microservice instance to the edge node , satisfying the constraint: , that is, each microservice instance can only be deployed to one edge node, and the agent outputs the deployment probability distribution of all edge nodes , where represents the probability of deploying the microservice instance to the edge node . By sampling this probability distribution, the actual deployment node is determined. When the instance deployment is completed, the agent performs request routing to determine the routing probability value of the current request path ; : The agent will obtain an immediate reward after making a decision . The goal of the agent is to optimize the deployment strategy as much as possible to obtain the maximum cumulative reward. Due to the increase in the request flow, the decision-making process of the agent also increases. To alleviate the problem of sparse rewards caused by long-term decisions, we give an immediate reward based on the cumulative delay of the request flow at the current deployment time. That is, the agent will obtain an immediate reward after making a decision , represents the reward obtained by deploying the th microservice instance in order . The reward is determined by the cumulative request flow delay and cost after orchestration, and is expressed as:
[0061] where, , , represents the total number of microservice instances to be deployed by the current request flow, represents the length of The expression is as follows:
[0062] where, Indicates the average response latency of the currently deployed request flow, which is the minimum average response latency at the current deployment moment in the historical training rounds; : The state transition function describes the process of state change after the agent executes an action. Specifically, the action of the agent is to deploy a microservice instance to a specific edge node, and the state value is updated according to the deployment of the instance on the edge node and the corresponding resource consumption. After the action the new state is represented by the following update:
[0063] where the function includes the update of the remaining resources of the edge node, the update of the information on the number of deployed microservice instances, and the update of the information on the microservice instances to be deployed.
[0064] Step S50, based on the number of microservices to be deployed, asynchronously deploy using a greedy strategy and update, according to the deployment result, all request routes in to obtain the orchestration result of the sub-service graph; In this embodiment, the entire orchestration process is based on reinforcement learning for microservice instance deployment, adaptively determines the request route in a probabilistic routing manner, and finally completes the AI service instance orchestration asynchronously.
[0065] First, split the sub-service call graph into multiple service chains, and then for the sub-chain corresponding to the current request, deploy the calculated microservice instances, and then calculate the routing path through probabilistic routing to obtain the temporary latency and deployment cost. After completing the orchestration of all microservice instances in the current sub-service graph, deploy the AI service instances based on the greedy strategy, and update the complete routing path and complete request latency of all request flows in the sub-service graph.
[0066] Among them, the orchestration strategy constitutes two time scales of large and small, that is , each large time period contains multiple small time periods , , where . Within each large time scale, deploy and orchestrate all instances in , and within each small time scale, for the request sub-chain in Arrange and gradually deploy microservice instances in sequence to complete the arrangement. Finally, on a large time scale, complete the deployment of AI service instances that have not been deployed and update the request routing.
[0067] Specifically, for the request flow in the service graph , deploy microservice instances. The call order is determined by . Within each small time scale , the RL Agent calculates, based on the policy network, the edge node selected for deployment of the current microservice instance at the current deployment time slot. According to the calculated number of all microservice instances, deploy them to the edge nodes in sequence. Meanwhile, to prevent the deployment policy from violating resource constraints, an additional action mask is introduced to mask the possible illegal actions of the agent at the current time slot. For the microservice instance , the reinforcement learning agent calculates the policy distribution based on the policy network as . The additional action mask prevents the deployment operation from violating the resource limits of the edge node, which is obtained based on the remaining resources of the edge node at the current time slot, that is ; for each edge node , define its mask value as follows:
[0068] Update the policy distribution according to the mask:
[0069] Sample the new policy distribution to obtain the deployment policy , update the remaining resources of the corresponding edge node, and perform a state transition .
[0070] Next, when the instance deployment is completed, the agent performs request routing to determine the routing probability value of the current request path , which is specifically: According to the deployment results of the microservice instances in the current , calculate the temporary routing path in the form of routing probability:
[0071] One of the call paths is represented as , and the calculation of the call path probability value is actually the product of the forward probabilities of each edge node on the path. Define that the probability of selecting the next edge node will be determined by the number of corresponding service instances deployed on the edge node. Similarly, the access splitting probability of the request also follows this policy, that is:
[0072]
[0073]
[0074] Among them represents the number of microservice instances deployed at the edge node , and represents the shunt arrival rate corresponding to the call path , which is obtained by calculation
[0075] As described above, the orchestration strategy constitutes two time scales, large and small. After all the microservice instances in the current time scale corresponding are deployed, update the corresponding instance number matrix and the arrival rate matrix :
[0076]
[0077] Based on the currently deployed results, that is, the instance number matrix and the arrival rate matrix of the edge node environment, calculate the residence service time of each microservice on the call path , and update the delay of the previously deployed request flow; the transmission bandwidth of the virtual link between two edge nodes is , and the amount of transmitted data per unit arrival rate between two types of services is , for the microservice instance located on the edge node , and the corresponding arrival rate is , its transmission delay is expressed as:
[0078] The propagation delay depends on the physical distance between two edge nodes , and the propagation speed of the signal in the air medium is expressed as , and the propagation delay is expressed as:
[0079] The communication time of adjacent microservice calls is given by the following formula:
[0080] When the microservice call location is adjacent to the AI service, the two types of microservices do not communicate directly. Therefore, the communication delay in this section is temporarily 0, that is , ; The temporary delay of is weighted by the delays of all call paths:
[0081] After the current orchestration is completed, calculate the reward and update the deployment trajectory of the microservices: , , where , and and are the calculation results of the policy network and value network of the agent at the current moment, respectively. The current small time scale ends and enters the next orchestration.
[0082] Next, after completing all in the current sub-service graph orchestration, asynchronously deploy the AI service instances that have not been deployed. Specifically, count the microservice instances located at the adjacent call locations of the AI service instances in the current sub-service graph , and then use the greedy strategy to iteratively select edge nodes, calculate the additional service time, communication time with adjacent microservice instances, and additional cost generated by deploying the AI service instances on the selected edge nodes. Among them, the additional service time , additional communication time , and additional deployment cost are given by the following formulas respectively:
[0083]
[0084]
[0085] Among them, , ; Since the additional service time is not affected by the deployment strategy, select the edge node that minimizes the sum of the additional communication time and additional cost for the actual deployment of the AI service instance, that is:
[0086] Under the current time scale , after the AI service instance is deployed in , its newly added deployment node set is represented as in all The routing, where one call path is updated to: , and the corresponding routing probability is also updated to satisfy , and the shunt request arrival rate of each corresponding call path is also updated to satisfy , and the call path set is updated to , satisfying , and and respectively represent the number of call paths before and after the update, and update accordingly The end-to-end delay, that is, the service delay of each will increase the AI service delay and communication time:
[0087] When the current large time scale ends, it enters the next sub-service graph scheduling.
[0088] Step S60: Collect the decision trajectories generated by the reinforcement learning agent during the entire scheduling process, perform agent training and neural network update, and obtain a fully trained agent.
[0089] In this embodiment, the experience sampled from is used to calculate the advantage function at each moment according to the GAE formula , where is the discount factor, is the trade-off coefficient between bias and variance, and then calculate the target value function at each moment:
[0090]
[0091] Calculate the policy loss and the value loss :
[0092]
[0093] where represents the ratio of the new policy to the old policy, , and limit it within the range of to ensure that the target policy update does not deviate too much from the sampling policy, and update the parameters with gradient ascent and gradient descent respectively:
[0094]
[0095] The agent is continuously trained with decisions until convergence.
[0096] The embodiment of the present application provides a hybrid service orchestration method for AI services and microservices based on the Maas paradigm in an edge computing scenario, aiming to solve the complexity problem caused by the mutual coupling of service deployment and request routing decision-making involved in heterogeneous service orchestration. For the edge computing scenario, this method designs a multi-service call graph orchestration strategy based on the deep reinforcement learning method to minimize the weighted loss of end-to-end response latency and service deployment cost. The method determines the number of AI service instances required for each service graph and the number of microservice instances of each type in each request chain under each service graph based on the idea of weighing service latency and deployment cost; establishes the orchestration process as a Markov decision process, introduces the concept of double time scales, and based on the reinforcement learning agent, sequentially orchestrates each request sub-chain in each sub-service graph in the system and continuously updates the edge node state value in real time, asynchronously deploys the AI service instances corresponding to the sub-service graph based on the greedy strategy and updates the request routing of each request; trains the reinforcement learning agent by collecting the decision-making trajectories of the orchestration strategy, continuously updates the neural network parameters of the agent, and finally realizes the globally optimal orchestration scheme.
[0097] In a second aspect, the embodiment of the present application also provides an AI service and microservice hybrid orchestration device.
[0098] In one embodiment, referring to Figure 2 , Figure 2 is a schematic diagram of the functional modules of an embodiment of the AI service and microservice hybrid orchestration device of the present application. As Figure 2 shown, the AI service and microservice hybrid orchestration device includes: A determination module 10, configured to determine for each sub-service graph the corresponding several types of microservice instances and one AI service instance , where is composed of request service chains that call the same AI service instance ; A quantity determination module 20, configured to determine the quantity to be deployed based on the request arrival rate corresponding to the sub-service graph, the processing capacity and the deployment cost; determine the quantity to be deployed based on the request arrival rate corresponding to the sub-service graph, the processing capacity of each type of microservice instance among the several types of microservice instances and the deployment cost; An orchestration module 30, configured to, based on the quantity to be deployed, use the reinforcement learning agent to Each in the microservice instances are deployed one by one in sequence and adaptively determine the temporary request routing in a probabilistic routing manner to obtain all in the temporary deployment and request results, is the one after removing of ; Based on the number of microservices to be deployed, adopt a greedy strategy to asynchronously deploy , and update all in the request routing to obtain the orchestration result of the sub-service graph ; The iteration module 40 is used to collect the decision-making trajectories generated by the reinforcement learning agent during the entire orchestration process, perform agent training and neural network update to obtain a fully trained agent.
[0099] Among them, the function implementation of each module in the above AI service and microservice hybrid orchestration device corresponds to each step in the above AI service and microservice hybrid orchestration method embodiment, and its function and implementation process will not be elaborated here one by one.
[0100] In a third aspect, an embodiment of the present application provides an AI service and microservice hybrid orchestration device. The AI service and microservice hybrid orchestration device can be a device with data processing functions such as a personal computer (PC), a laptop computer, a server, etc.
[0101] Referring to Figure 3 , Figure 3 is a schematic hardware structure diagram of the AI service and microservice hybrid orchestration device involved in the embodiment solution of the present application. In the embodiment of the present application, the AI service and microservice hybrid orchestration device may include a processor, a memory, a communication interface, and a communication bus.
[0102] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.
[0103] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, etc., which are used to implement the interconnection of components inside the AI service and microservice hybrid orchestration device, and interfaces for implementing the interconnection of the AI service and microservice hybrid orchestration device with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display screen (Display), a keyboard (Keyboard), etc.
[0104] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0105] The processor can be a general-purpose processor, which can call the AI service and microservice hybrid orchestration program stored in the memory and execute the AI service and microservice hybrid orchestration method provided by the embodiments of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the AI service and microservice hybrid orchestration program is called can refer to the various embodiments of the AI service and microservice hybrid orchestration method of the present application, which will not be elaborated here.
[0106] Those skilled in the art can understand that Figure 3 the hardware structure shown in
[0107] In a fourth aspect, the embodiments of the present application further provide a computer-readable storage medium.
[0108] The computer-readable storage medium of the present application stores an AI service and microservice hybrid orchestration program, and when the AI service and microservice hybrid orchestration program is executed by a processor, the steps of the AI service and microservice hybrid orchestration method as described above are implemented.
[0109] Among them, the method implemented when the AI service and microservice hybrid orchestration program is executed can refer to the various embodiments of the AI service and microservice hybrid orchestration method of the present application, which will not be elaborated here.
[0110] It should be noted that the serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0111] In the description of the specification, claims and the above-mentioned drawings of this application, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. Descriptions such as "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequential order, nor do they limit that "first", "second" and "third" are of different types.
[0112] In the description of the embodiments of this application, words such as "exemplary", "for example" or "for illustration" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary", "for example" or "for illustration" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example" or "for illustration" is intended to present relevant concepts in a specific manner.
[0113] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.
[0114] In some processes described in the embodiments of this application, a plurality of operations or steps appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of this application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.
[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions to enable a terminal device to execute the methods described in the various embodiments of this application.
[0116] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. An AI service and microservice hybrid orchestration method, characterized in that The AI service and microservice hybrid orchestration method includes: For each sub-service graph , determine the corresponding several microservice instances and an AI service instance , where is composed of a request service chain that calls the same AI service instance ; Based on the request arrival rate corresponding to the sub-service graph , processing capacity and determine the deployment cost of the number to be deployed; Based on the request arrival rate corresponding to the sub-service graph , for each microservice instance among the several microservice instances processing capacity and the deployment cost are determined the number to be deployed; Based on the number of microservices to be deployed, use a reinforcement learning agent to sequentially deploy each microservice instance in and adaptively determine the temporary request routing in a probabilistic routing manner to obtain the temporary deployment and request results of all Remove from ; Based on the number to be deployed, use the greedy strategy to deploy asynchronously and update according to the deployment result all request routes in to obtain the orchestration result of the sub-service graph; Collect the decision-making trajectories generated by the reinforcement learning agent during the entire orchestration process, perform agent training and neural network update to obtain a fully trained agent.
2. The AI service and microservice hybrid orchestration method according to claim 1, wherein Based on the request arrival rate corresponding to the sub-service graph , processing capacity and determine the deployment cost The number to be deployed includes: Based on the request arrival rate corresponding to the sub-service graph , processing capacity and the deployment cost to construct a function , the function uses the number to be deployed as the independent variable, and takes the value of the independent variable corresponding to the minimum function value of as the number to be deployed; Based on the request arrival rate corresponding to the sub-service graph , for each microservice instance among the several microservice instances processing capacity and deployment cost determination The number to be deployed includes: Based on the request arrival rate corresponding to the sub-service graph , for each microservice instance among the several microservice instances processing capacity and deployment cost to construct a function , the function takes the number of to-be-deployed as the independent variable, and takes the value of the independent variable corresponding to the minimum function value of as the number of to-be-deployed.
3. The AI service and microservice hybrid orchestration method according to claim 1, wherein The decision-making process of the reinforcement learning agent can be modeled by a Markov decision process MDP, and the design of the MDP quadruple is as follows: : At time , the state observed by the agent is represented as: Among them, represents an edge node remaining resource quantity represents the information on the number of microservice instances deployed at each edge node at the current moment includes the resource requirements for microservice instances to be deployed currently, the microservice instance numbers, and the information on adjacent calling microservice instances in the request service chain they are in, including the instance numbers and deployment locations; : At time , the action of the agent is to select the edge node for deploying the microservice instance at the current deployment time, and the action space is represented as a one-dimensional discrete vector: , where is a binary variable indicating whether to deploy the microservice instance to the edge node and satisfies the constraint: , that is, each microservice instance can only be deployed to one edge node. The agent outputs the deployment probability distribution of all edge nodes , where represents the probability of deploying the microservice instance to the edge node . By sampling this probability distribution, the actual deployment node is determined. When the instance deployment is completed, the agent performs request routing to determine the routing probability value of the current request path . : The agent makes a decision and then receives an immediate reward , indicating the reward obtained by the th microservice instance deployed in sequence, and the reward is determined by the cumulative request flow delay and cost after orchestration, expressed as: Among them, , , representing the total number of microservice instances to be deployed in the current request stream, represents the length of, The expression is as follows: Among them, represents the average response latency of the currently deployed request flow, which is the minimum average response latency at the current deployment moment during historical training rounds; : The state transition function describes the process of state change after the agent executes an action. Specifically, the action of the agent is to deploy a microservice instance to a specific edge node, and the state value is updated according to the deployment of the instance on the edge node and the corresponding resource consumption. After the action the new state is represented by the following update: Among them, the function includes the update of the remaining resources of edge nodes, the information update of the number of deployed microservice instances, and the information update of the microservice instances to be deployed.
4. The AI service and microservice hybrid orchestration method according to claim 3, wherein Introduce an additional action mask to shield the possible violation actions of the agent in the current time slot for the microservice instance , the reinforcement learning agent calculates the policy distribution based on the policy network as , the additional action mask prevents deployment operations from violating the edge node resource limits, which is obtained based on the remaining resources of the edge node in the current time slot, i.e., ; for each edge node , define its mask value as follows: Update the policy distribution according to the mask: Obtain the deployment policy by sampling the new policy distribution , update the remaining resources of the corresponding edge node, and perform state transition .
5. The AI service and microservice hybrid orchestration method according to claim 4, wherein When After the instance deployment is completed, the agent performs request routing to determine the routing probability value of the current request path including: including: According to the current microservice instance deployment result, calculate the temporary routing path in a routing probability manner: One of the call paths is represented as , and the call path probability value The calculation is actually the product of the forward probabilities of each edge node on the path. It is defined that the probability of selecting the next edge node will be determined by the number of corresponding service instances deployed on the edge node. Similarly, the access diversion probability of the same request Also follows this strategy, that is: Among them indicates the number of microservice instances deployed at the edge node, and indicates the shunt arrival rate corresponding to the call path, which is calculated.
6. The AI service and microservice hybrid orchestration method according to claim 5, characterized in that The scheduling strategy consists of two time scales, namely , and each large time period contains multiple small time periods , , among which , within each large time scale , all instances are deployed and scheduled, and within each small time scale among them, in is scheduled, and the microservice instances are deployed step by step in order and the scheduling is completed. Finally, at the large time scale, the deployment of the AI service instances that have not been deployed and the update of the request routing are completed. After the deployment of all microservice instances corresponding to the current time scale is completed, the corresponding instance number matrix and the arrival rate matrix and are updated: Calculate the residence service time of each microservice on the call path based on the instance number matrix and arrival rate matrix in the edge node environment , and update the latency of the previously deployed request flow; the transmission bandwidth of the virtual link between two edge nodes is , and the amount of transmitted data under the unit arrival rate between two types of services is , for the microservice instance located on the edge node , and the corresponding arrival rate is , its transmission latency is expressed as: The propagation delay depends on the physical distance between two edge nodes , and the propagation speed of the signal in the air medium is expressed as , and the propagation delay is expressed as: The communication time between adjacent microservice calls is given by the following formula: When the microservice call location is adjacent to the AI service, the two types of microservices do not communicate directly. Therefore, the communication delay in this section is temporarily 0, that is , ; The temporary delay of is weighted by the delays of all call paths: After the current Calculate the rewards after the scheduling ends and update the deployment track of the microservices: , , where , and and are the calculation results of the policy network and value network of the agent at the current moment respectively. When the current small time scale ends and enters the next scheduling.
7. The AI service and microservice hybrid orchestration method according to claim 6, wherein After completing the orchestration of all in the current sub-service graph asynchronously deploy the AI service instances that have not been deployed. Specifically, count the microservice instances in the current sub-service graph at the adjacent call positions of the AI service instances, and then use the greedy strategy to iteratively select edge nodes, calculate the additional service time, the communication time with adjacent microservice instance calls, and the additional cost incurred by the deployment after deploying the AI service instances on the selected edge nodes, where the additional service time , the additional communication time , and the additional deployment cost are respectively given by the following formulas: Among them, , ; Select the edge node that minimizes the sum of the additional communication time and the additional cost for the actual deployment of the AI service instance, that is: At the current time scale After the AI service instance is deployed under , the newly added deployment node set is represented as , and all the routes in are updated accordingly. One of the call paths is updated to: , and the corresponding routing probability is also updated to satisfy . The shunt request arrival rate of each corresponding call path is also updated to satisfy . The call path set is updated to to satisfy . and and represent the number of call paths before and after the update respectively. The end-to-end delay is updated accordingly. That is, the service delay of each will increase by the AI service delay and the communication time: When the current large time scale ends, enter the next sub-service graph orchestration.
8. An AI service and microservice hybrid orchestration device, characterized in that, The AI service and microservice hybrid orchestration device includes: Determination module, for each sub-service graph , determine corresponding several microservice instances and one AI service instance , where is composed of a request service chain that calls the same AI service instance ; A quantity determination module, configured to determine the quantity to be deployed based on the request arrival rate corresponding to the sub-service graph , processing capacity and the deployment cost determination of the quantity to be deployed; based on the request arrival rate corresponding to the sub-service graph , each of the several types of microservice instances processing capacity and the deployment cost determination of the quantity to be deployed; An orchestration module, based on the number of services to be deployed, uses a reinforcement learning agent to deploy each microservice instance in sequential order one by one, and adaptively determines a temporary request route in a probabilistic routing manner to obtain all temporary deployment and request results in ; Based on the number of services to be deployed, asynchronously deploys using a greedy strategy, and updates all request routes in to obtain the orchestration result of the sub-service graph; An iteration module for collecting the decision-making trajectories generated by the reinforcement learning agent during the entire orchestration process, performing agent training and neural network update to obtain a fully trained agent.
9. An AI service and microservice hybrid orchestration device, characterized in that, The AI service and microservice hybrid orchestration device includes a processor, a memory, and an AI service and microservice hybrid orchestration program stored on the memory and executable by the processor. When the AI service and microservice hybrid orchestration program is executed by the processor, it implements the steps of the AI service and microservice hybrid orchestration method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, An AI service and microservice hybrid orchestration program is stored on the computer-readable storage medium. When the AI service and microservice hybrid orchestration program is executed by a processor, it implements the steps of the AI service and microservice hybrid orchestration method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Heterogeneous multi-edge cloud collaborative micro-service deployment and routing joint optimization method and system
CN116915686A
Edge micro-service fine-grained deployment method and system based on reinforcement learning
CN117041330A
Joint optimization method and system for dynamic micro-service graph deployment and probability request routing
CN117692503A
Scaling method for a service, and related device
WO2021197364A1
Systems and methods for providing process automation and artificial intelligence, market aggregation, and embedded marketplaces for a transactions platform
WO2024025863A1
Cited By
Intelligent flow arrangement method based on fusion expert network and deep reinforcement learning
CN120835005A
AI service and micro service mixed arrangement method and system based on multi-edge cloud network
CN121585726A