AI service and micro service mixed arrangement method, device and equipment in cloud edge collaboration

By using deep reinforcement learning AMDQN algorithm and RAA and ICA algorithms in cloud-edge collaborative networks to dynamically adjust resources and instances, the mixed orchestration problem of AI services and microservices in heterogeneous environments is solved, and service response with low latency and load balancing is achieved.

CN120342877AActive Publication Date: 2025-07-18THREE GORGES HI TECH INFORMATION TECH CO LTD +1

Patent Information

Application Number
CN202510825423.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

It is difficult for existing network architectures to effectively implement hybrid orchestration of AI services and microservices in heterogeneous environments, especially in cloud-edge collaborative networks with heterogeneous computing and resource resources, which cannot meet the performance requirements of multi-scenario and multi-type user requests, especially in scenarios with delay-sensitive and computation-intensive requests.

Method used

By obtaining target basic information and edge user requests, determining the amount of resources to be allocated and the number of instances, using the AMDQN algorithm of deep reinforcement learning for resource allocation and deployment, combining RAA and ICA algorithms for dynamic adjustment of resources and instances, and achieving hybrid orchestration of AI services and microservices.

Benefits of technology

Dynamic scaling and adjustment of resources or instances of different types of services is achieved in a heterogeneous environment, adapting to different time slots or urban edge user request intensity, reducing service response delays and balancing network resource load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342877A_ABST
    Figure CN120342877A_ABST
Patent Text Reader

Abstract

The invention discloses an AI service and micro service mixed arrangement method, device and equipment in cloud edge collaboration, and the method comprises the steps: determining a to-be-allocated heterogeneous service resource quantity according to a target edge user request, a heterogeneous service resource quantity threshold value and full-size AI service operation intensity; carrying out resource allocation on the full-size AI service of a cloud computing layer in the current heterogeneous service orchestration network based on the heterogeneous service resource quantity to be allocated; determining the number of to-be-deployed instances of the lightweight AI service and the micro service according to the target edge user request; screening out a target AMDQN corresponding to the target basic information from an arrangement strategy database containing a plurality of basic AMDQNs having a mapping relationship with the geographic position-time slot combination; and controlling the target AMDQN to interact with the current heterogeneous service orchestration network, so as to deploy the lightweight AI service and the micro service corresponding to the number of the instances to be deployed in the current heterogeneous service orchestration network, thereby effectively realizing the mixed orchestration of the AI service and the micro service in the heterogeneous environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cloud-edge collaborative networks, and particularly to a method, apparatus, and device for hybrid orchestration of AI services and microservices in cloud-edge collaboration. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, AI services have been widely applied in fields such as smart homes, smart cities, and autonomous driving. However, these applications have extremely high requirements for service performance, including service response latency, resource utilization, and system robustness. However, due to the fact that existing network architectures generally focus on centralized computing, it is difficult to meet the performance requirements of multi-scenario and multi-type user requests. Especially in cloud-edge collaborative networks with computing and resource heterogeneity, the hybrid orchestration of AI services and microservices faces many challenges.

[0003] Currently, few related works or technologies focus on making valuable research on the hybrid orchestration of AI services and microservices. Among them, existing research mainly focuses on the deployment optimization of a single service type. For example, in the mobile edge computing (MEC) scenario, the microservice architecture is used to achieve horizontal scaling of microservice instances, or only the deployment and task offloading of AI services for cloud-edge-end collaboration are considered. However, these methods usually ignore the complex interaction characteristics between AI services and microservices in the service chain and are difficult to effectively handle scenarios where latency-sensitive requests and computationally intensive requests coexist in heterogeneous environments. Therefore, these methods cannot be used to implement the hybrid orchestration of AI services and microservices. Summary of the Invention

[0004] This application provides a method, apparatus, and device for hybrid orchestration of AI services and microservices in cloud-edge collaboration to effectively achieve the hybrid orchestration of AI services and microservices in a heterogeneous environment.

[0005] In a first aspect, an embodiment of this application provides a method for hybrid orchestration of AI services and microservices in cloud-edge collaboration, including the following steps: Obtain target basic information and a target edge user request corresponding to the current heterogeneous service orchestration network, where the target basic information includes a target geographical location and a target time slot; Determine the heterogeneous service resource amount to be allocated corresponding to the full-size AI service according to the target edge user request, a preset heterogeneous service resource amount threshold, and the operation intensity of the full-size AI service; Allocate resources to the full-size AI service in the cloud computing layer of the current heterogeneous service orchestration network based on the heterogeneous service resource amount to be allocated; Determine the number of instances to be deployed corresponding to the lightweight AI service and microservices according to the target edge user request; Select the base invalid action mask deep Q-network AMDQN corresponding to the target basic information from the preset orchestration policy database, where the orchestration policy database contains multiple base AMDQNs that have a mapping relationship with the geographical location-time slot combination as the target AMDQN; Control the target AMDQN to interact with the current heterogeneous service orchestration network to deploy lightweight AI services and microservices corresponding to the number of instances to be deployed in the current heterogeneous service orchestration network; Among them, the generation method of the base AMDQN is: based on historical edge user requests, service deployment information, network topology information, and the number of computing platforms in the base heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination, construct the state space, action space, and optimization objective function respectively, and determine the historical number of instances to be deployed corresponding to the lightweight AI services and microservices; construct a reward function according to the optimization objective function and the historical number of instances to be deployed; construct an initial AMDQN based on the state space, action space, and reward function and train it to generate the base AMDQN.

[0006] In a second aspect, an embodiment of the present application provides a hybrid orchestration device for AI services and microservices in cloud-edge collaboration, including: An information acquisition module, which is used to acquire target basic information and target edge user requests corresponding to the current heterogeneous service orchestration network, where the target basic information includes the target geographical location and the target time slot; A resource determination module, which is used to determine the amount of heterogeneous service resources to be allocated corresponding to the full-size AI service according to the target edge user request, the preset heterogeneous service resource amount threshold, and the operation intensity of the full-size AI service; A resource allocation module, which is used to allocate resources to the full-size AI services in the cloud computing layer of the current heterogeneous service orchestration network based on the amount of heterogeneous service resources to be allocated; An instance calculation module, which is used to determine the number of instances to be deployed corresponding to the lightweight AI services and microservices according to the target edge user request; A model screening module, which is used to select the base invalid action mask deep Q-network AMDQN corresponding to the target basic information from the preset orchestration policy database as the target AMDQN, where the orchestration policy database contains multiple base AMDQNs that have a mapping relationship with the geographical location-time slot combination; A hybrid orchestration module, which is used to control the target AMDQN to interact with the current heterogeneous service orchestration network to deploy lightweight AI services and microservices corresponding to the number of instances to be deployed in the current heterogeneous service orchestration network; A model generation module, which is used to construct a state space, an action space, and an optimization objective function based on historical edge user requests, service deployment information, network topology information, and the number of computing platforms in a basic heterogeneous service orchestration network corresponding to a preset geographical location-time slot combination, and determine the number of historical instances to be deployed corresponding to lightweight AI services and microservices; construct a reward function according to the optimization objective function and the number of historical instances to be deployed; construct an initial AMDQN based on the state space, the action space, and the reward function and train it to generate a basic AMDQN.

[0007] In a third aspect, an embodiment of the present application provides a device for hybrid orchestration of AI services and microservices in cloud-edge collaboration. The device for hybrid orchestration of AI services and microservices in cloud-edge collaboration includes a processor, a memory, and a cloud-edge collaboration AI service and microservice hybrid orchestration program stored on the memory and executable by the processor. When the cloud-edge collaboration AI service and microservice hybrid orchestration program is executed by the processor, the steps of the method for hybrid orchestration of AI services and microservices in cloud-edge collaboration as described above are implemented.

[0008] The beneficial effects brought by the technical solution provided by the embodiment of the present application include: Obtain target basic information including the target geographical location and target time slot and target edge user requests corresponding to the current heterogeneous service orchestration network; determine the heterogeneous service resource quantity to be allocated corresponding to the full-size AI service according to the target edge user requests, the heterogeneous service resource quantity threshold, and the operation intensity of the full-size AI service; allocate resources to the full-size AI service in the cloud computing layer of the current heterogeneous service orchestration network based on the heterogeneous service resource quantity to be allocated; determine the number of instances to be deployed corresponding to the lightweight AI service and microservices according to the target edge user requests; screen out the AMDQN corresponding to the target basic information from the orchestration policy database containing multiple basic AMDQNs having a mapping relationship with the geographical location-time slot combination as the target AMDQN; control the target AMDQN to interact with the current heterogeneous service orchestration network to deploy the lightweight AI service and microservices corresponding to the number of instances to be deployed in the current heterogeneous service orchestration network; wherein, the method for generating the basic AMDQN is: based on historical edge user requests, service deployment information, network topology information, and the number of computing platforms in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination, respectively construct a state space, an action space, and an optimization objective function, and determine the historical number of instances to be deployed corresponding to the lightweight AI service and microservices; construct a reward function according to the optimization objective function and the historical number of instances to be deployed; construct an initial AMDQN based on the state space, the action space, and the reward function and perform training to generate the basic AMDQN. Through the present application, dynamic scaling adjustment of heterogeneous resources or instances of different types of services can be achieved to adapt to the edge user request intensity in different time slots or cities, and the load balancing of network resources is realized through the AMDQN of deep reinforcement learning. At the same time, the service response delay is reduced to the greatest extent, thereby effectively realizing the hybrid orchestration of AI services and microservices in a heterogeneous environment. Brief Description of the Drawings

[0009] Figure 1 It is a schematic flowchart of an embodiment of the method for hybrid orchestration of AI services and microservices in cloud-edge collaboration of the present application; Figure 2 It is a schematic diagram of the heterogeneous service orchestration network architecture involved in the solution of the embodiment of the present application; Figure 3 It is a schematic diagram of the interaction framework between the AMDQN agent and the environment involved in the solution of the embodiment of the present application; Figure 4 It is a schematic diagram of the AMDQN algorithm architecture involved in the solution of the embodiment of the present application; Figure 5 It is a schematic diagram of the hybrid orchestration of AI services and microservices involved in the solution of the embodiment of the present application; Figure 6 It is a schematic diagram of the hardware structure of the device for hybrid orchestration of AI services and microservices in cloud-edge collaboration involved in the solution of the embodiment of the present application. Detailed implementation manners

[0010] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0011] First, some technical terms in this application are explained to facilitate the understanding of this application by those skilled in the art.

[0012] AI service: Services and applications using AI technology.

[0013] Full-size AI service: An AI service without model compression, which usually requires a large amount of heterogeneous service resources and is generally used to process computationally intensive requests such as LLMs (Large Language Models) inference tasks.

[0014] Lightweight AI service: Usually obtained from full-size AI services through model compression techniques (such as quantization, pruning, etc.). Deploying lightweight AI services requires fewer resources. And to ensure the stability of the service, lightweight AI services of the same type are generally obtained based on the same model compression method. Therefore, it is usually considered that lightweight AI services of the same type consume the same heterogeneous service resources and have the same inference rate.

[0015] Microservice: A service unit with an independent code library, database, and running environment that can be deployed and upgraded separately. Each microservice is responsible for a single function or business logic and interacts and collaborates with other services through an API (Application Programming Interface) to finally form a complete application to provide services for users.

[0016] To make the purpose, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.

[0017] In a first aspect, an embodiment of this application provides a method for hybrid orchestration of AI services and microservices in cloud-edge collaboration.

[0018] In one embodiment, refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the method for hybrid orchestration of AI services and microservices in cloud-edge collaboration of this application. As Figure 1As shown in the figure, the method for hybrid orchestration of AI services and microservices in cloud-edge collaboration includes: Step S10: Obtain the target basic information and target edge user requests corresponding to the current heterogeneous service orchestration network, where the target basic information includes the target geographical location and target time slot.

[0019] Exemplarily, it should be noted that the execution subject of this embodiment can be an AI service and microservice hybrid orchestration device with functions such as data processing, network communication, and program operation in a heterogeneous service orchestration network architecture, or other computer devices with similar functions, which are not limited herein.

[0020] It can be understood that the services in this embodiment include AI services and microservices, and the AI services include full-size AI services and lightweight AI services; among them, the full-size AI service set is , represents the total number of full-size AI services, and the minimum heterogeneous service resource threshold for deploying the full-size AI service set is , 、 and respectively represent the computing resource threshold, bandwidth resource threshold, and memory resource threshold corresponding to the full-size AI service , and in addition, the inference efficiency (i.e., the service speed for requests) of the full-size AI service set is expressed as ; the lightweight AI service set is , represents the total number of lightweight AI services, and the resource consumption for deploying 1 lightweight AI service is , 、 and respectively represent the computing resource consumption, bandwidth resource consumption, and memory resource consumption corresponding to the lightweight AI service , and in order to ensure the stability of the lightweight AI service, its inference efficiency is a determined value; the microservice set is , represents the total number of microservices. Microservices are mainly network functions in the application layer of the Open Systems Interconnection, such as firewalls, traffic shaping, and deep packet inspection, and the resource consumption for instantiating 1 microservice is , 、 and respectively represent the computing resource consumption, bandwidth resource consumption, and memory resource consumption corresponding to the microservice The corresponding computing resource consumption, bandwidth resource consumption, and memory resource consumption. In addition, the microservice service rate is also a definite value and follows an exponential distribution. It should be noted that for the sake of simplicity of description, parameters with the same meaning in the subsequent embodiments will not be elaborated on their meanings anymore.

[0021] In addition, for the heterogeneous service orchestration network, this embodiment also defines the resource allocation decision variable as the computing resources and bandwidth resources that a full-size AI service can obtain; defines the service deployment decision variable as the deployment quantities of the i-th lightweight AI service and the i-th microservice on the n-th computing platform respectively, and based on the service deployment decision variable, defines the binary decision variable indicating whether the i-th lightweight AI service and the i-th microservice are deployed on the n-th computing platform; defines the request routing decision variable indicating the routing probability that the edge user request is routed to the service on the computing platform after being processed and then routed to the service on the computing platform for further processing.

[0022] Refer to Figure 2 As shown, it should be understood that the architecture of the current heterogeneous service orchestration network is a heterogeneous service orchestration network architecture composed of four-layer network structures: user access layer, edge computing layer, cloud computing layer, and orchestration control layer. This architecture supports heterogeneous user requests, service types, service resources, and computing nodes, etc.; among them, the user access layer is constructed through the edge user request set. It should be noted that the edge user request set is , represents the total number of types of all requests, and each request can be represented by a triple as , represents the binary label of the request , that is, represents that the request is a latency-sensitive request, while represents that the request is a compute-intensive request; represents a set of services that need to be traversed in sequence to complete the request , which may be a pure microservice chain or a mixed call chain containing microservices and AI services, such as ; is for the request The arrival rate; it can be understood that the edge user requests include the type of request, the composition of the request, the arrival rate of the request, the required AI services and microservice types, service type preferences, and the number of concurrent requests, etc. The specific information it contains can be determined according to actual needs and is not limited here.

[0023] This embodiment constructs the edge computing layer and the cloud computing layer based on the communication relationship of the computing platform; among them, lightweight AI services and microservices are deployed through the edge computing platform in the edge computing layer, and a sensing module is embedded in the edge computing layer to sense the edge user requests through the sensing module, and then obtain the sensing results corresponding to the edge user requests; it should be noted that the implementation methods and principles of how to obtain the sensing results of edge user requests through the sensing module are well-known common knowledge in the art, so for the sake of simplicity of description, they will not be elaborated here. It should be noted that since full-size AI services have high requirements for heterogeneous service resources, and the cloud computing platform in the cloud computing layer has a large-scale computing cluster and advanced parallel processing hardware devices, only the cloud computing platform can deploy full-size AI services, and the full-size AI services deployed on the cloud computing platform can process computing-intensive requests such as big data analysis and LLMs inference tasks; in addition, since commercial AI services have very high requirements for the accuracy of services, such as financial analysis, atmospheric prediction, etc., and if lightweight AI services are used to process computing-intensive requests, it may cause a large amount of economic losses. Therefore, this embodiment stipulates that full-size AI services are only deployed on the cloud computing platform in the cloud computing layer, and there is only 1 full-size AI service for each type. Among them, the cloud computing platform and the edge computing platform can be connected through a high-latency public network, and different edge computing platforms can be connected through a high-speed dedicated line.

[0024] The orchestration control layer in this embodiment includes a time slot control system and an orchestration policy database, and the time slot control system further includes a time slot management module, a database access module, a proxy selection module, and an execution and feedback module, etc.; among them, the time slot management module is responsible for the scheduling and management of system time slots to determine the current time period and the status of time slots (idle, occupied, upcoming, etc.); the database access module obtains the current system time slot through the time slot management module and queries the orchestration policy database to select a suitable AMDQN (i.e., Deep Q-Network based on Invalid Action Mask) proxy; the proxy selection module selects a suitable AMDQN proxy according to the AMDQN proxy corresponding to the current time slot queried from the orchestration policy database and passes it to the subsequent execution and feedback module; once the AMDQN proxy is selected and executed, the execution and feedback module is responsible for monitoring its execution situation and feeding back the execution result (whether the proxy execution is successful, whether the proxy needs to be switched, etc.).

[0025] In this embodiment, for the current heterogeneous service orchestration network with the need for hybrid orchestration of AI services and microservices, obtain the corresponding target edge user requests and target basic information including the target geographical location and target time slot; wherein, the target edge user requests refer to all edge user requests received in the current heterogeneous service orchestration network; the target geographical location refers to the location where the current heterogeneous service network is located, for example, the target geographical location is City A; the target time slot refers to the time corresponding to the current heterogeneous service network, for example, the target time slot is aa:bb am on zz day of yy month in xx year. It should be understood that there is a mapping relationship between the geographical location-time slot and the AMDQN in this embodiment, that is, the AMDQN corresponding to the current heterogeneous service orchestration network can be screened out through the geographical location and time slot.

[0026] Step S20: Determine the heterogeneous service resource amount to be allocated corresponding to the full-size AI service according to the target edge user requests, the preset heterogeneous service resource amount threshold, and the operation intensity of the full-size AI service.

[0027] Exemplarily, it should be understood that since service providers are limited by the double cost thresholds of CAPEX (Capital Expenditure) and OPEX (Operating Expense), the heterogeneous resource pool of the cloud computing platforms they can schedule is limited, that is, the limited bandwidth resources and computing resources and other heterogeneous service resources in the heterogeneous resource pool form heterogeneous service resource constraints; therefore, in order to maximize resource utilization, this embodiment proposes a heterogeneous resource allocation algorithm RAA, that is, through the RAA algorithm, reasonable resource allocation is performed on the full-size AI services deployed on the cloud computing platform, so that the inference efficiency of the full-size AI services reaches the maximum under the condition of limited heterogeneous service resources. It can be seen that this embodiment will perform heterogeneous service resource allocation of full-size AI services through the RAA algorithm in the cloud computing layer to achieve maximum resource utilization.

[0028] Specifically, for the RAA algorithm, it needs to determine the request intensity of all edge users for the full-size AI service according to the perception results corresponding to the target edge user requests, and calculate the inference efficiency of the full-size AI service according to the request intensity and the heterogeneous service resource amount threshold; then determine the priority of allocating heterogeneous service resource amounts for each full-size AI service through the inference efficiency, request intensity, and operation intensity of the full-size AI service, and finally determine the heterogeneous service resource amount to be allocated for each full-size AI service according to this priority. It should be noted that the heterogeneous service resource amount threshold includes but is not limited to the computing resource amount threshold, bandwidth resource amount threshold, and memory resource amount threshold, and the specific value setting of the heterogeneous service resource amount threshold can be determined according to actual needs and is not limited here.

[0029] Further, in one embodiment, determining the heterogeneous service resource quantity to be allocated corresponding to the full-scale AI service according to the target edge user request, the preset heterogeneous service resource quantity threshold, and the operation intensity of the full-scale AI service includes: Obtaining a target perception result corresponding to the target edge user request based on the current heterogeneous service orchestration network; Calculating a first request intensity of the full-scale AI service based on the target perception result; Determining a first inference efficiency of the full-scale AI service according to the first request intensity; Calculating a first heterogeneous service resource quantity through the first inference efficiency, the gain coefficient of the bandwidth resource, the gain coefficient of the computing resource, and the operation intensity of the full-scale AI service; Calculating a second inference efficiency of the full-scale AI service according to the maximum value between the first heterogeneous service resource quantity and the preset heterogeneous service resource quantity threshold; Determining a first priority for allocating heterogeneous service resource quantities for each full-scale AI service based on the second inference efficiency, the first request intensity, and the operation intensity of the full-scale AI service; Determining the heterogeneous service resource quantity to be allocated for each full-scale AI service based on the first priority.

[0030] Exemplarily, in this embodiment, first, a target perception result corresponding to the target edge user request is obtained through the current heterogeneous service orchestration network, and then the target perception result is substituted into the following request intensity calculation formula of the full-scale AI service to obtain the first request intensity of each full-scale AI service from all edge users:

[0031] In the formula, represents the first request intensity of the full-scale AI service ; the binary label represents the type of the edge user request , where represents that the edge user request is a latency-sensitive request, while represents that the edge user request is a compute-intensive request; the binary variable represents whether a complete edge user request needs to go through the full-scale AI service inference, where represents no, while represents yes; represents the arrival rate of the edge user request .

[0032] Then, in order to enable the full-scale AI service To meet the request intensity of edge users, the first request intensity is taken as the first inference efficiency of the full-size AI service. On the premise of meeting the heterogeneous service resource constraints of the cloud computing platform, the RAA algorithm will allocate the bandwidth resources and computing resources that meet the minimum requirements for each full-size AI service. Specifically, substitute the first inference efficiency, the gain coefficient of bandwidth resources, the gain coefficient of computing resources, and the operation intensity of the full-size AI service into the following inference efficiency formula to solve for the first heterogeneous service resource quantity:

[0033] In the formula, represents the first inference efficiency of the full-size AI service, respectively represent the gain coefficients of bandwidth resources and computing resources, respectively represent the bandwidth resource quantity and computing resource quantity in the first heterogeneous service resource quantity, represents the full-size AI service operation intensity, which reflects the density of the computing amount of the full-size AI service relative to the memory access amount.

[0034] Then, judge the magnitude between the first heterogeneous service resource quantity and the heterogeneous service resource quantity threshold, and substitute the maximum value of the two into the above inference efficiency formula to recalculate the second inference efficiency full_AI_infer_eff of each full-size AI service. For example, if the bandwidth resource quantity in the first heterogeneous service resource quantity is greater than the bandwidth resource quantity threshold, substitute the bandwidth resource quantity in the first heterogeneous service resource quantity into the above inference efficiency formula; otherwise, substitute the bandwidth resource quantity threshold into the above inference efficiency formula. Similarly, if the computing resource quantity in the first heterogeneous service resource quantity is greater than the computing resource quantity threshold, substitute the computing resource quantity in the first heterogeneous service resource quantity into the above inference efficiency formula; otherwise, substitute the computing resource quantity threshold into the above inference efficiency formula.

[0035] Then, record the difference between each second inference efficiency full_AI_infer_eff and the first request intensity as infer_eff_equal, and the difference sequence infer_eff_equal_list can be obtained. For example, assume there are 4 full-size AI services to And the corresponding infer_eff_equal_list = [1, 0, 3, 0]; then perform a deduplication operation on infer_eff_equal_list and sort it in ascending order to obtain infer_eff_line_list = [0, 1, 3]; immediately afterwards, calculate the difference between the maximum value in infer_eff_line_list and each value to obtain infer_eff_lack_list = [3, 2, 0]; the RAA algorithm will count the full-size AI services corresponding to each value in infer_eff_line_list and sort them in descending order according to the operation intensity of the full-size AI services, denoted as line_corr_full_AI. For example, assuming , then line_corr_full_AI = , that is, the sorting in line_corr_full_AI represents the priorities of different full-size AI services.

[0036] Finally, allocate the corresponding heterogeneous service resources in infer_eff_lack_list to the full-size AI services in line_corr_full_AI in sequence until the remaining resources in the cloud computing platform are exhausted or the values in infer_eff_equal_list are equal, indicating that the process of allocating resources to the full-size AI services ends, thereby obtaining the heterogeneous service resource amounts to be allocated for each full-size AI service.

[0037] Step S30: Allocate resources to the full-size AI services in the cloud computing layer of the current heterogeneous service orchestration network based on the heterogeneous service resource amounts to be allocated.

[0038] Exemplarily, in this embodiment, after determining the heterogeneous service resource amounts to be allocated for each full-size AI service, allocate resources to the full-size AI services in the cloud computing layer of the current heterogeneous service orchestration network according to the heterogeneous service resource amounts to be allocated for each full-size AI service, so as to reasonably allocate computing resources and bandwidth resources for the full-size AI services, and then realize the orchestration of full-size AI services in a heterogeneous environment, thereby ensuring the best inference efficiency.

[0039] Step S40: Determine the number of instances to be deployed corresponding to the lightweight AI service and the microservice according to the target edge user request.

[0040] Exemplarily, it can be understood that since the lightweight AI services and microservices have relatively low resource requirements, they can be distributedly deployed on multiple edge computing platforms in the form of multiple instances. Among them, in this embodiment, a service instance calculation algorithm ICA is proposed to determine the number of instances to be deployed for the lightweight AI services and microservices. Specifically, for the ICA algorithm, it is necessary to determine the request intensity of all edge users for the lightweight AI services and microservices based on the perception results corresponding to the target edge user requests, and determine the basic number of instances for the lightweight AI services and microservices through the request intensity; then determine the priority of increasing the number of instances for the lightweight AI services and microservices through the basic number of instances; finally, determine the number of instances to be added for the lightweight AI services and microservices to be deployed based on the priority of increasing the number of instances to meet the high-concurrency user requests.

[0041] Further, in one embodiment, the determining the number of instances to be deployed corresponding to the lightweight AI services and microservices according to the target edge user requests includes: Obtaining a target perception result corresponding to the target edge user request based on the current heterogeneous service orchestration network; Calculating the second request intensity of all edge users for each lightweight AI service and the third request intensity of all edge users for each microservice based on the target perception result; Determining the first basic number of instances for each lightweight AI service according to the second request intensity, and determining the second basic number of instances for each microservice according to the third request intensity; Determining the second priority of increasing the number of instances for each lightweight AI service through the first basic number of instances, and determining the number of instances to be deployed for each lightweight AI service based on the first basic number of instances and the second priority; Determining the maximum delay reduction amount for adding 1 instance to each microservice according to the second basic number of instances, determining the third priority of increasing the number of instances for each microservice based on the sorting result of the maximum delay reduction amount, and determining the number of instances to be deployed for each microservice based on the second basic number of instances and the third priority.

[0042] Exemplarily, in this embodiment, first, a target perception result corresponding to the target edge user request is obtained through the current heterogeneous service orchestration network, and then the target perception result is respectively substituted into the following request intensity calculation formulas for lightweight AI services and request intensity calculation formulas for microservices to obtain the second request intensity of all edge users for each lightweight AI service and the third request intensity of all edge users for each microservice:

[0043] In the formula, represents the lightweight AI service The second request intensity; binary label Indicates the edge user request Type; binary variable Indicates the completion of a complete edge user request Whether lightweight AI service inference is required; Indicates the microservice The third request intensity; integer variable For a complete edge user request Needs to go through the microservice Number of processing times.

[0044] Then the ICA algorithm will determine the basic instance numbers of each lightweight AI service and each microservice according to the service intensity in the steady-state queuing network, that is, for each lightweight AI service The second request intensity and inference efficiency Substitute into the following service instance calculation formula for lightweight AI services to calculate the first basic instance number of each lightweight AI service, and then obtain the basic instance number sequence light_AI_req_list of lightweight AI services:

[0045] In the formula, Indicates the first basic instance number of the lightweight AI service

[0046] Similarly, substitute the third request intensity and inference efficiency of each microservice Into the following service instance calculation formula for microservices to calculate the second basic instance number of each microservice, and then obtain the basic instance number sequence ms_req_list of microservices:

[0047] In the formula, Indicates the second basic instance number of the microservice

[0048] Immediately afterwards, the ICA algorithm will sort the basic instance numbers in light_AI_req_list in descending order to obtain the priority list light_AI_prio_list of lightweight AI services, that is, the sorting of light_AI_prio_list represents the priority of increasing the instance numbers of different lightweight AI services; assume there are four lightweight AI services To ​​If light_AI_req_list = [3, 2, 5, 6], then after sorting in descending order according to the number of basic instances of each lightweight AI service, we get light_AI_prio_list = ; Then, on the premise of satisfying the heterogeneous service resource constraints of the cloud computing platform, the ICA algorithm will sequentially add 1 instance to each lightweight AI service in light_AI_prio_list until the heterogeneous service resource constraints are not satisfied, thereby obtaining the number of instances to be deployed for each lightweight AI service.

[0049] To improve the utility of services deployed in a multi-instance manner, the ICA algorithm will calculate the maximum delay reduction for adding 1 instance to each microservice, that is, for each microservice the second basic instance number is substituted into the following microservice priority calculation formula to obtain the score corresponding to the maximum delay reduction for each microservice, and then the microservice-side score sequence ms_score_list is obtained:

[0050] In the formula, represents the service queue delay corresponding to deploying instances of microservices in the current heterogeneous service orchestration network; it can be understood that the larger the score in ms_score_list, the higher the priority of adding an instance to its corresponding microservice; next, the ICA algorithm will add 1 instance to the microservice with the highest score in ms_score_list, recalculate the score of this microservice, and continuously repeat the above process until the resource constraints are not satisfied, indicating that the calculation process of microservice instances ends, thereby obtaining the number of instances to be deployed for each microservice.

[0051] Specifically, assume there are 3 microservices and the number of basic instances of each microservice calculated according to the microservice instance calculation formula is , then the maximum delay reduction (i.e., score) for adding 1 instance to each microservice calculated according to the microservice priority calculation formula is ; It can be seen that adding 1 instance to microservice can reduce the delay the most. Therefore, if the remaining resources in the current heterogeneous service orchestration network are only enough to deploy 1 instance, this embodiment should choose to add 1 instance to microservice , then the number of instances of the 3 microservices will correspondingly become ; Since in the previous step this embodiment only updated microservice The number of instances, so next only the microservices need to be recalculated score. Assume that the maximum latency reduction when its number of instances increases from 4 to 5 is 0.25 (that is, the recalculated score of the microservice is 0.25). Then the scores of the 3 microservices at this time become , which means that assuming that the remaining resources in the current heterogeneous service orchestration network can still deploy 1 more instance, it will preferentially choose to add an instance for the microservice ; continuously repeat the above process until the remaining resources in the network can no longer support adding any 1 microservice instance, and the number of to-be-deployed instances of each microservice can be determined.

[0052] Step S50: Screen out the basic invalid action mask depth Q-network AMDQN corresponding to the target basic information from the preset orchestration policy database. The orchestration policy database contains multiple basic AMDQNs having a mapping relationship with the geographical location-time slot combination; wherein, the generation method of the basic AMDQN is: based on the historical edge user requests, service deployment information, network topology information, and the number of computing platforms in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination, respectively construct a state space, an action space, and an optimization objective function, and determine the historical number of to-be-deployed instances corresponding to the lightweight AI service and the microservices; construct a reward function according to the optimization objective function and the historical number of to-be-deployed instances; construct an initial AMDQN based on the state space, the action space, and the reward function and train it to generate the basic AMDQN.

[0053] Exemplarily, in this embodiment, different basic AMDQN agents will be constructed for different geographical location-time slot combinations, and the basic AMDQN agents and the mapping relationship between the basic AMDQN agents and the geographical location-time slot combinations will be stored in the orchestration policy database, so that the time slot control system screens out the basic AMDQN agent corresponding to the combination of the target geographical location and the target time slot from the orchestration policy database as the target AMDQN agent, so as to maximize the cumulative reward through continuous interaction and repeated iteration between the target AMDQN agent and the heterogeneous service orchestration network. Its ultimate goal is to minimize the average response latency of all edge user requests to the greatest extent without violating constraints such as resources, routing, and services, while balancing the service resource load of the heterogeneous service orchestration network.

[0054] Among them, see Figure 3As shown in the figure, based on the service deployment information, network topology information, number of computing platforms, and perception results of historical edge user requests in the basic heterogeneous service orchestration network corresponding to the preset geographical location - time slot combination, the state space, action space, and optimization objective function of the AMDQN agent are determined. Then, the number of historical instances to be deployed corresponding to the lightweight AI service and microservices is determined through the ICA algorithm. Based on the optimization objective function and the number of historical instances to be deployed, the reward function can be constructed. Then, the initial AMDQN agent is constructed through the state space, action space, and reward function and trained until it converges to generate the basic AMDQN agent corresponding to the preset geographical location - time slot combination and transfer it to the orchestration policy database.

[0055] Among them, when traditional methods implement hybrid orchestration, they usually optimize service deployment and request routing separately, failing to fully utilize global resources for joint decision-making, which further restricts network performance. In this embodiment, during the training of the initial AMDQN agent, the request routing probability of the AMDQN agent is also determined based on the service deployment status of the basic heterogeneous service orchestration network. According to the communication delay between computing platforms and the aggregation degree of service instances on the computing platforms, different weights of the request routing probability are used to forward historical edge user requests from the parent computing platform to the next child computing platform, so as to plan a reasonable service path for each edge user request, thereby meeting the low-latency user requirements.

[0056] See Figure 3 and Figure 4 As shown in the figure, the AMDQN agent continuously interacts with the heterogeneous service orchestration network to obtain learning experience, and then updates the network structure and weight parameters. Among them, AMDQN mainly includes two Q networks designed based on the dueling network architecture, namely the online Q network and the target Q network . The former is responsible for generating the Q value of each action in the current state, guiding the action output and gradually updating its parameters during the training process; the latter provides a stable Q value estimate for the next state, which is used to calculate the TD target (Temporal Difference Target) to update the online Q network, thereby reducing the fluctuations during training. It should be noted that Figure 3 and Figure 4 The network environment in is a network environment composed of the hybrid orchestration of AI services and microservices. For example, it can be Figure 5 The network environment composed of the hybrid orchestration of AI services and microservices shown in the figure.

[0057] The following embodiments will combine Figure 4 to elaborate on the detailed process of the invalid action mask deep Q network (AMDQN) algorithm: Before each training starts, the online Q network With the target Q network The weight parameters of are randomly initialized, and at the initial stage of each training round, the heterogeneous service orchestration network is also initialized; then, the AMDQN agent will execute the initial solution of service deployment, that is, deploy 1 instance of each service in the heterogeneous service orchestration network in sequence to ensure the basic service capabilities of the heterogeneous service orchestration network; when the heterogeneous service orchestration network changes with the training process, the invalid action mask will be calculated first; it should be understood that the invalid action mask is based on a boolean operation mask, and significantly reduces the search space of the AMDQN agent by masking each deployment operation that does not meet the service resource constraints, enabling the AMDQN model to converge faster; the AMDQN agent randomly explores among the remaining valid actions and obtains a transition sample (state by executing an action , action , next state , and a flag indicating whether the current round of training is completed ) and stores it in the replay buffer ; once the learning condition is met, the AMDQN will randomly draw from a batch of transition samples of size and their corresponding importance sampling weights , calculate their target Q values and the current Q value respectively, then perform a gradient descent operation on the loss function to update the online Q network parameters , and finally perform a soft update on the target Q network parameters .

[0058] After multiple rounds of iterative training, it is necessary to evaluate the algorithm performance. Specifically, the environment can be initialized, and deployment actions determined by the greedy policy can be selected in multi-step interactions to obtain the optimal AI service and microservice hybrid orchestration strategy that the current algorithm can obtain; if the convergence determination condition is met, this pre-trained AMDQN model (i.e., the obtained basic AMDQN, that is, the AMDQN agent) will be saved to the orchestration strategy database, and it will be used as a control agent to provide the optimal solution for solving the AI service and microservice hybrid orchestration problem of the heterogeneous service orchestration network at the corresponding time slot.

[0059] Furthermore, in one embodiment, constructing a state space, an action space, and an optimization objective function based on historical edge user requests, service deployment information and network topology information in the basic heterogeneous service orchestration network corresponding to a preset geographical location-time slot combination, and the number of computing platforms respectively includes: Construct a state space based on the historical edge user requests, service deployment information, and network topology information in the basic heterogeneous service orchestration network corresponding to the preset geographical location - time slot combination; Construct an action space based on the number of computing platforms in the basic heterogeneous service orchestration network corresponding to the preset geographical location - time slot combination; Construct an optimization objective function based on the historical edge user requests and service deployment information in the basic heterogeneous service orchestration network corresponding to the preset geographical location - time slot combination.

[0060] Among them, the expression of the reward function is:

[0061] In the formula, represents the reward value at the th time step in the th training round, represents the total number of instances of the service to be deployed, represents the comprehensive metric of the performance of the heterogeneous service orchestration network at the th time step in the th training round, respectively represent the weight parameters of the influence of the time step and the training round on the iterative reward; represents the average service delay of the edge user request , represents the load - balancing index of the heterogeneous service orchestration network, respectively represent the weight parameters of the influence of the average service delay and the load - balancing index on the optimization objective function.

[0062] Exemplarily, in this embodiment, the average service delay (i.e., the average response delay) of the historical edge user request and the load - balancing index of the basic heterogeneous service orchestration network are respectively determined through the perception results corresponding to the historical edge user requests and the service deployment information in the basic heterogeneous service orchestration network; then, based on the average service delay of the historical edge user request and the load - balancing index of the basic heterogeneous service orchestration network, the optimization objective function is determined; among them, the expression of the optimization objective function is:

[0063] In the formula, is the average service delay of the historical edge user request , is the load - balancing index of the basic heterogeneous service orchestration network, They are the weight parameters of the average service delay and the load balancing index affecting the optimization objective function respectively, and their specific values can be determined according to actual requirements and are not limited here.

[0064] This embodiment will define the state space , it should be noted that the state space of AMDQN can be described as the sum of two parts, namely the network resource state of the heterogeneous service orchestration network S resource and the service deployment state S deployment , that is, the state space of AMDQN is defined as ; where the network resource state is , which includes the remaining situation of the heterogeneous service resources of each computing platform and the occupancy situation ; the service deployment state is , which are respectively the distribution of various deployed services on the computing platform and the location of the computing platform where the current service is deployed in the heterogeneous service orchestration network ; for example, assuming there are 3 computing platforms in the current network , then the communication delay matrix between the computing platforms is:

[0065] If the computing platform where the current service is deployed is , then calculate The sum of the communication delays to any next computing platform is , so The normalized communication delays to any next computing platform are respectively , so .

[0066] Based on the above-defined state space, this embodiment will construct the state space corresponding to the basic heterogeneous service orchestration network according to the historical edge user requests and the service deployment information and network topology information in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination.

[0067] Next, define the action space , it can be understood that each action of AMDQN is to deploy a specific service on the selected computing platform; in fact, for any service to be deployed, before AMDQN deploys an instance for it according to the state of the heterogeneous service orchestration network, it always obtains a one-dimensional probability list according to the environmental state, specifically:

[0068] In the formula, Represents the computing platform The probability of being selected for service deployment. AMDQN is based on The policy finally determines the computing platform for service deployment; it should be noted that due to the invalid action mask mechanism in AMDQN, AMDQN will, before outputting According to Filter out the computing platforms that do not have enough resources for service deployment, and set their corresponding To 0 and evenly distribute it to other computing platforms; for example, there are 3 computing platforms in the current network And the corresponding Assume that based on the invalid action mask mechanism, it is determined that Does not have enough resources for service deployment, then set To 0 and divide the original Corresponding 0.6 evenly to And Then get the new And output it.

[0069] In addition, once the service deployment is successful, AMDQN will immediately update the routing percentages of each service in the heterogeneous service orchestration network according to the routing probability update formula, the communication delay between computing platforms, and the different weights of the request routing probability based on the aggregation degree of service instances on the computing platforms, so as to obtain a routing probability table as shown in Figure 5 Thereby forwarding the user request from the parent computing platform to the next child computing platform according to the routing probability table. It can be seen that in this embodiment, the AMDQN control agent designed based on deep reinforcement learning jointly optimizes service deployment and request routing to balance the network resource load and at the same time minimize the service response delay to the greatest extent Among them, the routing probability update formula is:

[0070] In the formula, Represents the routing probability that the edge user request is routed to the service On the computing platform After being processed and then routed to the service On the computing platform For continued processing, Respectively represent the sets of full-size AI services, lightweight AI services, and microservices, Represents from the computing platform To The communication delay, Is the set of all computing platforms in the heterogeneous service orchestration network that deploy the service Represents the computing platform The service deployed on ​The quantity of represents the number of deployed services at the time step; represents the weight parameter that indicates the influence of the aggregation degree of communication delays between computing platforms on the request routing probability on the computing platform; represents the weight parameter that indicates the influence of the aggregation degree of service instances between computing platforms on the request routing probability on the computing platform; represents whether there is a service deployed on the computing platform.

[0071] Based on the action space defined above, in this embodiment, an action space corresponding to the basic heterogeneous service orchestration network will be constructed based on the number of computing platforms in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination.

[0072] Define the reward function , and the reward function of the AMDQN describes the feedback obtained by the AMDQN agent from the environment for each interaction with the heterogeneous service orchestration network. However, in the problem of hybrid orchestration of AI services and microservices, the comprehensive performance of the heterogeneous service orchestration network can only be accurately evaluated after both resource allocation and service deployment are completed, which means that using the reward of the AMDQN to evaluate the comprehensive performance of the heterogeneous service orchestration network is only accurate at the last step. Therefore, in order to improve the accuracy of each step of evaluation as much as possible, this embodiment adopts a new iterative reward shaping method (i.e., sparse reward shaping), which shapes the iterative rewards at different training stages through an iterative manner, thereby improving the training efficiency; among them, the expression of the reward function is:

[0073] In the formula, represents the reward value at the th time step in the th training round, represents the total number of instances of the service to be deployed, represents the comprehensive metric of the performance of the heterogeneous service orchestration network at the th time step in the th training round, respectively represent the weight parameters that indicate the influence of the time step and the training round on the iterative reward; represents the average service delay of the edge user request , represents the load balancing index of the heterogeneous service orchestration network, respectively represent the weight parameters that indicate the influence of the average service delay and the load balancing index on the optimization objective function.

[0074] It can be seen that based on the reward function defined above, in this embodiment, the reward function corresponding to the basic heterogeneous service orchestration network will be constructed according to the optimization objective function and the number of historical instances to be deployed.

[0075] Further, in one embodiment, constructing the optimization objective function based on the historical edge user requests and the service deployment information in the basic heterogeneous service orchestration network corresponding to the preset geographical location - time slot combination includes: Input the historical edge user requests into the basic heterogeneous service orchestration network to obtain historical perception results; Calculate the average service delay corresponding to the historical edge user requests according to the historical perception results; Calculate the load balancing index corresponding to the basic heterogeneous service orchestration network through the service deployment information in the basic heterogeneous service orchestration network; Determine the optimization objective function based on the average service delay and the load balancing index.

[0076] Among them, the calculation formula for the average service delay is:

[0077] In the formula, represents the average service delay of the edge user request , represents the edge user request the set of all request paths of represents the edge user request the th service path of represents the th computing platform whether there is an AI service deployed on it, represents the th computing platform whether there is a microservice deployed on it, represents the edge user request on the th computing platform the queue delay of the AI service on it, represents the edge user request on the th computing platform the processing delay of the microservice on it, represents the edge user request on the computing platform the AI service on Route to the computing platform after processing to the AI service on it The routing probability for continued processing, indicating the arrival rate of edge user requests of, indicating the completion of a full edge user request whether it is necessary to go through full-scale AI service inference, indicating the completion of a full edge user request whether it is necessary to go through lightweight AI service inference, indicating the unit time slot length, indicating the full-scale AI service operation intensity of, indicating the lightweight AI service operation intensity of, indicating the full-scale AI service inference efficiency of, indicating the lightweight AI service inference efficiency of, indicating the arrival rate of edge user requests assigned to a complete service path of the arrival rate on, indicating the edge user request at the th computing platform corresponding to the AI service in the service chain of the arrival rate on, indicating the average communication delay during cross-computing platform processing.

[0078] The calculation formula for the load balancing index is:

[0079] In the formula, is the load balancing index of the basic heterogeneous service orchestration network, respectively represent the sensitivity coefficients of the computing resource occupancy rate, bandwidth resource occupancy rate, and memory resource occupancy rate, respectively represent the variances of the computing resource occupancy rate, bandwidth resource occupancy rate, and memory resource occupancy rate, indicating the total number of computing platforms, res indicating the computing resource occupancy rate, bandwidth resource occupancy rate, or memory resource occupancy rate, indicating the average utilization rate of heterogeneous service resources, indicating the computing platform utilization rate of heterogeneous service resources on.

[0080] Exemplarily, in this embodiment, first, the historical perception result corresponding to the historical edge user request is obtained through the historical heterogeneous service orchestration network, and then the historical perception result is substituted into the following average service delay calculation formula to obtain the average service delay of the historical edge user request:

[0081] In the formula, represents the average service delay of the edge user request . represents the set of all request paths of the edge user request . represents the th service path of the edge user request . represents whether there is an AI service deployed on the th computing platform . represents whether there is a microservice deployed on the th computing platform . represents the queue delay on the AI service of the edge user request on the th computing platform . represents the processing delay on the microservice of the edge user request on the th computing platform . represents the routing probability that after the AI service of the edge user request on the computing platform is processed, it is routed to the AI service on the computing platform for continued processing . represents the arrival rate of the edge user request . represents whether a full-scale AI service inference is required to complete a complete edge user request . represents whether a lightweight AI service inference is required to complete a complete edge user request . represents the unit time slot length . represents the operation intensity of the full-scale AI service . represents the lightweight AI service The operation intensity, which reflects the density of the lightweight AI service computing volume relative to the memory access volume; Represents the inference efficiency of the full-size AI service which can be calculated by the RAA algorithm; Represents the inference efficiency of the lightweight AI service Represents the arrival rate of edge user requests distributed to a complete service path Represents the arrival rate of edge user requests at the th computing platform on the corresponding service chain the th service Refers to the full-size AI service or the lightweight AI service located at the th service in the service chain; Represents the average communication delay during cross-computing platform processing.

[0082] Next, substitute the service deployment information in the basic heterogeneous service orchestration network into the following load balancing index calculation formula to calculate the load balancing index of the basic heterogeneous service orchestration network:

[0083] In the formula, is the load balancing index of the basic heterogeneous service orchestration network, respectively represent the sensitivity coefficients of the computing resource occupancy rate, bandwidth resource occupancy rate, and memory resource occupancy rate, respectively represent the variances of the computing resource occupancy rate, bandwidth resource occupancy rate, and memory resource occupancy rate, represents the total number of computing platforms, res represents the heterogeneous service resources, that is, which is also the computing resource occupancy rate, bandwidth resource occupancy rate, or memory resource occupancy rate, represents the heterogeneous service resources the average utilization rate of represents the computing platform on the heterogeneous service resources

[0084] Finally, according to the average service delay of historical edge user requests The optimized objective function can be determined.

[0085] Step S60: Control the target AMDQN to interact with the current heterogeneous service orchestration network, so as to deploy lightweight AI services and microservices corresponding to the number of instances to be deployed in the current heterogeneous service orchestration network.

[0086] Exemplarily, in this embodiment, based on the time slot control system, a suitable target AMDQN agent is selected to continuously interact with the current heterogeneous service orchestration network until the hybrid orchestration of lightweight AI services and microservices is completed, that is, lightweight AI services and microservices corresponding to the number of instances to be deployed are deployed in the current heterogeneous service orchestration network, thereby effectively realizing the hybrid orchestration of AI services and microservices in a heterogeneous computing network such as user requests, service types, service resources, and computing nodes.

[0087] It can be seen that compared with most existing service orchestration technologies that only consider the deployment of single microservices or AI services and split service deployment and request routing as independent problems and then optimize them separately, this embodiment solves the problem of hybrid orchestration of AI services and microservices in a heterogeneous service orchestration network architecture by splitting the hybrid orchestration process of AI services and microservices and designing a three-stage algorithm based on heuristic and reinforcement learning methods; among them, the three-stage algorithm consists of a heuristic algorithm including the RAA algorithm and the ICA algorithm and the AMDQN reinforcement learning algorithm; the RAA algorithm is used to complete the resource allocation of full-size AI services on the cloud computing platform, the ICA algorithm is used to determine the deployment quantities of lightweight AI services and microservices, and the AMDQN algorithm deploys each service in turn on different computing platforms in the network according to the deployment quantities determined by the ICA, and updates the routing probability of each computing platform while deploying each service. In summary, the heterogeneous service orchestration network architecture in this embodiment provides technical support for the hybrid orchestration of AI services and microservices. The RAA algorithm and the ICA algorithm realize the dynamic scaling adjustment of heterogeneous resources or instances of different types of services to adapt to the edge user request intensity in different time slots or cities, and the AMDQN control agent designed based on deep reinforcement learning jointly optimizes service deployment and request routing to balance the network resource load while minimizing the service response delay to ultimately solve the problem of hybrid orchestration of AI services and microservices in a heterogeneous service orchestration network architecture.

[0088] In a second aspect, an embodiment of the present application further provides a device for hybrid orchestration of AI services and microservices in cloud-edge collaboration.

[0089] In one embodiment, the device for hybrid orchestration of AI services and microservices in cloud-edge collaboration includes: An information acquisition module, which is used to acquire target basic information and target edge user requests corresponding to the current heterogeneous service orchestration network, where the target basic information includes a target geographical location and a target time slot; A resource determination module, which is used to determine the heterogeneous service resource quantity to be allocated corresponding to the full-size AI service according to the target edge user request, a preset heterogeneous service resource quantity threshold, and the operation intensity of the full-size AI service; A resource allocation module, which is used to allocate resources to the full-size AI service in the cloud computing layer of the current heterogeneous service orchestration network based on the heterogeneous service resource quantity to be allocated; An instance calculation module, which is used to determine the number of instances to be deployed corresponding to the lightweight AI service and microservices according to the target edge user request; A model screening module, which is used to screen out the basic invalid action mask deep Q-network AMDQN corresponding to the target basic information from a preset orchestration policy database, and the orchestration policy database contains multiple basic AMDQNs having a mapping relationship with the geographical location-time slot combination as the target AMDQN; A hybrid orchestration module, which is used to control the interaction between the target AMDQN and the current heterogeneous service orchestration network to deploy the lightweight AI service and microservices corresponding to the number of instances to be deployed in the current heterogeneous service orchestration network; A model generation module, which is used to construct a state space, an action space, and an optimization objective function based on historical edge user requests, service deployment information, network topology information, and the number of computing platforms in the basic heterogeneous service orchestration network corresponding to a preset geographical location-time slot combination, and determine the number of historical instances to be deployed corresponding to the lightweight AI service and microservices; construct a reward function according to the optimization objective function and the number of historical instances to be deployed; construct an initial AMDQN based on the state space, the action space, and the reward function and perform training to generate a basic AMDQN.

[0090] Further, in one embodiment, the resource determination module is specifically used for: Obtaining a target perception result corresponding to the target edge user request based on the current heterogeneous service orchestration network; Calculating a first request intensity of the full-size AI service based on the target perception result; Determining a first inference efficiency of the full-size AI service according to the first request intensity; Calculating a first heterogeneous service resource quantity through the first inference efficiency, a gain coefficient of bandwidth resources, a gain coefficient of computing resources, and the operation intensity of the full-size AI service; Calculating a second inference efficiency of the full-size AI service according to the maximum value between the first heterogeneous service resource quantity and the preset heterogeneous service resource quantity threshold; Determine the first priority of allocating heterogeneous service resources for each full-size AI service based on the second inference efficiency, the first request intensity, and the operation intensity of the full-size AI service; Determine the heterogeneous service resource quantity to be allocated for each full-size AI service based on the first priority.

[0091] Further, in one embodiment, the instance calculation module is specifically configured to: Obtain a target perception result corresponding to a target edge user request based on the current heterogeneous service orchestration network; Calculate the second request intensity of all edge users for each lightweight AI service and the third request intensity of all edge users for each microservice based on the target perception result; Determine the first basic instance number of each lightweight AI service according to the second request intensity, and determine the second basic instance number of each microservice according to the third request intensity; Determine the second priority of increasing the instance number for each lightweight AI service through the first basic instance number, and determine the number of instances to be deployed for each lightweight AI service based on the first basic instance number and the second priority; Determine the maximum delay reduction amount of adding 1 instance for each microservice according to the second basic instance number, determine the third priority of increasing the instance number for each microservice based on the sorting result of the maximum delay reduction amount, and determine the number of instances to be deployed for each microservice based on the second basic instance number and the third priority.

[0092] Further, in one embodiment, the model generation module is specifically configured to: Construct a state space based on the historical edge user requests, the service deployment information and network topology information in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination; Construct an action space based on the number of computing platforms in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination; Construct an optimization objective function based on the historical edge user requests and the service deployment information in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination.

[0093] Further, in one embodiment, the model generation module is specifically further configured to: Input the historical edge user requests into the basic heterogeneous service orchestration network to obtain historical perception results; Calculate the average service delay corresponding to the historical edge user requests according to the historical perception results; Calculate the load balancing index corresponding to the basic heterogeneous service orchestration network through the service deployment information in the basic heterogeneous service orchestration network; Determine an optimization objective function based on the average service delay and the load balancing index.

[0094] Further, in one embodiment, the calculation formula for the average service delay is:

[0095] In the formula, represents the average service delay of the edge user request , represents the set of all request paths of the edge user request , represents the th service path of the edge user request represents the th computing platform where there is AI service deployed, represents the th computing platform where there is microservice deployed, represents the queue delay on the AI service of the th computing platform of the edge user request , represents the processing delay on the microservice of the th computing platform of the edge user request , represents the routing probability that after the AI service on the computing platform finishes processing the edge user request, it is routed to the AI service on the computing platform to continue processing, represents the arrival rate of the edge user request , represents whether a full-scale AI service inference is required to complete a complete edge user request , represents whether a lightweight AI service inference is required to complete a complete edge user request , represents the unit time slot length, represents the operation intensity of the full-scale AI service , represents the operation intensity of the lightweight AI service , represents the full-scale AI service , represents the lightweight AI service The operation intensity, represents the inference efficiency of the full-size AI service ; represents the inference efficiency of the lightweight AI service ; represents the arrival rate of edge user requests assigned to a complete service path ; represents the arrival rate of edge user requests at the th computing platform corresponding to the AI service in the service chain ; represents the average communication delay during cross-computing platform processing.

[0096] Furthermore, in one embodiment, the calculation formula of the load balancing index is:

[0097] In the formula, is the load balancing index of the basic heterogeneous service orchestration network, respectively represent the sensitivity coefficients of the computing resource occupancy rate, bandwidth resource occupancy rate, and memory resource occupancy rate, respectively represent the variances of the computing resource occupancy rate, bandwidth resource occupancy rate, and memory resource occupancy rate, represents the total number of computing platforms, res represents the computing resource occupancy rate, bandwidth resource occupancy rate, or memory resource occupancy rate, represents the average utilization rate of heterogeneous service resources, represents the computing platform the utilization rate of heterogeneous service resources on.

[0098] Furthermore, in one embodiment, the expression of the reward function is:

[0099] In the formula, represents the reward value at the th time step in the th training round, represents the total number of instances of the service to be deployed, represents the comprehensive measure of the performance of the heterogeneous service orchestration network at the th time step in the th training round, respectively represent the weight parameters of the influence of the time step and training round on the iterative reward; represents the edge user request The average service delay represents the load balancing index of the heterogeneous service orchestration network respectively represent the weight parameters reflecting the impact of the average service delay and the load balancing index on the optimization objective function.

[0100] Among them, the functional implementation of each module in the above-mentioned AI service and microservice hybrid orchestration device in cloud-edge collaboration corresponds to each step in the embodiment of the above-mentioned AI service and microservice hybrid orchestration method in cloud-edge collaboration, and its functions and implementation processes will not be elaborated here one by one.

[0101] In a third aspect, an embodiment of the present application provides an AI service and microservice hybrid orchestration device in cloud-edge collaboration. The AI service and microservice hybrid orchestration device in cloud-edge collaboration can be a device with data processing functions such as a personal computer (PC), a laptop, a server, etc.

[0102] Referring to Figure 6 , Figure 6 is a schematic hardware structure diagram of the AI service and microservice hybrid orchestration device involved in the solution of the embodiment of the present application. In the embodiment of the present application, the AI service and microservice hybrid orchestration device in cloud-edge collaboration may include a processor, a memory, a communication interface, and a communication bus.

[0103] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.

[0104] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, etc., which are used to implement the interconnection of components inside the AI service and microservice hybrid orchestration device in cloud-edge collaboration, as well as interfaces used to implement the interconnection of the AI service and microservice hybrid orchestration device in cloud-edge collaboration with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display, a keyboard, etc.

[0105] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0106] The processor can be a general - purpose processor, which can call the AI service and microservice hybrid orchestration program stored in the memory in cloud - edge collaboration and execute the AI service and microservice hybrid orchestration method provided in the embodiments of the present application. For example, the general - purpose processor can be a central processing unit (CPU). Among them, the method executed when the AI service and microservice hybrid orchestration program in cloud - edge collaboration is called can refer to the various embodiments of the AI service and microservice hybrid orchestration method in cloud - edge collaboration in the present application, which will not be elaborated here.

[0107] Those skilled in the art can understand that Figure 6 the hardware structure shown in does not constitute a limitation to the present application, and may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0108] Fourthly, the embodiments of the present application further provide a computer - readable storage medium.

[0109] The computer - readable storage medium of the present application stores an AI service and microservice hybrid orchestration program in cloud - edge collaboration. When the AI service and microservice hybrid orchestration program in cloud - edge collaboration is executed by a processor, it realizes the steps of the AI service and microservice hybrid orchestration method as described above.

[0110] Among them, the method realized when the AI service and microservice hybrid orchestration program in cloud - edge collaboration is executed can refer to the various embodiments of the AI service and microservice hybrid orchestration method in cloud - edge collaboration in the present application, which will not be elaborated here.

[0111] The terms "including" and "having" and any variations thereof in the specification, claims and drawings of the present application are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. The descriptions of terms such as "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second" and "third" are of different types.

[0112] In the description of the embodiments of this application, words such as "exemplary", "for example", or "for illustration purposes" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for illustration purposes" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for illustration purposes" is intended to present relevant concepts in a specific manner.

[0113] In the description of the embodiments of this application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "a plurality of" means two or more than two.

[0114] In some of the processes described in the embodiments of this application, there are multiple operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of this application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in order or in parallel, and these operations or steps may be combined.

[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device to execute the methods described in the embodiments of this application.

[0116] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.

Claims

1. A method for hybrid orchestration of AI services and microservices in cloud-edge collaboration, characterized in that, Including the following steps: Obtain the target basic information and the target edge user request corresponding to the current heterogeneous service orchestration network, where the target basic information includes the target geographical location and the target time slot; Determine the heterogeneous service resource amount to be allocated corresponding to the full-size AI service according to the target edge user request, the preset heterogeneous service resource amount threshold, and the operation intensity of the full-size AI service; Allocate resources to the full-size AI service in the cloud computing layer of the current heterogeneous service orchestration network based on the heterogeneous service resource amount to be allocated; Determine the number of instances to be deployed corresponding to the lightweight AI service and the microservice according to the target edge user request; Screen out the basic masked deep Q-network (AMDQN) corresponding to the target basic information as the target AMDQN from the preset orchestration policy database, where the orchestration policy database contains multiple basic AMDQNs having a mapping relationship with the geographical location-time slot combination; Control the target AMDQN to interact with the current heterogeneous service orchestration network to deploy the lightweight AI service and the microservice corresponding to the number of instances to be deployed in the current heterogeneous service orchestration network; Among them, the generation method of the basic AMDQN is: based on the historical edge user request, the service deployment information, the network topology information, and the number of computing platforms in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination, construct the state space, the action space, and the optimization objective function respectively, and determine the historical number of instances to be deployed corresponding to the lightweight AI service and the microservice; construct the reward function according to the optimization objective function and the historical number of instances to be deployed; Construct an initial AMDQN based on the state space, the action space, and the reward function and perform training to generate the basic AMDQN.

2. The AI service and microservice hybrid orchestration method in cloud-edge collaboration according to claim 1, characterized in that, The determining the heterogeneous service resource amount to be allocated corresponding to the full-size AI service according to the target edge user request, the preset heterogeneous service resource amount threshold, and the operation intensity of the full-size AI service includes: Obtain the target perception result corresponding to the target edge user request based on the current heterogeneous service orchestration network; Calculate the first request intensity of the full-size AI service based on the target perception result; Determine the first inference efficiency of the full-size AI service according to the first request intensity; Calculate the first heterogeneous service resource amount through the first inference efficiency, the gain coefficient of the bandwidth resource, the gain coefficient of the computing resource, and the operation intensity of the full-size AI service; Calculate the second inference efficiency of the full-size AI service according to the maximum value between the first heterogeneous service resource amount and the preset heterogeneous service resource amount threshold; Determine the first priority for allocating heterogeneous service resources to each full-size AI service based on the second inference efficiency, the first request intensity, and the operation intensity of the full-size AI service; Determine the heterogeneous service resource amount to be allocated for each full-size AI service based on the first priority.

3. The AI service and microservice hybrid orchestration method in cloud-edge collaboration according to claim 1, wherein, The determining the number of instances to be deployed corresponding to the lightweight AI service and the microservice according to the target edge user request includes: Obtain the target perception result corresponding to the target edge user request based on the current heterogeneous service orchestration network; Calculate the second request intensity of all edge users for each lightweight AI service and the third request intensity of all edge users for each microservice based on the target perception result; Determine the first basic instance number of each lightweight AI service according to the second request intensity, and determine the second basic instance number of each microservice according to the third request intensity; Determine the second priority for increasing the instance number of each lightweight AI service through the first basic instance number, and determine the number of instances to be deployed for each lightweight AI service based on the first basic instance number and the second priority; Determine the maximum delay reduction amount for increasing one instance of each microservice according to the second basic instance number, determine the third priority for increasing the instance number of each microservice based on the sorting result of the maximum delay reduction amount, and determine the number of instances to be deployed for each microservice based on the second basic instance number and the third priority.

4. The AI service and microservice hybrid orchestration method in cloud-edge collaboration according to claim 1, wherein Constructing a state space, an action space, and an optimization objective function based on historical edge user requests, service deployment information and network topology information in the basic heterogeneous service orchestration network corresponding to a preset geographical location-time slot combination, and the number of computing platforms respectively, includes: Construct a state space based on the historical edge user requests, service deployment information and network topology information in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination; Construct an action space based on the number of computing platforms in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination; Construct an optimization objective function based on historical edge user requests and service deployment information in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination.

5. The AI service and microservice hybrid orchestration method in cloud-edge collaboration according to claim 4, wherein The constructing an optimization objective function based on historical edge user requests and service deployment information in the basic heterogeneous service orchestration network corresponding to the preset geographical location-time slot combination includes: Input the historical edge user requests into the basic heterogeneous service orchestration network to obtain historical perception results; Calculate the average service delay corresponding to the historical edge user requests according to the historical perception results; Calculate the load balancing index corresponding to the basic heterogeneous service orchestration network through the service deployment information in the basic heterogeneous service orchestration network; Determine the optimization objective function based on the average service delay and the load balancing index.

6. The AI service and microservice hybrid orchestration method in cloud-edge collaboration according to claim 5, wherein The calculation formula of the average service delay is: Wherein, represents the average service delay of the edge user request ; represents the set of all request paths of the edge user request ; represents the th service path of the edge user request ; represents whether there is an AI service deployed on the th computing platform ; represents whether there is a microservice deployed on the th computing platform ; represents the queue delay on the AI service of the edge user request on the th computing platform ; represents the processing delay on the microservice of the edge user request on the th computing platform ; represents the routing probability that after the AI service of the edge user request on the computing platform is processed, it is routed to the AI service on the computing platform for continued processing ; represents the arrival rate of the edge user request ; represents whether a full-scale AI service inference is required to complete a complete edge user request ; represents whether a lightweight AI service inference is required to complete a complete edge user request ; represents the unit time slot length ; represents the operation intensity of the full-scale AI service ; represents the operation intensity of the lightweight AI service ; represents the inference efficiency of the full-scale AI service ; represents the inference efficiency of the lightweight AI service ; represents that the edge user request is assigned to a complete service path The arrival rate on indicating the edge user request at the th computing platform for the AI service corresponding to the service chain The arrival rate on indicating the average communication delay during cross-computing platform processing.

7. The AI service and microservice hybrid orchestration method in cloud-edge collaboration according to claim 5, wherein The calculation formula of the load balancing index is: In the formula, is the load balancing index of the basic heterogeneous service orchestration network, respectively represent the sensitivity coefficients of the computing resource occupancy rate, the bandwidth resource occupancy rate, and the memory resource occupancy rate, respectively represent the variances of the computing resource occupancy rate, the bandwidth resource occupancy rate, and the memory resource occupancy rate, represents the total number of computing platforms, res represents the computing resource occupancy rate, the bandwidth resource occupancy rate, or the memory resource occupancy rate, represents the average utilization rate of heterogeneous service resources, represents the computing platform the utilization rate of heterogeneous service resources on it.

8. The AI service and microservice hybrid orchestration method in cloud-edge collaboration according to claim 1, wherein, The expression of the reward function is: Wherein, represents the reward value at the -th time step in the -th training round, represents the total number of instances of the service to be deployed, represents the comprehensive metric of the performance of the heterogeneous service orchestration network at the -th time step in the -th training round, respectively represent the weight parameters of the influence of the time step and the training round on the iterative reward; represents the average service delay of the edge user request , represents the load balancing index of the heterogeneous service orchestration network, respectively represent the weight parameters of the influence of the average service delay and the load balancing index on the optimization objective function.

9. An AI service and microservice hybrid orchestration device in cloud-edge collaboration, characterized in that, Includes: An information acquisition module, which is used to acquire target basic information and target edge user requests corresponding to the current heterogeneous service orchestration network, and the target basic information includes a target geographical location and a target time slot; A resource determination module, which is used to determine the heterogeneous service resource amount to be allocated corresponding to the full-size AI service according to the target edge user requests, a preset heterogeneous service resource amount threshold, and the operation intensity of the full-size AI service; A resource allocation module, which is used to allocate resources to the full-size AI service in the cloud computing layer of the current heterogeneous service orchestration network based on the heterogeneous service resource amount to be allocated. An instance calculation module, which is used to determine the number of to-be-deployed instances corresponding to the lightweight AI service and microservices according to the target edge user request; A model screening module, which is used to screen out the basic invalid action mask deep Q-network AMDQN corresponding to the target basic information from a preset orchestration policy database, and the orchestration policy database contains multiple basic AMDQNs having a mapping relationship with the geographical location-time slot combination; A hybrid orchestration module, which is used to control the target AMDQN to interact with the current heterogeneous service orchestration network, so as to deploy the lightweight AI service and microservices corresponding to the number of to-be-deployed instances in the current heterogeneous service orchestration network; A model generation module, which is used to respectively construct a state space, an action space and an optimization objective function based on historical edge user requests, service deployment information in the basic heterogeneous service orchestration network corresponding to a preset geographical location-time slot combination, network topology information, and the number of computing platforms, and determine the number of historical to-be-deployed instances corresponding to the lightweight AI service and microservices; construct a reward function according to the optimization objective function and the number of historical to-be-deployed instances; Construct an initial AMDQN based on the state space, action space and reward function and train it to generate a basic AMDQN.

10. An AI service and microservice hybrid orchestration device in cloud-edge collaboration, characterized in that, The AI service and microservice hybrid orchestration device in cloud-edge collaboration includes a processor, a memory, and an AI service and microservice hybrid orchestration program in cloud-edge collaboration stored on the memory and executable by the processor. When the AI service and microservice hybrid orchestration program in cloud-edge collaboration is executed by the processor, the steps of the AI service and microservice hybrid orchestration method in cloud-edge collaboration as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Business management method and system, configuration server and edge computing device

    CN114979246A

  • Space-based network edge micro-service distributed arrangement system and method

    CN115622893A

  • Heterogeneous multi-edge cloud collaborative micro-service deployment and routing joint optimization method and system

    CN116915686A

  • Synchronization of artificial intelligence based microservices

    US20220272794A1

Cited By

  • Micro-service elastic telescopic arrangement method and arrangement device

    CN121078133A

  • Micro-service migration method, device and equipment and computer readable storage medium

    CN121644675A