A Heterogeneous Resource Cooperative Scheduling Method for Training AI Inference Services

By adopting the resource scheduling method of MILP model and greedy strategy in NG-RAN, the problems of heterogeneous resource management and AI inference service quality are solved, and efficient resource utilization and service processing efficiency are achieved.

CN119545437BActive Publication Date: 2025-06-27BEIJING JINLOU CENTURY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411658903.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-06-27
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

In NG-RAN supported by ECFN, how to effectively manage and schedule heterogeneous resources to support efficient AI inference services, especially in the connection between DU and CU, how to avoid resource waste and service quality decline.

Method used

A heterogeneous resource collaborative scheduling method based on the hybrid integer linear programming (MILP) model is proposed, combined with a collaborative heterogeneous resource scheduling algorithm (CHRS-GS) with greedy strategies, to maximize the number of successfully sent AI inference subtasks, minimize the utilization rate of heterogeneous resources and the backlog of AI inference services.

Benefits of technology

By optimizing resource allocation, the processing efficiency of AI inference services in NG-RAN is improved, resource waste and service backlog are reduced, and overall service quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119545437B_ABST
    Figure CN119545437B_ABST
Patent Text Reader

Abstract

The present invention discloses a heterogeneous resource collaborative scheduling method for training AI inference services, comprising the following steps: establishing a mixed integer linear programming model MILP based on the AI inference collaborative heterogeneous resource scheduling problem of distributed model training, and the ultimate goal of the proposed MILP is to maximize the number of successfully sent AI inference subtasks and minimize the usage rate of heterogeneous resources in the NG-RAN and the backlog of AI inference services; constructing the edge computing resource request, computing processing capacity, task backlog threshold of AI inference subtasks on the edge server, and complaint rate threshold of AI inference subtasks as constraint conditions; standardizing the deployment of DU-CU in the network architecture; adjusting the wavelength allocation in the network; and specifying the capacity requirements of the network. The present invention can achieve maximizing the number of successfully sent AI inference service subtasks, while minimizing the heterogeneous resource occupancy rate and the backlog of AI inference service subtasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to the field of mobile communication technologies, and particularly to a heterogeneous resource collaborative scheduling method for training AI inference services. Background Art:

[0002] The next-generation (NG) mobile communication is approaching, which will lead us into an unprecedented era with higher data transmission rates, shorter round-trip delays, more diverse service scenarios, and integration with advanced technologies such as artificial intelligence (AI) and machine learning (ML). These heterogeneous requirements pose new challenges to the capacity, latency, and network flexibility of the radio access network (RAN), accelerating the evolution of RAN towards the next-generation radio access network (NG-RAN). In the NG-RAN architecture, the baseband unit (BBU) - active antenna unit (AAU) is redefined as a three-layer network, including a distributed processing unit (DU), a central processing unit (CU), and an AAU. Functions sensitive to latency below the packet data convergence protocol (PDCP) layer are processed by the DU, while functions above the PDCP layer are managed by the CU. The connection between the DU and the CU is called the midhaul network, usually mediated by an optical network.

[0003] AI inference is the process of applying a trained AI model to new data for analysis and decision-making, and it has become a core technology in multiple industries. With the progress of computing power and big data technologies, the application scope of AI inference has expanded to fields such as medical diagnosis, financial risk assessment, autonomous driving, and intelligent manufacturing. Currently, the development direction of AI inference mainly focuses on real-time processing and low latency, model compression and optimization, multi-modal inference, etc. Despite significant progress, AI inference still faces challenges such as high computing resource requirements and model adaptability. AI inference based on distributed model training has significant performance advantages, including faster processing speeds and higher scalability, which benefit from parallel computing and resource optimization across multiple nodes.

[0004] NG-RAN has higher spectral efficiency and scalability, can handle more data traffic, and support a wider range of AI inference services. In addition, the optical network has excellent performance of high bandwidth and low latency. The RAN network has rich edge computing power resources that can be used for AI inference model training. The integrated optical network and NG-RAN are considered to be an efficient network for carrying AI inference services. Therefore, we propose a new network architecture based on an edge computing power network (ECFN) for distributed model training for AI inference.

[0005] However, the optimal utilization of heterogeneous resources and the quality of service (QoS) of AI inference services are crucial. Therefore, we mainly focus on efficient heterogeneous resource collaborative scheduling strategies to support the wide application of AI inference based on distributed model training while meeting the high service quality requirements of AI inference.

[0006] With the wide application of AI inference in various fields, the demand for computing resources and bandwidth resources is increasing continuously. Especially in the NG-RAN supported by ECFN, as the network architecture evolves into DU and CU, how to effectively manage and schedule these heterogeneous resources has become a key challenge. Currently, many resource scheduling methods cannot achieve the collaborative allocation of heterogeneous resources, resulting in improper resource allocation and further causing waste of computing and transmission resources. In addition, whether the AI inference subtasks are successfully sent will also affect the user's QoS, leading to task backlogs and affecting the service quality. Therefore, the number of successfully sent AI inference subtasks should also be considered. Therefore, it is urgent to explore heterogeneous resource collaborative scheduling strategies based on ECFN-supported NG-RAN to improve the efficiency of heterogeneous resources for AI inference services. Summary of the Invention:

[0007] The purpose of the present invention is to provide a heterogeneous resource collaborative scheduling method for AI inference services based on distributed model training to solve the deficiencies of the prior art.

[0008] The present invention is implemented by the following technical solutions: A heterogeneous resource collaborative scheduling method for training AI inference services includes the following steps:

[0009] A mixed-integer linear programming model MILP is established based on the heterogeneous resource scheduling problem of AI inference collaboration for distributed model training, and the ultimate goal of the proposed MILP is to maximize the number of successfully sent AI inference subtasks and minimize the utilization rate of heterogeneous resources in NG-RAN and the backlog of AI inference services.

[0010] Construct the edge computing resource request, computing processing capacity, task backlog threshold of AI inference subtasks on the edge server, and complaint rate threshold of AI inference subtasks as constraint conditions;

[0011] Standardize the deployment of DU-CU in the network architecture to ensure that the baseband processing function of each user service is only deployed on one GPP node, and specify the locations of RU, DU-CU, and MEC;

[0012] Adjust the wavelength allocation in the network, establish the necessary optical links from the access AAU to the MEC through the selected GPP nodes, and establish constraint conditions to ensure the continuity of each service request in the service flow and prevent the formation of loop paths in the network;

[0013] Specify the capacity requirements of the network, including restricting the total processing resources required for all services mapped to the GPP nodes not to exceed its available capacity, specifying that the links between the required endpoints must have sufficient bandwidth to meet service requests, specifying that the total processing resources required for all services mapped to the r-th node shall not exceed its available capacity, specifying that the edge data center m has sufficient capacity to meet all service requests, and specifying that each AI inference sub-task shall meet the total latency requirement.

[0014] Furthermore, it also includes resource allocation based on a collaborative heterogeneous resource scheduling algorithm with a greedy strategy, specifically:

[0015] Sort the computing power demand quantities in descending order to obtain the first list;

[0016] Sort the remaining edge computing power resources of the edge data centers in ascending order to obtain the second list;

[0017] Construct the remaining edge computing power resources of the edge data centers;

[0018] Allocate the computing power requests to the edge data centers;

[0019] Calculate the computing power resource usage, the number of successfully matched AI inference sub-tasks, the backlog of AI inference sub-tasks, and the joint optimization result of the three.

[0020] Furthermore, the step of allocating the computing power requests to the edge data centers is specifically:

[0021] Search from all users in the order of decreasing computing power demand quantity. If the computing resource request is positive, and the computing power of each edge data center meets the computing resources requested by each user, and at the same time the data processing capacity complaint rate of all users allocated to the j-th edge data center of the i-th computing power provider satisfies D i,j , then the f-th AI inference service of the k-th user is allocated to the j-th edge data center of the i-th computing power provider;

[0022] Assign the value of C i,j ′ - B k,f to C i,j ′. If there is not enough edge computing power in C i,j ′ for B k,f to use, and the data processing capacity complaint rate of all users allocated to the j-th edge data center of the i-th computing power provider satisfies D i,j , then allocate the C i,j ′ on the j-th edge data center of the i-th computing power provider to the f-th AI inference service of the k-th user;

[0023] Assign 0 to C i,j ′, (B k,f-C i,j ') is allocated to B k,f until all the computing power requests of B k,f are allocated to the edge data center and the algorithm ends.

[0024] Furthermore, the ultimate goal of the MILP is to maximize the number of successfully sent AI inference subtasks and minimize the utilization rate of heterogeneous resources in the NG-RAN and the backlog of AI inference services, which is specifically expressed by the formula:

[0025]

[0026] is to minimize the resource utilization rates of PP nodes, links, and edge data centers in the NG-RAN; where, V k,f,s,r is a boolean variable, which is 1 if the f-th type of AI inference service of the k-th user is processed by the s-th baseband function and passes through the r-th node; is the computing and processing requirement of the s-th segment of the baseband processing function of the f-th type of AI inference service of the k-th user, and U r is the computing power processing capacity of the r-th node; is a boolean variable, which is 1 if the f-th type of AI inference service of the k-th user occupies the w-th wavelength of the link (m, n) after being processed by the s-th baseband function; τ k,f,s is the bandwidth requirement of the f-th type of AI inference service of the k-th user after being processed by the s-th baseband function; U l is the bandwidth capacity of the l-th link; m k,f is a boolean variable, which is 1 if the f-th type of AI inference service of the k-th user uses the edge data center of the m-th center; η k,f is the computing power required for the f-th type of AI inference service of the k-th user; U m is the computing power processing capacity of the m-th edge data center;

[0027] is to maximize the number of successfully sent AI inference subtasks; where, S i,j is the computing power matching success rate of the j-th edge data center of the i-th computing power provider; is the computing power requirement (FLOPS / s) of the f-th type of AI inference service request of the k-th user to be allocated to the j-th edge data center of the i-th computing power provider;

[0028] is to minimize the backlog of AI inference services; where, σ i,j is the backlog of AI inference services (FLOPS / s) of the j-th edge data center of the i-th edge computing power provider; T is a task computing power processing cycle;

[0029] α, β, and γ are weight coefficients.

[0030] Furthermore, the construction of the edge computing resource request, computing processing power, the task backlog threshold of the AI inference subtask on the edge server, and the complaint rate threshold of the AI inference subtask are used as constraints, specifically:

[0031]

[0032] Among them, in formula (2), is the computing power requirement (FLOPS / s) of the f - type AI inference service request of the k - th user allocated to the j - th edge data center of the i - th computing power provider, B k,f is the f - type AI inference service request (FLOPS / s) of the k - th user; in formula (3), C i,j is the edge computing capacity processing power (FLOPS / s) of the j - th edge data center of the i - th computing power provider; in formula (4), ε is the backlog threshold (FLOPS) of the AI inference subtask in the edge data center; in formula (5), D i,j is the data processing capacity complaint rate (statistical value) of the j - th edge data center of the i - th computing power provider, J i,j is the remaining computable capacity resource amount (FLOPS) that can be appealed for the j - th edge data center of the i - th computing power provider.

[0033] Furthermore, the deployment of DU - CU in the network architecture is standardized to ensure that the baseband processing function of each user service is only deployed on one GPP node, and the positions of RU, DU - CU, and MEC are specified, specifically:

[0034]

[0035]

[0036] ∑ k,f,r Y k,f,r ≤∑ r,p G r,p ≤∑xF r ≤N r (8)

[0037]

[0038] These constraints also ensure the feasibility of photoelectric conversion and GPP node activation; among them, in formula (6), O k,f,s,r,p is a Boolean variable, which is 1 if the f - type AI inference service of the k - th user is deployed on GPPp of the r - th node after passing through the s - th baseband processing; in formula (7), λ k,f,ris a Boolean variable, which is 1 if the source node of the f-th AI service of the k-th user is deployed on r; in Equation (8), Y k,f,r is a Boolean variable, which is 1 if the OEO optical-electric conversion of the f-th AI inference service of the k-th user is performed on the r-th node; G r,p is a Boolean variable, which is 1 if the p-th GPP in the r-th node is activated; F r is a Boolean variable, which is 1 if the r-th node is activated; N r is the total number of nodes.

[0039] Furthermore, the wavelength allocation in the adjustment network is specifically as follows:

[0040]

[0041] if λ k,f,r = 0, k, f, j < sp, r (12)

[0042]

[0043] Equations (11) and (12) ensure the coordination between the DU-CU layout and the optical path allocation, and establish the necessary optical links from the access AAU to the MEC through the selected GPP nodes; where, O k,f,s,r,p is a Boolean variable, which is 1 if the f-th AI inference service of the k-th user is deployed on the GPPp on the r-th node after being processed by the s-th baseband; is a Boolean variable, which is 1 if the f-th AI inference service of the k-th user occupies the w-th wavelength of the link (m, n) after being processed by the s-th baseband function; N m,n is a Boolean variable, which is 1 if the m-th node is connected to the n-th node; λ k,f,r is a Boolean variable, which is 1 if the source node of the f-th AI service of the k-th user is deployed on r; Equation (13) ensures the continuity of each service request in the service flow and avoids any interruption during the transmission; where, M r is a Boolean variable, which is 1 if MEC is deployed on the r-th node; Equation (14) can prevent the formation of loop paths in the network.

[0044] Furthermore, the capacity requirements of the specified network are specifically as follows:

[0045]

[0046] Equation (15) ensures that the total processing resources required for all services mapped to the GPP nodes do not exceed their available capacity, where, O k,f,s,r,pis a Boolean variable, which is 1 if the f-th AI inference service of the k-th user is deployed on the GPPp of the r-th node after being processed by the s-th baseband; is the computing and processing requirement of the s-th segment of the baseband processing function of the f-th AI inference service of the k-th user; C p is the GPP computing power processing capacity (FLOPS / s) of the r-th node; Equation (16) stipulates that the link between the required endpoints must have sufficient bandwidth to meet the service request, where, is a Boolean variable, which is 1 if the f-th AI inference service of the k-th user occupies the w-th wavelength of the link (m, n) after being processed by the s-th baseband function; τ k,f,s is the bandwidth requirement of the f-th type of AI inference service of the k-th user after being processed by the s-th baseband function; U l is the bandwidth capacity of the l-th link; Equation (17) stipulates that the total sum of the processing resources required for all services mapped to the r-th node shall not exceed its available capacity, where, V k,f,s,r is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user passes through the r-th node after being processed by the s-th baseband function; is the computing and processing requirement of the s-th segment of the baseband processing function of the f-th AI inference service of the k-th user; U r is the computing power processing capacity of the r-th node; Finally, Equation (18) ensures that the edge data center m has sufficient capacity to meet all service requests; where, m k,f is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user uses the edge data center of the m-th center; η k,f is the computing power required for the f-th AI inference service of the k-th user; U m is the computing power processing capacity of the m-th edge data center; Equation (19) stipulates that each AI inference subtask should meet the total delay requirement; where, is a Boolean variable, which is 1 if the f-th AI inference service of the k-th user occupies the w-th wavelength of the link (m, n) after being processed by the s-th baseband function; L m,n is the link distance of (m, n); φ is the fiber propagation delay per kilometer (i.e., 5 μs / km); Y k,f,r is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user undergoes OEO optoelectronic conversion at the r-th node; T oeo is the delay of OEO and electronic switches; G r,p is a Boolean variable, which is 1 if the p-th GPP in the r-th node is activated; T v is the virtual function processing delay; t bv_oxcis the transmission delay of the variable bandwidth optical cross-connector; F r is a Boolean variable, which is 1 if the r-th node is activated; m k,f is a Boolean variable, which is 1 if the edge data center of the m-th center is used for the f-th type of AI inference service of the k-th user; λ k,f,r is a Boolean variable, which is 1 if the source node of the f-th type of AI service of the k-th user is deployed on r; T MEC is the delay for the edge data center to process the f-th type of AI inference service; T k,f is the total delay requirement of the f-th type of AI inference service of the k-th user for computing power.

[0047] Advantages of the present invention:

[0048] The present invention explores a heterogeneous resource collaborative scheduling strategy based on ECFN to improve the processing efficiency of AI inference services in NG-RAN, proposes a mixed integer linear programming mathematical model aiming to maximize the number of successfully sent AI inference service subtasks while minimizing the heterogeneous resource occupancy rate and the backlog of AI inference service subtasks. In addition, an efficient heuristic algorithm named CHRS-GS is proposed, which adopts a greedy strategy for resource allocation. Description of the drawings:

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0050] Figure 1 is the NG-RAN diagram supported by ECFN in the embodiment of the present invention; Specific implementation manners:

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0052] As Figure 1 shown, in the NG-RAN supported by ECFN of the present invention, a mixed integer linear programming (MILP) model is established for the AI inference collaborative heterogeneous resource scheduling problem based on distributed model training.

[0053] First, the parameter definitions of the MILP model are introduced in detail. For the constant parameters shown in Table I, the set of edge suppliers and the set of edge data centers are denoted as S and E, respectively. The set of users and the set of different types of AI inference are denoted as U and R, respectively, where k ∈ U and f ∈ R. In addition, ε is defined as the backlog threshold of the distributed model training tasks in the edge data centers.

[0054] For the variable parameters shown in Table II. The computing processing power (GOPS) of the j-th edge data center of the i-th computing power supplier is denoted as C i,j , where i ∈ S and j ∈ E. B k,f represents the computing resources (GOPS) required for the f-th AI inference service of the k-th user, where k ∈ U and f ∈ R. P k,f represents the complaint rate of the distributed model training tasks, which is a statistical value. J i,j and S i,j represent the remaining computing resources and the computing power matching success rate of the j-th edge data center of the i-th computing power supplier, respectively.

[0055] For the decision variables shown in Table III. A boolean variable is defined as If the computing power resources requested for the f-th AI inference service of the k-th user are allocated to the j-th edge data center of the i-th edge computing power supplier within the t-th time frame, then equals 1. is defined as a boolean variable, which equals 1 if the j-th edge data center of the i-th edge computing power supplier is used within the t-th time frame. is defined as a boolean variable, which equals 1 if the j-th edge data center of the i-th edge computing power supplier is used within the t-th time frame.

[0056] Table I

[0057]

[0058] Table II

[0059]

[0060]

[0061] Table III

[0062]

[0063]

[0064] The ultimate goal of the MILP proposed in this invention is to maximize the number of successfully transmitted AI inference subtasks and minimize the utilization rate of heterogeneous resources in the NG-RAN and the backlog of AI inference services, as shown in Formula (1).

[0065]

[0066] To minimize the resource utilization rate of PP nodes, links, and edge data centers in the NG-RAN; where V k,f,s,r is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user is processed by the s-th baseband function and passes through the r-th node; is the computing and processing requirement of the s-th segment of the baseband processing function for the f-th type of AI inference service of the k-th user, U r is the computing power processing capacity of the r-th node; is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user occupies the w-th wavelength of the link (m,n) after being processed by the s-th baseband function; τ k,f,s is the bandwidth requirement of the f-th type of AI inference service of the k-th user after being processed by the s-th baseband function; U l is the bandwidth capacity of the l-th link; m k,f is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user uses the edge data center of the m-th center; η k,f is the computing power required for the f-th type of AI inference service of the k-th user; U m is the computing power processing capacity of the m-th edge data center;

[0067] To maximize the number of successfully transmitted AI inference subtasks; where S i,j The computing power matching success rate of the j-th edge data center of the i-th computing power provider; The computing power requirement (FLOPS / s) of the f-th type of AI inference service request of the k-th user allocated to the j-th edge data center of the i-th computing power provider;

[0068] To minimize the backlog of AI inference services; where σ i,j The backlog of AI inference services (FLOPS / s) in the j-th edge data center of the i-th edge computing power provider; T is a task computing power processing cycle; α, β, γ are weight coefficients.

[0069] In addition, aspects such as computing power service distribution, DU-CU deployment, optical wavelength allocation, capacity, and latency constraints should also be considered.

[0070] Constraints such as edge computing resource requests, computing processing capabilities, task backlog thresholds for AI inference subtasks on edge servers, and complaint rate thresholds for AI inference subtasks need to be considered, as shown in formulas (2)-(5) respectively.

[0071]

[0072] Among them, in formula (2), is the computing power requirement (FLOPS / s) allocated to the j-th edge data center of the i-th computing power provider for the f-th type of AI inference service request of the k-th user, and B k,f is the f-th type of AI inference service request (FLOPS / s) of the k-th user; in formula (3), C i,j is the edge computing capacity processing capacity (FLOPS / s) of the j-th edge data center of the i-th computing power provider; in formula (4), ε is the backlog threshold (FLOPS) of AI inference subtasks in the edge data center; in formula (5), D i,j is the data processing capacity complaint rate (statistical value) of the j-th edge data center of the i-th computing power provider, and J i,j is the remaining computable capacity resource amount (FLOPS) that can be appealed in the j-th edge data center of the i-th computing power provider.

[0073] Formulas (6)-(10) comprehensively regulate the deployment of DU-CU in the network architecture. Specifically, they ensure that the baseband processing function of each user service is only deployed on one GPP node and specify the locations of RU, DU-CU, and MEC. In addition, these constraints also ensure the feasibility of optical-electric conversion and GPP node activation, which is crucial for the deployment of DU-CU. In summary, these constraints effectively limit and optimize the deployment of DU-CU, meet the system requirements, and improve the overall performance of the network.

[0074]

[0075] ∑ k,f,r Y k,f,r ≤∑ r,p G r,p ≤∑ r F r ≤N r (8)

[0076]

[0077] These constraints also ensure the feasibility of optical-electric conversion and GPP node activation; among them, in formula (6), O k,f,s,r,pis a Boolean variable, which is 1 if the f-th AI inference service of the k-th user is deployed on GPPp on the r-th node after being processed by the s-th baseband; in Equation (7), λ k,f,r is a Boolean variable, which is 1 if the source node of the f-th AI service of the k-th user is deployed on r; in Equation (8), Y k,f,R is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user has undergone OEO optoelectronic conversion on the r-th node; G r,p is a Boolean variable, which is 1 if the p-th GPP in the r-th node is activated; F r is a Boolean variable, which is 1 if the r-th node is activated; N r is the total number of nodes.

[0078] Equations (11)-(14) regulate the wavelength allocation process in the network. Equations (11) and (12) ensure the coordination between the DU-CU layout and the optical path allocation, and establish the necessary optical links from the access AAU to the MEC through the selected GPP nodes. This integration of the DU-CU layout and the optical path allocation is crucial for the seamless operation of the network. Then, Equation (13) guarantees the continuity of each service request in the service flow and avoids any interruption during the transmission. Finally, Equation (14) prevents the formation of loop paths in the network.

[0079]

[0080] if λ k,f,r = 0, k, f, j < sp, r (12)

[0081]

[0082] Equations (11) and (12) ensure the coordination between the DU-CU layout and the optical path allocation, and establish the necessary optical links from the access AAU to the MEC through the selected GPP nodes; where, O k,f,s,r,p is a Boolean variable, which is 1 if the f-th AI inference service of the k-th user is deployed on GPPp on the r-th node after being processed by the s-th baseband; is a Boolean variable, which is 1 if the f-th AI inference service of the k-th user occupies the w-th wavelength of the link (m, n) after being processed by the s-th baseband function; N m,n is a Boolean variable, which is 1 if the m-th node is connected to the n-th node; λ k,f,r is a Boolean variable, which is 1 if the source node of the f-th AI service of the k-th user is deployed on r; Equation (13) guarantees the continuity of each service request in the service flow and avoids any interruption during the transmission; where, M ris a Boolean variable, which is 1 if MEC is deployed at the r-th node; Equation (14) can prevent the formation of cyclic paths in the network.

[0083] Equations (15)-(18) specify the capacity requirements of the network. Equation (15) ensures that the total processing resources required for all services mapped to the GPP node do not exceed its available capacity. Equation (16) stipulates that the link between the required endpoints must have sufficient bandwidth to meet the service requests. Equation (17) stipulates that the total processing resources required for all services mapped to the r-th node do not exceed its available capacity. Finally, Equation (18) guarantees that the edge data center m has sufficient capacity to meet all service requests.

[0084]

[0085] Equation (19) stipulates that each AI inference subtask should meet the total latency requirement.

[0086]

[0087]

[0088] Equation (15) ensures that the total processing resources required for all services mapped to the GPP node do not exceed its available capacity, where, O k,f,s,r,p is a Boolean variable, which is 1 if the f-th AI inference service of the k-th user is deployed on GPPp at the r-th node after being processed by the s-th baseband; is the computing and processing requirement of the s-th segment of the baseband processing function of the f-th AI inference service of the k-th user; C p is the GPP computing power processing capacity (FLOPS / s) of the r-th node; Equation (16) stipulates that the link between the required endpoints must have sufficient bandwidth to meet the service requests, where, is a Boolean variable, which is 1 if the f-th AI inference service of the k-th user occupies the w-th wavelength of the link (m,n) after being processed by the s-th baseband function; τ k,f,s is the bandwidth requirement of the f-th type of AI inference service of the k-th user after being processed by the s-th baseband function; U l is the bandwidth capacity of the l-th link; Equation (17) stipulates that the total processing resources required for all services mapped to the r-th node do not exceed its available capacity, where, V k,f,s,r is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user passes through the r-th node after being processed by the s-th baseband function; is the computing and processing requirement of the s-th segment of the baseband processing function of the f-th AI inference service of the k-th user; U ris the computing power processing capacity of the r-th node; finally, formula (18) ensures that the edge data center m has sufficient capacity to meet all service requests; where, m k,f is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user uses the edge data center of the m-th center; η k,f is the computing power required for the f-th type of AI inference service of the k-th user; U m is the computing power processing capacity of the m-th edge data center; formula (19) stipulates that each AI inference subtask should meet the total delay requirement; where, is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user occupies the w-th wavelength of the link (m,n) after being processed by the s-th baseband function; L m,n is the link distance of (m,n); φ is the fiber propagation delay per kilometer (i.e., 5 μs / km); Y k,f,r is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user performs OEO optoelectronic conversion at the r-th node; T oeo is the delay of OEO and electronic switches; G r,p is a Boolean variable, which is 1 if the p-th GPP in the r-th node is activated; T v is the virtual function processing delay; t bv_oxc is the transmission delay of the variable bandwidth optical cross-connector; F r is a Boolean variable, which is 1 if the r-th node is activated; m k,f is a Boolean variable, which is 1 if the f-th type of AI inference service of the k-th user uses the edge data center of the m-th center; λ k,f,r is a Boolean variable, which is 1 if the source node of the f-th type of AI service of the k-th user is deployed on r; T MEC is the delay for the edge data center to process the f-th type of AI inference service; T k,f is the total delay requirement for the computing power of the f-th type of AI inference service of the k-th user.

[0089] The present invention also performs resource allocation based on a collaborative heterogeneous resource scheduling algorithm with a greedy strategy. The common optimization goal of the heuristic algorithm is to maximize the number of successfully sent AI inference service subtasks, minimize the usage rate of heterogeneous resources in NG-RAN and the backlog of AI inference subtasks. The proposed heuristic algorithm uses a greedy thinking strategy to allocate the AI inference service with the largest computing demand to the edge data center with the least computing resources. The pseudocode of the collaborative heterogeneous resource scheduling algorithm based on the greedy strategy (CHRS-GS) is shown in Algorithm Table Ⅳ.

[0090] The edge computing resource requests M for all AI inference services k,fSort in descending order according to the number of computing resource requirements to obtain a list X. Sort in ascending order according to the remaining computing resources N of the edge data center i,j to obtain list Y. Regard the remaining computing resources of the edge data center as (C i,j +σ i,j )(lines 1-3 of Table IV). Search all users in descending order of computing power demand. If the computing resource request is positive, and the computing power of each edge data center meets the computing resources requested by each user, and at the same time the data processing capacity complaint rate of all users assigned to the j-th edge data center of the i-th computing power provider satisfies D i,j , then the f-th AI inference service of the k-th user is assigned to the j-th edge data center of the i-th computing power provider. Then, assign the value of C i,j ′-B k,f to C i,j ′(lines 4-11 of Table IV). If there is not enough edge computing power in C i,j ′ for B k,f to use, and the data processing capacity complaint rate of all users assigned to the j-th edge data center of the i-th computing power provider satisfies D i,j , then assign C i,j ′ on the j-th edge data center of the i-th computing power provider to the f-th AI inference service of the k-th user. Then, assign 0 to C i,j ′, and assign (B k,f -C i,j ′) to B k,f . Until all computing power requests of B k,f are assigned to the edge data center, the algorithm ends (lines 12-21 of Table IV). Third, calculate the computing power resource usage, the number of successfully matched AI inference subtasks, the backlog of AI inference subtasks, and the joint optimization results of the three (line 22 of Table IV)

[0091] Table IV

[0092]

[0093]

[0094] In summary, the present invention explores a heterogeneous resource collaborative scheduling strategy based on ECFN to improve the processing efficiency of AI inference services in NG-RAN, proposes a mixed-integer linear programming mathematical model aiming to maximize the number of successfully sent AI inference service subtasks, while minimizing the heterogeneous resource occupancy rate and the backlog of AI inference service subtasks. In addition, an efficient heuristic algorithm called CHRS-GS is proposed, which uses a greedy strategy for resource allocation.

[0095] The present invention shows that ECFN has great potential in improving the execution efficiency of AI inference services, and also highlights the importance of optimizing resource scheduling in achieving high-performance network services. Future research can further explore how to apply these scheduling strategies to more complex network environments and consider more dynamic factors and user requirements to promote the development of intelligent networks.

[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for coordinating heterogeneous resources for training AI reasoning services, characterized in that: The following steps are involved: A mixed integer linear programming model MILP is established for the AI ​​inference collaborative heterogeneous resource scheduling problem based on distributed model training. The ultimate goal of the proposed MILP is to maximize the number of successfully sent AI inference subtasks and minimize the utilization rate of heterogeneous resources and the backlog of AI inference services in NG-RAN. Construct edge computing resource requests, computing processing capabilities, task backlog thresholds for AI reasoning subtasks on edge servers, and complaint rate thresholds for AI reasoning subtasks as constraints; Standardize the deployment of DU-CU in the network architecture, ensure that the baseband processing function of each user service is deployed on only one PP node, and specify the locations of RU, DU-CU and MEC; Adjust the wavelength allocation in the network, establish the necessary optical links from the access AAU to the MEC through the selected PP nodes, and establish constraints to ensure the continuity of each service request in the service flow and prevent the formation of loop paths in the network; Specify the capacity requirements of the network, including limiting the sum of processing resources required for all services mapped to the PP node to not exceed its available capacity, specifying that the links between the required endpoints must have sufficient bandwidth to meet service requests, specifying that the sum of processing resources required for all services mapped to the rth node must not exceed its available capacity, specifying that the edge data center m has sufficient capacity to meet all service requests, and specifying that each AI reasoning subtask should meet the total delay requirements.

2. According to claim 1, a heterogeneous resource collaborative scheduling method for training AI reasoning services is characterized in that: It also includes a cooperative heterogeneous resource scheduling algorithm based on a greedy strategy for resource allocation, specifically: Sort the computing power requirements in descending order to get the first list; Sort the remaining edge computing resources of the edge data center in ascending order to obtain a second list; Build the remaining edge computing resources of the edge data center; Allocate computing power requests to edge data centers; Calculate the usage of computing resources, the number of successfully matched AI reasoning subtasks, the backlog of AI reasoning subtasks, and the joint optimization results of the three.

3. According to claim 1, a heterogeneous resource collaborative scheduling method for training AI reasoning services is characterized in that: The ultimate goal of the MILP is to maximize the number of successfully sent AI reasoning subtasks and minimize the utilization of heterogeneous resources in NG-RAN and the backlog of AI reasoning services. The specific formula is expressed as follows: To minimize the resource usage of PP nodes, links and edge data centers in NG-RAN; V k,f,s,r is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user passes through the r-th node after being processed by the s-th baseband function, it is 1; The computation and processing requirements of the sth segment of baseband processing function serving the fth AI inference service for the kth user, U r is the computing power processing capacity of the rth node; is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user is processed by the s-th baseband function and occupies the w-th wavelength of the link (m, n), then it is 1; τ k,f,s The bandwidth requirement for the f-th AI reasoning service of the k-th user after being processed by the s-th baseband function; U l is the bandwidth capacity of the lth link; m k,f is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user uses the m-th edge data center, it is 1; η k,f The computing power required to serve the fth AI reasoning service for the kth user; U m is the computing power processing capacity of the mth edge data center; To maximize the number of successfully sent AI reasoning subtasks; where S i,j is the computing power matching success rate of the jth edge data center of the i-th computing power provider; The computing power demand allocated to the jth edge data center of the i-th computing power provider for the f-th AI reasoning service request of the k-th user; To minimize the backlog of AI reasoning services; where σ i,j is the AI ​​reasoning service backlog of the jth edge data center of the i-th edge computing provider; T is a task computing processing cycle; α, β, and γ are weight coefficients.

4. According to claim 1, a heterogeneous resource collaborative scheduling method for training AI reasoning services is characterized in that: The edge computing resource request, computing processing capacity, task backlog threshold of AI reasoning subtasks on edge servers, and complaint rate threshold of AI reasoning subtasks are constructed as constraints, specifically: Among them, in formula (2), The computing power demand of the jth edge data center allocated to the i-th computing power provider for the f-th AI reasoning service request of the k-th user, B k,f is the f-th AI reasoning service request of the k-th user; in formula (3), C i,j is the edge computing processing capacity of the jth edge data center of the i-th computing power provider; in formula (4), ε is the backlog threshold of the AI ​​reasoning subtask of the edge data center; in formula (5), D i,j is the data processing capacity complaint rate of the jth edge data center of the i-th computing power provider, J i,j The remaining computing power resources that can be appealed by the j-th edge data center of the i-th computing power provider.

5. According to claim 1, a method for coordinating heterogeneous resources for training AI reasoning services is characterized in that: The specification specifies the deployment of DU-CU in the network architecture, ensuring that the baseband processing function for each user service is deployed on only one PP, and specifies the locations of RU, DU-CU and MEC, specifically: These constraints also ensure the feasibility of photoelectric conversion and GPP node activation; where, in formula (6), O k,f,s,r,p is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user is deployed on the p-th GPP on the r-th node after being processed by the s-th baseband, it is 1; In formula (7), λ k,f,r is a Boolean variable, indicating that if the source node of the fth AI service of the kth user is deployed on the rth node, it is 1; in formula (8), Y k,f,r is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user performs OEO photoelectric conversion at the r-th node, it is 1; G r,p is a Boolean variable, which means that if the pth GPP in the rth node is activated, it is 1; F r is a Boolean variable, indicating that if the rth node is activated, it is 1; N r is the total number of nodes.

6. According to claim 1, a heterogeneous resource collaborative scheduling method for training AI reasoning services is characterized in that: The wavelength allocation in the adjustment network is specifically: Equations (11) and (12) ensure the coordination between DU-CU arrangement and optical path allocation, and establish the necessary optical link from access AAU to MEC through the selected GPP; where O k,f,s,r,p is a Boolean variable, which means that if the f-th AI inference service of the k-th user is deployed on the p-th GPP on the r-th node after being processed by the s-th baseband, it is 1; is a Boolean variable, indicating that if the fth AI reasoning service of the kth user is processed by the sth baseband function and occupies the wth wavelength of the link (m, n), then it is 1; N m,n is a Boolean variable, which means that if the mth node is connected to the nth node, it is 1; k,f,r is a Boolean variable, which means that if the source node of the fth AI service of the kth user is deployed on the rth node, it is 1; Formula (13) ensures the continuity of each service request in the service flow and avoids any interruption in the transmission process; where M r is a Boolean variable, which means that if MEC is deployed at the rth node, it is 1; Formula (14) can prevent the formation of a loop path in the network.

7. The method for coordinating heterogeneous resources for training AI reasoning services according to claim 1, characterized in that: The capacity requirements of the specified network are specifically: Formula (15) ensures that the sum of processing resources required by all services mapped to the GPP does not exceed its available capacity, where O k,f,s,r,p is a Boolean variable, which means that if the f-th AI inference service of the k-th user is deployed on the p-th GPP on the r-th node after being processed by the s-th baseband, it is 1; The computation and processing requirements of the sth segment of baseband processing function serving the fth AI inference service for the kth user; C p is the GPP computing power processing capacity FLOPS / s of the rth node; Formula (16) stipulates that the link between the required endpoints must have sufficient bandwidth to satisfy the service request, where is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user is processed by the s-th baseband function and occupies the w-th wavelength of the link (m, n), then it is 1; τ k,f,s The bandwidth requirement for the f-th AI reasoning service of the k-th user after being processed by the s-th baseband function; U l is the bandwidth capacity of the lth link; Formula (17) stipulates that the sum of the processing resources required for all services mapped to the rth node shall not exceed its available capacity, where V k,f,s,r is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user passes through the r-th node after being processed by the s-th baseband function, it is 1; The computation and processing requirements of the sth segment of baseband processing function serving the fth AI inference service for the kth user; U r is the computing power of the rth node; finally, formula (18) ensures that the edge data center m has sufficient capacity to meet all service requests; where m k,f is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user uses the m-th edge data center, it is 1; η k,f The computing power required to serve the fth AI reasoning service for the kth user; U m is the computing power of the mth edge data center; Formula (19) stipulates that each AI reasoning subtask should meet the total delay requirement; where, is a Boolean variable, indicating that if the fth AI reasoning service of the kth user is processed by the sth baseband function and occupies the wth wavelength of the link (m, n), then it is 1; L m,n is the link distance (m, n); φ is the optical fiber propagation delay per kilometer, 5μs / km; Y k,f,r is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user performs OEO photoelectric conversion at the r-th node, it is 1; T oeo is the delay between OEO and electronic switch; G r,p is a Boolean variable, which means that if the pth GPP in the rth node is activated, it is 1; T v Handle delays for virtual functions; t bv_oxc F is the transmission delay of the variable bandwidth optical cross-connector; r is a Boolean variable, indicating that if the rth node is activated, it is 1; m k,f is a Boolean variable, indicating that if the f-th AI reasoning service of the k-th user uses the m-th edge data center, it is 1; k,f,r is a Boolean variable, which means that if the source node of the fth AI service of the kth user is deployed on the rth node, it is 1; T MEC The latency of processing the fth type of AI inference service for the edge data center; T k,f The total latency requirement for computing power to serve the f-th AI inference class for the k-th user.

Citation Information

Patent Citations

  • AI inference task scheduling method and system oriented to multiple heterogeneous environments

    CN115756833A

  • Method for safely scheduling computing power resources in computing power network

    CN116723505A