AI reasoning task arrangement method and device for wireless access network, equipment and storage medium

By introducing artificial intelligence technology on the RAN side, we receive AI inference task guarantee strategies for AI inference models, and orchestrate AI inference tasks on computing nodes based on real-time communication status and computing capabilities, solving the problem of inefficient resource allocation and collaborative computing on the RAN side, and improving system performance.

CN120166409AActive Publication Date: 2025-06-17CHINA MOBILE COMM LTD RES INST +1

Patent Information

Application Number
CN202311733048.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-17
Estimated Expiration
2043-12-15

AI Technical Summary

Technical Problem

The prior art lacks unified perception scheduling capabilities on the RAN side, resulting in low allocation and collaborative computing efficiency of computing resources and communication resources and degradation of system performance.

Method used

Introduce artificial intelligence technology to receive AI inference task guarantee strategies for AI inference models on the RAN side, and perform near-real-time AI inference task collaborative arrangements on computing nodes based on real-time communication status and computing capabilities.

Benefits of technology

Improve the efficiency of RAN side resource allocation and collaborative computing, and improve system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166409A_ABST
    Figure CN120166409A_ABST
Patent Text Reader

Abstract

The invention discloses an AI reasoning task arrangement method, device and equipment of a wireless access network and a storage medium, an artificial intelligence technology is introduced to an RAN side, firstly, an AI reasoning task guarantee strategy of an AI reasoning model sent by a non-real-time wireless controller is received, and terminal information of at least one terminal is acquired, an optimization target of the wireless access network and feature information of an AI inference model are defined in the AI inference task guarantee strategy, and the terminal information comprises computing resource information and communication resource information; according to the method and the device, the AI reasoning task guarantee strategy is obtained, then the AI reasoning task arrangement scheme of the wireless access network is generated according to the AI reasoning task guarantee strategy and the terminal information, near-real-time AI reasoning task cooperative arrangement can be carried out on the computing nodes on the RAN side according to the real-time communication state and computing power, the RAN side resource allocation and cooperative computing efficiency is improved, and then the system performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and in particular, to an AI inference task orchestration method, apparatus, device, and storage medium for a radio access network. Background Art

[0002] With the continuous in-depth research on the open computing power in wireless local area networks, base stations, as an edge computing platform, have received increasing attention in the industry, and wireless local area networks are developing towards the integration of communication and computing. Although existing research provides solutions for balancing accuracy and latency in device edge computing collaboration, these studies mainly focus on edge servers and cannot be directly applied to the computing collaboration between base stations and terminals. The computing resources in base stations are shared by communication processing and computing tasks, and there is a competitive relationship between the two. Therefore, in the case of limited computing resources and communication bandwidth resources, the RAN (Radio Access Network) side needs to perform near-real-time adaptive orchestration of multi-base station and multi-terminal tasks according to the dynamically changing channel state and different information of different tasks.

[0003] The Near-RT RIC (Near-Real-Time RAN Intelligent Controller) in traditional O-RAN (Open-Radio Access Network) can collect near-real-time communication-related parameters through the standardized E2 interface to perform intelligent optimization of system performance. However, this is only the collection and optimization of communication-related parameters on the RAN side, lacking the unified perception and scheduling ability of computing resources and communication resources, resulting in low resource allocation and collaborative computing efficiency on the RAN side, thereby leading to a decline in system performance. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide an AI inference task orchestration method for a radio access network, which introduces artificial intelligence technology on the RAN side and can perform near-real-time AI inference task collaborative orchestration on the computing nodes on the RAN side according to the real-time communication state and computing power, improving the resource allocation and collaborative computing efficiency on the RAN side, and further enhancing the system performance.

[0005] To achieve the above purpose, the embodiments of the present invention provide an AI inference task orchestration method for a radio access network, which is applied to a near-real-time wireless controller, and the method includes:

[0006] Receive the AI inference task guarantee strategy of the AI inference model sent by the non-real-time wireless controller; wherein, the AI inference model is deployed in the wireless access network, and the optimization objectives of the wireless access network and the feature information of the AI inference model are defined in the AI inference task guarantee strategy;

[0007] Obtain the terminal information of at least one terminal; wherein, the terminal information includes computing resource information and communication resource information;

[0008] Generate an AI inference task orchestration plan for the wireless access network according to the AI inference task guarantee strategy and the terminal information.

[0009] As an improvement of the above solution, the generating an AI inference task orchestration plan for the wireless access network according to the AI inference task guarantee strategy and the terminal information includes:

[0010] According to the optimization objective of the wireless access network, determine the node to be allocated and the corresponding computing node for each terminal information, and select the corresponding splitting point and exit point from the feature information of the AI inference model for each terminal information; wherein, the splitting point and the exit point are one of the layers in the AI inference model;

[0011] Compose the AI inference task orchestration plan with the node to be allocated, its corresponding computing node, the splitting point and the exit point.

[0012] As an improvement of the above solution, after generating the AI inference task orchestration plan for the wireless access network according to the AI inference task guarantee strategy and the terminal information, the method further includes:

[0013] Send the AI inference task orchestration plan to the node to be allocated and the corresponding computing node, so that the node to be allocated and the corresponding computing node complete their respective inference calculation tasks.

[0014] As an improvement of the above solution, when the terminal is a user equipment, the obtaining the terminal information of at least one terminal includes:

[0015] Send a request message to the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment according to the request message; or,

[0016] Subscribe to the terminal information from the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment.

[0017] As an improvement of the above solution, when the terminal is a user equipment, the user equipment supports information interaction with the near real-time wireless controller; then, the obtaining the terminal information of at least one terminal includes:

[0018] Send a request message to the user equipment so that the user equipment reports terminal information according to the request message; or,

[0019] Subscribe to the terminal information from the user equipment so that the user equipment reports the terminal information.

[0020] To achieve the above object, an embodiment of the present invention further provides an AI inference task orchestration method for a radio access network, which is applied to a non-real-time radio controller. The method includes:

[0021] Obtain the model information of the AI inference model;

[0022] Generate an AI inference task guarantee policy according to the model information; wherein, the AI inference model is deployed in the radio access network, and the optimization objectives of the radio access network and the feature information of the AI inference model are defined in the AI inference task guarantee policy;

[0023] Send the AI inference task guarantee policy to a near-real-time radio controller so that the near-real-time radio controller generates an AI inference task orchestration plan for the radio access network according to the AI inference task guarantee policy and the terminal information; wherein, the terminal information includes computing resource information and communication resource information.

[0024] As an improvement of the above solution, the model information includes the feature information of the AI inference model and the performance guarantee parameters of the AI inference task; wherein, the performance guarantee parameters include but are not limited to the model inference accuracy and the number of inference calculations per unit time.

[0025] To achieve the above object, an embodiment of the present invention further provides an AI inference task orchestration device for a radio access network, including:

[0026] An AI inference task guarantee policy receiving module, configured to receive the AI inference task guarantee policy of the AI inference model sent by a non-real-time radio controller; wherein, the AI inference model is deployed in the radio access network, and the optimization objectives of the radio access network and the feature information of the AI inference model are defined in the AI inference task guarantee policy;

[0027] A terminal information obtaining module, configured to obtain the terminal information of at least one terminal; wherein, the terminal information includes computing resource information and communication resource information;

[0028] An AI inference task orchestration plan generating module, configured to generate an AI inference task orchestration plan for the radio access network according to the AI inference task guarantee policy and the terminal information.

[0029] To achieve the above object, an embodiment of the present invention further provides an AI inference task orchestration device for a radio access network, including:

[0030] A model information acquisition module, configured to acquire model information of an AI inference model;

[0031] An AI inference task guarantee policy generation module, configured to generate an AI inference task guarantee policy according to the model information; wherein, the AI inference model is deployed in the radio access network, and the optimization target of the radio access network and the feature information of the AI inference model are defined in the AI inference task guarantee policy;

[0032] An AI inference task guarantee policy sending module, configured to send the AI inference task guarantee policy to a near-real-time radio controller, so that the near-real-time radio controller generates an AI inference task orchestration scheme for the radio access network according to the AI inference task guarantee policy and terminal information; wherein, the terminal information includes computing resource information and communication resource information.

[0033] To achieve the above object, an embodiment of the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the AI inference task orchestration method for the radio access network as described in any of the above embodiments.

[0034] To achieve the above object, an embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the AI inference task orchestration method for the radio access network as described in any of the above embodiments.

[0035] Compared with the prior art, the AI inference task orchestration method, device, equipment, and storage medium for the radio access network disclosed by the present invention introduce artificial intelligence technology on the RAN side. First, it receives the AI inference task guarantee policy of the AI inference model sent by the non-real-time radio controller, and acquires the terminal information of at least one terminal. Among them, the optimization target of the radio access network and the feature information of the AI inference model are defined in the AI inference task guarantee policy, and the terminal information includes computing resource information and communication resource information; then, it generates an AI inference task orchestration scheme for the radio access network according to the AI inference task guarantee policy and the terminal information, and can perform near-real-time AI inference task collaborative orchestration on the computing nodes on the RAN side according to the real-time communication state and computing ability, improving the efficiency of resource allocation and collaborative computing on the RAN side, and thus enhancing the system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1It is a flowchart of an AI inference task orchestration method for a wireless access network provided by an embodiment of the present invention;

[0037] Figure 2 It is a schematic diagram of information interaction between a near-real-time wireless controller and a non-real-time wireless controller provided by an embodiment of the present invention;

[0038] Figure 3 It is a flowchart of a first method for obtaining terminal information of a user equipment provided by an embodiment of the present invention;

[0039] Figure 4 It is a flowchart of a second method for obtaining terminal information of a user equipment provided by an embodiment of the present invention;

[0040] Figure 5 It is a flowchart of another AI inference task orchestration method for a wireless access network provided by an embodiment of the present invention;

[0041] Figure 6 It is a structural block diagram of an AI inference task orchestration device for a wireless access network provided by an embodiment of the present invention;

[0042] Figure 7 It is a structural block diagram of another AI inference task orchestration device for a wireless access network provided by an embodiment of the present invention;

[0043] Figure 8 It is a structural block diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] With the rise of emerging applications such as XR (Extended Reality), autonomous driving, and industrial intelligent control, computational-intensive tasks related to deep learning, such as image recognition, have received more attention. However, the capabilities of mobile devices are limited and cannot meet their strict computational and latency requirements. Therefore, AI edge computing methods can be used to collaborate on computational tasks. AI edge computing methods can be implemented by constructing a deep neural network model. Since the convolutional layer in the deep neural network model structure performs dot product operations on the spatial dimensions of the input tensor to generate the feature map of the output tensor, the convolutional layer can be used as a model segmentation point, and multiple computing nodes can collaborate to complete the model inference. At the same time, resource waste can be reduced by adding a branch classifier to the model structure and exiting the inference early at the expense of accuracy. Based on the early exit mechanism and model segmentation technology, there are currently many studies on collaborative inference for edge computing nodes. By selecting appropriate segmentation points and exit points for AI inference tasks, the requirements for latency and inference accuracy can be met under the constraints of computational resources and system bandwidth. Therefore, in the embodiments of the present invention, AI technology is introduced on the RAN side to assist the terminal in completing computational tasks.

[0046] See Figure 1 , Figure 1 is a flowchart of an AI inference task orchestration method for a radio access network provided by an embodiment of the present invention. The AI inference task orchestration method for the radio access network is implemented by a near-real-time radio controller. The method includes:

[0047] S11. Receive the AI inference task guarantee policy of the AI inference model sent by the non-real-time radio controller;

[0048] S12. Obtain the terminal information of at least one terminal;

[0049] S13. Generate an AI inference task orchestration plan for the radio access network according to the AI inference task guarantee policy and the terminal information.

[0050] It should be noted that O-RAN includes a Near-Real-Time RAN Intelligent Controller (Near-RT RIC) and a Non-Real-Time RAN Intelligent Controller (Non-RT RIC). The terminal includes, but is not limited to, a base station, a centralized unit, a distributed unit DU (Distributed Unit), and a user equipment UE (User Equipment). The centralized unit includes CU (Centralized Unit), CU-CP (Centralized Unit-Control Plane), and CU-UP (Centralized Unit-User Plane). See Figure 2 , Figure 2 which is a schematic diagram of information interaction between the near-real-time radio controller and the non-real-time radio controller provided by an embodiment of the present invention. The non-real-time controller is connected to the near-real-time controller through an open and standardized A1 interface. The purpose of the non-real-time controller is to provide a corresponding machine learning model to support RAN intelligence, and for the non-real-time controller to provide an AI inference model and data for the near-real-time controller. In the O-RAN architecture, due to real-time requirements, when the near-real-time controller provides corresponding functions, it performs relevant operations by using an existing AI inference model. The AI inference model can perform one or more calculations among fault alarm analysis, coverage optimization, parameter optimization, spectrum analysis, inter-station coordination, mobility management, slice management, wireless positioning, and environmental perception recognition. The near-real-time controller and the terminal perform information interaction through the E2 interface. In the embodiment of the present invention, the functions of the E2 interface and the A1 interface in the O-RAN architecture are enhanced. The AI inference task guarantee policy generated by the non-real-time controller according to the model information of the AI inference model is sent to the near-real-time controller through the A1 interface, and then real-time terminal information is collected through the E2 interface. Near-real-time AI inference task collaborative orchestration is performed based on the terminal information and the AI inference task guarantee policy, improving the resource allocation and collaborative computing efficiency on the RAN side and enhancing the system performance.

[0051] Specifically, in step S11, the non-real-time radio controller generates the AI inference task guarantee policy according to the model information of the AI inference model; wherein, the model information includes the feature information of the AI inference model and the performance guarantee parameters of the AI inference task.

[0052] Exemplarily, the model information is obtained by the non-real-time radio controller from the SMO (Service Management and Orchestration Framework). For the source of the model information in the SMO, it can be directly configured by the operator or obtained through interaction with external applications. When obtaining the model information through interaction with external applications, relevant interfaces or APIs (Application Programming Interfaces) can be pre-designed, and connections with external applications can be established through authorization and authentication mechanisms to obtain the model information.

[0053] Exemplarily, the feature information of the AI inference model includes the AI inference model ID, number of model layers, types of each layer, output data volume of each layer, computational amount of each layer, split point location, exit point location, exit point accuracy, and computational amount of the exit point classifier, etc. Among them, the meaning corresponding to each piece of feature information can be referred to in Table 1, and specific examples can be referred to in Table 2.

[0054] Table 1 Feature Parameters and Their Corresponding Meanings

[0055]

[0056] Table 2 Example of Feature Information of the ResNet-18 Model

[0057]

[0058]

[0059] In Table 2, conv represents the convolutional layer, max pool represents the max pooling layer, avg pool represents the average pooling layer, fc represents the fully connected layer, Exit represents the exit point location. The ResNet-18 model has a total of four exit points, namely Exit1 located at the 5th layer, Exit2 located at the 9th layer, Exit3 located at the 13th layer, and Exit4 located at the 18th layer. The remaining layers can be used as split points. The exit point accuracy can evaluate the classification accuracy when selecting this layer as the exit point. The accuracies in the table are only examples.

[0060] Exemplarily, the performance guarantee parameters include but are not limited to model inference accuracy, number of inference calculations per unit time, model inference round-trip delay, and model spectral efficiency. Among them, the meaning corresponding to each performance guarantee parameter can be referred to in Table 3.

[0061] Table 3 Performance Guarantee Parameters and Their Corresponding Meanings

[0062]

[0063]

[0064] Further, after receiving the model information, the non-real-time wireless controller generates an AI inference task guarantee policy according to the model information. The AI inference task guarantee policy defines the optimization objectives of the wireless access network, the optimizable parameters, and the feature information of the AI inference model. The optimization objective of the wireless access network refers to the overall optimization objective of the guarantee policy in the current system. For example, maximizing the model inference accuracy in the system, minimizing the round-trip delay of model inference in the system, etc.; the optimizable parameters refer to the parameters that the near-real-time wireless controller can adjust for AI inference task guarantee, such as the selection of the splitting point and the exit point of the AI inference model. Assume that the optimization objective of the wireless access network is: maximizing the model inference accuracy in the system and minimizing the round-trip delay of model inference in the system. That is, after selecting the splitting point and the exit point, it is necessary to satisfy the maximum model inference accuracy and the minimum round-trip delay of model inference, which can be obtained by polling and calculating each splitting point and exit point.

[0065] Specifically, in step S12, the terminal information includes but is not limited to: computing resource information, communication resource information, current cell serving user information, and terminal preference information for AI inference tasks. Among them, the computing resource information includes but is not limited to: the utilization rate of the CPU (Central Processing Unit) of the base station and the terminal, CPU frequency, CPU core binding status, remaining CPU core number, FLOPS (Floating Point Operations Per Second) of CPU / GPU (Graphics Processing Unit) / NPU (Neural-network Processing Unit), GPU video memory capacity / remaining video memory, GPU cuda core number / remaining core number, etc.; the communication resource information includes but is not limited to: the total uplink / downlink bandwidth of the base station, remaining bandwidth, number of PRBs (Physical Resource Blocks), PRB utilization rate, etc., and may also include the channel condition of the terminal, the communication link delay from this base station to neighboring stations, etc. For example, the channel condition of the terminal includes SNR (Signal-to-Noise Ratio), RSSI (Received Signal Strength Indication), etc.; the current cell serving user information includes but is not limited to: user type identifier (non-AI user or AI user), AI inference task model identifier (such as: ResNet-18, VGG, etc.), etc.; the terminal preference information for AI inference tasks includes but is not limited to: preference for the selection of the splitting point for each AI inference task, etc.

[0066] It should be noted that when the terminal is a base station or the base station where the centralized unit and the distributed unit are located, the terminal information can be directly sent by the base station to the near-real-time radio controller. When the terminal is a user equipment, two methods for obtaining the terminal information of the user equipment are provided in the embodiments of the present invention. The first is to obtain it through the base station, and the second is to directly obtain it through the user equipment. The base station described in the embodiments of the present application may have various forms, such as a macro base station, a micro base station, a relay station, or an access point, etc. The base station can be an integrated base station, or can be a base station including a centralized unit CU and a distributed unit DU.

[0067] In the first implementation manner, when the terminal is a user equipment, the obtaining of the terminal information of at least one terminal includes: sending a request message to the base station where the user equipment camps, so that the base station reports the terminal information of the user equipment according to the request message; or subscribing to the terminal information from the base station where the user equipment camps, so that the base station reports the terminal information of the user equipment.

[0068] Exemplarily, referring to Figure 3 , the near-real-time radio controller subscribes to or requests the terminal information from the base station where the user equipment camps through the E2 interface, which can be periodic subscription or event-triggered reporting. Event triggers include, but are not limited to: base station computing power fluctuation, terminal computing power fluctuation, base station bandwidth fluctuation, terminal power fluctuation, terminal split mode preference change, etc. The base station collects the radio interface and terminal data of the user equipment according to the requirements, aggregates them into terminal information, and reports the terminal information of the user equipment to the near-real-time radio controller. In addition, the base station also synchronously reports its own terminal information to the near-real-time radio controller.

[0069] In the second implementation manner, when the terminal is a user equipment, the user equipment supports information interaction with the near-real-time radio controller; then, the obtaining of the terminal information of at least one terminal includes: sending a request message to the user equipment, so that the user equipment reports the terminal information according to the request message; or subscribing to the terminal information from the user equipment, so that the user equipment reports the terminal information.

[0070] Exemplarily, referring to Figure 4, the user equipment and the near-real-time radio controller enable both sides to support the corresponding interface application layer protocol based on the existing protocol stack, so as to realize the information interaction between the two sides. At this time, the near-real-time radio controller subscribes to or requests the terminal information from the user equipment through the E2 interface, which can be periodic subscription or event-triggered reporting. The user equipment collects the air interface and terminal data according to the requirements, aggregates them into terminal information, and reports the terminal information to the near-real-time radio controller. The near-real-time radio controller synchronously sends a request message to the base station or subscribes to the terminal information of the base station, so that the base station reports its own terminal information.

[0071] Specifically, in step S13, generating an AI inference task orchestration plan for the radio access network according to the AI inference task guarantee policy and the terminal information includes:

[0072] S131. According to the optimization objective of the radio access network, determine the nodes to be allocated and the corresponding computing nodes for each terminal information, and select the corresponding splitting point and exit point from the feature information of the AI inference model for each terminal information; wherein, the computing node is the base station, and the splitting point and the exit point are one of the layers in the AI inference model;

[0073] S132. Compose the AI inference task orchestration plan with the nodes to be allocated, their corresponding computing nodes, the splitting point and the exit point.

[0074] Exemplarily, the node to be allocated is a terminal that cannot meet the computing requirements by itself and needs to rely on the computing node to share the inference computing task; the computing node is a terminal that can receive the inference computing task of the node to be allocated. After receiving the terminal information sent by at least one terminal, the near-real-time radio controller determines the nodes to be allocated and the computing nodes, and will apply an optimization algorithm (such as the differential algorithm) to evaluate the computing nodes required by these nodes to be allocated, and select the splitting point and the exit point. For example, when it is determined according to the computing resource information that one of the nodes to be allocated requires a large amount of computing, more layers (that is, the number of layers between the splitting point and the exit point is relatively large) can be allocated to this node to be allocated, and vice versa; if it can be determined according to the communication resource information that this node to be allocated requires a computing node with a large remaining bandwidth, a computing node that meets the conditions will be preferentially allocated as the computing node of this node to be allocated; if the AI inference task model identifier is given in one of the terminal information at this time, when generating the AI inference task orchestration plan, the corresponding AI inference model will be selected according to the AI inference task model identifier for the inference task; if the splitting point selection preference for each AI inference task is given in one of the terminal information at this time, when generating the AI inference task orchestration plan, the splitting point will be preferentially allocated according to the splitting point selection preference of this node to be allocated.

[0075] Specifically, after generating an AI inference task orchestration plan for the radio access network according to the AI inference task guarantee policy and the terminal information, the method further includes:

[0076] S14. Sending the AI inference task orchestration plan to the node to be allocated and the corresponding computing node, so that the node to be allocated and the corresponding computing node complete their respective inference and calculation tasks.

[0077] Exemplarily, assume that the node to be allocated is a user equipment and the computing node is a base station. At this time, the terminal information of three user equipments and two base stations is collected. The two base stations and the three user equipments satisfy the following conditions: Base station A provides communication connection services for user equipment 1, base station B provides communication connection services for user equipment 2 and 3, both base stations A and B provide computing services, and the AI inference models are both ResNet-18 in Table 2. Assume that the orchestration plan at this time is to select a suitable computing node for user equipments 1 to 3, and select the splitting point and exit point for the AI inference task to be processed. The AI inference task orchestration plan is shown in Table 4 below.

[0078] Table 4 Example of AI inference task orchestration plan

[0079] User Equipment Computing Node Cutting Point Exit Point User Equipment 1 Base Station A Layer ID = 3 Exit2 User Equipment 2 Base Station A Layer ID = 8 Exit3 User Equipment 3 Base Station B Layer ID = 11 Exit3

[0080] Exemplarily, taking user equipment 1 as an example according to Table 4, the computing node selected for user equipment 1 at this time is base station A, that is, base station A and user equipment 1 jointly complete the operation task. The splitting point is the 3rd layer and the exit point is Exit2. Then the amount of calculation completed by base station A is the calculation amount from the 3rd layer to the 9th layer. According to the selected splitting point, user equipment 1 completes the inference and calculation of the previous layers. After the feature map generated at this splitting point layer is transmitted to base station A, base station A completes the inference and calculation from the 3rd layer to the 9th layer, and finally transmits the generated result back to user equipment 1. Therefore, the selection of the splitting point and the exit point determines the respective calculation amounts of the node to be allocated and the computing node, and the selection of the exit point also determines the accuracy of the model.

[0081] Compared with the prior art, the AI inference task orchestration method for a wireless access network disclosed by the present invention introduces artificial intelligence technology on the RAN side. First, it receives the AI inference task guarantee policy of the AI inference model sent by a non-real-time radio controller, and obtains the terminal information of at least one terminal. Among them, the optimization objectives of the wireless access network and the feature information of the AI inference model are defined in the AI inference task guarantee policy, and the terminal information includes computing resource information and communication resource information. Then, according to the AI inference task guarantee policy and the terminal information, an AI inference task orchestration plan for the wireless access network is generated. By selecting computing nodes, task splitting points, and exit points, near-real-time AI inference task collaboration orchestration can be performed on the computing nodes on the RAN side according to the real-time communication status and computing capabilities, improving the efficiency of resource allocation and collaborative computing on the RAN side, and thus enhancing the system performance.

[0082] See Figure 5 , Figure 5 is a flowchart of another AI inference task orchestration method for a wireless access network provided by an embodiment of the present invention. The AI inference task orchestration method for the wireless access network is implemented by a non-real-time radio controller, and the method includes:

[0083] S21. Obtain the model information of the AI inference model;

[0084] S22. Generate an AI inference task guarantee policy according to the model information;

[0085] S23. Send the AI inference task guarantee policy to a near-real-time radio controller, so that the near-real-time radio controller generates an AI inference task orchestration plan for the wireless access network according to the AI inference task guarantee policy and the terminal information.

[0086] Specifically, the model information includes the feature information of the AI inference model and the performance guarantee parameters of the AI inference task; among them, the performance guarantee parameters include, but are not limited to, the model inference accuracy and the number of inference calculations per unit time.

[0087] It should be noted that the specific working process of the AI inference task orchestration method for the wireless access network described in the embodiments of the present invention can refer to the above embodiments and will not be elaborated here.

[0088] See Figure 6 , Figure 6 is a structural block diagram of an AI inference task orchestration device 100 for a wireless access network provided by an embodiment of the present invention. The AI inference task orchestration device 100 for the wireless access network includes:

[0089] The AI inference task guarantee policy receiving module 11 is used to receive the AI inference task guarantee policy of the AI inference model sent by the non-real-time wireless controller; wherein, the AI inference model is deployed in the wireless access network, and the optimization objective of the wireless access network and the feature information of the AI inference model are defined in the AI inference task guarantee policy;

[0090] The terminal information acquisition module 12 is used to acquire the terminal information of at least one terminal; wherein, the terminal information includes computing resource information and communication resource information;

[0091] The AI inference task scheduling scheme generation module 13 is used to generate an AI inference task scheduling scheme for the wireless access network according to the AI inference task guarantee policy and the terminal information.

[0092] Specifically, the AI inference task scheduling scheme generation module 13 is specifically used for:

[0093] According to the optimization objective of the wireless access network, determine the node to be allocated and the corresponding computing node according to each terminal information, and select the corresponding splitting point and exit point from the feature information of the AI inference model according to each terminal information; wherein, the computing node is a base station, and the splitting point and the exit point are one of the layers in the AI inference model;

[0094] The AI inference task scheduling scheme is composed of the node to be allocated and its corresponding computing node, the splitting point and the exit point.

[0095] Specifically, the AI inference task scheduling device 100 for the wireless access network further includes:

[0096] The AI inference task scheduling scheme sending module is used to send the AI inference task scheduling scheme to the node to be allocated and the corresponding computing node, so that the node to be allocated and the corresponding computing node complete their respective inference calculation tasks.

[0097] Specifically, when the terminal is a user equipment, the terminal information acquisition module 12 is specifically used for: sending a request message to the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment according to the request message; or, subscribing to the terminal information from the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment.

[0098] Specifically, when the terminal is a user device, the user device supports information interaction with a near-real-time radio controller; then, the terminal information acquisition module 12 is specifically configured to: send a request message to the user device to cause the user device to report terminal information according to the request message; or, subscribe to the terminal information from the user device to cause the user device to report terminal information.

[0099] See Figure 7 , Figure 7 FIG. is a structural block diagram of another AI inference task orchestration device 200 for a radio access network provided by an embodiment of the present invention. The AI inference task orchestration device 200 for the radio access network includes:

[0100] A model information acquisition module 21, configured to acquire model information of an AI inference model;

[0101] An AI inference task guarantee policy generation module 22, configured to generate an AI inference task guarantee policy according to the model information; wherein, the AI inference model is deployed in a radio access network, and the AI inference task guarantee policy defines an optimization target of the radio access network and feature information of the AI inference model;

[0102] An AI inference task guarantee policy sending module 23, configured to send the AI inference task guarantee policy to a near-real-time radio controller, so that the near-real-time radio controller generates an AI inference task orchestration plan for the radio access network according to the AI inference task guarantee policy and terminal information; wherein, the terminal information includes computing resource information and communication resource information.

[0103] Specifically, the model information includes feature information of the AI inference model and performance guarantee parameters of the AI inference task; wherein, the performance guarantee parameters include, but are not limited to, model inference accuracy and the number of inference calculations per unit time.

[0104] Compared with the prior art, the AI inference task orchestration device for a radio access network disclosed in the present invention introduces artificial intelligence technology on the RAN side. First, it receives an AI inference task guarantee policy of an AI inference model sent by a non-real-time radio controller, and acquires terminal information of at least one terminal. Among them, the AI inference task guarantee policy defines an optimization target of the radio access network and feature information of the AI inference model, and the terminal information includes computing resource information and communication resource information; then, according to the AI inference task guarantee policy and the terminal information, it generates an AI inference task orchestration plan for the radio access network. By selecting computing nodes, task splitting points, and exit points, it can perform near-real-time AI inference task collaborative orchestration on the computing nodes on the RAN side according to the real-time communication state and computing capabilities, improving the efficiency of resource allocation and collaborative computing on the RAN side, and thus enhancing the system performance.

[0105] See Figure 8 , Figure 8 Figure 8 is a block diagram of an electronic device 300 provided by an embodiment of the present invention. The electronic device 300 includes a processor 31, a memory 32, and a computer program stored in the memory 32 and executable on the processor 31. When the processor 31 executes the computer program, the steps in the embodiments of the above-mentioned AI inference task orchestration method for each wireless access network are implemented, such as steps S11 to S13, S21 to S23.

[0106] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory 32 and executed by the processor 31 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 300.

[0107] The electronic device 300 may include, but is not limited to, a processor 31 and a memory 32. Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 300 and does not constitute a limitation on the electronic device 300. It may include more or fewer components than shown, or combine certain components, or different components. For example, the electronic device 300 may further include input / output devices, network access devices, buses, etc.

[0108] The processor 31 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor 31 is the control center of the electronic device 300, and connects various parts of the entire electronic device 300 through various interfaces and lines.

[0109] The memory 32 can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory 32 and invoking the data stored in the memory 32, the processor 31 realizes various functions of the electronic device 300. The memory 32 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 32 can include high-speed random access memory and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0110] Among them, if the modules / units integrated in the electronic device 300 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 31, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0111] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. An AI inference task orchestration method for a wireless access network, characterized in that, Applied to a near-real-time wireless controller, the method includes: Receiving an AI inference task guarantee strategy of an AI inference model sent by a non-real-time wireless controller; wherein, the AI inference model is deployed in a radio access network, and the optimization objectives of the radio access network and the feature information of the AI inference model are defined in the AI inference task guarantee strategy; Obtaining terminal information of at least one terminal; wherein, the terminal information includes computing resource information and communication resource information; Generating an AI inference task scheduling plan for the radio access network according to the AI inference task guarantee strategy and the terminal information.

2. The AI inference task orchestration method for a wireless access network according to claim 1, characterized in that, The generating an AI inference task scheduling plan for the radio access network according to the AI inference task guarantee strategy and the terminal information includes: According to the optimization objectives of the radio access network, determining the nodes to be allocated and the corresponding computing nodes according to each terminal information, and selecting the corresponding splitting point and exit point from the feature information of the AI inference model according to each terminal information; wherein, the splitting point and the exit point are one of the layers in the AI inference model; Composing the AI inference task scheduling plan with the nodes to be allocated and their corresponding computing nodes, the splitting point and the exit point.

3. The AI inference task orchestration method for a wireless access network according to claim 2, characterized in that, After generating an AI inference task scheduling plan for the radio access network according to the AI inference task guarantee strategy and the terminal information, the method further includes: Sending the AI inference task scheduling plan to the nodes to be allocated and the corresponding computing nodes, so that the nodes to be allocated and the corresponding computing nodes complete their respective inference calculation tasks.

4. The AI inference task orchestration method for a wireless access network according to claim 1, characterized in that, When the terminal is a user equipment, the obtaining terminal information of at least one terminal includes: Sending a request message to the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment according to the request message; or, Subscribing to the terminal information from the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment.

5. The AI inference task orchestration method for a wireless access network according to claim 1, characterized in that, When the terminal is a user equipment, the user equipment supports information interaction with the near-real-time wireless controller; then, the obtaining terminal information of at least one terminal includes: Sending a request message to the user equipment, so that the user equipment reports terminal information according to the request message; or, Subscribing to the terminal information from the user equipment, so that the user equipment reports terminal information.

6. An AI inference task orchestration method for a wireless access network, characterized in that, Applied to a non-real-time wireless controller, the method includes: Obtaining model information of an AI inference model; Generating an AI inference task guarantee strategy according to the model information; wherein, the AI inference model is deployed in a radio access network, and the optimization objectives of the radio access network and the feature information of the AI inference model are defined in the AI inference task guarantee strategy; Sending the AI inference task guarantee strategy to the near-real-time wireless controller, so that the near-real-time wireless controller generates an AI inference task scheduling plan for the radio access network according to the AI inference task guarantee strategy and terminal information; wherein, the terminal information includes computing resource information and communication resource information.

7. The AI inference task orchestration method for a wireless access network according to claim 6, characterized in that, The model information includes the feature information of the AI inference model and the performance guarantee parameters of the AI inference task; among them, the performance guarantee parameters include, but are not limited to, the model inference accuracy and the number of inference calculations per unit time.

8. An AI inference task orchestration device for a wireless access network, characterized in that, Comprising: An AI inference task guarantee policy receiving module, configured to receive the AI inference task guarantee policy of the AI inference model sent by the non-real-time wireless controller; wherein, the AI inference model is deployed in the wireless access network, and the optimization objective of the wireless access network and the feature information of the AI inference model are defined in the AI inference task guarantee policy; A terminal information acquisition module, configured to acquire the terminal information of at least one terminal; wherein, the terminal information includes computing resource information and communication resource information; An AI inference task scheduling scheme generation module, configured to generate an AI inference task scheduling scheme for the wireless access network according to the AI inference task guarantee policy and the terminal information.

9. An AI inference task orchestration device for a wireless access network, characterized in that, Comprising: A model information acquisition module, configured to acquire the model information of the AI inference model; An AI inference task guarantee policy generation module, configured to generate an AI inference task guarantee policy according to the model information; wherein, the AI inference model is deployed in the wireless access network, and the optimization objective of the wireless access network and the feature information of the AI inference model are defined in the AI inference task guarantee policy; An AI inference task guarantee policy sending module, configured to send the AI inference task guarantee policy to the near-real-time wireless controller, so that the near-real-time wireless controller generates an AI inference task scheduling scheme for the wireless access network according to the AI inference task guarantee policy and the terminal information; wherein, the terminal information includes computing resource information and communication resource information.

10. An electronic device, characterized in that Comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the AI inference task scheduling method for the wireless access network according to any one of claims 1 to 7.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the AI inference task scheduling method for the wireless access network according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Inference service deployment method and device, equipment and storage medium

    CN116301912A

  • Multi-level computing power network task scheduling method and device

    CN116708443A

  • Data-centric service-based network architecture

    US20210184989A1

  • Online reinforcement learning

    WO2022060777A1

  • Zero-touch deployment and orchestration of network intelligence in open ran systems

    WO2023172292A2

Cited By

  • Model reasoning method and device, network equipment, readable storage medium and program product

    CN120529342A