AI inference task scheduling method and device for radio access network, equipment and storage medium
By introducing AI technology on the RAN side, generating AI inference task orchestration schemes, and selecting appropriate splitting and exit points, the problem of unified perception and scheduling of computing and communication resources between base stations and terminals is solved, improving resource allocation and collaborative computing efficiency, and enhancing system performance.
Patent Information
- Application Number
- CN202311733048.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2043-12-15
AI Technical Summary
Existing technologies lack unified awareness and scheduling capabilities for computing and communication resources in computational collaboration between base stations and terminals, resulting in low efficiency in RAN-side resource allocation and collaborative computing, which fails to meet the computational and latency requirements of computationally intensive tasks.
Artificial intelligence technology is introduced on the RAN side. By receiving the task guarantee strategy and terminal information of the AI inference model, an AI inference task orchestration scheme is generated, and appropriate split points and exit points are selected to collaboratively complete the computing task.
It improves the efficiency of RAN-side resource allocation and collaborative computing, enhances system performance, and meets the computation and latency requirements of computationally intensive tasks.
Smart Images

Figure CN120166409B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of wireless communication, and in particular to an AI inference task scheduling method and device for a radio access network, an apparatus, and a storage medium. BACKGROUND
[0002] With the deepening of the research on the opening of computing capabilities in wireless local area networks, base stations as an edge computing platform are increasingly valued by the industry, and wireless local area networks are developing towards the integration of communication and computing. Although existing research provides solutions to balance accuracy and delay in device edge computing collaboration, these researches mainly focus on edge servers and cannot be directly applied to computing collaboration between base stations and terminals. The computing resources in the base station are shared by communication processing and computing tasks, and there is a competitive relationship between them. Therefore, under the condition of limited computing resources and communication bandwidth resources, the RAN(Radio Access Network, wireless access network) side needs to adaptively schedule multiple base station and terminal tasks in near real time according to the dynamically changing channel state and different information of different tasks.
[0003] The Near-RT RIC(Near-Real-Time RAN Intelligent Controller, near real-time controller) in the traditional O-RAN(Open-Radio Access Network, open radio access network) can collect near real-time communication-related parameters through the standardized E2 interface to intelligently optimize system performance. However, this is only the collection and optimization of communication-related parameters on the RAN side, and lacks unified perception and scheduling capabilities for computing resources and communication resources, resulting in low efficiency of resource allocation and collaborative computing on the RAN side, and thus leading to a decline in system performance. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide an AI inference task scheduling method for a radio access network, which introduces Artificial Intelligence technology on the RAN side, can perform near real-time AI inference task collaborative scheduling on the computing nodes on the RAN side according to real-time communication states and computing capabilities, improves the efficiency of resource allocation and collaborative computing on the RAN side, and thus improves the system performance.
[0005] To achieve the above purpose, the embodiments of the present application provide an AI inference task scheduling method for a radio access network, applied to a near real-time radio controller, the method comprising:
[0006] receive an AI inference task assurance policy of an AI inference model sent by a non-real-time wireless controller; wherein the AI inference model is deployed in a wireless access network, and the AI inference task assurance policy defines an optimization target of the wireless access network and feature information of the AI inference model;
[0007] obtain terminal information of at least one terminal; wherein the terminal information includes computing resource information and communication resource information;
[0008] generate an AI inference task scheduling scheme of the wireless access network according to the AI inference task assurance policy and the terminal information.
[0009] As an improvement of the above scheme, the AI inference task scheduling scheme of the wireless access network is generated according to the AI inference task assurance policy and the terminal information, including:
[0010] According to the optimization target of the wireless access network, determine the to-be-assigned node and the corresponding computing node according to each terminal information, and select the corresponding split point and exit point from the feature information of the AI inference model according to each terminal information; wherein the split point and the exit point are one layer in the AI inference model;
[0011] The to-be-assigned node and the corresponding computing node, the split point and the exit point constitute the AI inference task scheduling scheme.
[0012] As an improvement of the above scheme, after generating the AI inference task scheduling scheme of the wireless access network according to the AI inference task assurance policy and the terminal information, the method further includes:
[0013] Send the AI inference task scheduling scheme to the to-be-assigned node and the corresponding computing node, so that the to-be-assigned node and the corresponding computing node complete their respective inference computing tasks.
[0014] As an improvement of the above scheme, when the terminal is a user equipment, the terminal information of at least one terminal includes:
[0015] Send a request message to the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment according to the request message; or,
[0016] Subscribe to the terminal information of the user equipment to the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment.
[0017] As an improvement of the above scheme, when the terminal is a user equipment, the user equipment supports information interaction with a near-real-time wireless controller; then, the terminal information of at least one terminal includes:
[0018] sending a request message to the user equipment, so that the user equipment reports terminal information according to the request message; or
[0019] subscribing to the terminal information of the user equipment, so that the user equipment reports terminal information.
[0020] To achieve the above-mentioned purpose, the embodiment of the present application also provides an AI inference task scheduling method of a radio access network, applied to a non-real-time radio controller, and the method comprises:
[0021] obtaining model information of an AI inference model;
[0022] generating an AI inference task guarantee strategy according to the model information; wherein the AI inference model is deployed in a radio access network, and the AI inference task guarantee strategy defines an optimization target of the radio access network and feature information of the AI inference model;
[0023] sending the AI inference task guarantee strategy to a near-real-time radio controller, so that the near-real-time radio controller generates an AI inference task scheduling scheme of the radio access network according to the AI inference task guarantee strategy and terminal information; wherein the terminal information comprises computing resource information and communication resource information.
[0024] As an improvement of the above-mentioned scheme, the model information comprises feature information of the AI inference model and performance guarantee parameters of the AI inference task; wherein the performance guarantee parameters comprise but are not limited to model inference accuracy and inference calculation times per unit time.
[0025] To achieve the above-mentioned purpose, the embodiment of the present application also provides an AI inference task scheduling device of a radio access network, comprising:
[0026] an AI inference task guarantee strategy receiving module, configured to receive an AI inference task guarantee strategy of an AI inference model sent by a non-real-time radio controller; wherein the AI inference model is deployed in a radio access network, and the AI inference task guarantee strategy defines an optimization target of the radio access network and feature information of the AI inference model;
[0027] a terminal information obtaining module, configured to obtain terminal information of at least one terminal; wherein the terminal information comprises computing resource information and communication resource information;
[0028] an AI inference task scheduling scheme generating module, configured to generate an AI inference task scheduling scheme of the radio access network according to the AI inference task guarantee strategy and the terminal information.
[0029] To achieve the above object, the embodiment of the present application further provides an AI inference task arrangement device of a wireless access network, comprising:
[0030] a model information acquisition module, configured to acquire model information of an AI inference model;
[0031] an AI inference task guarantee strategy generation module, configured to generate an AI inference task guarantee strategy according to the model information; wherein the AI inference model is deployed in the wireless access network, and the AI inference task guarantee strategy defines an optimization target of the wireless access network and feature information of the AI inference model;
[0032] an AI inference task guarantee strategy sending module, configured to send the AI inference task guarantee strategy to a near-real-time wireless controller, so that the near-real-time wireless controller generates an AI inference task arrangement scheme of the wireless access network according to the AI inference task guarantee strategy and terminal information; wherein the terminal information comprises computing resource information and communication resource information.
[0033] To achieve the above object, the embodiment of the present application further provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the AI inference task arrangement method of the wireless access network according to any one of the above embodiments when executing the computer program.
[0034] To achieve the above object, the embodiment of the present application further provides a computer readable storage medium, comprising a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the AI inference task arrangement method of the wireless access network according to any one of the above embodiments when the computer program runs.
[0035] Compared with the prior art, the AI inference task arrangement method, device, equipment and storage medium of the wireless access network disclosed by the present application introduce artificial intelligence technology on the RAN side, first receive the AI inference task guarantee strategy of the AI inference model sent by the non-real-time wireless controller, and acquire terminal information of at least one terminal, wherein the AI inference task guarantee strategy defines the optimization target of the wireless access network and the feature information of the AI inference model, and the terminal information comprises computing resource information and communication resource information; then generate an AI inference task arrangement scheme of the wireless access network according to the AI inference task guarantee strategy and the terminal information, which can make near-real-time AI inference task cooperation arrangement for the computing nodes on the RAN side according to the real-time communication state and computing ability, improve the efficiency of RAN side resource allocation and cooperative computing, and further improve the system performance. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1is a flow chart of an AI inference task scheduling method of a wireless access network provided by an embodiment of the present application;
[0037] Figure 2 is a schematic diagram of information interaction of a near real-time wireless controller and a non-real-time wireless controller provided by an embodiment of the present application;
[0038] Figure 3 is a flow chart of a terminal information acquisition method of a first user equipment provided by an embodiment of the present application;
[0039] Figure 4 is a flow chart of a terminal information acquisition method of a second user equipment provided by an embodiment of the present application;
[0040] Figure 5 is a flow chart of another AI inference task scheduling method of a wireless access network provided by an embodiment of the present application;
[0041] Figure 6 is a structural block diagram of an AI inference task scheduling device of a wireless access network provided by an embodiment of the present application;
[0042] Figure 7 is a structural block diagram of another AI inference task scheduling device of a wireless access network provided by an embodiment of the present application;
[0043] Figure 8 is a structural block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0044] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.
[0045] With the rise of emerging applications such as XR (Extended Reality), autonomous driving, and industrial intelligent control, deep learning related algorithms such as image recognition and other computationally intensive tasks are receiving more attention. However, mobile devices have limited capabilities and cannot meet their strict computing requirements and delay requirements, so AI edge computing can be used to collaborate to complete computing tasks. AI edge computing can be implemented by constructing a deep neural network model. Since the convolution layer in the structure of the deep neural network model is a kernel that performs point multiplication operations on the spatial dimensions of the input tensor to generate the feature map of the output tensor, the convolution layer can be used as a splitting point for the model, and multiple computing nodes can be used to collaboratively complete the inference of the model. At the same time, branch classifiers can be added to the model structure to exit the inference early to reduce resource waste under the condition of sacrificing accuracy. Based on the early exit mechanism and model splitting technology, there are currently many studies on collaborative inference of edge computing nodes. By selecting appropriate splitting points and exit points for AI inference tasks, the requirements for latency and inference accuracy can be met under the constraints of computing resources and system bandwidth. Therefore, in the present embodiment, AI technology is introduced on the RAN side to assist terminals in completing computing tasks.
[0046] Referring to Figure 1 , Figure 1 is a flowchart of an AI inference task scheduling method of a wireless access network provided by an embodiment of the present application. The AI inference task scheduling method of the wireless access network is implemented by a near real-time radio controller. The method comprises:
[0047] S11, receiving an AI inference task guarantee strategy of an AI inference model sent by a non-real-time radio controller;
[0048] S12, obtaining terminal information of at least one terminal;
[0049] S13, generating an AI inference task scheduling scheme of a wireless access network according to the AI inference task guarantee strategy and the terminal information.
[0050] It is worth mentioning that the O-RAN includes a near-real-time controller (Near-RT RIC) and a non-real-time controller (Non-RT RIC), the terminal includes but is not limited to a base station, a centralized unit, a distributed unit DU (Distributed Unit) and a user equipment UE (User Equipment), and the centralized unit includes a CU (Centralized Unit), a CU-CP (Centralized Unit-Control Plane) and a CU-UP (Centralized Unit-User Plane). See Figure 2 , Figure 2 is a schematic diagram of information interaction of a near-real-time wireless controller and a non-real-time wireless controller provided by the embodiment of the application. The non-real-time controller is connected to the near-real-time controller through an open and standardized A1 interface. The non-real-time controller aims to provide a corresponding machine learning model to support RAN intelligence, and provide an AI inference model and data for the near-real-time controller. Due to the requirement of real-time in the O-RAN architecture, the near-real-time controller performs related operations by utilizing an existing AI inference model when providing corresponding functions. The AI inference model can perform one or more of fault alarm analysis, coverage optimization, parameter optimization, spectrum analysis, inter-station coordination, mobility management, slice management, wireless positioning and environment sensing identification. The near-real-time controller and the terminal perform information interaction through an E2 interface. In the embodiment of the application, the functions of the E2 interface and the A1 interface in the O-RAN architecture are enhanced. The AI inference task guarantee strategy generated by the non-real-time controller according to the model information of the AI inference model is issued to the near-real-time controller through the A1 interface. Real-time terminal information is collected through the E2 interface. According to the terminal information and the AI inference task guarantee strategy, near-real-time AI inference task collaboration scheduling is performed, the RAN side resource allocation and collaboration calculation efficiency are improved, and the system performance is improved.
[0051] Specifically, in step S11, the non-real-time wireless controller generates the AI inference task guarantee strategy according to model information of an AI inference model; wherein the model information includes feature information of the AI inference model and performance guarantee parameters of an AI inference task.
[0052] Exemplarily, the model information is obtained by the non-real-time wireless controller from a SMO (Service Management and Orchestration), and the source of the model information in the SMO can be directly configured by an operator or obtained by interacting with an external application. When the model information is obtained by interacting with the external application, a related interface or API (Application Programming Interface) can be designed in advance, and a connection is established with the external application through an authorization and authentication mechanism to obtain the model information.
[0053] Exemplarily, the feature information of the AI inference model includes an AI inference model ID, a model layer number, a layer type, a layer output data volume, a layer calculation volume, a split point position, an exit point position, an exit point accuracy, and an exit point classifier calculation volume. The meaning of each feature information can refer to Table 1, and a specific example can refer to Table 2.
[0054] Table 1: Feature parameters and their corresponding meanings
[0055]
[0056] Table 2: Feature information example of ResNet-18 model
[0057]
[0058]
[0059] In Table 2, conv represents a convolution layer, max pool represents a maximum pooling layer, avg pool represents an average pooling layer, fc represents a fully connected layer, Exit represents an exit point position, the ResNet-18 model has four exit points, Exit1 at the 5th layer, Exit2 at the 9th layer, Exit3 at the 13th layer, and Exit4 at the 18th layer, and the remaining layers can be used as split points. The exit point accuracy can evaluate the classification accuracy when selecting this layer as an exit point, and the accuracy in the table is only an example.
[0060] Exemplarily, the performance guarantee parameters include but are not limited to model inference accuracy, inference calculation number per unit time, model inference round trip delay, and model spectral efficiency. The meaning of each performance guarantee parameter can refer to Table 3.
[0061] Table 3: Performance guarantee parameters and their corresponding meanings
[0062]
[0063]
[0064] Further, the non-real-time wireless controller generates an AI inference task guarantee strategy according to the model information after receiving the model information, and the AI inference task guarantee strategy defines an optimization target of a wireless access network, an optimizable parameter, and feature information of an AI inference model. The optimization target of the wireless access network refers to the overall optimization target of the guarantee strategy in the current system, such as maximizing the model inference accuracy in the system or minimizing the model inference round trip delay in the system. The optimizable parameter refers to the parameter that can be adjusted by the near-real-time wireless controller for the AI inference task guarantee, such as the selection of the split point and the exit point of the AI inference model. Assuming that the optimization target of the wireless access network is to maximize the model inference accuracy in the system and to minimize the model inference round trip delay in the system, that is, after the split point and the exit point are selected, the model inference accuracy needs to be maximized and the model inference round trip delay needs to be minimized, which can be obtained by polling calculation for each split point and exit point.
[0065] Specifically, in step S12, the terminal information includes but is not limited to: computing resource information, communication resource information, current cell service user information, and terminal preference information for AI inference tasks. The computing resource information includes but is not limited to: CPU (Central Processing Unit) utilization rate, CPU frequency, CPU core binding state, remaining CPU core number, CPU / GPU (Graphics Processing Unit) / NPU (Neural-network Processing Unit) floating point operation per second (FLOPS), GPU video memory capacity / remaining video memory, GPU cuda core number / remaining core number, etc. of the base station and the terminal; the communication resource information includes but is not limited to: uplink / downlink total bandwidth, remaining bandwidth, PRB (Physical Resource Block) number, PRB utilization rate, etc. of the base station, and can also include channel conditions of the terminal, communication link delay from the base station to the neighboring station, etc. For example, the channel conditions of the terminal include SNR (Signal-to-Noise Ratio), RSSI (Received Signal Strength Indication), etc.; the current cell service user information includes but is not limited to: user type identification (non-AI user or AI user), AI inference task model identification (such as ResNet-18, VGG, etc.), etc.; and the terminal preference information for AI inference tasks includes but is not limited to: split point selection preference for each AI inference task, etc.
[0066] It is worth mentioning that when the terminal is a base station or a base station where the centralized unit and the distributed unit are located, the terminal information can be directly sent by the base station to the near real-time wireless controller. When the terminal is a user equipment, two ways of obtaining terminal information of the user equipment are provided in the embodiments of the application, the first way is to obtain through the base station, and the second way is to obtain directly through the user equipment. The base station described in the embodiments of the application can have various forms, such as a macro base station, a micro base station, a relay station or an access point, etc. The base station can be an integrated base station, or can be a base station including a centralized unit CU and a distributed unit DU.
[0067] In the first implementation, when the terminal is a user equipment, the obtaining of the terminal information of the at least one terminal comprises: sending a request message to a base station where the user equipment resides, so that the base station reports terminal information of the user equipment according to the request message; or, subscribing to the terminal information of the user equipment from the base station where the user equipment resides, so that the base station reports the terminal information of the user equipment.
[0068] For example, referring to Figure 3 , the near real-time wireless controller subscribes or requests the terminal information from the base station where the user equipment resides through an E2 interface, which can be periodic subscription or event-triggered reporting. The event-triggered reporting includes but is not limited to base station computing power fluctuation, terminal computing power fluctuation, base station bandwidth fluctuation, terminal power fluctuation, terminal splitting mode preference change, etc. The base station collects air interface and terminal data of the user equipment according to the demand, aggregates the terminal information, and reports the terminal information of the user equipment to the near real-time wireless controller. In addition, the base station also synchronously reports its own terminal information to the near real-time wireless controller.
[0069] In the second implementation, when the terminal is a user equipment, the user equipment supports information interaction with the near real-time wireless controller; then, the obtaining of the terminal information of the at least one terminal comprises: sending a request message to the user equipment, so that the user equipment reports terminal information according to the request message; or, subscribing to the terminal information of the user equipment, so that the user equipment reports the terminal information.
[0070] For example, referring to Figure 4The user equipment and the near real-time wireless controller support corresponding interface application layer protocols on the basis of an existing protocol stack, so as to realize information interaction between the two. At this time, the near real-time wireless controller subscribes or requests the terminal information from the user equipment through the E2 interface, which can be periodic subscription or event-triggered reporting. The user equipment collects air interface and terminal data according to the demand, aggregates the terminal information, and reports the terminal information to the near real-time wireless controller. The near real-time wireless controller synchronously sends a request message to the base station or subscribes the terminal information of the base station, so that the base station reports its own terminal information.
[0071] Specifically, in step S13, the AI inference task scheduling scheme of the wireless access network is generated according to the AI inference task guarantee policy and the terminal information, including:
[0072] S131, according to the optimization target of the wireless access network, determining the to-be-assigned node and the corresponding computing node according to each terminal information, and selecting the corresponding split point and exit point from the feature information of the AI inference model according to each terminal information; wherein the computing node is a base station, and the split point and the exit point are one layer in the AI inference model.
[0073] S132, the to-be-assigned node and the corresponding computing node, the split point and the exit point are used to form the AI inference task scheduling scheme.
[0074] For example, the to-be-assigned node is a terminal that cannot meet the computing demand by itself and needs to share the inference computing task with a computing node; the computing node is a terminal that can receive the inference computing task of the to-be-assigned node. After receiving the terminal information sent by at least one terminal, the near real-time wireless controller determines the to-be-assigned node and the computing node, applies an optimization algorithm (such as a differential algorithm), evaluates the computing nodes required by the to-be-assigned nodes, and selects the split point and the exit point. For example, according to the computing resource information, it can be determined that one of the to-be-assigned nodes needs a large amount of computation, and a large number of layers (i.e. the number of layers between the split point and the exit point is larger) can be allocated to this to-be-assigned node, otherwise a smaller number of layers can be allocated; according to the communication resource information, it can be determined that this to-be-assigned node needs a computing node with a larger remaining bandwidth, and the computing node that meets the condition is preferentially allocated as the computing node of this to-be-assigned node; if the AI inference task model identifier is given in one of the terminal information, the corresponding AI inference model will be selected for inference task according to the AI inference task model identifier when generating the AI inference task scheduling scheme; if the split point selection preference for each AI inference task is given in one of the terminal information, the split point will be preferentially allocated according to the split point selection preference of this to-be-assigned node when generating the AI inference task scheduling scheme.
[0075] Specifically, after generating the AI inference task scheduling scheme of the radio access network according to the AI inference task guarantee policy and the terminal information, the method further comprises:
[0076] S14, sending the AI inference task scheduling scheme to the to-be-assigned node and the corresponding computing node, so that the to-be-assigned node and the corresponding computing node complete respective inference calculation tasks.
[0077] For example, assuming that the to-be-assigned node is a user equipment and the computing node is a base station, at this time, the terminal information of three user equipments and two base stations is collected, and the two base stations and the three user equipments satisfy the following conditions: the base station A provides communication connection services for the user equipment 1, the base station B provides communication connection services for the user equipments 2 and 3, the base stations A and B both provide computing services, and the AI inference models are both ResNet-18 in Table 2. Assuming that at this time, the scheduling scheme is to select appropriate computing nodes for the user equipments 1-3, and to select the split point and the exit point processed by the AI inference task, the AI inference task scheduling scheme is shown in Table 4.
[0078] Table 4 AI inference task scheduling scheme example
[0079] User Equipment Computing Node Split Point Exit Point User Equipment 1 Base Station A Layer ID = 3 Exit 2 User Equipment 2 Base Station A Layer ID = 8 Exit 3 User Equipment 3 Base Station B Layer ID = 11 Exit 3
[0080] For example, according to Table 4, the user equipment 1 is taken as an example for illustration, at this time, the computing node selected for the user equipment 1 is the base station A, that is, the base station A and the user equipment 1 jointly complete the calculation task, the split point is the third layer, and the exit point is Exit2, so the calculation amount completed by the base station A is the calculation amount from the third layer to the ninth layer. According to the selected split point, the user equipment 1 completes the inference calculation of the previous layers, the feature map generated at the split point is transmitted to the base station A, the base station A then completes the inference calculation from the third layer to the ninth layer, and finally the generated result is transmitted back to the user equipment 1. Therefore, the selection of the split point and the exit point determines the respective calculation amounts of the to-be-assigned node and the computing node, and the selection of the exit point also determines the accuracy of the model.
[0081] Compared with the prior art, the AI inference task scheduling method of the wireless access network disclosed in the application introduces an artificial intelligence technology on the RAN side, first receives an AI inference task guarantee strategy of an AI inference model sent by a non-real-time wireless controller, and obtains terminal information of at least one terminal, wherein the AI inference task guarantee strategy defines an optimization target of the wireless access network and feature information of the AI inference model, and the terminal information includes computing resource information and communication resource information; then an AI inference task scheduling scheme of the wireless access network is generated according to the AI inference task guarantee strategy and the terminal information, through selection of a computing node, a task division point and an exit point, near-real-time AI inference task cooperative scheduling can be performed on the computing node on the RAN side according to real-time communication states and computing capabilities, the efficiency of resource allocation and cooperative computing on the RAN side is improved, and the system performance is further improved.
[0082] Referring to Figure 5 , Figure 5 is a flowchart of another AI inference task scheduling method of a wireless access network provided by an embodiment of the application, and the AI inference task scheduling method of the wireless access network is implemented by a non-real-time wireless controller, and the method comprises the following steps of:
[0083] S21, obtaining model information of an AI inference model;
[0084] S22, generating an AI inference task guarantee strategy according to the model information;
[0085] S23, sending the AI inference task guarantee strategy to a near-real-time wireless controller, so that the near-real-time wireless controller generates an AI inference task scheduling scheme of a wireless access network according to the AI inference task guarantee strategy and terminal information.
[0086] Specifically, the model information includes feature information of the AI inference model and performance guarantee parameters of an AI inference task; wherein the performance guarantee parameters include but are not limited to model inference accuracy and inference calculation times per unit time.
[0087] It should be noted that the specific working process of the AI inference task scheduling method of the wireless access network described in the embodiment of the application can refer to the above-mentioned embodiments, and will not be described here.
[0088] Referring to Figure 6 , Figure 6 is a structural block diagram of an AI inference task scheduling device 100 of a wireless access network provided by an embodiment of the application, and the AI inference task scheduling device 100 of the wireless access network comprises:
[0089] The AI inference task guarantee policy receiving module 11 is configured to receive an AI inference task guarantee policy of an AI inference model sent by a non-real-time wireless controller, wherein the AI inference model is deployed in a wireless access network, and the AI inference task guarantee policy defines an optimization target of the wireless access network and feature information of the AI inference model.
[0090] The terminal information obtaining module 12 is configured to obtain terminal information of at least one terminal, wherein the terminal information includes computing resource information and communication resource information.
[0091] The AI inference task arrangement scheme generating module 13 is configured to generate an AI inference task arrangement scheme of the wireless access network according to the AI inference task guarantee policy and the terminal information.
[0092] Specifically, the AI inference task arrangement scheme generating module 13 is specifically configured to:
[0093] determine, according to each terminal information, a to-be-assigned node and a corresponding computing node according to the optimization target of the wireless access network, and select, according to each terminal information, a corresponding split point and an exit point from the feature information of the AI inference model, wherein the computing node is a base station, and the split point and the exit point are one layer in the AI inference model.
[0094] The AI inference task arrangement scheme is composed of the to-be-assigned node and the corresponding computing node, the split point, and the exit point.
[0095] Specifically, the AI inference task arrangement device 100 of the wireless access network further includes:
[0096] The AI inference task arrangement scheme sending module is configured to send the AI inference task arrangement scheme to the to-be-assigned node and the corresponding computing node, so that the to-be-assigned node and the corresponding computing node complete respective inference computing tasks.
[0097] Specifically, when the terminal is a user equipment, the terminal information obtaining module 12 is specifically configured to: send a request message to a base station in which the user equipment is camped, so that the base station reports terminal information of the user equipment according to the request message; or, subscribe to the terminal information from the base station in which the user equipment is camped, so that the base station reports the terminal information of the user equipment.
[0098] Specifically, when the terminal is a user equipment, the user equipment supports information interaction with a near real-time radio controller; then, the terminal information acquisition module 12 is specifically configured to: send a request message to the user equipment, so that the user equipment reports terminal information according to the request message; or, subscribe to the terminal information from the user equipment, so that the user equipment reports terminal information.
[0099] Referring to Figure 7 , Figure 7 is another structure block diagram of an AI inference task scheduling apparatus 200 of a wireless access network provided by an embodiment of the present application, the AI inference task scheduling apparatus 200 of the wireless access network comprising:
[0100] a model information acquisition module 21 configured to acquire model information of an AI inference model;
[0101] an AI inference task guarantee strategy generation module 22 configured to generate an AI inference task guarantee strategy according to the model information; wherein the AI inference model is deployed in a wireless access network, and the AI inference task guarantee strategy defines an optimization target of the wireless access network and feature information of the AI inference model;
[0102] an AI inference task guarantee strategy sending module 23 configured to send the AI inference task guarantee strategy to a near real-time radio controller, so that the near real-time radio controller generates an AI inference task scheduling scheme of the wireless access network according to the AI inference task guarantee strategy and terminal information; wherein the terminal information comprises computing resource information and communication resource information.
[0103] Specifically, the model information comprises feature information of the AI inference model and performance guarantee parameters of an AI inference task; wherein the performance guarantee parameters comprise, but are not limited to, model inference accuracy and inference calculation times per unit time.
[0104] Compared with the prior art, the AI inference task scheduling apparatus of the wireless access network disclosed in the present application introduces artificial intelligence technology on the RAN side, first receives an AI inference task guarantee strategy of an AI inference model sent by a non-real-time radio controller, and acquires terminal information of at least one terminal, wherein the AI inference task guarantee strategy defines an optimization target of the wireless access network and feature information of the AI inference model, and the terminal information comprises computing resource information and communication resource information; then, an AI inference task scheduling scheme of the wireless access network is generated according to the AI inference task guarantee strategy and the terminal information, through selection of computing nodes, task division points and exit points, near real-time AI inference task collaborative scheduling of the computing nodes on the RAN side can be performed according to real-time communication states and computing capabilities, the efficiency of resource allocation and collaborative computing on the RAN side is improved, and the system performance is further improved.
[0105] Referring to Figure 8 , Figure 8 is a structural block diagram of an electronic device 300 provided by an embodiment of the present application, which includes a processor 31, a memory 32, and a computer program stored in the memory 32 and executable on the processor 31. The processor 31 implements the steps in the AI inference task scheduling method embodiments of various radio access networks described above when executing the computer program, such as steps S11-S13, S21-S23.
[0106] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 32 and executed by the processor 31 to complete the present application. The one or more modules / units can be a series of computer program instruction segments that can complete a specific function, which are used to describe the execution process of the computer program in the electronic device 300.
[0107] The electronic device 300 can include, but is not limited to, a processor 31, a memory 32. Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 300 and does not constitute a limitation on the electronic device 300, which can include more or fewer components than the diagram, or combine certain components, or different components, for example, the electronic device 300 can also include an input / output device, a network access device, a bus, etc.
[0108] The processor 31 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor 31 is the control center of the electronic device 300, which connects various parts of the entire electronic device 300 through various interfaces and lines.
[0109] The memory 32 can be used to store the computer programs and / or modules, and the processor 31 realizes various functions of the electronic device 300 by running or executing the computer programs and / or modules stored in the memory 32, and calling the data stored in the memory 32. The memory 32 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), and the like. In addition, the memory 32 can include a high-speed random access memory, and can also include a nonvolatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.
[0110] The modules / units integrated in the electronic device 300, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor 31 executes the computer program, the steps of the above-mentioned various method embodiments can be realized. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0111] The above is the preferred embodiment of the present application, and it should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements are also considered within the protection scope of the present application.
Claims
1. A method for AI inference task orchestration of a radio access network, the method comprising: The method applied to a near real-time wireless controller comprises: receiving an AI inference task guarantee strategy of an AI inference model sent by a non-real-time wireless controller; wherein the AI inference model is deployed in a wireless access network, and the AI inference task guarantee strategy defines an optimization target of the wireless access network and feature information of the AI inference model; obtaining terminal information of at least one terminal; wherein the terminal information comprises computing resource information and communication resource information; generating an AI inference task scheduling scheme of the wireless access network according to the AI inference task guarantee strategy and the terminal information.
2. The AI inference task orchestration method of the wireless access network according to claim 1, wherein, The generating of the AI inference task scheduling scheme of the wireless access network according to the AI inference task guarantee strategy and the terminal information comprises: determining a to-be-allocated node and a corresponding computing node according to each terminal information and selecting a corresponding split point and an exit point from the feature information of the AI inference model according to each terminal information according to the optimization target of the wireless access network; wherein the split point and the exit point are one layer in the AI inference model; composing the AI inference task scheduling scheme with the to-be-allocated node and the corresponding computing node, the split point and the exit point. 3.The AI inference task orchestration method of a wireless access network according to claim 2, wherein, After the AI inference task scheduling scheme of the wireless access network is generated according to the AI inference task guarantee strategy and the terminal information, the method further comprises: sending the AI inference task scheduling scheme to the to-be-allocated node and the corresponding computing node to enable the to-be-allocated node and the corresponding computing node to complete respective inference computing tasks. 4.The AI inference task orchestration method of a wireless access network of claim 1, wherein, When the terminal is a user equipment, the obtaining of the terminal information of at least one terminal comprises: sending a request message to a base station in which the user equipment resides to enable the base station to report terminal information of the user equipment according to the request message; or subscribing to the terminal information from the base station in which the user equipment resides to enable the base station to report terminal information of the user equipment. 5.The AI inference task orchestration method of a wireless access network according to claim 1, wherein, When the terminal is a user equipment, the user equipment supports information interaction with the near real-time wireless controller; then, the obtaining of the terminal information of at least one terminal comprises: sending a request message to the user equipment to enable the user equipment to report terminal information according to the request message; or subscribing to the terminal information from the user equipment to enable the user equipment to report terminal information.
6. An AI inference task orchestration method of a radio access network, the method comprising: The method applied to a non-real-time wireless controller comprises: obtaining model information of an AI inference model; generating an AI inference task guarantee strategy according to the model information; wherein the AI inference model is deployed in a wireless access network, and the AI inference task guarantee strategy defines an optimization target of the wireless access network and feature information of the AI inference model; sending the AI inference task guarantee strategy to a near real-time wireless controller to enable the near real-time wireless controller to generate an AI inference task scheduling scheme of the wireless access network according to the AI inference task guarantee strategy and terminal information; wherein the terminal information comprises computing resource information and communication resource information.
7. The AI inference task orchestration method of the wireless access network according to claim 6, wherein, The model information includes feature information of the AI inference model and performance guarantee parameters of an AI inference task; wherein the performance guarantee parameters include but are not limited to model inference accuracy and inference calculation times per unit time. 8.A device for AI inference task orchestration of a radio access network, characterized in that, Comprise: An AI inference task guarantee strategy receiving module configured to receive an AI inference task guarantee strategy of an AI inference model sent by a non-real-time wireless controller; wherein the AI inference model is deployed in a wireless access network, and the AI inference task guarantee strategy defines an optimization target of the wireless access network and feature information of the AI inference model; A terminal information obtaining module configured to obtain terminal information of at least one terminal; wherein the terminal information includes computing resource information and communication resource information; An AI inference task arrangement scheme generating module configured to generate an AI inference task arrangement scheme of the wireless access network according to the AI inference task guarantee strategy and the terminal information. 9.A device for AI inference task orchestration of a radio access network, characterized in that, Comprise: A model information obtaining module configured to obtain model information of an AI inference model; An AI inference task guarantee strategy generating module configured to generate an AI inference task guarantee strategy according to the model information; wherein the AI inference model is deployed in a wireless access network, and the AI inference task guarantee strategy defines an optimization target of the wireless access network and feature information of the AI inference model; An AI inference task guarantee strategy sending module configured to send the AI inference task guarantee strategy to a near-real-time wireless controller, so that the near-real-time wireless controller generates an AI inference task arrangement scheme of the wireless access network according to the AI inference task guarantee strategy and terminal information; wherein the terminal information includes computing resource information and communication resource information.
10. An electronic device, comprising: A device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the AI inference task arrangement method of the wireless access network according to any one of claims 1 to 7 when executing the computer program.
11. A computer readable storage medium, characterized in that, The computer readable storage medium comprises a stored computer program, wherein when the computer program runs, the device where the computer readable storage medium is located executes the AI inference task arrangement method of the wireless access network according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data-centric service-based network architecture
US20210184989A1
Online reinforcement learning
WO2022060777A1