Model inference method and apparatus, network device, readable storage medium, and program product
By dynamically adapting the executor through the wireless intelligent management and orchestration function layer, and utilizing cloud-distilled lightweight models for inference on the terminal, the problem of limited computing power of terminal devices is solved, thereby improving the efficiency and accuracy of model inference.
Patent Information
- Application Number
- CN202511030948.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Constraints on terminal devices in terms of computing power, storage space, and energy consumption pose challenges to directly running large-scale artificial intelligence models, while cloud computing relies on large models in the cloud for inference, resulting in limited efficiency.
The wireless intelligent management and orchestration functional layer obtains terminal model inference requirement information, dynamically adapts the executor (cloud, node or terminal), distills lightweight models from large models in the cloud, performs local inference on the terminal, and verifies the accuracy of the results by the functional layer.
It reduces reliance on cloud computing resources, lowers network transmission overhead, improves model inference efficiency and accuracy, adapts to terminal needs, and alleviates computational pressure.
Smart Images

Figure CN120529342B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical fields of wireless communication and artificial intelligence, and in particular to a model inference method and device, a network device, a computer readable storage medium and a computer program product. BACKGROUND
[0002] With the rapid development of artificial intelligence (AI) technology, the processing demand of end-side devices (terminals) on models is increasing. However, due to the constraints of terminals in terms of computing power, storage space and energy consumption, directly running large-scale artificial intelligence models (large models) on terminals often faces great challenges.
[0003] In related technologies, cloud computing can provide strong computing power support for the processing demand of terminal models, which mainly relies on large models in the cloud for inference, resulting in the problem of limited model inference efficiency. SUMMARY
[0004] Therefore, it is necessary to provide a model inference method, device, network device, computer readable storage medium and computer program product to solve the above technical problems.
[0005] In a first aspect, the present application provides a model inference method applied to a wireless intelligent management and orchestration function layer, comprising:
[0006] obtaining, by an access node of an intelligent radio access network, model inference demand information of a terminal;
[0007] determining an execution party of a model inference task according to the model inference demand information and resource state information of the intelligent radio access network; the model inference task is obtained according to the model inference demand information; the execution party includes a cloud, a node or the terminal; the node is a node in a node set of the intelligent radio access network; the node set includes the access node;
[0008] triggering the execution party to execute the model inference task by using a target model, so that the terminal obtains a target inference result; wherein the target model is distilled from a large model by the cloud according to the model inference demand information; and the target inference result is an inference result that meets a verification condition of the wireless intelligent management and orchestration function layer.
[0009] In a second aspect, the present application further provides a model inference device applied to a wireless intelligent management and orchestration function layer, comprising:
[0010] an information obtaining module, configured to obtain, by an access node of an intelligent radio access network, model inference demand information of a terminal;
[0011] an execution determining module, configured to determine an execution party of a model inference task according to the model inference requirement information and resource state information of the intelligent wireless access network, the model inference task being obtained according to the model inference requirement information, the execution party comprising a cloud, a node or the terminal, the node being a node in a node set of the intelligent wireless access network, and the node set comprising the access node;
[0012] a task processing module, configured to trigger the execution party to execute the model inference task by using a target model, so that the terminal obtains a target inference result, wherein the target model is obtained by the cloud from a large model distillation according to the model inference requirement information, and the target inference result is an inference result meeting a verification condition of the wireless intelligent management and orchestration function layer.
[0013] In a third aspect, the present application further provides a network device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0014] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the above method when executed by a processor.
[0015] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, and the computer program implements the steps of the above method when executed by a processor.
[0016] The model inference method, device, network equipment, computer readable storage medium and computer program product can obtain model inference requirement information of the terminal through the access node of the intelligent wireless access network, determine an execution party of the model inference task according to the model inference requirement information and resource state information of the intelligent wireless access network, and the execution party can include a cloud, a node or a terminal, wherein the node is a node in a node set of the intelligent wireless access network, and the node set includes the access node. The wireless intelligent management and arrangement function layer can trigger the execution party to execute the model inference task by using a target model, so that the terminal obtains a target inference result, wherein the target model is distilled from a large model by the cloud according to the model inference requirement information, and the target inference result is an inference result that meets the verification condition of the wireless intelligent management and arrangement function layer. The scheme can flexibly adapt the execution party for the model inference task according to the model inference requirement information and the resource state information of the intelligent wireless access network by the wireless intelligent management and arrangement function layer, the execution party can include the cloud, the node or the terminal, and the cloud is not needed to be relied on to execute the model inference task, the dependence on the cloud computing resource can be reduced, a large amount of data transmission between the terminal side and the cloud side can be reduced, the network transmission overhead is reduced, the computing power can be avoided to be concentrated on the terminal side, the efficiency of the model inference is improved as a whole, the target model is distilled from the large model by the cloud, so that the model inference requirement of the terminal can be adapted, the computing pressure of the model inference is also reduced, and the inference result obtained by the terminal needs to meet the verification condition of the wireless intelligent management and arrangement function layer, so that the accuracy of the model inference is ensured. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without creating any inventive labor.
[0018] Figure 1 An application environment diagram of the model inference method in an embodiment;
[0019] Figure 2 A schematic diagram of the wireless intelligent management and arrangement function layer in an embodiment;
[0020] Figure 3 A flowchart of the model inference method in an embodiment;
[0021] Figure 4 A timing diagram of the model inference method in an embodiment;
[0022] Figure 5 A structural block diagram of the model inference device in an embodiment;
[0023] Figure 6 Figure 1 is a diagram of an internal structure of a network device in one embodiment. DETAILED DESCRIPTION
[0024] For the purpose of making the purpose, technical scheme and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0025] The model inference method provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 The application environment can include a terminal (terminal side), a node of an intelligent radio access network (AI RAN, Artificial Intelligence Radio Access Network), a radio intelligent management and orchestration function layer (RAN AI Layer), and a cloud.
[0026] The terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, a projection device, etc. The portable wearable device can be a head-mounted device, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, etc.
[0027] With the development of intelligent radio access network (RAN, Radio Access Network), the intelligent radio access network provides a lower latency and intelligent computing collaboration solution for the terminal side by introducing AI computing power in the wireless network. The intelligent radio access network can apply artificial intelligence and machine learning technology to various aspects of wireless access network planning, deployment, management, and optimization to improve network performance, resource utilization efficiency, enhance network flexibility and intelligence, and better meet user demand for high-speed, stable, and low-latency network connections.
[0028] The radio intelligent management and orchestration function layer can intelligently and automatically manage and schedule the intelligent radio access network to improve network performance, optimize resource utilization, enhance service flexibility, and improve user experience. For example, Figure 2As shown, the wireless intelligent management and orchestration function layer can include a model management function module and a computing resource scheduling module. The model management function module can be a function module in the wireless intelligent management and orchestration function layer for managing artificial intelligence models (AI models), and can be used to distribute target models obtained from the cloud and the like. The computing resource scheduling module can be a component in the wireless intelligent management and orchestration function layer that implements efficient collaborative management and optimized allocation of computing, communication, and other resources in a wireless communication network. The wireless intelligent management and orchestration function layer can be connected to various intelligent wireless access networks, including but not limited to 3GPP RAN (wireless access network in 3GPP), cloud-native RAN (network architecture combining cloud-native technology and wireless access network RAN), or other RAN.
[0029] As shown in the application environment, Figure 1 a terminal can access a node of an intelligent wireless access network, and the node of the intelligent wireless access network accessed by the terminal is referred to as an access node. The intelligent wireless network can include a plurality of nodes, which can form a node set of the intelligent wireless access network (which can include nodes 1 to n), the node set can include the access node, and the nodes in the node set of the intelligent wireless network can be controlled and managed by the wireless intelligent management and orchestration function layer.
[0030] In related technologies, cloud computing can provide strong computing power support for the processing needs of the terminal's model, which mainly relies on large models in the cloud for inference, resulting in limited model inference efficiency.
[0031] To this end, the model inference method provided by the embodiments of the present application can flexibly adapt the execution party for the model inference task according to the model inference requirement information and the resource state information of the intelligent wireless access network by the wireless intelligent management and orchestration function layer. The execution party can include the cloud, the node or the terminal, without relying on the cloud to execute the model inference task, which can reduce the dependence on cloud computing resources, reduce the large amount of data transmission between the end side and the cloud, reduce network transmission overhead, and also avoid concentrating computing power on the end side, thereby improving the efficiency of model inference as a whole. The target model is distilled from the large model by the cloud, so that it can adapt to the model inference requirements of the terminal and also reduce the computing pressure of model inference. In addition, the inference result obtained by the terminal needs to meet the verification condition of the wireless intelligent management and orchestration function layer, thereby ensuring the accuracy of model inference.
[0032] In one exemplary embodiment, as shown in Figure 3 a model inference method is provided, which can be applied to the wireless intelligent management and orchestration function layer as in Figure 1 the method can include the following steps:
[0033] In step S301, the model inference requirement information of the terminal is obtained by the access node of the intelligent wireless access network.
[0034] In this step, the terminal can be connected to the access node of the intelligent wireless access network, and the model inference request is sent to the wireless intelligent management and arrangement function layer through the access node. The model inference request can include the model inference requirement information. The wireless intelligent management and arrangement function layer can receive the model inference request of the terminal through the access node of the intelligent wireless access network, and obtain the model inference requirement information according to the model inference request. The model inference requirement information is the requirement information provided by the terminal for the model inference, which can include the inference resource information of the terminal (such as the computing power, power, storage space, etc. of the terminal itself), the inference target required by the terminal (such as the need to identify a specific object, etc.), the source data provided by the terminal for the model inference (such as image data, text data, etc.), the inference delay required by the terminal (such as how long to get the inference result, etc.), the computing power resource required by the terminal, etc.
[0035] In step S302, the execution party of the model inference task is determined according to the model inference requirement information and the resource state information of the intelligent wireless access network.
[0036] The resource state information of the intelligent wireless access network refers to the state information of the resources of the intelligent wireless access network that can be used to execute the model inference, which can include the state information of the network resources and the computing power resources. The state information of the network resources can represent the state of the transmission and load of the intelligent wireless access network, and the state information of the computing power resources can represent the computing power of each node of the intelligent wireless access network.
[0037] In this step, the wireless intelligent management and arrangement function layer can determine the execution party of the model inference task according to the model inference requirement information and the resource state information of the intelligent wireless access network. The model inference task can be obtained by the wireless intelligent management and arrangement function layer according to the model inference requirement information, for example, the model inference task can be generated according to the inference target required by the terminal and the source data provided by the terminal for model inference. The execution party of the model inference task can include the cloud, the node of the intelligent wireless access network or the terminal, wherein the node can be a node in the node set of the intelligent wireless access network, and the node set includes the access node described above, that is, the execution party can be the access node. Therefore, the wireless intelligent management and arrangement function layer dynamically adapts the cloud, the node of the intelligent wireless access network or the terminal as the execution party of the model inference task according to the model inference requirement information and the resource state information of the intelligent wireless access network, for example, the wireless intelligent management and arrangement function layer can determine the cloud, the node of the intelligent wireless access network or the terminal as the execution party of the model inference task according to the inference resource information of the terminal, the inference delay required by the terminal, the computing resource required by the terminal, the state information of the network resource and the computing resource of the intelligent wireless access network.
[0038] In step S303, the execution party triggers the execution of the model inference task using the target model to make the terminal obtain the target inference result.
[0039] The target model is distilled from a large model by the cloud according to the model inference requirement information, and the target model adapted to the terminal can be distilled from a large model by the cloud according to the model inference requirement information, for example, the cloud can distill a lightweight small model adapted to the terminal as the target model from a large model according to the inference resource information of the terminal, the inference target required by the terminal and the source data provided by the terminal for model inference. In the field, a large model and a small model can be embodied according to the size, the number of parameters and the like of the model, for example, a large model generally has a large number of parameters (such as tens of billions or even tens of billions), the number of parameters of a small model is relatively small (may be tens of thousands to tens of millions), a large model generally has a complex neural network architecture (contains multiple hidden layers and a large number of neurons, the model file size is also large, may reach tens of gigabytes or even larger), and the model architecture of a small model is relatively simple (fewer layers and fewer neurons, the model file is usually smaller, a few megabytes to a few tens of megabytes).
[0040] In this step, the wireless intelligent management and orchestration function layer can assign the model inference task to the execution party after determining the execution party, and make the execution party obtain the target model obtained by the cloud, trigger the execution party to execute the model inference task using the target model, and make the terminal obtain the target inference result. The target inference result is an inference result that meets the verification condition of the wireless intelligent management and orchestration function layer. After the execution party executes the model inference task using the target model, it can obtain the corresponding inference result, which can be instructed by the wireless intelligent management and orchestration function layer whether to meet the verification condition. The verification condition can be set by the user. As an implementation manner, the verification condition can include whether the inference result meets the required inference target of the terminal, such as identifying animals in the image. The verification condition can include the name of the animal pointed to by the inference result. For example, the required inference target is to generate images of animals, characters, and other objects. The verification condition can include that the image quality evaluation value of the generated image meets the threshold condition. The image quality evaluation value of the generated image can be obtained by the image quality evaluation model deployed in the wireless intelligent management and orchestration function layer. The generated image can be input into the image quality evaluation model (which can be an artificial intelligence model such as a convolutional neural network) to obtain the image quality evaluation value. If the image quality evaluation value is greater than or equal to the set threshold (which can be set by the user), the image quality evaluation value of the generated image meets the threshold condition. If the image quality evaluation value is less than the set threshold, the image quality evaluation value of the generated image does not meet the threshold condition. As an implementation manner, if the inference result meets the verification condition, the inference result can be given to the terminal as the target inference result. If the inference result does not meet the verification condition, the target inference result can be obtained by re-inferencing after optimizing the target model and given to the terminal.
[0041] The model inference method of the embodiment can obtain the model inference requirement information of the terminal through the access node of the intelligent wireless access network, determine the execution party of the model inference task according to the model inference requirement information and the resource state information of the intelligent wireless access network, and the execution party can include the cloud, the node or the terminal, wherein the node is a node in the node set of the intelligent wireless access network, and the node set includes the access node. The wireless intelligent management and arrangement function layer can trigger the execution party to execute the model inference task by using the target model, so that the terminal obtains the target inference result, wherein the target model is distilled from the large model by the cloud according to the model inference requirement information, and the target inference result is the inference result that meets the verification condition of the wireless intelligent management and arrangement function layer. The scheme can flexibly adapt the execution party for the model inference task according to the model inference requirement information and the resource state information of the intelligent wireless access network by the wireless intelligent management and arrangement function layer, and the execution party can include the cloud, the node or the terminal. The model inference task does not need to rely on the cloud, which can reduce the dependence on the cloud computing resources, reduce the large amount of data transmission between the terminal and the cloud, reduce the network transmission overhead, avoid the concentration of computing power on the terminal side, and improve the efficiency of model inference as a whole. The target model is distilled from the large model by the cloud, which can adapt to the model inference requirement of the terminal and reduce the computing pressure of the model inference. In addition, the inference result obtained by the terminal needs to meet the verification condition of the wireless intelligent management and arrangement function layer, so as to ensure the accuracy of the model inference.
[0042] In an exemplary embodiment, after obtaining the model inference requirement information of the terminal through the access node of the intelligent wireless access network in step S301, the above method can further include the following steps:
[0043] The model inference requirement information is sent to the cloud. The model inference requirement information is used to distill the target model from the large model by the cloud.
[0044] In the embodiment, the wireless intelligent management and arrangement function layer can send the model inference requirement information to the cloud after obtaining the model inference requirement information of the terminal. The cloud can distill the target model adapted to the terminal from the large model according to the model inference requirement information. For example, the cloud can distill a light small model adapted to the terminal as the target model (a commodity identification model, a traffic sign identification model, etc.) according to the inference resource information (computing power, power, storage space, etc.) of the terminal, the inference target (identifying commodities or traffic signs, etc.) required by the terminal and the source data (image data, text data, etc.) provided by the terminal for model inference. The scheme of the embodiment can distill the target model from the large model by the cloud, so that the wireless intelligent management and arrangement function layer can flexibly distribute the model inference task to the cloud, the node of the intelligent wireless access network or the terminal for execution.
[0045] In an example embodiment, determining, according to the model inference requirement information and the resource state information of the intelligent wireless access network, an execution party of the model inference task of step S302 can include:
[0046] If it is determined according to the model inference requirement information that the terminal satisfies the task execution condition, the terminal is determined as the execution party of the model inference task; if it is determined according to the model inference requirement information that the terminal does not satisfy the task execution condition, a node in the node set of the intelligent wireless access network or the cloud is determined as the execution party according to the model inference requirement information and the resource state information of the intelligent wireless access network.
[0047] In the embodiment, the wireless intelligent management and orchestration function layer can first determine whether the terminal satisfies the task execution condition according to the model inference requirement information. The task execution condition can include conditions such as delay and computing power resources of model inference, and can be obtained according to the model inference requirement information (such as inference delay required by the terminal, computing power resources required by the terminal). For example, the wireless intelligent management and orchestration function layer can determine whether the terminal satisfies the task execution condition according to the computing power of the terminal itself. If it is determined that the terminal satisfies the task execution condition, the wireless intelligent management and orchestration function layer can determine the terminal as the execution party of the model inference task, and the terminal executes the model inference task using the target model to reduce network transmission delay. If it is determined that the terminal does not satisfy the task execution condition, the wireless intelligent management and orchestration function layer can further determine a node in the node set of the intelligent wireless access network or the cloud as the execution party according to the model inference requirement information and the resource state information of the intelligent wireless access network.
[0048] In an example embodiment, further, determining, according to the model inference requirement information and the resource state information of the intelligent wireless access network, a node in the node set of the intelligent wireless access network or the cloud as the execution party can include:
[0049] If it is determined according to the model inference requirement information and the resource state information of the intelligent wireless access network that there is a node in the node set of the intelligent wireless access network that satisfies the task execution condition, the execution party is determined according to the node that satisfies the task execution condition; if it is determined according to the model inference requirement information and the resource state information of the intelligent wireless access network that there is no node in the node set of the intelligent wireless access network that satisfies the task execution condition, the cloud is determined as the execution party.
[0050] In the embodiment, in the case that the terminal does not satisfy the task execution condition, the wireless intelligent management and arrangement function layer can further determine a node in the node set of the intelligent wireless access network or the cloud as the execution party according to the model inference requirement information and the resource state information of the intelligent wireless access network. The wireless intelligent management and arrangement function layer can determine whether there is a node in the node set of the intelligent wireless access network that satisfies the task execution condition according to the model inference requirement information and the resource state information of the intelligent wireless access network. For example, the wireless intelligent management and arrangement function layer can determine whether there is a node in the node set of the intelligent wireless access network that satisfies the task execution condition (the conditions of the model inference time delay and the computing power resource) according to the terminal demand inference time delay, the terminal demand computing power resource, the network resource state information (which can represent the state of the transmission and load of the intelligent wireless access network), and the computing power resource state information (which can represent the computing power of each node of the intelligent wireless access network). If there is a node that satisfies the task execution condition, the node can be determined as the execution party in the case that the number of the node is one, and the node can be determined as the execution party according to the distance between the node and the terminal in the case that the number of the node is more than one. The node closest to the terminal can be selected as the execution party to improve the efficiency of the model inference. If there is no node that satisfies the task execution condition, the wireless intelligent management and arrangement function layer can determine the cloud as the execution party to complete the model inference task.
[0051] In one of the example embodiments, in the case that the execution party includes the cloud, the step of triggering the execution party to execute the model inference task using the target model to make the terminal obtain the target inference result can include:
[0052] triggering the cloud to execute the model inference task using the target model; obtaining a first inference result obtained by the cloud in executing the model inference task, determining the first inference result as the target inference result satisfying the verification condition, and sending the target inference result to the terminal through the access node.
[0053] In the embodiment, in the case that the execution party includes the cloud, since the cloud has distilled the target model from the large model according to the model inference requirement information, the wireless intelligent management and arrangement function layer can assign the model inference task to the cloud to trigger the cloud to execute the model inference task by using the target model to obtain a first inference result, and the cloud sends the first inference result to the wireless intelligent management and arrangement function layer. The wireless intelligent management and arrangement function layer can determine whether the first inference result meets the verification condition according to the execution party of the model inference task, that is, the verification condition can include whether the model inference task is executed by the target execution party. In the case that the cloud is the target execution party, the wireless intelligent management and arrangement function layer can determine the first inference result obtained by the cloud in executing the model inference task as the target inference result meeting the verification condition, and send the target inference result to the terminal through the access node. Thus, in the case that neither the terminal nor the node of the intelligent wireless access network meets the task execution condition, the target inference result can be obtained by the cloud and given to the terminal.
[0054] In another exemplary embodiment, in the case that the execution party includes the node, the step of triggering the execution party to execute the model inference task by using the target model to make the terminal obtain the target inference result can include:
[0055] obtaining the target model obtained by the cloud; sending the target model to the node to make the node execute the model inference task by using the target model to obtain a second inference result, and sending the second inference result as the target inference result meeting the verification condition to the terminal through the access node.
[0056] In the embodiment, in the case that the execution party includes the node of the intelligent wireless access network, the wireless intelligent management and arrangement function layer can obtain the target model distilled from the large model according to the model inference requirement information, send the target model to the node, and assign the model inference task to the node of the intelligent wireless access network to trigger the node to execute the model inference task by using the target model to obtain a second inference result. The wireless intelligent management and arrangement function layer can also instruct the node to determine the second inference result as a target inference result satisfying the verification condition, and instruct the node to send the target inference result to the terminal through the access node. The wireless intelligent management and arrangement function layer can instruct the second inference result to satisfy the verification condition according to the execution party of the model inference task, that is, the verification condition can include whether the model inference task is the target execution party. The node of the intelligent wireless access network can be the target execution party. Thus, when the wireless intelligent management and arrangement function layer assigns the model inference task to the node of the intelligent wireless access network, it can instruct the node to determine the second inference result as the target inference result satisfying the verification condition and instruct the node to send the target inference result to the terminal through the access node. Thus, in the case that the terminal does not satisfy the task execution condition and the node of the intelligent wireless access network satisfies the task execution condition, the node of the intelligent wireless access network sends the target inference result to the terminal through the access node, avoids concentrating computing power in the cloud or on the terminal side, and improves the overall inference efficiency.
[0057] In another exemplary embodiment, in the case that the execution party includes the terminal, the step S303 of triggering the execution party to execute the model inference task by using the target model to make the terminal obtain the target inference result can include: obtaining the target model obtained by the cloud; sending the target model to the terminal through the access node to make the terminal execute the model inference task by using the target model to obtain a third inference result and return the third inference result through the access node; and verifying the third inference result to determine the target inference result, so that the terminal obtains the target inference result.
[0058] In the embodiment, when the execution party includes the terminal, the wireless intelligent management and orchestration function layer can obtain a target model distilled from a large model by the cloud according to model inference requirement information, send the target model to the terminal through the access node, and assign a model inference task to the terminal to trigger the terminal to execute the model inference task using the target model to obtain a third inference result, and return the third inference result to the wireless intelligent management and orchestration function layer through the access node. When the execution party includes the terminal, after the wireless intelligent management and orchestration function layer obtains the target model from the cloud, before sending the target model to the terminal through the access node, the wireless intelligent management and orchestration function layer can first determine whether the inference resource information of the terminal has changed. If not, the wireless intelligent management and orchestration function layer can send the target model to the terminal through the access node. If it has changed, and the changed inference resource information indicates that the target model is applicable to the terminal (such as an increase in the terminal's own computing power, power, storage space, etc.), the wireless intelligent management and orchestration function layer can send the target model to the terminal through the access node. If it has changed, and the changed inference resource information indicates that the target model is not applicable to the terminal (such as a decrease in the terminal's own computing power, power, storage space, etc.), the wireless intelligent management and orchestration function layer can send the changed inference resource information to the cloud for secondary distillation of the target model according to the changed inference resource information and return the secondary distilled target model to the wireless intelligent management and orchestration function layer. The wireless intelligent management and orchestration function layer can send the secondary distilled target model to the terminal through the access node for execution of the model inference task to obtain a third inference result. The wireless intelligent management and orchestration function layer verifies the third inference result to determine a target inference result that meets the verification condition, and then the wireless intelligent management and orchestration function layer can cause the terminal to obtain the target inference result. As an implementation manner, the verification condition can include whether the inference result meets the required inference target of the terminal, etc. For example, if the required inference target is to identify animals in an image, the verification condition can include whether the inference result points to the name of the animal, etc. As an implementation manner, if the third inference result meets the verification condition, the wireless intelligent management and orchestration function layer can instruct the terminal to take the third inference result as the target inference result, thereby obtaining the target inference result. If the third inference result does not meet the verification condition, the wireless intelligent management and orchestration function layer can optimize the target model through the cloud (such as increasing the number of network layers of the model, etc.), send the optimized target model to the terminal through the access node for re-inference, and the third inference result obtained by re-inference can also be verified again in the above manner until the third inference result meets the verification condition, etc. Thus, when the terminal meets the task execution condition, the model inference task is executed by the terminal and the inference result is verified by the wireless intelligent management and orchestration function layer, which improves the inference credibility and enhances the adaptability of the model in complex wireless environments.
[0059] In an example embodiment, further, in the case where the execution party comprises a terminal, the above-mentioned verifying the third inference result to determine the target inference result, and causing the terminal to obtain the target inference result, can comprise:
[0060] obtaining a reference inference result of a reference execution party; if it is judged according to the reference inference result and the third inference result that the third inference result satisfies the verification condition, sending indication information to the terminal through the access node; if it is judged according to the reference inference result and the third inference result that the third inference result does not satisfy the verification condition, taking the reference inference result as the target inference result and sending it to the terminal through the access node.
[0061] In the example embodiment, in the case where the execution party comprises a terminal, the wireless intelligent management and orchestration function layer can obtain the reference inference result of the reference execution party. The wireless intelligent management and orchestration function layer can select a node in the node set of the cloud or the intelligent wireless access network as the reference execution party, for example, a node closest to the wireless intelligent management and orchestration function layer can be selected as the reference execution party. The model inference task can be assigned to the node, and the reference execution party is triggered to execute the model inference task by using the target model to obtain and return the reference inference result. Thus, the wireless intelligent management and orchestration function layer can compare the reference inference result and the third inference result to judge whether the third inference result satisfies the verification condition. The verification condition can comprise that the third inference result is the same as the reference inference result. If the third inference result satisfies the verification condition, the wireless intelligent management and orchestration function layer can send indication information to the terminal through the access node. The indication information is used to instruct the terminal to determine the third inference result as the target inference result, and thus the terminal can obtain the target inference result. If the third inference result does not satisfy the verification condition, the intelligent management and orchestration function layer can take the reference inference result as the target inference result, and send the target inference result to the terminal through the access node, and thus the terminal can obtain the target inference result. The scheme of the example embodiment can ensure the accuracy of the end-side inference and reduce the misjudgment rate.
[0062] In an example embodiment, in combination with Figure 4 The model inference method of the present application is described.
[0063] In the related art, inference mainly relies on cloud-side large models. Even if a lightweight model is directly deployed on the terminal side, it is difficult to optimize and adjust according to the terminal-side computing power, network conditions and task requirements due to the lack of dynamic adaptation capability. Especially under different network environments and computing power conditions, the deployment and execution strategy of the model is difficult to adjust flexibly, resulting in limited inference efficiency and accuracy. In addition, the terminal-side inference result usually lacks a verification mechanism, which may affect the inference accuracy due to environmental changes, model drift and other factors, and lacks reliable model quality assurance.
[0064] To this end, the scheme of the present embodiment can implement model collaborative inference based on a cloud-intelligent radio access network (AI RAN)-terminal architecture. The cloud-side large model can generate a lightweight small model according to the terminal-side requirements through dynamic distillation, and utilize the wireless intelligent management and orchestration function layer (RAN AI Layer) for intelligent scheduling and inference optimization. The wireless intelligent management and orchestration function layer can not only reasonably allocate inference tasks according to the current network state and computing power, but also combine the computing power and storage capacity of the terminal side to perform direct deployment of the model and secondary distillation optimization of the model. At the same time, the terminal-side inference result can be verified through the wireless intelligent management and orchestration function layer to improve the accuracy and reliability of the inference. Thus, the scheme of the present embodiment combines the global optimization capability of the cloud side, the intelligent scheduling capability of the wireless intelligent management and orchestration function layer, and the low-power local inference capability of the terminal side, to realize an efficient, dynamically adaptive intelligent model collaborative optimization strategy.
[0065] Overall, the scheme of the present embodiment can include the following steps:
[0066] 1. Cloud-side intelligent model distillation: Based on the specific requirements of the terminal side, the cloud-side large model first performs intelligent distillation to generate a lightweight small model adapted to the terminal side, which can optimize the model structure to adapt to the computing power and storage capacity of the terminal side.
[0067] 2. Model scheduling and distribution: The small model obtained through distillation can be pushed to the wireless intelligent management and orchestration function layer, which can intelligently allocate according to the current network access state, node computing power, terminal device capability and other information, and can select the optimal node to complete the model inference task, or select the cloud side or the terminal side to complete the model inference task.
[0068] 3. Terminal-side intelligent inference: The terminal side can run the lightweight model locally for inference, achieving low-latency and efficient inference calculation, reducing dependence on cloud computing resources, and reducing network transmission overhead.
[0069] 4. Reasoning result verification and optimization: the reasoning result of the terminal side is uploaded to the wireless intelligent management and orchestration function layer, which can verify the accuracy by combining historical data or comparing the reasoning results of the nodes of the intelligent wireless access network, etc., to ensure the reliability of the reasoning result of the terminal side, and can also dynamically optimize the parameters of the lightweight small model according to the verification situation to improve the model generalization ability and reasoning accuracy.
[0070] Therefore, the scheme of the embodiment takes the wireless intelligent management and orchestration function layer as the intelligent scheduling and reasoning verification center, ensuring efficient deployment of the model, low latency reasoning and accuracy of the reasoning result.
[0071] In the scheme of the embodiment, multiple network devices such as the cloud, the wireless intelligent management and orchestration function layer, the nodes of the intelligent wireless access network and the terminal are involved, and through the interaction of these network devices, dynamic distillation, distribution, reasoning and verification of intelligent models are realized, thereby optimizing the reasoning efficiency and accuracy of the terminal side and improving the intelligent computing capability of the overall network.
[0072] In the above architecture of the embodiment, the cloud can undertake the training and distillation of intelligent models. The cloud can generate lightweight small models based on the large model and the computing capability, storage space and business demand of the terminal side, etc. The cloud can also push the distilled model to the wireless intelligent management and orchestration function layer. The cloud is not only an AI computing platform, but also cooperates with the wireless intelligent management and orchestration function layer to provide adaptive intelligent models for different terminals.
[0073] The wireless intelligent management and orchestration function layer can be deployed in an edge side data center and can schedule and control multiple nodes of the intelligent wireless access network. The wireless intelligent management and orchestration function layer can be responsible for receiving the model pushed from the cloud, and can intelligently decide the reasoning execution path of the model in combination with the load condition of the current intelligent wireless access network, the distribution of the computing power resources of the nodes and the computing capability of the terminal side, etc. When the computing capability of the terminal side is strong, the wireless intelligent management and orchestration function layer can directly deploy the small model to the terminal side to perform local reasoning. When the computing capability of the terminal side is limited, the wireless intelligent management and orchestration function layer can select to perform reasoning calculation in the nodes of the intelligent wireless access network and then send the reasoning result to the terminal side. In addition, the wireless intelligent management and orchestration function layer also has the terminal side reasoning result verification function, that is, after the terminal side performs reasoning, the reasoning result is returned to the wireless intelligent management and orchestration function layer for secondary checking to improve the reasoning accuracy and robustness.
[0074] In terms of network devices, the above architecture of the present embodiment can also involve radio access network (RAN) devices such as eNB (evolved Node B) / gNB (new generation Node B), RRU (remote radio unit), DU (distributed unit) / CU (centralized unit), etc. These devices provide efficient and low-latency intelligent computing services for the end side through 5G / 6G wireless access technology. The radio access network devices not only connect the end side, but also serve as computing and communication nodes of the wireless intelligent management and orchestration function layer, support dynamic scheduling of model inference tasks, and optimize data transmission paths using wireless channel sensing to reduce network latency and computing overhead.
[0075] In the present embodiment, the edge device can include smartphones, AR / VR devices, robots, autonomous driving systems, industrial Internet of Things terminals, etc. These end-side devices can dynamically receive lightweight models allocated by the wireless intelligent management and orchestration function layer according to their computing power, storage space, and business requirements, and perform local inference computation. When the end-side computing capability is limited, collaborative inference can be allocated by the wireless intelligent management and orchestration function layer, and business decisions can be made after verifying the inference results. In addition, data interaction between the end side and the wireless intelligent management and orchestration function layer can be performed through the 5G / 6G network to ensure the real-time and stability of intelligent inference.
[0076] In terms of network communication relationships, the cloud can communicate with the wireless intelligent management and orchestration function layer through the core network (5GC) or the SMO (service management and orchestration) of the O-RAN (open radio access network) architecture to push distilled models. The wireless intelligent management and orchestration function layer can obtain real-time resource state information of the intelligent radio access network to optimize model distribution strategies. The wireless intelligent management and orchestration function layer can determine the optimal model inference execution path and perform local inference or collaborative computation when necessary. The end side directly interacts with the intelligent radio access network (which has wireless access capabilities and intelligent business processing capabilities), receives lightweight models and performs inference, and uploads the inference results for secondary verification by the wireless intelligent management and orchestration function layer to ensure the accuracy of the inference results.
[0077] On the software architecture, in the embodiment, cloud AI computing framework, intelligent management and orchestration function, end-side intelligent model engine and other components can be included. The cloud AI computing framework can be responsible for training and distilling large models, and interacting with the wireless intelligent management and orchestration function layer through an API (application programming interface). The intelligent management and orchestration function can include computing power management, intelligent inference management, inference verification module and other functions, and can support dynamic scheduling and model verification of intelligent wireless access network side inference tasks. The end-side intelligent model engine can be a lightweight local inference framework, can be compatible with the model provided by the wireless intelligent management and orchestration function layer, and can support interactive optimization of inference results.
[0078] The scheme of the embodiment can be widely applied to intelligent robots, augmented reality (AR), autonomous driving, industrial Internet of Things (IoT) and other scenarios. For example, in the intelligent robot scenario, the robot can obtain an intelligent navigation model through the cloud, and the wireless intelligent management and orchestration function layer can push the optimized small model to the robot terminal after scheduling computing resources, so that the robot can perform efficient local inference and improve the accuracy of path planning. In the augmented reality (AR) scenario, the AR terminal can dynamically allocate computing tasks through the wireless intelligent management and orchestration function layer, reduce the local computing burden, and improve the real-time rendering performance. In the autonomous driving system, the large-scale AI model trained by the cloud can be adapted to the capability of the vehicle-mounted computing platform under the intelligent scheduling of the wireless intelligent management and orchestration function layer, and efficient inference can be performed at the end side to improve the intelligent decision-making capability of autonomous driving.
[0079] The following describes the application of the scheme of the embodiment in augmented reality (AR) and environmental perception with the intelligent glasses as the end-side device, and describes how the intelligent glasses use the scheme to perform scene recognition, model inference and optimization to improve the intelligent interaction experience. The process can be completed by the cloud, the wireless intelligent management and orchestration function layer, the intelligent wireless access network and the end-side intelligent glasses, and can include model distillation, intelligent scheduling, inference verification and other steps, combined with Figure 4 The method can include the following steps:
[0080] In step S401, the end-side intelligent glasses collect data. In step S402, the end-side intelligent glasses send a model inference request to the AI RAN node according to the collected data.
[0081] When the user wears the smart glasses and enters a new environment or has information recognition needs, the sensing system (including the camera, depth sensor, etc.) of the smart glasses will capture the image and sensing data of the current environment. The lightweight perception model running inside the smart glasses will perform preliminary object recognition, such as detecting road signs, buildings, and commodities, etc. However, due to the limited computing power of the end-side device, it may not be able to independently complete the semantic understanding of complex scenes. To further improve the recognition accuracy, the smart glasses can send a model inference request to the AI RAN node according to the collected data to request a more accurate model for inference calculation. The model inference request can include low-resolution images, key point data, computing power requirements, and model inference requirement information such as the computing power and storage space of the smart glasses of the current scene, so that the wireless intelligent management and arrangement function layer can evaluate the model inference strategy accordingly.
[0082] In step S403, the AI RAN node sends the model inference request to the computing resource scheduling module of the wireless intelligent management and arrangement function layer. In step S404, the computing resource scheduling module determines the execution party of the model inference task according to the model inference requirement information and the resource state information of the intelligent radio access network. In step S405, the computing resource scheduling module sends the execution party of the model inference task and the model inference request to the model management function module of the wireless intelligent management and arrangement function layer. In step S406, in the case where the terminal is the execution party of the model inference task, the model management function module hands over the model inference task to the terminal. In step S407, in the case where the node is the execution party of the model inference task, the model management function module hands over the model inference task to the node. In step S408, in the case where the cloud is the execution party of the model inference task, the model management function module hands over the model inference task to the cloud.
[0083] The computing resource scheduling module can evaluate the end-side computing power, determine whether the smart glasses can perform inference calculation of the complete model, and if the end-side computing power is insufficient, the AI RAN node or the cloud computing power needs to be used. The best inference path is selected: if the end-side computing is feasible, the model inference task can be allocated to the smart glasses for local execution to reduce network transmission delay. If the end-side computing power is insufficient and the AI RAN node can handle it, the nearest AI RAN node (such as gNB (next-generation base station) or vBBU (virtual baseband processing unit)) can be selected to perform inference calculation and return the result to the smart glasses. In the case of high load of the AI RAN node, more accurate inference can be performed in the cloud, and the processed inference result is returned to the wireless intelligent management and arrangement function layer and then sent to the smart glasses by the access node. The decision basis of the computing resource scheduling module can include real-time network load, computing power, power of the smart glasses, inference delay requirement, and other parameters to ensure that the inference task can be completed efficiently.
[0084] Step S409, the model management function module sends a model inference request to the cloud. Step S410, the cloud obtains a target model from the large model distillation according to the model inference requirement information.
[0085] The cloud-side large model library can dynamically distill a lightweight model as a target model according to the task requirements of the smart glasses. As an example, the distillation process can include:
[0086] According to the scene characteristics, the model is cut: if the user is in a shopping mall, a commodity recognition model can be cut; if the user is on the street, a traffic sign recognition model can be optimized.
[0087] Optimize the glasses computing power parameters: the calculation complexity of the model can be adjusted, and redundant calculations can be reduced, so that it can run efficiently on the smart glasses.
[0088] Generate a small model that fits: finally get a small model that fits the current environment as a target model, which can be pushed to the wireless intelligent management and arrangement function layer.
[0089] Step S411, in the case where the cloud does not need to perform the model inference task, the cloud sends the target model to the model management function module.
[0090] Step S412, in the case where the AI RAN node performs the model inference task, the model management function module sends the target model to the AI RAN node. Step S413, in the case where the terminal performs the model inference task, the model management function module sends the target model to the terminal.
[0091] Step S414, in the case where the cloud performs the model inference task, the cloud sends the inference result obtained by performing the model inference task to the wireless intelligent management and arrangement function layer. Step S415, the wireless intelligent management and arrangement function layer sends the inference result as a target inference result to the access node. Step S416, the access node sends the target inference result to the terminal. In the case where the AI RAN node performs the model inference task, the AI RAN node can send the inference result obtained by performing the model inference task as a target inference result to the access node, and the access node also sends the target inference result to the terminal.
[0092] Step S417, in the case that the terminal executes the model inference task, the terminal executes the model inference task by using the target model to obtain an inference result. Step S418, the terminal sends the obtained inference result to the access node. Step S419, the access node sends the inference result obtained by the terminal to the model management function module. Step S420, the model management function module verifies the inference result obtained by the terminal to determine the target model result. Step S421, the model management function module makes the terminal obtain the target model result. Wherein, the model management function module can verify the end-side inference result, that is, after the smart glasses complete the inference, the inference result is uploaded to the model management function module, the model management function module can compare the model result obtained by the terminal with the reference inference result obtained by the reference execution side such as the AI RAN node, to ensure the accuracy of the end-side inference and reduce the misjudgment rate. If the end-side inference result is different from the reference inference result, the target model of the smart glasses can be optimized for re-inference until the verification condition is met, or the reference inference result can be directly sent to the terminal as the target inference result.
[0093] Therefore, the smart glasses finally obtain the target inference result verified by the model management function module, and the target inference result can be superimposed on the user's field of view in an augmented reality (AR) manner. For example, in a shopping mall, the glasses recognize a certain product and display the price and discount information. On the street, the glasses recognize traffic signals and remind the user when it is safe to cross the road. In a museum, the glasses recognize exhibits and provide real-time historical introductions.
[0094] Through the scheme, the smart glasses can dynamically adapt to the best intelligent model in a low computing power and high latency environment, improve the inference efficiency and accuracy, and reduce the computing burden and power consumption of the device.
[0095] The scheme of the embodiment can solve the problems of limited end-side computing power, high network latency, and lack of verification mechanism for model inference in the related art, and achieve the following technical effects:
[0096] 1. Reduce the end-side computing pressure: through cloud distillation and AI RAN intelligent scheduling, avoid running large models directly on the end side, and realize low-power and efficient inference.
[0097] 2. Optimize network resource scheduling: use AI RAN to dynamically allocate inference tasks to avoid concentrating computing power in the cloud or on the end side, and improve overall computing efficiency.
[0098] 3. Improve inference accuracy and robustness: the end-side inference result is verified by the AI RAN to improve the inference credibility and enhance the adaptability of the model in complex wireless environments.
[0099] 4. Reduce network transmission cost: reduce the large amount of data transmission between the end side and the cloud, improve the real-time performance of inference, and be suitable for AI application scenarios with low delay and high reliability.
[0100] Specifically, compared with the traditional fixed small model deployment mode, the scheme of the embodiment can adaptively adjust the model according to different application scenarios to ensure the generalization ability and calculation efficiency of end-side inference. Compared with the traditional cloud computing mode, the scheme of the embodiment can fully utilize the calculation ability of the wireless intelligent management and arrangement function layer, reduce the calculation pressure of the end side, dynamically balance the calculation tasks between the end side, the intelligent wireless access network and the cloud, optimize the utilization rate of wireless network resources, improve the overall inference efficiency, effectively reduce the inference response time, meet the requirements of application scenarios such as smart glasses, augmented reality (AR), intelligent driving and other application scenarios with strict requirements on millisecond-level latency, and improve user experience. Compared with the mode of relying on a single model result, the scheme of the embodiment can ensure the reliability and accuracy of end-side inference through result verification of the wireless intelligent management and arrangement function layer, effectively reduce the risk of inference errors caused by limited end-side calculation ability. Compared with the fixed model mode of the end side, the scheme of the embodiment can optimize model deployment according to the terminal device situation, maximize the improvement of device endurance, reduce the demand for computing power, and at the same time guarantee the inference performance. Compared with the mode of AI computing mode relying on fixed architecture, the scheme of the embodiment can automatically adjust the calculation mode for different devices to adapt to a wider range of intelligent terminal application requirements. In summary, the scheme of the embodiment solves the deficiencies of related technologies in model adaptability, calculation efficiency, inference accuracy, low latency optimization and device endurance, and provides an efficient, reliable and low-energy AI computing solution for wireless intelligent terminals.
[0101] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0102] Based on the same inventive concept, the embodiments of the present application also provide a model inference device for implementing the model inference method described above. The implementation scheme for solving problems provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more model inference device embodiments provided below can refer to the limitations of the model inference method in the above text, which will not be repeated here.
[0103] In an exemplary embodiment, as shown in Figure 5 A model inference device is provided, which can be applied to a wireless intelligent management and orchestration function layer. The device 500 can include:
[0104] An information acquisition module 501 is configured to acquire model inference requirement information of a terminal through an access node of an intelligent wireless access network.
[0105] An execution determination module 502 is configured to determine an execution party of a model inference task according to the model inference requirement information and resource state information of the intelligent wireless access network. The model inference task is obtained according to the model inference requirement information. The execution party includes a cloud, a node, or the terminal. The node is a node in a node set of the intelligent wireless access network. The node set includes the access node.
[0106] A task processing module 503 is configured to trigger the execution party to execute the model inference task by using a target model, so that the terminal obtains a target inference result. The target model is distilled from a large model by the cloud according to the model inference requirement information. The target inference result is an inference result that meets a verification condition of the wireless intelligent management and orchestration function layer.
[0107] In an exemplary embodiment, the information acquisition module 501 is further configured to send the model inference requirement information to the cloud. The model inference requirement information is used to distill a target model from a large model by the cloud.
[0108] In an exemplary embodiment, the execution determination module 502 is configured to determine the terminal as the execution party of the model inference task if it is judged that the terminal meets a task execution condition according to the model inference requirement information. If it is judged that the terminal does not meet the task execution condition according to the model inference requirement information, a node in the node set of the intelligent wireless access network or the cloud is determined as the execution party according to the model inference requirement information and the resource state information of the intelligent wireless access network.
[0109] In an example embodiment, the execution determining module 502 is configured to determine the execution party according to a node satisfying the task execution condition, if it is determined according to the model inference requirement information and the resource state information of the intelligent wireless access network that there is a node satisfying the task execution condition in the node set of the intelligent wireless access network; and determine the cloud as the execution party, if it is determined according to the model inference requirement information and the resource state information of the intelligent wireless access network that there is no node satisfying the task execution condition in the node set of the intelligent wireless access network.
[0110] In an example embodiment, in the case that the execution party includes the cloud, the task processing module 503 is configured to trigger the cloud to execute the model inference task by using a target model; obtain a first inference result obtained by the cloud in executing the model inference task, determine the first inference result as a target inference result satisfying a verification condition, and send the target inference result to the terminal through the access node; or, in the case that the execution party includes the node, the task processing module 503 is configured to obtain a target model obtained by the cloud; send the target model to the node, so that the node executes the model inference task by using the target model to obtain a second inference result, and sends the second inference result as a target inference result satisfying a verification condition to the terminal through the access node; or, in the case that the execution party includes the terminal, the task processing module 503 is configured to obtain a target model obtained by the cloud; send the target model to the terminal through the access node, so that the terminal executes the model inference task by using the target model to obtain a third inference result and returns the third inference result through the access node; verify the third inference result to determine the target inference result, and make the terminal obtain the target inference result.
[0111] In an example embodiment, in the case that the execution party includes the terminal, the task processing module 503 is configured to obtain a reference inference result of a reference execution party; the reference execution party includes the cloud or a node in the node set of the intelligent wireless access network; the reference inference result is an inference result obtained by the reference execution party in executing the model inference task by using the target model; if it is determined according to the reference inference result and the third inference result that the third inference result satisfies the verification condition, send indication information to the terminal through the access node; the indication information is used to instruct the terminal to determine the third inference result as the target inference result; and if it is determined according to the reference inference result and the third inference result that the third inference result does not satisfy the verification condition, send the reference inference result as the target inference result to the terminal through the access node.
[0112] Each module in the model inference apparatus can be implemented by software, hardware, and combinations thereof, in whole or in part. The modules can be embedded in or independent of a processor in the network device in hardware form, or stored in a memory in the network device in software form, so as to be invoked and executed by the processor to perform operations corresponding to the modules.
[0113] In an exemplary embodiment, a network device is provided, which can serve as a wireless intelligent management and orchestration function layer. The network device can be a server or the like, and its internal structure diagram can be as shown in Figure 6 The network device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the network device is configured to provide computing and control capabilities. The memory of the network device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The input / output interface of the network device is configured to exchange information between the processor and external devices. The communication interface of the network device is configured to communicate with external devices through network connection. The computer program is executed by the processor to implement a model inference method.
[0114] Those skilled in the art can understand that Figure 6 The structure shown in the above
[0115] In an embodiment, a network device is also provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0116] In an embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0117] In an embodiment, a computer program product is provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0118] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0119] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments of each method. In the embodiments provided in the present application, any reference to memory, database or other medium can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.
[0120] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described above, however, any combination of these technical features is deemed to be within the scope of the present application.
[0121] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be pointed out that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A model inference method, comprising: The method is applied to a wireless intelligent management and orchestration function layer, and comprises the following steps: obtaining model inference requirement information of a terminal through an access node of an intelligent wireless access network; sending the model inference requirement information to a cloud end; the model inference requirement information is used for the cloud end to distill a small model adapted to the terminal from a large model as a target model; determining an execution party of a model inference task according to the model inference requirement information and resource state information of the intelligent wireless access network; the model inference task is obtained according to the model inference requirement information; the execution party comprises the cloud end, a node or the terminal; the node is a node in a node set of the intelligent wireless access network; the node set comprises the access node; wherein the execution party is determined in the order from the terminal to the node and then to the cloud end; triggering the execution party to execute the model inference task by using the target model, so that the terminal obtains a target inference result; the target inference result is an inference result meeting a verification condition of the wireless intelligent management and orchestration function layer; in the case where the execution party comprises the terminal, triggering the execution party to execute the model inference task by using the target model, so that the terminal obtains the target inference result, comprising: obtaining the target model obtained by the cloud end; sending the target model to the terminal through the access node, so that the terminal executes the model inference task by using the target model and returns a third inference result through the access node; verifying the third inference result to determine the target inference result, so that the terminal obtains the target inference result.
2. The method of claim 1, wherein, the determining of the execution party of the model inference task according to the model inference requirement information and the resource state information of the intelligent wireless access network comprises: if it is determined according to the model inference requirement information that the terminal meets a task execution condition, determining the terminal as the execution party of the model inference task; if it is determined according to the model inference requirement information that the terminal does not meet the task execution condition, determining a node in a node set of the intelligent wireless access network or the cloud end as the execution party according to the model inference requirement information and the resource state information of the intelligent wireless access network.
3. The method of claim 2, wherein, the determining of the node in the node set of the intelligent wireless access network or the cloud end as the execution party according to the model inference requirement information and the resource state information of the intelligent wireless access network comprises: if it is determined according to the model inference requirement information and the resource state information of the intelligent wireless access network that there is a node meeting the task execution condition in the node set of the intelligent wireless access network, determining the node meeting the task execution condition as the execution party; if it is determined according to the model inference requirement information and the resource state information of the intelligent wireless access network that there is no node meeting the task execution condition in the node set of the intelligent wireless access network, determining the cloud end as the execution party.
4. The method according to any one of claims 1 to 3, characterized in that, In a case where the execution party includes the cloud, triggering the execution party to perform the model inference task by using the target model to enable the terminal to obtain the target inference result includes: triggering the cloud to perform the model inference task by using the target model; obtaining a first inference result obtained by the cloud performing the model inference task, determining the first inference result as the target inference result satisfying the verification condition, and sending the target inference result to the terminal through the access node; Or, In a case where the execution party includes the node, triggering the execution party to perform the model inference task by using the target model to enable the terminal to obtain the target inference result includes: obtaining the target model obtained by the cloud; sending the target model to the node to enable the node to perform the model inference task by using the target model to obtain a second inference result, and sending the second inference result as the target inference result satisfying the verification condition to the terminal through the access node.
5. The method of claim 4, wherein, in a case where the execution party includes the terminal, verifying the third inference result to determine the target inference result to enable the terminal to obtain the target inference result includes: obtaining a reference inference result of a reference execution party; the reference execution party includes the cloud or a node in a node set of the intelligent wireless access network; the reference inference result is an inference result obtained by the reference execution party performing the model inference task by using the target model; if it is determined according to the reference inference result and the third inference result that the third inference result satisfies the verification condition, sending indication information to the terminal through the access node; the indication information is used to instruct the terminal to determine the third inference result as the target inference result; if it is determined according to the reference inference result and the third inference result that the third inference result does not satisfy the verification condition, determining the reference inference result as the target inference result and sending the target inference result to the terminal through the access node. The apparatus is applied to a wireless intelligent management and orchestration function layer, and includes:
6. A model inference apparatus characterized by comprising: an information obtaining module, configured to obtain model inference requirement information of a terminal through an access node of an intelligent wireless access network, and send the model inference requirement information to a cloud; the model inference requirement information is used to enable the cloud to distill a small model from a large model to obtain a lightweight small model adapted to the terminal as a target model; an execution determining module, configured to determine an execution party of a model inference task according to the model inference requirement information and resource state information of the intelligent wireless access network; the model inference task is obtained according to the model inference requirement information; the execution party includes the cloud, a node, or the terminal; the node is a node in a node set of the intelligent wireless access network; the node set includes the access node; and the execution party is determined in an order from the terminal to the node and then to the cloud. The task processing module is configured to trigger the execution party to execute a model inference task by using a target model, so that the terminal obtains a target inference result; wherein the target model is distilled from a large model by the cloud according to the model inference requirement information; the target inference result is an inference result that meets the verification condition of the wireless intelligent management and orchestration function layer; in the case where the execution party includes a terminal, triggering the execution party to execute a model inference task by using a target model, so that the terminal obtains a target inference result, includes: obtaining the target model obtained by the cloud; sending the target model to the terminal through an access node, so that the terminal executes a model inference task by using the target model and returns a third inference result through the access node; verifying the third inference result to determine the target inference result, so that the terminal obtains the target inference result. 7.A network device, comprising a memory and a processor, wherein the memory stores a computer program, and the network device is configured to perform the method according to any one of claims 1-6. The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 5.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 5.
9. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 5. The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Block chain-based edge intelligent system resource management method
CN119893588A