Model reasoning method and device, network equipment, readable storage medium and program product
Through the dynamic adaptation of the executor of wireless intelligent management and orchestration functional layer, the cloud-based distillation lightweight model and inference are used to perform inference at the terminal, the problem of limited computing power of terminal devices is solved and efficient and accurate model reasoning is achieved.
Patent Information
- Application Number
- CN202511030948.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Limitations in computing power, storage space and energy consumption of terminal devices lead to challenges in directly running large-scale artificial intelligence models, and cloud computing relies on cloud models for inference to limit efficiency.
The terminal model inference requirements information is obtained through the wireless intelligent management and orchestration functional layer, and dynamically adapt the executor (cloud, node or terminal), and the terminal performs local inference using cloud distillation and lightweight model, and ensures the accuracy of the results through intelligent scheduling and verification.
Reduce dependence on cloud computing resources, reduce network transmission overhead, improve model inference efficiency and accuracy, adapt to terminal needs, and reduce computing pressure.
Smart Images

Figure CN120529342A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of wireless communications and artificial intelligence technology, and in particular to a model reasoning method, apparatus, network equipment, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, the demand for model processing on end-devices (terminals) is increasing. However, due to limitations in computing power, storage space, and energy consumption, running large-scale AI models directly on terminals often faces significant challenges.
[0003] In related technologies, cloud computing can provide powerful computing power support for the processing needs of terminal models. However, it mainly relies on large models in the cloud for reasoning, which leads to the problem of limited model reasoning efficiency. Summary of the Invention
[0004] Based on this, it is necessary to provide a model reasoning method, apparatus, network device, computer-readable storage medium and computer program product to address the above technical problems.
[0005] In a first aspect, the present application provides a model reasoning method applied to the wireless intelligent management and orchestration function layer, including:
[0006] Obtaining model reasoning requirement information of the terminal through the access node of the intelligent wireless access network;
[0007] Determining an executor of a model reasoning task according to the model reasoning requirement information and the resource status information of the intelligent wireless access network; the model reasoning task is obtained according to the model reasoning requirement information; the executor includes a cloud, a node, or the terminal; the node is a node in a node set of the intelligent wireless access network; the node set includes the access node;
[0008] Trigger the executor to use the target model to perform the model reasoning task, so that the terminal obtains the target reasoning result; wherein, the target model is distilled from the large model by the cloud according to the model reasoning requirement information; the target reasoning result is the reasoning result that meets the verification conditions of the wireless intelligent management and orchestration function layer.
[0009] In a second aspect, the present application also provides a model reasoning device, which is applied to the wireless intelligent management and orchestration function layer, including:
[0010] An information acquisition module is used to obtain the model reasoning requirement information of the terminal through the access node of the intelligent wireless access network;
[0011] An execution determination module is configured to determine an executor of a model reasoning task based on the model reasoning requirement information and the resource status information of the intelligent wireless access network; the model reasoning task is obtained based on the model reasoning requirement information; the executor includes a cloud, a node, or the terminal; the node is a node in a node set of the intelligent wireless access network; the node set includes the access node;
[0012] A task processing module is used to trigger the executor to use the target model to perform the model reasoning task, so that the terminal obtains the target reasoning result; wherein, the target model is distilled from the large model by the cloud according to the model reasoning requirement information; the target reasoning result is an inference result that meets the verification conditions of the wireless intelligent management and orchestration function layer.
[0013] In a third aspect, the present application further provides a network device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0014] In a fourth aspect, the present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.
[0015] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.
[0016] The above-mentioned model reasoning method, apparatus, network equipment, computer-readable storage medium and computer program product, the wireless intelligent management and orchestration function layer can obtain the model reasoning requirement information of the terminal through the access node of the intelligent wireless access network, and determine the executor of the model reasoning task based on the model reasoning requirement information and the resource status information of the intelligent wireless access network. The executor may include the cloud, node or terminal, wherein the node is a node in the node set of the intelligent wireless access network, and the node set includes the access node; the wireless intelligent management and orchestration function layer can trigger the executor to use the target model to perform the model reasoning task, so that the terminal obtains the target reasoning result, wherein the target model is obtained by distillation from the large model by the cloud according to the model reasoning requirement information, and the target reasoning result is the reasoning result that meets the verification conditions of the wireless intelligent management and orchestration function layer. This solution allows the wireless intelligent management and orchestration functional layer to flexibly adapt the executor for the model reasoning task based on the model reasoning demand information and the resource status information of the intelligent wireless access network. The executor may include the cloud, node or terminal, and there is no need to rely on the cloud to perform model reasoning tasks. This can reduce dependence on cloud computing resources, reduce large amounts of data transmission between the terminal and the cloud, reduce network transmission overhead, and avoid the concentration of computing power on the terminal side, thereby improving the efficiency of model reasoning as a whole. The target model is distilled from the large model by the cloud, so that it can adapt to the model reasoning requirements of the terminal and reduce the computational pressure of model reasoning. The reasoning results obtained by the terminal must meet the verification conditions of the wireless intelligent management and orchestration functional layer to ensure the accuracy of model reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 An application environment diagram of a model reasoning method in one embodiment;
[0019] Figure 2 A schematic diagram of a wireless intelligent management and orchestration function layer in one embodiment;
[0020] Figure 3 1 is a flow chart of a model reasoning method in one embodiment;
[0021] Figure 4 is a timing diagram of a model reasoning method in one embodiment;
[0022] Figure 5 is a structural block diagram of a model reasoning device in one embodiment;
[0023] Figure 6 FIG. 4 is a diagram showing the internal structure of a network device in one embodiment. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0025] The model reasoning method provided in the embodiment of the present application can be applied to Figure 1 The application environment shown in the figure can include: terminals (on the device side), nodes of the artificial intelligence radio access network (AI RAN), the wireless intelligent management and orchestration function layer (RAN AI layer), and the cloud.
[0026] Terminals include, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices include smart speakers, smart TVs, smart air conditioners, smart car devices, and projectors. Portable wearable devices include head-mounted devices, which can be virtual reality (VR) devices or augmented reality (AR) devices.
[0027] With the development of intelligent radio access networks (RANs), intelligent RANs are introducing AI computing power into wireless networks, providing a lower-latency, intelligent computing collaboration solution for the end-user. Intelligent RANs can apply artificial intelligence and machine learning technologies to all aspects of RAN planning, deployment, management, and optimization, improving network performance, resource utilization, and enhancing network flexibility and intelligence to better meet user demands for high-speed, stable, and low-latency network connections.
[0028] Among them, the wireless intelligent management and orchestration function layer can perform intelligent and automated management and scheduling of intelligent wireless access networks, improve network performance, optimize resource utilization, enhance business flexibility and improve user experience. Figure 2As shown, the wireless intelligent management and orchestration functional layer can include a model management functional module and a general computing resource scheduling module. The model management functional module can be a functional module in the wireless intelligent management and orchestration functional layer for managing artificial intelligence models (AI models), and can be used to distribute target models obtained in the cloud. The general computing resource scheduling module can be a component in the wireless intelligent management and orchestration functional layer that implements efficient collaborative management and optimized allocation of computing, communication, and other resources in the wireless communication network. The wireless intelligent management and orchestration functional layer can be connected to various intelligent wireless access networks, including but not limited to 3GPP RAN (radio access network in the Third Generation Partnership Project 3GPP), cloud native RAN (a network architecture that combines cloud native technology with radio access network RAN), or other RANs.
[0029] like Figure 1 In the illustrated application environment, a terminal can access a node of a smart wireless access network. The node of the smart wireless access network accessed by the terminal is referred to as an access node. The smart wireless network may include multiple nodes, which may form a node set of the smart wireless access network (which may include nodes 1 through n). This node set may include access nodes. The nodes in the node set of the smart wireless network may be controlled and managed by the wireless intelligent management and orchestration function layer.
[0030] In related technologies, cloud computing can provide powerful computing power support for the processing needs of terminal models. However, it mainly relies on large models in the cloud for reasoning, which leads to the problem of limited model reasoning efficiency.
[0031] In this regard, the model reasoning method provided in the embodiment of the present application can be flexibly adapted by the wireless intelligent management and orchestration functional layer for the model reasoning task based on the model reasoning demand information and the resource status information of the intelligent wireless access network. The executor may include the cloud, node or terminal, and there is no need to rely on the cloud to perform the model reasoning task. It can reduce the dependence on cloud computing resources, reduce the large amount of data transmission between the terminal side and the cloud, reduce network transmission overhead, and avoid the concentration of computing power on the terminal side, thereby improving the efficiency of model reasoning as a whole. The target model is distilled from the large model by the cloud, so that it can adapt to the model reasoning requirements of the terminal and reduce the computational pressure of model reasoning. The reasoning results obtained by the terminal must meet the verification conditions of the wireless intelligent management and orchestration functional layer, thereby ensuring the accuracy of model reasoning.
[0032] In an exemplary embodiment, Figure 3 As shown, a model inference method is provided, which can be applied to Figure 1 The wireless intelligent management and orchestration function layer in the embodiment of the present invention may include the following steps:
[0033] Step S301: obtaining model reasoning requirement information of a terminal through an access node of an intelligent wireless access network.
[0034] In this step, the terminal can connect to an access node of the intelligent wireless access network and send a model inference request to the wireless intelligent management and orchestration function layer through the access node. This model inference request can include model inference requirement information. The wireless intelligent management and orchestration function layer can receive the terminal's model inference request through the access node of the intelligent wireless access network and obtain the model inference requirement information based on the model inference request. The model inference requirement information is the requirements provided by the terminal for model inference. This model inference requirement information may include the terminal's inference resource information (such as the terminal's own computing power, power, storage space, etc.), the terminal's required inference target (such as the need to identify a specific object), the terminal's source data for model inference (such as image data, text data, etc.), the terminal's required inference latency (such as how long it takes to obtain an inference result), the terminal's required computing power resources, and so on.
[0035] Step S302: determining an executor of the model reasoning task according to the model reasoning requirement information and the resource status information of the intelligent wireless access network.
[0036] The resource status information of the intelligent wireless access network refers to the status information of the resources of the intelligent wireless access network that can be used to perform model inference, and may include the status information of network resources and computing resources. The status information of network resources may indicate the transmission and load status of the intelligent wireless access network, and the status information of computing resources may indicate the computing power of each node of the intelligent wireless access network.
[0037] In this step, the wireless intelligent management and orchestration functional layer can determine the executor of the model inference task based on the model inference requirement information and the resource status information of the intelligent wireless access network. The model inference task can be obtained by the wireless intelligent management and orchestration functional layer based on the model inference requirement information. For example, the model inference task can be generated based on the inference target required by the terminal and the source data provided by the terminal for model inference. The executor of the model inference task can include a node or terminal in the cloud or the intelligent wireless access network. The node can be a node in a node set of the intelligent wireless access network, and the node set includes the aforementioned access node, i.e., the executor can be an access node. Thus, the wireless intelligent management and orchestration functional layer dynamically adapts the node or terminal in the cloud or the intelligent wireless access network as the executor of the model inference task based on the model inference requirement information and the resource status information of the intelligent wireless access network. For example, the wireless intelligent management and orchestration functional layer can determine the node or terminal in the cloud or the intelligent wireless access network as the executor of the model inference task based on the inference resource information of the terminal, the inference latency required by the terminal, the computing power resources required by the terminal, and the network resources and computing power resource status information of the intelligent wireless access network.
[0038] Step S303: triggering the execution party to use the target model to perform the model reasoning task, so that the terminal obtains the target reasoning result.
[0039] The target model is distilled from the large model by the cloud based on model inference requirement information. The target model adapted for the terminal can be obtained by distilling the large model based on the model inference requirement information. For example, the cloud can distill the large model based on the terminal's inference resource information, the terminal's required inference target, and the source data provided by the terminal for model inference to obtain a lightweight, adapted small model as the target model. In the relevant field, the terms large and small can be defined based on model scale and number of parameters. For example, large models generally have a massive number of parameters (e.g., billions or even trillions), while small models have relatively fewer parameters (ranging from tens of thousands to tens of millions). Large models generally have complex neural network architectures (including multiple hidden layers and a large number of neurons, resulting in large model file sizes, potentially reaching tens of gigabytes or even larger). Small models have relatively simple architectures (fewer layers and neurons, resulting in smaller model files, typically ranging from a few megabytes to tens of megabytes).
[0040] In this step, after determining the executor, the wireless intelligent management and orchestration function layer can assign the model inference task to the executor, enable the executor to obtain the target model obtained from the cloud, trigger the executor to perform the model inference task using the target model, and enable the terminal to obtain the target inference result. The target inference result is the inference result that meets the verification conditions of the wireless intelligent management and orchestration function layer. Among them, after the executor uses the target model to perform the model reasoning task, a corresponding reasoning result can be obtained. The wireless intelligent management and orchestration function layer can indicate whether the reasoning result meets the verification condition. The verification condition can be set by the user. As an implementation method, the verification condition may include whether the reasoning result meets the reasoning target required by the terminal, etc. For example, if the required reasoning target is to identify animals in an image, the verification condition may include the name of the animal pointed to by the reasoning result, etc.; for example, if the required reasoning target is to generate images of objects such as animals and people, the verification condition may include whether the image quality evaluation value of the generated image meets the threshold condition. The image quality evaluation value of the generated image can be obtained by an image quality evaluation model deployed in the wireless intelligent management and orchestration function layer. The generated image can be input into the image quality evaluation model (which can use an artificial intelligence model such as a convolutional neural network) to obtain the image quality evaluation value. If the image quality evaluation value is greater than or equal to a set threshold (which can be set by the user), the image quality evaluation value of the generated image meets the threshold condition. If the image quality evaluation value is less than the set threshold, the image quality evaluation value of the generated image does not meet the threshold condition. Regarding the method for the terminal to obtain the target reasoning result, as an implementation method, if the reasoning result meets the verification conditions, the reasoning result can be delivered to the terminal as the target reasoning result; if the reasoning result does not meet the verification conditions, the target reasoning result can be obtained by optimizing the target model and then re-reasoning, etc. and delivered to the terminal.
[0041] In the model inference method of this embodiment, the wireless intelligent management and orchestration function layer can obtain the model inference requirement information of the terminal through the access node of the intelligent wireless access network, and determine the executor of the model inference task based on the model inference requirement information and the resource status information of the intelligent wireless access network. The executor may include the cloud, a node or a terminal, wherein the node is a node in the node set of the intelligent wireless access network, and the node set includes the access node; the wireless intelligent management and orchestration function layer can trigger the executor to use the target model to perform the model inference task, so that the terminal obtains the target inference result, wherein the target model is distilled from the large model by the cloud according to the model inference requirement information, and the target inference result is the inference result that meets the verification conditions of the wireless intelligent management and orchestration function layer. This solution allows the wireless intelligent management and orchestration functional layer to flexibly adapt the executor for the model reasoning task based on the model reasoning demand information and the resource status information of the intelligent wireless access network. The executor may include the cloud, node or terminal, and there is no need to rely on the cloud to perform model reasoning tasks. This can reduce dependence on cloud computing resources, reduce large amounts of data transmission between the terminal and the cloud, reduce network transmission overhead, and avoid the concentration of computing power on the terminal side, thereby improving the efficiency of model reasoning as a whole. The target model is distilled from the large model by the cloud, so that it can adapt to the model reasoning requirements of the terminal and reduce the computational pressure of model reasoning. The reasoning results obtained by the terminal must meet the verification conditions of the wireless intelligent management and orchestration functional layer to ensure the accuracy of model reasoning.
[0042] In an exemplary embodiment, after obtaining the model reasoning requirement information of the terminal through the access node of the intelligent wireless access network in step S301, the above method may further include the following steps:
[0043] Send the model inference requirement information to the cloud. The model inference requirement information is used by the cloud to distill the target model from the large model.
[0044] In this embodiment, after obtaining the terminal's model inference requirement information, the wireless intelligent management and orchestration functional layer can send this information to the cloud. The cloud can then distill the large model based on the model inference requirement information to obtain a target model that is compatible with the terminal. For example, the cloud can distill the large model based on the terminal's inference resource information (computing power, power, storage space, etc.), the terminal's required inference goal (such as identifying products or traffic signs), and the source data provided by the terminal for model inference (such as image data or text data), to obtain a corresponding lightweight small model as the target model (such as a product recognition model or a traffic sign recognition model). In this embodiment, the cloud can distill the target model from the large model, allowing the wireless intelligent management and orchestration functional layer to flexibly allocate model inference tasks to the cloud, nodes in the intelligent wireless access network, or terminals for execution.
[0045] In an exemplary embodiment, determining the executor of the model reasoning task according to the model reasoning requirement information and the resource status information of the intelligent wireless access network in step S302 may include:
[0046] If the terminal is judged to meet the task execution conditions according to the model reasoning requirement information, the terminal is determined to be the executor of the model reasoning task; if the terminal is judged to not meet the task execution conditions according to the model reasoning requirement information, the node in the node set of the intelligent wireless access network or the cloud is determined to be the executor according to the model reasoning requirement information and the resource status information of the intelligent wireless access network.
[0047] In this embodiment, the wireless intelligent management and orchestration functional layer may first determine whether the terminal meets the task execution conditions based on the model inference requirement information. The task execution conditions may include conditions regarding model inference latency and computing resources. The task execution conditions can be obtained based on the model inference requirement information (e.g., the terminal's required inference latency and computing resources). For example, the wireless intelligent management and orchestration functional layer may determine whether the terminal meets the task execution conditions based on the terminal's own computing power. If the terminal is determined to meet the task execution conditions, the wireless intelligent management and orchestration functional layer may determine the terminal as the executor of the model inference task, and have the terminal perform the model inference task using the target model to reduce network transmission latency. If the terminal is determined not to meet the task execution conditions, the wireless intelligent management and orchestration functional layer may further determine a node in the node set of the intelligent wireless access network or the cloud as the executor based on the model inference requirement information and resource status information of the intelligent wireless access network.
[0048] In an exemplary embodiment, further, the above-mentioned determining a node in the node set of the smart wireless access network or the cloud as the executor based on the model reasoning demand information and the resource status information of the smart wireless access network may include:
[0049] If it is determined based on the model reasoning requirement information and the resource status information of the intelligent wireless access network that there are nodes in the node set of the intelligent wireless access network that meet the task execution conditions, the executor is determined based on the nodes that meet the task execution conditions; if it is determined based on the model reasoning requirement information and the resource status information of the intelligent wireless access network that there are no nodes in the node set of the intelligent wireless access network that meet the task execution conditions, the cloud is determined as the executor.
[0050] In this embodiment, if the terminal is determined to not meet the task execution conditions, the wireless intelligent management and orchestration function layer may further determine a node in the node set of the intelligent wireless access network or the cloud as the executor based on the model inference requirement information and the resource status information of the intelligent wireless access network. The wireless intelligent management and orchestration function layer may determine whether a node in the node set of the intelligent wireless access network meets the task execution conditions based on the model inference requirement information and the resource status information of the intelligent wireless access network. For example, the wireless intelligent management and orchestration function layer may determine whether a node in the node set of the intelligent wireless access network meets the task execution conditions (model inference latency and computing resource conditions) based on the inference latency required by the terminal, the computing power resources required by the terminal, network resource status information (which may indicate the transmission and load status of the intelligent wireless access network), and computing power resource status information (which may indicate the computing power of each node in the intelligent wireless access network). If a node meets the task execution conditions, then if there is only one node, then that node may be determined as the executor. If there are multiple nodes, then the executor may be determined based on the distance between the node and the terminal. The node closest to the terminal may be selected as the executor to improve model inference efficiency. If there is no node that meets the task execution conditions, the wireless intelligent management and orchestration function layer can determine the cloud as the executor to complete the model reasoning task.
[0051] In one exemplary embodiment, when the execution party includes a cloud, the triggering execution party in step S303 uses the target model to perform the model reasoning task so that the terminal obtains the target reasoning result, which may include:
[0052] Trigger the cloud to perform a model reasoning task using the target model; obtain a first reasoning result obtained by executing the model reasoning task on the cloud, determine the first reasoning result as a target reasoning result that meets the verification conditions, and send the target reasoning result to the terminal through the access node.
[0053] In this embodiment, when the executor includes the cloud, since the cloud has already distilled the target model from the large model based on the model inference requirement information, the wireless intelligent management and orchestration functional layer can assign the model inference task to the cloud, triggering the cloud to execute the model inference task using the target model to obtain a first inference result. The cloud then sends the first inference result to the wireless intelligent management and orchestration functional layer. The wireless intelligent management and orchestration functional layer can then determine whether the first inference result meets a verification condition based on the executor of the model inference task. That is, the verification condition can include whether the executor of the model inference task is the target executor, where the cloud can be set as the target executor. Thus, the wireless intelligent management and orchestration functional layer can determine the first inference result obtained by the cloud executing the model inference task as the target inference result that meets the verification condition, and send the target inference result to the terminal via the access node. Thus, even if neither the terminal nor the node of the intelligent wireless access network meets the task execution condition, the target inference result obtained by the cloud can be delivered to the terminal.
[0054] In another exemplary embodiment, when the executor includes a node, the triggering executor in step S303 uses the target model to perform the model reasoning task so that the terminal obtains the target reasoning result, which may include:
[0055] Obtain the target model obtained in the cloud; send the target model to the node so that the node uses the target model to perform the model reasoning task to obtain a second reasoning result, and send the second reasoning result as the target reasoning result that meets the verification conditions to the terminal through the access node.
[0056] In this embodiment, when the executor includes a node of the intelligent wireless access network, the wireless intelligent management and orchestration function layer can obtain a target model distilled from a large model based on model inference requirement information in the cloud, send the target model to the node, and assign a model inference task to the node of the intelligent wireless access network to trigger the node to execute the model inference task using the target model to obtain a second inference result. The wireless intelligent management and orchestration function layer can also instruct the node to determine the second inference result obtained as a target inference result that meets a verification condition, and instruct the node to send the target inference result to the terminal via the access node. The wireless intelligent management and orchestration function layer can indicate whether the second inference result meets the verification condition based on the executor of the model inference task, that is, the verification condition can include whether the executor of the model inference task is the target executor. The node of the intelligent wireless access network can be set as the target executor. Therefore, when assigning the model inference task to the node of the intelligent wireless access network, the wireless intelligent management and orchestration function layer can instruct the node to determine the second inference result obtained as the target inference result that meets the verification condition and instruct the node to send the target inference result to the terminal via the access node. Therefore, when the terminal does not meet the task execution conditions but the nodes of the intelligent wireless access network do, the nodes of the intelligent wireless access network send the target inference results to the terminal through the access node, avoiding the concentration of computing power on the cloud or terminal side and improving the overall inference efficiency.
[0057] In another exemplary embodiment, when the executor includes a terminal, the triggering executor of step S303 uses the target model to perform the model reasoning task so that the terminal obtains the target reasoning result, which may include: obtaining the target model obtained in the cloud; sending the target model to the terminal through the access node so that the terminal uses the target model to perform the model reasoning task to obtain and return the third reasoning result through the access node; verifying the third reasoning result to determine the target reasoning result, so that the terminal obtains the target reasoning result.
[0058] In this embodiment, when the executor includes a terminal, the wireless intelligent management and orchestration functional layer can obtain the target model distilled from the large model based on the model reasoning requirement information in the cloud, send the target model to the terminal through the access node, and assign the model reasoning task to the terminal to trigger the terminal to use the target model to execute the model reasoning task to obtain a third reasoning result, and return the third reasoning result to the wireless intelligent management and orchestration functional layer through the access node. In which, when the executor includes a terminal, after the wireless intelligent management and orchestration function layer obtains the target model from the cloud, before sending the target model to the terminal through the access node, it can first determine whether the terminal's reasoning resource information has changed. If there is no change, the wireless intelligent management and orchestration function layer can send the target model to the terminal through the access node. If there is a change, and the changed reasoning resource information indicates that the target model is applicable to the terminal (such as the terminal's own computing power, power, storage space increases, etc.), the wireless intelligent management and orchestration function layer can send the target model to the terminal through the access node. If there is a change, and the changed reasoning resource information indicates that the target model is not applicable to the terminal (such as the terminal's own computing power, power, storage space decreases, etc.), the wireless intelligent management and orchestration function layer can send the changed reasoning resource information to the cloud, so that the cloud can perform secondary distillation on the target model according to the changed reasoning resource information and return it to the wireless intelligent management and orchestration function layer. The wireless intelligent management and orchestration function layer can send the secondary distilled target model to the terminal through the access node for it to execute the model reasoning task to obtain a third reasoning result. The wireless intelligent management and orchestration functional layer verifies the third inference result to determine a target inference result that meets verification conditions. The wireless intelligent management and orchestration functional layer then enables the terminal to obtain the target inference result. As an implementation method, the verification conditions may include whether the inference result meets the inference target required by the terminal. For example, if the required inference target is to identify an animal in an image, the verification conditions may include whether the inference result indicates the name of the animal. As an implementation method, if the third inference result meets the verification conditions, the wireless intelligent management and orchestration functional layer may instruct the terminal to use the third inference result as the target inference result, thereby obtaining the target inference result. If the third inference result does not meet the verification conditions, the wireless intelligent management and orchestration functional layer may optimize the target model via the cloud (e.g., by increasing the number of network layers in the model), and send the optimized target model to the terminal via the access node for re-inference. The re-inferenced third inference result may also be re-verified using the aforementioned method until the third inference result meets the verification conditions. Thus, if the terminal meets the task execution conditions, the terminal executes the model inference task, and the wireless intelligent management and orchestration functional layer verifies the inference result, thereby improving the inference credibility and enhancing the model's adaptability in complex wireless environments.
[0059] In an exemplary embodiment, further, when the execution party includes a terminal, the above-mentioned verification of the third reasoning result to determine the target reasoning result so that the terminal obtains the target reasoning result may include:
[0060] Obtain the reference reasoning result of the reference executor; if it is judged that the third reasoning result meets the verification conditions based on the reference reasoning result and the third reasoning result, send an indication message to the terminal through the access node; if it is judged that the third reasoning result does not meet the verification conditions based on the reference reasoning result and the third reasoning result, use the reference reasoning result as the target reasoning result and send it to the terminal through the access node.
[0061] The reference executor may include a node in a node set of the cloud or intelligent wireless access network; the reference inference result is the inference result obtained by the reference executor using the target model to perform the model inference task. In this embodiment, when the executor includes a terminal, the wireless intelligent management and orchestration function layer can obtain the reference inference result of the reference executor. The wireless intelligent management and orchestration function layer can select a node in the node set of the cloud or intelligent wireless access network as the reference executor. For example, the wireless intelligent management and orchestration function layer can select a node in the node set of the cloud or intelligent wireless access network as the reference executor, assign the model inference task to the node, and trigger the reference executor to perform the model inference task using the target model to obtain and return the reference inference result. The wireless intelligent management and orchestration function layer can then compare the reference inference result with the third inference result to determine whether the third inference result meets a verification condition. The verification condition can include that the third inference result is identical to the reference inference result. If the third inference result meets the verification condition, the wireless intelligent management and orchestration function layer can send an indication to the terminal via the access node, instructing the terminal to determine the third inference result as the target inference result, thereby allowing the terminal to obtain the target inference result. If the third inference result does not meet the verification conditions, the intelligent management and orchestration function layer can use the reference inference result as the target inference result and send the target inference result to the terminal through the access node, so that the terminal can obtain the target inference result. The solution of this embodiment can ensure the accuracy of terminal-side inference and reduce the error rate.
[0062] In an exemplary embodiment, in combination Figure 4 The model reasoning method of this application is described.
[0063] Related technologies primarily rely on large cloud-based models for inference. Even when lightweight models are deployed directly on the edge (terminal), they lack dynamic adaptability, making it difficult to optimize and adjust based on edge computing power, network conditions, and task requirements. This is especially true in different network environments and computing power conditions, making it difficult to flexibly adjust model deployment and execution strategies, limiting inference efficiency and accuracy. Furthermore, edge-based inference results often lack verification mechanisms, which can affect inference accuracy due to factors such as environmental changes and model drift, resulting in a lack of reliable model quality assurance.
[0064] To this end, the solution of this embodiment enables model collaborative inference based on a cloud-intelligent radio access network (AI RAN)-end architecture. Large cloud models are dynamically distilled based on end-side requirements to generate lightweight smaller models. The RAN AI layer performs intelligent scheduling and inference optimization. The RAN AI layer not only rationally allocates inference tasks based on current network status and computing power, but also directly deploys models based on end-side computing power and storage capabilities, and performs secondary distillation optimization. Furthermore, end-side inference results can be verified by the RAN AI layer to improve inference accuracy and reliability. Thus, the solution of this embodiment combines the global optimization capabilities of the cloud, the intelligent scheduling capabilities of the RAN AI layer, and the low-power local inference capabilities of the end-side to implement an efficient and dynamically adaptive intelligent model collaborative optimization strategy.
[0065] In general, the solution of this embodiment may include the following steps:
[0066] 1. Cloud-based intelligent model distillation: Based on the specific needs of the client, the cloud-based large model can first be intelligently distilled to generate a lightweight small model adapted to the client. The model structure can be optimized to adapt to the computing power and storage capacity of the client.
[0067] 2. Model Scheduling and Distribution: The distilled small models can be pushed to the wireless intelligent management and orchestration function layer. The wireless intelligent management and orchestration function layer can intelligently allocate models based on information such as the current network access status, node computing power, and device capabilities. The optimal node can be selected to complete the model inference task, and the model inference task can also be completed in the cloud or on the device.
[0068] 3. Intelligent reasoning on the device: The device can run lightweight models locally for reasoning, achieving low-latency and efficient reasoning calculations, reducing dependence on cloud computing resources, and reducing network transmission overhead.
[0069] 4. Inference result verification and optimization: The inference results on the end side are uploaded to the wireless intelligent management and orchestration functional layer. The wireless intelligent management and orchestration functional layer can verify the accuracy by combining historical data or comparing the inference results of nodes in the intelligent wireless access network to ensure the reliability of the inference results on the end side. It can also dynamically optimize the parameters of the lightweight small model based on the verification results to improve the model generalization capability and inference accuracy.
[0070] Therefore, the solution of this embodiment uses the wireless intelligent management and orchestration function layer as the intelligent scheduling and reasoning verification center to ensure efficient deployment of the model, low-latency reasoning and accuracy of the reasoning results.
[0071] The solution of this embodiment involves multiple network devices, including the cloud, wireless intelligent management and orchestration functional layer, nodes and terminals of the intelligent wireless access network. Through the interaction of these network devices, dynamic distillation, distribution, reasoning and verification of intelligent models are achieved, thereby optimizing the reasoning efficiency and accuracy on the terminal side and improving the intelligent computing capabilities of the entire network.
[0072] In the above-described architecture of this embodiment, the cloud can undertake the training and distillation of intelligent models. Based on the large model and information such as the computing power, storage space, and business requirements of the terminal side, the cloud can generate lightweight small models. The cloud can also push the distilled model to the wireless intelligent management and orchestration functional layer. The cloud is not only a platform for AI computing, but also collaborates with the wireless intelligent management and orchestration functional layer to provide adaptive intelligent models for different terminals.
[0073] The wireless intelligent management and orchestration layer can be deployed in edge data centers, enabling the scheduling and control of multiple nodes within the intelligent wireless access network. The layer receives model push notifications from the cloud and intelligently determines the model's inference execution path based on information such as the current intelligent wireless access network load, node computing resource distribution, and device-side computing capabilities. When device-side computing capabilities are strong, the layer can deploy small models directly to the device, allowing the device to perform inference locally. When device-side computing capabilities are limited, the layer can perform inference calculations within the intelligent wireless access network nodes and then send the results to the device. Furthermore, the layer provides on-device inference result verification. After the device performs inference, the results are returned to the layer for secondary verification to improve inference accuracy and robustness.
[0074] In terms of network equipment, the architecture described above in this embodiment may also involve radio access network (RAN) equipment, such as eNBs (evolved base stations) / gNBs (new generation base stations), remote radio units (RRUs), distributed units (DUs) / centralized units (CUs). These devices provide efficient, low-latency intelligent computing services to the end-users via 5G / 6G wireless access technologies. The RAN equipment not only connects to the end-users but also serves as computing and communication nodes for the wireless intelligent management and orchestration layer. It supports dynamic scheduling of model inference tasks and optimizes data transmission paths using wireless channel awareness to reduce network latency and computing overhead.
[0075] In this embodiment, edge devices can include smartphones, AR / VR devices, robots, autonomous driving systems, and industrial IoT terminals. These edge devices can dynamically receive lightweight models assigned by the wireless intelligent management and orchestration layer based on their computing power, storage space, and business needs, and perform local inference calculations. When edge computing power is limited, the wireless intelligent management and orchestration layer can assign collaborative inference and make business decisions after verifying the inference results. Furthermore, data exchange between the edge and the wireless intelligent management and orchestration layer can be carried out over 5G / 6G networks, ensuring the real-time and stable performance of intelligent inference.
[0076] In terms of network communication, the cloud communicates with the wireless intelligent management and orchestration layer through the core network (5GC) or O-RAN (Open Radio Access Network) architecture's Service Management and Orchestration (SMO) to push the distilled model. The wireless intelligent management and orchestration layer obtains real-time resource status information from the intelligent radio access network to optimize the model distribution strategy. The wireless intelligent management and orchestration layer determines the optimal model inference execution path and performs local inference or collaborative computing when necessary. The device side directly interacts with the intelligent radio access network (which has both wireless access and intelligent service processing capabilities), receives the lightweight model, performs inference, and uploads the inference results. The wireless intelligent management and orchestration layer performs secondary verification to ensure the accuracy of the inference results.
[0077] In terms of software architecture, this embodiment may include components such as a cloud-based AI computing framework, intelligent management and orchestration functionality, and an on-device intelligent model engine. The cloud-based AI computing framework is responsible for training and distilling large models and interacts with the wireless intelligent management and orchestration layer through an API (Application Programming Interface). The intelligent management and orchestration layer may include computing power management, intelligent inference management, and an inference verification module, supporting dynamic scheduling and model verification of inference tasks on the intelligent wireless access network side. The on-device intelligent model engine is a lightweight local inference framework that is compatible with models provided by the wireless intelligent management and orchestration layer and supports interactive optimization of inference results.
[0078] The solution of this embodiment can be widely used in scenarios such as intelligent robots, augmented reality (AR), autonomous driving, and the Industrial Internet of Things (IoT). For example, in an intelligent robot scenario, the robot can obtain an intelligent navigation model through the cloud. After scheduling computing resources, the wireless intelligent management and orchestration function layer can push the optimized small model to the robot terminal, enabling it to efficiently reason locally and improve the accuracy of path planning. In an augmented reality (AR) scenario, the AR terminal can dynamically allocate computing tasks through the wireless intelligent management and orchestration function layer, reducing the local computing burden and improving real-time rendering performance. In an autonomous driving system, large-scale AI models trained in the cloud can be adapted according to the capabilities of the on-board computing platform and intelligently scheduled by the wireless intelligent management and orchestration function layer, and efficient reasoning can be performed on the terminal side to enhance the intelligent decision-making capabilities of autonomous driving.
[0079] The following uses smart glasses as the end-side device to illustrate the application of the solution of this embodiment in augmented reality (AR) and environmental perception, and describes how smart glasses use this solution to perform scene recognition, model reasoning and optimization to enhance the intelligent interactive experience. This process can be completed by the cloud, wireless intelligent management and orchestration function layer, intelligent wireless access network and end-side smart glasses in a collaborative manner, and can include steps such as model distillation, intelligent scheduling, reasoning verification, etc. Figure 4 , the method may include the following steps:
[0080] In step S401, the device-side smart glasses collect data. In step S402, the device-side smart glasses send a model inference request to the AI RAN node based on the collected data.
[0081] When a user wearing smart glasses enters a new environment or needs to identify information, the glasses' sensor system (including cameras, depth sensors, etc.) captures images and sensor data of the current environment. A lightweight perception model running within the smart glasses performs preliminary object recognition, such as detecting road signs, buildings, and products. However, due to the limited computing power of the device-side, it may not be able to independently complete the semantic understanding of complex scenes. To further improve recognition accuracy, the smart glasses can send a model inference request to the AI RAN node based on the collected data, requesting a more accurate model for inference calculations. This model inference request can include a low-resolution image of the current scene, key point data, computing power requirements, and the smart glasses' own computing power and storage space, among other model inference requirements. This information is used by the wireless intelligent management and orchestration layer to evaluate the model inference strategy.
[0082] In step S403, the AI RAN node sends a model inference request to the general computing resource scheduling module of the wireless intelligent management and orchestration functional layer. In step S404, the general computing resource scheduling module determines the executor of the model inference task based on the model inference requirement information and the resource status information of the intelligent wireless access network. In step S405, the general computing resource scheduling module sends the executor of the model inference task and the model inference request to the model management functional module of the wireless intelligent management and orchestration functional layer. In step S406, if the terminal is the executor of the model inference task, the model management functional module assigns the model inference task to the terminal. In step S407, if the node is the executor of the model inference task, the model management functional module assigns the model inference task to the node. In step S408, if the cloud is the executor of the model inference task, the model management functional module assigns the model inference task to the cloud.
[0083] The general computing resource scheduling module evaluates the device-side computing power to determine whether the smart glasses can perform the complete model inference calculations. If the device-side computing power is insufficient, AI RAN nodes or cloud computing capabilities are required. The module selects the optimal inference path: If device-side computing is feasible, the model inference task can be assigned to the smart glasses for local execution to reduce network transmission latency. If the device-side computing power is insufficient but the AI RAN node can handle it, the closest AI RAN node (such as a gNB (next-generation base station) or vBBU (virtual baseband unit)) can be selected to perform the inference calculations and return the results to the smart glasses. In situations such as high AI RAN node load, more accurate inference can be performed in the cloud. The processed inference results are returned to the wireless intelligent management and orchestration functional layer, and then sent to the smart glasses by the access node. The general computing resource scheduling module's decision-making can be based on parameters such as real-time network load, the smart glasses' computing power, power consumption, and inference latency requirements to ensure efficient completion of inference tasks.
[0084] In step S409, the model management module sends the model inference request to the cloud. In step S410, the cloud distills the target model from the large model according to the model inference requirement information.
[0085] The large model library in the cloud can be used to dynamically distill a lightweight model as the target model based on the task requirements of the smart glasses. As an example, the distillation process may include:
[0086] Crop the model based on scene features: if the user is in a shopping mall, you can crop a product recognition model; if the user is on the street, you can optimize the traffic sign recognition model.
[0087] Optimize parameters for glasses computing power: You can adjust the computational complexity of the model and reduce redundant calculations so that it can run efficiently on smart glasses.
[0088] Generate an adapted small model: Finally, a small model suitable for the current environment is obtained as the target model, which can be pushed to the wireless intelligent management and orchestration function layer.
[0089] In step S411, when the cloud does not need to perform the model inference task, the cloud sends the target model to the model management function module.
[0090] In step S412, when the AI RAN node performs a model inference task, the model management function module sends the target model to the AI RAN node. In step S413, when the terminal performs a model inference task, the model management function module sends the target model to the terminal.
[0091] In step S414, when the cloud performs a model inference task, the cloud sends the inference result obtained from performing the model inference task to the wireless intelligent management and orchestration functional layer. In step S415, the wireless intelligent management and orchestration functional layer sends the inference result as the target inference result to the access node. In step S416, the access node sends the target inference result to the terminal. When the AI RAN node performs a model inference task, the AI RAN node can send the inference result obtained from performing the model inference task as the target inference result to the access node, and the access node will also send the target inference result to the terminal.
[0092] In step S417, when the terminal performs a model inference task, the terminal uses the target model to perform the model inference task and obtains an inference result. In step S418, the terminal sends the obtained inference result to the access node. In step S419, the access node sends the obtained inference result to the model management function module. In step S420, the model management function module verifies the inference result obtained by the terminal to determine the target model result. In step S421, the model management function module enables the terminal to obtain the target model result. The model management function module can verify the inference result on the terminal side. Specifically, after the smart glasses complete inference, they upload the inference result to the model management function module. The model management function module can compare the model result obtained by the terminal with the reference inference result obtained by the reference execution party, such as the AI RAN node, to ensure the accuracy of the inference on the terminal side and reduce the false positive rate. If the inference result on the terminal side differs from the reference inference result, the target model of the smart glasses can be optimized and re-inferenced until the verification conditions are met. Alternatively, the reference inference result can be directly sent to the terminal as the target inference result.
[0093] The smart glasses then obtain a target inference result, verified by the model management module, and can overlay it onto the user's field of view using augmented reality (AR). For example, in a shopping mall, the glasses can identify a product and display its price and discount information. On the street, they can identify traffic signals and remind the user when to cross safely. In a museum, they can identify exhibits and provide real-time historical information.
[0094] Through this solution, smart glasses can dynamically adapt to the optimal intelligent model in an environment with low computing power and high latency requirements, improve reasoning efficiency and accuracy, and reduce the computing burden and power consumption of the device.
[0095] The solution of this embodiment can solve the problems of limited end-side computing power, high network latency, and lack of verification mechanism for model inference in related technologies, achieving the following technical effects:
[0096] 1. Reduce computing pressure on the edge: Through cloud distillation and AI RAN intelligent scheduling, the edge avoids running large models directly, achieving low power consumption and efficient inference.
[0097] 2. Optimize network resource scheduling: Use AI RAN to dynamically allocate inference tasks, avoid concentrating computing power in the cloud or on the device side, and improve overall computing efficiency.
[0098] 3. Improved inference accuracy and robustness: The device-side inference results are verified by AI RAN, improving inference credibility and enhancing the model's adaptability in complex wireless environments.
[0099] 4. Reduce network transmission costs: Reduce the large amount of data transmission between the end side and the cloud, improve the real-time performance of reasoning, and is suitable for low-latency, high-reliability AI application scenarios.
[0100] Specifically, compared to traditional fixed small model deployment methods, the solution of this embodiment can adaptively adjust the model according to different application scenarios, ensuring the generalization and computational efficiency of on-device reasoning. Compared to traditional cloud computing models, this embodiment can reduce on-device computing pressure by fully leveraging the computing power of the wireless intelligent management and orchestration layer. It also dynamically balances computing tasks between the device, the intelligent wireless access network, and the cloud, optimizing wireless network resource utilization and improving overall inference efficiency. It can also effectively reduce inference response time, meeting the needs of application scenarios such as smart glasses, augmented reality (AR), and intelligent driving that have strict millisecond-level latency requirements, thereby improving the user experience. Compared to models that rely solely on the results of a single model, the solution of this embodiment can ensure the reliability and accuracy of on-device reasoning through result verification at the wireless intelligent management and orchestration layer, effectively reducing the risk of inference errors caused by limited on-device computing power. Compared to fixed models on the device, the solution of this embodiment can optimize model deployment based on the terminal device, maximizing device battery life and reducing computing power requirements while ensuring inference performance. Compared to AI computing models that rely on fixed architectures, the solution of this embodiment can automatically adjust the computing model for different devices, making it suitable for a wider range of smart terminal application needs. In summary, the solution of this embodiment addresses the shortcomings of related technologies in model adaptability, computing efficiency, reasoning accuracy, low-latency optimization, and device endurance, providing an efficient, reliable, and low-energy AI computing solution for wireless smart terminals.
[0101] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0102] Based on the same inventive concept, the present application also provides a model inference device for implementing the aforementioned model inference method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more model inference device embodiments provided below can be found in the above-mentioned limitations of the model inference method and will not be repeated here.
[0103] In an exemplary embodiment, Figure 5 As shown, a model reasoning device is provided, which can be applied to the wireless intelligent management and orchestration function layer. The device 500 may include:
[0104] The information acquisition module 501 is used to obtain the model reasoning requirement information of the terminal through the access node of the intelligent wireless access network;
[0105] An execution determination module 502 is configured to determine an executor of a model reasoning task based on the model reasoning requirement information and the resource status information of the intelligent wireless access network; the model reasoning task is obtained based on the model reasoning requirement information; the executor includes a cloud, a node, or the terminal; the node is a node in a node set of the intelligent wireless access network; the node set includes the access node;
[0106] The task processing module 503 is used to trigger the executor to use the target model to perform the model reasoning task, so that the terminal obtains the target reasoning result; wherein, the target model is obtained by the cloud end by distilling from the large model according to the model reasoning requirement information; the target reasoning result is the reasoning result that meets the verification conditions of the wireless intelligent management and orchestration function layer.
[0107] In an exemplary embodiment, the information acquisition module 501 is further used to send the model reasoning requirement information to the cloud; the model reasoning requirement information is used by the cloud to distill the target model from the large model.
[0108] In an exemplary embodiment, the execution determination module 502 is used to determine that the terminal is the executor of the model reasoning task if it is determined that the terminal meets the task execution conditions according to the model reasoning requirement information; if it is determined that the terminal does not meet the task execution conditions according to the model reasoning requirement information, then determine that a node in the node set of the smart wireless access network or the cloud is the executor based on the model reasoning requirement information and the resource status information of the smart wireless access network.
[0109] In an exemplary embodiment, the execution determination module 502 is configured to determine the executor based on the node that meets the task execution condition if it is determined based on the model reasoning requirement information and the resource status information of the smart wireless access network that there is a node in the node set of the smart wireless access network that meets the task execution condition; and determine the cloud as the executor if it is determined based on the model reasoning requirement information and the resource status information of the smart wireless access network that there is no node in the node set of the smart wireless access network that meets the task execution condition.
[0110] In an exemplary embodiment, when the executor includes the cloud, the task processing module 503 is used to trigger the cloud to perform the model reasoning task using the target model; obtain a first reasoning result obtained by the cloud performing the model reasoning task, determine the first reasoning result as a target reasoning result that meets the verification condition, and send the target reasoning result to the terminal through the access node; or, when the executor includes the node, the task processing module 503 is used to obtain the target model obtained by the cloud; send the target model to the node so that the node performs the model reasoning task using the target model to obtain a second reasoning result, and send the second reasoning result as a target reasoning result that meets the verification condition to the terminal through the access node; or, when the executor includes the terminal, the task processing module 503 is used to obtain the target model obtained by the cloud; send the target model to the terminal through the access node so that the terminal performs the model reasoning task using the target model to obtain a third reasoning result and returns it through the access node; verify the third reasoning result to determine the target reasoning result, so that the terminal obtains the target reasoning result.
[0111] In an exemplary embodiment, when the executor includes the terminal, the task processing module 503 is used to obtain a reference reasoning result of a reference executor; the reference executor includes a node in the node set of the cloud or the intelligent wireless access network; the reference reasoning result is the reasoning result obtained by the reference executor using the target model to execute the model reasoning task; if it is judged that the third reasoning result meets the verification condition based on the reference reasoning result and the third reasoning result, an indication message is sent to the terminal through the access node; the indication message is used to instruct the terminal to determine the third reasoning result as the target reasoning result; if it is judged that the third reasoning result does not meet the verification condition based on the reference reasoning result and the third reasoning result, the reference reasoning result is used as the target reasoning result and is sent to the terminal through the access node.
[0112] Each module in the above-mentioned model inference device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the network device in hardware form, or can be stored in the memory of the network device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0113] In an exemplary embodiment, a network device is provided, which can serve as a wireless intelligent management and orchestration function layer. The network device can be a server, etc., and its internal structure diagram can be as follows: Figure 6 As shown. The network device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the network device is used to provide computing and control capabilities. The memory of the network device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the network device is used to exchange information between the processor and an external device. The communication interface of the network device is used to communicate with an external device through a network connection. When the computer program is executed by the processor, a model reasoning method is implemented.
[0114] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the network device to which the solution of the present application is applied. The specific network device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0115] In one embodiment, a network device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0116] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0117] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0118] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0119] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0120] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0121] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A model reasoning method, characterized in that: Applied to the wireless intelligent management and orchestration function layer, the method includes: Obtaining model reasoning requirement information of the terminal through the access node of the intelligent wireless access network; Determining an executor of a model reasoning task according to the model reasoning requirement information and the resource status information of the intelligent wireless access network; the model reasoning task is obtained according to the model reasoning requirement information; the executor includes a cloud, a node, or the terminal; the node is a node in a node set of the intelligent wireless access network; the node set includes the access node; Trigger the executor to use the target model to perform the model reasoning task, so that the terminal obtains the target reasoning result; wherein, the target model is distilled from the large model by the cloud according to the model reasoning requirement information; the target reasoning result is the reasoning result that meets the verification conditions of the wireless intelligent management and orchestration function layer.
2. The method according to claim 1, characterized in that After acquiring the model reasoning requirement information of the terminal through the access node of the intelligent wireless access network, the method further includes: The model reasoning requirement information is sent to the cloud; the model reasoning requirement information is used by the cloud to distill the target model from the large model.
3. The method according to claim 1, characterized in that The determining the executor of the model reasoning task according to the model reasoning requirement information and the resource status information of the intelligent wireless access network includes: If it is determined according to the model reasoning requirement information that the terminal meets the task execution condition, then the terminal is determined to be the executor of the model reasoning task; If it is determined according to the model reasoning requirement information that the terminal does not meet the task execution condition, a node in the node set of the smart wireless access network or a cloud is determined as the executor according to the model reasoning requirement information and the resource status information of the smart wireless access network.
4. The method according to claim 3, characterized in that The determining, based on the model reasoning requirement information and the resource status information of the intelligent wireless access network, a node in the node set of the intelligent wireless access network or a cloud as the executor includes: If it is determined based on the model reasoning requirement information and the resource status information of the intelligent wireless access network that a node that meets the task execution condition exists in the node set of the intelligent wireless access network, then determining the executor based on the node that meets the task execution condition; If it is determined based on the model reasoning requirement information and the resource status information of the intelligent wireless access network that there is no node in the node set of the intelligent wireless access network that meets the task execution condition, the cloud is determined as the executor.
5. The method according to any one of claims 1 to 4, characterized in that In a case where the executor includes the cloud, triggering the executor to perform the model reasoning task using the target model so that the terminal obtains the target reasoning result includes: triggering the cloud to perform the model reasoning task using the target model; obtaining a first reasoning result obtained by the cloud performing the model reasoning task, determining the first reasoning result as a target reasoning result that meets a verification condition, and sending the target reasoning result to the terminal via the access node; or, In a case where the executor includes the node, triggering the executor to perform the model reasoning task using the target model so that the terminal obtains the target reasoning result includes: obtaining the target model obtained on the cloud; sending the target model to the node so that the node performs the model reasoning task using the target model to obtain a second reasoning result, and sending the second reasoning result as the target reasoning result that meets the verification condition to the terminal via the access node; or, In the case where the executor includes the terminal, triggering the executor to use the target model to perform the model reasoning task so that the terminal obtains the target reasoning result includes: obtaining the target model obtained from the cloud; sending the target model to the terminal through the access node, so that the terminal uses the target model to perform the model reasoning task to obtain and return a third reasoning result through the access node; verifying the third reasoning result to determine the target reasoning result, so that the terminal obtains the target reasoning result.
6. The method according to claim 5, characterized in that In a case where the executing party includes the terminal, verifying the third reasoning result to determine the target reasoning result, so that the terminal obtains the target reasoning result, includes: Obtaining a reference reasoning result of a reference executor; the reference executor includes a node in the cloud or the node set of the intelligent wireless access network; the reference reasoning result is a reasoning result obtained by the reference executor performing the model reasoning task using the target model; If it is determined that the third reasoning result satisfies the verification condition based on the reference reasoning result and the third reasoning result, sending indication information to the terminal through the access node; the indication information is used to instruct the terminal to determine the third reasoning result as the target reasoning result; If it is determined based on the reference reasoning result and the third reasoning result that the third reasoning result does not meet the verification condition, the reference reasoning result is used as the target reasoning result and sent to the terminal through the access node.
7. A model reasoning device, characterized in that: Applied to the wireless intelligent management and orchestration function layer, the device includes: An information acquisition module is used to obtain the model reasoning requirement information of the terminal through the access node of the intelligent wireless access network; An execution determination module is configured to determine an executor of a model reasoning task based on the model reasoning requirement information and the resource status information of the intelligent wireless access network; the model reasoning task is obtained based on the model reasoning requirement information; the executor includes a cloud, a node, or the terminal; the node is a node in a node set of the intelligent wireless access network; the node set includes the access node; A task processing module is used to trigger the executor to use the target model to perform the model reasoning task, so that the terminal obtains the target reasoning result; wherein, the target model is distilled from the large model by the cloud according to the model reasoning requirement information; the target reasoning result is an inference result that meets the verification conditions of the wireless intelligent management and orchestration function layer.
8. A network device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Execution method and device of model reasoning task, equipment and storage medium
CN119865503A
Block chain-based edge intelligent system resource management method
CN119893588A
AI reasoning task arrangement method and device for wireless access network, equipment and storage medium
CN120166409A
Model task processing method and device, communication equipment, readable storage medium and program product
CN120166422A
Method and system for code offloading in mobile computing
WO2017067586A1