Model inference method and apparatus, and device

By exchanging information and conducting joint inference between the AI ​​units in the first and second devices, the problem of large model inference latency was solved, and computing power was offloaded and accuracy was improved.

WO2026103626A1PCT designated stage Publication Date: 2026-05-21VIVO MOBILE COMM CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2025-11-07
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

In existing technologies, model inference latency is relatively high, terminal computing power is limited, resulting in excessively long inference latency, and cloud inference has excessively high transmission latency, leading to high end-to-end response latency.

Method used

The first device sends information to the second device to identify the first AI unit and receives its inference results. It then performs joint inference with the second AI unit, utilizing the data and computing resources of the second device to offload computing power and improve accuracy.

Benefits of technology

It reduces inference latency, improves model inference accuracy, expands the range of extractable data features, and enhances the performance of joint inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025133363_21052026_PF_FP_ABST
    Figure CN2025133363_21052026_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of communications. Disclosed are a model inference method and apparatus, and a device. The model inference method in the embodiments of the present application comprises: a first device sending first information to a second device, wherein the first information is used for determining a first artificial intelligence (AI) unit; the first device receiving an inference result of the first AI unit that is sent by the second device, wherein the inference result of the first AI unit is an inference result obtained by using the first AI unit to perform inference on data collected by the second device; and the first device performing inference on the basis of the inference result of the first AI unit and a second AI unit, so as to obtain an inference result of the second AI unit, wherein the first AI unit is aligned with the second AI unit.
Need to check novelty before this filing date? Find Prior Art

Description

Model reasoning methods, devices and equipment

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411619882.X, filed on November 13, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application belongs to the field of communication technology, specifically relating to a model reasoning method, apparatus, and device. Background Technology

[0004] In related technologies, the application of artificial intelligence (AI) model inference is generally divided into terminal-local inference and cloud-based inference. Terminal-local inference involves data collected by the terminal for model inference. When the computational load of the task requiring inference is too large, the terminal's computing power is limited, leading to excessively long inference latency. Cloud-based inference involves model inference performed by a cloud server. However, because the cloud is far from the terminal, transmission latency is too high, resulting in high end-to-end response latency. Therefore, it is evident that the model inference methods in related technologies suffer from significant inference latency. Summary of the Invention

[0005] This application provides a model reasoning method, apparatus, and device that can solve the problem of large reasoning latency in model reasoning.

[0006] Firstly, a model reasoning method is provided, which includes:

[0007] The first device sends first information to the second device, the first information being used to identify the first artificial intelligence (AI) unit;

[0008] The first device receives the inference result of the first AI unit sent by the second device. The inference result of the first AI unit is the inference result obtained by using the first AI unit to infer the data collected by the second device.

[0009] The first device performs inference based on the inference result of the first AI unit and the second AI unit to obtain the inference result of the second AI unit. The first AI unit and the second AI unit are aligned.

[0010] Secondly, a model reasoning method is provided, which includes:

[0011] The second device receives the first information sent by the first device and determines the first AI unit based on the first information.

[0012] The second device uses the first AI unit to infer the data collected by the second device and obtain the inference result of the first AI unit;

[0013] The second device sends the inference result of the first AI unit to the first device;

[0014] The reasoning result of the first AI unit is used to perform reasoning to obtain the reasoning result of the second AI unit, and the first AI unit and the second AI unit are aligned.

[0015] Thirdly, a model reasoning device is provided, comprising:

[0016] The sending module is used to send first information to the second device, wherein the first information is used to identify the first artificial intelligence (AI) unit;

[0017] The receiving module is used to receive the inference result of the first AI unit sent by the second device. The inference result of the first AI unit is the inference result obtained by using the first AI unit to infer the data collected by the second device.

[0018] The processing module is used to perform inference based on the inference result of the first AI unit and the second AI unit to obtain the inference result of the second AI unit, wherein the first AI unit and the second AI unit are aligned.

[0019] Fourthly, a model reasoning device is provided, comprising:

[0020] The receiving module is used to receive first information sent by the first device and determine the first AI unit based on the first information;

[0021] The processing module is used to perform reasoning on the data collected by the second device using the first AI unit to obtain the reasoning result of the first AI unit;

[0022] The sending module is used to send the inference results of the first AI unit to the first device;

[0023] The reasoning result of the first AI unit is used to perform reasoning to obtain the reasoning result of the second AI unit, and the first AI unit and the second AI unit are aligned.

[0024] Fifthly, a model reasoning apparatus is provided, the apparatus being configured to perform the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.

[0025] In a sixth aspect, a first device is provided, the first device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.

[0026] In a seventh aspect, a first device is provided, including a processor and a communication interface, wherein,

[0027] A communication interface is used to send first information to a second device, wherein the first information is used to identify a first artificial intelligence (AI) unit.

[0028] The processor is configured to receive the inference result of the first AI unit sent by the second device, wherein the inference result of the first AI unit is an inference result obtained by inferring the data collected by the second device using the first AI unit;

[0029] The communication interface is also used to perform inference based on the inference result of the first AI unit and the second AI unit to obtain the inference result of the second AI unit, wherein the first AI unit and the second AI unit are aligned.

[0030] In an eighth aspect, a second device is provided, the second device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the second aspect.

[0031] Ninthly, a second device is provided, including a processor and a communication interface, wherein,

[0032] A communication interface is used to receive first information sent by a first device and determine a first AI unit based on the first information.

[0033] The processor is used to infer the data collected by the second device using the first AI unit to obtain the inference result of the first AI unit;

[0034] The communication interface is also used to send the inference results of the first AI unit to the first device;

[0035] The reasoning result of the first AI unit is used to perform reasoning to obtain the reasoning result of the second AI unit, and the first AI unit and the second AI unit are aligned.

[0036] In a tenth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.

[0037] Eleventhly, a wireless communication system is provided, comprising: a first device and a second device, wherein the first device is configured to perform the steps of the method as described in the first aspect, and the second device is configured to perform the steps of the method as described in the second aspect.

[0038] In a twelfth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run programs or instructions to implement the method as described in the first aspect, or to implement the method as described in the second aspect.

[0039] In a thirteenth aspect, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the method as described in the first aspect, or to implement the method as described in the second aspect.

[0040] In this embodiment, the first device sends first information to the second device, enabling the second device to identify the first AI unit and use the first AI unit to perform model inference on the data collected by the second device. This allows the second device to utilize data from its side for model inference, improving inference accuracy. The first device receives the inference result of the first AI unit sent by the second device and performs inference based on the inference result of the first AI unit and the second AI unit. Thus, both the first and second devices participate in model inference, utilizing the resources of both. The computing power of the second device can be used to help the first device, which has limited computing power, offload computing power, thereby reducing inference latency. In addition, the second device also collects data from its side for model inference, expanding the range of extractable data features and improving model inference accuracy. The first AI unit and the second AI unit are aligned, which improves the inference effect of joint inference using the first AI unit and the second AI unit. Attached Figure Description

[0041] Figure 1a is a block diagram of a wireless communication system applicable to an embodiment of this application;

[0042] Figure 1b is one of the architecture diagrams of a data plane protocol stack that can be applied to an embodiment of this application;

[0043] Figure 1c is a second architecture diagram of a data plane protocol stack applicable to an embodiment of this application;

[0044] Figure 2 is a flowchart of one of the model reasoning methods provided in the embodiments of this application;

[0045] Figure 3 is a flowchart of one of the model reasoning methods in related technologies;

[0046] Figure 4 is a flowchart of a model reasoning method in related technologies (the second one).

[0047] Figure 5 is a second flowchart of a model reasoning method provided in an embodiment of this application;

[0048] Figure 6 is a flowchart of a model reasoning method provided in an embodiment of this application;

[0049] Figure 7 is a flowchart of a model reasoning method provided in an embodiment of this application;

[0050] Figure 8 is a flowchart of a model reasoning method provided in an embodiment of this application;

[0051] Figure 9 is a flowchart of a model reasoning method provided in an embodiment of this application;

[0052] Figure 10 is a flowchart of a model reasoning method provided in an embodiment of this application;

[0053] Figure 11 is a flowchart of an embodiment of the model reasoning method provided in this application;

[0054] Figure 12 is a schematic diagram of one of the structures of a model inference device provided in an embodiment of this application;

[0055] Figure 13 is a second schematic diagram of the structure of a model inference device provided in an embodiment of this application;

[0056] Figure 14 is a schematic diagram of the structure of a communication device provided in an embodiment of this application;

[0057] Figure 15 is a schematic diagram of the structure of a terminal provided in an embodiment of this application;

[0058] Figure 16 is one of the structural schematic diagrams of a network-side device provided in an embodiment of this application;

[0059] Figure 17 is a second schematic diagram of the structure of a network-side device provided in an embodiment of this application. Detailed Implementation

[0060] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0061] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0062] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.

[0063] It is worth noting that the technologies described in this application are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in this application are often used interchangeably, and the described technologies can be used with the systems and radio technologies mentioned above, as well as with other systems and radio technologies. The following description describes New Radio (NR) systems for illustrative purposes, and the term NR is used in most of the following description; however, these technologies can also be applied to systems other than NR systems, such as 6th generation (6G) radio systems. th Generation 6G communication system.

[0064] Figure 1a shows a block diagram of a wireless communication system applicable to an embodiment of this application. The wireless communication system includes a terminal 11 and a network-side device 12. The terminal 11 can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM, or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the specific type of terminal 11 is not limited in this application embodiment. Network-side equipment 12 may include access network equipment or core network equipment, wherein access network equipment may also be referred to as Radio Access Network (RAN) equipment, radio access network function, or radio access network unit. Access network equipment may include base stations, Wireless Local Area Network (WLAN) access points (APs), or Wireless Fidelity (WiFi) nodes, etc.The term "base station" can be referred to as Node B (NB), Evolved Node B (eNB), Next Generation Node B (gNB), New Radio Node B (NR Node B), Access Point, Relay Base Station (RBS), Serving Base Station (SBS), Base Transceiver Station (BTS), Radio Base Station, Radio Transceiver, Basic Service Set (BSS), Extended Service Set (ESS), Home Node B (HNB), Home Evolved Node B, Transmit / Receive Point (TRP), or any other suitable term in the relevant field, as long as the same technical effect is achieved. The term "base station" is not limited to any specific technical terminology. It should be noted that this application embodiment only uses a base station in an NR system as an example for description and does not limit the specific type of base station.

[0065] Core network equipment, also known as core network nodes, core network functions, or core network elements, includes, but is not limited to, at least one of the following: Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), Home Subscriber Server (HSS), Centralized network configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (or L-NEF), and Binding Support. The core network functions include: BSF (Block Network Function), Application Function (AF), Location Management Function (LMF), Gateway Mobile Location Centre (GMLC), and Network Data Analytics Function (NWDAF). It should be noted that this application embodiment only uses core network equipment in the NR system as an example and does not limit the specific type of core network equipment. If the name of the core network equipment mentioned in this application embodiment changes in subsequent protocol versions (e.g., 6G), it will still be within the scope of protection of this application.

[0066] Optionally, the core network equipment can be implemented by one or more functional modules in a single device, or by multiple devices working together; this application does not specifically limit this. It is understood that the aforementioned functional modules can be network elements in hardware devices, software functional modules running on dedicated hardware, or virtualized functional modules instantiated on a platform (e.g., a cloud platform).

[0067] For ease of understanding, the following explains some aspects of the embodiments of this application:

[0068] 1. Data plane

[0069] The data plane is a protocol stack used for data collection and transmission within a mobile network, as shown in Figures 1b and 1c. Figure 1b shows an example of a data plane protocol stack for UE-RAN, and Figure 1c shows a data plane protocol architecture for User Equipment (UE), Radio Access Network (RAN), and Core Network. The data plane consists of core network data plane functions, RAN data plane functions, and UE data plane functions, providing end-to-end connectivity. The data plane is responsible for data control, including data collection coordination, data collection configuration, and data transmission configuration. It also handles functions such as data acquisition, data transmission, data preprocessing, data privacy and security, data analysis, data storage, and data services.

[0070] The model reasoning method, apparatus, and related equipment provided in this application will be described in detail below with reference to the accompanying drawings and through some embodiments and application scenarios.

[0071] Referring to Figure 2, which is a flowchart of a model inference method provided in an embodiment of this application, the model inference method includes the following steps:

[0072] Step 101: The first device sends first information to the second device, the first information being used to identify the first artificial intelligence (AI) unit;

[0073] Step 102: The first device receives the inference result of the first AI unit sent by the second device. The inference result of the first AI unit is the inference result obtained by using the first AI unit to infer the data collected by the second device.

[0074] Step 103: The first device performs inference based on the inference result of the first AI unit and the second AI unit to obtain the inference result of the second AI unit. The first AI unit and the second AI unit are aligned.

[0075] The first information may include at least one of the following: model information of the first AI unit; description information of the first AI unit. When the first information includes model information and description information of the first AI unit, the model information and description information of the first AI unit can be sent by a single signaling message or by different signaling messages; this embodiment does not limit this.

[0076] It's important to note that AI applications on the terminal generally fall into two categories: local terminal inference and cloud-based inference. The limitation of local terminal inference lies in the limited computing resources of the terminal. When the computational load of the task requiring inference is too large, the limited terminal computing power leads to excessively long inference latency, making it impossible to guarantee the overall operation timeline requirements. Cloud-based inference, on the other hand, suffers from excessive transmission latency due to the distance between the cloud and the terminal, thus also resulting in high end-to-end response latency.

[0077] In related technologies, as shown in Figure 3, data services can be provided by the network. The requesting party requests data from the network, which corresponds to the model input of the requesting party's AI unit. To respond, the network needs to collect data on its side. After receiving the response data, the requesting party performs inference using its AI unit and obtains the inference result. The requesting party can be a UE or a third-party OTT server. In this related technology, the network provides data services to the requesting party, and this data serves as the model input for the requesting party's AI unit. If the number of parameters in the AI ​​unit is large, but the requesting party's computing power is limited, there is a problem of excessively long inference latency for the AI ​​unit.

[0078] In another related technology, as shown in Figure 4, the network provides data services or AI services, and the requester requests the results of the AI ​​service. The requester requests the results of the AI ​​service from the network; for example, if the AI ​​service is environment reconstruction, the data is the result of environment reconstruction; if the AI ​​service is trajectory planning, the data is the result of trajectory planning. To obtain the results of the AI ​​service, the network can obtain the results based on model inference from the AI ​​unit or based on non-AI algorithms. The requester can be a UE or a third-party OTT server. In this related technology, the network directly provides the results of the AI ​​service to the requester, such as the results of environment reconstruction, trajectory planning, environmental prediction, and target monitoring. However, if the requester also has local information that can be used as model input, the requester's local data cannot be utilized, limiting the inference accuracy of the AI ​​unit; or if local information is shared with the network for inference, there are issues of excessive uplink overhead and leakage of local data privacy.

[0079] To ensure response latency for latency-sensitive services and improve the inference accuracy of AI units, embodiments of this application consider utilizing the computing resources and collectable data resources of the communication network to help reduce the inference latency of terminal AI applications and improve the inference accuracy of terminal AI applications.

[0080] In one implementation, a first device sends first information to a second device, receives the inference result of a first AI unit from the second device, and determines the inference result of a second AI unit based on the inference result of the first AI unit and the second AI unit. The first information is used to determine the first AI unit; for example, the first information includes the first AI unit. The first AI unit and the second AI unit are aligned.

[0081] In one implementation, the first AI unit and the second AI unit are sub-units of the target AI unit.

[0082] In one implementation, the first AI unit and the second AI unit are aligned, also known as matched or paired.

[0083] In one implementation, the first AI unit and the second AI unit are aligned. Optionally, the alignment of the first AI unit and the second AI unit means that the first AI unit and the second AI unit can be sub-units of the target AI unit, or that the first AI unit and the second AI unit can be different sub-units of the same AI unit.

[0084] In one implementation, the first AI unit and the second AI unit are aligned. Optionally, alignment of the first AI unit and the second AI unit means that when the first AI unit and the second AI unit perform joint reasoning, their reasoning accuracy meets a preset requirement. For example, the joint reasoning accuracy of the first AI unit and the second AI unit is greater than or equal to a preset threshold 1, or the error of the joint reasoning is less than or equal to a preset threshold 2.

[0085] In one implementation, the output of the first AI unit has no physical meaning.

[0086] In one implementation, the output of the first AI unit has physical meaning.

[0087] In one implementation, the first device may be a terminal (such as a UE) or an OTT server. An OTT (Over-The-Top) server is a third-party server whose content or services are built on top of basic telecommunications services.

[0088] In one implementation, the second device may be a network-side device.

[0089] In one implementation, the first device can be a network-side device. The second device can be a terminal (such as a UE) or an OTT server.

[0090] The AI ​​unit described in this application embodiment may also be referred to as an AI model, machine learning (ML) model, ML unit, AI structure, AI function, AI characteristic, machine learning model, neural network, neural network function, or neural network functionality, etc. Alternatively, the AI ​​unit may refer to a processing unit capable of implementing specific algorithms, formulas, processing flows, capabilities, etc., related to AI. The AI ​​unit may also be a processing method, algorithm, function, module, or unit for a specific dataset. Furthermore, the AI ​​unit may be a processing method, algorithm, function, module, or unit running on AI or ML-related hardware such as a graphics processing unit (GPU), neural network processing unit (NPU), tensor processing unit (TPU), or application-specific integrated circuit (ASIC). This application does not impose specific limitations in this regard. Optionally, the specific dataset includes the input and / or output of the AI ​​unit.

[0091] Optionally, the identifier of the AI ​​unit may be an AI model identifier, an AI structure identifier, an AI algorithm identifier, or an identifier of a specific dataset associated with the AI ​​unit, or an identifier of a specific scenario, environment, channel characteristics, or device related to the AI ​​or ML, or an identifier of a function, feature, capability, or module related to the AI ​​or ML. This application does not specifically limit this.

[0092] In this embodiment, the first device sends first information to the second device, enabling the second device to identify the first AI unit and use the first AI unit to perform model inference on the data collected by the second device. This allows the second device to utilize data from its side for model inference, improving inference accuracy. The first device receives the inference result of the first AI unit sent by the second device and performs inference based on the inference result of the first AI unit and the second AI unit. Thus, both the first and second devices participate in model inference, utilizing the resources of both. The computing power of the second device can be used to help the first device, which has limited computing power, offload computing power, thereby reducing inference latency. In addition, the second device also collects data from its side for model inference, expanding the range of extractable data features and improving model inference accuracy. The first AI unit and the second AI unit are aligned, which improves the inference effect of joint inference using the first AI unit and the second AI unit.

[0093] Optionally, the alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions:

[0094] The first AI unit and the second AI unit are sub-units of the target AI unit;

[0095] The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

[0096] The target AI unit may include a first AI unit and a second AI unit. The target AI unit may be composed of a first AI unit and a second AI unit. For example, the first AI unit may be the first 1 to K layers of the target AI unit, and the second AI unit may be the remaining layers of the target AI unit excluding the first AI unit.

[0097] In this embodiment, the first AI unit and the second AI unit are sub-units of the target AI unit, which can better align the AI ​​units performing model inference on the first device side and the second device side; by using the two sub-units of the target AI unit for joint inference, better inference results can be obtained.

[0098] In this embodiment, the reasoning accuracy of the joint reasoning of the first AI unit and the second AI unit meets the preset requirements. By aligning the AI ​​units that perform model reasoning on the first device side and the second device side with the reasoning accuracy of the AI ​​units, a better reasoning effect can be obtained.

[0099] Optionally, the first information includes at least one of the following:

[0100] The model information of the first AI unit; the description information of the first AI unit;

[0101] The model information of the first AI unit includes at least one of the following:

[0102] The executable file corresponding to the first AI unit;

[0103] Information used to indicate the reference model structure of the first AI unit;

[0104] Information used to indicate the order of model parameters of the first AI unit;

[0105] The parameter information of the first AI unit;

[0106] The identifier of the first AI unit;

[0107] Information used to indicate the parameter modification layer of the first AI unit;

[0108] Information used to indicate the layer corresponding to the first AI unit.

[0109] In one implementation, the first device sends model information of the first AI unit to the second device, which can be understood or replaced as the first device sending the first AI unit to the second device.

[0110] It should be understood that the second device obtains the executable file corresponding to the first AI unit. When the first AI unit needs to be used, it can run the executable file corresponding to the first AI unit, and the execution result of the executable file can be used as the reasoning result of the first AI unit.

[0111] The information indicating the reference model structure of the first AI unit can be understood or replaced as a reference model structure indication of the first AI unit. This reference model structure indication can be used to indicate a reference model structure, for example, a typical reference model structure predefined by the protocol, such as a fully connected network, convolutional neural network, transformer network, or Long Short-Term Memory (LSTM) network. The second device can use this reference model structure as the model structure of the first AI unit. Additionally, the first information may also include model parameters of the first AI unit, or the second device may pre-store model parameters of the first AI unit. The second device can determine the first AI unit based on its model structure and model parameters.

[0112] The information indicating the order of model parameters of the first AI unit can be understood or replaced as a model parameter order indicator for the first AI unit. This model parameter order indicator can indicate the order of model parameters; for example, it can indicate descending or ascending order by layer, or hierarchical indication of neurons, or joint indication of multiple layers of neurons. The second device can determine the model parameters of the first AI unit based on this model parameter order indicator. For example, the second device can determine the model parameters of the first AI unit by combining the parameter file of the first AI unit and the model parameter order indicator.

[0113] The parameter information of the first AI unit may include a parameter file for the first AI unit. The second device can determine the model parameters of the first AI unit based on the parameter information of the first AI unit. For example, the parameter file of the first AI unit may include the model parameters of the first AI unit. In addition, the first information may also include a reference model structure indication of the first AI unit, or the second device may use a preset model structure as the model structure of the first AI unit. The second device can determine the first AI unit through the model structure and model parameters of the first AI unit.

[0114] It should be understood that the second device can identify the first AI unit by its identifier. The second device may store multiple AI units, and the first AI unit can be identified from the stored AI units by its identifier; or, the second device may request the first AI unit from other devices by its identifier; and so on. This embodiment does not limit this.

[0115] The information indicating the parameter modification layer of the first AI unit can be understood or replaced as a modification layer indication for the first AI unit. The second device can determine the layer whose model parameters need to be modified based on this parameter modification layer indication information. For example, the first information also includes the parameter file of the first AI unit, and the parameter modification layer indication information indicates that the parameter modification layer is the Nth layer. The second device can modify the model parameters of the Nth layer of the stored AI unit to the parameters in the parameter file of the first AI unit, and use the modified AI unit as the first AI unit. The stored AI unit can be an AI unit already shared by the first device.

[0116] The information used to indicate the layer corresponding to the first AI unit can be understood or replaced as the layer indication corresponding to the first AI unit (e.g., the first 1 to K layers). For example, the second device can use the first 1 to K layers of stored AI units as the first AI unit. This stored AI unit can be an AI unit already shared by the first device.

[0117] In this embodiment, by sending the first information to the second device, the second device can determine the first AI unit through the first information. The first device sending the first information to the second device to determine the first AI unit can ensure that the first AI unit and the second AI unit are aligned. This ensures that the AI ​​units performing model inference on the first device side and the second device side are aligned, avoiding the signaling interaction required to align the models at both ends when the first device and the second device each provide AI units.

[0118] Optionally, the reasoning result of the first AI unit satisfies any one of the following conditions:

[0119] The reasoning result of the first AI unit does not have corresponding physical parameters;

[0120] The reasoning results of the first AI unit have corresponding physical parameters.

[0121] In one implementation, when the output of the first AI unit has corresponding physical parameters, the first device can parse the physical parameters corresponding to the output of the first AI unit, while the second device cannot parse the physical parameters corresponding to the output of the first AI unit; or,

[0122] When the output of the first AI unit has corresponding physical parameters, both the first device and the second device can parse the physical parameters corresponding to the output of the first AI unit.

[0123] The reasoning result of the first AI unit may include any one of the following:

[0124] The output of the first AI unit has no physical meaning.

[0125] The physical meaning of the output of the first AI unit cannot be interpreted by the second device, but the physical meaning can be understood by the first device.

[0126] The physical meaning of the output of the first AI unit can be interpreted by both the second and first devices.

[0127] In one implementation, the reasoning result of the first AI unit is the output result of the first AI unit. This output result has no physical meaning, which helps to protect the first AI unit from being used by devices other than the first device, and can be used to restrict the users of the first AI unit.

[0128] In one implementation, the reasoning result of the first AI unit is the output result of the first AI unit. The second device cannot interpret the physical meaning of the output result, but the first device can know the physical meaning of the output result. This is beneficial to protecting the first AI unit from being used by devices other than the first device. It can be used to restrict the users of the first AI unit. Compared with an output result without physical meaning, it is beneficial for the first device to perform verification of the transmission result.

[0129] In one implementation, the reasoning result of the first AI unit is the output result of the first AI unit. Both the second device and the first device can parse the physical meaning of the output result, which is beneficial for sharing the output result of the first AI unit with other requesters.

[0130] In this embodiment, the reasoning result of the first AI unit does not have corresponding physical parameters, so even if a device other than the first device obtains the reasoning result of the first AI unit, it cannot parse the reasoning result. This is beneficial to protect the first AI unit from being used by devices other than the first device, thereby limiting the users of the first AI unit.

[0131] In this embodiment, the reasoning result of the first AI unit has corresponding physical parameters, which is beneficial for the first device to verify whether the reasoning result of the first AI unit is accurate, and can better support sharing the first AI unit with other devices.

[0132] Optionally, the first device performs inference based on the inference result of the first AI unit and the second AI unit, including any one of the following:

[0133] The first device uses the reasoning result of the first AI unit as the model input of the second AI unit for reasoning;

[0134] The first device uses the reasoning results of the first AI unit and the data collected by the first device as model inputs for the second AI unit to perform reasoning.

[0135] In this embodiment, the first device uses the reasoning result of the first AI unit as the model input of the second AI unit for reasoning. The second device (such as a network) is less likely to have a perception blind spot and may have a better observation perspective than the first device. It can obtain better observation data than the first device, thereby improving the reasoning accuracy of the first AI unit and the second AI unit.

[0136] In this embodiment, the first device uses the inference result of the first AI unit and the data collected by the first device as the model input of the second AI unit for inference. This allows the device to simultaneously utilize the data collected locally by the second device and the first device for model inference, thereby improving the inference accuracy of the first AI unit and the second AI unit.

[0137] Optionally, the first information is carried by at least one of the following:

[0138] Control plane messages; User plane messages; Data plane messages.

[0139] In one implementation, the protocol can predefine which of the following three types of messages—control plane, user plane, and data plane—is used to transmit the first information. The network can then configure which of these three messages is used to transmit the first information, or the first device can autonomously choose which message to use. For example, if the first AI unit is small, the network can configure it to transmit the first information via control plane messages; if the first AI unit is large, the network can configure it to transmit the first information via either user plane or data plane messages, thus providing greater flexibility in transmitting the first information.

[0140] In one implementation, the protocol can predefine the transmission of first information via control plane messages.

[0141] In one implementation, the protocol can predefine the transmission of first information via user plane messages.

[0142] In one implementation, the protocol can predefine the transmission of first information via data plane messages.

[0143] In this embodiment, by transmitting the first information through at least one of control plane messages, user plane messages, and data plane messages, the first device can send the first information to the second device, and the transmission of the first information is highly flexible.

[0144] Optionally, the control plane message includes a first message, which is used to request the allocation of resources for AI services;

[0145] And / or,

[0146] The user plane message includes a second message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0147] And / or,

[0148] The data plane message includes a third message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0149] In this embodiment, transmitting first information via a first message enables the transmission of first information via control plane messages; transmitting first information via a second message enables the first device to send first information to the second device via user plane messages; and transmitting first information via a third message enables the first device to send first information to the second device via data plane messages.

[0150] Optionally, the second device includes a first network element, and the first device sends first information to the second device, including at least one of the following:

[0151] The first device sends first information to the first network element through the first entity, and the first message is a message transmitted between the first device and the first entity;

[0152] The first device sends the second message to the first network element;

[0153] The first device sends the third message to the first network element.

[0154] The second device can be a network-side device. This network-side device may include a first entity, a second network element, and the first network element. The first entity can be a newly established entity, such as an AI service management function or a collaborative control function; alternatively, the first entity can be co-located with an existing entity, such as co-located with a communication management function or a policy-based billing function. The second network element can be a network element used for service resource management; for example, the second network element can be an AI resource management function, a resource management function, or a computing power resource management function. The first network element can be a network element used for service processing and the implementation of service policy rules; for example, the first network element can be an AI resource node, a resource node, a computing power resource node, or a service resource node.

[0155] The first message carries the first information.

[0156] In one embodiment, the control plane message further includes a fourth message and a fifth message. The fourth message is used to request AI resources; the fifth message is used to configure AI resources. The fourth message carries the first information, and the fifth message carries the first information.

[0157] In one implementation, the first message is a message sent by the first device to the first entity.

[0158] In one implementation, the fourth message is a message sent by the first entity to the second network element.

[0159] In one implementation, the fifth message is a message sent by the second network element to the first network element.

[0160] In one embodiment, the second message is a message sent by the first device to the first network element.

[0161] In one embodiment, the third message is a message sent by the first device to the first network element.

[0162] In one implementation, the first device can transmit first information to the first network element through a first message, a fourth message, and a fifth message, and the first network element can use the first AI unit to perform reasoning to obtain the reasoning result of the first AI unit.

[0163] Optionally, when the first information includes the description information of the first AI unit, the description information of the first AI unit includes at least one of the following:

[0164] The effective duration of the first AI unit;

[0165] The configuration environment of the first AI unit;

[0166] The model input information of the first AI unit;

[0167] The model outputs relevant information from the first AI unit;

[0168] The computational load of the first AI unit;

[0169] The inference latency required by the first AI unit;

[0170] The required computing speed of the first AI unit;

[0171] The storage requirements of the first AI unit;

[0172] The function of the first AI unit;

[0173] Information used to indicate whether a second device is needed for data collection;

[0174] Used to obtain information about the application service AS of the first AI unit;

[0175] Wherein, the model input information of the first AI unit is associated with the function of the first AI unit; and / or, the model output information of the first AI unit is associated with the function of the first AI unit.

[0176] The configuration environment of the first AI unit may include the deep learning framework of the first AI unit (such as TensorFlow, PyTorch, etc.) or the library version supported by the first AI unit.

[0177] The model input information of the first AI unit may include: the model input information indication of the first AI unit, such as data items, data format, data precision, etc.

[0178] The model output information of the first AI unit may include: the model output information indication of the first AI unit, such as data items, data precision, data dimensions, storage format, etc.

[0179] The computational load of the first AI unit can include the plural form of floating-point operations (FLOPs). Specifically, it can include, for example, the number of floating-point operations (FLOPs), millions of floating-point operations (MFLOPs), giga-floating-point operations (GFLOPs), etc.

[0180] The inference latency required for the first AI unit is, for example, 1ms, 5ms, 10ms, 100ms, etc.

[0181] The required computing speed for the first AI unit includes, for example, floating-point operations per second (FLOPS), millions of floating-point operations per second (MFLOPS), and giga floating-point operations per second (GFLOPS).

[0182] The information used to obtain the application service AS of the first AI unit may include: AS address, AS name, and / or AS equipment manufacturer, etc.

[0183] In this embodiment, the first device sends the description information of the first AI unit to the second device. The description information of the first AI unit can be used to configure the environment of the first AI unit, allocate the required computing resources to the first AI unit, indicate how to collect the data required for model inference for the first AI unit, and / or determine and obtain the first AI unit from the first device on the OTT server or user plane. Thus, the second device can use the description information of the first AI unit to perform model inference using the first AI unit.

[0184] Optionally, the first AI unit may include at least one of the following functions:

[0185] Environment reconstruction; trajectory planning; driving decision-making (or navigation decision); object recognition; object detection; object tracking; intrusion detection.

[0186] Among them, environmental reconstruction is also known as Simultaneous Localization and Mapping (SLAM).

[0187] In addition, target recognition can also be described as object recognition; target detection can also be described as object detection; and target tracking can also be described as object tracking.

[0188] In this embodiment, a first AI unit and a second AI unit are used for joint inference in the environment reconstruction scenario. By using data collected from the second device (such as a network) as input when the first AI unit is inferring, and using data collected from the first device as model input when the second AI unit is inferring, the environment can be perceived from more angles during environment reconstruction, which can improve the effect of environment reconstruction. In addition, since the first AI unit is offloaded to the second device for inference, it is beneficial to reduce the overall inference latency of the first AI unit and the second AI unit when the computing power of the first device is limited, which can further improve the effect of environment reconstruction.

[0189] In this implementation, a first AI unit and a second AI unit are used for joint reasoning in the trajectory planning scenario. This allows for the perception of the environment from more angles during trajectory planning, thereby improving the trajectory planning effect. For example, if only the first device collects data for model reasoning, visual blind spots are likely to occur during trajectory planning, making it impossible to reconstruct dynamic obstacles in the blind spots. The planned trajectory is prone to collisions with dynamic obstacles in the blind spots. However, by integrating the environmental reconstruction data collected by the first and second devices, dynamic obstacles in the blind spots can be included, resulting in a safer trajectory planning outcome that avoids collisions.

[0190] In this implementation, the first AI unit and the second AI unit are used for joint reasoning in the driving decision-making scenario, similar to the trajectory planning scenario. This allows the environment to be perceived from more angles during driving decisions, making the driving process of the first device safer and more efficient, thereby improving the driving decision-making effect.

[0191] Similarly, in target recognition scenarios, joint reasoning using the first AI unit and the second AI unit can improve target recognition performance; or, in target detection scenarios, joint reasoning using the first AI unit and the second AI unit can improve target detection performance; or, in target tracking scenarios, joint reasoning using the first AI unit and the second AI unit can improve target tracking performance; or, in intrusion detection scenarios, joint reasoning using the first AI unit and the second AI unit can improve intrusion detection performance.

[0192] Optionally, when the function of the first AI unit includes environment reconstruction, the model input information of the first AI unit includes at least one of the following: a first data item, the data dimension of the first data item; the model output information of the first AI unit includes a second data item; and / or

[0193] When the function of the first AI unit includes trajectory planning, the model input information of the first AI unit includes at least one of the following: a third data item, wherein the data dimension of the third data item; the model output information of the first AI unit includes a fourth data item; and / or

[0194] When the function of the first AI unit includes driving decision-making, the model input information of the first AI unit includes at least one of the following: a fifth data item, wherein the data dimension of the fifth data item; the model output information of the first AI unit includes a sixth data item.

[0195] The first data item includes at least one of the following: chromaticity and luminance;

[0196] The data dimensions of the first data item include at least one of the following: number of channels, height, width, and number of frames;

[0197] The second data item includes at least one of the following: time point, obstacle type, three-dimensional coordinates of the obstacle, envelope of the obstacle, volume of the obstacle, moving speed of the obstacle, moving direction of the obstacle, and distance between the obstacle and the target object;

[0198] The third data item includes at least one of the following: chromaticity, luminance;

[0199] The data dimensions of the third data item include at least one of the following: number of channels, height, width, and number of frames;

[0200] The fourth data item includes at least one of the following: time point, speed, and direction of movement;

[0201] The fifth data item includes at least one of the following: chromaticity, luminance;

[0202] The data dimensions of the fifth data item include at least one of the following: number of channels, height, width, and number of frames;

[0203] The sixth data item includes at least one of the following: time point, throttle opening, braking force, and steering angle.

[0204] In related technologies, network-provided data services can improve the inference accuracy of terminal AI applications, but they do not solve the problem of limited terminal computing power. Network-provided AI services can reduce the inference latency of terminal AI applications, but they suffer from problems such as ineffective utilization of locally collected information, excessive uplink overhead, and leakage of terminal data privacy. This application proposes a method for providing AI services via a network, which helps to utilize information collected both locally on the terminal and from the network side, and can reduce the terminal's computing power and power consumption.

[0205] The following examples will provide further explanation:

[0206] The first device is the UE (User Equipment), which can also be an OTT (Over-The-Top) server.

[0207] Second device: Network-side device (including at least one of the following: first target network element, second target network element, third target network element, fourth target network element, fifth target network element, sixth target network element, seventh target network element, eighth target network element, ninth target network element, tenth target network element, and eleventh target network element).

[0208] The network-side devices are shown in Table 1:

[0209] Table 1

[0210] Example 1:

[0211] This example provides a segmentation model through a first device, which reduces the computing power requirements of the first device and can utilize data collected from both the first device and the network simultaneously. This avoids problems such as excessive AI application response latency due to limited terminal computing power or limited inference accuracy due to limited data sources.

[0212] As shown in Figure 5, the corresponding signaling flow is as follows:

[0213] Step (1): The first device sends first information to the second device. The first information includes the first AI unit.

[0214] The first AI unit in the first information includes at least one of the following:

[0215] The executable file corresponding to the first AI unit;

[0216] The reference model structure indication for the first AI unit;

[0217] The order of model parameters in the first AI unit;

[0218] The parameter file for the first AI unit;

[0219] Instructions of the AI ​​unit of the first device (i.e., the identifier of the first AI unit);

[0220] Modification layer instructions for the first AI unit;

[0221] The layer indicator corresponding to the first AI unit (e.g., the first 1 to K layers).

[0222] The first information enables the second device to acquire the first AI unit, thereby reducing the computing power requirements of the first device.

[0223] In one implementation, the method of sending the first AI unit includes:

[0224] Method 1: Direct instruction;

[0225] -Method 1a:

[0226] The first information includes: the executable file corresponding to the first AI unit;

[0227] -Method 1b:

[0228] The first piece of information includes:

[0229] The reference model structure indication for the first AI unit;

[0230] The order of model parameters in the first AI unit;

[0231] The parameter file for the first AI unit.

[0232] -Method 1c:

[0233] The first piece of information includes:

[0234] The reference model structure indication for the first AI unit;

[0235] The layer indicator corresponding to the first AI unit (e.g., the first 1 to K layers);

[0236] The order of model parameters in the first AI unit;

[0237] The parameter file for the first AI unit.

[0238] Method 2: Indirect Instruction

[0239] -Method 2a:

[0240] The first information includes: instructions from the AI ​​units that have been shared by the first device;

[0241] Method 2b:

[0242] The first piece of information includes:

[0243] Instructions from the AI ​​unit already shared by the first device;

[0244] The parameter file for the first AI unit;

[0245] Modification layer instructions for the first AI unit;

[0246] Method 2c:

[0247] The first piece of information includes:

[0248] Instructions from the AI ​​unit already shared by the first device;

[0249] The layer indicator corresponding to the first AI unit (e.g., the first 1 to K layers).

[0250] Optionally, the first information may further include third information, which includes at least one of the following:

[0251] The effective duration of the first AI unit;

[0252] Configure the environment, such as deep learning frameworks (TensorFlow, PyTorch, etc.) and supported library versions;

[0253] The first AI unit's model input information indicates, for example, data items, data format, data precision, etc.

[0254] The model output information of the first AI unit indicates, for example, data items, output data precision, data dimensions, storage format, etc.

[0255] The computational load of the first AI unit, such as the number of floating-point operations (FLOPs), millions of floating-point operations (MFLOPs), billions of floating-point operations (GFLOPs), etc.

[0256] Storage requirements for the first AI unit;

[0257] The function of the first AI unit.

[0258] Step (2): The second device performs model reasoning for the first AI unit based on the first information collection data.

[0259] The second device determines the data requirements to be collected based on the model input information of the first AI unit in the first information. The second device collects the corresponding data, uses it as the model input of the first AI unit, and obtains the reasoning result of the first AI unit.

[0260] Optionally, the second device determines the second information based on the model output information indication (such as output data precision, data dimension, and storage format) of the first AI unit.

[0261] Through step (2), data from the network side can be collected, thereby obtaining more information compared to local inference on the first device, thus improving the inference accuracy of the AI ​​unit and improving service accuracy.

[0262] Step (3): The second device feeds back second information to the first device, the second information including the reasoning result of the first AI unit.

[0263] The reasoning result of the first AI unit in the second information includes at least one of the following:

[0264] The output of the intermediate layer of the target AI unit (i.e., the first AI unit) has no physical meaning;

[0265] The physical meaning of the output of the first AI unit cannot be interpreted by the second device, but the physical meaning can be understood by the first device.

[0266] The physical meaning of the output of the first AI unit can be interpreted by both the second and first devices.

[0267] In one implementation, the reasoning result of the first AI unit is the output result of the intermediate layer of the target AI unit. This output result has no physical meaning, which helps to protect the first AI unit from being used by devices other than the first device and can be used to restrict the users of the first AI unit.

[0268] In one implementation, the reasoning result of the first AI unit is the output result of the first AI unit. The second device cannot interpret the physical meaning of the output result, but the first device can know the physical meaning of the output result. This is beneficial to protecting the first AI unit from being used by devices other than the first device. It can be used to restrict the users of the first AI unit. Compared with an output result without physical meaning, it is beneficial for the first device to perform verification of the transmission result.

[0269] In one implementation, the reasoning result of the first AI unit is the output result of the first AI unit. Both the second device and the first device can parse the physical meaning of the output result, which is beneficial for sharing the output result of the first AI unit with other requesters.

[0270] Optionally, the second information further includes at least one of the following:

[0271] Instructions from the first AI unit;

[0272] Timestamp;

[0273] The remaining effective time of the first AI unit.

[0274] Step (4): The first device performs model inference for the second AI unit based on the second information and obtains the inference result. The second AI unit and the first AI unit are aligned.

[0275] In one embodiment, the operation of the first device includes at least one of the following:

[0276] The first device uses the second information as the model input of the second AI unit, performs model reasoning on the second AI unit, and obtains the reasoning result;

[0277] The first device uses the second information and the data collected locally by the first device as model inputs to the second AI unit, performs model inference on the second AI unit, and obtains the inference result.

[0278] In this embodiment, the first device uses the second information as the model input of the second AI unit, performs model inference on the second AI unit, and obtains the inference result. The second device (such as the network) is less likely to have a perception blind spot and may have a better observation perspective than the first device, and can obtain better observation data than the first device, thereby improving the inference accuracy of the AI ​​unit.

[0279] In this embodiment, the first device uses the second information and the data collected locally by the first device as the model input of the second AI unit, performs model inference on the second AI unit, and obtains the inference result. The model inference can be performed using the data collected locally by the network side and the first device, which further improves the inference accuracy of the AI ​​unit.

[0280] Example 1 provides a process for providing AI services over a network, which can offload the computing power of the first device and improve the inference accuracy of the AI ​​unit by using data collected from the network side. In this process, the first AI unit sent by the first device to the second device can ensure that the first AI unit and the second unit are aligned. This ensures that the AI ​​unit inferred by the second device is aligned with the AI ​​unit inferred by the first device, avoiding the signaling interaction required to align the two end models when the first device and the second device each provide AI units.

[0281] Example 2:

[0282] Example 2 describes the sending process of the first AI unit using a typical scenario.

[0283] Example 2-1:

[0284] This example refines the proposed solution's process in conjunction with potential future external AI service processes, particularly including the transmission signaling (such as control plane interaction signaling) bearer for the first AI unit. This allows the transmission of the first AI unit to be achieved through additional interactions in the external AI service process.

[0285] As shown in Figure 6, the signaling flow is as follows:

[0286] Step (1): The first device sends a first target message (i.e., the aforementioned first message) to the first entity. The first target message is used to request the allocation of corresponding resources for the AI ​​service. Optionally, the first target message includes first information, and the first information includes a first AI unit. The resources include at least one of communication resources, data resources, computing power resources, and AI resources.

[0287] The first entity can be a newly established entity, such as a fifth target network element (e.g., AI service management function, or collaborative control function), or it can be co-located with an existing entity, such as co-located with a second target network element (e.g., communication management function), or co-located with a fourth target network element (e.g., policy billing function).

[0288] The first target message can be forwarded through the NAS interface.

[0289] The first AI unit includes at least one of the following:

[0290] The executable file corresponding to the first AI unit;

[0291] Reference model structure indication of the first AI unit (i.e., information used to indicate the reference model structure of the first AI unit);

[0292] The model parameter order indicator of the first AI unit (i.e., information used to indicate the order of model parameters of the first AI unit);

[0293] The parameter file of the first AI unit (i.e., the parameter information of the first AI unit);

[0294] Instructions of the AI ​​unit of the first device (i.e., the identifier of the first AI unit);

[0295] Modification layer indication of the first AI unit (i.e., information used to indicate the parameter modification layer of the first AI unit);

[0296] The layer indicator corresponding to the first AI unit (i.e., information used to indicate the layer corresponding to the first AI unit), for example, the first 1 to K layers.

[0297] Optionally, the first information may also include at least one of the following:

[0298] The effective duration of the first AI unit;

[0299] Configure the environment, such as deep learning frameworks (TensorFlow, PyTorch, etc.) and supported library versions;

[0300] The first AI unit's model input information indicators, such as data items, data format, data precision, etc.

[0301] The first AI unit's model output information indicates, such as output data precision, data dimensions, and storage format.

[0302] Step (2): The first entity and the fourth target network element (e.g., policy charging function) determine the AI ​​policy.

[0303] Step (3a): The first entity sends a second target message to the ninth target network element (e.g., data plane management function) to request the data source and determine the data source information.

[0304] The second target message is determined based on the first information, specifically, based on the model input information of the first AI unit.

[0305] Step (3a) is optional. Step (3a) must be performed under at least one of the following conditions:

[0306] The AI ​​session establishment request does not contain data source information;

[0307] The AI ​​session establishment request includes an indication that interaction with the data plane management function is required;

[0308] The AI ​​business management function cannot obtain data source information.

[0309] In one implementation, the AI ​​business management function selects the data plane management function instance, considering at least one of the following factors:

[0310] DNN; S-NSSAI; AI business identifier; supported AI services; service area information; data source information.

[0311] Step (3b): The ninth target network element (e.g., data plane management function) sends a third target message to the first entity and responds to the second target message to indicate the data source information.

[0312] Step (4): The first entity sends a fourth target message (i.e., the aforementioned fourth message) to the seventh target network element (AI resource management function). The fourth target message contains AI resource request information and may also include first information. The first information includes an AI unit. The seventh target network element (AI resource management function) determines a suitable AI resource node based on the AI ​​resource request information and the AI ​​strategy. The seventh target network element is the aforementioned second network element.

[0313] The first AI unit in the first information includes at least one of the following:

[0314] The executable file corresponding to the first AI unit;

[0315] The reference model structure indication for the first AI unit;

[0316] The order of model parameters in the first AI unit;

[0317] The parameter file for the first AI unit;

[0318] Instructions from the AI ​​unit of the first device;

[0319] Modification layer instructions for the first AI unit;

[0320] The layer indicator corresponding to the first AI unit (e.g., the first 1 to K layers).

[0321] The fourth target message is determined based on the first information. Specifically, it can be determined based on the model output information of the first AI unit. The fourth target message can be used to indicate the data output information of the seventh network.

[0322] The AI ​​resource request information includes at least one of the following:

[0323] Model information (such as the model or its download address);

[0324] Data source information (data source address);

[0325] AI business computational load (computing power size or type);

[0326] AS address.

[0327] Step (5): The seventh target network element (such as the AI ​​resource management function) selects the appropriate eighth target network element (such as the AI ​​resource node) based on the AI ​​resource request information.

[0328] Step (6a): The seventh target network element (such as the AI ​​resource management function) sends a fifth target message (i.e., the aforementioned fifth message) to the eighth target network element. The fifth target message includes AI resource configuration information and may also include a first AI unit, used to guide AI resource nodes to process and forward AI service data. The eighth target network element is the aforementioned first network element.

[0329] The AI ​​resource configuration information includes at least one of the following:

[0330] Data source information (such as data source address);

[0331] AI business detection information;

[0332] AI business processing rules;

[0333] AS address.

[0334] The first AI unit in the fifth target message includes at least one of the following:

[0335] The executable file corresponding to the first AI unit;

[0336] The reference model structure indication for the first AI unit;

[0337] The order of model parameters in the first AI unit;

[0338] The parameter file for the first AI unit;

[0339] Instructions from the AI ​​unit of the first device;

[0340] Modification layer instructions for the first AI unit;

[0341] The layer indicator corresponding to the first AI unit (e.g., the first 1 to K layers).

[0342] Step (6b): The eighth target network element sends a sixth target message to the seventh target network element in response to the AI ​​resource configuration. The seventh target network element establishes an AI bearer.

[0343] Step (7): The seventh target network element sends a seventh target message to the first entity in response to the fourth target message, or the seventh target message may be called an AI resource request response message. The seventh target message contains information about the eighth target network element (such as an AI resource node), and the information of the eighth target network element (such as an AI resource node) includes at least one of the following:

[0344] IP address; ID information; Fully Qualified Domain Name (FQDN).

[0345] Step (8): Communication resource request and communication bearer establishment.

[0346] Step (9): The first entity sends an eighth target message to the first device. The eighth target message is used in response to the first target message to notify the UE of the result of the AI ​​service request, and includes, but is not limited to, at least one of the following: communication resource node and AI resource node information.

[0347] Step (10): The eighth target network element (such as an AI resource node) obtains intranet data from the tenth target network element (such as a data source node).

[0348] The eighth target network element (such as an AI resource node) determines the data source information, including but not limited to at least one of the following methods:

[0349] Obtain from AI resource allocation information;

[0350] Obtain from local configuration.

[0351] Step (11): The eighth target network element (such as an AI resource node) performs model inference for the first AI unit.

[0352] The eighth target network element uses the data collected in step (10) as the model input of the first AI unit, performs model inference on the first AI unit, and obtains the inference result.

[0353] Step (12): The eighth target network element (such as an AI resource node) sends AI service data to the first device. The AI ​​service data includes second information. The second information includes the inference result of the first AI unit. The first AI data is carried by user plane signaling.

[0354] The reasoning result of the first AI unit in the second information includes at least one of the following:

[0355] The output of the intermediate layer of the target AI unit has no physical meaning.

[0356] The physical meaning of the output of the first AI unit cannot be interpreted by the second device, but the physical meaning can be understood by the first device.

[0357] The physical meaning of the output of the first AI unit can be interpreted by both the second and first devices.

[0358] Step (13): The first device performs model inference for the second AI unit based on the second information and obtains the inference result. The second AI unit and the first AI unit are aligned.

[0359] The operation of the first device includes at least one of the following:

[0360] The first device uses the second information as the model input of the second AI unit, performs model reasoning on the second AI unit, and obtains the reasoning result;

[0361] The first device uses the second information and the data collected locally by the first device as model inputs to the second AI unit, performs model inference on the second AI unit, and obtains the inference result.

[0362] Example 2-2:

[0363] Example 2-1 sends the first AI unit in the control plane. This example, however, sends the first AI unit in the user plane, thus reducing signaling overhead in the control plane.

[0364] As shown in Figure 7, the signaling flow is as follows:

[0365] Step (1): The first device sends a first target message to the first entity. The first target message is used to request the allocation of corresponding resources for the AI ​​service. The first target message contains third information. The resources include at least one of communication resources, data resources, computing resources, and AI resources.

[0366] The first entity can be a newly established entity, such as a fifth target network element (e.g., AI service management function, or collaborative control function), or it can be co-located with an existing entity, such as co-located with a second target network element (e.g., communication management function), or co-located with a fourth target network element (e.g., policy billing function).

[0367] The first target message can be forwarded through the NAS interface.

[0368] The third information is the description information of the first AI unit, including at least one of the following:

[0369] Configure the environment, such as deep learning frameworks (TensorFlow, PyTorch, etc.) and supported library versions;

[0370] The computational cost of the first AI unit, such as FLOPs, MFLOPs, GFLOPs, etc.

[0371] The inference latency required for the first AI unit;

[0372] The computing speed required for the first AI unit;

[0373] Storage requirements for the first AI unit;

[0374] The first AI unit's model input information indicators, such as data items, data format, data precision, etc.

[0375] The effective duration of the first AI unit (optional);

[0376] The first AI unit's model output information indicates, such as output data precision, data dimensions, and storage format (optional).

[0377] Step (2): The first entity and the fourth target network element (e.g., policy charging function) determine the AI ​​policy.

[0378] Step (3a): The first entity sends a second target message to the ninth target network element (e.g., data plane management function) to request the data source and determine the data source information.

[0379] The second target message is determined based on third information. For example, the relevant information for the data source request can be determined based on the model input information of the first AI unit.

[0380] In one implementation, the AI ​​business management function selects the data plane management function instance, considering at least one of the following factors:

[0381] DNN; S-NSSAI; AI business identifier; supported AI services; service area information; data source information.

[0382] Step (3b): The ninth target network element (e.g., data plane management function) sends a third target message to the first entity and responds to the second target message to indicate the data source information.

[0383] Step (4): The first entity sends a fourth target message to the seventh target network element (such as the AI ​​resource management function). The fourth target message contains AI resource request information and may also include third information. The seventh target network element (such as the AI ​​resource management function) determines a suitable AI resource node based on the AI ​​resource request information and the AI ​​strategy.

[0384] The third information is the description information of the first AI unit, including at least one of the following:

[0385] Configure the environment, such as deep learning frameworks (TensorFlow, PyTorch, etc.) and supported library versions;

[0386] The computational cost of the first AI unit, such as FLOPs, MFLOPs, GFLOPs, etc.

[0387] The inference latency required for the first AI unit;

[0388] The computing speed required for the first AI unit;

[0389] Storage requirements for the first AI unit;

[0390] The first AI unit's model input information indicators, such as data items, data format, data precision, etc.

[0391] The effective duration of the first AI unit (optional);

[0392] The first AI unit's model output information indicates, such as output data precision, data dimensions, and storage format (optional).

[0393] The fourth target message is determined based on the third information; specifically, it can be determined based on the model output information of the first AI unit. The fourth target message can be used to indicate the data output information of the seventh network.

[0394] In one implementation, the fourth target message includes information such as the configuration environment, the computational load of the first AI unit, and the storage requirements of the first AI unit, which helps the seventh target network element configure the computing power and storage resources required for the inference of the first AI unit.

[0395] The AI ​​resource request information includes at least one of the following:

[0396] Model information (such as the model or its download address);

[0397] Data source information (data source address);

[0398] AI business computational load (computing power size, type);

[0399] AS address.

[0400] Step (5): The seventh target network element (such as the AI ​​resource management function) selects the appropriate eighth target network element (such as the AI ​​resource node) based on the AI ​​resource request information.

[0401] Step (6a): The seventh target network element (such as the AI ​​resource management function) sends a fifth target message to the eighth target network element. The fifth target message includes AI resource configuration information and may also include third information, which is used to guide the AI ​​resource nodes to process and forward AI service data.

[0402] The AI ​​resource configuration information includes at least one of the following:

[0403] Data source information (such as data source address);

[0404] AI business detection information;

[0405] AI business processing rules;

[0406] AS address.

[0407] The third information is the description information of the first AI unit, including at least one of the following:

[0408] Configure the environment, such as deep learning frameworks (TensorFlow, PyTorch, etc.) and supported library versions;

[0409] The computational cost of the first AI unit, such as FLOPs, MFLOPs, GFLOPs, etc.

[0410] The inference latency required for the first AI unit;

[0411] The computing speed required for the first AI unit;

[0412] Storage requirements for the first AI unit;

[0413] The first AI unit's model input information indicators, such as data items, data format, data precision, etc.

[0414] The effective duration of the first AI unit (optional);

[0415] The first AI unit's model output information indicates, such as output data precision, data dimensions, and storage format (optional).

[0416] Step (6b): The eighth target network element sends a sixth target message to the seventh target network element in response to the AI ​​resource configuration. The seventh target network element establishes an AI bearer.

[0417] Steps (7) to (9): Same as steps (7) to (9) in Example 2-1.

[0418] Step (10): The first device sends the first AI service data (i.e., the aforementioned second message) to the eighth target network element (such as an AI resource node). The first AI service data includes the first AI unit. The first AI data is carried by user plane signaling.

[0419] The first AI unit in the first AI business data includes at least one of the following:

[0420] The executable file corresponding to the first AI unit;

[0421] The reference model structure indication for the first AI unit;

[0422] The order of model parameters in the first AI unit;

[0423] The parameter file for the first AI unit;

[0424] Instructions from the AI ​​unit of the first device;

[0425] Modification layer instructions for the first AI unit;

[0426] The layer indicator corresponding to the first AI unit (e.g., the first 1 to K layers).

[0427] Step (11): Same as step (10) in Example 2-1;

[0428] Step (12): Same as step (11) in Example 2-1;

[0429] Step (13): Same as step (12) in Example 2-1;

[0430] Step (14): Same as step (13) in Example 2-1.

[0431] Example 2-3:

[0432] Example 2-1 transmits the first AI unit in the control plane. This example, however, transmits the first AI unit in the data plane, thereby reducing signaling overhead in the control plane.

[0433] As shown in Figure 7, the signaling flow is as follows:

[0434] Steps (1) to (9): Same as steps (1) to (9) in Example 2-2.

[0435] Step (10): The first device sends the first AI service data (i.e., the aforementioned third message) to the eighth target network element (such as an AI resource node). The first AI service data includes the first AI unit. The first AI data is carried by data plane signaling.

[0436] Steps (11) to (12): Same as steps (11) to (12) in Example 2-2.

[0437] Step (13): The eighth target network element (such as an AI resource node, i.e., the aforementioned first network element) sends AI service data to the first device. The AI ​​service data includes second information. The second information includes the inference result of the first AI unit. The first AI data is carried by data plane signaling.

[0438] Step (14): Same as step (14) in Example 2-2.

[0439] Example 2-4:

[0440] This example describes how the first AI unit is obtained from the AS (via the control plane).

[0441] As shown in Figure 8, the signaling flow is as follows:

[0442] Steps (1) to (3): Same as steps (1) to (3) in Example 2-2 of Figure 7.

[0443] Step (4a): The first entity sends a ninth target message to the eleventh target network element (e.g., AS). The ninth target message is used to obtain the first AI unit related to the description information of the first AI unit from the AS. The ninth target information includes at least one of the following: AS address, description information of the first AI unit.

[0444] Step (4b): The first entity obtains the tenth target message from the eleventh target network element. The tenth target message is used to respond to the ninth target message, which includes the first AI unit.

[0445] Steps (5) to (14): Same as steps (4) to (13) in Example 2-1 of Figure 6.

[0446] Example 2-5:

[0447] This example describes how the first AI unit is obtained from the AS (obtained via the user plane).

[0448] As shown in Figure 9, the signaling flow is as follows:

[0449] Steps (1) to (3): Same as steps (1) to (3) in Example 2-2 of Figure 7.

[0450] Step (4): The first entity sends a ninth target message to the eleventh target network element (e.g., AS). The ninth target message is used to obtain the first AI unit related to the description information of the first AI unit from the AS. The ninth target information includes at least one of the following: AS address, description information of the first AI unit.

[0451] Steps (5) to (10): Same as steps (4) to (9) in Example 2-2 of Figure 7.

[0452] Step (11): The first entity obtains the tenth target message from the eleventh target network element. The tenth target message is used to respond to the ninth target message, and the ninth target message includes the first AI unit. The eleventh message is carried by user plane signaling.

[0453] Steps (12) to (15): Same as steps (11) to (14) in Example 2-2 of Figure 7.

[0454] Example 2-6:

[0455] This example describes how the first AI unit is obtained from the AS (via the data plane).

[0456] As shown in Figure 9, the signaling flow is as follows:

[0457] Steps (1) to (10): Same as steps (1) to (10) in Example 2-5 of Figure 8.

[0458] Step (11): The first entity obtains the tenth target message from the eleventh target network element. The tenth target message is used to respond to the ninth target message, and the ninth target message includes the first AI unit. The eleventh message is carried by data plane signaling.

[0459] Steps (12) to (15): Same as steps (11) to (14) in Example 2-3 of Figure 7.

[0460] Example 3:

[0461] This example describes a SLAM scenario that expands on the amount of interaction.

[0462] Examples 1 and 2 are processes that do not distinguish between use cases, while this example is for environment reconstruction, or the Simultaneous Localization and Mapping (SLAM) use case. It provides a detailed description of the interaction volume, thereby enabling specific typical use cases and avoiding the problem that specific use cases have special requirements that cannot be provided by the network.

[0463] As shown in Figure 10, the signaling flow is as follows:

[0464] Step (1): The first device sends a first target message to the first entity. The first target message is used to request the allocation of corresponding resources for the AI ​​service. The first target message contains third information. The resources include at least one of communication resources, data resources, computing resources, and AI resources.

[0465] The first entity can be a newly established entity, such as a fifth target network element (e.g., AI service management function, or collaborative control function), or it can be co-located with an existing entity, such as co-located with a second target network element (e.g., communication management function), or co-located with a fourth target network element (e.g., policy billing function).

[0466] The first target message can be forwarded through the NAS interface.

[0467] The third information is the description information of the first AI unit, including at least one of the following:

[0468] Configure the environment, such as deep learning frameworks (TensorFlow, PyTorch, etc.) and supported library versions;

[0469] The computational cost of the first AI unit, such as FLOPs, MFLOPs, GFLOPs, etc.

[0470] The inference latency required for the first AI unit;

[0471] The computing speed required for the first AI unit;

[0472] Storage requirements for the first AI unit;

[0473] The first AI unit's model input information indicators, such as data items, data format, data precision, etc.

[0474] The functions of the first AI unit;

[0475] The effective duration of the first AI unit (optional);

[0476] The first AI unit's model output information includes, for example, data items, output data precision, data dimensions, and storage format (optional).

[0477] Among them, the model input information indication and the model output information indication of the first AI unit are related to specific use cases.

[0478] The methods for indicating model input information or model output information include:

[0479] Method 1: Predefined

[0480] The agreement stipulates, or the first device and the first entity agree in advance on the data parameters, data dimensions, data precision, etc., corresponding to a certain AI function.

[0481] For example, when the function of the first AI unit is equal to SLAM, the first entity can know:

[0482] The parameters input to the model include chromaticity;

[0483] The dimensions of the model input include "number of channels, height, width, and frame", or "Channel x Height x Width x Frame".

[0484] Specifically, each dimension contains the following:

[0485] Channels: This refers to the number of channels in an image. Color images typically have three channels (red, green, and blue, i.e., RGB), while grayscale images have only one channel. In deep learning, additional channels are sometimes used to represent other information, such as depth information or heatmaps.

[0486] Height: Represents the height of the image, i.e., the number of pixels from the top to the bottom.

[0487] Width: Represents the width of the image, that is, the number of pixels from left to right.

[0488] Frame: Represents a frame in a video; each frame is an image.

[0489] The data size of the model input is 3x480x640x10, which means that the number of channels is 3 (three color channels (RGB)), the height is 480, the width is 640, and the number of frames is 10.

[0490] The data precision is 8 bits, meaning that each pixel has 8 bits of data.

[0491] Method 2: Predefined + partially explicit instructions

[0492] For example, the function of the first AI unit is SLAM, and the data precision is 10 bits.

[0493] When the function of the first AI unit is set to SLAM, the first entity can obtain a default reference configuration, as shown in Method 1. At this time, the default data precision is 8 bits. The data precision can be updated to 10 bits by explicitly indicating that the data precision is 10 bits.

[0494] For example, when the use case is environment refactoring:

[0495] The data items in the model input information indication of the first AI unit may include at least one of the following:

[0496] Color; brightness.

[0497] The data dimension in the model input information indication of the first AI unit may include at least one of the following:

[0498] Number of channels; height; width; number of frames.

[0499] Optionally, when the model output of the first AI unit has physical meaning:

[0500] The data items in the model output information indication of the first AI unit may include at least one of the following:

[0501] Time point; obstacle type; obstacle's 3D coordinates; obstacle's envelope; obstacle's volume; obstacle's movement speed; obstacle's movement direction; distance between the obstacle and the target object.

[0502] For example, when the use case is trajectory planning:

[0503] The data items in the model input information indication of the first AI unit may include at least one of the following:

[0504] Color; brightness.

[0505] The data dimension in the model input information indication of the first AI unit may include at least one of the following:

[0506] Number of channels; height; width; number of frames.

[0507] Optionally, when the model output of the first AI unit has physical meaning:

[0508] The data items in the model output information indication of the first AI unit may include at least one of the following:

[0509] Point in time; speed; direction of movement.

[0510] When the use case is a driving decision:

[0511] The data items in the model input information indication of the first AI unit may include at least one of the following:

[0512] Color; brightness.

[0513] The data dimension in the model input information indication of the first AI unit may include at least one of the following:

[0514] Number of channels; height; width; number of frames.

[0515] Optionally, when the model output of the first AI unit has physical meaning:

[0516] The data items in the model output information indication of the first AI unit may include at least one of the following:

[0517] Timing; throttle opening; braking force; steering angle.

[0518] Step (2): Same as Example 2-2.

[0519] Step (3a): The first entity sends a second target message to the ninth target network element (e.g., data plane management function) to request the data source and determine the data source information.

[0520] The second target message is determined based on third information. For example, the relevant information for the data source request can be determined based on the model input information of the first AI unit.

[0521] When the use case involves environment reconfiguration, trajectory planning, and control decision-making:

[0522] The data items in the model input information indication of the first AI unit may include at least one of the following:

[0523] Color; brightness.

[0524] The data dimension in the model input information indication of the first AI unit may include at least one of the following:

[0525] Number of channels; height; width; number of frames.

[0526] Step (3b): The ninth target network element (e.g., data plane management function) sends a third target message to the first entity and responds to the second target message to indicate the data source information.

[0527] Step (4): The first entity sends a fourth target message to the seventh target network element (such as the AI ​​resource management function). The fourth target message contains AI resource request information and may also include third information. The seventh target network element (such as the AI ​​resource management function) determines a suitable AI resource node based on the AI ​​resource request information and the AI ​​strategy.

[0528] The third information is the description information of the first AI unit, including at least one of the following:

[0529] Configure the environment, such as deep learning frameworks (TensorFlow, PyTorch, etc.) and supported library versions;

[0530] The computational cost of the first AI unit, such as FLOPs, MFLOPs, GFLOPs, etc.

[0531] The inference latency required for the first AI unit;

[0532] The computing speed required for the first AI unit;

[0533] Storage requirements for the first AI unit;

[0534] The first AI unit's model input information indicators, such as data items, data format, data precision, etc.

[0535] The effective duration of the first AI unit (optional);

[0536] The first AI unit's model output information indicates, such as output data precision, data dimensions, and storage format (optional).

[0537] The fourth target message is determined based on the third information. Specifically, the fourth target message can be determined based on the model output information of the first AI unit. The fourth target message can be used to indicate the data output information of the seventh network.

[0538] In one implementation, the fourth target message includes information such as the configuration environment, the computational load of the first AI unit, and the storage requirements of the first AI unit, which helps the seventh target network element configure the computing power and storage resources required for the inference of the first AI unit.

[0539] When the use case involves environment refactoring:

[0540] Optionally, when the model output of the first AI unit has physical meaning:

[0541] The data items in the model output information indication of the first AI unit may include at least one of the following:

[0542] Time point; obstacle type; obstacle's 3D coordinates; obstacle's envelope; obstacle's volume; obstacle's movement speed; obstacle's movement direction; distance between the obstacle and the target object.

[0543] When the use case is trajectory planning:

[0544] Optionally, when the model output of the first AI unit has physical meaning, the data items in the model output information indication of the first AI unit may include at least one of the following:

[0545] Point in time; speed; direction of movement.

[0546] When the use case is a driving decision:

[0547] Optionally, when the model output of the first AI unit has physical meaning, the data items in the model output information indication of the first AI unit may include at least one of the following:

[0548] Timing; throttle opening; braking force; steering angle.

[0549] Step (5): The seventh target network element (such as the AI ​​resource management function) selects the appropriate eighth target network element (such as the AI ​​resource node) based on the AI ​​resource request information.

[0550] Step (6a): The seventh target network element (such as the AI ​​resource management function) sends a fifth target message to the eighth target network element. The fifth target message includes AI resource configuration information and may also include third information, which is used to guide the AI ​​resource nodes to process and forward AI service data.

[0551] The AI ​​resource configuration information includes at least one of the following:

[0552] Data source information (such as data source address); AI business detection information; AI business processing rules; AS address.

[0553] The third information is the description information of the first AI unit, including at least one of the following:

[0554] Configuration environment: deep learning frameworks (TensorFlow, PyTorch, etc.), supported library versions, etc.;

[0555] The computational cost of the first AI unit, such as FLOPs, MFLOPs, GFLOPs, etc.

[0556] The inference latency required for the first AI unit;

[0557] The computing speed required for the first AI unit;

[0558] Storage requirements for the first AI unit;

[0559] The model input information for the first AI unit includes: data items, data format, data precision, etc.

[0560] The effective duration of the first AI unit (optional);

[0561] The model output information of the first AI unit indicates: output data precision, data dimensions, and storage format.

[0562] When the use case involves environment refactoring:

[0563] Optionally, when the model output of the first AI unit has physical meaning, the data items in the model output information indication of the first AI unit may include at least one of the following:

[0564] Time point; obstacle type; obstacle's 3D coordinates; obstacle's envelope; obstacle's volume; obstacle's movement speed; obstacle's movement direction; distance between the obstacle and the target object.

[0565] When the use case is trajectory planning:

[0566] Optionally, when the model output of the first AI unit has physical meaning, the data items in the model output information indication of the first AI unit may include at least one of the following:

[0567] Point in time; speed; direction of movement.

[0568] When the use case is a driving decision:

[0569] Optionally, when the model output of the first AI unit has physical meaning, the data items in the model output information indication of the first AI unit may include at least one of the following:

[0570] Timing; throttle opening; braking force; steering angle.

[0571] Step (6b): The eighth target network element sends a sixth target message to the seventh target network element in response to the AI ​​resource configuration. The seventh target network element establishes an AI bearer.

[0572] Steps (7) to (12): Same as Example 2-2.

[0573] Step (13): The eighth target network element (such as an AI resource node) sends AI service data to the first device. The AI ​​service data includes second information. The second information includes the inference results of the first AI unit.

[0574] The reasoning result of the first AI unit in the second information includes at least one of the following:

[0575] The output of the intermediate layer of the target AI unit has no physical meaning.

[0576] The physical meaning of the output of the first AI unit cannot be interpreted by the second device, but the physical meaning can be understood by the first device.

[0577] The physical meaning of the output of the first AI unit can be interpreted by both the second and first devices.

[0578] When the use case involves environment refactoring:

[0579] Optionally, when the model output of the first AI unit has physical meaning, the second information includes at least one of the following:

[0580] Time point; obstacle type; obstacle's 3D coordinates; obstacle's envelope; obstacle's volume; obstacle's movement speed; obstacle's movement direction; distance between the obstacle and the target object.

[0581] When the use case is trajectory planning:

[0582] Optionally, when the model output of the first AI unit has physical meaning, the second information includes at least one of the following:

[0583] Point in time; speed; direction of movement.

[0584] When the use case is a driving decision:

[0585] Optionally, when the model output of the first AI unit has physical meaning, the second information includes at least one of the following:

[0586] Timing; throttle opening; braking force; steering angle.

[0587] Step (14): Same as step (13) in Example 2-1.

[0588] This application provides a method for offloading terminal computing power while simultaneously utilizing network data to provide AI services, thereby reducing terminal inference latency and improving inference accuracy. By having the terminal provide a first AI unit for model inference via the network, the problem of extensive information interaction required when the terminal and network each provide AI units can be solved.

[0589] Referring to Figure 11, which is a flowchart of a model inference method provided in an embodiment of this application, the model inference method includes the following steps:

[0590] Step 201: The second device receives the first information sent by the first device, and determines the first AI unit based on the first information;

[0591] Step 202: The second device uses the first AI unit to perform reasoning on the data collected by the second device to obtain the reasoning result of the first AI unit;

[0592] Step 203: The second device sends the inference result of the first AI unit to the first device;

[0593] The reasoning result of the first AI unit is used to perform reasoning to obtain the reasoning result of the second AI unit, and the first AI unit and the second AI unit are aligned.

[0594] Optionally, the alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions:

[0595] The first AI unit and the second AI unit are sub-units of the target AI unit;

[0596] The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

[0597] Optionally, the first information includes at least one of the following:

[0598] The model information of the first AI unit; the description information of the first AI unit;

[0599] The model information of the first AI unit includes at least one of the following:

[0600] The executable file corresponding to the first AI unit;

[0601] Information used to indicate the reference model structure of the first AI unit;

[0602] Information used to indicate the order of model parameters of the first AI unit;

[0603] The parameter information of the first AI unit;

[0604] The identifier of the first AI unit;

[0605] Information used to indicate the parameter modification layer of the first AI unit;

[0606] Information used to indicate the layer corresponding to the first AI unit.

[0607] Optionally, the reasoning result of the first AI unit satisfies any one of the following conditions:

[0608] The reasoning result of the first AI unit does not have corresponding physical parameters;

[0609] The reasoning results of the first AI unit have corresponding physical parameters.

[0610] Optionally, the first information is carried by at least one of the following:

[0611] Control plane messages; User plane messages; Data plane messages.

[0612] Optionally, the control plane message includes at least one of the following:

[0613] The first message is used to request the allocation of resources for AI services;

[0614] The fourth message is used to request AI resources;

[0615] The fifth message is used for AI resource allocation;

[0616] And / or,

[0617] The user plane message includes a second message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0618] And / or,

[0619] The data plane message includes a third message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0620] Optionally, the second device includes a first network element, and the second device receives first information sent by the first device, including at least one of the following:

[0621] The first network element receives the fifth message sent by the second network element, the fifth message being determined based on the fourth message sent by the first entity, and the fourth message being determined based on the first message sent by the first device;

[0622] The first network element receives the second message sent by the first device;

[0623] The first network element receives the third message sent by the first device.

[0624] Optionally, when the first information includes the description information of the first AI unit, the description information of the first AI unit includes at least one of the following:

[0625] The effective duration of the first AI unit;

[0626] The configuration environment of the first AI unit;

[0627] The model input information of the first AI unit;

[0628] The model outputs relevant information from the first AI unit;

[0629] The computational load of the first AI unit;

[0630] The inference latency required by the first AI unit;

[0631] The required computing speed of the first AI unit;

[0632] The storage requirements of the first AI unit;

[0633] The function of the first AI unit;

[0634] Information used to indicate whether a second device is needed for data collection;

[0635] Used to obtain information about the application service AS of the first AI unit;

[0636] Wherein, the model input information of the first AI unit is associated with the function of the first AI unit; and / or, the model output information of the first AI unit is associated with the function of the first AI unit.

[0637] Optionally, the first AI unit may include at least one of the following functions:

[0638] Environment reconstruction; trajectory planning; driving decision; target recognition; target detection; target tracking; intrusion detection.

[0639] Optionally, when the function of the first AI unit includes environment reconstruction, the model input information of the first AI unit includes at least one of the following: a first data item, the data dimension of the first data item; the model output information of the first AI unit includes a second data item; and / or

[0640] When the function of the first AI unit includes trajectory planning, the model input information of the first AI unit includes at least one of the following: a third data item, wherein the data dimension of the third data item; the model output information of the first AI unit includes a fourth data item; and / or

[0641] When the function of the first AI unit includes driving decision-making, the model input information of the first AI unit includes at least one of the following: a fifth data item, wherein the data dimension of the fifth data item; the model output information of the first AI unit includes a sixth data item.

[0642] The first data item includes at least one of the following: chromaticity and luminance;

[0643] The data dimensions of the first data item include at least one of the following: number of channels, height, width, and number of frames;

[0644] The second data item includes at least one of the following: time point, obstacle type, three-dimensional coordinates of the obstacle, envelope of the obstacle, volume of the obstacle, moving speed of the obstacle, moving direction of the obstacle, and distance between the obstacle and the target object;

[0645] The third data item includes at least one of the following: chromaticity, luminance;

[0646] The data dimensions of the third data item include at least one of the following: number of channels, height, width, and number of frames;

[0647] The fourth data item includes at least one of the following: time point, speed, and direction of movement;

[0648] The fifth data item includes at least one of the following: chromaticity, luminance;

[0649] The data dimensions of the fifth data item include at least one of the following: number of channels, height, width, and number of frames;

[0650] The sixth data item includes at least one of the following: time point, throttle opening, braking force, and steering angle.

[0651] Optionally, the second device includes a first entity, and when the first information includes description information of the first AI unit, determining the first AI unit based on the first information includes:

[0652] The first entity obtains the model information of the first AI unit from the third network element based on the first information;

[0653] The first AI unit is determined based on the model information of the first AI unit;

[0654] The model information of the first AI unit includes at least one of the following:

[0655] The executable file corresponding to the first AI unit;

[0656] Information used to indicate the reference model structure of the first AI unit;

[0657] Information used to indicate the order of model parameters of the first AI unit;

[0658] The parameter information of the first AI unit;

[0659] The identifier of the first AI unit;

[0660] Information used to indicate the parameter modification layer of the first AI unit;

[0661] Information used to indicate the layer corresponding to the first AI unit.

[0662] In one implementation, the first entity can obtain the model information of the first AI unit from a third network element (e.g., an OTT server or AF) based on first information (such as the description information of the first AI unit).

[0663] It should be noted that this embodiment is an implementation of the second device corresponding to the embodiment shown in FIG2. The specific implementation can be found in the relevant description of the embodiment shown in FIG2. To avoid repeated description, this embodiment will not be repeated.

[0664] The model inference method provided in this application can be executed by a model inference device. This application uses the execution of the model inference method by a model inference device as an example to illustrate the model inference device provided in this application.

[0665] This application provides a model inference device. As an example, the model inference device can be a communication device or a component within a communication device, such as a chip. The communication device can be a terminal, a network-side device, or a server, etc. Exemplarily, the terminal can be, but is not limited to, the type of terminal 11 listed above, and the network-side device can be, but is not limited to, the type of network-side device 12 listed above. This application does not impose specific limitations.

[0666] The model inference device includes a receiving module, a transmitting module, and a processing module. These modules can be implemented in software or hardware. When implemented in hardware, the processing module can be implemented by a processor. For example, the processor can include general-purpose processors, special-purpose processors, such as a Central Processing Unit (CPU), microprocessor, Digital Signal Processor (DSP), Artificial Intelligence (AI) processor, Graphics Processing Unit (GPU), Application Specific Integrated Circuit (ASIC), Network Processor (NP), Field Programmable Gate Array (FPGA), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The receiving and transmitting modules can be implemented by a communication interface, which can include one or more of the following: transceiver, pins, circuits, bus, radio frequency unit, etc.

[0667] Specifically, referring to Figure 12, when the model inference device is a first device or a component of the first device, the model inference device 300 includes:

[0668] Sending module 301 is used to send first information to the second device, the first information being used to determine the first artificial intelligence (AI) unit;

[0669] The receiving module 302 is used to receive the inference result of the first AI unit sent by the second device. The inference result of the first AI unit is the inference result obtained by using the first AI unit to infer the data collected by the second device.

[0670] The processing module 303 is used to perform reasoning based on the reasoning result of the first AI unit and the second AI unit to obtain the reasoning result of the second AI unit, wherein the first AI unit and the second AI unit are aligned.

[0671] Optionally, the alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions:

[0672] The first AI unit and the second AI unit are sub-units of the target AI unit;

[0673] The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

[0674] Optionally, the first information includes at least one of the following:

[0675] The model information of the first AI unit; the description information of the first AI unit;

[0676] The model information of the first AI unit includes at least one of the following:

[0677] The executable file corresponding to the first AI unit;

[0678] Information used to indicate the reference model structure of the first AI unit;

[0679] Information used to indicate the order of model parameters of the first AI unit;

[0680] The parameter information of the first AI unit;

[0681] The identifier of the first AI unit;

[0682] Information used to indicate the parameter modification layer of the first AI unit;

[0683] Information used to indicate the layer corresponding to the first AI unit.

[0684] Optionally, the reasoning result of the first AI unit satisfies any one of the following conditions:

[0685] The reasoning result of the first AI unit does not have corresponding physical parameters;

[0686] The reasoning results of the first AI unit have corresponding physical parameters.

[0687] Optionally, the processing module is specifically used for any of the following:

[0688] The reasoning result of the first AI unit is used as the model input of the second AI unit for reasoning;

[0689] The reasoning results of the first AI unit and the data collected by the first device are used as model inputs for the second AI unit for reasoning.

[0690] Optionally, the first information is carried by at least one of the following:

[0691] Control plane messages; User plane messages; Data plane messages.

[0692] Optionally, the control plane message includes a first message, which is used to request the allocation of resources for AI services;

[0693] And / or,

[0694] The user plane message includes a second message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0695] And / or,

[0696] The data plane message includes a third message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0697] Optionally, the second device includes a first network element, and the transmitting module is specifically used for at least one of the following:

[0698] The first entity sends first information to the first network element, and the first message is a message transmitted between the first device and the first entity;

[0699] Send the second message to the first network element;

[0700] The third message is sent to the first network element.

[0701] Optionally, when the first information includes the description information of the first AI unit, the description information of the first AI unit includes at least one of the following:

[0702] The effective duration of the first AI unit;

[0703] The configuration environment of the first AI unit;

[0704] The model input information of the first AI unit;

[0705] The model outputs relevant information from the first AI unit;

[0706] The computational load of the first AI unit;

[0707] The inference latency required by the first AI unit;

[0708] The required computing speed of the first AI unit;

[0709] The storage requirements of the first AI unit;

[0710] The function of the first AI unit;

[0711] Information used to indicate whether a second device is needed for data collection;

[0712] Used to obtain information about the application service AS of the first AI unit;

[0713] Wherein, the model input information of the first AI unit is associated with the function of the first AI unit; and / or, the model output information of the first AI unit is associated with the function of the first AI unit.

[0714] Optionally, the first AI unit may include at least one of the following functions:

[0715] Environment reconstruction; trajectory planning; driving decision; target recognition; target detection; target tracking; intrusion detection.

[0716] Optionally, when the function of the first AI unit includes environment reconstruction, the model input information of the first AI unit includes at least one of the following: a first data item, the data dimension of the first data item; the model output information of the first AI unit includes a second data item; and / or

[0717] When the function of the first AI unit includes trajectory planning, the model input information of the first AI unit includes at least one of the following: a third data item, wherein the data dimension of the third data item; the model output information of the first AI unit includes a fourth data item; and / or

[0718] When the function of the first AI unit includes driving decision-making, the model input information of the first AI unit includes at least one of the following: a fifth data item, wherein the data dimension of the fifth data item; the model output information of the first AI unit includes a sixth data item.

[0719] The first data item includes at least one of the following: chromaticity and luminance;

[0720] The data dimensions of the first data item include at least one of the following: number of channels, height, width, and number of frames;

[0721] The second data item includes at least one of the following: time point, obstacle type, three-dimensional coordinates of the obstacle, envelope of the obstacle, volume of the obstacle, moving speed of the obstacle, moving direction of the obstacle, and distance between the obstacle and the target object;

[0722] The third data item includes at least one of the following: chromaticity, luminance;

[0723] The data dimensions of the third data item include at least one of the following: number of channels, height, width, and number of frames;

[0724] The fourth data item includes at least one of the following: time point, speed, and direction of movement;

[0725] The fifth data item includes at least one of the following: chromaticity, luminance;

[0726] The data dimensions of the fifth data item include at least one of the following: number of channels, height, width, and number of frames;

[0727] The sixth data item includes at least one of the following: time point, throttle opening, braking force, and steering angle.

[0728] Referring to Figure 13, when the model inference device is a second device or a component of a second device, the model inference device 400 includes:

[0729] The receiving module 401 is used to receive first information sent by the first device and determine the first AI unit based on the first information;

[0730] The processing module 402 is used to perform reasoning on the data collected by the second device using the first AI unit to obtain the reasoning result of the first AI unit;

[0731] The sending module 403 is used to send the reasoning result of the first AI unit to the first device;

[0732] The reasoning result of the first AI unit is used to perform reasoning to obtain the reasoning result of the second AI unit, and the first AI unit and the second AI unit are aligned.

[0733] Optionally, the alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions:

[0734] The first AI unit and the second AI unit are sub-units of the target AI unit;

[0735] The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

[0736] Optionally, the first information includes at least one of the following:

[0737] The model information of the first AI unit; the description information of the first AI unit;

[0738] The model information of the first AI unit includes at least one of the following:

[0739] The executable file corresponding to the first AI unit;

[0740] Information used to indicate the reference model structure of the first AI unit;

[0741] Information used to indicate the order of model parameters of the first AI unit;

[0742] The parameter information of the first AI unit;

[0743] The identifier of the first AI unit;

[0744] Information used to indicate the parameter modification layer of the first AI unit;

[0745] Information used to indicate the layer corresponding to the first AI unit.

[0746] Optionally, the reasoning result of the first AI unit satisfies any one of the following conditions:

[0747] The reasoning result of the first AI unit does not have corresponding physical parameters;

[0748] The reasoning results of the first AI unit have corresponding physical parameters.

[0749] Optionally, the first information is carried by at least one of the following:

[0750] Control plane messages; User plane messages; Data plane messages.

[0751] Optionally, the control plane message includes at least one of the following:

[0752] The first message is used to request the allocation of resources for AI services;

[0753] The fourth message is used to request AI resources;

[0754] The fifth message is used for AI resource allocation;

[0755] And / or,

[0756] The user plane message includes a second message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0757] And / or,

[0758] The data plane message includes a third message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0759] Optionally, the second device includes a first network element, and the receiving module is specifically used for at least one of the following:

[0760] The fifth message is received from the second network element, the fifth message being determined based on the fourth message sent by the first entity, and the fourth message being determined based on the first message sent by the first device.

[0761] Receive the second message sent by the first device;

[0762] Receive the third message sent by the first device.

[0763] Optionally, when the first information includes the description information of the first AI unit, the description information of the first AI unit includes at least one of the following:

[0764] The effective duration of the first AI unit;

[0765] The configuration environment of the first AI unit;

[0766] The model input information of the first AI unit;

[0767] The model outputs relevant information from the first AI unit;

[0768] The computational load of the first AI unit;

[0769] The inference latency required by the first AI unit;

[0770] The required computing speed of the first AI unit;

[0771] The storage requirements of the first AI unit;

[0772] The function of the first AI unit;

[0773] Information used to indicate whether a second device is needed for data collection;

[0774] Used to obtain information about the application service AS of the first AI unit;

[0775] Wherein, the model input information of the first AI unit is associated with the function of the first AI unit; and / or, the model output information of the first AI unit is associated with the function of the first AI unit.

[0776] Optionally, the first AI unit may include at least one of the following functions:

[0777] Environment reconstruction; trajectory planning; driving decision; target recognition; target detection; target tracking; intrusion detection.

[0778] Optionally, when the function of the first AI unit includes environment reconstruction, the model input information of the first AI unit includes at least one of the following: a first data item, the data dimension of the first data item; the model output information of the first AI unit includes a second data item; and / or

[0779] When the function of the first AI unit includes trajectory planning, the model input information of the first AI unit includes at least one of the following: a third data item, wherein the data dimension of the third data item; the model output information of the first AI unit includes a fourth data item; and / or

[0780] When the function of the first AI unit includes driving decision-making, the model input information of the first AI unit includes at least one of the following: a fifth data item, wherein the data dimension of the fifth data item; the model output information of the first AI unit includes a sixth data item.

[0781] The first data item includes at least one of the following: chromaticity and luminance;

[0782] The data dimensions of the first data item include at least one of the following: number of channels, height, width, and number of frames;

[0783] The second data item includes at least one of the following: time point, obstacle type, three-dimensional coordinates of the obstacle, envelope of the obstacle, volume of the obstacle, moving speed of the obstacle, moving direction of the obstacle, and distance between the obstacle and the target object;

[0784] The third data item includes at least one of the following: chromaticity, luminance;

[0785] The data dimensions of the third data item include at least one of the following: number of channels, height, width, and number of frames;

[0786] The fourth data item includes at least one of the following: time point, speed, and direction of movement;

[0787] The fifth data item includes at least one of the following: chromaticity, luminance;

[0788] The data dimensions of the fifth data item include at least one of the following: number of channels, height, width, and number of frames;

[0789] The sixth data item includes at least one of the following: time point, throttle opening, braking force, and steering angle.

[0790] Optionally, the second device includes a first entity, and when the first information includes the description information of the first AI unit, the processing module is specifically used for:

[0791] The first entity obtains the model information of the first AI unit from the third network element based on the first information;

[0792] The first AI unit is determined based on the model information of the first AI unit;

[0793] The model information of the first AI unit includes at least one of the following:

[0794] The executable file corresponding to the first AI unit;

[0795] Information used to indicate the reference model structure of the first AI unit;

[0796] Information used to indicate the order of model parameters of the first AI unit;

[0797] The parameter information of the first AI unit;

[0798] The identifier of the first AI unit;

[0799] Information used to indicate the parameter modification layer of the first AI unit;

[0800] Information used to indicate the layer corresponding to the first AI unit.

[0801] The model inference apparatus provided in this application embodiment can implement the various processes implemented in the method embodiments of Figures 2 to 11 and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0802] As shown in Figure 14, this application embodiment also provides a communication device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. For example, when the communication device 500 is a first device, when the program or instructions are executed by the processor 501, they implement the various steps of the above-described model inference method embodiment and achieve the same technical effect. When the communication device 500 is a second device, when the program or instructions are executed by the processor 501, they implement the various steps of the above-described model inference method embodiment and achieve the same technical effect. To avoid repetition, this will not be described again here.

[0803] This application also provides a terminal, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps in the method embodiment shown in FIG2. This terminal embodiment corresponds to the above-described first device-side method embodiment, and all implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and can achieve the same technical effect. The terminal may be the model inference device shown in FIG12. Specifically, FIG15 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of this application.

[0804] The terminal 600 includes, but is not limited to, at least some of the following components: radio frequency unit 601, network module 602, audio output unit 603, input unit 604, sensor 605, display unit 606, user input unit 607, interface unit 608, memory 609, and processor 610.

[0805] Those skilled in the art will understand that terminal 600 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to processor 610 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The terminal structure shown in Figure 15 does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0806] It should be understood that, in this embodiment, the input unit 604 may include a graphics processor 6041 and a microphone 6042. The graphics processor 6041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 606 may include a display panel 6061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. Other input devices 6072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0807] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 601 can transmit it to the processor 610 for processing; in addition, the radio frequency unit 601 can send uplink data to the network-side device. Typically, the radio frequency unit 601 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.

[0808] The memory 609 can be used to store software programs or instructions, as well as various data. The memory 609 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 609 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 609 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0809] Processor 610 may include one or more processing units; optionally, processor 610 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 610.

[0810] The radio frequency unit 601 is used to: send first information to the second device, wherein the first information is used to determine the first artificial intelligence (AI) unit;

[0811] The radio frequency unit 601 is also configured to receive the inference result of the first AI unit sent by the second device, wherein the inference result of the first AI unit is the inference result obtained by inferring the data collected by the second device using the first AI unit.

[0812] The processor 610 is configured to perform inference based on the inference result of the first AI unit and the second AI unit to obtain the inference result of the second AI unit, wherein the first AI unit and the second AI unit are aligned.

[0813] Optionally, the alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions:

[0814] The first AI unit and the second AI unit are sub-units of the target AI unit;

[0815] The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

[0816] Optionally, the first information includes at least one of the following:

[0817] The model information of the first AI unit; the description information of the first AI unit;

[0818] The model information of the first AI unit includes at least one of the following:

[0819] The executable file corresponding to the first AI unit;

[0820] Information used to indicate the reference model structure of the first AI unit;

[0821] Information used to indicate the order of model parameters of the first AI unit;

[0822] The parameter information of the first AI unit;

[0823] The identifier of the first AI unit;

[0824] Information used to indicate the parameter modification layer of the first AI unit;

[0825] Information used to indicate the layer corresponding to the first AI unit.

[0826] Optionally, the reasoning result of the first AI unit satisfies any one of the following conditions:

[0827] The reasoning result of the first AI unit does not have corresponding physical parameters;

[0828] The reasoning results of the first AI unit have corresponding physical parameters.

[0829] Optionally, the processor 610 is specifically used for any of the following:

[0830] The reasoning result of the first AI unit is used as the model input of the second AI unit for reasoning;

[0831] The reasoning results of the first AI unit and the data collected by the first device are used as model inputs for the second AI unit for reasoning.

[0832] Optionally, the first information is carried by at least one of the following:

[0833] Control plane messages; User plane messages; Data plane messages.

[0834] Optionally, the control plane message includes a first message, which is used to request the allocation of resources for AI services;

[0835] And / or,

[0836] The user plane message includes a second message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0837] And / or,

[0838] The data plane message includes a third message, which is used to send AI business data, and the AI ​​business data includes the first information.

[0839] Optionally, the second device includes a first network element, and the radio frequency unit 601 is specifically used for at least one of the following:

[0840] The first entity sends first information to the first network element, and the first message is a message transmitted between the first device and the first entity;

[0841] Send the second message to the first network element;

[0842] The third message is sent to the first network element.

[0843] Optionally, when the first information includes the description information of the first AI unit, the description information of the first AI unit includes at least one of the following:

[0844] The effective duration of the first AI unit;

[0845] The configuration environment of the first AI unit;

[0846] The model input information of the first AI unit;

[0847] The model outputs relevant information from the first AI unit;

[0848] The computational load of the first AI unit;

[0849] The inference latency required by the first AI unit;

[0850] The required computing speed of the first AI unit;

[0851] The storage requirements of the first AI unit;

[0852] The function of the first AI unit;

[0853] Information used to indicate whether a second device is needed for data collection;

[0854] Used to obtain information about the application service AS of the first AI unit;

[0855] Wherein, the model input information of the first AI unit is associated with the function of the first AI unit; and / or, the model output information of the first AI unit is associated with the function of the first AI unit.

[0856] Optionally, the first AI unit may include at least one of the following functions:

[0857] Environment reconstruction; trajectory planning; driving decision; target recognition; target detection; target tracking; intrusion detection.

[0858] Optionally, when the function of the first AI unit includes environment reconstruction, the model input information of the first AI unit includes at least one of the following: a first data item, the data dimension of the first data item; the model output information of the first AI unit includes a second data item; and / or

[0859] When the function of the first AI unit includes trajectory planning, the model input information of the first AI unit includes at least one of the following: a third data item, wherein the data dimension of the third data item; the model output information of the first AI unit includes a fourth data item; and / or

[0860] When the function of the first AI unit includes driving decision-making, the model input information of the first AI unit includes at least one of the following: a fifth data item, wherein the data dimension of the fifth data item; the model output information of the first AI unit includes a sixth data item.

[0861] The first data item includes at least one of the following: chromaticity and luminance;

[0862] The data dimensions of the first data item include at least one of the following: number of channels, height, width, and number of frames;

[0863] The second data item includes at least one of the following: time point, obstacle type, three-dimensional coordinates of the obstacle, envelope of the obstacle, volume of the obstacle, moving speed of the obstacle, moving direction of the obstacle, and distance between the obstacle and the target object;

[0864] The third data item includes at least one of the following: chromaticity, luminance;

[0865] The data dimensions of the third data item include at least one of the following: number of channels, height, width, and number of frames;

[0866] The fourth data item includes at least one of the following: time point, speed, and direction of movement;

[0867] The fifth data item includes at least one of the following: chromaticity, luminance;

[0868] The data dimensions of the fifth data item include at least one of the following: number of channels, height, width, and number of frames;

[0869] The sixth data item includes at least one of the following: time point, throttle opening, braking force, and steering angle.

[0870] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description in Figure 2 of the method embodiment and achieve the same or corresponding technical effects. To avoid repetition, it will not be described again here.

[0871] This application also provides a network-side device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps of the method embodiment shown in FIG11. This network-side device embodiment corresponds to the second device method embodiment described above. All implementation processes and methods of the above method embodiments can be applied to this network-side device embodiment and can achieve the same technical effect.

[0872] Specifically, this application embodiment also provides a network-side device, which can be the model inference device shown in FIG13. As shown in FIG16, the network-side device 700 includes: an antenna 701, a radio frequency device 702, a baseband device 703, a processor 704, and a memory 705. The antenna 701 is connected to the radio frequency device 702. In the uplink direction, the radio frequency device 702 receives information through the antenna 701 and sends the received information to the baseband device 703 for processing. In the downlink direction, the baseband device 703 processes the information to be transmitted and sends it to the radio frequency device 702. The radio frequency device 702 processes the received information and transmits it through the antenna 701.

[0873] The method executed by the network-side device in the above embodiments can be implemented in the baseband device 703, which includes a baseband processor.

[0874] The baseband device 703 may include at least one baseband board, on which multiple chips are disposed, as shown in FIG16. One of the chips is, for example, a baseband processor, which is connected to the memory 705 via a bus interface to call the program in the memory 705 and execute the network device operation shown in the above method embodiment.

[0875] The network-side device may also include a network interface 706, such as a Common Public Radio Interface (CPRI).

[0876] Specifically, the network-side device 700 in this application embodiment further includes: instructions or programs stored in memory 705 and executable on processor 704. Processor 704 calls the instructions or programs in memory 705 to execute the methods executed by each module shown in FIG13 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.

[0877] Specifically, this application also provides a network-side device. As shown in FIG17, the network-side device 800 includes a processor 801, a network interface 802, and a memory 803. The network-side device may be the model inference device shown in FIG13. The network interface 802 is, for example, a common public radio interface (CPRI).

[0878] Specifically, the network-side device 800 in this application embodiment further includes: instructions or programs stored in memory 803 and executable on processor 801. Processor 801 calls the instructions or programs in memory 803 to execute the methods executed by each module shown in FIG13 and achieve the same technical effect. To avoid repetition, it will not be described in detail here.

[0879] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described model inference method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0880] The processor mentioned above is the processor in the terminal or network-side device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.

[0881] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described model inference method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0882] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0883] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described model reasoning method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0884] This application also provides a wireless communication system, including a first device and a second device. The first device can be used to execute the steps of the model inference method applied to the first device as described above, and the second device can be used to execute the steps of the model inference method applied to the second device as described above.

[0885] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0886] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.

[0887] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.

Claims

1. A model reasoning method, comprising: The first device sends first information to the second device, the first information being used to identify the first artificial intelligence (AI) unit; The first device receives the inference result of the first AI unit sent by the second device. The inference result of the first AI unit is the inference result obtained by using the first AI unit to infer the data collected by the second device. The first device performs inference based on the inference result of the first AI unit and the second AI unit to obtain the inference result of the second AI unit. The first AI unit and the second AI unit are aligned.

2. The method of claim 1, wherein, The alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions: The first AI unit and the second AI unit are sub-units of the target AI unit; The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

3. The method of claim 1 or 2, wherein, The first information includes at least one of the following: The model information of the first AI unit; the description information of the first AI unit; The model information of the first AI unit includes at least one of the following: The executable file corresponding to the first AI unit; Information used to indicate the reference model structure of the first AI unit; Information used to indicate the order of model parameters of the first AI unit; The parameter information of the first AI unit; The identifier of the first AI unit; Information used to indicate the parameter modification layer of the first AI unit; Information used to indicate the layer corresponding to the first AI unit.

4. The method of any one of claims 1-3, wherein, The reasoning result of the first AI unit satisfies the following conditions: The reasoning results of the first AI unit have corresponding physical parameters.

5. The method of any one of claims 1-4, wherein, The first device performs inference based on the inference result of the first AI unit and the second AI unit, including any one of the following: The first device uses the reasoning result of the first AI unit as the model input of the second AI unit for reasoning; The first device uses the reasoning results of the first AI unit and the data collected by the first device as model inputs for the second AI unit to perform reasoning.

6. The method of any one of claims 1-5, wherein, The first information is carried by at least one of the following: Control plane messages; User plane messages; Data plane messages.

7. The method of claim 6, wherein, The control plane message includes a first message, which is used to request the allocation of resources for AI services; And / or, The user plane message includes a second message, which is used to send AI business data, and the AI ​​business data includes the first information. And / or, The data plane message includes a third message, which is used to send AI business data, and the AI ​​business data includes the first information.

8. The method of claim 7, wherein, The second device includes a first network element, and the first device sends first information to the second device, including at least one of the following: The first device sends first information to the first network element through the first entity, and the first message is a message transmitted between the first device and the first entity; The first device sends the second message to the first network element; The first device sends the third message to the first network element.

9. The method of claim 3, wherein, When the first information includes the description information of the first AI unit, the description information of the first AI unit includes at least one of the following: The effective duration of the first AI unit; The configuration environment of the first AI unit; The model input information of the first AI unit; The model outputs relevant information from the first AI unit; The computational load of the first AI unit; The inference latency required by the first AI unit; The required computing speed of the first AI unit; The storage requirements of the first AI unit; The function of the first AI unit; Information used to indicate whether a second device is needed for data collection; Used to obtain information about the application service AS of the first AI unit; Wherein, the model input information of the first AI unit is associated with the function of the first AI unit; and / or, the model output information of the first AI unit is associated with the function of the first AI unit.

10. A model reasoning method, comprising: The second device receives the first information sent by the first device and determines the first AI unit based on the first information. The second device uses the first AI unit to infer the data collected by the second device and obtain the inference result of the first AI unit; The second device sends the inference result of the first AI unit to the first device; The reasoning result of the first AI unit is used to perform reasoning to obtain the reasoning result of the second AI unit, and the first AI unit and the second AI unit are aligned.

11. The method of claim 10, wherein, The alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions: The first AI unit and the second AI unit are sub-units of the target AI unit; The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

12. The method of claim 10 or 11, wherein, The first information includes at least one of the following: The model information of the first AI unit; the description information of the first AI unit; The model information of the first AI unit includes at least one of the following: The executable file corresponding to the first AI unit; Information used to indicate the reference model structure of the first AI unit; Information used to indicate the order of model parameters of the first AI unit; The parameter information of the first AI unit; The identifier of the first AI unit; Information used to indicate the parameter modification layer of the first AI unit; Information used to indicate the layer corresponding to the first AI unit.

13. The method of any one of claims 10-12, wherein, The reasoning result of the first AI unit satisfies the following conditions: The reasoning results of the first AI unit have corresponding physical parameters.

14. The method of any one of claims 10-13, wherein, The first information is carried by at least one of the following: Control plane messages; User plane messages; Data plane messages.

15. The method of claim 14, wherein, The control plane messages include at least one of the following: The first message is used to request the allocation of resources for AI services; The fourth message is used to request AI resources; The fifth message is used for AI resource allocation; And / or, The user plane message includes a second message, which is used to send AI business data, and the AI ​​business data includes the first information. And / or, The data plane message includes a third message, which is used to send AI business data, and the AI ​​business data includes the first information.

16. The method of claim 15, wherein, The second device includes a first network element, and the second device receives first information sent by the first device, including at least one of the following: The first network element receives the fifth message sent by the second network element, the fifth message being determined based on the fourth message sent by the first entity, and the fourth message being determined based on the first message sent by the first device; The first network element receives the second message sent by the first device; The first network element receives the third message sent by the first device.

17. The method of claim 12, wherein, When the first information includes the description information of the first AI unit, the description information of the first AI unit includes at least one of the following: The effective duration of the first AI unit; The configuration environment of the first AI unit; The model input information of the first AI unit; The model outputs relevant information from the first AI unit; The computational load of the first AI unit; The inference latency required by the first AI unit; The required computing speed of the first AI unit; The storage requirements of the first AI unit; The function of the first AI unit; Information used to indicate whether a second device is needed for data collection; Used to obtain information about the application service AS of the first AI unit; Wherein, the model input information of the first AI unit is associated with the function of the first AI unit; and / or, the model output information of the first AI unit is associated with the function of the first AI unit.

18. The method of any one of claims 10-17, wherein, The second device includes a first entity. When the first information includes description information of the first AI unit, determining the first AI unit based on the first information includes: The first entity obtains the model information of the first AI unit from the third network element based on the first information; The first AI unit is determined based on the model information of the first AI unit; The model information of the first AI unit includes at least one of the following: The executable file corresponding to the first AI unit; Information used to indicate the reference model structure of the first AI unit; Information used to indicate the order of model parameters of the first AI unit; The parameter information of the first AI unit; The identifier of the first AI unit; Information used to indicate the parameter modification layer of the first AI unit; Information used to indicate the layer corresponding to the first AI unit.

19. A model reasoning device, comprising: The sending module is used to send first information to the second device, wherein the first information is used to identify the first artificial intelligence (AI) unit; The receiving module is used to receive the inference result of the first AI unit sent by the second device. The inference result of the first AI unit is the inference result obtained by using the first AI unit to infer the data collected by the second device. The processing module is used to perform inference based on the inference result of the first AI unit and the second AI unit to obtain the inference result of the second AI unit, wherein the first AI unit and the second AI unit are aligned.

20. The apparatus of claim 19, wherein, The alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions: The first AI unit and the second AI unit are sub-units of the target AI unit; The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

21. The apparatus of claim 19 or 20, wherein, The first information includes at least one of the following: The model information of the first AI unit; the description information of the first AI unit; The model information of the first AI unit includes at least one of the following: The executable file corresponding to the first AI unit; Information used to indicate the reference model structure of the first AI unit; Information used to indicate the order of model parameters of the first AI unit; The parameter information of the first AI unit; The identifier of the first AI unit; Information used to indicate the parameter modification layer of the first AI unit; Information used to indicate the layer corresponding to the first AI unit.

22. The apparatus of any one of claims 19-21, wherein, The processing module is specifically used for any one of the following: The reasoning result of the first AI unit is used as the model input of the second AI unit for reasoning; The reasoning results of the first AI unit and the data collected by the first device are used as model inputs for the second AI unit for reasoning.

23. A model reasoning device, comprising: The receiving module is used to receive first information sent by the first device and determine the first AI unit based on the first information; The processing module is used to perform reasoning on the data collected by the second device using the first AI unit to obtain the reasoning result of the first AI unit; The sending module is used to send the inference results of the first AI unit to the first device; The reasoning result of the first AI unit is used to perform reasoning to obtain the reasoning result of the second AI unit, and the first AI unit and the second AI unit are aligned.

24. The apparatus of claim 23, wherein, The alignment of the first AI unit and the second AI unit includes the first AI unit and the second AI unit satisfying at least one of the following conditions: The first AI unit and the second AI unit are sub-units of the target AI unit; The reasoning accuracy of the joint reasoning performed by the first AI unit and the second AI unit meets the preset requirements.

25. The apparatus of claim 23 or 24, wherein, The first information includes at least one of the following: The model information of the first AI unit; the description information of the first AI unit; The model information of the first AI unit includes at least one of the following: The executable file corresponding to the first AI unit; Information used to indicate the reference model structure of the first AI unit; Information used to indicate the order of model parameters of the first AI unit; The parameter information of the first AI unit; The identifier of the first AI unit; Information used to indicate the parameter modification layer of the first AI unit; Information used to indicate the layer corresponding to the first AI unit.

26. A communication device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the model inference method as claimed in any one of claims 1-9, or implementing the steps of the model inference method as claimed in any one of claims 10-18.

27. A readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the model inference method as claimed in any one of claims 1-9, or implement the steps of the model inference method as claimed in any one of claims 10-18.

28. A computer program / program product that, when executed by at least one processor, implements the steps of the model inference method as claimed in any one of claims 1-9, or implements the steps of the model inference method as claimed in any one of claims 10-18.