Model reasoning method and communication apparatus
Through inference negotiation between network devices and terminals, model inference tasks are dynamically allocated, and the problem of insufficient computing power of the terminal is solved, effective model inference collaboration and resource optimization are achieved, and user experience and privacy are improved.
Patent Information
- Application Number
- PCT/CN2024/127642
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-24
AI Technical Summary
The model inference effect or the delay caused by insufficient computing power of terminal devices is poor or the time delay is too large, which affects the user experience. At the same time, the network equipment performs model inference that may not guarantee the privacy of user data and serious waste of resources.
Through inference negotiation between network devices and terminals, model inference tasks are dynamically allocated, so that the terminal and network devices are each responsible for performing part of the inference of session-related models, achieving effective coordination and on-demand allocation.
This avoids the problem of poor model inference effect or excessive delay caused by insufficient computing power of terminals, improves user experience, optimizes resource utilization, and ensures user data privacy.
Smart Images

Figure CN2024127642_24072025_PF_FP_ABST
Abstract
Description
Model reasoning method and communication device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 16, 2024, with application number 202410063848.2 and application name “Model Reasoning Method and Communication Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communications, and more particularly, to a model reasoning method and a communication device. Background Art
[0003] Model inference refers to the process by which a model generates results. Users often require model inference when using applications or data channel (DC) applets. For example, face-swapping technology and voice translation in video streaming all involve model inference. Model inference generally requires high computing power. Insufficient computing power on the terminal can lead to poor model inference results or high latency, impacting the user experience.
[0004] Summary of the Invention
[0005] The present application provides a communication method and a communication device that can achieve effective collaboration between terminals and network devices and on-demand allocation of reasoning tasks.
[0006] In a first aspect, a communication method is provided. The method can be executed by a terminal, or by a component of the terminal (e.g., a processor, chip, or chip system), or by a logic module or software that implements all or part of the terminal's functions. For ease of description, the following description uses execution by a terminal as an example.
[0007] The method includes: receiving model inference information from a first network device, the model inference information being used by a terminal to execute inference of a first part model related to a session; executing inference of the first part model according to the model inference information, and sending the inference result of the first part model to a second network device, or receiving the inference result of the second part model related to the session from the second network device, and executing inference of the first part model according to the inference result of the second part model and the model inference information.
[0008] According to the model reasoning method provided in this application, reasoning negotiation can be carried out between the network device and the terminal, that is, the first network device can determine the division of labor of the model reasoning of the session, that is, determine which part of the model reasoning is the responsibility of the terminal, and which part of the model reasoning is the responsibility of the second network device. In this way, the reasoning tasks can be dynamically allocated between the terminal and the second network device through the first network device, realizing effective collaboration between the terminal and the second network device and on-demand allocation of model reasoning tasks. Furthermore, this method can avoid the problem of poor model reasoning effect or large latency due to insufficient computing power of the terminal, thereby improving user experience.
[0009] It should be noted that in this application, the division of labor for model reasoning for a session refers to the division of labor for reasoning about a model (referred to as the target model in this application) related to the session (for example, a service within the session). The reasoning tasks described in this application also refer to reasoning about this target model.
[0010] In one possible implementation, before receiving the model inference information from the first network device, the method further includes: sending inference capability information of the terminal to the first network device. For example, the inference capability information includes one or more of the following: available computing power for inference, inference accuracy, available cache for inference, an artificial intelligence (AI) / deep learning (ML) framework and version number used for inference, or a supported intermediate data compression algorithm.
[0011] Based on the above technical solution, after the first network device receives the reasoning capability information of the terminal, it can determine whether to split the reasoning task (or, the model related to the session or the reasoning of the model related to the session) based on the reasoning capability information of the terminal, and when it is determined to split the reasoning task, it can reasonably split the reasoning task based on the reasoning capability information of the terminal.
[0012] In a possible implementation, the available computing power for inference is determined by the terminal based on the configuration of its computing resources, or the available computing power for inference is determined by the terminal based on its remaining computing resources or available computing resources.
[0013] In a possible implementation, before receiving the model inference parameter information, the method further includes: sending the session information and / or the current connection bandwidth of the terminal to the first network device.
[0014] Based on the above technical solution, the first network device can determine whether to split the reasoning task based on the information of the session and / or the current connection bandwidth of the terminal, and if it is determined to split the reasoning task, it can reasonably split the reasoning task based on the reasoning capability information of the terminal.
[0015] In a possible implementation, the model reasoning information includes an identifier of the first partial model.
[0016] Based on the above technical solution, the terminal can determine to execute reasoning on the first part of the model according to the identifier of the first part of the model.
[0017] In one possible implementation, the model inference information also includes one or more of the following: the number of parameters of the first part model, the type of inference output data, the tensor shape of the intermediate data, the tensor structure of the intermediate data, or the compression algorithm of the intermediate data.
[0018] For example, based on the type of inference output data, the terminal can determine whether to send the inference result to the second network device. For example, based on the tensor shape and tensor structure of the intermediate data, the terminal can perform encoding or decoding. For example, based on the compression algorithm of the intermediate data, the terminal can compress the encoded data. By compressing the encoded data, the data transmission volume can be reduced, preventing the data transmission volume from exceeding the terminal's current connection bandwidth.
[0019] In a possible implementation, sending the reasoning capability information to the first network device includes: sending a first message to the first network device, where the first message includes the reasoning capability information, and the first message is a registration request message or a session call message.
[0020] Based on the above technical solution, the terminal can send the reasoning capability information of the terminal to the first network device during the registration process or the session call process.
[0021] In one possible implementation, the second network device is an artificial intelligence (AI) / deep learning (ML) media plane network element, and the first network device is an AI / ML control plane network element; or, the second network device is a media resource function network element, and the first network device is an application server; or, the second network device is an Internet service media server, and the first network device is an Internet service signaling server.
[0022] In a possible implementation, sending the reasoning capability information of the terminal to the first network device includes: when the second network device has model reasoning capability, sending the reasoning capability information to the first network device.
[0023] Based on the above technical solution, the terminal sends the reasoning capability information to the first network device only when it confirms that the second network device has the model reasoning capability, which can save unnecessary signaling overhead.
[0024] In one possible implementation, when the second network device has model reasoning capability, before sending the reasoning capability information to the first network device, the method further includes: obtaining capability indication information of the second network device, where the capability indication information indicates that the second network device has model reasoning capability.
[0025] Based on the above technical solution, the terminal can determine whether the second network device has model reasoning capabilities based on the capability indication information of the second network device. Furthermore, the terminal can send the terminal's reasoning capability information to the first network device, so that the first network device can determine whether to split the reasoning task based on the terminal's reasoning capability information, and if it is determined that the reasoning task is to be split, the reasoning task is reasonably split.
[0026] On the second aspect, a communication method is provided, which can be executed by a first communication device, or by a component of the first communication device (such as a processor, chip, or chip system, etc.), or by a logic module or software that can realize all or part of the functions of the first communication device.
[0027] The method includes: determining model inference information and sending the model inference information to a terminal, wherein the model inference information is used by the terminal to perform inference of a first part of a model related to a session.
[0028] According to the model reasoning method provided in this application, reasoning negotiation can be carried out between the network device and the terminal, that is, the first network device can determine the division of labor of the model reasoning of the session, that is, determine which part of the model reasoning is the responsibility of the terminal, and which part of the model reasoning is the responsibility of the second network device. In this way, the reasoning tasks can be dynamically allocated between the terminal and the second network device through the first network device, realizing effective collaboration between the terminal and the second network device and on-demand allocation of model reasoning tasks. Furthermore, this method can avoid the problem of poor model reasoning effect or large latency due to insufficient computing power of the terminal, thereby improving user experience.
[0029] In a possible implementation, the model reasoning information includes an identifier of the first partial model.
[0030] Based on the above technical solution, the terminal can determine to execute the inference of the first part model according to the identifier of the first part model.
[0031] In one possible implementation, the model inference information also includes one or more of the following: the number of parameters of the first part model, the type of inference output data, the tensor shape of the intermediate data, the tensor structure of the intermediate data, or the compression algorithm of the intermediate data.
[0032] In one possible implementation, determining model reasoning information includes determining the model reasoning information based on one or more of the following: reasoning capability information of the terminal, reasoning capability information of the second network device, information of the session, or the current connection bandwidth of the terminal, wherein the second network device is used to perform reasoning of the second part of the model related to the session.
[0033] For example, the model inference information of the terminal may include one or more of the following: available computing power for inference, inference accuracy, available cache for inference, AI / ML framework and version number used for inference, or supported intermediate data compression algorithm. For example, the model inference information of the second network device may include one or more of the following: available computing power for inference, inference accuracy, available cache for inference, AI / ML framework and version number used for inference, or supported intermediate data compression algorithm.
[0034] Based on the above technical solution, the first network device can determine whether to split the inference task, and if so, properly split the inference task. For example, the second network device can calculate the inference computing power requirement for the session based on the terminal's available inference computing power, information about the session, and the terminal's current connection bandwidth, and determine whether the inference computing power requirement exceeds the terminal's available inference computing power. Furthermore, the second network device can determine the inference division of labor for the session based on the available inference computing power of the second network device, and decide whether to compress the transmitted data based on the amount of intermediate data and the terminal's current connection bandwidth.
[0035] In a possible implementation manner, the method further includes: receiving reasoning capability information of the terminal from the terminal.
[0036] In a possible implementation, the reasoning capability information of the terminal is carried in a registration request message or a session call message.
[0037] In a possible implementation, before determining the model reasoning information, the method further includes: sending capability indication information of the second network device to the terminal, where the capability indication information indicates that the second network device has model reasoning capability.
[0038] Based on the above technical solution, the terminal can send the terminal's reasoning capability information to the first network device based on the capability indication information of the second network device when confirming that the second network device has the model reasoning capability, so that the first network device can determine whether to split the reasoning task based on the terminal's reasoning capability information, and reasonably split the reasoning task when it is determined to split the reasoning task.
[0039] In a possible implementation, the method further includes: sending an identifier of a second partial model related to the session to the second network device.
[0040] Based on the above technical solution, the second network device can determine to execute the inference of the second partial model according to the identifier of the second partial model.
[0041] In one possible implementation, sending an identifier of a second part model related to the session to a second network device includes: sending a resource request message to the second network device, the resource request message including the identifier of the second part model, the resource request message being used to request the second network device to reserve resources, the resources being used to perform reasoning of the second part model.
[0042] Based on the above technical solution, the second network device can determine that the second network device performs reasoning of the second part of the model according to the resource request message, and can reserve resources to perform reasoning of the second part of the model.
[0043] In a possible implementation manner, the resource request information further includes address information of the terminal.
[0044] In one possible implementation, sending an identifier of a second partial model related to the session to a second network device includes: sending indication information to the second network device, the indication information instructing the second network device to perform reasoning of the second partial model, the indication information including the identifier of the second partial model.
[0045] That is, the second network device may be instructed to execute the inference of the second partial model by carrying the identifier of the second partial model in the indication information.
[0046] On the third aspect, a communication method is provided, which can be executed by a second communication device, or by a component of the second communication device (such as a processor, chip, or chip system, etc.), or by a logic module or software that can realize all or part of the functions of the second communication device.
[0047] The method includes: receiving an identifier of a second part model from a first network device; receiving an inference result of the first part model from a terminal, and executing inference of the second part model based on the inference result of the first part model; or, executing inference of the second part model, and sending an inference result of the second part model obtained by executing the inference of the second part model to the terminal.
[0048] According to the model inference method provided in this application, the second network device and the terminal are each responsible for executing the inference of a part of the model. This method can avoid the problem of poor model inference effect or large delay caused by insufficient computing power of the terminal, and improve user experience.
[0049] In one possible implementation, receiving an identifier of a second part model from a first network device includes: receiving a resource request message from the first network device, the resource request message including the identifier of the second part model, the resource request message being used to request the second network device to reserve resources, the resources being used to perform reasoning of the second part model.
[0050] Based on the above technical solution, the second network device can determine that the second network device performs reasoning of the second part of the model according to the resource request message, and can reserve resources to perform reasoning of the second part of the model.
[0051] In a possible implementation manner, the resource request message further includes address information of the terminal.
[0052] In one possible implementation, receiving an identifier of a second partial model from a first network device includes: receiving indication information from the first network device, the indication information instructing the second network device to perform reasoning on the second partial model, the indication information including the identifier of the second partial model.
[0053] That is, the second network device may be instructed to execute the inference of the second partial model by carrying the identifier of the second partial model in the indication information.
[0054] In one possible implementation, the method further includes: sending inference capability information of the second network device to the first network device. For example, the model inference information of the second network device may include one or more of the following: available computing power for inference, inference accuracy, available cache for inference, the AI / ML framework and version number used for inference, or supported intermediate data compression algorithms.
[0055] Based on the above solution, the first network device can reasonably divide the reasoning tasks according to the reasoning capability information of the second network device.
[0056] In a fourth aspect, a communication device is provided, the device being configured to execute the method of any possible implementation of the first to third aspects. Specifically, the device may include units and / or modules, such as a processing unit and / or a communication unit, for executing the method of any possible implementation of the first to third aspects.
[0057] In one implementation, the apparatus is a communication device (e.g., a terminal, a first network device, or a second network device). When the apparatus is a communication device, the communication unit may be a transceiver or an input / output interface; and the processing unit may be at least one processor. Alternatively, the transceiver may be a transceiver circuit. Alternatively, the input / output interface may be an input / output circuit.
[0058] In another implementation, the apparatus is a chip, chip system, or circuit for a communication device (such as a terminal, a first network device, or a second network device). When the apparatus is a chip, chip system, or circuit for a communication device, the communication unit may be an input / output interface, interface circuit, output circuit, input circuit, pin, or related circuit on the chip, chip system, or circuit; and the processing unit may be at least one processor, processing circuit, or logic circuit.
[0059] In a fifth aspect, a communication device is provided, comprising: at least one processor configured to execute a computer program or instruction stored in a memory to perform the method of any possible implementation of aspects 1 to 3. Optionally, the device further comprises a memory configured to store the computer program or instruction. Optionally, the device further comprises a communication interface, through which the processor reads the computer program or instruction stored in the memory.
[0060] In one implementation, the apparatus is a communication device (such as a terminal, a first network device, or a second network device).
[0061] In another implementation, the apparatus is a chip, a chip system, or a circuit for a communication device (such as a terminal, a first network device, or a second network device).
[0062] In a sixth aspect, a processor is provided for executing the methods provided in the above aspects.
[0063] For the operations such as sending and acquiring / receiving involved in the processor, unless otherwise specified, or if they do not conflict with their actual functions or internal logic in the relevant descriptions, they can be understood as operations such as processor output and input, or as sending and receiving operations performed by the radio frequency circuit and antenna. This application does not limit this.
[0064] In a seventh aspect, a computer-readable storage medium is provided, which stores a program code executed by a user device, and the program code includes a method for executing any possible implementation of the first to third aspects above.
[0065] In an eighth aspect, a computer program product comprising instructions is provided, which, when run on a computer, enables the computer to execute the method in any possible implementation of the first to third aspects above.
[0066] In the ninth aspect, a communication system is provided, comprising an apparatus for executing the method in any possible implementation of the first aspect, an apparatus for executing the method in any possible implementation of the second aspect, and an apparatus for executing the method in any possible implementation of the third aspect.
[0067] In one example, the apparatus for executing the method in any possible implementation of the first aspect is a terminal; the apparatus for executing the method in any possible implementation of the second aspect is a first network device; and the apparatus for executing the method in any possible implementation of the third aspect is a second network device. Optionally, the first network device and the second network device are the same network device. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] FIG1 is a schematic diagram of a communication system applicable to an embodiment of the present application;
[0069] FIG2 is a schematic diagram of a communication system applicable to another embodiment of the present application;
[0070] FIG3 is a schematic diagram of a communication system applicable to another embodiment of the present application;
[0071] FIG4 is a schematic diagram of a model reasoning method 400 provided in an embodiment of the present application;
[0072] FIG5 is a schematic flow chart of a model reasoning method 500 provided in an embodiment of the present application;
[0073] FIG6 is a schematic flowchart of a model reasoning method 600 provided in an embodiment of the present application;
[0074] FIG7 is a schematic flow chart of a model reasoning method 700 provided in an embodiment of the present application;
[0075] FIG8 is a schematic flow chart of a model reasoning method 800 provided in an embodiment of the present application;
[0076] FIG9 is a schematic diagram of a communication device 900 provided in an embodiment of the present application;
[0077] FIG10 is a schematic diagram of a communication device 1000 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0078] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0079] In the description of this application, unless otherwise specified, " / " indicates that the objects associated with each other are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, in the description of this application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. In addition, to facilitate the clear description of the technical solutions of the embodiments of this application, in the embodiments of this application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0080] In the various method embodiments of the present application, the size of the serial number does not mean the order of execution. The order of execution should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0081] It is understood that, in this application, expressions such as "under...", "if...", "when...", "if...", and similar expressions may be used interchangeably. Furthermore, these expressions all imply that corresponding actions will be taken under certain objective circumstances, and do not limit the timeframe, require no judgment in implementation, or imply any other limitations.
[0082] It can be understood that in this application, the information indicated by the indication information is referred to as the information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated can also be indirectly indicated by indicating other information, wherein there is an association between the other information and the information to be indicated. It is also possible to indicate only a part of the information to be indicated, while the other parts of the information to be indicated are known or agreed in advance. For example, the indication of specific information can also be achieved with the help of the arrangement order of each information agreed in advance (such as specified in the protocol), thereby reducing the indication overhead to a certain extent.
[0083] It is understood that some optional features in the embodiments of the present application may, in certain scenarios, be implemented independently of other features, such as the solution on which they are currently based, to solve corresponding technical problems and achieve corresponding effects. They may also be combined with other features in certain scenarios as needed. Accordingly, the devices provided in the embodiments of the present application may also implement these features or functions accordingly, which will not be described in detail here.
[0084] It is understood that in the embodiments of the present application, "A network element" and "A" can be replaced with each other. For example, the MRF network element can be replaced with MRF, and the AS can be replaced with the AS network element. In addition, the services involved in the embodiments of the present application can refer to AL / ML services.
[0085] In this application, unless otherwise specified, the same or similar parts between the various embodiments can refer to each other. In the various embodiments in this application, and the various implementation methods / implementation methods / implementation methods in each embodiment, if there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments and the various implementation methods / implementation methods / implementation methods in each embodiment are consistent and can be referenced to each other. The technical features in different embodiments and the various implementation methods / implementation methods / implementation methods in each embodiment can be combined to form new embodiments, implementation methods, implementation methods, or implementation methods according to their inherent logical relationships. The implementation methods of this application described below do not constitute a limitation on the scope of protection of this application.
[0086] The technical solutions provided in this application can be applied to a variety of communication systems, such as: fifth generation (5G) or new radio (NR) systems, over the top (OTT) systems (or systems that provide various application services to users through the Internet), long term evolution (LTE) systems, LTE frequency division duplex (FDD) systems, LTE time division duplex (TDD) systems, etc. The technical solutions provided in this application can also be applied to future communication systems, such as the sixth generation mobile communication system. The technical solutions provided in this application can also be applied to device to device (D2D) communication, vehicle-to-everything (V2X) communication, machine to machine (M2M) communication, machine type communication (MTC), and Internet of Things (IoT) communication systems or other communication systems.
[0087] The terminal in the embodiments of the present application may also be referred to as user equipment (UE), access terminal, mobile device, user terminal, terminal device, wireless communication device, user agent or user device. For example, the terminal may be a mobile phone, a tablet computer, a laptop computer, a PDA, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc.
[0088] In the embodiments of the present application, a terminal is sometimes described as a UE, and the terms "UE" and "terminal" are interchangeable. A device for implementing the functions of a terminal may be a terminal, or may be a device capable of supporting the terminal in implementing the functions, such as a chip system or chip, which may be installed in a terminal. In the embodiments of the present application, a chip system may be composed of a chip, or may include a chip and other discrete components.
[0089] The network device in the embodiments of the present application may be a device for communicating with a terminal. As an example, the network device may be a device in a communication network for providing services to the terminal. In the embodiments of the present application, the communication network may include an operator's IP multimedia subsystem (IMS) network, an established communication network, or other communication networks, without limitation.
[0090] As an example, a network device may include at least one of the following: an artificial intelligence (AI) / deep learning (ML) control plane function (AI / ML control function, AI / ML-C) network element, an AI / ML media plane function (AI / ML media function, AI / ML-M) network element, an application server (AS), a media resource function network element, and an OTT server. The following briefly describes each device.
[0091] 1. AI / ML-C NE: This is the core network control plane NE for AI / ML services, providing AI / ML inference negotiation capabilities. For example, the AI / ML-C NE coordinates inference operations for AI / ML services between terminals and network-side devices (such as AI / ML-M NEs) and determines the division of labor for inference operations.
[0092] 2. AI / ML-M network element: The core network media plane network element for AI / ML services, which can be used to provide the inference function of AI / ML models.
[0093] It can be understood that AI / ML-C network elements and AI / ML-M network elements can be understood as network elements used to implement different functions, which can be combined into network slices as needed. AI / ML-C network elements and AI / ML-M network elements can be independent hardware devices, or they can be integrated into the same hardware device, or they can be network elements in hardware devices, or they can be software functions running on dedicated hardware, or they can be virtualized function modules instantiated on a platform (for example, a cloud platform). This application does not limit the specific form of the above network elements. It can also be understood that the above naming is only defined to facilitate the distinction between different functions and should not constitute any limitation to this application.
[0094] It can also be understood that the AI / ML-C network element and the AI / ML-M network element can be provided by the IMS operator or by a third party without limitation.
[0095] 3. AS: refers to the server that provides application layer services to terminals. AS is the AS in IMS.
[0096] 4. Media Resource Function Network Element: This element is used to process media data (such as session-related media data) (for example, to infer AI / ML models) and can participate in inference negotiation. As an example, the media resource function network element can be the multimedia resource function (MRF) network element in the 3GPP standard specification, which includes the multimedia resource function controller (MRFC) and the multimedia resource function processor (MRFP).
[0097] 5. OTT server: refers to a server used to provide over the top (OTT) services (also known as Internet services) above the operator network, also known as an Internet service server, which can provide reasoning negotiation functions, for example, coordinating the reasoning operations of AI / ML services between the terminal and the network side device, determining the division of labor for reasoning operations, etc. In addition, the OTT server can also be used to implement signaling processing and routing of OTT AI / ML services, as well as media processing of AI / ML services. As an example, the OTT server may include an OTT AI / ML service signaling server (or Internet service signaling server), an OTT AI / ML service media processing server (or Internet service media processing server), and an OTT AI / ML service routing server. Exemplarily, the OTT service signaling server is responsible for processing signaling or messages related to the AI / ML service, the OTT service media processing server is responsible for processing media data related to the AI / ML service, and the OTT service routing server is responsible for routing or forwarding signaling or messages related to the AI / ML service. The naming of the OTT server does not limit the scope of protection of the embodiments of the present application.
[0098] The network devices and terminals in this application can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; they can also be deployed on water; they can also be deployed in the air on aircraft, balloons, and satellites. The embodiments of this application do not limit the scenarios in which the network devices and terminals are located.
[0099] To facilitate understanding of the embodiments of the present application, a brief explanation of the terms involved in the present application is given.
[0100] 1. AI / ML model: Also referred to as model in this application. The AI / ML model is essentially a functional relationship, and this function does not have a corresponding analytical expression, but is represented by a nested combination of multiple linear or nonlinear operations. Among them, the AI / ML model can be expressed as y=f(x), where x represents the model input, y represents the model output, and f represents the function composed of the model structure and parameters. For example, the expression of a simple AI model - a single-layer fully connected neural network model is y=σ(Wx+b)
[0101] Where σ is the activation function. For example, σ can be the hyperbolic tangent function, i.e., tanh, where tanh(x) = [exp(x) - exp(-x)] / [exp(x) + exp(-x)]. If the input of tanh(x) is a vector, the tanh operation is performed element-by-element on each component of the vector. W and b are both trainable neural network parameters.
[0102] 2. Model parameters: These can be understood as adjustable variables within the model. They are derived through training algorithms and data and are used to adjust the model to suit different tasks and scenarios.
[0103] 3. Model inference: This is the process of reasoning about a model, also referred to as inference or AI / ML reasoning. Given data x, inputting data x into an AI / ML model and obtaining the output y = f(x). In other words, model inference is the process of obtaining the output of an AI / ML model, or generating results from an AI / ML model.
[0104] In one approach, the terminal can use native or installed AI / ML applications that perform reasoning based on the terminal's built-in reasoning framework (such as built-in in the GMS / HMS of the operating system).
[0105] In another method, the terminal uses the DC applet downloaded from the Internet, and the DC applet performs reasoning based on the WebNN built into the browser.
[0106] 4. Split reasoning: This refers to splitting the complete reasoning task related to a session (i.e., the reasoning task described later) between the terminal and the network according to a specific division of labor. The terminal and the network each perform reasoning for the AI / ML model they are responsible for.
[0107] 5. Data Channel (DC): The data channel described in the embodiments of this application may be, for example, an IMS data channel, i.e., a data channel within the IMS, which can be used to transmit data based on the Stream Control Transmission Protocol (SCTP). In other words, the data channel may be a logical channel or data connection based on SCTP for data transmission. In the embodiments of this application, the DC may be used to transmit signaling or data, such as reasoning division of labor interaction information and / or reasoning results.
[0108] 6. Computing power: In the embodiments of this application, computing power is used to measure the device's ability to reason about AI / ML models.
[0109] 7. Intermediate data: After executing split inference, the output data obtained by executing inference on part of the AI / ML model is not the final inference result.
[0110] 8. Operators: AI / ML models are composed of individual computing units, which we call operators. Operators correspond to the computational logic in each layer of the neural network.
[0111] 9. Tensor: It is a container for operator calculation data, including input data and output data.
[0112] 10. Tensor descriptor: It is a description of the input and output data of the tensor. The main attributes include name, shape, data type and data arrangement format.
[0113] 11. WebNN: Web neural network (WebNN) is a web-based application programming interface (API) defined by W3C (www.w3c.org) for hardware acceleration of neural network inference.
[0114] The above briefly explains the terms involved in this application, which will not be repeated in the following embodiments.
[0115] First, the network architecture applicable to this application is briefly introduced with reference to Figures 1 to 3.
[0116] Figure 1 is a schematic diagram of a communication system applicable to an embodiment of the present application. As shown in Figure 1, the communication system may include a UE, a network element in an IMS network, an AI / ML-C network element, and an AI / ML-M network element. Among them, the network elements in the IMS network may include, for example: a data channel signaling function (DCSF) network element, an AS, a data channel media function (DCMF) network element, an IMS access gateway (IMS-AGW), a proxy-call session control function (P-CSCF) network element, an interrogating-call session control function (I-CSCF) network element, and a serving-call session control function (S-CSCF) network element. The following is a brief introduction to each network element. For details not described in detail, please refer to the previous description.
[0117] 1. IMS-AGW: Mainly responsible for providing a media anchor point for accessing the IMS network entrance.
[0118] 2. P-CSCF network element: Located in the visited network, it is the entry point for terminals to access the IMS network and is mainly responsible for forwarding Session Initialization Protocol (SIP) signaling between the terminal and the home network.
[0119] 3. I-CSCF network element: Located in the home network, it is the unified entry point of the home network and is mainly responsible for allocating or querying S-CSCF network elements that serve terminals.
[0120] 4. S-CSCF network element: Located in the home network, it is the central node of the IMS network and is mainly responsible for terminal registration, authentication, session, routing and service triggering.
[0121] 5. DCSF network element: DC control plane network element, mainly responsible for providing DC management functions.
[0122] 6. DCMF network element: DC media plane network element, mainly responsible for providing DC media processing functions.
[0123] In one example, the functions of the DCMF are performed by a media function (MF) network element and / or an MRF network element.
[0124] In one example, the functions of the AI / ML-C network element and the AI / ML-M network element are provided by the DC AS. That is, the AI / ML-C network element and the AI / ML-M network element can be combined into one network element, which can be the DC AS. The DC AS is a service AS network element for DC services.
[0125] In one example, the AI / ML-C network element may be a DC AS, and the functions of the AI / ML-M network element may be implemented by a DCMF.
[0126] Figure 2 is a schematic diagram of another communication system applicable to an embodiment of the present application. As shown in Figure 2, the communication system may include a UE and network elements in an IMS network. The network elements in the IMS network may include, for example, an AS, an MRF network element, an IMS-AGW, a P-CSCF network element, an I-CSCF network element, and an S-CSCF network element. For an introduction to each network element, please refer to the previous description.
[0127] FIG3 is a schematic diagram of another communication system applicable to an embodiment of the present application. As shown in FIG3 , the communication system may include a UE and an OTT server. For an introduction to each network element, please refer to the previous description.
[0128] It can be understood that Figures 1 to 3 are only simplified schematic diagrams for ease of understanding, and the present application is not limited thereto. As an example, these communication systems may further include other UEs, or these communication systems may further include other UEs and communication networks to which other UEs belong. Taking Figure 1 as an example, for the sake of distinction, the UE in Figure 1 is referred to as UE#1, and the other UEs are referred to as UE#2. UE#2 can communicate with the IMS network in Figure 1. Further optionally, the communication system may also include UE#2 and the IMS network to which UE#2 belongs, and UE#2 can communicate with the IMS network to which UE#1 belongs through the IMS network to which UE#2 belongs. The network elements in the IMS network to which UE#2 belongs may refer to the network elements in the IMS network to which UE#1 belongs, and this is not limited.
[0129] It can also be understood that the network elements mentioned in the embodiments of the present application, such as P-CSCF network elements, I-CSCF network elements, MRF network elements, S-CSCF network elements and other functional network elements, can be understood as network elements for implementing different functions, which can be combined into network slices as needed. These network elements can be independent hardware devices, or they can be integrated into the same hardware device, or they can be network elements in hardware devices, or they can be software functions running on dedicated hardware, or they can be virtualized function modules instantiated on a platform (for example, a cloud platform). This application does not limit the specific form of the above network elements. It can also be understood that the above naming is only defined to facilitate the distinction between different functions and should not constitute any limitation to this application. This application does not exclude the possibility of adopting other names in 6G networks and other future networks.
[0130] When using applications or DC applets, users often require model inference. For example, face-swapping technology in video clips and voice translation all involve model inference. Model inference generally requires high computing power. If model inference is performed by the terminal, its inference capabilities may not meet complex model inference requirements due to limitations in software / hardware resources, size, weight, power consumption, and heat dissipation, thus impacting the user experience. If model inference is performed by network equipment, user data privacy cannot be guaranteed, latency-sensitive services cannot be met, and network equipment resources are significantly wasted.
[0131] In view of this, this application proposes a model reasoning method, in which network devices and terminals conduct reasoning negotiation, and dynamically allocate reasoning tasks between terminals and network devices according to different reasoning requirements, so as to achieve effective collaboration between terminals and network devices and on-demand allocation of reasoning tasks.
[0132] The above briefly explains the terms involved in this application, which will not be repeated in the following embodiments.
[0133] The following will describe in detail the method provided by the embodiment of the present application in conjunction with the accompanying drawings. It will be understood that the flowchart provided in this application mainly uses the terminal and the network device (for example, the first network device, the second network device, etc.) as the execution subject of the interaction diagram to illustrate the method, but this application does not limit the execution subject of the interaction diagram. For example, the terminal / network device in the flowchart may also be a chip, a chip system, or a processor that supports the terminal / network device to implement the method, or a logic module or software that can implement all or part of the terminal / network device functions.
[0134] FIG4 is a schematic diagram of a model inference method 400 provided in an embodiment of the present application. Each step in the method 400 is described below.
[0135] S410: The first network device determines model inference information #1 (ie, an example of model inference information).
[0136] S420: The first network device sends the model inference information #1 to the terminal. Correspondingly, the terminal receives the model inference information #1 from the first network device.
[0137] The model inference information #1 may be used for inference of the first part of the model related to the terminal execution session (ie, this session, also referred to as the first session).
[0138] It should be noted that the first part of the model can be part of a session-related model, and the reasoning of the second part of the model related to the session can be performed by a network device (e.g., a second network device). It should be understood that the reasoning of the model related to the session is the session-related reasoning task described above. The model related to the session is referred to as the target model in this application. Exemplarily, the session-related model described in this application can be used for an AI / ML service (sometimes referred to as a service in this article) in the session.
[0139] For example, a session-related model can be a session-related AI / ML model. The session-related model, i.e., the target model, can be divided into a first partial model and a second partial model. It should be noted that the first partial model is a submodel of the target model, for example, denoted as the first submodel, and the second partial model is another submodel of the target model, for example, denoted as the second submodel.
[0140] Among them, the first network device and the second network device can be logical functional devices, which can be deployed on the same physical device or on different physical devices, and there is no restriction on this. For example, the first network device is an AI / ML control plane network element (such as AI / ML-C network element), and the second network device is an AI / ML media plane network element (such as AI / ML-M network element); for another example, the first network device is an AS, and the second network device is a media resource function network element (such as an MRF network element); for another example, the first network device is an Internet service signaling server (such as an OTT signaling server), and the second network device is an Internet service media server (such as an OTT media server). The following text will explain in detail the possible processes in various situations in combination with Figures 5 to 8.
[0141] Specifically, the model reasoning information #1 may include one or more of the following: the identification of the first part model, the number of parameters of the first part model, the type of inference output data, the tensor shape of the intermediate data, the tensor structure of the intermediate data, or the compression algorithm of the intermediate data.
[0142] (1) Identification of the first part of the model: used to identify the first part of the model.
[0143] It can be understood that the identifier of the first part model is the identifier of the first sub-model.
[0144] It should be noted that the identifier of the first part of the model in the model reasoning information #1 is optional.
[0145] In an optional implementation, the first network device can dynamically determine the model to be inferred by the terminal and the second network device, and instruct the terminal to perform inference on the first model by including the identifier of the first model in the model inference information #1. This approach enables flexible division of inference tasks.
[0146] In an example of the above method, the model reasoning information #1 may also include an identifier of the second part model, that is, it may simultaneously indicate that the reasoning of the first part model is performed by the terminal and the reasoning of the second part model is performed by the second network device.
[0147] In another optional manner, the terminal may learn that it is executing the reasoning of the first part of the model through pre-negotiation, protocol provisions, or network pre-configuration, but the terminal does not know when to execute the reasoning of the first part of the model. In this case, the model reasoning information #1 may include trigger information, such as 1 bit, which is used to trigger the terminal to execute model reasoning. Accordingly, the terminal can execute the reasoning of the first part of the model based on the trigger information. Unlike the previous manner, the model reasoning information #1 in this manner may not include the identifier of the first part of the model, thereby saving signaling overhead.
[0148] (2) The number of parameters of the first sub-model: that is, the number of parameters of the first sub-model. The number of parameters of the first sub-model can be used to evaluate the model complexity and computational complexity of the first sub-model.
[0149] It should be noted that the number of parameters of the first part of the model in the model inference information #1 is optional, and the model inference information #1 may not include the number of parameters of the first part of the model.
[0150] (3) Type of inference output data: the type of output data obtained by the terminal executing the inference of the model, that is, the type of output data obtained by the inference of the first part of the model.
[0151] If the terminal performs reasoning on the first part of the model, and the output data of the reasoning on the first part of the model is used as the input for reasoning on the second part of the model performed by the second network device, the type of the inference output data is intermediate data. If the second network device performs reasoning on the second part of the model, and the output data of the second part of the model is used as the input for reasoning on the first part of the model performed by the terminal, the type of the inference output data is text, audio, or other types (the specific type corresponds to the service corresponding to the target model).
[0152] Based on the type of the inference output data, the terminal can determine the method for sending the inference result. For example, if the type of the inference output data is intermediate data, the inference result is sent using DC, which can ensure the transmission reliability of the inference result. If the type of the inference output data is text or audio, the inference result is sent using the real-time transport protocol (RTP), which can ensure the real-time nature of the inference result. It should be noted that the two concepts of inference output data and inference result involved here will be explained in S440a and S440b below.
[0153] It should be noted that the model reasoning information #1 is a possible reasoning negotiation result of the target model or the session between the terminal and the first network device (for example, recorded as negotiation result 1). Another possible negotiation result of the reasoning negotiation of the target model (for example, recorded as negotiation result 2) is that the terminal performs the reasoning of the target model, that is, the target model is not split, and the terminal performs the reasoning of the entire target model. For negotiation result 2, the first network device may also send corresponding information to the terminal (for example, recorded as reasoning model information #3) to indicate the negotiation result. The reasoning model information #3 may not include the type of the inference output data, or the reasoning model information #3 may include the type of the inference output data, and the type of the inference output data is text or audio (the specific type corresponds to the business corresponding to the target model).
[0154] (4) The tensor shape of intermediate data:
[0155] If the terminal performs inference on the first partial model, and the output data of the inference on the first partial model is used as input to inference on the second partial model performed by the second network device, the tensor shape of the intermediate data represents the tensor shape of the output data of the inference on the first partial model. If the second network device performs inference on the second partial model, and the output data of the inference on the second partial model is used as input to inference on the first partial model performed by the terminal, the tensor shape of the intermediate data represents the tensor shape of the output data of the inference on the second partial model.
[0156] (5) Tensor structure of intermediate data:
[0157] If the terminal performs inference on the first partial model, and the output data of the inference on the first partial model serves as input to inference on the second partial model performed by the second network device, the tensor structure of the intermediate data represents the tensor structure of the output data of the inference on the first partial model. If the second network device performs inference on the second partial model, and the output data of the inference on the second partial model serves as input to inference on the first partial model performed by the first network device, the tensor structure of the intermediate data represents the tensor structure of the output data of the inference on the second partial model.
[0158] In this application, the tensor shape of the intermediate data and the tensor structure of the intermediate data are used to encode and / or decode the intermediate data. How to use the tensor shape of the intermediate data and the tensor structure of the intermediate data will be described below in S440a to S460a and S440b to S460b.
[0159] It should be noted that the tensor shape and / or tensor structure of the intermediate data in the model inference information #1 are optional. In one optional implementation, the first network device may determine the tensor shape and / or tensor structure of the intermediate data and include it in the model inference information #1. In another optional implementation, the terminal may obtain the tensor shape and / or tensor structure of the intermediate data through pre-negotiation, protocol provisions, or network pre-configuration.
[0160] In addition, as described above, the negotiation result sent by the first network device may be reasoning model information #3. In this case, reasoning model information #3 may not include (4) and (5).
[0161] (6) Compression algorithm for intermediate data:
[0162] If the terminal performs reasoning on the first partial model, and the output data of the reasoning on the first partial model serves as input to reasoning on the second partial model performed by the second network device, then the compression algorithm for the intermediate data represents the compression algorithm for the encoded data obtained based on the output of the first partial model. If the second network device performs reasoning on the second partial model, and the output data of the reasoning on the second partial model serves as input to reasoning on the first partial model performed by the terminal, then the compression algorithm for the intermediate data represents the compression algorithm for the encoded data obtained based on the output of the second partial model by the second network device. For more details on the compression algorithm for the intermediate data, please refer to S440a to S460a.
[0163] Exemplarily, the compression algorithm of the intermediate data may be Feature Compression for Video Coding for Machines (FC_VCM) or None, where None indicates no compression.
[0164] It should be noted that the compression algorithm for the intermediate data in model inference information #1 is optional. In one optional implementation, the first network device may determine the compression algorithm for the intermediate data and include it in model inference information #1. In another optional implementation, the terminal may obtain the compression algorithm for the intermediate data through pre-negotiation, protocol provisions, or network pre-configuration.
[0165] In addition, as described above, the negotiation result sent by the first network device may be the reasoning model information #3. In this case, the reasoning model information #3 may not include (6).
[0166] S430: The first network device sends model inference information #2 to the second network device. Correspondingly, the second network device receives the model inference information #2 from the first network device.
[0167] Among them, model reasoning information #2 is used by the second network device to perform reasoning on the second part of the model.
[0168] Optionally, before S430, the first network device may first determine model reasoning information #2, and then send model reasoning information #2 to the second network device. For example, the first network device may also determine model reasoning information #2 when determining model reasoning information #1.
[0169] Model reasoning information #2 may include one or more of the following: the identifier of the second part model, the number of parameters of the second part model, the type of inference output data, the tensor shape of the intermediate data, the tensor structure of the intermediate data, or the compression algorithm of the intermediate data.
[0170] (1) Identification of the second part model: used to identify the second part model.
[0171] It can be understood that the identifier of the second part model is the identifier of the second sub-model.
[0172] It should be noted that the identifier of the second part of the model in the model reasoning information #1 is optional.
[0173] In an optional implementation, the second network device can dynamically determine the model to be inferred by the terminal and the second network device, and instruct the second network device to perform inference on the second model by including the identifier of the second model in the model inference information #2. This approach can achieve flexible division of labor for inference tasks.
[0174] In an example of the above method, the model reasoning information #2 may also include an identifier of the first part of the model, that is, it may simultaneously indicate that the reasoning of the first part of the model is performed by the terminal and the reasoning of the second part of the model is performed by the second network device.
[0175] In another optional manner, the second network device may learn that it is executing the reasoning of the second part of the model through pre-negotiation, protocol provisions, or network pre-configuration, but the second network device does not know when to execute the reasoning of the second part of the model. In this case, the model reasoning information #2 may include trigger information, such as 1 bit, which is used to trigger the second network device to execute model reasoning. Accordingly, the second network device may execute the reasoning of the second part of the model based on the trigger information. Unlike the previous manner, the model reasoning information #2 in this manner may not include the identifier of the second part of the model, thereby saving signaling overhead.
[0176] It should be understood that the second network device can download the second partial model from the network based on the identifier of the second partial model, and then perform reasoning on the second partial model. Alternatively, the second network device can pre-store multiple models, and based on the identifier of the second partial model, the second network device can locally search for the second partial model and then perform reasoning on the second partial model.
[0177] (2) The number of parameters of the second sub-model: that is, the number of parameters of the second sub-model. The number of parameters of the second sub-model can be used to evaluate the model complexity and computational complexity of the second sub-model.
[0178] It should be noted that the number of parameters of the second part of the model in the model inference information #2 is optional, and the model inference information #2 may not include the number of parameters of the second part of the model.
[0179] (3) Type of inference output data: the type of output data obtained by the second network device executing the inference of the model, that is, the type of output data of the inference of the second part of the model.
[0180] If the second network device performs reasoning on the second model, and the output of the reasoning on the second model serves as the input for reasoning on the first model performed by the terminal, the type of the inference output data is intermediate data. If the terminal performs reasoning on the first model, and the output of the reasoning on the first model serves as the input for reasoning on the second model performed by the second network device, the type of the inference output data is text, audio, or other types (the specific type corresponds to the service corresponding to the target model).
[0181] Based on the type of the inference output data, the second network device can determine how to send the inference result. For example, if the inference output data is intermediate data, the inference result is sent using DC, which ensures reliable transmission of the inference result. If the inference output data is text or audio, for example, the inference result is sent using RTP, which ensures real-time transmission of the inference result. It should be noted that the concepts of inference output data and inference result involved here will be explained in S440a and S440b below.
[0182] It should be noted that model inference information #2 is the negotiation result 1 described above. For the negotiation result 2 described above, the first network device may also send corresponding information (e.g., recorded as inference model information #4) to the second network device to indicate the negotiation result. Since the second network device does not perform model inference, inference model information #4 may not include the type of inference output data.
[0183] (4) The tensor shape of intermediate data:
[0184] If the terminal performs inference on the first partial model, and the output data of the inference on the first partial model is used as input to inference on the second partial model performed by the second network device, the tensor shape of the intermediate data represents the tensor shape of the output data of the inference on the first partial model. If the second network device performs inference on the second partial model, and the output data of the inference on the second partial model is used as input to inference on the first partial model performed by the terminal, the tensor shape of the intermediate data represents the tensor shape of the output data of the inference on the second partial model.
[0185] (5) Tensor structure of intermediate data:
[0186] If the terminal performs inference on the first partial model, and the output data of the inference on the first partial model serves as input to inference on the second partial model performed by the second network device, the tensor structure of the intermediate data represents the tensor structure of the output data of the inference on the first partial model. If the second network device performs inference on the second partial model, and the output data of the inference on the second partial model serves as input to inference on the first partial model performed by the first network device, the tensor structure of the intermediate data represents the tensor structure of the output data of the inference on the second partial model.
[0187] In this application, the tensor shape of the intermediate data and the tensor structure of the intermediate data are used to encode and / or decode the intermediate data. How to use the tensor shape of the intermediate data and the tensor structure of the intermediate data will be described below in S440a to S460a and S440b to S460b.
[0188] It should be noted that the tensor shape and / or tensor structure of the intermediate data in model inference information #2 are optional. In one optional implementation, the first network device may determine and include the tensor shape and / or tensor structure of the intermediate data in model inference information #2. In another optional implementation, the second network device may obtain the tensor shape and / or tensor structure of the intermediate data through pre-negotiation, protocol provisions, or network pre-configuration.
[0189] In addition, as described above, the negotiation result sent by the first network device may be reasoning model information #4. In this case, the reasoning model information #4 may not include (4) and (5).
[0190] (6) Compression algorithm for intermediate data:
[0191] If the terminal performs reasoning on the first partial model, and the output data of the reasoning on the first partial model serves as input to reasoning on the second partial model performed by the second network device, then the compression algorithm for the intermediate data represents the compression algorithm for the encoded data obtained based on the output of the first partial model. If the second network device performs reasoning on the second partial model, and the output data of the reasoning on the second partial model serves as input to reasoning on the first partial model performed by the terminal, then the compression algorithm for the intermediate data represents the compression algorithm for the encoded data obtained based on the output of the second partial model by the second network device. For more detailed information about the compression algorithm for the intermediate data, please refer to S440a to S460a and S440b to S460b.
[0192] Exemplarily, the compression algorithm of the intermediate data may be Feature Compression for Video Coding for Machines (FC_VCM) or None, where None indicates no compression.
[0193] It should be noted that the compression algorithm for the intermediate data in model inference information #2 is optional. In one optional implementation, the first network device can determine the compression algorithm for the intermediate data and include it in model inference information #2. In another optional implementation, the second network device can obtain the compression algorithm for the intermediate data through pre-negotiation, protocol provisions, or network pre-configuration.
[0194] In addition, as described above, the negotiation result sent by the first network device may be the reasoning model information #4. In this case, the reasoning model information #4 may not include (6).
[0195] Exemplarily, the first network device may send model inference information #2 to the second network device via a resource request message, such as sending an identifier of the second partial model and / or the trigger information (also referred to as indication information). The resource request message is used to request the first network device to reserve resources, where the resources are used to perform inference of the second partial model, such as computing power for performing inference of the second partial model.
[0196] It should be understood that the first network device may reserve resources in advance for executing the inference of the second part of the model, or may reserve resources after receiving the resource request message.
[0197] Exemplarily, the first network device may also send the terminal's address information to the second network device. For example, the terminal's address information may be sent via the resource request message. The terminal's address information is used to connect the terminal to the second network device, so that the terminal can send inference results of the first partial model to the second network device and receive inference results of the second partial model from the second network device.
[0198] The method 400 further includes S440a to S460a or S440b to S460b described below.
[0199] S440a, the terminal performs reasoning of the first part of the model according to the model reasoning information #1.
[0200] Take, for example, a target model used for a speech-to-text service in a video call. If the voice information corresponding to this service is input to the terminal, the terminal will then perform inference on the first partial model based on the model inference information #1. For example, based on the identifier of the first partial model in the model inference information #1, the terminal may download the first partial model and perform inference on the first partial model to obtain the inferred output data of the first partial model. For another example, the terminal may pre-store the first partial model. Upon receiving the model inference information #1, the terminal performs inference on the first partial model to obtain the inferred output data of the first partial model. Next, based on the type of the inferred output data in the model inference information #1, the terminal determines that the inferred output data of the first partial model is intermediate data. Then, based on the tensor shape and tensor structure of the intermediate data in the model inference information #1, the terminal encodes the inferred output data of the first partial model to obtain encoded data (referred to as first encoded data). Optionally, if the model inference information #1 includes a compression algorithm for the intermediate data, the terminal may compress the first encoded data using the intermediate data compression algorithm in the model inference information #1 to obtain compressed first encoded data. By compressing the first coded data, the data transmission volume can be reduced, thereby preventing the data transmission volume from exceeding the connection bandwidth of the terminal.
[0201] For ease of description, in this application, the inference output data of the first part model, the first encoded data, and the compressed first encoded data are collectively referred to as the inference result of the first part model.
[0202] S450a: The terminal sends the inference result of the first partial model to the second network device. Correspondingly, the second network device receives the inference result of the first partial model.
[0203] For example, the terminal sends the inference result of the first part of the model to the second network device through the DC.
[0204] S460a, the second network device performs reasoning of the second partial model based on the reasoning result of the first partial model and model reasoning information #2.
[0205] For example, if the inference result of the first partial model in S450a is compressed first encoded data, the second network device can decompress the compressed first encoded data based on the compression algorithm of the intermediate data in model inference information #2 to obtain the first encoded data. Next, the second network device decodes the first encoded data based on the tensor shape and tensor structure of the intermediate data in model inference information #2 to obtain the inference output data of the first partial model. Then, based on the identifier of the second partial model in model inference information #2, the second network device can use the inference output data of the first partial model as the input of the second partial model, perform inference on the second partial model, and obtain the inference output data of the second partial model.
[0206] For example, if the inference result of the first partial model in S450a is first encoded data, the second network device decodes the received inference result of the first partial model based on the tensor shape and tensor structure of the intermediate data in model inference information #2 to obtain the inference output data of the first partial model. Then, based on the identifier of the second partial model in model inference information #2, the second network device can use the inference output data of the first partial model as the input of the second partial model, perform inference on the second partial model, and obtain the inference output data of the second partial model.
[0207] After obtaining the inference output data of the second model, the second network device can send the inference output data of the second model to the other end of the session based on the type of inference output data in model inference information #2. It is understood that if the target model is used for speech-to-text services in a video call, the inference output data of the second model will be the text corresponding to the voice information input to the terminal. It should be noted that in this article, the terminal refers to one end of the session, and the other end of the session refers to another terminal or server.
[0208] S440b: The second network device performs inference of the second part of the model.
[0209] Let's take the example of a target model used for a speech-to-text service in a video call. If the voice information corresponding to this service is input to the other end of the conversation, such as another terminal, and the other end of the conversation then sends the voice information to a second network device, then after receiving the voice information, the second network device may perform inference on the second model based on the identifier of the second model in the model inference information #2, obtaining the inference output data of the second model. Next, the second network device determines that the inference output data of the second model is intermediate data based on the type of the inference output data in the model inference information #2. Then, based on the tensor shape and tensor structure of the intermediate data in the model inference information #2, it encodes the inference result of the second model based on the tensor shape and tensor structure of the intermediate data in the model inference information #2, obtaining encoded data (referred to as "second encoded data"). Optionally, if the model inference information #2 includes a compression algorithm for the intermediate data, the second network device may compress the second encoded data using the compression algorithm for the intermediate data in the model inference information #2, obtaining compressed second encoded data. Compression of the second encoded data can reduce data transmission volume, preventing the data transmission volume from exceeding the terminal's connection bandwidth.
[0210] For ease of description, in this application, the inference output data of the second part model, the second encoded data and the compressed second encoded data are collectively referred to as the inference result of the second part model.
[0211] S450b: The second network device sends the inference result of the second partial model to the terminal. Correspondingly, the terminal receives the inference result of the second partial model from the second network device.
[0212] For example, the second network device sends the inference result of the second part of the model to the terminal through the DC.
[0213] S460b, the terminal performs reasoning of the first part of the model based on the reasoning result of the second part of the model and the model reasoning information #1.
[0214] For example, if the inference result of the second partial model in S450b is compressed second encoded data, the terminal can decompress the compressed second encoded data according to the compression algorithm of the intermediate data in the model inference information #1 to obtain the second encoded data. Next, the terminal decodes the second encoded data according to the tensor shape and tensor structure of the intermediate data in the model inference information #1 to obtain the inference output data of the second partial model. Then, based on the identifier of the first partial model in the model inference information #1, the terminal can use the inference output data of the second partial model as the input of the first partial model, execute inference on the first partial model, and obtain the inference output data of the first partial model.
[0215] For example, if the inference result of the second partial model in S450b is the second encoded data, the terminal decodes the received inference result of the second partial model based on the tensor shape and tensor structure of the intermediate data in the model inference information #1 to obtain the inference output data of the second partial model. Then, based on the identifier of the first partial model in the model inference information #1, the terminal can use the inference output data of the second partial model as the input of the first partial model, perform inference on the first partial model, and obtain the inference output data of the first partial model.
[0216] After obtaining the inference output data of the second model, the second network device can send the inference output data of the second model to the other end of the session based on the type of inference output data in model inference information #1. It will be understood that if the target model is used for speech-to-text services in a video call, the inference output data of the first model will be the text corresponding to the voice information input to the second network device.
[0217] According to the model reasoning method provided in this application, reasoning negotiation can be performed between the network device and the terminal, that is, the first network device can determine the division of labor of the model reasoning of the session, that is, determine which part of the model reasoning is the responsibility of the terminal, and which part of the model reasoning is the responsibility of the second network device. In this way, the reasoning tasks can be dynamically allocated between the terminal and the second network device through the first network device, thereby achieving effective collaboration between the terminal and the second network device and on-demand allocation of model reasoning tasks. Furthermore, this method can avoid the problem of poor model reasoning effect or large delay caused by insufficient computing power of the terminal, and can also avoid the delay and resource waste problems that may be caused by the second network device performing model reasoning alone, thereby improving user experience.
[0218] The following describes possible implementations of S410.
[0219] In some embodiments, S410 specifically includes: the first network device determines the model reasoning information #1 of the terminal through one or more of the following: the reasoning capability information of the terminal, the reasoning capability information of the second network device, the information of the session, or the current connection bandwidth of the terminal.
[0220] Similarly, the model reasoning information #2 in S430 may also be determined based on the reasoning capability information of the terminal, the reasoning capability information of the second network device, the information of the session, and / or the current connection bandwidth of the terminal.
[0221] 1. The terminal's reasoning capability information, the session information, and the network bandwidth currently provided by the terminal
[0222] (1) Terminal reasoning capability information: indicates the terminal's reasoning capability.
[0223] Exemplarily, the inference capability information of the terminal may include one or more of the following: the available computing power for inference of the terminal, the inference accuracy, the AI / ML framework and version number used for inference, and the supported intermediate data compression algorithm.
[0224] a. The terminal's available computing power for reasoning indicates the available computing power or remaining computing power of the terminal for executing model reasoning. As an example, the terminal's available computing power for reasoning is determined based on the configuration of its own computing resources (such as software and hardware resources), or the terminal's available computing power for reasoning is determined by the terminal based on the terminal's available or remaining computing resources. The available or remaining computing resources of the terminal can be determined by the terminal based on the terminal's available or remaining computing resources when the session is established or during the session establishment process (including the initial session establishment and the updated session establishment).
[0225] For example, the terminal's available computing power for inference can be defined based on the available power consumption limited by the graphics processing unit (GPU). In this example, the value of the terminal's available computing power for inference can be a specific wattage. For example, if the terminal is an AR glasses, and the available power consumption of the AR glasses is approximately 0.5-2 watts, then the terminal's available computing power for inference can be defined as 0.5 watts. For example, if the terminal is an all-in-one VR machine, and the available power consumption of the all-in-one VR machine is approximately 3-7 watts, then the terminal's available computing power for inference can be defined as 3 watts.
[0226] For example, the terminal's available computing power for inference can also be defined based on the GPU's remaining available computing power. In this example, the value of the terminal's available computing power for inference can be a specific floating point operations per second (TFLOPS). For example, if the terminal is an AR glasses, the available computing power for inference of the AR glasses is 0.2 TFLOPS.
[0227] b. Inference accuracy, such as INT8 and INT16.
[0228] d. The AI / ML framework and version number used for inference, such as TensorRT v2.0.
[0229] e. Supported intermediate data compression algorithms, such as FC_VCM.
[0230] (2) Session information: including the service type corresponding to the session. For example, the service type is automatic speech recognition (ASR) 1 or ASR2, where the difference between ASR1 and ASR2 is that one corresponds to the terminal executing the inference of the first part of the model and sending the inference result of the first part of the model to the second network device, while the other corresponds to the terminal executing the inference of the first part of the model based on the inference result of the second part of the model from the second network device.
[0231] (3) The current connection bandwidth of the terminal: for example, 256K.
[0232] In one implementation, before S410 , the method may further include: the terminal sending the terminal's reasoning capability information, session information, and / or the terminal's current connection bandwidth to the first network device.
[0233] In one example, when the terminal determines that the second network device has model reasoning capability, the terminal sends the terminal's reasoning capability information, session information, and / or the terminal's current connection bandwidth to the first network device. For example, the terminal receives the capability indication information from the first network device (for example, it can be the signaling sent by the first network device to the terminal during the registration process), and the capability indication information indicates that the second network device has model reasoning capability. The capability indication information may, for example, include the reasoning capability information of the second network device, and the terminal determines that the second network device has model reasoning capability based on the reasoning capability information of the second network device. Alternatively, the terminal assumes that the second network device has model reasoning capability.
[0234] Exemplarily, the terminal may send the terminal's reasoning capability information to the first network device during the registration process, or the terminal may send the terminal's reasoning capability information, session information, and / or the terminal's current connection bandwidth to the first network device during the call establishment process. For example, the terminal sends a registration request message or a session call message to the first network device. The session call message may be a session call request message or a session call response message. The registration request message may include the terminal's reasoning capability information, and the session call message may include the terminal's reasoning capability information, session information, and / or the terminal's current connection bandwidth. The session information, the terminal's current connection bandwidth, and the terminal's reasoning capability information may be carried in the same signaling or in different signaling.
[0235] It will be appreciated that the above description uses the example of a terminal sending its reasoning capability information to a first network device. This application does not limit the manner in which the first network device obtains the terminal's reasoning capability information. For example, the terminal may also send information used to determine the terminal's reasoning capability information to the first network device, and the first network device may determine the terminal's reasoning capability information based on this information.
[0236] Alternatively, the session information may not be sent by the terminal, but may be generated by the first network device based on application service logic and session parameters. Application service logic refers to the service type used by the application, such as automatic speech recognition (ASR). Session parameters refer to parameters related to the current session, such as video resolution.
[0237] 2. Reasoning capability information of the second network device
[0238] The reasoning capability information of the second network device indicates the reasoning capability of the second network device.
[0239] In one implementation, before S410 , the method may further include: the second network device sending reasoning capability information of the second network device to the first network device.
[0240] Exemplarily, the reasoning capability information of the second network device may include one or more of the following: the available computing power for reasoning of the second network device, the accuracy of reasoning, the available cache for reasoning, the AI / ML framework and version number used for reasoning, and the supported intermediate data compression algorithm. The parameters listed here are similar to the parameters contained in the reasoning capability information of the terminal described above. For details, please refer to the above description of the parameters contained in the reasoning capability information of the terminal, which will not be repeated here. It should be understood that the terminal and the second network device use the same measurement method to measure the available computing power for reasoning.
[0241] Illustratively, the second network device may periodically send the reasoning capability information of the second network device to the first network device during normal operation.
[0242] The following is an example of how to determine model reasoning information #1 and / or model reasoning information #2 based on the reasoning capability information of the terminal, the reasoning capability information of the second network device, the information of the session, and / or the current connection bandwidth of the terminal.
[0243] In one possible approach, the first network device can determine the first partial model and the second partial model based on the terminal's reasoning capability information, the second network device's reasoning capability information, and information about the session. It can be understood that determining the first partial model is equivalent to determining the identifier of the first partial model and the number of parameters of the first partial model. Determining the second partial model is equivalent to determining the identifier of the second partial model and the number of parameters of the second partial model. This is illustrated below with reference to Examples 1 and 2.
[0244] Example 1
[0245] The first network device determines whether the AI / ML framework and version number used for reasoning in the terminal's reasoning capability information are identical or compatible with the AI / ML framework and version number used for reasoning in the reasoning capability information of the second network device. If so, the first network device may then determine the inference accuracy corresponding to the model corresponding to the service type corresponding to the session, i.e., the inference accuracy corresponding to the target model. If the inference accuracy in the terminal's reasoning capability information is identical to the inference accuracy in the reasoning capability information of the second network device, the inference accuracy corresponding to the target model is either the inference accuracy in the terminal's reasoning capability information or the inference accuracy in the reasoning capability information of the second network device. If the inference accuracy in the terminal's reasoning capability information is different from the inference accuracy in the reasoning capability information of the second network device, the inference accuracy corresponding to the target model is the lower of the inference accuracy in the terminal's reasoning capability information and the inference accuracy in the reasoning capability information of the second network device. The first network device may then determine the inference computing power of the session based on the information of the session (specifically, the model corresponding to the service type corresponding to the session, i.e., the target model) and the inference accuracy corresponding to the target model. If the inference computing power of the session exceeds the available inference computing power of the terminal, the first network device determines whether the excess inference computing power exceeds the available inference computing power of the second network device. If the excess inference computing power does not exceed the available inference computing power of the second network device, the target model can be split into the first and second parts based on the principle of fully utilizing the available inference computing power of the terminal. This solution can fully utilize the available inference computing power of the terminal, thereby minimizing service latency.
[0246] For example, the inference computing power of the session is 2 watts, and the inference available computing power of the terminal is 0.5 watts, then the above-mentioned excess inference computing power is 1.5 watts. The inference available computing power of the second network device is 10 watts, which is greater than 1.5 watts, so that the target model can be split into a first part model and a second part model according to the principle of making full use of the inference available computing power of the terminal. For example, as shown in Table 1, the second network device stores a variety of splitting schemes for the target model, each of which can split the target model into two sub-models (denoted as: sub-model A and sub-model B), and the two parts of the model corresponding to each splitting scheme correspond to a certain inference computing power. According to Table 1, it can be determined that if the target model is split into a first part model and a second part model according to the principle of making full use of the inference available computing power of the terminal, then the first part model is sub-model 1 and the second part model is sub-model 2.
[0247] Table 1
[0248] Example 2
[0249] The first network device can determine the inference computing power of the session according to the method described in Example 1. If the inference computing power of the session exceeds the available inference computing power of the second network device, then for the exceeded inference computing power, determine whether it exceeds the available inference computing power of the terminal. If the exceeded inference computing power does not exceed the available inference computing power of the terminal, the target model can be split into a first part model and a second part model according to the principle of fully utilizing the available inference computing power of the second network device. This solution fully utilizes the available inference computing power of the second network device, can save the power consumption of the terminal, or can save the computing resources of the terminal so that the terminal can perform other processing in parallel.
[0250] For example, if the inference computing power of the session is 2 watts and the available inference computing power of the second network device is 1.8 watts, the excess inference computing power is 0.2 watts. The available inference computing power of the terminal is 0.5 watts, which is greater than 0.2 watts. Therefore, the target model can be split into the first and second partial models based on the principle of fully utilizing the available inference computing power of the second network device. For example, using Table 1 again, the first partial model is sub-model 3, and the second partial model is sub-model 4.
[0251] For example, in an embodiment of the present application, the inference computing power of the session can be the power consumption or equivalent TFLOPS of the session inference, which can be estimated based on the target model matching baseline data. For example, the power consumption or equivalent TFLOPS of the session inference can be benchmarked using a GPU testing tool. The test uses a unified input including resolution, algorithm, AI framework and other settings, and tests the power consumption or equivalent TFLOPS of different chips executing inference for each AI / ML business as the benchmark data.
[0252] In one possible manner, the first network device may determine the type of the inferred output data according to the session information, specifically the service type corresponding to the session in the session information.
[0253] For example, if the service type is ASR1 as mentioned above, the type of the inference output data in model inference information #1 is intermediate data, and the type of the inference output data in model inference information #2 is text. If the service type is ASR2 as mentioned above, the type of the inference output data in model inference information #1 is text, and the type of the inference output data in model inference information #2 is intermediate data.
[0254] In one possible approach, when the output data of the inference of the first partial model is intermediate data, the first network device can determine the tensor shape and tensor structure of the intermediate data in model inference information #1 and model inference information #2 based on the first partial model. The tensor shape and tensor structure of the intermediate data in model inference information #1 and model inference information #2 are the same.
[0255] When the output data of the inference of the second partial model is intermediate data, the first network device may determine the tensor shape and tensor structure of the intermediate data in model inference information #1 and model inference information #2 based on the second partial model. The tensor shape and tensor structure of the intermediate data in model inference information #1 and model inference information #2 are the same.
[0256] Those skilled in the art will understand that the output data of the model has a corresponding tensor shape and tensor structure. Determining the model is equivalent to determining the tensor shape and tensor structure of the output data of the model.
[0257] In one possible manner, the first network device may determine a compression algorithm for the intermediate data in the model reasoning information #1 and the model reasoning information #2 according to the connection bandwidth of the terminal.
[0258] For example, if the terminal's connection bandwidth is capable of transmitting the first encoded data or the second encoded data described above, then it is determined not to compress the intermediate data; otherwise, it is determined to compress the intermediate data. In the scenario where it is determined to compress the intermediate data, the specific compression algorithm can be determined based on the data volume of the first encoded data or the data volume of the second encoded data, as well as the terminal's connection bandwidth. The data volume of the first encoded data can be estimated based on the first partial model, and the data volume of the second encoded data can be estimated based on the second partial model.
[0259] It should be understood that the above-mentioned schemes for determining model reasoning information #1 and / or model reasoning information #2 are merely illustrative and should not constitute any limitation to the present application.
[0260] For ease of understanding, the following describes possible processes applicable to embodiments of the present application. Any terms or concepts that appear below that are the same as those described above can be referred to in the descriptions of the corresponding terms or concepts above and will not be repeated here. It is understood that the following description primarily uses the example of performing inference negotiation in a session call as an example. UE#1 in the following description is the terminal described above, and UE#2 is assumed to be the other end of the session.
[0261] Figure 5 is a schematic flowchart of a model inference method 500 provided in an embodiment of the present application. This method 500 can be used to implement the solution of method 400 described above. In this method 500, it is assumed that the first network device and the second network device are the same, i.e., the network device in Figure 5 ; or that the network device in Figure 5 is a system consisting of the first network device and the second network device. The steps of this method 500 are described below.
[0262] S501, UE#1 sends reasoning capability information of UE#1 to a network device.
[0263] UE#1 reports its reasoning capability information to the network device so that during subsequent session establishment, the network device can perform reasoning negotiation between UE#1 and the network device based on UE#1's reasoning capability information. For details about the parameters included in UE#1's reasoning capability information, please refer to the relevant description above and will not be repeated here.
[0264] S502: The network device saves the reasoning capability information of UE#1.
[0265] Optionally, method 500 further includes S503 and S504.
[0266] S503: The network device sends the reasoning capability information of the network device to UE#1. For details about the reasoning capability information of the network device, please refer to the relevant description above and will not be repeated here.
[0267] S504, UE#1 determines, based on the reasoning capability information of the network device, that the network device has model reasoning capability.
[0268] S503 and S504 are optional steps, that is, the network device does not need to provide the reasoning capability information of the network device to UE#1, and UE#1 may assume that the network device has the model reasoning capability.
[0269] The following describes the reasoning negotiation process using four scenarios.
[0270] Scenario 1: UE#1 initiates an initial session call.
[0271] S505, UE#1 sends a session call request to the network device.
[0272] It should be noted that the reasoning capability information of UE#1 reported by UE#1 in S501 may also be sent in S505. That is, method 500 may not include S501 and S502, and in S502, the session call request includes the reasoning capability information of UE#1. For example, in one possible scenario, method 500 does not include S501 and S502, but includes S503 and S504, and S505 may be executed after S504, that is, UE#1 determines that the network device has model reasoning capability based on the reasoning capability information of the network device before reporting the reasoning capability information of UE#1 to the network device. In another possible scenario, method 500 does not include S501 to S504, that is, UE#1 assumes that the network device has model reasoning capability and reports the reasoning capability information of UE#1 to the network device.
[0273] Optionally, the session call request may further include information about the session and / or the current connection bandwidth of UE#1.
[0274] S506: The network device determines model inference information #1 according to the session call request.
[0275] For example, the network device may determine model reasoning information #1 based on one or more of the following: the network device's reasoning capability information, UE #1's reasoning capability information, session information, and UE #1's current connection bandwidth. For details about model reasoning information #1 and how the network device determines model reasoning information #1, please refer to the previous description.
[0276] It should be noted that the session information may be carried in a session call request, or may be determined by the network device itself. For example, the network device may determine the session information based on application service logic and session parameters.
[0277] S507: The network device sends a session call response to UE#1. The session call response carries the inference negotiation result of the session, that is, model inference information #1.
[0278] Scenario 2: UE#1 receives an initial session call.
[0279] S508: The network device sends a session call request to UE#1.
[0280] Optionally, the session call request is used to request the reasoning capability information of UE#1. Further, the session call request is also used to request the information of the session and / or the current connection bandwidth of UE#1.
[0281] S509, UE#1 sends a session call response to the network device.
[0282] If the session call request in S508 is used to request UE#1's reasoning capability information, the session call response in S509 carries UE#1's reasoning capability information. If the session call request in S508 is also used to request information about the session and / or UE#1's current connection bandwidth, the session call response in S509 carries the session information and / or UE#1's current connection bandwidth.
[0283] It is understandable that the reasoning capability information of UE#1 reported by UE#1 in S501 may also be sent in S509. That is, method 500 may not include S501, and in S509, the session call response includes the reasoning capability information of UE#1.
[0284] S510: The network device determines model inference information #1 according to the session call response.
[0285] For example, the network device may determine model reasoning information #1 based on one or more of the following: the network device's reasoning capability information, UE#1's reasoning capability information, session information, and UE#1's current connection bandwidth. For details on how the network device determines model reasoning information #1, please refer to the previous description.
[0286] It should be noted that the session information may be carried in a session call response, or may be determined by the network device itself. For example, the network device may determine the session information based on application service logic and session parameters.
[0287] S511: The network device sends the session call response to UE#1, carrying the inference negotiation result of the session, that is, model inference information #1.
[0288] Scenario 3: UE#1 initiates a session update call.
[0289] S512, UE#1 sends a session update call request to the network device.
[0290] S513: The network device re-determines the model inference information #1 according to the session update call request. It should be understood that the model inference information #1 here may be different from the model inference information #1 determined before the session update.
[0291] Exemplarily, the session update call request may include updated: the reasoning capability information of UE#1, the information of the session, and / or the current connection bandwidth of UE#1.
[0292] In an embodiment of the present application, if the model reasoning task needs to be redistributed due to a change in the processing requirements of the session, the network device can, in step S513, for example, redetermine the model reasoning information #1 based on the updated reasoning capability information of UE#1, the information of the session and the current connection bandwidth of UE#1.
[0293] S514: The network device sends the re-determined model inference information #1 to UE#1.
[0294] S512 to S514 are similar to S505 to S507 and will not be repeated here.
[0295] Scenario 4: UE#1 receives a session update call.
[0296] S515: The network device sends a session update call request to UE#1.
[0297] S516, UE#1 sends a session update call response to the network device.
[0298] S517: The network device re-determines model inference information #1 based on the session update call response. It should be understood that the model inference information #1 here may be different from the model inference information #1 determined before the session update.
[0299] Exemplarily, the session update call response may include updated: UE#1's reasoning capability information, information about the session, and / or UE#1's current connection bandwidth, and re-determine model reasoning information #1.
[0300] In an embodiment of the present application, if the model reasoning task needs to be redistributed due to a change in the processing requirements of the session, the network device can, in step S517, for example, redetermine the model reasoning information #1 based on the updated reasoning capability information of UE#1, the information of the session and the current connection bandwidth of UE#1.
[0301] S518: The network device sends the re-determined model inference information #1 to UE#1.
[0302] S515 to S518 are similar to S508 to S511 and will not be repeated here.
[0303] The above exemplary description is made in combination with four scenarios. Regardless of the above scenario, after determining the model inference information #1, the following steps S519 to S522 or S523 to S526 may be included.
[0304] S519, UE#1 performs reasoning of the first part of the model according to model reasoning information #1.
[0305] S520, UE#1 sends the inference result of the first part of the model to the network device.
[0306] S521, the network device performs reasoning of the second part of the model according to the reasoning result of the first part of the model to obtain the reasoning result of the second part of the model.
[0307] It should be understood that S519 to S521 are the same as S440a to S460a in the method 400, and reference may be made to S440a to S460a.
[0308] S522: The network device sends the inference output data of the second part of the model obtained in S521 to UE#2.
[0309] S523, UE#2 sends media data to the network device. For example, the media data may be audio or voice information. For example, if the target model is used for speech-to-text services, the media data may be voice information.
[0310] S524: The network device performs inference of the second part of the model on the media data.
[0311] S525: The network device sends the inference result of the second part of the model to UE#1.
[0312] S526, UE#1 performs reasoning of the first part of the model based on the reasoning result of the second part of the model and the model reasoning information #1, and obtains the output data of the reasoning of the first part of the model.
[0313] It should be understood that S524 to S526 are the same as S440b to S460b in method 400, and reference may be made to S440b to S460b.
[0314] Based on the above technical solution, the terminal reports its reasoning capability information to the network device, and the network device also notifies the terminal of its reasoning capability information. This means that the terminal and network device can exchange their respective reasoning capability information. Through a session call request, the terminal and network device can negotiate reasoning. After negotiation is complete, the terminal and network device each execute reasoning for their respective models. Furthermore, if the terminal or network device initiates a session update request, the terminal and network device can renegotiate reasoning.
[0315] It is understood that in the solution shown in FIG5 , the division of labor for model reasoning is determined by the network device (e.g., S506 or S510 ). Alternatively, the division of labor for model reasoning may be determined by the terminal. For example, UE#1 may determine the model reasoning to be performed by the network device and the parameters required for performing the model reasoning based on the reasoning capability information of UE#1 and the reasoning capability information of the network device, and notify the network device of the model reasoning to be performed by the network device and the parameters required for performing the model reasoning.
[0316] Figure 6 is a schematic flow chart of a model reasoning method 600 provided in an embodiment of the present application. The method 600 can be used to implement the solution of the above-mentioned method 400. The method 600 can complete the reasoning negotiation between the terminal and the network based on the DC, and the reasoning negotiation can be performed, for example, after the DC is established. In the method 600, the first network device is an AI / ML-C network element, the second network device is an AI / ML-M network element, and the IMS includes one or more network elements. For example, reference can be made to the architecture shown in Figure 1, which will be uniformly represented by "IMS" below. As an example, the method 600 shown in Figure 6 can be used for the architecture of Figure 1. The steps in the method 600 are described below.
[0317] S601, UE#1 sends the reasoning capability information of UE#1 to the IMS.
[0318] In one possible implementation, UE#1 sends its reasoning capability information to the IMS during registration with the IMS. For example, UE#1 sends a SIP REGISTER message to the IMS, and the SIP REGISTER message carries the reasoning capability information of UE#1. For example, the reasoning capability information of UE#1 can be carried in a header field of SIP signaling (e.g., a SIP REGISTER message).
[0319] As an example, the reasoning capability information of UE#1 is in the following format:
[0320] AI-Capability:AI-Inference=1.2TFlops; AI-Precision=INT16; Cache=1G; AI-Framework=WebNN v1.2; Intermediate-compression=FC_VCM
[0321] Among them, "AI-Inference" indicates the computing power available for inference of UE#1, which is 1.2TFlops in this example. "AI-Precision" indicates the inference precision of UE#1, which is INT16 in this example. "Cache" indicates the cache available for inference of UE#1, which is 1GB in this example. "AI-Framework" indicates the AI / ML framework and version number used by UE#1 for inference, which is WebNN v1.2 in this example. Generally speaking, for native applications or installed applications, inference is based on the terminal's built-in AI framework, while for DC applications, inference is performed based on WebNN. "Intermediate-compression" indicates the intermediate data compression algorithm supported by UE#1, which is FC_VCM in this example.
[0322] S602: The IMS sends the reasoning capability information of UE#1 to the AI / ML-C network element.
[0323] For example, after successfully registering and authenticating UE#1, the IMS forwards the reasoning capability information of UE#1 to the AI / ML-C network element. As an example, the IMS sends an HTTP message to the AI / ML-C network element, and the HTTP message carries the reasoning capability information of UE#1.
[0324] S603: The AI / ML-C network element saves the reasoning capability information of UE#1.
[0325] Optionally, the method 600 further includes S604 to S606.
[0326] S604: The AI / ML-C network element sends the reasoning capability information of the AI / ML-M network element to the IMS.
[0327] In one possible implementation, the AI / ML-C network element sends an HTTP response message (such as a 200 message) to the IMS. The HTTP response message includes the reasoning capability information of the AI / ML-M network element.
[0328] This application does not limit how the AI / ML-C network element obtains the reasoning capability information of the AI / ML-M network element. For example, the information may be sent by the AI / ML-M network element to the AI / ML-C network element, or may be pre-configured on the AI / ML-C network element.
[0329] S605 , the IMS sends the reasoning capability information of the AI / ML-M network element to UE#1.
[0330] S606, UE#1 determines that the AI / ML-M network element has model reasoning capability.
[0331] Steps S604 to S606 are optional. Specifically, if steps S604 to S606 are executed, UE#1 can determine whether the AI / ML-M network element has model reasoning capability based on the reasoning capability information of the AI / ML-M network element. If steps S604 to S606 are not executed, UE#1 can assume that the AI / ML-M network element has model reasoning capability.
[0332] S607: Audio and video media channels are established and DC is established.
[0333] When initiating or receiving a session, UE#1 first establishes an audio or video media channel with the IMS and a DC between UE#1 and the IMS. The DC between UE#1 and the IMS can be used to transmit signaling or data between UE#1 and the IMS, such as information related to the inference division of labor. It is understood that the signaling between UE#1 and the IMS mentioned below can be transmitted via this DC.
[0334] The following describes the reasoning negotiation process in combination with two scenarios.
[0335] Scenario 1: UE#1 initiates inference negotiation.
[0336] S6111: UE#1 sends an inference negotiation request message to the IMS. The inference negotiation request message is used to perform inference negotiation with the AI / ML-C network element.
[0337] Optionally, the reasoning negotiation request information may include reasoning capability information of UE#1.
[0338] For example, the reasoning capability information of UE#1 reported by UE#1 in S601 may also be sent in S6111. That is, method 600 may not include S601 and S602, and instead send the reasoning capability information of UE#1 via the reasoning negotiation request message. If the reasoning capability information of UE#1 was reported in S601, the reasoning capability information of UE#1 may not be reported in S6111, or the reasoning capability information of UE#1 may be reported again in S6111.
[0339] Optionally, the inference negotiation request information may further include information about the session and / or the current connection bandwidth of UE#1.
[0340] In one possible scenario, UE#1 assumes that the AI / ML-M network element has model reasoning capability. Therefore, UE#1 carries the information of the session and / or the current connection bandwidth of UE#1 in the reasoning negotiation request information.
[0341] In another possible scenario, UE#1 determines whether the AI / ML-M network element has the model reasoning capability based on the reasoning capability information of the AI / ML-M network element received in S603. If UE#1 determines that the AI / ML-M network element has the model reasoning capability, UE#1 carries the information of the session and / or the current connection bandwidth of UE#1 in the reasoning negotiation request information. If UE#1 determines that the AI / ML-M network element does not have the model reasoning capability, UE#1 does not carry the information of the session and / or the current connection bandwidth of UE#1 in the reasoning negotiation request information. For the case where the AI / ML-M network element does not have the model reasoning capability, as an example, UE#1 can perform the model reasoning by itself. The embodiments of the present application mainly introduce the case where the AI / ML-M network element has the model reasoning capability.
[0342] As an example, the inference negotiation request message may be in the following format:
[0343] AI-Negotiation-Request:AI-Inference=1.2TFlops; AI-Precision=INT16; Cache=1G; AI-Framework=WebNN v1.2; AI-Model-Type=NLP; Bandwidth=256K; Intermediate-compression=FC_VCM
[0344] For the parameters other than “Bandwidth” and “AI-Model-Type”, please refer to the description in S601.
[0345] "AI-Model-Type" is session information, indicating the model type. In this example, the model type is NLP, which indicates the natural language processing service model. "Bandwidth" indicates UE#1's current connection bandwidth. In this example, it is 256 KB, indicating that UE#1's current connection bandwidth is 256 KB.
[0346] S6112: The IMS sends an HTTP request message to the AI / ML-C network element.
[0347] After receiving the reasoning negotiation request from UE#1, the IMS may convert the reasoning negotiation request into an HTTP request message and send it to the AI / ML-C network element. The HTTP request message may include the parameters carried in the reasoning negotiation request in S6111. For example, the HTTP request message may include UE#1's reasoning capability information, and may further include information about the session and / or UE#1's current connection bandwidth.
[0348] S6113: The AI / ML-C network element determines the inference division information based on the inference negotiation request information.
[0349] Exemplarily, the AI / ML-C network element may determine the reasoning division of labor information based on one or more of the following: the reasoning capability information of the AI / ML-M network element, the reasoning capability information of UE#1, the information of the session, and the current connection bandwidth of UE#1.
[0350] It should be understood that model inference information #1 is one example of inference division information. That is, if the AI / ML-C network element determines that UE #1 will perform inference for the first part of the model, the inference division information is model inference information #1. For more information about inference division information and how to determine it, please refer to the previous description of model inference information #1.
[0351] It should be noted that the session information can be carried in an HTTP request message or determined by the AI / ML-C network element itself. For example, the AI / ML-C network element can determine the session information based on application business logic and session parameters.
[0352] S6114: The AI / ML-C network element sends inference division information to the IMS.
[0353] In one possible implementation, the AI / ML-C network element sends an HTTP response message to the IMS. The HTTP response message carries the inference negotiation result of the current session, that is, the inference division of labor information.
[0354] As an example, the inference division of labor information is in the following form:
[0355] AI-Negotiation-Response: Result=success; AI-Model-Identifier=model_xxx; AI-Model-Param=11M; Output-Data-Type=intermediate; Output-Tensor-Shape=[1,64,64]; Output-Tensor-Structure=Pytorch2.0; Intermediate-Compression=FC_VCM
[0356] The "Result" field indicates the negotiation result. In this example, a value of "success" indicates a successful negotiation, meaning that UE#1 and the AI / ML-M network element each perform inference on a portion of the model. It should be understood that this field could also indicate a failure, meaning that the negotiation was unsuccessful, meaning that the AI / ML-M network element is not authorized to perform inference on a portion of the model, and UE#1 performs inference on all models.
[0357] "AI-Model-Identifier" indicates the identifier of the model for which the terminal is responsible for inference. If the AI / ML-C network element determines that the AI / ML-M network element will perform inference for all models, that is, if there is no division of labor for model inference or no splitting of the target model, this field may not be included in the inference division information. It should be understood that in one example, the value of this field is the identifier of the first part of the model.
[0358] "AI-Model-Param" indicates the number of parameters for the model identified by "AI-Model-Identifier." In this example, this field is 11M, where M represents one million. If the AI / ML-C network element determines that the AI / ML-M network element will perform inference for all models, this field may not be included in the inference division information. It should be understood that, in one example, the value of this field is the number of parameters for the first part of the model.
[0359] "Output-Data-Type" indicates the type of the inference output data. If it is split inference, that is, the AI / ML-M network element determines that UE#1 and the AI / ML-M network element are each responsible for the inference of a part of the model, then this field is intermediate, indicating that the type of the inference output data is intermediate data. For example, if the AI / ML-M network element determines that UE#1 is responsible for the inference of all models, then the Output-Data-Type = result. If the AI / ML-M network element determines that the AI / ML-M network element is responsible for the inference of all models, then the Output-Data-Type = media.
[0360] Output-Tensor-Shape indicates the tensor shape of the intermediate data. In this example, this field is [1,64,64], indicating a three-dimensional tensor with 1 element in the first dimension, 64 elements in the second dimension, and 64 elements in the third dimension. If you are not performing split inference, this field is not required.
[0361] Output-Tensor-Structure represents the tensor structure of intermediate data. In this example, this field is PyTorch 2.0.
[0362] "Intermediate-Compression" indicates the compression algorithm for intermediate data. Its value can be VC_FCM or None. In this example, it is FC_VCM. If Intermediate-Compression = None, no compression is performed. If split inference is not used, this field can be omitted. In one example, if "Intermediate-Compression" is not present, no compression is performed.
[0363] S6115, IMS sends reasoning division of labor information to UE#1.
[0364] After receiving the reasoning division information sent by the AI / ML-C network element, the IMS forwards the reasoning division information to UE#1.
[0365] In a possible implementation, the IMS sends a reasoning negotiation response to UE#1, where the reasoning negotiation response carries the reasoning division of labor information, or the reasoning division of labor information is the reasoning negotiation response.
[0366] S6116: The AI / ML-C network element requests resources from the AI / ML-M network element.
[0367] When AI / ML-C determines that the AI / ML-M network element performs model inference, for example, when the inference division of labor information is model inference information #1, it can apply for resources from the AI / ML-M network element, that is, request the AI / ML-M network element to perform the inference of the second part of the model. In one possible implementation method, the AI / ML-C network element sends an HTTP request message to the AI / ML-M network element. The HTTP request message is used to apply to the AI / ML-M network element for resources related to executing the inference of the second part of the model. In one example, the HTTP request message can be the resource request message described above, or the resource request message described above can be the HTTP request message. In one example, the HTTP request message can include model inference information #2.
[0368] Exemplarily, the AI / ML-C network element applies for resources from the AI / ML-M network element, including: the AI / ML-C network element requests the AI / ML-M network element to create or allocate endpoint resources, and accordingly, the AI / ML-M network element creates or allocates the corresponding resource endpoint (also called media endpoint) and returns the information of the corresponding resource endpoint to the AI / ML-C network element. For example, after receiving the above-mentioned HTTP request, the AI / ML-M network element creates or allocates two resource endpoints, namely a first resource endpoint and a second resource endpoint. The attributes of the first resource endpoint include the local connection address of the first resource endpoint, and the attributes of the second resource endpoint include the local connection address of the second resource endpoint. The local connection address of the first resource endpoint and the local connection address of the second resource endpoint are used to establish a communication connection between UE#1 and the AI / ML-M network element and to establish a communication connection between UE#2 and the AI / ML-M network element, respectively.
[0369] Exemplarily, the AI / ML-C network element requesting resources from the AI / ML-M network element may further include: the AI / ML-C network element requesting the AI / ML-M network element to reserve or allocate computing resources, and accordingly, the AI / ML-M network element reserves or allocates corresponding computing resources, such as memory or CPU time slots. For example, the AI / ML-C network element sends an inference computing power requirement to the AI / ML-M network element, where the inference computing power requirement indicates the computing power consumed by the inference of the model that the AI / ML-M network element is responsible for (such as the inference of the second part model). In this way, the AI / ML-M network element can reserve or allocate corresponding computing power resources based on the inference computing power requirement.
[0370] Optionally, the resource request result returned by the AI / ML-M network element to the AI / ML-C network element may include result indication information and / or resource endpoint information. The result indication information may be used to notify the AI / ML-C network element whether the AI / ML-M network element will execute the inference of the model (e.g., the second part model) for which the AI / ML-M network element is responsible. Exemplarily, the resource endpoint information includes the local connection address of the first resource endpoint and the local connection address of the second resource endpoint.
[0371] In one possible scenario, if the resource endpoint information returned by the AI / ML-M network element to the AI / ML-C network element includes the local connection address of the first resource endpoint, the AI / ML-C network element may send the local connection address of the first resource endpoint to UE#1, so that UE#1 can communicate with the AI / ML-M network element via the local connection address of the first resource endpoint, such as UE#1 sending the inference result obtained after UE#1 performs model inference to the AI / ML-M network element. For example, the resource endpoint information and inference division information sent to UE#1 can be carried in one signaling or message, or in different signaling or messages, without limitation. In addition, if the resource endpoint information returned by the AI / ML-M network element to the AI / ML-C network element includes the local connection address of the second resource endpoint, the AI / ML-C network element may also send the local connection address of the second resource endpoint to UE#2, so that UE#2 can communicate with the AI / ML-M network element through the local connection address of the second resource endpoint, such as UE#2 obtaining the inference output data obtained by the AI / ML-M network element after the AI / ML-M network element performs model inference from the AI / ML-M network element. The AI / ML-C network element may send the local connection address of the second resource endpoint directly to UE#2, or may send the local connection address of the second resource endpoint to UE#2 through the network device to which UE#2 belongs, without limitation.
[0372] It can be understood that the execution order of S6114 and S6116 is not limited.
[0373] For example, S6114 is executed first, followed by S6116. That is, after the AI / ML-C network element determines the inference division of labor information, it directly sends the inference division of labor information to UE#1 via the IMS. In this case, the local connection address and the inference division of labor information can be carried in different signaling. That is, the AI / ML-C network element first sends the inference division of labor information to UE#1 via the IMS, and then, after receiving the local connection address from the AI / ML-M network element, sends the local connection address to UE#1 via the IMS.
[0374] For another example, S6116 is executed first, followed by S6114. That is, after the AI / ML-C network element determines the inference division of labor information, it can first request resources from the AI / ML-M network element, and then send the inference division of labor information to UE#1 via the IMS. In this case, after the AI / ML-C network element receives the local connection address from the AI / ML-M network element, it sends the inference division of labor information and the local connection address to UE#1 via the IMS. As an example, the local connection address and the inference division of labor information can be carried in a single signaling.
[0375] Scenario 2: UE#1 initiates re-inference negotiation.
[0376] S6121, UE#1 sends a re-negotiation request message to the IMS.
[0377] In certain situations, such as when the inference requirements of the inference service change during a session and / or the terminal's available inference computing power changes, the terminal may initiate renegotiation. In these situations, UE#1 can send a renegotiation request message to the IMS to renegotiate inference. For example, the renegotiation request message may include information about UE#1's inference capabilities and UE#1's current connection bandwidth, and may further include information about the session.
[0378] S6122: The IMS sends an HTTP message to the AI / ML-C network element.
[0379] The HTTP message may include parameters in the re-reasoning negotiation request information.
[0380] S6123: The AI / ML-C network element determines the reasoning division of labor information based on the HTTP message.
[0381] In an embodiment of the present application, if the model reasoning task needs to be redistributed, for example, due to changes in reasoning requirements and / or changes in the terminal's available computing power for reasoning, the AI / ML-C network element can determine new reasoning division information in S6123 based on the updated relevant information of UE#1, such as the updated reasoning capability information of UE#1.
[0382] S6124: The AI / ML-C network element sends inference division information to the IMS.
[0383] S6125, IMS sends reasoning division of labor information to UE#1.
[0384] It should be understood that S6121 to S6125 are similar to S6111 to S6115. For details, please refer to S6111 to S6115 and will not be repeated here.
[0385] S6126: The AI / ML-C network element requests the AI / ML-M network element to update inference resources.
[0386] When the AI / ML-C network element determines that the AI / ML-M network element has updated its inference based on the inference division information, it can apply to the AI / ML-M network element to update the rendering resources, that is, to request the AI / ML-M network element to be responsible for the inference model after the inference division is updated. In one possible implementation, the AI / ML-C network element sends an HTTP message to the AI / ML-M network element. The HTTP message is used to apply for resource modification. For details, please refer to the description in S6116, which will not be repeated here. It can be understood that in this step, the AI / ML-M network element can recreate the endpoint resources or reuse the endpoint resources created or allocated in S6116; the AI / ML-M network element can reuse the computing power resources reserved or allocated in S6116, or reallocate the computing power resources. The reallocated computing power resources may be different from the computing power resources reserved or allocated in S6116.
[0387] The above examples illustrate two scenarios. Regardless of the scenario, after determining the reasoning division of labor information, the following steps S6211 to S6214 or S6221 to S6224 may be included. It should be noted that it is assumed below that the reasoning division of labor information instructs UE#1 to perform reasoning on the first part of the model, while the reasoning on the second part of the model is performed by the AI / ML-M network element.
[0388] S6211, UE#1 performs reasoning of the first part of the model according to the reasoning division of labor information.
[0389] This step may refer to S440a and will not be described again here.
[0390] S6212, UE#1 sends the inference result of the first part of the model to the AI / ML-M network element.
[0391] In one possible implementation, UE#1 sends the inference result of the first part of the model to the IMS (such as DCS-M) through Application DC, and the IMS forwards the inference result of the first part of the model to the AI / ML-M network element.
[0392] S6213: The AI / ML-M network element performs reasoning of the second part of the model based on the reasoning result of the first part of the model, and obtains output data of the reasoning of the second part of the model.
[0393] Steps S6211 to S6213 are similar to S440a to S460a in method 400, and reference may be made to S440a to S460a.
[0394] S6214: The AI / ML-M network element sends the inference output data of the second part of the model to UE#2.
[0395] S6221: The AI / ML-M network element receives media data from UE#2. For example, if the target model is used for speech-to-text services, the media data is speech information.
[0396] S6222: The AI / ML-M network element performs inference of the second part of the model on the media data.
[0397] S6223: The AI / ML-M network element sends the inference result of the second part of the model to UE#1.
[0398] In one possible implementation, the AI / ML-M network element sends the inference result of the second part of the model to the IMS through the Application DC, and the IMS forwards the inference result of the second part of the model to UE#1.
[0399] S6224, UE#1 performs reasoning of the first part of the model based on the reasoning results of the second part of the model and the reasoning division of labor information.
[0400] Steps S6221 to S6224 are similar to S440b to S470b in method 400, and reference may be made to S440b to S470b.
[0401] Based on the above technical solution, a terminal can report its reasoning capability information to the AI / ML-C network element during IMS registration. The AI / ML-C network element can also notify the terminal of the AI / ML-M network element's reasoning capability information. If the terminal determines that the AI / ML-M network element has model reasoning capabilities, it can include the terminal's capability information, session information, and / or the terminal's current connection bandwidth in a DC establishment request to request reasoning negotiation. The AI / ML-C network element can determine the reasoning division of labor based on the terminal's capability information, session information, and / or the terminal's current connection bandwidth. For example, the AI / ML-C network element can calculate the computing power required for the reasoning task and determine whether it exceeds the terminal's available computing power. If so, the AI / ML-C network element can also combine the AI / ML-M network element's reasoning capability information to determine the reasoning division of labor for the session and notify the terminal of the reasoning negotiation results. When the AI / ML-M network element participates in model reasoning, it requests reasoning resources from the AI / ML-M network element to enable it to perform model reasoning. The terminal and AI / ML-M unit then perform inference on their respective models. Furthermore, if, during a session, due to circumstances such as changes in inference requirements, the inference task may need to be redistributed, the terminal can re-initiate inference negotiation based on the updated inference requirements.
[0402] It is understood that in the solution shown in Figure 6, the division of labor is determined by the AI / ML-C network element (such as S6113, 6123, etc.). In one possible solution, the division of labor can also be determined by the terminal. For example, UE#1 can determine the reasoning division of labor information based on UE#1's reasoning capability information, session information, and UE#1's current connection bandwidth. Optionally, before executing S6211 or S6221, UE#1 can also send the reasoning division of labor information to the AI / ML-C network element.
[0403] Figure 7 is a schematic flow chart of a model reasoning method 700 provided in an embodiment of the present application. The method 700 can be used to implement the solution of the above-mentioned method 400. The method 700 can complete the reasoning negotiation between the terminal and the network based on IMS SIP signaling, and the reasoning negotiation can be completed, for example, during the call establishment process. In the method 700, the first network device is an AS, the second network device is an MRF network element, and the MRF network element can include, for example, MRFC and MRFP. The IMS Core includes one or more network elements, for example, reference can be made to the architecture shown in Figure 2, and "IMS core" is uniformly used below. As an example, the method 700 shown in Figure 7 can be used for the architecture of Figure 2. The steps in the method 700 are explained below.
[0404] S701, UE#1 sends the reasoning capability information of UE#1 to the IMS core network (IMS core).
[0405] This step is similar to S601, and you can refer to S601, which will not be repeated here.
[0406] S702: The IMS core sends the reasoning capability information of UE#1 to the AS.
[0407] For example, after the IMS core successfully registers and authenticates UE#1, it forwards the reasoning capability information of UE#1 to the AS. As an example, the IMS core sends a SIP REGISTER message to the AS, and the SIP REGISTER message carries the reasoning capability information of UE#1.
[0408] S703, the AS saves the reasoning capability information of UE#1.
[0409] Optionally, the method 700 further includes S704 to S706.
[0410] S704, the AS sends the reasoning capability information of the MRF network element to the IMS core.
[0411] This application does not limit how the AS obtains the reasoning capability information of the MRF network element. For example, it can be sent by the MRF network element to the AS, or it can be pre-configured on the AS.
[0412] S705 , the IMS core sends the reasoning capability information of the MRF network element to UE#1.
[0413] S706, UE#1 determines that the MRF network element has model reasoning capability.
[0414] Steps S704 to S706 are optional. Specifically, if steps S704 to S706 are executed, UE#1 can determine whether the MRF network element has model reasoning capability based on the reasoning capability information of the MRF network element. If steps S704 to S706 are not executed, UE#1 can assume that the MRF network element has model reasoning capability.
[0415] The following describes the reasoning negotiation process using four scenarios.
[0416] Scenario 1: UE#1 initiates call establishment.
[0417] S7111, UE#1 sends an INVITE message to the IMS core.
[0418] It is understandable that the reasoning capability information of UE#1 reported by UE#1 in S701 may also be sent in S7111. That is, method 700 may not include 701, and in S7111, the INVITE message includes the reasoning capability information of UE#1.
[0419] Optionally, the INVITE message may also include information about the session and / or the current connection bandwidth of UE#1.
[0420] In one possible scenario, UE#1 assumes that the MRF network element has model reasoning capability, and therefore, UE#1 carries the session information and / or UE#1's current connection bandwidth in the INVITE message.
[0421] Another possible scenario is that UE#1 determines whether the MRF network element has the model reasoning capability based on the reasoning capability information of the MRF network element received in S705. If UE#1 determines that the MRF network element has the model reasoning capability, UE#1 carries the information of the session and / or the current connection bandwidth of UE#1 in the INVITE message. If UE#1 determines that the MRF network element does not have the model reasoning capability, UE#1 does not carry the information of the session and / or the current connection bandwidth of UE#1 in the INVITE message. For the case where the MRF network element does not have the model reasoning capability, as an example, UE#1 can perform the model reasoning by itself. The embodiments of the present application mainly introduce the case where the MRF network element has the model reasoning capability.
[0422] S7112: The IMS core sends an INVITE message to the AS.
[0423] After receiving the INVITE message, the IMS core can transparently transmit it to the AS.
[0424] S7113: The AS determines the inference division of labor information based on the INVITE message.
[0425] For example, the AS can determine the reasoning division of labor information based on one or more of the following: the reasoning capability information of the MRF network element, the reasoning capability information of UE#1, information about the session, and the current connection bandwidth of UE#1. For more information about model reasoning information #1 and how to determine model reasoning information #1, please refer to the relevant description above.
[0426] It should be understood that model reasoning information #1 is one example of reasoning division of labor information. That is, if the AS determines that UE #1 will perform reasoning for the first part of the model, the reasoning division of labor information is model reasoning information #1. For more information about reasoning division of labor information and how to determine it, please refer to the previous description related to model reasoning information #1.
[0427] It should be noted that the session information may be carried in the INVITE message, or may be determined by the AS itself. For example, the AS may determine the session information based on application service logic and session parameters.
[0428] S7114: AS sends inference division information to the IMS core.
[0429] In a possible implementation, the AS sends an 18X For INVITE message or a 200 For INVITE message to the IMS core. The 18X For INVITE message or the 200 For INVITE message carries the inference negotiation result of the current session, that is, the inference division of labor information.
[0430] In another possible implementation, the AS sends a provisional response ACKnowledgement (PRACK) For INVITE message or an ACKnowledgement (ACK) For INVITE message to the IMS core. The PRACK For INVITE message or the ACK For INVITE message carries the reasoning negotiation result of this session, that is, the reasoning division of labor information.
[0431] As an example, the inference division of labor information is in the following form:
[0432] AI-Negotiation-Response: Result=success; AI-Model-Identifier=model_xxx; AI-Model-Param=11M; Output-Data-Type=intermediate; Output-Tensor-Shape=[1,64,64]; Output-Tensor-Structure=Pytorch2.0; Intermediate-Compression=FC_VCM
[0433] For details about the fields listed here, please refer to the relevant description in S6115, which will not be repeated here.
[0434] S7115: IMS core sends reasoning division of labor information to UE#1.
[0435] After receiving the inference division information sent by the AS, the IMS core forwards the inference division information to UE#1.
[0436] In a possible implementation manner, the IMS core sends an 18X For INVITE message or a 200 For INVITE message to UE#1, where the 18X For INVITE message or the 200 For INVITE message carries the reasoning division of labor information.
[0437] S7116: AS applies for resources from the MRF network element.
[0438] When the AS determines that the MRF network element executes the reasoning of the model based on the model reasoning information #1, for example, when the reasoning division of labor information is model reasoning information #1, it can apply for resources from the MRF network element, that is, request the MRF network element to execute the reasoning of the second part of the model. In one possible implementation method, the AS sends an INVITE message to the MRF network element, and the INVITE message is used to apply to the MRF network element for resources related to executing the reasoning of the second part of the model. In one example, the INVITE message can be the resource request message described above, or the resource request message described above can be the INVITE message. In one example, the INVITE message can include model reasoning information #2.
[0439] Exemplarily, the AS's application for resources from the MRF includes: the AS requests the MRF network element to create or allocate endpoint resources, and accordingly, the MRF network element creates or allocates corresponding resource endpoints (also called media endpoints) and returns the information of the corresponding resource endpoints to the AS. For example, after receiving the above-mentioned INVITE message, the MRF network element creates or allocates two resource endpoints, namely a first resource endpoint and a second resource endpoint, wherein the attributes of the first resource endpoint include the local connection address of the first resource endpoint, and the attributes of the second resource endpoint include the local connection address of the second resource endpoint, wherein the local connection address of the first resource endpoint and the local connection address of the second resource endpoint are used to establish a communication connection between UE#1 and the MRF network element and to establish a communication connection between UE#2 and the MRF network element, respectively.
[0440] Exemplarily, the AS's request for resources from the MRF may further include: the AS requests the MRF network element to reserve or allocate computing resources, and the MRF network element accordingly reserves or allocates corresponding computing resources, such as memory or CPU time slots. For example, the AS sends an inference computing power requirement to the MRF network element, where the inference computing power requirement indicates the computing power consumed by the inference of the model for which the MRF network element is responsible, such as the inference of the second part model. In this way, the MRF network element can reserve or allocate corresponding computing power resources based on the inference computing power requirement.
[0441] Optionally, the resource application result returned by the MRF network element to the AS may include result indication information and / or resource endpoint information. The result indication information can be used to notify the AS whether the MRF network element will execute the reasoning of the model that the MRF network element is responsible for reasoning (for example, the second part model). Further, optionally, if the resource application result is used to notify the AS: the MRF network element will execute the reasoning of the model (for example, the second part model) for the MRF network element. Exemplarily, the resource endpoint information includes the local connection address of the first resource endpoint and the local connection address of the second resource endpoint.
[0442] In one possible scenario, if the information of the resource endpoint returned by the MRF network element to the AS includes the local connection address of the first resource endpoint, the AS may send the local connection address of the first resource endpoint to UE#1, so that UE#1 can communicate with the MRF network element through the local connection address of the first resource endpoint, such as UE#1 sending the inference result obtained after the MRF network element model is inferred to the MRF network element. For example, the resource endpoint information and the inference division information sent to UE#1 can be carried in one signaling or message, or in different signaling or messages, without limitation. In addition, if the information of the resource endpoint returned by the MRF network element to the AS includes the local connection address of the second resource endpoint, the AS may also send the local connection address of the second resource endpoint to UE#2, so that UE#2 can communicate with the MRF network element through the local connection address of the second resource endpoint, such as the inference output data obtained by UE#2 after executing the model inference from the MRF network element. The AS may directly send the local connection address of the second resource endpoint to UE#2, or may send the local connection address of the second resource endpoint to UE#2 through the network device to which UE#2 belongs, without limitation.
[0443] It can be understood that the execution order of S7114 and S7116 is not limited.
[0444] For example, SS7114 is executed first, followed by SS7116. That is, after the AS determines the inferred division of labor information, it directly sends this inferred division of labor information to UE#1 via the IMS core. In this case, the local connection address and the inferred division of labor information can be carried in different signaling. That is, the AS first sends the inferred division of labor information to UE#1 via the IMS core, and then, after receiving the local connection address from the MRF network element, sends the local connection address to UE#1 via the IMS core.
[0445] For another example, S7116 is executed first, followed by S7114. That is, after the AS determines the inference division of labor information, it can first request resources from the MRF network element, and then send the inference division of labor information to UE#1 via the IMS core. In this case, after the AS receives the local connection address from the MRF network element, it sends the inference division of labor information and the local connection address to UE#1 via the IMS core. As an example, the local connection address and inference division of labor information can be carried in a single signaling.
[0446] Scenario 2: UE#1 receives a call setup request.
[0447] S7121: The AS sends an INVITE message to the IMS core.
[0448] Optionally, the INVITE message is used to request the reasoning capability information of UE#1. Alternatively, the INVITE message is used to request session information and / or the current connection bandwidth of UE#1.
[0449] S7122: The IMS core sends an INVITE message to UE#1.
[0450] If the INVITE message in S7121 is used to request UE#1's reasoning capability information, then the INVITE message in S7122 is used to request UE#1's reasoning capability information. If the INVITE message in S7121 is used to request session information and / or UE#1's current connection bandwidth, then the INVITE message in S7122 is used to request session information and / or UE#1's current connection bandwidth.
[0451] S7123, UE#1 sends a 200 For INVITE message to the IMS core.
[0452] The 200 For INVITE message may carry parameters requested by the INVITE message, such as UE#1's reasoning capability information, session information, and / or UE#1's current connection bandwidth.
[0453] It is understandable that the reasoning capability information of UE#1 reported by UE#1 in S701 may also be sent in S7123. That is, method 700 may not include S701, and in S7123, the 200 For INVITE message includes the reasoning capability information of UE#1.
[0454] The 200 For INVITE message in S7123 may also be replaced by an 18X For INVITE message, and there is no restriction on the specific type of the message.
[0455] At S7124, the IMS core sends a 200 For INVITE message to the AS.
[0456] The 200 For INVITE message in S7124 may also be replaced by an 18X For INVITE message, and there is no restriction on the specific type of the message.
[0457] At S7125, the AS determines the model division information based on the 200 For INVITE message.
[0458] Exemplarily, the AS may determine the reasoning division of labor information based on one or more of the following: reasoning capability information of the MRF network element, reasoning capability information of UE#1, information of the session, and current connection bandwidth of UE#1.
[0459] It should be understood that model reasoning information #1 is one example of reasoning division of labor information. That is, if the AS determines that UE #1 will perform reasoning for the first part of the model, the reasoning division of labor information is model reasoning information #1. For more information about reasoning division of labor information and how to determine it, please refer to the previous description related to model reasoning information #1.
[0460] It should be noted that the session information may be carried in the 200 For INVITE message, or may be determined by the AS itself. For example, the AS may determine the session information based on application service logic and session parameters.
[0461] S7126: AS sends inference division information to the IMS core.
[0462] S7127: IMS core sends inference division information to UE#1.
[0463] S7128: AS applies for resources from the MRF network element.
[0464] S7126 to S7128 are similar to S7114 to S7116 and are not described here again.
[0465] Scenario 3: UE#1 initiates a call update.
[0466] S7131, UE#1 sends a REINVITE message to the IMS core.
[0467] In some cases, such as changes in the inference requirements of the inference service and / or the terminal's available inference computing power during a session, the terminal may initiate a call update. In these cases, UE#1 can send a REINVITE message to the IMS core to renegotiate inference. For example, the REINVITE message may include information about UE#1's inference capabilities and its current connection bandwidth, as well as information about the session.
[0468] S7132: The IMS core sends a REINVITE message to the AS.
[0469] S7133: The AS determines the model division of labor information based on the REINVITE message.
[0470] In an embodiment of the present application, if the model reasoning task needs to be redistributed, for example, due to changes in reasoning requirements and / or changes in the terminal's available computing power for reasoning, the AS can determine new reasoning division information in S7135 based on the updated relevant information of UE#1, such as the updated reasoning capability information of UE#1.
[0471] S7134: AS sends inference division information to the IMS core.
[0472] S7135: IMS core sends inference division information to UE#1.
[0473] It should be understood that S7121 to S7125 are similar to S7111 to S7115. For details, please refer to S7111 to S7115 and will not be repeated here.
[0474] S7136: AS applies to the MRF network element for updating inference resources.
[0475] When AS determines that the MRF network element side has updated the reasoning based on the reasoning division information, it can apply to the MRF network element to update the reasoning resources, that is, request the MRF network element to update the reasoning division model that the MRF network element is responsible for reasoning. One possible implementation method is that AS sends a REINVITE message to the MRF network element, and the REINVITE message is used to apply for resource modification. For details, please refer to the description in S7116, which will not be repeated here. It can be understood that in this step, the MRF network element can recreate the endpoint resources, or reuse the endpoint resources created or allocated in S7116; the MRF network element can reuse the computing power resources reserved or allocated in S7116, or reallocate the computing power resources, and the reallocated computing power resources can be different from the computing power resources reserved or allocated in S7116.
[0476] Scenario 4: UE#1 receives a call update.
[0477] S7141: The AS sends a REINVITE message to the IMS core.
[0478] S7142: The IMS core sends a REINVITE message to UE#1.
[0479] S7143, UE#1 sends a 200 For REINVITE message to the IMS core.
[0480] At S7144, the IMS core sends a 200 For REINVITE message to the AS.
[0481] At S7145, the AS determines the inference division of labor information based on the 200 For REINVITE message.
[0482] In an embodiment of the present application, if the model reasoning task needs to be redistributed, for example, due to changes in reasoning requirements and / or changes in the terminal's available computing power for reasoning, the AS can determine new reasoning division information in S7145 based on the updated relevant information of UE#1, such as the updated reasoning capability information of UE#1.
[0483] S7146: AS sends inference division information to the IMS core.
[0484] In a possible implementation manner, the AS sends an ACK For REINVITE message to the IMS core, where the ACK For REINVITE message includes inference division of labor information.
[0485] S7147: IMS core sends inference division information to UE#1.
[0486] In a possible implementation manner, the IMS core sends an ACK For REINVITE message to UE#1, where the ACK For REINVITE message includes inference division of labor information.
[0487] It can be understood that S7141 to S7147 are similar to S7121 to S7127 and will not be repeated here.
[0488] S7148: AS applies to the MRF network element for updating inference resources.
[0489] When AS determines that the MRF network element side has updated the reasoning based on the reasoning division information, it can apply to the MRF network element to update the reasoning resources, that is, request the MRF network element to update the reasoning division model after the MRF network element is responsible for reasoning. One possible implementation method is that AS sends a REINVITE message to the MRF network element, and the REINVITE message is used to apply for resource modification. For details, please refer to the description in S7116, which will not be repeated here. It can be understood that in this step, the MRF network element can recreate the endpoint resources, or reuse the endpoint resources created or allocated in S7128; the MRF network element can reuse the computing power resources reserved or allocated in S7128, or reallocate the computing power resources, and the reallocated computing power resources can be different from the computing power resources reserved or allocated in S7128.
[0490] The above exemplary descriptions are based on four scenarios. Regardless of the above scenario, after determining the reasoning division of labor information, the following steps S7211 to S7214 or S7221 to S7224 may be included. It should be noted that it is assumed below that the reasoning division of labor information instructs UE#1 to perform reasoning of the first part of the model, while the reasoning of the second part of the model is performed by the MRF network element.
[0491] S7211, UE#1 performs reasoning of the first part of the model according to the reasoning division of labor information.
[0492] S7212, UE#1 sends the inference result of the first part of the model to the MRF network element.
[0493] In one possible implementation, UE#1 sends the inference result of the first part of the model to the MRF network element through the IMS core.
[0494] S7213, the MRF network element performs reasoning of the second part of the model based on the reasoning result of the first part of the model.
[0495] Steps S7211 to S7213 are similar to S440a to S460a in method 400, and reference may be made to S440a to S460a.
[0496] S7214, the MRF network element sends the inference result of the second part of the model to UE#2.
[0497] S7221: The MRF network element receives media data from UE#2. For example, if the target model is used for speech-to-text services, the media data is speech information.
[0498] S7222: The MRF network element performs inference of the second part of the model on the media data.
[0499] S7223, the MRF network element sends the inference result of the second part of the model to UE#1.
[0500] In one possible implementation, the MRF network element sends the inference result of the second part of the model to UE#1 through the IMS core.
[0501] S7224, UE#1 performs reasoning of the first part of the model based on the reasoning results of the second part of the model and the reasoning division of labor information.
[0502] Steps S7222 to S7223 are similar to S440b to S460b in method 400, and reference may be made to S440b to S460b.
[0503] Based on the above technical solution, a terminal can report its reasoning capability information to the AS during IMS registration, and the AS can also notify the terminal of the MRF network element's reasoning capability information. If the terminal determines that the MRF network element has model reasoning capabilities, it can request reasoning negotiation in a call request, including the terminal's capability information, session information, and / or the terminal's current connection bandwidth. The AS can determine the reasoning division of labor based on the terminal's capability information, session information, and / or the terminal's current connection bandwidth. For example, the AS can calculate the computing power required for the reasoning task and determine whether it exceeds the terminal's available computing power. If so, the AS can also combine the MRF network element's reasoning capability information to determine the reasoning division of labor for the session and notify the terminal of the negotiation results. When the MRF network element participates in model reasoning, the AS requests reasoning resources from the MRF network element to enable it to perform model reasoning. The terminal and AS then each perform reasoning for the model they are responsible for. Furthermore, if, during a session, due to circumstances such as changes in reasoning requirements, the reasoning task may need to be re-divided, the terminal can re-initiate reasoning negotiation based on the updated reasoning requirements.
[0504] It is understood that in the solution shown in Figure 7, the division of labor is determined by the AS (e.g., S7113, 7125, etc.). Alternatively, the division of labor can be determined by the terminal. For example, UE#1 can determine the inference division of labor information based on UE#1's inference capability information, session information, and UE#1's current connection bandwidth. Optionally, before executing S7211 or S7221, UE#1 can also send the inference division of labor information to the AS network element.
[0505] Figure 8 is a schematic flow chart of a model inference method 800 provided in an embodiment of the present application. This method 800 can be used to implement the solution of the above-mentioned method 400. This method 800 can complete the inference negotiation between the terminal and the network based on HTTP / HTTPS messages. The inference negotiation can be completed, for example, during the call establishment process. In this method 800, it is assumed that the first network device and the second network device are OTT servers. As an example, the method 800 shown in Figure 8 can be used in the architecture of Figure 3. The steps in this method 800 are described below.
[0506] S801, UE#1 sends the reasoning capability information of UE#1 to the OTT server.
[0507] In one possible implementation, UE#1 sends reasoning capability information to the OTT server during registration with the OTT server. For example, UE#1 sends an HTTP request message to the OTT server, and the HTTP request message carries the reasoning capability information of UE#1.
[0508] As an example, the format of the reasoning capability information of UE#1 may be in JSON format, as follows:
[0509] Content-Type:application / json
[0510] {
[0511] “AI-Capability”:
[0512] "AI-Inference": "1.2TFlops", / / indicates the available computing power for inference of UE#1
[0513] "AI-Precision": "INT16", / / indicates the inference accuracy of UE#1
[0514] "Cache":"1G", / / indicates the inference available cache of UE#1
[0515] "AI-Framework": "WebNN v1.2", / / indicates the AI / ML framework and version number used by UE#1 for reasoning
[0516] "Intermediate-compression": "FC_VCM" / / indicates the intermediate data compression algorithm supported by UE#1
[0517] }
[0518] As another example, the format of the reasoning capability information of UE#1 may be in XML format, as follows:
[0519] Content-Type:application / xml
[0520] <ai-capability>
[0521] <AI-Inference Value="1.2TFlops”> / / Indicates the available computing power for inference of UE#1
[0522] <AI-Precision Value="INT16”> / / Indicates the inference accuracy of UE#1
[0523] <Cache Value="1G”> / / Indicates the inference available cache of UE#1
[0524] <AI-Framework Value="WebNN v1.2”> / / Indicates the AI / ML framework and version number used by UE#1 for reasoning.
[0525] <Intermediate-compression Value="FC_VCM”> / / Indicates the intermediate data compression algorithm supported by UE#1
[0526]
[0527] For the above parameters, please refer to the relevant descriptions in the previous article and will not be repeated here.
[0528] S802, the OTT server saves the reasoning capability information of UE#1.
[0529] For example, after the OTT server successfully registers and authenticates UE#1, it can save (or record) the reasoning capability information of UE#1.
[0530] Optionally, the method 800 further includes S803 to S804.
[0531] S803, the OTT server sends the reasoning capability information of the OTT server to UE#1.
[0532] In one possible implementation, the OTT server sends an HTTP response message to UE#1, where the HTTP response message includes the reasoning capability information of the OTT server.
[0533] S804, UE#1 determines that the OTT server has model reasoning capabilities.
[0534] UE#1 can determine whether the OTT server has model reasoning capabilities based on the OTT server's reasoning capability information. Steps S803 to S804 are optional, meaning the OTT server does not need to provide UE#1 with its reasoning capability information. UE#1 can assume that the OTT server has model reasoning capabilities.
[0535] The following describes the reasoning negotiation process using four scenarios.
[0536] Scenario 1: UE#1 initiates inference negotiation.
[0537] S8111, UE#1 sends an HTTP request message to the OTT server.
[0538] It is understandable that the reasoning capability information of UE#1 reported by UE#1 in S801 may also be sent in S8111. That is, method 800 may not include 801, and in S8111, the HTTP request message includes the reasoning capability information of UE#1.
[0539] Optionally, the HTTP request message may further include information about the session and / or the current connection bandwidth of UE#1.
[0540] In one possible scenario, UE#1 assumes that the OTT server has model reasoning capabilities, and therefore, UE#1 includes information about the session and / or UE#1's current connection bandwidth in the HTTP request message.
[0541] In another possible scenario, UE#1 determines whether the OTT server has model reasoning capabilities based on the reasoning capability information of the OTT server received in S803. If UE#1 determines that the OTT server has model reasoning capabilities, UE#1 carries the information of the session and / or UE#1's current connection bandwidth in the HTTP request message. If UE#1 determines that the OTT server does not have model reasoning capabilities, UE#1 does not carry the information of the session and / or UE#1's current connection bandwidth in the HTTP request message. For the case where the OTT server does not have model reasoning capabilities, as an example, UE#1 can perform model reasoning by itself. The embodiments of the present application mainly introduce the case where the OTT server has model reasoning capabilities.
[0542] As an example, the format of the parameters carried in the HTTP request message may be in JSON format, as follows:
[0543] Content-Type:application / json
[0544] {
[0545] “AI-Capability”:
[0546] "AI-Inference": "1.2TFlops", / / indicates the available computing power for inference of UE#1
[0547] "AI-Precision": "INT16", / / indicates the inference accuracy of UE#1
[0548] "Cache":"1G", / / indicates the inference available cache of UE#1
[0549] "AI-Framework": "WebNN v1.2", / / indicates the AI / ML framework and version number used by UE#1 for reasoning
[0550] "Intermediate-compression": "FC_VCM", / / indicates the intermediate data compression algorithm supported by UE#1
[0551] “AI-Model-Type”: “NLP”, / / indicates the model type
[0552] "Bandwidth": "256K" / / indicates the current connection bandwidth of UE#1
[0553] }
[0554] As another example, the format of the parameters carried in the HTTP request message may be in XML format, as follows:
[0555] Content-Type:application / xml
[0556] <ai-capability>
[0557] <AI-Inference Value="1.2TFlops"> / / Indicates the available inference computing power of UE#1
[0558] <AI-Precision Value="INT16"> / / Indicates the inference precision of UE#1
[0559] <Cache Value="1G"> Indicates the available inference cache of UE#1
[0560] <AI-Framework Value="WebNN v1.2"> / / Indicates the AI / ML framework and version number used for inference by UE#1
[0561] <Intermediate-compression Value="FC_VCM"> / / Indicates the intermediate data compression algorithm supported by UE#1
[0562] <AI-Model-Type Value="NLP"> / / Indicates the model type
[0563] <Bandwidth Value="256K" / / Indicates the current connection bandwidth of UE#1
[0564] Regarding the above parameters, reference can be made to the relevant descriptions in the previous text, which will not be elaborated here.
[0565] S8112, The OTT server determines the inference division information according to the HTTP request message.
[0566] Exemplarily, the OTT server can determine the inference division information according to one or more of the following: the inference capability information of the OTT server, the inference capability information of UE#1, the information of this session, and the current connection bandwidth of UE#1. Regarding model inference information #1 and how to determine model inference information #1, reference can be made to the relevant descriptions in the previous text.
[0567] It should be understood that model inference information #1 is a case of inference division information. That is, if the OTT server determines that UE#1 performs the inference of the first part of the model, the inference division information is model inference information #1. Regarding the inference division information and how to determine the inference division information, reference can be made to the relevant descriptions in the previous text related to model inference information #1.
[0568] It should be noted that the information of this session can be carried by the HTTP request message or determined by the OTT server itself. For example, the OTT server can determine the information of this session according to the application business logic and session parameters.
[0569] S8113, the OTT server sends reasoning division of labor information to UE#1.
[0570] In one possible implementation, the OTT server sends an HTTP response message to UE#1. The HTTP response message carries the inference negotiation result of this session, that is, the inference division of labor information.
[0571] As an example, the format of the reasoning division of labor information may be in JSON format, as follows:
[0572] Content-Type:application / json
[0573] {
[0574] "AI-Negotiation-Response":{
[0575] "Result":"success",
[0576] "AI-Model-Identifier":"model_xxx",
[0577] "AI-Model-Param":"11M",
[0578] "Output-Data-Type":"intermediate",
[0579] "Output-Tensor-Shape":"[1,64,64]",
[0580] "Output-Tensor-Structure":"Pytorch2.0",
[0581] "Intermediate-Compression":"FC_VCM"}
[0582] }
[0583] As another example, the format of the reasoning division of labor information may be in XML format, as follows:
[0584] Content-Type:application / xml
[0585] <ai-negotiation-response>
[0586] <Result Value="success”>
[0587] <AI-Model-Identifier Value="model_xxx”>
[0588] <AI-Model-Param Value="11M”>
[0589] <Output-Data-Type Value="intermediate”>
[0590] <Output-Tensor-Shape Value="[1,64,64]”>
[0591] <Output-Tensor-Structure Value="Pytorch2.0”>
[0592] <Intermediate-Compression Value="FC_VCM”>
[0593]
[0594] For the fields involved in the above two examples, please refer to the relevant description in S6114, which will not be repeated here.
[0595] Scenario 2: UE#1 receives inference negotiation.
[0596] S8121, the OTT server sends an HTTP request message to UE#1.
[0597] Optionally, the HTTP request message is used to request the reasoning capability information of UE#1. Alternatively, the HTTP request message is used to request session information and / or the current connection bandwidth of UE#1.
[0598] S8122, UE#1 sends an HTTP response message to the OTT server.
[0599] The HTTP response message may carry the parameters requested by the IHTTP request message, such as the reasoning capability information of UE#1, session information, and / or the current connection bandwidth of UE#1.
[0600] It is understandable that the reasoning capability information of UE#1 reported by UE#1 in S801 may also be sent in S8122. That is, method 800 may not include S801, and in S8122, the HTTP response message includes the reasoning capability information of UE#1.
[0601] S8123, the OTT server determines the model division of labor information based on the HTTP response message.
[0602] Exemplarily, the OTT server may determine the reasoning division of labor information based on one or more of the following: the reasoning capability information of the OTT server, the reasoning capability information of UE#1, the information of the session, and the current connection bandwidth of UE#1.
[0603] It should be understood that model reasoning information #1 is one example of reasoning division of labor information. That is, if the OTT server determines that UE #1 will perform reasoning for the first part of the model, the reasoning division of labor information is model reasoning information #1. For more information about reasoning division of labor information and how to determine it, please refer to the previous description related to model reasoning information #1.
[0604] It should be noted that the session information can be carried in an HTTP response message, or can be determined by the OTT server itself. For example, the OTT server can determine the session information based on application business logic and session parameters.
[0605] S8124, the OTT server sends model division information to UE#1.
[0606] In one possible implementation, the OTT server sends an HTTP message to UE#1, where the HTTP message includes model division information.
[0607] S8124 is similar to S8113 and is not described here.
[0608] Scenario 3: UE#1 initiates inference renegotiation (or model inference renegotiation).
[0609] S8131, UE#1 sends a session update call request to the OTT server.
[0610] In a possible implementation, UE#1 sends an HTTP message to the OTT server, where the HTTP message includes a session update call request, or the HTTP message is used to request a session update call.
[0611] In some cases, such as changes in the inference requirements of the inference service and / or the terminal's available inference computing power during a session, the terminal may initiate a call update. In these cases, UE#1 can send a session update request to the OTT server to renegotiate inference. For example, the session update request may include information about UE#1's inference capabilities and its current connection bandwidth, and may also include information about the session.
[0612] S8132: The OTT server updates the call request based on the session and determines the model division of labor information.
[0613] In an embodiment of the present application, if the model reasoning task needs to be redistributed, for example, due to changes in the reasoning requirements and / or changes in the terminal's available computing power for reasoning, the OTT server can determine new reasoning division information in S8132 based on the updated relevant information of UE#1, such as the updated reasoning capability information of UE#1.
[0614] S8133, the OTT server sends reasoning division of labor information to UE#1.
[0615] It should be understood that S8131 to S8133 are similar to S8111 to S8113 and will not be repeated here.
[0616] Scenario 4: UE#1 receives inference renegotiation (or model inference renegotiation).
[0617] S8141, the OTT server sends a session update call request to UE#1.
[0618] S8142, UE#1 sends a session update call response to the OTT server.
[0619] S8143, the OTT server updates the call response based on the session and determines the inference division of labor information.
[0620] In an embodiment of the present application, if the model reasoning task needs to be redistributed, for example, due to changes in reasoning requirements and / or changes in the terminal's available computing power for reasoning, the OTT server can determine new reasoning division information in S8143 based on the updated relevant information of UE#1, such as the updated reasoning capability information of UE#1.
[0621] S8144, the OTT server sends inference division of labor information to UE#1.
[0622] It should be understood that S8141 to S8144 are similar to S8121 to S8124 and will not be repeated here.
[0623] The above examples illustrate four scenarios. Regardless of the scenario, after determining the inference division information, the following steps S8211 to S8214 or S8221 to S8224 may be included. It should be noted that the following assumes that the inference division information instructs UE#1 to perform inference on the first model, while the OTT server performs inference on the second model.
[0624] S8211, UE#1 performs reasoning of the first part of the model according to the reasoning division of labor information.
[0625] S8212, UE#1 sends the inference result of the first part of the model to the OTT server.
[0626] S8213, the OTT server performs reasoning of the second part of the model based on the reasoning result of the first part of the model.
[0627] Steps S8211 to S8213 are similar to S440a to S460a in method 400, and reference may be made to S440a to S460a.
[0628] S8214, the OTT server sends the inference result of the second part of the model to UE#2.
[0629] S8221: The OTT server receives media data from UE#2. For example, if the target model is used for speech-to-text services, the media data is voice information.
[0630] S8222: The OTT server performs inference on the second part of the model on the media data.
[0631] S8223, the OTT server sends the inference result of the second part of the model to UE#1.
[0632] In one possible implementation, the MRF network element sends the inference result of the second part of the model to UE#1 through the IMS core.
[0633] S8224, UE#1 performs reasoning of the first part of the model based on the reasoning results of the second part of the model and the reasoning division of labor information.
[0634] Steps S8222 to S8224 are similar to S440b to S460b in method 400, and reference may be made to S440b to S460b.
[0635] It will be appreciated that the above method 800 is primarily illustrated using an OTT server as an example. As previously described, an OTT server may include a signaling server (or OTT signaling server), a media processing server (or OTT media server, or OTT media processing server), and a routing server, with different servers performing different functions. For example, the first network device is a signaling server, and the second network device is a media processing server. For details, please refer to the description of method 600 or 700, which will not be repeated here.
[0636] Based on the above technical solution, when the terminal registers with the OTT server, the terminal and the OTT server can exchange their respective reasoning capability information. When the terminal determines that the OTT server has the model reasoning capability, it can conduct reasoning negotiation with the OTT server through the session call request. For example, the OTT server can calculate the computing power requirements of the reasoning task and determine whether it exceeds the terminal's available reasoning computing power. If it exceeds the terminal's available reasoning computing power, the OTT server can also combine its own reasoning capability information to decide the reasoning division of this session and notify the terminal of the reasoning negotiation results. Then, the terminal and the OTT server respectively perform the reasoning of the model they are responsible for. In addition, if due to some circumstances during the session, such as changes in reasoning requirements, the reasoning task may need to be redivided, the terminal can re-initiate the reasoning negotiation and re-negotiate the reasoning based on the updated reasoning requirements.
[0637] It should be noted that in some of the above embodiments, the description of sending the inference output data to the other end of the session (e.g., another terminal (such as UE#2) or a server) is merely illustrative. In one design, the inference output data may be sent to the other end of the session, or data encoded with the inference output data, or data processed by other means, may be sent to the other end of the session. The specific method for processing the inference output data or sending the inference output data to the other end of the session can be determined based on specific requirements, and reference can be made to the prior art.
[0638] It should also be noted that in some of the above embodiments, this application describes operations related to the inference output data or the received data obtained based on the inference output data (e.g., encoding, decoding, compression, or decompression). It should be understood that these descriptions are merely illustrative and do not constitute any limitation on this application. For example, in some implementations, it may be necessary to perform other direct or indirect processing on the inference output data or the received data obtained based on the inference output data.
[0639] It is understood that in the solution shown in FIG8 , the division of labor is determined by the OTT server (e.g., S8112, S8123, etc.). Alternatively, the division of labor may be determined by the terminal. For example, UE#1 may determine the division of labor based on its reasoning capability information, session information, and UE#1's current connection bandwidth.
[0640] It should be understood that the examples in Figures 5 to 8 of the embodiments of the present application are merely intended to facilitate understanding of the embodiments of the present application by those skilled in the art, and are not intended to limit the embodiments of the present application to the specific scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or variations based on the examples in Figures 5 to 8, and such modifications or variations also fall within the scope of the embodiments of the present application.
[0641] It should also be understood that the steps in Figures 5 to 8 are merely illustrative and not intended to be strict limitations. Furthermore, the sequence numbers of the aforementioned processes do not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation of the embodiments of this application.
[0642] It is also understood that in some of the above embodiments, network elements in existing network architectures are mainly used as examples for illustrative purposes. It should be understood that the embodiments of the present application do not limit the specific form of the network elements. For example, network elements that can achieve the same functions in the future are applicable to the embodiments of the present application.
[0643] It can also be understood that in some of the above embodiments, reasoning on the first partial model and executing reasoning on the first partial model can be replaced with each other, and similarly, reasoning on the second partial model and executing reasoning on the second partial model can be replaced with each other.
[0644] It can also be understood that in some of the above embodiments, sending media data can also be replaced by sending data.
[0645] It should also be understood that the message names mentioned in the various embodiments of this application do not limit the scope of protection of the embodiments of this application. For example, the REINVITE message and the 200 For REINVITE message are both exemplary, and the 200 For REINVITE message can be understood as a response message to the REINVITE message. For another example, the HTTP request message and HTTP response message are named for differentiation; both are HTTP messages. It should be understood that, using the example of A sending a message to B, any message that can be used between A and B is applicable to the embodiments of this application.
[0646] It is also understood that in some of the above embodiments, sending messages is mentioned multiple times. Taking A sending a message to B as an example, A sending a message to B may include A sending the message directly to B or A sending the message to B through other devices or network elements, and there is no limitation on this.
[0647] It can also be understood that in the above-mentioned various method embodiments, the methods and operations implemented by the device can also be implemented by components of the device (such as chips or circuits), without limitation.
[0648] Corresponding to the methods provided in the above method embodiments, embodiments of the present application also provide corresponding apparatuses, which include modules for executing the corresponding methods in the above method embodiments. The modules may be software, hardware, or a combination of software and hardware. It is understood that the technical features described in the above method embodiments are also applicable to the following apparatus embodiments.
[0649] Figure 9 is a schematic diagram of a communication device 900 provided in an embodiment of the present application. The device 900 includes a transceiver unit 910 and a processing unit 920. The transceiver unit 910 can be used to implement corresponding communication functions. The transceiver unit 910 can also be referred to as a communication interface or a communication unit. The processing unit 920 can be used to perform processing operations, such as determining a rendering mode.
[0650] Optionally, the device 900 also includes a storage unit, which can be used to store instructions and / or data, and the processing unit 920 can read the instructions and / or data in the storage unit so that the device implements the actions of the network device (for example, the first network device or the second network device) or the terminal in the aforementioned method embodiments.
[0651] In the first design, the device 900 can be the terminal in the aforementioned embodiment (such as the terminal in Figure 4, the UE in Figures 5 to 8), or it can be a component of the terminal (such as a chip). The device 900 can implement the steps or processes executed by the terminal in the above method embodiment. Among them, the transceiver unit 910 can be used to perform the operations related to transceiving of the terminal in the above method embodiment (such as the operations of sending and / or receiving data or messages). For example, the transceiver unit 910 can be used to perform the operations of the terminal and / or receiving data or messages in Figure 4, and can also be used to perform the operations of sending and / or receiving data or messages of the UE in Figures 5 to 8. The processing unit 920 can be used to perform the processing-related operations of the terminal in the above method embodiment, or operations other than transceiving (such as operations other than sending and / or receiving data or messages). For example, the processing unit 920 can be used to perform S440a or S440b in Figure 4. For another example, the processing unit 920 can be used to perform the processing operations of the UE in Figures 5 to 8.
[0652] In one possible implementation, the transceiver unit 910 is used to receive model inference information from the first network device, and the model inference information is used for the terminal to perform inference of the first part model related to the session; the processing unit 920 is used to perform inference of the first part model based on the model inference information; the transceiver unit 910 is also used to send the inference result of the first part model to the second network device.
[0653] In one possible implementation, the transceiver unit 910 is used to receive model inference information from a first network device, and the model inference information is used by the terminal to perform inference of a first part model related to a session; the transceiver unit 910 is also used to receive an inference result of a second part model related to the session from the second network device; and the processing unit 920 is used to perform inference of the first part model based on the inference result of the second part model and the model inference information.
[0654] Optionally, the transceiver unit 910 is further configured to send the reasoning capability information of the terminal to the first network device.
[0655] Optionally, the reasoning capability information includes one or more of the following: available computing power for reasoning, reasoning accuracy, available cache for reasoning, AI / ML framework and version number used for reasoning, and supported intermediate data compression algorithms.
[0656] Optionally, the available computing power for inference is determined by the terminal based on the configuration of its computing resources, or the available computing power for inference is determined by the terminal based on its remaining computing resources or available computing resources.
[0657] Optionally, the transceiver unit 910 is further configured to send the session information and / or the current connection bandwidth of the terminal to the first network device.
[0658] Optionally, the model reasoning information includes an identifier of the first partial model.
[0659] Optionally, the model inference information also includes one or more of the following: the number of parameters of the first part model, the type of inference output data, the tensor shape of the intermediate data, the tensor structure of the intermediate data, or the compression algorithm of the intermediate data.
[0660] Optionally, the transceiver unit 910 is specifically configured to send a first message to the first network device, where the first message includes the reasoning capability information, and the first message is a registration request message or a session call message.
[0661] Optionally, the second network device is an artificial intelligence AI / deep learning ML media plane network element, and the first network device is an AI / ML control plane network element; or, the second network device is a media resource function network element, and the first network device is an application server; or, the second network device is an Internet service media server, and the first network device is an Internet service signaling server.
[0662] Optionally, the transceiver unit 910 is specifically configured to send the reasoning capability information to the first network device when the second network device has model reasoning capability.
[0663] Optionally, the transceiver unit 910 or the processing unit 920 is further configured to obtain capability indication information of the second network device, where the capability indication information indicates that the second network device has model reasoning capability.
[0664] In the second design, the device 900 can be the first network device in the aforementioned embodiment (such as the first network device in Figure 4, the network device in Figure 5, the XR-C network element in Figure 6, the AS in Figure 7, and the OTT server in Figure 8), or it can be a component of the first network device (such as a chip). The device 900 can implement the steps or processes corresponding to those performed by the first network device in the above method embodiment. Among them, the transceiver unit 910 can be used to perform the transceiver-related operations of the first network device in the above method embodiment (such as the operation of sending and / or receiving data or messages), and can also be used to perform the operation of the network device in Figure 5 to send and / or receive data or messages, and can also be used to perform the operation of the XR-C network element in Figure 6 to send and / or receive data or messages, and can also be used to perform the operation of the AS in Figure 7 to send and / or receive data or messages, and can also be used to perform the operation of the OTT server in Figure 8 to send and / or receive data or messages. The processing unit 920 can be used to perform operations related to data and / or information processing of the first network device in the above method embodiment, or operations other than sending and receiving (such as operations other than sending and / or receiving data or messages). For example, the processing unit 920 can be used to perform operations related to data and / or information processing of the first network device in Figure 4, and can also be used to perform operations related to data and / or information processing of the network device in Figure 5, and can also be used to perform operations related to data and / or information processing of the XR-C network element in Figure 6, and can also be used to perform operations related to data and / or information processing of the AS in Figure 7, and can also be used to perform operations related to data and / or information processing of the OTT server in Figure 8.
[0665] In one possible implementation, the processing unit 920 is used to determine model inference information, which is used by the terminal to perform inference on the first part of the model related to the session; and the transceiver unit 910 is used to send the model inference information to the terminal.
[0666] Optionally, the model reasoning information includes an identifier of the first partial model.
[0667] Optionally, the model inference information also includes one or more of the following: the number of parameters of the first part model, the type of inference output data, the tensor shape of the intermediate data, the tensor structure of the intermediate data, or the compression algorithm of the intermediate data.
[0668] Optionally, the processing unit 920 is specifically used to determine the model inference information based on one or more of the following: the reasoning capability information of the terminal, the reasoning capability information of the second network device, the information of the session, or the current connection bandwidth of the terminal, wherein the second network device is used to perform the inference of the second part of the model related to the session.
[0669] Optionally, the transceiver unit 910 is further configured to receive reasoning capability information of the terminal from the terminal.
[0670] Optionally, the reasoning capability information of the terminal is carried in a registration request message or a session call message.
[0671] Optionally, the transceiver unit 910 is further configured to send capability indication information of the second network device to the terminal, where the capability indication information indicates that the second network device has model reasoning capability.
[0672] Optionally, the transceiver unit 910 is further configured to send an identifier of the second partial model related to the session to the second network device.
[0673] Optionally, the transceiver unit 910 is specifically used to send a resource request message to the second network device, the resource request message including the identifier of the second part model, the resource request message is used to request the second network device to reserve resources, and the resources are used to perform reasoning of the second part model.
[0674] Optionally, the resource request information further includes address information of the terminal.
[0675] Optionally, the transceiver unit 910 is specifically configured to send indication information to the second network device, wherein the indication information instructs the second network device to perform inference of the second partial model, and the indication information includes an identifier of the second partial model.
[0676] In the third design, the device 900 can be the second network device in the aforementioned embodiment (such as the second network device in Figure 4, the network device in Figure 5, the XR-M network element in Figure 6, the MRF network element in Figure 7, and the OTT server in Figure 8), or it can be a component of the second network device (such as a chip). The device 900 can implement the steps or processes corresponding to those performed by the second network device in the above method embodiment. Among them, the transceiver unit 910 can be used to perform the transceiver-related operations of the second network device in the above method embodiment (such as the operation of sending and / or receiving data or messages), and can also be used to perform the operation of the network device in Figure 5 to send and / or receive data or messages, and can also be used to perform the operation of the XR-M network element in Figure 6 to send and / or receive data or messages, and can also be used to perform the operation of the MRF network element in Figure 7 to send and / or receive data or messages, and can also be used to perform the operation of the OTT server in Figure 8 to send and / or receive data or messages. The processing unit 920 can be used to perform operations related to data and / or information processing of the second network device in the above method embodiment, or operations other than sending and receiving (such as operations other than sending and / or receiving data or messages). For example, the processing unit 920 can be used to perform operations related to data and / or information processing of the second network device in Figure 4, and can also be used to perform operations related to data and / or information processing of the network device in Figure 5, and can also be used to perform operations related to data and / or information processing of the XR-M network element in Figure 6, and can also be used to perform operations related to data and / or information processing of the MRF network element in Figure 7, and can also be used to perform operations related to data and / or information processing of the OTT server in Figure 8.
[0677] In one possible implementation, the transceiver unit 910 is used to receive the identifier of the second part model from the first network device; the transceiver unit 910 is also used to receive the inference result of the first part model from the terminal; and the processing unit 920 is used to perform inference of the second part model based on the inference result of the first part model.
[0678] A possible implementation method is that the transceiver unit 910 is used to receive the identifier of the second part model from the first network device; the processing unit 920 is used to perform reasoning of the second part model; the transceiver unit 910 is also used to send the inference result of the second part model obtained by executing the reasoning of the second part model to the terminal.
[0679] Optionally, the transceiver unit 910 is specifically used to receive a resource request message from the first network device, the resource request message including the identifier of the second part model, the resource request message being used to request the second network device to reserve resources, and the resources being used to execute reasoning of the second part model.
[0680] Optionally, the resource request message further includes address information of the terminal.
[0681] Optionally, the transceiver unit 910 is specifically configured to receive instruction information from the first network device, where the instruction information instructs the second network device to perform inference of the second partial model, and the instruction information includes an identifier of the second partial model.
[0682] It can be understood that the specific process of each unit executing the above corresponding steps has been described in detail in the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0683] It can also be understood that the device 900 here is embodied in the form of a functional unit. The term "unit" here can refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combined logic circuit and / or other suitable components that support the described functions. In an optional example, those skilled in the art can understand that the device 900 can be specifically a terminal in the above embodiment (such as the terminal in Figure 4, the UE in Figures 5 to 8), and can be used to execute the various processes and / or steps corresponding to the terminal in the above method embodiments, for example, the steps performed by the terminal in Figure 4 and the steps performed by the UE in Figures 5 to 8. Alternatively, the apparatus 900 may be specifically the first network device in the above-mentioned embodiments (such as the first network device in FIG. 4 , the network device in FIG. 5 , the XR-C network element in FIG. 6 , the AS in FIG. 7 , and the OTT server in FIG. 8 ), and may be used to execute the various processes and / or steps corresponding to the first network device in the above-mentioned method embodiments, for example, the steps executed by the first network device in FIG. 4 , the steps executed by the network device in FIG. 5 , the steps executed by the XR-C network element in FIG. 6 , the steps executed by the AS in FIG. 7 , and the steps executed by the OTT server in FIG. 8 . Alternatively, the apparatus 900 may be specifically the second network device in the above-mentioned embodiments (such as the second network device in FIG. 4 , the network device in FIG. 5 , the XR-M network element in FIG. 6 , the MRF network element in FIG. 7 , and the OTT server in FIG. 8 ), and may be used to execute the various processes and / or steps corresponding to the second network device in the above-mentioned method embodiments, for example, the steps executed by the second network device in FIG. 4 , the steps executed by the network device in FIG. 5 , the steps executed by the XR-M network element in FIG. 6 , the steps executed by the MRF network element in FIG. 7 , and the steps executed by the OTT server in FIG. 8 . To avoid repetition, I will not go into details here.
[0684] The apparatus 900 of each of the above-mentioned schemes has the function of implementing the corresponding steps performed by the terminal in the above-mentioned method (such as the terminal in Figure 4, the UE in Figures 5 to 8); or, the apparatus 900 of each of the above-mentioned schemes has the function of implementing the corresponding steps performed by the first network device in the above-mentioned method (such as the first network device in Figure 4, the network device in Figure 5, the XR-C network element in Figure 6, the AS in Figure 7, and the OTT server in Figure 8); or, the apparatus 900 of each of the above-mentioned schemes has the function of implementing the corresponding steps performed by the second network device in the above-mentioned method (such as the second network device in Figure 4, the network device in Figure 5, the XR-M network element in Figure 6, the MRF network element in Figure 7, and the OTT server in Figure 8). The functions can be implemented by hardware, or the corresponding software can be implemented by hardware. The hardware or software includes one or more modules corresponding to the above functions; for example, the transceiver unit can be replaced by a transceiver (for example, the sending unit in the transceiver unit can be replaced by a transmitter, and the receiving unit in the transceiver unit can be replaced by a receiver), and other units, such as the processing unit, can be replaced by a processor to respectively perform the sending and receiving operations and related processing operations in each method embodiment.
[0685] In addition, the transceiver unit 910 may also be a transceiver circuit (for example, may include a receiving circuit and a sending circuit), and the processing unit may be a processing circuit.
[0686] It is understood that the apparatus in FIG9 may be the device in the aforementioned embodiment, or may be a chip or chip system, such as a system on chip (SoC). The transceiver unit may be an input / output circuit or a communication interface; the processing unit may be a processor, microprocessor, or integrated circuit integrated on the chip. This is not limited here.
[0687] Figure 10 is a schematic diagram of another communication device 1000 provided in an embodiment of the present application. The device 1000 includes a processor 1010, which is configured to execute computer programs or instructions stored in a memory 1020, or read data stored in the memory 1020, to perform the methods described in the above method embodiments. Optionally, there are one or more processors 1010.
[0688] Optionally, as shown in FIG10 , the apparatus 1000 further includes a memory 1020 for storing computer programs or instructions and / or data. The memory 1020 may be integrated with the processor 1010 or may be separately provided. Optionally, there may be one or more memories 1020 .
[0689] Optionally, as shown in Figure 10, the apparatus 1000 further includes a transceiver 1030, which is configured to receive and / or transmit signals. For example, the processor 1010 is configured to control the transceiver 1030 to receive and / or transmit signals.
[0690] As a solution, the apparatus 1000 is used to implement the operations performed by the terminal (such as the terminal in FIG. 4 and the UE in FIG. 5 to FIG. 8 ) in the above method embodiments.
[0691] As another solution, the device 1000 is used to implement the operations performed by the first network device (such as the first network device in Figure 4, the network device in Figure 5, the XR-C network element in Figure 6, the AS in Figure 7, and the OTT server in Figure 8) in the above method embodiments.
[0692] As another solution, the device 1000 is used to implement the operations performed by the second network device (such as the second network device in Figure 4, the network device in Figure 5, the XR-M network element in Figure 6, the MRF network element in Figure 7, and the OTT server in Figure 8) in the above method embodiments.
[0693] It is understood that the processor mentioned in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0694] It can also be understood that the memory mentioned in the embodiments of the present application can be a volatile memory and / or a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM). For example, RAM can be used as an external cache. By way of example and not limitation, RAM includes the following forms: static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0695] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) can be integrated into the processor.
[0696] It should also be noted that the memory described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0697] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions for implementing the methods executed by a terminal (such as the terminal in FIG. 4 and the UE in FIG. 5 to FIG. 8 ) in the above-mentioned method embodiments.
[0698] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions for implementing the methods executed by the first network device (such as the first network device in Figure 4, the network device in Figure 5, the XR-C network element in Figure 6, the AS in Figure 7, and the OTT server in Figure 8) in the above-mentioned method embodiments.
[0699] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions for implementing the methods executed by the second network device (such as the second network device in Figure 4, the network device in Figure 5, the XR-M network element in Figure 6, the MRF network element in Figure 7, and the OTT server in Figure 8) in the above-mentioned method embodiments.
[0700] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed by a computer, implement the methods performed by a terminal (such as the terminal in FIG. 4 and the UE in FIG. 5 to FIG. 8 ) in the above-mentioned method embodiments.
[0701] An embodiment of the present application also provides a computer program product comprising instructions, which, when executed by a computer, implement the methods performed by the first network device (such as the first network device in Figure 4, the network device in Figure 5, the XR-C network element in Figure 6, the AS in Figure 7, and the OTT server in Figure 8) in the above-mentioned method embodiments.
[0702] An embodiment of the present application also provides a computer program product comprising instructions, which, when executed by a computer, implement the methods performed by the second network device (such as the second network device in Figure 4, the network device in Figure 5, the XR-M network element in Figure 6, the MRF network element in Figure 7, and the OTT server in Figure 8) in the above-mentioned method embodiments.
[0703] An embodiment of the present application also provides a communication system, including at least one of the aforementioned terminal (such as the terminal in Figure 4, the UE in Figures 5 to 8), a first network device (such as the first network device in Figure 4, the network device in Figure 5, the XR-C network element in Figure 6, the AS in Figure 7, and the OTT server in Figure 8), and a second network device (such as the second network device in Figure 4, the network device in Figure 5, the XR-M network element in Figure 6, the MRF network element in Figure 7, and the OTT server in Figure 8).
[0704] The explanation of the relevant contents and beneficial effects of any of the above-mentioned devices can be referred to the corresponding method embodiments provided above, which will not be repeated here.
[0705] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0706] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. For example, the computer can be a personal computer, a server, or a network device, etc. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)). For example, the aforementioned available medium includes, but is not limited to, various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0707] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A model inference method, characterized in that, Including: Receiving model inference information from a first network device, where the model inference information is used for the terminal to perform inference on a first part of a model related to a session; Performing inference on the first part of the model according to the model inference information and sending the inference result of the first part of the model to a second network device, or Receiving the inference result of a second part of the model related to the session from the second network device and performing inference on the first part of the model according to the inference result of the second part of the model and the model inference information.
2. The method according to claim 1, characterized in that, Before receiving the model inference information from the first network device, the method further includes: Sending the inference capability information of the terminal to the first network device.
3. The method according to claim 2, characterized in that, The inference capability information includes one or more of the following: available computing power for inference, inference accuracy, available cache for inference, artificial intelligence (AI) / deep learning (ML) framework and version number used for inference, and supported intermediate data compression algorithm.
4. The method according to claim 3, characterized in that, The available computing power for inference is determined by the terminal based on the configuration of its computing resources, or the available computing power for inference is determined by the terminal based on its remaining computing resources or available computing resources.
5. The method according to any one of claims 1 to 4, characterized in that Before receiving the model inference information from the first network device, the method further includes: Sending the information of the session and / or the current connection bandwidth of the terminal to the first network device.
6. The method according to any one of claims 1-5, characterized in that, The model inference information includes the identifier of the first part of the model.
7. The method according to claim 6, wherein The model inference information further includes one or more of the following: the number of parameters of the first part of the model, the type of output data of the inference, the tensor shape of the intermediate data, the tensor structure of the intermediate data, or the compression algorithm of the intermediate data.
8. The method according to claim 2, wherein The sending the inference capability information to the first network device includes: Sending a first message to the first network device, where the first message includes the inference capability information, and the first message is a registration request message or a session call message.
9. The method according to any one of claims 1-8, characterized in that, The second network device is an AI / ML media plane network element, and the first network device is an AI / ML control plane network element; or, the second network device is a media resource function network element, and the first network device is an application server; or, the second network device is an Internet service media server, and the first network device is an Internet service signaling server.
10. The method according to claim 2, characterized in that The sending the inference capability information of the terminal to the first network device includes: When the second network device has model inference capability, sending the inference capability information to the first network device.
11. The method according to claim 10, wherein Before sending the inference capability information to the first network device when the second network device has model inference capability, the method further includes: Obtaining capability indication information of the second network device, where the capability indication information indicates that the second network device has model inference capability.
12. A model inference method, characterized in that, Including: Determining model inference information, where the model inference information is used for the terminal to perform inference on a first part of a model related to a session; Sending the model inference information to the terminal.
13. The method according to claim 12, characterized in that, The model inference information includes the identifier of the first part of the model.
14. The method according to claim 13, wherein The model inference information further includes one or more of the following: the number of parameters of the first part of the model, the type of the inference output data, the tensor shape of the intermediate data, the tensor structure of the intermediate data, or the compression algorithm of the intermediate data.
15. The method according to any one of claims 12-14, characterized in that, The determining of the model inference information includes: determining the model inference information according to one or more of the following: the inference capability information of the terminal, the inference capability information of the second network device, the information of the session, or the current connection bandwidth of the terminal, where the second network device is used to perform the inference of the second part of the model related to the session.
16. The method according to claim 15, characterized in that, The method further includes: receiving the inference capability information of the terminal from the terminal.
17. The method according to claim 16, wherein The inference capability information of the terminal is carried in a registration request message or a session call message.
18. The method according to any one of claims 12-17, characterized in that, Before the determining of the model inference information, the method further includes: sending, to the terminal, capability indication information of the second network device, where the capability indication information indicates that the second network device has model inference capability.
19. The method according to any one of claims 12 - 18, characterized in that, The method further includes: sending, to the second network device, the identifier of the second part of the model related to the session.
20. The method according to claim 19, wherein The sending, to the second network device, the identifier of the second part of the model related to the session includes: sending, to the second network device, a resource request message, where the resource request message includes the identifier of the second part of the model, and the resource request message is used to request the second network device to reserve resources for performing the inference of the second part of the model.
21. The method according to claim 20, characterized in that, The resource request information further includes the address information of the terminal.
22. The method according to claim 19, wherein The sending, to the second network device, the identifier of the second part of the model related to the session includes: sending, to the second network device, indication information, where the indication information indicates that the second network device performs the inference of the second part of the model, and the indication information includes the identifier of the second part of the model.
23. A model inference method, characterized in that including: receiving the identifier of the second part of the model from the first network device; receiving the inference result of the first part of the model from the terminal, and performing the inference of the second part of the model according to the inference result of the first part of the model; or, performing the inference of the second part of the model, and sending, to the terminal, the inference result of the second part of the model obtained by performing the inference of the second part of the model.
24. The method according to claim 23, wherein The receiving the identifier of the second part of the model from the first network device includes: receiving a resource request message from the first network device, where the resource request message includes the identifier of the second part of the model, and the resource request message is used to request the second network device to reserve resources for performing the inference of the second part of the model.
25. The method according to claim 24, wherein The resource request message further includes the address information of the terminal.
26. The method according to claim 23, wherein The receiving the identifier of the second part of the model from the first network device includes: receiving indication information from the first network device, where the indication information indicates that the second network device performs the inference of the second part of the model, and the indication information includes the identifier of the second part of the model.
27. A communication device, characterized in that, including a unit or module that performs any of the following: The method according to any one of claims 1-11, the method according to any one of claims 12-22, or the method according to any one of claims 23-26.
28. A communication device, characterized in that, Comprising at least one processor, the at least one processor being configured to execute a computer program stored in a memory, such that the apparatus performs any one of the following: The method according to any one of claims 1-11, the method according to any one of claims 12-22, or the method according to any one of claims 23-26.
29. The device according to claim 28, characterized in that, The apparatus further comprises the memory and / or a communication interface, the communication interface being coupled to the processor, the communication interface being configured to input and / or output information.
30. A communication system, characterized in that, Comprising: A terminal, a first network device, and a second network device, wherein the terminal is configured to perform the method according to any one of claims 1-11, the first network device is configured to perform the method according to any one of claims 12-22, and the second network device is configured to perform the method according to any one of claims 23-26.
31. A communication system, characterized in that, Comprising: A first network device and a second network device, wherein the first network device is configured to perform the method according to any one of claims 12-22, and the second network device is configured to perform the method according to any one of claims 23-26.
32. A computer-readable storage medium, characterized in that, Comprising a computer program, when the computer program runs on a computer, causing the computer to perform any one of the following: The method according to any one of claims 1-11, the method according to any one of claims 12-22, or the method according to any one of claims 23-26.
33. A computer program product, characterized in that, The computer program product comprises computer program code, characterized in that: When the computer program code runs on a computer, causing the computer to implement any one of the following: The method according to any one of claims 1-11, the method according to any one of claims 12-22, or the method according to any one of claims 23-26.
Citation Information
Patent Citations
Model reasoning method and communication device
CN120343006A
Splitting reasoning method and device
CN116776976A
Federated learning method, federated learning system, first device, and third device
WO2022222152A1
Model segmentation assisting method and apparatus, and readable storage medium
WO2023116259A1
Data processing method and apparatus, and communication device
WO2023125879A1