Electronic device, method and storage medium for model reasoning

CN120283239APending Publication Date: 2025-07-08SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080498.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-28
Filing Date
2023-11-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In wireless communication systems, during the AI/ML model reasoning process, a single terminal device has limited computing resources, making it difficult to effectively share the computing load of complex model reasoning, resulting in excessive resource consumption and affecting efficiency and delay.

Method used

By splitting the AI/ML model into multiple sub-parts and having multiple participant devices jointly perform model inference, wireless networks are used to allocate resources to participant devices to achieve effective transmission of model inference information and sharing of computing load.

Benefits of technology

It reduces the resource requirements of a single device, improves the efficiency and latency performance of model inference, and enhances the computing power and communication quality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283239A_ABST
    Figure CN120283239A_ABST
Patent Text Reader

Abstract

The invention relates to an electronic device, a method and a storage medium for model reasoning. Various embodiments for AI / ML model reasoning are described. In one embodiment, an electronic device includes processing circuitry configured to form split information of at least a first portion of an AI / ML model based on respective state information of a first terminal device and other one or more terminal devices, the split information specifies that split model reasoning is to be performed by a plurality of participant devices for a plurality of sub-portions of the at least first portion of the AI / ML model; and causing the wireless network to allocate resources for transmitting model reasoning information to at least one of the plurality of participant devices based on the split information.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method and storage medium for model reasoning Technical Field

[0001] The present disclosure generally relates to wireless communication systems and methods, including techniques for performing artificial intelligence (AI) / machine learning (ML) model inference in wireless communication systems. Background Art

[0002] AI / ML technologies are being applied to a wide range of applications across multiple industries, significantly improving productivity. For example, in wireless communication systems, mobile devices (such as smartphones, smart cars, drones, and mobile robots) are increasingly replacing traditional algorithms (such as speech recognition, machine translation, image recognition, video processing, and user behavior prediction) with AI / ML models to enable a variety of applications. These applications include, for example, enhanced photography, intelligent personal assistants, VR / AR, video gaming, video analysis, personalized shopping recommendations, autonomous driving / navigation, smart home appliances, mobile robotics, mobile healthcare, and mobile finance.

[0003] AI / ML models can be trained and used to perform model inference for specific AI / ML tasks. During model inference, real-world input is passed through the AI / ML model, and a prediction for the task is output. For example, the input can be the pixels of an image or the sampled amplitude of an audio wave. Accordingly, the output of the AI / ML model can be the probability that an image contains a specific object or the probability that an audio sequence contains a specific word. It should be understood that the results of model inference are related to the complexity of the AI / ML model, and the complexity of the AI / ML model is related to the resources consumed by model inference.

[0004] Summary of the Invention

[0005] The first aspect of the present disclosure relates to a model inference method in a wireless communication system, comprising: determining an AI / ML model corresponding to an AI / ML task of a first terminal device; forming splitting information of at least a first part of the AI / ML model based on corresponding status information of the first terminal device and one or more other terminal devices, wherein the splitting information specifies that split model inference is to be performed by multiple participant devices for multiple sub-parts of the at least first part of the AI / ML model; and based on the splitting information, enabling a wireless network to allocate resources for transmitting model inference information to at least one of the multiple participant devices. The first aspect of the present disclosure also relates to an electronic device. The electronic device includes a processing circuit configured to execute the method according to the first aspect. In an embodiment, the electronic device can be used for a terminal device or a network endpoint.

[0006] The second aspect of the present disclosure relates to a model inference method in a wireless communication system, comprising: obtaining split information of at least a first part of an AI / ML model, wherein the AI / ML model corresponds to an AI / ML task of a first terminal device, and the split information specifies that split model inference is to be performed by multiple participant devices for multiple sub-parts of the at least first part of the AI / ML model; and based on the split information, allocating resources for transmitting model inference information to at least one of the multiple participant devices. The second aspect of the present disclosure also relates to an electronic device for a base station. The electronic device includes a processing circuit configured to execute the method according to the second aspect.

[0007] The third aspect of the present disclosure relates to a model reasoning method in a wireless communication system, comprising: receiving an instruction to perform model reasoning from a first terminal device, the instruction including indication information of a corresponding sub-part of an AI / ML model and indication information of a downstream participant device; receiving a resource allocation for transmitting model reasoning information, wherein the resource allocation indicates resources for a wireless link (e.g., including a direct link) with an upstream participant device and the downstream participant device; based on the instruction and the resource allocation, receiving first intermediate data from the upstream participant device via a wireless link with the upstream participant device; inputting the first intermediate data into the corresponding sub-part of the AI / ML model to obtain second intermediate data; and based on the instruction and the resource allocation, sending the second intermediate data to the downstream participant device via a wireless link with the downstream participant device. The third aspect of the present disclosure also relates to an electronic device for a terminal device. The electronic device includes a processing circuit configured to execute a method according to the third aspect.

[0008] A fourth aspect of the present disclosure relates to a computer-readable storage medium having executable instructions stored thereon, which, when executed by one or more processors, implement the operations of the methods according to various embodiments of the present disclosure.

[0009] A fifth aspect of the present disclosure relates to a computer program product comprising instructions which, when executed by a computer, enable implementation of the method according to various embodiments of the present disclosure.

[0010] The above summary is provided to summarize some exemplary embodiments in order to provide a basic understanding of various aspects of the subject matter described herein. Therefore, the above features are merely examples and should not be construed as narrowing the scope or spirit of the subject matter described herein in any way. Other features, aspects, and advantages of the subject matter described herein will become apparent from the detailed description described below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] A better understanding of the present disclosure may be obtained when the following detailed description of the embodiments is considered in conjunction with the accompanying drawings. The same or similar reference numerals are used in the various drawings to represent the same or similar components. The accompanying drawings, together with the following detailed description, are incorporated into and form a part of this specification and are used to illustrate the embodiments of the present disclosure and to explain the principles and advantages of the present disclosure. In particular:

[0012] FIG1 shows an example block diagram of a communication system according to an embodiment of the present disclosure.

[0013] FIG2A shows an example of an AI / ML model according to an embodiment of the present disclosure.

[0014] FIG2B shows an example of a split AI / ML model according to an embodiment of the present disclosure.

[0015] FIG3A shows an exemplary electronic device for a terminal device or a network endpoint according to an embodiment of the present disclosure.

[0016] FIG3B shows an exemplary electronic device for a terminal device according to an embodiment of the present disclosure.

[0017] FIG3C illustrates an exemplary electronic device for a base station according to an embodiment of the disclosure.

[0018] 3D to 3F illustrate an exemplary process for split model inference according to an embodiment of the present disclosure.

[0019] FIG4 illustrates example operations for splitting an AI / ML model according to an embodiment of the present disclosure.

[0020] 5A and 5B illustrate examples of split AI / ML models according to an embodiment of the present disclosure.

[0021] 6A to 6C illustrate examples of split information according to an embodiment of the present disclosure.

[0022] 7A and 7B illustrate example operations for distributing split information according to an embodiment of the present disclosure.

[0023] 8A to 8D illustrate example operations for performing split model inference according to an embodiment of the present disclosure.

[0024] FIG9A shows an example signaling process for allocating transmission resources to participant devices for model inference according to an embodiment of the present disclosure.

[0025] FIG9B illustrates example operations for allocating transmission resources to participant devices for model inference according to an embodiment of the present disclosure.

[0026] FIG10 shows an example method for resource allocation in model reasoning according to an embodiment of the present disclosure.

[0027] FIG11 shows an example method for resource allocation in model inference according to an embodiment of the present disclosure.

[0028] FIG12 illustrates an example method for model reasoning according to an embodiment of the present disclosure.

[0029] FIG13 shows an example block diagram of a computer that can be implemented as a terminal device or a network endpoint according to an embodiment of the present disclosure.

[0030] FIG14 is a block diagram illustrating a first example of a schematic configuration of a gNB to which the technology of the present disclosure may be applied.

[0031] FIG15 is a block diagram illustrating a second example of a schematic configuration of a gNB to which the technology of the present disclosure may be applied.

[0032] FIG. 16 is a block diagram illustrating an example of a schematic configuration of a smartphone to which the technology of the present disclosure can be applied.

[0033] FIG. 17 is a block diagram showing an example of a schematic configuration of a car navigation device to which the technology of the present disclosure can be applied.

[0034] FIG18A shows an example of layer-level computation and communication resource evaluation for the AlexNet model.

[0035] FIG18B shows an example of layer-level computation and communication resource evaluation for the VGG-16 model.

[0036] While the embodiments described in this disclosure may be susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and are herein described in detail. However, it should be understood that the drawings and detailed description thereof are not intended to limit the embodiments to the particular forms disclosed, but on the contrary, the intent is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the claims. DETAILED DESCRIPTION

[0037] The following describes representative applications of various aspects of the apparatus and method of the present disclosure. The description of these examples is only to add context and help understand the described embodiments. Therefore, it is clear to those skilled in the art that the embodiments described below can be implemented without some or all of the specific details. In other cases, well-known process steps are not described in detail to avoid unnecessarily obscuring the described embodiments. Other applications are also possible, and the solutions of the present disclosure are not limited to these examples.

[0038] In general, all terms used herein will be interpreted according to their ordinary meaning in the relevant technical field, unless different meanings and / or implications are clearly given in the context of use. Unless clearly otherwise specified, references to elements, devices, components, units and operations etc. are intended to be openly interpreted as at least one instance in elements, devices, components, units and operations. The operation of any method disclosed herein does not have to be performed in the precise order disclosed, unless the operation is clearly or implicitly described as being after or before another operation. Any feature of any embodiment disclosed herein can be applied to any other appropriate embodiment. Similarly, any advantage of any embodiment can be applied to any other embodiment, and vice versa. Other purposes, features and advantages of the embodiment will become clear from the following description.

[0039] Communication System Example

[0040] Figure 1 shows an example block diagram of a communication system according to an embodiment of the present disclosure. It should be noted that Figure 1 only shows one of many types and possible arrangements of communication systems; the features of the present disclosure can be implemented in any of the various systems as needed.

[0041] As shown in Figure 1, the communication system 100 includes a base station 120 and terminal devices 110A, 110B to 110N. The base station 120 and the terminal devices 110A to 110N can be configured to perform uplink and downlink communications via a Uu interface. The terminal devices 110A to 110N can be configured to perform sidelink communications via a PC5 interface. Accordingly, the base station 120 can allocate transmission resources to the uplink and downlink as well as the sidelink based on the transmission requirements and resource conditions of the specific terminal device. In addition, the base station 120 can be configured to communicate with a network 130 (e.g., a core network of a cellular service provider, a telecommunications network such as a public switched telephone network (PSTN), and / or the Internet). Therefore, the base station 120 can facilitate communication between the terminals 110A to 110N and / or between the terminals 110A to 110N and the network 130, and the terminal devices 110A to 110N can communicate directly within the effective communication range of the sidelink.

[0042] Based on service requirements, use cases, and / or available spectrum, base station 120 can be configured to employ various radio access technologies (RATs). In FIG1 , the coverage area of ​​base station 120 can be referred to as a cell, and base station 120 and other similar base stations (not shown) can provide continuous or nearly continuous communication signal coverage to terminals 110A to 110N over a wide geographic area.

[0043] As shown in Figure 1, the communication system 100 includes a cloud 140, a mobile edge computing (MEC) 150, and an Internet data center (IDC) 160. The cloud 140 can provide services such as IaaS, PaaS, and SaaS to terminal devices via the network 130. Computing resources (e.g., servers) can be deployed in the cloud 140 and MEC 150 to support the computing needs of communication services (e.g., communication and computing convergence services). Generally speaking, the cloud 140 can be deployed on a remote server, and the MEC 150 can be located at a base station, a central office, or any aggregation point in the network. Therefore, compared to the cloud 140, the MEC 150 is closer to the terminal device, which helps reduce network congestion, lower latency, and improve the user's quality of experience (QoE). The IDC 160 can provide hosting services, enabling operation and maintenance of various devices (including computing devices) that centrally collect, store, process, and transmit data over the Internet. In this disclosure, devices in base station 120, cloud 140, MEC 150, IDC 160, and any similar entities in the network may be referred to as network endpoints.

[0044] In the present disclosure, a base station may be a 5G NR base station or a 5G LTE-A base station, such as a gNB and ng-eNB. A gNB can provide NR user plane and control plane protocols for terminal devices. An ng-eNB is a node defined for compatibility with 4G LTE communication systems. It may be an upgrade of the evolved Node B (eNB) of the LTE radio access network and provides Evolved Universal Terrestrial Radio Access (E-UTRA) user plane and control plane protocols for UEs. Furthermore, examples of base stations may include, but are not limited to, at least one of a base transceiver station (BTS) and a base station controller (BSC) in a GSM system; at least one of a radio network controller (RNC) and a Node B in a WCDMA system; an access point (AP) in a WLAN or WiMAX system; and corresponding network nodes in future or developing communication systems. Some of the functions of a base station herein may also be implemented as an entity that controls communications in D2D, M2M, and V2X scenarios, or as an entity that performs spectrum coordination in cognitive radio communication scenarios.

[0045] In the present disclosure, terminal devices may have the full breadth of their usual meanings, for example, terminal devices may be mobile stations (MS), user equipment (UE), etc. Terminal devices may be implemented as, for example, mobile phones, handheld devices, media players, computers, laptops, tablet computers, on-board units (OBU) or vehicles, roadside units (RSU), wearable devices, Internet of Things (IoT) devices, or virtually any type of wireless device. In some cases, terminal devices may communicate using multiple wireless communication technologies. For example, terminal devices may be configured to communicate using one or more of GSM, UMTS, CDMA2000, WiMAX, LTE, LTE-A, WLAN, NR, Bluetooth, etc.

[0046] Split AI / ML models and model reasoning

[0047] Artificial intelligence (AI) is the science and engineering of building intelligent machines capable of performing tasks similar to humans. Subfields of AI include machine learning (ML), which gives computers the ability to learn without being explicitly programmed. Specifically, ML algorithms can be trained to learn to handle new problems without having to create specialized programs to solve each new problem. ML algorithms include, for example, decision trees, K-means clustering, and Bayesian networks. For example, after training the model using data samples, these algorithms can be used for classification and prediction. In the field of ML, neural networks (NNs) are commonly used as models.

[0048] For specific AI / ML tasks, a variety of alternative AI / ML models can be set for model reasoning. For example, for image recognition tasks, alternative AI / ML models that can be used for model reasoning include the AlexNet model, the VGG-16 model, the ResNet-152 model, and the GoogleNet model. The sizes of these models range from tens of megabytes to hundreds of megabytes. In one implementation, the terminal device can download the specific model configuration of a specific model from the network in real time when needed, or the terminal device can semi-statically download the specific model configuration of a specific model from the network through high-level configuration. In one implementation, in order to reduce the amount of data for transmitting the model configuration, the specific model configuration of the configured specific model can be written into the terminal device (such as a chip). The specific model configuration of a specific model may include various model parameters, such as the number of layers of the model, the number and weights of neurons in each layer, and the connection relationship between neurons in the layers.

[0049] FIG2A illustrates an AI / ML model according to an embodiment of the present disclosure. As an example, AI / ML model 200A in FIG2A is a neural network model. As shown in FIG2A , AI / ML model 200A includes multiple layers, including an input layer 201, an output layer 206, and intermediate layers (or hidden layers) 202 through 205. Each layer has a certain number of neurons, each with a specific weight, and neurons in different layers are connected. When an input value 220 is input into AI / ML model 200A, the corresponding value is first received by the neurons in input layer 201 and propagated to the neurons in intermediate layer 202 through connections with neurons in the next layer. The neurons in intermediate layer 202 calculate the weighted sum of the output values ​​of the neurons in the previous layer and output this weighted sum to the neurons in the next intermediate layer 203 through connections with neurons in the next layer. This process continues in this manner until the neurons in output layer 206 calculate the weighted sum of the output values ​​of the neurons in the previous layer and output an inference result 240 for the input value 220.

[0050] In the example of FIG2A , the AI / ML model 200A has four intermediate layers 202 to 205. Depending on the application requirements, the intermediate layers can be of any number, and the present disclosure need not limit this. The AI / ML model 200A is composed of a series of fully connected layers (i.e., all outputs are connected to all inputs), which is called a multilayer perceptron (MLP) model. As a further example, the neural network model also includes a convolutional neural network (CNN) and a recurrent neural network (RNN) model. Although the following description refers more to the multilayer perceptron model, embodiments of the present disclosure can be applied to various other types of AI / ML models.

[0051] Generally speaking, the more complex the AI / ML model is, the greater the amount of computation and storage required for model reasoning through the AI / ML model. Taking a neural network model as an example, the more layers a neural network model has and the richer the connections between neurons, the greater the amount of computation and storage required for model reasoning through the neural network model. Typically, the computing resources used by a single terminal device (e.g., 110A) to support model reasoning are limited. In an embodiment of the present disclosure, other terminal devices (e.g., 110B, 110N) and / or base stations (e.g., 120) may participate in the reasoning process of the AI / ML model to share the resource consumption of a single terminal device (e.g., 110A). For example, the AI / ML model may be split into multiple parts, and each participant device may perform model reasoning only for the corresponding part of the AI / ML model (rather than the entire AI / ML model).

[0052] Figure 2B shows a split AI / ML model according to an embodiment of the present disclosure. As an example, the split AI / ML model 200B is obtained by splitting the AI / ML model 200A. As shown in Figure 2B, the AI / ML model 200A is split into three parts through two splitting points (i.e., layers 203 and 204). Specifically, part I includes the input layer 201 and the intermediate layers 202-203, part II includes the intermediate layers 203-204, and part III includes the intermediate layers 204-205 and the output layer 206. It should be understood that the split parts of the AI / ML model can be other appropriate numbers, and there can be multiple ways to split the AI / ML model, as described in detail below in conjunction with Figures 5A and 5B.

[0053] In an embodiment of the present disclosure, model reasoning performed by multiple participant devices on a split AI / ML model may be referred to as split model reasoning. For example, when a specific AI / ML task of a terminal device 110A originally requires the use of an AI / ML model 200A, the entire AI / ML model 200A may be split into an AI / ML model 200B based on the status information of the terminal device 110A and other terminal devices, and the terminal device 110A and other participant devices (including the terminal device and / or the base station 120) may jointly perform model reasoning for the AI / ML model 200B. Specifically, the terminal device 110A inputs the input value corresponding to the AI / ML task into part I and obtains intermediate data 221 through reasoning. Then, the terminal device 110A transmits the intermediate data 221 to its downstream participant device, such as the terminal device 110B. At the terminal device 110B, the intermediate data 221 is input into part II and obtains intermediate data 222 through reasoning. Then, the terminal device 110B transmits the intermediate data 222 to its downstream participant device, such as the base station 120. At the base station 120, the intermediate data 222 is input into part III and the result data 240 is obtained through inference. Then, the base station 120 can return the result data 240 to the terminal device 110A. It should be understood that the number of participant devices performing model inference can be other appropriate numbers.

[0054] In the above split model reasoning, the terminal device 110A directly related to the AI / ML task and model reasoning can be called the main participant device, and the terminal device 110B and the base station 120 that assist in performing model reasoning can be called auxiliary participant devices. On the one hand, the main participant device performs model reasoning for the first split part (i.e., part I), and the input values ​​corresponding to the AI / ML task that may involve privacy can be input locally into the AI / ML model on the main participant device, thereby avoiding data leakage and improving security. On the other hand, each participant device only needs to perform model reasoning for part I, II or III, thereby reducing the resource requirements of complex model reasoning on a single device.

[0055] In an embodiment of the present disclosure, for model inference with terminal device 110A as the primary participant device, AI / ML model splitting (e.g., splitting AI / ML model 200A into AI / ML model 200B) can be performed by terminal device 110A, base station 120, or any network endpoint related to AI / ML. In addition, auxiliary participant devices may include other terminal devices and network endpoints (including base station 120 or any device with computing power, such as devices in cloud 140, MEC 150, and IDC 160). In some embodiments, auxiliary participant devices include only other terminal devices. In some embodiments, auxiliary participant devices include only network endpoints. In some embodiments, auxiliary participant devices may include both other terminal devices and network endpoints.

[0056] In the case where a network endpoint (e.g., base station 120) is required to participate in model reasoning, the AI / ML model may be pre-split into parts corresponding to the main terminal device 110A and the base station 120, respectively. For example, the AI / ML model 200A may be pre-split into layers 201-204 corresponding to the terminal device 110A and layers 204-206 corresponding to the base station 120. In this case, in order to share the computational load of model reasoning performed by the terminal device 110A, model splitting according to an embodiment of the present disclosure may include splitting at least a portion of the AI / ML model into multiple sub-parts (e.g., splitting layers 201-204 into parts I and II) so that other terminal devices can participate in model reasoning for this part.

[0057] In some embodiments, the primary participant device can be a user device, and the secondary participant devices can include various vehicles. For example, a vehicle can have wireless communication capabilities and AI / ML model inference capabilities. Compared to a user device, a vehicle may have greater computing power and more power reserves, making it suitable for assisting other devices in performing split model inference. In one embodiment, the degree to which the primary and secondary participant devices participate in model inference can be controlled based on the nature of the vehicle. For example, the vehicle can be a public vehicle such as a taxi or bus. Accordingly, the user device needs to perform model inference on a larger model portion (for example, portion I in Figure 2B may be larger) to avoid leaking private data to public vehicles. For another example, the vehicle can be a friend's or one's own private vehicle. Accordingly, while meeting certain privacy requirements, the user device can perform model inference on a suitably smaller model portion (for example, portion I in Figure 2B may be smaller), thereby further leveraging the vehicle's role in assisting model inference. In some cases, the user can even provide local data directly to the vehicle and instruct it to perform model inference, without performing model inference itself.

[0058] It should be understood that split model reasoning requires the transmission of model reasoning information, such as intermediate data and result data, between multiple participant devices. Furthermore, the latency of transmitting model reasoning information should be reasonable to ensure that the entire model reasoning process is completed within the specified time period. In an embodiment of the present disclosure, the base station 120 can allocate resources to multiple participant devices for transmitting model reasoning information between the multiple participant devices, thereby facilitating the execution of split model reasoning.

[0059] Example electronic device

[0060] 3A shows an example electronic device 300 for a terminal device or network endpoint according to an embodiment of the present disclosure. The terminal device may correspond to a primary participant device (e.g., terminal device 110A), and the network endpoint includes, for example, a device in base station 120, cloud 140, MEC 150, or IDC 160.

[0061] The electronic device 300 may include various units to implement various embodiments of AI / ML model splitting and model reasoning according to the present disclosure. In the example of Figure 3A, the electronic device 300 includes an AI / ML task control unit 302 and a transceiver unit 304. For example, the AI / ML task control unit 302 can be configured to split the AI / ML model (e.g., AI / ML model 200), and the transceiver unit 304 can be configured to perform communication with other devices. The various operations described below in conjunction with the terminal device or network endpoint and in conjunction with AI / ML model splitting can be implemented by units 302 to 304 of the electronic device 300 or other possible units.

[0062] In one embodiment, the AI / ML task control unit 302 may generate splitting information for at least the first portion of the AI / ML model 200A based on the corresponding status information of the terminal device 110A and one or more other terminal devices. This splitting information may, for example, specify that split model inference be performed on multiple sub-portions of at least the first portion of the AI / ML model 200A by multiple participant devices. In some examples, at least the first portion of the AI / ML model 200A may correspond to a portion or the entire AI / ML model 200A. Accordingly, the entire AI / ML model 200A may be split into sub-portions I, II, and III; or, if portion III is pre-split, only layers 201-204 of the AI / ML model 200A may be split into sub-portions I and II. In one example, the AI / ML model 200A is a model corresponding to a specific AI / ML task. For example, the AI / ML task may be image recognition, and the AI / ML model 200A is a model trained to recognize content in images.

[0063] In one embodiment, the AI / ML task control unit 302 can, based on the split information, cause the wireless network to allocate resources for transmitting model inference information to at least one of the multiple participant devices. The allocated resources can be used for direct links and / or uplinks and downlinks.

[0064] In one embodiment, the transceiver unit 304 can receive status information from multiple terminal devices for use in splitting the AI / ML model 200A. The transceiver unit 304 can also send a resource allocation request to the network (e.g., the base station 120 or its resource allocation unit) so that the wireless network allocates resources for transmitting model inference information to at least one of the multiple participant devices. The transceiver unit 304 can also be configured to control or execute operations related to signaling or message transmission and reception.

[0065] In an embodiment, the electronic device 300 may be implemented at a chip level, or may be implemented at a device level by including other external components (eg, wired or wireless links). The electronic device 300 may function as a communication device as a whole.

[0066] 3B shows an example electronic device 310 for a terminal device according to an embodiment of the present disclosure. The terminal device may correspond to a primary or secondary participant device.

[0067] The electronic device 310 may include various units to implement various embodiments of AI / ML model reasoning according to the present disclosure. In the example of Figure 3B, the electronic device 310 includes an AI / ML task execution unit 311 and a transceiver unit 314. For example, the AI / ML task execution unit 311 may be configured to perform model reasoning for a sub-part (e.g., part I or II) of an AI / ML model (e.g., AI / ML model 200B). The transceiver unit 314 may be configured to perform communication with a base station or other device, such as transmitting model reasoning information. The various operations described below in conjunction with the terminal device or AI / ML model reasoning may be implemented by units 311 and 314 of the electronic device 310 or other possible units.

[0068] In one embodiment, the AI / ML task execution unit 311 is configured to perform model inference on sub-part I of the split AI / ML model 200B. For example, the AI / ML task execution unit 311 may input input values ​​corresponding to a specific application into sub-part I to obtain intermediate data 221. Accordingly, the transceiver unit 314 may be configured to provide the intermediate data 221 to downstream participant devices (e.g., via a direct link).

[0069] In one embodiment, the AI / ML task execution unit 311 is configured to perform model inference on sub-part II of the split AI / ML model 200B. For example, the transceiver unit 314 can be configured to receive intermediate data from an upstream participant device, and the AI / ML task execution unit 311 can input the intermediate data into sub-part II to obtain intermediate data 222. Accordingly, the transceiver unit 314 can be configured to provide the intermediate data 222 to the downstream participant device (e.g., via a direct link).

[0070] In one embodiment, the AI / ML task execution unit 311 is configured to perform model inference on sub-part III of the split AI / ML model 200B. For example, the transceiver unit 314 can be configured to receive intermediate data from an upstream participant device, and the AI / ML task execution unit 311 can input the intermediate data into sub-part III to obtain result data 240. Accordingly, the transceiver unit 314 can be configured to provide the result data 240 to the primary participant device (e.g., via a direct link).

[0071] Optionally, electronic device 310 may further include an AI / ML task control unit 312. The AI / ML task control unit 312 may be configured to decompose an AI / ML model (e.g., AI / ML model 200A). The operation of the AI / ML task control unit 312 is similar to that of the AI / ML task control unit 302 and may be further understood with reference to the description of electronic device 300.

[0072] In an embodiment, the electronic device 310 may be implemented at a chip level, or may be implemented at a device level by including other external components (eg, a radio link, an antenna, etc.) The electronic device 310 may function as a communication device as a whole.

[0073] FIG3C shows an example electronic device 320 for a base station according to an embodiment of the present disclosure. The base station may correspond to base station 120.

[0074] The electronic device 320 may include various units to implement various embodiments of allocating transmission resources to facilitate AI / ML model reasoning according to the present disclosure. In the example of Figure 3C, the electronic device 320 includes a resource allocation unit 321 and a transceiver unit 324. For example, the resource allocation unit 321 can be configured to allocate resources for transmitting model reasoning information to at least one of a plurality of participant devices. The transceiver unit 304 is configured to perform communications with other network endpoints and / or terminal devices. The various operations described below in conjunction with the base station or resource allocation can be implemented by units 321 and 324 of the electronic device 320 or other possible units.

[0075] In one embodiment, the resource allocation unit 321 may obtain split information for at least a first portion of the AI / ML model. For example, the transceiver unit 324 may receive the split information for at least the first portion of the AI / ML model from a terminal device or a network endpoint. The AI / ML model corresponds to an AI / ML task of a primary participant device, and the split information specifies that split model inference is to be performed across multiple participant devices for multiple sub-portions of at least the first portion of the AI / ML model. Based on the split information, the resource allocation unit 321 may allocate resources for transmitting the model inference information to at least one of the multiple participant devices.

[0076] Optionally, electronic device 320 may further include an AI / ML task control unit 322. AI / ML task control unit 322 may be configured to decompose an AI / ML model (e.g., AI / ML model 200A). The operation of AI / ML task control unit 312 is similar to that of AI / ML task control unit 302 and may be further understood with reference to the description of electronic device 300.

[0077] In an embodiment, the electronic device 320 may be implemented at a chip level, or may be implemented at a device level by including other external components (eg, a radio link, an antenna, etc.) The electronic device 320 may function as a communication device as a whole.

[0078] It should be noted that the above-mentioned units are only logical modules divided according to the specific functions implemented by them, rather than being used to limit the specific implementation method. For example, they can be implemented in software, hardware or a combination of software and hardware. In actual implementation, the above-mentioned units can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). Among them, the processing circuit can refer to various implementations of a digital circuit system, an analog circuit system or a mixed signal (a combination of analog and digital) circuit system that performs functions in a computing system. The processing circuit may include, for example, circuits such as an integrated circuit (IC), an application-specific integrated circuit (ASIC), part or circuit of a separate processor core, an entire processor core, a separate processor, a programmable hardware device such as a field programmable gate array (FPGA), and / or a system including multiple processors.

[0079] Overall process example

[0080] Figures 3D through 3F illustrate exemplary processes for split model inference according to an embodiment of the present disclosure. Processes 330A and 330C are described below in conjunction with terminal devices 110A through 110N and base station 120, where terminal device 110A is the primary participant in model inference, and other devices can act as secondary participants in different scenarios.

[0081] As shown in Figure 3D, at 331, each terminal device reports its own status information (e.g., indicating computing status and / or communication status) to the base station 120. In one embodiment, the report can be periodic or event-based (e.g., in response to a request from the base station 120). At 332, the terminal device 110A, which is the main participant in a specific AI / ML task, sends a model inference request to the base station 120. For example, the request can include AI / ML task indication information or corresponding AI / ML model indication information. At 333, upon receiving the model inference request, the base station 120 generates split information for the split model inference and allocates transmission resources to assist in the execution of the model inference. For example, generating the split information can include the base station 120 determining the model split point between the base station and the terminal device side, and determining the model split point between the terminal devices 110A and 110B participating in the inference. In one example, the model split point between the base station and the terminal device side can be pre-configured for a specific AI / ML model. Generating the split information can also include generating a service flow for each participant device to perform model inference based on the split point. In one example, the service flow may be terminal device 110A->terminal device 110B->base station 120->terminal device 110A. Further, the base station 120 may allocate direct link resources between the terminal devices 110A and 110B based on the splitting information so as to transmit intermediate data between the two terminal devices, allocate uplink resources between the terminal device 110B and the base station 120, and downlink resources between the base station 120 and the terminal device 110A so as to transmit intermediate data and result data between the terminal device and the base station 120. At 334 and 335, the base station 120 transmits an inference indication message to the terminal devices 110A and 110B participating in the model reasoning, respectively. The inference indication message may indicate the split parts and transmission resource allocation corresponding to each terminal device. At 336, multiple participant devices perform the split model reasoning together based on the corresponding split parts and transmission resource allocation. After the model reasoning is completed, the allocated transmission resources may be released.

[0082] In process 330A, the base station 120 is responsible for splitting the AI / ML model and forming split information, and the base station 120 allocates transmission resources to the participant devices to assist in the execution of the split model reasoning. In some embodiments, other network endpoints (such as devices in the cloud 140, MEC 150, or IDC 160) may be responsible for splitting the AI / ML model and forming split information, and the base station 120 is still responsible for allocating transmission resources to the participant devices. In such an embodiment, the base station 120 is required to forward the received status information of each terminal device and the inference request of the terminal device 110A to the network endpoint. The network endpoint can form the split information similar to the base station 120 and forward it to the base station 120 so that the base station 120 can similarly allocate transmission resources. It should be understood that the subsequent operations can be similar to process 330A.

[0083] In the following process 330B, terminal device 110A, serving as the primary participant device for a specific AI / ML task, is responsible for generating split information, and base station 120 allocates transmission resources to participant devices to assist in the execution of split model inference. As shown in Figure 3E, at 341, terminal device 110A sends a model inference request to base station 120. For example, this request may include AI / ML task indication information or corresponding AI / ML model indication information. At 342, upon receiving the model inference request, base station 120 may determine the model split point between itself and the terminal device based on the AI / ML task indication information or the corresponding AI / ML model indication information, and transmit the result to terminal device 110A. In one example, this split point may be pre-configured for a specific AI / ML model. Once the model split point between base station 120 and the terminal device is determined, at 343, terminal device 110A may negotiate with other terminal devices to participate in the split model inference. For example, terminal device 110A may similarly send model inference requests to other terminal devices. When determining that it can participate in split model reasoning based on the AI / ML task indication information or the corresponding AI / ML model indication information, terminal devices 110B, 110N, etc. can report their own status information to terminal device 110A. Then, terminal device 110A can, for example, determine terminal device 110B as a participant device based on the status information and form split information for model reasoning. For example, forming the split information can include determining the model split point between terminal devices 110A and 110B participating in the reasoning. Forming the split information can also include forming a service flow for each participant device to perform model reasoning based on the split point. In one example, the service flow can be terminal device 110A->terminal device 110B->base station 120->terminal device 110B->terminal device 110A. At 345, terminal device 110A sends the terminal device-side split information to base station 120. At 346, upon receiving the split information, base station 120 allocates transmission resources to assist in the execution of the split model reasoning. For example, the base station 120 can allocate direct link resources between the terminal devices 110A and 110B based on the splitting information so as to transmit intermediate data and result data between the two terminal devices, and allocate uplink and downlink resources between the terminal device 110B and the base station 120 so as to transmit intermediate data and result data between the terminal device and the base station 120. At 347 and 348, the base station 120 transmits an inference indication message to the terminal devices 110A and 110B participating in the model reasoning, respectively. The inference indication message can indicate the split parts and transmission resource allocation corresponding to each terminal device. At 349, multiple participant devices perform the split model reasoning together based on the corresponding split parts and transmission resource allocation. After the model reasoning is completed, the allocated transmission resources can be released.

[0084] In the following process 330C, only terminal devices participate in model inference. Terminal device 110A, as the primary participant device for a specific AI / ML task, is responsible for generating split information and allocating transmission resources for other terminal devices. As shown in Figure 3F, at 351, terminal device 110A may negotiate with other terminal devices to participate in split model inference. For example, terminal device 110A may similarly send a model inference request to other terminal devices. Upon determining that they can participate in split model inference based on the AI / ML task indication information or the corresponding AI / ML model indication information, terminal devices 110B, 110N, etc. may report their status information to terminal device 110A. At 352, terminal device 110A may, for example, determine terminal devices 110B and 110N as participant devices based on the status information and generate split information for model inference. For example, generating the split information may include determining a model split point between the terminal devices participating in the inference. Generating the split information may also include generating a service flow for each participant device to perform model inference based on the split point. In one example, the service flow can be terminal device 110A->terminal device 110B->terminal device 110N->terminal device 110B->terminal device 110A, or terminal device 110A->terminal device 110B->terminal device 110N->terminal device 110A. Terminal device 110A also allocates transmission resources in an autonomous manner to assist in the execution of split model reasoning. For example, terminal device 110A can allocate direct link resources between terminal devices based on split information so as to transmit intermediate data and / or result data between the two terminal devices. At 353 and 354, terminal device 110A transmits an inference indication message to terminal devices 110B and 110N participating in model reasoning, respectively. The inference indication message can indicate the split parts and transmission resource allocation corresponding to each terminal device. At 355, multiple participant devices perform split model reasoning together based on the corresponding split parts and transmission resource allocation. After the model reasoning is completed, the allocated direct link resources can be released.

[0085] AI / ML model splitting example

[0086] Figure 4 illustrates example operations for splitting an AI / ML model according to an embodiment of the present disclosure. Example operations 400 may be performed, for example, by a primary participant device (e.g., terminal device 110A), a base station (e.g., 120), or other network endpoints (e.g., devices in cloud 140, MEC 150, or IDC 160).

[0087] As shown in Figure 4, the example operation 400 includes obtaining corresponding status information of multiple terminal devices (402). The multiple terminal devices include a main participant device (i.e., terminal device 110A) and one or more other terminal devices. For example, each terminal device can periodically send (e.g., broadcast) its own status information so that the device performing model splitting can receive the status information. It should be noted that the status information can indicate a state associated with model reasoning. For example, the status information can indicate the computing state of the corresponding terminal device, such as including at least one of the usage status of computing resources (such as CPU, GPU), the usage status of storage resources (such as RAM), or the power level. Additionally or alternatively, the status information can indicate the communication state of the corresponding terminal device, such as including channel status information reflecting at least one of the channel quality, number of transmission layers, or data rate of the uplink and downlink, direct link. The terminal device can perform channel estimation by receiving pilot information or reference information from a base station or other terminal devices.

[0088] In one embodiment, the primary participant device (e.g., terminal device 110A) can obtain status information of one or more other terminal devices. For example, terminal device 110A can obtain capability information of other terminal devices, indicating whether the corresponding terminal device supports model reasoning for participating in the split. Terminal device 110A can then receive corresponding status information only from terminal devices that support participation (e.g., including terminal device 110B). This can save power consumption of terminal device 110A associated with listening to status information broadcasts.

[0089] As an example, in the case where other terminal devices are required to participate in the split model reasoning, the terminal device 110A can learn whether other terminal devices support the model reasoning for participating in the split through the direct link UE capability transfer (Sidelink UE capability transfer) process. Taking the direct link UE capability transfer process with the terminal device 110B as an example, the terminal device 110A can send a UECapabilityEnquirySidelink message to the terminal device 110B to inquire about the capabilities of the terminal device 110B. In response to receiving the UECapabilityEnquirySidelink message, the terminal device 110B can reply to the terminal device 110A with a UECapabilityInformationSidelink message, which includes capability information indicating whether the terminal device 110B supports the model reasoning for participating in the split. Additionally, the UECapabilityInformationSidelink message may include the types of models supported by the terminal device 110B, such as the AlexNet model and the VGG-16 model for image recognition. Additionally or alternatively, the UECapabilityInformationSidelink message may include the real-time computing status of the terminal device 110B reflecting the current operating status and / or the overall computing capability reflecting the configuration status, so that the terminal device 110A can determine the specific manner of enabling the terminal device 110B to participate in the split model reasoning based on the configuration status of the terminal device 110B and / or more accurately based on the current operating status of the terminal device 110B.

[0090] Example operation 400 includes splitting at least a first portion of the AI / ML model to form split information (404) of at least the first portion of the AI / ML model. Taking AI / ML model 200A as an example, the at least first portion may include layers 201-203, layers 201-204, or layers 201-206. For example, in response to determining that the computing state of terminal device 110A is insufficient to support reasoning for at least the first portion of AI / ML model 200A, split operation 404 is performed. Of course, if the computing state of terminal device 110A indicates that the corresponding resources are sufficient, split operation 404 can also be performed so that terminal device 110A can have remaining computing resources for other tasks or operations.

[0091] In one embodiment, the splitting operation 404 may include determining one or more terminal devices whose computing status and / or communication status is better than a specific threshold. It should be noted that a computing status better than the threshold may indicate that the corresponding terminal device has computing resources, storage resources, and / or power to perform model reasoning; a communication status better than the threshold may indicate that the corresponding terminal device has suitable channel quality and / or data rate to transmit model reasoning information between participant devices. In one example, some or all of the terminal devices determined based on the threshold are determined as participant devices along with the main participant device (i.e., terminal device 110A).

[0092] In one embodiment, once the participant devices are identified, at least the first portion of the AI / ML model 200A can be split into multiple sub-parts based on the participant devices and their status information. For example, the number of sub-parts to be split can be determined based on the number of participant devices. Taking the splitting of layers 201-204 of the AI / ML model 200A as an example, given that the participant devices include terminal devices 110A and 110B, it can be determined that layers 201-204 need to be split into two sub-parts. For another example, the split points for forming multiple sub-parts can be set based on the computing status of the participant devices, so that the inference workload of the sub-parts is consistent with the computing status of the participant devices. This can help each participant device handle a model inference workload that matches its own computing resources, storage resources, and / or power consumption. It should be understood that the scope of the split parts needs to be reasonably determined for the master participant device to ensure that data corresponding to the AI / ML task, which may have privacy implications, is retained locally on the master participant device to prevent data leakage to downstream participant devices. Figure 5A shows another example of a split AI / ML model according to an embodiment of the present disclosure. For example, both AI / ML model 200B in FIG. 2B and AI / ML model 500A in FIG. 5A are obtained by splitting layers 201-204 of AI / ML model 200A (e.g., base station 120 needs to perform model reasoning for section III). In AI / ML model 200B, the split point is at layer 203. This requires terminal device 110A to perform model reasoning for layers 201-203, while terminal device 110B only performs model reasoning for layers 203-204. In AI / ML model 500A, the split point is at layer 202. Accordingly, terminal device 110A only performs model reasoning for layers 201-202, while terminal device 110B needs to perform model reasoning for layers 202-204. The splitting method for layers 201-204 can be determined based on the computing states of terminal devices 110A and 110B.

[0093] FIG5B shows another example of a split AI / ML model according to an embodiment of the present disclosure. In this example, the split AI / ML model 500B includes four parts I-IV. In this example, the base station 120 needs to perform model reasoning for part IV. In one embodiment, the terminal devices 110A, 110B, and 110N are determined as participant devices through the split operation 404. Accordingly, layers 201 to 204 form three sub-parts I-III through two split points (i.e., layers 202 and 203).

[0094] Figure 18 A shows an example of computing and communication resource evaluation based on layers for the AlexNet model. The AlexNet model is a CNN model for image recognition. As shown in Figure 18 A, the architecture of the AlexNet model includes an input layer (denoted as input in the figure), a convolutional layer (denoted as conv in the figure), a relu layer (denoted as relu in the figure), a cross-channel normalization layer (denoted as norm in the figure), a pooling layer (denoted as pool in the figure), a fully connected layer (denoted as fc in the figure), a dropout layer (denoted as drop in the figure), a softmax layer (denoted as softmax in the figure) and an argmax layer (denoted as argmax in the figure). Figure 18 B shows an example of computing and communication resource evaluation based on layers for the VGG-16 model. The VGG-16 model is another CNN model for image recognition. As shown in Figure 18 B, the architecture of the VGG-16 model is similar to that of the AlexNet model.

[0095] The split AlexNet model or VGG-16 model can be analyzed based on the computational and data characteristics of each layer in the model. As shown in Figures 18A and 18B, the size of the intermediate data transmitted from one layer to the next depends on the location of the split point. Therefore, for a specific image frame rate, the data rate required for one participant device to transmit intermediate data to a downstream participant device is related to the split point of the model. For example, assuming that images (with a resolution of 227×227) in a 30-frame-per-second video stream need to be classified, for the AlexNet model, the data rates corresponding to different split points range from 4.8Mbit / s to 65Mbit / s, and for the VGG-16 model, the data rates corresponding to different split points range from 24Mbit / s to 720Mbit / s.

[0096] Taking the AlexNet model as an example, in an embodiment, a communication state threshold of 4.8 Mbit / s can be set for the data rate. Based on the specific scenario of the split model reasoning, this data rate can be set for at least one of the uplink or the through link. Accordingly, through the split operation 404, multiple terminal devices with data rates higher than 4.8 Mbit / s can be determined. Some or all of these multiple terminal devices can be identified as participant devices along with the main participant device (i.e., terminal device 110A).

[0097] Once the participant devices are determined, the split points for multiple sub-parts can be determined based on the data rates of the direct links and uplinks of the participant devices, so that the data rates required for transmission to downstream devices corresponding to the split points are compatible with the data rates of the participant devices. For example, terminal device 110B, which has the highest direct link data rate (e.g., 42 Mbits) with terminal device 110A, can be determined as the downstream participant device of terminal device 110A, and alternative split point 2 can be determined as the split point. Terminal device 110N, which has an uplink data rate greater than 4.8 Mbit / s, can be determined as the downstream participant device of terminal device 110B, and alternative split point 3 can be determined as the split point. This can facilitate the transmission of intermediate data between multiple participant devices with a smaller delay, thereby completing the entire model inference process within a time period acceptable to the user.

[0098] In some embodiments, the split information of the AI / ML model may include 1) indication information of the multiple sub-parts split, and 2) information of the participant device that performs model reasoning. Figure 6A shows a first example of split information according to an embodiment of the present disclosure. In Figure 6A, the split information corresponding to the split AI / ML models 200B, 500A and 500B are shown in sequence. In this example, the indication information is expressed by the specific layer number of the split sub-part. Taking the split information of the AI / ML model 200B as an example, the "Indication Information" column indicates that sub-part I includes layers 1 to 3 of the complete AI / ML model 200A, and sub-part II includes layers 3 to 4. The "Executor" column indicates that the model reasoning of sub-part I is performed by participant 1, and the model reasoning of sub-part II is performed by participant 2.

[0099] Figure 6B shows a second example of split information according to an embodiment of the present disclosure. In Figure 6B, the split information corresponding to the split AI / ML models 200B, 500A and 500B are shown in sequence. In this example, the indication information is expressed by the split points that form the sub-parts. Taking the split information of the AI / ML model 200B as an example, the "Indication Information" column indicates that sub-part I is formed by a single split point located at the 3rd layer of the complete AI / ML model 200A (that is, sub-part I is the first sub-part), and sub-part II is formed by the two split points located at the 3rd and 4th layers. The "Executor" column indicates that the model reasoning of sub-part I is performed by participant 1, and the model reasoning of sub-part II is performed by participant 2.

[0100] Figure 6C shows a third example of split information according to an embodiment of the present disclosure. In Figure 6C, the split information corresponding to the split AI / ML models 200B, 500A and 500B are also shown in sequence. In this example, the indication information is expressed through the model configuration of the sub-part. Still taking the split information of the AI / ML model 200B as an example, the "Indication Information" column includes the specific model configurations of sub-part I and sub-part II, including the number and weights of neurons in each layer and the connection relationship between neurons in the layers. Similarly, the "Executor" column indicates that the model reasoning of sub-part I is performed by participant 1, and the model reasoning of sub-part II is performed by participant 2.

[0101] It should be understood that in some embodiments, the index of the model portion targeted for model inference (e.g., layer number, split point) can be notified to the participant device via sub-portion indication information (e.g., as shown in Figures 6A and 6B), allowing the participant device to determine the model configuration of the model portion based on the index of the model portion and the overall model configuration (e.g., complete model 200A). Since it is often necessary to perform the same or similar AI / ML tasks, each participant device can have the same model configuration of the AI / ML model (e.g., model 200A) locally, and the model configurations of multiple participant devices can be updated synchronously. As previously described, the local AI / ML model can be written to the participant device or semi-statically configured to the participant device. Therefore, the participant device can determine the specific model configuration of the corresponding sub-portion based on the sub-portion index and the overall model configuration. Alternatively, in some embodiments, the specific model configuration of the model portion targeted for model inference can be notified to the participant device via sub-portion indication information (e.g., as shown in Figure 6C). Once the specific model configuration of the corresponding sub-portion is determined, the participant device can input input values ​​or intermediate data into the model portion to obtain the corresponding output data.

[0102] It should be understood that, through the participant device information (e.g., the "Executor" column), the split information specifies the order in which the participant devices execute model reasoning for the corresponding sub-parts one by one, thereby being able to represent the service flow for model reasoning. For example, the three split information in Figure 6A represent the service flows of "Participant 1->Participant 2", "Participant 1->Participant 2", and "Participant 1->Participant 2->Participant 3", respectively. In some embodiments, the participant device information can be used to notify a specific participant device of at least the information of the downstream participant device, so that the participant device knows how to transmit the intermediate data it generates.

[0103] Figure 7A illustrates a first example operation for distributing split information according to an embodiment of the present disclosure. Example operation 700A is described below with reference to AI / ML model 500B. In Figure 7A, terminal device 110A corresponds to participant 1 and is the primary participant device, while terminal devices 110B and 110N correspond to participants 2 and 3, respectively, and are secondary participant devices. In this example, the primary participant device forms and distributes the split information (e.g., as shown in Figures 6A and 6B).

[0104] As shown in FIG7A , example operation 700A includes the terminal device 110A notifying the corresponding participants of the split information of the AI / ML model based on the information in the "Executor" column of the split information. Specifically, at 712 , the terminal device 110A notifies the terminal device 110B, which is the participant 2, of the split information for participant 2; at 714 , the terminal device 110A notifies the terminal device 110N, which is the participant 3, of the split information for participant 3.

[0105] In some embodiments, after forming the split information (e.g., in FIG6A and FIG6B ), the primary participant device may provide the split information to the base station 120. In this way, the base station 120 may distribute the split information to the corresponding participant devices through operations similar to 700A.

[0106] FIG7B illustrates a second example operation for distributing split information according to an embodiment of the present disclosure. Still referring to the split AI / ML model 500B, example operation 700B is described, where terminal device 110A is the primary participant device, and terminal devices 110B and 110N are secondary participant devices. In this example, the split information (e.g., in FIG6A and FIG6B ) is distributed by a control device 701, which may be generated by the control device 701 or received from another device. The control device 701 may be a base station 120 or other network endpoint.

[0107] As shown in FIG7B , example operation 700B includes the control device 701 notifying the corresponding participant devices of the sub-parts of the AI / ML model based on the performer information in the split information. Specifically, at 722 , the control device notifies the terminal device 110A, which is participant 1, of the split information for participant 1; at 724 , the control device notifies the terminal device 110B, which is participant 2, of the split information for participant 2; and at 726 , the control device notifies the terminal device 110N, which is participant 3, of the split information for participant 3.

[0108] In operations 700A and 700B, the split information for a specific participant may include indication information of the corresponding sub-part, so that the participant can determine the specific model configuration of the sub-part. In one embodiment, the indication information may at least indicate the index information of the corresponding sub-part. Taking operation 700A as an example, at 712, terminal device 110A may notify terminal device 110B, which is participant 2, of the index information of sub-part II; at 714, terminal device 110A notifies terminal device 110N, which is participant 3, of the index information of sub-part III. In this way, terminal devices 110B and 110N can determine the specific model configuration of the corresponding sub-part based on the overall model configuration (i.e., 200A) and the index information.

[0109] In one embodiment, the instruction information may indicate the specific model configuration of the corresponding sub-part. Still taking operation 700 as an example, at 712, terminal device 110A may notify terminal device 110B, which is participant 2, of the specific model configuration of sub-part II. At 714, terminal device 110A may notify terminal device 110N, which is participant 3, of the specific model configuration of sub-part III.

[0110] Split model inference example

[0111] Figures 8A through 8D illustrate example operations for performing split model inference according to an embodiment of the present disclosure. The split model inference operations are described below in conjunction with split AI / ML model 200B. In this example operation, terminal device 110A is the primary participant device in model inference, e.g., the corresponding AI / ML task is initiated by terminal device 110A. Other devices are auxiliary participant devices in model inference.

[0112] As shown in FIG8A , operation 800A includes, at 812, the terminal device 110A performing model inference for part I. For part I, when the input value 220 is input, the neurons in the input layer 201 receive the corresponding value and propagate the value to the neurons in the intermediate layer 202. The neurons in the intermediate layer 202 calculate the weighted sum of the output values ​​of the neurons in the input layer 201 and propagate the value to the neurons in the intermediate layer 203. The neurons in the intermediate layer 203 calculate the weighted sum of the output values ​​of the neurons in the intermediate layer 202, and the weighted sum forms the intermediate data 221. At 822, the terminal device 110A transmits the intermediate data 221 to the downstream participant device, namely, the terminal device 110B.

[0113] At 814, terminal device 110B performs model inference for Part II. Specifically, upon receiving intermediate data 221, terminal device 110B provides intermediate data 221 to corresponding neurons in intermediate layer 204 via neurons in intermediate layer 203. The neurons in intermediate layer 204 calculate the weighted sum of the output values ​​of the neurons in intermediate layer 203, and this weighted sum forms intermediate data 222. At 824, terminal device 110B transmits intermediate data 222 to downstream participant devices.

[0114] In one embodiment, the downstream participant device is a network endpoint 801 (e.g., a base station 120, a cloud server, an MEC server, or a device in an IDC). That is, for AI / ML tasks on a terminal device, the terminal device and the network endpoint jointly perform model reasoning to share the computing load of the terminal device through the relatively sufficient computing resources on the network endpoint side. In one embodiment, the downstream participant device is another terminal device 110N. That is, for AI / ML tasks on a single terminal device, multiple terminal devices jointly perform model reasoning to share the computing load of the single terminal device through only the computing resources of multiple terminal devices.

[0115] At 816, the network endpoint 801 or terminal device 110N performs model inference for the final section III. Specifically, upon receiving the intermediate data 222, the network endpoint 801 or terminal device 110N provides the intermediate data 222 to the corresponding neurons in the intermediate layer 205 via the neurons in the intermediate layer 204. The neurons in the intermediate layer 205 calculate the weighted sum of the output values ​​of the neurons in the intermediate layer 204 and propagate it to the neurons in the output layer 206. The neurons in the output layer 206 calculate the weighted sum of the output values ​​of the neurons in the intermediate layer 205, and this weighted sum forms the inference result 240 for the input value 220.

[0116] At 826 , the network endpoint 801 or the terminal device 110N transmits the inference result 240 to the terminal device 110A. At this point, the terminal device 110A obtains the inference result for its AI / ML task.

[0117] In the example of Figure 8B, the terminal device 110A is the main participant device for model reasoning, and the terminal device 110B and the network endpoint 801 are auxiliary participant devices. The same operations in Figure 8B as those in Figure 8A are shown with the same reference numerals, and these operations can be understood in conjunction with the description of Figure 8A. Only the differences between operation 800B and operation 800A are described here. Specifically, after completing the model reasoning for Part II, at 844, the intermediate data 222 is transmitted to the terminal device 110A by the terminal device 110B, and at 844', it is forwarded to the network endpoint 801 by the terminal device 110A. In this example, the uplink and downlink communications with the network endpoint 801 are performed by the main participant device. This is advantageous in a case where other terminal devices (such as 110B) do not have good uplink communication.

[0118] In the example of FIG8C , terminal device 110A is the primary participant device for model reasoning, and terminal device 110B and network endpoint 801 are secondary participant devices. Operations in FIG8C that are identical to those in FIG8A are denoted by the same reference numerals, and these operations can be understood in conjunction with the description of FIG8A . Only the differences between operation 800C and operation 800A are described here. Specifically, after completing model reasoning for Section III, at 866 , result data 240 is transmitted by network endpoint 801 to terminal device 110B, and at 866 ′, is forwarded by terminal device 110B to terminal device 110A. In this example, uplink and downlink communications with network endpoint 801 are performed by terminal device 110B, acting as a secondary participant device. This is advantageous when terminal device 110A, acting as the primary participant device, does not have good uplink and downlink communications.

[0119] In the example of FIG8D , terminal device 110A is the primary participant device for model reasoning, network endpoint 801 is the secondary participant device, and terminal device 110B acts as a relay device between terminal device 110A and network endpoint 801. Specifically, at 872 , terminal device 110A performs model reasoning for Part I. At 882 , intermediate data 222 is transmitted from terminal device 110A to terminal device 110B, and at 882 ′, terminal device 110B forwards it to network endpoint 801. At 816 , network endpoint 801 performs model reasoning for Part III. At 886 , result data 240 is transmitted from network endpoint 801 to terminal device 110B, and at 886 ′, terminal device 110B forwards it to terminal device 110A. In this example, terminal device 110B, which does not participate in model reasoning, acts as a relay device between the primary and secondary participant devices. This is advantageous when terminal device 110A, the primary participant device, lacks good uplink and downlink communication.

[0120] 8A to 8D only illustrate split model inference operations performed by three participant devices. It should be understood that in the presence of more participant devices, split model inference can be performed in a manner similar to operations 800A to 800D.

[0121] Resource Allocation Example

[0122] Figure 9A illustrates an example signaling flow for allocating transmission resources to participant devices for model inference according to an embodiment of the present disclosure. Signaling flow 900A is described with reference to a context similar to Figure 7B , i.e., for a split AI / ML model 500B, with terminal device 110A serving as the primary participant device and terminal devices 110B and 110N serving as secondary participant devices.

[0123] As shown in Figure 9A, signaling process 900A includes, at 902, terminal device 110A sending a resource allocation request to base station 120. In one embodiment, the resource allocation request may include at least the split information of model 500B. For example, terminal device 110A may indicate that it is requesting transmission resources for each terminal device by providing at least the split information to base station 120, thereby assisting in the execution of model reasoning. Once the transmission resources for each terminal device are determined based on the split information, base station 120 notifies each terminal device of resource allocation information at 904-906, indicating the resource allocation for at least one of the direct link, uplink, and downlink. In one embodiment, the resource allocation information may be sent to each terminal device together with the split information for each participant, such as in Figure 7B. In one embodiment, the resource allocation request may correspond to the inference request at 331 or the split information at 345, or may be sent together with them.

[0124] Alternatively or additionally, the base station 120 may allocate transmission resources to the terminal devices 110A-110N in response to receiving split information of the model 500B from other network endpoints or determining it by itself.

[0125] 9B illustrates example operations for allocating transmission resources to participant devices for model inference according to an embodiment of the present disclosure. Example operations 900B may be performed by base station 120.

[0126] As shown in FIG9B , example operations 900B include, at 912, base station 120 determining split information for sub-parts of the AI / ML model (e.g., as shown in FIG6A and FIG6B ). In one embodiment, the split information is formed by base station 120. In one embodiment, the split information is formed by a terminal device (e.g., 110A) or other network endpoint (e.g., a device in cloud 140, MEC 150, or IDC 160). For example, after forming the split information, terminal device 110A or other network endpoint sends a resource allocation request to base station 120, and the resource allocation request may include the split information.

[0127] At 914, the base station 120 allocates transmission resources for the direct link and / or uplink and downlink to the participant devices based on the split information. Specifically, the base station 120 can determine the transmission requirements for the intermediate data and result data based on the executor information in the split information. Taking the three participants in Figure 8A as an example, based on the service flow consisting of terminal device 110A->terminal device 110B->terminal device 110N, the transmission requirements for the intermediate data and result data that can be determined are shown in Table 1. Based on the service flow consisting of terminal device 110A->terminal device 110B->base station 120, the transmission requirements for the intermediate data and result data that can be determined are shown in Table 2.

[0128] In the example of Table 2, the transmission of intermediate data and result data involves the uplink and downlink between a specific terminal device and the base station 120. In the case where there are multiple terminal devices participating in model reasoning, corresponding resources can be allocated to terminal devices (rather than specific terminal devices) with better uplink and downlink communication quality to transmit corresponding model reasoning information, and the model reasoning information can be further transmitted between intermediate devices through a direct link. For example, in Table 3, the intermediate data generated by the terminal device 110B can alternatively be transmitted via the uplink between the terminal device 110A and the base station 120. In this way, it is possible to avoid the inability to transmit intermediate data or result data of model reasoning due to poor communication quality between a single terminal device and the base station 120, which ultimately makes it impossible to complete the model reasoning.

[0129] In one embodiment, transmission resources may be allocated to corresponding participant devices based on the expected output data volume (e.g., the amount of data within a time period) of the model inference of the corresponding sub-part of the AI / ML model.

[0130] Once resource allocation is complete, base station 120 may send resource allocation information to the corresponding participant devices. For example, the resource allocation information may indicate resource allocation for at least one of the direct link, uplink, and downlink. Accordingly, the operations of transmitting model inference information at 822, 824, and 826 in FIG8A may be based on the resource allocation for the direct link, uplink, and / or downlink allocated by base station 120.

[0131] Table 1

[0132] Table 2

[0133] Table 3

[0134] Example Method

[0135] Figure 10 shows an example method for resource allocation in model reasoning according to an embodiment of the present disclosure. The method can be performed by, for example, electronic devices 300, 310, or 320. As shown in Figure 10, method 1000 may include forming split information of at least a first part of the AI / ML model based on corresponding state information of a first terminal device and one or more other terminal devices (box 1002). The split information specifies that split model reasoning is to be performed by multiple participant devices for multiple sub-parts of the at least first part of the AI / ML model. For example, the AI / ML model corresponds to an AI / ML task. Additionally, method 1000 may include determining an AI / ML model corresponding to the AI / ML task of the first terminal device. As shown in Figure 10, method 1000 may also include, based on the split information, causing the wireless network to allocate resources for transmitting model reasoning information to at least one of the multiple participant devices (box 1004). Further details of the method can be understood with reference to the above description of each electronic device, terminal device, or network endpoint.

[0136] In one embodiment, method 1000 also includes: obtaining corresponding status information of the first terminal device and one or more other terminal devices, wherein the status information is associated with model reasoning, and wherein the status information indicates the computing status and / or communication status of the corresponding terminal device, the computing status includes at least one of computing resource usage status, storage resource usage status or power status, and the communication status includes at least one of channel quality or data rate.

[0137] In one embodiment, the splitting information includes indication information of the multiple sub-parts and information of participant devices that perform model inference, and forming the splitting information includes: determining the terminal device and the first terminal device whose corresponding status is better than a threshold among the other one or more terminal devices as the multiple participant devices; based on the corresponding status information of the multiple participant devices, splitting the at least the first part of the AI / ML model into the multiple sub-parts.

[0138] In one embodiment, the multiple sub-parts correspond to the multiple participant devices, the model reasoning workload of the multiple sub-parts matches the computing state of the corresponding participant devices, and the communication state of the multiple participant devices can support the transmission of model reasoning information.

[0139] In one embodiment, the AI / ML model comprises a neural network model, and the at least first portion of the AI / ML model comprises one or more first layers of the AI / ML model, or comprises all layers of the AI / ML model; and / or

[0140] The model inference information includes model inference intermediate data and / or model inference result data.

[0141] In one embodiment, causing the wireless network to allocate resources for transmitting model inference information includes sending the split information of the at least first portion of the AI / ML model to a base station.

[0142] In one embodiment, method 1000 further includes: sending instructions for executing model inference to corresponding participant devices based on the split information, the instructions including indication information of corresponding sub-parts of the AI / ML model and indication information of downstream devices.

[0143] In one embodiment, the electronic device is implemented as a network endpoint or a part of a network endpoint, and the network endpoint includes a cloud server and / or an edge server.

[0144] In one embodiment, the electronic device is implemented as a first terminal device or a part of the first terminal device, and method 1000 also includes: receiving resource allocation information for the first terminal device from the base station, wherein the resource allocation information indicates resource allocation for at least one of a direct link, an uplink, and a downlink.

[0145] In one embodiment, method 1000 further includes: inputting local data into a first sub-portion of the AI / ML model to obtain first intermediate data; and providing the first intermediate data to the first participant device via a direct link with the first participant device based on resource allocation for the direct link.

[0146] In one embodiment, method 1000 further includes: based on resource allocation for the through link, receiving second intermediate data output by the second participant device via the through link with the second participant device; based on resource allocation for the uplink, sending the second intermediate data to the network via the uplink; or based on resource allocation for the downlink, receiving an inference result corresponding to the AI / ML model from the network via the downlink.

[0147] Figure 11 shows an example method for resource allocation in model reasoning according to an embodiment of the present disclosure. The method can be performed by an electronic device 320. As shown in Figure 11, method 1100 may include obtaining split information of at least a first part of the AI / ML model (box 1102). The AI / ML model corresponds to the AI / ML task of the first terminal device, and the split information specifies that split model reasoning is to be performed by multiple participant devices for multiple sub-parts of the at least first part of the AI / ML model. As shown in Figure 10, method 1100 may also include allocating resources for transmitting model reasoning information to at least one of the multiple participant devices based on the split information (box 1104). Further details of the method can be understood with reference to the above description of the electronic device 320 or the base station.

[0148] In one embodiment, the method 1100 further includes: forming the splitting information based on corresponding status information of the first terminal device and one or more other terminal devices; or receiving the splitting information from the first terminal device or a network endpoint.

[0149] In one embodiment, the splitting information includes indication information of the multiple sub-parts and information of participant devices that perform model inference, and forming the splitting information includes: determining the terminal device and the first terminal device whose corresponding status is better than a threshold among the other one or more terminal devices as the multiple participant devices; based on the corresponding status information of the multiple participant devices, splitting the at least the first part of the AI / ML model into the multiple sub-parts.

[0150] In one embodiment, the multiple sub-parts correspond to the multiple participant devices, the model reasoning workload of the multiple sub-parts matches the computing state of the corresponding participant devices, and the communication state of the multiple participant devices can support the transmission of model reasoning information.

[0151] In one embodiment, method 1100 further includes: sending instructions for executing model inference to corresponding participant devices based on the split information, the instructions including indication information of corresponding sub-parts of the AI / ML model and indication information of downstream devices.

[0152] In one embodiment, method 1100 further includes: allocating resources to the corresponding participant device based on the expected output data volume of the model inference of the corresponding sub-part of the AI / ML model; and sending resource allocation information to the corresponding participant device, wherein the resource allocation information indicates resource allocation for at least one of a direct link, an uplink, and a downlink.

[0153] Figure 12 shows an example method for model reasoning according to an embodiment of the present disclosure. The method can be performed by an electronic device 310. As shown in Figure 12, method 1200 may include receiving an instruction to perform model reasoning from a first terminal device (box 1202), the instruction including indication information of a corresponding sub-portion of the AI / ML model and indication information of a downstream participant device. Method 1200 may also include receiving a resource allocation for transmitting model reasoning information, wherein the resource allocation indicates resources for a wireless link (e.g., including a direct link) with an upstream participant device and a downstream participant device (box 1204). Method 1200 may also include receiving first intermediate data from an upstream participant device via a wireless link with the upstream participant device based on the instruction and resource allocation (box 1206). Method 1200 may also include inputting the first intermediate data into a corresponding sub-portion of the AI / ML model, obtaining second intermediate data, and sending the second intermediate data to the downstream participant device via a wireless link with the downstream participant device based on the instruction and resource allocation (box 1208). Further details of the method can be understood with reference to the above description of the electronic device 310 or terminal device.

[0154] The above describes various exemplary electronic devices and methods according to the embodiments of the present disclosure. It should be understood that the operations or functions of these electronic devices can be combined with each other to implement more or fewer operations or functions than described. The operational steps of each method can also be combined with each other in any appropriate order to similarly implement more or fewer operations than described.

[0155] It should be understood that the machine-executable instructions in the machine-readable storage medium or program product according to the embodiments of the present disclosure can be configured to perform operations corresponding to the above-mentioned device and method embodiments. When referring to the above-mentioned device and method embodiments, the embodiments of the machine-readable storage medium or program product are clear to those skilled in the art and are therefore not described again. Machine-readable storage media and program products for carrying or including the above-mentioned machine-executable instructions also fall within the scope of the present disclosure. Such storage media may include, but are not limited to, floppy disks, optical disks, magneto-optical disks, memory cards, memory sticks, and the like. In addition, it should be understood that the above-mentioned series of processes and devices may also be implemented by software and / or firmware.

[0156] In addition, it should be understood that the above series of processes and devices can also be implemented through software and / or firmware. In the case of implementation through software and / or firmware, the program constituting the software is installed from a storage medium or network to a computer with a dedicated hardware structure, such as the general-purpose computer 1300 shown in Figure 13. When various programs are installed, the computer can perform various functions, etc. Figure 13 shows an example block diagram of a computer that can be implemented as a terminal device or network endpoint according to an embodiment of the present disclosure.

[0157] 13 , a central processing unit (CPU) 1301 executes various processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage section 1308 to a random access memory (RAM) 1303. In the RAM 1303, data required when the CPU 1301 executes various processes and the like is also stored as needed.

[0158] The CPU 1301, the ROM 1302, and the RAM 1303 are connected to one another via a bus 1304. An input / output interface 1305 is also connected to the bus 1304.

[0159] The following components are connected to the input / output interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet.

[0160] A drive 1310 is also connected to the input / output interface 1305 as needed. A removable medium 1311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1310 as needed so that a computer program read therefrom is installed in the storage section 1308 as needed.

[0161] In the case of realizing the above-described series of processing by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 1311 .

[0162] Those skilled in the art will appreciate that such storage media are not limited to the removable medium 1311 shown in FIG. 13 , which stores programs therein and is distributed separately from the device to provide the programs to users. Examples of the removable medium 1311 include magnetic disks (including floppy disks (registered trademark)), optical disks (including compact disk read-only memories (CD-ROMs) and digital versatile disks (DVDs)), magneto-optical disks (including minidiscs (MDs) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be a ROM 1302, a hard disk included in the storage section 1308, or the like, in which the programs are stored and distributed to users together with the device containing them.

[0163] Application examples according to the present disclosure will be described below with reference to FIG. 14 to FIG. 17 .

[0164] Application examples for base stations

[0165] First application example

[0166] FIG14 is a block diagram illustrating a first example of a schematic configuration of a gNB to which the techniques of this disclosure may be applied. gNB 1400 includes multiple antennas 1410 and a base station device 1420. Base station device 1420 and each antenna 1410 may be connected to each other via an RF cable. In one implementation, gNB 1400 (or base station device 1420) may correspond to electronic device 300A described above.

[0167] Each antenna 1410 includes a single or multiple antenna elements (such as multiple antenna elements included in a multiple-input multiple-output (MIMO) antenna) and is used for base station device 1420 to transmit and receive wireless signals. As shown in Figure 14, gNB 1400 may include multiple antennas 1410. For example, multiple antennas 1410 may be compatible with multiple frequency bands used by gNB 1400.

[0168] The base station device 1420 includes a controller 1421 , a memory 1422 , a network interface 1423 , and a wireless communication interface 1425 .

[0169] The controller 1421 may be, for example, a CPU or DSP, and operates various higher-layer functions of the base station device 1420. For example, the controller 1421 generates data packets based on the data in the signal processed by the wireless communication interface 1425 and transmits the generated packets via the network interface 1423. The controller 1421 may bundle data from multiple baseband processors to generate bundled packets and transmit the generated bundled packets. The controller 1421 may have logic functions for performing control such as radio resource control, radio bearer control, mobility management, admission control, and scheduling. This control may be performed in conjunction with a nearby gNB or core network node. The memory 1422 includes RAM and ROM and stores programs executed by the controller 1421 and various types of control data (such as terminal lists, transmission power data, and scheduling data).

[0170] The network interface 1423 is a communication interface for connecting the base station device 1420 to the core network 1424. The controller 1421 can communicate with the core network node or another gNB via the network interface 1423. In this case, the gNB 1400 and the core network node or other gNB can be connected to each other via a logical interface (such as an S1 interface and an X2 interface). The network interface 1423 can also be a wired communication interface or a wireless communication interface for wireless backhaul. If the network interface 1423 is a wireless communication interface, the network interface 1423 can use a higher frequency band for wireless communication than the frequency band used by the wireless communication interface 1425.

[0171] The wireless communication interface 1425 supports any cellular communication scheme, such as Long Term Evolution (LTE) and LTE-Advanced, and provides wireless connectivity to terminals located in the cell of the gNB 1400 via the antenna 1410. The wireless communication interface 1425 may typically include, for example, a baseband (BB) processor 1426 and RF circuitry 1427. The BB processor 1426 can perform various signal processing functions, such as encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and performs various types of signal processing for layers such as Layer 1 (L1), Medium Access Control (MAC), Radio Link Control (RLC), and Packet Data Convergence Protocol (PDCP). In place of the controller 1421, the BB processor 1426 may perform some or all of the aforementioned logical functions. The BB processor 1426 may be a memory storing communication control programs, or a module including a processor configured to execute programs and associated circuitry. Program updates can modify the functionality of the BB processor 1426. This module may be a card or blade inserted into a slot in the base station device 1420. Alternatively, it may be a chip mounted on the card or blade. Meanwhile, the RF circuit 1427 may include, for example, a mixer, a filter, and an amplifier, and transmits and receives wireless signals via the antenna 1410. Although FIG14 shows an example in which one RF circuit 1427 is connected to one antenna 1410, the present disclosure is not limited to this illustration, and one RF circuit 1427 may be connected to multiple antennas 1410 at the same time.

[0172] As shown in Figure 14 , the wireless communication interface 1425 may include multiple BB processors 1426. For example, multiple BB processors 1426 may be compatible with multiple frequency bands used by gNB 1400. As shown in Figure 14 , the wireless communication interface 1425 may include multiple RF circuits 1427. For example, multiple RF circuits 1427 may be compatible with multiple antenna elements. While Figure 14 illustrates an example in which the wireless communication interface 1425 includes multiple BB processors 1426 and multiple RF circuits 1427, the wireless communication interface 1425 may also include a single BB processor 1426 or a single RF circuit 1427.

[0173] Second application example

[0174] FIG15 is a block diagram illustrating a second example of a schematic configuration of a gNB to which the techniques of this disclosure can be applied. A gNB 1530 includes multiple antennas 1540, a base station device 1550, and an RRH 1560. The RRH 1560 and each antenna 1540 can be connected to each other via an RF cable. The base station device 1550 and the RRH 1560 can be connected to each other via a high-speed line such as an optical fiber cable. In one implementation, the gNB 1530 (or base station device 1550) herein may correspond to the electronic device 300A described above.

[0175] Each antenna 1540 includes a single or multiple antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for RRH 1560 to transmit and receive wireless signals. As shown in Figure 15, gNB 1530 may include multiple antennas 1540. For example, multiple antennas 1540 may be compatible with multiple frequency bands used by gNB 1530.

[0176] Base station device 1550 includes a controller 1551, a memory 1552, a network interface 1553, a wireless communication interface 1555, and a connection interface 1557. Controller 1551, memory 1552, and network interface 1553 are the same as controller 1421, memory 1422, and network interface 1423 described with reference to FIG.

[0177] The wireless communication interface 1555 supports any cellular communication scheme (such as LTE and LTE-Advanced) and provides wireless communication to terminals located in the sector corresponding to the RRH 1560 via the RRH 1560 and the antenna 1540. The wireless communication interface 1555 may generally include, for example, a BB processor 1556. The BB processor 1556 is identical to the BB processor 1426 described with reference to FIG. 14 , except that the BB processor 1556 is connected to the RF circuit 1564 of the RRH 1560 via the connection interface 1557. As shown in FIG. 15 , the wireless communication interface 1555 may include multiple BB processors 1556. For example, multiple BB processors 1556 may be compatible with multiple frequency bands used by the gNB 1530. Although FIG. 15 illustrates an example in which the wireless communication interface 1555 includes multiple BB processors 1556, the wireless communication interface 1555 may also include a single BB processor 1556.

[0178] The connection interface 1557 is an interface for connecting the base station device 1550 (wireless communication interface 1555) to the RRH 1560. The connection interface 1557 may also be a communication module for connecting the base station device 1550 (wireless communication interface 1555) to the RRH 1560 for communication in the high-speed line.

[0179] The RRH 1560 includes a connection interface 1561 and a wireless communication interface 1563 .

[0180] The connection interface 1561 is an interface for connecting the RRH 1560 (wireless communication interface 1563) to the base station device 1550. The connection interface 1561 may also be a communication module for communication in the above-mentioned high-speed line.

[0181] The wireless communication interface 1563 transmits and receives wireless signals via the antenna 1540. The wireless communication interface 1563 may generally include, for example, an RF circuit 1564. The RF circuit 1564 may include, for example, a mixer, a filter, and an amplifier, and transmits and receives wireless signals via the antenna 1540. Although FIG. 15 shows an example in which one RF circuit 1564 is connected to one antenna 1540, the present disclosure is not limited to this illustration, and one RF circuit 1564 may be connected to multiple antennas 1540 simultaneously.

[0182] As shown in FIG15 , the wireless communication interface 1563 may include multiple RF circuits 1564. For example, multiple RF circuits 1564 may support multiple antenna elements. Although FIG15 shows an example in which the wireless communication interface 1563 includes multiple RF circuits 1564, the wireless communication interface 1563 may also include a single RF circuit 1564.

[0183] Application examples for terminal devices

[0184] First application example

[0185] 16 is a block diagram illustrating an example of a schematic configuration of a smartphone 1600 to which the techniques of the present disclosure may be applied. The smartphone 1600 includes a processor 1601, a memory 1602, a storage device 1603, an external connection interface 1604, a camera 1606, a sensor 1607, a microphone 1608, an input device 1609, a display 1610, a speaker 1611, a wireless communication interface 1612, one or more antenna switches 1615, one or more antennas 1616, a bus 1617, a battery 1618, and an auxiliary controller 1619. In one implementation, the smartphone 1600 (or processor 1601) herein may correspond to the electronic device 300B described above.

[0186] The processor 1601 may be, for example, a CPU or a system on a chip (SoC), and controls the functions of the application layer and other layers of the smartphone 1600. The memory 1602 includes RAM and ROM, and stores data and programs executed by the processor 1601. The storage device 1603 may include storage media such as semiconductor memories and hard disks. The external connection interface 1604 is an interface for connecting external devices (such as memory cards and universal serial bus (USB) devices) to the smartphone 1600.

[0187] The camera 1606 includes an image sensor (such as a charge coupled device (CCD) and a complementary metal oxide semiconductor (CMOS)) and generates a captured image. The sensor 1607 may include a group of sensors such as a measurement sensor, a gyroscope sensor, a geomagnetic sensor, and an acceleration sensor. The microphone 1608 converts the sound input to the smartphone 1600 into an audio signal. The input device 1609 includes, for example, a touch sensor, a keypad, a keyboard, a button, or a switch configured to detect a touch on the screen of the display device 1610, and receives an operation or information input from the user. The display device 1610 includes a screen (such as a liquid crystal display (LCD) and an organic light emitting diode (OLED) display) and displays the output image of the smartphone 1600. The speaker 1611 converts the audio signal output from the smartphone 1600 into sound.

[0188] The wireless communication interface 1612 supports any cellular communication scheme (such as LTE and LTE-Advanced) and performs wireless communication. The wireless communication interface 1612 may generally include, for example, a BB processor 1613 and an RF circuit 1614. The BB processor 1613 may perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and perform various types of signal processing for wireless communication. Meanwhile, the RF circuit 1614 may include, for example, a mixer, a filter, and an amplifier, and transmit and receive wireless signals via an antenna 1616. The wireless communication interface 1612 may be a chip module on which the BB processor 1613 and the RF circuit 1614 are integrated. As shown in FIG16 , the wireless communication interface 1612 may include multiple BB processors 1613 and multiple RF circuits 1614. Although FIG16 shows an example in which the wireless communication interface 1612 includes multiple BB processors 1613 and multiple RF circuits 1614, the wireless communication interface 1612 may also include a single BB processor 1613 or a single RF circuit 1614.

[0189] In addition, in addition to the cellular communication scheme, the wireless communication interface 1612 can support other types of wireless communication schemes, such as a short-range wireless communication scheme, a near field communication scheme, and a wireless local area network (LAN) scheme. In this case, the wireless communication interface 1612 may include a BB processor 1613 and an RF circuit 1614 for each wireless communication scheme.

[0190] Each of the antenna switches 1615 switches the connection destination of the antenna 1616 between a plurality of circuits (eg, circuits for different wireless communication schemes) included in the wireless communication interface 1612 .

[0191] Each of the antennas 1616 includes a single or multiple antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals via the wireless communication interface 1612. As shown in FIG16 , the smartphone 1600 may include multiple antennas 1616. Although FIG16 shows an example in which the smartphone 1600 includes multiple antennas 1616, the smartphone 1600 may also include a single antenna 1616.

[0192] In addition, the smartphone 1600 may include an antenna 1616 for each wireless communication scheme. In this case, the antenna switch 1615 may be omitted from the configuration of the smartphone 1600.

[0193] The bus 1617 connects the processor 1601, the memory 1602, the storage device 1603, the external connection interface 1604, the camera 1606, the sensor 1607, the microphone 1608, the input device 1609, the display device 1610, the speaker 1611, the wireless communication interface 1612, and the auxiliary controller 1619. The battery 1618 supplies power to the various blocks of the smartphone 1600 shown in FIG16 via feeders, which are partially shown as dashed lines in the figure. The auxiliary controller 1619 operates the minimum necessary functions of the smartphone 1600, for example, in sleep mode.

[0194] Second application example

[0195] 17 is a block diagram illustrating an example of a schematic configuration of a car navigation device 1720 to which the techniques of the present disclosure may be applied. Car navigation device 1720 includes a processor 1721, a memory 1722, a global positioning system (GPS) module 1724, a sensor 1725, a data interface 1726, a content player 1727, a storage medium interface 1728, an input device 1729, a display device 1730, a speaker 1731, a wireless communication interface 1733, one or more antenna switches 1736, one or more antennas 1737, and a battery 1738. In one implementation, car navigation device 1720 (or processor 1721) herein may correspond to electronic device 300B described above.

[0196] The processor 1721 may be, for example, a CPU or an SoC, and controls a navigation function and other functions of the car navigation device 1720. The memory 1722 includes a RAM and a ROM, and stores data and programs executed by the processor 1721.

[0197] The GPS module 1724 uses GPS signals received from GPS satellites to measure the position (such as latitude, longitude, and altitude) of the car navigation device 1720. The sensor 1725 may include a group of sensors such as a gyroscope sensor, a geomagnetic sensor, and an air pressure sensor. The data interface 1726 is connected to, for example, the vehicle network 1741 via a terminal not shown, and obtains data generated by the vehicle (such as vehicle speed data).

[0198] The content player 1727 reproduces content stored in a storage medium (such as a CD or DVD) inserted into the storage medium interface 1728. The input device 1729 includes, for example, a touch sensor, button, or switch configured to detect a touch on the screen of the display device 1730, and receives operations or information input from the user. The display device 1730 includes a screen such as an LCD or OLED display and displays images of the navigation function or reproduced content. The speaker 1731 outputs sounds of the navigation function or reproduced content.

[0199] The wireless communication interface 1733 supports any cellular communication scheme (such as LTE and LTE-Advanced) and performs wireless communication. The wireless communication interface 1733 may generally include, for example, a BB processor 1734 and an RF circuit 1735. The BB processor 1734 may perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and perform various types of signal processing for wireless communication. Meanwhile, the RF circuit 1735 may include, for example, a mixer, a filter, and an amplifier, and transmit and receive wireless signals via an antenna 1737. The wireless communication interface 1733 may also be a chip module on which the BB processor 1734 and the RF circuit 1735 are integrated. As shown in Figure 17, the wireless communication interface 1733 may include multiple BB processors 1734 and multiple RF circuits 1735. Although Figure 17 shows an example in which the wireless communication interface 1733 includes multiple BB processors 1734 and multiple RF circuits 1735, the wireless communication interface 1733 may also include a single BB processor 1734 or a single RF circuit 1735.

[0200] In addition, in addition to the cellular communication scheme, the wireless communication interface 1733 can support other types of wireless communication schemes, such as short-range wireless communication schemes, near field communication schemes, and wireless LAN schemes. In this case, for each wireless communication scheme, the wireless communication interface 1733 can include a BB processor 1734 and an RF circuit 1735.

[0201] Each of the antenna switches 1736 switches a connection destination of the antenna 1737 between a plurality of circuits included in the wireless communication interface 1733 , such as circuits for different wireless communication schemes.

[0202] Each of the antennas 1737 includes a single or multiple antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals via the wireless communication interface 1733. As shown in FIG17 , the car navigation device 1720 may include multiple antennas 1737. Although FIG17 shows an example in which the car navigation device 1720 includes multiple antennas 1737, the car navigation device 1720 may also include a single antenna 1737.

[0203] In addition, the car navigation device 1720 may include an antenna 1737 for each wireless communication scheme. In this case, the antenna switch 1736 may be omitted from the configuration of the car navigation device 1720.

[0204] The battery 1738 supplies power to the respective blocks of the car navigation device 1720 shown in Fig. 17 via a feeder line, which is partially shown as a dotted line in the figure. The battery 1738 accumulates the power supplied from the vehicle.

[0205] The technology of the present disclosure may also be implemented as an in-vehicle system (or vehicle) 1740 including a car navigation device 1720, an in-vehicle network 1741, and one or more blocks of a vehicle module 1742. The vehicle module 1742 generates vehicle data (such as vehicle speed, engine speed, and fault information) and outputs the generated data to the in-vehicle network 1741.

[0206] It should be understood that the technical solutions of the present disclosure can be implemented through the following example implementations.

[0207] 1. An electronic device comprising a processing circuit, wherein the processing circuit is configured to:

[0208] Determining an AI / ML model corresponding to the AI / ML task of the first terminal device;

[0209] Based on the corresponding state information of the first terminal device and the other one or more terminal devices, forming splitting information of at least a first part of the AI / ML model, the splitting information specifying that split model inference is to be performed on multiple sub-parts of the at least first part of the AI / ML model by multiple participant devices; and

[0210] Based on the split information, the wireless network allocates resources for transmitting model reasoning information to at least one of the plurality of participant devices.

[0211] 2. The electronic device according to clause 1, wherein the processing circuit is further configured to: obtain corresponding state information of the first terminal device and one or more other terminal devices, wherein the state information is associated with model reasoning, and

[0212] The status information indicates the computing status and / or communication status of the corresponding terminal device, the computing status includes at least one of the computing resource usage status, the storage resource usage status or the power level, and the communication status includes at least one of the channel quality or the data rate.

[0213] 3. The electronic device according to clause 2, wherein the split information includes information indicating the plurality of sub-parts and information about a participant device performing model inference, and forming the split information comprises:

[0214] Determine the terminal devices whose corresponding status is better than a threshold and the first terminal device among the other one or more terminal devices as the multiple participant devices;

[0215] The at least first portion of the AI / ML model is split into the plurality of sub-portions based on respective state information of the plurality of participant devices.

[0216] 4. An electronic device according to clause 3, wherein the multiple sub-parts correspond to the multiple participant devices, the model reasoning workload of the multiple sub-parts matches the computing state of the corresponding participant devices, and the communication state of the multiple participant devices can support the transmission of model reasoning information.

[0217] 5. The electronic device according to clause 1, wherein:

[0218] The AI / ML model comprises a neural network model, and the at least first portion of the AI / ML model comprises one or more first layers of the AI / ML model, or comprises all layers of the AI / ML model; and / or

[0219] The model inference information includes model inference intermediate data and / or model inference result data.

[0220] 6. An electronic device as described in clause 4, wherein causing the wireless network to allocate resources for transmitting model inference information includes sending the split information of the at least first part of the AI / ML model to a base station.

[0221] 7. The electronic device of clause 6, wherein the processing circuit is further configured to:

[0222] Based on the split information, an instruction for executing model inference is sent to the corresponding participant device, where the instruction includes indication information of the corresponding sub-part of the AI / ML model and indication information of the downstream device.

[0223] 8. The electronic device of clause 7, wherein the electronic device is implemented as a network endpoint or a portion of a network endpoint, the network endpoint comprising a cloud server and / or an edge server.

[0224] 9. The electronic device of clause 7, wherein the electronic device is implemented as a first terminal device or a part of a first terminal device, and the processing circuit is further configured to:

[0225] receiving resource allocation information for a first terminal device from the base station, wherein the resource allocation information indicates resource allocation for at least one of a direct link, an uplink, and a downlink; or

[0226] Based on the split information, a direct link resource for transmitting model inference information is allocated to at least one of the multiple participant devices.

[0227] 10. The electronic device of clause 9, wherein the processing circuit is further configured to:

[0228] Inputting local data into a first sub-portion of the AI / ML model to obtain first intermediate data; and

[0229] Based on the resource allocation for the through link, first intermediate data is provided to the first participant device via the through link with the first participant device.

[0230] 11. The electronic device of clause 9, wherein the processing circuit is further configured to:

[0231] sending the first intermediate data to the second participant device via the direct link with the second participant device based on the resource allocation for the direct link;

[0232] receiving, based on the resource allocation for the through link, second intermediate data output by the second participant device via the through link with the second participant device;

[0233] sending the second intermediate data to the network via the uplink based on the resource allocation for the uplink;

[0234] receiving, based on resource allocation for the through link, an inference result forwarded by the second participant device via the through link with the second participant device; or

[0235] Based on the resource allocation for the downlink, an inference result corresponding to the AI / ML model is received from the network via the downlink.

[0236] 12. An electronic device for a base station, comprising a processing circuit, wherein the processing circuit is configured to:

[0237] Obtaining split information of at least a first portion of an AI / ML model, wherein the AI / ML model corresponds to an AI / ML task of a first terminal device, the split information specifying that split model inference is to be performed by a plurality of participant devices on a plurality of sub-portions of the at least first portion of the AI / ML model; and

[0238] Based on the split information, resources for transmitting model inference information are allocated to at least one of the plurality of participant devices.

[0239] 13. The electronic device of clause 12, wherein the processing circuit is further configured to:

[0240] forming the splitting information based on corresponding status information of the first terminal device and one or more other terminal devices; or

[0241] The splitting information is received from a first terminal device or a network endpoint.

[0242] 14. The electronic device of clause 13, wherein the split information includes information indicating the plurality of sub-parts and information about a participant device that performs model inference, and forming the split information comprises:

[0243] Determine the terminal devices whose corresponding status is better than a threshold and the first terminal device among the other one or more terminal devices as the multiple participant devices;

[0244] The at least first portion of the AI / ML model is split into the plurality of sub-portions based on respective state information of the plurality of participant devices.

[0245] 15. An electronic device according to clause 14, wherein the multiple sub-parts correspond to the multiple participant devices, the model reasoning workload of the multiple sub-parts matches the computing state of the corresponding participant devices, and the communication state of the multiple participant devices can support the transmission of model reasoning information.

[0246] 16. The electronic device of clause 13, wherein the processing circuit is further configured to:

[0247] Based on the split information, an instruction for executing model inference is sent to the corresponding participant device, where the instruction includes indication information of the corresponding sub-part of the AI / ML model and indication information of the downstream device.

[0248] 17. The electronic device of clause 16, wherein the processing circuit is further configured to:

[0249] Allocating resources to the corresponding participant devices based on the expected output data volume of the model inference of the corresponding sub-part of the AI / ML model; and

[0250] Resource allocation information is sent to corresponding participant devices, wherein the resource allocation information indicates resource allocation for at least one of a direct link, an uplink, and a downlink.

[0251] 18. An electronic device for a second terminal device, comprising a processing circuit, wherein the processing circuit is configured to:

[0252] Receiving an instruction to perform model inference from a first terminal device, the instruction including indication information of a corresponding sub-part of the AI / ML model and indication information of a downstream participant device;

[0253] receiving a resource allocation for transmitting model reasoning information, wherein the resource allocation indicates resources for wireless links with an upstream participant device and the downstream participant device;

[0254] receiving first intermediate data from the upstream participant device via a wireless link with the upstream participant device based on the instruction and the resource allocation;

[0255] Inputting the first intermediate data into the corresponding sub-part of the AI / ML model to obtain second intermediate data; and

[0256] Based on the instruction and the resource allocation, second intermediate data is sent to the downstream participant device via a wireless link with the downstream participant device.

[0257] 19. A model reasoning method in a wireless communication system, comprising:

[0258] Determining an AI / ML model corresponding to the AI / ML task of the first terminal device;

[0259] Based on the corresponding state information of the first terminal device and the other one or more terminal devices, forming splitting information of at least a first part of the AI / ML model, the splitting information specifying that split model inference is to be performed on multiple sub-parts of the at least first part of the AI / ML model by multiple participant devices; and

[0260] Based on the split information, the wireless network allocates resources for transmitting model reasoning information to at least one of the plurality of participant devices.

[0261] 20. A model reasoning method in a wireless communication system, comprising:

[0262] Obtaining split information of at least a first portion of an AI / ML model, wherein the AI / ML model corresponds to an AI / ML task of a first terminal device, the split information specifying that split model inference is to be performed by a plurality of participant devices on a plurality of sub-portions of the at least first portion of the AI / ML model; and

[0263] Based on the split information, resources for transmitting model inference information are allocated to at least one of the plurality of participant devices.

[0264] 21. A model inference method in a wireless communication system, comprising: a second terminal device:

[0265] Receiving an instruction to perform model inference from a first terminal device, the instruction including indication information of a corresponding sub-part of the AI / ML model and indication information of a downstream participant device;

[0266] receiving a resource allocation for transmitting model reasoning information, wherein the resource allocation indicates resources for wireless links with an upstream participant device and the downstream participant device;

[0267] receiving first intermediate data from the upstream participant device via a wireless link with the upstream participant device based on the instruction and the resource allocation;

[0268] Inputting the first intermediate data into the corresponding sub-part of the AI / ML model to obtain second intermediate data; and

[0269] Based on the instruction and the resource allocation, second intermediate data is sent to the downstream participant device via a wireless link with the downstream participant device.

[0270] 22. A computer program product comprising instructions which, when executed by a computer, cause the method according to any one of clauses 19 to 21 to be carried out.

[0271] The exemplary embodiments of the present disclosure are described above with reference to the accompanying drawings, but the present disclosure is certainly not limited to the above examples. Those skilled in the art may obtain various changes and modifications within the scope of the appended claims, and it should be understood that these changes and modifications will naturally fall within the technical scope of the present disclosure.

[0272] For example, a plurality of functions included in one unit in the above embodiments may be implemented by separate devices. Alternatively, a plurality of functions implemented by a plurality of units in the above embodiments may be implemented by separate devices, respectively. In addition, one of the above functions may be implemented by a plurality of units. Needless to say, such a configuration is included in the technical scope of the present disclosure.

[0273] In this specification, the steps described in the flowchart include not only processing executed in time series in the order described, but also processing executed in parallel or individually rather than necessarily in time series. In addition, even in the steps processed in time series, it goes without saying that the order can be changed as appropriate.

[0274] Although the present disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions and transformations can be made without departing from the spirit and scope of the present disclosure as defined by the appended claims. Moreover, the terms "comprises," "comprising," or any other variations thereof in the embodiments of the present disclosure are intended to cover non-exclusive inclusions, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. In the absence of further restrictions, an element defined by the statement "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

Claims

1. An electronic device comprising a processing circuit, wherein the processing circuit is configured to: Determining an AI / ML model corresponding to the AI / ML task of the first terminal device; Based on the corresponding state information of the first terminal device and the other one or more terminal devices, forming splitting information of at least a first part of the AI / ML model, the splitting information specifying that split model inference is to be performed on multiple sub-parts of the at least first part of the AI / ML model by multiple participant devices; and Based on the split information, the wireless network allocates resources for transmitting model reasoning information to at least one of the plurality of participant devices.

2. The electronic device according to claim 1, wherein The processing circuit is further configured to: obtain corresponding state information of the first terminal device and one or more other terminal devices, wherein the state information is associated with the model reasoning, and The status information indicates the computing status and / or communication status of the corresponding terminal device, the computing status includes at least one of the computing resource usage status, the storage resource usage status or the power level, and the communication status includes at least one of the channel quality or the data rate.

3. The electronic device according to claim 2, wherein The split information includes indication information of the multiple sub-parts and information of participant devices that perform model reasoning, and forming the split information includes: Determine the terminal devices whose corresponding status is better than a threshold and the first terminal device among the other one or more terminal devices as the multiple participant devices; The at least first portion of the AI / ML model is split into the plurality of sub-portions based on respective state information of the plurality of participant devices.

4. The electronic device according to claim 3, wherein The multiple sub-parts correspond to the multiple participant devices, the model reasoning workloads of the multiple sub-parts match the computing states of the corresponding participant devices, and the communication states of the multiple participant devices can support the transmission of model reasoning information.

5. The electronic device according to claim 1, wherein The AI / ML model comprises a neural network model, and the at least first portion of the AI / ML model comprises one or more first layers of the AI / ML model, or comprises all layers of the AI / ML model; and / or The model inference information includes model inference intermediate data and / or model inference result data. The electronic device according to claim 4 , wherein: Causing the wireless network to allocate resources for transmitting model inference information includes sending the split information of the at least first portion of the AI / ML model to a base station.

7. The electronic device according to claim 6, wherein: The processing circuit is further configured to: Based on the split information, an instruction for executing model inference is sent to the corresponding participant device, where the instruction includes indication information of the corresponding sub-part of the AI / ML model and indication information of the downstream device.

8. The electronic device according to claim 7, wherein: The electronic device is implemented as a network endpoint or a part of a network endpoint, and the network endpoint includes a cloud server and / or an edge server.

9. The electronic device according to claim 7, wherein: The electronic device is implemented as a first terminal device or a part of the first terminal device, and the processing circuit is further configured to: receiving resource allocation information for a first terminal device from the base station, wherein the resource allocation information indicates resource allocation for at least one of a direct link, an uplink, and a downlink; or Based on the split information, a direct link resource for transmitting model inference information is allocated to at least one of the multiple participant devices.

10. The electronic device according to claim 9, wherein The processing circuit is further configured to: Inputting local data into a first sub-portion of the AI / ML model to obtain first intermediate data; and Based on the resource allocation for the through link, first intermediate data is provided to the first participant device via the through link with the first participant device.

11. The electronic device according to claim 9, wherein The processing circuit is further configured to: sending the first intermediate data to the second participant device via the direct link with the second participant device based on the resource allocation for the direct link; receiving, based on the resource allocation for the through link, second intermediate data output by the second participant device via the through link with the second participant device; sending the second intermediate data to the network via the uplink based on the resource allocation for the uplink; receiving, based on resource allocation for the through link, an inference result forwarded by the second participant device via the through link with the second participant device; or Based on the resource allocation for the downlink, an inference result corresponding to the AI / ML model is received from the network via the downlink.

12. An electronic device for a base station, comprising a processing circuit, wherein the processing circuit is configured to: Obtaining split information of at least a first portion of an AI / ML model, wherein the AI / ML model corresponds to an AI / ML task of a first terminal device, the split information specifying that split model inference is to be performed by a plurality of participant devices on a plurality of sub-portions of the at least first portion of the AI / ML model; and Based on the split information, resources for transmitting model inference information are allocated to at least one of the plurality of participant devices.

13. The electronic device according to claim 12, wherein: The processing circuit is further configured to: forming the splitting information based on corresponding status information of the first terminal device and one or more other terminal devices; or The splitting information is received from a first terminal device or a network endpoint.

14. The electronic device according to claim 13, wherein: The split information includes indication information of the multiple sub-parts and information of participant devices that perform model reasoning, and forming the split information includes: Determine the terminal devices whose corresponding status is better than a threshold and the first terminal device among the other one or more terminal devices as the multiple participant devices; The at least first portion of the AI / ML model is split into the plurality of sub-portions based on respective state information of the plurality of participant devices.

15. The electronic device according to claim 14, wherein The multiple sub-parts correspond to the multiple participant devices, the model reasoning workloads of the multiple sub-parts match the computing states of the corresponding participant devices, and the communication states of the multiple participant devices can support the transmission of model reasoning information.

16. The electronic device according to claim 13, wherein: The processing circuit is further configured to: Based on the split information, an instruction for executing model inference is sent to the corresponding participant device, where the instruction includes indication information of the corresponding sub-part of the AI / ML model and indication information of the downstream device.

17. The electronic device according to claim 16, wherein: The processing circuit is further configured to: Allocating resources to the corresponding participant devices based on the expected output data volume of the model inference of the corresponding sub-part of the AI / ML model; and Resource allocation information is sent to corresponding participant devices, wherein the resource allocation information indicates resource allocation for at least one of a direct link, an uplink, and a downlink.

18. An electronic device for a second terminal device, comprising a processing circuit, wherein the processing circuit is configured to: Receiving an instruction to perform model inference from a first terminal device, the instruction including indication information of a corresponding sub-part of the AI / ML model and indication information of a downstream participant device; receiving a resource allocation for transmitting model reasoning information, wherein the resource allocation indicates resources for wireless links with an upstream participant device and the downstream participant device; receiving first intermediate data from the upstream participant device via a wireless link with the upstream participant device based on the instruction and the resource allocation; Inputting the first intermediate data into a corresponding sub-part of the AI / ML model to obtain second intermediate data; as well as Based on the instruction and the resource allocation, second intermediate data is sent to the downstream participant device via a wireless link with the downstream participant device.

19. A model reasoning method in a wireless communication system, comprising: Determining an AI / ML model corresponding to the AI / ML task of the first terminal device; Based on corresponding state information of the first terminal device and one or more other terminal devices, forming splitting information of at least a first part of the AI / ML model, the splitting information specifying that split model inference is to be performed on multiple sub-parts of the at least first part of the AI / ML model by multiple participant devices; as well as Based on the split information, the wireless network allocates resources for transmitting model reasoning information to at least one of the plurality of participant devices.

20. A model reasoning method in a wireless communication system, comprising: Obtaining split information of at least a first portion of an AI / ML model, wherein the AI / ML model corresponds to an AI / ML task of a first terminal device, the split information specifying that split model inference is to be performed on a plurality of sub-portions of the at least first portion of the AI / ML model by a plurality of participant devices; as well as Based on the split information, resources for transmitting model inference information are allocated to at least one of the plurality of participant devices.

21. A model inference method in a wireless communication system, comprising: Receiving an instruction to perform model inference from a first terminal device, the instruction including indication information of a corresponding sub-part of the AI / ML model and indication information of a downstream participant device; receiving a resource allocation for transmitting model reasoning information, wherein the resource allocation indicates resources for wireless links with an upstream participant device and the downstream participant device; receiving first intermediate data from the upstream participant device via a wireless link with the upstream participant device based on the instruction and the resource allocation; Inputting the first intermediate data into a corresponding sub-part of the AI / ML model to obtain second intermediate data; as well as Based on the instruction and the resource allocation, second intermediate data is sent to the downstream participant device via a wireless link with the downstream participant device.

22. A computer program product comprising instructions which, when executed by a computer, cause the method according to any one of claims 19 to 21 to be implemented.