Method and apparatus for scheduling for distributed inference

By determining UE capabilities and dividing inference tasks into subtasks, the central node optimizes scheduling for distributed inference, enhancing performance and reducing latency in wireless communication networks.

WO2025227545A1PCT designated stage Publication Date: 2025-11-06HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/109793
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-03
Filing Date
2024-08-05
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Low-cost devices in wireless communication networks are unreliable for distributed inference tasks due to low individual capability and unsatisfactory inference performance, necessitating improved communication efficiency for collaborative work.

Method used

A central node determines configuration information based on UE capabilities, dividing inference tasks into subtasks and assigning them to multiple devices, using topologies like ping-pong, chain, or star topologies, and scheduling these devices to perform distributed inference efficiently.

Benefits of technology

This approach enhances inference performance by reducing latency and DCI transmissions, improving scheduling efficiency, and ensuring predictable data transmissions among UEs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109793_06112025_PF_FP_ABST
    Figure CN2024109793_06112025_PF_FP_ABST
Patent Text Reader

Abstract

There is provided methods, apparatus and systems for scheduling for distributed inference. According to embodiments, based on collected device capabilities from a plurality of devices that are to perform a distributed inference, a central node can divide an inference task into multiple inference subtasks and assign these tasks to the plurality of devices. Furthermore, according to embodiments, also based on the collected device capabilities and divided inference subtasks, a central node can configure the distributed inference network or network topology associated with the devices, such that these devices can work jointly to perform the distributed inference task. According to embodiments, the plurality of devices associated with the distributed inference task can be transmitted their respective configuration information from the central node. The configuration information can include details regarding the inference subtask being performed and information indicative of reception of input for the inference subtask and the transmission of output upon completion of the inference subtask.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR SCHEDULING FOR DISTRIBUTED INFERENCE

[0001] RELATED APPLICATION

[0002] This application claims the benefit of priority to U.S. provisional patent application serial no. 63 / 642, 410, entitled “Method and Apparatus for Scheduling for Distributed Inference” and filed on May 3, 2024. The entire content of the above application is hereby incorporated by reference.TECHNICAL FIELD

[0003] The present disclosure pertains to the field of communication networks, and in particular to methods, apparatus and systems for scheduling for distributed inference.BACKGROUND

[0004] The next-generation wireless communication networks are expected to deliver services beyond data transmissions. The application scenarios will become more diverse, including artificial intelligence (AI) , machine learning (ML) and sensing functionalities. A difference from existing AI and sensing devices is that they will be carried out jointly by the wireless network and user equipment. To realize this promise, advanced distributed algorithms, together with new protocols, are being introduced.

[0005] In some examples of a computation known as distributed inference, where the task of a giant deep neural network (DNN) at a centralized data center is decomposed and assigned to many local devices. In this way, the computation and traffic load of each single device is greatly alleviated, and a difficult inference task can be collectively solved by multiple low-cost devices powered by battery. Such inference frameworks are just some examples, future wireless-enabled sensing or joint communications and sensing applications also require the collection, encoding, transmission and decoding of data collected from physical layer, rather than higher layers.

[0006] Still, challenges arise as these low-cost devices may be unreliable, both in terms of communication and computation. Their individual capability is relatively low and the resulting inference performance may be unsatisfactory. To enhance the inference performance for providing quality of service (QoS) , many devices may be required to work collaboratively. Communication efficiency that is tailored to new application scenarios is one area that is subject to improvement.

[0007] Therefore, there is a need for methods, apparatus and systems that mitigate or obviate one or more limitations in the state of the art.SUMMARY

[0008] Aspects of the present disclosure provide methods, apparatuses and systems for scheduling for distributed inference.

[0009] According to embodiments, there is provided a method for scheduling an inference task as a distributed inference task, the inference task associated with a plurality of user equipments (UE) . The method includes determining, by a central node, configuration information indicative of the distributed inference task, the determining based at least in part on capability information associated with each of the plurality of UEs. The method further includes transmitting, by the central node to each of the UEs, the configuration information indicative of the distributed inference task.

[0010] In some embodiments, the method further includes receiving, by the central node, respective capability information from each of the plurality of UEs.

[0011] In some embodiments, the capability information includes one or more of inference time, supported artificial intelligence  / machine learning (AI / ML) algorithms and neural network parameters. In some embodiments, the determining includes determining a topology for association with the plurality of UEs for performing the distributed inference task. In some embodiments, the topology is selected from the group comprising a ping-pong topology, a chain topology, a ring topology and a star topology.

[0012] In some embodiments, the inference task is divided into inference subtasks, wherein each of the inference subtasks are based at least in part on the topology and the capabilities of each of the plurality of UEs. In some embodiments, for each of the plurality of UEs, the configuration information includes information indicative of both reception of input for the inference subtask associated with a particular UE and transmission of output upon completion of the inference subtask associated with the particular UE. In some embodiments, the configuration information includes information indicative of multiple inference tasks being performed by the plurality of UEs. In some embodiments, each of the multiple inference tasks has an inference process ID (IID) and wherein each inference subtask associated with a particular inference task has a subtask assignment index (SAI) .

[0013] In some embodiments, upon completion of a configuration phase, transmitting, by the central node to each of the plurality of UEs, a schedule indicative of performance of the distributed inference task. In some embodiments, each of the plurality of UEs are scheduled individually. In some embodiments, the plurality of UEs are jointly scheduled.

[0014] According to embodiments, there is provided an apparatus for scheduling an inference task as a distributed inference task, the apparatus comprising a processor and memory for storing instructions, the instructions when executed by the processor configure the apparatus to perform one or more of the above or further defined elsewhere herein.

[0015] According to embodiments, there is provided a system for scheduling an inference task as a distributed inference task, the system comprising a central node and a plurality of user equipments (UEs) , each including a processor and memory for storing instructions. The instructions when executed by a respective processor configure the central node to determine configuration information indicative of the distributed inference task, the determining based at least in part on capability information associated with each of the plurality of UEs and transmit, to each of the UEs, the configuration information indicative of the distributed inference task. The instructions when executed by a respective processor configure the plurality of UEs to receive, from the central node, the configuration information indicative of the distributed inference task.

[0016] According to some embodiments, the instructions when executed by the respective processor further configure each of the plurality of UEs to transmit, to the central node, respective capability information.

[0017] According to some embodiments, during the determining, the instructions when executed by the respective processor further configure the central node to determine a topology for association with the plurality of UEs for performing the distributed inference task.

[0018] According to some embodiments, upon completion of a configuration phase, the instructions when executed by the respective processor further configure the central node to transmit, to each of the plurality of UEs, a schedule indicative of performance of the distributed inference task.BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Further features and advantages of the present invention will become apparent from the following detailed description, taken in combination with the appended drawings, in which:

[0020] FIG. 1 illustrates a communication system in which embodiments of the present disclosure can be implemented.

[0021] FIG. 2 illustrates other aspects of a communication system in which embodiments of the present disclosure can be implemented.

[0022] FIG. 3 illustrates apparatuses in communication with one another, in which embodiments of the present disclosure can be implemented.

[0023] FIG. 4 illustrates further aspects of an apparatus according to embodiments of the present disclosure.

[0024] FIG. 5 illustrates further aspects of an apparatus according to embodiments of the present disclosure.

[0025] FIG. 6 illustrates a system for performing redundant inferences.

[0026] FIG. 7 illustrates a correlated coded inference framework.

[0027] FIG. 8 illustrates an example inference scenario which is associated with various embodiments of the present disclosure.

[0028] FIG. 9 illustrates a distributed computing network which is associated with various embodiments of the present disclosure.

[0029] FIG. 10 illustrates aspects of a link between network nodes, according to embodiments of the present disclosure.

[0030] FIG. 11 illustrates example scenarios involving data transfers between nodes, which can be performed according to embodiments of the present disclosure.

[0031] FIG. 12 illustrates an example of a decoder layer in a transformer neural network, according to embodiments of the present disclosure.

[0032] FIG. 13 illustrates a ping-pong topology for application with embodiments of the present disclosure.

[0033] FIG. 14 illustrates an example resource assignment for the topology illustrated in FIG. 13.

[0034] FIG. 15 illustrates a chain topology for application with embodiments of the present disclosure.

[0035] FIG. 16 illustrates an example resource assignment for the topology illustrated in FIG. 15.

[0036] FIG. 17 illustrates a ring topology for application with embodiments of the present disclosure.

[0037] FIG. 18 illustrates an example resource assignment for the topology illustrated in FIG. 17.

[0038] FIG. 19 illustrates another example resource assignment for the topology illustrated in FIG. 17.

[0039] FIG. 20 illustrates a star topology for application with embodiments of the present disclosure.

[0040] FIG. 21 illustrates an example resource assignment for the topology illustrated in FIG. 20.

[0041] FIG. 22 illustrates a method to deploy a distributed inference network in a cellular network, according to embodiments of the present disclosure.

[0042] FIG. 23 illustrates a chain topology for distributed inference, wherein this topology is being used for three jointly scheduled inference processes, according to a embodiment of the present disclosure.

[0043] FIG. 24 illustrates an example topology for application with embodiments of the present disclosure.

[0044] FIG. 25 illustrates a timing sequence when the UEs immediately transmit immediately after the particular UE completes its inference subtask, according to embodiments of the present disclosure.

[0045] FIG. 26 illustrates a timing sequence when a UE needs to wait for another period of time after it completes its inference subtask, according to embodiments of the present disclosure.

[0046] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION

[0047] Some wireless communication protocols adopted the limited idea of multi-node collaborative AI / ML algorithms. In 5.5G, it has already been proposed that base station (BS) and user equipment (UE) can jointly perform channel state information (CSI) compression and decompression based on two AI / ML components (e.g., neural networks) . For example, a UE-side AI compresses the CSI and a BS-side AI decompresses and recovers the CSI. Such a two-sided model requires the cooperation of BS and UE.

[0048] According to some embodiments, it can be generalized that a distributed inference framework can allow the participation and cooperation of more network entities in future generation communication systems. In addition to CSI compression / decompression, future generation networks are expected to perform general AI / ML tasks in a distributed way.

[0049] According to embodiments, for distributed inference tasks, data transmissions between UE workers and BS can be coordinated seamlessly so that less waiting time will occur between data transmissions and inference. For example, once UE1 has finished its own inference subtask, it can immediately transmit the intermediate results to UE2 for further inference. Therefore, even at the beginning of the entire inference task, the BS can be able to know all the transmissions that are required to complete the task.

[0050] As such, all transmissions are quite predictable and the scheduling information can be dramatically compressed. Specifically, a BS may jointly schedule a sequence of consecutive transmissions for a group of UE workers, instead of schedule each transmission for each UE worker separately. This can improve scheduling efficiency by reducing the number of downlink control information (DCI) transmissions and thus can reduce physical downlink control channel (PDCCH) blind detections. The overall inference latency can also be reduced.

[0051] According to embodiments, based on collected device capabilities from a plurality of devices that are to perform a distributed inference, a central node (for example a base station) can divide a training or inference task into multiple inference subtasks and assign these tasks to the plurality of devices. Furthermore, according to embodiments, based on the collected device capabilities and divided inference subtasks, a central node can configure the distributed inference network or network topology associated with the devices, such that these devices can work jointly to perform a distributed inference task. According to embodiments, the plurality of devices associated with the distributed inference task can be transmitted their respective configuration  information from the central node. The configuration information can include details regarding the inference subtask being performed and information indicative of reception of input for the inference subtask and the transmission of output upon completion of the inference subtask. In some embodiments, the transmission of the configuration information can be performed through the transmission of downlink control information (DCI) , however other transmission formats may be used as would be readily understood.

[0052] According to embodiments, multiple distributed inference tasks can be performed by one or more of the UEs at the same time. As such in order to keep track of each of the inference processes and the inference subtasks associated with each of the inference processes and the related transmissions. According to embodiments, once the inference process ID (IID) and subtask assignment index (SAI) are configured, the whole inference network or topology is known to the UEs (e.g. chain topology, ring topology, star topology etc. ) . The remaining scheduling information includes the time and frequency resource assignment for each transmission among the UEs.

[0053] FIG. 1 illustrates a communication system in which embodiments of the present disclosure can be implemented. Referring to FIG. 1, as an illustrative example, a simplified schematic illustration of a communication system is provided. The communication system 100 may comprise a radio access network 120. The radio access network (RAN) 120 may be a future generation radio access network, or a legacy (e.g. 5th generation (5G) , 4th generation (4G) , 3th generation (3G) or 2nd generation (2G) ) radio access network. In some implementations, radio access refers to a next generation air interface of standards which may comprise both terrestrial networks (TNs) and non-terrestrial networks (NTNs) , and more details will be described below. One or more communication electronic device (ED) 110a, 110b, 110c, 110d, 110e, 110f, 110g, 110h, 110i, 110j (generically referred to as 110) may be interconnected to one another or connected to one or more network nodes 170a, 170b (generically referred to as 170) in the RAN 120. A core network (CN) 130 may be a part of the communication system and may be dependent or independent of the radio access technology used in the communication system 100. The communication system 100 may also comprise a public switched telephone network (PSTN) 140, the internet 150, and other networks 160.

[0054] Referring to FIG. 1, as an illustrative example, a simplified schematic illustration of a communication system is provided. The communication system 100 may comprise a radio access network 120. The radio access network (RAN) 120 may be a future generation radio access network, or a legacy (e.g. 5th generation (5G) , 4th generation (4G) , 3th generation (3G) or 2nd generation (2G) ) radio access network. In some implementations, radio access refers to a next generation air interface of standards which may comprise both terrestrial networks (TNs) and non-terrestrial networks (NTNs) , and more details will be described below. One or more communication electronic device (ED) 110a, 110b, 110c, 110d, 110e, 110f, 110g, 110h, 110i, 110j (generically referred to as 110) may be interconnected to one another or connected to one or more network nodes 170a, 170b (generically referred to as 170) in the RAN 120. A core network (CN) 130 may be a part of the communication system and may be dependent or independent of the radio access technology used in the communication system 100. The communication system 100 may also comprise a public switched telephone network (PSTN) 140, the internet 150, and other networks 160.

[0055] In general, the communication system 100 enables communication of multiple wireless or wired elements. The communication system 100 may provide content, such as voice, data, video, and / or text, via broadcast, multicast, groupcast, unicast, etc. The communication system 100 may operate by sharing resources, such as carrier spectrum bandwidth, among its constituent elements.

[0056] The communication system 100 may provide a wide range of communication services and applications including enhanced mobile broadband (eMBB) services, ultra-reliable low-latency communication (URLLC) services, massive machine type communication (mMTC) services, integrated sensing and communication (ISAC) , immersive communication, massive communication, hyper reliable and low-latency communication, ubiquitous connectivity, integrated AI and communication, and other services that can be provided by a future generation communication system. The communication system 100 may provide other services and applications such as earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery and mobility, etc.

[0057] The communication system 100 may include a terrestrial communication system (or network) and / or a non-terrestrial communication system (or network) . The communication system 100 may provide a high degree of availability and robustness through a joint operation of a terrestrial communication system and a non-terrestrial communication system. For example,  integrating a non-terrestrial communication system (or components thereof) into a terrestrial communication system can result in a heterogeneous network comprising multiple layers. The heterogeneous network may achieve better overall performance through efficient multi-link joint operation, more flexible functionality sharing, and faster physical layer link switching between terrestrial networks and non-terrestrial networks. The terrestrial communication system and the non-terrestrial communication system could be considered sub-systems of the communication system 100.

[0058] FIG. 2 illustrates other aspects of a communication system in which embodiments of the present disclosure can be implemented. In particular, FIG. 2 illustrates another example for communication system 100. As described earlier, the communication system 100 may include ED 110a, 110b, 110c, 110d (generically referred to as ED 110) , RAN 120a, 120b, and one or more of a CN 130, a PSTN 140, the internet 150, and other networks 160. In addition, the communication system 100 may also include a non-terrestrial network (NTN) 120c. The RANs 120a, 120b may include respective network nodes 170a, 170b such as base stations 170a, 170b, which may be generically referred to as terrestrial network (TN) devices or terrestrial transmit and receive points (T-TRPs) 170a, 170b (generically referred to as 170) . As referred to herein, the terms “TRP” and “base station” may be used interchangeably unless explicitly noted otherwise in a given example or section. For brevity, this disclosure may primarily refer to base station; however, absent an explicit limitation, references to TRP are merely non-limiting instances of interchangeable use. The T-TRPs 170a, 170b may be base stations mounted on a building or tower. In one implementation, the NTN 120c includes a RAN node such as base station 172, which may be generically referred to as an NTN device, a non-terrestrial node, a non-terrestrial network device, a non-terrestrial base station, or a non-terrestrial transmit and receive point (NT-TRP) 172.

[0059] In some implementations, the NT-TRP 172 is not attached to the ground, for example, in the case of an airborne base station. An airborne base station may be implemented using communication equipment supported or carried by a flying device. For example, a flying device may include an airborne platform (e.g. a blimp or an airship) , balloon, drone (e.g. quadcopter) , and other types aerial vehicles. In some implementations, an airborne base station may be supported or carried by an unmanned aerial system (UAS) or an unmanned aerial vehicle (UAV) , such as a drone. An airborne base station may be a moveable or mobile base station that can be flexibly deployed in different locations to meet network demand. A satellite base station is another example of a non-terrestrial base station. A satellite base station may be implemented using communication equipment supported or carried by a satellite. A satellite base station may also be referred to as an orbiting base station. High altitude platforms is yet another example of a non-terrestrial base station, including international mobile telecommunication base stations.

[0060] As referred to herein, and unless specified otherwise, a “TRP” may also refer to a T-TRP or a NT-TRP, a “T-TRP” may also refer to a “TN TRP” , and a “NT-TRP” may also refer to a “NTN TRP” . The NTN 120c may be considered to be a radio access network (RAN) , with operational aspects in common with the RANs 120a, 120b. The NTN 120c may include at least one NTN device and at least one corresponding terrestrial network device, the at least one NTN device may function as a transport layer device and the at least one corresponding terrestrial network device may function as a RAN node, which communicates with the ED 110 via the non-terrestrial network device. In addition, there may be a NTN gateway in the ground (i.e., referred as a terrestrial network device) also function as a transport layer device to communicate with both the NTN device and the RAN node. The RAN node may communicate with the ED 110 via the NTN device and the NTN gateway. In some implementations, the NTN gateway and the RAN node may be located in the same device.

[0061] A base station (also referred to TRP as stated above) 170 may be a network element in radio access network responsible for radio transmission and reception in one or more cells to or from the user equipment. Base station 170 may be known by other names in some implementations, such as a base transceiver station (BTS) , a radio base station, a network node, a network device, a device on the network side, a transmit / receive node, a Node B, an evolved NodeB (eNodeB or eNB) , a Home eNodeB, a next Generation NodeB (gNB) , a transmission point (TP) , a site controller, an access point (AP) , a wireless router, a relay station, a terrestrial node, a terrestrial network device, a terrestrial base station, a positioning node, among other possibilities. The base station 170 may be a macro base station (BS) , a pico BS, a relay node, a donor node, or the like, or combinations thereof. When a base station 170 performs (or is configured to perform) a method described herein, it may be interpreted as the base station, one or more modules (or units) in the base station, a circuit or chip, or a combination thereof, may perform the method. For example, the circuit or chip may include a modem chip, also referred to as a baseband chip, a system on chip (SoC) including a modem core, system in package (SIP) ) , and the like, and may be responsible for one or more communication functions in the base station.

[0062] The EDs 110a-110d and TRPs 170a-170b, 172 are examples of communication equipment that can be configured to implement some or all of the operations and / or embodiments described herein. The T-TRP 170a forms part of the RAN 120a, which may include other TRPs, and / or other devices. Also, the TRP 170b forms part of the RAN 120b, which may include other TRPs, and / or devices. Each TRP 170a, 170b may transmit and / or receive wireless signals within a particular geographic region or area, sometimes referred to as a “cell” or “coverage area” . The TRPs 170a-170b may be responsible for allocating and  / or configuring resources and transmission and / or reception in a set of cells. A cell may be a Radio network object that can be uniquely identified from a (cell) identification that is broadcasted over a geographical region or area from base stations associated with the cell. A Cell can be either FDD or TDD mode. A cell may also refer to the carrier frequencies within the DL / UL carrier bandwidth resources of a single standalone carrier or a component carrier in a carrier aggregation mode. A cell may be further divided into cell sectors, and a base station 170a-170b may, for example, employ multiple transceivers to provide service to multiple sectors. In some implementations, there may be established pico or femto cells where the radio access technology supports such. In some implementations, multiple transceivers could be used for each cell, for example using multiple-input multiple-output (MIMO) technology. The number of RAN 120a-120b shown is exemplary only. Any number of RAN may be contemplated when devising the communication system 100.

[0063] Any base station may be a single element, as shown, or multiple elements, distributed in the corresponding RAN, or otherwise. In some implementations, a plurality of RAN nodes coordinate to assist the ED 110 in implementing radio access, and different RAN nodes separately implement different functions of the base station. For example, the RAN node may be a central unit (CU) , a distributed unit (DU) , a CU-control plane (CP) , a CU-user plane (UP) , or a radio unit (RU) etc. The CU and the DU may be separately deployed, or may be included in a same element (i.e., a baseband unit (BBU) ) . The RU may be included in a radio frequency device or a radio frequency unit (i.e., a remote radio unit (RRU) , an active antenna unit (AAU) , or a remote radio head (RRH) ) . In different systems, the CU (or the CU-CP and the CU-UP) , the DU, or the RU may also have different names, but a person skilled in the art may understand meanings thereof. For example, in an open radio access network (ORAN) system, a CU may also be referred to as an open CU (O-CU) , a DU may also be referred to as an open DU (O-DU) , and a CU-CP may also be referred to as an open CU-CP (O-CU-CP) . The CU-UP may also be referred to as an open CU-UP (O-CU-UP) , and the RU may also be referred to as an open RU (O-RU) . Any one of the CU (or the CU-CP, the CU-UP) , the DU, and the RU may be implemented by using a software module, a hardware module, or a combination of a software module and a hardware module.

[0064] Further, communication (s) between different devices / apparatuses in various embodiments of this application may refer to direct communication between different devices / apparatuses (that is, no forwarding is required by another device / apparatuses) , or may refer to communication (s) between different devices / apparatuses via another device / apparatus (that is, forwarding is required by another device / apparatus) . Alternatively, it may refer to that a functional unit inside the device / apparatus uses another functional unit in the device / apparatus to communicate with another device / apparatus. In other words, "sending (or transmitting) information to... (an ED or a base station) " in this application may be understood as that a destination endpoint of the information is an ED or a base station. It may include sending / transmitting information directly or indirectly to an ED or a base station. Similarly, "receiving information from... (an ED or a base station) " may be understood as that a source endpoint of the information is an ED or a base station, and may include directly or indirectly receiving information from an ED or a base station. Necessary processing such as format conversion, digital-to-analog conversion, amplification, and filtering may be performed on the information between the source endpoint that sends the information and the destination endpoint. However, the destination endpoint may understand valid information from the source endpoint. Similar descriptions in this application may be understood similarly. Details are not described herein again. In the present disclosure, the terms "send" and "transmit" may be used interchangeably in embodiments of this application.

[0065] The ED 110 is used to connect persons, objects, machines, etc. The ED 110 may be widely used in various scenarios including, for example, cellular communications, device-to-device (D2D) , vehicle to everything (V2X) , peer-to-peer (P2P) , machine-to-machine (M2M) , MTC, internet of things (IoT) , virtual reality (VR) , augmented reality (AR) , mixed reality (MR) , metaverse, digital twin, industrial control, self-driving, remote medical, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery and mobility, etc.

[0066] Each ED 110 represents any suitable end user device for wireless operation and may include such devices (or may be referred to but not limited to) as a user equipment (UE) or a user device or a terminal device, a wireless transmit / receive unit (WTRU) , a mobile station, a fixed or mobile subscriber unit, a cellular telephone, a station (STA) , a MTC device, a personal digital assistant (PDA) , a smartphone, a laptop, a computer, a tablet, a wireless sensor, a consumer electronics device, a smart book, a vehicle, a car, a truck, a bus, a train, or an IoT device, wearable devices (such as a watch, a pair of glasses, head mounted equipment, etc. ) , an industrial device, or an apparatus in (e.g. module, modem, or chip) or comprising the forgoing devices, among other possibilities. Future generation EDs 110 may be referred to using other terms. When an ED 110 performs (or is configured to perform) a method described herein, it may be interpreted as the ED, one or more module (or units) in the ED, a circuit or chip, or a combination thereof, may perform the method. For example, the circuit or chip may include a modem chip, also referred to as a baseband chip, a system on chip (SoC) including a modem core, or system in package (SIP) ) , and the like, and may be responsible for one or more communication functions in the ED.

[0067] Each ED 110 connected to TRPs 170a-170b, and / or TRPs 172 can be dynamically or semi-statically turned-on (i.e., established, activated, or enabled) , turned-off (i.e., released, deactivated, or disabled) and / or configured in response to one of more of:connection availability and connection necessity.

[0068] Any ED 110 may be alternatively or additionally configured to interface, access, or communicate with any TRPs 170a, 170b and 172, the Internet 150, the CN 130, the PSTN 140, the other networks 160, or any combination of the preceding. In some examples, ED 110a may communicate an uplink (UL) and / or downlink (DL) transmission over a terrestrial air interface 190a with station-TRP 170a. In some examples, the EDs 110a, 110b, 110c, and 110d may also communicate directly with one another via one or more sidelink (SL) air interfaces 190b. In some examples, ED 110d may communicate an UL and / or DL transmission over a non-terrestrial air interface 190c with NT-TRP 172.

[0069] An air interface (e.g., 190a, 190b, 190c) generally includes a number of components and associated parameters that collectively specify how a transmission is to be sent and / or received over a wireless communications link between two or more communicating devices such as ED and base station. For example, an air interface may include one or more components defining the waveform (s) , frame structure (s) , multiple access scheme (s) , protocol (s) , coding scheme (s) and / or modulation scheme (s) for conveying information (e.g., data) over a wireless communications link. The air interfaces 190a and 190b may use similar communication technology, such as any suitable radio access technology.

[0070] The non-terrestrial air interface 190c can enable communication between the ED 110d and one or multiple NT-TRPs 172 via a wireless link or simply a link. For some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection between a group of EDs 110 and one or multiple NT-TRPs 172 for multicast transmission.

[0071] The TRPs 170a-170b, 172 may communicate with one another over one or more air interfaces 190e, 190f using wireless communication links (e.g., radio frequency (RF) , microwave, infrared (IR) , etc. ) or wired communication links. The air interfaces 190e, 190f may utilize any suitable radio access technology, and may be substantially similar to the air interfaces 190a, 190c over which the EDs 110a-110d communicate with one or more of the TRP 170a-170b, 172 or they may be substantially different. For example, the communication system 100 may implement one or more channel access methods, such as code division multiple access (CDMA) , time division multiple access (TDMA) , frequency division multiple access (FDMA) , orthogonal FDMA (OFDMA) , or single-carrier FDMA (SC-FDMA) .

[0072] The RANs 120a and 120b are in communication with the CN 130 to provide the EDs 110a 110b, and 110c with various services such as voice, data, and other services. The RANs 120a and 120b and / or the CN 130 may be in direct or indirect communication with one or more other RANs (not shown) , which may or may not be directly served by CN 130, and may or may not employ the same radio access technology as RAN 120a, RAN 120b or both. The CN 130 may also serve as a gateway access between (i) the RANs 120a and 120b or EDs 110a 110b, and 110c or both, and (ii) other networks (such as the PSTN 140, the Internet 150, and the other networks 160) . In addition, some or all of the EDs 110a 110b, and 110c may include functionality for communicating with different wireless networks over different wireless links using different wireless technologies and / or protocols. Instead of wireless communication (or in addition thereto) , the EDs 110a 110b, and 110c may communicate via wired communication channels to a service provider or switch (not shown) , and to the Internet 150. PSTN 140 may include circuit  switched telephone networks for providing plain old telephone service (POTS) . Internet 150 may include a network of computers and subnets (intranets) or both, and incorporate protocols, such as internet protocol (IP) , transmission control protocol (TCP) , user datagram protocol (UDP) . EDs 110a 110b, and 110c may be multimode devices capable of operation according to multiple radio access technologies, and incorporate multiple transceivers necessary to support such.

[0073] In addition, the communication system 100 may comprise a sensing agent (not shown) to manage the sensed data from ED 110 and / or any one of TRPs 170 a-170b, 172. In one implementation, the sensing agent may be part of any one of TRPs 170 a-b, 172. In another implementation, the sensing agent is a separate node that can communicate with the CN 130 and / or the RAN 120 (e.g., any one of TRPs 170 a-b, 172) .

[0074] FIG. 3 illustrates an example of an apparatus 310 wirelessly communicating with apparatus 320 in a communication system (e.g., the communication system 100) , in which embodiments of the present disclosure can be implemented. The apparatus 310 may be an electronic device (e.g. ED 110) . The apparatus 320 may be a network node (e.g. network node 170) such as T-TRP 170 or a NT-TRP 172. Although there is only one apparatus 310, and one apparatus 320 shown in the figure, the number of apparatus 310 and / or 320 could be one or more. For example, one ED 110 may be served by only one T-TRP 170 (or one NT-TRP 172) , by more than one T-TRP 170 (or more than one NT-TRP 172) . One ED 110 may be served by one or more T-TRP 170 and one or more NT-TRP172. Similarly, one T-TRP 170 (or one NT-TRP172) may serve one or more ED 110.

[0075] Apparatus 310 includes at least one processor 210. Only one processor 210 is illustrated to avoid congestion in the drawing. The apparatus 310 may further include a transmitter 201 and a receiver 203 coupled to one or more antennas 204. Only one antenna 204 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas 204 may alternatively be panels. The transmitter 201 and the receiver 203 may be integrated, e.g. as a transceiver. The transceiver is configured to modulate data or other content for transmission by at least one antenna 204 or network interface controller (NIC) . The transceiver is also configured to demodulate data or other content received by the at least one antenna 204. Each transceiver includes any suitable structure for generating signals for wireless or wired transmission and / or processing signals received wirelessly or by wire. Each antenna 204 includes any suitable structure for transmitting and / or receiving wireless or wired signals. The apparatus 310 may include at least one memory 208. Only the transmitter 201, receiver 203, processor 210, memory 208, and antenna 204 is illustrated for simplicity, but the apparatus 310 may include one or more other components. In the present disclosure, the transceiver (or transmitter 201 and / or receiver203) may be viewed as an interface circuit.

[0076] The memory 208 stores instructions used to perform operations described herein. The memory 208 may also stores data used, generated, or collected by the apparatus 310. For example, the memory 208 could store software instructions or modules configured to implement some or all of the functionality and / or embodiments described herein and that are executed by one or more processor 210.

[0077] The apparatus 310 may further include one or more input / output devices (not shown) or interfaces. The input / output devices or interfaces permit interaction with a user or other devices in the network. Each input / output device or interface includes any suitable structure for providing information to or receiving information from a user, and / or for network interface communications. Suitable structures include, for example, a speaker, microphone, keypad, keyboard, display, touch screen, etc.

[0078] The processor 210 may perform (or control the apparatus 310 to perform) operations (or methods) described herein as being performed by the apparatus 310. For example, the processor 210 performs or controls the apparatus 310 to perform receiving transport blocks (TBs) , using a resource for decoding of one of the received TBs, releasing the resource for decoding of another of the received TBs, and / or receiving configuration information configuring a resource. In detail, the operation may include those operations related to preparing a transmission for UL transmission to the apparatus 320; those operations related to processing DL transmissions received from the apparatus 320; and those operations related to processing SL transmission to and from another apparatus 310. Processing operations related to preparing a transmission for UL transmission may include operations such as encoding, modulating, transmit beamforming, and generating symbols for transmission. Processing operations related to processing DL transmissions may include operations such as receive beamforming, demodulating and decoding received symbols. Processing operations related to processing SL transmissions may include operations such as transmit / receive beamforming, modulating / demodulating and encoding / decoding symbols. Depending upon the embodiment, a DL transmission may be received by the receiver 203, possibly using receive beamforming, and the processor 210 may extract signaling from the DL transmission  (e.g. by detecting and / or decoding the signaling) . An example of signaling may be a reference signal transmitted by the apparatus 320. In some implementations, the processor 210 implements the transmit beamforming and / or the receive beamforming based on the indication of beam direction, e.g. beam angle information (BAI) , received from the apparatus 320. In some implementations, the processor 210 may perform operations relating to network access (e.g. initial access) and / or downlink synchronization, such as operations relating to detecting a synchronization sequence, decoding and obtaining the system information, etc. In some implementations, the processor 210 may perform channel estimation, e.g. using a reference signal received from the apparatus 320.

[0079] Although not illustrated, the processor 210 may form part of the transmitter 201 and / or part of the receiver 203. Although not illustrated, the memory 208 may form part of the processor 210.

[0080] The processor 210, the processing components of the transmitter 201, and the processing components of the receiver 203 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory (e.g. in the memory 208) .

[0081] The apparatus 320 includes one or more processors 260 (only one processor 260 is illustrated in the figure) . The apparatus 320 may further include at least one transmitter 252 and at least one receiver 254 coupled to one or more antennas 256. Only one antenna 256 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas 256 may alternatively be panels. The transmitter 252 and the receiver 254 may be integrated as a transceiver. The apparatus 320 may further include at least one memory 258. The apparatus 320 may further include scheduler 253. Only the transmitter 252, receiver 254, processor 260, memory 258, antenna 256 and scheduler 253 are illustrated for simplicity, but the apparatus 320 may include one or more other components. In the present disclosure, the transceiver (or transmitter 252 and / or receiver 254) may be viewed as an interface circuit.

[0082] In some implementations, the parts of the apparatus 320 may be distributed. For example, some of the modules of the apparatus 320 may be located remote from the equipment that houses the antennas 256 for the apparatus 320 (thereby also can be viewed as one of more nodes) , and may be coupled to the equipment that houses the antennas 256 over a communication link (not shown) sometimes known as front haul, such as common public radio interface (CPRI) . Therefore, in some implementations, the term apparatus 320 may also refer to nodes on the network side that perform processing operations, such as determining the location of the apparatus 310, resource allocation (scheduling) , message generation, and encoding / decoding, and that are not necessarily part of the equipment that houses the antennas 256 of the apparatus 320. The nodes may also be coupled to other apparatus 320s. In some implementations, the apparatus 320 may actually be a plurality of nodes that are operating together to serve the apparatus 310, e.g. through the use of coordinated multipoint transmissions, or the use of ORAN system as described above in the application.

[0083] The processor 260 performs operations including those related to: preparing a transmission for DL transmission to the apparatus 310, processing an UL transmission received from the apparatus 310, preparing a transmission for backhaul transmission to another apparatus 320, and processing a transmission received over backhaul from another apparatus 320. Processing operations related to preparing a transmission for DL or backhaul transmission may include operations such as encoding, modulating, precoding (e.g. multiple input multiple output (MIMO) precoding) , transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the UL or over backhaul may include operations such as receive beamforming, demodulating received symbols, and decoding received symbols. The processor 260 may also perform operations relating to network access (e.g. initial access) and / or DL synchronization, such as generating the content of synchronization signal blocks (SSBs) , generating the system information, etc. In some implementations, the processor 260 also generates an indication of beam direction, e.g. BAI, which may be scheduled for transmission by a scheduler 253 which will be described below. In some implementations, the processor 276 implements the transmit beamforming and / or receive beamforming based on beam direction information (e.g. BAI) received from another apparatus 320. The processor 260 performs other network side processing operations described herein, such as determining the location of the apparatus 310, determining where to deploy another apparatus 320, etc. In some implementations, the processor 260 may generate signaling, e.g. to configure one or more parameters of the apparatus 310 and / or one or more parameters of another apparatus 320. Any signaling generated by the processor 260 is sent by the transmitter 252. In some implementations, the apparatus 320 implements physical layer processing. In some implementations, the apparatus 320 may implement higher layer functions such as functions at the medium access control (MAC) or radio link control (RLC) layer in addition to physical layer processing. The apparatus 320 may further comprise scheduler 253 coupled to the processor 260 or integrated in the processor 260. The scheduler 253 may be included within or operated separately  from the apparatus 320a. The scheduler 253 may schedule UL, DL, SL, and / or backhaul transmissions, including issuing scheduling grants and / or configuring scheduling-free (e.g., “configured grant” ) resources.

[0084] The apparatus 320a may further includes a memory 258 storing instructions used to perform operations described herein. The memory 258 may also stores data used, generated, or collected by the apparatus 320a. For example, the memory 258 could store software instructions or modules configured to implement some or all of the functionality and / or embodiments described herein and that are executed by the processor 260.

[0085] Although not illustrated, the processor 260 may form part of the transmitter 252 and / or part of the receiver 254. Also, although not illustrated, the processor 260 may implement the scheduler 253. Although not illustrated, the memory 258 may form part of the processor 260.

[0086] The processor 260, the scheduler 253, the processing components of the transmitter 252, and the processing components of the receiver 254 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory, e.g. in the memory 258.

[0087] The apparatus 320 and / or the apparatus 310 may include other components, but these have been omitted for the sake of clarity.

[0088] Note that “signaling” , as used herein, may alternatively be called control signaling, control message, control information, or message for simplicity. Signaling between a base station (e.g., the TRP 170a-b, 172) and a UE or sensing device (e.g., ED 110) , or signaling between a different UE or sensing device (e.g., between ED 110a and ED110b) may be carried in physical layer signaling (also called as dynamic signaling) , which is transmitted in a physical layer control channel. For DL, the physical layer signaling may be known as downlink control information (DCI) which is transmitted in a physical downlink control channel (PDCCH) . For UL, the physical layer signaling may be known as uplink control information (UCI) which is transmitted in a physical uplink control channel (PUCCH) . For SL, signaling between different UEs or sensing devices (e.g., between ED 110a and ED110b) may be known as SL control information (SCI) which is transmitted in a physical sidelink control channel (PSCCH) . Signaling may be carried in a higher layer (e.g., higher than physical layer) signaling, which is transmitted in a physical layer data channel, e.g. in a physical downlink shared channel (PDSCH) for downlink signaling, in a physical uplink shared channel (PUSCH) for uplink signaling, and in a physical sidelink shared channel (PSSCH) for SL signaling. Higher layer signaling may also be called static signaling, or semi-static signaling. Higher layer signaling may be radio resource control (RRC) protocol signaling or media access control -control element (MAC-CE) signaling. Signaling may be included in a combination of physical layer signaling and higher layer signaling.

[0089] It should be noted that in present application, “information” , when different from “message” , may be carried in one single message, or be carried in more than one separate message.

[0090] FIG. 4 illustrates further aspects of an example of an apparatus 410 according to embodiments of the present disclosure. The apparatus 410 may be a communication device or an apparatus implemented in a communication device such as ED 110 or TRPs 170a-170b, 172. For example, the apparatus implemented in a communication device may be an integrated circuit, which in some contexts may be known by other colloquial names, such as chip, modem, modem chip, baseband chip, or baseband processor. In some implementations, one or more integrated circuits can be packaged into a system-on-chip, a system-in-package, or a multi-chip module. The apparatus may comprise one or more integrated circuits or comprise one or more integrated circuits and other discrete components. In some implementations, the apparatus 410 may be a module in ED 110, or apparatus 310. In some implementations, the apparatus 410 may be a module in one of TRPs 170a-170b, 172, or apparatus 320.

[0091] In an example, the apparatus 410 may include one or more processors / processor cores 411, and an interface circuit 412. The apparatus 410 may further include a memory 413. The one or more processors / processor cores 411 are configured to process signals and execute one or more communication protocols. The memory 413 is configured to store at least a part of corresponding computer program instructions and / or data. In an example, the one or more processors (or processor cores) 411 execute the computer program instructions stored in the memory 413 to implement related operations (for example, inputting, outputting, receiving, and transmitting) in the foregoing method embodiments. In some implementations, the memory 413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 413 is configured to store all of the corresponding computer program instructions and / or data for execution by the one or more processors / processor cores 411. In  some implementations, the memory 413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 413 is configured to store a part of the corresponding computer program instructions and / or data. For example, the part of the corresponding computer program instructions and / or data include computer program instructions and / or data that need to be currently executed by the one or more processors / processor cores 411. Thus, the memory 413 may store different parts of computer program instructions and / or data for a plurality times for the one or more processors (or processor cores) 411 to perform related operations in the foregoing method embodiments. As a communication interface, the interface circuit 412 is configured to implement communication with another component. For example, the interface circuit 412 may communicate a signal with other apparatus / system such as a radio frequency processing apparatus, or processor system. Optionally, to reduce a load of the processor core, a baseband signal processing circuit 414 may be also disposed to implement processing of at least a part of baseband signals, including signal demodulation, modulation, encoding, decoding, or the like.

[0092] Apparatus 410 may be processor 210 (or 260) in apparatus 310 (or 320) , in some scenario, or included in processor 210 (or 260) in apparatus 310 (or 320) in some scenario. apparatus 410 may be or include a baseband chip. In some implementations, the apparatus 410 may be independently packaged into a chip. In some implementations, the apparatus 310 (or 320) includes different types of chips. The apparatus 410 may be packaged into a processor chip (for example, a SoC chip or an SIP chip) with the different types of chips. In some implementations, the apparatus 410 may be packaged into a chip with some or all of circuits of a radio frequency processing system that may further included in the apparatus 310 (or 320) .

[0093] FIG. 5 illustrates further aspects of an example of apparatus 510, according to embodiments of the present disclosure. Apparatus 510 may include corresponding modules or units configured to implement methods and / or embodiments described herein. In some implementations, the apparatus 510 includes a processing unit 512 and a communication unit 513. Optionally, the apparatus 510 may further include a storage unit 511 configured to store apparatus program code (or instructions) and / or data.

[0094] The apparatus 510 may be an ED side apparatus, for example, an ED or a module in an ED, or a circuit or a chip responsible for a communication function in an ED. In some implementations, apparatus 510 may be implemenated as apparatus 310, accoridngly, the processing unit 512 is implemented as processor 210, the communication unit 513 is implemented as transmitter 201 and / or receiver 203, and the storage unit 511 is implementated as memory 208.

[0095] The apparatus 510 may be a base station side apparatus, for example, a base station or a module in a base station, or a circuit or a chip responsible for a communication function in a base station. In some implementations, apparatus 510 may be implemenated as apparatus 320, accoridngly, the processing unit 512 is implemented as processor 260 (the scheduler 253 may also be included) , the communication unit 513 is implemented as transmitter 252 and / or receiver 254, and the storage unit 511 is implementated as memeory 258.

[0096] In some implementations, when the apparatus 510 is an ED 110 or a module in an ED 110, a function of the apparatus 510 may be implemented by one or more processors. Specifically, the processor may include a modem chip, or a system on chip SoC chip or an SIP chip that includes a modem core. A function of the communication unit 513 may be implemented by a transceiver circuit.

[0097] In some implementations, when the apparatus 510 is a circuit or a chip that is responsible for a communication function in a ED 110, for example, a modem chip, a system on chip SoC chip or an SIP chip that includes a modem core, a function of the processing unit 512 may be implemented by a circuit system that is in the chip and that includes one or more processors or processor cores. A function of the communication unit 513 may be implemented by an interface circuit or a data transceiver circuit on the foregoing chip.

[0098] It may be understood that division into the units in the foregoing apparatus is merely logical function division. Each function may correspond to one functional unit, or two or more functions may be integrated into one functional unit. In actual implementation, all or some of the units may be integrated into one physical entity, or may be distributed in different physical entities. In addition, the foregoing functional units may be implemented in a form of hardware, may be implemented in a form of software, or may be implemented in a form of a combination of hardware and software. Whether a function is performed in a form of hardware or software depends on particular applications and design constraint conditions of the technical solutions. A person skilled in the art may use different methods to implement the described functions for each particular application, but it should not be considered that the implementation goes beyond the scope of this application.

[0099] In an example, a functional unit in any one of the foregoing apparatuses may be configured as one or more integrated circuits for implementing the foregoing methods, for example, one or more application-specific integrated circuits (application-specific integrated circuits, ASICs) , one or more central processing units (central processing units, CPUs) , one or more microprocessors (microcontroller units, MCUs) , one or more digital signal processors (digital signal processors, DSP) , one or more field programmable gate arrays (field programmable gate arrays, FPGAs) , or a combination of at least two of these integrated circuit forms.

[0100] In an example, the storage unit may include a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, and / or a register.

[0101] A processor, a processor system, a application processor, a baseband processor, a processor circuit, or a processor core may be collectively referred to as a processor. The processor may include one or a combination of a central processing unit (central processing unit, CPU) , a digital signal processor (digital signal processor, DSP) , a microprocessor (microprocessor unit, MPU) , a microcontroller (microcontroller unit, MCU) , a graphics processing unit (graphics processing unit, GPU) , a field programmable gate array (field programmable gate array, FPGA) , an artificial intelligence processor (artificial intelligence processor, AI processor) , or a neural network processing unit (neural network processing unit, NPU) .

[0102] The memory may include one or more of the following storage media: a random access memory (random access memory, RAM) , a static random access memory (static RAM, SRAM) , a dynamic random access memory (dynamic RAM, DRAM) , a phase-change memory (phase-change memory, PCM) , a resistive random access memory (resistive RAM, ReRAM) , a magnetoresistive random access memory (magnetoresistive RAM, MRAM) , a ferroelectric random access memory (ferroelectric RAM, FRAM) , a cache (cache) , a register (register) , a read-only memory (read-only memory, ROM) , a flash memory (flash memory) , an erasable programmable read-only memory (erasable programmable ROM, EPROM) , a hard disk (hard disk) , and the like. In an example, the computer program instructions used to execute the foregoing embodiments may be stored in a non-volatile memory, for example, at least a part of the memory 1060 (for example, one or more of a ROM, a flash memory, an EPROM, or a hard disk) . When the terminal runs, a part or all of corresponding computer program instructions may be loaded to a memory that has a higher transmission speed with the processor, for example, at least a part of the memory 1036 and / or the memory 10312 (for example, one or more of a RAM, an SRAM, a DRAM, a PCM, a RERAM, an MRAM, a FRAM, a cache (cache) , or a register) , so that the processor executes the computer program instructions to perform the steps in the foregoing method embodiments.

[0103] FIG. 6 illustrates a system for performing redundant inferences. The first version of coded inference illustrated in FIG. 6 deploys some redundant inference units (RIU) on additional devices. These RIUs take a linear combination of the original “systematic” inputs and are trained to output linear combinations of the original “systematic” outputs. Theoretically, given the DNN’s ability to learn from training data, an RIU is expected to output parity results that can be used to recover any lost systematic output. This method enables non-linear distributed coded inference, but comes at a cost of increased complexity at the RIU. In fact, the RIU processes more complex input and output a larger amount of inference result.

[0104] The second version of coded inference solves the abovementioned exponential training / inference complexity by imposing AI / ML-invariant transformations when generating redundant inputs. In this way, DNN of a “redundant inference job” should be able to recognize its input without additional training, and DNNs for SIU can be used for RIU without modification. In other words, the deployment of inference tasks does not need to be done in a case-by-case manner with time-consuming training. Considering that inference applications will be more specialized and diverse in future generation communication systems, this approach may save time and budget for the service providers.

[0105] The aforementioned methods may be viewed as similar to error correction coding that originally deals with binary inputs. The adaptation to inference tasks can be two-fold:

[0106] ● The inputs are now image / audio / video / point cloud, etc; and

[0107] ● The outputs are either class / label or quantity, or both.

[0108] As seen, these are no longer binary input / output that follows classic statistical assumptions. They are not independent and memoryless, which are typically taken for granted in classic channel coding.

[0109] ● Instead, the inputs for DNNs are highly correlated and dependent. For example, the following correlations are ubiquitous. With respect to spatial correlation, objects in the real world co-exist for a reason. For example, in one picture, a zebra is  more likely to co-appear with a giraffe, but less likely to co-appear with a whale. With respect to temporal correlation, events in the real world are causal in the sense that one leads to another. For example, in both audio and video clips, the events / objects observed in adjacent time frames are dependent.

[0110] If we find a way to properly exploit these correlations, the inference performance can be further enhanced. A goal of correlated coded inference can be to find a way to inherent correlation to enhance coded inference performance.

[0111] As shown in FIG. 7, the correlated coded inference framework further exploits the inherent correlation (also redundancy introduced by nature) in the source input. The two kinds of redundancies are jointly used to improve the performance of the inference algorithm. A consideration for embodiments is the ubiquitous statistical resemblance between independent data sets.

[0112] Applying Kschischang’s characterization of marginal probability, we can get a refined estimate of an object or an event. A factor graph is defined between random variables (variable nodes) and their relationships (check nodes) . In our case, a variables node is an object or an event, and the check nodes represent their joint probability.

[0113] The marginal probability can be written in the following form:

[0114] where fA (x1) and fB (x2) are the probabilities of class / object 1 and class / object 2 given standalone observation after executing any inference algorithm, and fC (x1, x2) is a joint probability of co-appearance of class / object 1 and class / object 2 given the statistic from training data before executing any inference algorithm.

[0115] The correlated coded inference framework applies to a wide range of system architectures. For example, the application can be joint sensing or detection in a wireless network. The SIUs and RIUs, as well as the correlation exploitation, are implemented in different terminal devices, user equipment (UEs) or IoT devices. The coded object recovery is implemented in a base station or an access point. As used herein, a base station or gNB can refer to a transmitter device which may receive feedback from a receiver device. A UE may refer to a receiver device which is operatively coupled to the transmitter device. The method can also be employed to fully distributed networks where all the steps are carried out in the terminal devices. The terminal devices can be fixed cameras or mobile phones, in-vehicle sensors, etc.

[0116] It has been realized that there is a focus on the design of distributed AI / ML algorithms, rather than a communication protocol to implement such algorithms.

[0117] According to embodiments there is provided a general-purpose and forward-compatible AI / ML framework for distributed inference.

[0118] According to embodiments, data content and types can be processed, but they can at least involve some physical layer data (PHY data) , which may include:

[0119] ● Extended channel state information (CSI) for massive multi-in multi-out (MIMO) ;

[0120] ● AI information (e.g. dataset, weights, gradients, inference / fine-tuning results) ;

[0121] ● Sensing information (radio frequency (RF) map, sensing results) .

[0122] According to embodiments, there is provided a network configured with devices equipped with AI / ML functions to perform distributed inference. Information that can be required to enable distributed inference among a plurality of devices, can include:

[0123] ● The input / output associated with each device and the timing associated with the input and output thereof; and

[0124] ● Network topology of the plurality of devices involved in the distributed inference.

[0125] According to embodiments, the above-mentioned configuration of the network for distributed inference can depend on at least the following factors associated with each of the particular devices associated with the network:

[0126] ● Device requests;

[0127] ● Device processing capability; and

[0128] ● Device network position and channel conditions associated therewith.

[0129] FIG. 8 illustrates an example inference scenario which is associated with various embodiments of the present disclosure. The example illustrated in FIG. 8 is of a distributed inference scenario. A difference from current inference frameworks is that the  inference on a large-scale neural network is implemented on multiple devices, rather than a single device. To do so, prompts, tokens, weights, datasets and even gradients are transmitted over the air. This is basically a tradeoff between communication and storage / computation. Advantages of this configuration can features that:

[0130] ● Each device computes part of inference task and each device stores only part of weights and internal variables. This can result in smaller storage and lower complexity than carrying out inference in a single network; and

[0131] ● To further reduce communication overhead, each device can perform inference using the same part of the neural network each time, such that it stores the same part of weights and internal variables.

[0132] The applications for a distributed inference scenario can be associated with a fleet of self-driving cars or a swarm of UAVs, for example.

[0133] FIG. 9 illustrates a distributed computing network which is associated with various embodiments of the present disclosure. In general, as illustrated in FIG. 9, a distributed network may contain:

[0134] ● A group of network nodes jointly perform an algorithm, where each node may take input from another node’s output; and

[0135] ● A computing graph determines the topology of nodes, and also includes data transfer path.

[0136] Advantageous effects of distributed inference can include:

[0137] ● Offload computation load to multiple network nodes, whereas a node can be a BS or a UE; and

[0138] ● Utilize the high-throughput, low-latency and low-power communication links in future generation communication systems.

[0139] FIG. 10 illustrates aspects of a link between network nodes, according to embodiments of the present disclosure. Several types of data may be transmitted between network nodes, for example as illustrated in FIG. 10, for distributed AI. These types of data that are transferred between network nodes can include the transfer of datasets between network nodes, the transfer of weights between network nodes and the transfer of gradients between network nodes.

[0140] ● Transfer of datasets between nodes. In this case, a node can independently train its own AI / ML algorithms based on the dataset, with an objective of minimizing the error of algorithm output with respect to the dataset. The AL / ML algorithms are proprietary and each vendor can develop its own solution;

[0141] ● Transfer of weights between nodes. In this case, a node does not train any AI / ML algorithm, and the neural network architecture and functions are determined by the node (e.g., BS) that transmits the weights to the node. The receiving node can implement the exact weights has no freedom for proprietary implementation;

[0142] ● Transfer of gradients between nodes. In this case, multiple nodes jointly train an AI / ML algorithm.

[0143] FIG. 11 illustrates example scenarios involving data transfers between nodes, which can be performed according to embodiments of the present disclosure. In the following examples illustrated in FIG. 11, node O can be a base station, and node A and node B can be user devices.

[0144] Based on the above network architecture, example embodiments include solutions for conducting distributed inference in wireless networks. The main parameters that can be considered with respect to a distributed inference can include one or more of: 1) a device reports inference capabilities thereof; 2) division of the AI / ML task into subtasks; 3) distributed inference network topologies; 4) distributed inference procedures; and 5) joint scheduling for the distributed inference.

[0145] Device Reports Inference Capability

[0146] According to embodiments, device capabilities are to be reported to a central node (such as BS) . The central node can pick or select nodes (e.g. devices) to form a distributed inference network according to device capabilities. These capabilities can include, but not limited to, the algorithm processing time and the support algorithm. Having regard to algorithm processing time and the support algorithm, capabilities associated therewith can include:

[0147] ● Algorithm processing time:

[0148] ○ Inference time;

[0149] ○ Training time;

[0150] ● Supported algorithm:

[0151] ○ Neural network parameters: maximum number of network layers, maximum width of each layer, maximum parameter size, and parameter quantization width;

[0152] ○ Neural network types: transformers, recurrent neural network  / long short-term memory (RNN / LSTM) , convolution neural network (CNN) , multilayer perceptron (MLP) , etc.;

[0153] ○ Neural network modules: embedding, positioning encoding, decoding layers, linear layers, softmax layers, etc.;

[0154] According to embodiments, the definition of device inference capabilities can be configured in a table, for example TABLE 1, wherein the reporting of the device’s inference capabilities can be referenced with respect to this predefined table. For example, the four UE processing capabilities that are presented in TABLE 1 can be indexed and reported by using {00, 01, 10, 11} . In this type of capability reporting, the device can provide the necessary information to a central node with an associated reduction of the amount of information being transmitted by the UE to the central node.

[0155] TABLE 1

[0156] Divide an AI / ML Task to Subtasks

[0157] According to embodiments, based on the collected device capabilities, a central node (such as BS) can divide a training or inference task into multiple subtasks and assign these tasks to multiple network nodes or devices (which in some instances may include the central node itself) .

[0158] ● A neural network is sliced into multiple segments, each including a subset of neurons, e.g., certain layers associated with the neural network, which can be associated with a subtask. It to be understood that;

[0159] ● Each network node undertaking a subtask can be called a worker;

[0160] ● Each worker is assigned a subset (e.g., several layers or segments that include a subset of neurons) of a neural network and associated with a particular subtask;

[0161] ● The input for a worker is the input data to its assigned subset (e.g. segment) of the neural network;

[0162] ● The output for a worker is the output data that is generated by its assigned subset (e.g. segment) of the neural network.

[0163] According to embodiments, configuration information may be sent to the network node (e.g. worker) performing the inference subtask, this configuration information can include but not limited to the following information:

[0164] ● Input layer index;

[0165] ● Input vectors and their sizes;

[0166] ● Output layer index;

[0167] ● Output vectors and their sizes;

[0168] ● Structure of the neural network layer associated to the subtask, including:

[0169] ○ Type of layers, and their sizes;

[0170] ○ Connections between the layers.

[0171] FIG. 12 illustrates an example of a decoder layer in a transformer neural network, according to embodiments of the present disclosure. As illustrated in FIG. 12, each decoder layer performs the subtasks that have been assigned thereto. For example, network node 1210 performs the subtasks associated with decoder layer n, wherein an example of these decoder layer tasks have been expanded as 1220. Input to this decoder layer 1240 can include query (Q) , key (K) and value (V) which is output from the previous decoder layer n-1. Output 1230 from the subtasks of the neural network, which associated with decoder layer n are the resulting input to the subsequent decoder layer n+1.

[0172] Distributed Inference Network Topologies

[0173] According to embodiments, based on the collected device capabilities and divided subtasks, a central node (such as BS) can configure the distributed inference network such that these devices (each defined as a worker) can work seamlessly to jointly perform a distributed inference task. For example, for each worker, configurations about its position in the inference network (which can include input, output, and which other worker (s) it receives input from, and with other worker (s) it sends output to) can be signaled to the particular device (e.g. worker) by a central node. In addition, for each worker, the algorithm to be performed (which can be in the form of weights or dataset associated with the inference subtask) can be provided to the particular worker by a central node. Furthermore, for each worker, the time and frequency resources via which the worker receives input for the inference subtask and by which the worker transmits inference subtask output can be assigned by a central node. As such, according to some embodiments, the central node provides details or configuration data to each of the workers (e.g. network nodes doing each portion of the inference) with information indicative of the particular inference subtask to be performed and the network resources by which the worker will receive input for the inference subtask and network resources by which the worker will transmit as output to a subsequent worker the results of the particular inference subtask.

[0174] FIG. 13 illustrates a ping-pong topology for application with embodiments of the present disclosure. As illustrated in FIG. 13, when a ping-pong topology is used, the two nodes directly transmit and receive from each other according to the subtask assignment and associated configuration from the central node. In this way, the complexity of performing the inference is offloaded to two nodes or workers. As illustrated, the central node transmits the respective control information (e.g. indicative of the configuration of the distributed inference network, the subtasks of the inference being performed (e.g. algorithm) and the resource assignment for receiving input for performing the inference subtask and resource assignment for transmitting output upon completion of the inference subtask. This central node transmits the node 1 control information 1320 to node 1 and the node 2 control information 1310 to node 2. As illustrated, control information traffic from the central node to the nodes is indicated by dashed lines and the data traffic between the nodes is indicated by solid line.

[0175] FIG. 14 illustrates an example resource assignment for the topology illustrated in FIG. 13. It is to be readily understood that while DCI scheduling is illustrated in FIG. 14, other forms of transmission or provision of control information for the performing the inference subtask can be used. For example, radio resource channel (RRC) signaling may be suitable or other signaling mechanism as would be readily understood. An example of DCI 1400 scheduling and resulting resource assignment is illustrated in FIG. 14, wherein transmissions between the nodes, including the central node and other nodes, can be performed via the physical downlink shared channel (PDSCH) , physical uplink shared channel (PUSCH) or the physical sidelink shared channel (PSSCH) or other suitable channel as would be readily understood. As illustrated in FIG. 14, the central node transmits the first input (① 1410) to node 1, then the two nodes (nodes 1 and node 2) exchange information (② 1420, 1440, ③ 1430, 1450) twice via sidelink communication before transmitting the final output (④ 1460) to the central node.

[0176] FIG. 15 illustrates a chain topology for application with embodiments of the present disclosure. As illustrated, the central node transmits the respective control information (e.g. indicative of the configuration of the distributed inference network, the subtasks of the inference being performed (e.g. algorithm) and the resource assignment for receiving input for performing the inference subtask and resource assignment for transmitting output upon completion of the inference subtask. In the example illustrated in FIG. 15, when a chain topology is used, a node directly receives input for the inference subtask to be performed thereby from a previous node and transmits the output of the subtask to the next node via sidelink communication according to the subtask assignment and associated configuration that was received from the central node. Initial input 1510 for inference is received at node 1 and then outputs of a subtask inference are sequentially transmitted to the subsequent node as input therefor (1520, 1530, 1540, 1550) and the final inference output 1560 from node 5 is the resulting inference output. As illustrated, control information traffic 1502 from the central node to the nodes is indicated by dashed lines and the data traffic 1504 between the nodes is indicated by solid line.

[0177] FIG. 16 illustrates an example resource assignment for the topology illustrated in FIG. 15. It is to be readily understood that while DCI scheduling is illustrated in FIG. 16, other forms of transmission or provision of control information for the performing the inference subtask can be used. For example, radio resource channel (RRC) signaling may be suitable or other signaling mechanism as would be readily understood. An example of DCI scheduling and resulting resource assignment is  illustrated in FIG. 16, wherein transmissions between the nodes, including the central node and other nodes, can be performed via the PDSCH, PUSCH or PSSCH or other suitable channel as would be readily understood.

[0178] FIG. 17 illustrates a ring topology for application to with embodiments of the present disclosure. As illustrated, the central node transmits the respective control information (e.g. indicative of the configuration of the distributed inference network, the subtasks of the inference being performed (e.g. algorithm) and the resource assignment for receiving input for performing the inference subtask and resource assignment for transmitting output upon completion of the inference subtask. In the example illustrated in FIG. 17, when a ring topology is used, a node directly receives from a previous node and transmits to the next node via sidelink communication according to the subtask assignment and associated configuration from the central node. As illustrated, control information traffic 1710 from the central node to the nodes is indicated by dashed lines and the data traffic 1720 between the nodes is indicated by solid line. Initial input 1710 for inference is received at node 1 from the central node and then outputs of a subtask inference are sequentially transmitted to the subsequent node as input therefor (1720, 1730, 1740, 1750, 1760, 1770, 1780, 1790) and the final inference output 1795 from node 1 is the resulting inference output which is transmitted to the central node. As illustrated, control information traffic 1502 from the central node to the nodes is indicated by dashed lines and the data traffic 1504 between the nodes is indicated by solid line.

[0179] FIG. 18 illustrates an example resource assignment for the topology illustrated in FIG. 17. It is to be readily understood that while DCI scheduling is illustrated in FIG. 18, other forms of transmission or provision of control information for the performing the inference subtask can be used. For example, RRC signaling may be suitable or other signaling mechanism as would be readily understood... An example of DCI scheduling and resulting resource assignment is illustrated in FIG. 18, wherein transmissions between the nodes, including the central node and other nodes, can be performed via the PDSCH, PUSCH or PSSCH or other suitable channel as would be readily understood.

[0180] FIG. 19 illustrates another example resource assignment for the topology illustrated in FIG. 17. It is to be readily understood that while DCI scheduling is illustrated in FIG. 19, other forms of transmission or provision of control information for the performing the inference subtask can be used. For example, RRC signaling may be suitable or other signaling mechanism as would be readily understood. The data flow can also circle the ring more than once (see FIG. 19) before generating the final output. As illustrated in FIG. 19, after the performance of the inference subtasks by each of the nodes, the output from node 1 is subsequently used as input for a second round of subtask inference by the nodes in the ring configuration, wherein the subsequent output from node 1 is transmitted to the central node as the final inference.

[0181] As defined in the previous examples, there is a reliance on sidelink communication in order to transfer intermediate data (e.g. output from inference subtasks) between to nodes (e.g. UEs or workers) . However, in some instances sidelink communication may not be allowed, and as such the nodes will have to configured or instructed by the control information to communicate with the central node only. In this configuration, each node would receive input for its respective inference subtask from the central node and transmit output or inference output from the performed subtask to the central node for subsequent transmission as input to the next node.

[0182] FIG. 20 illustrates a star topology for application to with embodiments of the present disclosure. As illustrated, the central node transmits the respective control information (e.g. indicative of the configuration of the distributed inference network, the subtasks of the inference being performed (e.g. algorithm) and the resource assignment for receiving input for performing the inference subtask and resource assignment for transmitting output upon completion of the inference subtask. In the example illustrated in FIG. 20, when a star topology is used, every node can receive from the central node and transmit to the central node according to the subtask assignment and associated configuration from the central node. This star-like topology can be more compatible with current standards. As illustrated, control information traffic from the central node to the nodes is indicated by dashed lines and the data traffic between the nodes is indicated by solid line.

[0183] FIG. 21 illustrates an example resource assignment for the topology illustrated in FIG. 20. It is to be readily understood that while DCI scheduling is illustrated in FIG. 21, other forms of transmission or provision of control information for the performing the inference subtask can be used. For example, RRC signaling may be suitable or other signaling mechanism as would be readily understood. An example of DCI scheduling and resulting resource assignment is illustrated in FIG. 21, wherein  transmissions between the nodes, including the central node and other nodes, can be performed via the PDSCH, PUSCH or other suitable channel as would be readily understood.

[0184] Distributed Inference Procedures

[0185] According to embodiments, in a cellular network, the procedures to deploy a distributed inference network may contain two phases, the first phase being a configuration or installation phase, and the second phase being an inference phase.

[0186] According to embodiments, the two phases, namely the configuration phase 2202 and the inference phase 2204 and examples of detailed steps associated with each phase, are illustrated in FIG. 22 and further discussed below.

[0187] According to embodiments, the configuration phase 2202 includes a plurality of steps. It is to be realized that the configuration phase is illustrated using a base station (BS) 2205, which may also be termed the central node. A BS can take on a variety of different device formats or device configurations, provided that the BS is capable of providing the desired functionality. It is to also be realized that a UE, which may also be termed the worker, can take on a variety of different device formats or device configurations, provided that the UE is capable of providing the desired functionality. The configuration phase 2202 can include:

[0188] 1. Each UE 2210 (e.g. worker) reports its respective capability 2220 to the BS 2205, wherein these capabilities can including inference time, supported algorithms and neural network parameters;

[0189] 2. Based on the AI / ML application, which may include an evaluation by the BS of the distributed inference configuration based on the UE capabilities, the BS transmits 2222 AI / ML configurations and algorithms to the UEs, including input and output layers and their sizes. It is to be understood that in some instances this transmission 2222 by the BS may be individual transmissions to each of the UEs, or may be configured as one or more broadcasts to some or all of the UEs involved in the distributed inference) ;

[0190] 3. In some embodiments, the dependencies among the UEs may also be made known to the UEs, such that each UE knows from which UE (s) it will receive input and to which UE (s) to send output upon completion of the inference subtask.

[0191] 4. BS transmits 2224 information about neural network architecture, including the weights and / or dataset to each of the UE;

[0192] 5. Each of the UEs configure 2226 themselves according to the information received from the BS, (for example, each of the UEs installs the required portion of the AI / ML associated with their particular inference subtask) and once ready to perform the inference subtask, each of the UEs sends 2228 an uplink message to the BS to indicate their respective readiness for execution of their respective inference subtask.

[0193] According to embodiments, the inference phase 2204 includes a plurality steps. It is to be understood that UE1 2211 and UE2 2212 illustrated in FIG. 22 are UEs 2210 that have performed the configuration phase 2202 as discussed above. The inference phase 2204 can include:

[0194] 1. In some embodiments, a prompt or command is triggered by either the BS or one of the UEs to start a distributed inference task;

[0195] 2. BS schedules 2240 the transmission on the time, frequency and spatial resources for performing the distributed inference task. It is to be understood that this transmission can provide the identification of particular communication resources for the reception and transmission of information to and from the UEs. In some embodiments, the scheduling can be performed individually with each of the UEs. In some embodiments, in order to improve scheduling efficiency (e.g., reduce physical downlink control channel (PDCCH) blind detection) , the scheduling for multiple UEs may be done jointly instead of separately. According to some embodiments, because the data transmissions between UEs and BS should be well coordinated to reduce overall inference latency, transmissions are quite predictable and the scheduling information can be dramatically compressed. The details associated with this are discussed in further detail elsewhere herein.

[0196] 3. According to the above configurations and scheduling, BS / UE starts data transmission and coordinated distributed inference. In some embodiments, the BS can send 2242 a prompt or trigger to the UEs for initiating the inference. In some embodiments, the prompt or trigger is sent solely to the UE that is performing the first inference subtask.

[0197] 4. The UE1 2211 performs receives input and subsequent performs 2244 the initial inference subtask. The output is subsequently transmitted 2246 to the UE2 2212 and this output from UE1 is used as input for the inference subtask to be performed by UE2. This sequence of execution of the inference subtask and subsequent transmission of results to the  next UE will continue and be defined by the configuration information (e.g. topology of the UEs performing the distributed inference) that was provided to each of the UEs during the configuration phase 2202.

[0198] Joint Scheduling for Distributed Inference

[0199] According to embodiments, multiple distributed inference tasks can be performed by one or more of the UEs at the same time. As such in order to keep track of each of the inference processes and the inference subtasks associated with each of the inference processes and the related transmissions, the RRC or DCI or other signally configuration used for the transmission of the configuration information, may include the following information.

[0200] ● Inference process ID (IID) , wherein each IID represents an independent inference task;

[0201] ● Subtask assignment index (SAI) , wherein each SAI represents an inference subtask in an inference task;

[0202] ● Time resource assignment for a transmission;

[0203] ● Frequency resource assignment for a transmission.

[0204] According to embodiments, once the IIDs and SAIs are configured, the whole inference network or topology is known to the UEs (e.g. chain topology, ring topology, star topology etc. ) . The remaining scheduling information includes the time and frequency resource assignment for each transmission among the UEs.

[0205] FIG. 23 illustrates a chain topology for distributed inference, wherein this topology is being used for three jointly scheduled inference processes, according to a embodiment of the present disclosure. As illustrated, there are three inference processes, namely IID=1, IID=2 and IID=3, the processes of which is initiated at different nodes of the chain topology. For node 1, it receives configuration information that includes the identification of the one inference process and inference subtask associated therewith 2340. As such, node 1 solely receives input 2300 for single inference process. For node 2, it receives configuration information that includes the identification of the two inference processes and respective inference subtasks associated therewith 2342. Node 2 receives input 2310 for the second inference process and the output from node 1 which is the input for the inference subtask associated with the first inference process. For node 3, it receives configuration information that includes the identification of the three inference processes and respective inference subtasks associated therewith 2344. Node 3 receives input 2320 for the third inference process and the two outputs from node 2 which are used as the input for the respective inference subtasks associated with the first inference process and the second inference subtask. For node 4, it receives configuration information that includes the identification of the three inference processes and respective inference subtasks associated therewith 2346. Node 4 receives three outputs from node 3 which are used as the input for the respective inference subtasks associated with the first inference process, second inference subtask and third inference subtask. For node 5, it receives configuration information that includes the identification of the three inference processes and respective inference subtasks associated therewith 2348. Node 5 receives three outputs from node 4 which are used as the input for the respective inference subtasks associated with the first inference process, second inference subtask and third inference subtask. The outputs from node 5 are the first inference process results 2305, second inference process results 2315 and third inference process results 2325.

[0206] It has been realized that for distributed inference tasks, data transmissions between the UEs (e.g. workers) and BS (e.g. central node) are usually well coordinated seamlessly so that there is less waiting time that will occur between data transmissions and resulting inference. For example, having regard to FIG. 22, once UE1 has finished its own inference subtask, it can immediately transmit the intermediate results to UE2 for further inference. Therefore, even at the beginning of the entire inference task, the BS should be able to determine all of the transmissions that are required to complete the overall inference task.

[0207] In other words, all transmissions are quite predictable and the scheduling information can be significantly compressed. Specifically, a BS may jointly schedule a sequence of consecutive transmissions for a group of UEs, instead of separately scheduling each transmission for each UE. This can improve scheduling efficiency by reducing the number of configuration information transmissions (e.g. DCI transmissions) and thus can reduce PDCCH blind detections. In this manner, the overall inference latency can also be reduced.

[0208] According to embodiments, in order to perform joint scheduling, a group of UEs will first be signaled by BS about their relationship and thus these UEs will know they are in the same group for distributed inference. Moreover, each UE will also know, for a given inference task, which UE will be assigned the previous subtask and which UE will be assigned the subsequent subtask.

[0209] According to embodiments, given that the inference time for each UE to process each inference subtask is known or can be determined, the time gap between a UE’s reception of the input from a previous UE and transmission of the output to a subsequent UE is therefore know. This information may also be broadcast by BS to the whole UE group that is performing the inference process) .

[0210] For example, if a UE knows a previous UE’s reception timing and the inference time, the UE will be able to derive its own timing with respect to receptions and transmissions. Therefore, in some embodiments this UE does not need to be separately scheduled or even scheduled at all. The BS schedules the first UE (e.g. worker) for each inference process, and broadcasts this scheduling information to the whole group of UEs. As such all UEs will be able to derive their own timing for receptions and transmissions based on this broadcast information in light of the fact they have already received the configuration information associated with the inference process.

[0211] However, it is to be understood that an exception to this can occur if the channel quality indicator (CQI) changes significantly. In this instance, the BS can adjusts resource allocation accordingly. In this case, the BS can signal a new DCI or other signaling format in order to update the scheduling for the group of UEs.

[0212] In the example illustrated in FIG. 24, which is a ping-pong topology as also illustrated in FIG. 13, the BS first configures the inference task by sending at least the following configuration information which relate to the inference subtask sequences and the inference  / waiting times associated with the inference process.

[0213] ● Subtask sequences:

[0214] ● SAI=1, 3, 5 for UE1

[0215] ● SAI=2, 4 for UE2

[0216] ● This implies a transmission sequence of ① 2410, ② 2415, ③ 2420, ② 2425, ③ 2430 and ④ 2440

[0217] ● Inference / waiting time:

[0218] ● Inference time 1, 2510, for UE1

[0219] ● Inference time 2, 2520, for UE2

[0220] ● Additional waiting time after SAI=2 for UE2 2610

[0221] In this example, when joint scheduling, the BS schedules the time and frequency resource assignments for the first transmission from the BS to the UE1. FIG. 25 illustrates a timing sequence when the UEs immediately transmit immediately after the particular UE completes its inference subtask. As such, in light of the configuration information transmitted to UE1, no further scheduling information would be required for the next 5 transmissions. FIG. 26 illustrates a timing sequence when a UE needs to wait for another period of time after it completes its inference subtask. In this example, the BS can further indicate this “additional waiting time” in the configuration information, for example the DCI, to indicate that a timing offset is to be applied in all subsequent transmissions after SAI=2. In this manner the UEs are capable of determining the timing associated with respect to each of the inference subtasks in relation to the inference process.

[0222] According to embodiments, there is provided a new wireless network based solution for distributed algorithms which can include, devices reporting inference capabilities, dividing an AI / ML task into a collection of subtasks, distributed inference network topologies, distributed inference procedures and joint scheduling for distributed inference.

[0223] According to embodiments, protocols are provided for applications such as distributed inference and distributed sensing, which can facilitate integration into future generation wireless communication systems or other communication networks.

[0224] According to embodiments, protocols to divide an AI / ML task to subtasks are provided, which can reduce the storage and complexity load of each single device.

[0225] According to embodiments, distributed inference network topologies are provided, which can flexibly adapt to various network topologies.

[0226] According to embodiments, distributed inference procedures are provided, which can be suitable for standard specification.

[0227] According to embodiments, joint scheduling for distributed inference is provided, which can provide efficient scheduling with less overhead and lower latency.

[0228] According to embodiments, the methods can apply to a wide range of AI / ML algorithms and applications. Although the methods discussed herein are applied to deep neural networks (DNNs) , it would be readily understood that the methods of the instant application can also apply to some AI / ML algorithms such as logistic regression and support vector machine (SVM) among others. Object detection can be envisioned as an example, as object detection contains both classification and regression. In practice, the methods disclosed herein can be tailored or generalized to other applications as well.

[0229] According to embodiments, the methods discussed herein can also apply to various communication networks, such as 5G, Wi-Fi, and future generation wireless communication systems. However, it would be readily understood that in some cases, a network is not necessary. For example, with respect to a correlated inference, the algorithm can be implemented on a single device.

[0230] Embodiments can apply to a wide range of system architectures, including cellular networks, Wi-Fi networks and ad-hoc networks.

[0231] Embodiments can be implemented in next-generation mobile and wireless network services, cloud and edge computing services, and sensing services among other services as would readily understood. The methods can be useful as long as the inference task is not implemented solely in a data center. It is contemplated that possible applications of the methods of the instant application can be associated with networked inference, environment sensing and autonomous driving, etc.

[0232] The use of embodiments may be easily detected because the correlation information and the collaborative inference procedures can be specified by standards and protocols.

[0233] According to embodiments, each unit that carries out the inference task is an inference unit, their inputs can be images, audio, video and their outputs are classification or regression results. In some embodiments, the inference units can be further divided into systematic inference units (SIU) and redundant inference units (RIU) . SIUs can work in the same way as inference units, and their inputs can be called systematic inputs. RIUs take redundant input which is generated from the systematic inputs such that the RIUs can directly perform ML / AI algorithms without additional training. The outputs of both SIUs and RIUs can be used to recover the inference results.

[0234] It is to be understood that the specific technical terms used herein are defined around distributed inference, where devices in different locations individually and collaboratively perform machine learning algorithms to accomplish inference tasks.

[0235] In the present disclosure, the terms “a” , “an” and “one” are defined to mean “at least one” , that is, these terms do not exclude a plural number of items, unless stated otherwise.

[0236] In the present disclosure, terms such as “substantially” , “generally” and “about” , which modify a value, condition or characteristic of a feature of an exemplary embodiment, should be understood to mean that the value, condition or characteristic is defined within tolerances that are acceptable for the proper operation of this exemplary embodiment for its intended application.

[0237] In the present disclosure, unless stated otherwise, the terms “connected” and “coupled” , and derivatives and variants thereof, refer herein to any structural or functional connection or coupling, either direct or indirect, between two or more elements. For example, the connection or coupling between the elements can be acoustical, mechanical, optical, electrical, thermal, logical, or any combinations thereof.

[0238] In the present disclosure, expressions such as “match” , “matching” and “matched” , including variants and derivatives thereof, are intended to refer herein to a condition in which two or more elements are either the same or within some predetermined tolerance of each other. That is, these terms are meant to encompass not only “exactly” or “identically” matching the two elements but also “substantially” , “approximately” or “subjectively” matching the two or more elements, as well as providing a higher or best match among a plurality of matching possibilities.

[0239] In the present disclosure, the expression “based on” is intended to mean “based at least partly on” , that is, this expression can mean “based solely on” or “based partially on” , and so should not be interpreted in a limited manner. More particularly, the expression “based on” could also be understood as meaning “depending on” , “representative of” , “indicative of” , “associated with” or similar expressions.

[0240] In the present disclosure, the terms "system" and "network" may be used interchangeably in embodiments of this application. "At least one" means one or more, and "a plurality of" means two or more. The term "and / or" describes an association relationship of associated objects, and indicates that three relationships may exist. For example, A and / or B may indicate the following three cases: Only A exists, both A and B exist, and only B exists, where A and B may be singular or plural. The character  " / " usually indicates an "or" relationship between associated objects. "At least one of the following items (pieces) " or a similar expression thereof indicates any combination of these items, including a single item (piece) or any combination of a plurality of items (pieces) . For example, "at least one of A, B, or C" includes A, B, C, A and B, A and C, B and C, or A, B, and C, and "at least one of A, B, and C" may also be understood as including A, B, C, A and B, A and C, B and C, or A, B, and C. In addition, unless otherwise specified, ordinal numbers such as "first" and "second" in embodiments of this application are used to distinguish between a plurality of objects, and are not used to limit a sequence, a time sequence, priorities, or importance of the plurality of objects.

[0241] A person skilled in the art should understand that embodiments of this application may be provided as a method, an appartus (or system) , computer-readable storage medium, or a computer program product. Therefore, this application may use a form of a hardware-only embodiment, a software-only embodiment, or an embodiment with a combination of software and hardware. Moreover, this application may use a form of a computer program product that is implemented on one or more computer-usable storage media (including but not limited to a disk memory, an optical memory, and the like) that include computer-usable program code.

[0242] This application is described with reference to the flowcharts and / or block diagrams of the method, the device (system) , and the computer program product according to this application. It should be understood that computer program instructions may be used to implement each process and / or each block in the flowcharts and / or the block diagrams and a combination of a process and / or a block in the flowcharts and / or the block diagrams. The computer program instructions may be provided for a general-purpose computer, a dedicated computer, an embedded processor, or a processor of another programmable data processing device to generate a machine, so that the instructions executed by the computer or the processor of the another programmable data processing device generate an apparatus for implementing a specific function in one or more procedures in the flowcharts and / or in one or more blocks in the block diagrams.

[0243] The computer program instructions may alternatively be stored in a computer-readable memory that can indicate a computer or another programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate an artifact that includes an instruction apparatus. The instruction apparatus implements a specific function in one or more procedures in the flowcharts and / or in one or more blocks in the block diagrams.

[0244] The computer program instructions may alternatively be loaded onto a computer or another programmable data processing device, so that a series of operations and steps are performed on the computer or the another programmable device, so that computer-implemented processing is generated. Therefore, the instructions executed on the computer or the another programmable device provide steps for implementing a specific function in one or more procedures in the flowcharts and / or in one or more blocks in the block diagrams.

[0245] It is clear that a person skilled in the art can make various modifications and variations to this application without departing from the scope of this application. This application is intended to cover these modifications and variations of this application provided that they fall within the scope of protection defined by the following claims and their equivalent technologies.

[0246] Acronyms, Abbreviations and Initialisms

Claims

1.A method comprising:determining, by a central node, configuration information indicative of a distributed inference task, the determining based at least in part on capability information associated with each of a plurality of UEs;transmitting, by the central node to each of the UEs, the configuration information indicative of the distributed inference task.2.The method according to claim 1, further comprising receiving, by the central node, respective capability information from each of the plurality of UEs.3.The method according to claim 1 or 2, wherein the capability information includes one or more of inference time, supported artificial intelligence or machine learning (AI / ML) algorithms, and neural network parameters.4.The method according to claim 1, wherein the determining includes determining a topology for association with the plurality of UEs for performing the distributed inference task.5.The method according to claim 4, wherein the topology is selected from a group comprising a ping-pong topology, a chain topology, a ring topology and a star topology.6.The method according to claim 4 or 5, wherein the distributed inference task is associated with an inference task divided into inference subtasks, wherein each of the inference subtasks are based at least in part on the topology and the capabilities of each of the plurality of UEs.7.The method according to claim 6, wherein for each of the plurality of UEs, the configuration information includes information indicative of both reception of input for the inference subtask associated with a particular UE and transmission of output upon completion of the inference subtask associated with the particular UE.8.The method according to claim 6, wherein the configuration information includes information indicative of multiple inference tasks being performed by the plurality of UEs.9.The method according to claim 8, wherein each of the multiple inference tasks has an inference process ID (IID) and wherein each inference subtask associated with a particular inference task has a subtask assignment index (SAI) .10.The method according to any one of claims 1 to 9, wherein upon completion of a configuration phase, the method further comprises transmitting, by the central node to each of the plurality of UEs, a schedule indicative of performance of the distributed inference task.11.The method according to claim 10, wherein each of the plurality of UEs are scheduled individually.12.The method according to claim 10, wherein the plurality of UEs are jointly scheduled.13.A method comprising:receiving configuration information indicative of a distributed inference task, the configuration information based at least in part on capability information associated with each of a plurality of UEs.14.An apparatus comprising a processor configured to cause the apparatus to perform the method of any one of claims 1 to 13.15.A computer program comprising instructions that, when executed, cause a computer to perform the method of any one of claims 1 to 13.16.A system for scheduling an inference task as a distributed inference task, the system comprising a central node configured to perform the method of any one of claims 1 to 12 and a user equipment configured to perform the method of claim 13.

Citation Information

Patent Citations

  • Reliable edge accelerated reasoning task allocation method in Internet of Vehicles environment

    CN116360996A

  • Model scheduling method, device and equipment based on GPU (Graphics Processing Unit) resources and medium

    CN116385255A

  • Unmanned aerial vehicle group cooperative reasoning method and system based on model segmentation

    CN116805195A

  • Generic Reasoner Distribution Method

    US20130262366A1