Method and apparatus for artificial intelligence inference on a user equipment

By partitioning LLMs into segments and distributing them via Base Stations, the method addresses computational and latency challenges, enabling secure and efficient AI inference on UEs, particularly for IoT devices.

WO2025251390A1PCT designated stage Publication Date: 2025-12-11HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/107485
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-04
Filing Date
2024-07-25
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

The deployment of large language models (LLMs) on resource-constrained User Equipment (UEs) faces challenges due to computational limitations and latency requirements, necessitating efficient deployment strategies that address privacy and security concerns.

Method used

A method and apparatus that utilize Base Stations (BSs) as central hubs for AI model distribution, partitioning LLMs into segments, and transmitting them to UEs over dedicated links, allowing local inference while managing resource constraints.

Benefits of technology

Enables efficient and secure AI inference on UEs, reducing latency and meeting real-time responsiveness needs for IoT devices, while protecting sensitive data through model partitioning and distributed computing techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024107485_11122025_PF_FP_ABST
    Figure CN2024107485_11122025_PF_FP_ABST
Patent Text Reader

Abstract

There is provided a method and apparatus for performing local inference on resource constrained devices. Network elements store a plurality of Artificial Intelligence (AI) models. Client devices may request AI models for performing local inference. Based on the channel conditions and the capabilities of the client device, a selected AI model is partitioned into segments at the network element and each segment is successively transmitted to the client device, such that the client device may execute one segment at a time, using an output of a segment as the input for the next segment.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR ARTIFICIAL INTELLIGENCE INFERENCE ON A USER EQUIPMENT

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The present application claims priority claims priority from U.S. Patent Application No. 63 / 656,063, filed June 4th, 2024 and incorporated herein by reference.

[0003] FIELD OF THE DISCLOSURE

[0004] The present disclosure relates to wireless communications, edge computing, and distributed computing. Specifically, the present disclosure relates to a method and apparatus for allowing Artificial Intelligence (AI) inference to be performed on a User Equipment (UE) .BACKGROUND

[0005] Large language models (LLMs) are a type of artificial intelligence (AI) model that have revolutionized the field of natural language processing (NLP) . These models are trained on massive amounts of text data, allowing them to learn complex patterns and relationships within human language. This enables LLMs to perform a wide range of tasks with remarkable proficiency, including generating human-quality text, translating languages, writing different kinds of creative content, and answering your questions in an informative way.

[0006] The power of LLMs lies in their ability to understand and process information within the context of the surrounding text. This contextual awareness allows them to generate coherent and meaningful responses, mimicking human-like communication. Unlike traditional NLP models that rely on predefined rules or limited datasets, LLMs learn directly from the data, continuously improving their abilities as they are exposed to more information.

[0007] The development of LLMs has opened doors to exciting possibilities across various industries. They are enhancing communication by powering translation tools and improving the accuracy of speech recognition systems. LLMs are also transforming information retrieval, enabling users to find relevant information quickly and efficiently. In healthcare, LLMs are assisting with tasks such as medical diagnosis and summarizing patient records.

[0008] However, the deployment of LLMs also presents challenges. Their large size and computational requirements often necessitate cloud-based processing, raising concerns about latency, privacy, security, and accessibility. Therefore, ongoing research focuses on optimizing LLMs for efficient deployment on edge devices and exploring distributed computing solutions to overcome these challenges.

[0009] The demand for utilizing LLMs is rapidly increasing, extending beyond human users to encompass a vast network of Internet-of-Things (IoT) devices. However, performing LLM inference directly on these devices is often impractical due to their limited computational capabilities. This challenge has given rise to the need for online LLM inference facilitated by advanced wireless systems, particularly in the context of the upcoming era next generation wireless networks and beyond.

[0010] Unlike human users who may tolerate some latency, IoT devices often require real-time responsiveness with millisecond-level delays. This is crucial for applications such as autonomous driving, industrial automation, and smart city infrastructure, where instantaneous decision-making is essential. As a result, a significant portion of future data traffic is expected to be driven by AI-powered tasks, particularly online LLM inference for real-time applications.

[0011] This shift towards online LLM inference has profound implications for future LLM deployment strategies. Traditional cloud-based solutions may not be able to meet the stringent latency requirements of IoT devices. This necessitates exploring alternative  approaches, such as deploying LLMs at the network edge, closer to the end devices. Additionally, efficient model partitioning and distributed computing techniques will be crucial for dividing the computational load among edge servers and user devices, ensuring scalability and low latency.

[0012] Furthermore, the massive data exchange involved in online LLM inference raises concerns about privacy and security. Developing privacy-preserving techniques, such as federated learning and differential privacy, will be essential to protect sensitive user data while enabling collaborative inference. As we move towards the era, addressing these challenges will be paramount for unlocking the full potential of online LLM inference and enabling a new generation of intelligent, interconnected applications.SUMMARY

[0013] The present disclosure proposes an approach for deploying LLMs in UEs, effectively bridging the gap between the increasing demand for AI and the limited capabilities of UEs.

[0014] According to a first aspect, there is provided a method at a network element, comprising receiving a request from a user equipment (UE) , the request comprising at least one of a UE identifier, a model identifier, UE capabilities, or channel state information (CSI) , searching for an artificial intelligence (AI) model based on the request, when the AI model is identified, establishing a dedicated link between the network element and the UE, partitioning the identified AI model into a plurality of segments, and transmitting each segment of the plurality of segments to the UE over the dedicated link.

[0015] According to the first aspect, in one possible design, the method further comprises allocating resources for transmission of the plurality of segments to the UE, and indicating the resources to the UE.

[0016] According to the first aspect, in another possible design, the method further comprises, transmitting to the UE a first indication comprising information on a selected AI model stored at the network element.

[0017] According to the first aspect, in another possible design, the method further comprises, receiving a plurality of requests from a plurality of UEs, each of the plurality of requests comprising the model identifier, said transmitting comprises multicasting each segment of the plurality of segments to each of the plurality of UEs.

[0018] According to the first aspect, in another possible design, the method further comprises receiving a second indication from at least one UE of the plurality of UEs that a segment transmission comprised an error, determining that a number of the at least one UEs is below a threshold, and continuing the multicasting of the plurality of segments to each of the plurality of UEs.

[0019] According to a second aspect, there is provided a network element comprising a processor and a communications subsystem, wherein the network element is configured to receive a request from a user equipment (UE) , the request comprising at least one of a UE identifier, a model identifier, UE capabilities, or channel state information (CSI) , search for an artificial intelligence (AI) model based on the request, when the AI model is identified, establish a dedicated link between the network element and the UE, partition the identified AI model into a plurality of segments, and transmit each segment of the plurality of segments to the UE over the dedicated link.

[0020] According to the second aspect, in one possible design, the network element is further configured to allocate resources for transmission of the plurality of segments to the UE, and indicate the resources to the UE.

[0021] According to the second aspect, in another possible design, the network element is further configured to transmit to the UE a first indication comprising information on a selected AI model stored at the network element.

[0022] According to the second aspect in another possible design, the network element is further configured to receive a plurality of requests from a plurality of UEs, each of the plurality of requests comprising the model identifier, said transmitting comprises multicasting each segment of the plurality of segments to each of the plurality of UEs.

[0023] According to the second aspect, in another possible design, the network element is further configured to receive a second indication from at least one UE of the plurality of UEs that a segment transmission comprised an error, determine that a number of the at least one UEs is below a threshold, and continue the multicasting of the plurality of segments to each of the plurality of UEs.

[0024] According to a third aspect, there is provided a computer readable medium having stored thereon executable code for execution by a processor of a network element, the executable code comprising instructions for performing any embodiment of the first aspect.

[0025] According to a fourth aspect, there is provided a method at a User Equipment (UE) comprising, transmitting a request to a network element, the request comprising a UE identifier, a model identifier, UE capabilities, and Channel State Information (CSI) , establishing a dedicated link with the network element, receiving a first segment of an Artificial Intelligence (AI) model corresponding to the model identifier over the dedicated link, creating an input token, computing the first segment with the input token and storing a first output token, receiving a subsequent segment of the AI model over the dedicated link, and repeating the steps of receiving the subsequent segment and computing the subsequent segment until all segments of the AI model are computed, wherein an output for a given segment is used as an input for a next segment.

[0026] According to the fourth aspect, in one possible design, the method further comprises receiving a first indication from the network element, the first indication indicating resources for transmissions of each segment of the AI model.

[0027] According to the fourth aspect, in another possible design, the method further comprises receiving a second indication from the network element, the second indication comprising information on a selected AI model stored at the network element.

[0028] According to the fourth aspect, in another possible design, the method further comprises discarding a segment from memory after computation of the segment.

[0029] According to the fourth aspect, in another possible design, the method further comprises discarding an output token after using it as an input for a segment.

[0030] According to the fourth aspect, in another possible design, the method further comprises transmitting to the network element an intermediate output of at least one segment of the AI model for temporary storage.

[0031] According to a fifth aspect, there is provided a User Equipment (UE) , comprising a processor and a communications subsystem, the UE being configured to transmit a request to a network element, the request comprising a UE identifier, a model identifier, UE capabilities, and Channel State Information (CSI) , establish a dedicated link with the network element, receive a first segment of an Artificial Intelligence (AI) model corresponding to the model identifier over the dedicated link, create an input token, compute the first segment with the input token and store a first output token, receive a subsequent segment of the AI model over the dedicated link, and repeat the steps of receiving the subsequent segment and computing the subsequent segment until all segments of the AI model are computed, wherein an output for a given segment is used as an input for a next segment.

[0032] According to the fifth aspect, in one possible design, the UE is further configured to receive a first indication from the network element, the first indication indicating resources for transmissions of each segment of the AI model.

[0033] According to the fifth aspect, in another possible design, the UE is further configured to receive a second indication from the network element, the second indication comprising information on a selected AI model stored at the network element.

[0034] According to the fifth aspect, in another possible design, the UE is further configured to discard a segment from memory after computation of the segment.

[0035] According to the fifth aspect, in another possible design, the UE is further configured to discard an output token after using it as an input for a segment.

[0036] According to the fifth aspect, in another possible design, the UE is further configured to transmit to the network element an intermediate output of at least one segment of the AI model for temporary storage.

[0037] According to a sixth aspect, there is provided a computer readable medium having stored thereon executable code for execution by a processor of a User Equipment (UE) , the executable code comprising instructions for performing the method of any embodiment of the fourth aspect.

[0038] According to a seventh aspect, there is provided a communication apparatus, configured to perform the method of any embodiment of the first aspect, or the method of any embodiment of the fourth aspect.

[0039] According to the seventh aspect, in one possible design, the communication apparatus comprises a receiving unit configured to receive a request from a user equipment (UE) , the request comprising at least one of a UE identifier, a model identifier, UE capabilities, or channel state information (CSI) , a processing unit configured to search for an artificial intelligence (AI) model based on the request, when the AI model is identified, establish a dedicated link between the network element and the UE, and partition the identified AI model into a plurality of segments, and a transmitting unit configured to transmit each segment of the plurality of segments to the UE over the dedicated link.

[0040] According to the seventh aspect, in another possible design, the communication apparatus comprises a transmitting unit to transmit a request to a network element, the request comprising a UE identifier, a model identifier, UE capabilities, and Channel State Information (CSI) , a receiving unit configured to receive a first segment of an Artificial Intelligence (AI) model corresponding to the model identifier over a dedicated link, and a subsequent segment of the AI model over the dedicated link, a processing unit configured to establish a dedicated link with the network element, create an input token, compute the first segment with the input token and storing a first output token, compute the subsequent segment with the first output token and store a subsequent output token, where the steps of receiving the subsequent segment and computing the subsequent segment are repeated until all segments of the AI model are computed, wherein an output for a given segment is used as an input for a next segment.

[0041] According to the seventh aspect, in another possible design, the communication apparatus comprises one or more processors configured to search for an artificial intelligence (AI) model based on the request, when the AI model is identified, establish a dedicated link between the network element and the UE, and partition the identified AI model into a plurality of segments, an interface circuit, configured to receive a request from a user equipment (UE) , the request comprising at least one of a UE identifier, a model identifier, UE capabilities, or channel state information (CSI) , and transmit each segment of the plurality of segments to the UE over the dedicated link.

[0042] According to the seventh aspect, in another possible design, the communication apparatus comprises one or more processors configured to establish a dedicated link with the network element, create an input token, compute the first segment with the input token and storing a first output token, and compute the subsequent segment with the first output token and storing a subsequent output token, and an interface circuit, configured to transmit a request to a network element, the request comprising a UE identifier, a model identifier, UE capabilities, and Channel State Information (CSI) , and receive a first segment of an Artificial Intelligence (AI) model corresponding to the model identifier over a dedicated link, and a subsequent segment of the AI model over the dedicated link, where the steps of receiving the subsequent segment and computing the subsequent segment are repeated until all segments of the AI model are computed, wherein an output for a given segment is used as an input for a next segment.

[0043] According to the seventh aspect, in another possible design, the interface circuit comprises one or more transceivers.

[0044] According to an eight aspect, there is provided an apparatus comprising one or more processors, and memories storing instructions which, when executed by the one or more processors, cause the apparatus to perform the method of any one embodiment of the first aspect or the fourth aspect.

[0045] According to a ninth aspect, there is provided a communication system comprising a first communication apparatus configured to perform the method of any embodiment of the first aspect, and a second communication apparatus configured to perform the method of any embodiment of the fourth aspect.BRIEF DESCRIPTION OF THE DRAWINGS

[0046] FIG. 1 illustrates a top-level overview of a process for deploying an AI model on a UE according to at least some implementations of the present disclosure.

[0047] FIG. 2 shows a simplified schematic illustration of a communication system according to at least some implementations of the present disclosure.

[0048] FIG. 3 further details a communication system according to at least some implementations of the present disclosure.

[0049] FIG. 4 illustrates an example of an apparatus according to at least some implementations of the present disclosure.

[0050] FIG. 5 illustrates an example of an apparatus according to at least some implementations of the present disclosure.

[0051] FIG. 6 depicts the various units or modules within an apparatus, according to at least some implementations of the present disclosure.

[0052] FIG. 7 illustrates a process for performing AI inference at a cloud data center according to at least some implementations of the present disclosure.

[0053] FIG. 8 illustrates a framework for centralized management of AI models according to at least some implementations of the present disclosure.

[0054] FIG. 9 illustrates a process for AI model announcements according to at least some implementations of the present disclosure.

[0055] FIG. 10 illustrates a process for AI model announcements according to at least some implementations of the present disclosure.

[0056] FIG. 11 illustrates a process for requesting AI inference according to at least some implementations of the present disclosure.

[0057] FIG. 12 illustrates a process for processing an AI inference request according to at least some implementations of the present disclosure.

[0058] FIG. 13 illustrates a process for processing an AI inference request according to at least some implementations of the present disclosure.

[0059] FIG. 14 illustrates a process for performing cyclical inference of model segments according to at least some implementations of the present disclosure.

[0060] FIG. 15 illustrates a method for performing local AI inference according to at least some implementations of the present disclosure.

[0061] FIG. 16 is a block diagram of a computing device that may be used for implementing the described methods according to at least some implementations of the present disclosure.DETAILED DESCRIPTION

[0062] In order to provide ubiquitous AI capabilities, it is desired to perform inference directly on the UE to address latency and privacy issues. However, the diverse AI models with increasing resource requirements, as well as the inherent limitations of the hardware for UEs, pose a significant challenge for such deployment. The proposed solution uses Base Stations (BSs) as a central hub for a diverse collection of AI models. When a UE requests an inference task, the BS allocates resources and schedules tasks to guide the UE to complete local inference. Some processes for achieving this solution include:

[0063] UE requests local inference from a BS: This request may include the UE's ID, capabilities, model requirements, and channel state information (CSI) to assist the BS in estimating the UE’s performance. To facilitate the efficient and reliable transmission  of such inference requests from the UE to the BS, future networks may provide specific physical and MAC layer protocols for transmitting such requests.

[0064] BS attempts to match the UE’s local inference request: By considering the UE's capabilities, channel conditions, and model version, and by selecting its model repository, a suitable LLM match may be found, otherwise, the BS may reject the request, prompting the UE to explore alternative solutions or wait for improved channel conditions.

[0065] BS establishes a dedicated link to support the UE’s local inference: A determination, potentially negotiated between the BS and the UE, of the optimal strategy for model partitioning and data transmission is made (model transfer or model delivery) .

[0066] By specific model partitioning, even resource-constrained UEs may effectively perform local AI inference. In this process, the BS cyclically transmits model segments to the UE, allowing the UE to conduct local inference using each segment without exceeding its computational capacity.

[0067] FIG. 1 illustrates a top-level overview of a process for deploying an AI model on a UE. UE 10 and base station 20 communicate with each other through a communication link 30. UE 10 may send a request 31 for local AI inference, comprising a UE ID, a UE capability, model requirements, and channel state information. Base station 20 may then attempt to satisfy the request by finding a suitable AI model, as illustrated by block 33, and if successful, base station 20 may then establish a dedicated link 32 for supporting local AI inference at UE 10. Otherwise, base station 20 may reject the request or provide alternative options to UE 10, as illustrated by arrow 34. When a suitable AI model is found, UE 10 may request the model from base station 20, as illustrated by arrow 35, and base station 20 may then partition the model as illustrated by block 36 and cyclically transfer segments of the AI model to UE 10, as illustrated by arrow 37, for UE 10 to perform local inference. This process shall be described in greater detail below.

[0068] Operating Environment

[0069] FIG. 2, is a schematic illustration of an example communication system according to an implementation of the present disclosure, there is shown a communication system 100 that includes a radio access network (RAN) 120, one or more communication electronic devices (EDs) 110a, 110b, 110c, 110d, 110e, 110f, 110g, 110h, 110i, 110j (collectively referred to as 110) , a core network 130, a Public Switched Telephone Network (PSTN) 140, the Internet 150, and other networks 160. The RAN 120 may include, but is not limited to, a future generation RAN, or a legacy RAN such as, but not limited to, 5th generation (5G) , 4th generation (4G) , 3rd generation (3G) or 2nd generation (2G) radio access network. The RAN 120 may be, for example, an Evolved Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access Network (E-UTRAN) , a NextGen RAN (NG RAN) , or some other type of RAN. Examples of RAN 120 based on the evolution of telecommunications standards include, but is not limited to, GSM (Global System for Mobile Communications) and CDMA (Code Division Multiple Access) for 2G, UMTS (Universal Mobile Telecommunications System) based on WCDMA (Wideband Code Division Multiple Access) and CDMA2000 for 3G, LTE (Long-Term Evolution) and WiMAX (Worldwide Interoperability for Microwave Access) for 4G, and NR (New Radio) for 5G. In some implementations, The RAN 120 may use any radio access technology (RAT) in the wireless interface between the one or more EDs 110 and the RAN 120. In some implementations, the term “radio access” may refer to the future generation air interface standards which may include both terrestrial networks (TNs) and non-terrestrial networks (NTNs) . These networks will be described in greater detail below in conjunction with various implementations. The one or more communication EDs 110 (also referred to as “user equipment” ) are configured to connect (e.g., communicatively couple) with each other or to one or more network nodes 170a, 170b (collectively referred to as 170) in the RAN 120. The core network (CN) 130 is a part of the communication system 100 and consists of network nodes (e.g., 170a, 170b) which provide support for the network features and telecommunication services. In some implementations, the CN 130 may be dependent on the RAT used in the communication system 100. In other implementations, the CN 130 may be access-agnostic, i.e., the CN 130 may be independent of the RAT used in the communication system 100. There are different types of CN 130, for different 3GPP system generations. For example, the CN 130 is the Evolved Packet Core (EPC) in 4G, also known as the Evolved Packet System (EPS) . In another example, the CN 130 is the 5G Core (5GC) which was developed as part of the 5G System (5GS) . The CN 130 also enables integration of different 3GPP and non-3GPP access types. In some implementations and referring to FIG. 2, the  CN 130 also provides the interface towards external networks that may include the PSTN 140, the Internet 150, and other networks 160 in the communication system 100.

[0070] In general, the communication system 100 facilitates interaction between multiple wireless or wired elements. The communication system 100 may transmit different types of content, such as voice, data, video, and / or text, through different transmission methods such as, but not limited to, broadcast, multicast, groupcast, and unicast. Additionally, the communication system 100 operates by allocating and / or sharing resources, such as carrier spectrum bandwidth, among its constituent elements.

[0071] The communication system 100 may provide a wide range of communication services and applications including, but not limited to, Enhanced Mobile Broadband (eMBB) services, Ultra-Reliable Low-Latency Communication (URLLC) services, Massive Machine Type Communication (mMTC) services, Integrated Sensing And Communication (ISAC) , immersive communication, Ultra-massive Machine-Type Communication (uMTC) , hyper reliable and low-latency communication, ubiquitous connectivity, integrated AI and communication, and other services that can be provided by a future generation communication system. The communication system 100 may provide other services and applications such as, but not limited to, earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery and mobility and the like.

[0072] The communication system 100 may include a terrestrial communication system (or network) and / or a non-terrestrial communication system (or network) . The communication system 100 may provide a high degree of availability and robustness through a joint operation of the terrestrial communication system and the non-terrestrial communication system. For example, integrating a non-terrestrial communication system (or components thereof) into a terrestrial communication system can result in a heterogeneous network comprising multiple layers. The heterogeneous network may achieve better overall performance through efficient multi-link joint operation, more flexible functionality sharing, and faster physical layer link switching between terrestrial networks and non-terrestrial networks. The terrestrial communication system and the non-terrestrial communication system could be considered as sub-systems of the communication system 100.

[0073] FIG. 3 illustrates another example communication system 100 according to an implementation of the present disclosure, there is shown the communication system 100 includes EDs 110a, 110b, 110c, 110d (collectively referred to as ED 110) , RANs 120a, 120b, one or more CNs 130, a PSTN 140, the Internet 150, and other networks 160. Additionally, the communication system 100 may also include a non-terrestrial network (NTN) 120c. The RANs 120a and 120b may include network nodes 170a and 170b respectively. Examples of network nodes 107a, 107b include base stations, which can be generally referred to as terrestrial network (TN) devices or terrestrial transmit and receive points (T-TRPs) 170a and 170b (collectively referred to as 170) . In this context, the terms "TRP" and "base station" are used interchangeably unless otherwise specified. For simplicity, this disclosure primarily refers to network nodes as base stations; however, unless explicitly stated otherwise, references to TRP are considered non-limiting and interchangeable. The T-TRPs 170a, 170b may be base stations mounted on a building or tower. In one implementation, the NTN 120c includes a RAN node such as a base station 172, which may be generally referred to as an NTN device, a non-terrestrial node, a non-terrestrial network device, a non-terrestrial base station, or a non-terrestrial transmit and receive point (NT-TRP) 172.

[0074] In some implementations, the NT-TRP 172 is not attached to the ground, for example, as in the case of an airborne base station. An airborne base station may be implemented using communication equipment supported or carried by a flying device. For example, a flying device may include, but is not limited to, an airborne platform (such as a blimp or an airship) , balloon, drone (such as quadcopter) , and other types of aerial vehicles. In some implementations, an airborne base station may be supported or carried by an unmanned aerial system (UAS) or an unmanned aerial vehicle (UAV) , such as a drone. An airborne base station may be a moveable or mobile base station that can be flexibly deployed in different locations to meet network demand. A satellite base station is another example of a non-terrestrial base station. A satellite base station may be implemented using communication equipment supported or carried by a satellite. A satellite base station may also be referred to as an orbiting base station. High altitude platforms are yet another example of non-terrestrial base stations, including international mobile telecommunication base stations.

[0075] As referred to herein, and unless specified otherwise, a “TRP” may also refer to a T-TRP or an NT-TRP, a “T-TRP” may also refer to a “TN TRP” , and an “NT-TRP” may also refer to an “NTN TRP” . The NTN 120c may be considered a RAN, sharing operational aspects with RANs 120a, 120b. The NTN 120c may include at least one NTN device and at least one corresponding terrestrial network device. The at least one NTN device may function as a transport layer device and the at least one corresponding terrestrial network device may function as a RAN node, communicating with the ED 110 via the NTN device. Additionally, there may be an NTN gateway on the ground (referred to as a terrestrial network device) that also functions as a transport layer device facilitating communication with both the NTN device and the RAN node. The RAN node may communicate with the ED 110 via the NTN device and the NTN gateway. In some implementations, the NTN gateway and the RAN node may be located within the same device.

[0076] A base station 170 (also referred to as a TRP as stated above) is a network element within a radio access network responsible for radio transmission and reception in one or more cells to or from the ED (such as auser equipment) . In different implementations, the base station 170 may also be known as a base transceiver station (BTS) , a radio base station, a network node, a network device, a device on the network side, a transmit / receive node, a Node B, an evolved NodeB (eNodeB or eNB) , a Home eNodeB, a next Generation NodeB (gNB) , a transmission point (TP) , a site controller, an access point (AP) , a wireless router, a relay station, a terrestrial node, a terrestrial network device, a terrestrial base station, a non-terrestrial node, a non-terrestrial network device, a non-terrestrial base station, and a positioning node, among other possibilities. The base station 170 may be a macro base station (BS) , a pico BS, a relay node, a donor node, or combinations thereof. When the base station 170 performs (or is configured to perform) a method described herein, it may be interpreted as the base station itself, one or more modules (or units) in the base station, a circuit or chip, or a combination thereof, performing the method. For example, the circuit or chip may include a modem chip, also referred to as a baseband chip, a system on chip (SoC) including a modem core, system in package (SIP) ) , and the like, and may be responsible for one or more communication functions within the base station.

[0077] The EDs 110a-110d and TRPs 170a-170b, 172 are examples of communication equipment configured to implement some or all of the operations and / or implementations described herein. The T-TRP 170a forms part of the RAN 120a, which may include other TRPs, and / or other devices. Also, the TRP 170b forms part of the RAN 120b, which may include other TRPs, and / or devices. Each TRP 170a, 170b may transmit and / or receive wireless signals within a particular geographic region or area, sometimes referred to as a “cell” or a “coverage area” . The TRPs 170a-170b may be responsible for allocating and / or configuring resources and transmission and / or reception in a set of cell (s) . A cell is a radio network object that can be uniquely identified by a cell identification that is broadcasted over a geographical region or area from base stations associated with the cell. A cell can work in either FDD or TDD mode. A cell may be further divided into cell sectors, and a base station 170a-170b may, for example, employ one or more transceivers to provide services to one or more sectors. Some implementations, may include pico or femto cells if supported by the radio access technology. In some implementations, one or more transceivers could be used for each cell, such as with Multiple-Input Multiple-Output (MIMO) technology. The number of RANs 120a-120b shown is merely an example. Any number of RANs may be contemplated when designing the communication system 100.

[0078] A base station may be a single element, as shown in the figures, or multiple elements distributed throughout the corresponding RAN, or otherwise configured. In some implementations, a plurality of RAN nodes coordinate to assist the ED 110 in implementing radio access, and different RAN nodes separately implement and handle different functions of the base station. For example, the RAN node may be a central unit (CU) , a distributed unit (DU) , a CU-control plane (CP) , a CU-user plane (UP) , or a radio unit (RU) etc. The CU and the DU may be separately deployed, or included within the same element (i.e., a baseband unit (BBU) ) . The RU may be included in a radio frequency device or a radio frequency unit (i.e., a remote radio unit (RRU) , an active antenna unit (AAU) , or a remote radio head (RRH) ) . In different systems, the CU (or the CU-CP and the CU-UP) , the DU, or the RU may be known by different names, but their functions are understood by person skilled in the art. For example, in an open radio access network (ORAN) system, a CU may be referred to as an open CU (O-CU) , a DU may be referred to as an open DU (O-DU) , and a CU-CP may be referred to as an open CU-CP (O-CU-CP) . The CU-UP may also be referred to as an open CU-UP (O-CU-UP) , and the RU may also be referred  to as an open RU (O-RU) . Any one of the CU (or the CU-CP, the CU-UP) , the DU, and the RU may be implemented using a software module, a hardware module, or a combination of a software module and a hardware module.

[0079] Furthermore, communication between different devices / apparatuses in various implementations of this disclosure may refer to direct communication (that is, without the need of forwarding by another device / apparatus) , or may refer to communication (s) between different devices / apparatuses via another device / apparatus (that is, requiring forwarding by another device / apparatus) . Alternatively, such communication (s) may involve one functional unit inside a device / apparatus using another functional unit within the device / apparatus to communicate with another device / apparatus. In other words, phrases such as "sending (or transmitting) information to. . . (an ED or a base station) " in this disclosure may be understood as a destination endpoint of the information being an ED or a base station, including, sending / transmitting information directly or indirectly to an ED or a base station. Similarly, phrases like "receiving information from. . . (an ED or a base station) " may be understood as a source endpoint of the information being an ED or a base station, including directly or indirectly receiving information from an ED or a base station. Between the source endpoint that sends the information and the destination endpoint, necessary processing such as, but not limited to, format conversion, digital-to-analog conversion, amplification, and filtering may be performed on the information. However, the destination endpoint may understand valid information from the source endpoint. A similar understanding applies to other descriptions in this disclosure without reiterating details already described. In the present disclosure, the terms "send" and "transmit" may be used interchangeably in different implementations of this disclosure.

[0080] The ED 110 is used to connect people, objects, machines, and other entities. The ED 110 may be widely used in various scenarios including, but not limited to, cellular communications, device-to-device (D2D) , vehicle to everything (V2X) , peer-to-peer (P2P) , machine-to-machine (M2M) , MTC, internet of things (IoT) , virtual reality (VR) , augmented reality (AR) , mixed reality (MR) , metaverse, digital twin, industrial control, self-driving, remote medical, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, and autonomous delivery and mobility.

[0081] Each ED 110 represents any suitable end user device for wireless operation and may include such devices (or may be referred to as, but not limited to) a user equipment (UE) or a user device or a terminal device, a wireless transmit / receive unit (WTRU) , a mobile station, a fixed or mobile subscriber unit, a cellular telephone, a station (STA) , an MTC device, a personal digital assistant (PDA) , a smartphone, a laptop, a computer, a tablet, a wireless sensor, a consumer electronics device, a smart book, a vehicle, a car, a truck, a bus, a train, or an IoT device, wearable devices (such as a watch, a pair of glasses, head mounted equipment, etc. ) , an industrial device, or an apparatus (such as a module, modem, or chip) in the forgoing devices, among other possibilities. Future generation EDs 110 may be referred to by other terms. When an ED 110 performs (or is configured to perform) a method described herein, it may be interpreted as the ED itself, one or more modules (or units) in the ED, a circuit or chip, or a combination thereof, performing the method. For example, the circuit or chip may include a modem chip, also referred to as a baseband chip, a system on chip (SoC) including a modem core, or system in package (SIP) ) , and the like, and may be responsible for one or more communication functions in the ED.

[0082] Each ED 110 connected to TRPs 170a-170b, and / or TRPs 172 can be dynamically or semi-statically turned-on (i.e., established, activated, or enabled) , turned-off (i.e., released, deactivated, or disabled) and / or configured in response to one of more of: connection availability and connection necessity.

[0083] Any ED 110 may be alternatively or additionally configured to interface, access, or communicate with any of the TRPs 170a, 170b and 172, the Internet 150, the CN 130, the PSTN 140, the other networks 160, or any combination thereof. In some examples, the ED 110a may communicate an uplink (UL) and / or downlink (DL) transmission over a terrestrial air interface 190a with station-TRP 170a. In some examples, the EDs 110a, 110b, 110c, and 110d may also communicate directly with one another via one or more sidelink (SL) air interfaces 190b. In some examples, the EDs 110a, 110d may communicate using an UL and / or DL transmission over a non-terrestrial air interface 190c with NT-TRP 172.

[0084] An air interface (such as, for example, 190a, 190b, 190c) generally includes a number of components and associated parameters that collectively specify how a transmission is to be sent and / or received over a wireless communications link between two or more communicating devices such as EDs and base station (s) . For example, an air interface may include one or more components defining the waveform (s) , frame structure (s) , multiple access scheme (s) , protocol (s) , coding scheme (s) and / or modulation scheme (s) for conveying information (such as, data) over a wireless communications link. The air interfaces 190a and 190b may use similar communication technology, that may include any suitable radio access technology.

[0085] The non-terrestrial air interface 190c can enable communication between the EDs 110a, 110d and one or more NT-TRPs 172 via a wireless link or simply a link. For some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection between a group of EDs 110 and one or more NT-TRPs 172 for multicast transmission.

[0086] The TRPs 170a-170b, 172 may communicate with one another over one or more air interfaces 190e, 190f using wireless communication links (such as radio frequency (RF) , microwave, infrared (IR) , etc. ) or wired communication links. The air interfaces 190e, 190f may utilize any suitable radio access technology, and may be substantially similar to the air interfaces 190a, 190c over which the EDs 110a-110d communicate with one or more of the TRP 170a-170b, 172 or they may be substantially different. For example, the communication system 100 may implement one or more channel access methods, such as Time Division Multiple Access (TDMA) , Frequency Division Multiple Access (FDMA) , Code Division Multiple Access (CDMA) , Single Carrier Frequency Division Multiple Access (SC-FDMA) , Low Density Signature Multicarrier Code Division Multiple Access (LDS-MC-CDMA) , Non-Orthogonal Multiple Access (NOMA) , Pattern Division Multiple Access (PDMA) , Lattice Partition Multiple Access (LPMA) , Resource Spread Multiple Access (RSMA) , and Sparse Code Multiple Access (SCMA) .

[0087] The RANs 120a and 120b are in communication with the CN 130 to provide the EDs 110a, 110b, and 110c with various services such as voice, data, multimedia, and other services. The RANs 120a and 120b and / or the CN 130 may be in direct or indirect communication with one or more other RANs (not shown) , which may or may not be directly served by the CN 130, and may employ different radio access technologies from RAN 120a and / or RAN 120b. The CN 130 may also serve as a gateway access between (i) the RANs 120a and 120b and / or the EDs 110a, 110b, and 110c, and (ii) other networks (such as the PSTN 140, the Internet 150, and the other networks 160) . In addition, some or all of the EDs 110a, 110b, and 110c may include functionality for communicating with different wireless networks over different wireless links using different wireless technologies and / or protocols. For example, the EDs 110a, 110b, and 110c communicate using different cellular communications protocols, such as, but not limited to, a Global System for Mobile Communications (GSM) protocol, a code-division multiple access (CDMA) network protocol, a Push-to-Talk (PTT) protocol, a PTT over Cellular (POC) protocol, a Universal Mobile Telecommunications System (UMTS) protocol, a 3GPP Long Term Evolution (LTE) protocol, a fifth generation (5G) protocol, a New Radio (NR) protocol, and the like. Instead of wireless communication (or in addition thereto) , the EDs 110a, 110b, and 110c may communicate using wired communication channels to a service provider or switch (not shown) , and / or to the Internet 150. The PSTN 140 may include circuit switched telephone networks for providing plain old telephone service (POTS) . The Internet 150 may include a network of computers and subnets (intranets) or both, and incorporate protocols, such as internet protocol (IP) , transmission control protocol (TCP) , user datagram protocol (UDP) . EDs 110a, 110b, and 110c may be multimode devices capable of operation according to multiple radio access technologies, and may incorporate one or multiple transceivers necessary to support such.

[0088] In addition, the communication system 100 may comprise a sensing agent (not shown) to manage the sensed data from ED 110 and / or any one of TRPs 170a, 170b, 172. In one implementation, the sensing agent may be part of any one of TRPs 170a, 170b, 172. In another implementation, the sensing agent is a separate node that can communicate with the CN 130 and / or the RAN 120 (such as any one of TRPs 170a, 170b, 172) .

[0089] FIG. 4 is a schematic illustration showing an apparatus 310 wirelessly communicating with another apparatus 320 within a communication system (e.g., the communication system 100) according to an implementation of the present disclosure. The apparatus 310 may be an electronic device (such as ED 110) . The apparatus 320 may be a network node (e, g., the network node 170) such as T- TRP 170 or an NT-TRP 172. Although only one apparatus 310, and one apparatus 320 are shown in the figure, the number of apparatus 310 and / or number of apparatus 320 can vary, potentially including one or more of each. For example, a single ED 110 may be served by a single T-TRP 170 (or a single NT-TRP 172) , or by multiple T-TRPs 170 (or multiple NT-TRPs 172) . Similarly, a single ED 110 may be served by one or more T-TRPs 170 and one or more NT-TRPs 172. Similarly, a single T-TRP 170 (or a single NT-TRP 172) may serve one or more EDs 110.

[0090] The apparatus 310 may include one or more processors 210. For clarity and to avoid overcrowding the illustration, only a single processor 210 is illustrated. The apparatus 310 may further include a transmitter 201 and a receiver 203 coupled to one or more antennas 204. For clarity, only a single antenna 204 is illustrated. One, some, or all of the antennas 204 may alternatively be panels. In some implementations, the transmitter 201 and the receiver 203 are separate from each other. In other implementations, the transmitter 201 and the receiver 203 may be integrated into a single unit, for example, as a transceiver. The transceiver is configured to modulate data or other content for transmission by the one or more antennas 204 or a network interface controller (NIC) . The transceiver may also be configured to demodulate data or other content received by the one or more antennas 204. A transceiver may include any suitable structure for generating signals for wireless or wired transmission and / or for processing signals received through wireless or wired communication. Each antenna 204 includes any suitable structure for transmitting and / or receiving wireless or wired signals. The apparatus 310 may include a memory 208. In some implementations, the apparatus 310 may include multiple memories 208. Only a single transmitter 201, receiver 203, processor 210, memory 208, and antenna 204 is illustrated for simplicity, but the apparatus 310 may include one or more other components. In some implementations of the present disclosure, the transceiver (or transmitter 201 and / or receiver 203) may be viewed as an interface circuit.

[0091] The memory 208 is configured to store instructions used to perform operations described herein. The memory 208 may also be configured to store data that is used, generated, or collected by the apparatus 310. For example, the memory 208 can store software instructions or modules configured to implement some or all of the functionalities and / or operations described herein and that which are executed by the one or more processors 210.

[0092] The apparatus 310 may further include one or more input / output devices (not shown) or interfaces. The input / output devices or interfaces facilitate interaction with a user or other devices in the network. Each input / output device or interface includes suitable components for facilitating transmission of information to a user and reception of information from a user, and for various network interface communications. Such components may include, but are not limited to, a speaker, microphone, keypad, keyboard, display, touch screen, and the like.

[0093] The processor 210 may be configured to perform (or control the apparatus 310 to perform) operations (or methods) described herein as being performed by the apparatus 310. For example, the processor 210 performs or controls the apparatus 310 to perform the operations of: a) receiving one or more transport blocks (TBs) , b) using a resource for decoding at least one of the received TBs, c) releasing the resource for decoding another of the received TBs, and / or d) receiving configuration information configuring a resource. Specifically, the operations may include tasks related to: preparing a transmission for UL transmission to the apparatus 320, processing DL transmissions received from the apparatus 320, and handling SL transmission to and from another apparatus 310. Processing operations related to preparing a transmission for UL transmission may include operations such as, but not limited to, encoding, modulating, transmit beamforming, and generating symbols for transmission. Processing operations related to processing DL transmissions may include operations such as, but not limited to, receive beamforming, demodulating and decoding received symbols. Processing operations related to processing SL transmissions may include operations such as, but not limited to, transmit / receive beamforming, modulating / demodulating and encoding / decoding symbols. Depending upon the implementation, a DL transmission may be received by the receiver 203, possibly using receive beamforming, and the processor 210 may extract signaling from the DL transmission (such as by detecting and / or decoding the signaling) . An example of signaling may be a reference signal transmitted by the apparatus 320. In some implementations, the processor 210 implements the transmit beamforming and / or the receive beamforming based on the indication of beam direction, such as beam angle information (BAI) , received from the apparatus 320. In some  implementations, the processor 210 may be configured to perform operations relating to network access (such as initial access) and / or downlink synchronization, which includes operations for detecting a synchronization sequence, decoding and obtaining the system information, and the like. In some implementations, the processor 210 may perform channel estimation, such as using a reference signal received from the apparatus 320.

[0094] Although not illustrated, in some implementations, the processor 210 may either be a part of the transmitter 201 or a part of the receiver 203 or a part of both the transmitter 201 and the receiver 203. Although not illustrated, in some implementations, the memory 208 may be a part of the processor 210.

[0095] The processor 210, along with the processing components of the transmitter 201 and the receiver 203 may each be implemented by one or more processors that may the same or different. These processors are configured to execute instructions stored in a memory (such as in the memory 208) .

[0096] The apparatus 320 includes one or more processors 260 (only one processor 260 is illustrated) . The apparatus 320 may further include one or more transmitters 252 and one or more receivers 254 coupled to one or more antennas 256. Only a single antenna 256 is illustrated to avoid clutter in the illustration. One, some, or all of the antennas 256 may alternatively be panels. In some implementations, the transmitter 252 and the receiver 254 are separate from each other. In other implementations, the transmitter 252 and the receiver 254 may be integrated into a single unit such as, for example, as a transceiver. The apparatus 320 may further include a memory 258. In some implementations, the apparatus 320 may include multiple memories 258. The apparatus 320 may further include a scheduler 253. Only a single transmitter 252, receiver 254, processor 260, memory 258, antenna 256 and scheduler 253 are illustrated for simplicity, however the apparatus 320 may include one or more other components. In the present disclosure, in some implementations, the transceiver (or transmitter 252 and / or receiver 254) may be viewed as an interface circuit.

[0097] In some implementations, various components of the apparatus 320 may be distributed. For example, some of the modules of the apparatus 320 may be located remotely from the equipment housing the antennas 256 for the apparatus 320 (and therefore also can be viewed as one or more nodes) . These modules, which can be considered as one or more nodes, may be coupled to the equipment that houses the antennas 256 over a communication link (not shown) , sometimes referred to as front haul, such as the Common Public Radio Interface (CPRI) . Therefore, in some implementations, the term apparatus 320 may also refer to network-side nodes that perform processing operations such as, but not limited to, determining the location of the apparatus 310, resource allocation (scheduling) , message generation, and encoding / decoding, and that which are not necessarily part of the equipment that houses the antennas 256 of the apparatus 320. The nodes may also be coupled to other apparatuses 320. In some implementations, the apparatus 320 may actually be a plurality of nodes that are operating together to serve the apparatus 310, such as through the use of coordinated multipoint transmissions, or through the use of ORAN system as described above in the disclosure.

[0098] The processor 260 is configured to perform operations including those related to: preparing a transmission for DL transmission to the apparatus 310, processing an UL transmission received from the apparatus 310, preparing a transmission for backhaul transmission to another apparatus 320, and processing a transmission received over backhaul from another apparatus 320. Processing operations related to preparing a transmission for DL or backhaul transmission may include operations such as, but not limited to, encoding, modulating, precoding (such as MIMO precoding) , transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the UL or over backhaul may include operations such as, but not limited to, receive beamforming, demodulating received symbols, and decoding received symbols. The processor 260 may also be configured to perform operations relating to network access (such as initial access) and / or DL synchronization, such as generating the content of synchronization signal blocks (SSBs) , generating the system information, and the like. In some implementations, the processor 260 is further configured to generate an indication of beam direction, such as BAI, which may be scheduled for transmission by the scheduler 253 which will be described below. In some implementations, the processor 260 implements the transmit beamforming and / or receive beamforming based on beam direction information (such as BAI) received from another apparatus 320. The processor 260 is configured to perform other network side processing operations described herein, such as, but not  limited to, determining the location of the apparatus 310, determining where to deploy another apparatus 320, and the like. In some implementations, the processor 260 may generate signaling data, to configure one or more parameters of the apparatus 310 and / or one or more parameters of another apparatus 320. Any signaling data generated by the processor 260 is sent by the transmitter 252. In some implementations, the apparatus 320 implements physical layer processing. In some implementations, the apparatus 320 may perform higher layer functions such as those at the Medium Access Control (MAC) or Radio Link Control (RLC) layers in addition to physical layer processing. In the apparatus 320, the scheduler 253 may be coupled to the processor 260 or integrated within the processor 260. In some implementations, the scheduler 253 may be integrated within the apparatus 320 or may be operated separately from the apparatus 320. The scheduler 253 may schedule UL, DL, SL, and / or backhaul transmissions, including issuing scheduling grants and / or configuring scheduling-free (such as “configured grant” ) resources.

[0099] The apparatus 320 may further include a memory 258 that is configured to store instructions for performing the operations described herein. The memory 258 may also store data that is used, generated, or collected by the apparatus 320. For example, the memory 258 can store software instructions or modules configured to implement some or all of the functionalities and / or implementations described herein and that which are executed by the processor 260.

[0100] Although not illustrated, the processor 260 may be implemented as part of the transmitter 252 and / or a part of the receiver 254. Although not illustrated, in some implementations, the processor 260 may implement the scheduler 253 and the memory 258 may be implemented as part of the processor 260.

[0101] The processor 260, the scheduler 253, the processing components of the transmitter 252, and the processing components of the receiver 254 may each be implemented by the same or different processors that are configured to execute instructions stored in a memory, such as in the memory 258.

[0102] The apparatus 320 and / or the apparatus 310 may include other components, not shown or described herein for the sake of clarity.

[0103] Note that the term “signaling” , as used herein, may alternatively be referred to as control signaling, control message, control information, or message for simplicity. Signaling between a base station (such as the TRP 170a, 170b, 172) and a UE or sensing device (such as ED 110) , or signaling between a different UE or sensing device (such as between ED 110a and ED 110b) may be carried in physical layer signaling (also called as dynamic signaling) , which is transmitted in a physical layer control channel. For DL, the physical layer signaling may be known as downlink control information (DCI) which is transmitted in a physical downlink control channel (PDCCH) . For UL, the physical layer signaling may be known as uplink control information (UCI) which is transmitted in a physical uplink control channel (PUCCH) . For SL, signaling between different UEs or sensing devices (such as between ED 110a and ED 110b) may be known as SL control information (SCI) which is transmitted in a physical sidelink control channel (PSCCH) . Signaling may be carried in a higher layer (such as higher than physical layer) signaling, which is transmitted in a physical layer data channel, such as in a physical downlink shared channel (PDSCH) for downlink signaling, in a physical uplink shared channel (PUSCH) for uplink signaling, and in a physical sidelink shared channel (PSSCH) for SL signaling. Higher layer signaling may also be called static signaling, or semi-static signaling. The higher layer signaling may include radio resource control (RRC) protocol signaling or media access control -control element (MAC-CE) signaling. Signaling may be included in a combination of physical layer signaling and higher layer signaling.

[0104] It should be noted that in the present disclosure, “information” , when different from “message” , may be carried within a single message, or may be carried in multiple separate messages.

[0105] FIG. 5 illustrates an example apparatus 410 according to an implementation of the present disclosure. The apparatus 410 may be a communication device or an apparatus implemented in a communication device such as the ED 110 or the TRPs 170a, 170b, 172. For example, the apparatus 410 implemented in an ED may be an integrated circuit, which in some instances may be referred to as a chip, a modem, a modem chip, a baseband chip, or a baseband processor. In some implementations, one or more integrated  circuits can be packaged into a system-on-chip, a system-in-package, or a multi-chip module. The apparatus 410 can include one or more integrated circuits and other discrete components. In some implementations, the apparatus 410 may be a module within the ED 110, or within the apparatus 310. In some implementations, the apparatus 410 may be a module within one of the TRPs 170a, 170b, 172, or the apparatus 320.

[0106] In an example, the apparatus 410 may include one or more processors 411, and an interface circuit 412. The apparatus 410 may further include a memory 413. The one or more processors 411 are configured to process signals and execute one or more communication protocols. The memory 413 is configured to store at least a part of corresponding computer program instructions and / or data. In an example, the one or more processors 411 execute the computer program instructions stored in the memory 413 to implement related operations (for example, inputting, outputting, receiving, and transmitting) in the method embodiments disclosed herein. In some implementations, the memory 413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 413 is configured to store all of the corresponding computer program instructions and / or data for execution by the one or more processors 411. In some implementations, the memory 413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 413 is configured to store a part of the corresponding computer program instructions and / or data. For example, the part of the corresponding computer program instructions and / or data may include computer program instructions and / or data that need to be currently executed by the one or more processors 411. Thus, the memory 413 may store different parts of computer program instructions and / or data for a plurality of times for the one or more processors 411 to perform related operations in the method embodiments disclosed herein. As a communication interface, the interface circuit 412 is configured to implement communication with another component. For example, the interface circuit 412 may communicate a signal with other apparatus / system such as a radio frequency processing apparatus, or processor system. The communication includes transmitting signal (or data, information) to another component or device, or receives signal from another component or device. “transmitting” includes outputting the signal to a component or device that is directly or indirectly coupled to the interface circuit (transmitting unit) . “receiving” includes inputting or obtaining a signal from a component or device that is directly or indirectly couped to the interface circuit (receiving unit) . Optionally, to reduce a load of the one or more processors, a baseband signal processing circuit 414 may be also disposed to implement processing of at least a part of baseband signals, including signal demodulation, modulation, encoding, decoding, or the like.

[0107] The apparatus 410 may be the processor 210 (or 260) within the apparatus 310 (or 320) , in some scenarios, or may be included within the processor 210 (or 260) within the apparatus 310 (or 320) in some scenarios. The apparatus 410 may be a baseband chip or may include a baseband chip. In some implementations, the apparatus 410 may be independently packaged into a chip. In some implementations, the apparatus 310 (or 320) includes different types of chips. The apparatus 410 may be packaged into a processor chip (for example, an SoC chip or an SIP chip) with the different types of chips. In some implementations, the apparatus 410 may be packaged into a chip with some or all of circuits of a radio frequency processing system that may further be included in the apparatus 310 (or 320) .

[0108] FIG. 6 illustrates example apparatus 510 according to an implementation of the present disclosure. The apparatus 510 may include corresponding modules or units configured to implement methods and / or implementations described herein. In some implementations, the apparatus 510 includes a processing unit 512 and a communication unit 513. Optionally, the apparatus 510 may further include a storage unit 511 configured to store apparatus program code (or instructions) and / or data.

[0109] The apparatus 510 may be an ED side apparatus, for example, an ED or a module in an ED, or a circuit or a chip responsible for a communication function in an ED. In some implementations, apparatus 510 may be the apparatus 310. The processing unit 512 may be the processor 210. The communication unit 513 may comprise a receiving unit and / or a transmitting unit. The receiving unit and / or the transmitting unit may be the transmitter 201 and / or the receiver 203 respectively. The storage unit 511 may be the memory 208.

[0110] The apparatus 510 may be a base station side apparatus, for example, a base station or a module in a base station, or a circuit or a chip responsible for a communication function in a base station. In some implementations, apparatus 510 may be apparatus  320. The processing unit 512 may be the processor 260 (the scheduler 253 may also be included) . The communication unit 513 may comprise a receiving unit and / or a transmitting unit. The receiving unit and / or the transmitting unit may be the transmitter 252 and / or the receiver 254 respectively. The storage unit 511 may be the memory 258.

[0111] In some implementations, when the apparatus 510 is an ED 110 or a module in an ED 110, a function of the apparatus 510 may be implemented by one or more processors. Specifically, the processor may include a modem chip, or a system on chip (SoC) chip or an SIP chip that includes a modem core. A function of the communication unit 513 may be implemented by a transceiver circuit.

[0112] In some implementations, when the apparatus 510 is a circuit or a chip that is responsible for a communication function in an ED 110, such as a modem chip, a system on chip (SoC) chip or an SIP chip that includes a modem core -a function of the processing unit 512 may be implemented by a circuit system within the chip which includes one or more processors. A function of the communication unit 513 may be implemented by an interface circuit or a data transceiver circuit on the chip.

[0113] It may be understood that the units in the apparatus 510 may be logical or functional. Each function may correspond to one functional unit, or two or more functions may be integrated into a single functional unit. In actual implementation, all or some of the units may be integrated into a single physical entity, or may be distributed across different physical entities. In addition, the functional units may be implemented in the form of hardware, software, or a combination of hardware and software. Whether a function is implemented in the form of hardware or software depends on particular applications and design constraint conditions of the technical solutions. A person skilled in the art may use different methods to implement the described functions for specific applications, but it should not be considered that the implementation goes beyond the scope of this disclosure.

[0114] In an example, a functional unit in any one of the apparatuses may be configured as one or more integrated circuits for implementing the methods disclosed herein, for example, as one or more application-specific integrated circuits (application-specific integrated circuits, ASICs) , one or more central processing units (CPUs) , one or more microprocessors or microprocessor units (MPUs) , one or more microcontrollers or microcontroller units (MCUs) , one or more digital signal processors (DSPs) , one or more field programmable gate arrays (FPGAs) , or a combination of these.

[0115] In an example, the storage unit 511 may include a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, and / or a register.

[0116] A processor may be referred to as a processor system, an application processor, a baseband processor, a processor circuit, or a processor core. The processor may include one or a combination of one or more central processing units (CPUs) , one or more digital signal processors (DSPs) , one or more microprocessors (microprocessor units, MPUs) , one or more microcontrollers (microcontroller units, MCUs) , one or more graphics processing units (GPUs) , one or more field programmable gate arrays (FPGAs) , one or more artificial intelligence processors (AI processors) , or one or more neural network processing units (NPUs) .

[0117] Memory or a storage unit may include one or more of the following storage media: a random access memory (RAM) , a static random access memory (static RAM, SRAM) , a dynamic random access memory (dynamic RAM, DRAM) , a phase-change memory (PCM) , a resistive random access memory (resistive RAM, ReRAM) , a magnetoresistive random access memory (magnetoresistive RAM, MRAM) , a ferroelectric random access memory (ferroelectric RAM, FRAM) , a cache, a register, a read-only memory (ROM) , a flash memory (flash memory) , an erasable programmable read-only memory (erasable programmable ROM, EPROM) , a hard disk, and the like. In an example, computer program instructions used to execute embodiments may be stored in a non-volatile memory, for example, at least a part of a memory or storage unit (for example, one or more of a ROM, a flash memory, an EPROM, or a hard disk) . When a terminal runs, a part or all of corresponding computer program instructions may be loaded to a memory that has a higher transmission speed with the processor, for example, at least a part of a memory or a storage unit (for example, one or more of a RAM, an SRAM, a DRAM, a PCM, a RERAM, an MRAM, a FRAM, a cache, or a register) , so that the processor executes the computer program instructions to perform the steps in the method embodiments disclosed herein.

[0118] In the present disclosure, the terms “a” or “an” are defined to mean “at least one” , that is, these terms do not exclude a plural number of items, unless stated otherwise.

[0119] In the present disclosure, terms such as “substantially” , “generally” and “about” , which modify a value, condition or characteristic of a feature of an example embodiment, should be understood to mean that the value, condition or characteristic is defined within tolerances that are acceptable for the proper operation of the example embodiment for its intended application.

[0120] In the present disclosure, unless stated otherwise, the terms “connected” and “coupled” , and derivatives and variants thereof, refer herein to any structural or functional connection or coupling, either direct or indirect, between two or more elements. For example, the connection or coupling between the elements can be acoustical, mechanical, optical, electrical, thermal, logical, or any combinations thereof.

[0121] In the present disclosure, expressions such as “match” , “matching” and “matched” , including variants and derivatives thereof, are intended to refer herein to a condition in which two or more elements are either the same or within some predetermined tolerance of each other. That is, these terms are meant to encompass not only “exactly” or “identically” matching the two elements but also “substantially” , “approximately” or “subjectively” matching the two or more elements, as well as providing a higher or best match among a plurality of matching possibilities.

[0122] In the present disclosure, the expression “based on” is intended to mean “based at least partly on” , that is, this expression can mean “based solely on” or “based partially on” , and so should not be interpreted in a limited manner. More particularly, the expression “based on” could also be understood as meaning “depending on” , “representative of” , “indicative of” , “associated with” or similar expressions.

[0123] In the present disclosure, the terms "system" and "network" may be used interchangeably in different embodiments of this application. "At least one" means one or more, and "a plurality of" means two or more. The term "and / or" describes an association relationship of associated objects, and indicates that three relationships may exist. For example, A and / or B may indicate the following three cases: Only A exists, both A and B exist, and only B exists, where A and B may be singular or plural. The character " / " indicates an "or" relationship between associated objects. "At least one of the following items (pieces) " or a similar expression thereof indicates any combination of these items, including a single item (piece) or any combination of a plurality of items (pieces) . For example, "at least one of A, B, or C" includes: only A; only B; only C; A and B; A and C; B and C; or A, B, and C, and "at least one of A, B, and C" may also be understood as including: only A; only B; only C; A and B; A and C; B and C; or A, B, and C. In addition, unless otherwise specified, ordinal numbers such as "first" and "second" in embodiments of this application are used to distinguish between a plurality of objects, and are not used to limit a sequence, a time sequence, priorities, or importance of the plurality of objects.

[0124] A person skilled in the art should understand that embodiments of this application may be provided as a method, an apparatus (or system) , computer-readable storage medium, or a computer program product. Therefore, this application may use a form of a hardware-only embodiment, a software-only embodiment, or an embodiment with a combination of software and hardware. Moreover, this application may use a form of a computer program product that is implemented on one or more computer-usable storage media (including but not limited to a disk memory, an optical memory, and the like) that include computer-usable program code.

[0125] This application is described with reference to the flowcharts and / or block diagrams of the method, the device (system) , and the computer program product according to this application. It should be understood that computer program instructions may be used to implement each process and / or each block in the flowcharts and / or the block diagrams and a combination of a process and / or a block in the flowcharts and / or the block diagrams. The computer program instructions may be provided for a general-purpose computer, a dedicated computer, an embedded processor, or a processor of another programmable data processing device and enable a machine to execute the instructions. When executed by any computer or the processor of a programmable data processing device, the instructions cause the apparatus to implement specific functions as described in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams. The computer program instructions may alternatively be stored in a computer-readable memory that can indicate  a computer or another programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate an artifact that includes an instruction apparatus. The instruction apparatus implements a specific function in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams.

[0126] The computer program instructions may alternatively be loaded onto a computer or another programmable data processing device, so that a series of operations and steps are performed on the computer or another programmable device, so that computer-implemented processing is generated. Therefore, the instructions executed on the computer or on another programmable device provide steps for implementing specific functions as described in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams.

[0127] It is clear that a person skilled in the art can make various modifications and variations to this application without departing from the scope of this disclosure. This disclosure is intended to cover these modifications and variations of this application provided that they fall within the scope of protection defined by the following claims and their equivalent technologies.

[0128] Current Cloud-based LLM Deployment

[0129] The current paradigm for commercial LLM inference relies heavily on remote data centers equipped with powerful GPUs. Users, whether through chat interfaces or APIs, send their input data over the internet to these data centers. The LLMs process the information and return the results, again via the internet. While this approach offers scalability and centralized management, it suffers from several drawbacks.

[0130] One major concern is latency. The round-trip communication between users and distant data centers introduces significant delays, hindering real-time applications that demand immediate responses. This is particularly problematic for time-sensitive tasks such as autonomous driving or industrial control, where milliseconds matter. Additionally, transmitting sensitive data over long distances raises privacy and security risks, making the system vulnerable to potential breaches and unauthorized access.

[0131] To address the latency issue, some propose deploying dedicated GPUs for inference at the base station level, closer to the users. While this approach reduces communication delays, it introduces new challenges. Packing a large number of GPUs at base stations significantly increases energy consumption, not only for computation but also for cooling purposes. This leads to higher operational costs and environmental concerns. Moreover, concentrating expensive GPU resources at base stations raises infrastructure costs and creates potential bottlenecks, as the demand for LLM inference continues to grow.

[0132] The Growing Demand for AI on Resource-Constrained Devices

[0133] The desire for ubiquitous AI capabilities is rapidly expanding beyond human users to encompass a vast network of user devices, primarily composed of non-human entities like IoT sensors and actuators. These devices, often with limited processing power and storage, seek to leverage AI for tasks ranging from environmental monitoring and industrial automation to smart infrastructure management. However, a significant challenge arises due to the inherent limitations of their hardware.

[0134] Adding to the complexity, the need for LLM inference on user devices is often sporadic, not continuous. Furthermore, the specific LLM or other large model required may vary depending on the task at hand, necessitating the ability to update and switch between different models. This dynamic demand creates additional challenges for deploying LLMs directly on user devices, especially considering their limited storage capacity. Furthermore, equipping these devices with powerful GPUs, which are often not fully utilized, may not be a cost-effective or energy-efficient solution.

[0135] This challenge is further amplified by the rapid advancement of LLMs. Despite continuous improvements in device hardware, the sheer size and complexity of LLMs, often requiring tens of gigabytes of storage, make it impractical for them to handle LLM inference directly. This is especially true for older generation devices with even more constrained resources. As a result, alternative solutions are needed to bridge the gap between the growing demand for AI and the limited capabilities of user devices.

[0136] Resource-constrained Device Requesting Local Inference from BS

[0137] The future landscape envisions a massive influx of ubiquitous IoT devices seamlessly integrated into our environment. These devices will rely on wireless access to tap into the power of AI, as hardware and power constraints prevent them from hosting dedicated, high-performance AI chips for inference tasks. However, the demand for AI functionality on these devices is often sporadic, not continuous, requiring an adaptable solution.

[0138] Adding to the complexity is the rapid evolution of AI models, especially LLMs and other deep learning architectures, which makes it impractical to fixate on a specific model architecture with predetermined layers and neuron counts. An ideal solution will be future-proof to accommodate upcoming advancements and efficiently support new AI architectures. The immense size of LLMs presents a significant hurdle, as even standard GPUs struggle to accommodate their storage requirements. The sheer size of these models poses a challenge for the limited memory capacity of IoT devices.

[0139] Cloud-based AI models are illustrated with respect to FIG. 7. Specifically, as seen in FIG. 7, a resource-constrained device 701 may send a query to a cloud data center 702 over a network 703. The cloud data center 702 may then process the query through an LLM, and provide the results back to the resource-constrained device 701 over network 703.

[0140] While utilizing cloud-based AI models offers scalability and efficiency, performing inference directly on the user equipment (UE) provides a crucial advantage: enhanced data privacy and security. By keeping sensitive data localized within the device and processing it directly on-device, the risk of data leakage or unauthorized access during transmission or processing is significantly mitigated. This is particularly important for applications dealing with personal or confidential information, where maintaining data privacy is paramount. Local inference empowers users to leverage the power of AI while retaining control over their data, fostering trust and confidence in AI-driven services.

[0141] According to at least some embodiments of the present disclosure, a network comprises base stations acting as a central hub for a diverse collection of AI models, each designed for specific functionalities and applications. These pre-trained and validated models are strategically distributed to base stations within the core network, bringing AI capabilities closer to the end users.

[0142] This approach allows for centralized management and security of the models, ensuring they are certified, registered, and adhere to necessary regulations, as illustrated in FIG. 8. Each model 801 is provided with metadata 802 that may comprise a unique identifier and a detailed description outlining its characteristics, including size, parameters, and resource requirements. This information enables efficient resource allocation and task scheduling within the BS network. Finally, multiple versions 803 of a model can be made available, each representing different levels of optimization through techniques like quantization and pruning. This allows for flexibility in choosing the most suitable model version based on the specific capabilities and constraints of the user device requesting AI functionality. This model information is provided to core network 804, and in some embodiments, to regulatory bodies.

[0143] According to at least some embodiments, the base station periodically broadcasts information about the available AI models it can serve to a UE or any resource-constrained IoT device. This broadcast could be incorporated into existing system information messages or dedicated AI model announcement signals within the physical and MAC layers. The information may include the model ID, version, size, resource requirements, and other relevant details. To efficiently broadcast available AI model information from the BS to UE, a network may consider incorporating this information into existing system information messages or introducing dedicated AI model announcement signals.

[0144] In at least some embodiments, as illustrated in FIG. 9, the PHY (physical) layer may utilize existing physical channels dedicated for broadcasting system information. This could involve time-division multiplexing (TDM) or frequency-division multiplexing (FDM) schemes to allocate specific slots or sub-carriers for AI model announcements. The MAC layer may extend existing system information messages to include fields for AI model details such as model ID, version, size, and performance characteristics, and then may schedule the AI model announcements periodically within the system information broadcast cycle, ensuring timely updates for the UEs.

[0145] Specifically, as seen in FIG. 9, a base station 902 may allocate TDM or FDM resources for broadcasting system information, as illustrated by block 910. Base station 902 may then prepare information for a new AI model to be announced, and populate an existing system information message with this information, as illustrated by block 920. This information may include, without limitation, model ID, version, size, and performance characteristics. The new AI model information may then be provided in a message from base station 902 to UE 901 on the allocated resource, as illustrated by arrow 930. In at least some implementations, this information may be provided in a new medium access layer (MAC) message format. UE 902 may then decode message 930 to obtain information on the new model, as illustrated by block 940. The process of FIG. 9 may be performed at regular intervals, or when a new AI model becomes available.

[0146] In at least some other embodiments, as illustrated in FIG. 10, the PHY layer may allocate specific physical channels solely for AI model announcements. This could involve utilizing unused spectrum or dynamically sharing resources with other control channels based on demand. And it may also implement synchronization mechanisms to ensure UEs can efficiently detect and decode the AI model announcement signals. The MAC layer may introduce a new MAC message type specifically for AI model announcements. This allows for flexibility in designing the message format and content.

[0147] Specifically, as seen in FIG. 10, a base station 1002 may allocate resources for broadcasting a dedicated AI model announcement message, as illustrated by block 1010. Base station 1002 may then prepare information for a new AI model to be announced, and populate a dedicated AI model announcement message with this information, as illustrated by block 1020. This information may include, without limitation, model ID, version, size, and performance characteristics. The new AI model information may then be provided in a message from base station 1002 to UE 1001 on the allocated resource, as illustrated by arrow 1030. In at least some implementations, this information may be provided in a new MAC message format. UE 1002 may then decode message 1030 to obtain information on the new model, as illustrated by block 1040. In at least some implementations, this information may be provided in a new MAC message format. The process of FIG. 10 may be performed at regular intervals, or when a new AI model becomes available.

[0148] When a UE requires AI inference, it initiates a request to the BS, outlining its capabilities (processing power and storage) , desired inference performance level (considering different model versions and their trade-offs) , and acceptable latency. This request may further include the UE's ID, channel state information (CSI) for assessing the current link quality, and potentially other relevant parameters. This information is important as it anticipates the substantial data exchange required for model transfer and inference execution.

[0149] To enable efficient and reliable communication of AI inference requests from the UE to the BS, networks may provide specific physical and MAC layer protocols. The PHY layer may allocate dedicated physical channels for UE requests to avoid contention with other data traffic. This could involve dynamically sharing resources based on demand. Alternatively, the PHY layer may implement a random-access scheme where UEs contend for channel access using techniques like slotted ALOHA. This approach is more suitable for sporadic request patterns. Accordingly, a network may introduce a new MAC control message type specifically designed for AI inference requests. This message would include fields for:

[0150] ● UE ID: Unique identifier of the requesting UE;

[0151] ● Capabilities: Processing power, storage capacity, and supported operations of the UE;

[0152] ● Model Requirements: Desired model ID, version preferences, and acceptable performance trade-off;

[0153] ● Latency Constraints: Maximum tolerable delay for inference execution;

[0154] ● Channel State Information (CSI) : Feedback on the current channel quality between the UE and BS.

[0155] The MAC layer may perform grant-based scheduling. If dedicated channels are used, the base station can schedule request transmissions from UEs based on factors such as priority, channel conditions, and current network load; or it may execute  random access with collision resolution by implementing efficient collision resolution mechanisms to handle situations where multiple UEs attempt to send requests simultaneously.

[0156] Reference is now made to FIG. 11 which illustrates a dedicated channel and protocol for UEs to send AI inference requests. Specifically, as seen in FIG. 11, UE 1101 may utilize physical channels for AI inference requests, via dynamic resource sharing or by using a scheme like ALOHA, as illustrated by block 1110. UE 1101 may then prepare a request for AI inference, as illustrated by block 1120. Specifically, this request may comprise parameters including, but not limited to, a UE identifier, UE capabilities, such as processing power, storage capacity, and supported operations, model performance level, latency constraints, and channel conditions such as a CSI for the channel between UE 1101 and base station 1102. UE 1101 may then transmit the AI inference request to base station 1102 on a dedicated physical channel, as illustrated by arrow 1130.

[0157] BS matching the UE’s local inference request

[0158] Upon receiving a request from the UE for a local AI inference, the BS attempts to match the UE's requirements with its available model repository. If a suitable match is found, considering both the UE's capabilities and the channel conditions, the process proceeds to the next stage. However, if no appropriate model is available or the channel quality is insufficient to support the required data transfer, the BS may reject the request, prompting the UE to explore alternative solutions or wait for improved conditions. This communication exchange could involve dedicated signaling within the MAC layer or utilize existing control channels with extensions to accommodate AI-specific information.

[0159] For instance, after receiving an AI inference request from a UE, the BS undertakes a series of steps to determine the feasibility of fulfilling the request.

[0160] Step 1 -Model Matching: The BS maintains a database of available AI models, each with associated metadata including:

[0161] ● Model ID and version;

[0162] ● Size and computational requirements;

[0163] ● Performance characteristics (for example, accuracy and latency) ;

[0164] ● Supported operations and data types;

[0165] ● Security and trust information, such as model provenance, integrity verification data, or certifications.

[0166] The BS runs a matching algorithm that compares the UE's capabilities and request parameters with the available models. This algorithm considers factors such as:

[0167] ● UE capabilities (for example, processing power, storage capacity, and supported operations) ;

[0168] ● Model requirements (for example, desired model ID, version preferences, and performance needs) ;

[0169] ● Latency constraints: Maximum acceptable delay for inference execution;

[0170] ● Channel conditions: Current link quality based on the received CSI;

[0171] ● Security requirements of the UE and the requested application.

[0172] Step 2 -Decision Making and Response: If a suitable model is found that satisfies the UE's requirements and the channel conditions are adequate for data transfer, the BS proceeds to the next stage, which may involve model optimization, partitioning, and transmission scheduling.

[0173] If no appropriate model is available or the channel quality is deemed insufficient to support the required data transfer within the latency constraints, the BS sends a rejection message to the UE. This message could include:

[0174] ● A reason for the rejection (for example, no matching model, poor channel conditions) ;

[0175] ● Alternative options (for example suggesting a different model version, requesting the UE to try again later) .

[0176] A network may implement new MAC control messages for conveying the matching result and any additional information (e.g., reasons for rejection, alternative options) . It may designate specific channels or time slots for this communication to  ensure timely and reliable delivery by, for example, extending existing 5G NR MAC control messages (e.g., resource allocation messages) to include fields for AI-specific information like model matching results and rejection reasons. This approach reduces overhead but may require careful design to ensure backward compatibility and efficient utilization of the existing control channels.

[0177] According to some embodiments, some strategy to balance load is taken by distributing the model matching workload across multiple base station processing units or edge servers to prevent bottlenecks and ensure scalability; by caching frequently requested models at the base station to expedite future requests and reduce response times; and by dynamically updating model repository with new models or updated versions to offer the latest AI capabilities to UEs.

[0178] BS Establishes Dedicated Link to Support the UE’s Local Inference

[0179] Once the BS successfully matches a suitable AI model with the UE’s request and capabilities, both parties engage in a more granular negotiation phase. This involves determining the optimal strategy for model partitioning and data transmission (model transfer or model delivery) .

[0180] In at least some embodiments, the BS and the UE decide on the granularity of model partitioning. This could involve dividing the model into layers, with options such as transmitting one layer per transmission or grouping multiple layers together. If a single layer exceeds the UE's capacity, further division might be necessary, potentially splitting complex layers like transformers based on headers or other internal structures. A resulting piece of partition is called as segment.

[0181] Next, the BS and UE may agree upon the transmission schedule and resource allocation. This includes determining the time slot or frequency resources for transmitting each model partition (segments) . Options include utilizing specific transmission time intervals (TTIs) , allocating a certain number of OFDM symbols within a TTI, or even dedicating entire frames or superframes for model transfer depending on the size and urgency of the request. Additionally, the BS and UE may negotiate strategies for dynamically adapting the transmission schedule and resource allocation based on changes in channel conditions or UE availability.

[0182] Then, based on the current channel conditions and feedback from the UE (e.g., CSI) , both parties select an appropriate modulation and coding scheme (MCS) and transport block (TB) size (s) to ensure reliable and efficient data transmission. This adaptive approach ensures optimal performance despite potential fluctuations in the wireless channel.

[0183] FIG. 12 and FIG. 13 illustrate the connection process between UE and BS when UE requests local AI inference from a base station.

[0184] Specifically, in FIG. 12, the base station permits the connection and establishes a dedicated link to UE by identifying a suitable match that meets the UE's requirements. Subsequently, a negotiation regarding model partitioning and data transmission strategy takes place, leading to an agreed resource allocation. Furthermore, based on the current CSI feedback from UE, the base station may configure parameters like channel coding and modulation schemes to guarantee effective data communication between UE and BS.

[0185] As seen in FIG. 12, UE 1201 may send an AI inference request to base station 1202, as illustrated by arrow 1210. As discussed above, the AI inference request 1210 may include parameters for the AI inference. Base station 1202 may then try to match the request, including the request parameters, with a model stored in its model repository, as illustrated by block 1220. Upon finding a match, base station 1202 may then allocate resources for transmission of the model to UE 1201, as illustrated by arrow 1230. UE 1201 may then provide base station 1202 with channel conditions, e.g., current CSI, as illustrated by arrow 1240, and base station 1202 may use the channel conditions to determine an appropriate MCS and TB size for transmission of the model, as illustrated by arrow 1250.

[0186] In FIG. 13, the BS initially rejects the connection due to a mismatched model version. Upon receiving the suggested model version from the BS, the UE adjusts and resends the inference requests. Once a suitable match is found, the connection is established, allowing the UE to synchronize with the BS to ensure alignment in communication frequency and timing.

[0187] As seen in FIG. 13, UE 1301 may send an AI inference request to base station 1302, as illustrated by arrow 1310. As discussed above, the AI inference request may include parameters for the AI inference. Base station 1302 may then try to match the  request, including the request parameters, with a model stored in its model repository, as illustrated by block 1320. In the implementation illustrated in FIG. 13, base station 1302 rejects request 1310 as no suitable match was found, and provides UE 1301 with a model suggestion, as illustrated by arrow 1330. UE 1301 may then respond to model suggestion 1330 with a new request 1340. In some implementations, the new request 1340 may be based on model suggestion 1330.

[0188] In the implementation illustrated in FIG. 13, base station 1302 accepts new request 1340 and allocate resources for transmission of the model to UE 1301, as illustrated by arrow 1350. UE 1301 may then provide base station 1302 with channel conditions for example current CSI, as illustrated by arrow 1360, and base station 1302 may use the channel conditions to determine an appropriate MCS and TB size for transmission of the model, as illustrated by arrow 1370.

[0189] The PHY layer may allocate dedicated physical channels or sub-carriers or sub-bands for model transfer to ensure consistent performance and avoid interference with other data traffic.

[0190] The MAC layer may further divide the model partitions into smaller data units suitable for transmission within the MAC layer payload, which contains MAC headers to include information such as:

[0191] ● Model ID and partition (segment) number, to identify the specific model and its corresponding segment;

[0192] ● Sequence number, to ensure correct ordering and reassembly of data units at the UE;

[0193] ● Error detection codes, to detect and correct transmission errors.

[0194] The BS grants transmission opportunities to the UE for each data unit based on the agreed-upon schedule and resource allocation. The BS may maintain separate queues for model transfer data units to prioritize them over other data traffic. The BS may also implement mechanisms for the UE to acknowledge successful reception of data units and provide feedback on channel quality to enable adaptive MCS selection at the BS.

[0195] All these designs are to control congestion to prevent network overload and ensure fair resource allocation among multiple UEs requesting model transfers, to protect the confidentiality and integrity of the transmitted model data, and to ensure accurate time and frequency synchronization between the BS and UE for efficient data transmission and reception.

[0196] Privacy-Preserving Inference through Segmented Execution

[0197] To enable efficient and privacy-preserving AI inference on resource-constrained devices, the inference process is divided into segments, each corresponding to a portion of the complete AI model. Upon receiving segment- (i) , the UE performs local inference using this specific portion of the model. This process, illustrated in FIG. 14, ensures that sensitive user data remains confined within the device, mitigating privacy concerns.

[0198] As seen in FIG. 14, base station 1402 may provide UE 1401 with a first segment of an AI model, as illustrated by arrow 1410. Upon receiving the first segment, UE 1401 may perform inference on the first segment, as illustrated by block 1411. In at least some implementations, UE 1401 may create an input token to be submitted as an input for the first segment. When the inference on the first segment is completed, UE 1401 may store an output for the first segment as illustrated by block 1412. This step may also involve caching data specific to the first segment. UE 1401 may also discard parameters for the first segment, as illustrated by block 1413. This helps free memory on UE 1401 for processing subsequent segments.

[0199] Base station 1402 may then provide UE 1401 with the next segment, as illustrated by arrow 1420. Upon receiving the second segment, UE 1401 may perform inference on the second segment, as illustrated by block 1421. The input for the next segment is the output of the previous segment, and the output stored previously at block 1412 may be provided to the second segment as an input. When inference on the second segment is completed, UE 1401 may store the output of the second segment, as illustrated by block 1422. As in the case of the first segment, this may involve caching data specific to the second segment. UE 1401 may also discard parameters for the second segment, as well as the output of the first segment and any cached data that is specific to the first segment, as illustrated by block 1423.

[0200] Optionally, after performing inference on a segment, UE 1401 may provide base station 1402 with cached data and intermediate outputs, as illustrated with arrow 1430. This helps free storage space at UE 1401, and is particularly suited to devices with low storage capacity. Base station 1402, upon receiving the cached data and intermediate outputs, may perform dynamic cache management on behalf of UE 1401, and provide back to UE 1401, at an opportune time, the cached data and intermediate outputs, as illustrated by arrow 1440.

[0201] Generally, once the inference on segment- (i) is complete, the UE may discard the model parameters associated with that segment to free up storage space. However, the UE retains the inference output and any related cache data specific to segment- (i) . This cached output plays a crucial role as it serves as the input for the subsequent segment, segment- (i+1) , when it arrives. After processing segment- (i+1) and generating its output, the UE can safely discard the inference output from segment- (i) , further optimizing storage utilization.

[0202] The BS transmits the model segments in a cyclical manner, guaranteeing that each segment eventually reaches the UE. Upon receiving segment- (i) for the second time, which may have updated weight parameters compared to the previous transmission of segment- (i) (for example, segment- (i) may have similar functionality but different weight parameters compared to the previous segment-(i) ) , the UE leverages the previously stored cache, including the intermediate outputs and layer-specific information, to the inference process. This cyclical approach and efficient caching mechanism ensure that the UE only needs to execute inference on one segment at a time, consuming storage resources equivalent to a single segment.

[0203] Specifically, for storage-constrained UEs, the relevant cached data and intermediate outputs for a specific segment- (i) can be sent to the BS upon completion of segment- (i) . At this point, the UE can discard all parameters and data associated with that segment to free up storage space. The BS will manage these data and timely transmit them to the UE when needed for processing subsequent segments, ensuring the efficiency of inference.

[0204] Consequently, even resource-constrained devices can effectively perform local AI inference while maintaining data privacy and minimizing communication overhead. For storage-constrained UEs, the BS may temporarily store intermediate results or cached data from the UE, reducing the UE’s storage burden.

[0205] Efficient Multicast for Simultaneous AI Model Delivery

[0206] In many scenarios, multiple identical devices, such as IoT devices, within a cell may require simultaneous access to the same AI model. For instance, consider a network of traffic cameras deployed along a road, all needing to run the same object detection model. In such cases, a one-to-one unicast approach between the BS and each individual UE becomes inefficient and resource-intensive.

[0207] To address this, a multicast approach can be adopted where the BS transmits the model segments to a group of UEs concurrently. This significantly reduces the overall transmission time and optimizes resource utilization compared to individual unicast sessions. However, managing errors and ensuring reliable delivery become more complex in a multicast scenario.

[0208] Unlike unicast where any issues trigger an immediate stop-and-fix response between the BS and the affected UE, multicast introduces a tolerance mechanism. Before initiating the transmission, a predetermined threshold is established, representing an acceptable percentage of UEs experiencing issues. As long as the number of UEs encountering problems remains below this threshold, the BS continues the multicast transmission without interruption. This approach balances efficiency with reliability, acknowledging that occasional errors with a small subset of devices may be tolerable in scenarios where timely delivery to the majority is crucial.

[0209] The specific threshold can be dynamically adjusted based on the application requirements, network conditions, and the number of UEs in the multicast group. Additionally, error correction techniques like forward error correction (FEC) can be employed to further enhance reliability without interrupting the multicast stream.

[0210] By implementing efficient multicast protocols, scenarios with multiple identical IoT devices can be handled efficiently, enabling timely and synchronized AI model delivery while maintaining a balance between reliability and efficiency.

[0211] Optimizing Downlink Throughput for Efficient Model Transfer

[0212] Given the substantial size of AI models, even when partitioned, and the stringent latency requirements of UE-based inference, maximizing downlink throughput is beneficial. While dedicated physical channels can be allocated for model transfer, as mentioned earlier, further enhancements at the MAC layer can significantly boost efficiency.

[0213] Next generation systems may introduce novel protocols specifically designed for high-throughput applications like AI model delivery. One promising candidate is the adoption of Polar codes for AI model transmission in downlink data channels. These codes offer excellent error correction capabilities and can be efficiently decoded using successive cancellation (SC) decoders. SC decoders are currently the most energy-efficient and area-efficient approach for achieving gigabit-per-second throughput, making them ideal for handling the large data volumes associated with AI models.

[0214] To further optimize throughput, systems may move away from traditional rate matching techniques like puncturing and repetition. Instead, they might employ codewords with lengths that are powers of two, aligning naturally with the inherent structure of Polar codes. This eliminates the need for complex rate matching procedures, streamlining the encoding and decoding processes and further enhancing throughput.

[0215] Additional strategies for optimizing downlink throughput may include:

[0216] ● Link Aggregation: Combining multiple physical channels or carriers to create a wider bandwidth pipe for data transmission;

[0217] ● Advanced Modulation Schemes: Utilizing higher-order modulation schemes (e.g., 256-QAM) to pack more information into each transmitted symbol, increasing spectral efficiency;

[0218] ● MIMO Techniques: Ultra-massive Employing multiple-input multiple-output (MIMO) technologies to transmit multiple data streams simultaneously, further boosting data rates.

[0219] By implementing these advanced techniques at both the physical and MAC layers, networks can achieve unprecedented downlink throughput, enabling the efficient and timely delivery of large AI models to resource-constrained UEs and facilitating real-time AI inference at the network edge. Furthermore, advanced techniques such as channel coding, MIMO techniques, and others are considered to enhance multicast reliability and improve downlink throughput.

[0220] An overview of a method according to at least some implementations of the present disclosure is illustrated with respect to FIG. 15.

[0221] As seen in FIG. 15, a base station 1502 may provide UE 1501 with a model announcement, as illustrated by arrow 1510. At some later point, UE 1501 may send a request for an AI model to base station 1502, as illustrated by arrow 1520.

[0222] Base station 1502 may then perform a model search within its model repository to identify an AI model matching the request from UE 1501, as illustrated by block 1530. Upon finding a matching model, base station 1502 establishes a dedicated link with UE 1501, as illustrated by arrow 1540, and partitions the AI model into a plurality of segments, as illustrated by block 1550.

[0223] Base station 1502 may then provide a first segment to UE 1501, as illustrated by arrow 1560. UE 1501 then performs inference for the first segment, as illustrated by block 1570, and stores the output of the first segment, as illustrated by block 1580. For the first segment, the input may comprise an input token created by UE 1501.

[0224] At block 1590, it is determined whether all segments of the AI model have been processed. If there are remaining segments to be processed, base station 1502 provides the next segment, as illustrated by arrow 1560, and blocks 1570, 1580, and 1590, are repeated. Subsequent segments use the output of the previous segment as an input.

[0225] When it is determined, at block 1590, that all segments have been processed, the process ends.

[0226] The above functionality may be implemented on any one or combination of computing devices. FIG. 16 is a block diagram of a computing device 1600 that may be used for implementing the devices and methods disclosed herein. Specific devices  may utilize all of the components shown, or only a subset of the components, and levels of integration may vary from device to device. Furthermore, a device may contain multiple instances of a component, such as multiple processing units, processors, memories, transmitters, receivers, etc. The computing device 1600 may comprise a processor 1610, memory 1620, a mass storage device 1640, and peripherals 1630. Peripherals 1630 may comprise, amongst others one or more input / output devices, such as a speaker, microphone, mouse, touchscreen, keypad, keyboard, printer, display, network interfaces, and the like. Communications between processor 1610, memory 1620, mass storage device 1640, and peripherals 1630 may occur through one or more buses 1650.

[0227] The bus 1650 may be one or more of any type of several bus architectures including a memory bus or memory controller, a peripheral bus, video bus, or the like. The processor 1610 may comprise any type of electronic data processor. The memory 1620 may comprise any type of system memory such as static random-access memory (SRAM) , dynamic random-access memory (DRAM) , synchronous DRAM (SDRAM) , read-only memory (ROM) , a combination thereof, or the like. In an embodiment, the memory 1620 may include ROM for use at boot-up, and DRAM for program and data storage for use while executing programs.

[0228] The mass storage device 1640 may comprise any type of storage device configured to store data, programs (e.g. instructions or code) , and other information and to make the data, programs, and other information accessible via the bus. The mass storage device 1640 may comprise, for example, one or more of a solid-state drive, hard disk drive, a magnetic disk drive, an optical disk drive, or the like. The memory 1620 or mass storage 1640 may store instructions, which when executed by a processor or processing unit, cause or configure the computing device 1600 to perform any of the methods described herein.

[0229] Computing device 1600 may further comprise a communications subsystem 1660 for communicating with other computing devices or for connecting computing device 1600 to a computer network. Communications subsystem 1660 may comprise one or more network interfaces (not shown) , which may comprise wired links, such as an Ethernet cable or the like, and / or wireless links to access nodes or different networks. The network interface allows the processing unit to communicate with remote units via the networks. For example, the network interface may provide wireless communication via one or more transmitters / transmit antennas 1670 and one or more receivers / receive antennas 1670. In an embodiment, the processing unit is coupled to a local-area network or a wide-area network, for data processing and communications with remote devices, such as other processing units, the Internet, remote storage facilities, or the like.

[0230] Computing device may further comprise a power source 1680.

[0231] The present disclosure may be implemented on a computing device such as exemplary computing device 1600. Computing device 1600 may be a network element of a telecommunications network, such that the network element may be connected to other network elements of the telecommunication network, where all network elements form the telecommunication network. The network element may also receive communications from client devices connected to the telecommunication network and provide services to such client devices.

[0232] Through the descriptions of the preceding embodiments, the teachings of the present disclosure may be implemented by using hardware only or by using a combination of software and hardware. Software or other computer executable instructions for implementing one or more embodiments, or one or more portions thereof, may be stored on any suitable computer readable storage medium. The computer readable storage medium may be a tangible or in transitory / non-transitory medium such as optical (e.g., CD, DVD, Blu-Ray, etc. ) , magnetic, hard disk, volatile or non-volatile, solid state, or any other type of storage medium known in the art.

[0233] Additional features and advantages of the present disclosure will be appreciated by those skilled in the art.

[0234] The structure, features, accessories, and alternatives of specific embodiments described herein and shown in the Figures are intended to apply generally to all of the teachings of the present disclosure, including to all of the embodiments described and illustrated herein, insofar as they are compatible. In other words, the structure, features, accessories, and alternatives of a specific embodiment are not intended to be limited to only that specific embodiment unless so indicated.

[0235] Moreover, the previous detailed description is provided to enable any person skilled in the art to make or use one or more embodiments according to the present disclosure. Various modifications to those embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of the teachings provided herein. Thus, the present methods, systems, and or devices are not intended to be limited to the embodiments disclosed herein. The scope of the claims should not be limited by these embodiments, but should be given the broadest interpretation consistent with the description as a whole. Reference to an element in the singular, such as by use of the article "a" or "an" is not intended to mean "one and only one" unless specifically so stated, but rather "one or more" . All structural and functional equivalents to the elements of the various embodiments described throughout the disclosure that are known or later come to be known to those of ordinary skill in the art are intended to be encompassed by the elements of the claims.

[0236] Furthermore, nothing herein is intended as an admission of prior art or of common general knowledge. Furthermore, citation or identification of any document in this application is not an admission that such document is available as prior art, or that any reference forms a part of the common general knowledge in the art. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

[0237] In the foregoing description, numerous details are set forth to provide an understanding of the subject disclosed herein. However, implementations may be practiced without some of these details. Other implementations may include modifications and variations from the details discussed above. It is intended that the appended claims cover such modifications and variations.

Claims

1.A method at a network element, comprising:receiving a request from a user equipment (UE) , the request comprising at least one of a UE identifier, a model identifier, UE capabilities, or channel state information (CSI) ;searching for an artificial intelligence (AI) model based on the request;when the AI model is identified, establishing a dedicated link between the network element and the UE;partitioning the identified AI model into a plurality of segments; andtransmitting each segment of the plurality of segments to the UE over the dedicated link.2.The method of claim 1, further comprising:allocating resources for transmission of the plurality of segments to the UE; andindicating the resources to the UE.3.The method of claim 1, further comprising, transmitting to the UE a first indication comprising information on a selected AI model stored at the network element.4.The method of claim 1, further comprising:receiving a plurality of requests from a plurality of UEs, each of the plurality of requests comprising the model identifier; andwherein said transmitting comprises multicasting each segment of the plurality of segments to each of the plurality of UEs.5.The method of claim 4, further comprising:receiving a second indication from at least one UE of the plurality of UEs that a segment transmission comprised an error;determining that a number of the at least one UEs is below a threshold; andcontinuing the multicasting of the plurality of segments to each of the plurality of UEs.6.A network element comprising:a processor; anda communications subsystem;wherein the network element is configured to:receive a request from a user equipment (UE) , the request comprising at least one of a UE identifier, a model identifier, UE capabilities, or channel state information (CSI) ;search for an artificial intelligence (AI) model based on the request;when the AI model is identified, establish a dedicated link between the network element and the UE;partition the identified AI model into a plurality of segments; andtransmit each segment of the plurality of segments to the UE over the dedicated link.7.The network element of claim 6, further configured to:allocate resources for transmission of the plurality of segments to the UE; andindicate the resources to the UE.8.The network element of claim 6, further configured to, transmit to the UE a first indication comprising information on a selected AI model stored at the network element.9.The network element of claim 6, further configured to:receive a plurality of requests from a plurality of UEs, each of the plurality of requests comprising the model identifier; andwherein said transmitting comprises multicasting each segment of the plurality of segments to each of the plurality of UEs.10.The network element of claim 9, further configured to:receive a second indication from at least one UE of the plurality of UEs that a segment transmission comprised an error;determine that a number of the at least one UEs is below a threshold; andcontinue the multicasting of the plurality of segments to each of the plurality of UEs.11.A non-transitory computer-readable storage medium having instructions stored thereon which, when executed by an apparatus, cause the apparatus to perform the method of any one of claims 1 to 5.12.A method at a User Equipment (UE) , comprising:transmitting a request to a network element, the request comprising a UE identifier, a model identifier, UE capabilities, and Channel State Information (CSI) ;establishing a dedicated link with the network element;receiving a first segment of an Artificial Intelligence (AI) model corresponding to the model identifier over the dedicated link;creating an input token;computing the first segment with the input token and storing a first output token;receiving a subsequent segment of the AI model over the dedicated link;computing the subsequent segment with the first output token and storing a subsequent output token; andrepeating the steps of receiving the subsequent segment and computing the subsequent segment until all segments of the AI model are computed, wherein an output for a given segment is used as an input for a next segment.13.The method of claim 12, further comprising:receiving a first indication from the network element, the first indication indicating resources for transmissions of each segment of the AI model.14.The method of claim 12, further comprising, receiving a second indication from the network element, the second indication comprising information on a selected AI model stored at the network element.15.The method of claim 12, further comprising, discarding a segment from memory after computation of the segment.16.The method of claim 12, further comprising, discarding an output token after using it as an input for a segment.17.The method of claim 12, further comprising transmitting to the network element an intermediate output of at least one segment of the AI model for temporary storage.18.A user equipment comprising:a processor; anda communications subsystem;wherein the user equipment is configured to:transmit a request to a network element, the request comprising a UE identifier, a model identifier, UE capabilities, and Channel State Information (CSI) ;establish a dedicated link with the network element;receive a first segment of an Artificial Intelligence (AI) model corresponding to the model identifier over the dedicated link;create an input token;compute the first segment with the input token and storing a first output token;receive a subsequent segment of the AI model over the dedicated link;compute the subsequent segment with the first output token and store a subsequent output token; andrepeat the steps of receiving the subsequent segment and computing the subsequent segment until all segments of the AI model are computed, wherein an output for a given segment is used as an input for a next segment.19.The user equipment of claim 18, further configured to:receive a first indication from the network element, the first indication indicating resources for transmissions of each segment of the AI model.20.The user equipment of claim 18, further configured to receive a second indication from the network element, the second indication comprising information on a selected AI model stored at the network element.21.The user equipment of claim 18, further configured to, discard a segment from memory after computation of the segment.22.The user equipment of claim 18, further configured to, discard an output token after using it as an input for a segment.23.The user equipment of claim 18, further configured to transmit to the network element an intermediate output of at least one segment of the AI model for temporary storage.24.A non-transitory computer-readable storage medium having instructions stored thereon which, when executed by an apparatus, cause the apparatus to perform the method of any one of claims 12 to 17.25.A communication apparatus, configured to perform the method according to any one of claims 1 to 5 or any one of claims 12 to 17.26.The communication apparatus of claim 25, comprising:a receiving unit configured to receive a request from a user equipment (UE) , the request comprising at least one of a UE identifier, a model identifier, UE capabilities, or channel state information (CSI) ; anda processing unit configured to:search for an artificial intelligence (AI) model based on the request;when the AI model is identified, establish a dedicated link between the network element and the UE; andpartition the identified AI model into a plurality of segments; anda transmitting unit configured to transmit each segment of the plurality of segments to the UE over the dedicated link.27.The communication apparatus of claim 25, comprising:a transmitting unit configured to:transmit a request to a network element, the request comprising a UE identifier, a model identifier, UE capabilities, and Channel State Information (CSI) ;a receiving unit configured to:receive a first segment of an Artificial Intelligence (AI) model corresponding to the model identifier over a dedicated link, and a subsequent segment of the AI model over the dedicated link;a processing unit configured to:establish a dedicated link with the network element;create an input token;compute the first segment with the input token and storing a first output token; andcomputing the subsequent segment with the first output token and storing a subsequent output token;where the steps of receiving the subsequent segment and computing the subsequent segment are repeated until all segments of the AI model are computed, wherein an output for a given segment is used as an input for a next segment.28.The communication apparatus of claim 25, comprising:one or more processors configured to:search for an artificial intelligence (AI) model based on the request;when the AI model is identified, establish a dedicated link between the network element and the UE; andpartition the identified AI model into a plurality of segments; andan interface circuit configured to:receive a request from a user equipment (UE) , the request comprising at least one of a UE identifier, a model identifier, UE capabilities, or channel state information (CSI) ; andtransmit each segment of the plurality of segments to the UE over the dedicated link.29.The communication apparatus of claim 25, comprising:one or more processors configured to:establish a dedicated link with the network element;create an input token;compute the first segment with the input token and storing a first output token; andcompute the subsequent segment with the first output token and storing a subsequent output token; andan interface circuit configured to:transmit a request to a network element, the request comprising a UE identifier, a model identifier, UE capabilities, and Channel State Information (CSI) ;receive a first segment of an Artificial Intelligence (AI) model corresponding to the model identifier over a dedicated link, and a subsequent segment of the AI model over the dedicated link;where the steps of receiving the subsequent segment and computing the subsequent segment are repeated until all segments of the AI model are computed, wherein an output for a given segment is used as an input for a next segment.30.The communication apparatus of claim 28 or claim 29, wherein the interface circuit comprises one or more transceivers.31.An apparatus comprising:one or more processors; andone or more memories storing instructions which, when executed by the one or more processors, cause the apparatus to perform the method of any one of claims 1 to 5 or any one of claims 12 to 17.32.A communication system, wherein the communication system comprises a first communication apparatus configured to perform the method of any one of claims 1 to 5 and a second communication apparatus configured to perform the method of any one of claims 12 to 17.

Citation Information

Patent Citations

  • AI / ML service device for use in NG-RAN

    CN116939749A

  • Token-Based Validation Method for Segmented Content Delivery

    US20140115724A1

  • Model coordination method and apparatus

    US20230004839A1

  • Terminal device, network device, and method for ai model transfer

    WO2024093151A1