Method and apparatus of quantization configuration for artificial intelligence (AI) / machine learning (ML) models

By training AI/ML models at UE or base stations and employing quantization methods, the method optimizes data transmission and inference processes, addressing inefficiencies in 5G NR technology and enhancing communication performance.

WO2026059542A1PCT designated stage Publication Date: 2026-03-19MEDIATEK INC +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

There is a need for improved techniques in 5G New Radio (NR) technology to enhance AI/ML model training and inference processes in wireless communication systems, particularly in quantization methods for efficient data transmission and processing between user equipment (UE) and base stations.

Method used

The method involves training AI/ML models at the UE or base station, performing quantization and de-quantization of data samples, and executing encoders or decoders to optimize data transmission and inference stages, utilizing various quantization methods such as scalar and vector quantization.

Benefits of technology

This approach enhances the efficiency and effectiveness of AI/ML model training and inference processes, reducing data overhead and improving communication performance in wireless networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024045924_19032026_PF_FP_ABST
    Figure US2024045924_19032026_PF_FP_ABST
Patent Text Reader

Abstract

In an aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The method may be performed by a UE. In certain configurations, the UE collects data samples for an artificial intelligence (AI) / machine learning (ML) model at the UE and a base station. The AI / ML model is trained at the UE or at the base station in a training stage. The UE performs, according to a quantization method, quantization of the data samples to obtain quantized data samples. The UE transmits the quantized data samples to the base station. The UE executes an encoder or a decoder of the trained AI / ML model in an inference stage. The quantization method may be a latent quantization method for a latent space, and the data samples are latent vectors of Channel State Information (CSI) samples measured by the UE.
Need to check novelty before this filing date? Find Prior Art

Description

MTK Ref. No. MUSI-24-0043PCTMETHOD AND APPARATUS OF QUANTIZATION CONFIGURATION FOR ARTIFICIAL INTELLIGENCE (Al) / MACHINE LEARNING (ML) MODELSBACKGROUNDField

[0001] The present disclosure relates generally to communication systems, and more particularly, to techniques of methods and apparatuses of quantization configuration for artificial intelligence (AI) / machine learning (ML) models.Background

[0002] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.

[0003] Wireless communication systems are widely deployed to provide various telecommunication services such as telephony, video, data, messaging, and broadcasts. Typical wireless communication systems may employ multiple-access technologies capable of supporting communication with multiple users by sharing available system resources. Examples of such multiple-access technologies include code division multiple access (CDMA) systems, time division multiple access (TDMA) systems, frequency division multiple access (FDMA) systems, orthogonal frequency division multiple access (OFDMA) systems, single-carrier frequency division multiple access (SC-FDMA) systems, and time division synchronous code division multiple access (TD-SCDMA) systems.

[0004] These multiple access technologies have been adopted in various telecommunication standards to provide a common protocol that enables different wireless devices to communicate on a municipal, national, regional, and even global level. An example telecommunication standard is 5G New Radio (NR). 5G NR is part of a continuous mobile broadband evolution promulgated by Third Generation Partnership Project (3 GPP) to meet new requirements associated with latency, reliability, security, scalability (e.g., with Internet of Things (IoT)), and other requirements. Some aspects of 5GNRmay be based on the 4G Long Term Evolution (LTE) standard. There exists a need for further improvements in 5G NR technology. These improvements may also be applicable to other multi-access technologies and the telecommunication standards that employ these technologies.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCTSUMMARY

[0005] The following presents a simplified summary of one or more aspects in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements of all aspects nor delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that is presented later.

[0006] In an aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The method may be performed by a UE. In certain configurations, the UE collects data samples for training an artificial intelligence (AI) / machine learning (ML) model at the UE and a base station. The AI / ML model is trained at the UE or at the base station in a training stage. The UE performs, according to a quantization method, quantization of the data samples to obtain quantized data samples. The UE transmits the quantized data samples to the base station. The UE executes an encoder or a decoder of the trained AI / ML model in an inference stage.

[0007] In another aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The method may be performed by a base station. In certain configurations, the base station receives, from a UE, quantized data samples. The base station performs, according to a quantization method, de-quantization of the quantized data samples to obtain de-quantized data samples for training of an AI / ML model at the UE or the base station in a training stage. The base station executes an encoder or a decoder of the trained AI / ML model in an inference stage.

[0008] To the accomplishment of the foregoing and related ends, the one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the annexed drawings set forth in detail certain illustrative features of the one or more aspects. These features are indicative, however, of but a few of the various ways in which the principles of various aspects may be employed, and this description is intended to include all such aspects and their equivalents.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCTBRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG. l is a diagram illustrating an example of a wireless communications system and an access network.

[0010] FIG. 2 is a diagram illustrating a base station in communication with a UE in an access network.

[0011] FIG. 3 illustrates an example logical architecture of a distributed access network.

[0012] FIG. 4 illustrates an example physical architecture of a distributed access network.

[0013] FIG. 5 is a diagram showing an example of a DL-centric slot.

[0014] FIG. 6 is a diagram showing an example of an UL-centric slot.

[0015] FIG. 7 is a diagram illustrating a CSI compression cycle in an AI / ML model between a UE and a base station.

[0016] FIG. 8 is a diagram illustrating an example uniform scalar quantization method.

[0017] FIG. 9 is a diagram illustrating an example non-uniform scalar quantization method described by a direct specification of the I / O relation.

[0018] FIG. 10 is a diagram illustrating an example non-uniform scalar quantization method described by shaping I / O of a uniform scalar quantization.

[0019] FIG. 11 is a diagram illustrating an example vector quantization method.

[0020] FIG. 12 is a diagram illustrating an example segmented vector quantization method.

[0021] FIG. 13 is a diagram illustrating a data collection and training process using a training type I approach for an AI / ML model between a UE and a base station, where the training is performed at the UE.

[0022] FIG. 14 is a diagram illustrating a data collection and training process using a training type I approach for an AI / ML model between a UE and a base station, where the training is performed at the base station.

[0023] FIG. 15 is a diagram illustrating a data collection and training process using a training type II approach for an AI / ML model between a UE and a base station, where joint training is performed at both the UE and the base station.

[0024] FIG. 16 is a diagram illustrating an example data collection process with forward pass and backward propagation procedures in the data collection and training process in FIG. 15.

[0025] FIG. 17 is a diagram illustrating latent quantization and gradient quantization of the forward pass and backward propagation procedures in the data collection and training process in FIG. 16.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0026] FIG. 18 is a diagram illustrating an example quantification configuration of the UE and the base station in the data collection and training process in FIG. 15.

[0027] FIG. 19 is a diagram illustrating a data collection and training process using a training type III approach for an AI / ML model between a UE and a base station, where sequential separate training is performed firstly at the UE and then at the base station.

[0028] FIG. 20 is a diagram illustrating an example quantification configuration of the UE and the base station in the data collection and training process in FIG. 19.

[0029] FIG. 21 is a diagram illustrating an example quantification configuration of two UEs and the base station in the data collection and training process in FIG. 19.

[0030] FIG. 22 is a diagram illustrating a data collection and training process using a training type III approach for an AI / ML model between a UE and a base station, where sequential separate training is performed firstly at the base station and then at the UE.

[0031] FIG. 23 is a diagram illustrating an example quantification configuration of the UE and the base station in the data collection and training process in FIG. 22.

[0032] FIG. 24 is a table showing the configuring entity for the quantization method in each of the training type approaches when the encoder is at the UE and the decoder is at the base station.

[0033] FIG. 25 is a table showing the configuring entity for the quantization method in each of the training type approaches when the encoder is at the base station and the decoder is at the UE.

[0034] FIG. 26 is a flow chart of a method (process) for wireless communication of a UE.

[0035] FIG. 27 is a flow chart of a method (process) for wireless communication of a base station.DETAILED DESCRIPTION

[0036] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well known structures and components are shown in block diagram form in order to avoid obscuring such concepts.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0037] Several aspects of telecommunications systems will now be presented with reference to various apparatus and methods. These apparatus and methods will be described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as “elements”). These elements may be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.

[0038] By way of example, an element, or any portion of an element, or any combination of elements may be implemented as a “processing system” that includes one or more processors. Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, systems on a chip (SoC), baseband processors, field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functionality described throughout this disclosure. One or more processors in the processing system may execute software. Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software components, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0039] Accordingly, in one or more example aspects, the functions described may be implemented in hardware, software, or any combination thereof. If implemented in software, the functions may be stored on or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media includes computer storage media. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise a random-access memory (RAM), a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the aforementioned types of computer-readable media, or any other medium that can be used to storeLL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT computer executable code in the form of instructions or data structures that can be accessed by a computer.

[0040] FIG. l is a diagram illustrating an example of a wireless communications system and an access network 100. The wireless communications system (also referred to as a wireless wide area network (WWAN)) includes base stations 102, UEs 104, an Evolved Packet Core (EPC) 160, and another core network 190 (e.g., a 5G Core (5GC)). The base stations 102 may include macrocells (high power cellular base station) and / or small cells (low power cellular base station). The macrocells include base stations. The small cells include femtocells, picocells, and microcells.

[0041] The base stations 102 configured for 4G LTE (collectively referred to as Evolved Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access Network (E-UTRAN)) may interface with the EPC 160 through backhaul links 132 (e.g., SI interface). The base stations 102 configured for 5G NR (collectively referred to as Next Generation RAN (NG-RAN)) may interface with core network 190 through backhaul links 184. In addition to other functions, the base stations 102 may perform one or more of the following functions: transfer of user data, radio channel ciphering and deciphering, integrity protection, header compression, mobility control functions (e.g., handover, dual connectivity), inter cell interference coordination, connection setup and release, load balancing, distribution for non-access stratum (NAS) messages, NAS node selection, synchronization, radio access network (RAN) sharing, multimedia broadcast multicast service (MBMS), subscriber and equipment trace, RAN information management (RIM), paging, positioning, and delivery of warning messages. The base stations 102 may communicate directly or indirectly (e.g., through the EPC 160 or core network 190) with each other over backhaul links 134 (e.g., X2 interface). The backhaul links 134 may be wired or wireless.

[0042] The base stations 102 may wirelessly communicate with the UEs 104. Each of the base stations 102 may provide communication coverage for a respective geographic coverage area 110. There may be overlapping geographic coverage areas 110. For example, the small cell 102’ may have a coverage area 110’ that overlaps the coverage area 110 of one or more macro base stations 102. A network that includes both small cell and macrocells may be known as a heterogeneous network. A heterogeneous network may also include Home Evolved Node Bs (eNBs) (HeNBs), which may provide service to a restricted group known as a closed subscriber group (CSG). The communication links 120 between the base stations 102 and the UEs 104 may include LL Docket 1007576.315WO1 ,MTK Ref. No. MUSI-24-0043PCT uplink (UL) (also referred to as reverse link) transmissions from a UE 104 to a base station 102 and / or downlink (DL) (also referred to as forward link) transmissions from a base station 102 to a UE 104. The communication links 120 may use multiple-input and multiple-output (MIMO) antenna technology, including spatial multiplexing, beamforming, and / or transmit diversity. The communication links may be through one or more carriers. The base stations 102 / UEs 104 may use spectrum up to 7 MHz (e.g., 5, 10, 15, 20, 100, 400, etc. MHz) bandwidth per carrier allocated in a carrier aggregation of up to a total of Yx MHz (x component carriers) used for transmission in each direction. The carriers may or may not be adjacent to each other. Allocation of carriers may be asymmetric with respect to DL and UL (e.g., more or fewer carriers may be allocated for DL than for UL). The component carriers may include a primary component carrier and one or more secondary component carriers. A primary component carrier may be referred to as a primary cell (PCell) and a secondary component carrier may be referred to as a secondary cell (SCell).

[0043] Certain UEs 104 may communicate with each other using device-to-device (D2D) communication link 158. The D2D communication link 158 may use the DL / UL WWAN spectrum. The D2D communication link 158 may use one or more sidelink channels, such as a physical sidelink broadcast channel (PSBCH), a physical sidelink discovery channel (PSDCH), a physical sidelink shared channel (PSSCH), and a physical sidelink control channel (PSCCH). D2D communication may be through a variety of wireless D2D communications systems, such as for example, FlashLinQ, WiMedia, Bluetooth, ZigBee, Wi-Fi based on the IEEE 802.11 standard, LTE, or NR.

[0044] The wireless communications system may further include a Wi-Fi access point (AP) 150 in communication with Wi-Fi stations (STAs) 152 via communication links 154 in a 5 GHz unlicensed frequency spectrum. When communicating in an unlicensed frequency spectrum, the STAs 152 / AP 150 may perform a clear channel assessment (CCA) prior to communicating in order to determine whether the channel is available.

[0045] The small cell 102’ may operate in a licensed and / or an unlicensed frequency spectrum. When operating in an unlicensed frequency spectrum, the small cell 102’ may employ NR and use the same 5 GHz unlicensed frequency spectrum as used by the Wi-Fi AP 150. The small cell 102’, employing NR in an unlicensed frequency spectrum, may boost coverage to and / or increase capacity of the access network.

[0046] A base station 102, whether a small cell 102’ or a large cell (e.g., macro base station), may include an eNB, gNodeB (gNB), or another type of base station. Some base LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT stations, such as gNB 180 may operate in a traditional sub 6 GHz spectrum, in millimeter wave (mmW) frequencies, and / or near mmW frequencies in communication with the UE 104. When the gNB 180 operates in mmW or near mmW frequencies, the gNB 180 may be referred to as an mmW base station. Extremely high frequency (EHF) is part of the RF in the electromagnetic spectrum. EHF has a range of 30 GHz to 300 GHz and a wavelength between 1 millimeter and 10 millimeters. Radio waves in the band may be referred to as a millimeter wave. Near mmW may extend down to a frequency of 3 GHz with a wavelength of 100 millimeters. The super high frequency (SHF) band extends between 3 GHz and 30 GHz, also referred to as centimeter wave. Communications using the mmW / near mmW radio frequency band (e.g., 3 GHz - 300 GHz) has extremely high path loss and a short range. The mmW base station 180 may utilize beamforming 182 with the UE 104 to compensate for the extremely high path loss and short range.

[0047] The base station 180 may transmit a beamformed signal to the UE 104 in one or more transmit directions 108a. The UE 104 may receive the beamformed signal from the base station 180 in one or more receive directions 108b. TheUE 104 may also transmit a beamformed signal to the base station 180 in one or more transmit directions. The base station 180 may receive the beamformed signal from the UE 104 in one or more receive directions. The base station 180 / UE 104 may perform beam training to determine the best receive and transmit directions for each of the base station 180 / UE 104. The transmit and receive directions for the base station 180 may or may not be the same. The transmit and receive directions for the UE 104 may or may not be the same.

[0048] The EPC 160 may include a Mobility Management Entity (MME) 162, other MMEs 164, a Serving Gateway 166, a Multimedia Broadcast Multicast Service (MBMS) Gateway 168, a Broadcast Multicast Service Center (BM-SC) 170, and a Packet Data Network (PDN) Gateway 172. The MME 162 may be in communication with a Home Subscriber Server (HSS) 174. The MME 162 is the control node that processes the signaling between the UEs 104 and the EPC 160. Generally, the MME 162 provides bearer and connection management. All user Internet protocol (IP) packets are transferred through the Serving Gateway 166, which itself is connected to the PDN Gateway 172. The PDN Gateway 172 provides UE IP address allocation as well as other functions. The PDN Gateway 172 and the BM-SC 170 are connected to the IP Services 176. The IP Services 176 may include the Internet, an intranet, an IP LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCTMultimedia Subsystem (IMS), a PS Streaming Service, and / or other IP services. The BM-SC 170 may provide functions for MBMS user service provisioning and delivery. The BM-SC 170 may serve as an entry point for content provider MBMS transmission, may be used to authorize and initiate MBMS Bearer Services within a public land mobile network (PLMN), and may be used to schedule MBMS transmissions. The MBMS Gateway 168 may be used to distribute MBMS traffic to the base stations 102 belonging to a Multicast Broadcast Single Frequency Network (MBSFN) area broadcasting a particular service, and may be responsible for session management (start / stop) and for collecting eMBMS related charging information.

[0049] The core network 190 may include an Access and Mobility Management Function (AMF) 192, other AMFs 193, a location management function (LMF) 198, a Session Management Function (SMF) 194, and a User Plane Function (UPF) 195. The AMF 192 may be in communication with a Unified Data Management (UDM) 196. The AMF 192 is the control node that processes the signaling between the UEs 104 and the core network 190. Generally, the SMF 194 provides QoS flow and session management. All user Internet protocol (IP) packets are transferred through the UPF 195. The UPF 195 provides UE IP address allocation as well as other functions. The UPF 195 is connected to the IP Services 197. The IP Services 197 may include the Internet, an intranet, an IP Multimedia Subsystem (IMS), a PS Streaming Service, and / or other IP services.

[0050] The base station may also be referred to as a gNB, Node B, evolved Node B (eNB), an access point, a base transceiver station, a radio base station, a radio transceiver, a transceiver function, a basic service set (BSS), an extended service set (ESS), a transmit reception point (TRP), or some other suitable terminology. The base station 102 provides an access point to the EPC 160 or core network 190 for a UE 104. Examples of UEs 104 include a cellular phone, a smart phone, a session initiation protocol (SIP) phone, a laptop, a personal digital assistant (PDA), a satellite radio, a global positioning system, a multimedia device, a video device, a digital audio player (e.g., MP3 player), a camera, a game console, a tablet, a smart device, a wearable device, a vehicle, an electric meter, a gas pump, a large or small kitchen appliance, a healthcare device, an implant, a sensor / actuator, a display, or any other similar functioning device. Some of the UEs 104 may be referred to as loT devices (e.g., parking meter, gas pump, toaster, vehicles, heart monitor, etc.). The UE 104 may also be referred to as a station, a mobile station, a subscriber station, a mobile unit, a LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT subscriber unit, a wireless unit, a remote unit, a mobile device, a wireless device, a wireless communications device, a remote device, a mobile subscriber station, an access terminal, a mobile terminal, a wireless terminal, a remote terminal, a handset, a user agent, a mobile client, a client, or some other suitable terminology.

[0051] Although the present disclosure may reference 5G New Radio (NR), the present disclosure may be applicable to other similar areas, such as LTE, LTE-Advanced (LTE-A), Code Division Multiple Access (CDMA), Global System for Mobile communications (GSM), or other wireless / radio access technologies.

[0052] FIG. 2 is a block diagram of a base station 210 in communication with a UE 250 in an access network. In the DL, IP packets from the EPC 160 may be provided to a controller / processor 275. The controller / processor 275 implements layer 3 and layer 2 functionality. Layer 3 includes a radio resource control (RRC) layer, and layer 2 includes a packet data convergence protocol (PDCP) layer, a radio link control (RLC) layer, and a medium access control (MAC) layer. The controller / processor 275 provides RRC layer functionality associated with broadcasting of system information (e.g., MIB, SIBs), RRC connection control (e.g., RRC connection paging, RRC connection establishment, RRC connection modification, and RRC connection release), inter radio access technology (RAT) mobility, and measurement configuration for UE measurement reporting; PDCP layer functionality associated with header compression / decompression, security (ciphering, deciphering, integrity protection, integrity verification), and handover support functions; RLC layer functionality associated with the transfer of upper layer packet data units (PDUs), error correction through ARQ, concatenation, segmentation, and reassembly of RLC service data units (SDUs), re-segmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto transport blocks (TBs), demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling, and logical channel prioritization.

[0053] The transmit (TX) processor 216 and the receive (RX) processor 270 implement layer 1 functionality associated with various signal processing functions. Layer 1, which includes a physical (PHY) layer, may include error detection on the transport channels, forward error correction (FEC) coding / decoding of the transport channels, interleaving, rate matching, mapping onto physical channels, modulation / demodulation of physical channels, and MIMO antenna processing. The LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCTTX processor 216 handles mapping to signal constellations based on various modulation schemes (e.g., binary phase-shift keying (BPSK), quadrature phase-shift keying (QPSK), M-phase-shift keying (M-PSK), M-quadrature amplitude modulation (M-QAM)). The coded and modulated symbols may then be split into parallel streams. Each stream may then be mapped to an OFDM subcarrier, multiplexed with a reference signal (e.g., pilot) in the time and / or frequency domain, and then combined together using an Inverse Fast Fourier Transform (IFFT) to produce a physical channel carrying a time domain OFDM symbol stream. The OFDM stream is spatially precoded to produce multiple spatial streams. Channel estimates from a channel estimator 274 may be used to determine the coding and modulation scheme, as well as for spatial processing. The channel estimate may be derived from a reference signal and / or channel condition feedback transmitted by the UE 250. Each spatial stream may then be provided to a different antenna 220 via a separate transmitter 218TX. Each transmitter 218TX may modulate an RF carrier with a respective spatial stream for transmission.

[0054] At the UE 250, each receiver 254RX receives a signal through its respective antenna 252. Each receiver 254RX recovers information modulated onto an RF carrier and provides the information to the receive (RX) processor 256. The TX processor 268 and the RX processor 256 implement layer 1 functionality associated with various signal processing functions. The RX processor 256 may perform spatial processing on the information to recover any spatial streams destined for the UE 250. If multiple spatial streams are destined for the UE 250, they may be combined by the RX processor 256 into a single OFDM symbol stream. The RX processor 256 then converts the OFDM symbol stream from the time-domain to the frequency domain using a Fast Fourier Transform (FFT). The frequency domain signal comprises a separate OFDM symbol stream for each subcarrier of the OFDM signal. The symbols on each subcarrier, and the reference signal, are recovered and demodulated by determining the most likely signal constellation points transmitted by the base station 210. These soft decisions may be based on channel estimates computed by the channel estimator 258. The soft decisions are then decoded and deinterleaved to recover the data and control signals that were originally transmitted by the base station 210 on the physical channel. The data and control signals are then provided to the controller / processor 259, which implements layer 3 and layer 2 functionality.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0055] The controller / processor 259 can be associated with a memory 260 that stores program codes and data. The memory 260 may be referred to as a computer- readable medium. In the UL, the controller / processor 259 provides demultiplexing between transport and logical channels, packet reassembly, deciphering, header decompression, and control signal processing to recover IP packets from the EPC 160. The controller / processor 259 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.

[0056] Similar to the functionality described in connection with the DL transmission by the base station 210, the controller / processor 259 provides RRC layer functionality associated with system information (e.g., MIB, SIBs) acquisition, RRC connections, and measurement reporting; PDCP layer functionality associated with header compression / decompression, and security (ciphering, deciphering, integrity protection, integrity verification); RLC layer functionality associated with the transfer of upper layer PDUs, error correction through ARQ, concatenation, segmentation, and reassembly of RLC SDUs, re- segmentation of RLC data PDUs, and reordering of RLC data PDUs; and MAC layer functionality associated with mapping between logical channels and transport channels, multiplexing of MAC SDUs onto TBs, demultiplexing of MAC SDUs from TBs, scheduling information reporting, error correction through HARQ, priority handling, and logical channel prioritization.

[0057] Channel estimates derived by a channel estimator 258 from a reference signal or feedback transmitted by the base station 210 may be used by the TX processor 268 to select the appropriate coding and modulation schemes, and to facilitate spatial processing. The spatial streams generated by the TX processor 268 may be provided to different antenna 252 via separate transmitters 254TX. Each transmitter 254TX may modulate an RF carrier with a respective spatial stream for transmission. The UL transmission is processed at the base station 210 in a manner similar to that described in connection with the receiver function at the UE 250. Each receiver 218RX receives a signal through its respective antenna 220. Each receiver 218RX recovers information modulated onto an RF carrier and provides the information to a RX processor 270.

[0058] The controller / processor 275 can be associated with a memory 276 that stores program codes and data. The memory 276 may be referred to as a computer- readable medium. In the UL, the controller / processor 275 provides demultiplexing between transport and logical channels, packet reassembly, deciphering, header LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT decompression, control signal processing to recover IP packets from the UE 250. IP packets from the controller / processor 275 may be provided to the EPC 160. The controller / processor 275 is also responsible for error detection using an ACK and / or NACK protocol to support HARQ operations.

[0059] New radio (NR) may refer to radios configured to operate according to a new air interface (e.g., other than Orthogonal Frequency Divisional Multiple Access (OFDMA)-based air interfaces) or fixed transport layer (e.g., other than Internet Protocol (IP)). NR may utilize OFDM with a cyclic prefix (CP) on the uplink and downlink and may include support for half-duplex operation using time division duplexing (TDD). NR may include Enhanced Mobile Broadband (eMBB) service targeting wide bandwidth (e.g. 80 MHz beyond), millimeter wave (mmW) targeting high carrier frequency (e.g. 60 GHz), massive MTC (mMTC) targeting nonbackward compatible MTC techniques, and / or mission critical targeting ultra-reliable low latency communications (URLLC) service.

[0060] A single component carrier bandwidth of 100 MHz may be supported. In one example, NR resource blocks (RBs) may span 12 sub-carriers with a sub-carrier bandwidth of 60 kHz over a 0.25 ms duration or a bandwidth of 30 kHz over a 0.5 ms duration (similarly, 50MHz BW for 15kHz SCS over a 1 ms duration). Each radio frame may consist of 10 subframes (10, 20, 40 or 80 NR slots) with a length of 10 ms. Each slot may indicate a link direction (i.e., DL or UL) for data transmission and the link direction for each slot may be dynamically switched. Each slot may include DL / UL data as well as DL / UL control data. UL and DL slots for NR may be as described in more detail below with respect to FIGs. 5 and 6.

[0061] The NR RAN may include a central unit (CU) and distributed units (DUs). A NR BS (e.g., gNB, 5G Node B, Node B, transmission reception point (TRP), access point (AP)) may correspond to one or multiple BSs. NR cells can be configured as access cells (ACells) or data only cells (DCells). For example, the RAN (e.g., a central unit or distributed unit) can configure the cells. DCells may be cells used for carrier aggregation or dual connectivity and may not be used for initial access, cell sei ection / re sei ection, or handover. In some cases DCells may not transmit synchronization signals (SS) in some cases DCells may transmit SS. NR BSs may transmit downlink signals to UEs indicating the cell type. Based on the cell type indication, the UE may communicate with the NR BS. For example, the UE mayLL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT determine NR BSs to consider for cell selection, access, handover, and / or measurement based on the indicated cell type.

[0062] FIG. 3 illustrates an example logical architecture of a distributed RAN 300, according to aspects of the present disclosure. A 5G access node 306 may include an access node controller (ANC) 302. The ANC may be a central unit (CU) of the distributed RAN. The backhaul interface to the next generation core network (NG- CN) 304 may terminate at the ANC. The backhaul interface to neighboring next generation access nodes (NG-ANs) 310 may terminate at the ANC. The ANC may include one or more TRPs 308 (which may also be referred to as BSs, NR BSs, Node Bs, 5G NBs, APs, or some other term). As described above, a TRP may be used interchangeably with “cell.”

[0063] The TRPs 308 may be a distributed unit (DU). The TRPs may be connected to one ANC (ANC 302) or more than one ANC (not illustrated). For example, for RAN sharing, radio as a service (RaaS), and service specific ANC deployments, the TRP may be connected to more than one ANC. A TRP may include one or more antenna ports. The TRPs may be configured to individually (e.g., dynamic selection) or jointly (e.g., joint transmission) serve traffic to a UE.

[0064] The local architecture of the distributed RAN 300 may be used to illustrate fronthaul definition. The architecture may be defined that support fronthauling solutions across different deployment types. For example, the architecture may be based on transmit network capabilities (e.g., bandwidth, latency, and / or jitter). The architecture may share features and / or components with LTE. According to aspects, the next generation AN (NG- AN) 310 may support dual connectivity with NR. The NG- AN may share a common fronthaul for LTE and NR.

[0065] The architecture may enable cooperation between and among TRPs 308. For example, cooperation may be preset within a TRP and / or across TRPs via the ANC 302. According to aspects, no inter- TRP interface may be needed / present.

[0066] According to aspects, a dynamic configuration of split logical functions may be present within the architecture of the distributed RAN 300. The PDCP, RLC, MAC protocol may be adaptably placed at the ANC or TRP.

[0067] FIG. 4 illustrates an example physical architecture of a distributed RAN 400, according to aspects of the present disclosure. A centralized core network unit (C- CU) 402 may host core network functions. The C-CU may be centrally deployed. C- CU functionality may be offloaded (e.g., to advanced wireless services (AWS)), in an LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT effort to handle peak capacity. A centralized RAN unit (C-RU) 404 may host one or more ANC functions. Optionally, the C-RU may host core network functions locally. The C-RU may have distributed deployment. The C-RU may be closer to the network edge. A distributed unit (DU) 406 may host one or more TRPs. The DU may be located at edges of the network with radio frequency (RF) functionality.

[0068] FIG. 5 is a diagram 500 showing an example of a DL-centric slot. The DL-centric slot may include a control portion 502. The control portion 502 may exist in the initial or beginning portion of the DL-centric slot. The control portion 502 may include various scheduling information and / or control information corresponding to various portions of the DL-centric slot. In some configurations, the control portion 502 may be a physical DL control channel (PDCCH), as indicated in FIG. 5. The DL-centric slot may also include a DL data portion 504. The DL data portion 504 may sometimes be referred to as the payload of the DL-centric slot. The DL data portion 504 may include the communication resources utilized to communicate DL data from the scheduling entity (e.g., UE or BS) to the subordinate entity (e.g., UE). In some configurations, the DL data portion 504 may be a physical DL shared channel (PDSCH).

[0069] The DL-centric slot may also include a common UL portion 506. The common UL portion 506 may sometimes be referred to as an UL burst, a common UL burst, and / or various other suitable terms. The common UL portion 506 may include feedback information corresponding to various other portions of the DL-centric slot. For example, the common UL portion 506 may include feedback information corresponding to the control portion 502. Non-limiting examples of feedback information may include an ACK signal, a NACK signal, a HARQ indicator, and / or various other suitable types of information. The common UL portion 506 may include additional or alternative information, such as information pertaining to random access channel (RACH) procedures, scheduling requests (SRs), and various other suitable types of information.

[0070] As illustrated in FIG. 5, the end of the DL data portion 504 may be separated in time from the beginning of the common UL portion 506. This time separation may sometimes be referred to as a gap, a guard period, a guard interval, and / or various other suitable terms. This separation provides time for the switch-over from DL communication (e.g., reception operation by the subordinate entity (e.g., UE)) to UL communication (e.g., transmission by the subordinate entity (e.g., UE)). One of LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT ordinary skill in the art will understand that the foregoing is merely one example of a DL-centric slot and alternative structures having similar features may exist without necessarily deviating from the aspects described herein.

[0071] FIG. 6 is a diagram 600 showing an example of an UL-centric slot. The UL-centric slot may include a control portion 602. The control portion 602 may exist in the initial or beginning portion of the UL-centric slot. The control portion 602 in FIG. 6 may be similar to the control portion 502 described above with reference to FIG. 5. The UL-centric slot may also include an UL data portion 604. The UL data portion 604 may sometimes be referred to as the pay load of the UL-centric slot. The UL portion may refer to the communication resources utilized to communicate UL data from the subordinate entity (e.g., UE) to the scheduling entity (e.g., UE or BS). In some configurations, the control portion 602 may be a physical DL control channel (PDCCH).

[0072] As illustrated in FIG. 6, the end of the control portion 602 may be separated in time from the beginning of the UL data portion 604. This time separation may sometimes be referred to as a gap, guard period, guard interval, and / or various other suitable terms. This separation provides time for the switch-over from DL communication (e.g., reception operation by the scheduling entity) to UL communication (e.g., transmission by the scheduling entity). The UL-centric slot may also include a common UL portion 606. The common UL portion 606 in FIG. 6 may be similar to the common UL portion 506 described above with reference to FIG. 5. The common UL portion 606 may additionally or alternatively include information pertaining to channel quality indicator (CQI), sounding reference signals (SRSs), and various other suitable types of information. One of ordinary skill in the art will understand that the foregoing is merely one example of an UL-centric slot and alternative structures having similar features may exist without necessarily deviating from the aspects described herein.

[0073] In some circumstances, two or more subordinate entities (e.g., UEs) may communicate with each other using sidelink signals. Real-world applications of such sidelink communications may include public safety, proximity services, UE-to- network relaying, vehicle-to-vehicle (V2V) communications, Internet of Everything (loE) communications, loT communications, mission-critical mesh, and / or various other suitable applications. Generally, a sidelink signal may refer to a signal communicated from one subordinate entity (e.g., UE1) to another subordinate entity LL Docket 1007576.315WO1 , , toMTK Ref. No. MUSI-24-0043PCT(e.g., UE2) without relaying that communication through the scheduling entity (e.g., UE or BS), even though the scheduling entity may be utilized for scheduling and / or control purposes. In some examples, the sidelink signals may be communicated using a licensed spectrum (unlike wireless local area networks, which typically use an unlicensed spectrum).

[0074] In the 5G NR standard, Al is being explored to enhance the air interface, which is the communication link between devices (e.g., UEs) and base stations (e.g., gNBs). Channel State Information (CSI) compression is a study item in NR Release 18.

[0075] FIG. 7 is a diagram illustrating a CSI compression cycle in an AI / ML model between a UE and a base station. As shown in FIG. 7, the CSI compression cycle 700 involves processes on both the UE 710 and the gNB 720. On the UE 710, the process includes: (1) channel estimation, in which the channel is measured / estimated by the UE to obtain the CSI; (2) pre-processing, in which an optional pre-processor 712 of the UE 710 translates the CSI obtained to an intermediate domain; (3) compression, in which the pre-processed CSI is compressed by an AI / ML-based encoder 714 to be fed back to the gNB 720; and (4) quantization, in which the compressed CSI is quantized by a quantizer 716 into a bit stream, such that the quantized CSI may be transmitted by the UE 710 to the gNB 720. On the gNB 720, the process correspondingly includes: (1) de-quantization, in which the quantized CSI is de-quantized or recovered by a dequantizer 722 from the bit stream; (2) de-compression, in which the CSI feedback is de-compressed by an AEML-based decoder 724; (3) post-processing, in which an optional post-processor 726, if needed, reverts the effects of the pre-processing by the UE 710 on spatial information; and (4) precoding, in which the reconstructed CSI is used for precoding purposes in the transmission. The CSI compression cycle 700 ensures efficient feedback and accurate signal reconstruction for enhanced communication performance.

[0076] Quantization in the CSI compression involves discretizing a continuous space into a finite number of representative points, and each representative point is indexed with a finite number of bits. In certain configurations, the quantization may serve to compactly represent latent vectors during the forward pass of training and inference stages, to compactly represent gradient vectors during backpropagation of the training stage, and to facilitate data collection by lowering the overhead through compacting CSI samples.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0077] In certain configurations, the quantization method used in the CSI compression may include the scalar quantization (SQ) method and the vector quantization (VQ) method. In the SQ method, each element on a vector (e.g., a latent vector or a gradient vector) is mapped to a discretized value, and each discretized element is represented by a finite number of bits.

[0078] FIG. 8 is a diagram illustrating an example uniform scalar quantization method. As shown in FIG. 8, when the uniformed SQ 800 is used, the encoder 810 converts the CSI 805 to a vector 820 (e.g., a latent vector), and each element on the vector 820 is mapped to a discretized value 830, with each discretized element 830 being represented by a finite number of bits. The input (i.e., the elements of the vector 820) and the output (i.e., the discretized elements 830) may be shown in a two-dimensional space, where the continued space of the input is uniformly discretized within the input range 840 and the output range 850. In this case, the uniformed SQ method may be described by SQ descriptors representing all aspects of SQ. Specifically, the SQ descriptors include the output levels, the input intervals, the length of each input interval, the step size of the output levels, the number of the input intervals, the number of the output levels, the input range and the output range.

[0079] In certain configurations, the uniformed SQ method may be described by a subset of the SQ descriptors, without using all of the SQ descriptors. For example, in one embodiment, the uniform SQ may be described with the descriptors of the quantization input intervals and the quantization output levels, which fully describe the I / O relations. In another embodiment, the uniform SQ may be described with the descriptors of the input intervals, the number of the output levels, and the output range. In a further embodiment, the uniform SQ may be described with the descriptors of the output levels, the number of the input intervals, and the input range. In yet another embodiment, the uniform SQ may be described with the descriptors of the numbers of the output levels and the input intervals, the input range and the output range. In yet a further embodiment, the uniform SQ may be described with the descriptors of the length of each input interval, the step size of the output levels, the input range and the output range. The description of the uniform SQ in each embodiment may be used as the configuration information of the quantization method to be exchanged between the UE and the base station (e.g., gNB). Thus, the UE and / or the base station may fully establish the desired uniform SQ upon receiving one of the descriptions.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0080] FIG. 9 is a diagram illustrating an example non-uniform scalar quantization method described by a direct specification of the I / O relation. Specifically, in comparison with the uniform SQ 800 as shown in FIG. 8, the non-uniform SQ 900 in the two- dimensional space is discretized in a non-uniform way within the input range 840 and the output range 850. In this case, each input interval and each output level may have unique values of the SQ descriptors. Thus, the non-uniform SQ 900 may be described by a direct specification of the VO relation, with the descriptors of the quantization input intervals and the quantization output levels that fully describe the I / O relations.

[0081] FIG. 10 is a diagram illustrating an example non-uniform scalar quantization method described by shaping I / O of a uniform scalar quantization. In comparison to the non- uniform SQ 900, which is described by a direct specification of the I / O relation, the non-uniform SQ is described by shaping the I / O relation of the non-uniform SQ 1010 to a uniform SQ 1020 with a compressing function (-,0) 1030 and an expanding function1040. In certain configurations, examples of the compressing and expanding functions may include, without being limited thereto, p-Law functions, A- Law functions, Piece-wise linear functions. Similar to the uniform SQ 800 as shown in FIG. 8, the uniform SQ 1020 may be described by a subset of the SQ descriptors.

[0082] FIG. 11 is a diagram illustrating an example vector quantization method. In the VQ method, a vector (e.g., a latent vector or a gradient vector) is mapped to a new discretized vector. Specifically, in the VQ method, the VQ 1100 is described by a codebook (CB) and a quantization rule, such that when the encoder 1110 converts the CSI samples 1105 to a vector 1120 (e.g., a latent vector), the vector 1120 may be mapped to one of pre-defined vectors 1130 in the CB based on the quantization rule. In certain configurations, the quantization rule may include factors such as the distance, the min / max distance over a dim, etc.

[0083] In certain configurations, the challenges of the VQ method as shown in FIG. 11 exist in that the codebook design for VQ is computationally expensive, and its complexity increases with the dimension of the input vectors. Furthermore, the number of the CSI samples in the training dataset should exceed the number of the representative points. Therefore, a segmented VQ framework is proposed in order to help with dimension reduction, thus relaxing the excessive need for training data.

[0084] FIG. 12 is a diagram illustrating an example segmented vector quantization method. Specifically, in the segmented VQ method 1200, in the VQ method, when the encoder 1210 converts the CSI samples 1205 to a vector 1220 (e.g., a latent vector), the vector LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT1220 is firstly broken down into N smaller segments 1230, and each segment 1230 is then mapped to one of pre-defined segmented vector (also referred to as the “codeword”) 1240 in a corresponding CB (e.g., CB z, where z = 1, 2, . . . , N) based on a corresponding quantization rule (e.g., quantization rule z, where z = 1, 2, . . . , N). In certain configurations, in the training stage, the segmented VQ method 1200 involves designing the corresponding VQ codebook for each segment of the training vectors in order to obtain the corresponding CB. In the inference stage, the same segmentation is applied to the inference vector, such that each segment 1230 to the inference vector may be assigned with a corresponding codeword 1240 based on the corresponding CB and quantization rule. The codewords 1240 of all segments 1230 are then concatenated to obtain a final codeword. Thus, the final codeword may be mapped into a bit steam as the quantized data.

[0085] In the segmented VQ method, the segmented VQ 1200 may be described by the segmentation information of the vectors and the full descriptions of all CBs and quantization rules corresponding to the segments. Specifically, the segmentation information of the vectors may include a start, an end and a length of each segment of the latent vectors.

[0086] In certain configurations, the data collection and training process for an AI / ML model between a UE and a base station (i.e., gNB) may be performed using different training approaches, including (i) the training type I approach, in which the training is performed at a single entity (i.e., either the UE or the gNB) with the model being transferred after the training process; (ii) the training type II approach, in which joint training is performed at different entities (i.e., both UE and gNB) without model transfer; and (iii) the training type III approach, in which sequential separate training is performed (e.g., UE-first or gNB-first).

[0087] FIG. 13 is a diagram illustrating a data collection and training process using a training type I approach for an AI / ML model between a UE and a base station, where the training is performed at the UE. Specifically, in the process 1300, the UE 1302 measures the CSI samples, and uses the CSI samples as input and target CSI samples to perform training of the AI / ML model in the training stage. Once the training is complete, the UE 1302 transmits the decoder of the AI / ML model to the base station 1304 and executes the encoder at the UE 1302 in the inference stage.

[0088] As shown in FIG. 13, at operation 1310, the UE 1302 performs data collection to collect data samples (e.g., CSI). Specifically, the UE 1302 measures / estimates the LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCTCSI in order to obtain the CSI samples. For example, the UE 1302 may estimate the CSI in the latent space to obtain the latent vectors of the CSI samples. It should be noted that the data collection operation 1310 may be a continuous operation to obtain batches of the latent vectors of the CSI samples. At operation 1320, the UE 1302 performs the training process of the whole AI / ML model using the CSI samples in the training stage. In other words, the training process is conducted solely by the UE 1302 throughout the training stage, which leads to a relatively large storage requirement for the UE 1302. In this case, the UE 1302, as the sole entity to perform training the AI / ML model, may determine a latent quantization method being used for the AI / ML model, and configure the quantizer of the AI / ML model during the training stage according to the latent quantization method. In certain configurations, the latent quantization method may be a uniform SQ method, a non-uniform SQ method, a VQ method or a segmented VQ method.

[0089] Once the training process is complete, at operation 1330, the base station 1304 sends a request to the UE 1302 for the decoder of the trained AI / ML model. Upon receiving the request, at operation 1340, the UE 1302 transmits the decoder to the base station 1304. At operation 1350, the base station 1304 deploys the decoder received. Simultaneously, at operation 1360, the UE 1302 transmits the configuration information of the latent quantization method to the base station 1304. Specifically, the configuration information may include the latent description of the latent quantization method (e.g., uniform / non-uniform SQ method, VQ method, segmented VQ method). At operation 1370, the base station 1304 configures the de-quantizer according to the latent description in the configuration information received, such that the quantization / de-quantization of the UE 1302 and the base station 1304 are aligned. Thus, the UE 1302 and the base station 1304 may perform corresponding operations (e.g., CSI compression, image uploading, etc.) using the AI / ML model in the inference stage. For example, at operation 1380, the UE 1302 may generate (e.g., by pre-processing, compression and quantization) quantized CSI samples and transmit the quantized CSI samples to the base station 1304. Upon receiving the quantized CSI samples, at operation 1390, the base station performs corresponding operations (e.g., de-quantization, de-compression, post-processing, etc.) to the quantized CSI samples, and at operation 1395, the base station 1304 transmits corresponding data back to the UE 1302.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0090] In the process 1300, the UE 1302 measures the CSI samples and transmits them to the base station 1304 for training purposes. In certain configurations, the UE 1302 may also measure assistant information and share the assistant information with the base station 1304 as well.

[0091] In the process 1300, the operation 1360 (i.e., the UE 1302 transmitting the configuration information to the base station 1304) is performed after the operations 1330-1350 (i.e., the base station 1304 requesting, receiving and deploying the decoder). However, the operation 1360 may be performed in the inference stage simultaneously with or prior to the operations 1330-1350. In other words, the UE 1302 may transmit the configuration information to the base station 1304 before or simultaneously when transmitting the decoder to the base station 1304.

[0092] FIG. 14 is a diagram illustrating a data collection and training process using a training type I approach for an AI / ML model between a UE and a base station, where the training is performed at the base station. Specifically, in the process 1400, the UE 1402 measures the CSI samples and transmits them to the base station 1404 (e.g., gNB) for training purposes, and the base station 1404 uses the CSI samples received as input and target CSI samples to perform training of the AI / ML model.

[0093] As shown in FIG. 14, at operation 1410, the UE 1402 performs data collection to collect data samples (e.g., CSI). Specifically, the UE 1402 measures / estimates the CSI in order to obtain the CSI samples. For example, the UE 1402 may estimate the CSI in the latent space to obtain the latent vectors of the CSI samples. At operation 1415, the UE 1402 transmits the data samples collected to the base station 1404, such that the base station 1404 collects the data samples. It should be noted that the data collection operation 1410 may be a continuous operation to obtain batches of the latent vectors of the CSI samples, and the UE 1402 may have to transmit each batch of the latent vectors to the base station 1404 in multiple operations 1415, which leads to a large overhead for the base station 1404. At operation 1420, the base station 1404 performs the training process of the whole AI / ML model using the CSI samples in the training stage. In other words, the training process is conducted solely by the base station 1404 throughout the training stage. In this case, the base station 1404, as the sole entity to perform training the AI / ML model, may determine a latent quantization method being used for the AI / ML model, and configure the de-quantizer of the AI / ML model during the training stage according to the latent quantization method. In certainLL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT configurations, the latent quantization method may be a uniform SQ method, a non- uniform SQ method, a VQ method or a segmented VQ method.

[0094] Once the training process is complete, at operation 1430, the UE 1302 sends a request to the base station 1404 for the encoder of the trained AI / ML model. Upon receiving the request, at operation 1440, the base station 1404 transmits the encoder to the UE 1402. At operation 1450, the UE 1402 deploys the encoder received. Simultaneously, at operation 1460, the base station 1404 transmits the configuration information of the latent quantization method to the UE 1402. Specifically, the configuration information may include the latent description of the latent quantization method (e.g., uniform / non-uniform SQ method, VQ method, segmented VQ method). At operation 1470, the UE 1402 configures the quantizer according to the latent description in the configuration information received, such that the quantization / de-quantization of the UE 1402 and the base station 1404 are aligned. Thus, the UE 1402 and the base station 1404 may perform corresponding operations (e.g., CSI compression, image uploading, etc.) using the AI / ML model in the inference stage. For example, at operation 1480, the UE 1402 may generate (e.g., by pre-processing, compression and quantization) quantized CSI samples and transmit the quantized CSI samples to the base station 1404. Upon receiving the quantized CSI samples, at operation 1490, the base station performs corresponding operations (e.g., de-quantization, decompression, post-processing, etc.) to the quantized CSI samples, and at operation 1495, the base station 1404 transmits corresponding data back to the UE 1402.

[0095] In the process 1400, the UE 1402 measures the CSI samples and transmits them to the base station 1404 for training purposes. In certain configurations, the UE 1402 may also measure assistant information and share the assistant information with the base station 1304 as well.

[0096] In the process 1400, the operation 1460 (i.e., the base station 1404 transmitting the configuration information to the UE 1402) is performed after the operations 1430- 1450 (i.e., the UE 1402 requesting, receiving and deploying the encoder). However, the operation 1460 may be performed in the inference stage simultaneously with or prior to the operations 1430-1450. In other words, the base station 1404 may transmit the configuration information to the UE 1402 before or simultaneously when transmitting the encoder to the UE 1402.

[0097] FIG. 15 is a diagram illustrating a data collection and training process using a training type II approach for an AI / ML model between a UE and a base station, where joint LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT training is performed at both the UE and the base station. Specifically, in the process 1500, multiple UEs 1502 (labeled as 1502-1 to 1502-N) and multiple base stations 1504 (labeled as 1504-1 to 1504-N) are provided. In a pre-training data collection stage, the UEs 1502 measure the CSI samples and share them with the base stations 1504 as target CSI samples of the training stage. In the training stage, the UEs 1502 use the collected CSI samples as the input CSI samples (and optionally the assistant information), and the base stations 1504 use the CSI samples as the target CSI samples (and optionally the assistant information) for supervised learning.

[0098] As shown in FIG. 15, at operation 1510, each of the UEs 1502-1 to 1502-N respectively performs data collection to collect data samples (e.g., CSI). Specifically, each UE 1502 measures / estimates the CSI in order to obtain the CSI samples. For example, each UE 1502 may estimate the CSI in the latent space to obtain the latent vectors of the CSI samples. At operation 1515, each UE 1502 respectively transmits the data (e.g., latent vectors of the CSI samples) to the base stations 1504-1 to 1504- N for training purposes, thus completing the pre-training data collection stage. In certain configurations, the UEs 1502 may also measure assistant information and share the assistant information with the base stations 1504 as well.

[0099] Once the base stations 1504-1 to 1504-N, at operation 1520, the training stage may start, in which each UE 1502 uses the collected CSI samples as the input CSI samples to train the corresponding encoder, and each base station 1504 uses the collected CSI samples to train the corresponding decoder. This leads to a large storage requirement for the UEs 1502 and a large overhead for the base stations 1504. Once the training stage is complete, at operation 1530, the inference stage may start, in which the UEs 1502 and the base stations 1504 may perform corresponding operations (e.g., CSI compression, image uploading, etc.) using the AI / ML model in the inference stage. The operations in the inference stage may be similar to the operations 1380 to 1395 in the process 1300 and the operations 1480 to 1495 in the process 1400, and are thus not further elaborated herein.

[0100] FIG. 16 is a diagram illustrating an example data collection process with forward pass and backward propagation procedures in the data collection and training process in FIG. 15. Specifically, in the process 1600, multiple UEs 1602 (labeled as 1602-1 to 1602-N) and multiple base stations 1604 (labeled as 1604-1 to 1604-N) are provided, and the pre-training data collection stage include forward pass (FP) and backpropagation (BP) procedures, allowing the base stations 1604 to collect latent LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT vectors from the UEs 1602 and the UEs 1602 to collect gradient vectors from the base stations 1604.

[0101] As shown in FIG. 16, at operation 1610, the UEs 1602-1 to 1602-N respectively perform FP to collect CSI samples, and estimate the CSI in the latent space to obtain the latent vectors. At operation 1615, the UEs 1602 respectively transmit the latent vectors to the base stations 1604-1 to 1604-N, allowing the base stations 1604 to collect the latent vectors from the UEs 1602. Specifically, each base station 1604 receives a latent vector per every input CSI sample during the FP. At operation 1620, the base stations 1604 respectively perform FP to process the latent vectors collected / received. At operation 1625, the base stations 1604 respectively perform BP to generate gradient vectors. At operation 1630, the base stations 1604 respectively transmit the gradient vectors to the UEs 1602, allowing the UEs 1602 to collect the gradient vectors from the base stations 1604. Specifically, each UE 1604 receives the gradient vectors once per every batch of the latent vectors transmitted in the FP. At operation 1640, the UEs 1602 respectively perform BP by backward propagation of the gradient vectors received and update the corresponding parameters of the AI / ML model once per every gradient vector received. In the process 1600, there is a large overhead in both FP and BP (especially in the FP).

[0102] In the process 1600, both the FP and BP loops cross through two entities (i.e., the UEs 1602 and the base station 1604), leading to the large overhead. Thus, quantization may apply in both latent and gradient spaces.

[0103] The process 1600 as shown in FIG. 16 is an unfrozen training process, in which both entities (i.e., UEs 1602 and base stations 1604) in the training process may update their corresponding parameters in the AI / ML model. In certain configurations, a frozen UE process may be utilized, in which the UEs 1602 are frozen (i.e., parameters at the UE side are not updated). In this case, only the part of the AI / ML model at the base stations 1604 (e.g., decoders) are trained, and the BP procedure may be redundant. In certain configurations, a frozen gNB process may be utilized, in which the base stations 1604 are frozen (i.e., parameters at the gNB side are not updated). In this case, only the part of the AI / ML model at the UEs 1602 (e.g., encoders) are trained.

[0104] FIG. 17 is a diagram illustrating latent quantization and gradient quantization of the forward pass and backward propagation procedures in the data collection and training process in FIG. 16. As shown in FIG. 17, the UE 1702 has the encoder 1710, and the LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT base station 1704 (i.e., gNB) has the decoder 1720. To reduce the overhead in the latent space, latent quantization is applied the FP procedure with the latent quantizer 1730 at the UE 1702 and the latent de-quantizer 1740 at the base station 1704. Similarly, to reduce the overhead in the gradient space, gradient quantization is applied the BP procedure with the gradient quantizer 1750 at the base station 1704 and the gradient de-quantizer 1760 at the UE 1702, such that the data (i.e., gradient vectors) in the BP procedure goes through the latent de-quantizer 1740, the gradient quantizer 1750, the gradient de-quantizer 1760 and the latent quantizer 1730.

[0105] In the process 1700, both latent quantization and gradient quantization apply, and each may adopt a corresponding quantization method (e.g., a uniform SQ method, a non- uniform SQ method, a VQ method or a segmented VQ method). It should be noted that the gradient quantization method is not necessarily the same as the latent quantization method. In other words, the gradient quantization may adopt one quantization method, while the latent quantization may adopt another quantization method. Each of the

[0106] FIG. 18 is a diagram illustrating an example quantification configuration of the UE and the base station in the data collection and training process in FIG. 15. In the process 1800, at operation 1810, the base station 1804 transmits the configuration information of the gradient quantification method (e.g., descriptions of the gradient quantification method) to the UE 1802. At operation 1820, the UE 1802 configures the gradient de-quantizer according to the descriptions of the gradient quantification method. At operation 1830, the UE 1802 transmits the configuration information of the latent quantification method (e.g., descriptions of the latent quantification method) to the base station 1804. At operation 1840, the base station 1804 configures the latent de-quantizer according to the descriptions of the latent quantification method. Once both the latent and gradient quantification methods are configured at both the UE 1802 and the base station 1804, the UE 1802 and the base station 1804 may perform the FP procedure at operation 1850 and then the BP procedure at operation 1860.

[0107] In the process 1800, the configuration of the gradient quantification method (i.e., operations 1810 and 1820) is performed prior to the configuration of the latent quantification method (i.e., operations 1830 and 1840). In certain configurations, the configurations of the latent and gradient quantification methods may be performed concurrently or in any sequential order, as long as both latent and gradient quantifications are configured prior to the FP / BP procedures. LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0108] FIG. 19 is a diagram illustrating a data collection and training process using a training type III approach for an AI / ML model between a UE and a base station, where sequential separate training is performed firstly at the UE and then at the base station. Specifically, in the process 1900, multiple UEs 1902 (labeled as 1902-1 to 1902-N) and multiple base stations 1904 (labeled as 1904-1 to 1904-N) are provided, with the UEs 1902 performing the data collection and training stages firstly, and then generating corresponding data for the base stations 1904 to perform training.

[0109] As shown in FIG. 19, at operation 1910, each of the UEs 1902-1 to 1902-N respectively performs data collection to collect data samples (e.g., CSI). Specifically, each UE 1902 measures / estimates the CSI in order to obtain the CSI samples. For example, each UE 1902 may estimate the CSI in the latent space to obtain the latent vectors of the CSI samples. Optionally, the UEs 1902 may also measure assistant information. At operation 1920, each UE 1902 respectively performs training for the corresponding AI / ML model (particularly the encoder, as the encoder will run on the UE 1902). In this case, each UE 1902, as the sole entity to perform training the AI / ML model, may determine a latent quantization method being used for the AI / ML model, and configure the latent quantizer of the AI / ML model during the training stage (operation 1920) according to the latent quantization method. In certain configurations, the latent quantization method may be a uniform SQ method, a non- uniform SQ method, a VQ method or a segmented VQ method.

[0110] At operation 1930, each UE 1902 generates a dataset including the target CSI samples and the corresponding latent vectors. This leads to a large storage requirement in the data collection (operation 1910) and data generation (operation 1930) stages for the UEs 1902. At operation 1940, each UE 1902 respectively transmits the dataset (e.g., the latent vectors of the CSI samples) generated to the base stations 1904-1 to 1904- N for training purposes, allowing the base stations 1904 to collect the datasets from the UEs 1902. This leads to a large overhead in the data collection (operation 1940) stage for the base stations 1904. At operation 1950, each base station 1904 performs training of the decoder of the AI / ML model, thus completing the whole training stage for both the UEs 1902 and the base stations 1904. At operation 1960, the inference stage may start, in which the UEs 1902 and the base stations 1904 may perform corresponding operations (e.g., CSI compression, image uploading, etc.) using the AI / ML model in the inference stage. The operations in the inference stage may beLL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT similar to the operations 1380 to 1395 in the process 1300 and the operations 1480 to 1495 in the process 1400, and are thus not further elaborated herein.

[0111] FIG. 20 is a diagram illustrating an example quantification configuration of the UE and the base station in the data collection and training process in FIG. 19. In the process 2000, at operation 2010, the UE 2002 trains the encoder of the AI / ML model. Specifically, in the training process, the UE 2002 configures the latent quantizer with a specific latent quantification method. At operation 2020, the UE 2002 transmits the configuration information of the latent quantification method (e.g., descriptions of the latent quantification method) to the base station 2004. At operation 2030, the base station 2004 configures the latent de-quantizer according to the descriptions of the latent quantification method. Once the latent quantification method is configured at both the UE 2002 and the base station 2004, at operation 2040, the UE 2002 may transmit the dataset (i.e., target latent vectors of the CSI samples) to the base station 2004. At operation 2050, the base station 2004 may perform training to the decoder with the dataset.

[0112] FIG. 21 is a diagram illustrating an example quantification configuration of two UEs and the base station in the data collection and training process in FIG. 19. Specifically, in the process 2100, two UEs 2102 and 2106 as well as a base station 2104 are provided, allowing the two UEs 2102 and 2106 to configure respective latent quantification methods for the base station 2104.

[0113] As shown in FIG. 21 , at operation 2010, the UE 2102 trains the encoder of the AI / ML model. Specifically, in the training process, the UE 2102 configures the latent quantizer with a specific latent quantification method. Similarly, at operation 2120, the UE 2106 trains the encoder of the AI / ML model and configures the latent quantizer with a specific latent quantification method. It should be noted that the latent quantification methods adopted by the UEs 2102 and 2106 may be the same quantification method, or may be different quantification methods.

[0114] At operation 2130, the UE 2102 transmits the configuration information of the latent quantification method (e.g., descriptions of the latent quantification method) to the base station 2104. At operation 2140, the base station 2104 configures the latent dequantizer for the UE 2102 according to the descriptions of the latent quantification method. Similarly, at operation 2150, the UE 2106 transmits the configuration information of the latent quantification method (e.g., descriptions of the latent quantification method) to the base station 2104. At operation 2160, the base station LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT2104 configures the latent de-quantizer for the UE 2106 according to the descriptions of the latent quantification method.

[0115] Once the latent quantification methods are configured at both the UE 2102 and the base station 2104 and both the UE 2106 and the base station 2104, at operation 2170, the UE 2102 may transmit the dataset (i.e., latent vectors of the CSI samples) to the base station 2104. Similarly, at operation 2180, the UE 2106 may transmit the dataset (i.e., latent vectors of the CSI samples) to the base station 2104. Upon receiving all the datasets from the UEs 2102 and 2106, at operation 2190, the base station 2104 may perform training procedure to train the decoder.

[0116] FIG. 22 is a diagram illustrating a data collection and training process using a training type III approach for an AI / ML model between a UE and a base station, where sequential separate training is performed firstly at the base station and then at the UE. Specifically, in the process 2200, multiple UEs 2202 (labeled as 2202-1 to 2202-N) and multiple base stations 2204 (labeled as 2204-1 to 2204-N) are provided, with the base stations 2204 performing the data collection and training stages firstly, and then generating corresponding data for the UEs 2202 to perform training.

[0117] As shown in FIG. 22, at operation 2210, each of the UEs 2202-1 to 2202-N respectively performs data collection to collect data samples (e.g., CSI). Specifically, each UE 2202 measures / estimates the CSI in order to obtain the CSI samples. For example, each UE 2202 may estimate the CSI in the latent space to obtain the latent vectors of the CSI samples. Optionally, the UEs 2202 may also measure assistant information. At operation 2215, the UEs 2202 respectively transmit the data to the base stations 2204-1 to 2204-N, allowing the base stations to collect the data (i.e., latent vectors of the CSI samples). This leads to a large overhead in the data collection (operation 2210 for the UEs 2202, and operation 2215 for the base stations 2204) stages. At operation 2220, each base station 2204 respectively performs training for the corresponding AI / ML model (particularly the decoder, as the decoder will run on each base station 2204). In this case, each base station 2204, as the sole entity to perform training the AI / ML model, may determine a latent quantization method being used for the AI / ML model, and configure the latent de-quantizer of the AI / ML model during the training stage (operation 2220) according to the latent quantization method. In certain configurations, the latent quantization method may be a uniform SQ method, a non-uniform SQ method, a VQ method or a segmented VQ method.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0118] At operation 2230, each base station 2204 generates a dataset including the input CSI samples and the corresponding latent vectors. This leads to a large storage requirement in the data generation (operation 2230) stages for the base stations 2204. At operation 2240, each base station 2204 respectively transmits the dataset (e.g., the latent vectors of the CSI samples) generated to the UEs 2202 for training purposes, allowing the UEs 2202 to collect the datasets from the base stations 2204. This leads to a large overhead in the data collection (operation 2240) stage for the UEs 2202. At operation 2250, each UE 2202 performs training of the encoder of the AI / ML model, thus completing the whole training stage for both the UEs 2202 and the base stations 2204. At operation 2260, the inference stage may start, in which the UEs 2202 and the base stations 2204 may perform corresponding operations (e.g., CSI compression, image uploading, etc.) using the AI / ML model in the inference stage. The operations in the inference stage may be similar to the operations 1380 to 1395 in the process 1300 and the operations 1480 to 1495 in the process 1400, and are thus not further elaborated herein.

[0119] FIG. 23 is a diagram illustrating an example quantification configuration of the UE and the base station in the data collection and training process in FIG. 22. In the process 2300, at operation 2310, the base station 2304 trains the decoder of the AI / ML model. Specifically, in the training process, the base station 2304 configures the latent de-quantizer with a specific latent quantification method. At operation 2320, the base station 2304 transmits the configuration information of the latent quantification method (e.g., descriptions of the latent quantification method) to the UE 2302. At operation 2330, the UE 2302 configures the latent quantizer according to the descriptions of the latent quantification method. Once the latent quantification method is configured at both the UE 2302 and the base station 2304, at operation 2340, the base station 2304 may transmit the dataset (i.e., input latent vectors of the CSI samples) to the UE 2302. At operation 2350, the UE 2302 may perform training to the decoder with the dataset.

[0120] FIG. 24 is a table showing the configuring entity for the quantization method in each of the training type approaches when the encoder is at the UE and the decoder is at the base station. Specifically, as described in the aforementioned embodiments, the UE has the encoder and the base station (i.e., gNB) has the decoder. Thus, in each training type approach, the first entity using the quantization method in the training procedure is the entity responsible for configuring the quantization. Thus, the entity LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT responsible for configuring the quantization provides a full description of the quantization method to the other entity in order to align the entities. As shown in FIG. 24, in the training type 1 and the training type 3 approaches, only latent quantization is performed. Thus, the UE is the configuring entity when the UE firstly performs the training process (e.g., UE-training in training type 1 and UE-first in training type 3), and the base station is the configuring entity when the base station (gNB) firstly performs the training process (e.g., gNB-training in training type 1 and gNB-first in training type 3). In the training type 2 approach, if unfrozen training is performed, the UE is the configuring entity for latent quantization, and the base station (gNB) is the configuring entity for gradient quantization. If training is performed in the frozen UE approach, the UE is the configuring entity for latent quantization, and there is no gradient quantization (since the UE is frozen). If training is performed in the frozen gNB approach, the base station (gNB) is the configuring entity for both latent and gradient quantization (since the gNB is frozen).

[0121] Although the aforementioned embodiments all recite the encoder at the UE and the decoder at the base station, in certain configurations, it is possible that the UE has the decoder and the base station (i.e., gNB) has the encoder.

[0122] FIG. 25 is a table showing the configuring entity for the quantization method in each of the training type approaches when the encoder is at the base station and the decoder is at the UE. Thus, in each training type approach, the first entity using the quantization method in the training procedure is the entity responsible for configuring the quantization. Thus, the entity responsible for configuring the quantization provides a full description of the quantization method to the other entity in order to align the entities. As shown in FIG. 25, in the training type 1 and the training type 3 approaches, only latent quantization is performed. Thus, the UE is the configuring entity when the UE firstly performs the training process (e.g., UE-training in training type 1 and UE-first in training type 3), and the base station is the configuring entity when the base station (gNB) firstly performs the training process (e.g., gNB-training in training type 1 and gNB-first in training type 3). In the training type 2 approach, if unfrozen training is performed, the base station (gNB) is the configuring entity for both latent and gradient quantization. If training is performed in the frozen UE approach, the UE is the configuring entity for latent quantization, and there is no gradient quantization (since the UE is frozen). If training is performed in the frozen gNB approach, the base station (gNB) is the configuring entity for both latent and LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT gradient quantization (since the gNB is frozen). In this case, the only difference exists in the training type 2 unfrozen approach, where the base station (gNB) serves as the configuring entity for both latent and gradient quantization.

[0123] FIG. 26 is a flow chart of a method (process) for wireless communication of a UE. The method may be performed by a UE (e.g., UE 710, 1302, 1402, 1502, 1602, 1702, 1802, 1902, 2002, 2102, 2202 or 2302). At procedure 2610, the UE collects data samples for training an AI / ML model at the UE and a base station. The AI / ML model is trained at the UE or at the base station in a training stage. At procedure 2620, the UE performs, according to a quantization method, quantization of the data samples to obtain quantized data samples. At procedure 2630, the UE transmits the quantized data samples to the base station. At procedure 2640, the UE executes an encoder or a decoder of the trained AI / ML model in an inference stage.

[0124] In certain embodiments, the UE compresses, by the encoder, the CSI samples to obtained compressed CSI samples, and the quantized data samples are obtained by performing quantization of the compressed CSI samples.

[0125] In certain embodiments, when the UE is the configuring entity for the latent quantization method, the UE configures the latent quantization method at the UE. The UE transmits configuration information of the latent quantization method to the base station.

[0126] In certain embodiments, when the UE is not the configuring entity for the latent quantization method, the UE receives, from the base station, configuration information of the latent quantization method. The UE configures the latent quantization method according to the configuration information received from the base station.

[0127] In certain embodiments, the UE further receives, from the base station, configuration information of a gradient quantization method. The UE configures the gradient quantization method for further training the AI / ML model with gradient vectors according to the configuration information of the gradient quantization method received from the base station. The UE receives, from the base station, quantized gradient vectors. The UE performs, according to the gradient quantization method, de-quantization of the quantized gradient vectors to obtain the gradient vectors. The UE performs BP with gradient vectors and updating parameters of the Al or ML model.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0128] FIG. 27 is a flow chart of a method (process) for wireless communication of a base station. The method may be performed by a base station (e.g., gNB, base station 720, 1304, 1404, 1504, 1604, 1704, 1804, 1904, 2004, 2104, 2204 or 2304). At procedure 2710, the base station receives, from a UE, quantized data samples. At procedure 2720, the base station performs, according to a quantization method, de-quantization of the quantized data samples to obtain de-quantized data samples for training of an AI / ML model at the UE or the base station in a training stage. At procedure 2730, the base station executes an encoder or a decoder of the trained AI / ML model in an inference stage.

[0129] In certain embodiments, when the base station is not the configuring entity for the latent quantization method, the base station receives, from the UE, configuration information of the latent quantization method. The base station configures the latent quantization method according to the configuration information received from the UE.

[0130] In certain embodiments, when the base station is the configuring entity for the latent quantization method, the base station configures the latent quantization method at the base station. The base station transmits configuration information of the latent quantization method to the UE.

[0131] In certain embodiments, the base station configures a gradient quantization method. The base station transmits configuration information of the gradient quantization method to the UE. The base station performs BP to obtain gradient vectors for further training the AI / ML model. The base station performs, according to the gradient quantization method, quantization of the gradient vectors to obtain quantized gradient vectors. The base station transmits the quantized gradient vectors to the UE.

[0132] In certain embodiments, the latent quantization method and / or the gradient quantization method may be a uniform SQ method, a non-uniform SQ method, a VQ method or a segmented VQ method.

[0133] It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed is an illustration of exemplary approaches. Based upon design preferences, it is understood that the specific order or hierarchy of blocks in the processes / flowcharts may be rearranged. Further, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order, and are not meant to be limited to the specific order or hierarchy presented.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT

[0134] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language claims, wherein reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more.” The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects. Unless specifically stated otherwise, the term “some” refers to one or more. Combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ include any combination of A,B, and / or C, and may include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as “at least one of A, B, or C,” “one or more of A, B, or C,” “at least one of A, B, and C,” “one or more of A, B, and C,” and “A, B, C, or any combination thereof’ may be A only, B only, C only, A and B, A and C, B andC, or A and B and C, where any such combinations may contain one or more member or members of A, B, or C. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims. The words “module,” “mechanism,” “element,” “device,” and the like may not be a substitute for the word “means.” As such, no claim element is to be construed as a means plus function unless the element is expressly recited using the phrase “means for.”LL Docket 1007576.315WO1

Claims

MTK Ref. No. MUSI-24-0043PCTCLAIMSWHAT IS CLAIMED IS:

1. A method of wireless communication of a user equipment (UE), comprising: collecting data samples for an artificial intelligence (AI) / machine learning (ML) model at the UE and a base station, wherein the AI / ML model is trained at the UE or at the base station in a training stage; performing, according to a quantization method, quantization of the data samples to obtain quantized data samples; transmitting the quantized data samples to the base station; and executing an encoder or a decoder of the trained AI / ML model in an inference stage.

2. The method of claim 1, wherein the quantization method is a latent quantization method for a latent space, and the data samples are latent vectors of Channel State Information (CSI) samples measured by the UE.

3. The method of claim 2, further comprising: compressing, by the encoder, the CSI samples to obtained compressed CSI samples, wherein the quantized data samples are obtained by performing quantization of the compressed CSI samples.

4. The method of claim 2, further comprising: configuring the latent quantization method at the UE; and transmitting configuration information of the latent quantization method to the base station.

5. The method of claim 2, further comprising: receiving, from the base station, configuration information of the latent quantization method; and configuring the latent quantization method according to the configuration information received from the base station.LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT6. The method of claim 2, wherein the latent quantization method is a uniform scalar quantization (SQ) method described by a subset of SQ descriptors, and the SQ descriptors include: output levels, input intervals, a length of each of the input intervals, a step size of the output levels, a number of the input intervals, a number of the output levels, an input range, and an output range.

7. The method of claim 2, wherein the latent quantization method is a non-uniform scalar quantization (SQ) method described by a direct specification of an input / output (I / O) relation or a shaping I / O of a uniform SQ, wherein the direction specification of the I / O relation includes quantization input intervals and quantization output levels, and wherein the shaping I / O of the uniform SQ includes: a compressing function f (-,0); an expanding functiona description of the uniform SQ described by a subset of SQ descriptors.

8. The method of claim 2, wherein the latent quantization method is a vector quantization (VQ) method described by: a codebook (CB); and a quantization rule, wherein each of the latent vectors is mapped to one of pre-defined vectors in the CB based on the quantization rule.

9. The method of claim 2, wherein the latent quantization method is a segmented vector quantization (VQ) method described by: segmentation information of the latent vectors; a plurality of codebooks (CBs); and a plurality of quantization rules, LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT wherein for each of the latent vectors, each of segments is mapped to one of predefined codeword in a corresponding one of the CBs based on a corresponding one of the quantization rules.

10. The method of claim 2, further comprising: receiving, from the base station, configuration information of a gradient quantization method; configuring the gradient quantization method for further training the AI / ML model with gradient vectors according to the configuration information of the gradient quantization method received from the base station; receiving, from the base station, quantized gradient vectors; performing, according to the gradient quantization method, de-quantization of the quantized gradient vectors to obtain the gradient vectors; and performing backpropagation (BP) with the gradient vectors and updating parameters of the Al or ML model.

11. The method of claim 10, wherein the gradient quantization method is a uniform scalar quantization (SQ) method, a non-uniform SQ method, a vector quantization (VQ) method or a segmented VQ method.

12. A method of wireless communication of a base station, comprising: receiving, from a user equipment (UE), quantized data samples; performing, according to a quantization method, de-quantization of the quantized data samples to obtain de-quantized data samples for of an artificial intelligence (AI) / machine learning (ML) model at the UE and the base station, wherein the AI / ML model is trained at the UE or at the base station in a training stage; and executing an encoder or a decoder of the trained AI / ML model in an inference stage.

13. The method of claim 12, wherein the quantization method is a latent quantization method for a latent space, and the data samples are latent vectors of Channel State Information (CSI) samples measured by the UE.

14. The method of claim 13, further comprising:LL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT receiving, from the UE, configuration information of the latent quantization method; and configuring the latent quantization method according to the configuration information received from the UE.

15. The method of claim 13, further comprising: configuring the latent quantization method at the base station; and transmitting configuration information of the latent quantization method to the UE.

16. The method of claim 13, wherein the latent quantization method is a uniform scalar quantization (SQ) method described by a subset of SQ descriptors, and the SQ descriptors include: output levels, input intervals, a length of each of the input intervals, a step size of the output levels, a number of the input intervals, a number of the output levels, an input range, and an output range.

17. The method of claim 13, wherein the latent quantization method is a non- uniform scalar quantization (SQ) method described by a direct specification of an input / output (EO) relation or a shaping EO of a uniform SQ, wherein the direction specification of the EO relation includes quantization input intervals and quantization output levels, and wherein the shaping EO of the uniform SQ includes: a compressing function f (-,0); descriptions of an expanding functionand a description of the uniform SQ described by a subset of SQ descriptors.

18. The method of claim 13, wherein the latent quantization method is a vector quantization (VQ) method described by: a codebook (CB); andLL Docket 1007576.315WO1MTK Ref. No. MUSI-24-0043PCT a quantization rule, wherein each of the latent vectors is mapped to one of pre-defined vectors in the CB based on the quantization rule.

19. The method of claim 13, wherein the latent quantization method is a segmented vector quantization (VQ) method described by: segmentation information of the latent vectors; a plurality of codebooks (CBs); and a plurality of quantization rules, wherein for each of the latent vectors, each of segments is mapped to one of predefined codeword in a corresponding one of the CBs based on a corresponding one of the quantization rules.

20. The method of claim 13, further comprising: configuring a gradient quantization method; transmitting configuration information of the gradient quantization method to the UE; performing backpropagation (BP) to obtain gradient vectors for further training the AI / ML model; performing, according to the gradient quantization method, quantization of the gradient vectors to obtain quantized gradient vectors; and transmitting the quantized gradient vectors to the UE, wherein the gradient quantization method is a uniform scalar quantization (SQ) method, a non-uniform SQ method, a vector quantization (VQ) method or a segmented VQ method.LL Docket 1007576.315WO1