Systems and methods for distributed federated training collaboration for channel state information feedback

WO2026169265A1PCT designated stage Publication Date: 2026-08-13APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-02
Publication Date
2026-08-13

Smart Images

  • Figure US2025027458_13082026_PF_FP_ABST
    Figure US2025027458_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for distributed federated training collaboration for channel state information (CSI) feedback are described herein. A base station receives, from a user equipment (UE), a bitstream generated by an encoder at the UE using dataset CSI; provides the bitstream as an input to a decoder at the base station to generate a reconstructed CSI; performs a loss function calculation for the decoder based on differences between the dataset CSI and the reconstructed CSI; identifies, based on the loss function calculation, decoder gradients for the decoder, wherein the decoder gradients include UE-specific decoder gradients corresponding to the UE; adjusts the decoder based on the decoder gradients; and sends, to the UE, the UE-specific decoder gradients. Such operations as performed across sets of multiple UEs are described. Corresponding UE behaviors, including the calculation and application of UE-specific encoder gradients using the provided UE-specific decoder gradients, are also described.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR DISTRIBUTED FEDERATED TRAINING COLLABORATION FOR CHANNEL STATE INFORMATION FEEDBACKTECHNICAL FIELD

[0001] This application relates generally to wireless communication systems, including wireless communication systems employing the use of AI / ML models to perform CSI compression / decompression.BACKGROUND

[0002] Wireless mobile communication technology uses various standards and protocols to transmit data between a base station and a wireless communication device. Wireless communication system standards and protocols can include, for example, 3rd Generation Partnership Project (3GPP) Long Term Evolution (LTE) (e.g., 4G), 3GPP New Radio (NR) (e.g., 5G), and Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard for Wireless Local Area Networks (WLAN) (commonly known to industry groups as Wi-Fi®).

[0003] As contemplated by the 3GPP, different wireless communication systems' standards and protocols can use various radio access networks (RANs) for communicating between a base station of the RAN (which may also sometimes be referred to generally as a RAN node, a network node, or simply a node) and a wireless communication device known as a user equipment (UE). 3GPP RANs can include, for example. Global System for Mobile communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE) RAN (GERAN). Universal Terrestrial Radio Access Network (UTRAN), Evolved Universal Terrestrial Radio Access Network (E-UTRAN), and / or Next-Generation Radio Access Network (NG-RAN).

[0004] Each RAN may use one or more radio access technologies (RATs) to perform communication between the base station and the UE. For example, the GERAN implements GSM and / or EDGE RAT, the UTRAN implements Universal Mobile Telecommunication System (UMTS) RAT or other 3 GPP RAT, the E-UTRAN implements LTE RAT (sometimes simply referred to as LTE). and NG-RAN implements NR RAT (sometimes referred to herein as 5G RAT, 5GNR RAT, or simply NR). In certain deployments, the E-UTRAN may also implement NR RAT. In certain deployments, NG-RAN may also implement LTE RAT.14920-2526-7773.1 P70625WO1

[0005] A base station used by a RAN may correspond to that RAN. One example of an E-UTRAN base station is an Evolved Universal Terrestrial Radio Access Network (E-UTRAN) Node B (also commonly denoted as evolved Node B, enhanced Node B, eNodeB, or eNB). One example of an NG-RAN base station is a next generation Node B (also sometimes referred to as a g Node B or gNB).

[0006] A RAN provides its communication services with external entities through its connection to a core network (CN). For example, E-UTRAN may utilize an Evolved Packet Core (EPC) while NG-RAN may utilize a 5G Core Network (5GC).

[0007] Frequency bands for 5G NR may be separated into two or more different frequency ranges. For example, Frequency Range 1 (FR1) may include frequency bands operating in sub-6 gigahertz (GHz) frequencies, some of which are bands that may be used by previous standards, and may potentially be extended to cover new spectrum offerings from 410 megahertz (MHz) to 7125 MHz. Frequency Range 2 (FR2) may include frequency bands from 24.25 GHz to 52.6 GHz. Note that in some systems, FR2 may also include frequency bands from 52.6 GHz to 71 GHz (or beyond). Bands in the millimeter wave (mmWave) range of FR2 may have smaller coverage but potentially higher available bandwidth than bands in FR1. Skilled persons will recognize these frequency ranges, which are provided by way of example, may change from time to time or from region to region.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0008] To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number refer to the figure number in which that element is first introduced.

[0009] FIG. 1 illustrates a diagram showing an example of a two-sided AI / ML model for CSI compression and decompression, according to embodiments discussed herein.

[0010] FIG. 2 illustrates a diagram for an inter-vendor training collaboration procedure for training an AI / ML model for CSI compression / decompression that relies on standardized data / a standardized dataset format and a dataset exchange between a base station and a UE.

[0011] FIG. 3 illustrates a diagram for a mechanism for training collaboration with distributed gradient exchange.24920-2526-7773.1 P70625WO1

[0012] FIG. 4 illustrates a method of a base station, according to embodiments discussed herein.

[0013] FIG. 5 illustrates a method of a UE, according to embodiments discussed herein.

[0014] FIG. 6 illustrates an example architecture of a wireless communication system, according to embodiments disclosed herein.

[0015] FIG. 7 illustrates a system for performing signaling between a wireless device and a network device, according to embodiments disclosed herein.DETAILED DESCRIPTION

[0016] Various embodiments are described with regard to a UE. However, reference to a UE is merely provided for illustrative purposes. The example embodiments may be utilized with any electronic component that may establish a connection to a network and is configured with the hardware, software, and / or firmware to exchange information and data with the network. Therefore, the UE as described herein is used to represent any appropriate electronic component.

[0017] Downlink channel state information (CSI) may be sent from a UE to a base station through feedback channels. The base station may use the CSI feedback to, for example, reduce interference and increase throughput for massive multiple input multiple output (MIMO) communication.

[0018] It may be understood that such feedback represents a relatively large amount of signaling overhead. In various wireless communication systems, vector quantization or codebook-based feedback may be used with the expectation of reducing this overhead. The feedback quantities resulting from these approaches, however, scale linearly with the number of transmit antennas. Accordingly, these approaches may still, for some cases (e.g., when hundreds or thousands of centralized or distributed transmit antennas are used) represent a high level of signaling overhead.

[0019] Accordingly, in various wireless communication systems, artificial intelligence (AI) / machine learning (ML)-based CSI encoding / compression and decoding / decompression mechanisms may be used in order to reduce signaling overhead associated with the transmission of CSI from the UE to the base station.

[0020] FIG. 1 illustrates a diagram 100 showing an example of a two-sided AI / ML model 102 for CSI compression and decompression, according to embodiments discussed herein. The AI / ML model 102 includes the (logical) UE side 104 (a portion of34920-2526-7773.1 P70625WO1the AI / ML model 102 that exists at a UE 106) and the (logical) network side 108 (a portion of the AI / ML model 102 that exists at, for example, a base station of a network 110), as illustrated.

[0021] The UE side 104 of the AI / ML model 102 that operates at the UE 106 includes an encoder 112. The encoder 112 is configured to accept an input 116 and to provide a bitstream 118 that is based on that input 116 as output. As illustrated, in some cases, the input 116 may include CSI, such as a raw downlink (DL) channel estimate or precoder information (such as a codebook-based indication of a precoder W and / or a set of eigenvectors corresponding to a precoder W).

[0022] The bitstream 118 is then transmitted from the UE 106 to the network 110.

[0023] The network side 108 of the AI / ML model 102 that operates at the network 110 (e.g., a base station of the network 110) includes a decoder 114. The decoder 114 is configured to accept the bitstream 118 as input and to decode information therein as output 120 to the network 110 for further processing. In this way, the information from the input 116 is made known to the network 110. Accordingly, in some cases, the output 120 may be understood to include CSI, such as a raw DL channel estimate or precoder information (such as a codebook-based indication of a precoder W and / or a set of eigenvectors corresponding to a precoder W), corresponding to the format of the input 116.

[0024] Based on the encoding mechanism used by the encoder 112, the bitstream 118 may be smaller than the raw or data as presented in the input 116. The encoder 112 may thus be understood to “compress” the input 116 into the bitstream, which is correspondingly understood to represent “compressed” information.

[0025] The result is that the transmission of the bitstream 118 to from the UE 106 to the network 110 results in the use of fewer radio resources than an alternative case where the input 116 is itself sent from the UE 106 to the network side 108 without such encoding / compression.

[0026] The output 120 may correspondingly be referred to variously herein as “decoded,” “decompressed,” “recovered,” etc.

[0027] In various wireless communication system deployments, it may be anticipated that different devices (base stations, UEs, etc.) deployed and operating within the wireless communication system come from various different sources / vendors / manufacturers. Accordingly, various mechanisms to provide for “inter- 44920-2526-7773.1 P70625WO1vendor” training collaboration for the training of a two-sided AI / ML model for CSI compression / decompression in the case of, for example, a UE that is from a first vendor and a base station that is from a second vendor are proposed.

[0028] In a first option for inter-vendor training collaboration, a fully standardized reference model (having, e.g., a standardized structure and standardized parameters) may be defined for use at each of the base station and the UE.

[0029] In a second option for inter-vendor training collaboration, a standardized reference model having a standardized structure may be defined for use at each of the base station and the UE. Then, a parameter exchange between base station and the UE is carried out that equips both the UE and the base station with compatible / consistent parameters for the present use of an AI / ML model.

[0030] In a third option for inter-vendor training collaboration, a standardized AI / ML model format is defined for use at both the UE and the base station. Then, a reference model exchange between the UE and the base station is used to train an AI / ML model to be used by the UE and the base station.

[0031] In a fourth option for inter-vendor training collaboration, standardized data / a standardized dataset format is defined for use at both a UE and a base station. Then, a dataset exchange that uses this standardized data / dataset format facilitates the training of an AI / ML model that is to be used at both the UE and the base station.

[0032] Preliminarily, it is noted that FIG. 1 as described above introduces its AI / ML model 102 for CSI compression / decompression in terms of an encoder 112 that is found at a UE 106 and a decoder 114 that is located at the base station of a network 110. This discussion corresponds to the active use of an already-trained AI / ML model for CSI compression / decompression purposes. Various embodiments herein relate to mechanisms for AI / ML model training with the goal of developing a decoder / encoder pairing corresponding to an AI / ML model such as the model 102 described in FIG. 1.

[0033] In such training contexts, it will be understood that, in addition to an encoder at a UE and a decoder at the network, other entities associated with the AI / ML model could be used / developed in an intermediate fashion (e.g., to facilitate the training of the AI / ML model). For example, there may be one or more decoders at a UE and / or one or more encoders at a base station that are associated with training the AI / ML model (e.g., as will now be discussed in FIG. 2).54920-2526-7773.1 P70625WO1

[0034] FIG. 2 illustrates a diagram 200 for an inter-vendor training collaboration procedure for training an AI / ML model for CSI compression / decompression that relies on standardized data / a standardized dataset format and a dataset exchange between a base station and a UE.

[0035] As illustrated, at the network 202, a joint training 206 for a network-side encoder Ei 208 and a network-side decoder Di 210 is performed.

[0036] Then, a dataset 212 is sent from the network 202 to the UE 204. The dataset 212 includes target CSI and CSI feedback information. The CSI feedback information may include model parameters for the AI / ML model at the network 202 (as represented by the network-side encoder Ei 208 and / or the network-side decoder Di 210) Note that the dataset 212 may be a partial dataset.

[0037] Training at the UE 204 then proceeds according to one of a first alternative 214 and a second alternative 216.

[0038] In the first alternative 214, the UE 204 performs training in multiple steps. In a first step 224, the UE 204 trains a UE-side decoder D2218 using the dataset 212 (e.g., the target CSI and the CSI feedback information) received from the network 202. Note that as part of this process, the UE 204 may train the UE-side decoder D2218 in such a way that the UE-side decoder D2218 has a different structure than the structure of the network-side decoder Di 210

[0039] Then, in a second step 226, the UE 204 freezes the model for the UE-side decoder D2218 and performs a joint training procedure with the UE-side decoder D2218 to train a UE-side encoder E2220. Note that as part of this process, the UE 204 may train the UE-side encoder E2220 in such a way that the UE-side encoder E2220 has a different structure than the structure of the network-side encoder Ei 208.

[0040] In a second alternative 216. upon receiving the dataset 212 (e.g.. the target CSI and the CSI feedback information), the UE 204 uses the dataset 212 to train a UE-side encoder E2222. Note that as part of this process, the UE 204 may train UE-side encoder E2222 in such a way that the UE-side encoder E2222 has a different structure than the structure of the network-side encoder Ei 208.

[0041] Accordingly, it may be understood that in cases corresponding to such examples, a network vendor first trains a pair of (CSI generation model (a network-side encoder Ei 208), CSI reconstruction model (network-side decoder Di 210)) Then, a corresponding dataset 212 is sent to a UE 204 using the standardized dataset format.64920-2526-7773.1 P70625WO1

[0042] Note that this may occur with respect to multiple UEs, each of which may be different (e.g.. different chipsets, different vendors, etc.). It has been identified that there may be some inherent differences as to operation of these UEs from different vendors. For example, different UEs may use different UE-sided conditions and / or different UE-specific configurations (e.g., different analog and / or signal processing procedures) that, all else being equal, cause materially different model input distributions / characteristics corresponding to a CSI generated by a reference signal measurement at each of those UEs.

[0043] Accordingly, it should be understood that an input CSI V perceived at a UE could be UE-dependent corresponding to an inferencing stage, and may be different from a CSI V that is assumed at the network. For example, the CSI V assumed at the network may be identified according to a UE agnostic analysis / formula, and not according to a result of any analysis that takes into account particular UE-sided conditions and / or UE-specific configurations of any given UE.

[0044] Thus, a training dataset from the network side that is delivered to and used at a UE side without considering / accounting for particular UE-sided conditions / UE-specific configurations can result in mismatched data distribution / characteristics between the training dataset and the inference dataset that is used to generate input data for the UE-side / UE-part model inference. This mismatch may result in performance degradation of the AI / ML model.

[0045] Accordingly, embodiments discussed herein relate mechanisms for avoiding such performance impacts due to mismatches between network-side data distributions and corresponding UE-sided inferencing data distributions. In other words, embodiments herein relate to defined mechanisms for AI / ML model training that are intended to alleviate the effects of mismatches in the way that different devices measure / use / understand CSI.

[0046] It has also been identified that an over-the-air dataset exchange from the network side of network vendor to each UE at a UE side is non-trivial. This non-trivial size relates to various particular concerns on the feasibility and / or complexity of corresponding training methods. For example, over-the-air delivery dataset exchanges can result in / represent: large Uu signaling overhead; large UE power consumption; large latency, such that the AI / ML model trained by using the delivered information may become outdated or have a relatively short model-in-use time; and / or that over-the-air74920-2526-7773.1 P70625WO1delivery requires mechanisms to control for which part(s) of the information (dataset, parameters, performance target, etc.) that is to be transmitted over which base station(s) to which UE(s) in the network. Various possible mechanisms for addressing these concerns may increase a standardization complexity for the wireless communication system, and may require offline cross-vendor collaboration efforts.

[0047] Accordingly, embodiments discussed herein relate mechanisms for minimizing and / or avoiding these feasibility and / or complexity concerns.

[0048] Finally, it has also been identified that one way to account for the fact that the network side typically operates with multiple UEs from multiple vendors would be for the network to employ a unique decoder model for all the UE encoder models used at the active UEs. However, this would mean that the network undertakes the complexity of training and maintaining a different network side decoder for each UE encoder (or at least for each UE vendor encoder represented within the active UEs).

[0049] Accordingly, embodiments discussed herein relate mechanisms for, alternatively, the training of a single AI / ML model decoder at the network that is sufficiently accurate when used with each / any of the various UE encoders used at the UEs of different vendors.

[0050] FIG. 3 illustrates a diagram 300 for a mechanism for training collaboration with distributed gradient exchange. As illustrated, a base station trains its decoder 302 to work with each of a first encoder 304 at a first UE, a second encoder 306 at a second UE, and so on through the k-th encoder 308 at a k-th UE.

[0051] Note that as part of the mechanism, each UE trains its own encoder model with a local set of dataset CSI V that are pertinent to that UE. Accordingly, it is understood that the first UE trains the first encoder 304 based on a first dataset CSI Vi 310, the second UE trains the second encoder 306 based on a second dataset CSI V2312, and so on until the k-th UE trains the k-th encoder 308 based on the k-th dataset CSI Vk314, as shown. The individual dataset CSI V used by each UE reflects that UE's particular radio frequency (RF) processing and / or other implementation characteristics.

[0052] Note also that each UE may provide the base station with its dataset CSI V that is used to train on that UE. In other words, each UE sends its dataset CSI V to the network for crowdsourcing (so, in the illustrated case, it is understood that the base station is informed of each of the first dataset CSI Ei 310, the second dataset CSI V2312,84920-2526-7773.1 P70625WO1and so on until the k-th dataset CSI Vk314). Each of these particular dataset CSI V as provided to the base station indicates the UE source from which it came.

[0053] Preliminarily, each UE performs a forward pass of its local dataset CSI V through its respective encoder. Accordingly, the first UE performs a forward pass through the first encoder 304 using the first dataset CSI V₁ 310, the second UE performs a forward pass through the second encoder 306 using the second dataset CSI V₂ 312, and so on until the k-th UE performs a forward pass through the k-th encoder 308 using the k-th dataset CSI Vk314, as illustrated.

[0054] Each of these forward passes produces a latent message c (e.g., a bitstream corresponding to the nature of a bitstream 118 that is an output of an encoder 112, as was described in relation to FIG. 1). Accordingly, as shown, the result of the forward pass at the first UE is the first latent message ci 316, the result of the forward pass at the second UE is the second latent message c₂ 318, and so on until the A-th UE. where the result of the forward pass is the k-th latent message ck320.

[0055] Each of the UEs then transmits its latent message c to the base station. The base station applies these latent messages c (the latent messages ci...ck 322) along the forward path of its decoder 302 to generate a reconstructed CSI V corresponding to each of the latent messages c (the reconstructed CSIs V1... Vk324).

[0056] The base station then calculates a loss function 326 using the dataset CSIs Vi... Vk (the first dataset CSI V₁ 310, the second dataset CSI V₂ 312, and so on through the k-th dataset CSI Vk314) as provided from the UEs as compared with the reconstructed CSIs V1... Vk324 as generated by the decoder 302.

[0057] The loss function 326 may use, for example, one or more of: a mean squared error (MSE) calculation that calculates the average squared difference between predicted and actual values; a mean absolute error (MAE) calculation that calculates the average absolute difference between predictions and actual values; a binary cross-entropy (BCE) calculation as may be used for binary classification tasks, where the result is a probability between 0 and 1; and / or a hinge loss determination as may be used with support vector machines (SVMs) for margin-based classification purposes.

[0058] Based on the result of this loss function, the base station computes a set of decoder gradients ∇ℒ(ΦDec) 328 with respect to the parameters used by the decoder 302. These decoder gradients ∇ℒ(ΦDec) 328 may be understood to be compensatory for the 94920-2526-7773.1 P70625WO1decoder 302 with respect to the result of the loss function 326 (in the sense that they are meant to be applied to the decoder 302 to adjust for / minimize the loss calculated by the loss function 326).

[0059] Note that the decoder gradients328 are made up of UE-specific decoder gradients. For example, the decoder gradients ∇ℒ(ΦDec) 328 may include the first UE-specific decoder gradients ∇ℒ(ΦDec,1) 330 for the first UE, the second UE-specific decoder gradients ∇ℒ(ΦDec,2) 332 for the second UE, and so on until the k-th UE-specific decoder gradients ∇ℒ(ΦDec,k) 334 for the k-th UE.

[0060] The decoder gradients E£(< J’£)ec) 328 may be calculated according to configurable batches. For example, if a configured batch size is 100, and if a dataset CSI size for any given UE is of size 1000, then there may be 1000 / 100 = 10 gradients computed corresponding to a dataset CSI from a UE. Then, if there are k = 10 UEs, there would be a total of 10 * 10 = 100 gradients within the set of decoder gradients V£(^>Dec) 328.

[0061] As illustrated, the base station transmits the first UE-specific decoder gradients E£(< I’£)ec l) 330 to the first UE, transmits the second UE-specific decoder gradients E£(< PDeCj2) 332 to the second UE, and so on until it transmits the -th UE-specific decoder gradients E£(Dec k) 334 to the £-th UE.

[0062] The base station also adjusts its decoder 302 corresponding to each of these gradients transmitted to each of these UEs. So, corresponding to the diagram 300, it will be understood that the base station adjusts the decoder 302 based on each of the first UE-specific decoder gradients ∇ℒ(ΦDec,1) 330, the second UE-specific decoder gradients332, and so on until the k-th UE-specific decoder gradients ∇ℒ(ΦDec,k) 334. Accordingly, the decoder 302 becomes more convergent to a solution / state that is accurate with respect to each of the first encoder 304, the second encoder 306, and so on until the k-th encoder 308. Note that these adjustments may be understood in general terms as a use of the set of gradients ∇ℒ(ΦDec) 328 to adjust the decoder 302 to compensate for the result of the loss function 326.

[0063] Upon receiving its set of UE-specific gradients, each UE performs a backward pass through its encoder to compute gradients with respect to / for its encoder. For example, as illustrated, the first UE performs a backward pass through the first encoder 304 using the first UE-specific decoder gradients E£(< I’£)ec l) 330 to generate the first104920-2526-7773\1 P70625WO1UE-specific encoder gradients ∇ℒ(ΘEnc,1) 336, the second UE performs a backward pass through the second encoder 306 using the second UE-specific encoder gradients332 to generate the second UE-specific encoder gradients ∇ℒ(ΘEnc,2) 338, and so on until the Uth UE performs a backward pass through the / -th encoder 308 using the Ulh UE-specific encoder gradients ∇ℒ(ΦDec,k) 334 to generate the Uth UE-specific decoder gradients ∇ℒ(ΘEnc,k) 340.

[0064] Then, each UE adjusts its own encoder using the calculated encoder gradients. For example, the first UE adjusts the first encoder 304 based on the first UE-specific encoder gradients ∇ℒ(ΘEnc,1) 336, the second UE adjusts the second encoder 306 based on the second UE-specific encoder gradients ∇ℒ(ΘEnc,2) 338, and so on until the Uth UE adjusts the Uth encoder 308 based on the Uth UE-specific encoder gradients ∇ℒ(ΘEnc,k) 340. Note that because each set of these UE-specific encoder gradients is ultimately sourced from corresponding decoder gradients that are to be applied at the decoder 302, these adjustments drive each individual encoder to be in closer convergence with the decoder 302 at the base station.

[0065] The procedure of FIG. 3 as just described may be repeated / looped across multiple iterations. Each iteration may involve each participating UE l...k as discussed. Accordingly, it will be understood that network parameters used for the decoder 302 are iteratively adjusted on the ensemble of the UEs 1.. A, thereby enabling the decoder 302 to converge to a relatively simple, single decoder model that is interoperable with all the UE encoders. Correspondingly, all UE decoders for UEs 1... C (e g., the first encoder 304, the second encoder 306, and so on until the Uth encoder 308) also are understood to iteratively update / converge to the decoder 302 as a result of this iterative looping.

[0066] Note that as part of the described procedure, the base station may conclude, based on a calculated value of a loss function 326, that the decoder 302 and the various encoders at the UEs 1... C (the first encoder 304, the second encoder 306, and so on until the / c-th encoder 308) are sufficiently accurate / converged. In such a case, the base station may send each of the UEs l...k an indication that convergence has been reached.Accordingly, the base station and each of the UEs...k break out of the iterations / looping of the procedure described here in relation to the diagram 300 and proceed to normal inferencing operation on a going forward basis.

[0067] FIG. 4 illustrates a method 400 of a base station, according to embodiments discussed herein. The method 400 includes receiving 402, from a first UE, a first114920-2526-7773\1 P70625WO1bitstream generated by a first encoder at the first UE using a first dataset CSI. The method 400 further includes providing 404 the first bitstream as a first input to a decoder at the base station to generate a first reconstructed CSI. The method 400 further includes performing 406 a first loss function calculation for the decoder based on differences between the first dataset CSI and the first reconstructed CSI. The method 400 further includes identifying 408, based on the first loss function calculation, first decoder gradients for the decoder, wherein the first decoder gradients include first UE-specific decoder gradients corresponding to the first UE. The method 400 further includes adjusting 410 the decoder based on the first decoder gradients. The method 400 further includes sending 412, to the first UE, the first UE-specific decoder gradients.

[0068] In some embodiments, the method 400 further includes receiving, from the first UE, the first dataset CSI.

[0069] In some embodiments, the method 400 further includes receiving, from the first UE, a second bitstream generated by the first encoder at the first UE using the first dataset CSI; providing the second bitstream as a second input to the decoder at the base station to generate a second reconstructed CSI; performing a second loss function calculation for the decoder based on differences between the first dataset CSI and the second reconstructed CSI; identifying, based on the second loss function calculation, second decoder gradients for the decoder, wherein the second decoder gradients include second UE-specific decoder gradients corresponding to the first UE; adjusting the decoder based on the second decoder gradients; and sending, to the first UE, the second UE-specific decoder gradients.

[0070] In some embodiments, the method 400 further includes receiving, from the first UE, a second bitstream generated by the first encoder at the first UE using the first dataset CSI; providing the second bitstream as a second input to the decoder at the base station to generate a second reconstructed CSI; performing a second loss function calculation for the decoder based on differences between the second dataset CSI and the second reconstructed CSI; determining, based on the second loss function calculation, that there is a convergence between the decoder at the base station and the first encoder at the first UE; and sending, to the first UE, an indication of the convergence.

[0071] In some embodiments, the method 400 further includes receiving, from a second UE, a second bitstream generated by a second encoder at the second UE using a second dataset CSI; providing the second bitstream as a second input to the decoder at the base124920-2526-7773.1 P70625WO1station to generate a second reconstructed CSI, wherein the first loss function calculation for the decoder is further performed based on differences between the second dataset CSI and the second reconstructed CSI, and the first decoder gradients identified based on the first loss function calculation further include second UE-specific decoder gradients corresponding to the second UE; and sending, to the second UE, the second UE-specific decoder gradients.

[0072] In some of these embodiments, the method 400 further includes receiving, from the first UE, the first dataset CSI; and receiving, from the second UE, the second dataset CSI.

[0073] In some of these embodiments, the method 400 further includes receiving, from the first UE, a third bitstream generated by the first encoder at the first UE using the first dataset CSI; providing the third bitstream as a third input to the decoder at the base station to generate a third reconstructed CSI; performing a second loss function calculation for the decoder based on differences between the first dataset CSI and the third reconstructed CSI; identifying, based on the second loss function calculation, second decoder gradients for the decoder, wherein the second decoder gradients include third UE-specific decoder gradients corresponding to the first UE; adjusting the decoder based on the second decoder gradients; and sending, to the first UE, the third UE-specific decoder gradients. In some such circumstances, the method 400 further includes receiving, from the second UE, a fourth bitstream generated by the second encoder at the second UE using the second dataset CSI; providing the fourth bitstream as a fourth input to the decoder at the base station to generate a fourth reconstructed CSI, wherein: the second loss function calculation for the decoder is further performed based on differences between the second dataset CSI and the fourth reconstructed CSI, and the second decoder gradients identified based on the second loss function calculation further include fourth UE-specific decoder gradients corresponding to the second UE; and sending, to the second UE, the fourth UE-specific decoder gradients.

[0074] In some of these embodiments, the method 400 further includes receiving, from the first UE. a third bitstream generated by the first encoder at the first UE using the first dataset CSI; providing the third bitstream as a third input to the decoder at the base station to generate a third reconstructed CSI; receiving, from the second UE, a fourth bitstream generated by the second encoder at the second UE using the second dataset CSI; providing the fourth bitstream as a fourth input to the decoder at the base station to134920-2526-7773\1 P70625WO1generate a fourth reconstructed CSI; performing a second loss function calculation for the decoder based on differences between the first dataset CSI and the third reconstructed CSI and differences between the second dataset CSI and the fourth reconstructed CSI; determining, based on the second loss function calculation, that there is a convergence between the decoder at the base station and the first encoder at the first UE and between the decoder at the base station and the second encoder at the second UE; and sending, to each of the first UE and the second UE. an indication of the convergence.

[0075] FIG. 5 illustrates a method 500 of a UE, according to embodiments discussed herein. The method 500 includes providing 502 a dataset CSI as a first input to an encoder at the UE to generate a first bitstream. The method 500 further includes sending 504, to a base station, the first bitstream. The method 500 further includes receiving 506, from the base station, first UE-specific decoder gradients corresponding to the UE. The method 500 further includes performing 508 a first backward pass through the encoder using the first UE-specific decoder gradients to generate first encoder gradients for the encoder. The method 500 further includes adjusting 510, method 500 adjusts the encoder based on the first encoder gradients.

[0076] In some embodiments, the method 500 further includes sending, to the base station, the dataset CSI.

[0077] In some embodiments, the method 500 further includes providing the dataset CSI as a second input to the encoder to generate a second bitstream; sending, to the base station, the second bitstream; receiving, from the base station, second UE-specific decoder gradients corresponding to the UE; performing a second backward pass through the encoder using the second UE-specific decoder gradients to generate second encoder gradients for the encoder; and adjusting the encoder based on the second encoder gradients.

[0078] In some embodiments, the method 500 further includes providing the dataset CSI as a second input to the encoder to generate a second bitstream; sending, to the base station, the second bitstream; and receiving, from the base station, an indication that there is a convergence between a decoder at the base station and the encoder at the UE.

[0079] FIG. 6 illustrates an example architecture of a wireless communication system 600, according to embodiments disclosed herein. The following description is provided for an example wireless communication system 600 that operates in conjunction with the144920-2526-7773\1 P70625WO1LTE system standards and / or 5G or NR system standards as provided by 3GPP technical specifications.

[0080] As shown by FIG. 6, the wireless communication system 600 includes UE 602 and UE 604 (although any number of UEs may be used). In this example, the UE 602 and the UE 604 are illustrated as smartphones (e.g., handheld touchscreen mobile computing devices connectable to one or more cellular networks), but may also comprise any mobile or non-mobile computing device configured for wireless communication.

[0081] The UE 602 and UE 604 may be configured to communicatively couple with a RAN 606. In embodiments, the RAN 606 may be NG-RAN, E-UTRAN, etc. The UE 602 and UE 604 utilize connections (or channels) (shown as connection 608 and connection 610, respectively) with the RAN 606, each of which comprises a physical communications interface. The RAN 606 can include one or more base stations (such as base station 612 and base station 614) that enable the connection 608 and connection 610.

[0082] In this example, the connection 608 and connection 610 are air interfaces to enable such communicative coupling, and may be consistent with RAT(s) used by the RAN 606, such as, for example, an LTE and / or NR.

[0083] In some embodiments, the UE 602 and UE 604 may also directly exchange communication data via a sidelink interface 616. The UE 604 is shown to be configured to access an access point (shown as AP 618) via connection 620. By way of example, the connection 620 can comprise a local wireless connection, such as a connection consistent with any IEEE 802.11 protocol, wherein the AP 618 may comprise a Wi-Fi® router. In this example, the AP 618 may be connected to another network (for example, the Internet) without going through a CN 624.

[0084] In embodiments, the UE 602 and UE 604 can be configured to communicate using orthogonal frequency division multiplexing (OFDM) communication signals with each other or with the base station 612 and / or the base station 614 over a multicarrier communication channel in accordance with various communication techniques, such as, but not limited to, an orthogonal frequency division multiple access (OFDMA) communication technique (e.g., for downlink communications) or a single carrier frequency division multiple access (SC-FDMA) communication technique (e.g., for uplink and ProSe or sidelink communications), although the scope of the embodiments is154920-2526-7773.1 P70625WO1not limited in this respect. The OFDM signals can comprise a plurality of orthogonal subcarriers.

[0085] In some embodiments, all or parts of the base station 612 or base station 614 may be implemented as one or more software entities running on server computers as part of a virtual network. In addition, or in other embodiments, the base station 612 or base station 614 may be configured to communicate with one another via interface 622. In embodiments where the wireless communication system 600 is an LTE system (e.g., when the CN 624 is an EPC), the interface 622 may be an X2 interface. The X2 interface may be defined between two or more base stations (e.g., two or more eNBs and the like) that connect to an EPC. and / or between two eNBs connecting to the EPC. In embodiments where the wireless communication system 600 is an NR system (e.g., when CN 624 is a 5GC), the interface 622 may be an Xn interface. The Xn interface is defined between two or more base stations (e.g., two or more gNBs and the like) that connect to 5GC. between a base station 612 (e.g.. a gNB) connecting to 5GC and an eNB. and / or between two eNBs connecting to 5GC (e.g., CN 624).

[0086] The RAN 606 is shown to be communicatively coupled to the CN 624. The CN 624 may comprise one or more network elements 626, which are configured to offer various data and telecommunications services to customers / subscribers (e.g., users of UE 602 and UE 604) who are connected to the CN 624 via the RAN 606. The components of the CN 624 may be implemented in one physical device or separate physical devices including components to read and execute instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium).

[0087] In embodiments, the CN 624 may be an EPC, and the RAN 606 may be connected with the CN 624 via an SI interface 628. In embodiments, the SI interface 628 may be split into two parts, an SI user plane (Sl-U) interface, which carries traffic data between the base station 612 or base station 614 and a serving gateway (S-GW), and the SI -MME interface, which is a signaling interface between the base station 612 or base station 614 and mobility management entities (MMEs).

[0088] In embodiments, the CN 624 may be a 5GC, and the RAN 606 may be connected with the CN 624 via an NG interface 628. In embodiments, the NG interface 628 may be split into two parts, an NG user plane (NG-U) interface, which carries traffic data between the base station 612 or base station 614 and a user plane function (UPF). and the SI control plane (NG-C) interface, which is a signaling interface between the164920-2526-7773.1 P70625WO1base station 612 or base station 614 and access and mobility management functions (AMFs).

[0089] Generally, an application server 630 may be an element offering applications that use internet protocol (IP) bearer resources with the CN 624 (e.g., packet switched data services). The application server 630 can also be configured to support one or more communication services (e.g., VoIP sessions, group communication sessions, etc.) for the UE 602 and UE 604 via the CN 624. The application server 630 may communicate with the CN 624 through an IP communications interface 632.

[0090] FIG. 7 illustrates a system 700 for performing signaling 734 between a wireless device 702 and a network device 718, according to embodiments disclosed herein. The system 700 may be a portion of a wireless communications system as herein described. The wireless device 702 may be, for example, a UE of a wireless communication system. The network device 718 may be. for example, a base station (e.g.. an eNB or a gNB) of a wireless communication system.

[0091] The wireless device 702 may include one or more processor(s) 704. The processor(s) 704 may execute instructions such that various operations of the wireless device 702 are performed, as described herein. The processor(s) 704 may include one or more baseband processors implemented using, for example, a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a controller, a field programmable gate array (FPGA) device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein.

[0092] The wireless device 702 may include a memory 706. The memory 706 may be a non-transitory computer-readable storage medium that stores instructions 708 (which may include, for example, the instructions being executed by the processor(s) 704). The instructions 708 may also be referred to as program code or a computer program. The memory 706 may also store data used by, and results computed by, the processor(s) 704.

[0093] The wireless device 702 may include one or more transceiver(s) 710 that may include RF transmitter circuitry and / or receiver circuitry that use the antenna(s) 712 of the wireless device 702 to facilitate signaling (e.g., the signaling 734) to and / or from the wireless device 702 with other devices (e.g., the network device 718) according to corresponding RATs.174920-2526-7773.1 P70625WO1

[0094] The wireless device 702 may include one or more antenna(s) 712 (e.g., one, two, four, or more). For embodiments with multiple antenna(s) 712, the wireless device 702 may leverage the spatial diversity of such multiple antenna(s) 712 to send and / or receive multiple different data streams on the same time and frequency resources. This behavior may be referred to as, for example, multiple input multiple output (MIMO) behavior (referring to the multiple antennas used at each of a transmitting device and a receiving device that enable this aspect). MIMO transmissions by the wireless device 702 may be accomplished according to precoding (or digital beamforming) that is applied at the wireless device 702 that multiplexes the data streams across the antenna(s) 712 according to known or assumed channel characteristics such that each data stream is received with an appropriate signal strength relative to other streams and at a desired location in the spatial domain (e.g., the location of a receiver associated with that data stream). Certain embodiments may use single user MIMO (SU-MIMO) methods (where the data streams are all directed to a single receiver) and / or multi user MIMO (MU-MIMO) methods (where individual data streams may be directed to individual (different) receivers in different locations in the spatial domain).

[0095] In certain embodiments having multiple antennas, the wireless device 702 may implement analog beamforming techniques, whereby phases of the signals sent by the antenna(s) 712 are relatively adjusted such that the (joint) transmission of the antenna(s) 712 can be directed (this is sometimes referred to as beam steering).

[0096] The wireless device 702 may include one or more interface(s) 714. The interface(s) 714 may be used to provide input to or output from the wireless device 702. For example, a wireless device 702 that is a UE may include interface(s) 714 such as microphones, speakers, a touchscreen, buttons, and the like in order to allow for input and / or output to the UE by a user of the UE. Other interfaces of such a UE may be made up of transmitters, receivers, and other circuitry (e.g., other than the transceiver(s) 710 / antenna(s) 712 already described) that allow for communication between the UE and other devices and may operate according to known protocols (e.g., Wi-Fi®, Bluetooth®, and the like).

[0097] The wireless device 702 may include a distributed gradient exchange training module 716. The distributed gradient exchange training module 716 may be implemented via hardware, software, or combinations thereof. For example, the distributed gradient exchange training module 716 may be implemented as a processor,184920-2526-7773.1 P70625WO1circuit, and / or instructions 708 stored in the memory 706 and executed by the processor(s) 704. In some examples, the distributed gradient exchange training module 716 may be integrated within the processor(s) 704 and / or the transceiver(s) 710. For example, the distributed gradient exchange training module 716 may be implemented by a combination of software components (e.g., executed by a DSP or a general processor) and hardware components (e.g., logic gates and circuitry) within the processor(s) 704 or the transceiver(s) 710.

[0098] The distributed gradient exchange training module 716 may be used for various aspects of the present disclosure, for example, aspects of FIG. 4. The distributed gradient exchange training module 716 may configure the wireless device 702 to provide a dataset CSI as a first input to an encoder at the UE to generate a first bitstream; send, to a base station, the first bitstream; receive, from the base station, first UE-specific decoder gradients corresponding to the UE; perform a first backward pass through the encoder using the first UE-specific decoder gradients to generate first encoder gradients for the encoder; and adjust the encoder based on the first encoder gradients.

[0099] The network device 718 may include one or more processor(s) 720. The processor(s) 720 may execute instructions such that various operations of the network device 718 are performed, as described herein. The processor(s) 720 may include one or more baseband processors implemented using, for example, a CPU, a DSP, an ASIC, a controller, an FPGA device, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein.

[0100] The network device 718 may include a memory 722. The memory 722 may be a non-transitory computer-readable storage medium that stores instructions 724 (which may include, for example, the instructions being executed by the processor(s) 720). The instructions 724 may also be referred to as program code or a computer program. The memory 722 may also store data used by, and results computed by, the processor(s) 720.

[0101] The network device 718 may include one or more transceiver(s) 726 that may include RF transmitter circuitry and / or receiver circuitry that use the antenna(s) 728 of the network device 718 to facilitate signaling (e.g., the signaling 734) to and / or from the network device 718 with other devices (e.g., the wireless device 702) according to corresponding RATs.

[0102] The network device 718 may include one or more antenna(s) 728 (e.g., one, two, four, or more). In embodiments having multiple antenna(s) 728, the network device 718194920-2526-7773.1 P70625WO1may perform MIMO, digital beamforming, analog beamforming, beam steering, etc., as has been described.

[0103] The network device 718 may include one or more interface(s) 730. The interface(s) 730 may be used to provide input to or output from the network device 718. For example, a network device 718 that is a base station may include interface(s) 730 made up of transmitters, receivers, and other circuitry (e.g., other than the transceiver(s) 726 / antenna(s) 728 already described) that enables the base station to communicate with other equipment in a core network, and / or that enables the base station to communicate with external networks, computers, databases, and the like for purposes of operations, administration, and maintenance of the base station or other equipment operably connected thereto.

[0104] The network device 718 may include a distributed gradient exchange training module 732. The distributed gradient exchange training module 732 may be implemented via hardware, software, or combinations thereof. For example, the distributed gradient exchange training module 732 may be implemented as a processor, circuit, and / or instructions 724 stored in the memory 722 and executed by the processor(s) 720. In some examples, the distributed gradient exchange training module 732 may be integrated within the processor(s) 720 and / or the transceiver(s) 726. For example, the distributed gradient exchange training module 732 may be implemented by a combination of software components (e.g., executed by a DSP or a general processor) and hardware components (e.g., logic gates and circuitry) within the processor(s) 720 or the transceiver(s) 726.

[0105] The distributed gradient exchange training module 732 may be used for various aspects of the present disclosure, for example, aspects of FIG. 5. The distributed gradient exchange training module 732 may configure the network device 718 to receive, from a UE, a bitstream generated by an encoder at the UE using a dataset CSI; provide the bitstream as an input to a decoder at the base station to generate a reconstructed CSI; perform a loss function calculation for the decoder based on differences between the dataset CSI and the reconstructed CSI; identify, based on the loss function calculation, decoder gradients for the decoder, wherein the decoder gradients include UE-specific decoder gradients corresponding to the UE; adjust the decoder based on the decoder gradients; and send, to the UE, the UE-specific decoder gradients.204920-2526-7773.1 P70625WO1

[0106] Embodiments contemplated herein include an apparatus comprising means to perform one or more elements of the method 500. This apparatus may be. for example, an apparatus of a UE (such as a wireless device 702 that is a UE, as described herein).

[0107] Embodiments contemplated herein include one or more non -transitory computer-readable media comprising instructions to cause an electronic device, upon execution of the instructions by one or more processors of the electronic device, to perform one or more elements of the method 500. This non-transitory computer-readable media may be, for example, a memory of a UE (such as a memory 706 of a wireless device 702 that is a UE, as described herein).

[0108] Embodiments contemplated herein include an apparatus comprising logic, modules, or circuitry to perform one or more elements of the method 500. This apparatus may be, for example, an apparatus of a UE (such as a wireless device 702 that is a UE, as described herein).

[0109] Embodiments contemplated herein include an apparatus comprising: one or more processors and one or more computer-readable media comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more elements of the method 500. This apparatus may be, for example, an apparatus of a UE (such as a wireless device 702 that is a UE, as described herein).

[0110] Embodiments contemplated herein include a signal as described in or related to one or more elements of the method 500.

[0111] Embodiments contemplated herein include a computer program or computer program product comprising instructions, wherein execution of the program by a processor is to cause the processor to carry out one or more elements of the method 500. The processor may be a processor of a UE (such as a processor(s) 704 of a wireless device 702 that is a UE, as described herein). These instructions may be, for example, located in the processor and / or on a memory of the UE (such as a memory 706 of a wireless device 702 that is a UE, as described herein).

[0112] Embodiments contemplated herein include an apparatus comprising means to perform one or more elements of the method 400. This apparatus may be. for example, an apparatus of a base station (such as a network device 718 that is a base station, as described herein).

[0113] Embodiments contemplated herein include one or more non-transitory computer-readable media comprising instructions to cause an electronic device, upon 214920-2526-7773\1 P70625WO1execution of the instructions by one or more processors of the electronic device, to perform one or more elements of the method 400. This non-transitory computer-readable media may be, for example, a memory of a base station (such as a memory 722 of a network device 718 that is a base station, as described herein).

[0114] Embodiments contemplated herein include an apparatus comprising logic, modules, or circuitry to perform one or more elements of the method 400. This apparatus may be, for example, an apparatus of a base station (such as a network device 718 that is a base station, as described herein).

[0115] Embodiments contemplated herein include an apparatus comprising: one or more processors and one or more computer-readable media comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform one or more elements of the method 400. This apparatus may be. for example, an apparatus of a base station (such as a network device 718 that is a base station, as described herein).

[0116] Embodiments contemplated herein include a signal as described in or related to one or more elements of the method 400.

[0117] Embodiments contemplated herein include a computer program or computer program product comprising instructions, wherein execution of the program by a processing element is to cause the processing element to carry out one or more elements of the method 400. The processor may be a processor of a base station (such as a processor(s) 720 of a network device 718 that is a base station, as described herein). These instructions may be, for example, located in the processor and / or on a memory of the base station (such as a memory 722 of a network device 718 that is a base station, as described herein).

[0118] For one or more embodiments, at least one of the components set forth in one or more of the preceding figures may be configured to perform one or more operations. techniques, processes, and / or methods as set forth herein. For example, a baseband processor as described herein in connection with one or more of the preceding figures may be configured to operate in accordance with one or more of the examples set forth herein. For another example, circuitry associated with a UE, base station, network element, etc. as described above in connection with one or more of the preceding figures may be configured to operate in accordance with one or more of the examples set forth herein.224920-2526-7773\1 P70625WO1

[0119] Any of the above described embodiments may be combined with any other embodiment (or combination of embodiments), unless explicitly stated otherwise. The foregoing description of one or more implementations provides illustration and description, but is not intended to be exhaustive or to limit the scope of embodiments to the precise form disclosed. Modifications and variations are possible in light of the above teachings or may be acquired from practice of various embodiments.

[0120] Embodiments and implementations of the systems and methods described herein may include various operations, which may be embodied in machine-executable instructions to be executed by a computer system. A computer system may include one or more general-purpose or special-purpose computers (or other electronic devices). The computer system may include hardware components that include specific logic for performing the operations or may include a combination of hardware, software, and / or firmware.

[0121] It should be recognized that the systems described herein include descriptions of specific embodiments. These embodiments can be combined into single systems, partially combined into other systems, split into multiple systems or divided or combined in other ways. In addition, it is contemplated that parameters, attributes, aspects, etc. of one embodiment can be used in another embodiment. The parameters, attributes, aspects, etc. are merely described in one or more embodiments for clarity, and it is recognized that the parameters, attributes, aspects, etc. can be combined with or substituted for parameters, attributes, aspects, etc. of another embodiment unless specifically disclaimed herein.

[0122] It is well understood that the use of personally identifiable information should follow privacy policies and practices that are generally recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy of users. In particular, personally identifiable information data should be managed and handled so as to minimize risks of unintentional or unauthorized access or use, and the nature of authorized use should be clearly indicated to users.

[0123] Although the foregoing has been described in some detail for purposes of clarity, it will be apparent that certain changes and modifications may be made without departing from the principles thereof. It should be noted that there are many alternative ways of implementing both the processes and apparatuses described herein. Accordingly, the present embodiments are to be considered illustrative and not restrictive, and the234920-2526-7773.1 P70625WO1description is not to be limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.244920-2526-7773.1 P70625WO1

Claims

1. CLAIMS1. A method of a base station, comprising:receiving, from a first user equipment (UE), a first bitstream generated by a first encoder at the first UE using a first dataset channel state information (CSI);providing the first bitstream as a first input to a decoder at the base station to generate a first reconstructed CSI;performing a first loss function calculation for the decoder based on differences between the first dataset CSI and the first reconstructed CSI;identifying, based on the first loss function calculation, first decoder gradients for the decoder, wherein the first decoder gradients include first UE-specific decoder gradients corresponding to the first UE;adjusting the decoder based on the first decoder gradients; andsending, to the first UE. the first UE-specific decoder gradients.

2. The method of claim 1, further comprising receiving, from the first UE, the first dataset CSI.

3. The method of claim 1, further comprising:receiving, from the first UE, a second bitstream generated by the first encoder at the first UE using the first dataset CSI;providing the second bitstream as a second input to the decoder at the base station to generate a second reconstructed CSI;performing a second loss function calculation for the decoder based on differences between the first dataset CSI and the second reconstructed CSI;identifying, based on the second loss function calculation, second decoder gradients for the decoder, wherein the second decoder gradients include second UE-specific decoder gradients corresponding to the first UE;adjusting the decoder based on the second decoder gradients; andsending, to the first UE, the second UE-specific decoder gradients.

4. The method of claim 1, further comprising:receiving, from the first UE, a second bitstream generated by the first encoder at the first UE using the first dataset CSI;254920-2526-7773.1 P70625WO1providing the second bitstream as a second input to the decoder at the base station to generate a second reconstructed CSI;performing a second loss function calculation for the decoder based on differences between a second dataset CSI and the second reconstructed CSI;determining, based on the second loss function calculation, that there is a convergence between the decoder at the base station and the first encoder at the first UE; andsending, to the first UE, an indication of the convergence.

5. The method of claim 1, further comprising:receiving, from a second UE. a second bitstream generated by a second encoder at the second UE using a second dataset CSI;providing the second bitstream as a second input to the decoder at the base station to generate a second reconstructed CSI, wherein:the first loss function calculation for the decoder is further performed based on differences between the second dataset CSI and the second reconstructed CSI, andthe first decoder gradients identified based on the first loss function calculation further include second UE-specific decoder gradients corresponding to the second UE; andsending, to the second UE, the second UE-specific decoder gradients.

6. The method of claim 5, further comprising:receiving, from the first UE, the first dataset CSI; andreceiving, from the second UE, the second dataset CSI.

7. The method of claim 5, further comprising:receiving, from the first UE, a third bitstream generated by the first encoder at the first UE using the first dataset CSI;providing the third bitstream as a third input to the decoder at the base station to generate a third reconstructed CSIperforming a second loss function calculation for the decoder based on differences between the first dataset CSI and the third reconstructed CSI;264920-2526-7773\1 P70625WO1identifying, based on the second loss function calculation, second decoder gradients for the decoder, wherein the second decoder gradients include third UE-specific decoder gradients corresponding to the first UE;adjusting the decoder based on the second decoder gradients; andsending, to the first UE, the third UE-specific decoder gradients.

8. The method of claim 7, further comprising:receiving, from the second UE, a fourth bitstream generated by the second encoder at the second UE using the second dataset CSI;providing the fourth bitstream as a fourth input to the decoder at the base station to generate a fourth reconstructed CSI, wherein:the second loss function calculation for the decoder is further performed based on differences between the second dataset CSI and the fourth reconstructed CSI, andthe second decoder gradients identified based on the second loss function calculation further include fourth UE-specific decoder gradients corresponding to the second UE; andsending, to the second UE, the fourth UE-specific decoder gradients.

9. The method of claim 5, further comprising:receiving, from the first UE, a third bitstream generated by the first encoder at the first UE using the first dataset CSI;providing the third bitstream as a third input to the decoder at the base station to generate a third reconstructed CSI;receiving, from the second UE, a fourth bitstream generated by the second encoder at the second UE using the second dataset CSI;providing the fourth bitstream as a fourth input to the decoder at the base station to generate a fourth reconstructed CSI;performing a second loss function calculation for the decoder based on differences between the first dataset CSI and the third reconstructed CSI and differences between the second dataset CSI and the fourth reconstructed CSI;determining, based on the second loss function calculation, that there is a convergence between the decoder at the base station and the first encoder at the first UE274920-2526-7773.1 P70625WO1and between the decoder at the base station and the second encoder at the second UE; andsending, to each of the first UE and the second UE, an indication of the convergence.

10. A method of a user equipment (UE). comprising:providing a dataset channel state information (CSI) as a first input to an encoder at the UE to generate a first bitstream;sending, to a base station, the first bitstream;receiving, from the base station, first UE-specific decoder gradients corresponding to the UE;performing a first backward pass through the encoder using the first UE-specific decoder gradients to generate first encoder gradients for the encoder; andadjusting the encoder based on the first encoder gradients.

11. The method of claim 10, further comprising sending, to the base station, the dataset CSI.

12. The method of claim 10, further comprising:providing the dataset CSI as a second input to the encoder to generate a second bitstream;sending, to the base station, the second bitstream;receiving, from the base station, second UE-specific decoder gradients corresponding to the UE;performing a second backward pass through the encoder using the second UE-specific decoder gradients to generate second encoder gradients for the encoder; and adjusting the encoder based on the second encoder gradients.

13. The method of claim 10, further comprising:providing the dataset CSI as a second input to the encoder to generate a second bitstream;sending, to the base station, the second bitstream; andreceiving, from the base station, an indication that there is a convergence between a decoder at the base station and the encoder at the UE.

14. An apparatus comprising means to perform the method of any of claim 1 to claim 13.284920-2526-7773.1 P70625WO115. A computer-readable media comprising instructions to cause an electronic device, upon execution of the instructions by one or more processors of the electronic device, to perform the method of any of claim 1 to claim 13.

16. An apparatus comprising logic, modules, or circuitry to perform the method of any of claim 1 to claim 13.

17. A baseband processor for a user equipment (UE) that is configured to cause the UE to perform one or more elements of any one of claim 10 to claim 13.

18. A baseband processor for a base station that is configured to cause the base station to perform one or more elements of any one of claim 1 to claim 9.294920-2526-7773.1 P70625WO1