Apparatus, method and computer program
By deploying AI/ML models in the communication system and updating model parameters using the differences between codeword sets, the problems of model overfitting and inefficiency of CSI feedback caused by environmental drift are solved, and more efficient CSI recovery and model adaptability are achieved.
Patent Information
- Application Number
- CN202280100896.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-10
- Publication Date
- 2025-05-16
AI Technical Summary
When existing communication systems process channel information between user equipment and base stations, it is difficult for existing communication systems to effectively adapt to environmental drift, resulting in inefficient model overfitting and CSI feedback.
By deploying an AI/ML model in the user equipment and the base station, the difference between the first codeword set and the second codeword set is used to determine the environment drift, and then the model parameters are updated to ensure that the model adapts to the current environment.
Improves recovery accuracy and efficiency of CSI feedback, reduces the risk of model overfitting, and reduces dependence on original uncompressed CSI data.
Smart Images

Figure CN120019596A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to apparatus, methods and computer programs for communication systems, and particularly, but not exclusively, to apparatus, methods and computer programs related to codewords providing channel information. Background Art
[0002] A communication system may be considered as a facility that enables communication between two or more communication devices or provides communication devices with access to a data network.
[0003] The communication system may be a wireless communication system. Examples of wireless communication systems include public land mobile networks (PLMNs) operating based on radio access technology standards such as provided by 3GPP (3rd Generation Partnership Project) or ETSI (European Telecommunications Standards Institute), satellite communication systems, and different wireless local area networks, such as wireless local area networks (WLANs). Wireless communication systems operating based on radio access technologies may typically be divided into cells and are therefore often referred to as cellular systems.
[0004] The communication system and associated equipment typically operate according to one or more radio access technologies defined in a given standard specification, such as a standard provided by 3GPP or ETSI, which sets out what the various entities associated with the communication system and the communication equipment accessing or connected to the communication system are allowed to do and how it should be achieved. The communication protocols and / or parameters to be used by the communication equipment to access or connect to the communication system are also typically defined in the standard. Examples of standards are the so-called LTE (Long Term Evolution) and 5G (5th Generation) standards provided by 3GPP. Summary of the invention
[0005] According to one aspect, a method is provided, comprising: determining a difference based on information related to a first codeword set and information related to a second codeword set, wherein the first codeword set is received from a user equipment and provides information about a channel between the user equipment and a base station, and the user equipment generates the first codeword set using a first model trained using a first training data set.
[0006] The second set of codewords may be obtained from the stored set of data.
[0007] The stored data set may include a first training data set.
[0008] The method may include triggering determination of a difference in response to a system level indicator exceeding a threshold, and updating the first model using the difference.
[0009] The method may include determining whether the first model is to be updated based on the difference.
[0010] Determining whether the first model is to be updated may include comparing the difference to a threshold value.
[0011] Updating the first model includes: updating the neural network of the first model.
[0012] The method may include updating the first model by training the neural network of the first model using a back-propagation algorithm to determine one or more updated parameters for a layer of the neural network of the first model.
[0013] The method may include causing one or more updated parameters to be sent to the user device to update the first model on the user device.
[0014] The one or more updated parameters may be gradients for a layer of the neural network of the first model.
[0015] Updating the model may include retraining the model using an updated training data set.
[0016] Codewords can provide channel state information.
[0017] Codewords can provide channel information in a multiple-input multiple-output environment.
[0018] The method may include training a first model using a first training data set to provide encoding in a user device, and causing the first model to be provided to the user device.
[0019] The method may include training a second model to provide decoding in a base station, the training of the second model using the first training data set.
[0020] The method may include training a second model using an output of the first model to provide decoding in a base station.
[0021] The method may include determining a reconstruction loss based on an input to the first model and an output from the second model, and updating the first model according to the difference and the reconstruction loss.
[0022] The method may include determining the difference based on a measure of a distance between a distribution of the first codeword and a distribution of the second codeword.
[0023] The method may include updating the second model when it is determined that the first model is to be updated.
[0024] The method may be executed by a device. The device may be arranged in a base station or be a base station.
[0025] According to another aspect, an apparatus is provided, comprising: a component for determining a difference based on information related to a first codeword set and information related to a second codeword set, wherein the first codeword set is received from a user equipment and provides information about a channel between the user equipment and a base station, and the user equipment generates the first codeword set using a first model trained using a first training data set.
[0026] The second set of codewords may be obtained from the stored set of data.
[0027] The stored data set may include a first training data set.
[0028] The apparatus may include means for triggering determination of a difference in response to a system level indicator exceeding a threshold and updating the first model using the difference.
[0029] The apparatus may include means for determining whether the first model is to be updated based on the difference.
[0030] Determining whether the first model is to be updated may include comparing the difference to a threshold value.
[0031] Updating the first model includes: updating the neural network of the first model.
[0032] The apparatus may include means for updating the first model by training the neural network of the first model using a back-propagation algorithm to determine one or more updated parameters for a layer of the neural network of the first model.
[0033] The apparatus may include means for causing one or more updated parameters to be sent to the user equipment to update the first model on the user equipment.
[0034] The one or more updated parameters may be gradients for a layer of the neural network of the first model.
[0035] Updating the model may include retraining the model using an updated training data set.
[0036] Codewords can provide channel state information.
[0037] Codewords can provide channel information in a multiple-input multiple-output environment.
[0038] The apparatus may include means for training a first model using a first training data set to provide encoding in a user device, and causing the first model to be provided to the user device.
[0039] The method may include means for training a second model to provide decoding in the base station, the training of the second model using the first training data set.
[0040] The apparatus may include means for training a second model using an output of the first model to provide decoding in a base station.
[0041] The apparatus may include means for determining a reconstruction loss based on an input to the first model and an output from the second model, and updating the first model according to the difference and the reconstruction loss.
[0042] The apparatus may include means for determining a difference based on a measure of a distance between a distribution of first codewords and a distribution of second codewords.
[0043] The apparatus may include means for updating the second model when it is determined that the first model is to be updated.
[0044] The device may be arranged in a base station or be the base station.
[0045] According to another aspect, an apparatus is provided, comprising a circuit configured to determine a difference based on information related to a first set of codewords and information related to a second set of codewords, wherein the first set of codewords is received from a user equipment and provides information about a channel between the user equipment and a base station, and the user equipment generates the first set of codewords using a first model trained using a first training data set.
[0046] The second set of codewords may be obtained from the stored set of data.
[0047] The stored data set may include a first training data set.
[0048] The circuitry may be configured to trigger determination of a difference in response to the system level indicator exceeding a threshold, and to update the first model using the difference.
[0049] The circuitry may be configured to determine, based on the difference, whether the first model is to be updated.
[0050] Determining whether the first model is to be updated may include comparing the difference to a threshold value.
[0051] Updating the first model includes: updating the neural network of the first model.
[0052] The circuitry may be configured to update the first model by training the neural network of the first model using a back-propagation algorithm to determine one or more updated parameters for a layer of the neural network of the first model.
[0053] The circuitry may be configured to cause one or more updated parameters to be sent to the user device to update the first model on the user device.
[0054] The one or more updated parameters may be gradients for a layer of the neural network of the first model.
[0055] Updating the model may include retraining the model using an updated training data set.
[0056] Codewords can provide channel state information.
[0057] Codewords can provide channel information in a multiple-input multiple-output environment.
[0058] The circuit may be configured to train a first model using a first training data set to provide encoding in a user device, and cause the first model to be provided to the user device.
[0059] The circuit may be configured to train a second model to provide decoding in a base station, the training of the second model using the first training data set.
[0060] The circuitry may be configured to train a second model using an output of the first model to provide decoding in a base station.
[0061] The circuit may be configured to determine a reconstruction loss based on an input to the first model and an output from the second model, and to update the first model according to the difference and the reconstruction loss.
[0062] The circuitry may be configured to determine the difference based on a measure of a distance between a distribution of the first codewords and a distribution of the second codewords.
[0063] The circuit may be configured to update the second model when it is determined that the first model is to be updated.
[0064] The device may be arranged in a base station or be the base station.
[0065] According to another aspect, a device is provided, comprising: at least one processor; and at least one memory storing instructions, which, when executed by the at least one processor, cause the device to at least perform: determining a difference based on information related to a first codeword set and information related to a second codeword set, wherein the first codeword set is received from a user equipment and provides information about a channel between the user equipment and a base station, and the user equipment generates the first codeword set using a first model trained using a first training data set.
[0066] The second set of codewords may be obtained from the stored set of data.
[0067] The stored data set may include a first training data set.
[0068] The apparatus may be caused to trigger determination of a difference in response to the system level indicator exceeding a threshold, and to update the first model using the difference.
[0069] The apparatus may be caused to determine, based on the difference, whether the first model is to be updated.
[0070] Determining whether the first model is to be updated may include comparing the difference to a threshold value.
[0071] Updating the first model includes: updating the neural network of the first model.
[0072] The apparatus may be caused to update the first model by training the neural network of the first model using a back-propagation algorithm to determine one or more updated parameters for a layer of the neural network of the first model.
[0073] The apparatus may be caused to cause one or more updated parameters to be sent to the user equipment to update the first model on the user equipment.
[0074] The one or more updated parameters may be gradients for a layer of the neural network of the first model.
[0075] Updating the model may include retraining the model using an updated training data set.
[0076] Codewords can provide channel state information.
[0077] Codewords can provide channel information in a multiple-input multiple-output environment.
[0078] The apparatus may be caused to: train a first model using a first training data set to provide encoding in a user device, and cause the first model to be provided to the user device.
[0079] The apparatus may be configured to: train a second model to provide decoding in a base station, the training of the second model using the first training data set.
[0080] The apparatus may be caused to: train a second model using an output of the first model to provide decoding in a base station.
[0081] The apparatus may be caused to determine a reconstruction loss based on an input to the first model and an output from the second model, and to update the first model according to the difference and the reconstruction loss.
[0082] The apparatus may be caused to determine the difference based on a measure of a distance between a distribution of the first codewords and a distribution of the second codewords.
[0083] The apparatus may be caused to update the second model when it is determined that the first model is to be updated.
[0084] The device may be arranged in a base station or be the base station.
[0085] According to another aspect, a method is provided, comprising: generating a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; and receiving an update to the first model from the base station, wherein the update comprises one or more updated parameters of a layer of a neural network for the first model.
[0086] The first model may receive a set of channel information, which is encoded by the first model to generate a corresponding first codeword.
[0087] The one or more updated parameters may include gradients for a layer of the neural network of the first model.
[0088] Codewords can provide channel state information.
[0089] Codewords can provide channel information in a multiple-input multiple-output environment.
[0090] The method may be performed by a device. The device may be arranged in a user equipment or be a user equipment.
[0091] According to another aspect, an apparatus is provided, comprising: a component for generating a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; and a component for receiving an update to the first model from the base station, wherein the update comprises one or more updated parameters of a layer of a neural network for the first model.
[0092] The first model may receive a set of channel information, which is encoded by the first model to generate a corresponding first codeword.
[0093] The one or more updated parameters may include gradients for a layer of the neural network of the first model.
[0094] Codewords can provide channel state information.
[0095] Codewords can provide channel information in a multiple-input multiple-output environment.
[0096] The apparatus may be arranged in a user equipment or be a user equipment.
[0097] According to another aspect, an apparatus is provided, comprising a circuit configured to: generate a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; and receive an update to the first model from the base station, wherein the update comprises one or more updated parameters of a layer of a neural network for the first model.
[0098] The first model may receive a set of channel information, which is encoded by the first model to generate a corresponding first codeword.
[0099] The one or more updated parameters may include gradients for a layer of the neural network of the first model.
[0100] Codewords can provide channel state information.
[0101] Codewords can provide channel information in a multiple-input multiple-output environment.
[0102] The apparatus may be arranged in a user equipment or be a user equipment.
[0103] According to another aspect, an apparatus is provided, comprising: at least one processor; and at least one memory storing instructions, which instructions, when executed by the at least one processor, cause the apparatus to at least: generate a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; and receive an update to the first model from the base station, wherein the update comprises one or more updated parameters of a layer of a neural network for the first model.
[0104] The first model may receive a set of channel information, which is encoded by the first model to generate a corresponding first codeword.
[0105] The one or more updated parameters may include gradients for a layer of the neural network of the first model.
[0106] Codewords can provide channel state information.
[0107] Codewords can provide channel information in a multiple-input multiple-output environment.
[0108] The apparatus may be arranged in a user equipment or be a user equipment.
[0109] According to another aspect, a method is provided, comprising: generating a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; and receiving an update to the first model from the base station.
[0110] The first model may receive a set of channel information, which is encoded by the first model to generate a corresponding first codeword.
[0111] Updating may include updating the neural networks of the first model and the second model.
[0112] The update may include one or more updated parameters of the layers of the neural networks for the first model and the second model.
[0113] To update the first model in the user equipment, one or more updated parameter gradients for the layers of the neural network of the first model may be sent from the base station to the user equipment.
[0114] Codewords can provide channel state information.
[0115] Codewords can provide channel information in a multiple-input multiple-output environment.
[0116] The method may be performed by a device. The device may be arranged in a user equipment or be a user equipment.
[0117] According to another aspect, an apparatus is provided, comprising: means for generating a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; and means for receiving an update to the first model from the base station.
[0118] The first model may receive a set of channel information, which is encoded by the first model to generate a corresponding first codeword.
[0119] Updating may include updating the neural network of the first model.
[0120] The update may include one or more updated parameters for a layer of the neural network of the first model.
[0121] The one or more updated parameters may include gradients for a layer of the neural network of the first model.
[0122] Codewords can provide channel state information.
[0123] Codewords can provide channel information in a multiple-input multiple-output environment.
[0124] The apparatus may be arranged in a user equipment or be a user equipment.
[0125] According to another aspect, an apparatus is provided that includes circuitry configured to: generate a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; and receive an update to the first model from the base station.
[0126] The first model may receive a set of channel information, which is encoded by the first model to generate a corresponding first codeword.
[0127] Updating may include updating the neural network of the first model.
[0128] The update may include one or more updated parameters for a layer of the neural network of the first model.
[0129] The one or more updated parameters may include gradients for a layer of the neural network of the first model.
[0130] Codewords can provide channel state information.
[0131] Codewords can provide channel information in a multiple-input multiple-output environment.
[0132] The apparatus may be arranged in a user equipment or be a user equipment.
[0133] According to another aspect, an apparatus is provided, comprising: at least one processor; and at least one memory storing instructions, which, when executed by the at least one processor, causes the apparatus to at least: generate a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; and receive an update to the first model from the base station.
[0134] The first model may receive a set of channel information, which is encoded by the first model to generate a corresponding first codeword.
[0135] Updating may include updating the neural network of the first model.
[0136] The update may include one or more updated parameters for a layer of the neural network of the first model.
[0137] The one or more updated parameters may include layer gradients for the neural network of the first model.
[0138] Codewords can provide channel state information.
[0139] Codewords can provide channel information in a multiple-input multiple-output environment.
[0140] The apparatus may be arranged in a user equipment or be a user equipment.
[0141] According to another aspect, there is provided a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform any of the methods previously set forth.
[0142] According to another aspect, there is provided a computer program comprising instructions which, when executed, cause any of the methods previously set out to be performed.
[0143] According to an aspect, there is provided a computer program comprising computer executable code which, when executed, causes any of the methods set out previously to be performed.
[0144] According to one aspect, a computer-readable medium is provided, comprising program instructions stored thereon for executing at least one of the above methods.
[0145] According to an aspect, a non-transitory computer-readable medium comprising program instructions is provided, which, when executed by an apparatus, causes the apparatus to perform any of the methods previously set forth.
[0146] According to an aspect, there is provided a non-transitory computer readable medium comprising program instructions which, when executed, cause any of the methods set forth previously to be performed.
[0147] According to one aspect, a non-volatile tangible storage medium is provided, comprising program instructions stored thereon for executing at least one of the above methods.
[0148] In the above, many different aspects have been described. It should be understood that further aspects can be provided by combining any two or more of the above aspects.
[0149] Various other aspects are also described in the following detailed description and appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0150] Some example embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0151] Figure 1 A schematic diagram of a 5G system is shown;
[0152] Figure 2 A schematic diagram of the device is shown;
[0153] Figure 3 A schematic representation of a user equipment is shown;
[0154] Figure 4 A schematic representation of the autoencoder architecture is shown;
[0155] Figure 5a shows the distribution of the kth feature learned from the dataset during the training set;
[0156] Figure 5b shows the distribution of the kth feature in the data in the deployment environment;
[0157] Figure 6 shows a schematic representation of the autoencoder architecture of some embodiments;
[0158] Figure 7 Methods of some embodiments are shown;
[0159] Figure 8 Schematically illustrates the encoder and decoder of some embodiments;
[0160] Figure 9a and Figure 9b Some simulation results are shown;
[0161] Fig.10 Another method of some embodiments is shown;
[0162] Fig.11 Another method of some embodiments is shown; and
[0163] Fig.12A schematic representation of a non-volatile memory medium storing instructions that, when executed by a processor, allow the processor to perform Figure 7 , Fig.10 and Fig.11 One or more steps of any method. DETAILED DESCRIPTION
[0164] In the following, certain embodiments are explained with reference to a communication device capable of communicating via a wireless cellular system and a mobile communication system serving such a communication device. Figure 1 , Figure 2 and Figure 3 Certain general principles of wireless communication systems, their access systems and communication devices are briefly explained to help understand the underlying technology of the described examples.
[0165] Figure 1 A schematic representation of a communication system operating based on the 5th generation radio access technology, commonly referred to as a 5G system (5GS), is shown. The 5GS may be a (radio) access network ((R)AN), a 5G core network (5GC), one or more application functions (AFs), and one or more data networks (DNs). User equipment may access or connect to one or more DNs via the 5GS.
[0166] The 5G(R)AN may include one or more base stations or radio access network (RAN) nodes, such as gNodeB (gNB). A base station (BS) or RAN node may include one or more distributed units connected to a central unit.
[0167] 5GC may include various network functions such as access and mobility management function (AMF), session management function (SMF), authentication server function (AUSF), user data management (UDM), user plane function (UPF), network data analysis function (NWDAF) and / or network open function (NEF). The operations performed by each of the various network functions of 5G are described in 3GPP TS23.501 and TS23.502 version 16 by way of example only.
[0168] Figure 2An example of an apparatus 200 is shown. The apparatus 200 may be provided as a radio access node such as a base station. The apparatus 200 may have at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause one or more functions to be performed. In this example, the apparatus may include at least one random access memory (RAM) 211a and / or at least one read-only memory (ROM) 211b and / or at least one processor 212, 213 and / or an input / output interface 214. At least one processor 212, 213 may be coupled to the RAM 211a and the ROM 211b. At least one processor 212, 213 may be configured to execute appropriate software code 215. The software code 215 may, for example, allow one or more steps to be performed to perform one or more aspects of the present disclosure.
[0169] Figure 3 An example of a communication device 300 is shown. The communication device 300 may be any device capable of sending and receiving radio signals. Non-limiting examples of the communication device 300 include user equipment (such as Figure 1 The communication device 300 may be a user device (user device) as shown), a mobile station (MS) or a mobile device (such as a mobile phone or so-called "smart phone"), a computer provided with a wireless interface card or other wireless interface facilities (e.g., a USB software dongle), a personal data assistant (PDA) or a tablet provided with wireless communication capabilities, a machine type communication (MTC) device, a cellular Internet of Things (CIoT) device, or any combination of these, etc. The communication device 300 can send or receive, for example, radio signals carrying communications. The communication can be one or more of voice, electronic mail (email), text message, multimedia, data, machine data, etc.
[0170] The communication device 300 may receive radio signals via an appropriate means for receiving over the air or radio interface 307, and may transmit radio signals via an appropriate means for transmitting radio signals. Figure 3 In the embodiment of the present invention, the transceiver device is schematically designated by a block 306. The transceiver device 306 may be provided, for example, by a radio component and an associated antenna arrangement. The antenna arrangement may be arranged inside or outside the mobile device and may include a single antenna or multiple antennas. The antenna arrangement may be an antenna array including multiple antenna elements.
[0171] The communication device 300 may be provided with at least one processor 301 and / or at least one ROM 302a and / or at least one RAM 302b and / or other possible components 303 for software and hardware-assisted execution of the tasks it is designed to perform, including control of access and communication to access systems (such as 5G RAN and other communication devices). At least one processor 301 is coupled to RAM 302b and ROM 302a. At least one processor 301 may be configured to execute instructions of software code 308. Execution of the instructions of software code 308 may, for example, allow the communication device 300 to perform one or more operations. Software code 308 may be stored in ROM 302a. It should be understood that in other embodiments, any other suitable memory may be used alternatively or additionally with the above-mentioned ROM and / or RAM examples.
[0172] The at least one processor 301 , the at least one ROM 302a and / or the at least one RAM 302b may be provided on a suitable circuit board, in an integrated circuit and / or in a chipset. This feature is indicated by reference numeral 304 .
[0173] The communication device 300 may optionally have a user interface, such as a keypad 305, a touch-sensitive screen or pad, a combination thereof, etc. Optionally, the communication device may have one or more of a display, a speaker, and a microphone.
[0174] In the following examples, the term UE or User Equipment is used. This term encompasses any examples of the previously discussed communication device 300 and / or any other communication device.
[0175] An example of a wireless communication system is an architecture standardized by the Third Generation Partnership Project (3GPP). Current radio access technologies standardized by 3GPP are generally referred to as 5G or NR. Other radio access technologies standardized by 3GPP include Long Term Evolution (LTE) or LTE Advanced Pro of Universal Mobile Telecommunications System (UMTS). A wireless communication system generally includes an access network, such as a radio access network operating based on a radio access technology, which includes a base station or a radio access network node. A wireless communication system may also include other types of access networks, such as a wireless local area network (WLAN) and / or a WiMAX (Worldwide Interoperability for Microwave Access) network. It should be understood that the example embodiments may also be used with standards for future radio access technologies such as 6G or higher.
[0176] Downlink channel state information (CSI) is used by a base station (BS) to obtain channel response and precoding for beamforming in the downlink of a massive multiple-input multiple-output (MIMO) system. In a MIMO system, downlink CSI is first estimated by a user equipment (UE) using a pilot signal and then sent back to the BS as feedback. However, due to the large number of antennas in a massive MIMO system, CSI feedback (based on, for example, codebook-based methods and compressed sensing methods) may consume bandwidth. In such a MIMO system, this method may be relatively complex.
[0177] Figure 4 FIG. 4 shows a CSI feedback enhancement method based on AI / ML (artificial intelligence / machine learning). Figure 4 As shown, an auto-encoder architecture is provided. A set of CSI values is determined at UE 400 for the MIMO system to provide a CSI data set 404. The original CSI data set 404 is compressed by an encoder 406 at UE 400 into codewords 412, which are then sent to BS / gNB 402. Upon receiving the CSI feedback codewords, a decoder at BS / gNB 402 can reconstruct the CSI to provide a reconstructed CSI data set 410. The CSI feedback codewords 412 compress representative information of the input CSI data. These compressed CSI feedback codewords constitute a space called a feature space, where the statistical characteristics of the codewords can be represented by the distribution of each feature / dimension.
[0178] Enhancing CSI feedback may improve performance. For example, there may be overhead reduction, CSI recovery accuracy improvement (leading to better performance), and / or prediction enhancement.
[0179] This may be in the context of different gNB-UE cooperation levels that need to be supported.
[0180] The encoder in the UE and the decoder in the gNB / BS can use AI / ML models. The adaptability of AI / ML models should be considered when designing CSI feedback solutions that support AI / ML. In this regard, refer to Figure 5a and Figure 5b . Figure 5a The distribution of the k-th feature learned from the dataset during the training set is shown. Figure 5b shows the distribution of the kth feature in the data in the deployment environment. Figure 5a and Figure 5b It can be seen from the comparison that there may be distribution drift in the feature space due to environmental drift.
[0181] One option could be to retrain the model to fit the current environment. Figure 4As shown, the encoder and decoder are deployed in the UE and gNB / BS respectively, and will be jointly (re)trained in the gNB / UE before deployment. This may require a large amount of uncompressed raw downlink CSI data to be transmitted from the UE back to the gNB. This may require relatively large resource expenditures for air data services as well as data storage.
[0182] Some embodiments may address issues related to overfitting to pre-trained models. Figure 5a and Figure 5b As discussed, changes in the RF (Radio Frequency) propagation environment may cause a drift in the CSI distribution. A pre-trained model for CSI feedback compression with a previous channel distribution may not be suitable for the new environment. This is known as the overfitting problem.
[0183] Some embodiments may address issues related to traffic density when sending raw uncompressed CSI for model retraining. In order to retrain the model to adapt to the updated propagation environment, UE-to-gNB data transmission of raw uncompressed CSI data may be traffic intensive. Furthermore, constant monitoring of channel state changes may make traffic density a constant issue in CSI feedback transmission.
[0184] Some embodiments may transfer the CSI feedback compression model to new environments in an unsupervised learning manner without requiring any labeled CSI live data (ie, retraining the model with original uncompressed CSI data).
[0185] Referring to the schematic diagram of the embodiment Figure 6 .
[0186] UE 600 has an encoder 606. The encoder has a trained neural network or AI / ML model. The trained neural network or AI / ML model can be downloaded from gNB / BS 602. Live CSI data from a deployment live environment t 605 is compressed by the encoder 606 on the UE into a vector z t 607 and is sent to the gNB.
[0187] The gNB / BS 602 trains the encoder / decoder NN or AI / ML model. The gNB / BS uses a pre-stored dataset 612X s Data x s The NN or AI model is trained as input. The NN or ML / AI model has an encoder part 614 and a decoder part 618. The output of the decoder part 618 is a reconstructed data set 620X s . Pre-stored datasets 612X s Data x sis provided as input to the encoder part of the NN or AI / ML model. The encoder part of the NN or AI / ML model provides a codeword output z s , which is input to the decoder part of the model. The encoder part 614 and the decoder part 618 are trained so that the data output by the decoder matches the input of the encoder. The trained encoder part of the NN or ML / AI model is downloaded by the UE.
[0188] It should be understood that the NN or ML / AI model for the encoder / decoder can be implemented using any suitable deep network architecture, such as fully connected (FC) layers, convolutional layers, long short-term memory (LSTM) networks, etc.
[0189] In some embodiments, a “signature difference” metric is calculated or determined to assess the z (received from the UE) s and z t The feature difference indicates the importance of environmental drift between training and deployment. In some embodiments, the feature difference is calculated by the domain adaptation module Adp(z t , z s ) is implemented by functional block 610. The domain adaptation module can monitor feature differences.
[0190] The domain adaptation module may use any suitable technique.
[0191] For example, the domain adaptation module may use a deep adaptation method, such as a difference-based formula. The difference-based formula may be MMD (maximum mean difference). In some embodiments, the method is not implemented by a NN (neural network)-based method.
[0192] In another example, the domain adaptation module may use a domain adversarial method. The domain adversarial method may use a NN-based domain classifier.
[0193] In another example, the domain adaptation module may use a difference-based n method. In this example, the module may be implemented by at least one processor and at least one memory.
[0194] The domain adaptation module can be used for model monitoring and / or model fine-tuning. In the domain adaptation module, the difference between the pre-stored environment and the drift environment can be obtained by the compression vector z from the pre-stored environment. s and the compression vector z from the drift environment t The input to the domain adaptation module is the pre-stored CSI codeword dataset and the live codeword dataset.
[0195] Using the determined differences, environmental drift can be detected in the model monitoring mode. By minimizing the differences, the model can be fine-tuned to make environmental drift indistinguishable in the model fine-tuning mode.
[0196] The difference may be determined in any suitable manner.
[0197] In some embodiments, a deep adaptive method may be used to determine the difference. The deep adaptive method may be based on the difference between the pre-stored environment distribution and the live environment distribution. The distribution of different environments is determined from the codeword data sets of different environments.
[0198] In some embodiments, measures of distance between distributions are used to determine differences.
[0199] Some examples of measuring distance include Kullback-Leibler divergence (KL divergence), Jensen-Shannon divergence (JS divergence), maximum mean divergence (MMD), and / or Wasserstein distance, among others.
[0200] The Wasserstein distance or Kantorovich-Rubinstein metric is a distance function defined between probability distributions on a given metric space. In this case, a distribution over a pre-stored environment distribution and a live environment distribution.
[0201] The Kullback-Leibler divergence (KL divergence) is a measure of the distance between two probability distributions. It is sometimes called relative entropy. In this case, the distributions over the pre-stored environment distribution and the live environment distribution.
[0202] Jensen-Shannon divergence is a way to measure the similarity between two probability distributions. In this case, the distribution over the pre-stored environment distribution and the live environment distribution.
[0203] If the distance exceeds a certain threshold, it can be considered that the environment drift is large and the pre-learned model is not suitable for the current propagation environment.
[0204] The Maximum Mean Difference (MMD) is a statistical test used to determine whether two given distributions are identical.
[0205] In the following examples, MMD is used as an example in deep adaptive methods. MMD is defined as measuring the difference between two distributions. In practice, an empirical estimate of MMD is used by computing the empirical expectation of samples X and Y. Where F is a class of functions f:X→R.
[0206] When implemented, the MMD can be determined using kernel embedding techniques as MMD(p,q)=E xx′ k(x,x′)+E yy′ k(y,y′)-2E xy k(x,y) where k(·,·) can be any general kernel such as Gaussian x, x′~p, and y, y′~q.
[0207] In some embodiments, a domain adversarial method can be used to determine the difference. The domain adversarial method distinguishes whether the data is from the pre-learned environment or from the live environment based on the domain classifier. If the result is indifferent, it means that there is little difference between the pre-stored environment and the live environment. If the result is different, it means that there is a difference between the pre-stored environment and the live environment. A gradient reversal layer is added to the classifier, which is intended to promote the indifference of features relative to environmental drift.
[0208] Therefore, if the feature difference is less than the threshold, this indicates that the deployment environment has a relatively high similarity with the training dataset and the current model is suitable. If the difference is greater than the threshold, retraining can be activated.
[0209] Alternatively or additionally, environmental drift may be determined based on one or more system-level indicators. The system-level indicator may include one or more of: a key performance indicator (KPI); a suitable parameter; and / or a suitable metric. The KPI may be, for example, downlink throughput. If the indicator drops below a threshold (or rises above a threshold), retraining may be activated.
[0210] In some embodiments, downlink throughput is used as a system-level metric to indicate whether there is significant environmental drift. Once the downlink throughput is below a threshold, Figure 6 The adaptive loss (loss 2) calculated in the domain adaptation module is activated for model retraining.
[0211] If retraining is required, the loss L of the feature difference determined by the domain adaptation module adp and the reconstruction loss L in the pre-stored training dataset rec The sum is used to form the total loss L. The three ML blocks (i.e., domain adaptation module, encoder, and decoder) can update their NN parameters relative to the gradient of L. It should be noted that in the case where retraining is initiated based on system-level KPIs, it is still possible to update the feature difference L based on the feature difference L. adp (rather than system-level KPIs) to calculate the loss function itself.
[0212] The reconstruction loss L may be determined in any suitable mannerrec The reconstruction loss L rec It can be viewed as a measure of how similar (or different) the input CSI data is to the model and the output reconstructed CSI data provided by the model.
[0213] For example, the reconstruction loss L rec can be set to the cosine similarity between the input CSI data and the output reconstructed CSI data, which can be expressed as where w i is the input raw CSI vector for frequency unit i, is the output CSI vector of frequency bin i, N is the total number of frequency bins, and E{·} denotes an averaging operation over multiple samples.
[0214] The adaptation loss may be presented depending on the method used by the domain adaptation module.
[0215] For example, for deep adaptation methods, the adaptation loss can be defined as the distance between the pre-stored environment distribution in the domain adaptation module and the live environment distribution. L adp = distance(t,s) where distance(·,·) can be any of the distances mentioned above. For example, if MMD is used as the distance to measure the difference between the compressed feature spaces t and s, the above formula can be rewritten as distance(t,s)=MMD(t,s).
[0216] For example, for domain adversarial methods, the adaptive loss can be defined as Where L BCE (·,label) is the binary cross entropy (BCE) loss, where 1 and 0 are labels representing different environments.
[0217] The total loss can be the sum of the reconstruction loss and the adaptation loss, given as L=L rec +L adp
[0218] The NN is trained using stochastic gradient descent with back-propagation by minimizing the total loss. It should be noted that to update the encoder in the UE, according to the theory of the back-propagation algorithm, the gNB only needs to send the gradients of the last layer in the encoder to the UE. In this case, the size of the gradients of the last layer in the encoder can be smaller (e.g., the size of each input CSI sample is denoted as N, the compression ratio is denoted as γ, and the size of the nodes of the first layer of the decoder will be γ·N).
[0219] In the absence of the need for live raw uncompressed CSI data, the ML-based CSI compression and recovery model of some embodiments can improve its recovery accuracy in a deployment environment. In some embodiments, "overfitting" of the training data set can be avoided.
[0220] The methods of some embodiments do not require labeled training data. Due to model adaptation, this can significantly reduce air transmission overhead.
[0221] Some embodiments may provide methods for consistent training-deployment feature difference monitoring.
[0222] In some embodiments, by minimizing the distance between the pre-stored environment distribution and the live environment distribution in the domain adaptation module, representations from both environments are learned and the model can be applied in the live environment with minimal loss in reconstruction accuracy.
[0223] refer to Figure 7 , which illustrates the methods of some embodiments.
[0224] In step 1, the autoencoder model is deployed on the gNB for CSI feedback reconstruction.
[0225] With pre-stored CSI data x s as input to train the model.
[0226] The output of the model is the reconstructed CSI data
[0227] Determine the reconstructed CSI data The corresponding input CSI data x s The reconstruction error is expressed as the reconstruction loss L rec .
[0228] In step 2, the live CSI data x t As input, the encoder is deployed on the UE. The output compressed CSI vector z _t Sent to gNB.
[0229] In step 3, the compressed CSI vectors z_ s and z_ t It is fed into the domain adaptation module to determine the difference between the pre-stored environment and the drifted environment. This difference is the adaptation loss L adp .
[0230] In step 4, if the difference exceeds a threshold indicating a relatively large change in the propagation environment, the total loss L = L rec +L adpOtherwise, the loss L = L rec In another embodiment, the difference may be determined based on one or more indicators such as those discussed previously. If the indicator meets the criteria, then L adp Added to the loss term to initiate adaptation.
[0231] In step 5, the NN on gNB and UE is fine-tuned by minimizing the total loss L.
[0232] Now some example simulations are described. A simulation dataset is generated for a link-level feature vector based CSI feedback study according to 3GPP TR 38.901. The dataset configuration is given below.
[0233] CDLC30 represents a CDLC channel model with a 30 ns delay spread, and CDLC300 represents a CDLC channel model with a 300 ns delay spread.
[0234] In the simulation, two cases are tested: 52 resource blocks (RB), CDLC30->CDLC300 48RB, CDLC300->CDLC30
[0235] For each case, Figure 4 The model shown (i.e., without the domain adaptation module) is used as a baseline. In the baseline scheme, the model is trained using pre-stored data and tested on both pre-stored data and live data. The sample numbers used for model training and testing are presented in the table below.
[0236] In the simulation, each sample consists of 832 real numbers, which corresponds to a large eigenvector concatenated from 13 subbands as follows: w=[w1,w2,…,w 13 ] Among them, w k (1≤k≤13) is the eigenvector for the kth subband channel. Each w k It has been processed into the following format: w k =[Re{w k,1},Im{w k,1},Re{w k,2},Im{w k,2},…,Re{w k,32},Im{w k,32}] where Re{.} and Im{.} are the real and imaginary parts.
[0237] An example NN architecture is used for the proposed scheme. First, the encoder input size is denoted as N (N=832 in this simulation), and the compression ratio is denoted as γ (γ=1 / 64 in this simulation).
[0238] Figure 8 An example encoder and decoder are shown in FIG. It should be understood that the encoder and decoder of an embodiment may have more Figure 8 The examples shown may have fewer or more layers. Figure 8 Different embodiments may use one or more different layers, or alternatively one or more layers are shown. The number of neurons in each layer is provided as an example.
[0239] The encoder 800 has three fully connected FC layers 802, 804, and 806. Each FC layer is followed by a batch normalization BN / activation layer 808, 810, and 812. In this example, the activation function in the neural network implementation is a leaky ReLu (rectified linear unit) function.
[0240] The input N is received by the first FC layer 802, and the output N.γ of the encoder is provided by the third BN / activation layer 812. In this example, the first FC layer 802 has 4N neurons, the second FC layer 804 has 4N neurons, and the third FC layer 802 has N.γ neurons.
[0241] The decoder 801 has three fully connected FC layers 814, 816, and 818. Each FC layer is followed by a batch normalization BN / activation layer 820, 822, and 824. In this example, the activation function is leaky ReLu.
[0242] The input is the output of the encoder -N.γ. This output is received by the first FC layer 814, and the output of the decoder N is provided by the third BN / activation layer 824. In this example, the first FC layer 814 has N.γ neurons, the second FC layer 816 has 4N neurons, and the third FC layer 818 has 4N neurons.
[0243] The encoder and decoder will mirror each other in terms of layers.
[0244] In some embodiments, since the CSI is fed back in the form of a bit stream, a quantizer 830 may be used. Figure 8 , where the output of the encoder is input to the quantizer, and the output of the quantizer is input to the decoder. The quantizer can be implemented by uniform quantization or non-uniform quantization. In this example, uniform quantization is used, and the quantization can be written as: Where s is the output of the encoder, s q is the output of the quantizer, and B is the number of quantization bits. In this example, MMD is used as a domain adaptation module to minimize the difference between pre-stored data and live data.
[0245] refer to Figure 9a and Figure 9b , which shows a plot of cosine similarity versus time (epoch). Figure 9a It is 52RB, CDLC30->CDLC300, and Figure 9b It is 48RB, CDLC300->CDLC30. In the following, source refers to pre-stored data, and target refers to live data. Figure 9a , the figure marked 900 is the source with domain adaptation, the figure marked 902 is the target with domain adaptation, and the figure marked 904 is the source without domain adaptation (baseline). Figure 9b , the one labeled 906 is the target with domain adaptation, the one labeled 908 is the target without domain adaptation, and the one labeled 910 is the source without domain adaptation.
[0246] exist Figure 9a In , the source is CDLC30 and the target is CDLC300. The delay spread of CDLC30 is equal to 30ns, making its CSI pattern "flatter / easier" than the CSI pattern of CDLC300. Therefore, regardless of whether CDLC30 is the source domain or the target domain, the model generally presents better CSI feedback accuracy in CDLC30 than in CDLC300.
[0247] like Figure 9a and Figure 9b As shown, it can be observed that the unsupervised learning method of some embodiments can enhance the CSI feedback reconstruction accuracy in a live environment. Figure 9a and Figure 9b As shown, respective lines 902 and 906 present higher CSI feedback than respective lines 904 and 908 (with domain adaptation). This indicates that some embodiments may enhance CSI feedback reconstruction accuracy in a live environment.
[0248] refer to Fig.10 , which illustrates the methods of some embodiments.
[0249] The method may be performed by an apparatus. The apparatus may be in a base station or may be a base station.
[0250] The apparatus may comprise suitable circuitry for providing the method.
[0251] Alternatively or additionally, the apparatus may include at least one processor and at least one memory storing instructions, which when executed by the at least one processor cause the apparatus to at least provide the following method.
[0252] Alternatively or additionally, the device may be, for example, Figure 2 discussed.
[0253] The method may be provided by computer program code or computer executable instructions.
[0254] The method may include: as reference A1, determining a difference based on information related to a first codeword set and information related to a second codeword set, wherein the first codeword set is received from a user equipment and provides information about a channel between the user equipment and a base station, and the user equipment generates the first codeword set using a first model trained using a first training data set.
[0255] It should be understood that Fig.10 The methods outlined in may be modified to include any of the previously described features.
[0256] refer to Fig.11 , which illustrates another approach of some embodiments.
[0257] The method may be performed by an apparatus. The apparatus may be in a user equipment or may be a user equipment.
[0258] The apparatus may comprise suitable circuitry for providing the method.
[0259] Alternatively or additionally, the apparatus may include at least one processor and at least one memory storing instructions, which when executed by the at least one processor cause the apparatus to at least provide the following method.
[0260] Alternatively or additionally, the device may be, for example, Figure 3 discussed.
[0261] The method may be provided by computer program code or computer executable instructions.
[0262] The method may include: as referenced to B1, using a first model to generate a first codeword, the first codeword providing information about a channel between the user equipment and the base station.
[0263] The method may include, as with reference B2, receiving an update to the first model from the base station, wherein the update includes one or more updated parameters of a layer of the neural network for the first model.
[0264] It should be understood that Fig.11 The methods outlined in may be modified to include any of the previously described features.
[0265] Fig.12 A schematic representation of a non-volatile memory medium 900a or 900b storing instructions and / or parameters is shown, which when executed by a processor allows the processor to perform one or more steps of the method of any embodiment. The non-volatile memory medium can be a computer disk (CD) or digital versatile disk (DVD), schematically marked as 900a, or a universal serial bus (USB) memory stick, schematically marked as 900b. Computer instructions or codes can be downloaded and stored in one or more memories. The memory medium can store instructions and / or parameters 902, which when executed by a processor allows the processor to perform one or more steps of the method of an embodiment.
[0266] The computer program code may be downloaded and stored in one or more memories of the device.
[0267] Note that while the above describes exemplifying embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention.
[0268] Note that although some embodiments have been described with respect to 5G networks, similar principles may be applied with respect to the standard.
[0269] Thus, although certain embodiments are described above by way of example with reference to certain example architectures for wireless networks, technologies and standards, the embodiments may be applied to any other suitable form of communication system in addition to the communication system shown and described herein.
[0270] It is also noted herein that while the above describes example embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the present invention.
[0271] As used herein, “at least one of: ” and “at least one of ” and similar expressions, where a list of two or more elements is combined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all of the elements.
[0272] In general, various embodiments may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the disclosure may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the disclosure is not limited thereto. Although various aspects of the disclosure may be illustrated and described as block diagrams, flow charts, or using some other graphical representation, it is well understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0273] As used in this application, the term "circuitry" may refer to one or more or all of the following: (a) hardware circuit implementations only (such as implementations only in analog and / or digital circuitry), and (b) a combination of hardware circuitry and software, for example (where applicable): (i) a combination of analog and / or digital hardware circuitry and software / firmware, and (ii) any portion of a hardware processor (including a digital signal processor), software and memory with software that work together to enable a device such as a mobile phone or server to perform various functions; and (c) A hardware circuit and / or processor that requires software (eg, firmware) for operation, such as a microprocessor or portion of a microprocessor, but the software may not be present when not required for operation.
[0274] This definition of circuitry applies to all uses of the term in this application, including any claims. As another example, as used in this application, the term circuitry also covers an implementation of only a hardware circuit or processor (or multiple processors) or a portion of a hardware circuit or processor and its (or its) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in a server, cellular network device, or other computing or network device.
[0275] Embodiments of the present disclosure may be implemented by computer software executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products, including software routines, small applications and / or macros) may be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer executable components that are configured to perform an embodiment when the program is running. One or more computer executable components may be at least one software code or part thereof.
[0276] Further, in this regard, it should be noted that any block of the logic flow as in the accompanying drawings may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants CDs. Physical media are non-transitory media.
[0277] As used herein, the term "non-transitory" is a limitation of the medium itself (ie, tangible, not a signal), not a limitation on data storage persistence (eg, RAM versus ROM).
[0278] The memory may be of any type suitable for the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. As non-limiting examples, the data processor may be of any type suitable for the local technical environment and may include one or more of the following: a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an FPGA, a gate-level circuit, and a processor based on a multi-core processor architecture.
[0279] Embodiments of the present disclosure may be practiced in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available to convert logic level designs into semiconductor circuit designs that are ready to be etched and formed on semiconductor substrates.
[0280] The scope of protection sought by various embodiments of the present disclosure is set forth by the claims. Embodiments and features described in this specification that do not fall within the scope of the claims, if any, should be interpreted as examples that aid in understanding the various embodiments of the present disclosure.
[0281] It should be noted that different claims with different claim scopes may be pursued in related applications such as divisional applications or continuation applications.
[0282] The foregoing description has provided a complete and informative description of exemplary embodiments of the present disclosure by way of non-limiting examples. However, when read in conjunction with the accompanying drawings and the appended claims, various modifications and adaptations may become apparent to those skilled in the relevant art in view of the foregoing description. However, all such and similar modifications of the teachings of the present disclosure will still fall within the scope of the present invention as defined in the appended claims. In fact, there is another embodiment comprising a combination of one or more embodiments with any other embodiment previously discussed.
Claims
1. A method comprising: The difference is determined based on information related to a first set of codewords and information related to a second set of codewords, wherein the first set of codewords is received from a user equipment and provides information about a channel between the user equipment and a base station, and the user equipment generates the first set of codewords using a first model trained using a first training data set.
2. The method of claim 1, wherein the second set of codewords is obtained from a stored data set. The method of claim 2 , wherein the stored data set comprises the first training data set.
4. A method according to any one of the preceding claims, comprising: Determination of the difference is triggered in response to a system level indicator exceeding a threshold, and the difference is used to update the first model.
5. The method according to any one of claims 1 to 4, comprising: Based on the difference, it is determined whether the first model is to be updated.
6. The method of claim 5, wherein determining whether the first model is to be updated comprises: The difference is compared to a threshold value.
7. The method according to any one of the preceding claims, wherein updating the first model comprises: The neural network of the first model is updated.
8. The method according to claim 7, comprising: The first model is updated by training the neural network of the first model using a back-propagation algorithm to determine one or more updated parameters for a layer of the neural network of the first model.
9. The method according to claim 8, comprising: The one or more updated parameters are caused to be sent to the user device to update the first model on the user device.
10. The method of claim 8 or 9, wherein the one or more updated parameters are gradients for the layers of the neural network of the first model.
11. A method according to any preceding claim, wherein the codeword provides channel state information.
12. A method according to any preceding claim, wherein the codeword provides channel information in a multiple-input multiple-output environment.
13. A method according to any one of the preceding claims, comprising: The first model is trained using the first training data set to provide encoding in the user device, and the first model is provided to the user device.
14. The method according to claim 13, comprising: A second model is trained to provide decoding in the base station, the training of the second model using the first training data set.
15. The method according to claim 14, comprising: The second model is trained using the output of the first model to provide decoding in the base station.
16. The method according to claim 13 or 14, comprising: A reconstruction loss is determined based on an input to the first model and an output from the second model, and the first model is updated according to the difference and the reconstruction loss.
17. A method according to any one of the preceding claims, comprising: The difference is determined based on a measure of a distance between a distribution of the first codewords and a distribution of the second codewords.
18. An apparatus comprising: at least one processor; as well as At least one memory stores instructions, which, when executed by the at least one processor, cause the apparatus to at least perform the method according to any one of claims 1 to 17.
19. A method comprising: generating a first codeword using a first model, the first codeword providing information about a channel between a user equipment and a base station; as well as An update to the first model is received from the base station, wherein the update comprises one or more updated parameters for a layer of a neural network of the first model.
20. The method of claim 18, wherein the first model receives a set of channel information, the set of channel information being encoded by the first model to generate the corresponding first codeword.
21. The method of claim 19 or 20, wherein the one or more updated parameters are gradients for the layers of the neural network of the first model.
22. A method according to any one of claims 19 to 21, wherein the codeword provides channel state information.
23. The method according to any one of claims 19 to 22, wherein the codeword provides channel information in a multiple-input multiple-output environment.
24. An apparatus comprising: at least one processor; as well as At least one memory stores instructions, which, when executed by the at least one processor, cause the apparatus to at least perform the method according to any one of claims 19 to 23.