Independent training mode for channel state information feedback
By training the codebook and codeword index on the network device side and performing CSI feedback compression on the terminal device side, the problem of high CSI feedback resource consumption in frequency division duplex MIMO systems is solved, achieving performance enhancement and resource optimization of CSI feedback.
Patent Information
- Application Number
- CN202380096907.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2025-11-14
AI Technical Summary
In frequency division duplex MIMO systems, the original CSI feedback consumes a large amount of uplink resources. Existing AI/ML-based CSI feedback enhancement methods have failed to effectively solve the bridging problem of vector quantization-aware CSI feedback, especially the indistinguishability between quantizers and dequantizers for floating-point latent vectors and fixed-point codewords under a separate training framework.
A CSI feedback-based standalone training method based on vector quantization is proposed. The network device trains the model to determine the codebook and codeword index, and provides the training dataset and codebook to the terminal device. The terminal device performs CSI feedback compression based on the received data and codebook to realize the conversion of floating-point latent vectors to fixed-point codewords.
The CSI feedback system performance has been enhanced, resource consumption has been reduced, the accuracy and efficiency of CSI feedback have been improved, it is applicable to the scalability of different UEs, and the resource consumption of data transmission has been reduced.
Smart Images

Figure CN120958736A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure generally relate to the telecommunications field, and more specifically to apparatus, methods, devices, and computer-readable storage media for a separate training method for channel state information (CSI) feedback. Background Technology
[0002] Accurate CSI feedback is crucial for efficient transmission in Frequency Division Duplex (FDD) Multiple-Input Multiple-Output (MIMO) systems, but transmitting the raw CSI to the base station consumes significant uplink resources. With the tremendous success of AI-based feature extraction and image compression, CSI feedback enhancement using AI / machine learning (ML) methods has been discussed and studied. Summary of the Invention
[0003] Typically, the exemplary embodiments of this disclosure provide a solution for a separate training approach for CSI feedback.
[0004] In a first aspect, an apparatus is provided. The apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: receive from a network device a training dataset associated with CSI, a codebook of the network device, and a first set of codeword indices associated with the codebook; and determine CSI feedback compression of the apparatus based on a first codeword and a second codeword, the first and second codewords being generated based on the training dataset, the codebook of the network device, and the first set of codeword indices.
[0005] In a second aspect, an apparatus is provided, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: determine latent vectors based on a training dataset associated with CSI; determine a first set of codeword indices based on the latent vectors and the apparatus's codebook; and transmit the training dataset, the first set of codeword indices, and the codebook to a terminal device.
[0006] In a third aspect, a method is provided. The method includes: receiving, at a terminal device, a training dataset associated with CSI, a codebook of the network device, and a first set of codeword indices associated with the codebook from a network device; and determining CSI feedback compression of the terminal device based on the first codeword and a second codeword, the first codeword and the second codeword being generated based on the training dataset, the codebook of the network device, and the first set of codeword indices.
[0007] In a fourth aspect, a method is provided. This method includes: determining latent vectors at a network device based on a training dataset associated with CSI; determining a first set of codeword indices based on the latent vectors and the network device's codebook; and sending the training dataset, the first set of codeword indices, and the codebook to a terminal device.
[0008] In a fifth aspect, an apparatus is provided, comprising: components for receiving from a network device a training dataset associated with CSI, a codebook of the network device, and a first set of codeword indices associated with the codebook; and components for determining CSI feedback compression based on a first codeword and a second codeword, the first codeword and the second codeword being generated based on the training dataset, the codebook of the network device, and the first set of codeword indices.
[0009] In a sixth aspect, an apparatus is provided, comprising: components for determining latent vectors based on a training dataset associated with CSI; components for determining a first set of codeword indices based on the latent vectors and a codebook of the apparatus; and components for transmitting the training dataset, the first set of codeword indices, and the codebook to a terminal device.
[0010] In a seventh aspect, a computer-readable medium is provided having instructions stored thereon, which, when executed by at least one processor of a device, cause the device to perform at least the method according to a third or fourth aspect.
[0011] Other features and advantages of the embodiments of this disclosure will also become apparent when read in conjunction with the accompanying drawings, which illustrate the principles of the embodiments of this disclosure by way of example. Attached Figure Description
[0012] The embodiments disclosed herein are presented in an exemplary sense, and their advantages are explained in more detail below with reference to the accompanying drawings.
[0013] Figure 1 An example environment in which example embodiments of this disclosure may be implemented is shown; Figure 2 Examples of the overall structure for CSI compression and reconstruction according to some exemplary embodiments of this disclosure are shown; Figure 3 Signaling diagrams illustrating examples of processes according to some exemplary embodiments of this disclosure are shown; Figure 4 Examples of models for training on the network device side according to some exemplary embodiments of this disclosure are shown; Figure 5 Examples of models for training on the terminal device side according to some exemplary embodiments of the present disclosure are shown; Figure 6Examples of training models for simulating processes according to some exemplary embodiments of the present disclosure are shown; Figure 7 The convergence performance of a simulation process according to some example embodiments of this disclosure is illustrated; Figure 8 A flowchart is shown as an example method for a separate training approach for CSI feedback according to some example embodiments of this disclosure; Figure 9 A flowchart is shown as an example method for a separate training approach for CSI feedback according to some example embodiments of this disclosure; Figure 10 A simplified block diagram of a device suitable for implementing example embodiments of the present disclosure is shown; and Figure 11 A block diagram of an example computer-readable medium according to some embodiments of the present disclosure is shown.
[0014] In all the accompanying drawings, the same or similar reference numerals may denote the same or similar elements. Detailed Implementation
[0015] The principles of this disclosure will now be described with reference to some exemplary embodiments. It should be understood that these embodiments are described for illustrative purposes only and to assist those skilled in the art in understanding and implementing this disclosure, and do not impose any limitation on the scope of this disclosure. The embodiments described herein can be implemented in various ways other than those described below.
[0016] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein may have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0017] References to "an embodiment," "embodiment," "example embodiment," etc., in this disclosure indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment includes that particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Additionally, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is believed that the influence of such feature, structure, or characteristic on other embodiments is within the knowledge of those skilled in the art, whether explicitly described or not.
[0018] It should be understood that although the terms “first,” “second,” etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0019] As used herein, “at least one of the following: ” and “at least one of ” and similar wording, where a list of two or more elements is connected by “and” or “or”, indicates at least any one element, or at least any two or more elements, or at least all elements.
[0020] As used herein, unless explicitly stated otherwise, “responding to A” does not indicate that the step is performed immediately after “A” occurs, but may include one or more intermediate steps.
[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that, when used herein, the terms “comprising,” “including,” “having,” “having,” “containing,” and / or “comprising” specify the presence of the stated features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0022] As used in this application, the term "circuit" may refer to one or more or all of the following: (a) Hardware circuit implementation only (such as implementation with purely analog and / or digital circuits), and (b) A combination of hardware circuitry and software, such as (if applicable): (i) A combination of analog and / or digital hardware circuitry with software / firmware, and (ii) Any part of a hardware processor with software (including digital signal processors, software, and memory, which work together to enable a device such as a mobile phone or server to perform various functions); and (c) The operation requires software (e.g., firmware) for the operation of (multiple) hardware circuits and / or (multiple) processors, such as (multiple) microprocessors or parts thereof, but the software may be absent when the operation does not require the software.
[0023] This definition of "circuit" applies to all uses of the term in this application, including in any claim. As another example, as used herein, the term "circuit" also covers only hardware circuitry or a processor (or multiple processors), or a portion of hardware circuitry or a processor and its accompanying software and / or firmware implementations. For example, where applicable to a particular claim element, the term "circuit" also covers baseband integrated circuits or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices.
[0024] As used herein, the term "communication network" refers to a network that conforms to any suitable communication standard, such as New Radio (NR), Long Term Evolution (LTE), LTE-A Advanced (LTE-A), Wideband Code Division Multiple Access (WCDMA), High-Speed Packet Access (HSPA), Narrowband Internet of Things (NB-IoT), Enhanced Machine-Type Communication (eMTC), etc. Furthermore, communication between terminal devices and network devices in a communication network can be performed according to any suitable generation of communication protocol, including but not limited to first-generation (1G), second-generation (2G), 2.5G, 2.75G, third-generation (3G), fourth-generation (4G), 4.5G, fifth-generation (5G) communication protocols and / or any other currently known or future-developed protocols. Embodiments of this disclosure can be applied to a variety of communication systems. Given the rapid development of communications, there will certainly be future types of communication technologies and systems that embody the nature of this disclosure. The scope of this disclosure should not be considered limited to the systems described above.
[0025] As used herein, the terms “network equipment,” “radio network equipment,” and / or “radio access network equipment” refer to a node in a communications network through which terminal equipment accesses the network and receives services. Depending on the terminology and technology applied, network equipment can refer to a base station (BS) or access point (AP), such as a Node B (NodeB or NB), an evolved Node B (eNodeB or eNB), an NR NB (also known as a gNB), a Remote Radio Unit (RRU), a Radio Head (RH), a Remote Radio Head (RRH), a relay, an Integrated Access and Backhaul (IAB) node, a low-power node (such as a femtosecond or picosecond), a non-terrestrial network (NTN) or non-terrestrial network equipment (such as satellite network equipment, low Earth orbit (LEO) satellites, and geostationary orbit (GEO) satellites), spacecraft network equipment, etc. In some example embodiments, a LEO (RAN) split architecture includes a centralized unit (CU) and a distributed unit (DU). In some other example embodiments, a portion or all of the radio access network equipment may be mounted on an airborne or space-based NTN vehicle.
[0026] The term "terminal device" refers to any terminal device capable of wireless communication. By way of example and not limitation, a terminal device may also be referred to as a communication device, user equipment (UE), subscriber station (SS), portable subscriber station, mobile station (MS), or access terminal (AT). Terminal devices may include, but are not limited to: mobile phones, cellular phones, smartphones, Voice over IP (VoIP) phones, wireless local loop phones, tablets, wearable terminal devices, personal digital assistants (PDAs), portable computers, desktop computers, image capture terminal devices (such as digital cameras), gaming terminal devices, music storage and playback devices, in-vehicle wireless terminal devices, wireless endpoints, mobile stations, laptop embedded devices (LEEs), laptop-mounted devices (LMEs), USB dongles, smart devices, wireless customer premises equipment (CPEs), Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronic devices, devices operating on commercial and / or industrial wireless networks, etc. The terminal device may also correspond to the mobile terminal (MT) portion of an IAB node (e.g., a relay node). In the following description, the terms "terminal device," "communication device," "terminal," "user equipment," and "UE" are used interchangeably.
[0027] As used herein, the terms “resource,” “transmission resource,” “resource block,” “physical resource block” (PRB), “uplink resource,” or “downlink resource” can refer to any resource used to perform communication, such as communication between terminal devices and network devices, including resources in the time domain, frequency domain, spatial domain, code domain, or any other resources used to implement communication. In the following, unless explicitly stated otherwise, resources in both the frequency and time domains will be used as examples of transmission resources used to describe some exemplary embodiments of this disclosure. Note that the exemplary embodiments of this disclosure are equally applicable to other resources in other domains.
[0028] Figure 1 An example communication network 100 in which embodiments of the present disclosure may be implemented is shown. For example... Figure 1 As shown, the communication network 100 may include a terminal device 110. In the following text, the terminal device 110 may also be referred to as a UE.
[0029] The communication network 100 may also include a network device 120. Hereinafter, the network device 120 may also be referred to as a gNB or eNB. The terminal device 110 can communicate with the network device 120 within the coverage area of the cell 102 managed by the network device 120.
[0030] It should be understood that Figure 1 The number of network devices and terminal devices shown is given for illustrative purposes and does not imply any limitation. Communication network 100 may include any suitable number of network devices and terminal devices.
[0031] In some example embodiments, the link from network device 120 to terminal device 110 may be referred to as a downlink (DL), while the link from terminal device 110 to network device 120 may be referred to as an uplink (UL). In the DL, network device 120 is a transmitting (TX) device (or transmitter), and terminal device 110 is a receiving (RX) device (or receiver). In the UL, terminal device 110 is a TX device (or transmitter), and network device 120 is an RX device (or receiver).
[0032] Communication in communication environment 100 can be implemented according to any suitable communication protocol, including but not limited to cellular communication protocols such as first-generation (1G), second-generation (2G), third-generation (3G), fourth-generation (4G), fifth-generation (5G), and sixth-generation (6G), wireless local network communication protocols such as IEEE 802.11, and / or any other currently known or future-developed protocols. Furthermore, communication can utilize any suitable wireless communication technology, including but not limited to: Code Division Multiple Access (CDMA), Frequency Division Multiple Access (FDMA), Time Division Multiple Access (TDMA), Frequency Division Duplex (FDD), Time Division Duplex (TDD), Multiple-Input Multiple-Output (MIMO), Orthogonal Frequency Division Multiple Access (OFDM), Discrete Fourier Transform Extended OFDM (DFT-s-OFDM), and / or any other currently known or future-developed technologies.
[0033] As mentioned above, CSI feedback enhancement through the use of AI / ML methods has been discussed and studied. To this end, the 3rd Generation Partnership Project (3GPP) approved a new research project (SI) in Release 18 to explore the benefits of enhancing CSI feedback through AI / ML technologies.
[0034] AI / ML research project use cases can focus on CSI feedback enhancements, such as overhead reduction, improved accuracy, and prediction; beam management, such as beam prediction in the time and / or spatial domains for overhead and latency reduction, and improved beam selection accuracy; and positioning accuracy enhancements for different scenarios, including scenarios with heavy non-line-of-sight (NLOS) conditions.
[0035] Within the AI-based CSI feedback framework, quantizer optimization is categorized into two types in 3GPP Release 18: quantization-unaware training and quantization-aware training. In the quantization-unaware training framework, the quantizer formula is not involved in the training process, which can introduce end-to-end performance degradation due to misalignment between the encoder / decoder and the quantizer. Conversely, quantization-aware training integrates the quantizer formula into the encoder-decoder training to achieve superior overall key performance indicators (KPIs) with good alignment across all modules.
[0036] It was agreed that bilateral model use cases can be used for CSI compression, and that quantization methods, including quantization non-perceptual training, quantization perceptual training, quantization methods including uniform and non-uniform quantization, scalar and vector quantization, and associated parameters (e.g., quantization resolution, etc.), as well as how to use quantization methods to evaluate and study the quantization of CSI feedback, can be applied.
[0037] Regarding training frameworks, 3GPP has defined three types: Type 1, Type 2, and Type 3. Type 1 and Type 2 training rely on end-to-end gradient propagation from the decoder back to the encoder to update both the encoder and decoder simultaneously. Type 3 training, on the other hand, assumes that UE-side training for the encoder and NW-side training for the decoder are performed in separate training sessions. Therefore, NW-side decoder training will not require detailed model parameters of the UE encoder for its own training, and vice versa.
[0038] However, a major challenge in AI-based quantization-aware CSI feedback lies in the indistinguishability of quantizers and dequantizers bridging floating-point latent vectors and fixed-point codewords. There are currently no studies or reports on vector quantization-aware CSI feedback within a separately trained framework.
[0039] In contrast to quantization-aware CSI feedback with scalar / vector quantization, which only has two learnable functions (i.e., encoder and decoder), quantization-aware CSI feedback with vector quantization is typically characterized by the joint learning of three functions (i.e., encoder, codebook, and decoder) simultaneously at a single entity; this is the Type 1 training category. When extending to Type 3 with separate training, the current experience / methods of Type 1 cannot be directly applied to the Type-3 extension because of constraints. The codebook used for vector quantization is pre-determined by the network (NW) in an NW-first training manner, so the encoder on the UE side needs to be learned independently while conforming to the predefined codebook. Therefore, a detailed approach, including the training process, required dataset, and feasible loss functions, may still be necessary.
[0040] Therefore, this disclosure proposes a mechanism for a separate training method for CSI feedback based on vector quantization. In this solution, the network device can determine its codebook and a first set of codeword indices by training a model based on a training dataset associated with CSI, and provides the training dataset, the network device's codebook, and the first set of codeword indices to the terminal device. The terminal device can then train CSI feedback compression based on the received training dataset, the network device's codebook, and the first set of codeword indices.
[0041] In this way, a mechanism for a separate training method for CSI feedback based on vector quantization is realized, and system performance enhancements can be achieved for CSI feedback.
[0042] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0043] Now for reference Figure 2 This illustrates an example of an overall structure for CSI compression and reconstruction according to some exemplary embodiments of this disclosure. The terminal device 110 side entity may include an encoder 112 and a quantizer 113, wherein the encoder 112 compresses the raw CSI data 111 into a latent feature vector, and the quantizer 113 converts the floating-point latent vector into a fixed-point codeword 114 for feedback. Conversely, the network device 120 side entity may include an inverse quantizer 121 and a decoder 122, wherein the inverse quantizer 121 transforms the fixed-point codeword 114 into a latent vector, and the decoder 122 reconstructs the raw CSI data 123.
[0044] Now for reference Figure 3 This illustrates signaling diagram 300 for communication according to some example embodiments of the present disclosure. For example... Figure 3 As shown, signaling diagram 300 involves terminal device 110 and network device 120. For discussion purposes, refer to... Figure 1 To describe signaling diagram 300.
[0045] During the training phase, network device 120 may collect (302) a training dataset (e.g., multiple stored historical CSI data collected from terminal device 110) and perform (304) a training process based on an AI / ML model to reconstruct the CSI feedback data. In this phase, the codebook of the network device, the decoder at the network device, and the hypothesis encoder may be trained together based on the training dataset associated with the CSI.
[0046] For network device-side training, vector quantization-aware training can be achieved using a vector quantization variational autoencoder (VQ-VAE) scheme, which includes the following modules: The encoder can compress the CSI data into latent feature vectors. ξ A segmenter can divide a long latent vector. ξ Divided into multiple segments ξ k Vector quantizers can quantize the latent vectors of each segment. ξ k In the case of the identifier codebook C The codeword index z is generated using the nearest codebook vector in the reference codebook; the vector dequantizer can generate each segment codeword using the reference codebook and the codeword index z. v k The combiner can combine all code fields. v k Cascaded into complete codewords v And the decoder can convert the codewords v Used as input to reconstruct CSI data .
[0047] Figure 4 Examples of models for training on the network device side are shown according to some exemplary embodiments of this disclosure.
[0048] As shown in the figure, the model 400 used for training on the network device side may include a hypothesis encoder 401, a segmenter 402, a quantizer 403, and a codebook 404 of the network device 120 (specified as...). C ), dequantizer 405, combiner 406 and decoder 407.
[0049] In some embodiments, the training process of network device 120 can be based on vector quantization-aware training.
[0050] During training, a training dataset (specified as 𝜒) associated with CSI can be provided to encoder 401. For example, the training dataset associated with CSI can be all or a subset of a pre-stored CSI dataset.
[0051] Encoder 401 takes the training dataset 𝜒 as input and outputs a compressed latent feature vector. ξ The most common deep network architectures (including fully connected (FC) layers, convolutional layers, transformer networks, and long short-term memory (LSTM) networks) can be used as encoder networks.
[0052] If the dimension of the latent vectors is too large for codebook learning, then segmenter 402 will... ξ Divided into sizes of Each segment ξ k All these segments are fed in parallel or sequentially into quantizer 403 to produce codeword index z.
[0053] Assume there exists a system of dimensions. D of BThe codebook 404, composed of codebook vectors, is specified as... C The vector quantizer 403 first measures each codebook vector. v i With the given paragraph ξ k The distance between them. Then, output the index element with the minimum distance. This operation can be written as:
[0054] in, This represents a distance metric, which can be Euclidean distance, cosine similarity, etc. Corresponding to W All index elements of each segment constitute the codeword index used for feedback. .
[0055] Vector dequantizer 405 maintains the same codebook 404 as quantizer 403 (specified as...). C Therefore, it is possible to refer to the codeword index z given the codeword index z. C To generate segmented codewords .
[0056] Combiner 406 cascades all segmented codewords To form a complete codeword v .
[0057] Decoder 407 will v Used as input to reconstruct the training dataset The decoder 407 can employ various deep network architectures, such as fully connected (FC) layers, convolutional layers, transformer networks, and long short-term memory (LSTM) networks.
[0058] Based on whether the codebook 404 is updated along with the autoencoder during training, as described above, there are two options for the loss function used for training.
[0059] For a learnable codebook, in order to train the encoder, codebook, and decoder in an end-to-end manner, the loss function can be written as:
[0060] in, Represents the true CSI χ With decoder output χ The reconstruction loss is calculated between the reconstructed CSIs at each location. The reconstruction loss can be measured using either the normalized mean square error (NMSE) or cosine similarity. As mentioned above, the piecewise latent vectors... ξ k Mapped to the nearest codebook vector This serves as the input to the decoder. Therefore, the second term, known as the quantization loss, optimizes the codebook to be as close as possible to the encoder output. In this case, the stopping gradient operator is used. The encoder output is treated as a constant. Specifically, the codebook can be updated using a dictionary learning algorithm with L2 norm error. The last term is the commitment loss, which further optimizes the encoder so that the output... ξ k Dedicated to codebooks.
[0061] For a fixed codebook, since there is no need to update the codebook, the quantization loss of the learnable codebook as described above can be eliminated, while retaining the other two losses. Therefore, the loss function can be written as follows:
[0062] In some other embodiments, the training process of network device 120 can be based on vector quantization-aware training. For example... Figure 4 The training model shown can also be used for vector quantization-aware training. The difference between vector quantization-aware training and vector quantization-unaware training is that the training of the autoencoder and the optimization of the vector quantizer are independent of each other.
[0063] During this training phase, the latent vector ξ The vectors are directly fed into decoder 407 for CSI reconstruction, thus quantization error is disregarded. After training the autoencoder, vector quantizer 403 is optimized by referencing the characteristics of these latent vectors, which involves the following steps.
[0064] First, sufficient latent vectors generated by the encoder can be collected as a dataset V. Then, clustering algorithms such as k-means or random initialization can be used to initialize the dataset with... D Each dimension and B The codebook consists of 404 vectors. Each latent vector in V can then be divided into W segments, and each segment can be assigned to its nearest codebook vector. The codebook can be updated by minimizing the difference between the original latent vector and its quantized representation.
[0065] In CSI feedback scenarios, differences can be measured using Euclidean distance or cosine similarity. Various strategies exist for updating the codebook, including K-means clustering and neural network-based methods.
[0066] After training the NW-side model, an appropriate shared dataset, along with the associated codebook, needs to be shared with the UE side for encoder training, without needing to know the NW-side model. (Return to reference) Figure 3 Network device 120 can provide terminal device 110 with (306) training dataset (specified as χ) and codebook. The codeword index z(χ) associated with the training dataset χ is also referred to below as the first set of codeword indices. In the following text, the training dataset (specified as χ) and the codeword index z(χ) associated with the training dataset χ can also be referred to as the shared dataset.
[0067] There are two benefits to providing such parameters to terminal device 110. First, it utilizes a common codebook shared by the NW-side entity and the UE-side entity. C First, the codeword spaces derived from different UEs can be well aligned, which helps the common NW-side decoder to reconstruct the code. Second, a universal codebook is utilized. C Fixed-point codeword index z is higher than floating-point high-dimensional codeword. v It uses far fewer resources.
[0068] Then, the terminal device 110 can, based on the received training dataset (specified as χ) and the codebook associated with the training dataset χ, C The codeword index z(χ) is used to perform the (308) training process on the terminal device 110.
[0069] For training on the terminal device side, the inverse quantizer can reproduce the segmented codewords using the reference codeword index z and codebook C. The combiner concatenates all segmented codewords into a complete codeword. v The encoder can compress the training dataset x into a format as close as possible to the reference codewords. v The code .
[0070] Figure 5 Examples of models for training on the terminal device side are shown according to some exemplary embodiments of this disclosure.
[0071] As shown in the figure, the model 500 used for training on the terminal device side may include an inverse quantizer 501, a combiner 502, and an encoder 503.
[0072] The inverse quantizer 501 in the terminal device-side entity takes the codeword index z as input. Each element in z... z k Corresponding to the codebook vector, it is a segmented codeword. The dequantizer 501 can be accessed via the reference codebook. C Sequentially or in parallel z k Mapped to Then, combiner 502 concatenates all these segments into a complete codeword. v This serves as a reference for encoder training. In short, the dequantizer 501 and combiner 502 only implement lookup and cascading functions, therefore requiring no additional information.
[0073] Encoder 503 takes the training dataset χ as input and outputs its associated codeword vector. The most common deep network architectures (including fully connected (FC) layers, convolutional layers, transformer networks, and long short-term memory (LSTM) networks) can be used as encoder networks.
[0074] The encoder is trained to minimize the encoder output in a supervised manner. With reference code v Differences in encoder output Fit to reference codeword v The loss function can be any metric used to evaluate the distance between two vectors, such as mean squared error (MSE) or generalized cosine similarity (GCS). In our simulation, we use MSE loss.
[0075] The above embodiments explain how to train the network device-side module and the terminal device-side module separately. In summary, the NW-side entity uses the VQ-VAE training structure and a pre-stored set of CSI datasets to train the hypothesis encoder, decoder, and codebook. After training, the codebook... C All NN parameters of the encoder and decoder are ultimately determined in the NW side entity.
[0076] Then, the NW-side entities generate a shared dataset and the final codebook. C It is shared with the UE. The shared dataset consists of the original training dataset X and the corresponding codeword index Z(X).
[0077] Given a pre-trained encoder and codebook, the training dataset χ is first compressed into latent vectors. ξ ,Then ξ Divided into W Each segment Given by B vectors The codebook C , W Each segment is mapped to its nearest codebook vector. , where each vector belongs to ,like Figure 4 As shown. Simultaneously, the export instruction indicates its location in the codebook. C Index vector of positions within Within the VQ-VAE framework, by v Instead of flattening represented by ξ It is fed into the decoder. For brevity, given a training dataset χ, the quantized latent vector v(χ) can be called a codeword, and the index vector z(χ) can be called a codeword index.
[0078] The NW-side entity does not share the codewords themselves with the UE, but instead uses the final codebook. C The shared dataset (X, Z(X)) is shared with the UE side for encoder training. (X, Z(X)) is a large-scale dataset. A set of pairs.
[0079] This design has two main advantages. First, it utilizes a common codebook shared by the NW-side entity and the UE-side entity. C First, the codeword spaces derived from different UEs can be well aligned, which facilitates the reconstruction of the common NW-side decoder. Second, a universal codebook is utilized. C Fixed-point codeword index z is higher than floating-point high-dimensional codeword. v It uses far fewer resources.
[0080] Then, the UE-side entity first refers to the given codeword index. codebook C To reproduce the typing v The encoder is then trained to minimize in a supervised manner. With reference code v The difference between them is used to compress the training dataset χ into codewords. Here, the UE-side encoder can employ a completely different NN structure than the NW-side encoder. The loss function can be any suitable metric for evaluating the distance between two vectors, such as mean squared error (MSE) or generalized cosine similarity (GCS).
[0081] Based on well-trained model 500 on the terminal device side, such as Figure 3 As shown, during the implementation phase (which can be called the inference phase), the NW-side entity and the UE can share the same codebook. C Terminal device 110 can obtain actual CSI data and generate a second set of codeword indexes based on the CSI data and a well-trained model (e.g., model 500) at terminal device 110, and provide the second set of codeword indexes to network device 120 (310).
[0082] It should be understood that network device 120 can provide training datasets (specified as χ) and codebooks to multiple terminal devices. C And the codeword index z(χ) associated with the training dataset χ. The multi-terminal device 110 can train itself based on the received training parameters.
[0083] Based on the received second set of codeword indexes and a well-trained model (e.g., model 400) at network device 120, network device 120 can reconstruct the CSI data obtained at the terminal device.
[0084] Simulations based on the solution disclosed herein can be discussed as follows. The configuration for generating a shared dataset is given below.
[0085] In this simulation, 100,000 feature vector samples are used as CSI data for both training and validation. Of these, 80,000 samples are used for training and 20,000 samples are used for validation. Each sample consists of N = 728 real numbers, corresponding to a large feature vector composed of 12 concatenated subbands, as shown below:
[0086] in It is the first k The feature vector of the sub-band channel.
[0087] Each x k It has been processed into the following format:
[0088] in It consists of the real part and the imaginary part.
[0089] In this simulation, SGCS is used as a performance metric between the original CSI data and the reconstructed CSI data.
[0090] First, the superiority of the proposed NW-priority vector quantization-aware single-training scheme is demonstrated. The detailed model structure of the proposed scheme is as follows: Figure 6 As shown, it may include an encoder 601 and a decoder 602. To achieve fair comparison, the number of overhead bits after quantization is set to 36 for all comparison schemes. For vector quantization-based schemes, such as... Figure 6 As shown, the number of codebook vectors B and the number of segments W in codebook 603 are 64 and 6 respectively, therefore the number of overhead bits is... For the scalar quantization-aware scheme, the output (i.e., the latent vector) has a dimension of 12, and each latent vector element is quantized to 8 levels (i.e., 3 bits), so the number of overhead bits is also... .
[0091] Figure 7 The simulation results shown present the convergence performance of different schemes in SGCS. Two observations can be made from the simulation results: (1) the NW-first VQ-aware standalone training scheme (curve 701) can achieve approximate performance with the type 1 joint training scheme (less than 0.8% degradation); (2) the vector quantization-based training scheme (curves 702 and 703) can achieve a significant performance gain (more than 8% improvement) over the scalar quantization-based scheme.
[0092] Simulation results demonstrate the advantages of the proposed NW-priority vector quantization-based aware training method, including: the proposed method for CSI feedback enhancement can achieve performance comparable to the Type 1 joint training method; the proposed method significantly outperforms the scalar quantization-based scheme; and there is no performance degradation even when the UE-side NN structure is not aligned with the NW-side NN structure.
[0093] The solution proposed in this disclosure presents a mechanism for separate training of CSI feedback. This disclosure considers a separate training framework for vector quantization-aware CSI feedback enhancement, which not only leverages the performance advantages of vector quantization-aware training but also avoids the exposure of proprietary NN models in practical deployments; that is, it implements quantization-aware CSI feedback with vector quantization within a (type-3) separate training framework.
[0094] Furthermore, by utilizing this universal codebook learned and provided by the NW, different UEs may be well-constrained when learning their individual encoders in a common quantization feature space, which contributes to the good reconstruction of the common NW-side decoder. The codebook and decoder learned in the NW are universal, and this approach is scalable for new UEs, meaning that the NW-side decoder can adapt to new UEs without retraining.
[0095] Finally, and very importantly, when using a universal codebook, the codeword index z(χ) in the shared dataset is an integer index vector, which consumes fewer resources than the codeword index. v The number of datasets is much smaller, thus improving the efficiency of dataset distribution.
[0096] Figure 8 A flowchart of an example method 800 for a separate training approach for CSI feedback, according to some example embodiments of this disclosure, is shown. Method 800 can be implemented as follows: Figure 1 The terminal device 110 shown is implemented here. For discussion purposes, reference will be made to... Figure 1 Description method 800.
[0097] At 810, terminal device 110 receives from network device a training dataset associated with CSI, network device codebook, and a first set of codeword indices associated with codebook.
[0098] At 820, terminal device 110 determines CSI feedback compression based on a first codeword and a second codeword, the first codeword and the second codeword being generated based on a training dataset, the network device's codebook, and a first set of codeword indices.
[0099] In some example embodiments, the terminal device can generate a first codeword by selecting a vector corresponding to a codeword index group from a codebook; generate a second codeword by compressing a training dataset at the encoder of the terminal device used to perform CSI feedback compression; and train the encoder at the terminal device by using a loss function to minimize the difference between the second codeword and the first codeword.
[0100] In some example implementations, the loss function is based on mean squared error or generalized cosine similarity.
[0101] In some example embodiments, the terminal device may obtain CSI data; generate a second set of codeword indexes associated with the codebook from the CSI data based on the determined CSI feedback compression; and send the second set of codeword indexes to the network device.
[0102] Figure 9 A flowchart of an example method 900 for a separate training approach for CSI feedback, according to some example embodiments of the present disclosure, is shown. Method 900 can be implemented as follows: Figure 1 The network devices shown are implemented in 120 locations. For discussion purposes, references will be made to... Figure 1 Description method 900.
[0103] At 910, network device 120 determines potential vectors based on the training dataset associated with CSI.
[0104] At position 920, network device 120 determines the first set of codeword indices based on the potential vector and the network device's codebook.
[0105] At position 930, network device 120 sends the training dataset, the first set of codeword indexes, and the codebook to the terminal device.
[0106] In some example embodiments, network device 120 may generate a set of segmented latent vectors from latent vectors; determine a plurality of segmented codewords from a codebook, each segmented codeword corresponding to a segmented latent vector; and determine a first set of codeword indices based on the plurality of segmented codewords.
[0107] In some example embodiments, network device 120 may receive a second set of codeword indexes associated with a codebook from a terminal device, wherein the second set of codeword indexes is generated by performing CSI feedback compression on CSI data at the terminal device; and reconstruct the CSI data associated with the CSI based on the received second set of codeword indexes and the codebook of the network device.
[0108] In some example embodiments, network device 120 may generate codewords based on the received second set of codeword indexes and the network device's codebook; and reconstruct CSI data associated with CSI by decoding the generated codewords.
[0109] In some example embodiments, the means capable of performing method 800 (e.g., implemented at terminal device 110) may include components for performing the corresponding steps of method 800. These components may be implemented in any suitable form. For example, the components may be implemented in a circuit or software module.
[0110] In some example embodiments, the apparatus includes: components for receiving from a network device a training dataset associated with CSI, a codebook of the network device, and a first set of codeword indices associated with the codebook; and components for CSI feedback compression based on a first codeword and a second codeword, the first codeword and the second codeword being generated based on the training dataset, the codebook of the network device, and the first set of codeword indices.
[0111] In some example embodiments, the apparatus includes: a component for generating a first codeword by selecting a vector corresponding to a codeword index set from a codebook; a component for generating a second codeword by compressing a training dataset at an encoder of a device for performing CSI feedback compression; and a component for minimizing the difference between the second codeword and the first codeword using a loss function at an encoder of a training device.
[0112] In some example implementations, the loss function is based on mean squared error or generalized cosine similarity.
[0113] In some example embodiments, the apparatus includes: components for acquiring CSI data; components for generating a second set of codeword indices associated with a codebook from the CSI data based on determined CSI feedback compression; and components for sending the second set of codeword indices to a network device.
[0114] In some example embodiments, the means capable of performing method 900 (e.g., implemented at network device 120) may include components for performing corresponding steps of method 900. These components may be implemented in any suitable form. For example, the components may be implemented in a circuit or software module.
[0115] In some example embodiments, the apparatus includes: components for determining latent vectors based on a training dataset associated with CSI; components for determining a first set of codeword indices based on the latent vectors and the codebook of the apparatus; and components for transmitting the training dataset, the first set of codeword indices, and the codebook to a terminal device.
[0116] In some example embodiments, the apparatus includes: components for generating a set of segmented latent vectors from latent vectors; components for determining a plurality of segmented codewords from a codebook, each segmented codeword corresponding to a segmented latent vector; and components for determining a first set of codeword indices based on the plurality of segmented codewords.
[0117] In some example embodiments, the apparatus includes: components for receiving a second set of codeword indices associated with a codebook from a terminal device, wherein the second set of codeword indices is generated by performing CSI feedback compression on CSI data at the terminal device; and components for reconstructing CSI data associated with CSI based on the received second set of codeword indices and the codebook of the apparatus.
[0118] In some example embodiments, the apparatus includes: components for generating codewords based on a received second set of codeword indexes and the apparatus's codebook; and components for reconstructing CSI data associated with the CSI by decoding the generated codewords.
[0119] Figure 10 This is a simplified block diagram of a device 1000 suitable for implementing exemplary embodiments of the present disclosure. The device 1000 can be provided to implement a communication device, for example, as... Figure 1 The terminal device 110 or network device 120 shown. As shown, device 1000 includes one or more processors 1010, one or more memories 1020 coupled to processor 1010, and one or more communication modules 1040 coupled to processor 1010.
[0120] Communication module 1040 is used for bidirectional communication. Communication module 1040 has one or more communication interfaces to facilitate communication with one or more other modules or devices. The communication interface can represent any interface necessary for communication with other network elements. In some example embodiments, communication module 1040 may include at least one antenna.
[0121] As a non-limiting example, processor 1010 can be any type suitable for a local technology network and can include one or more of the following: general-purpose computer, special-purpose computer, microprocessor, digital signal processor (DSP), and processor based on a multi-core processor architecture. Device 1000 can have multiple processors, such as application-specific integrated circuit chips that are time-dependent on a clock synchronized with the main processor.
[0122] Memory 1020 may include one or more non-volatile memories and one or more volatile memories. Examples of non-volatile memories include, but are not limited to, read-only memory (ROM) 1024, electrically programmable read-only memory (EPROM), flash memory, hard disk, optical disc (CD), digital video disc (DVD), optical disc, laser disc, and other magnetic and / or optical storage. Examples of volatile memories include, but are not limited to, random access memory (RAM) 1022 and other volatile memories that will not persist during power outages.
[0123] Computer program 1030 includes computer-executable instructions that are executed by an associated processor 1010. The instructions of program 1030 may include instructions for performing operations / actions of some example embodiments of this disclosure. Program 1030 may be stored in memory (e.g., ROM 1024). Processor 1010 can perform any suitable actions and processes by loading program 1030 into RAM 1022.
[0124] The exemplary embodiments of this disclosure can be implemented by program 1030, enabling device 1000 to execute as described in the reference. Figures 2 to 9 Any process discussed in this disclosure. Exemplary embodiments of this disclosure may also be implemented by hardware or by a combination of software and hardware.
[0125] In some example embodiments, program 1030 may be tangibly contained in a computer-readable medium, which may be included in device 1000 (such as in memory 1020) or in other storage devices accessible by device 1000. Device 1000 may load program 1030 from the computer-readable medium into RAM 1022 for execution. In some example embodiments, the computer-readable medium may include any type of non-transitory storage medium, such as ROM, EPROM, flash memory, hard disk, CD, DVD, etc. The term "non-transitory" as used herein refers to a limitation on the medium itself (i.e., tangible, not tactile), rather than a limitation on the persistence of data storage (e.g., RAM vs. ROM).
[0126] Figure 11 An example of a computer-readable medium 1100 is shown, which may be in the form of a CD, DVD, or other optical storage disc. The computer-readable medium 1100 has a program 1030 stored thereon.
[0127] Generally, the various embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. Although various aspects of the embodiments of this disclosure are shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as non-limiting examples, the blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0128] Some exemplary embodiments of this disclosure also provide at least one computer program product tangibly stored on a computer-readable medium, such as a non-transitory computer-readable medium. The computer program product includes computer-executable instructions, such as those included in a program module that execute in a device on a target physical or virtual processor to perform any of the methods described above. Typically, a program module includes routines, programs, libraries, objects, classes, components, data structures, etc., that perform a particular task or implement a particular abstract data type. In various embodiments, the functionality of a program module can be combined or split among program modules as needed. The machine-executable instructions for a program module can execute within a local or distributed device. In a distributed device, the program module can reside in both local and remote storage media.
[0129] Program code used to perform the methods of this disclosure may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a stand-alone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0130] In the context of this disclosure, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc.
[0131] Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable media can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination thereof. More specific examples of computer-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0132] Furthermore, although operations are described in a specific order, this should not be construed as requiring such operations to be performed in the specific order shown or sequentially, or to perform all shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure, but rather as descriptions of features that may be specific to particular embodiments. Unless explicitly stated otherwise, certain features described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, unless explicitly stated otherwise, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0133] Although this disclosure has been described in language specific to structural features and / or methodological actions, it should be understood that the disclosure as defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are disclosed as exemplary forms for implementing the claims.
Claims
1. An apparatus comprising: At least one processor; as well as At least one memory storing instructions that, when executed by the at least one processor, cause the device to at least: Receive from the network device a training dataset associated with Channel State Information (CSI), the codebook of the network device, and a first set of codeword indices associated with the codebook; as well as The CSI feedback compression of the device is determined based on the first codeword and the second codeword, which are generated based on the training dataset, the codebook of the network device, and the first set of codeword indices.
2. The apparatus of claim 1, wherein the apparatus is further caused to: The first codeword is generated by selecting a vector corresponding to the codeword index group from the codebook; The second codeword is generated by compressing the training dataset at the encoder of the device used to perform the CSI feedback compression; and The encoder at the device is trained by using a loss function to minimize the difference between the second codeword and the first codeword.
3. The apparatus of claim 2, wherein the loss function is based on mean square error or generalized cosine similarity.
4. The apparatus of claim 1, wherein the apparatus further causes: Obtain CSI data; Based on the determined CSI feedback compression, a second set of codeword indices associated with the codebook is generated from the CSI data; and Send the second set of codeword indexes to the network device.
5. An apparatus comprising: At least one processor; as well as At least one memory storing instructions that, when executed by the at least one processor, cause the device to at least: Potential vectors are determined based on a training dataset associated with Channel State Information (CSI). The first set of codeword indices is determined based on the potential vector and the codebook of the device; as well as The training dataset, the first set of codeword indexes, and the codebook are sent to the terminal device.
6. The apparatus of claim 5, wherein the apparatus is further caused to: Generate a set of piecewise potential vectors from the potential vectors; Determine multiple segmented codewords, each corresponding to a segmented potential vector, from the codebook; and The index of the first group of codewords is determined based on the multiple segmented codewords.
7. The apparatus of claim 5, wherein the apparatus further causes: Receive a second set of codeword indexes associated with the codebook from the terminal device, wherein the second set of codeword indexes is generated by performing CSI feedback compression on the CSI data at the terminal device; and Based on the received second set of codeword indexes and the codebook of the device, the CSI data associated with the CSI is reconstructed.
8. The apparatus of claim 7, wherein the apparatus further causes: Generate codewords based on the received second set of codeword indexes and the codebook of the device; and The CSI data associated with the CSI is reconstructed by decoding the generated codewords.
9. A method comprising: At the terminal device, a training dataset associated with Channel State Information (CSI), the codebook of the network device, and a first set of codeword indices associated with the codebook are received from the network device. as well as The CSI feedback compression of the terminal device is determined based on the first codeword and the second codeword, wherein the first codeword and the second codeword are generated based on the training dataset, the codebook of the network device, and the first set of codeword indexes.
10. A method comprising: Potential vectors are determined at network devices based on a training dataset associated with Channel State Information (CSI). The first set of codeword indices is determined based on the potential vector and the codebook of the network device; as well as The training dataset, the first set of codeword indexes, and the codebook are sent to the terminal device.
11. An apparatus comprising: A component for receiving from a network device a training dataset associated with Channel State Information (CSI), the codebook of the network device, and a first set of codeword indices associated with the codebook; as well as A component for determining CSI feedback compression of the device based on a first codeword and a second codeword, the first codeword and the second codeword being generated based on the training dataset, the codebook of the network device, and the first set of codeword indices.
12. An apparatus comprising: A component used to determine potential vectors based on a training dataset associated with Channel State Information (CSI); A component for determining a first set of codeword indices based on the latent vector and the codebook of the device; as well as A component for sending the training dataset, the first set of codeword indexes, and the codebook to a terminal device.
13. A computer-readable medium comprising instructions that, when executed by a device, cause the device to perform at least the method of claim 9 or the method of claim 10.