Methods and systems for CSI compression using machine learning with regularization
Physics-informed regularizers in two-sided machine learning models address CSI feedback challenges by improving accuracy and interoperability, reducing complexity for efficient CSI compression in wireless communication systems.
Patent Information
- Application Number
- PCT/CN2024/108155
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-10
- Filing Date
- 2024-07-29
- Publication Date
- 2025-11-13
AI Technical Summary
Existing wireless communication systems face challenges in optimizing CSI feedback payload reduction while ensuring accurate reconstruction at the base station and maintaining interoperability between UE and network vendors, particularly due to varying model architectures and resource constraints.
Implementing physics-informed regularizers in two-sided machine learning models for CSI compression, using properties like sparsity in angular and delay domains, to enhance accuracy, interoperability, and reduce complexity.
Improves CSI reconstruction accuracy, reduces model complexity, and enhances interoperability, making it suitable for real-time applications on resource-constrained devices.
Smart Images

Figure CN2024108155_13112025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR CSI COMPRESSION USING MACHINE LEARNING WITH REGULARIZATION
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. provisional patent application no. 63 / 645,439, entitled “METHOD AND APPARATUS FOR INTEROPERABLE CHANNEL STATE INFORMATION FEEDBACK USING PHYSICS-AWARE MACHINE LEARNING MODELS” , filed May 10, 2024, the entirety of which is hereby incorporated by reference.TECHNICAL FIELD
[0003] The present disclosure relates to wireless communications, in particular methods and systems for compression of channel station information (CSI) .BACKGROUND
[0004] Channel State Information (CSI) feedback is one of the mechanisms in modern wireless communication systems that help to provide valuable insights into the propagation environment between the transmitter and receiver. This information is useful for optimizing various aspects of the communication link, such as beamforming, precoding, and link adaptation, and may ultimately lead to improved spectral efficiency and overall system performance.
[0005] There is an interest in reducing the CSI feedback payload, in order to reduce the overhead for CSI feedback. At the same time, there remains a challenge in how to optimize CSI compression at the user equipment (UE) in a way that enables the compressed CSI to be successfully reconstructed at the base station (BS) so that the BS can use the CSI for subsequent tasks.SUMMARY
[0006] In various examples, the present disclosure describes methods, apparatuses, network entities and computer-readable media that enable two-sided machine learning-based compression of channel state information (CSI) using encoder and decoder models, where a physics-based regularizer is used during training. A physics-based regularizer refers to a regularizer that is defined based on a physical property of the wireless communication system, such as an expected sparsity of a precoding matrix in an angular domain, delay domain and / or time domain.
[0007] In various examples, the present disclosure describes methods, apparatuses, network entities and computer-readable media that include an indicator (which represent the physics-based regularizer or may be the physics-based regularizer itself) in the CSI feedback. The indicator included in the CSI feedback may be used to detect any inconsistencies in the reconstructed CSI, for example. This may be useful to ensure interoperability of the encoder and decoder models, and may provide a mechanism for error detection.
[0008] Examples of the present disclosure may provide a technical advantage in that the accuracy and / or generalizability of the trained encoder and decoder models may be improved. Use of a physics-informed regularizer enables integration of physical constraints during training and helps to constrain the solution space, leading to more accurate and physically consistent outputs from the trained encoder and decoder models, even with limited training data or in inference scenarios outside the training domain.
[0009] Examples of the present disclosure may also provide a technical advantage in that the interpretability of the encoder and decoder models may be improved. Use of a physics-informed regularizer may help to provide better insights into the relationship between inputs and outputs of the trained models based on established physical principles. This may help to improve model validation and analysis.
[0010] Examples of the present disclosure may also provide a technical advantage in that the data efficiency of the encoder and decoder models may be improved, meaning the models may achieve similar or better performance than conventional models, even using smaller training datasets. This is because the physics-based regularizer may help to prevent overfitting and may help to reduce the need for large amounts of training data.
[0011] Examples of the present disclosure may also provide a technical advantage in that the complexity of the encoder and decoder models may be reduced, which may enable the trained models to be more suitable for real-time applications and deployment on resource-constrained devices (such as UEs) .
[0012] Examples of the present disclosure may also provide a technical advantage in that the interoperability of encoder and decoder models trained by different vendors may be improved, by enabling separately-developed models to share a common physics-informed regularizer. Interoperability may be checked during training as well as after training is complete, by comparing regularizer values. Further, regularizer values may be included in feedback information, to help ensure interoperability of encoder and decoder models during inference.
[0013] Some embodiments of the present disclosure may include one or more of the following features, which can be separately adopted, or work together as a complete solution.
[0014] In an example of a first aspect, the present disclosure describes a method in a wireless communication system, the method including: obtaining a precoding matrix; compressing the precoding matrix into a low-dimensional representation using a first model; and transmitting channel state information (CSI) feedback including the low-dimensional representation and an indicator representing the precoding matrix, the indicator representing a characteristic of the precoding matrix.
[0015] In an example of the preceding example of the first aspect, the indicator may represent a regularizer value computed according to a defined formula that represents the characteristic of the precoding matrix based on a physical property of the wireless communication system.
[0016] In an example of any of the preceding examples of the first aspect, the characteristic of the precoding matrix may be a sparsity of the precoding matrix in an angular domain, a delay domain and / or a time domain.
[0017] In an example of any of the preceding examples of the first aspect, the transmitted CSI feedback may include multiple indicators representing the precoding matrix.
[0018] In an example of the preceding example of the first aspect, each indicator may represent a respective characteristic of the precoding matrix.
[0019] In an example of any of the preceding examples of the first aspect, the indicator included in the CSI feedback may be a quantized indicator.
[0020] In an example of any of the preceding examples of the first aspect, the method may further include: receiving parameters of the first model from a control signal or over a data channel.
[0021] In an example of any of the preceding examples of the first aspect, the first model may be an encoder model.
[0022] In an example of a second aspect, the present disclosure describes a method in a wireless communication system, the method including: receiving channel state information (CSI) feedback including a low-dimensional representation of a corresponding precoding matrix and a received indicator representing a characteristic of the corresponding precoding matrix; decompressing the low-dimensional representation using a second model to obtain a reconstructed precoding matrix; and performing a consistency verification by comparing a reconstructed indicator determined from the reconstructed precoding matrix against the received indicator.
[0023] In an example of the preceding example of the second aspect, the method may include: detecting an inconsistency based on a difference between the reconstructed indicator and the received indicator exceeding a threshold; and performing a corrective action.
[0024] In an example of the preceding example of the second aspect, the corrective action may include at least one of: transmitting a request for retransmission of CSI feedback; performing another decompression of the low-dimensional representation using a different second model to obtain a different reconstructed precoding matrix; or transmitting a request for codebook-based CSI feedback.
[0025] In an example of a preceding example of the second aspect, the method may include: detecting that the reconstructed precoding matrix is a successful reconstruction based on a difference between the reconstructed indicator and the received indicator being within a threshold; and performing a link operation and / or configuring a channel estimation operation based on the reconstructed precoding matrix.
[0026] In an example of the preceding example of the second aspect, performing the link operation and / or configuring the channel estimation operation may include at least one of: performing a beamforming operation; performing a link adaptation operation; configuring a set of resources for the channel estimation operation; or configuring a time duration of the channel estimation operation.
[0027] In an example of any of the preceding examples of the second aspect, the CSI feedback may include multiple received indicators.
[0028] In an example of the preceding example of the second aspect, each received indicator may represent a respective characteristic of the corresponding precoding matrix.
[0029] In an example of some of the preceding examples of the second aspect, performing the consistency verification may include comparing a combination of multiple reconstructed indicators computed from the reconstructed precoding matrix against a combination of the multiple received indicators.
[0030] In an example of the preceding example of the second aspect, the multiple reconstructed indicators may be combined using defined weights and the multiple received indicators may be combined using the defined weights.
[0031] In an example of any of the preceding examples of the second aspect, the method may further include: transmitting an indication of a performance of the second model based on comparison of the received indicator with the reconstructed indicator.
[0032] In an example of any of the preceding examples of the second aspect, the reconstructed indicator may represent a regularizer value computed according to a defined formula that represents the characteristic of the precoding matrix based on a physical property of the wireless communication system.
[0033] In an example of the preceding example of the second aspect, the characteristic of the corresponding precoding matrix may be a sparsity of the corresponding precoding matrix in an angular domain, a delay domain and / or a time domain.
[0034] In an example of any of the preceding examples of the second aspect, the method may include: receiving parameters of the second model from a control signal or over a data channel.
[0035] In an example of any of the preceding examples of the second aspect, the second model may be a decoder model.
[0036] In an example third aspect, the present disclosure describes a method in a wireless communication system, the method including: obtaining a training dataset including precoding matrix data; and training an autoencoder which includes training a first model of the autoencoder to perform a precoding matrix compression task or training a second model of the autoencoder to perform a precoding matrix decompression task, wherein the autoencoder is trained to minimize an error between an output of the autoencoder and the precoding matrix data, and is also trained to minimize a regularizer value representing a characteristic of each precoding matrix.
[0037] In an example of the preceding example of the third aspect, the method may include: during the training of the autoencoder, performing validation of a partly-trained first model or a partly-trained second model of a partly-trained autoencoder by: obtaining a validation dataset including precoding matrix data; obtaining a regularizer score representing regularizer values computed using outputs generated by the partly-trained autoencoder from the validation dataset; and transmitting the regularizer score.
[0038] In an example of the preceding example of the third aspect, the validation may be performed in response to a received signal.
[0039] In an example of a preceding example of the third aspect, the validation may be performed in accordance with a periodicity determined among a plurality of network entities of the wireless communication system.
[0040] In an example of some of the preceding examples of the third aspect, the validation dataset may be common to two or more network entities of the wireless communication system, the validation dataset may be associated with a defined formula for obtaining the regularizer score, and the regularizer score may be transmitted to at least one other network entity of the wireless communication system.
[0041] In an example of some of the preceding examples of the third aspect, the method may include: receiving a signal to end training, wherein the signal is received in response to transmitting the regularizer score.
[0042] In an example of some of the preceding examples of the third aspect, the method may include: receiving a reference score associated with the validation dataset; and comparing the regularizer score with the reference score to evaluate the partly-trained autoencoder.
[0043] In an example of some of the preceding examples of the third aspect, the method may include: terminating training based on the regularizer score satisfying a target threshold.
[0044] In an example of some of the preceding examples of the third aspect, the method may include: terminating training in response to a received control signal.
[0045] In an example of any of the preceding examples of the third aspect, the method may include: performing evaluation of a trained first model or a trained second model of a trained autoencoder by: obtaining a test dataset including precoding matrix data; obtaining a final regularizer score based on regularizer values computed using outputs generated by the trained autoencoder from the test dataset; and transmitting the final regularizer score.
[0046] In an example of the preceding example of the third aspect, the method may include: receiving another regularizer score from another network entity; and comparing the final regularizer score with the received another regularizer score to determine interoperability of the trained first model or the trained second model with the another network entity.
[0047] In an example of any of the preceding examples of the third aspect, the autoencoder may be trained to minimize multiple regularizer values together.
[0048] In an example of the preceding examples of the third aspect, the regularizer score may be computed based on multiple regularizer values computed using outputs generated by the partly-trained autoencoder from the validation dataset, and the multiple regularizer values may be combined using defined weights.
[0049] In an example of any of the preceding examples of the third aspect, the regularizer value may be computed according to a defined formula representing the characteristic of the precoding matrix based on a physical property of the wireless communication system.
[0050] In an example of the preceding example of the third aspect, the characteristic of the precoding matrix may be a sparsity of the precoding matrix in an angular domain, a delay domain, and / or a time domain.
[0051] In an example of some of the preceding examples of the third aspect, the defined formula may be a differentiable formula.
[0052] In an example of the preceding example of the third aspect, the defined formula may be based on a discrete Fourier transform (DFT) of the precoding matrix.
[0053] In an example of any of the preceding examples of the third aspect, the first model may be an encoder model and / or the second model may be a decoder model.
[0054] In another example aspect, the present disclosure describes an apparatus including: a memory; and a processor configured to execute instructions stored in the memory to cause the apparatus to carry out any of the preceding examples of the first aspect.
[0055] In another example aspect, the present disclosure describes a base station including: a memory; and a processor configured to execute instructions stored in the memory to cause the base station to carry out any of the preceding examples of the second aspect.
[0056] In another example aspect, the present disclosure describes a network entity including: a memory; and a processor configured to execute instructions stored in the memory to cause the network entity to carry out any of the preceding examples of the third aspect.
[0057] In another example aspect, the present disclosure describes a non-transitory computer readable medium having machine-executable instructions stored thereon, wherein the instructions, when executed by an apparatus, cause the apparatus to perform any of the preceding examples of the first aspect.
[0058] In another example aspect, the present disclosure describes a non-transitory computer readable medium having machine-executable instructions stored thereon, wherein the instructions, when executed by a base station, cause the base station to perform any of the preceding examples of the second aspect.
[0059] In another example aspect, the present disclosure describes a non-transitory computer readable medium having machine-executable instructions stored thereon, wherein the instructions, when executed by a network entity, cause the network entity to perform any of the preceding examples of the third aspect.
[0060] In another example aspect, the present disclosure describes a communication apparatus configured to perform any of the preceding examples of the first aspect, the second aspect or the third aspect.
[0061] In any of the preceding examples, the communication apparatus may include an obtaining unit configured to perform the obtaining, a processing unit configured to perform the compressing, and a transmitting unit configured to perform the transmitting.
[0062] In any of the preceding examples, the communication apparatus may include one or more processors configured to perform the obtaining and compressing; and an interface circuit configured to perform the transmitting.
[0063] In any of the preceding examples, the communication apparatus may include a receiving unit configured to perform the receiving; and a processing unit configured to perform the decompressing and comparing.
[0064] In any of the preceding examples, the communication apparatus may include: one or more processors configured to perform the decompressing and comparing; and an interface circuit configured to perform the receiving.
[0065] In any of the preceding examples, the communication apparatus may include an obtaining unit configured to perform the obtaining; and a processing unit configured to perform the training.
[0066] In any of the preceding examples, the communication apparatus may be a network entity.
[0067] In any of the preceding examples, the network entity may be a datacenter. In any of the preceding examples, the network entity may service one or more user equipment (UEs) . In any of the preceding examples, the network entity may service one or more base stations (BSs) .
[0068] In another example aspect, the present disclosure describes a computer program characterized in that, when the computer program is run on a computer, the computer is caused to execute any of the preceding example aspects.BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Reference will now be made, by way of example, to the accompanying drawings which show example embodiments of the present application, and in which:
[0070] FIG. 1 illustrates an example of how a physics-based regularizer may be used for two-sided model-based CSI feedback, in accordance with examples of the present disclosure;
[0071] FIG. 2 is a simplified schematic illustration of a communication system, which may be used to implement examples of the present disclosure;
[0072] FIG. 3 is a block diagram illustrating an example of a communication system, which may be used to implement examples of the present disclosure;
[0073] FIG. 4 is a block diagram illustrating an example of the devices of a communication system, which may be used to implement examples of the present disclosure;
[0074] FIG. 5A is a block diagram illustrating example units or modules in a device, which may be used to implement examples of the present disclosure;
[0075] FIG. 5B is a block diagram illustrating an example apparatus, which may be used to implement examples of the present disclosure;
[0076] FIG. 6 illustrates an example of CSI compression and reconstruction using encoder and decoder models of an autoencoder, in accordance with examples of the present disclosure;
[0077] FIG. 7 illustrates an example of how encoder and decoder models may be trained and used in inference for CSI feedback, in accordance with examples of the present disclosure;
[0078] FIGS. 8A and 8B together illustrate an example incorporating a physics-based regularizer in training of an autoencoder for CSI compression, in accordance with examples of the present disclosure;
[0079] FIG. 9 illustrates another example incorporating a physics-based regularizer in training of an autoencoder for CSI compression, in accordance with examples of the present disclosure;
[0080] FIG. 10 illustrates another example incorporating a physics-based regularizer in training of an autoencoder for CSI compression, in accordance with examples of the present disclosure;
[0081] FIG. 11 illustrates an example method for performing a mid-training validation of encoder and decoder models between a UE-vendor and a NW-vendor, in accordance with examples of the present disclosure;
[0082] FIG. 12 illustrates an example method for providing CSI feedback including a physics-based regularizer, in accordance with examples of the present disclosure;
[0083] FIG. 13 is a flowchart illustrating an example method for performing CSI feedback using an encoder model to compress the information in the CSI feedback, in accordance with examples of the present disclosure;
[0084] FIG. 14 is a flowchart illustrating an example method for receiving CSI feedback and decompressing information in the CSI feedback using a decoder model, in accordance with examples of the present disclosure; and
[0085] FIG. 15 is a flowchart illustrating an example method for training an encoder model or decoder model for compression or decompression of CSI feedback, in accordance with examples of the present disclosure.
[0086] Similar reference numerals may have been used in different figures to denote similar components.DETAILED DESCRIPTION
[0087] The concept of a two-sided model has emerged as a potential solution for efficient CSI feedback and utilization. The two-sided model involves a pair of artificial intelligence (AI) / machine learning (ML) models: a CSI generation model (encoder model) at the user equipment (UE) and a CSI reconstruction model (decoder model) at the base station (e.g., gNodeB (gNB) ) . The UE model compresses the CSI and transmits it to the gNB, where the network model reconstructs the CSI for subsequent processing and optimization tasks. This approach aims to reduce the overhead associated with CSI feedback while maintaining sufficient accuracy for effective link adaptation and beamforming.
[0088] However, there are challenges such as ensuring interoperability between UE and network vendors, as each may employ different model architectures and training datasets. Additionally, the complexity of existing AI / ML models for CSI compression raises concerns regarding processing demands and power consumption on UEs having limited computing resources.
[0089] In various examples, the present disclosure describes methods, apparatuses, network entities and computer-readable media that enable two-sided machine learning-based compression of CSI. Examples disclosed herein implement a regularizer representing a physical property of the wireless communication system during both training and inference phases. As discussed further below, use of such a physics-based regularizer may be useful for improving interoperability, improving accuracy and / or reducing complexity of encoder and decoder models, among other advantages.
[0090] The most recent study on artificial intelligence (AI) / machine learning (ML) for New Release (NR) air interface explored various aspects of the two-sided model for CSI compression, including model architectures, training collaboration types, performance evaluation, and potential specification impact. The study demonstrated the feasibility of achieving significant performance gains compared to legacy codebook-based approaches like eType II. However, several challenges remain to be addressed before widespread adoption of this technology.
[0091] One key challenge lies in ensuring interoperability between UE and network vendors, as each may employ different model architectures and training datasets. This necessitates exploring standardization options for reference models, datasets, or model structures, while minimizing the complexity of inter-vendor collaboration. Additionally, the complexity of existing AI / ML models for CSI compression raises concerns regarding processing demands and power consumption on UEs, prompting investigations into techniques for reducing model complexity without compromising performance. In the present disclosure, the term UE vendor may refer to a network entity (which may be a server, a computing system, etc. ) that services UEs, and the term network vendor (sometimes simplified to NW vendor) may refer to a network entity (which may be a server, a computing system, etc. ) that services other network entities such as base stations.
[0092] Furthermore, the 3rd Generation Partnership Project (3GPP) Release-18 (Rel-18) study primarily focused on spatial and frequency domain compression. The most recent study aims to extend this to the temporal domain, leveraging historical CSI information to further improve compression efficiency. This introduces additional considerations regarding prediction methods, handling non-ideal uplink control information (UCI) feedback, and managing the accumulated CSI information at both the UE and gNodeB (gNB) . Addressing these challenges is crucial for enabling robust and efficient CSI compression in dynamic wireless environments.
[0093] Physics-Informed Machine Learning
[0094] Physics-informed machine learning is a rapidly evolving field that seeks to bridge the gap between data-driven and physics-based modeling approaches. Traditional machine learning methods excel at learning complex patterns from data but often lack interpretability and generalizability, especially when faced with limited data or scenarios outside the training domain. On the other hand, physics-based models, derived from first principles and physical laws, offer strong interpretability and generalizability but can be challenging to develop and computationally expensive to solve, particularly for complex systems.
[0095] Physics-informed machine learning aims to leverage the strengths of both worlds. By incorporating physical laws and domain knowledge into the machine learning process, these hybrid approaches can achieve several benefits, which may include one or more of:
[0096] ● Improved Accuracy and Generalizability: Integrating physical constraints can guide the learning process and constrain the solution space, leading to more accurate and physically consistent predictions, even with limited data or in scenarios outside the training domain. This is particularly valuable for complex systems where purely data-driven models may struggle to capture the underlying physics.
[0097] ● Enhanced Interpretability: Incorporating physical knowledge can make the models more interpretable by providing insights into the relationship between inputs and outputs based on established physical principles. This enables a deeper understanding of the system and facilitates model validation and analysis.
[0098] ● Data Efficiency: Physics-informed models can often achieve similar or better performance with smaller datasets compared to purely data-driven approaches. This is because the physical constraints act as a form of regularization, preventing overfitting and reducing the need for large amounts of data.
[0099] ● Reduced Complexity: By leveraging prior knowledge about the system, physics-informed models can potentially be designed with smaller sizes and lower computational complexity, making them more suitable for real-time applications and deployment on resource-constrained devices.
[0100] Several techniques have been developed for incorporating physical knowledge into machine learning models, including:
[0101] ● Physics-informed neural networks (PINNs) : These neural networks embed physical laws as additional constraints within the loss function, ensuring that the learned model respects the underlying physics. This can be achieved by incorporating differential equations, conservation laws, or other physical principles.
[0102] ● Hybrid modeling: This approach combines physics-based models with data-driven components to leverage the strengths of both. For example, a physics-based model can be used to capture the dominant dynamics of a system, while a neural network can be used to learn the residual dynamics or complex interactions that are difficult to model explicitly.
[0103] ● Scientific Machine Learning: This broad field encompasses various methods that combine scientific computing with machine learning. This includes techniques for incorporating physical knowledge into model design, data generation, and uncertainty quantification.
[0104] Physics-informed machine learning has found applications in diverse fields, including Fluid dynamics, Material science, Biomechanics, Engineering design and so on.
[0105] By bridging the gap between data-driven and physics-based approaches, physics-informed machine learning holds immense promise for advancing scientific discovery and engineering design, ultimately leading to a deeper understanding of complex systems and more efficient solutions to real-world problems.
[0106] This disclosure presents a novel approach to CSI compression that leverages the principles of physics-informed machine learning. By incorporating knowledge about the physical properties of wireless channels, such as sparsity in the angular and delay domains, examples of the present disclosure may enhance the efficiency and accuracy of AI / ML models for CSI compression. This approach offers several benefits, which may include smaller model sizes, reduced data requirements, faster convergence during training, and improved interoperability between UE and network vendors, among others.
[0107] The disclosed method involves utilizing physics-based regularizers, such as the L1 norm of the Discrete Fourier Transform (DFT) of the beamforming matrix in the angular and delay domains, to guide the learning process and promote sparsity in the model's representation. This leads to more efficient compression and better reconstruction accuracy. Furthermore, the inclusion of the regularizer value in the CSI feedback enables the gNB (or more generally a base station or other network entity) to detect potential errors during reconstruction, enhancing the robustness of the system.
[0108] By embracing physics-informed machine learning, it may be possible to unlock new possibilities for CSI compression, paving the way for more efficient, reliable, and interoperable communication in future wireless networks. This approach holds the potential to significantly improve spectral efficiency, reduce overhead, and contribute to the advancement of 5G and beyond.
[0109] FIG. 1 illustrates an example of how a physics-based regularizer may be used in two-sided CSI compression, as disclosed herein. For example, at the UE 10, CSI may be encoded using an encoder model (e.g., implemented using a trained neural network) in which physical properties such as an expected sparsity of the beamforming matrix (or precoder matrix) in the angular and delay domains are represented by physics-based regularizers. Sparsity refers to the characteristic that the precoder matrix contains elements that are mostly zero (e.g., the number of non-zero elements of the precoder matrix may be fewer than half of the total number of elements in the matrix, or the number of non-zero elements of the precoder matrix may be approximately no more than the number of rows or columns of the matrix) . The physics-based regularizers may be computed using the L1 norm of the DFT of the beamforming matrix, for example. The compressed CSI that is outputted by the encoder model may be provided as feedback to the base station (BS) 12 together with the regularizer value. At the BS, a decoder model (e.g., implemented using another trained neural network) may be used to decode (or decompress) the CSI. The decoder model has been trained based on the same physics-based regularizers as the encoder model. For example, the equations defining the physics-based regularizers may be defined in standards.
[0110] In various examples of the present disclosure, the encoder model may be referred to as a first model and the decoder model may be referred to as a second model. The term “first model” may be understood to encompass an encoder model or any other machine learning model that is designed to perform the task of CSI compression as disclosed herein; in various examples. Similarly, the term “second model” may be understood to encompass a decoder model or any other machine learning model that is designed to perform the task of CSI reconstruction as disclosed herein.
[0111] To assist in understanding the present disclosure, reference is now made to FIG. 2.
[0112] Referring to FIG. 2, as an illustrative example without limitation, a simplified schematic illustration of a communication system is provided. The communication system 100 comprises a radio access network 120. The radio access network 120 may be a next generation radio access network, or a legacy (e.g. 5G, 4G, 3G or 2G) radio access network. One or more communication electronic devices (ED) 110a, 110b, 110c, 110d, 110e, 110f, 110g, 110h, 110i, 110j (generically referred to as 110) may be interconnected to one another or connected to one or more network nodes (170a, 170b, generically referred to as 170) in the radio access network 120. A core network 130 may be a part of the communication system and may be dependent or independent of the radio access technology used in the communication system 100. Also the communication system 100 comprises a public switched telephone network (PSTN) 140, the internet 150, and other networks 160.
[0113] FIG. 3 illustrates an example communication system 100. In general, the communication system 100 enables multiple wireless or wired elements to communicate data and other content. The purpose of the communication system 100 may be to provide content, such as voice, data, video, and / or text, via broadcast, multicast, groupcast, unicast, etc. The communication system 100 may operate by sharing resources, such as carrier spectrum bandwidth, between its constituent elements. The communication system 100 may include a terrestrial communication system and / or a non-terrestrial communication system. The communication system 100 may provide a wide range of communication services and applications (such as earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery and mobility, etc. ) . The communication system 100 may provide a high degree of availability and robustness through a joint operation of a terrestrial communication system and a non-terrestrial communication system. For example, integrating a non-terrestrial communication system (or components thereof) into a terrestrial communication system can result in what may be considered a heterogeneous network comprising multiple layers. Compared to conventional communication networks, the heterogeneous network may achieve better overall performance through efficient multi-link joint operation, more flexible functionality sharing, and faster physical layer link switching between terrestrial networks and non-terrestrial networks.
[0114] The terrestrial communication system and the non-terrestrial communication system can be considered sub-systems of the communication system. In the example shown in FIG. 2, the communication system 100 includes electronic devices (ED) 110a, 110b, 110c, 110d (generically referred to as ED 110) , radio access networks (RANs) 120a, 120b, a non-terrestrial communication network 120c, a core network 130, a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160. The RANs 120a, 120b include respective base stations (BSs) 170a, 170b, which may be generically referred to as terrestrial transmit and receive points (T-TRPs) 170a, 170b. The non-terrestrial communication network 120c includes an access node 172, which may be generically referred to as a non-terrestrial transmit and receive point (NT-TRP) 172.
[0115] Any ED 110 may be alternatively or additionally configured to interface, access, or communicate with any T-TRP 170a, 170b and NT-TRP 172, the Internet 150, the core network 130, the PSTN 140, the other networks 160, or any combination of the preceding. In some examples, ED 110a may communicate via an uplink and / or downlink transmission over a terrestrial air interface 190a with T-TRP 170a. In some examples, the EDs 110a, 110b, 110c, and 110d may also communicate directly with one another via one or more sidelink air interfaces 190b. In some examples, ED 110d may communicate via an uplink and / or downlink transmission over a non-terrestrial air interface 190c with NT-TRP 172.
[0116] The air interfaces 190a and 190b may use similar communication technology, such as any suitable radio access technology. For example, the communication system 100 may implement one or more channel access methods, such as code division multiple access (CDMA) , space division multiple access (SDMA) , time division multiple access (TDMA) , frequency division multiple access (FDMA) , orthogonal FDMA (OFDMA) , or single-carrier FDMA (SC-FDMA, also known as discrete Fourier transform spread OFDMA, DFT-s-OFDMA) in the air interfaces 190a and 190b. The air interfaces 190a and 190b may utilize other higher dimension signal spaces, which may involve a combination of orthogonal and / or non-orthogonal dimensions.
[0117] The non-terrestrial air interface 190c can enable communication between the ED 110d and one or multiple NT-TRPs 172 via a wireless link or simply a link. For some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection between a group of EDs 110 and one or multiple NT-TRPs 172 for multicast transmission.
[0118] The RANs 120a and 120b are in communication with the core network 130 to provide the EDs 110a 110b, and 110c with various services such as voice, data, and other services. The RANs 120a and 120b and / or the core network 130 may be in direct or indirect communication with one or more other RANs (not shown) , which may or may not be directly served by core network 130, and may or may not employ the same radio access technology as RAN 120a, RAN 120b or both. The core network 130 may also serve as a gateway access between (i) the RANs 120a and 120b or EDs 110a 110b, and 110c or both, and (ii) other networks (such as the PSTN 140, the Internet 150, and the other networks 160) . In addition, some or all of the EDs 110a 110b, and 110c may include functionality for communicating with different wireless networks over different wireless links using different wireless technologies and / or protocols. Instead of wireless communication (or in addition thereto) , the EDs 110a 110b, and 110c may communicate via wired communication channels to a service provider or switch (not shown) , and to the Internet 150. PSTN 140 may include circuit switched telephone networks for providing plain old telephone service (POTS) . Internet 150 may include a network of computers and subnets (intranets) or both, and incorporate protocols, such as Internet Protocol (IP) , Transmission Control Protocol (TCP) , User Datagram Protocol (UDP) . EDs 110a 110b, and 110c may be multimode devices capable of operation according to multiple radio access technologies, and incorporate multiple transceivers necessary to support such.
[0119] FIG. 4 illustrates another example of an ED 110 and a base station 170a, 170b and / or 170c. The ED 110 is used to connect persons, objects, machines, etc. The ED 110 may be widely used in various scenarios including, for example, cellular communications, device-to-device (D2D) , vehicle to everything (V2X) , peer-to-peer (P2P) , machine-to-machine (M2M) , machine-type communications (MTC) , internet of things (IoT) , virtual reality (VR) , augmented reality (AR) , mixed reality (MR) , metaverse, digital twin, industrial control, self-driving, remote medical, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery and mobility, etc.
[0120] Each ED 110 represents any suitable end user device for wireless operation and may include such devices (or may be referred to) as a user equipment / device (UE) , a wireless transmit / receive unit (WTRU) , a mobile station, a fixed or mobile subscriber unit, a cellular telephone, a station (STA) , a machine type communication (MTC) device, a personal digital assistant (PDA) , a smartphone, a laptop, a computer, a tablet, a wireless sensor, a consumer electronics device, a smart book, a vehicle, a car, a truck, a bus, a train, or an IoT device, wearable devices (such as a watch, a pair of glasses, head mounted equipment, etc. ) , an industrial device, or an apparatus in (e.g. communication module, modem, or chip) or comprising the forgoing devices, among other possibilities. Future generation EDs 110 may be referred to using other terms. The base station 170a and 170b is a T-TRP and will hereafter be referred to as T-TRP 170. Also shown in FIG. 3, a NT-TRP will hereafter be referred to as NT-TRP 172. Each ED 110 connected to T-TRP 170 and / or NT-TRP 172 can be dynamically or semi-statically turned-on (i.e., established, activated, or enabled) , turned-off (i.e., released, deactivated, or disabled) and / or configured in response to one of more of: connection availability and connection necessity.
[0121] The ED 110 includes a transmitter 201 and a receiver 203 coupled to one or more antennas 204. Only one antenna 204 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas 204 may alternatively be panels. The transmitter 201 and the receiver 203 may be integrated, e.g. as a transceiver. The transceiver is configured to modulate data or other content for transmission by at least one antenna 204 or network interface controller (NIC) . The transceiver is also configured to demodulate data or other content received by the at least one antenna 204. Each transceiver includes any suitable structure for generating signals for wireless or wired transmission and / or processing signals received wirelessly or by wire. Each antenna 204 includes any suitable structure for transmitting and / or receiving wireless or wired signals.
[0122] The ED 110 includes at least one memory 208. The memory 208 stores instructions and data used, generated, or collected by the ED 110. For example, the memory 208 could store software instructions or modules configured to implement some or all of the functionality and / or embodiments described herein and that are executed by one or more processing unit (s) (e.g., a processor 210) . Each memory 208 includes any suitable volatile and / or non-volatile storage and retrieval device (s) . Any suitable type of memory may be used, such as random access memory (RAM) , read only memory (ROM) , hard disk, optical disc, subscriber identity module (SIM) card, memory stick, secure digital (SD) memory card, on-processor cache, and the like.
[0123] The ED 110 may further include one or more input / output devices (not shown) or interfaces (such as a wired interface to the Internet 150 in FIG. 2) . The input / output devices or interfaces permit interaction with a user or other devices in the network. Each input / output device or interface includes any suitable structure for providing information to or receiving information from a user, and / or for network interface communications. Suitable structures include, for example, a speaker, microphone, keypad, keyboard, display, touch screen, etc.
[0124] The ED 110 includes the processor 210 for performing operations including those operations related to preparing a transmission for uplink transmission to the NT-TRP 172 and / or the T-TRP 170; those operations related to processing downlink transmissions received from the NT-TRP 172 and / or the T-TRP 170; and those operations related to processing sidelink transmission to and from another ED 110. Processing operations related to preparing a transmission for uplink transmission may include operations such as encoding, modulating, transmit beamforming, and generating symbols for transmission. Processing operations related to processing downlink transmissions may include operations such as receive beamforming, demodulating and decoding received symbols. Depending upon the embodiment, a downlink transmission may be received by the receiver 203, possibly using receive beamforming, and the processor 210 may extract signaling from the downlink transmission (e.g. by detecting and / or decoding the signaling) . An example of signaling may be a reference signal transmitted by the NT-TRP 172 and / or by the T-TRP 170. In some embodiments, the processor 210 implements the transmit beamforming and / or the receive beamforming based on the indication of beam direction, e.g. beam angle information (BAI) , received from the T-TRP 170. In some embodiments, the processor 210 may perform operations relating to network access (e.g. initial access) and / or downlink synchronization, such as operations relating to detecting a synchronization sequence, decoding and obtaining the system information, etc. In some embodiments, the processor 210 may perform channel estimation, e.g. using a reference signal received from the NT-TRP 172 and / or from the T-TRP 170.
[0125] Although not illustrated, the processor 210 may form part of the transmitter 201 and / or part of the receiver 203. Although not illustrated, the memory 208 may form part of the processor 210.
[0126] The processor 210, the processing components of the transmitter 201, and the processing components of the receiver 203 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory (e.g. in the memory 208) . Alternatively, some or all of the processor 210, the processing components of the transmitter 201, and the processing components of the receiver 203 may each be implemented using dedicated circuitry, such as a programmed field-programmable gate array (FPGA) , an application-specific integrated circuit (ASIC) , or a hardware accelerator such as a graphics processing unit (GPU) or an artificial intelligence (AI) accelerator.
[0127] The T-TRP 170 may be known by other names in some implementations, such as a base station, a base transceiver station (BTS) , a radio base station, a network node, a network device, a device on the network side, a transmit / receive node, a Node B, an evolved NodeB (eNodeB or eNB) , a Home eNodeB, a next Generation NodeB (gNB) , a transmission point (TP) , a site controller, an access point (AP) , a wireless router, a relay station, a terrestrial node, a terrestrial network device, a terrestrial base station, a base band unit (BBU) , a remote radio unit (RRU) , an active antenna unit (AAU) , a remote radio head (RRH) , a central unit (CU) , a distributed unit (DU) , a positioning node, among other possibilities. The T-TRP 170 may be a macro BS, a pico BS, a relay node, a donor node, or the like, or combinations thereof. The T-TRP 170 may refer to the forgoing devices or refer to apparatus (e.g. a communication module, a modem, or a chip) in the forgoing devices.
[0128] In some embodiments, the parts of the T-TRP 170 may be distributed. For example, some of the modules of the T-TRP 170 may be located remote from the equipment that houses the antennas 256 for the T-TRP 170, and may be coupled to the equipment that houses the antennas 256 over a communication link (not shown) sometimes known as front haul, such as common public radio interface (CPRI) . Therefore, in some embodiments, the term T-TRP 170 may also refer to modules on the network side that perform processing operations, such as determining the location of the ED 110, resource allocation (scheduling) , message generation, and encoding / decoding, and that are not necessarily part of the equipment that houses the antennas 256 of the T-TRP 170. The modules may also be coupled to other T-TRPs. In some embodiments, the T-TRP 170 may actually be a plurality of T-TRPs that are operating together to serve the ED 110, e.g. through the use of coordinated multipoint transmissions.
[0129] The T-TRP 170 includes at least one transmitter 252 and at least one receiver 254 coupled to one or more antennas 256. Only one antenna 256 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas 256 may alternatively be panels. The transmitter 252 and the receiver 254 may be integrated as a transceiver. The T-TRP 170 further includes a processor 260 for performing operations including those related to:preparing a transmission for downlink transmission to the ED 110, processing an uplink transmission received from the ED 110, preparing a transmission for backhaul transmission to the NT-TRP 172, and processing a transmission received over backhaul from the NT-TRP 172. Processing operations related to preparing a transmission for downlink or backhaul transmission may include operations such as encoding, modulating, precoding (e.g. multiple input multiple output (MIMO) precoding) , transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the uplink or over backhaul may include operations such as receive beamforming, demodulating received symbols, and decoding received symbols. The processor 260 may also perform operations relating to network access (e.g. initial access) and / or downlink synchronization, such as generating the content of synchronization signal blocks (SSBs) , generating the system information, etc. In some embodiments, the processor 260 also generates an indication of beam direction, e.g. BAI, which may be scheduled for transmission by a scheduler 253. The processor 260 performs other network-side processing operations described herein, such as determining the location of the ED 110, determining where to deploy the NT-TRP 172, etc. In some embodiments, the processor 260 may generate signaling, e.g. to configure one or more parameters of the ED 110 and / or one or more parameters of the NT-TRP 172. Any signaling generated by the processor 260 is sent by the transmitter 252. Note that “signaling” , as used herein, may alternatively be called control signaling. Signaling may be transmitted in a physical layer control channel, e.g. a physical downlink control channel (PDCCH) , in which case the signaling may be known as dynamic signaling. Signaling transmitted in a downlink physical layer control channel may be known as Downlink Control Information (DCI) . Signalling transmitted in an uplink physical layer control channel may be known as Uplink Control Information (UCI) . Signaling transmitted in a sidelink physical layer control channel may be known as Sidelink Control Information (SCI) . Signaling may be included in a higher-layer (e.g., higher than physical layer) packet transmitted in a physical layer data channel, e.g. in a physical downlink shared channel (PDSCH) , in which case the signaling may be known as higher-layer signaling, static signaling, or semi-static signaling. Higher-layer signaling may also refer to Radio Resource Control (RRC) protocol signaling or Media Access Control –Control Element (MAC-CE) signaling.
[0130] The scheduler 253 may be coupled to the processor 260. The scheduler 253 may be included within or operated separately from the T-TRP 170. The scheduler 253 may schedule uplink, downlink, sidelink, and / or backhaul transmissions, including issuing scheduling grants and / or configuring scheduling-free (e.g., “configured grant” ) resources. The T-TRP 170 further includes a memory 258 for storing information and data. The memory 258 stores instructions and data used, generated, or collected by the T-TRP 170. For example, the memory 258 could store software instructions or modules configured to implement some or all of the functionality and / or embodiments described herein and that are executed by the processor 260.
[0131] Although not illustrated, the processor 260 may form part of the transmitter 252 and / or part of the receiver 254. Also, although not illustrated, the processor 260 may implement the scheduler 253. Although not illustrated, the memory 258 may form part of the processor 260.
[0132] The processor 260, the scheduler 253, the processing components of the transmitter 252, and the processing components of the receiver 254 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory, e.g. in the memory 258. Alternatively, some or all of the processor 260, the scheduler 253, the processing components of the transmitter 252, and the processing components of the receiver 254 may be implemented using dedicated circuitry, such as a programmed FPGA, a hardware accelerator (e.g., a GPU or AI accelerator) , or an ASIC.
[0133] Although the NT-TRP 172 is illustrated as a drone only as an example, the NT-TRP 172 may be implemented in any suitable non-terrestrial form, such as satellites and high altitude platforms, including international mobile telecommunication base stations and unmanned aerial vehicles, for example. Also, the NT-TRP 172 may be known by other names in some implementations, such as a non-terrestrial node, a non-terrestrial network device, or a non-terrestrial base station. The NT-TRP 172 includes a transmitter 272 and a receiver 274 coupled to one or more antennas 280. Only one antenna 280 is illustrated to avoid congestion in the drawing. One, some, or all of the antennas may alternatively be panels. The transmitter 272 and the receiver 274 may be integrated as a transceiver. The NT-TRP 172 further includes a processor 276 for performing operations including those related to: preparing a transmission for downlink transmission to the ED 110, processing an uplink transmission received from the ED 110, preparing a transmission for backhaul transmission to T-TRP 170, and processing a transmission received over backhaul from the T-TRP 170. Processing operations related to preparing a transmission for downlink or backhaul transmission may include operations such as encoding, modulating, precoding (e.g. MIMO precoding) , transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the uplink or over backhaul may include operations such as receive beamforming, demodulating received symbols, and decoding received symbols. In some embodiments, the processor 276 implements the transmit beamforming and / or receive beamforming based on beam direction information (e.g. BAI) received from the T-TRP 170. In some embodiments, the processor 276 may generate signaling, e.g. to configure one or more parameters of the ED 110. In some embodiments, the NT-TRP 172 implements physical layer processing, but does not implement higher layer functions such as functions at the medium access control (MAC) or radio link control (RLC) layer. As this is only an example, more generally, the NT-TRP 172 may implement higher layer functions in addition to physical layer processing.
[0134] The NT-TRP 172 further includes a memory 278 for storing information and data. Although not illustrated, the processor 276 may form part of the transmitter 272 and / or part of the receiver 274. Although not illustrated, the memory 278 may form part of the processor 276.
[0135] The processor 276, the processing components of the transmitter 272, and the processing components of the receiver 274 may each be implemented by the same or different one or more processors that are configured to execute instructions stored in a memory, e.g. in the memory 278. Alternatively, some or all of the processor 276, the processing components of the transmitter 272, and the processing components of the receiver 274 may be implemented using dedicated circuitry, such as a programmed FPGA, a hardware accelerator (e.g., a GPU or AI accelerator) , or an ASIC. In some embodiments, the NT-TRP 172 may actually be a plurality of NT-TRPs that are operating together to serve the ED 110, e.g. through coordinated multipoint transmissions.
[0136] The T-TRP 170, the NT-TRP 172, and / or the ED 110 may include other components, but these have been omitted for the sake of clarity.
[0137] One or more steps of the embodiment methods provided herein may be performed by corresponding units or modules, according to FIG. 5A. FIG. 5A illustrates units or modules in a device, such as in the ED 110, in the T-TRP 170, or in the NT-TRP 172. For example, a signal may be transmitted by a transmitting unit or by a transmitting module. A signal may be received by a receiving unit or by a receiving module. A signal may be processed by a processing unit or a processing module. Other steps may be performed by an artificial intelligence (AI) or machine learning (ML) module. The respective units or modules may be implemented using hardware, one or more components or devices that execute software, or a combination thereof. For instance, one or more of the units or modules may be a circuit such as an integrated circuit. Examples of an integrated circuit includes a programmed FPGA, a GPU, or an ASIC. For instance, one or more of the units or modules may be logical such as a logical function performed by a circuit, by a portion of an integrated circuit, or by software instructions executed by a processor. It will be appreciated that where the modules are implemented using software for execution by a processor for example, the modules may be retrieved by a processor, in whole or part as needed, individually or together for processing, in single or multiple instances, and that the modules themselves may include instructions for further deployment and instantiation.
[0138] FIG. 5B illustrates an example apparatus 410 according to an implementation of the present disclosure. The apparatus 410 may be a communication device or an apparatus implemented in a communication device such as the ED 110 or the TRPs 170a, 170b, 172. For example, the apparatus 410 implemented in an ED may be an integrated circuit, which in some instances may be referred to as a chip, a modem, a modem chip, a baseband chip, or a baseband processor. In some implementations, one or more integrated circuits can be packaged into a system-on-chip, a system-in-package, or a multi-chip module. The apparatus 410 can include one or more integrated circuits and other discrete components. In some implementations, the apparatus 410 may be a module within the ED 110. In some implementations, the apparatus 410 may be a module within one of the TRPs 170a, 170b, 172.
[0139] In an example, the apparatus 410 may include one or more processors 411, and an interface circuit 412. The apparatus 410 may further include a memory 413. The one or more processors 411 are configured to process signals and execute one or more communication protocols. The memory 413 is configured to store at least a part of corresponding computer program instructions and / or data. In an example, the one or more processors 411 execute the computer program instructions stored in the memory 413 to implement related operations (for example, inputting, outputting, receiving, and transmitting) in the method embodiments disclosed herein. In some implementations, the memory 413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 413 is configured to store all of the corresponding computer program instructions and / or data for execution by the one or more processors 411. In some implementations, the memory 413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 413 is configured to store a part of the corresponding computer program instructions and / or data. For example, the part of the corresponding computer program instructions and / or data may include computer program instructions and / or data that need to be currently executed by the one or more processors 411. Thus, the memory 413 may store different parts of computer program instructions and / or data for a plurality times for the one or more processors 411 to perform related operations in the method embodiments disclosed herein. As a communication interface, the interface circuit 412 is configured to implement communication with another component. For example, the interface circuit 412 may communicate a signal with another apparatus or system, such as a radio frequency processing apparatus or another processor. The signal may include or carry information intended as a payload, such as user data, control information, etc. The signal may also include or carry information useful to a receiver, but not necessarily as a payload, such as a pilot signal or reference signal. Communicating the signal may include transmitting the signal to another component or device. Communicating the signal may additionally or alternatively include receiving the signal from another component or device. Transmitting the signal may include outputting the signal to a component or device that is directly or indirectly coupled to the interface circuit 412. Receiving the signal may include inputting or obtaining the signal from a component or device that is directly or indirectly couped to the interface circuit 412. Optionally, to reduce a load of the one or more processors, a baseband signal processing circuit 414 may be also disposed to implement processing of at least a part of baseband signals, including signal demodulation, modulation, encoding, decoding, or the like.
[0140] The apparatus 410 may be the processor 210 (or 260 or 276) within the ED 110 (or T-TRP 170 or NT-TRP 172) , in some scenarios, or may be included within the processor 210 (or 260 or 276) within the ED 110 (or T-TRP 170 or NT-TRP 172) in some scenarios. The apparatus 410 may be a baseband chip or may include a baseband chip. In some implementations, the apparatus 410 may be independently packaged into a chip. In some implementations, the ED 110 (or T-TRP 170 or NT-TRP 172) includes different types of chips. The apparatus 410 may be packaged into a processor chip (for example, an SoC chip or an SIP chip) with the different types of chips. In some implementations, the apparatus 410 may be packaged into a chip with some or all of circuits of a radio frequency processing system that may further be included in the ED 110 (or T-TRP 170 or NT-TRP 172) .
[0141] Furthermore, communication between different devices / apparatuses in various implementations of this disclosure may refer to direct communication (that is, without the need of forwarding by another device / apparatus) , or may refer to communication (s) between different devices / apparatuses via another device / apparatus (that is, requiring forwarding by another device / apparatus) . Alternatively, such communication (s) may involve one functional unit inside a device / apparatus using another functional unit within the device / apparatus to communicate with another device / apparatus. In other words, phrases such as "sending (or transmitting) information to... (an ED or a base station) " in this disclosure may be understood as a destination endpoint of the information being an ED or a base station, including sending / transmitting information directly or indirectly to an ED or a base station. Similarly, phrases like "receiving information from... (an ED or a base station) " may be understood as a source endpoint of the information being an ED or a base station, including directly or indirectly receiving information from an ED or a base station. Between the source endpoint that sends the information and the destination endpoint, necessary processing such as, but not limited to, format conversion, digital-to-analog conversion, amplification, and filtering may be performed on the information. However, the destination endpoint may understand valid information from the source endpoint. A similar understanding applies to other descriptions in this disclosure without reiterating details already described. In the present disclosure, the terms "send" and "transmit" may be used interchangeably in different implementations of this disclosure. )
[0142] Additional details regarding the EDs 110, the T-TRP 170, and the NT-TRP 172 are known to those of skill in the art. As such, these details are omitted here. In various examples, the present disclosure makes reference to a gNB or a BS as an example T-TRP; it should be understood that this is only exemplary and is not intended to be limiting.
[0143] Efficient CSI compression is helpful for reducing feedback overhead and enabling robust communication in 5G and beyond. Several approaches have been developed over the years, ranging from traditional codebook-based methods to more recent AI / ML-driven techniques, some of which are discussed below. Although the present disclosure may make reference to certain generations of technology it should be understood that this is not intended to be limiting.
[0144] Some example CSI compression methods include the following:
[0145] ● Codebook-based Feedback: This approach involves predefining a set of CSI representations, or codewords, and selecting the codeword that best matches the actual CSI. Examples include:
[0146] ○ Type I / II Codebook: These codebooks, specified in 3GPP standards, represent precoding matrices with different spatial characteristics, such as beamwidths and pointing angles. While offering low complexity, they can suffer from limited accuracy, especially in rich scattering environments.
[0147] ○ eType II Codebook: This enhanced codebook introduces additional parameters to represent the channel’s frequency selectivity, leading to improved accuracy compared to Type I / II but with increased complexity.
[0148] ● Doppler Domain Codebook: This codebook leverages the temporal correlation of the channel to further enhance compression efficiency, particularly for scenarios with high Doppler spread.
[0149] ● Quantization Techniques: These techniques reduce the precision of CSI representation to save bandwidth, employing methods like:
[0150] ● Scalar Quantization (SQ) : Each CSI coefficient is quantized independently using a predefined quantization step size.
[0151] ● Vector Quantization (VQ) : Groups of CSI coefficients are jointly quantized using a codebook of representative vectors. This can achieve higher compression ratios than SQ but requires careful codebook design.
[0152] Some example AI / ML-based CSI compression techniques include the following:
[0153] ● Autoencoders: These neural networks learn to compress and reconstruct CSI data by encoding it into a lower-dimensional latent representation. They offer flexibility and can potentially achieve higher compression ratios than traditional methods.
[0154] FIG. 6 illustrates an example of AI / ML-based CSI compression using autoencoders. As shown in FIG. 6, the encoder side 14 performs dimensionality reduction on input data (e.g., CSI data) , thus compressing the input data from an original size to a low-dimensional representation (which may be referred to as a precoding matrix indicator (PMI) mapping) in the latent space. By “low-dimensional” , it is meant that the input data is reduced in dimensionality (e.g., reduction from data in three-dimensions to data in only two-dimensions) thus obtaining a representation that is smaller in size. In the case where the input data is a form of CSI, the low-dimensional representation obtained from the compression may be referred to as a PMI mapping because the low-dimensional representation can be used to determine (or “map” ) a precoding matrix. The decoder side 16 performs dimensionality increase on the low-dimensional representation, thus reconstructing the data back to the original size.
[0155] Each CSI compression method discussed above presents its own set of challenges and trade-offs, which may include:
[0156] ● Codebook-based methods: Limited accuracy, especially for complex channel environments.
[0157] ● Quantization techniques: Loss of information due to reduced precision.
[0158] ● AI / ML-based methods: High complexity, potential for overfitting, and concerns regarding interoperability.
[0159] Ongoing research aims to address these challenges and improve the efficiency of CSI compression, by focusing on Standardization efforts to facilitate interoperability.
[0160] Using AI / ML to solve the challenges of compressing CSI data is appealing. But data-driven solutions face difficulties that need careful thought. One major issue is the large size and high computing needs of the resulting models. Training on huge datasets often leads to models too big and heavy for devices like phones, where processing power and battery life are limited. This creates a slowdown for real-time use, increasing delays and making it hard to actually use these promising techniques.
[0161] The training process itself is also challenging. Reaching convergence can take a long time and need a lot of computing resources and careful optimization. The massive amount of data makes this worse, leading to very long training times and a frustratingly slow path to optimal performance. Additionally, the quality and nature of the data itself can greatly impact how well the AI / ML models work. Biases or limitations in the training data can cripple the model's ability to generalize, resulting in poor performance when faced with real-world situations that differ from the training set.
[0162] Channel data detected locally in the real-world is often sparser when the training set is more localized. This suggests that smaller, more efficient models can be trained on data specific to a particular cell, site, or scenario, potentially achieving better compression and reducing computing needs. However, this brings new challenges, such as how to ensure rapid convergence with limited, localized data. Another challenge is how to quickly identify if a dataset has flaws that might ruin the entire training process.
[0163] Overcoming these challenges would be of benefit to efficient and insightful AI / ML-based CSI compression. Developing new training methods that work well with smaller datasets and incorporating robust data quality checks are desirable. By tackling these hurdles, the true potential of AI / ML may be unlocked and streamlined, adaptable communication may be achieved for current and future wireless networks.
[0164] FIG. 7 illustrates example training and inference of a generic two-sided AI / ML model for CSI compression. In this example, training data (in this case, precoder matrices V) are stored in a central database 702 (or datacenter) and accessible to both NW and UE vendors, for NW-side vendor training and UE-side vendor training, respectively. Each matrix V may have dimensions NBS x r, for example, where NBS represents the number of antennas at the BS and r represents the rank resulting from singular value decomposition (SVD) used to generate V. The NW-vendor implements an autoencoder model consisting of an encoder model, denoted ENCNW 704, and a decoder model, denoted DECNW 706. The input to the encoder ENCNW 704 is the precoder matrix V and the output of the encoder ENCNW 704 is a low-dimensional representation (which may be referred to as a PMI mapping) denoted S where S=ENC (V; θ) , where θ denotes the parameters of the encoder ENCNW 704. The input to the decoder DECNW 706 is S and the output of the decoder DECNW 706 is V’ where V’=DEC (s; Φ) , where Φ denotes the parameters of the decoder DECNW 706. The goal of the NW-side vendor training is to minimize some loss function (e.g., minimization of a squared error, which may be expressed as min (|V’-V|2) . In a similar manner, UE-side vendor training trains another autoencoder model (consisting of an encoder model, denoted ENCUE 708, and a decoder model, denoted DECUE 710) by minimizing some loss function (e.g., min (|V’-V|2) ) . After training, the trained UE-side encoder ENCUE 708 is used with the trained NW-side decoder DECNW 706 during inference.
[0165] Bringing Physics into CSI Feedback
[0166] This disclosure provides example techniques for CSI feedback compression using AI / ML, inspired by successes in other fields. Examples of the present disclosure may be referred to as "physics-informed machine learning" and may be based on using what is already known about the real world to make AI smarter and more efficient.
[0167] For decades, researchers have been studying CSI and wireless channels. It is known that these channels are governed by the laws of electromagnetism, a well-established branch of physics. This means there is already a lot of prior knowledge about how wireless signals behave and how channels change over time and space.
[0168] However, current AI / ML algorithms for CSI compression do not really take advantage of this knowledge. Instead, they rely solely on the brute force of deep neural networks to learn patterns from data. While this can work, it often leads to large, complex models that require extensive training and massive amounts of data.
[0169] Examples disclosed herein enable existing physical knowledge to be incorporated into the training process. This is where physics-informed machine learning comes in. By incorporating physical laws and principles into the AI models, various benefits may be achieved, which may include one or more of:
[0170] ● Smaller Models: By leveraging prior knowledge, the complexity of AI models may be reduced, making them more suitable for deployment on UEs with limited resources.
[0171] ● Less Data Needed: Physical constraints can guide the learning process and reduce the reliance on large datasets, speeding up training and improving efficiency.
[0172] ● Faster Convergence: With the help of physics, AI models can learn more effectively, leading to faster convergence during training.
[0173] ● Improved Interoperability: If different vendors incorporate the same physical principles into their models, it can promote interoperability and simplify collaboration.
[0174] By bringing physics into the domain of AI / ML two-sided model for CSI compression, examples of this disclosure will develop new AI / ML-based CSI compression techniques that are smarter, more efficient, and easier to implement in real-world networks.
[0175] Physics-Based Regularization
[0176] The integration of physics-based knowledge into AI / ML models holds immense promise for revolutionizing CSI compression techniques. By leveraging the inherent properties of wireless channels, the training process can be guided towards models that are not only more efficient but also promote interoperability between different vendors.
[0177] One such property, particularly relevant in the context of future systems with massive MIMO, is the sparsity of the channel in the angular domain at the gNB. With a large number of antennas deployed at the gNB, the angle of departure (AoD) domain exhibits sparsity, meaning only a few dominant directions contribute significantly to the channel. This characteristic arises from the fixed nature of the gNB antenna array and its specific deployment location, resulting in a stable and persistent sparsity pattern.
[0178] The UE, on the other hand, operates with fewer antennas and dynamic movement. After estimating the downlink MIMO channel H using CSI-RS resources, the UE performs SVD decomposition: H = UΣVH. The V matrix, representing the gNB's beamforming vectors, becomes the focal point for CSI compression. AI / ML methods, commonly employing autoencoders, learn to efficiently compress and reconstruct V by training an encoder (s= ENC (V; θ) ) and a decoder (V’=DEC (s; Φ) ) . The encoder transforms V into a compact representation, also known as the PMI mapping (s, latent layer in autoencoder) , which is then transmitted over the air interface to the gNB. The gNB decoder reconstructs the V matrix from the received PMI mapping (also referred to as a low-dimensional representation) , enabling accurate channel state information for beamforming and link adaptation.
[0179] Note that various physical laws can be integrated into the model training. In this disclosure, a few examples are presented, which may be implemented as alternatives as well as in combination. The examples may help to establish guidelines for the application of physical laws through regularization techniques.
[0180] Example #1:
[0181] Traditional autoencoder training focuses solely on minimizing the square error between the original V matrix and the reconstructed V': min (|V-DEC (ENC (V; θ) ; Φ) |2) . However, by incorporating knowledge of the AoD domain sparsity, a physics-based regularizer may be introduced into the training process. Since each column of V represents a beamforming vector, it will exhibit sparsity in the angular domain, with only a few dominant peaks corresponding to the significant directions of signal propagation. This leads to an additional constraint that minimizes the L1 norm of the DFT of V (min (|V’Wangle|) , where Wangle is the Fourier Transformation matrix) , effectively encouraging the model to learn a representation that reflects the underlying sparsity of the channel in the angular domain. That is, the DFT of the beamforming vector is expected to exhibit sparsity in the spatial frequency domain because the beamforming is expected to be focused on a few specific angles. Thus, by introducing the constraint min (|V’Wangle|) , the model is trained to learn a representation where the beamforming vectors are sparse in the angular domain.
[0182] FIGS. 8A and 8B illustrate an example incorporating physics-based regularization into autoencoder training for CSI compression, based on example #1 described above. This example illustrates how an encoder 808 and decoder 810 of an autoencoder may be trained at a network entity, for example at a UE-vendor or a NW-vendor (e.g., as illustrated in FIG. 7) . The autoencoder, comprising the encoder 808 and decoder 810, may be trained using a physics-based regularization approach using the regularizer described in example #1. In addition to minimizing a loss function, in this case min (|V’-V|2) as discussed previously, the encoder 808 and decoder 810 are trained to minimize a regularization constraint. In this example, the regularizer value to be minimized represents the sparsity of the precoding matrix V in the angular domain. For example, the encoder 808 and decoder 810 may be trained to minimize min (|V’Wangle|) .
[0183] Example #2:
[0184] This disclosure also describes additional regularizers that can further enhance CSI compression. Beyond the sparsity in the angular domain, another aspect to consider is the frequency-dependent behavior of the channel, which manifests as sparsity in the delay domain.
[0185] When considering CSI from multiple subbands or Resource Blocks (RBs) , a correlation exists along the frequency axis. Each subcarrier within a subband or RB corresponds to a specific wavelength, capturing different multipath components with varying delays. This inherently leads to a sparse representation in the delay domain, as only a limited number of propagation paths with distinct delays contribute significantly to the channel.
[0186] To capture this sparsity, a 2D-DFT may be employed. For instance, consider three V matrices measured at different subbands: V1, V2, and V3. These can be concatenated to form a single matrix V = [V1 V2 V3] . Applying a DFT to each column of V transforms it into the angular domain, resulting in V’Wangle. Subsequently, applying a DFT to each row of V’Wangle transforms it into the delay domain, resulting in Wdelay V’Wangle. This final matrix exhibits sparsity, reflecting the limited number of dominant propagation paths with distinct delays. In some examples, a DFT can be applied directly to each row of V to transform it into the delay domain without going through the angular domain transformation. As such, a regularizer may be defined based on a sparsity of the precoding matrix in the delay domain only. This disclosure encompasses regularizers that characterize sparsity in both the angular and delay domains as well as regularizers that characterize sparsity in the delay domain only or in the angular domain only, among other possibilities.
[0187] Therefore, this disclosure introduces an additional physics-based regularizer that minimizes the L1 norm of Wdelay V’Wangle: min (|Wdelay V’Wangle |) .
[0188] FIG. 9 illustrates another example incorporating physics-based regularization into autoencoder training for CSI compression, based on example #2 described above. This example illustrates how an encoder 908 and decoder 910 of an autoencoder may be trained at a network entity, for example at a UE-vendor or a NW-vendor, using a physics-based regularization approach. As described above, the precoding matrix V may be obtained by concatenating subband-specific precoding matrices V1, V2 and V3 (obtained from measurements at subband1, subband2 and subband3, respectively) . The encoder 908 and decoder 910 are trained to minimize loss function, in this case min (|V’-V|2) as discussed previously, and also to minimize a regularization constraint that is in this example min (|Wdelay V’Wangle |) . This regularizer value may represent a characteristic of the precoding matrix in both the angular domain as well as the delay domain, in particular an expected sparsity of the precoding matrix in both the angular and delay domains.
[0189] Example #3:
[0190] In practical wireless systems, CSI rarely undergoes abrupt changes within short time intervals. Instead, it exhibits a gradual evolution, with the CSI at a given time instant closely resembling the CSI of the subsequent instant. This temporal correlation presents an opportunity for further enhancing CSI compression by exploiting the sparsity in the timing domain.
[0191] Similar to the delay-domain approach, this sparsity may be captured by employing a 2D-DFT. Consider a scenario that concatenates V matrices measured at consecutive time instances: V1, V2, and V3, forming a combined matrix V = [V1 V2 V3] . Applying a DFT to each column of V transforms it into the angular domain, resulting in VWangle. Subsequently, applying a DFT to each row of V’Wangle transforms it into the timing domain, yielding WtimeV’Wangle. This matrix exhibits sparsity, reflecting the gradual evolution of CSI over time and the limited number of dominant temporal variations. Consequently, we can introduce a new physics-based regularizer that minimizes the L1 norm of WtimeV’Wangle: min (|WtimeV’Wangle |) . In some examples, applying a DFT directly to each row of V would transform it into the time domain without involving the angular domain. As such, a regularizer may be defined based on a sparsity of the precoding matrix in the time domain only. This disclosure encompasses regularizers that capture sparsity in both the angular and time domains, as well as regularizers that capture sparsity in the time domain only or in the angular domain only, among other possibilities.
[0192] FIG. 10 illustrates another example incorporating physics-based regularization into autoencoder training for CSI compression, based on example #3 described above. This example illustrates how an encoder 1008 and decoder 1010 of an autoencoder may be trained at a network entity, for example at a UE-vendor or a NW-vendor, using a physics-based regularization approach. As described above, the precoding matrix V may be obtained by concatenating time-specific precoding matrices V1, V2 and V3 (obtained from measurements at time1, time2 and time3, respectively) . The encoder 1008 and decoder 1010 are trained to minimize loss function, in this case min(|V’-V|2) as discussed previously, and also to minimize a regularization constraint that is in this example min(|WtimeV’Wangle|) . This regularizer value may represent a characteristic of the precoding matrix in both the angular domain as well as the time domain, in particular an expected sparsity of the precoding matrix in both the angular and time domains.
[0193] Example #4:
[0194] While DFT offers a readily accessible and efficient approach to capturing sparsity in the angular, delay, and timing domains, other transformations present potential alternatives. Discrete Cosine Transform (DCT) or wavelet transforms, for example, could be explored for their ability to represent specific channel characteristics. However, DFT's inherent stability and ease of implementation through Fast Fourier Transform (FFT) algorithms make it a compelling choice. Moreover, its clear and unambiguous representation facilitates communication and understanding, minimizing the risk of confusion during implementation and collaboration.
[0195] DFT stands out due to a crucial property: differentiability. This characteristic is essential for enabling the backpropagation algorithm, the backbone of training deep neural networks. DFT possesses a closed-form expression for its derivative, allowing for efficient calculation of gradients during backpropagation. Alternative transformations, such as DCT or wavelet transforms, may lack closed-form expressions for their derivatives. This necessitates employing numerical differentiation techniques during backpropagation, which can be computationally expensive and introduce numerical instability, potentially hindering the training process. However, such challenges does not exclude the use of DCT or wavelet transforms for defining a physics-based regularizer in accordance with the present disclosure. For example, challenges of using numerical differential techniques for backpropagation during training may be acceptable in certain scenarios.
[0196] Example #5:
[0197] This method is open for higher dimensional transformations, such as 3D-DFT, to capture the combined sparsity across angular, delay, and timing domains simultaneously. This approach holds the potential for even greater compression efficiency, as the resulting representation would be even sparser. However, it comes with trade-offs. Training and inference would require accommodating a large number of V matrices, increasing computational costs and potentially hindering real-time performance. Additionally, specifying the order of operations for higher-dimensional DFTs can become complex, necessitating careful coordination between vendors to ensure interoperability. This disclosure encompasses the use of physics-based regularizers that combine two or more domains, including the use of 2D-DFT, 3D-DFT or even higher dimensions.
[0198] It should be understood that examples #1-#5 described above are exemplary and are not intended to be limiting. The present disclosure is not intended to be limited to the example regularizer definitions described above. Other regularizers may be defined to represent other physical characteristics of the wireless communication network, including other expected characteristics of the precoding matrix not limited to sparsity. Further, examples described herein may be implemented alone or in combination.
[0199] Applying Regularization in Offline Two-Sided Model Training
[0200] The offline training phase of two-sided models for CSI compression typically involves collaboration between UE and network (NW) vendors, where both parties utilize a common dataset to train their respective autoencoder models. Each vendor, however, may employ different model architectures, training algorithms, and computational resources. This raises the question of how to ensure compatibility between the UE's encoder and the network's decoder, particularly when considering the incorporation of physics-based regularization. In the present disclosure, a UE-vendor may refer to a network entity that services one or more UEs, and a NW-vendor may refer to a network entity that services one or more other network entities such as one or more BSs. A UE-vendor may be responsible for developing and training an encoder model, and deploying the trained encoder model to the one or more UEs. A NW-vendor may be responsible for developing and training a decoder mode, and deploying the trained decoder model to the one or more BSs. Typically, the UE-vendor and NW-vendor are separate entities, however this is not intended to be limiting. In some examples, a single network entity may have the role of both UE-vendor and NW-vendor (i.e., servicing both UEs and BSs) . Typically, the UE-vendor is separate from the UE (s) serviced by the UE-vendor, and the NW-vendor is separate from the BS (s) serviced by the NW-vendor, but this is not intended to be limiting. Further, some examples described herein refer to a central entity that may coordinate training among multiple network entities. Such a central entity may coordinate training among only UE-vendors, among only NW-vendors, or among both UE-and NW-vendors. Further, such a central entity may itself be a UE-vendor or NW-vendor in some examples.
[0201] One approach to achieving compatibility is to define and enforce the same regularization constraints during the training process for both the UE and network models. By incorporating the same physical principles and sparsity-inducing regularizers, both vendors can ensure their models learn representations that are consistent and aligned with the underlying channel characteristics. This promotes interoperability and allows the UE's encoder to effectively communicate with the network's decoder, facilitating accurate CSI reconstruction at the gNB.
[0202] The agreement on a common set of regularizers becomes a litmus test for compatibility. If, after training on the same dataset, the resulting models exhibit different regularization constraints, it suggests that they may not be suitable for joint operation. This incompatibility could lead to performance degradation and hinder the effectiveness of the CSI compression scheme.
[0203] Here is an example of how regularization alignment can be used to assess compatibility between UE and network vendor models during offline training.
[0204] Consider the case where both vendors agree to incorporate the AoD sparsity regularizer, min (|VWangle|) , into their respective autoencoder models. This regularizer encourages the models to learn representations that reflect the inherent sparsity in the AoD angular domain of the channel at the gNB.
[0205] One example implementation is Mid-training Check. Although referred to as “mid-training check” , it should be understood that the check described herein may be performed at any time during training and / or this check may be carried out multiple times during training. At some predetermined point (s) during the training process, both vendors can extract the trained weights from their respective encoder and decoder models. Using these weights, they can calculate the value of the AoD sparsity regularizer for a common set of validation data. If the calculated values are close, it suggests that both models are learning similar representations and are likely compatible. Conversely, a significant discrepancy in the regularizer values indicates potential incompatibility, requiring further investigation and possibly adjustments to the training process.
[0206] FIG. 11 illustrates an example method 1100 for performing a mid-training check for regularizer compatibility between a UE-vendor and NW-vendor (which may be respective network entities) . Each vendor performs training of a respective model. For example, training of an encoder model may be performed by the UE-vendor and training of a decoder model may be performed by the NW-vendor (it should be understood that the encoder model may be trained as part of an autoencoder at the UE-vendor, and similarly the decoder model may be trained as part of an autoencoder at the NW-vendor) .
[0207] At operation 1102, training of the encoder model and the decoder model may be performed up to a predetermined point (e.g., up to a predetermined number of training rounds, up to a predetermined amount of training time, up to a predetermined amount of training data, etc. ) . In some examples, training at the UE-vendor and / or the NW-vendor may be performed until a signal to perform a mid-training check is received (e.g., the UE-vendor may send a signal to the NW-vendor or vice versa, or some other network entity may send a signal to both the UE-vendor and the NW-vendor to perform a mid-training check) .
[0208] To perform the mid-training check, the trained weights of the encoder model (e.g., the weights of the encoder model in a partly-trained autoencoder) are extracted by the UE-vendor at operation 1104 and the trained weights of the decoder model (e.g., the weights of the decoder model in a partly-trained autoencoder) are extracted by the NW-vendor at operation 1108.
[0209] At operation 1112, a common set of validation data may be exchanged between the UE-vendor and the NW-vendor. In some examples, instead of exchanging validation data between the UE-vendor and NW-vendor, a common set of validation data may be provided to both the UE-vendor and the NW-vendor by some other network entity (e.g., a central datacenter) . The validation data may, for example, a set of precoding matrices V stored at a central datacenter.
[0210] At operation 1106, the UE-vendor calculates a regularizer score (in this example, based on the AOD sparsity regularizer) from the validation data. For example, the partly-trained autoencoder may be used to encode each precoding matrix V into a low-dimensional representation (also referred to as a PMI mapping) , then decode the low-dimensional representation into the reconstructed precoding matrix V’. Then the regularizer value min(|V’Wangle|) may be calculated using the reconstructed precoding matrix V’. The regularizer values calculated over the set of validation data may be combined (e.g., averaged) into a regularizer score. In a similar manner, at operation 1110, the NW-vendor calculates a regularizer score from the same set of validation data and using the same regularizer definition (e.g., using the AOD sparsity regularizer described above) .
[0211] The regularizer scores calculated by the UE-vendor and the NW-vendor may be exchanged with each other and / or communicated to another network entity (e.g., a central datacenter) . At operation 1114, a determination is made whether the regularizer scores from the UE-vendor and the NW-vendor are similar to each other (e.g., the regularizer scores may be considered to be similar to each other if they are within 10%of each other, or if they are within a predefined margin of each other) . If the regularizer scores are similar, then at operation 1116 it is determined that the partly-trained encoder model at the UE-vendor and the partly-trained decoder model at the NW-vendor are compatible with each other. The training process at the UE-vendor and the NW-vendor may continue.
[0212] If the regularizer scores are dissimilar, then at operation 1118 it is determined that the partly-trained encoder model at the UE-vendor and the partly-trained decoder model at the NW-vendor are potentially incompatible with each other. At operation 1120, this incompatibility may be further investigated and / or the training process may be adjusted (e.g., by adjusting the number of training rounds, adjusting hyperparameters, etc. ) .
[0213] The mid-training check described above may be performed repeatedly throughout the training process (e.g., at regular intervals, such as at every 100 rounds of training) . If the mid-training check repeatedly determines a potential incompatibility, the training may be aborted.
[0214] Another example implementation is End-of-training Evaluation. After completing the training process, both vendors can perform a more thorough evaluation of their models using a common test dataset. Consistent performance across these metrics, particularly a low value for the sparsity regularizer, indicates strong compatibility and suggests that the UE encoder and network decoder can effectively work together in real-world deployments. The end-of-training evaluation may be performed in a manner similar to the mid-training check described above. The end-of-training evaluation may be performed in addition to the mid-training check (e.g., even after the encoder and decoder models are determined to be compatible during a mid-training check, this may be confirmed with an end-of-training evaluation) or instead of the mid-training check (e.g., the encoder and decoder models may not be checked for compatibility until both models have completed training) .
[0215] If the mid-training check reveals a significant difference in the regularizer values, both vendors may need to continue training their models for a longer duration or with adjusted hyperparameters to improve convergence and alignment.
[0216] In cases where compatibility issues persist, vendors may need to explore modifications to their model architectures or training algorithms to better align with the agreed-upon regularization constraints.
[0217] If the issue stems from limitations in the training dataset, vendors may consider augmenting the dataset with additional samples that better represent the diversity of real-world channel conditions.
[0218] Regularization in the Inference Phase
[0219] The role of physics-based regularization extends beyond the training phase, acting as a safeguard for reliable CSI reconstruction during inference. Once the UE vendor deploys the encoder on the UE and the network vendor deploys the decoder at the base station, the agreed-upon regularizer continues to play a crucial role in ensuring accurate and consistent CSI feedback and reconstruction.
[0220] After the UE encoder compresses the CSI data (V matrix or precoding matrix) , the UE calculates the regularizer value based on the specified domain or combination of domains. In some examples, this value is then included (e.g., included as an indicator) as part of the CSI feedback transmitted over the air interface to the gNB. At the gNB, the decoder reconstructs the CSI from the received PMI mapping (also referred to as a low-dimensional representation) and independently calculates the regularizer value (also referred to as the indicator) based on the same predefined rules.
[0221] The gNB then compares the received regularizer value (also referred to as the received indicator) from the UE with the locally calculated value (also referred to as the reconstructed indicator) . If the values are similar, it provides a strong indication that the CSI reconstruction was successful and the decoded representation accurately reflects the underlying channel characteristics. Conversely, a significant discrepancy between the values suggests a potential issue with the CSI feedback or reconstruction process. This could be caused by transmission errors, model mismatch, or changes in the channel environment that deviate from the assumptions used during training.
[0222] Therefore, standards may explicitly define the specifications for calculating and reporting the regularizer value during inference. This includes specifying the relevant domains or combination of domains (e.g., angular, delay, timing) , the quantization precision for the regularizer value, and the order of operations for multi-domain regularizers to ensure consistency between the UE and gNB calculations.
[0223] Here is an example: consider a system where both the UE and NW vendors have agreed to incorporate sparsity regularization in both the AoD and delay domains during training and inference. They have chosen to utilize the following loss function, which incorporates a physics-informed regularizer:
[0224] α × min (|V-V’|2) + β × min (|WdelayV’Wangle|)
[0225] where: α and β are weighting factors that determine the relative importance of the loss term and the regularization term; V is the original CSI matrix containing beamforming vectors for multiple subbands; V’ is the reconstructed CSI matrix after decoding the PMI mapping (also referred to as the low-dimensional representation) ; Wangle is the Fourier matrix for transforming the data into the angular domain; Wdelay is the Fourier matrix for transforming the data into the delay domain.
[0226] The UE measures the downlink channel and obtains the V matrix for multiple subbands. The UE encoder compresses V into a PMI mapping (also referred to as a low-dimensional representation) , which is then transmitted to the gNB. Before transmission, the UE calculates the regularizer value based on the agreed-upon formula: | WdelayV’Wangle |. This value is quantized and included as part of the CSI feedback (e.g., included as an indicator in the CSI feedback) .
[0227] The gNB receives the PMI mapping (also referred to as the low-dimensional representation) and the quantized regularizer value (also referred to as the quantized indicator) . The network decoder reconstructs the V' matrix from the PMI mapping (also referred to as the low-dimensional representation) . The gNB then independently calculates the regularizer value (also referred to as the reconstructed indicator) based on the same formula (where V’ is used in place of V in the above formula) and compares it with the received value from the UE.
[0228] If the received and calculated regularizer values (also referred to as the received and reconstructed indicators) are within a predefined tolerance range, the gNB considers the CSI reconstruction successful and proceeds with beamforming and link adaptation based on the reconstructed V' matrix.
[0229] If a significant discrepancy exists between the values, it indicates a potential issue with the CSI feedback or reconstruction. The gNB may take corrective actions, such as requesting retransmission of the CSI feedback, switching to a different model, or falling back to a legacy codebook-based approach.
[0230] FIG. 12 illustrates an example method 1200 for using regularization in the inference phase for CSI feedback. In this example, the UE implements an encoder model (e.g., implemented using a trained neural network) that has been trained by the UE-vendor and the gNB implements a decoder model (e.g., implemented using another trained neural network) that has been trained by the NW-vendor.
[0231] At operation 1202, the UE measures a downlink channel and obtains the precoding matrix V for multiple subbands (e.g., by measuring a respective precoding matrix V1, V2, etc. for respective subband1, subband2, etc. and concatenating the measured matrices together into a single precoding matrix V) .
[0232] At operation 1204, the UE compresses the precoding matrix V into a low-dimensional representation that is also referred to as a PMI mapping (or PMI map) , using the encoder model. At operation 1206, the UE also calculates an indicator based on V using the regularizer definition that is agreed upon between the UE and the gNB. In this example, the indicator represents a regularizer value calculated using the formula |WdelayVWangle|. In some examples, the regularizer value may be directly used as the indicator; in other examples, the indicator may be a representation of the regularizer value (e.g., the indicator may be a rounded estimate of the regularizer value, or may indicate a numerical range of the regularizer value, among other possibilities) . Optionally, at operation 1208, the UE may quantize the regularizer value (e.g., using a predetermined quantization method) to obtain a quantized indicator.
[0233] At operation 1210, the UE transmits CSI feedback to the gNB. In this example, the CSI feedback includes the quantized regularizer value (which is an indicator) and also the PMI mapping (which is a low-dimensional representation of the precoding matrix V) .
[0234] At operation 1212, the gNB uses the decoder model to reconstruct V’ from the received PMI mapping (also referred to as the low-dimensional representation) . At operation 1214, the gNB calculates a reconstructed regularizer value (also referred to as the reconstructed indicator) using the reconstructed precoding matrix V’, according to the agreed upon regularizer definition, in this case using the formula |WdelayV’Wangle| (where the reconstructed precoding matrix V’ has been used in place of the original precoding matrix V) .
[0235] At operation 1216, the gNB determines whether the reconstructed regularizer value (or reconstructed indicator) , which is calculated using the reconstructed precoding matrix V’, is similar to the received regularizer value (or received indicator) that was received from the CSI feedback. The reconstructed regularizer value (or reconstructed indicator) may be considered to be similar to the received regularizer value (or received indicator) if the difference between the two values is within a threshold (e.g., within 10%of each other or within a predetermined margin of each other) .
[0236] If the reconstructed regularizer value (or reconstructed indicator) and the received regularizer value (or received indicator) are not similar, then at optional operation 1218 corrective action may be taken by the gNB. Corrective action may include one or more of: requesting retransmission, switching to a different decoder model, or falling back to using a legacy codebook-based approach to CSI feedback, among other possibilities.
[0237] If the reconstructed regularizer value (or reconstructed indicator) and the received regularizer value (or received indicator) are similar to each other, then at optional operation 1220 the gNB may determine that reconstruction of the CSI feedback is successful and that the reconstructed precoding matrix V’ is correct. At optional operation 1222 the reconstructed precoding matrix V’ may be used to perform beamforming and / or link adaption operations.
[0238] The inclusion of the regularizer value in the CSI feedback allows the gNB to detect potential errors in the reconstruction process, enhancing the overall robustness of the CSI compression scheme. By leveraging sparsity in both the angular and delay domains, the models can achieve better compression ratios, leading to reduced feedback overhead and improved spectral efficiency. The agreement on a common regularization scheme promotes interoperability between UE and network vendors, facilitating collaboration and standardization efforts.
[0239] The examples described herein may demonstrate how physics-based regularization can be effectively applied to improve the reliability and efficiency of CSI compression, contributing to the advancement of next-generation and future wireless communication systems.
[0240] FIG. 13 is a flowchart illustrating an example method 1300 for performing CSI feedback using a first model to compress the information in the CSI feedback. The first model may be referred to as an encoder model. The first model may be any machine learning model designed to perform the task of CSI compression as disclosed herein. The method 1300 may be implemented at a UE or other apparatus configured to provide CSI feedback, for example. The method 1300 may be performed by a functional unit, such as a processor or any other suitable unit. The first model may be implemented using a trained neural network (e.g., trained by a UE-vendor and deployed to the UE) , where a physics-based regularizer as disclosed herein was used during training.
[0241] At an operation 1302, the UE may receive parameters of a first model (also referred to as an encoder model) . For example, the UE may receive the parameters from a network entity such as a UE-vendor. The UE-vendor may use any suitable signaling to deploy the first model (which may be implemented by a trained neural network) to the UE. For example, a control message (e.g., dedicated RRC message) may contain information about the first model, such as an identifier of the first model, an indication of the size of the first model (e.g., number of model parameters) and / or method by which the first model will be deployed to the UE (e.g., via control messages, over a data channel, etc. ) . The parameters of the first model (e.g., the parameters or weights of the trained neural network) may be deployed to the UE via control messages (e.g., using RRC signaling) where the model is relatively small (e.g., smaller number of parameters) ; a data channel (e.g., user plane data channel) may be suitable where the model is larger (e.g., larger number of parameters) . In some examples, parameters of the first model may be sent in segments, with appropriate error detection and / or correction mechanisms to help ensure more reliable deployment. Additionally, synchronization signals (e.g., from the UE-vendor or some other network entity) may be used to ensure that the first model deployed at the UE is aligned and up-to-date with a second model (also referred to as a decoder model) deployed at a network entity such as a base station.
[0242] As disclosed herein, the first model may have been trained using a regularizer that represents a characteristic of a precoding matrix based on a physical property of the wireless communication system. For example, the regularizer may be defined based on an expected or known sparsity of the precoding matrix in an angular domain, a delay domain and / or a time domain.
[0243] At an operation 1304, a precoding matrix is obtained. In some examples, the UE may obtain the precoding matrix in response to a control message. The control message may indicate relevant parameters such as an indication of the CSI-RS resources the UE should use for channel estimation, an indication of the duration or number of time instances over which the UE should perform channel estimation, etc. In some examples, the precoding matrix may be obtained by taking measurements at multiple different subbands or RBs and concatenating the measurements to form a single precoding matrix. In some examples, the precoding matrix may be obtained by taking measurements at multiple different time instances and concatenating the measurements to form a single precoding matrix.
[0244] At an operation 1306, the precoding matrix is compressed into a low-dimensional representation, which may also be referred to as a PMI mapping, using the first model.
[0245] At an operation 1308, CSI feedback is transmitted (e.g., to a base station or other network entity) . The transmitted CSI feedback includes the low-dimensional representation obtained at the operation 1306 and also an indicator based on the precoding matrix. The indicator represents a characteristic of the precoding matrix (e.g., an expected or known sparsity of the precoding matrix in the angular domain, delay domain and / or time domain) , based on a physical property of the wireless communication system. For example, the indicator may be a regularizer value computed from the precoding matrix, or may represent a regularizer value computed from the precoding matrix. The indicator may represent the regularizer value (e.g., may represent an approximation of the regularizer value) , may be directly proportional to the regularizer value, or may be a value equal to the regularizer value, among other possibilities. In some examples, the indicator may be referred to as a regularizer value, a value (e.g., a value derived from the precoding matrix and / or based on a physical property of the wireless communication system) , information (e.g., information representing a characteristic of the precoding matrix and / or the wireless communication system) , etc. Performing the operation 1308 may include performing operation 1310 and / or operation 1312.
[0246] At an operation 1310, the indicator is determined from the precoding matrix. For example, the indicator may represent a regularizer value that may be computed according to a defined formula (e.g., defined in a standard) based on an expected characteristic of the precoding matrix, such as an expected sparsity of the precoding matrix in the angular domain, delay domain and / or time domain, as discussed previously. It should be noted that the definition of the indicator (e.g., including the defined formula for computing the regularizer value) should be known and common to both the UE and the BS.
[0247] In some examples, multiple indicators may be computed, where each indicator represents a respective different characteristic of the precoding matrix (e.g., representing regularizer values computed according to a respective different formula) . For example, a first indicator may represent a regularizer value that is computed using a formula representing the sparsity characteristic of the precoding matrix in the angular domain (e.g., using the formula |VWangle|) and a second indicator value may represent another regularizer value that is computed using another formula representing the sparsity characteristic of the precoding matrix in the time domain (e.g., using the formula |WdelayVWangle|) . Optionally, if multiple indicators are computed, the multiple indicators may be combined into a single combined indicator. This combination can be achieved using a weighted average, where the weights may be predefined (e.g., according to a standard) . Other, possibly more complex combination techniques may be used, such as non-linear functions or conditional combinations based on specific channel conditions, among other possibilities. For indicators representing multi-dimensional regularizers involving transformations like DFT, the order of operations for these transformations should also be defined.
[0248] Optionally, at an operation 1312, the indicator (or multiple indicators) may be quantized. The quantization may be performed in accordance with quantization parameters defined in a control message, in some examples. In some examples, quantization methods may be defined in a standard.
[0249] Operation 1308 (which may include operations 1310 and / or 1312) may thus be performed to transmit CSI feedback including both the low-dimensional representation of the precoding matrix as well as the indicator representing a characteristic of the precoding matrix based on a physical property of the wireless communication system. Inclusion of the indicator as part of the CSI feedback may help to maintain accuracy of the transmitted information, and may help to ensure compatibility of first and second models. The indicator may be useful to detect incompatibility, enabling corrective actions to be taken, as discussed below with respect to FIG. 14.
[0250] FIG. 14 is a flowchart illustrating an example method 1400 for receiving CSI feedback and decompressing information in the CSI feedback using a second model. The second model may be referred to as a decoder model. The second model may be any machine learning model designed to perform the task of CSI decompression (or reconstruction) as disclosed herein. The method 1400 may be implemented at a BS (e.g., a gNB) or other network entity configured to receive CSI feedback, for example. The method 1400 may be performed by a functional unit, such as a processor or any other suitable unit. The second model may be implemented using a trained neural network (e.g., trained by a NW-vendor and deployed to the BS) , where a physics-based regularizer as disclosed herein was used during training.
[0251] At an operation 1402, the BS may receive parameters of a second model (also referred to as a decoder model) . For example, the BS may receive the parameters from a network entity such as a NW-vendor. The NW-vendor may use any suitable signaling to deploy the second model (which may be implemented by a trained neural network) to the BS. For example, a control message (e.g., dedicated RRC message) may contain information about the second model, such as an identifier of the second model, an indication of the size of the second model (e.g., number of model parameters) and / or method by which the second model will be deployed to the BS (e.g., via control messages, over a data channel, etc. ) . The parameters of the second model (e.g., the parameters or weights of the trained neural network) may be deployed to the BS via control messages (e.g., using RRC signaling) where the model is relatively small (e.g., smaller number of parameters) ; a data channel (e.g., user plane data channel) may be suitable where the model is larger (e.g., larger number of parameters) . In some examples, parameters of the second model may be sent in segments, with appropriate error detection and / or correction mechanisms to help ensure more reliable deployment. Additionally, synchronization signals (e.g., from the NW-vendor or some other network entity) may be used to ensure that the second model deployed at the BS is aligned and up-to-date with a first model (also referred to as an encoder model) deployed at a UE that is served by the BS.
[0252] As disclosed herein, the second model may have been trained using a regularizer that represents a characteristic of a precoding matrix based on a physical property of the wireless communication system. For example, the regularizer may be defined based on an expected or known sparsity of the precoding matrix in an angular domain, a delay domain and / or a time domain.
[0253] At an operation 1404, CSI feedback is received (e.g., from a UE) . The CSI feedback includes a low-dimensional representation of a corresponding precoding matrix (also referred to as a PMI mapping) and also includes an indicator that represents a characteristic of the corresponding precoding matrix (e.g., an expected or known sparsity of the corresponding precoding matrix in the angular domain, delay domain and / or time domain) , for example based on a physical property of the wireless communication system. As discussed previously, the indicator may represent a regularizer value that is based on a physical property of the wireless communication system.
[0254] At an operation 1406, the low-dimensional representation is decompressed using the second model to obtain a reconstructed precoding matrix.
[0255] At an operation 1408, consistency verification (also referred to as error detection) is performed. This may include performing an operation 1410, an operation 1412 and / or an operation 1414. As described below, consistency verification may be carried out based on the indicator received in the CSI feedback. Thus, the inclusion of the indicator in the CSI feedback provides a mechanism for identifying errors and discrepancies, and may indicate a possible incompatibility between the first model (deployed at a UE) and second model (deployed at a BS) . This may enable corrective action to be taken. The protocols to be used for verification and correction, including defined thresholds for detection of inconsistencies and possible corrective actions, may be defined in a standard.
[0256] At an operation 1410, consistency verification may be performed by comparing a reconstructed indicator, based on the reconstructed precoding matrix, with the received indicator (received from the CSI feedback) . The reconstructed indicator may represent a regularizer value computed according to a defined formula (e.g., defined in a standard) based on an expected characteristic of the precoding matrix, such as an expected sparsity of the precoding matrix in the angular domain, delay domain and / or time domain, as discussed previously. It should be noted that the definition of the indicator (e.g., including a defined formula for the regularizer value) should be known and common to both the UE and the BS (e.g., defined in a standard common to both the UE and the BS) .
[0257] In some examples, multiple indicators may be received in the CSI feedback and multiple reconstructed indicators may be computed, where each indicator represents a respective different characteristic of the precoding matrix (e.g., based on different regularizer values computed according to a respective different formula) . For example, a first reconstructed indicator may represent a regularizer value computed representing the sparsity characteristic of the precoding matrix in the angular domain (e.g., using the formula |VWangle|) and a second reconstructed indicator may represent a different regularizer value computed representing the sparsity characteristic of the precoding matrix in the time domain (e.g., using the formula |WdelayVWangle|) . Optionally, if multiple reconstructed indicators are computed, the multiple reconstructed indicators may be combined into a combined reconstructed indicator, for example using a weighted average where the weights may be predefined (e.g., according to a standard) or using other combination techniques as discussed previously with respect to FIG. 13. The definitions of the multiple indicators and the optional techniques and / or weights for combining the multiple indicators should be known and common to both the UE and the BS (e.g., defined in a standard common to both the UE and the BS) .
[0258] The reconstructed indicator may be compared with the received indicator to determine a difference between reconstructed and received indicators. If there are multiple received indicators, then corresponding multiple reconstructed indicators may be computed and each reconstructed indicator may be compared with the corresponding received indicator . For example, if a first indicator is received that represents sparsity of a corresponding precoding matrix in the angular domain (e.g., representing a regularizer value computed according to a first formula) and a second indicator is received that represents sparsity of a corresponding precoding matrix in the delay domain (e.g., representing a regularizer value computed according to a second formula) , then the first received indicator may be compared with a corresponding first reconstructed indicator based on the first formula and the second received indicator may be compared with a corresponding second reconstructed indicator based on the second formula. In some examples, the multiple received indicators may be combined (e.g., using predefined weights and / or techniques) into a single combined indicator and the multiple reconstructed indicators may be similarly combined (e.g., using the same predefined weights and / or techniques) into a combined reconstructed indicator for comparison.
[0259] If the difference exceeds a threshold (e.g., difference exceeding 10%or exceeding a defined margin of error) , then at an operation 1412 an inconsistency may be detected and corrective action may be performed. Corrective action may include, for example, requesting retransmission of the CSI feedback, switching to a different second model, or falling back to use a legacy codebook-based approach to CSI feedback, among other possibilities. If the difference is within the threshold (e.g., difference is within 10%or is within a defined margin of error) , then at an operation 1414 it may be detected that the reconstruction of the precoding matrix was successful. The reconstructed precoding matrix may then be used for performing link operations, such as beamforming and link adaptation, among other possibilities.
[0260] In some examples, at the operation 1414, the reconstructed indicator may be used to configure channel estimation operations. For example, the specific CSI-RS resources and the duration of channel estimation to be used for channel estimation by the UE may be dynamically adjusted by the BS based on the reconstructed indicator. For example, if the indicator indicates a high degree of sparsity in a particular angular or delay domain, the BS may configure the UE to estimate the channel using a reduced set of CSI-RS resources focused on the dominant directions or delays, thereby reducing the overhead associated with channel estimation.
[0261] Optionally, at an operation 1416, the BS may transmit (e.g., to a network entity such as the NW-vendor or other network entity) an indication of the performance of the second model based on the comparison performed at the operation 1410. For example, an indication of the performance may be transmitted to a central entity responsible for monitoring the performance of trained models across multiple BSs. In some examples, repeated evaluation of model performance during inference may enable corrective actions to be carried out (e.g., retraining of the model) by the central entity or the BS in real-time during inference if performance falls below a predefined threshold. The frequency and methodology for performing such performance checks during inference may be defined by a standard, which may further define how performance data is collected, transmitted and analyzed, for example.
[0262] In some examples, the BS may not transmit the indication of performance at the operation 1416. Instead, the BS may monitor the performance of the trained second model in real time, either locally or by transmitting relevant data to a central entity. The BS may take action (which may include transmitting an indication of performance at the operation 1416) if the performance of the trained model falls below a predefined threshold.
[0263] FIG. 15 is a flowchart illustrating an example method 1500 for training a first model (also referred to as an encoder model) for compression of CSI feedback or training a second model (also referred to as a decoder model) decompression of CSI feedback. The method 1500 may be implemented at a network entity such as a UE-vendor that services UEs, a NW-vendor that services BSs or other network entity, for example. The method 1500 may be performed by a functional unit, such as a processor or any other suitable unit. The first and second models may be trained as part of an autoencoder, and may each be implemented using a respective neural network (e.g., trained by a NW-vendor and deployed to the BS) .
[0264] At an operation 1502, a training dataset is obtained. The training dataset may, for example, be sampled from a database of precoding matrices V stored at a central datacenter. The training dataset may be obtained from the central datacenter in response to a request for training data.
[0265] At an operation 1504, the model is trained to perform a precoding matrix compression task (to obtain the first model) or to perform a precoding matrix decompression task (to obtain the second model) . It should be noted that because the first or second model may be trained as part of an autoencoder that includes both the first and second models (e.g., as illustrated in FIG. 7) , training a first model and training a second model may be performed in the same manner. To train the model (regardless of first or second model) , the aim is to minimize the loss function, which may be based on the error between an output of the autoencoder and the ground-truth precoding matrix in the training data.
[0266] As discussed previously with respect to FIG. 7, the goal of training may be to minimize a loss function defined as |V’-V|2, where V is the ground-truth precoding matrix in the training data and V’ is the reconstructed precoding matrix outputted by the autoencoder.
[0267] In addition to training based on the loss function, the goal of training is also to minimize a regularizer value that represents a characteristic of the precoding matrix based on a physical property of the wireless communication system. For example, a regularizer value may be defined by a formula (e.g., defined in a standard) that is based on an expected or known sparsity of the precoding matrix in an angular domain, delay domain and / or time domain, as disclosed herein. The regularizer value may be defined to be differentiable (E. g., based on a DFT of the precoding matrix) , which may enable easier and / or more efficient training (e.g., to enable backpropagation without requiring numerical approximations) .
[0268] In some examples, multiple regularizer values may be minimized during training, where each regularizer value may represent a respective characteristic of the precoding matrix. If there are multiple regularizer values, a set of predetermined weights may be used to combine (e.g., using a weighted average) the multiple regularizer values into a single combined regularizer value to be minimized.
[0269] In some examples, validation may be performed during training (also referred to as a mid-training check or mid-training validation) . At an operation 1506, during training (but prior to a final round of training) , validation of the partly-trained model may be performed. In some examples, validation during training may be performed in response to a received signal (e.g., a trigger signal may be received from another network entity, such as a central datacenter, a UE-vendor, a NW-vendor or other network entity) or may be performed in response to reaching a predefined condition (e.g., performed every 100 rounds of training) . This may involve performing operations 1508, 1510, 1512 and / or 1514.
[0270] The mid-training validation may be performed across multiple network entities that are concurrently training respective models. In some examples, a central entity (e.g., central datacenter) may be responsible for scheduling or coordinating mid-training validation across the multiple network entities. In other examples, there may not be any centralized scheduling or control of validation and instead the multiple network entities may communicate among themselves to determine when to perform mid-training validation.
[0271] In a centralized example, if validation during training is in response to a control signal from a central entity that is responsible for coordination among multiple network entities, the central entity may periodically or at intervals transmit a control signal or a common validation dataset to the multiple network entities to trigger the mid-training validation. The central entity may also be responsible for providing a common validation dataset to the multiple network entities. The central entity may receive regularizer scores (e.g., determined from the mid-training validation as described below) from the multiple network entities and may adjust the periodicity or interval of mid-training validation based on the returned regularizer scores and / or other predefined criteria. For example, if the regularizer scores from multiple network entities are consistently close to a target value or fall within a predefined acceptable range (e.g., indicating good model convergence and compatibility) , the central entity may increase the interval between validations, reducing the frequency of checks which may help to save resources. Conversely, if the regularizer scores show significant discrepancies or deviate from the target, the central entity may shorten the interval, triggering more frequent validations to closely monitor the training process and identify potential issues early on.
[0272] In a decentralized example, validation may be performed during training based on a consensus among two or more network entities that are concurrently performing training of respective models. For example, there may be protocols that enable network entities to vote on the periodicity of the mid-training validation in absence of a central entity that schedules or coordinates the validation. Each network entity may broadcast a selected periodicity for performing mid-training validation from among several predefined options. Then each network entity may tally all broadcasted votes and determine the majority preference. A protocol to be used by network entities for this voting may be defined in a standard, for example including predefined options for periodicity and the mechanisms to be used for broadcasting and tallying votes. Network entities may exchange validation data with each other to ensure that each network entity is using a common validation dataset.
[0273] At an operation 1508, a validation dataset may be obtained. As noted above, the validation dataset may be obtained from a central entity or from other network entities. In some examples, the validation dataset may be received without being requested, and the receipt of the validation dataset may be the signal that validation is to be performed (e.g., the central entity may provide the validation dataset to the network entity to cause the network entity to perform mid-training validation) . The validation dataset may be common to multiple network entities and may be used to ensure that models trained by different network entities meet a consistent quality level and / or level of interoperability. Together with obtaining the validation dataset, parameters for computing the regularizer value may also be obtained (e.g., identification of formula to use for computing the regularizer value, a set of weights for combining multiple regularizer values into a combined regularizer value, etc. ) . In some examples, the parameters for computing the regularizer value may be defined in a standard and may not need to be obtained with the validation dataset. In some examples, the validation dataset may be obtained with an associated reference score (e.g., a ground-truth reference score associated with the validation dataset, which may be used for later comparison) and / or an indication of the method for comparison against the reference score.
[0274] At an operation 1510, a regularizer score may be obtained. The regularizer score may be obtained by obtaining a regularizer value for each precoding matrix in the validation dataset and combining the regularizer scores into a single regularizer score. For example, each precoding matrix in the validation dataset may be inputted to the partly-trained autoencoder and the reconstructed precoding matrix obtained at the output of the partly-trained autoencoder may be used to compute a regularizer value. The regularizer values obtained from all the precoding matrices in the validation dataset may then be combined (e.g., averaged) to obtain a single regularizer score. Other techniques for combining the regularizer values may be used.
[0275] In examples where the validation dataset is obtained together with a reference score associated with the validation dataset, at an operation 1512, the regularizer score may be compared with a reference score. This comparison may enable a determination whether the training process requires adjusting. For example, if the regularizer score differs from the reference score by more than a defined amount (e.g., greater than 10%difference, or greater than a defined margin of error) , this may indicate that the training process requires adjusting (e.g., more rounds of training are required, hyperparameters need to be adjusted, more varied training data is required, etc. ) .
[0276] At an operation 1514, the regularizer score may be transmitted. If a comparison with a reference score was performed at the operation 1512, the result of the comparison may also be transmitted.
[0277] In a centralized example, the regularizer score may be transmitted to a central entity. The central entity may collect regularizer scores from all network entities concurrently undergoing training. The central entity may determine that training at a particular network entity (or all network entities) should terminate early. This decision may be based on various factors, such as the regularizer scores consistently meeting a predefined target threshold (e.g., indicating good convergence and compatibility) , or persistent discrepancies in regularizer scores (e.g., suggesting model incompatibility) . The central entity may also employ predefined criteria, such as time or resource limits, to trigger early termination. If the central entity decides to terminate training early, the central entity may transmit communications to inform all participating network entities of this decision.
[0278] In a decentralized example, the regularizer score may be shared with other network entities that are concurrently undergoing training. The network entities may reach a consensus among themselves whether training can be terminated early. This can be achieved through a voting process, such as where network entities broadcast messages indicating a vote to terminate after each instance of mid-training validation if their respective local regularizer score meets predefined criteria, such as exceeding a target threshold or aligning with scores from other entities. Training may be terminated if all network entities or a predefined majority vote to terminate.
[0279] The operation 1506 may be performed multiple times during training (e.g., at regular intervals or in response to an external signal) .
[0280] Training of the model may be terminated when a termination condition is reached, such as reaching a maximum number of training rounds or receiving a signal (from a central entity or from another network entity that is also concurrently undergoing training) to terminal training. As described above, in some examples training may be terminated early based validation results from on the mid-training validation.
[0281] In some examples, evaluation of the trained model may additionally or alternatively be performed after completion of training (i.e., after a final round of training has been completed) . The evaluation (or testing) that is performed on the trained model (also referred to as end-of-training evaluation) may be similar to the validation performed during training (e.g., performed at the operation 1506) . However, the end-of-training evaluation is typically more comprehensive, aiming to provide a final, rigorous assessment of the trained model's performance and suitability for deployment. It often involves a larger and more diverse test dataset, including data representing challenging or extreme scenarios typically not encountered during training. The evaluation also may encompass a broader set of metrics, such as reconstruction accuracy on the held-out test data, generalization ability to unseen channel conditions, computational efficiency, and robustness to noise and interference, among other possibilities. The evaluation of the trained model may involve performing operations 1518, 1520, 1522 and / or 1524.
[0282] Similar to the validation performed during training, the evaluation performed after completion of training may be performed across multiple network entities that have completed training on respective models. In some examples, a central entity (e.g., central datacenter) may be responsible for scheduling or coordinating end-of-training evaluation across the multiple network entities. In other examples, similar to the decentralized approach for mid-training validation, there may not be any centralized scheduling or control of evaluation, and instead, the multiple network entities may communicate among themselves to perform the end-of-training evaluation.
[0283] In a centralized example, termination of training may be controlled by a central entity. Alternatively, the network entity that is performing the training may terminate training if an internally determined termination condition is met (e.g., a maximum number of rounds of training is reached, based on the model converging or based on the results of mid-training validation as described above) and notify the central entity. End-of-training evaluation may be automatically triggered by the central entity after termination of training. The central entity may synchronize end-of-training evaluation across multiple network entities. The central entity may also be responsible for providing a common test dataset to the network entities. The central entity may receive regularizer scores (e.g., determined from the end-of-training evaluation as described below) from the multiple network entities and may approve or reject a trained model for deployment based on the corresponding regularizer score resulting from the end-of-training evaluation.
[0284] In a decentralized example, each network entity may determine whether to terminate training (e.g., based on a termination condition such as reaching a maximum number of rounds of training, based on the model converging, or based on the results of mid-training validation as described above) . Similar to the mid-training validation described above, multiple network entities may exchange test data with each other to ensure that each network entity is using a common test dataset for the end-of-training evaluation.
[0285] At an operation 1518, a test dataset may be obtained. The test dataset may be obtained in a similar manner to obtaining the validation dataset at the operation 1508. In some examples, the test dataset may be received without being requested, and the receipt of the test dataset may be the signal that end-of-training evaluation is to be performed (e.g., the test dataset may be automatically provided by the central entity after training is terminated) . The test dataset may be common to multiple network entities, and may be used to ensure that models trained by different network entities meet a consistent quality level and / or level of interoperability. Together with obtaining the test dataset, parameters for computing the regularizer value may also be obtained (e.g., identification of formula to use for computing the regularizer value, a set of weights for combining multiple regularizer values into a combined regularizer value, etc. ) . In some examples, the parameters for computing the regularizer value may be defined in a standard and may not need to be obtained with the test dataset. In some examples, the test dataset may be obtained with an associated reference score (e.g., a ground-truth reference score associated with the test dataset, which may be used for later comparison) and / or an indication of the method for comparison against the reference score.
[0286] At an operation 1520, a final regularizer score may be obtained. The final regularizer score may represent the performance of the trained model (whereas the regularizer score obtained during mid-training validation may represent the performance of the partly-trained model) . Similar to the operation 1510 described above, the final regularizer score may be obtained by obtaining a regularizer value for each precoding matrix in the test dataset and combining the regularizer scores into one final regularizer score. For example, each precoding matrix in the test dataset may be inputted to the trained autoencoder and the reconstructed precoding matrix obtained at the output of the trained autoencoder may be used to compute a regularizer value. The regularizer values obtained from all the precoding matrices in the test dataset may then be combined (e.g., averaged) to obtain the final regularizer score. Other techniques for combining the regularizer values may be used.
[0287] In examples where the test dataset is obtained together with a reference score associated with the test dataset, at an operation 1522, the final regularizer score may be compared with a reference score. This comparison may enable a determination whether the trained model is suitable for deployment. For example, if the final regularizer score differs from the reference score by more than a defined amount (e.g., greater than 10%difference, or greater than a defined margin of error) , this may indicate that the trained model has poor performance and should not be deployed or requires retraining.
[0288] At an operation 1524, the final regularizer score may be transmitted. If a comparison with a reference score was performed at the operation 1522, the result of the comparison may also be transmitted.
[0289] In a centralized example, the final regularizer score may be transmitted to a central entity. The central entity may collect final regularizer scores from all network entities that have completed training. The central entity may, based on similarity of final regularizer scores, determine which models can work together. For example, if the final regularizer score from a first network entity that is a UE-vendor is similar to the final regularizer score from a second network entity that is a NW-vendor, this may indicate that the trained first model from the UE-vendor is likely to be interoperable with the trained second model from the NW-vendor.
[0290] In a decentralized example, each network entity may share (e.g., via broadcast) its final regularizer score with other network entities that have completed training. Each network entity may independently determine which models from other network entities are likely to be interoperable with its own model. For example, each network entity may perform pairwise comparisons between its final regularizer score and the final regularizer scores received from other network entities. The network entity may determine that its own model is likely to be interoperable with another given model based on the two corresponding final regularizer scores being similar (e.g., the difference between the two corresponding final regularizer scores being within a predefined interoperability threshold for regularizer score differences) .
[0291] In various examples, the present disclosure has described techniques to use machine-learning based encoder and decoder models for CSI compression and decompression, using a physics-based regularizer during training of the models as well as during inference.
[0292] A physics-based regularizer may be a regularizer that represents a characteristic of the precoding matrix based on the physical property of the wireless communication system (e.g., reflecting physical laws of radio channels) , such as an expected sparsity of the precoding matrix in the angular domain, delay domain and / or time domain. The physics-based regularizer may be used in training of encoder and decoder models, and may be used for an indicator that is included as part of the CSI feedback during inference. In some examples, a regularizer may be defined such that it is differentiable (e.g., defined using DFT) , which may be more efficient for training using backpropagation. In some examples, multiple regularizers may be used in combination, representing respective characteristics of the precoding matrix and based on same or different channel properties.
[0293] A regularizer as disclosed herein may be defined in standards, including specifications for calculating and reporting an indicator representing the regularizer value during inference (e.g., including the relevant domains, quantization precision if applicable and / or order of operations for a multi-domain regularizer) . If multiple indicators are used, the technique for combining multiple indicators into a single combined indicator (e.g., a set of weights to combine multiple indicators using a weighted average, non-linear functions, conditional combinations, etc. ) may be defined in standards. The use of a standardized indicator representing regularizer values for encoder and decoder models may help to ensure interoperability between different vendors.
[0294] The methods for performing mid-training validation and end-of-training evaluation may be defined by standards, to help promote interoperability and consistency of models developed and trained by different network entities in the wireless communication system. For example, the definition of specific physics-based regularizers, methods of comparing regularizer values or regularizer scores and / or data to be used for validation or evaluation may be defined in a standard.
[0295] Examples disclosed herein include both centralized and decentralized implementations. In a centralized implementation, a central entity (e.g., a central datacenter) may be responsible for coordinating and scheduling validation and evaluation processes, including sending signals to initiate mid-training validation and / or end-of-training evaluation, comparing regularizer scores and / or making decisions for early termination of training and / or approval of trained models for deployment, among other responsibilities. The protocols used for communications between the central entity and network entities (e.g., UE-vendors and NW-vendors) may be defined by a standard, including specification of the data formats, signaling methods and / or security measures to protect the integrity and confidentiality of the transmitted data.
[0296] Examples of the present disclosure may result in improved accuracy and generalizability of the trained models, may help to improve interoperability of encoder and decoder models trained by different network entities (e.g., encoder model trained by a UE-vendor and decoder model trained by a different NW-vendor) . Examples of the present disclosure may also enable development of encoder / decoder models with reduced complexity, which may be beneficial for deployment of the trained model to resource-constrained devices (e.g., deployment to a UE having limited resources, such as a handheld device having limited battery power and limited processing power) . The use of a physics-based regularizer may also allow models to be efficiently trained using smaller datasets.
[0297] The present disclosure encompasses various embodiments, including not only method embodiments, but also other embodiments such as apparatus embodiments and embodiments related to non-transitory computer readable storage media. Embodiments may incorporate, individually or in combinations, the features disclosed herein.
[0298] Although this disclosure refers to illustrative embodiments, this is not intended to be construed in a limiting sense. Various modifications and combinations of the illustrative embodiments, as well as other embodiments of the disclosure, will be apparent to persons skilled in the art upon reference to the description.
[0299] Features disclosed herein in the context of any particular embodiments may also or instead be implemented in other embodiments. Method embodiments, for example, may also or instead be implemented in apparatus, system, and / or computer program product embodiments. In addition, although embodiments are described primarily in the context of methods and apparatus, other implementations are also contemplated, as instructions stored on one or more non-transitory computer-readable media, for example. Such media could store programming or instructions to perform any of various methods consistent with the present disclosure.
[0300] The terms "system" and "network" may be used interchangeably in embodiments of this application. "At least one" means one or more, and "a plurality of" means two or more. The term "and / or" describes an association relationship of associated objects, and indicates that three relationships may exist. For example, A and / or B may indicate the following three cases: Only A exists, both A and B exist, and only B exists, where A and B may be singular or plural. The character " / " usually indicates an "or" relationship between associated objects. "At least one of the following items (pieces) " or a similar expression thereof indicates any combination of these items, including a single item (piece) or any combination of a plurality of items (pieces) . For example, "at least one of A, B, or C" includes A, B, C, A and B, A and C, B and C, or A, B, and C, and "at least one of A, B, and C" may also be understood as including A, B, C, A and B, A and C, B and C, or A, B, and C. In addition, unless otherwise specified, ordinal numbers such as "first" and "second" in embodiments of this application are used to distinguish between a plurality of objects, and are not used to limit a sequence, a time sequence, priorities, or importance of the plurality of objects.
[0301] It should be understood that examples of the present disclosure may be embodied as a method, an apparatus, a non-transitory computer readable medium, a processing module, a chipset, a system chip or a computer program, among others. An apparatus may include a transmitting module configured to carry out transmitting steps described above and a receiving module configured to carry out receiving steps described above. An apparatus may include a processing module, processor or processing unit configured to control or cause the apparatus to carry out examples disclosed herein.
[0302] Although the present disclosure describes methods and processes with steps in a certain order, one or more steps of the methods and processes may be omitted or altered as appropriate. One or more steps may take place in an order other than that in which they are described, as appropriate.
[0303] Although the present disclosure is described, at least in part, in terms of methods, a person of ordinary skill in the art will understand that the present disclosure is also directed to the various components for performing at least some of the aspects and features of the described methods, be it by way of hardware components, software or any combination of the two. Accordingly, the technical solution of the present disclosure may be embodied in the form of a software product. A suitable software product may be stored in a pre-recorded storage device or other similar non-volatile or non-transitory computer readable medium, including DVDs, CD-ROMs, USB flash disk, a removable hard disk, or other storage media, for example. The software product includes instructions tangibly stored thereon that enable a processing device (e.g., a personal computer, a server, or a network device) to execute examples of the methods disclosed herein. The machine-executable instructions may be in the form of code sequences, configuration information, or other data, which, when executed, cause a machine (e.g., a processor or other processing device) to perform steps in a method according to examples of the present disclosure.
[0304] The present disclosure may be embodied in other specific forms without departing from the subject matter of the claims. The described example embodiments are to be considered in all respects as being only illustrative and not restrictive. Selected features from one or more of the above-described embodiments may be combined to create alternative embodiments not explicitly described, features suitable for such combinations being understood within the scope of this disclosure.
[0305] All values and sub-ranges within disclosed ranges are also disclosed. Also, although the systems, devices and processes disclosed and shown herein may comprise a specific number of elements / components, the systems, devices and assemblies could be modified to include additional or fewer of such elements / components. For example, although any of the elements / components disclosed may be referenced as being singular, the embodiments disclosed herein could be modified to include a plurality of such elements / components. The subject matter described herein intends to cover and embrace all suitable changes in technology.
Claims
1.A method in a wireless communication system, the method comprising:obtaining a precoding matrix;compressing the precoding matrix into a low-dimensional representation using a first model; andtransmitting channel state information (CSI) feedback including the low-dimensional representation and an indicator based on the precoding matrix, the indicator representing a characteristic of the precoding matrix.2.The method of claim 1, wherein the indicator represents a regularizer value computed according to a defined formula that represents the characteristic of the precoding matrix based on a physical property of the wireless communication system.3.The method of claim 1 or 2, wherein the characteristic of the precoding matrix is a sparsity of the precoding matrix in an angular domain, a delay domain and / or a time domain.4.The method of any one of claims 1 to 3, wherein the transmitted CSI feedback includes multiple indicators representing the precoding matrix.5.The method of claim 4, wherein each indicator represents a respective characteristic of the precoding matrix.6.The method of any one of claims 1 to 5, wherein the indicator included in the CSI feedback is a quantized indicator.7.The method of any one of claims 1 to 6, further comprising:receiving parameters of the first model from a control signal or over a data channel.8.The method of any one of claims 1 to 7, wherein the first model is an encoder model.9.A method in a wireless communication system, the method comprising:receiving channel state information (CSI) feedback including a low-dimensional representation of a corresponding precoding matrix and a received indicator representing a characteristic of the corresponding precoding matrix;decompressing the low-dimensional representation using a second model to obtain a reconstructed precoding matrix; andperforming a consistency verification by comparing a reconstructed indicator determined from the reconstructed precoding matrix against the received indicator.10.The method of claim 9, further comprising:detecting an inconsistency based on a difference between the reconstructed indicator and the received indicator exceeding a threshold; andperforming a corrective action.11.The method of claim 10, wherein the corrective action includes at least one of:transmitting a request for retransmission of CSI feedback;performing another decompression of the low-dimensional representation using a different second model to obtain a different reconstructed precoding matrix; ortransmitting a request for codebook-based CSI feedback.12.The method of claim 9, further comprising:detecting that the reconstructed precoding matrix is a successful reconstruction based on a difference between the reconstructed indicator and the received indicator being within a threshold; andperforming a link operation and / or configuring a channel estimation operation based on the reconstructed precoding matrix.13.The method of claim 12, wherein performing the link operation and / or configuring the channel estimation operation includes at least one of:performing a beamforming operation;performing a link adaptation operation;configuring a set of resources for the channel estimation operation; orconfiguring a time duration of the channel estimation operation.14.The method of any one of claims 9 to 13, wherein the CSI feedback includes multiple received indicators.15.The method of claim 14, wherein each received indicator represents a respective characteristic of the corresponding precoding matrix, and each indicator represents a same or different physical property of the wireless communication system.16.The method of claim 14 or 15, wherein performing the consistency verification comprises comparing a combination of multiple reconstructed indicators determined from the reconstructed precoding matrix against a combination of the multiple received indicators.17.The method of claim 16, wherein the multiple reconstructed indicators are combined using defined weights and the multiple received indicators are combined using the defined weights.18.The method of any one of claims 9 to 17, further comprising:transmitting an indication of a performance of the second model based on comparison of the received indicator with the reconstructed indicator.19.The method of any one of claims 9 to 18, wherein the reconstructed indicator represents a regularizer value computed according to a defined formula that represents the characteristic of the precoding matrix based on a physical property of the wireless communication system.20.The method of claim 19, wherein the characteristic of the corresponding precoding matrix is a sparsity of the corresponding precoding matrix in an angular domain, a delay domain and / or a time domain.21.The method of any one of claims 9 to 20, further comprising:receiving parameters of the second model from a control signal or over a data channel.22.The method of any one of claims 9 to 21, wherein the second model is a decoder model.23.A method in a wireless communication system, the method comprising:obtaining a training dataset including precoding matrix data; andtraining an autoencoder including training a first model of the autoencoder to perform a precoding matrix compression task or training a second model of the autoencoder to perform a precoding matrix decompression task, wherein the autoencoder is trained to minimize an error between an output of the autoencoder and the precoding matrix data, and is also trained to minimize a regularizer value representing a characteristic of each precoding matrix.24.The method of claim 23, further comprising:during the training of the autoencoder, performing validation of a partly-trained first model or a partly-trained second model of a partly-trained autoencoder by:obtaining a validation dataset including precoding matrix data;obtaining a regularizer score based on regularizer values computed using outputs generated by the partly-trained autoencoder from the validation dataset; andtransmitting the regularizer score.25.The method of claim 24, wherein the validation is performed in response to a received signal.26.The method of claim 24, wherein the validation is performed in accordance with a periodicity determined among a plurality of network entities of the wireless communication system.27.The method of any one of claims 24 to 26, wherein the validation dataset is common to two or more network entities of the wireless communication system, wherein the validation dataset is associated with a defined formula for obtaining the regularizer score, and wherein the regularizer score is transmitted to at least one other network entity of the wireless communication system.28.The method of any one of claims 24 to 27, further comprising:receiving a signal to end training, wherein the signal is received in response to transmitting the regularizer score.29.The method of any one of claims 24 to 28, further comprising:receiving a reference score associated with the validation dataset; andcomparing the regularizer score with the reference score to evaluate the partly-trained autoencoder.30.The method of any one of claims 24 to 29, further comprising:terminating training based on the regularizer score satisfying a target threshold.31.The method of any one of claims 24 to 29, further comprising:terminating training in response to a received control signal.32.The method of any one of claims 23 to 30, further comprising:performing evaluation of a trained first model or a trained second model of a trained autoencoder by:obtaining a test dataset including precoding matrix data;obtaining a final regularizer score based on regularizer values computed using outputs generated by the trained autoencoder from the test dataset; andtransmitting the final regularizer score.33.The method of claim 32, further comprising:receiving another regularizer score from another network entity; andcomparing the final regularizer score with the received another regularizer score to determine interoperability of the trained first model or the trained second model with the another network entity.34.The method of any one of claims 23 to 33, wherein the autoencoder is trained to minimize multiple regularizer values together.35.The method of claim 34 when dependent on claim 24, wherein the regularizer score is computed based on multiple regularizer values computed using outputs generated by the partly-trained autoencoder from the validation dataset, wherein the multiple regularizer values are combined using defined weights.36.The method of any one of claims 23 to 35, wherein the regularizer value is computed according to a defined formula representing the characteristic of the precoding matrix based on a physical property of the wireless communication system.37.The method of claim 36, wherein the characteristic of the precoding matrix is a sparsity of the precoding matrix in an angular domain, a delay domain, and / or a time domain.38.The method of claim 36 or 37, wherein the defined formula is a differentiable formula.39.The method of claim 38, wherein the defined formula is based on a discrete Fourier transform (DFT) of the precoding matrix.40.The method any one of claims 23 to 39, wherein the first model is an encoder model and / or the second model is a decoder model.41.An apparatus comprising:a memory; anda processor configured to execute instructions stored in the memory to cause the apparatus to carry out the method of any one of claims 1 to 8.42.A base station comprising:a memory; anda processor configured to execute instructions stored in the memory to cause the base station to carry out the method of any one of claims 8 to 22.43.A network entity comprising:a memory; anda processor configured to execute instructions stored in the memory to cause the base station to carry out the method of any one of claims 23 to 40.44.The network entity of claim 43, wherein the network entity is a central datacenter.45.The network entity of claim 43, wherein the network entity services one or more user equipment (UEs) .46.The network entity of claim 43, wherein the network entity services one or more base stations (BSs) .47.A non-transitory computer readable medium having machine-executable instructions stored thereon, wherein the instructions, when executed by an apparatus, cause the apparatus to perform the method of any one of claims 1 to 8.48.A non-transitory computer readable medium having machine-executable instructions stored thereon, wherein the instructions, when executed by a base station, cause the base station to perform the method of any one of claims 9 to 22.49.A non-transitory computer readable medium having machine-executable instructions stored thereon, wherein the instructions, when executed by a network entity, cause the network entity to perform the method of any one of claims 23 to 40.50.The non-transitory computer readable medium of claim 49, wherein the network entity is a datacenter.51.The non-transitory computer readable medium of claim 49, wherein the network entity services one or more user equipment (UEs) .52.The non-transitory computer readable medium of claim 49, wherein the network entity services one or more base stations (BSs) .53.A communication apparatus, configured to perform the method according to any one of claims 1 to 8.54.The communication apparatus of claim 53, further comprising:an obtaining unit configured to perform the obtaining;a processing unit configured to perform the compressing; anda transmitting unit configured to perform the transmitting.55.The communication apparatus of claim 53, further comprising:one or more processors configured to perform the obtaining and compressing; andan interface circuit configured to perform the transmitting.56.A communication apparatus, configured perform the method according to any one of claims 9 to 22.57.The communication apparatus of claim 56, further comprising:a receiving unit configured to perform the receiving; anda processing unit configured to perform the decompressing and comparing.58.The communication apparatus of claim 56, further comprising:one or more processors configured to perform the decompressing and comparing; andan interface circuit configured to perform the receiving.59.A communication apparatus, configured to perform the method according to any one of claims 23 to 40.60.The communication apparatus of claim 59, further comprising:an obtaining unit configured to perform the obtaining; anda processing unit configured to perform the training.61.The communication apparatus of claim 59, wherein the communication apparatus is a network entity.62.The communication apparatus of claim 61, wherein the network entity is a datacenter.63.The communication apparatus of claim 61, wherein the network entity services one or more user equipment (UEs) .64.The communication apparatus of claim 61, wherein the network entity services one or more base stations (BSs) .65.A computer program characterized in that, when the computer program is run on a computer, the computer is caused to execute the method of any one of claims 1 to 40.
Citation Information
Patent Citations
Channel estimation method and device
CN111656742A
Hybrid precoding and feedback method based on deep learning
CN114844541A
Systems, methods, and apparatus for artificial intelligence and machine learning for a physical layer of communication system
US20230131694A1
Techniques for channel state information and channel compression switching
WO2022227081A1
Cited By
Remote data transmission method and system based on deep sea detection sonar
CN121842337A