Model interoperability for machine learning based communication channel information reporting

By specifying the mapping between model information and channel information, and using machine learning models to generate and reconstruct channel information representations, the model interoperability problem in wireless communication systems is solved, and efficient channel feedback and resource utilization are achieved.

CN120880927APending Publication Date: 2025-10-31SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510557210.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-21
Filing Date
2025-04-29
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In wireless communication systems, existing technologies struggle to achieve model interoperability between wireless devices and network-side devices, resulting in insufficient accuracy of channel feedback information and high resource consumption.

Method used

By specifying model information, the mapping between channel information and resources, the capability declaration of wireless devices, and the data transmission scheme, machine learning models are used to generate and reconstruct representations of channel information, reducing the amount of cooperation and improving feedback accuracy.

Benefits of technology

This enables efficient collaboration among wireless device vendors, reduces resource consumption for channel feedback information, and improves the accuracy of channel feedback and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120880927A_ABST
    Figure CN120880927A_ABST
Patent Text Reader

Abstract

An apparatus for model interoperability of channel information reporting based on machine learning may include: a receiver configured to receive a reference signal using a channel; a transmitter configured to transmit a representation related to the channel; and a processing circuit configured to determine channel information based on the reference signal, generate a representation based on the channel information using a model, and transmit model information using the receiver or transmitter to specify the model. The model information may include an identifier of the model. The model information may include structural information of the model. The model information may include information about a type of input of the model. The model information may include information about a format of an input of the model. The model information may include mapping information for mapping the channel information to an input of the model. The mapping information may include information for the first sub-band and information for the second sub-band.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority and interest in U.S. Provisional Patent Application No. 63 / 640,195, filed April 29, 2024; U.S. Provisional Patent Application No. 63 / 697,935, filed September 23, 2024; and U.S. Provisional Patent Application No. 63 / 750,260, filed January 27, 2025, all of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to communication systems, and more specifically, to model interoperability and collaboration for machine learning-based communication channel information reporting. Background Technology

[0004] In some wireless communication systems, a user equipment (UE) can determine the channel conditions to report to the base station by performing channel measurements based on reference signals transmitted by the base station through the channel. The UE can use the channel measurements to calculate feedback in the form of a precoding matrix that the UE can report to the base station. The base station can then apply the precoding matrix to subsequent downlink transmissions through the channel, which can improve downlink transmission performance.

[0005] However, sending feedback information about channel conditions can consume relatively large amounts of resources as overhead. To reduce the amount of data used to send feedback information, some wireless communication systems can use one or more types of codebooks to enable the receiving device to send implicit and / or explicit channel condition feedback to the transmitting device. However, the use of codebooks may not provide feedback with sufficient accuracy. Furthermore, the use of codebooks may still involve transmitting a large amount of overhead data on the UL channel.

[0006] To overcome this problem, feedback schemes can use one or more machine learning (ML) models to transmit channel information of the wireless communication system. For example, the UE can use a first ML model to generate a compressed representation of the channel information, which can reduce the amount of resources required to transmit information to the base station. The base station can use a second ML model to reconstruct the channel information.

[0007] The problem with this approach is that it may involve some form of collaboration between the UE-side devices (and / or the vendors of the UE-side devices) and the network-side devices (and / or the vendors of the network-side devices) to achieve interoperability between models deployed on the UE side and the network (e.g., base station) side. For example, devices and / or vendors may need to collaborate to train an encoder model at the UE and a corresponding decoder model at the base station. However, without a technology (e.g., a standard) for exchanging information about models, data transmission, training, capabilities, etc., collaboration may be difficult or impossible. Summary of the Invention

[0008] To overcome these problems, this paper describes systems and methods for achieving interoperability of models in wireless communication systems, for example by facilitating collaboration between wireless devices and / or suppliers of wireless devices.

[0009] Some of the systems and methods described herein involve model information that can be used to specify a model. For example, model information may include model structure that specifies the model type (e.g., convolutional neural network (CNN), multilayer perceptron (MLP), etc.), input and / or output features, input and / or output shapes, number of layers, activation functions, etc. As another example, model information may specify the type and / or format of model inputs and / or outputs (e.g., target channel information) that can be specified in the spatial domain, frequency domain, angular domain, and / or delay domain.

[0010] Some of the additional systems and methods described in this paper involve mapping between channel information and resources in a model. For example, a model description may include a mapping between target channel information (e.g., target channel state information (CSI)) across resources such as subbands and time slots to the model's inputs and / or outputs.

[0011] Some of the additional systems and methods described herein relate to schemes for indicating the ability of a wireless device, such as a UE, to support one or more models. For example, a UE may declare the ability to support one or more specified models (e.g., a reference model). A network-side device (e.g., a base station) may configure the UE using one or more supported models that the UE can apply to inference operations. As another example, a UE may declare the ability to support one or more specified model structures (e.g., a reference model structure). A network-side device may configure the UE using one or more supported model structures, and one or more parameters (e.g., weights) of the model structure may be transmitted to the UE.

[0012] Some of the additional systems and methods described herein relate to schemes for instructing wireless devices, such as UEs, to modify model operations. For example, a UE may declare the ability to run a model (e.g., a reference model) as specified (e.g., directly), or to perform one or more operations, such as tuning, to modify the model's performance. As another example, a UE may declare the ability to run a model based on a specified model structure (e.g., a reference model structure) as specified (e.g., based on training using field data or other data), or to perform one or more operations, such as engineering, to modify the model's performance.

[0013] Some of the additional systems and methods described herein relate to schemes for transmitting data to a wireless device, such as a UE, to train a model. For example, the dataset may be specified to include the type and / or format (e.g., precoding matrix or channel matrix) of each data point. Additionally or alternatively, the dataset may be specified to include information mapping the data points to resources such as subbands, time slots, etc., in the time, spatial, and / or frequency domains.

[0014] Some of these methods can provide improvements because they reduce the amount of collaboration between wireless device vendors. Additionally or alternatively, some of these methods can provide improvements because they enable wireless device vendors to collaborate in a specified manner.

[0015] Some of the additional systems and methods described herein relate to schemes for performing inference in embodiments where CSI compression (e.g., in the time-space-frequency (TSF) domain) may have predictive information. For example, a UE may use a CSI prediction model to predict some CSIs and apply the predicted CSIs to an encoder (e.g., compression) model. In some such embodiments described herein, the UE may receive an indication that the output of the CSI prediction model may not be directly reported to the network side. This indication may be provided, for example, through a reporting quantity, a link between a first reporting configuration and a second reporting configuration, and / or otherwise. Additionally or alternatively, the UE may be configured to report the output of the CSI prediction model to the network side based on, for example, an indication from the network side. In some embodiments, the UE may determine the amount of processing capability from a specification and / or report the amount of processing capability.

[0016] Some of the additional systems and methods described herein relate to schemes for collecting training data in embodiments where CSI compression (e.g., in the TSF domain) can have predictive information. For example, in some embodiments, the encoder (e.g., compression) model can be trained using the target CSI of the measurement used to measure the time slot and the predicted CSI of the time slot used to predict the time slot as input. As another example, in some embodiments, the encoder (e.g., compression) model can be trained using the target CSI of the measurement used to measure both the time slot and the predicted time slot as input. As yet another example, in some embodiments, the prediction model and the encoder (e.g., compression) model can be trained jointly.

[0017] An apparatus may include: a receiver configured to receive a reference signal using a channel; a transmitter configured to transmit a channel-related representation; and processing circuitry configured to determine channel information based on the reference signal, use a model to generate a representation based on the channel information, and use the receiver or transmitter to transmit model information to specify the model. The model information may include an identifier for the model. The model information may include structural information about the model. The model information may include information about the type of input to the model. The model information may include information about the format of the input to the model. The model information may include mapping information for mapping the channel information to the input to the model. The mapping information may include information for a first subband and information for a second subband. The mapping information may include information for a first time slot and information for a second time slot.

[0018] An apparatus may include: a receiver configured to receive a reference signal using a channel; a transmitter configured to transmit a channel-related representation; and processing circuitry configured to determine channel information based on the reference signal, use a model to generate a representation based on the channel information, and use the receiver or transmitter to transmit model-related apparatus capability information. The capability information may include information about the model supported by the apparatus. The capability information may include information about the structure of the model supported by the apparatus. The capability information may include information about the inference operations of the model. The processing circuitry may be configured to transmit the capability information using a channel information reporting configuration. The capability information may include information about the type of cooperation supported by the apparatus. The capability information may include information about the apparatus's ability to modify the model. The capability information may include information about the amount of time associated with modifying the model.

[0019] An apparatus may include: a receiver configured to receive a reference signal using a channel; a transmitter configured to transmit a channel-related representation; and processing circuitry configured to determine channel information based on the reference signal, use a model to generate a representation based on the channel information, and use the receiver or transmitter to transmit dataset information to specify a dataset for the model. The dataset information may include the type of information in the dataset. The dataset information may include the format of the information in the dataset. The dataset information may include mapping information for the dataset.

[0020] An apparatus may include: a receiver configured to receive a reference signal using a channel; a transmitter configured to transmit a channel-related representation; and processing circuitry configured to determine channel information based on the reference signal, use one or more models to generate a channel-related prediction based on the channel information, and use one or more models to generate a representation based on the prediction. The processing circuitry may be configured to receive a reporting configuration including an indication and to report a representation based on the indication. The indication may include a number of reports. The reporting configuration may be a first reporting configuration including a link to a second reporting configuration. The link may include a configuration identifier. The first reporting configuration may be a prediction configuration, and the second reporting configuration may be a compression configuration. The processing circuitry may be configured to receive a reporting configuration including an indication and to report a prediction based on the indication. The processing circuitry may be configured to combine the prediction and representation to generate a combined report and to report the combined report. The processing circuitry may be configured to determine a quantity of processing capability from a specification. The processing circuitry may be configured to report a quantity of processing capability of the apparatus.

[0021] An apparatus may include a receiver and processing circuitry, wherein the processing circuitry is configured to perform one or more operations, the one or more operations including: using the receiver to receive one of one or more reference signals for a measurement time slot of a channel; determining measured channel information based on one of the one or more reference signals for the measurement time slot of the channel; using a first model to generate predicted channel information for one or more prediction time slots of the channel based on the measured channel information; and collecting training data for a second model based on one of the one or more reference signals for the measurement time slots. The processing circuitry may be configured to perform one or more operations, wherein the one or more operations include: training a second model using the training data and the predicted channel information. The training data may be first training data, and the processing circuitry may be configured to perform one or more operations, the one or more operations including: using the receiver to receive reference signals for the prediction time slots of the channel; and collecting second training data for the second model based on the reference signals for the prediction time slots. The processing circuitry may be configured to perform one or more operations, the one or more operations including training a second model using the first training data and the second training data. The processing circuitry may be configured to perform one or more operations, the one or more operations including: jointly training a first model and a second model using the training data and the predicted channel information. The second model can be a generative model, and the processing circuitry can be configured to perform training based on measured channel information and the output of a reconstructed model configured to receive the output of the generative model. The processing circuitry can also be configured to perform training based on measured channel information and predicted channel information used for predicting time slots. Attached Figure Description

[0022] The accompanying drawings are not necessarily drawn to scale, and throughout the drawings, for illustrative purposes, elements with similar structures or functions are generally indicated by the same reference numerals or portions thereof. The drawings are intended only to facilitate the description of the various embodiments described herein. The drawings do not depict every aspect of the teachings disclosed herein and do not limit the scope of the claims. To prevent obscurity, not all components, connections, etc., may be shown, and not all components may have reference numerals. However, the pattern of component configuration can be readily discerned from the drawings. The drawings, together with the specification, illustrate exemplary embodiments of this disclosure and, together with the specification, serve to explain the principles of this disclosure.

[0023] Figure 1 An embodiment of a wireless communication device according to the present disclosure is shown.

[0024] Figure 2 Another embodiment of a wireless communication device according to the present disclosure is shown.

[0025] Figure 3 An embodiment of the dual-model training scheme according to this disclosure is shown.

[0026] Figure 4 An embodiment of a system having a pair of models to provide channel information feedback according to this disclosure is shown.

[0027] Figure 5 An example embodiment of a system for reporting downlink physical layer information according to this disclosure is shown.

[0028] Figure 6 An example embodiment of a system for reporting uplink physical layer information according to this disclosure is shown.

[0029] Figure 7 An example embodiment of a system for reporting downlink physical layer channel state information according to this disclosure is shown.

[0030] Figure 8 An embodiment of the learning process of a machine learning model according to the present disclosure is shown.

[0031] Figure 9 An example embodiment of a method for jointly training a pair of encoder and decoder models according to the present disclosure is shown.

[0032] Figure 10 An example embodiment of a method for training a model using the latest shared values, according to this disclosure, is shown.

[0033] Figure 11 An example embodiment of a dual-model training scheme with preprocessing and postprocessing according to the present disclosure is shown.

[0034] Figure 12 An embodiment of a system for using a dual-model scheme according to the present disclosure is shown.

[0035] Figure 13 An example embodiment of a user equipment (UE) according to this disclosure is shown.

[0036] Figure 14 An example embodiment of a base station according to this disclosure is shown.

[0037] Figure 15 An embodiment of a method for providing physical layer information feedback according to the present disclosure is shown.

[0038] Figure 16 An embodiment of a system having a pair of models to provide channel information feedback according to this disclosure is shown.

[0039] Figure 17 An example embodiment of a pair of models that can be used for joint compression of channel information and channel quality information according to this disclosure is shown.

[0040] Figure 18An embodiment of a system according to this disclosure having a pair of models based on channel information of one or more sub-bands is shown.

[0041] Figure 19 An embodiment of a pair of models that can be used for channel quality indicator compression across subbands according to this disclosure is shown.

[0042] Figure 20 An embodiment of a system having a pair of models to provide channel information compression according to this disclosure is shown.

[0043] Figure 21 A first embodiment of a pair of models that can be used to implement a compression scheme according to this disclosure is shown.

[0044] Figure 22 A second embodiment of a pair of models that can be used to implement a compression scheme according to this disclosure is shown.

[0045] Figure 23 A third embodiment of a pair of models that can be used to implement a compression scheme according to this disclosure is shown.

[0046] Figure 24 An embodiment of a communication system is shown that uses channel constraint information and artificial intelligence and / or machine learning to provide channel information feedback in accordance with this disclosure.

[0047] Figure 25 An example embodiment of a communication system according to this disclosure is shown, in which a decoder model is shared with a UE configured with a constrained subspace.

[0048] Figure 26 An example embodiment of a communication system according to this disclosure is shown in which a decoder model is not shared with a UE configured with a constrained subspace.

[0049] Figure 27 An embodiment of a system according to this disclosure that uses constrained subspace information as input to a model is shown.

[0050] Figure 28 An embodiment of a communication system according to this disclosure is shown, which has data collection capabilities for providing channel information feedback using one or more models.

[0051] Figure 29 An example embodiment of a communication system according to this disclosure is shown, which has data collection capabilities for one or more models that can be used to provide channel information feedback.

[0052] Figure 30 An example embodiment of a communication system according to this disclosure is shown, which has data collection capabilities for one or more models that can be used to provide channel information feedback for NCJT schemes.

[0053] Figure 31 An embodiment of a communication system with model training collaboration functionality according to this disclosure is shown.

[0054] Figure 32 An example embodiment of a communication system with model training collaboration functionality according to this disclosure is shown.

[0055] Figure 33 An embodiment of a scheme for mapping channel information and / or reconstructed channel information to model inputs and / or outputs according to the present disclosure is shown.

[0056] Figure 34 An embodiment of a time constraint for channel information processing in a communication system according to the present disclosure is shown.

[0057] Figure 35 An embodiment of a timeline for non-periodic CSI reporting based on a periodic reference signal, according to this disclosure, is shown.

[0058] Figure 36 An embodiment of a timeline for CPU occupancy in a P / SP CSI report based on a periodic reference signal, according to this disclosure, is shown.

[0059] Figure 37 An embodiment of a timeline for CPU occupancy reporting based on a periodic reference signal, according to this disclosure, is shown.

[0060] Figure 38 An embodiment of a communication system according to the present disclosure is shown, which has functions related to channel information compression using predictive information.

[0061] Figure 39 An example embodiment of a communication system having functions related to channel information compression using predictive information, according to this disclosure, is shown.

[0062] Figure 40 This is a block diagram of an electronic device in a network environment according to embodiments of the present disclosure.

[0063] Figure 41 A system comprising a UE and a base station communicating with each other, according to an embodiment of the present disclosure, is shown. Detailed Implementation

[0064] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of this disclosure. However, those skilled in the art will understand that various aspects of the disclosure may be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits have not been described in detail so as not to obscure the subject matter of this disclosure.

[0065] Throughout this specification, references to "an embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment disclosed herein. Therefore, the phrases "in one embodiment," "in an embodiment," or "according to an embodiment" (or other phrases with similar meanings) appearing in various places throughout this specification may not necessarily all refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this regard, as used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" should not be construed as necessarily preferred or advantageous over other embodiments. Additionally, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Furthermore, depending on the context of the discussion herein, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. Similarly, hyphenated terms (e.g., "two-dimensional", "pre-determined", "pixel-specific", etc.) can occasionally be used interchangeably with their non-hyphenated counterparts (e.g., "two-dimensional", "predetermined", "pixel specific", etc.), and uppercase entries (e.g., "counter clock", "row select", "pixout", etc.) can be used interchangeably with their non-uppercase counterparts (e.g., "counter clock", "row select", "pixout", etc.). This occasional interchangeability should not be considered inconsistent with each other.

[0066] Furthermore, depending on the context of this discussion, singular terms may include their corresponding plural forms, and plural terms may include their corresponding singular forms. It should also be noted that the various figures shown and discussed herein (including component diagrams) are for illustrative purposes only and are not drawn to scale. For example, the dimensions of some elements may be enlarged relative to others for clarity. Additionally, reference numerals are repeated in the figures where deemed appropriate to indicate corresponding and / or similar elements.

[0067] The terminology used herein is for the purpose of describing some exemplary embodiments only and is not intended to limit the claimed subject matter. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprising” and / or “including” as used in this specification specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0068] It will be understood that when an element or layer is referred to as being on, "connected to," or "coupled to" another element or layer, it can be directly on, connected to, or coupled to another element or layer, or there may be intermediate elements or layers. Conversely, when an element is referred to as being "directly on," "directly connected to," or "directly coupled to" another element or layer, there are no intermediate elements or layers. The same reference numerals always indicate the same elements. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0069] As used herein, the terms “first,” “second,” etc., serve as labels for nouns that follow them and do not imply any kind of ordering (e.g., spatial, temporal, logical, etc.) unless explicitly defined as such. Furthermore, the same reference numerals may be used across two or more figures to indicate parts, components, blocks, circuits, units, or modules having the same or similar functions. However, this usage is merely for simplicity of description and ease of discussion; it does not imply that the construction or architectural details of such components or units are identical across all embodiments, or that such commonly referenced parts / modules are the only way to implement some of the exemplary embodiments disclosed herein.

[0070] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject pertains. It will be further understood that terms (such as those defined in common dictionaries) shall be interpreted as having a meaning consistent with their meaning in the context of the relevant field and shall not be interpreted in an idealized or overly formalized sense unless expressly defined herein.

[0071] As used herein, the term "module" indicates any combination of software, firmware, and / or hardware configured to provide the functionality described herein in conjunction with modules. For example, software may be embodied as a software package, code, and / or instruction set or instructions, and the term "hardware," as used in any implementation described herein, may include, for example, alone or in any combination, a component assembly, hardwired circuitry, programmable circuitry, state machine circuitry, and / or firmware storing instructions executed by the programmable circuitry. Modules may be embodied collectively or individually as circuitry forming part of a larger system, such as, but not limited to, integrated circuits (ICs), system-on-a-chip (SoCs), component assemblies, etc.

[0072] Overview

[0073] In some wireless communication systems, transmitting devices may rely on receiving devices to provide feedback information about channel conditions, enabling the transmitting device to transmit more efficiently to the receiving device over the channel. For example, in a 5G New Radio (NR) system, a base station (e.g., a gNodeB or gNB) can transmit a reference signal to a user equipment (UE) via a downlink (DL) channel. The UE can measure the reference signal to determine the channel conditions on the DL channel. Subsequently, the UE can transmit feedback information (e.g., channel state information (CSI)) indicating the channel conditions on the DL channel to the base station via an uplink (UL) channel. The base station can use the feedback information to improve its transmission to the UE over the DL channel, for example, by using beamforming.

[0074] However, sending feedback information about channel conditions can consume relatively large amounts of resources as overhead. To reduce the amount of data used to transmit feedback information, some wireless communication systems can use one or more types of codebooks to enable the receiving device to send implicit and / or explicit channel condition feedback to the transmitting device. For example, in 5G NR systems, a Type I codebook can be used to provide implicit CSI feedback to the gNB in ​​the form of an index, where the index can point to a predefined PMI selected by the UE based on the DL channel conditions. The gNB can then use the PMI for beamforming in the DL channel. As another example, a Type II codebook can be used to provide explicit CSI feedback, where the UE can derive the PMI, which can be fed back to the gNB, where the gNB can use the PMI for beamforming in the DL channel. However, the use of a Type I codebook may not provide CSI feedback with sufficient accuracy. Furthermore, the use of a Type II codebook may still involve transmitting a large amount of overhead data on the UL channel.

[0075] The feedback scheme according to this disclosure can use artificial intelligence (AI), machine learning (ML), deep learning, etc. (any or all of which can be individually and / or collectively referred to as machine learning or ML) to generate representations of physical layer information of a wireless communication system. For example, in some embodiments, the feedback scheme can use an ML model to generate a representation of feedback information for channel conditions (e.g., a representation of a channel matrix, precoding matrix, etc.). The representation can be a compressed, encoded, or otherwise modified form of the feedback information, depending on the implementation details, which can reduce the resources involved in transmitting feedback information between devices.

[0076] The feedback scheme according to this disclosure can also use machine learning to reconstruct physical layer information from representations. For example, in some embodiments, the feedback scheme can use an ML model to reconstruct feedback information or an approximation of feedback information from a representation of feedback information for channel conditions. For convenience, the ML model can be simply referred to as the model.

[0077] A model that generates a representation of an input (e.g., physical layer information, such as feedback information for channel conditions) can be called a generative model. A model that reconstructs the input or an approximation of the input from the representation of the input can be called a reconstructive model. The output of the reconstructive model can be called the reconstructed input. Therefore, the reconstructed input can be an input applied to the generative model, or an approximation, estimate, prediction, etc., of the input applied to the generative model. The generative model and the corresponding reconstructive model can be collectively referred to as a pair of ML models or a pair of models. In some embodiments, the generative model can be implemented as an encoder model, and / or the reconstructive model can be implemented as a decoder model. Therefore, the encoder model and the decoder model can also be referred to as a pair of ML models or a pair of models.

[0078] For the purpose of distinguishing a model from one or more other models, any model can be referred to as the first model, the second model, model A, model B, etc., and the labels used for models are not intended to suggest the type of model unless otherwise apparent from the context. For example, in the context of a pair of models, if model A indicates a generating model, then model B may indicate a refactoring model.

[0079] A node can refer to a base station, a UE, or any other device that can use one or more ML models as disclosed herein. Additional examples of nodes may include a UE-side server, a base station-side server (e.g., a gNB-side server), an eNodeB, a primary node, a secondary node, etc., whether logical, physical, or a combination thereof. For the purpose of distinguishing a node from one or more other nodes, any node may be referred to as a first node, a second node, node A, node B, etc., and the labels used for nodes are not intended to suggest the type of node unless otherwise apparent from the context. For example, in some embodiments, a first node may refer to a UE, and a second node may refer to a base station. However, in some other embodiments, a first node may refer to a first UE, and a second node may refer to a second UE configured to perform sidelink communication with the first UE.

[0080] In some example embodiments, the first node may use a first model (e.g., a generative model) to encode the channel matrix, precoding matrix, etc., to generate a feature vector that can be sent to the second node. The second node may use a second model (e.g., a reconstruction model) to decode the feature vector to reconstruct the original information (e.g., the channel matrix, precoding matrix, etc.) or an approximation of the original information.

[0081] According to some embodiments of this disclosure, a dual-model training scheme can be implemented, wherein the models can be trained in pairs. For example, a reconstruction model can be used to train a generative model, and / or a generative model can be used to train a reconstruction model. In some example implementations, a pair of models can be configured to implement an autoencoder, wherein an encoder model (e.g., for a first node) can be trained using a decoder model (e.g., for a second node).

[0082] In some embodiments, a first model (e.g., a generative model) that can be used for inference by the first node can be trained using a second model (e.g., a reconstructed model) that can actually be used for inference by the second node. Training can be performed by the first node, the second node, and / or any other means, such as a server that can train models (e.g., offline) and transmit one or more of the trained models to one or more of the nodes for inference.

[0083] Alternatively or additionally, a second model can be used to train the first model, wherein the second model can provide a certain amount of matching between the first and second models, even if the second model is not an actual model that can be used for inference by the second node. Alternatively or additionally, a reference model used for the second model can be used to train the first model. Alternatively or additionally, a second model can be used to train the first model, wherein the second model can be configured with values ​​for weights, hyperparameters, etc., that can be initialized to predetermined values, randomized values, etc.

[0084] In some embodiments, a pair of models may be trained simultaneously, sequentially (e.g., alternating between training the first model while freezing the second model, and then training the second model while freezing the first model) and / or similarly using the same or different training datasets.

[0085] In some embodiments, a node can use a quantizer to convert a representation of physical layer information into a form that can be more easily transmitted over a communication channel. For example, a quantizer can convert a real-number (e.g., integer) representation of physical layer information into a binary bitstream, which can then be applied to a polar encoder or other means for transmission over a physical uplink or downlink channel. Similarly, a node can use a dequantizer to convert a bitstream into a representation of physical layer information that can be used to reconstruct the physical layer information. In some embodiments, a quantizer or dequantizer can be considered part of an ML model. For example, a generative model may include an encoder and a corresponding quantizer, and / or a reconstructive model may include a corresponding dequantizer and a decoder.

[0086] According to some embodiments of this disclosure, one or more frameworks can be implemented for training models and / or transferring models between nodes. For example, in a first type of framework, a first node (node ​​A) can jointly train a pair of models (model A and model B). Node A can use the trained model A for inference and transfer the trained model B to a second node (node ​​B), whereby the second node can use the trained model B for inference. In a variation of the first type of framework, node A can transfer the trained model A to node B, and node B can use the trained model A to train its own model B for inference.

[0087] In the second type of framework, a reference model can be established as Model A for node A, and node B can then use the reference model as Model A to train Model B (e.g., assuming node A will use the reference model as Model A for inference). Node A can then use the reference model as Model A without further training, or node A can continue training the reference model to use as Model A. In some embodiments, multiple reference models can be established for Model A, and node B can train one or more versions of Model B corresponding to one or more of the reference models used for Model A. In embodiments with multiple reference models for Model A, node B can train one or more versions of Model B based on the multiple reference models used for Model A, and node B can indicate to node A which version of Model B it has selected to use, which version or versions of Model B provide the best performance, etc. Based on the indication from node B, node A can continue with the reference model corresponding to the Model B indicated by node B, or node A can choose any other model to use as Model A.

[0088] In the third type of framework, node A can start with model A, where model A can be in any initial state, such as pre-trained (e.g., offline training), untrained but configured with initial values, etc. Node B can start with model B, where model B can also be in any initial state. In some embodiments, node A and / or node B can have models that match each other (e.g., trained together) before training their own models. One or both nodes can train their respective models over a period of time, and then one or both nodes can share the trained model values ​​and / or the trained model with the other node. See below for reference. Figure 10 To describe the example embodiment in more detail, a first node (e.g., UE) and a second node (e.g., base station) may have a pair of models (e_0, ​​d_0), where e_0 may be an encoder model in an initial state at the UE, and d_0 may be a decoder model in an initial state at the base station. In a variant of the third type of framework, one or both nodes may train their respective models over one or more additional time periods, and one or both nodes may share trained model values ​​and / or the trained model with the other node, for example, at the end of each time period, at the end of an alternating time period, etc.

[0089] In any framework disclosed herein, when a model is transmitted to or from a node, the corresponding quantizer or dequantizer may be transmitted along with the model.

[0090] In some embodiments, training data can be collected based on a resource window (e.g., a window of time and / or frequency resources). For example, a node can be configured to collect training data (e.g., channel estimation) for a specific range of frequencies (e.g., subcarriers, subbands, etc.) and a specific range of times (e.g., symbols, time slots, frames, etc.). The size of the window can be determined, for example, based on the amount of training data that the node can store in memory. The collected training data can be used for online training by one or more nodes or saved for offline training.

[0091] In some embodiments, preprocessing and / or postprocessing can enable a pair of models to operate more efficiently. For example, domain knowledge of one or more inputs (e.g., frequency domain knowledge) can be used to perform preprocessing operations on at least a portion of one or more inputs to generate inputs of one or more transforms. Inputs of one or more transforms can be applied to a generative model to generate a representation of the inputs of one or more transforms. The representation of the inputs of one or more transforms can be applied to a reconstruction model, wherein the reconstruction model can generate reconstructed inputs of transforms (e.g., inputs of one or more transforms or approximations thereof). Domain knowledge can also be used to perform postprocessing operations (e.g., the inverse of preprocessing operations) on the inputs of reconstructed transforms to recover the original one or more inputs or approximations thereof. Depending on the implementation details, transform inputs and / or outputs (e.g., based on domain knowledge) can leverage one or more correlations between elements of one or more inputs, thereby reducing the processing burden, memory usage, power consumption, etc., of the generative and / or reconstruction models.

[0092] In some embodiments, the processing time of the model can be provided to the node. For example, if the node is configured to perform online training of the model (e.g., using a training dataset provided to or collected by the node), the node can be expected to update the model within a predetermined number of symbols or other time metrics.

[0093] According to some embodiments of this disclosure, a scheme in which multiple pairs of models can be trained, deployed, and / or activated for use by one or more nodes (e.g., by a pair of nodes) can be implemented. For example, different pairs of trained models can be activated to handle different channel environments, different matrix dimensions (e.g., for channel matrix, precoding matrix, etc.), etc. In some embodiments, a pair of models can be activated via signaling (e.g., RRC signaling, MAC-CE signaling, etc.). In some embodiments, the first node (e.g., gNB) can also instruct the second node (e.g., UE) to switch or deactivate the currently active model, for example via RRC, MAC CE, or dynamic signaling. A pair of models can be activated to train one or more of the models, use one or more of the models for inference, etc.

[0094] According to some embodiments of this disclosure, one or more formats for representing feedback information can be implemented, wherein the information can be generated by a generative model at a first node and sent to a second node for reconstruction. For example, the format for representing feedback information can be established as a type of uplink control information (UCI). The format may involve one or more types of coding (e.g., polar coding, low-density parity-check (LDPC) coding, etc.), which may depend on, for example, the type of physical channel used to transmit the UCI.

[0095] In some embodiments, CSI compression performance can be improved using AI and / or ML, for example, by leveraging one or more correlations in the time, frequency, and / or spatial domains, and / or by defining a training dataset across time, frequency, and / or space.

[0096] According to some embodiments of this disclosure, a first node (e.g., a UE) can determine precoding information that can be used by a second node (e.g., a base station). The precoding information can be used, for example, to determine channel quality information for a channel that can be used with the precoding information. Depending on the implementation details, enabling the first node to determine the precoding information used by the second node can reduce or eliminate mismatches between the precoding information and the channel quality information that can be determined based on the precoding information. For example, the base station can share a decoder model with the UE, wherein the UE can use the model to determine the precoding matrix used by the base station. As another example, the UE can train a decoder (e.g., a reference decoder described below) to reconstruct the precoding matrix used by the base station.

[0097] In some embodiments according to this disclosure, a pair of encoder and decoder models can be trained to jointly compress channel information (e.g., channel matrix, precoding matrix, etc.) and channel quality information (e.g., CQI) that can be determined based on the channel information, thereby reducing or eliminating mismatches between the channel information and the channel quality information. For example, the encoder and decoder can be trained using a training dataset, wherein the training dataset may include precoding information that can be matched with corresponding channel quality information. Depending on the implementation details, this can reduce or eliminate mismatches between the precoding information (e.g., precoding matrix) and the channel quality information (e.g., CQI) that can be determined based on the precoding information.

[0098] According to some embodiments of this disclosure, one or more machine learning models can be used to compress channel information across one or more subbands. For example, a first node can use an encoder model to generate a representation of the channel information across multiple subbands by combining (e.g., concatenating) the channel information (e.g., channel quality information) of multiple subbands into a vector and compressing said vector. A second node can use a decoder to reconstruct the channel information across the multiple subbands from said representation. Depending on the implementation details, compressing channel information across one or more subbands can improve performance, reduce complexity, etc.

[0099] According to some embodiments of this disclosure, one or more decoder models can be used to implement one or more compression schemes for generating representations of channel information and / or reporting channel information, wherein, in some implementations, the channel information may include precoded information. Depending on the implementation details, such embodiments may emulate codebook schemes while providing improved performance and / or flexibility, reduced complexity, etc. In some example embodiments, a pair of machine learning models (e.g., an encoder and a decoder) can be configured and / or trained to generate precoded information from channel information. For example, a first node (e.g., a UE) can apply channel information (e.g., one or more reference signal measurements) to an encoder, wherein the encoder can use a compression scheme to generate a representation (e.g., a codeword) based on the channel information. A second node (e.g., a base station) can apply the representation to a decoder, wherein the decoder can construct precoded information (e.g., a precoded matrix) based on the representation. In other example embodiments, a pair of machine learning models can receive any type of information, such as channel state information (e.g., channel quality information), precoded information (e.g., a precoded matrix), rank information, etc., as input, apply one or more compression schemes, and provide any type of information that can be used to determine the precoded information as output.

[0100] One or more implementations of one or more compression schemes can be achieved using one or more decoder models, providing spatial compression, frequency compression, combinations of spatial compression and frequency compression, etc. Depending on the implementation details, the compression scheme can provide separate compression for individual subbands, combined compression for individual subbands, and / or combinations thereof.

[0101] For example, a first node (e.g., a UE) may use one or more encoders to spatially compress channel information of one or more subbands to generate one or more representations (e.g., individual representations) of the channel information of one or more subbands. A second node (e.g., a base station) may use one or more decoders to recover the channel information of one or more subbands (e.g., using one or more individual representations).

[0102] As another example, the first node may use one or more spatial encoders and one or more frequency encoders to provide separate spatial compression and combined frequency compression for channel information of one or more subbands to generate one or more representations (e.g., a single representation) of the channel information of one or more subbands. The second node may use one or more spatial decoders and one or more frequency decoders to recover the channel information of one or more subbands from the one or more representations using separate spatial decompression and combined frequency decompression.

[0103] As another example, the first node can use a joint spatial and frequency encoder to provide combined spatial and frequency compression of channel information for one or more subbands to generate a combined representation of the channel information for one or more subbands. The second node can use a joint spatial and frequency decoder to recover the channel information for one or more subbands from the combined representation.

[0104] Therefore, in some embodiments, a decoder model may indicate one or more decoder models, an encoder model may indicate one or more encoder models, and a pair of models may indicate one or more decoder models and one or more encoder models.

[0105] Depending on the implementation details, embodiments that use one or more decoder models to generate precoded information and / or other information that can be used to determine the precoded information may provide improved performance and / or flexibility, reduced complexity, etc.

[0106] This disclosure includes numerous inventive principles related to artificial intelligence and machine learning for the physical layer of communication systems. These principles may have independent utility and may be embodied individually, and not every embodiment can utilize every principle. Furthermore, the principles may be embodied in various combinations, some of which may amplify the benefits of the individual principles in a synergistic manner.

[0107] For illustrative purposes, some embodiments may be described in the context of specific implementation details and / or applications (such as compression, decompression, and / or transmission of channel feedback information between one or more UEs, base stations (e.g., gNBs) in a 5G NR system). However, the principles of the invention are not limited to these details and / or applications and can be applied to any other context in which physical layer information can be processed and / or transmitted between wireless devices, regardless of whether any device can be a base station, UE, peer device, etc., and regardless of whether the channel can be a UL channel, DL channel, peer channel, etc. Furthermore, the principles of the invention can be applied to any type of wireless communication system capable of processing and / or exchanging physical layer information, such as other types of cellular networks (e.g., 4G LTE, 6G, and / or any future generation of cellular networks), Bluetooth, Wi-Fi, etc.

[0108] Machine learning models at the physical layer

[0109] Figure 1 An embodiment of a wireless communication apparatus according to the present disclosure is illustrated. Apparatus 101 may include a machine learning model 103, wherein the machine learning model 103 may receive physical layer information 105 as input and generate a representation 107 of the physical layer information as output. In some embodiments, apparatus 101 may transmit the representation 107 of the physical layer information to one or more other apparatuses, as indicated by arrow 109.

[0110] The representation 107 of the physical layer information can be a compressed, encoded, encrypted, mapped, or otherwise modified form of the physical layer information 105. Depending on the implementation details, modifying the physical layer information 105 by the machine learning model 103 to generate the representation 107 of the physical layer information can reduce the resources involved in transmitting the physical layer information 105 between devices.

[0111] The machine learning model 103 can be implemented using one or more of any type of AI and / or ML model, including neural networks (e.g., deep neural networks), linear regression, logistic regression, decision trees, linear discriminant analysis, Naive Bayes, support vector machines, learned vector quantization, etc. The machine learning model 103 can also be implemented, for example, using a generative model.

[0112] Physical layer information 105 may include any information related to the operation of the physical layer of the wireless communication device. For example, physical layer information 105 may include information related to one or more physical layer channels, signals, beams, etc. (e.g., state information, precoding information, etc.). Examples of physical layer channels may include one or more of the following: Physical Broadcast Channel (PBCH), Physical Random Access Channel (PRACH), Physical Downlink Control Channel (PDCCH), Physical Downlink Shared Channel (PDSCH), Physical Uplink Shared Channel (PUSCH), Physical Uplink Control Channel (PUCCH), Physical Sidelink Shared Channel (PSSCH), Physical Sidelink Control Channel (PSCCH), Physical Sidelink Feedback Channel (PSFCH), etc. Examples of physical layer signals may include one or more of the following: Primary Synchronization Signal (PSS), Secondary Synchronization Signal (SSS), Channel State Information Reference Signal (CSI-RS), Tracking Reference Signal (TRS), Sounding Reference Signal (SRS), etc.

[0113] Figure 2Another embodiment of a wireless communication device according to the present disclosure is shown. Device 202 may include a machine learning model 204, wherein the machine learning model 204 may receive a representation 208 of physical layer information as input and generate a reconstruction 206 of the representation 208 based on its physical layer information as output. In some embodiments, device 202 may receive the representation 208 of physical layer information from one or more other devices, as indicated by arrow 210.

[0114] Reconstruction 206 (which can be called the input to reconstruction) can be an approximation, estimation, prediction, etc., of the physical layer information upon which 208 can be based. Reconstruction 206 can be a form of decompression, decoding, decryption, reverse mapping, or other modification of the physical layer information upon which 208 can be based.

[0115] Machine learning model 204 can be implemented using one or more of any type of AI and / or ML model, including neural networks (e.g., deep neural networks), linear regression, logistic regression, decision trees, linear discriminant analysis, Naive Bayes, support vector machines, learned vector quantization, etc. Machine learning model 204 can be implemented, for example, using a reconstructed model.

[0116] The reconstructed physical layer information 206 may include any information related to the operation of the physical layer of the wireless communication device, such as, as described above regarding... Figure 1 The embodiments shown describe one or more channels, signals, etc.

[0117] Although not limited to any particular purpose, but in each Figure 1 and Figure 2 The wireless communication devices 101 and 202 shown can be used together to facilitate the transmission of physical layer information from between devices. For example, in some embodiments, device 101 can be implemented as a UE, where model 103 is implemented as a generation model, and device 202 can be implemented as a base station, where model 204 is implemented as a reconstruction model. In such an embodiment, generation model 103 can generate representation 107 by compressing physical layer information 105 (e.g., associated with a DL channel from the base station to the UE). The UE can transmit representation 107 to the base station (e.g., using a UL channel). The base station can input the representation (indicated as 208) into reconstruction model 204, where reconstruction model 204 can generate reconstructed physical layer information 206. The base station can use the reconstructed physical layer information 206, for example, to facilitate DL transmission from the base station to the UE. Depending on the implementation details, transmitting physical layer information 105 in the form of compressed representation 107 can reduce the amount of UL resources associated with transmitting physical layer information 105.

[0118] Dual-model training

[0119] Figure 3 An embodiment of the dual-model training scheme according to this disclosure is shown. Figure 3 The embodiment 300 shown can, for example, be with Figure 1 and Figure 2 It may be used with one or more models shown or any other embodiments disclosed herein.

[0120] refer to Figure 3 Training data 311 can be applied to generative model 303, which generates a representation 307 of the training data. Reconstruction model 304 can generate a reconstruction 312 of the training data based on the representation 307. In some embodiments, generative model 303 may include a quantizer for converting representation 307 into a quantized form (e.g., a bitstream) that can be transmitted via a communication channel. Similarly, in some embodiments, reconstruction model 304 may include a dequantizer, which converts the quantized representation 307 (e.g., a bitstream) into a form that can be used to generate the reconstruction training data 312.

[0121] The generative model 303 and the reconstructed model 304 can be trained as a pair, for example, by providing training feedback 314 to the generative model 303 and / or the reconstructed model 304 using a loss function 313. The training feedback 314 can be implemented, for example, using gradient descent, backpropagation, etc. In embodiments where one or both of the generative model 303 and the reconstructed model 304 can be implemented using one or more neural networks, the training feedback 314 can update one or more values ​​of weights, hyperparameters, etc., in the generative model 303 and / or the reconstructed model 304.

[0122] In some embodiments, loss function 313 (which may be implemented, for example, at least partially using reconstruction loss) can be operated to train generative model 303 and reconstructed model 304, thereby generating reconstructed training data 312 that approximates the original training data 311. This can be achieved, for example, by reducing or minimizing the loss output of loss function 313.

[0123] For example, if training data 311 is represented as x and reconstructed training data 312 is represented as x^, then the generative model 303 can be represented by the function f(x), and the reconstructed model 304 can be represented by the function g(f(x)), therefore, x^ = g(f(x)). The loss function 313 can be represented as L(x, x^). Therefore, in some embodiments, training the pair of models 303 and 304 may involve reducing or minimizing L by using training feedback 314.

[0124] While not limited to any particular type of representation 307 of the training data, in some embodiments, the pair of models 303 and 304 may seek to reduce the dimensionality of the representation 307 of the training data relative to the original training data 311. For example, generative model 303 may be trained to generate feature vectors that can identify or separate one or more features (e.g., latent features) of the training data, where the training data can reduce the overhead associated with storing and / or transmitting representation 307. Similarly, reconstruction model 304 may be trained to reconstruct the original training data 311 or an approximation thereof based on representation 307.

[0125] Once trained, the generative model 303 and / or reconstructed model 304 can be used, for example, in... Figure 1 and Figure 2 The reasoning is based on one or both of the wireless communication devices 101 and 202 shown herein, or in any other embodiment disclosed herein. Furthermore, regarding... Figure 3 The described dual-model training scheme can be used with one or more frameworks, as disclosed herein, for training models and / or transferring models between wireless devices. About Figure 3 The training described can be performed anywhere, for example, at wireless device 101, at wireless device 202, at another location (e.g., at a server located away from both devices 101 and 202), or at any combination of such locations. Furthermore, once trained, one or both of the generated model 303 and / or reconstructed model 304 can be transferred to another location for inference. In some embodiments, once trained, one of the models can be discarded, and the remaining model can be used, for example, as a pair with a separately trained model.

[0126] Machine learning model for channel information feedback

[0127] Figure 4 An embodiment of a system having a pair of models to provide channel information feedback according to this disclosure is shown. Figure 4 The system 400 shown can be used to implement any of the devices, models, training schemes, etc. disclosed herein, or can be implemented using any of the devices, models, training schemes, etc. disclosed herein, including... Figure 1 , Figure 2 and Figure 3 Those shown.

[0128] refer to Figure 4System 400 may include a first wireless device 401 and a second wireless device 402. The first wireless device 401 may be configured to receive transmissions from the second wireless device 402 via channel 415. To improve the effectiveness (e.g., efficiency, reliability, bandwidth, etc.) of transmissions via channel 415, the first wireless device 401 may provide feedback to the second wireless device 402 in the form of channel information 405, wherein the channel information 405 may be obtained, for example, by measuring one or more signals (e.g., reference signals) transmitted by the second wireless device 402 via channel 415.

[0129] The first wireless device 401 can use a first machine learning model 403 (which in this example can be implemented as a generative model) to generate a representation 407 of the channel information 405. The first wireless device 401 can, for example, use another channel, signal, etc. 416 to transmit the representation 407 to the second wireless device 402. The representation 407 can be a compressed, encoded, encrypted, mapped, or otherwise modified form of the channel information 405. Depending on the implementation details, modifying the channel information 405 by the machine learning model 403 to generate the representation 407 can reduce the resources involved in transmitting the channel information 405 to the second wireless device 402.

[0130] The second wireless device 402 can apply the representation 407 of the channel information to a second machine learning model 404, which in this example can be implemented as a reconstruction model. The reconstruction model 404 can generate a reconstruction 406 of the channel information 405. The reconstruction 406 (which can be referred to as the input to the reconstruction) can be the representation 407 based on the channel information 405, or an approximation, estimate, prediction, etc., of the channel information 405. The reconstruction 406 can be a form of decompression, decoding, decryption, reverse mapping, or other modification of the channel information 405. The second wireless device 402 can use the channel information 405 to improve its method of transmitting to the first wireless device 401 via channel 415.

[0131] Figure 4 The system 400 shown is not limited to any particular device (e.g., UE, base station, peer device, etc.), application (e.g., 4G, 5G, 6G, Wi-Fi, Bluetooth, etc.), and / or implementation details. However, for the purpose of illustrating some inventive principles, some example embodiments can be described in the context of a 5G NR system, in which the UE can receive different DL signals from the gNB.

[0132] Uplink and downlink transmission

[0133] In an NR system, the UE can receive DL transmissions from the gNB, including various information. For example, the UE can receive user data from the gNB in ​​a specific configuration of time and frequency resources called the Physical Downlink Shared Channel (PDSCH). The Multiple Access (MAC) layer at the gNB can provide user data intended to be delivered to the corresponding MAC layer at the UE side. The UE's physical (PHY) layer can receive the physical signals received on the PDSCH and apply them as input to the PDSCH processing chain, where the output of the PDSCH processing chain can be fed as input to the UE's MAC layer. Similarly, the UE can receive control data from the gNB using the Physical Downlink Control Channel (PDCCH). The control data can be called Downlink Control Information (DCI) and can be converted into PDCCH signals through the PDCCH processing chain at the gNB side.

[0134] The UE can transmit UL signals to the gNB using the Physical Uplink Shared Channel (PUSCH) and the Physical Uplink Control Channel (PUCCH) respectively to transmit user data and control information. The PUSCH can be used by the UE MAC layer to deliver data to the gNB. The PUCCH can be used to transmit control information, which can be called uplink control information (UCI), where UCI can be converted into PUCCH signals through the PUCCH processing chain on the UE side.

[0135] Channel state information

[0136] In an NR system, a UE may include a Channel State Information (CSI) generator, which can calculate a Channel Quality Indicator (CQI), a Precoding Matrix Indicator (PMI), a CSI Reference Signal Resource Indicator (CRI), and / or a Rank Indicator (RI), any one or all of which can be reported to one or more gNBs serving the UE. The CQI may be associated with a modulation and coding scheme (MCS) used for adaptive modulation and coding and / or frequency-selective resource allocation, the PMI may be used for channel-dependent closed-loop multiple-input multiple-output systems, and the RI may correspond to the number of useful transport layers.

[0137] In NR systems, CSI generation can be performed based on the CSI reference signal (CSI-RS) transmitted by the gNB. The UE can use the CSI-RS to measure downlink channel conditions and, for example, perform channel estimation and / or noise variance estimation to generate the CSI by measurement based on the CSI-RS signal.

[0138] In NR systems, CSI can be reported to the serving gNB using a Type I codebook, which can provide implicit CSI feedback to the gNB in ​​the form of an index pointing to a predefined PMI. Alternatively or additionally, CSI can be reported to the serving gNB using a Type II codebook, which can provide explicit CSI feedback. The UE can determine one or more dominant eigenvectors or singular vectors based on DL channel conditions. The UE can then use the dominant eigenvector or singular vector to derive the PMI, which can be fed back to the gNB, whereby the gNB can use the PMI for beamforming in the DL channel.

[0139] The use of a codebook can provide sufficient performance, for example, in embodiments with a limited number of antenna ports and / or users. However, in systems with a larger number of antenna ports and / or users (e.g., multiple-input multiple-output (MIMO) systems), and especially when using frequency division duplex (FDD), the relatively low resolution of a Type I codebook may not provide sufficiently accurate CSI feedback. Furthermore, the use of a Type II codebook may still involve transmitting a significant amount of overhead data on the UL channel.

[0140] Depending on the implementation details, some embodiments of the machine learning-based channel information feedback scheme of this disclosure can enable the UE to send full CSI information to the gNB while reducing the overhead associated with UL transmission to the gNB. Furthermore, the principles of this invention are not limited to the UE sending CSI to the gNB, but can be applied to any situation where the first device can send channel information feedback to the second device (e.g., reporting channel conditions of the uplink channel from the UE to the gNB, reporting channel conditions of the sidelink channel between UEs, etc.).

[0141] Example Implementation

[0142] Figure 5An example embodiment of a system for reporting downlink physical layer information according to this disclosure is shown. System 500 may include UE 501 (which may be designated as Node B) and gNB 502 (which may be designated as Node A). gNB 502 may transmit a transmission of DL signal 517 (e.g., reference signal (RS) transmission) to UE 501, from which UE 501 may extract measurement 518. UE 501 may include model 503, which may be configured, for example, as an encoder, to encode measurement 518 into a feature vector associated with the DL physical layer. The encoded measurement may then be quantized by quantizer 519 and sent back to gNB 502 as UL signal 520 (e.g., bitstream). In some embodiments, the description of the model at a node may also include quantizer and / or dequantizer descriptions, for example, mapping channel information (e.g., real CSI codewords) at the output of the encoder model to a function of quantized values ​​or bitstream, and vice versa at the decoder model at another node. gNB 502 can apply the received UL signal 520 to dequantizer 521 to generate an equivalent feature vector, wherein the equivalent feature vector can be fed into model 504 to extract information 522 related to the DL physical layer (e.g., necessary or optional information).

[0143] Figure 6 Example embodiments of a system for reporting uplink physical layer information according to this disclosure are shown. In some aspects, Figure 6 The system 600 shown in the figure can be similar to Figure 5 The system shown is 500, but system 600 can be configured to report uplink physical layer information instead of downlink physical layer information.

[0144] Specifically, system 600 may include gNB 601 (which may be designated as Node B) and UE 602 (which may be designated as Node A). UE 602 may transmit a UL signal 617 (e.g., a reference signal (RS) transmission) to gNB 601, from which gNB 601 can extract a measurement 618. gNB 601 may include model 603, which may be configured, for example, as an encoder, to encode the measurement 618 into a feature vector associated with the UL physical layer. The encoded measurement may then be quantized by quantizer 619 and sent back to UE 602 as a DL signal 620 (e.g., a bitstream). UE 602 may apply the received DL signal 620 to dequantizer 621 to generate an equivalent feature vector, which may be fed to model 604 to extract information 622 (e.g., necessary or optional information) associated with the UL physical layer.

[0145] Figure 7 An example embodiment of a system for reporting downlink physical layer channel state information according to this disclosure is shown. Details depend on the implementation. Figure 7 The system 700 shown can enable a gNB or other base station to retrieve complete CSI information from the UE (e.g., as opposed to codebook-based pointers, precoded matrix indicators, etc.), while using an ML model to compress the CSI (e.g., compress it to a relatively small number of bits), thereby reducing the uplink resource overhead involved in transmitting the CSI.

[0146] System 700 may include UE 701 and gNB 702. gNB 702 may transmit a DL reference signal 717, such as CSI-RS or demodulation reference signal (DMRS), which enables UE 701 to determine CSI 718 for DL ​​channel 715. UE 701 may include an ML model 703, wherein the ML model 703 may be configured as an encoder to encode CSI 718 into a feature vector. UE 701 may also include a quantizer 719, wherein the quantizer 719 may quantize the feature vector to a bitstream that can be transmitted to gNB 702 using UL signal 720. gNB 702 may include a dequantizer 721, which may reconstruct the feature vector from the bitstream. The feature vector may then be fed into an ML model 704, wherein the ML model 704 may be configured as a decoder to reconstruct an estimate 722 of CSI 718.

[0147] In some embodiments, the performance metric f(H,H^) can be used to evaluate the accuracy of the design, configuration, and / or training of the encoder model 703, decoder model 704, quantizer 719, and / or dequantizer 721. For example, the performance metric f(H,H^) can be implemented as a measurement of the error between channel estimates as follows:

[0148] f(H,H^)=‖HH∧‖^2 / ‖H‖^2(Equation 1)

[0149] Here, H and H∧ can represent channel estimation (e.g., CSI) at UE 701 and gNB 702, respectively. Such performance metrics can be useful, for example, for evaluating the accuracy of channel state information extracted by gNB 702.

[0150] Additionally or alternatively, system 700 can be configured such that UE 701 can use DL reference signal 717 to determine the precoding matrix based on current channel conditions. The precoding matrix can then be encoded into features by encoder model 703, quantized by quantizer 719, and transmitted to gNB 702 using UL signal 720. At gNB 702, dequantizer 721 can recover the feature vectors, which can be applied to decoder model 704 to reconstruct an estimate of the precoding matrix. For example, for a channel implementation H, a suitable precoding matrix can be implemented as a set S of singular vectors using singular value decomposition (SVD) of H, where H can be given as H = SΣD, where Σ can be a diagonal matrix and D can be a unitary matrix. In such an embodiment, encoder model 703, decoder model 704, quantizer 719, and / or dequantizer 721 can be configured such that gNB 702 can extract the set (e.g., matrix) S of singular vectors, and a performance metric can be implemented accordingly. Although Figure 7 The embodiments shown report downlink physical layer information, but other embodiments may be configured to report uplink physical layer information, sidelink physical information, etc., using similar principles according to this disclosure.

[0151] Model development, training and operation

[0152] Artificial intelligence (AI), machine learning (ML), deep learning, etc. (as described above, any or all of which can be individually and / or collectively referred to as machine learning or ML) can provide techniques for inferring one or more functions (e.g., complex functions) from data according to this disclosure. In a machine learning process, samples of data can be provided to an ML model, which can then apply one of various machine learning techniques to learn how to determine one or more functions using the provided data samples. For example, a machine learning process can allow an ML model to learn a function f(x) of the data sample input x. As described above, an ML model can also be referred to as a model.

[0153] In some embodiments, the machine learning process (which may also be referred to as the development process) may be carried out in one or more phases (which may also be referred to as periods), such as training, validation, testing, and / or inference (which may also be referred to as application phases). Some embodiments may omit one or more of these phases and / or include one or more additional phases. In some embodiments, all or part of one or more phases may be combined into one phase, and a phase may be divided into multiple phases. Furthermore, the order of phases or their parts may be changed.

[0154] During the training phase, a model can be trained to perform one or more target tasks. The training phase may involve using a training dataset, which may include i) data samples and ii) the outcome of a function f(x) for each sample in the training dataset. During the training phase, one or more training techniques may enable the model to learn approximate relationships (e.g., approximate functions), where the approximate relationship may represent a function f(x) or closely follow a function f(x).

[0155] During the validation phase, the model can be tested (e.g., after initial training) to evaluate its suitability for one or more target tasks. If the validation results are unsatisfactory, the model can undergo further training. If the validation phase provides successful results, the training phase can be considered successfully completed.

[0156] During the testing phase, the trained model can be tested to evaluate its suitability for one or more target tasks. In some embodiments, the training of the model may not proceed to the testing phase unless training is complete and successful results are verified.

[0157] During the inference phase, a trained model (e.g., in real-world applications) is used to perform one or more target tasks.

[0158] During the testing and / or inference phases, the model can use the approximate function learned during the training phase to determine the function value f(x) for other data samples that are different from the samples in the training phase.

[0159] In some embodiments, the success and / or performance of the machine learning process may involve using a sufficiently large training dataset, wherein the training dataset may contain enough information about the function f(x) and thus enable the model to obtain an acceptablely close approximation of the function f(x) through the training phase.

[0160] Figure 8 An embodiment of the learning process of a machine learning model according to the present disclosure is shown. Process 800 may begin at operation 823, where the training process may be initialized. For example, the structure of the model may be determined, the values ​​of the model (e.g., neural network weights, hyperparameters, etc.) may be initialized, and a training dataset with a sufficient number of samples may be constructed.

[0161] At operation 824, the training dataset can be used to train the initial model to determine the configuration of the candidate trained model, for example, by updating the values ​​of neural network weights, hyperparameters, etc., using gradient descent, backpropagation, etc.

[0162] In some embodiments, there can be an interrelationship between the construction of the training dataset and the training phase. For example, the training phase can involve a relatively long duration to complete, and the duration can depend on the number of samples in the training dataset. The duration can further depend on the type of training. For example, for full training and / or initial training, the model can be initialized, and training can be performed using a large dataset that may consist of many samples (e.g., samples that may not have been previously used to train the model). As another example, for partial training and / or update training, the model can be previously trained (or partially trained), and events (e.g., acquisition of new data samples, performance degradation of the model, model update events, etc.) can prompt modifications or adaptations to the model. In the case of partial training and / or update training, the model can be trained using a modified dataset that may differ from the large training dataset used for full training and / or initial training. For example, the modified training dataset can be a subset of the full dataset used for initial training, a newly acquired set of data samples, or a combination thereof.

[0163] At operation 825, the trained candidate model can be validated. In some embodiments, validation phase 825 can be performed iteratively with training phase 824. For example, if candidate model validation phase 825 fails, it can return to training phase 824, where training phase 824 can generate new candidate models. In some embodiments, different criteria (e.g., classification accuracy, minimum mean squared error (MMSE), etc.) can be established for determining validation success or failure.

[0164] In some embodiments, failing candidate models may not be allowed to return to training phase 824 (e.g., after more than a threshold number of failures, or if the performance criterion fails to pass the threshold within a certain duration or a certain number of validation steps), and the method may terminate at operation 826. However, if the performance of the candidate model using the validation data is determined to be acceptable (e.g., based on the criteria used to determine success or failure), the validation can be considered successful, and the trained candidate model can be passed to the testing phase at operation 827.

[0165] At operation 827, the performance of the trained model candidate that has passed the validation phase can be evaluated. The criteria used to declare the model's successful testing and / or failure during the testing phase of development can be similar to those used in the validation phase. However, one or more parameters used with the criteria during the testing phase (e.g., the number of steps, performance thresholds, etc.) may or may not differ from the parameters used in the validation phase.

[0166] If the test is successful, the model can be designated as the final model, and the process can proceed to operation 828. In some embodiments, if the model testing phase fails, the process can return to the training phase at operation 824 for further training. However, in some embodiments, further training may not be allowed (e.g., based on standards similar to those used during the validation phase 825), and the process can terminate at operation 826.

[0167] Model training and deployment framework

[0168] According to some embodiments of this disclosure, one or more frameworks for training and / or deploying models can be implemented. In some embodiments of the frameworks disclosed herein, one or more models trained and / or developed by a node can be tested against one or more reference models of a node, for example, to evaluate the model's compliance with one or more potential test cases that can be specified for a corresponding application.

[0169] In any embodiment of the framework disclosed herein, the quantizer function may be differentiable with a substantially zero derivative (e.g., a probability of 1) over some or all of the quantizer range (e.g., substantially over the entire range). Depending on the implementation details, this may result in backpropagation providing little or no update to the encoder weights. Thus, in some embodiments, the quantizer function may be approximated using a differentiable function (e.g., a reference to a differentiable quantizer function), wherein the differentiable function may be referred to as f_(quantizer,approx)(x) during the training phase, while the actual quantizer function may be used during the inference phase. Similarly, the dequantizer function may be approximated using a differentiable function (e.g., a reference to a differentiable dequantizer function), wherein the differentiable function may be referred to as f_(dequantizer,approx)(x) during the training phase, while the actual dequantizer function may be used during the inference phase. In some embodiments, the quantizer or dequantizer function used in conjunction with the model may be considered as part of the complete description of the corresponding model and may be transmitted with and as part of the model. Therefore, using any framework disclosed herein, if the first node shares a trained model with the second node (e.g., if node A trains a pair of models (model A and model B), then the trained model B is sent to node B), the first model can also share one or both of the approximate quantizer function f_(quantizer,approx)(x) and / or the approximate dequantizer function f_(dequantizer,approx)(x) with the second node, for example, via RRC signaling.

[0170] While the framework disclosed herein is not limited to any particular application and / or implementation details, in some embodiments, and depending on the implementation details, the framework can be used to train and / or test models that can reduce CSI feedback overhead.

[0171] Joint training framework

[0172] In some embodiments, a pair of models (e.g., model A and model B) can be jointly trained by one of two nodes (node ​​A or node B), and the trained model from the non-training node can be transferred to the non-training node (e.g., if joint training is performed by node A, then trained model B can be transferred to node B) for inference. For example, in the context of CSI compression, a base station can perform joint training of a pair of encoder and decoder models, and then pass the encoder model to the UE. The encoder model can also be referred to as an encoder, and the decoder model can also be referred to as a decoder.

[0173] In some embodiments of the joint training framework, further training (e.g., fine-tuning) of one or both of the trained models can be performed by nodes, where the models can be used for inference (e.g., to improve or optimize one or both of the models). In some embodiments, further training can be based on online data that can be obtained by one or more of the nodes, for example, during ongoing communication.

[0174] In some embodiments of the joint training framework, training nodes can use corresponding quantizer and / or dequantizer functions (e.g., approximate and / or differentiable quantizer and / or dequantizer functions) to train one or both models. Nodes receiving the trained models can also receive and use the corresponding quantizer and / or dequantizer functions for further training, validation, testing, inference, etc.

[0175] In some implementations, joint training of the model by nodes can produce a model that can be jointly matched with the target task, and thus can provide improved or optimized performance. Depending on the details of the implementation, such performance improvement can outweigh any communication overhead associated with sending the model to different nodes, and / or any mismatch between the model and / or nodes, such as that caused by joint training at a node that may be produced by a different manufacturer than the other node.

[0176] In a variant of the joint training framework, a node (e.g., a base station) can jointly train a pair of encoder and decoder models. The encoder and decoder pair can be trained, for example, using the reference differentiable quantizer and dequantizer functions described above. The base station can then share the trained decoder model with the UE, for example, via RRC signaling. However, the base station may or may not share the trained encoder model with the UE. If the base station shares the trained encoder model with the UE, the UE can use the trained encoder model as a reference encoder model. If the base station does not share the trained encoder model with the UE, the UE can build a reference encoder model based on, for example, randomly initialized weights, weights selectable for the UE implementation, or any other basis.

[0177] The UE can then use its trained decoder model received from the base station to train a reference encoder model. The reference encoder model can be trained online (which may indicate training that can be performed during operation). In some implementations, online training can be performed in operation (which may indicate training performed using training data (e.g., channel estimates H) that can be collected during operation). Therefore, the UE can use channel estimates H that can be collected over time to train the reference encoder model. The collected channel estimates can be used as new training datasets, for example, at specific points during training. Furthermore, the collected channel estimates can be stored for future online and / or offline training by the UE or any other device.

[0178] The UE can then use the trained encoder model for inference. The UE can also share the trained encoder model with the base station. This training process can continue as more training samples (e.g., channel estimates H) are collected by the UE and used for training.

[0179] Figure 9An example embodiment of a method for jointly training a pair of encoder and decoder models according to this disclosure is shown. At operation 929, the base station can use a training dataset, which may be referred to as Enc_ref and Dec_ref, to jointly train a pair of reference encoder and decoder models. At operation 930, the base station may share the reference decoder model Dec_ref with the UE. At operation 931, the base station decides whether to share the reference encoder model with the UE. If the base station shares the reference encoder with the UE, then at operation 932, the UE can use the shared reference encoder as a reference encoder for training. If the base station does not share the reference encoder with the UE, then at operation 933, the UE can, for example, use random weights, use weights based on the UE implementation, etc., to build the reference encoder model. At operation 934, the UE can train the reference encoder model at time point t_i. For example, time point t_i can be determined as the time since the UE has performed and collected sufficient channel estimates. This can be implemented, for example, as shown in Algorithm 1, where, for each time point t_i, the UE may have already collected a new online training set S_i during operation, where S_i may include channel estimates from t_(i-1) to t_i, and N may be the maximum number of online training sets on the UE side. In some implementations, after completing Algorithm 1, the UE may share the trained encoder model with the base station.

[0180] Algorithm 1

[0181]

[0182] Any training and deployment framework disclosed herein can be used with any type and / or combination of devices and any type of model and / or physical layer information. For example, even in Figure 8 In the illustrated embodiment, the base station performs initial joint training, and the encoder and decoder can be trained and used with channel estimation. However, in other embodiments, joint training can be performed by the UE or any other device, and the model can be trained and used with a precoding matrix or any other type of physical layer information.

[0183] In some embodiments, nodes such as UEs or base stations may collect new training data within a window (e.g., an explicit time and / or frequency window). The collected data can be used, for example, to construct a training dataset that the node can use to train a model. The window may be configured with start and / or end times that may be determined, for example, by the base station. In some embodiments, a timeline for determining the data collection window may be measured, for example, from one or more CSI-RS resources.

[0184] Alternatively or additionally, online training can be performed as follows: A first node (base station) may have a first model, and a second node may have a second model that can be paired with the first model. In some embodiments, the first and second nodes may operate in a connectivity mode (e.g., RRC connectivity mode). One or both of the nodes may have already obtained their models by sharing them with another node. In this example, one of the nodes may be a base station, and the other node may be a UE.

[0185] The base station can configure the UE using a predetermined online training dataset, and both nodes can use the predetermined online training dataset to update their respective models. When a node updates its own model, it can be assumed that another model at the other node is frozen. In some embodiments, one or more online training datasets can be specified (e.g., as part of a specification and / or provided to the UE and / or base station by a third node). Once the first node updates its first model (e.g., encoder or decoder), it can share the updated first model with the second node, and the second node can begin training its second model assuming the first model is frozen. The models can continue to alternate between periodically training and freezing their models, e.g., until an end time is reached.

[0186] Although the embodiments disclosed above can be described in the context that the UE can update and use the encoder, in some embodiments with online training, the UE and the base station (or any two other nodes with a pair of models, such as two UEs configured for sideband communication) can collect new training data and use it to update their own models (e.g., encoders or decoders) or both models (e.g., both encoders and decoders that can be configured as, for example, autoencoders).

[0187] In some embodiments, the first node can share newly collected training data (e.g., channel matrix) with the second node by sending the training data as data or control information. For example, the UE can generate a binary representation of one or more channel matrices and send the representation using PUSCH or PUCCH that follows the normal procedures for uplink transmission, coding, modulation, etc.

[0188] Alternatively or additionally, the UE can use its currently trained encoder to encode the channel matrix it has already acquired. The UE can send the encoded channel matrix (which may be referred to as CSI codewords) to the base station. The base station can then use its currently trained decoder to recover the channel matrix. The base station can then include the recovered channel matrix in a new training dataset, which can be used for further (e.g., online) training at the base station.

[0189] In some embodiments, in addition to exchanging training data between nodes, one or more nodes may also share their latest trained model (e.g., encoder and / or decoder) with another node. The sharing of training data and / or model may be performed at intervals (e.g., the model may be sent when it is updated), wherein the intervals may manage the communication overhead involved in such sharing.

[0190] In a framework with online model training, nodes can use memory buffers to store collected physical layer information (e.g., the CSI matrix from time t_(i-1) to time t_i). Depending on the implementation details, a node may collect some or all of the new training dataset before it can begin updating the model with new training data. However, if a node uses a dedicated memory buffer to store the training dataset, and the interval between time t_(i-1) and time t_i exceeds a certain value, the amount of memory involved in storing the training data may exceed the available buffer size because the number of CSI-RS within the time window may become too large. Furthermore, even if the interval between time t_(i-1) and time t_i is generally short enough to prevent buffer overflow, the node may encounter some reference signals with relatively short periodicity (e.g., a relatively large number of reference signals (e.g., CSI-RS) can be configured in the window), and therefore, the CSI matrix collected based on the reference signals may exceed the available dedicated memory buffer.

[0191] In some embodiments, a node may declare a data buffer capacity, which may be related, for example, to the size of the training dataset constructed from the collected training data. Depending on the implementation details, this can reduce or prevent problems of exceeding the capacity of the memory buffer used for new training data. For example, a node may declare or be assigned a predetermined memory buffer capacity based on: (1) the time interval for obtaining training data and / or updating the model based on the obtained training data (e.g., the maximum time interval); (2) the maximum number of reference signals (e.g., CSI-RS) expected by the node within the time window for constructing the training set; or (3) the shortened periodicity of the reference signals (e.g., CSI-RS) used to construct the training dataset. A node may be configured such that one or more reference signals and / or time windows that could potentially violate the predetermined memory buffer capacity are considered error cases.

[0192] Alternatively or additionally, default behavior can be defined when a violation of a node's predetermined memory buffer capacity occurs. For example, if the configuration of the reference signal and / or time window violates the node's memory buffer capacity, the node may update the model by storing and / or using only a subset of the collected training data. For example, if the UE reports a maximum of N_max CSI-RS within a window, and the gNB configures a larger number of N_(CSI-RS) CSI-RS within a window, the UE may update the model using only N_max CSI-RS out of the N_(CSI-RS) CSI-RS. How the UE selects which CSI-RS to use can be determined based on the UE implementation and / or according to one or more configured and / or fixed rules (e.g., the UE may use the latest N_max resources out of the N_(CSI-RS) resources).

[0193] In some embodiments, the size of the buffer used to collect training data can be implemented based on nodes, for example, without regard to specifications. For example, if the UE's training data buffer overflows, the UE can stop storing newly collected data (e.g., matrices) and continue using the data in the buffer to update the model. In some embodiments, once the model is updated, the UE can refresh the buffer and then start collecting new training data again.

[0194] In some embodiments, the UE may use a shared buffer to store new training data. Examples of shared buffers may include one or more buffers already used to store other channels, such as a PDSCH buffer, a main CCE LLR buffer, etc. In such embodiments, shared buffer space may be used based on availability, as it may already be fully or partially occupied for other dedicated uses. In some embodiments, the buffer for collected training data may be implemented on a node-based basis.

[0195] Training framework with reference model

[0196] In some frameworks according to this disclosure, a reference model can be established as Model A for node A, and node B can then use the reference model as Model A to train Model B (e.g., assuming node A will use the reference model as Model A for inference). Node A can then use the reference model as Model A without further training, or node A can continue training the reference model to use as Model A. In some embodiments, one or more Model A's can be provided and / or assigned to node A, and node B can use Model A at node B to train one or more Model B's, wherein Model A can be assumed to be one or more of node A's reference models. For example, node B can use Model A, assumed to be a reference model assigned to Model A, to train a first version of Model B. Node B can also assume Model A is another reference model to train a second version of Model B, and so on.

[0197] A reference model can be established, for example, through specifications, signaling (e.g., RRC signaling from the base station to the UE after the UE is connected by RRC).

[0198] In some embodiments, node B can notify node A which reference models it has selected to use to train different versions of model A. If only one reference model is available for model A, no communication is required because the reference model can be implicitly known. Node B can notify node A of one of a plurality of reference models available for training versions of model B; this model may correspond to, for example, the reference model that provides the best performance. Alternatively or additionally, node B can notify node A of a subset of the plurality of reference models; this subset may include the set of reference models with the best performance.

[0199] Regardless of any signaling from node B to node A, node A may or may not indicate to node B which reference model it has selected. Indicating the reference model can be useful, for example, to establish a common understanding between node A and node B, while not indicating the reference model can reduce signaling overhead. In implementations with multiple reference models, if a subset of the optimal performance models includes only one reference model (e.g., only one reference model is indicated from node B to node A as the optimal performance reference model), then node A may not provide an indication to node B, because node A's selection can be implicitly known by node B.

[0200] Once a reference model is established for node A, node A can use the reference model as model A or continue training the reference model. Depending on the implementation details, using the reference model as model A (e.g., with little or no further training or tuning) can provide a relatively high level of matching (e.g., best match) between the two models because node B can train model B assuming the reference model is used for model A. If multiple reference models exist that node B uses to train different versions of model B, node A can use any of the versions of trained model B; this can involve establishing a shared understanding between node A and node B about which version of trained model B will be used (e.g., node B can inform node A which version of trained model B is being used, or node A can inform node B which model to use).

[0201] Node A can continue training model A instead of using the reference model as model A without further training. This could be beneficial, for example, if the reference model is not suitable for the current network state (e.g., a wireless environment if the model is intended for CSI compression and decompression). Therefore, allowing node A to further train (e.g., tune or optimize) model A allows the model to fit the current network state. However, changing model A from a reference model assumed by node B when training model B can lead to a potential mismatch between the two models, which in turn can result in performance degradation.

[0202] In some embodiments, model A can be trained to overcome this potential mismatch. For example, to train model A, node B can send model B to node A, so the training of model A can be based on the actual model used by node B as model B.

[0203] If multiple trained versions of model B exist, node B can transmit a subset of these trained versions of model B, and node A can train multiple corresponding models A for each transmitted version of model B. In such an embodiment, models A and B can communicate to establish a shared understanding about which pair of models A and B can be selected. Depending on the implementation details, sharing multiple versions of model A can allow node A and / or node B to improve (e.g., optimize) performance by selecting the best pair of models A and B that may offer the best performance among the transmitted models. Alternatively, to reduce communication overhead, node B can transmit one of the multiple versions of model B, and node A can train a model A corresponding to the transmitted version of model B.

[0204] Alternatively or additionally, if node A continues training model A, node A may train experimental versions of model B to mimic the actual model B used by node B. The level of similarity between the experimental model B and the actual model B may depend on the design and / or architecture of model B, the training dataset used to train the experimental versions of model B, and / or the training procedure used to train the experimental versions of model B (e.g., initialization of weights, hyperparameters, etc.). If multiple trained versions of model B already exist, node A may train multiple corresponding experimental versions of model B. Alternatively, node A may train multiple models A using experimental versions of model B corresponding to each of the available reference models for model A; this can be particularly useful because it allows node A to train model A before being notified by node B which reference models(s) have been selected. In such embodiments, models A and B can communicate to establish a shared understanding of which pair of models A and B can be selected for use.

[0205] To further reduce the mismatch between the experimental model B and the actual model B, node B can share some auxiliary information with node A. Depending on the implementation details, sharing auxiliary information can help node A train the experimental model B in a manner that will produce an experimental model B similar to the actual model B. Examples of auxiliary information may include initialization values ​​(e.g., a random seed used by node B to train the actual model B, initial network weights, etc.), one or more optimization algorithms, one or more algorithms for feature selection, one or more algorithms for data preprocessing, information about the type of neural network (e.g., recurrent neural network (RNN), convolutional neural network (CNN), etc.), information about the structure of the model (e.g., the number of layers, the number of nodes per layer, etc.), information about the training dataset, etc. The use of this information can be enforced (e.g., via specification) or left to the implementation of the node.

[0206] In some embodiments, reference models for node A and / or node B may be specified (e.g., in a specification), for example, for testing purposes. Such embodiments may not involve any indication of which model is used by node A and / or node B. For example, when the gNB uses one or more reference models, the UE may be expected to meet one or more performance specifications. Depending on the implementation details, this can provide deployment guidelines regarding which models the nodes should use to achieve appropriate performance for machine learning tasks. In some embodiments, one or more performance requirements may be established for machine learning CSI compression tasks, for example, as part of a specification.

[0207] In some embodiments of the framework with a reference model, nodes can use corresponding quantizer and / or dequantizer functions (e.g., approximate and / or differentiable quantizer and / or dequantizer functions) to train any model including the reference model, and any model can also use corresponding quantizer and / or dequantizer functions for further training, validation, testing, inference, etc.

[0208] Training framework with the latest shared values

[0209] In some frameworks according to this disclosure, node A can start with model A, where model A can be in any initial state (e.g., pre-trained (e.g., offline training), untrained but configured with initial values, etc.). Node B can also start with model B, which can also be in any initial state. One or both nodes can train their respective models over a period of time (which may be referred to as a training cycle or iteration), and then one or both nodes may or may not share the trained model values ​​and / or the trained model with the other node. In some embodiments, a new training dataset may be provided directly or indirectly to one or both nodes, for example, at the beginning or end of a cycle. Nodes A and B can train their respective models using their latest knowledge of the weights of the model at the other node, for example, without any model swapping.

[0210] The first node can assume that the model at the second node (e.g., the decoder at the base station) is frozen to train its model using the latest weights (e.g., fed back from the second node) (e.g., the UE can train the encoder). The first node can train its model and update its weights to, for example, a maximum number of times (e.g., K_e times for the encoder), and then share the updated model weights with the second node. The same procedure can be implemented at the second node. Specifically, once the second node has received the updated model weights from the first node, the second node can assume that the model weights of the model at the first node are frozen at the latest state of the model shared by the first node, train its model and update its weights to the maximum number of times (e.g., K_d times for the decoder). The second node can then share its updated model weights with the first node. Thus, the first and / or second nodes may have trained their respective models to the maximum number of times and then shared the updated model values ​​with another node (this can be referred to as sharing a period or iteration).

[0211] In a variation of this framework, after one or more nodes share model state information (e.g., weights) with another node, for example, at the end of a sharing cycle, one or both nodes can begin another sharing cycle. For example, assuming the values ​​of the model on the other node are frozen to the latest values ​​shared by the other node, both nodes can train their models. At some point in time, or after a certain number of training cycles performed by the first and / or second node (e.g., at the end of another sharing cycle), one or both nodes can stop training and share their latest trained model with each other node. In some embodiments, at the beginning, a shared model (e.g., a fully shared model that can be initialized, for example, through offline training, handshakes, etc.) can be used as the initial value for the latest shared weights.

[0212] Figure 10 An example embodiment of a method for training a model using the latest shared values ​​according to this disclosure is shown. For illustrative purposes, it can be described in the context of a UE having an encoding model for CSI and a base station having a decoding model for CSI. Figure 10 The method shown is applicable to any type of node and / or physical layer information, but the principle can be applied to any type of node and / or physical layer information.

[0213] refer to Figure 10 At the start of the first shared period 1035-1, the encoder model can be in the initial state e_0, and the decoder model can be in the initial state d_0, as shown at shared point 1036-0. Both the encoder and decoder models in the initial states (e_0, ​​d_0) can be provided to the UE and the base station. Therefore, both the UE and the base station start with encoder and decoder models in the same initial states. The UE can then perform M training periods (e.g., training its encoder M times while its decoder model remains in the initial state d_0). While the UE is performing M training periods, the base station can perform N training periods (e.g., training its encoder N times while its encoder model remains in the initial state e_0).

[0214] For example, after the first training cycle performed by the UE, the encoder and decoder model of the UE can have state (e_1, d_0), after the second training cycle performed by the UE, the encoder and decoder model of the UE can have state (e_2, d_0), and so on, until after the Mth training cycle, the encoder and decoder model of the UE can have state (e_M, d_0).

[0215] Similarly, after the first training cycle performed by the base station, the encoder and decoder models of the base station can have states (e_0, ​​d_1), after the second training cycle performed by the base station, the encoder and decoder models of the base station can have states (e_0, ​​d_2), and so on, until after the Nth training cycle, the encoder and decoder models of the base station can have states (e_0, ​​d_N).

[0216] At the sharing point 1036-1 at the end of the sharing period 1035-1, the UE can send its trained encoder model to the base station, and the base station can send its trained decoder model to the UE. Therefore, both the UE and the base station can have encoder and decoder models with states (e_M, d_N).

[0217] In some embodiments, the UE and / or base station may stop training at this point and begin inference using their trained encoder and decoder models. However, in some other embodiments, one or both of the UE and / or base station may begin another shared cycle 1035-2. For example, the UE may then perform P training cycles by training its encoder up to P times while its decoder model remains in state d_N, and the base station may perform Q training cycles by training its decoder up to Q times while its encoder model remains in state e_M.

[0218] At the sharing point 1036-2 at the end of the sharing period 1035-2, the UE can send its trained encoder model to the base station, and the base station can send its trained decoder model to the UE. Therefore, both the UE and the base station can have encoder and decoder models with states (e_P, d_Q). The UE and / or the base station can execute any number of sharing periods, and any number of training periods per sharing period.

[0219] Figure 10A particular case of the illustrated embodiment is when M or N = 0, if M >> N, or if N >> M. For example, in the case of N = 0 and M > 0, the base station may not update the decoder model during the sharing period (e.g., no training period may be performed). However, the UE may train its encoder up to M times before sharing its encoder with the base station. Similarly, in the case of M = 0 and N > 0, the UE may not update its encoder model, while the base station may update its decoder model up to N times before sharing its decoder model with the UE. Depending on the implementation details, one or more of these particular cases may be advantageous, for example, if one of the nodes has difficulty or no access to the training dataset for online training at the node. In this case, the node that has access to (or easier access to) the training data can continue online training, which can enable the node continuing training to provide the trained model to another node that has no access to or limited access to the training data.

[0220] Special cases can also be implemented in an interleaved and / or alternating manner. For example, two nodes can begin with one variable M or N equal to zero and the other variable greater than zero. Once the model with an applicable model has been updated the number of times determined by the non-zero variable and shared with the other node, the non-zero variable can take a zero value, while the other variable becomes non-zero. This process can continue with M and N taking zero values ​​alternately. Such an interleaved training process allows the first node (e.g., UE or gNB) to train its model (e.g., encoder or decoder) multiple times while the model at the second node is frozen. Then, after the first node shares its trained model with the second node, the second node can train its model multiple times while the model at the first node is frozen, and so on.

[0221] In some embodiments, the values ​​of non-zero variables can affect the performance of a trained pair of models. For example, if the time between shared points is relatively large, the trained model (e_M, d_N) may have relatively poor performance, for instance, if each node has been trained assuming that the model weights on the other side may be significantly different from the model that will be shared at the next shared point. Figure 10 In the illustrated embodiment, the base station may assume the encoder has weights e_0 to train its decoder, and may later pair the trained decoder with a new encoder model e_M that may have deviated significantly from e_0. Therefore, in some embodiments, sharing the model at a relatively high frequency can improve the performance of the trained model.

[0222] Within any framework disclosed herein, nodes can use corresponding quantizer and / or dequantizer functions (e.g., approximate and / or differentiable quantizer and / or dequantizer functions) to train any model, including the reference model, and any model can also use corresponding quantizer and / or dequantizer functions for further training, validation, testing, inference, etc.

[0223] Using any framework disclosed herein, one or more nodes can transmit collected training data and / or datasets (e.g., channel estimation, precoding matrices, etc.) to another device that can train one or more models. For example, a UE and / or base station can upload collected online training data and / or one or more models to a server (e.g., a cloud-based server), whereby the server can use the uploaded training data to train one or more models and download one or more trained models to the UE and / or base station.

[0224] Any framework disclosed herein can be modified such that a first type of node can train a model for another type of node and share the trained model with multiple instances of a second type of node. For example, a base station can train an encoder for its decoder and share the trained encoder with multiple UEs. One or more UEs can apply the shared encoder at the UE to compress CSI and / or use the shared encoder for further online training. Furthermore, any framework disclosed herein can be implemented in a system where one of the nodes is not a base station (e.g., two UEs or other peer devices are configured for sidelink communication). In such an implementation, a UE can train a decoder for its encoder and share the trained decoder with one or more other UEs, where the one or more other UEs can use the trained decoder for direct inference and / or as a source of initial values ​​(e.g., weights) for further online training.

[0225] Model sharing mechanism

[0226] In some embodiments, nodes can use any type of communication mechanism (such as one or more uplink and / or downlink channels, signals, etc.) to transmit models, weights, etc. For example, when encoder model sharing is triggered, the UE can use one or more MAC Control Elements (MAC CE) PUSCHs to send the encoder model and / or weights to the gNB. Similarly, the gNB can use one or more MAC CE PDSCHs to send the decoder model and / or weights to one or more UEs.

[0227] Depending on the implementation details, the complete set of shared weights may be inefficient because the model may be relatively large and may consume a relatively large amount of downlink and / or uplink resources for sharing.

[0228] Some embodiments may establish a collection of one or more quantized models, which may be referred to as a model book. When a model is trained by a node, if sharing is requested, the node can map its model to one of the quantized models in the model book. One or more models in the model book can be shared among nodes. Instead of sending the pattern, the node can send the index of the mapped model from the model book. Depending on the implementation details, this can reduce the communication resources associated with model sharing.

[0229] In some embodiments, once the set of parameters of the model is known, the final result of the training can be known deterministically. For example, given (1) a training set, (2) an initial random seed for determining the initial weights, (3) optimizer parameters (e.g., fully defined optimizer parameters), and / or a training procedure, the trained model at the end of a certain number of training epochs (e.g., training cycles) can be uniquely determined. These parameters may be referred to as, for example, minimum description parameters. If the size of the minimum description is smaller than the size of the model's weights, nodes can share the minimum description parameters instead of the weights. Depending on the implementation details, this can reduce the communication overhead associated with sharing the model.

[0230] In some embodiments, one or more values ​​of the model (e.g., CSI encoding at a node and / or weights of the decision model) can be arranged in a vector W (e.g., a vector of weight elements). A dedicated compression autoencoder model (e.g., an encoder-decoder model pair) can be trained to compress W using an encoder at one node and a decoder at another node. If sharing of the CSI model is triggered and / or requested, a node can construct the vector W of the CSI model, encode it using the model compression encoder, and send the encoded vector to another node. The other node can use the model compression decoder to recover the weight vector W. Depending on the implementation details, this can reduce the communication overhead associated with sharing the model.

[0231] Online training processing time

[0232] In embodiments where nodes can perform online training of a model, a resource allowance (e.g., a quota for processing time, processing resources, etc.) may be provided to the nodes for training. Such a quota may be provided for training using an online training dataset that can be collected by the nodes (e.g., channel estimation based on measurements performed by the nodes) or an online training dataset that is RRC-configured (or reconfigured) or MAC-CE activated. A resource processing time quota ensures that the node has sufficient time to update the model using the online training dataset before the node is expected to have completed its update, e.g., to share the updated model with another node. However, in some embodiments, a processing time quota may be provided to the nodes regardless of whether the nodes are expected to share the trained model after processing.

[0233] For example, in an embodiment where the UE can collect an online training dataset by calculating channel estimates, the UE may be provided with a time horizon determined by N_(AIML,upadte) symbols starting from the end of the last symbol of the latest CSI-RS used for the online training set (e.g., to update the encoder model). If the UE is configured to report the updated model to another node (e.g., gNB), it may not be expected that the UE will report the model to the gNB earlier than N_(AIML,report) symbols starting from the last symbol of the latest CSI-RS in the training set.

[0234] As another example, in an embodiment where the UE can use an online training dataset for RRC configuration (or reconfiguration) or MAC-CE activation of the UE to perform online training of the encoder, it is not expected that the UE will update and / or report its encoder more than N symbols earlier than the latest symbol that has completed the corresponding RRC (re)configuration or has received the MAC-CE activation command.

[0235] Domain-based preprocessing

[0236] For compression purposes, a machine learning encoder can receive an input signal and generate a set of output features sufficient for a decoder to reconstruct the input signal. Under maximum compression, the output features can be expected to be independent of each other; otherwise, they can be further compressed.

[0237] While a pair of machine learning models may be able to generate feature vectors from inputs and reconstruct inputs from feature vectors, in some embodiments according to this disclosure, one or more preprocessing and / or postprocessing operations may be performed on the inputs of the generating model and / or the outputs of the reconstructing model. Depending on the implementation details, this may provide one or more potential benefits, such as reducing the processing burden and / or memory usage of one or both models, improving the accuracy and / or efficiency of one or both models, etc.

[0238] In some embodiments, preprocessing and / or postprocessing can be based on domain knowledge of the input signal. In some embodiments, preprocessing and / or postprocessing can provide the encoder with auxiliary information from the domain knowledge, depending on the implementation details, which can reduce the processing burden on the encoder. For example, if the vector to be compressed by the encoder can be characterized as a low-pass signal with relatively small variations, a Discrete Fourier Transform (DFT) and / or Inverse DFT (IDFT) can be performed to analyze the frequency domain representation of the vector. If the DC component of the DFT vector is greater than (e.g., significantly greater than) other components, it can indicate that the signal has low variations, and therefore preprocessing can be performed before compression by the encoder (and postprocessing after decompression by the decoder) to reduce the burden on the encoder / decoder pair.

[0239] In some embodiments, performing transformations and / or inverse transformations such as DFT and / or IDFT can provide a machine learning model with a clearer understanding of the level of correlation between the elements of the input vector. For example, in some embodiments (e.g., utilizing any framework disclosed herein), the CSI matrix can be input to a preprocessor, which can apply transformations (e.g., DFT / IDFT, Discrete Cosine Transform (DCT) / Inverse DCT (IDCT), etc.) to all or part of the input (e.g., on different CSI-RS ports). The transformed signal can then be input to an encoder and compressed. On the decoder side, the decoder's output can be applied to the inverse operator of the preprocessor's transformation (which can be implemented, for example, using a postprocessor) to generate a reconstructed input signal.

[0240] Figure 11 An example embodiment of a dual-model training scheme with preprocessing and post-processing according to this disclosure is shown. In some aspects, Figure 11 The embodiment 1100 shown can be similar to Figure 3 The embodiments shown are examples of this, and similar components can be identified using a designator ending with the same number. However, Figure 11The illustrated embodiment may include a preprocessor 1137 and a postprocessor 1138. The preprocessor 1137 may apply any type of transformation to the training data 1111 before it is applied to the generating model 1103. Similarly, the postprocessor 1138 may apply any type of inverse transformation (e.g., the inverse of the transformation applied by the preprocessor 1137) to the output of the reconstructed model 1104 to generate the final reconstructed training data 1112.

[0241] In some embodiments, the loss function 1113 used to train models 1103 and 1104 can be defined between the input of the generating model 1103 and the output of the reconstructed model 1104, as shown by solid lines 1139 and 1140. However, in some embodiments, the loss function 1113 can be defined between the input of the preprocessor 1137 and the output of the postprocessor 1138, as shown by dashed lines 1141 and 1142. Once models 1103 and 1104 are trained as follows... Figure 11 Once trained as shown, they can be used for inference.

[0242] While the principles related to preprocessing and / or postprocessing are not limited to any particular implementation details, for the purpose of illustrating the principles of the invention, an example embodiment of a scheme for preprocessing and postprocessing CSI matrices based on domain knowledge can be implemented as follows. Using a channel matrix of size N_rx×N_tx, for each pair of (i,j)RX and TX antennas, the channel elements corresponding to pairs of all resource elements (REs) within the time and frequency window can be concatenated to obtain a combination matrix H_(i,j) of size M×N, where M and N can be the number of CSI-RS subcarriers and orthogonal frequency division multiplexing (OFDM) symbols in the window. In some embodiments, it can be assumed that the matrix is ​​complex. In an example embodiment of the preprocessing scheme, H_(i,j) can be transformed, for example, using a DFT matrix. If U_freq and U_time are M×M and N×N DFT matrices, respectively, then matrix H_(i,j) can be transformed into X_(i,j) as follows:

[0243] X_(i,j)=U_freq^*H_(i,j)H_time (Equation 2)

[0244] It can be called the Delayed Doppler Representation (DDR) of H_(i,j). The matrix H_(i,j) can be reconstructed from the DDR as follows:

[0245] H_(i,j)=U_freq X_(i,j)U_time^* (Equation 3)

[0246] In some embodiments, the use of DDR transformation can result in a sparse X matrix, which in turn can reduce the complexity of learning and inference.

[0247] In some embodiments, preprocessing and / or postprocessing transformations can be used to transform the original training set into the corresponding DDR matrix. In such embodiments, CSI compression can then compress the transformed training set. Therefore, preprocessing (e.g., DDR transformation) can be performed on the UE side, while postprocessing (e.g., inverse DDR for recovering H) can be performed on the gNB side.

[0248] In some embodiments, the loss function can be defined based on the transform matrix (e.g., between the transform matrix from the input to the encoder and the transform matrix from the decoder output, such as...). Figure 11 (As shown).

[0249] In some embodiments, matrix H can be constructed based on the union of the individual CSI matrices in the time and / or frequency domains for each spatial channel (e.g., each transmit antenna (port) and each receive antenna (port) pair). One or more models can be trained and tested for each spatial channel. In some embodiments, H can be constructed based on the channel matrix of the RE, for example, where each matrix can have a size N_r × N_t, where N_r and N_t can be the number of receive antennas at the UE and the number of transmit antennas at the gNB, respectively.

[0250] CSI matrix formula

[0251] In some embodiments, the CSI information of REs or RE groups that the UE can compress can be referred to as the CSI matrix. Based on the analysis of a multiple-input multiple-output (MIMO) channel, the capacity distribution can be a Gaussian distribution with potentially different power allocations across transmit antennas. If the channel matrix is ​​decomposed into H_(r×t)=U∑V^H, it can be determined by first setting... (where x is iid) to obtain the capacity of the realized distribution. A Gaussian random vector with zero mean and unit variance is then used to calculate the capacity of each element in the following terms.

[0252]

[0253] Multiply by the power allocation given by the water-filling algorithm. The power allocation for the i-th channel can be P_i, i = 1, ..., t, where P_i can be obtained from the singular value decomposition of the channel matrix H. Therefore, the information used at gNB (e.g., the complete information required at gNB) can be the right singular value matrix V and the singular values ​​themselves. Thus, in some embodiments, the CSI matrix can be formulated (e.g., defined) as any one or more of the following: (a) the CSI matrix can be formulated as the channel matrix H; (b) the CSI matrix can be formulated as a concatenation of V and singular values; and / or (c) the CSI matrix can be formulated as matrix U.

[0254] In some embodiments, the UE may be configured to report any CSI matrix described above via RRC (re)configuration, MAC-CE command, or dynamically via DCI. In embodiments where the CSI matrix is ​​formulated according to (b) (e.g., a concatenation of V and singular values), the UE may also be configured to report only singular values.

[0255] For model training purposes, when the UE is configured to report a specific CSI matrix, the training set and / or loss function can be formulated based on the applicable CSI matrix. For example, when the UE is configured to report V, the training set may include V matrices obtained from the estimated channel matrix, and the loss can be formulated based on the V matrix input to the encoder and the reconstructed V matrix output by the decoder.

[0256] Node capabilities

[0257] Implementations of any framework disclosed herein may involve the use of resources such as memory, processing, and / or communication resources, for example, to store new training data, share models between nodes, and train and / or apply specific types of neural network architectures, such as CNNs or RNNs for the model. Different nodes, such as the UE, may have different capabilities for implementing neural networks. For example, the UE may support or may be able to support CNNs but not RNNs. In some embodiments, the UE may report its ability to support specific types of neural network architectures (e.g., network types such as CNNs, RNNs, etc.) and / or any other aspect reflecting its constraints and capabilities in applying encoder models.

[0258] In some embodiments, a node (e.g., a UE) may use a list to report its capabilities and / or limitations, wherein the list may include any number of the following: (a) one or more network types, such as CNN, RNN, a specific type of RNN, Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), transformer, etc.; (b) one or more aspects related to the size of the model, such as the number of layers, the number of input and / or output channels of a CNN, the number of hidden states of an RNN, etc.; and / or (c) any other type of structural constraints.

[0259] Depending on its reported capabilities, a node (e.g., a UE) may not be expected to train or test encoder models that violate any constraints reported by the node and / or require capabilities beyond those declared by the node. In some embodiments, this can be ensured regardless of the applicable framework and / or location of the model's training / inference. For example, if the framework is implemented such that the gNB can pre-train encoders and decoders and share them with the UE, the UE may not expect the encoder model to violate its capabilities. As another example, if one or more pairs of encoders and decoders are trained offline and specified in an applicable specification (e.g., the NR specification), the UE may not expect the applicable model to violate its capabilities. In some embodiments, the UE may report the ability to activate one or more models via signaling and may declare which specific encoder / decoder pairs, or individual encoders or decoders, it can support. The gNB can then indicate to the UE which encoder / decoder pair can be applied to the UE. This indication may be provided, for example, via system information in the DCI, RRC configuration, dynamic signaling, etc.

[0260] Adjusted through online training

[0261] In some frameworks, it may be expected that nodes such as UEs can train their encoder models or both encoder and decoder models online, for example, by collecting new training data (e.g., samples) during operation or by offline provisioning and updating based on one or more models. If a node only updates the encoder model, then encoder model tuning and / or optimization may also depend on decoder weights and / or the model, since the loss function may also depend on decoder weights. In such implementations, nodes may also declare the ability to handle constraints on the decoder model, even if the encoder can be used on the gNB side. One or more of these constraints may be applied as follows: (1) a node may declare one or more online training features and / or fine-tuning as a capability; (2) a node reporting the ability to support online training may further report constraints on the support structure of the encoder model, which may be, for example, any constraints mentioned in (a) to (c) above; (3) a node reporting the ability to support online training may further report constraints on the support structure of the decoder model, which may be, for example, any constraints mentioned in (a) to (c) above; and / or (4) if multiple models including encoder and decoder pairs are specified in the specification, a node may declare the ability to indicate which pairs of encoders and decoders or which individual encoders and / or decoders can be supported.

[0262] Many-pair model

[0263] In some embodiments, multiple pairs of models (e.g., encoder / decoder pairs) can be trained and / or deployed to operate (e.g., simultaneously) on two nodes (e.g., UE and gNB). Each pair can differ from the other in a) both encoder and decoder, b) encoder only, or c) decoder only. In some embodiments, multiple pairs of models can be configured (e.g., optimized) to handle different scenarios that can be specified to handle different channel environments, which in turn can lead to different distributions of training and / or test datasets.

[0264] Multiple model pairs can be used, for example, to adapt to different dimensions in the training data. For example, the dimension of the CSI matrix can be determined based on the number of CSI-RS ports. In some embodiments, if the UE is to report a first CSI matrix H_1 and a second matrix H_2 with different dimensions, a single encoder and decoder pair can be used to process matrices of different sizes. In such an embodiment, the encoder and decoder can be trained as follows: The matrix input to the encoder can be shaped to a fixed size by appending zeros to a configuration that can be mutually understood between the UE and the gNB. Therefore, the training set can initially include matrices of different sizes, which can be modified by appending zeros as described above to convert the matrices to a fixed matrix size. Communication mechanisms can be implemented to enable the gNB and the UE to share the same understanding regarding, for example, the size of the CSI matrix requested in the UE's report. Depending on the implementation details, the matrix redimensionalization technique can be applied to any matrix size.

[0265] Alternatively or additionally, multiple pairs of models (e.g., encoder / decoder pairs) can be trained, where different pairs of models can be configured to handle different CSI matrix sizes.

[0266] In some embodiments, and depending on the implementation details, multiple pairs of models can be implemented without increasing complexity. For example, in the case of multiple pairs of models, if the CSI report includes a CSI matrix corresponding to a specific number of CSI-RS ports, the inference time for calculating the CSI report may be smaller in the case of multiple pairs than in the case of a single pair. Furthermore, if each RRC configuration or MAC-CE activation includes a CSI report for a specific case corresponding to a specific pair, the UE can load an applicable model into the modem while keeping one or more other models in the UE controller. Depending on the implementation details, this can reduce the use of the modem's internal memory. In embodiments with multiple pairs, different pairs can be categorized according to any of the following configurations: (a) Each pair of models can be configured to handle a specific CSI matrix size. For example, a pair of models can receive a CSI matrix based on CSI-RS estimates having a specific number of ports and also associated with a specific number of receive antennas of the UE. The UE can report the number of its receive antennas to the gNB in ​​one report, or report them separately for different numbers of CSI-RS ports. (b) Each pair of models can be configured to handle different distributions of training and / or test datasets. (c) Each pair of models can be configured to handle different channel environments for the training and / or testing datasets.

[0267] Training set association and model-to-configuration

[0268] In embodiments where multiple model pairs can be configured to handle different scenarios, nodes (e.g., UEs or base stations) can be configured with different training datasets, for example, different training datasets for specific scenarios or specific model pairs (e.g., encoder / decoder pairs). Therefore, the UE and / or base station can have different training datasets, for example, for each different pair. Once a trigger occurs for a node (e.g., UE or gNB), the node can also be signaled to indicate which model pair should be trained. For example, using online training, the gNB can instruct the UE to begin training a specific model pair. If online training is performed during operation by collecting new datasets, an association can be provided between the CSI-RS and the encoder / decoder pair, for example, via the number of CSI-RS ports.

[0269] Once multiple pairs of models have been trained and are ready for deployment during the inference phase, a node (e.g., a UE) may need to know which pair to use to encode the channel matrix. For example, each pair may be associated with a certain dimension of the CSI matrix to be encoded. This dimension can be referred to as the input dimension of the encoder model. In some embodiments, the UE may determine the pair of models used to encode the CSI matrix as follows: CSI-RS can be implicitly or explicitly associated with a pair of models. The UE uses the pair of models associated with the CSI-RS to encode the CSI matrix. Using implicit association, CSI-RS can be mapped to a specific pair based on the number of CSI-RS ports and / or the number of receive antennas at the UE. Thus, if the dimension of the CSI matrix obtained from the CSI-RS is equal to the input dimension of a pair of models, then the CSI-RS can be mapped to that pair of models. If multiple pairs have the same eligible input dimension, a reference pair can be selected, for example, based on rules that can be established between the UE and the gNB. Using explicit association, the CSI-RS for which it reports the CSI matrix can be indicated via RRC configuration or dynamically in the DCI, for example, using a pair index.

[0270] In any of the embodiments described above, if the UE is signaled to report the CSI matrix via a pair of models with different input dimensions, the UE may append zeros to match the size of the matrix to the input dimensions. However, the UE may not expect to be signaled to report the CSI matrix using a pair of models with input dimensions smaller than the CSI matrix.

[0271] Compression with reduced model size

[0272] In some embodiments, a pair of models can be configured as an autoencoder to compress the CSI matrix and / or utilize redundancy and / or correlation between CSI matrix elements. If the CSI matrix is ​​reported per RE, the correlation may simply be the spatial correlation between different paths between different pairs of transmit antennas (e.g., CSI-RS ports) and receive antennas. Depending on the implementation details, the amount of this correlation may be limited, and therefore the autoencoder may not be able to adequately compress the CSI matrix.

[0273] In some embodiments, the compression capability of the autoencoder may be related to the amount of redundancy and / or correlation (which may be referred to as spatial correlation) between the elements of the CSI matrix. Since wireless channels can also be correlated in the time and / or frequency domains, time-domain and / or frequency correlations may also exist. Therefore, estimated channels for multiple OFDM symbols and / or multiple resource elements (REs), multiple resource blocks (RBs), or multiple subbands can be used as input to a single training sample. For example, a channel matrix corresponding to multiple REs can be specified as input to the autoencoder. In one such method according to this disclosure, the UE may be configured via RRC to specify time and / or frequency resource bundles for forming the training dataset.

[0274] Depending on the implementation details, the compression performance of an autoencoder can be improved by compressing the CSI of multiple REs across different frequency and / or time resources. Therefore, a combined CSI matrix of multiple REs can be input into a time and frequency window. The combined CSI matrix can then be obtained from the individual CSI matrices of the REs in the cascaded window. Depending on the implementation details, the combined CSI matrix may be more likely to have significant correlations between its elements due to the time and frequency flatness of the channel. Therefore, if a model takes the combined CSI matrix as input, it may be able to compress it to a higher degree than multiple models operating on individual RE matrices. In some embodiments, the UE can be configured with time and / or frequency windows and one or more configurations that can indicate which REs the UE can employ to determine the combined CSI matrix. Such a configuration can be used in both the training and / or testing phases to obtain the combined matrix.

[0275] Input size reduction via a subset of the CSI matrix

[0276] In some embodiments, the autoencoder can encode the CSI matrices of different REs within certain time and frequency windows. If the channel causes the correlation between elements of the channel matrix to be absent or weak in certain domains (e.g., time or frequency), the set of elements in the union of the CSI matrices can be partitioned into subsets with relatively strong intra-subset element correlation and relatively weak inter-subset element correlation. For example, if the autoencoder wants to compress four CSI matrices of four REs on the same OFDM symbol, the matrices can be represented as follows:

[0277]

[0278] If the correlation in the frequency domain is strong and there is little or no correlation in the spatial domain (i.e., between the elements of a matrix), the autoencoder can be configured to compress a vector of length 4, and the autoencoder can be applied four times on the following subsets: subset 1 (a_1,b_1,c_1,d_1); subset 2 (a_2,b_2,c_2,d_2); subset 3 (a_3,b_3,c_3,d_3); and subset 4 (a_4,b_4,c_4,d_4).

[0279] The CSI matrix can then be reconstructed at the decoder by, for example, reconstructing four vectors using the same decoder. As mentioned above, subsets can be selected such that they can utilize one or more correlations in one or more domains. To further illustrate, in the example above, if there is correlation between elements in the spatial domain, the subset selection described above can prevent the network from using correlation to further compress the CSI matrix. Instead, the following subset selections can allow the utilization of correlations in both the frequency and spatial domains: subset 1 (a_1,a_2,b_1,b_2); subset 2 (c_1,c_2,d_1,d_2); subset 3 (a_3,a_4,b_3,b_4); and subset 4 (c_3,c_4,d_3,d_4).

[0280] In some embodiments, the following framework can be used as a basis for a reduced model size with an input dimension of N_features. (1) The UE can be configured to report a CSI matrix of M REs, where the M REs may be on the same or different OFDM symbols and may be within a time and / or frequency window. Each CSI matrix may have N elements. (2) The UE may divide the M×N elements into (M×N) / N_features subsets. Common rules may be established between the UE and the gNB for subset selection. (3) An autoencoder (e.g., a single autoencoder) may be used to compress and recover the N_features elements in each subset. In the example subsets described above, M = N = 4, and N_features = 4.

[0281] Input size reduced by resource element selection

[0282] The size of the encoder network can be reduced by decreasing the size of the combined input matrix. In some embodiments, the size of the combined matrix can be reduced by (a) removing certain elements of each RE matrix if two REs exist in a window of CSI matrices H_1 and H_2 with the same dimensions. The combined matrix can be constructed to have the same dimensions as H_1 or H_2, but by selectively picking (i,j) elements from H_1 or H_2. Alternatively or additionally, the size of the combined matrix can be reduced by (b) constructing a matrix that excludes certain REs from the window of CSI matrices.

[0283] These examples are shown in Table 1, which shows a window with two REs and two CSI matrices. Using method (a), the combined matrix can be constructed as shown in Table 1, while using method (b), the combined matrix can be constructed by selecting one of the two matrices.

[0284] Table 1

[0285]

[0286] Number of CSI-RS ports and training set

[0287] The channel matrix estimated from the CSI-RS with N_port ports and the reception via N_r receive antennas at the UE can have a dimension of N_r × M_port. An autoencoder can be used to remove redundancy from the matrix. In some embodiments, if the training set includes matrices of different dimensions (e.g., corresponding to CSI-RS with different numbers of ports), identifying and / or removing redundancy patterns may be more difficult. Therefore, in some embodiments, the training set for any method disclosed herein may consist only or primarily of matrices of the same dimension and / or associated with the same number of CSI-RS ports. Thus, it may not be desirable for the UE to be configured with a training set or CSI reporting and measurement configuration that results in a training set with dimensions different from those of the training dataset matrix.

[0288] UCI format

[0289] Using any framework disclosed herein, the output of a generative model (e.g., an ML encoder) can be considered a type of UCI (which may be referred to as, for example, Artificial Intelligence Machine Learning (AIML) CSI). In some embodiments, the AIML CSI can be obtained from a CSI report and a measurement configuration with associated CSI-RS resources and reporting settings. In some embodiments, the AIML CSI can be transmitted to the gNB via PUCCH or PUSCH (e.g., following Rel-15 behavior). Thus, the format of the representation of feedback information for physical layer information can be established as a type of uplink UCI. The format may involve one or more types of coding (e.g., polar coding, low-density parity-check (LDPC) coding, etc.), which may depend on, for example, the type of physical channel used to transmit the UCI. In some embodiments, the type of uplink UCI (e.g., AIML CSI) transmitted using PUCCH can use polar coding, while transmission using PUSCH can use LDPC coding. Furthermore, the CSI can be quantized before encoding. Thus, the AIML CSI can be quantized into a bit stream (0s and 1s) and input to a polar encoder or an LDPC encoder.

[0290] Adaptability to different network providers

[0291] When a UE connects to a network, it may not know which network vendor created the network it connects to. Since different vendors may employ different training techniques and / or network architectures for their machine learning models, the availability of this information can affect the training model on the UE side. Therefore, in some embodiments, a network indication or AI / ML index can be provided to the UE via system information (e.g., via one of the SIBs). The UE can then use this information to adapt its training to a specific network vendor configuration.

[0292] ML model lifecycle management

[0293] In some ML applications, the performance of ML models may degrade over time and may not be fully executed during the duration the application was trained. Therefore, ML models can be updated frequently to accommodate temporal changes that may occur in their operating environment, such as statistical variations in the wireless channel in the case of CSI compression.

[0294] According to some embodiments of this disclosure, a management framework can be provided to enable efficient and / or timely updates of one or more ML models with acceptable overhead. To facilitate such a framework, one embodiment may implement model monitoring, wherein nodes can track the performance of ML models. In some embodiments, this may involve model monitoring, wherein nodes can track the performance of ML models. Model monitoring may be based on one or more performance metrics as follows: (1) Task-based metrics can be used to (e.g., directly) evaluate the performance of tasks performed by the ML model. For example, these metrics can include accuracy, mean squared error (MSE) performance, etc. (2) System-based metrics can be used to track the overall performance of the system, such as correct decoding of transmissions, or other system-level key performance indicators (KPIs) that can provide a less direct measure of the performance of the ML models used by nodes in the system.

[0295] When the performance of an ML model is deemed unacceptable based on agreed and / or configured metrics, the management framework may initiate an ML model update procedure. Performance can be considered unacceptable, for example, (1) when the performance of the ML model is unacceptable according to one or more agreed and / or configured metrics; and / or (2) if the performance of the ML model is unacceptable for a specific duration greater than a threshold time.

[0296] The threshold used to determine unacceptable performance can be implemented as a configured and / or specified parameter. The duration can be i) measured cumulatively, for example, any duration of unacceptable performance can be added to a global counter and the global counter value compared to the threshold, or ii) measured continuously, for example, only continuous durations of unacceptable performance greater than the threshold can be considered.

[0297] When the performance of the ML model is deemed unacceptable, the management framework can trigger an update process that can be implemented in one of the following ways: (1) The management framework can require the complete training process, such as regarding... Figure 8 As described. In this case, the ML model can be retrained from scratch, or it can be retrained from the current ML model. Training in this case can use the entire training dataset, with or without additional data samples that may have been recently acquired. (2) The management framework may require partial training, in which the ML model can be retrained from the current ML model and may use newly acquired data samples.

[0298] Performance metrics in training and testing

[0299] To evaluate the performance of different models for CSI compression tasks, some embodiments of this disclosure can focus on aspects of CSI compression. In such embodiments, different models can be compared based on their respective compressed CSI matrices and their ability to recover CSI matrices, such that the recovered matrix is ​​as close as possible to the true CSI matrix. The determination of the closeness can be related to the operation of the gNB utilizing the CSI matrix. For example, if the gNB calculates the SVD of the channel matrix as H_(r×t)=U∑V^H and uses the right singular vector V to determine the pre-encoder, the closeness between V at the encoder input and the recovered V at the decoder output can be determined.

[0300] In some embodiments, the proximity measure between two matrices can be implemented on an element-wise basis, and an average can be taken over some or all elements to provide a single loss value. Alternatively or additionally, having one or a few erroneous elements in a matrix can be just as detrimental as having many erroneous elements. In this case, the loss function can be determined on a matrix-wise basis (e.g., the maximum value of the element-wise errors over all elements of the matrix).

[0301] In some embodiments, the performance of the CSI encoder and decoder models can also be evaluated in conjunction with other blocks of the system. For example, if the block error rate (BLER) is used as a system performance metric, comparisons between different CSI models can be based on their obtained BLERs. Other system KPIs (such as throughput, resource utilization, etc.) can also be used for this purpose.

[0302] In embodiments where BLER is the metric of interest, configuring the gNB to use the information provided by the CSI matrix may affect system performance. For example, assuming the CSI model is perfect in the sense that the channel matrix transmitted by the UE is fully recovered at the gNB, and the channel matrix indicates a rank-1 channel, decoding may fail if the gNB schedules a rank-2 PDSCH. Therefore, to establish a link between the compression capability of the CSI model and system performance, assumptions about the gNB operation can be used. In some embodiments, the gNB of the processing function f_gNB can be defined as obtaining the decoder output, e.g., H^, and providing the resulting BLER estimate as BLER = f_gNB(H^) (or BLER = f_gNB(H,H^)). Then, both CSI compression and gNB operation aspects can be considered to define the loss function during training. For example, the loss can be defined as a weighted sum of two terms, as follows:

[0303] loss(H,H^)=α.loss_CSI(H,H^)+β.BLER (Equation 6)

[0304] Here, α and β are hyperparameters used for training.

[0305] Reliability of uplink channel

[0306] In some embodiments, it can be assumed that the encoder output (also referred to as the CSI codeword) is available on the decoder side without errors. Therefore, the CSI codeword can be transmitted over the uplink channel via PUCCH or PUSCH with infinite reliability, ensuring that PUCCH and / or PUSCH decoding will not fail. However, in some instances, during the inference phase, for example, when PUSCH / PUCCH decoding fails, the CSI codeword can be delivered to the gNB (decoder) with one or more errors. In this case, a noisy version of the CSI codeword may be available at the decoder. The impact of imperfections in the uplink channel during the training phase can be modeled as follows.

[0307] For each training example input to the encoder, the CSI codeword at the encoder's output can be represented as x. Considering the imperfections of the uplink channel, the decoder's input y can be modeled as...

[0308] y = x + ω (Equation 7)

[0309] Here, ω is additive noise that can be used to model the residual error after decoding the uplink channel. Additive noise can be generated during the training phase as follows.

[0310] In Method 1, ω can be modeled as a Gaussian random vector with zero mean and variance σ^2. The variance can be indicated to the UE via RRC configuration or left to the UE implementation. In Method 2, the channel between x and y can be modeled by performing PUCCH and / or PUSCH decoding for each training example to obtain the residual error vector ω, which is then assumed to be added to x to obtain y.

[0311] In terms of collaborative learning

[0312] Using federated learning (FL), a global model at the server can be learned by being learned individually at multiple nodes connected to the server, and the learned model can be shared with the server. The server can then perform one or more operations on the received model to obtain the final model. For example, privacy considerations and / or requirements that nodes do not share their data with the server may prompt such an arrangement.

[0313] In the CSI compression use case, the server can be considered a gNB, and different UEs connected to the gNB can be considered model update nodes. Different UEs can have different training sets with the same or different distributions. If the distributions are the same, each UE can update its model using its own training set and share the model with the gNB. The gNB can then perform one or more operations, such as averaging the models to obtain a final model. The gNB can share the obtained final model with UEs that share its model. The final model can be expected to outperform the individual received models because it is trained based on the union of all training sets across all participating UEs. Therefore, FL can be used to improve CSI compression performance. In cases where different distributions are available at different UEs, FL can help capture distributions that a particular UE has not yet seen by using a model shared by UEs that have already seen the distribution. In any case, FL can be used to obtain a model that takes into account the different environments observed by the UEs.

[0314] Using the FL framework according to this disclosure, a gNB can configure a group of UEs in an FL group. UEs in the same FL group can be configured to have the same encoder and / or decoder (e.g., autoencoder or AE) architecture. Therefore, their encoders and decoders may differ only in the weights actually trained, but have the same configuration in terms of the number of layers, the number of units, activation functions, and other parameters defining the network structure.

[0315] For UEs in the group, the input sizes of the encoder and decoder models can be the same or similar. The encoder inputs can also have the same or similar meanings for the UEs. For example, the encoder input for a UE (e.g., all UEs) can be a channel matrix or a singular value matrix V. The gNB can instruct the UEs to update their models via RRC, DCI, or MAC CE commands and share the updates with the gNB. In some embodiments, not all UEs in the group participate in the update process simultaneously. The gNB can send information about the training, hyperparameters, and / or other aspects of the FL via the group common (GC) DCI, where UEs in the same FL group can have a specific portion of the DCI configured via RRC.

[0316] Additional Examples

[0317] Figure 12 An embodiment of a system for using a dual-model scheme according to the present disclosure is shown. Figure 12 The embodiments shown can be described in the context of one or more test models, but the same or similar embodiments can also be used to utilize any model disclosed herein (e.g., utilizing a trained model). Figure 3 The generated model 303 and / or reconstructed model 304 shown in the figure are used for verification, reasoning, etc.

[0318] refer to Figure 12 System 1200 may include a first node (node ​​A) having a generation model 1203 and a second node (node ​​B) having a reconstruction model 1204. Test data 1211 may be applied to generation model 1203, wherein generation model 1203 can generate a representation 1207 of the test data. Reconstruction model 1204 can generate a reconstruction 1212 of the test data based on the representation 1207. In some embodiments, generation model 1203 may include a quantizer to convert representation 1207 into a quantized form (e.g., a bitstream) that can be transmitted via a communication channel. Similarly, in some embodiments, reconstruction model 1204 may include a dequantizer, wherein the dequantizer can convert the quantized representation 1207 (e.g., a bitstream) into a form that can be used to generate the reconstructed test data 1212.

[0319] The generative model 1203 and the reconstructed model 1204 can be obtained in any manner, including using any framework described herein. For example, using a joint training framework, the generative model 1203 and the reconstructed model 1204 can be trained as a pair at node A, where node A can send the reconstructed model 1204 to node B. Other embodiments may use a training framework with a reference model, a training framework with the latest shared values, or any other framework and / or technique to obtain and / or train the generative model 1203 and the reconstructed model 1204.

[0320] Figure 13 An example embodiment of a user equipment (UE) according to this disclosure is shown. Figure 13 The embodiment 1300 shown may include a radio transceiver 1302 and a controller 1304, wherein the controller 1304 may control the operation of the transceiver 1302 and / or any other component in the UE 1300. The UE 1300 may be used to implement, for example, any of the functions described in this disclosure, including determining channel information based on one or more reference signals from a base station, generating a representation of the channel information based on channel conditions using a machine learning model, transmitting the representation of the channel information, collecting training data (e.g., during a window), performing preprocessing and / or postprocessing (e.g., for a CSI matrix), deploying and / or activating one or more pairs of ML models, etc.

[0321] Transceiver 1302 can send / receive one or more signals to / from a base station and may include interface units for such transmission / reception. For example, transceiver 1302 may receive one or more signals from a base station and / or may send a representation of channel information to the base station on a UL channel.

[0322] Controller 1304 may include, for example, one or more processors 1306 and memory 1308, wherein memory 1308 may store instructions for one or more processors 1306 to execute code to implement any of the functions described in this disclosure. For example, controller 1304 may be configured to implement one or more machine learning models as disclosed herein, and to determine channel information based on one or more reference signals from a base station, generate a representation of the channel information using the machine learning model based on channel conditions, transmit the representation of the channel information, collect training data (e.g., during a window), perform preprocessing and / or postprocessing (e.g., for a CSI matrix), deploy and / or activate one or more pairs of ML models, etc.

[0323] Additionally or alternatively, controller 1304 may be configured to implement any of the functions disclosed herein related to the use of constraint information for the channel, including, for example, providing the UE with a constrained subspace for the channel and / or enabling the UE to ensure that the precoder determined using the reconstruction model is orthogonal to the constrained subspace. For example, controller 1304 may be configured to implement respectively in Figure 24 , Figure 25 and / or Figure 26 Shown and / or about Figure 24 , Figure 25 and / or Figure 26 Some or all of the described CSI determination logic 2452, precoding determination logic 2552 and / or precoding determination logic 2652.

[0324] Additionally or alternatively, controller 1304 may be configured to implement any functions disclosed herein regarding data collection for use with artificial intelligence and / or machine learning models, including, for example, establishing data collection sessions and / or declaring data collection capabilities. Controller 1304 may be configured, for example, to implement in Figure 28 Shown and / or about Figure 28 The data collection logic 2866 and / or data collection logic 2867 described are also described. Figure 29 Shown and / or about Figure 29 Some or all of the data collection logic 2966 and / or data collection logic 2967 described.

[0325] Additionally or alternatively, controller 1304 may be configured to implement any functions disclosed herein related to data collection for incoherent joint transmission, including, for example, reporting data collection capabilities and / or dataset descriptions. Controller 1304 may be configured to, for example, implement... Figure 30 Shown and / or about Figure 30 The data collection logic described includes some or all of 3066, 3067A and / or 3067B.

[0326] Figure 14 An example embodiment of a base station according to this disclosure is shown. Figure 14 The illustrated embodiment 1400 may include a radio transceiver 1402 and a controller 1404, wherein the controller 1404 may control the operation of the transceiver 1402 and / or any other component in the base station 1400. The base station 1400 may be used to implement, for example, any of the functions described in this disclosure, including transmitting one or more reference signals to the UE on the DL channel, reconstructing the representation of channel information, performing preprocessing and / or postprocessing (e.g., for the CSI matrix), deploying and / or activating one or more pairs of ML models, etc.

[0327] Transceiver 1402 can send / receive one or more signals to / from a user equipment and may include interface elements for such transmission / reception. For example, transceiver 1402 can send one or more reference signals to the UE on a DL channel and / or receive precoded information from the UE on a UL channel.

[0328] Controller 1404 may include, for example, one or more processors 1406 and memory 1408, wherein memory 1408 may store instructions for the one or more processors 1406 to execute code to implement any base station functions described herein. For example, controller 1404 may be used to implement one or more machine learning models as disclosed herein, as well as to transmit one or more reference signals to the UE on the DL channel, reconstruct a representation of channel information, perform preprocessing and / or postprocessing (e.g., for the CSI matrix), deploy and / or activate one or more pairs of ML models, etc.

[0329] Additionally or alternatively, controller 1404 may be configured to implement any of the functions disclosed herein related to the use of constraint information for the channel, including, for example, providing the UE with a constrained subspace for the channel and / or enabling the UE to ensure that the precoder determined using the reconstruction model is orthogonal to the constrained subspace.

[0330] Additionally or alternatively, controller 1404 may be configured to implement any of the functions disclosed herein regarding data collection for use with artificial intelligence and / or machine learning models, including, for example, establishing data collection sessions and / or declaring data collection capabilities.

[0331] Additionally or alternatively, controller 1404 may be configured to implement any of the functions disclosed herein related to data collection for noncoherent joint transmission, including, for example, reporting data collection capabilities and / or dataset descriptions.

[0332] exist Figure 13 and Figure 14In the illustrated embodiments, transceivers 1302 and 1402 can be implemented using various components to receive and / or transmit RF signals, such as amplifiers, filters, modulators and / or demodulators, A / D and / or DA converters, antennas, switches, phase shifters, detectors, couplers, conductors, transmission lines, etc. Controllers 1304 and / or 1404 can be implemented using hardware, software, and / or any combination thereof. For example, all or part of the hardware implementation may include combinational logic, sequential logic, timers, counters, registers, gate arrays, amplifiers, synthesizers, multiplexers, modulators, demodulators, filters, vector processors, complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), systems-on-a-chip (SoCs), system-in-package (SIPs), multi-chip modules, state machines, data converters such as ADCs and DACs, etc. All or part of the software implementation may include one or more processor cores, memory, program and / or data storage devices, etc., which may be located locally and / or remotely, and may be programmed to execute instructions to perform one or more functions of the controller. Some embodiments may include one or more processors (such as microcontrollers), CPUs (such as complex instruction set computer (CISC) processors (such as x86 processors) and / or reduced instruction set computer (RISC) processors (such as ARM processors) that execute instructions stored in any type of memory, graphics processing unit (GPU), neural processing unit (NPU), tensor processing unit (TPU), etc.)

[0333] Figure 15 An embodiment of a method for providing physical layer information feedback according to the present disclosure is shown. The method may begin at operation 1502. At operation 1504, the method may determine physical layer information for the wireless device at the wireless device. At operation 1506, the method may use a machine learning model to generate a representation of the physical layer information. At operation 1508, the method may transmit the representation of the physical layer information from the user equipment from the wireless device. The method may end at operation 1510.

[0334] exist Figure 15 The components and / or operations shown in the embodiments illustrated and in any of the embodiments disclosed herein are merely exemplary. Some embodiments may involve various additional components and / or operations not shown, and some embodiments may omit certain components and / or operations. Furthermore, in some embodiments, the arrangement of components and / or the temporal order of operations may vary. Although some components may be shown as separate components, in some embodiments, some components shown separately may be integrated into a single component, and / or some components shown as a single component may be implemented using multiple components.

[0335] Precoding and channel information mismatch

[0336] In an NR system, the UE can measure downlink channel conditions based on reference signals (e.g., CSI-RS or DMRS) transmitted by the gNB. The UE can use these channel measurements to determine (e.g., calculate) channel information that it can report to the gNB, such as precoding matrices and / or channel quality indicators (CQIs). If the gNB applies the reported precoding matrices to subsequent transmissions, the UE can calculate a precoding matrix that results in a good (e.g., optimally available) equivalent channel for downlink transmissions. If the gNB uses the reported precoding matrices, the CQI that the UE can calculate based on the precoding matrices can indicate the expected channel quality. For example, the reported CQI can be used to select modulation order, code rate, etc., for subsequent transmissions performed by the gNB using the reported precoding matrices.

[0337] As mentioned above, CQI can be calculated based on the precoding matrix. Therefore, accurately determining channel quality information (e.g., CQI) may involve or require the UE to know or assume the precoding information (e.g., the precoding matrix) applied by the gNB. In an NR system where the precoding matrix can be determined using a codebook (e.g., using the RI and / or PMI reported by the UE), the UE can know the precoding matrix applied by the gNB to the CSI-RS.

[0338] However, in systems using machine learning frameworks to report channel information, the UE may be unaware of or assume the precoding information applied by the gNB. For example, in some embodiments, model pairs (e.g., the encoder at the UE and the decoder at the gNB) can be trained such that when the channel matrix is ​​applied as input to the encoder at the UE, the decoder at the gNB can directly output the precoding matrix. Such embodiments may, for example, use... Figure 7 The configuration shown is implemented in which the channel state information 718 can be implemented as a channel matrix determined by measuring the reference signal 717 and applied to the machine learning encoder 703 at the UE 701. In such an embodiment, the reconstruction of the channel state information 722 can be implemented as a precoding matrix, wherein the precoding matrix can be obtained as the output of the machine learning decoder 704 at the gNB 702 (e.g., directly).

[0339] For example, you can use Figure 3The configuration shown is used to train such an embodiment, wherein training data 311 may include data pairs, each pair of which may include a channel matrix and a corresponding precoding matrix computed by the UE based on the channel matrix. During training, the channel matrix may be applied as input to the generative model 303, and the corresponding precoding matrix may be used as a training objective (e.g., for the loss function 313), thereby prompting the reconstructed model 304 to generate output data 312 as the precoding matrix.

[0340] However, in such an embodiment, the UE may not be able to access the trained decoder used by the gNB. For example, see reference... Figure 7 The encoder 703 and decoder 704 can be trained by the network, where the network may only transmit encoder 703 to UE 701 and only transmit decoder 702 to gNB 702. As another example, gNB 702 can train encoder 703 and decoder 704 and only transmit encoder 703 to UE 701. Therefore, the UE may not be able to determine the precoding matrix that it can report to gNB (e.g., in the form of the output of encoder 703), which gNB 702 can reconstruct using decoder 704. (Although gNB 702 can reconstruct the precoding matrix, in some systems, it may not be necessary to use the reconstructed precoding matrix or any precoding matrix previously reported by UE 701.) Therefore, UE 701 may not have access to the precoding matrix used to calculate channel quality information. Depending on the implementation details, this may lead to a mismatch between the precoding information and the channel quality information reported by the UE to gNB. In some embodiments, as used in this context, a mismatch may indicate a situation where the channel quality information reported to the wireless device may not be adequately based on the corresponding precoding information reported to the wireless device.

[0341] In some embodiments according to this disclosure, a first wireless device (e.g., a UE) may determine precoding information (e.g., a precoding matrix) used by a second wireless device (e.g., a gNB) to enable the first wireless device to determine channel quality information (e.g., CQI) based on the precoding information. Depending on the implementation details, this may reduce or eliminate mismatches between the precoding information and the channel quality information reported by the first wireless device to the second wireless device.

[0342] Figure 16 An embodiment of a system having a model pair to provide channel information feedback according to the present disclosure is shown. Figure 16 The system 1600 shown can be used to implement any apparatus, model, training scheme, etc. disclosed herein, or can be implemented using any apparatus, model, training scheme, etc. disclosed herein. System 1600 may include components that can be used with… Figure 4 and / or Figure 7 One or more elements (e.g., components, operations, etc.) in the embodiments shown are similar, wherein similar elements may be indicated by reference numerals that end with and / or contain the same numbers, letters, etc.

[0343] exist Figure 16 In the system 1600 shown, a first wireless device (e.g., UE) 1601 can receive a signal (e.g., a reference signal) 1617 from a second wireless device (e.g., a base station) 1602 via channel 1615, for example, to enable the first wireless device 1601 to determine the channel conditions of channel 1615. The second wireless device 1602 can apply precoding information (e.g., a precoding matrix) 1645 to subsequent transmissions (e.g., PDCCH, PDSCH, etc.).

[0344] The first wireless device 1601 may include precoding determination logic 1643. Additionally or alternatively, the second wireless device 1602 may include sharing logic 1644. The precoding determination logic 1643 and / or the sharing logic 1644 may be implemented individually and / or collaboratively to enable the first wireless device 1601 to determine (e.g., compute) precoding information 1645 (or its reconstruction), wherein the precoding information 1645 (or its reconstruction) further enables the first wireless device 1601 to determine channel quality information 1659. Depending on the implementation details, enabling the first wireless device 1601 to determine the precoding information 1645 applied by the second wireless device 1602 can help reduce or eliminate mismatches between the precoding information 1645 (or its reconstruction) and the channel quality information 1659 reported by the first wireless device 1601 to the second wireless device 1602.

[0345] In some example embodiments, the first wireless device 1601 may use a reconstruction model (e.g., a decoder model) to generate a reconstruction 1646 of the precoded information 1645. For example, shared logic 1644 at the second wireless device 1602 may transmit a reconstruction model 1604 (e.g., information about model size, dimensions, weights, etc.) to the first wireless device 1601, wherein the transmitted model may be indicated as shared model 1604A.

[0346] Such an embodiment can be implemented, for example, using a generative model 1603 implemented with an encoder and a reconstructive model 1604 implemented with a decoder. The second wireless device 1602 can train the encoder 1603 and the decoder 1604 (e.g., in an autoencoder configuration). Shared logic 1644 at the second wireless device 1602 can, for example, use over-the-air (OTA) transmission to transmit both the encoder 1603 and the decoder 1604A to the first wireless device 1601. The first wireless device 1601 can use the encoder 1603 to generate a representation 1607 based on channel information 1605, where the channel information 1605 may include, for example, channel measurements (e.g., a channel matrix). The second wireless device 1602 can use the decoder 1604 to generate precoding information (e.g., a precoding matrix) 1645 from the representation 1607. However, since the first wireless device 1601 can also receive the decoder 1604A, it can use the decoder 1604A to generate reconstructed precoding information (e.g., a precoding matrix) 1646, which can be the same as or similar to the precoding information 1645. The first wireless device 1601 can then use the reconstructed precoding information 1646 to perform calculation 1660 to determine channel quality information 1659.

[0347] The first wireless device 1601 may transmit channel quality information 1659 to the second wireless device 1602 in any suitable manner. For example, the first wireless device 1601 may combine the channel quality information 1659 with a representation 1607 generated by the generative model 1603 (e.g., by appending the channel quality information 1659 to the representation 1607) and transmit the combined channel quality information 1659 and representation 1607 to the second wireless device 1602 using another channel (e.g., an uplink channel), signal, etc. 1616. In some embodiments, the second wireless device 1602 may remove the channel quality information 1659 from the combined information before applying the representation 1607 to the reconstruction model 1604.

[0348] As another example, the first wireless device 1601 can receive the reconstructed model 1604A from a server on a wireless network on which both the first wireless device 1601 and the second wireless device 1602 can operate (e.g., using OTA transmission). In such an embodiment, the reconstructed model 1604A can be identified and / or registered to the network, and the network can activate the reconstructed model 1604A using a model identifier. If the reconstructed model 1604A is shared by the second wireless device 1602 and the first wireless device 1601, then the second wireless device 1602 can share the reconstructed model 1604A corresponding to the activated generated model 1603.

[0349] In some additional example embodiments, the first wireless device 1601 may use a locally trained reconstruction model 1647 to generate a reconstruction 1646 of the precoded information 1645. For example, the first wireless device 1601 may train a reference model (e.g., a reference decoder) 1647 that can match the reconstruction model 1604 used by the second wireless device 1602. The first wireless device 1601 may then use the locally trained model 1647 to generate (e.g., reconstruct) the precoded information 1646, and then the first wireless device 1601 may use the precoded information 1646 to determine (e.g., compute) channel quality information (e.g., CQI) 1659. One or more of these operations may be individually and / or collaboratively controlled, supported, etc., by the precoded determination logic 1643.

[0350] In some embodiments, the first wireless device 1601 may implement the reference model as a reference decoder model that the first wireless device 1601 can refine, for example, by fine-tuning (which may be the same as or different from the decoder model used by the second wireless device 1602). Alternatively or additionally, the first wireless device 1601 may implement the reference model using its own encoder model, wherein the first wireless device 1601 may use the encoder model to train the decoder model, which may be different from the decoder model used by the second wireless device 1602, but the first wireless device 1601 may still use the decoder model to determine the precoding matrix, which may then be used to determine the channel quality information to be reported to the second wireless device 1602.

[0351] In some further example embodiments where the first wireless device 1601 (e.g., UE) can report channel information 1605 (e.g., channel matrix) to the second wireless device 1602, a training dataset including precoding information (e.g., one or more precoding matrices) 1646 calculated by the first wireless device 1601 can be used. Therefore, the first wireless device 1601 can know the corresponding precoding information 1645 obtained at the second wireless device 1602. However, since the number of columns in the precoding matrix can correspond to the rank reported using RI, the second wireless device 1602 may not know the size of the precoding matrix, and consequently, the size of the corresponding UCI payload. In such embodiments, the first wireless device 1601 can report enough matrices (e.g., a complete N_T×N_T matrix) so that the second wireless device 1602 can determine the precoding information (e.g., the precoding matrix) 1646. For example, the second wireless device 1602 may subsequently use a specific number of the first columns of the matrix as a precoding matrix, wherein the number of the first columns can be determined by the rank as reported by the RI (e.g., determined to be equal to the rank as reported by the RI). In some embodiments implemented by the first wireless device 1601 using a UE, the UE may use existing algorithms to compute the precoding matrix and then compute the CQI based on the precoding matrix. One or more of these operations may be individually and / or collaboratively controlled, supported, etc., by the precoding determination logic 1643.

[0352] According to some additional embodiments of this disclosure, mismatches between precoding information and channel quality information can be reduced or eliminated by jointly processing (e.g., compressing) channel information (e.g., a channel information matrix such as a channel matrix (H), a precoding matrix (P), etc.) and channel quality information (e.g., CQI) that can be determined based on the channel information. For example, in some embodiments, a generative model and a reconstruction model (e.g., in an autoencoder configuration) can be trained to jointly compress and / or decompress channel information and channel quality information that can be determined based on the channel information.

[0353] Figure 17 An example embodiment of a model pair that can be used for joint compression of channel information and channel quality information according to this disclosure is shown. Figure 17 The embodiments shown can be used, for example, to implement Figure 4 The generative model 403 and / or reconstructed model 404 shown, and Figure 7 The encoder 703 and / or decoder 704 are shown.

[0354] refer to Figure 17The generative model can be implemented using the machine learning encoder 1703, and the reconstructed model can be implemented using the machine learning decoder 1704. The precoding matrix and CQI (for i = 1, ..., N, indicated by Precoder) for each subband i are... i and CQI i (where N can represent the number of subbands) can be used as input to encoder 1703, which can generate a joint representation 1707 of the input. Based on the joint representation 1707, a reconstructed precoding matrix and CQI (referred to as RecPrecoder for i = 1, ..., N) can be generated for each subband i. i and RecCQI i ), which is the output of decoder 1704.

[0355] The encoder 1703 and decoder 1704 can be trained using training data, wherein, for one or more (e.g., each) subband i, the training data may include target channel information (e.g., target CQI), wherein the target channel information may be, or can be based on, the precoding matrix for the subband and the CQI calculated by the UE based on the corresponding precoding matrix. The UE may use one or more existing algorithms for calculating the precoding matrix and / or CQI to calculate the precoding matrix and / or the corresponding CQI for the subband. Once the model pair is trained (e.g., when used for inference), the reconstructed precoding output RecPrecoder... i It can be compared with the corresponding reconstructed CQI output RecCQI. i Matching, for example, is possible because precoding information and corresponding CQIs can be matched in the training data used to train model pairs 1703 and 1704. Alternatively or additionally, the training data used to train model pairs 1703 and 1704 can be generated by any other source (e.g., network, base station, etc.), wherein any other source can provide training data that may include a sufficient match between precoding information (e.g., precoding matrix) and channel quality information (e.g., CQI). One or more of these operations can be controlled, supported, etc., individually and / or collaboratively by joint processing logic that may be located at the UE, gNB, network, or any other location or combination thereof.

[0356] The training of encoder 1703 and decoder 1704 can be performed at the UE, at the base station, at the network (e.g., a server), or at any other location, and the trained encoder 1703 and decoder 1704 can be transmitted to the location where they will be used for inference. For example, encoder 1703 can be transmitted to the UE (if it has not been trained at the UE), and decoder 1704 can be transmitted to the base station (if it has not been trained at the base station).

[0357] although Figure 17 The illustrated embodiments can be described as inputs and / or outputs within the context of one or more precoding matrices and / or CQIs, but any type of channel information, precoding information, etc., can be used as inputs and / or outputs. For example, in some embodiments, for each subband, a channel matrix can be applied as input to encoder 1703 in addition to or as a replacement for the precoding matrix. In such embodiments, for each subband, a reconstruction of the channel matrix can be reconstructed by decoder 1704 as output, in addition to or as a replacement for the reconstructed precoding matrix. In some embodiments where the channel matrix can be applied as a replacement for the precoding matrix as input to encoder 1703, for each subband, the UE can compute a precoding matrix as an intermediate result, from which a corresponding CQI applicable to encoder 1703 is computed.

[0358] Furthermore, although it can be described in the context of multiple subbands Figure 17 The embodiments shown can be applied, but the search principle can be used to train model pairs to jointly compress channel information and precoding information of multiple and / or single frequency bands, sub-bands, etc.

[0359] In some embodiments, joint compression (e.g., for joint reporting) can provide one or more other potential benefits beyond potentially providing matching between channel information components such as CQI and precoded information, depending on the implementation details. For example, using different ML models (e.g., encoders and / or decoders) to report each channel information quantity (e.g., RI, CQI, PMI, etc.) increases the time, complexity, etc., involved in training and / or inference for different models. However, embodiments of this disclosure can use joint compression (and / or joint reporting) on ​​values ​​such as RI, CQI, PMI, channel matrix, precoded matrix, etc., depending on the implementation details, which can reduce the time, complexity, etc., involved in training and / or inference. Some embodiments can increase the size of the model used for joint compression and / or reporting based on, for example, different distributions of channel information quantities to improve compression performance. In some embodiments, a wireless device (e.g., a UE) can be configured with different encoders and / or decoders, wherein each encoder, each decoder, each pair of models, etc., can be associated with an identifier (e.g., a model ID) and mapped to a specific combination of channel information quantities. In some embodiments, more than one model can be mapped to the same set of quantities, wherein different encoders, different decoders, different pairs of models (identified by different model IDs) can be used to address different channel environments, configurations, etc.

[0360] CQI compression across subbands

[0361] In NR systems, the UE can report CQI compressed using a differential scheme, where compression can be based on the CQI of different subbands deviating from the broadband value. For example, instead of using 4*N_sb bits (where N_sb can indicate the number of subbands) to report CQI, the UE can use 2 bits per subband to report both the broadband CQI value and the differential CQI value. However, this reporting scheme may not provide sufficient compression in some implementations, especially in systems where frequency division duplex (FDD) can be used and / or where there may be an increased antenna size with a relatively large amount of feedback reporting.

[0362] The scheme for compressing and / or reporting channel information using machine learning according to this disclosure can compress a combination of channel information from a set of subbands. Such embodiments can be advantageous, for example, for compressing and / or reporting quantities such as CQI, whereby the UE can be configured to report subband CQIs for one or more subbands (e.g., each subband or a subset of subbands) used for channel information reporting. Depending on the implementation details, the scheme for compressing and / or reporting channel information on a subband basis using machine learning according to this disclosure can leverage one or more correlations (e.g., in the time domain, frequency domain, and / or spatial domain) to improve performance and / or flexibility, reduce complexity, etc.

[0363] Figure 18 An embodiment of a system according to this disclosure having a model pair for providing channel information based on one or more subbands is shown. Figure 18 The system 1800 shown can be used to implement any apparatus, model, training scheme, etc. disclosed herein, or can be implemented using any apparatus, model, training scheme, etc. disclosed herein.

[0364] System 1800 may include components that can be connected with Figure 4 and / or Figure 16 In the embodiments shown, one or more elements (e.g., components, operations, etc.) are similar, wherein similar elements may be indicated by reference numerals ending with and / or containing the same numbers, letters, etc. However, in Figure 18 In the illustrated system 1800, the first wireless device 1801 may include subband compression logic 1848. Additionally or alternatively, the second wireless device 1802 may include subband decompression logic 1849. The subband compression logic 1848 and / or the subband decompression logic 1849 may be implemented individually and / or collaboratively to enable the first wireless device 1801 to compress a combination of channel information for a set (e.g., one or more) subbands.

[0365] For example, at a first wireless device 1801 (e.g., a UE), subband compression logic 1848 can concatenate the values ​​of channel information 1805 from multiple subbands into a vector, wherein the vector can be compressed by a generation model 1803 to generate a representation 1807 of the values ​​of the channel information 1805 from multiple subbands. At a second wireless device 1802 (e.g., a base station), a reconstruction model 1804 can recover the vector or an approximation of the vector, wherein subband decompression logic 1849 can divide the vector or an approximation of the vector into individual subbands to generate reconstructed channel information 1806 or an approximation of the channel information from multiple subbands.

[0366] Figure 19 An embodiment of a model pair that can be used for CQI compression across subbands according to this disclosure is shown. Figure 19 The illustrated embodiments can be used, for example, to implement Figure 18 The generation model 1803 and / or reconstruction model 1804 are shown. For illustrative purposes, when used with system 1900, the first wireless device 1801 and the second wireless device 1802 can be implemented as a UE and a gNB, respectively.

[0367] refer to Figure 19 A generative model can be implemented using a machine learning encoder 1903, and a reconstructed model can be implemented using a machine learning decoder 1904. CQI vectors (CQI1, CQI2, ..., CQI...) N The CQI vector (CQI1, CQI2, ..., CQIN) can include the CQI values ​​of one or more (e.g., each) of subbands 1, 2, ..., N (where N can indicate the number of subbands). N This can be achieved, for example, by using the individual CQI values ​​CQI1, CQI2, ..., CQI of the cascaded subbands. N To form the CQI vector, it can be applied to encoder 1903, such that each value CQI1, CQI2, ..., CQI... N It can be a single input to encoder 1903, such as Figure 19 As shown. Alternatively, or additionally, the entire CQI vector can be applied as a single input to encoder 1903. Encoder 1903 can use the CQI vector and / or individual CQI values ​​to generate a representation 1907 of channel information values ​​for multiple subbands.

[0368] Decoder 1904 can receive representation 1907 and generate the reconstructed CQI values ​​RCQI1, RCQI2, ..., RCQI of subbands 1, 2, ..., N. N As individual outputs and / or as reconstructed vectors (RCQI1, RCQI2, ..., RCQI... N It can be divided (e.g., by sub-band decompression logic) to provide individual outputs of one or more sub-bands.

[0369] In some embodiments, CQI values ​​are CQI1, CQI2, ..., CQI N One or more of these can be provided to the CQI table as inputs (e.g., discrete inputs) (e.g., CQI indices, which can be scalar values ​​between, for example, 0 and 15). For example, if the CQI table has sixteen rows, the CQI values ​​are CQI. i (Where i = 1, 2, ..., N) can be a 4-bit value. Alternatively or additionally, the CQI values ​​are CQI1, CQI2, ..., CQI... N One or more of these values ​​can be provided as continuous values ​​(e.g., arbitrary real values) of coding rate and / or modulation order. In embodiments using one or more continuous values ​​as CQI values, a modulation and coding scheme (MCS) table with relatively high resolution can be used, for example, to accommodate a finer-grained indication of the gNB.

[0370] Depending on the implementation details, as mentioned above... Figure 18 and Figure 19 The aforementioned scheme for compressing and / or reporting channel information (e.g., CQI) on a subband basis can leverage one or more correlations in the input data (e.g., between elements of a CQI vector) to improve performance and / or flexibility, reduce complexity, etc. For example, encoder 1903 may be able to compress a CQI vector with multiple elements having the same or similar values ​​to a greater extent than a CQI vector with a variety of different values.

[0371] In some embodiments, one or more of the encoder 1903, decoder 1904, generative model 1803, and / or reconstructive model 1804 may be implemented using a machine learning model, wherein the machine learning model may be adapted to work with inputs having different relevance properties. Thus, for example, the model for compressing CQI values ​​across subbands may be specialized (e.g., a specialized model as opposed to a general ML model) to improve or optimize compression performance based on, for example, the range of input values, vector and / or matrix dimensions, weights, etc.

[0372] Compression of channel information

[0373] In NR systems, the UE can compress the precoding matrix in the spatial and / or frequency domains using a codebook. For example, the UE can (e.g., to the gNB) report the RI and the PMI associated with the RI. The rank indicated by the RI and PMI can be used together to determine the precoding matrix. For example, the PMI can be selected from a set of supporting matrices that may be referred to as the codebook. In some embodiments, the codebook can be implemented as a mapping table between a set of "i" indices and precoding matrices. For example, the UE can use indices (i_1,1,i_1,2,i_1,3,i_2) with a Type 1 single-plane codebook, where the Type 1 single-plane codebook can uniquely define a precoding matrix of a given rank. However, codebook compression schemes may not provide sufficient compression in some implementations, especially in systems where frequency division duplex (FDD) can be used and / or where there may be an increased antenna size with a relatively large amount of feedback reporting.

[0374] The scheme for using machine learning to compress and / or report channel information according to this disclosure may use one or more decoder models to generate precoded information and / or other information that can be used, for example, to determine the precoded information. In some embodiments, such a scheme may mimic a codebook scheme and / or implement a hierarchical compression mechanism. In some embodiments, such a scheme may implement one or more encoders and / or decoders having specific structures for compression in the time domain, frequency domain, and / or spatial domain. Depending on the implementation details, such a scheme may improve performance and / or flexibility, reduce complexity, and so on.

[0375] Figure 20 An embodiment of a system having a model pair for providing channel information compression according to the present disclosure is shown. Figure 20 The system 2000 shown can be used to implement any apparatus, model, training scheme, etc. disclosed herein, or can be implemented using any apparatus, model, training scheme, etc. disclosed herein. System 2000 may include components that can be used with… Figure 4 , Figure 16 and / or Figure 18 One or more elements (e.g., components, operations, etc.) in the embodiments shown are similar, wherein similar elements may be indicated by reference numerals that end with and / or contain the same numbers, letters, etc.

[0376] exist Figure 20In the system 2000 shown, a first wireless device (e.g., a UE) 2001 can receive a signal (e.g., a reference signal) 2017 from a second wireless device (e.g., a base station) 2002 via channel 2015. The second wireless device 2002 can apply precoding information (e.g., a precoding matrix) 2045 to the signal 2017. The first wireless device 2001 may include compression scheme logic 2050. Additionally or alternatively, the second wireless device 2002 may include compression scheme logic 2051. Compression scheme logic 2050 and / or 2051 may implement one or more schemes individually and / or collaboratively to enable the first wireless device 2001 and the second wireless device 2002 to implement one or more compression schemes for generating a representation 2016 of channel information 2005.

[0377] Figure 20 The embodiments shown are not limited to any particular form of channel information 2005 applied to the generation model 2003 and / or any particular form of reconstructed channel information 2006 generated by the reconstruction model 2004. For example, in some example embodiments, the generation model 2003 and the reconstruction model 2004 may be implemented using an encoder and a decoder (e.g., configured as an autoencoder), respectively, wherein the encoder and decoder may be configured and / or trained to generate precoded information from channel information (such as a channel matrix) using one or more compression schemes. In such embodiments, a first radio device 2001, which may be implemented as a UE, may perform one or more CSI-RS measurements to obtain channel information, such as one or more channel matrices, feature vectors, etc. The UE may apply the channel information to the encoder to generate codewords that the UE may transmit to a second radio device 2002, wherein the second radio device 2002 may be implemented as a base station (e.g., a gNB). The base station may apply the codewords to the decoder to construct a precoded matrix. In such embodiments, codewords may emulate one or more codebook indices, and / or decoders may emulate the codebook; however, depending on the implementation details, a scheme with one or more decoders as described above may improve performance and / or flexibility, reduce complexity, etc.

[0378] However, in other embodiments, the reconstruction of channel information 2005 and / or channel information 2006 can be implemented in any other form. For example, instead of generating a precoding matrix as the output of generator model 2003 based on the channel information applied as input to generator model 2003, generator model 2003 and reconstruction model 2004 can be trained such that reconstruction model 2004 can reconstruct the channel information or its approximation applied as input to generator model 2003. The second wireless device 2002 can then use the reconstructed channel information to determine (e.g., compute) the precoding matrix. In other embodiments, the reconstruction of channel information 2005 and / or channel information 2006 can be implemented using CQI, RI, and / or any other type of information.

[0379] In embodiments where the encoder at the UE and the decoder at the base station can be configured and / or trained to generate precoded information such as a precoded matrix from channel information such as a channel matrix, the decoder's output (and post-processing, if any) can be a precoded matrix. In some such embodiments, the UE may also be configured with a decoder, wherein the decoder is trained to provide an output as a precoded matrix. The decoder at the UE may be obtained, for example, from the base station, wherein the base station may share a decoder model including weights, input and / or output dimensions, etc. In some embodiments, the UE and the base station may have a common understanding of the type of decoder output (e.g., how precoded information can be constructed from the decoder output).

[0380] In embodiments where post-processing schemes can be applied after the decoder, a pre-processing scheme can be shared with the UE. The UE can be provided with one or more implicit and / or explicit instructions on how to construct precoding information (e.g., a precoding matrix) from the decoder's output (and any post-processing, if any). For example, if the decoder's output dimension is N_t*v, where v is the rank indicated by RI, the precoder can be constructed by shaping the output vector into an N_t×v matrix on a column-by-column and / or row-by-row basis.

[0381] The following description is used for Figure 20 Some embodiments of the training and / or testing scheme of the system shown are illustrated. For illustrative purposes, embodiments of the training and / or testing scheme may be described in the context of a system in which the first wireless device 2001 and the second wireless device 2002 may be implemented using a UE and a gNB, respectively, and the generation model 2003 and the reconstruction model 2004 may be implemented using an encoder and a decoder, respectively. However, the principles are not limited to these or any other implementation details.

[0382] In a first embodiment of the training and / or testing scheme, the encoder and decoder can be jointly designed, configured, trained, etc., based on a training dataset (e.g., by the UE vendor and gNB vendor, respectively). The training dataset can include a set (X, Y), where X can refer to one or more CSI-RS measurements, such as the channel matrix of a set of CSI-RS measurements obtained, for example, within a time and frequency window, and Y can refer to a precoding matrix that can be computed, for example, by the UE. The decoder can be shared with the UE (e.g., transmitted to the UE) for use during inference, for example, to generate a precoding matrix that the UE can use for CSI-RS measurements. During inference, the UE can apply the encoder in a manner similar to that used during training.

[0383] In a second embodiment of the training and / or testing scheme, the gNB can train the decoder using a training dataset that may or may not be available from the UE. The gNB can share the decoder with the UE, including one or more weights, input and / or output dimensions, post-processing, etc. In some embodiments, the encoder training, as well as one or more input types and / or dimensions, can be determined by the UE based on, for example, its implementation. When training the encoder, the UE may seek to minimize the loss function between the output of the decoder P_dec and the desired target pre-encoder P_target, which can be computed by the UE. In some embodiments, the gNB can configure one or more time and / or frequency windows for the UE, and the CSI-RS used within the windows, which can be used as training data to train the encoder. The UE can then use the trained encoder for inference.

[0384] In a third embodiment of the training and / or testing scheme, the gNB can train the decoder based on a training dataset that may or may not be obtained from the UE, and can share the decoder with the UE, including one or more weights, input and / or output dimensions, post-processing, etc. In some embodiments, the training of the encoder and one or more input types and / or dimensions can be determined by the UE based on, for example, the implementation of the UE. However, in this type of embodiment, the gNB can also share a training dataset for training the encoder with the UE, including one or more inputs and / or labels. The UE can then use the shared training dataset to train the encoder for use during inference.

[0385] In any of the three embodiments of the training and / or testing schemes described above, because the precoding matrix domain can depend on the reported rank, the UE can apply one or more different models for one or more different reported ranks. For example, if the UE reports a first rank as the RI, the UE can use a first encoder model trained for a first decoder, and if the UE reports a second rank as the RI, the UE can use a second encoder model trained for a second decoder model.

[0386] In any of the three embodiments of the training and / or testing schemes described above, depending on the implementation details, the decoder may be characterized as performing a role similar to that of the NR codebook; however, compared to the use of the codebook, it has the potential to improve performance and / or flexibility, reduce complexity, etc., by using a machine learning decoder model.

[0387] Figure 21 A first embodiment of a model pair that can be used to implement a compression scheme according to the present disclosure is shown. Figure 21 The embodiment 2100 shown can be used, for example, to implement Figure 20 The generative model 2003 and the reconstructive model 2004 are shown in the figure.

[0388] refer to Figure 21 Embodiment 2100 may include a generation model 2103 and a reconstruction model 2104 that can be configured, for example, to operate as an autoencoder. The generation model 2103 (which may be located, for example, at the UE) may include one or more spatial encoders 2153-1 to 2153-N, wherein the spatial encoders 2153-1 to 2153-N may be arranged to receive one or more spatial encoder inputs SEI-1 to SEI-N corresponding to channel information of subbands 1, 2, ..., N, respectively, where N may indicate the number of subbands. The one or more spatial encoders 2153-1 to 2153-N may generate one or more spatial compression outputs SEO1 to SEO-N based on the corresponding encoder inputs SEI-1 to SEI-N, respectively.

[0389] The compressed outputs SEO-1 to SEO-N can be transmitted to a reconstruction model 2104 (which may be located, for example, at a base station), in which they can be applied as sub-band decoder inputs SDI-1 to SDI-N to one or more spatial decoders 2154-1 to 2154-N, respectively. The one or more spatial decoders 2154-1 to 2154-N can spatially decompress the decoder inputs SDI-1 to SDI-N to generate one or more decompressed outputs SDO-1 to SDO-N, respectively.

[0390] While not limited to any particular implementation details, in some embodiments, the encoder inputs SEI-1 to SEI-N can be implemented as channel matrices H_1^i (where i = 1, 2, ..., N) for the corresponding subbands, and the spatial encoder outputs SEO-1 to SEO-N can be implemented as representations v_1^i of the precoding matrices P_1^i for the corresponding subbands. Then, the spatial decoders 2154-1 to 2154-N can generate decompressed outputs SDO-1 to SDO-N from the representation v_1^i as reconstructed precoding matrices P_1^i. Therefore, the reporting of precoding information can be performed independently in different subbands.

[0391] In some embodiments, spatial encoders 2153-1 to 2153-N can perform spatial compression operations in parallel, for example, in embodiments where spatial encoders 2153-1 to 2153-N can be implemented using more than one set of hardware (e.g., a separate processor, circuitry, etc. for each encoder). In some other embodiments, one or more of spatial encoders 2153-1 to 2153-N can perform spatial compression operations sequentially, for example, in embodiments where spatial encoders 2153-1 to 2153-N can be implemented using fewer than N hardware instances (e.g., a single processor, circuitry, etc. for all N encoders). Similarly, spatial decoders 2154-1 to 2154-N can perform spatial decompression operations in parallel or sequentially, depending on, for example, the number of hardware groups used to implement the decoder.

[0392] Figure 22 A second embodiment of a model pair that can be used to implement a compression scheme according to the present disclosure is shown. Figure 22 The embodiment 2200 shown can be used, for example, to implement Figure 20 The generative model 2003 and the reconstructed model 2004 are shown. Example 2200 may include models that can be used with... Figure 21 One or more elements (e.g., components, operations, etc.) in the embodiments shown are similar, wherein similar elements may be indicated by reference numerals that end with and / or contain the same numbers, letters, etc.

[0393] refer to Figure 22Embodiment 2200 may include a generation model 2203 and a reconstruction model 2204, which may be configured to operate, for example, as an autoencoder. The generation model 2203 (which may be located, for example, at the UE) may include one or more spatial encoders 2253-1 to 2253-N, wherein the spatial encoders 2253-1 to 2253-N may be arranged to receive one or more spatial encoder inputs SEI-1 to SEI-N corresponding to channel information of subbands 1, 2, ..., N, respectively, where N may indicate the number of subbands. The one or more spatial encoders 2253-1 to 2253-N may generate one or more spatial compression outputs SEO1 to SEO-N based on the corresponding encoder inputs SEI-1 to SEI-N, respectively.

[0394] The generative model 2203 may also include a frequency encoder 2254, which may be arranged to receive one or more spatially compressed outputs SEO-1 to SEO-N and generate spatially and frequency-compressed representations FEO of the sub-band encoder inputs SEI-1 to SEI-N.

[0395] The compressed representation FEO can be transmitted to a reconstruction model 2204 (which may be located, for example, at a base station), where it can be used as input to a frequency decoder 2256. The frequency decoder 2256 can generate one or more spatially decompressed outputs SDI-1 to SDI-N corresponding to subbands 1, 2, ..., N, respectively. These spatially decompressed outputs SDI-1 to SDI-N can be decompressed in the frequency domain but still compressed in the spatial domain. The N outputs of the frequency decoder 2256 can be used as inputs SDI-1 to SDI-N to one or more spatial decoders 2254-1 to 2254-N. The one or more spatial decoders 2254-1 to 2254-N can spatially decompress the decoder inputs SDI-1 to SDI-N to generate one or more decompressed outputs SDO-1 to SDO-N, which can be decompressed in both the frequency and spatial domains. Therefore, spatial compression and decompression of precoded information can be performed independently in different subbands, while frequency compression and decompression can be performed simultaneously across subbands.

[0396] While not limited to any particular implementation details, in some embodiments, the inputs SEI-1 to SEI-N of the spatial encoders 2253-1 to 2253-N can be implemented as channel matrices H_1^i (where i = 1, 2, ..., N) for the corresponding subbands, and the spatial encoder outputs SEO-1 to SEO-N can be implemented as spatially compressed representations v_1^i of the precoding matrices P_1^i for the corresponding subbands. The frequency encoder 2255 can compress the representation v_1^i in the frequency domain to generate an output FEO as a representation vector v, which can be compressed in both the frequency and spatial domains. The frequency decoder 2256 can decompress the representation vector v in the frequency domain to reconstruct the spatially compressed representation v_1^i, which can be used as one or more inputs SDI-1 to SDI-N to one or more spatial decoders 2154-1 to 2154-N. Then, the spatial decoders 2154-1 to 2154-N can generate decompressed outputs SDO-1 to SDO-N from the representation v_1^i as the reconstructed precoding matrix P^_1^i.

[0397] Therefore, in Figure 22 In the illustrated embodiment, the generation model 2203 can operate in two stages. In the first stage, precoded information can be compressed independently at each subband via spatial domain compression. In this stage, N encoder models can be used, for example, one encoder model per subband. The models can be the same or different, depending on various parameters, such as those associated with different subbands within the channel. In the second stage, the spatially compressed precoded information of the subbands can be jointly compressed in the frequency domain using a single coding model. In some embodiments, the outputs SEO-1 to SEO-N of one or more spatial encoders 2253-1 to 2253-N can represent the precoded matrix for each subband without using inter-subband correlations. Subsequently, the frequency decoder 2256 can perform frequency compression on the spatially compressed precoded matrices for each subband and compress them into a vector v, where the vector v can include information for representing all precoded matrices. Therefore, Figure 22 The embodiments shown can provide frequency domain decompression and a decoder model for each subband that can reconstruct the precoding matrix for each subband.

[0398] In some embodiments, spatial encoders 2253-1 to 2253-N and / or frequency encoders 2255 may perform compression operations in parallel, for example, in embodiments where spatial encoders 2253-1 to 2253-N and frequency encoders 2255 may be implemented using more than one set of hardware (e.g., a separate processor, circuitry, etc. for each encoder). In some other embodiments, one or more of spatial encoders 2253-1 to 2253-N and / or frequency encoders 2255 may perform compression operations sequentially, for example, in embodiments where spatial encoders 2253-1 to 2253-N and / or frequency encoders 2255 may be implemented using fewer than N+1 hardware instances (e.g., a single processor, circuitry, etc. for all N+1 encoders). Similarly, depending on, for example, the number of hardware sets used to implement the decoder, spatial decoders 2254-1 to 2254-N and / or frequency decoders 2256 may perform spatial decompression operations in parallel or sequentially.

[0399] Figure 23 A third embodiment of a model pair that can be used to implement a compression scheme according to this disclosure is shown. Figure 23 The embodiment 2300 shown can be used, for example, to implement Figure 20 The generative model 2003 and the reconstructed model 2004 are shown. Example 2300 may include models that can be used with... Figure 21 and / or Figure 22 One or more elements (e.g., components, operations, etc.) in the embodiments shown are similar, wherein similar elements may be indicated by reference numerals that end with and / or contain the same numbers, letters, etc.

[0400] refer to Figure 23 Embodiment 2300 may include a generation model 2303 and a reconstruction model 2304 that can be configured, for example, to operate as an autoencoder. The generation model 2303 (which may be located, for example, at the UE) may include a joint spatial and frequency encoder 2357, wherein the joint spatial and frequency encoder 2357 can receive one or more encoder inputs SFEI-1 to SFEI-N corresponding to channel information of subbands 1, 2, ..., N, respectively, where N may indicate the number of subbands. The joint spatial and frequency encoder 2357 can generate an output SFEO, which may be a spatial and frequency compressed representation of the encoder inputs SFEI-1 to SFEI-N.

[0401] Spatial and frequency compression representation SFE0 can be transmitted to reconstruction model 2304 (which may be located, for example, at a base station), where it can be used as input to joint spatial and frequency decoder 2358. Joint spatial and frequency decoder 2358 can generate one or more decoder outputs SFDO-1 to SFDO_N corresponding to subbands 1, 2, ..., N, respectively, which can be decompressed in both the spatial and frequency domains. Therefore, spatial and frequency compression and decompression can be performed simultaneously across subbands.

[0402] While not limited to any particular implementation details, in some embodiments, the inputs SFEI-1 to SFEI-N can be implemented as channel matrices H_1^i (where i = 1, 2, ..., N) corresponding to the sub-bands, and the output SFEO can be implemented as a representation vector v, where the representation vector v can be a spatially and frequency-compressed representation of the N precoding matrices corresponding to the N sub-bands. The joint spatial and frequency decoder 2358 can decompress the representation vector v in the spatial and frequency domains to recover the reconstructed precoding matrix P_1^i. Therefore, in some embodiments, in Figure 23 In the illustrated embodiment 2300, the joint spatial and frequency encoder 2357 can be characterized as jointly computing the precoding matrices P_1^N and compressing them into vector v, and the joint spatial and frequency decoder 2358 can be characterized as restoring the matrices to P_1^N.

[0403] In some embodiments, spatial correlation can refer to the correlation across one or more transmit (Tx) antenna ports. Any of the embodiments described above can implement spatial compression individually (e.g., on a per-layer basis) or jointly across multiple (e.g., all) layers of the channel. For example, in some embodiments, a precoding matrix can be applied as input to the encoder. For example, a 4x3 precoding matrix can have three layers (e.g., columns), where each layer has a precoding vector of length four (e.g., four rows). In some embodiments, spatial compression can be jointly implemented across multiple layers (e.g., all layers) by applying the entire matrix to the encoder. Alternatively or additionally, spatial compression can be implemented on a per-layer basis, for example, by applying layers (e.g., columns) of the matrix to the encoder one at a time (e.g., a vector (column) with four elements (rows) at a time). The number of layers can be indicated, for example, by RI.

[0404] In an NR system, one or more of the following constraints can be applied to channel information reporting. Assuming the following dependencies between CQI parameters (if reported), the UE can calculate the CSI parameters (if reported): LI can be calculated conditioned on the reported CQI, PMI, RI, and CRI; CQI can be calculated conditioned on the reported PMI, RI, and CRI; PMI can be calculated conditioned on the reported RI and CRI; and / or RI can be calculated conditioned on the reported CRI.

[0405] In some embodiments, one or more of the ML-based compression and / or reporting schemes disclosed herein may implement one or more constraints, wherein, depending on the implementation details, the one or more constraints may be similar to the constraints described above. Therefore, in some embodiments, if the UE reports a CQI, the CQI for one or more subbands (e.g., each subband) can be calculated based on the reported precoding information (e.g., a precoding matrix) for that subband. For example, utilizing information about... Figure 21 The described embodiments allow for the calculation of the CQI for a sub-band based on a precoding matrix P^_i, where the precoding matrix P^_i may involve a decoder shared with the UE. In embodiments where a decoder may not be shared with the UE, the UE may calculate the CQI based on information about each sub-band #i. As another example, utilizing information about... Figure 22 and Figure 23 In the described embodiments, if a decoder is shared with the UE, the UE can calculate the CQI for subband #i based on subband PMIi or decoder output P^_i. In some embodiments, for the purpose of specifying this condition, the precoding information (e.g., PMI) for each subband can indicate the subband encoder output (or input) or the corresponding decoder output, depending on the implementation details, which may assume that each of them represents a subband precoding matrix.

[0406] In an NR system, the bit width of the reported channel information can depend on one or more RRC parameters. In some embodiments, the bit width used to report precoding information (e.g., PMI) can additionally depend on the rank reported by the RI. Within the framework of this disclosure for using machine learning to report channel information, if the RI, CQI, and encoder codewords (e.g., the encoder output representing channel information) are all transmitted in the same PUCCH, one or more mechanisms can be implemented to ensure a shared understanding between the UE and the base station regarding the size of the UCI payload. Because the bit width used for precoding information (e.g., PMI) can depend on the RI, in some embodiments, the precoding information can be reported separately in a PUCCH other than the one carrying the RI. In some other embodiments, the UCI payload can be appended with one or more elements (e.g., zeros), where the number of elements can depend on the indicated rank. For example, if N max If N(v) is the maximum UCI payload size at all possible supported ranks, then for a given rank v used to report a payload size utilizing N(v), the UE can use N... max-v A zero is added to the UCI payload.

[0407] Processing time

[0408] In some embodiments (e.g., as part of lifecycle management (LCM), the base station may instruct the UE to update the currently active model (e.g., fine-tune a new dataset), switch to a new model, activate a new model, deactivate a model, etc. When updating the model, if the UE is instructed to update the encoder model based on an online training set (which may be collected via channel estimation, RRC configuration, MAC-CE activation, etc.), the UE may use (e.g., require) a minimum amount of time to update the model. The UE may use (e.g., require) a minimum amount of time regardless of whether the UE will share its updated model with the base station.

[0409] In some embodiments, if the online training set is collected by the UE via online channel estimation, it is not expected that the UE will update the model before the following time period expires, wherein the time period can be represented as, for example, a plurality of symbols starting from the end of the last symbol of the latest CSI-RS used for the online training set (which can be referred to as, for example, N_(AIML,upadte) symbols). In some embodiments, if the UE is configured to report the updated model to the base station, it is not expected that the UE will report the model to the gNB earlier than the following time period, wherein the time period can be represented as, for example, a plurality of symbols starting from the last symbol of the latest CSI-RS used in the training set (which can be referred to as, for example, N_(AIML,report) symbols).

[0410] In some embodiments, if the training set is configured to the UE by RRC, the UE may not be expected to update, or update and report the encoder earlier than the following time amount, wherein the time amount may be represented as, for example, a number of symbols (which may be referred to as, for example, N symbols) starting from the latest symbol that may have triggered the update command.

[0411] In some embodiments, when a UE is instructed to switch to a new model, a processing time similar to that provided for updating the model can be provided to the UE. For example, if the UE receives a command (e.g., via DCI, MAC CE, RRC, etc.) to switch to a new model, it may not be expected that the UE will switch to the model before a time period (which may involve activating the model), wherein the time period may be represented as a number of symbols (which may be referred to as, for example, N_(AIML,switch) symbols) starting from the end of the last symbol (e.g., the last symbol of the PDCCH) to which the switching command can be delivered to the UE.

[0412] In some embodiments, a similar processing time may be provided to the UE to activate the model. For example, if the UE receives a command (e.g., via DCI, MAC CE, RRC, etc.) to activate and / or deactivate the model, it may not be expected that the UE activates and / or deactivates the model before the following time period, wherein the time period may be represented as a plurality of symbols (which may be referred to as, for example, N_(AIML,activate) and / or N_(AIML,deactivate) symbols) starting from the end of the last symbol (e.g., the last symbol of the PDCCH) that the handover command can be delivered to the UE.

[0413] In some embodiments, one or more of the processing times described above can be described in time units based on the subcarrier spacing (SCS) numbers of the cell delivering the command and / or the cells on which the test model can be sent (e.g., the cells on which the corresponding CSI report can be sent). In some embodiments, the SCS of the cells where the model is active can also be considered. For example, multiple symbols (such as those mentioned in the methods described above) used to represent the time quantity can be described based on the minimum SCS among the SCS values.

[0414] In some embodiments, the UE may declare (e.g., based on a query performed by the base station) one or more processing time capabilities, such as PDSCH and / or PUSCH processing time capabilities. Additionally or alternatively, the UE may also declare any processing time capabilities disclosed herein. For example, in some embodiments, the UE may declare a minimum amount of time or a minimum number of symbols that it can support (e.g., can use or needs) for N_(AIML,upadte), N_(AIML,switch), and N_(AIML,activate) / N_(AIML,deactivate). For example, the UE may declare to the base station its capability for model switching with 20 symbols (e.g., minimum processing time), which could mean that the UE may not be expected to switch to the new model before 20 symbols after the end symbol of the switching command.

[0415] Reports using individual models

[0416] Within the AI / ML channel information reporting framework, if the UE reports the channel matrix, the gNB can infer the channel quality, including rank, eigenvectors, and / or eigenvalues, and thus the SVD precoding matrix. However, since performance can depend on the UE's downlink signal processing, including the UE's ability to support ranks and its UE-specific precoding information (e.g., precoding matrix) calculation algorithm, the UE may also report one or more other values, such as RI and / or CQI, within the AI / ML CSI reporting framework. For illustrative purposes, some embodiments of AI / ML (which may be referred to as ML)-based reporting schemes can be described in the context of systems that perform AI / ML CSI reporting by reporting CSI matrices (e.g., channel matrix, precoding matrix, etc.) via an autoencoder; however, the principles are not limited to use with an autoencoder.

[0417] In some embodiments, in addition to CSI matrix reporting, the UE can also report RI and CQI via existing reporting or an AI / ML framework. Using an AI / ML framework, the UE can be configured with subbands to report the CSI matrix for each subband or subset of subbands based on a bitmap. Some examples of RI, PMI, and / or CQI reporting are described below.

[0418] In NRC SI reporting, a UE can report only one RI for all subbands specified in the CSI reporting configuration. The same approach can also be used for some embodiments of the AI / ML CSI reporting framework. In some additional embodiments, one or more RI values ​​corresponding to one or more subbands can be reported by concatenating one or more RI values ​​corresponding to one or more (e.g., each) subbands into a vector and compressing that vector, depending on the implementation details; this can be similar to CSI matrix compression. Since the relevance properties of the RI vector can differ from those of the CSI matrix, a different ML model (e.g., a network) can be used to compress the RI vector than the ML model that can be used for CSI matrix compression. For example, the ML model used for RI reporting can have input and / or output dimensions, weights, etc., that can be customized for compressing the RI vector.

[0419] CSI matrix reporting can be performed according to any scheme disclosed herein. In some embodiments, the UE can be configured to report the CSI matrix for each subband or subset of subbands according to the CSI reporting configuration.

[0420] CQI reporting based on an AI / ML framework can be performed using any of the schemes disclosed in this article.

[0421] In some embodiments, if the UE is configured to report one or more of the three quantities discussed above (e.g., RI, CSI, and / or CQI) separately in different PUCCHs, then to reduce the input size of the second model (CSI matrix reporting model), the UE can use temporal domain multiplexing (TDM) to report different quantities (e.g., the three quantities) at different times via the PUCCH. For example, the UE can report RI in the first PUCCH. The UE can report a CSI matrix with dimensions that may depend on the reported RI in the second PUCCH, for example, because the gNB can only decode the second PUCCH after successfully decoding the first PUCCH. Depending on the implementation details, decoding of the first and third PUCCHs can be performed at any time, but the second PUCCH can only be decoded after the first PUCCH is decoded. One or more of these aspects may be based on assumptions that PUCCH decoding may involve the decoder model of the autoencoder running at the gNB.

[0422] Model ID-based lifecycle management

[0423] According to this disclosure, some communication systems can implement LCM schemes based on model IDs. In some embodiments of model-ID-based LCM, one or more UE capabilities can be based on multiple models and / or functions. In some model-ID-based LCM schemes, models can be identified and / or registered to the network. However, it may be redundant for the UE to identify two different models that are logically identical or similar and correspond to the same or similar physical models. Therefore, some embodiments may assume that each logical model identified by the base station has a distinct physical model. Alternatively, a logical model may correspond to multiple different physical models (e.g., the UE may have low-complexity models and high-complexity models for a given artificial intelligence and / or machine learning features and / or functions). Once a model is identified, the UE can report to the base station which models it supports. For example, due to the nature of model identification and capability reporting, some embodiments can support a relatively large number of models in function-based LCM systems and / or model-ID-based LCM.

[0424] However, one potentially problematic aspect is the possibility of having multiple physical models for a single logical model. In some embodiments, this type of condition can be handled in the UE capability report because the UE can be aware of the condition and can adjust the capability report to suit it.

[0425] In some embodiments, a UE capability report may include one or more details that enable the UE to report to the base station which models among the identified models it can support. In a reporting (e.g., signaling) scheme for such details, a portion of the report (e.g., a first portion) may include the UE vendor ID. Additionally or alternatively, another portion of the report (e.g., a second portion) may include one or more indices of supported models identified by the UE vendor. For example, a bitmap of length N may be used (e.g., by the vendor) to identify N models to the base station, such that each of the N elements of the bitmap can correspond to an identified model (e.g., in ascending or descending order of model ID). From the perspective of the UE and / or the base station, once the vendor is identified, a local logical ID from 0 to N-1 can be assigned to the identified model. The UE can then use the bitmap to declare support for the identified model.

[0426] In addition to indicating one or more supported model IDs to the base station, the UE can also declare (e.g., as a capability report) the maximum number of models it can support that are active simultaneously (e.g., N). active,max Then, the UE can anticipate that the base station will not activate more than N simultaneously. active,max One model.

[0427] In some embodiments, to accommodate the additional complexity arising from running multiple physical models for the same identified logical model, the UE may also report a scalar value α for each supported model IDi. i >1, where α can be identified to the network during model identification. The UE may not expect the base station to be active more than N simultaneously. active,max One model can be considered where the scaling value α is taken into account. e To calculate the number of activated models. Therefore, if the base station activation has a set If the model has an ID, then the number of activated models can be calculated as: Number of activated models = ∑ i∈S α e (Equation 8)

[0428] UE can expect it to be no greater than N active,max .

[0429] Some additional aspects of this disclosure relate to the coexistence of metadata and / or other RRC configurations. When a model is identified, a type of information called metadata can be disclosed to the network. In different embodiments, metadata may include different types of information. For example, in some embodiments, metadata may include some (e.g., all) information that can be used by the model for inference operations. In such embodiments, a base station may activate a specific model that may be similar or equivalent to perform some (e.g., all) RRC configurations that can be used by the model for inference operations. When a base station activates a model ID, it may have the effect of activating a specific CSI reporting configuration. In some NR implementations, CSI reporting configurations may include information about subband configurations, codebook types, etc., and all of this information may be included in the metadata. Such an approach can reduce RRC overhead, but may reduce the flexibility of the base station because, for example (1) some (e.g., each) functions and / or RRC configurations can be done via model identification, and / or (2) unlike some NR implementations, the base station may not configure metadata.

[0430] Some additional aspects of this disclosure relate to model delivery and / or transmission. In some communication systems, model delivery (e.g., downloading a model from a server) may not be part of the specification. However, the specification may include model delivery (e.g., transmitting a model to a UE covered by the specification using signaling provided by the specification). One or more of the following techniques can be implemented to transmit a model to a UE according to some embodiments of this disclosure.

[0431] Some embodiments can transmit one or more weights of a known model structure. For example, to transmit a model to a UE, a model structure can be established between the UE and the base station (e.g., by specifying a pool of structures such as standard structures), each of which may have an ID. The UE can declare its support for certain structures as UE capability signaling. To transmit the model, the base station can indicate the structure ID of the specified structure to the UE and then use a downlink channel (e.g., PDSCH) to send N weights. The mapping between vectorized weights and network structures can be specified in the specification. The base station can also indicate the model ID when transmitting network weights to the UE.

[0432] Some implementations may transmit the model as a model ID. For example, the UE and base station vendor may develop multiple models offline and store them on a server and / or in a registry. The base station can then indicate the applicable model ID to the UE. The UE can then use the model ID to retrieve the model from the server.

[0433] Some embodiments may utilize the indication of a dataset ID and / or an additional ID (e.g., an additional condition ID) to enable model delivery. For example, a base station may indicate a dataset ID or an additional condition ID to the UE. The UE may know which model to use based on the mapping between the model trained during the training phase and the dataset ID. Therefore, in some embodiments, one or more IDs intended for different purposes may be used as model IDs.

[0434] Any techniques discussed above related to model ID-based LCM can be applied in the context of CSI compression, CSI prediction, etc.

[0435] Some additional aspects of this disclosure relate to model identification schemes for CSI feedback in systems using Non-Coherent Joint Transport (NCJT). For example, utilizing NCJT CSI reporting for a UE receiving transmissions from multiple Transmitter-Receiver Points (TRPs), the UE can report CSIs for different TRP (e.g., base station) transmission assumptions. The UE can be configured to report X CSIs (where X = 0, 1, 2) associated with a single TRP measurement assumption and one CSI associated with an NCJT measurement assumption. These different assumptions may result in different precoding matrices to be reported. Since the UE can report different precoder matrices for different transmission schemes, the AIML encoder / decoder pair (e.g., autoencoder) can be different for different assumptions. One or more of the following techniques for NCJT CSI reporting can be implemented according to some embodiments of this disclosure.

[0436] Regarding model IDs, if the NCJT CSI report corresponds to M different hypotheses, the corresponding model IDs can include M individual model IDs (e.g., pairing IDs) for the M different autoencoders. Therefore, model IDs can be represented as M tuples (ID#1, ..., ID#M). Additionally or alternatively, it can be assumed that the UE can use the same model for any sTRPCSI. For example, if the UE reports CSIs for three hypotheses: sTRP#1, sTRP#2, and NCJT transmission, the IDs can take the form of a pair (ID for the model used for sTRP, ID for the model used for NCJT).

[0437] In some embodiments, the LCM of the model used for NCJT CSI feedback can be implemented in a manner similar to that used for a single TRP report, with the following differences: (1) for model activation, switching, deactivation and / or other model ID-based LCMs, the base station can indicate the model in the form of NCJT (e.g., tuples); and / or (2) the CSI report configuration for NCJT can include IEs indicating the model ID in the form of tuples.

[0438] CSI payload size based on the number of models and / or features

[0439] Using a relatively small number of encoder models (e.g., a single model or only a few models) to report different numbers of subbands and / or different CSI payload sizes (e.g., encoder outputs) can help maintain relatively low hardware complexity at the UE. Thus, a single CSI generation model can be implemented using a relatively large number of N output features to accommodate the maximum number of expected features that might be used. However, in some cases, a relatively small number of subbands can be reported, and therefore, the actual number M of features transmitted may be less than N, resulting in potentially unnecessarily low compression rates.

[0440] In this scenario, embodiments of this disclosure can provide the UE and the base station with a shared understanding of which M features are transmitted and which NM features are punctured. This enables the base station to correctly apply the received features to the N corresponding input features of the CSI reconstruction model. The generative model at the UE and the reconstruction model at the base station can be trained such that the punctured features are set to fixed values ​​(e.g., zero) in the reconstruction model.

[0441] During inference operations, the base station can indicate to the UE (e.g., statically and / or dynamically, via RRC, DCI, MAC CE signaling, etc.) which set of features the UE can transmit. In one example implementation (which can be relatively efficient), an order can be established among the sets of features (e.g., M features), and the base station and / or the UE can indicate the number of features to be transmitted (e.g., N).

[0442] In some embodiments, once the model pair is activated, the base station can provide the UE with information about the CSI payload (e.g., via RRC, DCI, MAC CE signaling, etc.). Payload information can be provided, for example, using an index.

[0443] One type of CSI payload information may include, for example, the CSI payload size represented by the number of features. For example, if a generation mode (e.g., an encoder) generates a CSI output codeword with a length of N, the base station may indicate an actual CSI payload of size M, which may be less than, equal to, or greater than N. In such an embodiment, the UE may transmit only M features from the encoder output (e.g., after appropriate quantization).

[0444] Another type of CSI payload information may include a set of features to be transmitted. For example, an index may determine which subset of M features from a set of N features the UE can transmit. In some embodiments, the index may correspond to a bitmap of length N, where a first binary value (e.g., 1 or 0) may indicate that the feature has been transmitted, and a second binary value (e.g., 0 or 1) may indicate that the feature has not been transmitted. In some embodiments, the UE may transmit M features in ascending order of their indices. Additionally or alternatively, the set of features to be transmitted may be indicated using the starting index and length (number) of the features to be transmitted. This may simplify the implementation depending on the implementation details. CSI payload information may indicate, for example, a starting index 10 and a length 20 in a CSI codeword having N = 30 features. In this case, the CSI codeword may include a set of features (x_1, x_2, ..., x_10, x_11, ..., x_29, x_30), which the UE may transmit after quantizing the features to be transmitted (x_10, ..., x_29).

[0445] Another type of CSI payload information may include one or more of quantization methods, granularity, etc., which can instruct the UE on how to quantize the M transmitted CSI codeword features.

[0446] Codebook subset constraints

[0447] As described above, the UE can use a codebook to provide CSI feedback to the base station, indicating the PMI selected by the UE based on DL channel conditions measured in response to a reference signal. The base station can then use the PMI for beamforming in the DL channel. However, in some cases, the use of a particular PMI may cause beamforming interference (e.g., with neighboring cells). Therefore, some communication systems can implement a codebook subset constraint (CBSR) feature, which can constrain the UE to use specific portions of the codebook for channel estimation and / or feedback. For example, the base station can constrain the UE to report a specific PMI by setting specific base vector coefficients to zero (e.g., using a bitmap). Thus, the UE can select beams and / or coefficients from an unconstrained set, thereby ensuring that the UE can only report PMIs orthogonal to the constrained PMI (e.g., non-interfering). However, in systems that use artificial intelligence and / or machine learning to generate and / or compress CSI, the base vectors of the precoder for reporting may not exist. Therefore, the UE may not be able to determine the constrained PMI.

[0448] Constrained subspace information using AI / ML CSI feedback

[0449] In the wireless communication system according to this disclosure, channel-related constraint information can be provided to the UE. This enables the UE to ensure that the precoder or other CSI feedback determined using artificial intelligence and / or machine learning can result in beamforming that is orthogonal (e.g., non-interfering) to the constrained subspace used for the channel.

[0450] For example, in a system that uses linear transformations (such as angular delay transforms) to determine the precoder from the output of the decoder model, the base station can configure the UE using constrained subspaces and / or decoder output values ​​(e.g., coefficients that can be zero at specific indices). In embodiments that share the decoder model with the UE, the UE can generate a CSI, where the CSI causes the decoder output to be zero at a specified index. In embodiments that do not share the decoder model with the UE, information can be provided to the UE to ensure that the encoder input is of the same type and / or format as the decoder output. The UE can then apply subspace constraints (e.g., specific indices corresponding to zero coefficients) to the encoder input, which can result in the decoder output having zero coefficients at the corresponding indices.

[0451] As another example, in a system that uses the output of the decoder model as the CSI (e.g., the decoder model generates a pre-encoder matrix as output), the base station can utilize a constrained subspace to configure the UE. In embodiments where the decoder model is shared with the UE, the UE can ensure that the decoder output is orthogonal to the constrained subspace. In embodiments where the decoder model is not shared with the UE, the UE can ensure that the encoder input is orthogonal to the constrained subspace, which in turn ensures that the decoder output is sufficiently orthogonal to the constrained subspace.

[0452] As another example, the system can use information about the constrained subspace as input to the encoder model (e.g., explicit input). For instance, the constrained subspace can be used as input to the model to train the encoder model in a way that allows the output of the corresponding decoder model to reside in a subspace orthogonal to the constrained subspace.

[0453] For convenience, precoders, PMIs, coefficients, indices (multiple indices), matrices (multiple matrices), vectors, subcarriers, subbands, subspaces, subchannels, and / or any other elements that contribute directly or indirectly to orthogonal transmission can be called and / or characterized as orthogonal. Orthogonality can be determined, for example, in the time domain, spatial domain, frequency domain, and / or code domain.

[0454] Figure 24 An embodiment of a communication system is shown that uses channel constraint information and artificial intelligence and / or machine learning to provide channel information feedback in accordance with this disclosure. Figure 24 The system 2400 shown can be used to implement any apparatus, model, training scheme, etc. disclosed herein, or can be implemented using any apparatus, model, training scheme, etc. disclosed herein. System 2400 may include one or more elements (e.g., components, operations, etc.) that may be similar to... Figure 4 , Figure 7 Elements in the embodiments shown in the accompanying drawings and / or other figures, wherein similar elements may be indicated by reference numerals ending with and / or containing the same numbers, letters, etc.

[0455] exist Figure 24 In the system 2400 shown, a first wireless device (e.g., UE) 2401 can receive a reference signal 2417 from a second wireless device (e.g., base station) 2402 via channel 2415, for example, to enable the first wireless device 2401 to determine the channel conditions of channel 2415. The second wireless device 2402 can apply precoding information (e.g., precoding matrix) 2445 to some transmissions (such as PDCCH, PDSCH, etc.) and some reference signals (such as CSI-RS).

[0456] The first wireless device 2401 can perform channel measurements based on a reference signal 2417 and channel conditions to generate channel information (e.g., channel estimation, channel matrix, precoding information, etc.) 2405, which can be transmitted to the second wireless device 2402, for example, in the form of representation 2407, wherein representation 2407 can be generated using a generative model 2403. The first wireless device 2401 can transmit representation 2407 to the second wireless device 2402, for example, using another channel (e.g., an uplink channel), signals, etc. 2416.

[0457] The first wireless device 2401 can also receive constraint information 2461 related to channel 2415 from the second wireless device 2402. The constraint information 2461 enables the first wireless device 2401 to ensure that the channel information 2405 it generates conforms to one or more channel constraint requirements indicated by the constraint information 2461. For example, the constraint information 2461 may include information about a constrained subspace used for precoding information. The CSI determination logic 2452 at the first wireless device 2401 can implement one or more schemes, as described in more detail below, to ensure that the precoding information generated by the first wireless device 2401 is orthogonal to the constrained subspace. Therefore, CSI determination logic 2452 can use constraint information 2461 to interact with channel information 2405, generation model 2403, and / or representation 2407 to ensure that representation 2407 leads to reconstruction 2406 (e.g., precoding information 2445 such as a precoding matrix) after being applied to reconstruction model 2404 at second wireless device 2402, wherein second wireless device 2402 can use reconstruction 2406 to transmit a transmission orthogonal to the constrained subspace indicated by constraint information 2461 via channel 2415.

[0458] In some embodiments, the generating model 2403 and the reconstructing model 2404 can be trained (e.g., in an autoencoder configuration) to jointly compress and / or decompress channel information using any training scheme (including one or more joint training and / or deployment frameworks disclosed herein).

[0459] Figure 25 An example embodiment of a communication system according to this disclosure is shown, wherein a decoder model is shared with a UE configured with a constrained subspace. Figure 25 The system 2500 shown can be used, for example, to implement Figure 24 The system 2400 shown includes a system in which similar elements can be indicated by reference numerals ending with and / or containing the same numbers, letters, etc. For example, refer to... Figure 25For illustrative purposes, the first communication device can be implemented using UE 2501, the second communication device can be implemented using base station 2502, the constraint information can be implemented using constrained subspace 2561, the CSI determination logic can be implemented using precoding determination logic 2552, the channel information can be implemented using precoding information 2505, the generation model can be implemented using encoder model 2503 (which may also be called an encoder), and / or the reconstruction model can be implemented using decoder model 2504 (which may also be called a decoder).

[0460] In various embodiments, the constrained subspace 2561 may be represented, for example, by one or more of the following: a matrix, a column space of a matrix, a vector, an indication of one or more decoder outputs that can be set to a specific value (e.g., zero) (e.g., using one or more indices), and / or a combination thereof.

[0461] In various embodiments, the precoding information 2505 may be represented, for example, by one or more of the following: a precoder (e.g., a precoding matrix, vector, etc.), a set of coefficients from which the final (e.g., reconstructed) precoding information 2506 (e.g., based on angular delay transform and / or any other linear transform) can be computed, and / or a plurality of thereof and / or a combination thereof.

[0462] UE 2501 can be configured with a constrained subspace 2561, for example, via RRC, DCI, MAC CE signaling, etc.

[0463] The encoder model 2503 and decoder model 2504 can be trained (e.g., in an autoencoder configuration) to jointly compress and / or decompress precoded information 2505 using any training scheme (including one or more joint training and / or deployment frameworks disclosed herein).

[0464] The sharing logic 2544 at base station 2502 can share decoder model 2504 by transmitting information about model size, dimensions, weights, etc., to UE 2501, whereby UE 2501 can use the transmitted information to implement shared model 2504A. The sharing logic 2544 can, for example, use over-the-air transmission to transmit information about decoder model 2504.

[0465] Because UE 2501 can access a shared version 2504A of the decoder model 2504 used by base station 2502, UE 2501 can observe version 2506A of the reconstructed precoding information 2506 generated at base station 2502. In such embodiments, and depending on the implementation details, few restrictions (if any) can be placed on the inputs that UE 2501 can apply to encoder model 2503. For example, in some embodiments, as long as the inputs applied to encoder model 2503 make the reconstructed precoding information 2506A generated at shared decoder model 2504A orthogonal to the constrained subspace 2561, the reconstructed precoding information 2506 generated by decoder model 2504 at base station 2502 should also be orthogonal to the constrained subspace 2561.

[0466] Therefore, in some embodiments, the precoding determination logic 2552 can be freely implemented with any scheme for determining the precoding information 2505, and the result representation 2507 generated by the encoder model 2503 and sent to the base station 2502 (e.g., using uplink channels, signals, etc. 2516) should make the reconstructed precoding information 2546 at the base station 2502 orthogonal to the constrained subspace 2561.

[0467] Figure 26 An example embodiment of a communication system according to this disclosure is shown, wherein the decoder model is not shared with a UE configured with a constrained subspace. Figure 26 The system 2600 shown can be used, for example, to implement Figure 24 In the system 2400 shown, similar elements can be indicated by reference numerals ending with and / or containing the same numbers, letters, etc. For example, for illustrative purposes, UE 2601, base station 2602, constrained subspace 2661, precoding determination logic 2652, precoding information 2605, encoder model 2603, and / or decoder model 2604 can be indicated by reference numerals, letters, etc. Figure 25 The implementation is as shown in the embodiments. Additionally, encoder model 2603 and decoder model 2604 can be trained (e.g., in an autoencoder configuration) to jointly compress and / or decompress precoded information 2605 using any training scheme (including one or more joint training and / or deployment frameworks disclosed herein).

[0468] However, in Figure 26In the system 2600 shown, because the decoder model 2604 may not be shared with the UE 2601, the precoding determination logic 2652 may not be able to determine (e.g., directly using the shared model) what reconstructed precoding information 2606 can be generated by the decoder model 2604 at the base station 2602 in response to one or more specific inputs applied to the encoder model 2603.

[0469] Figure 26 The system 2600 shown can implement a subspace-constrained scheme in which one or more inputs 2663 of encoder model 2603 can be coordinated with one or more outputs 2664 of decoder model 2604. The scheme can then utilize joint training of encoder model 2603 and decoder model 2604 (e.g., in an autoencoder configuration) to ensure that the output 2664 of decoder model 2604 is orthogonal to the constrained subspace 2661.

[0470] For example, base station 2602 can send information about the configuration of the output 2664 of decoder 2604 to UE 2601. Precoding determination logic 2652 at UE 2601 can use this information to apply input 2663 to encoder 2603, where input 2663 has the same or similar configuration as the output 2664 of decoder 2604. If precoding determination logic 2652 selects an input 2663 to encoder 2603 that is orthogonal to the constrained subspace 2661, then joint training of encoder model 2603 and decoder model 2604 can ensure that the output 2664 of decoder 2604 is also orthogonal to the constrained subspace 2661. Depending on the implementation details, this can help overcome the lack of a shared decoder at UE 2601.

[0471] In some embodiments, if the output 2664 of the decoder 2604 is within an acceptable error range of the constrained subspace 2661, or within an acceptable error range of the input 2663 to the encoder 2603 which is orthogonal to the constrained subspace 2661, then the output 2664 of the decoder 2604 can be considered orthogonal to the constrained subspace 2661.

[0472] In some embodiments, information regarding the configuration of the output 2664 of the decoder 2604 can be sent to the UE 2601 via RRC, DCI, MAC CE signaling, etc.

[0473] In some example embodiments described below, the type of CSI can indicate the nature (e.g., one or more characteristics) of the CSI information used as input to a generative model (e.g., an encoder) and / or output from a reconstructive model (e.g., a decoder). For example, CSI can be implemented as a pre-encoder type of CSI, a channel matrix type of CSI, etc. Furthermore, the format can indicate the dimensions, order of elements, etc., in vectors and / or matrices. For example, if the input to the encoder and the output from the corresponding decoder have the same format, it can indicate that they are matrices of the same dimension, where elements (i,j) can represent the same entity, e.g., both are second features of the latent space, or both are second elements of the pre-encoder matrix of the same subband.

[0474] Report CSI in the angular delay domain or using a linear transformation.

[0475] In some embodiments, the decoder output can be represented as The target CSI precoder (which can be based on the decoder output) (to calculate) can be expressed as Furthermore, the constrained subspace can be represented as a matrix S. Precoder It can be orthogonal to the constrained subspace S (if) Therefore, a target CSI precoder orthogonal to the constrained subspace S is generated. This can involve determining one or more inputs to be applied to an encoder-decoder pair (e.g., an autoencoder), where the encoder-decoder pair can cause the decoder to generate an input that is orthogonal to the constrained subspace S (e.g., ) pre-encoder The output. Furthermore, as shown below, in some embodiments, in order to generate the pre-encoder... Make Decoder output The specific coefficient must be zero.

[0476] In some embodiments, the decoder output may be represented as one or more coefficients of a set of basis vectors (e.g., unlike a decoder that can directly generate the target CSI in the form of a precoder or channel matrix). The following analysis can be applied to such embodiments to generate outputs that result in orthogonality to the constrained subspace S (e.g., ) pre-encoder Output

[0477] For each layer, the spatial frequency domain CSI matrix, channel, or precoder matrix Angular delay transformation It can be defined as

[0478] H = F c H a F d (Equation 9)

[0479] Among them, F c It is N tx ×N tx DFT matrix, and F d It is N sb ×N sb DFT matrix. This can be replaced by F. M×N Let's write it, where N c It is M×N tx DFT submatrix, and F d By removing With appropriate zero rows and columns, and thus having M <N tx and N <N sb N sb ×NDFT submatrix. The following analysis can assume, without loss of generality, that...

[0480] As mentioned above, the decoder output can be... To represent. Therefore, It can represent the set of coefficients from which the final output CSI can be calculated (e.g., from which coefficients can be calculated). Calculate, for example, the precoder matrix The target CSI (channel matrix, etc.) can be calculated based on, for example, a linear transformation with specific basis vectors.

[0481] The CSI matrix can be a precoder matrix. It can be achieved, for example, via angular delay transformation. Obtain from the decoder output

[0482]

[0483] or any other linear transformation Among them, F c and F d This can be a basis matrix (e.g., a DFT matrix). Assuming we need... With matrix The given constrained subspaces are orthogonal, and the constraints can be applied directly to the decoder output, for example, by setting specific coefficients to specific values ​​(e.g., zero).

[0484] As mentioned above, precoder It can be orthogonal to the constrained subspace (if) Substitute equation 10 into... To obtain Multiply And pay attention to ST F c It can be a binary matrix that has more columns than rows (it can be called a fat matrix), and the output is... Specific coefficients can be set (e.g., forced) to zero to satisfy orthogonality constraints. Therefore, the base station can indicate to the UE one or more indices corresponding to the coefficients of the decoder's output, where one or more indices should be zero to ensure the pre-encoder... Orthogonality with the constrained subspace.

[0485] In some example embodiments where CSI can be reported in the angular delay domain or using a linear transformation, the decoder used by the base station may not be available to the UE. The above information regarding... Figure 26 An example embodiment of such an implementation is shown and / or described using a subspace constraint scheme. In such an embodiment, base station 2602 can configure UE 2601 using a constrained subspace 2661, wherein the constrained subspace 2661 is implemented using a set of constrained subspaces or precoder vectors (e.g., DFT beams) in the form of a matrix S. Since base station 2602 can know the output of decoder 2604... With pre-encoder matrix The relationship between the base station 2602 and the base station 2601 is such that the base station 2602 can also use the set of indexes corresponding to the coefficients that the output of the decoder should be zero to configure the UE 2601.

[0486] The zero condition can then be implemented at the output of decoder 2604 using one or more of the following techniques: (1) UE 2601 can be relied upon to report the CSI codeword that causes the decoder output to satisfy the zero condition (e.g., representation 2607 output by encoder 2603). Therefore, UE 2601 can select any input type available with its implementation, as long as it satisfies the decoder output requirements. (2) Information about the configuration of the output 2664 of decoder 2604 can be provided to the UE. The precoding determination logic 2652 at UE 2601 can use this information to ensure that it uses the input 2663 of encoder 2603, which has the same configuration (e.g., type, format, etc.) as the output 2664 of decoder 2604. Therefore, the output of decoder 2604 can have the same dimensions as the input 2663 of encoder 2603. UE 2601 can force a zero value at the encoder input at an index corresponding to the index of the coefficient for which the decoder output should be zero. Since the decoder output can be a reconstruction (e.g., reconstruction 2606) of the encoder input (e.g., pre-encoded information 2605) with a specific reconstruction error, it can generally be expected that the zero condition applied to the input of encoder 2603 will be replicated at the output of decoder 2604 up to the specific error.

[0487] In some other example embodiments where CSI can be reported in the angular delay domain or using a linear transformation, the decoder used by the base station can be shared with the UE. Such embodiments can use the above-mentioned... Figure 25 The subspace constraint scheme shown and / or described is used for implementation. In such an embodiment, base station 2502 can configure UE 2501 using constrained subspace 2561, wherein constrained subspace 2561 is implemented using a set of constrained subspaces or precoder vectors (e.g., DFT beams) in the form of matrix S. Since base station 2502 can know the output of decoder 2504... With pre-encoder matrix The relationship between the base station 2502 and the UE 2501 can be configured using the set of indexes corresponding to the coefficients whose output of the decoder should be zero.

[0488] Then, one or more of the following techniques can be used to achieve the zero condition at the output of decoder 2504. (1) UE 2501 can be relied upon to report the CSI codeword that causes the decoder output to satisfy the zero condition (e.g., representation 2507 output by encoder 2503). However, since UE 2501 has access to the shared decoder model 2504A, there can be no restriction on the configuration of the input applied to encoder 2503. Therefore, UE 2501 can choose any input available with its implementation, as long as the input causes the output generated at decoder 2504 (or reference decoder) to satisfy the zero-value condition at the desired index. (2) UE 2501 can train its encoder 2503 such that the decoder output has zero values ​​for the coefficients at the indicated index. In such a training framework, the encoder 2503 at UE 2501 can be trained using information about the constrained subspace (e.g., the subspace itself and / or one or more constrained coefficients, e.g., binary vectors with zeros at constrained indices of the decoder output and / or the constrained subspace).

[0489] In some alternative embodiments, it may not be expected that the UE ensures the orthogonality of the CSIs it can report to a base station with a constrained subspace. For example, the UE can use the encoded representation of the reported CSIs as a way to cause the decoder at the base station to generate a reconstructed output. The base station sends the codewords to report CSI in a normal manner. When the base station outputs from the decoder... Constructing the preencoder matrix In this case, the base station can force one or more coefficients of the decoder output to zero to enforce orthogonality with the constrained subspace. (The original decoder output at these indices may or may not already be zero.) Therefore, the UE is not expected to report CSI codewords that result in zero values ​​at the desired indices in the decoder output. In such an embodiment, the indices may or may not be indicated to the UE by the base station.

[0490] Reporting CSI in the spatial-frequency domain

[0491] CSI feedback (e.g., precoder matrix) can be reported in the space-frequency domain (e.g., directly). In some embodiments, describing the output encoder via a vector space may be impossible or inconvenient. In such embodiments, the following analysis can be used to ensure that the reported pre-encoder matrix is ​​orthogonal to the constrained subspace (e.g., () scheme.

[0492] The decoder output can be represented as Here, e can represent the error vector (e.g., its magnitude). The constrained subspace can be represented by a matrix. It is represented by a column space.

[0493] In some embodiments of the orthogonality scheme, the UE can derive one or more precoder vectors by projecting one or more precoder vectors onto a space orthogonal to the constrained subspace, such that w T S = 0. The following equation can be used to confirm that this property is retained in the output:

[0494]

[0495] In some embodiments, encoder-decoder pairs (e.g., autoencoders) can be expected to perform compression and / or decompression with a relatively small mean square error (MSE). Therefore, properties can also be preserved at the output, possibly with the following errors:

[0496]

[0497] It is possible to reduce (e.g., minimize) the MSE such that the right-hand side can be less than a given ∈:

[0498]

[0499] Assume the output precoder vector is normalized for power (e.g., or N tx When the reconstructed precoder vector With the basis vectors of the constrained subspace s iWhen i = 1, ..., Z are orthogonal, the inner product can be reduced (e.g., minimized).

[0500] Therefore, some embodiments can generate CSI feedback (e.g., pre-encoder matrix) orthogonal to the constrained subspace by ensuring orthogonality at the encoder input and allowing for some relatively small errors at the corresponding decoder output. ).

[0501] In this context, CSI can be reported in the space-frequency domain (e.g., as a pre-encoder matrix). In some example embodiments, the decoder used by the base station may not be available to the UE. This can be explained using the information above regarding... Figure 26 Such embodiments are implemented using subspace constraint schemes shown and / or described. In such embodiments, base station 2602 may configure UE 2601 using constrained subspace 2661, wherein constrained subspace 2661 is implemented using a set of constrained subspaces or precoder vectors (e.g., DFT beams) in the form of matrix S.

[0502] Then one or more of the following techniques can be used to ensure the pre-encoder matrix of the report. Orthogonality with the constrained subspace 2661. (1) It is possible to rely on UE 2601 to report what causes the output of decoder 2604 to satisfy the zero condition (e.g., The CSI codeword of the encoder 2603 (e.g., representation 2607 output by the encoder 2603). Therefore, the UE 2601 can select any input type available with its implementation, as long as it meets the decoder output requirements. (2) Information indicating that the input of the encoder 2603 and the output of the decoder 2604 are configured as pre-encoder vectors can be provided to the UE. The input of the encoder 2603 can be determined by the UE 2601, such that the encoder (e.g., pre-encoder matrix) is configured as a pre-encoder vector. The output of the decoder is orthogonal to the constrained subspace. In some embodiments, the UE 2601 can ensure that the residual orthogonality (or reconstruction) error at the decoder output is less than a specific value. For example, the UE can report CSI feedback such that for some ∈, the orthogonality error is...

[0503] CSI can be reported in the spatial-frequency domain (e.g., as a pre-encoder matrix). In some other example embodiments, the decoder used by the base station can be shared with the UE. Such embodiments can use the methods described above regarding... Figure 25The subspace constraint scheme shown and / or described is implemented. In such an embodiment, base station 2502 can configure UE 2501 using constrained subspace 2561, wherein constrained subspace 2561 is implemented using a set of constrained subspaces or precoder vectors (e.g., DFT beams) in the form of matrix S. Because UE 2501 can access shared decoder model 2504A, there may be no constraints on the configuration of the inputs applied to encoder 2503. Therefore, UE 2501 can select any input available with its implementation, as long as the input results in an output (e.g., precoder matrix) generated at decoder 2504. The orthogonality is found with the constrained subspace. In some embodiments, the CSI feedback can be reported using UE 2501, such that for some ∈, the orthogonality error is obtained.

[0504] The above about Figure 24 , Figure 25 and / or Figure 26 In the described embodiments, one or more models can be trained without using information about the constrained subspace. Therefore, an autoencoder can be trained without any assumptions about the constrained subspace. In such a framework, one or more models can be trained using examples of training data that can implicitly include different possibilities of subset constraints, but information about the constrained subspace can be omitted from the explicit training input. Nevertheless, constraints can be ensured during the inference phase based on the consistency between the training and inference phases regarding how the UE determines subset constraints for a given configuration of the encoder input.

[0505] Training and inference using constrained subspace information

[0506] In some embodiments, constrained subspace information can be used as input to the generative model (e.g., the encoder of an autoencoder pair of the model) (e.g., explicit training input and / or explicit inference input).

[0507] Figure 27 An embodiment of a system according to this disclosure that uses constrained subspace information as input to a model is shown. Figure 27 The system 2700 shown herein can be used to implement any of the embodiments disclosed herein that use artificial intelligence and / or machine learning to generate and / or compress CSI, or can be implemented using any of the embodiments disclosed herein that use artificial intelligence and / or machine learning to generate and / or compress CSI, including those shown and / or described in other figures, wherein similar elements may be indicated by reference numerals ending with and / or containing the same numbers, letters, etc.

[0508] refer to Figure 27 System 2700 may include a first node (node ​​A) having a generative model 2703 and a second node (node ​​B) having a reconstructed model 2704. System 2700 is illustrated in the inference phase of operation, where input data 2762 applied to the generative model (e.g., encoder) 2703 may include channel information (e.g., precoding information) 2705 and constraint information (e.g., constraint subspace information) 2761. However, in the training phase of operation, input data 2762 may include training data instead of channel information 2705.

[0509] Generative model 2703 can generate a representation 2707 of input data 2762, wherein reconstruction model 2704 can use representation 2707 to generate a reconstruction (e.g., precoding information) 2706 of channel information 2705. Because generative model 2703 and reconstruction model 2704 can be trained using constraint information 2761, the reconstruction 2706 of channel information can lie outside the constrained subspace indicated by constraint information 2761. In some embodiments, Figure 27 The scheme shown can be called and / or characterized as constraint perception.

[0510] Constraint information 2761 may include, for example, the constrained subspace itself, one or more constraints on the coefficients of the basis vector space (e.g., indices indicating coefficients that should be zero), etc. In some embodiments, channel information (e.g., precoding information) 2705 and / or reconstruction 2706 may be similar to those quantities implemented in a system without constraint-aware operations, but have the benefit of ensuring that reconstruction 2706 (e.g., the reported precoder) is outside the constrained subspace because constraint information 2761 is included in the input data 2762 used for training and / or inference.

[0511] In some embodiments, the generating model 2703 may include a quantizer to convert the representation 2707 into a quantized form (e.g., a bitstream) that can be transmitted via a communication channel. Similarly, in some embodiments, the reconstructing model 2704 may include a dequantizer, wherein the dequantizer can convert the quantized representation 2707 (e.g., a bitstream) into a form that can be used to generate training data 2712 for reconstruction.

[0512] The generative model 2703 and the reconstructed model 2704 can be obtained in any manner, including using any framework described herein. For example, using a joint training framework, the generative model 2703 and the reconstructed model 2704 can be trained as a pair at node B, where node B can send the reconstructed model 2704 to node A. Other embodiments may use a training framework with a reference model, a training framework with the latest shared values, or any other framework and / or technique to obtain and / or train the generative model 2703 and the reconstructed model 2704.

[0513] In some embodiments, the loss can be computed at the output of the reconstructed model 2704 using labels that may include constraint information. Therefore, it can be ensured that the labels (which may be referred to as true labels) reside in the constrained subspace.

[0514] In some embodiments, when the output of the reconstruction model 2704 is not orthogonal to the constrained subspace determined by the constraint information 2761, the loss function can utilize a penalty that can be introduced into the loss function. In such an embodiment, Figure 27 The scheme shown can report acceptable or optimal CSI codewords (e.g., encoder outputs and / or PMI indices) within a given reconstruction model 2704 (e.g., outside a constrained subspace).

[0515] Additionally or alternatively, the loss function can be implemented as a regular loss function that ensures the tightness of the output of the reconstructed model 2704 with the true generated model 2703. Since tightness (e.g., in the sense of MSE) reduces the norm of the orthogonality error, it can ensure that the output lies within an allowed subspace (e.g., it can satisfy one or more orthogonality conditions (e.g., ...). ).

[0516] In some embodiments, during the inference phase, the base station may utilize constraint information (e.g., a constrained subspace) 2761 to configure the UE, wherein the constraint information 2761 may be used (e.g., directly) as input to the generative model 2703. Then, using the constraint information 2761 for training and / or inference operations may facilitate reconstruction (e.g., pre-encoded information) 2706 within a permitted subspace.

[0517] In embodiments where the base station shares reconstruction model 2704 with the UE, the base station can configure the UE using a set of constrained subspaces or precoder vectors (e.g., DFT beams) in the form of a matrix S, and the type of output CSI can be indicated as a precoder vector. In such embodiments, the UE can be relied upon to report CSI codewords (e.g., representation 2707 output by generation model 2703) such that the output of reconstruction model 2704 is orthogonal to the constrained subspace. In some such embodiments, the UE 2701 can ensure that the residual orthogonality (or reconstruction) error at the decoder output is less than a specific value. For example, the UE can report CSI feedback such that for ∈ , the orthogonality error is less than a certain value.

[0518] In embodiments where the reconstructed model 2704 is not shared with the UE, the base station may configure the UE using a set of constrained subspaces or precoder vectors (e.g., DFT beams) in the form of a matrix S, and may then use one or more of the following techniques to ensure that the reported CSI (e.g., precoder) is orthogonal to the constrained subspace: (1) The input configuration of the generation model 2703 and the output configuration of the reconstructed model 2704 may be indicated (or assumed by the UE in the absence of an indication received via the network) as the precoder. The UE may report CSI codewords (e.g., representation 2707 output by the generation model 2703) such that the precoder output by the reconstructed model 2704 is orthogonal to the constrained subspace. (2) The UE may assume that the input configuration of the generation model 2703 and the output configuration of the reconstructed model 2704 are the same or similar (e.g., precoder). The UE may determine the input of the generation model 2703 to be orthogonal to the constrained subspace, and therefore the output of the reconstructed model 2704 will also be orthogonal. (3) The UE may train a reference model to be used as the generation model 2703. The UE can then report the CSI codeword so that the corresponding reference used to reconstruct model 2704 also satisfies the orthogonality condition.

[0519] Data collection for AIML models used in CSI reporting

[0520] In some wireless communication systems that use artificial intelligence and / or machine learning models, the UE may collect data for various purposes (e.g., use cases), such as model training, model inference, model performance monitoring, etc. Examples of data types collected for one or more of these purposes may include channel information (e.g., a channel matrix that can be input to an encoder at the UE), CSI feedback (e.g., the output of an encoder model at the UE, which can be sent to a base station and used as input to a decoder model at the base station), target CSI (e.g., a pre-encoder or channel matrix that can be output from a decoder model at the base station), one or more gradients (e.g., which can be used for backpropagation with type 2 training), etc.

[0521] Data collection at the UE may involve using UE processing resources to perform downlink channel measurements (e.g., channel estimation) based on CSI-RS signals transmitted by the base station. In one method for collecting data, the base station may configure the UE using a CSI reporting configuration, wherein the UE can typically use the CSI reporting configuration to measure, calculate, and report the CSI to the base station.

[0522] However, under different circumstances, different devices (e.g., UE or base station) may need to collect different types of data. Therefore, when a CSI reporting configuration is configured, the UE may calculate one or more types of CSI it may not need to collect and / or report to the base station, unnecessarily burdening the UE's processing resources. Another potential problem is that the UE may not be able to and / or be configured to support one or more CSI-RS configurations for data collection. Furthermore, the base station may not be able to determine the UE's capabilities and / or configuration for data collection. An additional potential problem is that the UE may waste resources preparing uplink channels to send CSIs it does not need to report. Therefore, the UE may need to operate differently when collecting data compared to when reporting CSIs according to a normal CSI reporting configuration. Another potential problem is that the base station may use a CSI reporting configuration to configure the UE where the CSI reporting configuration may not generate one or more types of data that the UE may need to collect for its purposes (such as training a local model, using the local model to perform inference, etc.). Furthermore, the UE may not be able to determine whether or when the base station expects it to collect data.

[0523] To overcome these potential problems, the systems and methods according to this disclosure enable a UE to establish a data collection session with a base station to request the collection of one or more specific types of data and / or indicate the ability to support one or more data collection configurations.

[0524] Figure 28 An embodiment of a communication system according to this disclosure is shown, wherein the communication system has data collection capabilities for one or more models that can be used to provide channel information feedback. System 2800 may include one or more elements (e.g., components, operations, etc.) that may be similar to Figure 4 , Figure 7 Elements in the embodiments shown in the accompanying drawings and / or other figures, wherein similar elements may be indicated by reference numerals ending with and / or containing the same numbers, letters, etc.

[0525] exist Figure 28In the system 2800 shown, a first wireless device (e.g., UE) 2801 can receive a reference signal 2817 from a second wireless device (e.g., base station) 2802 via channel 2815, for example, to enable the first wireless device 2801 to determine the channel conditions of channel 2815. The second wireless device 2802 can apply precoding information (e.g., precoding matrix) 2845 to some transmissions (such as PDCCH, PDSCH, etc.) and some reference signals (such as CSI-RS).

[0526] The first wireless device 2801 can perform channel measurements based on a reference signal 2817 and channel conditions to generate channel information (e.g., channel estimation, channel matrix, precoding information, etc.) 2805. The first wireless device 2801 can transmit the channel information to the second wireless device 2802, for example, in the form of a representation 2807 that can be generated using a generative model 2803. The first wireless device 2801 can transmit the representation 2807 to the second wireless device 2802, for example, using another channel (e.g., an uplink channel), signals, etc. 2816. The second wireless device 2802 can include a reconstruction model 2804 that it can use to generate a reconstruction 2806 of the channel information 2805.

[0527] In some embodiments, the generating model 2803 and the reconstructing model 2804 may be trained (e.g., in an autoencoder configuration) to jointly compress and / or decompress channel information 2805 using any training scheme (including one or more joint training and / or deployment frameworks disclosed herein).

[0528] The first wireless device 2801 and the second wireless device 2802 may each include data collection logic 2866 and 2867, wherein data collection logic 2866 and 2867 may individually and / or jointly implement some or all of the data collection functions disclosed herein. For example, data collection logic 2866 and / or 2867 may enable the first wireless device 2801 to establish a data collection session between the first wireless device and the second wireless device 2802. For example, the first wireless device 2801 may send a request to the second wireless device 2802 to begin a data collection session. The request may be accompanied by information indicating one or more types of data to be collected and / or one or more reference signal configurations to be used for data collection. The second wireless device 2802 may send an instruction to the first wireless device 2801 to begin a data collection session in response to a request from the first wireless device 2801 or at the initiation of the second wireless device 2802.

[0529] As another example, data collection logic 2866 and / or 2867 may enable the first wireless device 2801 to declare one or more data collection capabilities 2868 to the second wireless device 2802. In some embodiments, the first wireless device 2801 may declare the capability to support one or more data collection configurations, wherein the one or more data collection configurations may include information about the purpose of data collection, one or more types of data to be collected, and / or information about corresponding supported reference signal configurations that can be used to collect data.

[0530] Figure 29 An example embodiment of a communication system according to this disclosure is shown, wherein the communication system has data collection capabilities for one or more models that can be used to provide channel information feedback. Figure 29 The system 2900 shown may include components that can be used with... Figure 4 , Figure 7 , Figure 28 One or more elements (e.g., components, operations, etc.) that are similar to those in the embodiments shown in the accompanying drawings and / or other drawings, wherein similar elements may be indicated by reference numerals that end with and / or contain the same numbers, letters, etc. Figure 29 The system 2900 shown can be used, for example, to implement a data collection scheme for CSI compression in the spatial-frequency (SF) domain and / or time-space-frequency (TSF) domain as described below, as well as other data collection schemes according to this disclosure.

[0531] For the purpose of illustrating some example embodiments of data collection for CSI compression, Figure 29 The system 2900 shown is illustrated with some example implementation details. Specifically, the first communication device can be implemented using UE 2901, the second communication device can be implemented using base station 2902, the reference signal can be implemented using CSI-RS signal 2917, the channel information can be implemented using channel matrix 2905, the generation model can be implemented using encoder model 2903 (which may also be called an encoder), the representation of the channel information can be implemented using CSI feedback 2907, the reconstruction model can be implemented using decoder model 2904 (which may also be called a decoder), and the reconstruction of the channel information can be implemented using target CSI 2906 (e.g., a pre-encoder matrix, which can be used to determine one or more coefficients of the pre-encoder matrix through a combination of basis vectors, etc.). However, Figure 29 The principles of the system 2900 shown are not limited to these or any other implementation details.

[0532] Some or all of the data collection functions disclosed herein may be implemented, at least in part, individually and / or collectively, by data collection logics 2966 and 2967 located at UE 2901 and base station 2902, respectively. In some embodiments, UE 2901 may determine target CSI 2906 (e.g., the output of decoder model 2904 at base station 2902) by using, for example, a shared copy of a decoder model, a locally trained decoder model, a reference decoder model, etc.

[0533] In some embodiments, UE 2901 may send request 2969 to base station 2902 to initiate a data collection session. Request 2969 may be accompanied by information indicating one or more types of data to be collected, and / or one or more CSI-RS configurations to be used for data collection. Request 2969 may be sent to base station 2902, for example, using a dedicated scheduling request (SR).

[0534] Base station 2902 may initiate a data collection session by sending an indication (e.g., command, trigger, signal, activation, etc.) 2970 to UE 2901 in response to a request 2969 from UE 2901 or at the initiation of base station 2902. For example, base station 2902 may initiate a data collection session if UE 2901 has previously declared its ability to support one or more specific data collection operations (e.g., the ability to collect one or more specific types of data using one or more specific CSI-RS configurations). Base station 2902 may send the indication 2970 to UE 2901 to begin the data collection session using, for example, RRC, MAC CE, DCI, etc.

[0535] In some embodiments, base station 2902 may configure UE 2901 for data collection using a CSI reporting configuration that can be instructed to be used for data collection. When such a CSI reporting configuration is configured, UE 2901 may avoid reporting some or all of the CSIs to base station 2902. Depending on the implementation details, UE 2901 may also avoid performing one or more calculations for CSI feedback 2907, target CSI 2906 (e.g., precoded information), etc. The CSI reporting configuration may be instructed to be used for data collection, for example, using information elements (IEs) (e.g., data collection use cases). Additionally or alternatively, base station 2902 may configure UE 2901 to use one or more CSI-RS resource sets for data collection purposes (e.g., when CSI reporting by UE 2901 is not expected). In such a configuration, one or more CSI-RS resource sets may be coordinated (e.g., matched) with the purpose of the collected data, the type of data to be collected, etc.

[0536] In some embodiments, UE 2901 may declare the capability 2968 to support one or more data collection configurations. For example, a data collection configuration (which may also be referred to as a feature set) may include information about the purpose of data collection (e.g., CSI compression (for training, inference, and / or monitoring), CSI prediction, etc.), one or more types of data to be collected (e.g., channel matrix 2905, CSI feedback 2907, and / or target CSI 2906), and / or information about the corresponding supported CSI-RS configurations that can be used to collect data (e.g., CSI-RS subbands, number of CSI-RS ports, supported SCS, time-domain window information for CSI compression, etc.).

[0537] In some embodiments, various identifiers (IDs) can be used to indicate the type of information to be collected, the use case of the collected data, etc. For example, each type of collected data (e.g., channel matrix 2905, CSI feedback 2907, target CSI 2906, etc.) can be indicated by a corresponding data type ID number. Therefore, the data collection configuration (e.g., sent as part of a declaration of data collection capability 2968) can include a set of numbers indicating the types of data that UE 2901 can collect. As another example, each use case for the collected data (e.g., purpose, scenario, etc.) can be indicated by a corresponding use case ID. Examples of use cases that may have corresponding use case IDs include CSI compressio...

Claims

1. An apparatus for model interoperability in machine learning-based channel information reporting, comprising: The receiver is configured to use the channel to receive a reference signal; The transmitter is configured to transmit a representation associated with the channel; and The processing circuit is configured as follows: The channel information is determined based on the reference signal; The model is used to generate the representation based on the channel information; and The receiver or the transmitter is used to transmit model information to specify the model.

2. The apparatus according to claim 1, wherein, The model information includes the model's identifier.

3. The apparatus according to claim 1, wherein, The model information includes the structural information of the model.

4. The apparatus according to claim 1, wherein, The model information includes information about the type of input to the model.

5. The apparatus according to claim 1, wherein, The model information includes information about the format of the input to the model.

6. The apparatus according to claim 1, wherein, The model information includes mapping information used to map the channel information to the input of the model.

7. The apparatus according to claim 6, wherein, The mapping information includes: Information used for the first sub-band; and Information used for the second subband.

8. The apparatus according to claim 6, wherein, The mapping information includes: Information used for the first time slot; and Information used for the second time slot.

9. An apparatus for model interoperability in machine learning-based channel information reporting, comprising: The receiver is configured to use the channel to receive a reference signal; The transmitter is configured to transmit a representation associated with the channel; and The processing circuit is configured as follows: The channel information is determined based on the reference signal; The model is used to generate the representation based on the channel information; and The receiver or the transmitter is used to transmit capability information of the device related to the model.

10. The apparatus according to claim 9, wherein, The capability information includes information about the models supported by the device.

11. The apparatus according to claim 9, wherein, The capability information includes information about the model structure supported by the device.

12. The apparatus according to claim 9, wherein, The capability information includes information about the inference operations used for the model.

13. The apparatus according to claim 9, wherein, The processing circuitry is also configured to transmit the capability information using a channel information report configuration.

14. The apparatus according to claim 9, wherein, The capability information includes information about the type of collaboration supported by the device.

15. The apparatus according to claim 9, wherein, The capability information includes information about the device's ability to modify the model.

16. The apparatus according to claim 15, wherein, The capability information includes information about the amount of time associated with modifying the model.

17. An apparatus for model interoperability of machine learning-based channel information reporting, comprising: The receiver is configured to use the channel to receive a reference signal; The transmitter is configured to transmit a representation associated with the channel; and The processing circuit is configured as follows: Channel information is determined based on the reference signal; The model is used to generate the representation based on the channel information; and The receiver or the transmitter is used to transmit dataset information to specify the dataset used for the model.

18. The apparatus according to claim 17, wherein, The dataset information includes the types of information in the dataset.

19. The apparatus according to claim 17, wherein, The dataset information includes the format of the information in the dataset.

20. The apparatus according to claim 17, wherein, The dataset information includes mapping information for the dataset.