Communication method and communication apparatus
By acquiring the dataset and model from the network side on the terminal side and combining them with performance metrics for model training, the problem of model training mismatch between terminal devices and network devices is solved, thus improving the performance of CSI feedback.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-04-02
AI Technical Summary
In existing technologies, the AI model training dataset or model mismatch between terminal devices and network devices leads to CSI feedback performance not meeting requirements or even becoming unusable.
The terminal side acquires the dataset and model sent by the network side, and combines performance metrics to train the model, including methods such as supervised learning, self-supervised learning, and knowledge distillation, to optimize model parameters and improve CSI feedback performance.
This enables the rapid and efficient training of compliant models on the terminal side, improving the accuracy and efficiency of CSI feedback.
Smart Images

Figure CN2025122482_02042026_PF_FP_ABST
Abstract
Description
Communication method and communication apparatus
[0001] This application claims priority to the Chinese patent application No. 202411403188.4, filed on September 30, 2024, with the State Intellectual Property Office of China, and the Chinese patent application No. 202411403188.4 has the title of “Communication method and communication apparatus”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of wireless communication, and in particular to a communication method and a communication apparatus. BACKGROUND
[0003] At present, there are some solutions that use an artificial intelligence (AI) model to compress and reconstruct (or recover) channel state information (CSI). The sender (such as a terminal device) of a CSI report can compress the CSI by using the AI model and send the obtained CSI report. The receiver (such as a network device) of the CSI report can reconstruct the channel measurement result based on the received CSI report by using the AI model.
[0004] In order to support the docking and model development of the two-end AI model (such as the sender and the receiver of the CSI report), a dataset or a model needs to be provided for the sender and / or the receiver to train the model. However, the training of the receiver and the dataset or model of the sender may not match, resulting in that the obtained model does not meet the requirements or even cannot be used. SUMMARY
[0005] The present application provides a communication method and a communication apparatus, which are beneficial to train a model that meets the requirements and improve the feedback performance of the CSI.
[0006] In a first aspect, a communication method is provided. The method can be applied to a terminal device, such as the terminal device itself, or a component deployed in the terminal device, such as a circuit or a chip (such as a modem chip, also known as a baseband chip, or a system on chip (SoC) chip or a system in package (SIP) chip containing a modem core, etc.) inside the terminal device, etc. Alternatively, the method can be applied to a logic module or software, etc. capable of implementing all or part of the functions of the terminal device. Alternatively, the method can be performed by a first apparatus, which can be a terminal device, or a component in the terminal device, such as a circuit or a chip (such as a modem chip, also known as a baseband chip, or a SoC chip or a SIP chip containing a modem core, etc.) inside the terminal device, or a logic module or software, etc. capable of implementing part or all of the functions of the terminal device. The present application does not limit this.
[0007] Exemplarily, the method comprises: obtaining a data set and / or a model, the data set and / or the model being used for model training; obtaining a performance indicator corresponding to the obtained data set and / or the obtained model; and the performance indicator corresponding to an inference task of a first model, the first model being obtained based on the model training of the data set and / or the model, and the inference task being an inference of compressing CSI.
[0008] The data set can be sent by a network side to a terminal side. The terminal side obtaining the data set and / or the model can comprise: a terminal device of the terminal side receiving the data set and / or the model from a network device of the network side; or a host or a cloud server of an OTT system of the terminal side receiving the data set and / or the model from an intelligent network element of the network side; or the terminal device of the terminal side obtaining the data set and / or the model received from the intelligent network element from a host or a cloud server of an OTT system; or the host or the cloud server of the OTT system of the terminal side obtaining the data set and / or the model received from the network device from the terminal device.
[0009] Optionally, the method further comprises: receiving the data set and / or the model.
[0010] Based on the above technical solution, the terminal side can obtain the data set and / or model for model training, and obtain the performance indicator, which corresponds not only to the data set and / or model obtained by the terminal side, but also to the inference task of the first model. In other words, after obtaining the data set and / or model and the performance indicator, the terminal side can determine how to perform model training and the performance indicator that should be met. Therefore, model training can be performed accordingly. In this way, the terminal side can obtain a required model through model training, thereby improving the performance of CSI feedback.
[0011] In combination with the first aspect, in some implementations of the first aspect, the obtaining the performance indicator comprises: receiving first information, the first information being used to indicate the performance indicator.
[0012] That is, the network side can send the first information used to indicate the performance indicator to the terminal side. Since the network side can send the data set and / or model for model training to the terminal side, the network side can send the performance indicator corresponding to the data set and / or model to the terminal side, so that the terminal device performs corresponding model training according to the received data set and / or model and the performance indicator determined based on the first information, to obtain a required model.
[0013] In combination with the first aspect, in some implementations of the first aspect, the obtaining the data set and / or model comprises: obtaining a first data set, the first data set comprising one or more of the following data: first target CSI, first CSI feedback information, or first reconstructed CSI; wherein the first CSI feedback information is obtained by compressing and quantizing the first target CSI, and the first reconstructed CSI is obtained by dequantizing and decompressing the first CSI feedback information.
[0014] In a possible design, the first data set comprises: the first target CSI and the first CSI feedback information.
[0015] Optionally, the first target CSI and the first CSI feedback information are used for model training of the first model.
[0016] The terminal side can take the first target CSI as the input of the first model, and take the compressed CSI obtained by dequantizing the first CSI feedback information as the label, to perform model training of the first model. Exemplarily, the training can be regarded as supervised learning for the first CSI feedback information (or the compressed CSI obtained by dequantizing the first CSI feedback information).
[0017] Based on the above design of the first data set, the terminal side can perform model training of the first model, which has fewer steps and a simpler training process, and thus is conducive to the terminal side to obtain the first model through training more quickly.
[0018] Optionally, the performance indicator corresponds to an inference task of the first model, and the performance indicator is a performance requirement to be met by the inference task of the first model.
[0019] Optionally, the performance indicator includes one or more ranges to be met: a mean square error (MSE) between information obtained by compressing the first target CSI using the first model and information obtained by dequantizing the first CSI feedback information; a normalized mean square error (NMSE) between the information obtained by compressing the first target CSI using the first model and the information obtained by dequantizing the first CSI feedback information; a mean absolute error (MAE) between the information obtained by compressing the first target CSI using the first model and the information obtained by dequantizing the first CSI feedback information; or a weighted sum of two or more of the MSE, the NMSE, and the MAE.
[0020] The difference between the output of the first model and the label is characterized by the NMSE, the MSE, the MAE, or a weighted sum of two or more of the above. The terminal side can optimize the parameters of the first model based on the difference to reduce the difference, thereby obtaining a first model with higher accuracy for model inference.
[0021] Optionally, the first target CSI and the first CSI feedback information are used for model training of a second model, the second model is used for model training of a third model to obtain the first model, and the third model includes the first model.
[0022] The third model can be used to compress a target CSI to obtain compressed CSI, and can be used to decompress information obtained by dequantizing CSI feedback information to obtain reconstructed CSI. The compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information. In other words, the third model can include the first model and the second model described above.
[0023] Based on the above design of the first data set, the terminal side can perform model training of the first model and the second model, and the training process can monitor the performance from the entire process of CSI compression and reconstruction. Moreover, the model training can be trained using terminal side data, which can improve the performance of CSI feedback.
[0024] Optionally, the performance indicator corresponds to an inference task of the first model, and the performance indicator includes a performance requirement to be met by the third model in compressing and reconstructing the second target CSI, where the first model is configured to perform inference of compressing CSI.
[0025] Optionally, the performance indicator includes one or more of the following ranges to be met: MSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information by the second model and the first target CSI; NMSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information by the second model and the first target CSI; MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information by the second model and the first target CSI; MSE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; NMSE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; MAE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; GCS between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; SGCS between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; or a weighted sum of multiple ones of the following: MSE, NMSE, MAE, MSE, NMSE, MAE, GCS, or SGCS between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information by the second model and the first target CSI, and the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI.
[0026] The difference between the output of the second model and the input of the first model is characterized based on the NMSE, the MSE, the MAE, or a weighted sum of two or more of the above. The terminal side can optimize the parameters of the first model based on the difference to reduce the difference between the two, so as to obtain a first model with higher accuracy for model inference.
[0027] In another possible design, the first data set includes: first CSI feedback information and first reconstructed CSI; the first CSI feedback information and the first reconstructed CSI are used for model training of the second model, the second model is used for model training of the third model, and the third model includes the first model.
[0028] The second model can be used to decompress the information obtained by dequantizing the first CSI feedback information. The third model can be used to compress the target CSI to obtain compressed CSI, and can be used to decompress the information obtained by dequantizing the CSI feedback information to obtain reconstructed CSI. The compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information. In other words, the third model can include the first model and the second model described above.
[0029] The terminal side can first train the second model. The terminal side can use the information obtained by dequantizing the first CSI feedback information (i.e., the compressed CSI with quantization loss) as the input of the second model, and use the first reconstructed CSI as the label to perform model training of the second model. Thereafter, the terminal side can fix the parameters of the second model, and train the third model to obtain the first model.
[0030] The first possible way for the terminal side to perform model training of the third model is that the terminal side can use the second target CSI of the terminal side as the input of the first model, use the output of the first model as the input of the second model, and perform model training to make the reconstructed CSI output by the second model tend to be close to the second target CSI input to the first model, thereby obtaining the first model. Illustratively, this training can be regarded as self-supervised learning for the second target CSI.
[0031] The second possible way for the terminal side to perform model training of the third model is that the terminal side can also use the first reconstructed CSI as the input of the first model, use the output of the first model as the input of the second model, and perform model training to make the reconstructed CSI output by the second model tend to be close to the first reconstructed CSI input to the first model, thereby obtaining the first model. Illustratively, this training can be regarded as self-supervised learning for the first reconstructed CSI.
[0032] Based on the above design of the first data set, the terminal side can perform model training of the first model and the second model. The training process can monitor the performance from the whole process of CSI compression and reconstruction. Moreover, the model training can be trained by using the data of the terminal side, and can improve the performance of CSI feedback.
[0033] Optionally, the performance indicator corresponds to an inference task of the first model, and includes a performance requirement that should be met by compression and reconstruction of the second target CSI by the third model, where the first model is used to perform inference of compressing CSI.
[0034] Corresponding to the first manner of performing model training of the third model at the terminal side, optionally, the performance indicators include one or more of the following ranges to be met: MSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; NMSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; MSE between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; NMSE between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; MAE between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; GCS between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; SGCS between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; or a weighted sum of multiple ones of MSE, NMSE, MAE, GCS, or SGCS between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI.
[0035] Corresponding to the second manner of performing model training of the third model at the terminal side, optionally, the performance indicators include one or more of the following ranges to be met: MSE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; NMSE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; MAE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; GCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; SGCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; or a weighted sum of multiple ones of MSE, NMSE, MAE, GCS, or SGCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI.
[0036] The difference between the output of the second model and the input of the first model is characterized based on the NMSE, the MSE, the MAE, the GCS or the SGCS, or a weighted sum of two or more of the above. The terminal side can optimize the parameters of the first model based on the difference to reduce the difference, thereby obtaining a first model with higher accuracy for model inference. In addition, more performance indicators are used to monitor model training, and more comprehensive performance can be monitored.
[0037] In yet another possible design, the first data set includes: the first target CSI, the first CSI feedback information, and the first reconstructed CSI; wherein the first target CSI and the first reconstructed CSI are used for training of a third model to obtain the first model, and the third model includes the first model.
[0038] The third model can be used to compress the first target CSI to obtain compressed CSI, and can be used to decompress the information obtained by dequantizing the CSI feedback information to obtain reconstructed CSI; and the compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information. In other words, the third model can include the first model and the second model described above.
[0039] The terminal side can perform model training of the third model by taking the first target CSI as the input of the third model and taking the first reconstructed CSI as the label. Exemplarily, the training can be supervised learning for the first reconstructed CSI.
[0040] Based on the above design of the first data set, the terminal side can perform model training of the first model and the second model. The training process can monitor the performance from the entire process of CSI compression and reconstruction, thereby monitoring more comprehensive performance.
[0041] Optionally, the performance indicator corresponds to an inference task of the first model, and includes: the performance indicator is a performance requirement to be met by compression and reconstruction of the first target CSI by the third model, wherein the first model is used to perform inference of compressing CSI.
[0042] Optionally, the performance indicator includes one or more ranges to be satisfied: the MSE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; the NMSE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; the MAE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; the GCS between the reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; the SGCS between the reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; or a weighted sum of a plurality of the MSE, the NMSE, the MAE, the GCS, or the SGCS.
[0043] The difference between the output and the input of the third model is characterized based on the NMSE, the MSE, the MAE, the GCS, the SGCS, or a weighted sum of two or more of the above. The terminal side can optimize the parameters of the first model based on the difference to reduce the difference, so that a first model with higher accuracy can be obtained for model inference. In addition, more performance indicators are used to monitor model training, and more comprehensive performance can be monitored.
[0044] In combination with the first aspect, in some implementations of the first aspect, the obtaining the dataset and / or the model comprises: obtaining a fourth model, the fourth model being used for model training of the first model.
[0045] The fourth model can be used to compress the second target CSI.
[0046] In a possible design, the fourth model is a teacher model of the first model, and is used to perform model training of a student model of the first model to obtain the first model.
[0047] That is, the model training is performed in a manner of knowledge distillation. The student model of the first model obtained through training is the first model.
[0048] The model training is performed in a manner of knowledge distillation, which can reduce the memory and computational complexity of the model, and also enables the student model to be trained to have similar performance as the teacher model.
[0049] In another possible design, the fourth model is used to generate a second dataset, the second dataset being used for model training of the first model, and the second dataset including compressed CSI obtained by compressing the second target CSI.
[0050] That is, the second data set is generated by the fourth model sent by the network side, and the model training of the first model on the terminal side is performed according to the second data set.
[0051] The terminal side generates the second data set by the model from the network side, and the model training of the first model is performed according to the second data set. The steps of the model training are less, the training process is simple, and it is beneficial for the terminal side to obtain the first model through training faster.
[0052] Optionally, the performance indicator corresponds to the inference task of the first model, and the performance indicator includes a performance requirement that should be met by the inference task of the first model.
[0053] Optionally, the performance indicator includes one or more ranges that should be met: the MSE between the compressed CSI obtained by compressing the second target CSI by the first model and the compressed CSI obtained by compressing the second target CSI by the fourth model; the NMSE between the compressed CSI obtained by compressing the second target CSI by the first model and the compressed CSI obtained by compressing the second target CSI by the fourth model; the MAE between the compressed CSI obtained by compressing the second target CSI by the first model and the compressed CSI obtained by compressing the second target CSI by the fourth model; or the weighted sum of two or more of the MSE, the NMSE, and the MAE between the compressed CSI obtained by compressing the second target CSI by the first model and the compressed CSI obtained by compressing the second target CSI by the fourth model.
[0054] The difference between the output of the first model and the label is characterized by the NMSE, the MSE, the MAE, or the weighted sum of two or more of the above, and the parameters of the first model are optimized based on the difference to reduce the difference, so that a first model with high accuracy can be obtained for model inference.
[0055] In combination with the first aspect, in some implementations of the first aspect, the obtaining of the data set and / or the model includes: obtaining a fourth model, the fourth model being used for model training of a third model to obtain the first model, the third model including the first model.
[0056] The fourth model is used for compressing the second target CSI; the third model is used for compressing the second target CSI to obtain compressed CSI; and the third model can be used for decompressing the information obtained by dequantizing the CSI feedback information to obtain reconstructed CSI. The compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information. In other words, the third model can include the first model and the second model described above.
[0057] The terminal side can first train the first model according to the fourth model, and then use the trained first model in the third model to perform model training of the third model to obtain the first model. Exemplarily, this process can be regarded as supervised learning of the third model.
[0058] Based on the above scheme, the terminal side not only performs separate model training on the first model and monitors the process of CSI compression by the first model, but also applies the first model to the third model to monitor the whole process of CSI compression and reconstruction, which can further improve the performance of the trained first model.
[0059] Optionally, the performance indicator corresponds to an inference task of the first model, and includes: a requirement that should be met by the inference task of the first model, and / or a performance requirement that should be met by compression and reconstruction of the second target CSI by the third model, wherein the first model is used to perform inference of CSI compression.
[0060] Optionally, the performance indicator includes one or more ranges that should be met: MSE between information obtained by compressing the second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; NMSE between information obtained by compressing the second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; MAE between information obtained by compressing the second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; GCS between reconstructed CSI obtained by processing the second target CSI by the third model and the second target CSI; SGCS between reconstructed CSI obtained by processing the second target CSI by the third model and the second target CSI; or a weighted sum of multiple ones of MSE, NMSE, MAE, GCS and SGCS between information obtained by compressing the second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model, and between reconstructed CSI obtained by processing the second target CSI by the third model and the second target CSI.
[0061] The information obtained by compressing the second target CSI by the first model can be compressed CSI, and the information obtained by compressing the second target CSI by the fourth model can also be compressed CSI.
[0062] Different performance indicators are indicated to monitor different processes of possible training of the terminal side to ensure the performance of CSI feedback.
[0063] With reference to the first aspect, in some implementations of the first aspect, the obtaining the dataset or the model comprises: obtaining a second model, the second model being used for model training of a third model to obtain the first model, the third model comprising the first model.
[0064] The second model can be used for decompression of the compressed CSI. The third model is used for compression of the target CSI to obtain the compressed CSI, and can be used for decompression of the information obtained by dequantization of the CSI feedback information to obtain the reconstructed CSI. The compressed CSI corresponds to the information obtained by dequantization of the CSI feedback information. In other words, the third model can comprise the first model and the second model described above.
[0065] The terminal side can perform model training of the third model according to the second model. As described above, the third model comprises the first model and the second model, and thus the first model can be obtained by model training of the third model. Exemplarily, this process can be regarded as self-supervised learning of the third model.
[0066] The terminal side can apply the second model sent by the network side to the third model to perform model training of the first model. This training process has fewer steps and is relatively simple, which is conducive to the terminal side to obtain the first model more quickly. Moreover, this model training can be performed using terminal side data, which can improve the feedback performance of the CSI.
[0067] Optionally, the performance indicator corresponds to an inference task of the first model, and the performance indicator comprises a performance requirement that should be met by compression and reconstruction of the second target CSI by the third model, wherein the first model is used to perform inference of compression of the CSI.
[0068] Optionally, the performance indicator comprises one or more of the following ranges that should be met: GCS between the reconstructed CSI obtained by processing of the second target CSI by the third model and the second target CSI; SGCS between the reconstructed CSI obtained by processing of the second target CSI by the third model and the second target CSI; or a weighted sum of the GCS and the SGCS between the reconstructed CSI obtained by processing of the second target CSI by the third model and the second target CSI.
[0069] The SGCS, the GCS, or the weighted sum of the two, is used to represent the difference between the output and the input of the third model, and the parameters of the third model are optimized based on the difference to reduce the difference between the two, so that a first model with high precision can be obtained for model inference.
[0070] With reference to the first aspect, in some implementations of the first aspect, the method further includes: performing the inference task of the first model in a case where the performance indicator can be satisfied.
[0071] The inference task of the first model is performed in a case where the performance indicator is satisfied, so that the model performance on the terminal side is guaranteed, which is conducive to the business requirements.
[0072] With reference to the first aspect, in some implementations of the first aspect, the method further includes: in a case where the performance indicator cannot be satisfied, sending second information, the second information being used for one or more of the following: indicating that the performance indicator cannot be satisfied; requesting to change the performance indicator; requesting to change the data set; requesting to change the model; or requesting to close the model training.
[0073] In this way, the network side can respond in a case where the performance indicator cannot be satisfied, such as closing the model training, or changing the performance indicator, or changing the model, or changing the data set, etc. Since the response of the network side can be made after learning the status of the terminal side, the model, data set, performance indicator, etc. issued can better adapt to the capability of the terminal side, thereby facilitating to improve the performance of the double-end model.
[0074] With reference to the first aspect, in some possible implementations of the first aspect, before receiving the data set and / or the model, the method further includes: sending capability information, the capability information being used to indicate the computing capability supported by the terminal side.
[0075] The terminal side reports the capability information to the network side, which facilitates the network side to learn the computing capability of the terminal side, and then can issue a data set and / or a model that can adapt to the capability of the terminal side, and a corresponding performance indicator, thereby helping the terminal side to obtain a model that meets the requirements through model training, and being conducive to obtaining improved feedback performance of CSI.
[0076] In a second aspect, a communication method is provided, which can be applied to a network device, such as the network device itself, or a component deployed in the network device, such as a circuit or a chip (such as a modem chip, also known as a baseband chip, or a SoC chip or a SIP chip containing a modem core, etc.) inside the network device; or, it can also be applied to a logic module or software, etc. that can realize all or part of the functions of the network device. Or it can also be said that the communication method can be executed by a first apparatus, which can be a network device; or a component in the network device, such as a circuit or a chip (such as a modem chip, also known as a baseband chip) inside the network device, or a SoC chip or a SIP chip containing a modem core, etc.; or a logic module or software that can realize part or all of the functions of the network side, etc. The present application does not limit this.
[0077] Exemplarily, the method comprises: sending a data set and / or a model, the data set and / or the model being used for model training; and sending first information, the first information being used for indicating a performance index, the performance index corresponding to an inference task of a first model, the first model being obtained based on model training of the data set and / or the model, the inference task being an inference of compressing CSI.
[0078] With reference to the second aspect, in some implementations of the second aspect, the sending the data set and / or the model comprises: sending a first data set, the first data set comprising one or more of the following data: first target CSI, first CSI feedback information, or first reconstructed CSI; wherein the first CSI feedback information is obtained by compressing and quantizing the first target CSI, and the first reconstructed CSI is obtained by dequantizing and decompressing the first CSI feedback information.
[0079] In a possible design, the first data set comprises: the first target CSI and the first CSI feedback information, the first target CSI and the first CSI feedback information being used for model training of the first model.
[0080] Optionally, the first target CSI and the first CSI feedback information are used for model training of the first model.
[0081] Optionally, the performance index corresponding to the inference task of the first model comprises: the performance index being a performance requirement to be met by performing the inference task of the first model.
[0082] Optionally, the performance index comprises one or more of the following ranges to be met: a MSE between information obtained by compressing the first target CSI by using the first model and information obtained by dequantizing the first CSI feedback information; a NMSE between information obtained by compressing the first target CSI by using the first model and information obtained by dequantizing the first CSI feedback information; a MAE between information obtained by compressing the first target CSI by using the first model and information obtained by dequantizing the first CSI feedback information; or a weighted sum of a plurality of the following: the MSE, the NMSE, or the MAE between information obtained by compressing the first target CSI by using the first model and information obtained by dequantizing the first CSI feedback information.
[0083] Optionally, the first target CSI and the first CSI feedback information are used for model training of a second model, the second model being used for model training of a third model to obtain the first model, the third model comprising the first model.
[0084] Optionally, the performance indicator includes one or more of the following ranges to be satisfied: MSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information using the second model and the first target CSI; NMSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information using the second model and the first target CSI; MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information using the second model and the first target CSI; MSE between the reconstructed CSI obtained by processing the first target CSI using the third model and the first target CSI; NMSE between the reconstructed CSI obtained by processing the first target CSI using the third model and the first target CSI; MAE between the reconstructed CSI obtained by processing the first target CSI using the third model and the first target CSI; GCS between the reconstructed CSI obtained by processing the first target CSI using the third model and the first target CSI; SGCS between the reconstructed CSI obtained by processing the first target CSI using the third model and the first target CSI; or a weighted sum of a plurality of the following: MSE, NMSE, MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information using the second model and the first target CSI; MSE, NMSE, MAE, GCS, or SGCS between the reconstructed CSI obtained by processing the first target CSI using the third model and the first target CSI.
[0085] In another possible design, the first data set includes: first CSI feedback information and first reconstructed CSI; the first CSI feedback information and the first reconstructed CSI are used for model training of a second model, which is used for model training of the first model.
[0086] Optionally, the performance indicator includes one or more of the following ranges to be satisfied: MSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; NMSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; MSE between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; NMSE between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; MAE between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; GCS between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; SGCS between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI; or a weighted sum of multiple ones of the MSE, the NMSE, the MAE, the GCS, or the SGCS between the reconstructed CSI obtained by processing the second target CSI through the first model and the second model and the second target CSI.
[0087] Optionally, the performance indicator includes one or more of the following ranges to be satisfied: MSE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; NMSE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; MAE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; GCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; SGCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; or a weighted sum of multiple ones of the MSE, the NMSE, the MAE, the GCS, or the SGCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI.
[0088] In yet another possible design, the first data set includes: the first target CSI, the first CSI feedback information, and the first reconstructed CSI; wherein the first target CSI and the first reconstructed CSI are used for training of the third model to obtain the first model, and the third model includes the first model.
[0089] The third model is configured to compress the first target CSI to obtain compressed CSI, and to decompress information obtained by dequantizing the CSI feedback information to obtain reconstructed CSI. The compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information.
[0090] Optionally, the performance indicator includes one or more ranges to be satisfied: MSE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; NMSE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; MAE between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; GCS between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; SGCS between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI; and weighted sum of multiple ones of the MSE, the NMSE, the MAE, the GCS, or the SGCS between the reconstructed CSI obtained by processing the first target CSI by the third model and the first target CSI.
[0091] In combination with the second aspect, in some implementations of the second aspect, the sending the data set and / or the model comprises: sending a fourth model, the fourth model being configured for model training of the first model.
[0092] The fourth model is configured to compress the second CSI.
[0093] In one possible design, the fourth model is a teacher model of the first model, and is configured to perform model training of a student model of the first model to obtain the first model.
[0094] In another possible design, the fourth model is configured to generate a second data set, the second data set being configured for model training of the first model, and the second data set including compressed CSI obtained by compressing the second target CSI.
[0095] Optionally, the performance indicator includes one or more of the following ranges to be satisfied: MSE between the compressed CSI obtained by compressing the second target CSI using the first model and the compressed CSI obtained by compressing the second target CSI using the fourth model; NMSE between the compressed CSI obtained by compressing the second target CSI using the first model and the compressed CSI obtained by compressing the second target CSI using the fourth model; MAE between the compressed CSI obtained by compressing the second target CSI using the first model and the compressed CSI obtained by compressing the second target CSI using the fourth model; GCS between the reconstructed CSI obtained by processing the second target CSI using the third model and the second target CSI; SGCS between the reconstructed CSI obtained by processing the second target CSI using the third model and the second target CSI; or a weighted sum of a plurality of the following: MSE, NMSE, MAE, GCS, and SGCS between the reconstructed CSI obtained by processing the second target CSI using the third model and the second target CSI, and information obtained by compressing the second target CSI using the first model and information obtained by compressing the second target CSI using the fourth model.
[0096] With reference to the second aspect, in some implementations of the second aspect, the sending the dataset and / or the model comprises: sending a second model, the second model being used for model training of a third model to obtain the first model, the third model comprising the first model. Optionally, the performance indicator includes one or more of the following ranges to be satisfied: GCS between the reconstructed CSI obtained by processing the second target CSI using the third model and the second target CSI; SGCS between the reconstructed CSI obtained by processing the second target CSI using the third model and the second target CSI; or a weighted sum of the GCS and the SGCS between the reconstructed CSI obtained by processing the second target CSI using the third model and the second target CSI.
[0097] With reference to the second aspect, in some implementations of the second aspect, the method further comprises: receiving second information, the second information being used for one or more of the following: indicating that the performance indicator cannot be satisfied; requesting to replace the performance indicator; requesting to replace the dataset; requesting to replace the model; or requesting to close the model training.
[0098] With reference to the second aspect, in some possible implementations of the second aspect, before the sending the dataset and / or the model, the method further comprises: receiving capability information, the capability information being used to indicate a computing capability supported at a terminal side.
[0099] In the implementations of the second aspect, the more detailed descriptions of the first model, the second model, the third model and the fourth model, and the descriptions of the correspondence between the performance indicators and the inference tasks of the first model can be referred to the descriptions in the first aspect, and will not be repeated.
[0100] It should be understood that the method provided by the second aspect corresponds to the first aspect, and the descriptions and technical effects of the implementations of the second aspect can be referred to the descriptions of the first aspect, and will not be repeated.
[0101] In the third aspect, an apparatus is provided. The apparatus can include function modules corresponding to the method described in any possible implementation of the first aspect, or function modules corresponding to the method described in any possible implementation of the second aspect. The modules can be hardware circuits, software, or a combination of hardware circuits and software.
[0102] In one design, the apparatus can include a processing module and a communication module. The communication module can be configured to perform the sending and receiving actions performed by the terminal side in the method described in the first aspect, and the processing module can be configured to perform the processing-related actions performed by the terminal side in the method described in the first aspect.
[0103] In one design, the apparatus can be a terminal device, or an apparatus, module, circuit, or chip configured to be deployed in a terminal device, or an apparatus that can be used in conjunction with a terminal device, such as an OTT host or a cloud server.
[0104] In one design, the apparatus can include a processing module and a communication module. The communication module can be configured to perform the sending and receiving actions performed by the network side in the method described in the second aspect, and the processing module can be configured to perform the processing-related actions performed by the network side in the method described in the second aspect.
[0105] In one design, the apparatus can be a network device, or an apparatus, module, circuit, or chip configured to be deployed in a network device, or an apparatus that can be used in conjunction with a network device, such as a smart network element deployed with a radio access network (RAN) intelligent controller (RIC).
[0106] In the fourth aspect, an apparatus is provided, including a processor and a storage medium, the storage medium storing instructions that, when executed by the processor, cause the method in the first aspect or any possible implementation of the first aspect to be implemented, or cause the method in the second aspect or any possible implementation of the second aspect to be implemented.
[0107] In a fifth aspect, there is provided an apparatus comprising processing circuitry for processing data and / or information to cause the method in the first aspect or any possible implementation of the first aspect, or to cause the method in the second aspect or any possible implementation of the second aspect, to be performed.
[0108] The processing circuitry can include one or more processors, or all or a portion of the circuitry of the one or more processors for controlling or processing functions.
[0109] Optionally, the apparatus can further include a memory for storing a program or instructions, and the processor is configured to execute the program or instructions to cause the method in the first aspect or any possible implementation of the first aspect, or to cause the method in the second aspect or any possible implementation of the second aspect, to be performed.
[0110] Optionally, the apparatus can further include the transceiver circuitry, or an input / output interface.
[0111] In a sixth aspect, there is provided a chip comprising processing circuitry for executing a program or instructions to cause the method in the first aspect or any possible implementation of the first aspect, or to cause the method in the second aspect or any possible implementation of the second aspect, to be performed.
[0112] Optionally, the chip can further include a memory for storing the program or instructions.
[0113] Optionally, the chip can further include the transceiver circuitry, or an input / output interface.
[0114] In a seventh aspect, there is provided a computer readable storage medium comprising instructions, which when executed by a processor, cause the method in the first aspect or any possible implementation of the first aspect, or the method in the second aspect or any possible implementation of the second aspect, to be performed.
[0115] In an eighth aspect, there is provided a computer program product comprising computer program code or instructions, which when executed by a processor, cause the method in the first aspect and any possible implementation of the first aspect, or the method in the second aspect or any possible implementation of the second aspect, to be performed.
[0116] In a ninth aspect, a communication system is provided, which comprises the apparatus of the first aspect and any possible implementation manner of the first aspect, or comprises the apparatus of the second aspect and any possible implementation manner of the second aspect.
[0117] It should be understood that the third aspect to the ninth aspect of the present application correspond to the technical solutions of the first aspect to the second aspect of the present application, and the beneficial effects achieved by each aspect and the corresponding possible implementation manners are similar, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0118] FIG. 1 is a schematic diagram of a communication system suitable for the communication method according to the embodiments of the present application;
[0119] FIG. 2 is a schematic diagram of another communication system suitable for the communication method according to the embodiments of the present application;
[0120] FIG. 3 is a schematic diagram of a possible application framework in the communication system;
[0121] FIG. 4 is a schematic diagram of another possible application framework in the communication system;
[0122] FIG. 5 is a schematic diagram of CSI feedback using an auto-encoder (AE) model according to the embodiments of the present application;
[0123] FIG. 6 shows an example of a neuron structure;
[0124] FIG. 7 is a schematic diagram of a deep neural network (DNN);
[0125] FIG. 8 is a schematic diagram of data set interfacing between the network side and the terminal side according to the embodiments of the present application;
[0126] FIG. 9 is a schematic diagram of model interfacing between the network side and the terminal side according to the embodiments of the present application;
[0127] FIG. 10 is a schematic flowchart of the communication method according to the embodiments of the present application;
[0128] FIG. 11A and FIG. 11B are schematic diagrams of a processing flow at the terminal side according to the embodiments of the present application;
[0129] FIG. 12A and FIG. 12B are schematic diagrams of another processing flow at the terminal side according to the embodiments of the present application;
[0130] FIG. 13 is a schematic diagram of another processing flow at the terminal side according to the embodiments of the present application;
[0131] FIG. 14 is a schematic diagram of another processing flow at the terminal side according to the embodiments of the present application;
[0132] FIG. 15 is a schematic diagram of another process flow at the terminal side according to an embodiment of the present application;
[0133] FIG. 16 is a schematic diagram of another process flow at the terminal side according to an embodiment of the present application;
[0134] FIG. 17 is a schematic diagram of another process flow at the terminal side according to an embodiment of the present application;
[0135] FIG. 18 is a schematic diagram of another process flow at the terminal side according to an embodiment of the present application;
[0136] FIG. 19 is a schematic diagram of another process flow at the terminal side according to an embodiment of the present application;
[0137] FIG. 20 is another schematic flowchart of a communication method according to an embodiment of the present application;
[0138] FIG. 21 and FIG. 22 are schematic block diagrams of a communication apparatus according to an embodiment of the present application. DETAILED DESCRIPTION
[0139] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0140] For the convenience of understanding the embodiments of the present application, the following points are first explained:
[0141] First, in the present application, the terminal side can also be referred to as the UE side, including: a terminal device, a component (such as a circuit or a chip inside the terminal device, etc.) deployed in the terminal device, a device (such as a host of an OTT system or a cloud server) deployed outside the terminal device or a component (such as a circuit or a chip inside the device, etc.) deployed in the device outside the terminal device. The network side (NW side) includes: a network device in communication with the terminal device, a component (such as a circuit or a chip inside the network device with near-real-time RAN intelligent control function, etc.) deployed in the network device, a device (such as an intelligent network element, such as an intelligent network element with near-real-time RAN intelligent control function) deployed outside the network device or a component (such as a circuit or a chip inside the intelligent network element, etc.) deployed in the intelligent network element. Among them, the network device can include: an access network device, a core network device or an operation administration and maintenance (OAM).
[0142] Secondly, in the present application, indication includes direct indication (also referred to as explicit indication) and indirect indication (also referred to as implicit indication). Directly indicating information A means including the information A; indirectly indicating information A can mean indicating the information A by the correspondence between the information A and information B and directly indicating the information B; or indicating the information A by a preset rule that can be used to determine A according to B and directly indicating the information B. The correspondence between the information A and the information B and the preset rule can be predefined, pre-stored, pre-burned, or pre-configured.
[0143] Thirdly, in the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it, but does not rule out the case that the associated objects before and after it represent an "and" relationship. The specific meaning can be understood in combination with the context. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent: a, b, c; a and b; a and c; b and c; or a and b and c. Where a, b, and c can be single or multiple.
[0144] Fourthly, in the present application, the use of prefixes such as "first", "second", and the like is only for the convenience of distinguishing and describing different things belonging to the same name category, and does not constrain the order, size, or quantity of the things. For example, "first target CSI" and "second target CSI" are only different target CSIs, and do not limit the quantity, size relationship, or priority relationship of the target CSIs; for another example, "first model" and "second model" are only different models, and do not limit the quantity, size relationship, or priority relationship of the models; for another example, "first information" and "second information" are only different indication information, and do not limit the quantity, time sequence, size relationship, or priority relationship of the information.
[0145] Fifth, in the present application, "sending" and "receiving" represent the direction of signal transmission. For example, "sending information to XX" can be understood as the destination of the information is XX, which can include direct transmission through the air interface, and also includes indirect transmission through the air interface by other units or modules. "Receiving information from YY" can be understood as the source of the information is YY, which can include direct reception from YY through the air interface, and also includes indirect reception from YY through the air interface from other units or modules. "Sending" can also be understood as the "output" of the chip interface, and "receiving" can also be understood as the "input" of the chip interface. In other words, sending and receiving can be between devices, such as between terminal devices and computing nodes, or within devices, such as between components, modules, chips, software modules or hardware modules within devices through buses, wires or interfaces.
[0146] Sixth, in the embodiments of the present application, "when", "if" and "when" all refer to the device will make corresponding processing under certain objective circumstances, not limited to time, and also does not require the device to have a judgment action when it is implemented, nor does it mean that there are other limitations.
[0147] Seventh, in the present application, "example", "exemplarily", "for example" or "such as" are used to represent as an example, illustration or explanation. Any embodiment or design scheme described as "example", "exemplarily", "for example" or "such as" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of "example", "exemplarily", "for example" or "such as" is intended to present the relevant concept in a specific way.
[0148] The technical solutions provided in the present application can be applied to various communication systems, for example, a 5th generation (5G) or new radio (NR) system, a long term evolution (LTE) system, an LTE frequency division duplex (FDD) system, an LTE time division duplex (TDD) system, a wireless local area network (WLAN) system, a satellite communication system, a future communication system, or a fusion system of multiple systems, and the like. The technical solutions provided in the present application can also be applied to device to device (D2D) communication, vehicle-to-everything (V2X) communication, machine to machine (M2M) communication, machine type communication (MTC), and an internet of things (IoT) communication system or other communication systems.
[0149] A network element in a communication system can send a signal to another network element or receive a signal from another network element. The signal can include information, signaling, data, and the like. The network element can also be replaced by an entity, a network entity, a device, a communication device, a communication module, a node, a communication node, and the like. For example, the communication system can include at least one terminal device and at least one network device. The network device can send a downlink signal to the terminal device, and / or the terminal device can send an uplink signal to the network device. It can be understood that the terminal device in the present disclosure can be replaced by a first network element, and the network device can be replaced by a second network element, both of which perform the corresponding communication method in the present disclosure.
[0150] FIG. 1 is a schematic diagram of a communication system suitable for a communication method according to an embodiment of the present application. As shown in FIG. 1, the communication system 100A can include at least one access network device, such as the access network device 110 shown in FIG. 1, and can also include at least one terminal device, such as the terminal device 120 and the terminal device 130 shown in FIG. 1. The access network device 110 and the terminal devices (such as the terminal device 120 and the terminal device 130) can communicate through a wireless link. The communication devices in the communication system, for example, the access network device 110 and the terminal device 120, can communicate through multi-antenna technology.
[0151] In a wireless communication network, such as a mobile communication network, the services supported by the network are increasingly diverse, and thus the requirements to be met are increasingly diverse. For example, the network needs to be able to support ultra-high rates, ultra-low latencies, and / or ultra-large connections. This feature makes network planning, network configuration, and / or resource scheduling increasingly complex. In addition, as the functions of the network become increasingly powerful, such as supporting increasingly high frequency spectrums, supporting high-order multiple input multiple output (MIMO) technology, supporting beam forming (BF), supporting beam management, and other new technologies, network energy saving has become a hot research topic. These new requirements, new scenarios, and new features bring unprecedented challenges to network planning, operation and maintenance, and efficient operation. To meet this challenge, artificial intelligence technology can be introduced into the wireless communication network, thereby realizing network intelligence. In order to support AI technology in the wireless network, an AI node can also be introduced into the network.
[0152] FIG. 2 is a schematic diagram of another communication system suitable for the communication method of the embodiments of the present application. Compared with the communication system 100A shown in FIG. 1, the communication system 100B shown in FIG. 2 further includes an AI network element 140. The AI network element 140 is configured to perform AI-related operations, such as constructing a training data set or training an AI model. The AI network element can also be referred to simply as an intelligent network element.
[0153] In a possible implementation, the access network device 110 can send data related to the training of the AI model to the AI network element 140, and the AI network element 140 constructs a training data set and trains an AI model. For example, the data related to the training of the AI model can include data reported by the terminal device. The AI network element 140 can send the result of the AI model-related operation to the access network device 110 and forward it to the terminal device through the access network device 110. For example, the result of the AI model-related operation can include at least one of the following: a trained AI model, an evaluation result or a test result of the model, and the like. Illustratively, part of the trained AI model can be deployed on the access network device 110, and the other part can be deployed on the terminal device 120 and / or the terminal device 130. Alternatively, the trained AI model can be deployed on the access network device 110. Or, the trained AI model can be deployed on the terminal device 120 and / or the terminal device 130.
[0154] It should be understood that FIG. 2 is only used as an example to illustrate that the AI network element 140 is directly connected with the access network device 110, and in other scenarios, the AI network element 140 can also be connected with a terminal device. Alternatively, the AI network element 140 can be connected with both the access network device 110 and the terminal device. Alternatively, the AI network element 140 can also be connected with the access network device 110 through a third-party network element. The connection relationship between the AI network element and other network elements is not limited in the embodiments of the present application. For example, the AI network element 140 can also be arranged as a module in the access network device and / or the terminal device, for example, in the access network device 110 or the terminal device shown in FIG. 1.
[0155] It should be noted that FIG. 1 and FIG. 2 are only simplified schematic diagrams for example and understanding, for example, the communication system can also include other devices, such as wireless relay devices and / or wireless backhaul devices, which are not shown in FIG. 1 and FIG. 2. In actual application, the communication system can include multiple access network devices, and can also include multiple terminal devices. The number of access network devices and terminal devices included in the communication system is not limited in the embodiments of the present application.
[0156] In the embodiments of the present application, the terminal device can also be referred to as a UE, an access terminal, a user unit, a user station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent or a user equipment.
[0157] The terminal device can be a device providing voice / data, for example, a handheld device with wireless connection function, a vehicle-mounted device, etc. At present, some examples of terminals are: mobile phone, tablet computer, notebook computer, palm computer, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), handheld device with wireless communication function, computing device or other processing device connected to a wireless modem, wearable device, terminal device in a 5G network, or terminal device in a future evolved public land mobile network (PLMN), etc. The embodiments of the present application are not limited thereto.
[0158] By way of example and not limitation, in the embodiments of the present application, the terminal device can also be a wearable device. The wearable device can also be referred to as a wearable smart device, which is a general term for devices that are designed and developed by applying wearable technology to daily wear, such as glasses, gloves, watches, clothing, and shoes. The wearable device is a portable device that is directly worn on the body or integrated into the user's clothes or accessories. The wearable device is not only a hardware device, but also a device that realizes powerful functions through software support and data interaction and cloud interaction. The general wearable smart device includes devices with full functions, large size, and the ability to realize complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, and devices that focus on a certain application function and need to be used in cooperation with other devices, such as smart phones, such as various smart wristbands and smart jewelry for monitoring vital signs.
[0159] In the embodiments of the present application, the apparatus for implementing the function of the terminal device can be a terminal device, or an apparatus capable of supporting the terminal device to implement the function, for example, a chip system, which can be installed in the terminal device or used in matching with the terminal device. In the embodiments of the present application, the chip system can be composed of a chip, or can include the chip and other discrete devices. In the embodiments of the present application, only the apparatus for implementing the function of the terminal device is taken as an example for illustration, and the scheme of the embodiments of the present application is not limited thereto.
[0160] The network device in the embodiments of the present application can be a device for communicating with the terminal device, which can include an access network device, for example, the access network device can be a base station. The access network device in the embodiments of the present application can refer to a RAN node (or device) for accessing the terminal device to the wireless network. The base station can be broadly covered by the following various names, or be replaced by the following names, such as: Node B (NodeB), evolved Node B (eNB), next generation Node B (gNB), relay station, access point, transmitting and receiving point (TRP), transmitting point (TP), primary station, secondary station, motor slide retainer (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), radio unit (RU), positioning node, etc. The base station can be a macro base station, a micro base station, a relay node, a donor node or the like, or a combination thereof. The base station can also refer to a communication module, modem or chip for being arranged in the foregoing devices or apparatuses. The base station can also be a mobile switching center, a device assuming the function of a base station in D2D, V2X, M2M communication, a device assuming the function of a base station in a future communication system, etc. The base station can support networks of the same or different access technologies. Optionally, the RAN node can also be a server, a wearable device, a vehicle or a vehicle-mounted device, etc. For example, the access network device in the V2X technology can be a road side unit (RSU). The embodiments of the present application do not limit the specific technology and specific device form adopted by the access network device.
[0161] A base station can be fixed, or mobile. For example, a helicopter or unmanned aerial vehicle can be configured to act as a mobile base station, one or more cells can move according to the location of the mobile base station. In other examples, a helicopter or unmanned aerial vehicle can be configured to act as a device that communicates with another base station.
[0162] In some deployments, the access network device mentioned in the embodiments of the present application can be a device including a CU, or a DU, or a device including a CU and a DU, or a control plane CU node (central unit-control plane (CU-CP)) and a user plane CU node (central unit-user plane (CU-UP)) and a DU node. For example, the access network device can include a gNB-CU-CP, a gNB-CU-UP and a gNB-DU.
[0163] In some deployments, a plurality of RAN nodes cooperate to assist a terminal to implement wireless access, and different RAN nodes respectively implement part of the functions of a base station. For example, the RAN node can be a CU, a DU, a CU-CP, a CU-UP, or an RU, etc. The CU and the DU can be separately arranged, or can also be included in the same network element, for example, in a BBU. The RU can be included in a radio frequency device or a radio frequency unit, for example, included in an RRU, an AAU or an RRH.
[0164] The RAN node can support one or more types of fronthaul interface, different fronthaul interfaces respectively corresponding to DUs and RUs having different functions. If the fronthaul interface between the DU and the RU is a common public radio interface (CPRI), the DU is configured to implement one or more of baseband functions, and the RU is configured to implement one or more of radio frequency functions. If the fronthaul interface between the DU and the RU is another interface, relative to the CPRI, one or more of the partial baseband functions of the downlink and / or uplink, such as, for the downlink, one or more of precoding, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / add cyclic prefix (CP), are moved from the DU to the RU for implementation, and for the uplink, one or more of digital beamforming (BF), or fast Fourier transform (FFT) / remove cyclic prefix (CP) are moved from the DU to the RU for implementation. In a possible implementation, the interface can be an enhanced common public radio interface (eCPRI). Under the eCPRI architecture, the splitting manner between the DU and the RU is different, corresponding to different categories (Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, F.
[0165] Taking eCPRI Cat A as an example, for downlink transmission, with layer mapping as the cut, the DU is configured to implement layer mapping and one or more functions (i.e., one or more of encoding, rate matching, scrambling, modulation, layer mapping) before layer mapping, and other functions (e.g., one or more of resource element (RE) mapping, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / adding CP) after layer mapping are implemented in the RU. For uplink transmission, with de-RE mapping as the cut, the DU is configured to implement de-mapping and one or more functions (i.e., one or more of decoding, de-rate matching, de-scrambling, de-modulation, inverse discrete Fourier transform (IDFT), channel equalization, de-RE mapping) before de-mapping, and other functions (e.g., one or more of digital BF or FFT / CP removal) after de-mapping are implemented in the RU. It can be understood that the function description of the DU and the RU corresponding to various types of eCPRI can refer to the eCPRI protocol, and will not be described here.
[0166] In a possible design, the processing unit in the BBU for implementing baseband functions is referred to as a base band high (BBH) unit, and the processing unit in the RRU / AAU / RRH for implementing baseband functions is referred to as a base band low (BBL) unit.
[0167] In different systems, the CU (or CU-CP and CU-UP), DU or RU can also have different names, but those skilled in the art can understand their meanings. For example, in an open RAN (ORAN) architecture, the CU can also be referred to as an open-CU (O-CU), the DU can also be referred to as an open-DU (O-DU), the CU-CP can also be referred to as an open-CU-CP (O-CU-CP), the CU-UP can also be referred to as an open-CU-UP (O-CU-UP), and the RU can also be referred to as an open-RU (O-RU). Any of the CU (or CU-CP, CU-UP), DU and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.
[0168] In the embodiments of the present application, the apparatus for implementing the function of the network device can be a network device, or can be an apparatus capable of supporting the network device to implement the function, such as a chip system, a hardware circuit, a software module, or a hardware circuit plus a software module. The apparatus can be installed in the network device or used in combination with the network device. In the embodiments of the present application, only the apparatus for implementing the function of the network device is taken as an example for description, and the scheme of the embodiments of the present application is not limited in this way.
[0169] The network device and / or the terminal device can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; can also be deployed on water surface; and can also be deployed on airplanes, balloons and satellites in the air. The scenarios in which the network device and the terminal device are located are not limited in the embodiments of the present application. In addition, the terminal device and the network device can be hardware devices, or can be software functions running on special hardware, software functions running on general hardware, such as virtualized functions instantiated on a platform (for example, a cloud platform), or entities including special or general hardware devices and software functions. The specific forms of the terminal device and the network device are not limited in the present application.
[0170] Optionally, the AI node can be deployed in one or more of the following positions in the communication system: an access network device, a terminal device, or a network element of a core network, and the like. Optionally, the AI node can also be deployed separately, for example, in a position other than any of the above devices, such as a host or a cloud server of an over the top (OTT) system. The AI node can communicate with other devices in the communication system, which can be one or more of the following: an access network device, a terminal device, or a network element of a core network, and the like.
[0171] It can be understood that the number of AI nodes is not limited in the present application. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on functions, such as different AI nodes being responsible for different functions.
[0172] It can also be understood that the AI node can be a device independent of each other, or can be integrated into the same device to implement different functions, or can be a network element in a hardware device, or can be a software function running on special hardware, or can be a virtualized function instantiated on a platform (for example, a cloud platform), and the specific forms of the AI node are not limited in the present application.
[0173] The AI node can be an AI network element or an AI module.
[0174] FIG. 3 is a schematic diagram of a possible application framework in a communication system. As shown in FIG. 3, the network elements in the communication system are connected through interfaces (e.g., next generation (NG) interface, Xn interface), or air interfaces. The NG interface is an interface between a radio access network and a 5G core network. The Xn interface is an interface between access network devices, and the air interface is an interface between an access network device and a terminal device. One or more AI modules (only one is shown in FIG. 3 for clarity) are deployed in one or more of the network element nodes, such as a core network device, an access network node (RAN node), a terminal, or an OAM device. The access network node can be a single RAN node or can include multiple RAN nodes, such as a CU and a DU. The CU and / or the DU can also be provided with one or more AI modules. Optionally, the CU can also be split into a CU-CP and a CU-UP. One or more AI models are deployed in the CU-CP and / or the CU-UP.
[0175] The AI module is used to implement a corresponding AI function. The AI modules deployed in different network elements can be the same or different. The AI module can implement different functions according to different parameter configurations of the model of the AI module. The model of the AI module can be configured based on one or more of the following parameters: a structural parameter (such as at least one of the number of neural network layers, the width of the neural network, the connection relationship between layers, the weight of a neuron, the activation function of a neuron, or the bias in the activation function), an input parameter (such as the type of the input parameter and / or the dimension of the input parameter), or an output parameter (such as the type of the output parameter and / or the dimension of the output parameter). The bias in the activation function can also be referred to as the bias of the neural network.
[0176] One AI module can have one or more models. One model can infer an output including one parameter or multiple parameters. The learning process, the training process, or the inference process of different models can be deployed in different nodes or devices, or can be deployed in the same node or device.
[0177] The network device can be a network device provided with one or more AI modules. The network device can include one or more of the core network device, the access network node (RAN node), or the operation administration and maintenance (OAM) shown in FIG. 3. For example, the AI module can be a RAN intelligent controller (RIC) shown in FIG. 4, such as a near-real time RIC (near-RT RIC) or a non-real time RIC (Non-RT RIC). For example, the near-real time RIC is provided in the RAN node (for example, in the CU, the DU), and the non-real time RIC is provided in the OAM, the cloud server, the core network device, or other access network device. The RIC can obtain a subset of data from multiple terminal devices from the RAN node (for example, the CU, the CU-CP, the CU-UP, the DU, and / or the RU), reorganize the subset of data into a training data set, and train based on the training data set. For example, the near-real time RIC and the non-real time RIC can also be provided as a network element, respectively, and the access network device can be the near-real time RIC or the non-real time RIC.
[0178] FIG. 4 is a schematic diagram of another possible application framework in a communication system. In addition to the access network node (CUs, DUs, and RUs are shown in the figure) and the terminal, the communication system shown in FIG. 4 also includes a RIC. For example, the RIC can be an AI module shown in FIG. 3, which can be used to implement AI-related functions. The RIC includes a near-real time RIC and a non-real time RIC. The non-real time RIC mainly processes non-real time information, such as data that is not sensitive to latency, which can be seconds. The near-real time RIC mainly processes near-real time information, such as data that is relatively sensitive to latency, which can be tens of milliseconds.
[0179] The near-real time RIC is used for model training and inference. For example, the near-real time RIC is used to train an AI model and perform inference using the AI model. The near-real time RIC can obtain network side and / or terminal side information from the RAN node (for example, the CU, the CU-CP, the CU-UP, the DU, and / or the RU) and / or the terminal. The information can be used as training data or inference data. Optionally, the near-real time RIC can submit the inference result to the RAN node and / or the terminal. Optionally, the CU and the DU, and / or the DU and the RU can exchange the inference result. For example, the near-real time RIC submits the inference result to the DU, and the DU sends it to the RU.
[0180] The non-real-time RIC is also used for model training and inference. For example, for training an AI model, inference is performed using the model. The non-real-time RIC can obtain network-side and / or terminal-side information from the RAN node (for example, one or more of a CU, a CU-CP, a CU-UP, a DU, or an RU) and / or a terminal. This information can be used as training data or inference data, and the inference result can be delivered to the RAN node and / or the terminal. Alternatively, the inference result can be exchanged between the CU and the DU, and / or between the DU and the RU, for example, the non-real-time RIC delivers the inference result to the DU, which then sends it to the RU.
[0181] The near-real-time RIC and the non-real-time RIC can also be separately provided as a network element. Alternatively, the near-real-time RIC and the non-real-time RIC can also be part of other devices, for example, the near-real-time RIC is provided in the RAN node (for example, in the CU or the DU), and the non-real-time RIC is provided in the OAM, the cloud server, the core network device, or other network devices.
[0182] With the development of wireless communication technology, more and more services are supported, and higher requirements are put forward for the communication system in terms of system capacity, communication delay, and the like. Among them, a massive MIMO system can achieve spatial diversity gain and significantly increase system capacity by configuring a large-scale antenna array at the transceiver end. For example, an access network device can simultaneously send data to multiple terminal devices using the same time-frequency resource, that is, multi-user MIMO (MU-MIMO), or the access network device can also simultaneously send multiple data streams to the same terminal device, that is, single-user MIMO (SU-MIMO). The data between the multiple terminal devices or the multiple data streams of the same terminal device are spatially multiplexed, so it becomes a key direction in the evolution of communication systems.
[0183] The access network device needs to obtain the CSI of the downlink channel to determine the configuration of the downlink data channel of the terminal device, such as the resource, the modulation and coding scheme (MCS), and the precoding.
[0184] Taking precoding as an example, in massive MIMO, an access network device needs to precode downlink data by using a precoding matrix. The access network device can use precoding technology to realize spatial division multiplexing between terminal devices or between data streams, that is, data between different terminal devices or between different data streams of the same terminal device is isolated in space, so as to reduce interference between different terminal devices or between different data streams and improve the signal to interference plus noise ratio (SINR) of the terminal device. In order to calculate the precoding matrix, the access network device needs to obtain the CSI of the downlink channel, and determine the precoding matrix according to the CSI.
[0185] In a time division duplexing (TDD) system, since the uplink and downlink channels are reciprocal, the access network device can obtain the uplink CSI by measuring the uplink reference signal, and then infer the more accurate downlink CSI, for example, use the uplink CSI as the downlink CSI. However, in a frequency division duplexing (FDD) system, the uplink and downlink reciprocity cannot be guaranteed, and the downlink CSI is obtained by the terminal device measuring the downlink reference signal, such as measuring the channel state information reference signal (CSI-RS) or the synchronization signal block (SSB) to obtain the downlink CSI. Therefore, the terminal device needs to generate a CSI report in a manner of predefinition by a protocol or configuration by the access network device, and feed back the generated CSI report to the access network device, so that the access network device obtains the downlink CSI.
[0186] In the FDD system, one important part of the CSI feedback is the precoding matrix indicator (PMI), that is, 0-1 bits are used to quantize the channel matrix or the precoding matrix in the CSI. The design of the PMI (also referred to as the codebook design) is a basic problem in a mobile communication system. The traditional codebook design method is to predefine (agree) a series of precoding matrices and corresponding numbers in the protocol, and these precoding matrices are called code words. The linear combination of the predefined code words or a plurality of predefined code words can approximate the channel matrix or the precoding matrix. Therefore, the terminal device can feed back one or more of the numbers corresponding to the code words and the weighting coefficients to the access network device through the PMI, so as to reconstruct the channel matrix or the precoding matrix by the access network device.
[0187] With the increasing size of the antenna array of the MIMO system, the number of supportable antenna ports increases, and the dimension of the corresponding channel matrix and precoding matrix grows. In order to enable the terminal device to estimate (or measure) the downlink channel, the access network device increases the overhead of the reference signal. At the same time, the error of approximating the large-scale channel matrix and precoding matrix with a limited number of predefined codebooks increases. One method to improve the accuracy of channel reconstruction is to increase the number of codebooks in the codebook, but this will also increase the overhead of the CSI feedback (including one or more of the codebook corresponding number and the weighting coefficient), thereby reducing the available resources for data transmission and causing a loss of system capacity. In summary, it is necessary to study how to more effectively compress the channel information and how to more effectively reconstruct the channel according to the feedback information without increasing the overhead of the reference signal and the overhead of the CSI feedback. There is a correlation between different elements in the downlink channel matrix between the access network device and the terminal device, and there is a correlation between the downlink channel matrices of different time slots. For example, the correlation between different elements in the channel matrix means that there is a set of bases (which can be represented by matrices U1 and U2), and when the channel matrix H is projected onto the set of bases, a sparse equivalent channel H' can be obtained, that is, H' = U1H H U2 is a sparse matrix, where the superscript H represents the conjugate transpose operation. In theory, only the non-zero elements in H' need to be estimated and fed back through the reference signal to reconstruct the channel matrix H. Therefore, there is a compression space for the overhead of the reference signal and the CSI feedback. However, the traditional CSI feedback scheme, such as the above-mentioned feedback method based on the codebook, does not fully utilize the channel compression space, and the channel compression process can cause a large amount of information loss. The method of machine learning (such as deep learning (DL)) has stronger nonlinear feature extraction capability, so it can more effectively extract the correlation between channel matrices, and thus can more effectively compress the channel information and more effectively reconstruct the channel information according to the feedback information compared with the traditional scheme.
[0188] To facilitate understanding of the embodiments of the present application, the terms involved in the present application are first explained.
[0189] 1、Channel state information (CSI): The meaning of CSI is broader than that in the traditional scheme, including but not limited to channel quality indication (CQI), precoding matrix indicator (PMI), rank indicator (RI), CSI-RS resource indicator (CRI), and can also include one or more of the following: channel response information (such as channel response matrix, frequency domain channel response information, time domain channel response information), weight information corresponding to the channel response, reference signal receiving power (RSRP), reference signal receiving quality (RSRQ) or signal to interference plus noise ratio (SINR) and the like.
[0190] Several CSI-related terms involved in the present application are as follows:
[0191] Target CSI: It can also be referred to as full amount of CSI information, pre-compression CSI information, original CSI, original CSI information, etc. In the following, the target CSI is represented by V for convenience of differentiation and description.
[0192] CSI feedback information: It can also be referred to as feedback information of channel measurement result, feedback information of channel information, feedback information of CSI. The CSI feedback information can be obtained by compressing and quantizing the target CSI.
[0193] Compressed CSI: It can also be referred to as compressed information of CSI, compressed information of channel information, compressed information of CSI, compressed channel information, etc. In the following, the compressed CSI obtained by compression is represented by C, and the compressed CSI obtained by de-quantization of the CSI feedback information is represented by . It can be understood that the compressed CSI obtained by de-quantization of the CSI feedback information is compressed CSI with quantization loss.
[0194] Reconstructed CSI: can also be referred to as recovered CSI, reconstructed CSI, recovered CSI, reconstructed information of channel information, decompressed CSI, decompressed channel information, etc. For convenience of differentiation and description below, the reconstructed CSI is represented by .
[0195] Among the several CSI-related terms listed above, the CSI feedback information is the result obtained by compressing the target CSI. In another implementation, the CSI feedback information can also be the result obtained by compressing and quantizing the target CSI, in which case, the CSI feedback information can also include compressed CSI.
[0196] 2. AI model: a function model that maps a certain dimension of input to a certain dimension of output, and the model parameters of which can be obtained through machine learning (ML) training. For example, f(x) = ax 2 +b is a quadratic function model, which can be regarded as an AI model, and a and b correspond to the parameters of the model, which can be obtained through machine learning training.
[0197] 3. AE model: can generally refer to a network structure composed of two AI models, such as an encoder and a decoder. Each model can be an AI model. The AE model can also be referred to as a bilateral model, a double-end model, a cooperative model, etc. The encoder and the decoder of the AE are usually trained together and can be used together.
[0198] In this application, the feedback of the CSI can be implemented based on the AI model of the AE. FIG. 5 is a schematic diagram of CSI feedback using an AE model according to an embodiment of the present application.
[0199] As shown in the figure, the terminal side compresses the target CSI through the encoder, and the network side reconstructs the CSI through the decoder. Exemplarily, the terminal side can take the target CSI (i.e., V) as the input of the encoder, and the encoder can compress the target CSI to obtain the compressed CSI (i.e., C). The terminal side can quantize the compressed CSI to obtain the CSI feedback information. The terminal side can send the CSI feedback information to the network side, such as through a CSI report.
[0200] The network side can first dequantize the CSI feedback information to obtain the compressed CSI with quantization loss (i.e., C ). The network side can take the compressed CSI as the input of the decoder, and the decoder can decompress the CSI based on the compressed CSI to obtain the reconstructed CSI (i.e., ).
[0201] Wherein, the quantizer used for performing quantization can be predefined, such as protocol predefined, or can be indicated by the network side, which is not limited in the present application.
[0202] In another implementation, the encoder on the terminal side can also compress and quantize the target CSI, and output the CSI feedback information. The decoder on the network side can also dequantize and decompress the CSI feedback information to obtain the reconstructed CSI.
[0203] It should be noted that the encoder on the terminal side can be deployed inside the terminal device, or in other devices outside the terminal device, such as the aforementioned OTT host or cloud server, etc. The decoder on the network side can be deployed inside the network device, or in other devices outside the network device, such as the aforementioned intelligent network element.
[0204] It should be understood that although the encoder and the decoder are shown in the figure, this is only a model division from the functional point of view, and the encoder can also be called the first model, and the decoder can also be called the second model. In addition, the present application does not limit the number of models contained in the AE model.
[0205] It should also be understood that the AE model is only one possible model for implementing the above functions, and should not constitute any limitation on the present application. The AE model can also be replaced by other AI models capable of achieving the same or similar functions.
[0206] 4. Model training: By selecting a suitable loss function, the model parameters are trained using an optimization algorithm to minimize the difference between the predicted value of the model and the true value (ground truth) (or target value, label).
[0207] The model training methods involved in the present application include supervised learning, self-supervised learning and knowledge distillation. The following will explain the above methods one by one.
[0208] Supervised learning: also called supervised learning. According to the collected sample value and sample label, the mapping relationship between the sample value and the sample label is learned by using the machine learning algorithm, and the learned mapping relationship is expressed by using the machine learning model. The process of training the machine learning model is the process of learning the mapping relationship. For example, in signal detection, the noisy received signal is the sample, and the true constellation point corresponding to the signal is the label. The machine learning expects to learn the mapping relationship between the sample and the label through training, that is, to learn a signal detector. During training, the model parameters are optimized by calculating the error between the predicted value of the model and the true label. Once the mapping relationship is learned, the learned mapping can be used to predict the sample label of each new sample. The learned mapping relationship of supervised learning can include linear mapping and nonlinear mapping. According to the type of label, the learned task can be divided into classification task and regression task.
[0209] Self-supervised learning: a kind of unsupervised learning. Unsupervised learning is to use the collected sample value to explore the internal pattern of the sample by using the algorithm. Self-supervised learning is to use the sample itself as a supervision signal, that is, the model learns the mapping relationship from the sample to the sample. During training, the model parameters are optimized by calculating the error between the predicted value of the model and the sample itself. Self-supervised learning can be used in signal compression and decompression recovery applications. Common algorithms include autoencoders and generative adversarial networks.
[0210] Knowledge distillation: generally, a large model is usually a single complex network or a collection of several networks, which has good performance and generalization ability. Because the small model has limited expression ability due to its small network size, the knowledge learned by the large model can be used to guide the training of the small model, which is the process of knowledge distillation. Knowledge distillation can make the small model have the same performance as the large model, but the number of parameters is reduced and the inference delay is shortened, thereby realizing model compression and acceleration; and directly training a small model with a large amount of data is not easy to obtain good performance, but training a large model with a large amount of data and then distilling knowledge from the large model to the small model can obtain good performance; in addition, using knowledge distillation can also realize the integration and migration of data sets in different fields.
[0211] Knowledge distillation adopts a teacher-student mode, that is, the teacher model is used to assist the training of the student model. Among them, the teacher model is a complex large model, and the student model is a simple small model. Because the teacher model has strong learning ability, it can transfer the knowledge learned by it to the student model with relatively weak learning ability, so as to enhance the generalization ability of the student model. The complex and heavy but effective teacher model does not go online, and is simply a tutor role. The flexible and lightweight student model is really deployed online to perform prediction tasks.
[0212] 5. Performance metric function: can be used to describe the difference between the predicted value and the true value of the model. When the value of the performance metric function reaches the range required by the performance indicator, the model training can be ended.
[0213] The performance metric function can be a loss function, which can also be called an objective function, a cost function, etc., used to measure the difference between the predicted value and the true value. The higher the output value (loss) of the loss function, the greater the difference, and the process of model training can be to minimize this value. The loss function involved in the embodiments of the present application, for example, has MAE, MSE and NMSE.
[0214] The performance metric function can also be used to measure the similarity between the predicted value and the true value. The performance metric function used to measure the similarity between the predicted value and the true value in the embodiments of the present application, for example, has GCS and SGCS.
[0215] The calculation method of each performance metric function can refer to the formula below:
[0216] wherein, is the predicted value, including N values: H is the true value, including N values: w1, w2, …, w N .
[0217] In another implementation, the above-mentioned GCS and SGCS can also be replaced by the loss function as follows: the difference between GCS and 1, the difference between SGCS and 1, in this case, the process of model training is also the process of minimizing this difference.
[0218] 6. Model file and model parameter: the model file and / or model parameter can be used to determine the model. The model file can be used to indicate the model structure, which includes, for example, but is not limited to, feedforward neural network (FNN), convolutional neural network (CNN) and recurrent neural network (RNN). The model file can have a fixed format, such as a standard pre-defined format, or a format agreed upon by both ends of the interface. The model parameter can refer to the parameter in the neural network model, for example, including but not limited to, the number of layers of the neural network, the type and weight of neurons in each neural network layer, etc. The present application does not limit the way of issuing the parameters of the reference model.
[0219] Take DNN as an example. The idea behind DNN comes from the neuronal structure of the brain. Each neuron can perform a weighted summation operation on its inputs and then use the result of the weighted summation operation to generate an output through a nonlinear function. Figure 6 shows an example of a neuron structure. The input of the neuron shown in Figure 6 is x = [x0 x1 … x N-1 The weights corresponding to the inputs are w = [w0 w1 … w] N-1 The bias of the weighted summation is b. The nonlinear function f() can take many forms; for example, the nonlinear function f() can be the maximum value function max{0, x}. Then the effect of a neuron's execution is... Where N is a positive integer; n is a positive integer greater than or equal to 0 and less than or equal to (N-1).
[0220] A DNN typically has multiple neural network layers, including an input layer, one or more hidden layers, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. Each layer contains multiple neurons. Layers are fully connected; that is, any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. The input layer processes the received numerical values (i.e., the DNN's input) through neurons and then passes them to the hidden layers. Similarly, the hidden layers pass the computation results to the final output layer, producing the DNN's output. Figure 7 shows an example of a DNN. The DNN model shown in Figure 7 has three neural network layers: an input layer, a hidden layer, and an output layer.
[0221] It should be understood that the examples in conjunction with Figures 6 and 7 above are shown for ease of understanding only and should not constitute any limitation on this application. This application does not limit the structure and parameters used in the AI model.
[0222] One of the model structure or model parameters can be predefined, while the other can be sent by the sender (e.g., the network side). Alternatively, both the model structure and model parameters can be sent by the sender (e.g., the network side). This application does not impose any restrictions on this.
[0223] In this embodiment of the application, sending a model may refer to sending a model file and / or model parameters, and receiving a model may refer to receiving a model file and / or model parameters.
[0224] 7. Two-way connection: This refers to the connection between the sender (e.g., the network side) and the receiver (e.g., the terminal side). It can be a connection between datasets or between models. A dataset is used in machine learning for model training, validation, and testing; the quantity and quality of the dataset affect the effectiveness of machine learning. In this embodiment, the dataset can be used for model training.
[0225] The following takes the inference task of CSI compression as an example, and the data set docking and model docking are explained respectively.
[0226] The docking of the data set mainly refers to that the sender provides the data set for the receiver for model training of the receiver. FIG. 8 is a schematic diagram of data set docking between the network side and the terminal side.
[0227] As shown in FIG. 8, the network side can first obtain the data set by joint training. Exemplarily, the network side can use a virtual CSI compression model (such as an encoder) to compress the target CSI, and then use a CSI reconstruction model (such as a decoder) to decompress the target CSI. That is, the target CSI is the input of the CSI compression model, and the output of the CSI compression model can be compressed CSI. The compressed CSI can be the input of the CSI reconstruction model, and the CSI reconstruction model can output reconstructed CSI. In this way, the data set can include one or more of the following: target CSI or reconstructed CSI.
[0228] Optionally, the compressed CSI can be further quantized to obtain CSI feedback information, which can be used for CSI reporting in specific implementation. Correspondingly, the data set can further include one or more of the following: target CSI, CSI feedback information or reconstructed CSI. The quantizer can be predefined, such as protocol predefined, or sent by the sender. Optionally, the data set further includes the quantizer. The quantizer can also be replaced by a quantization method.
[0229] The network side can send the data set to the terminal side. The terminal side can perform model training according to the received data set to obtain the CSI compression model of the terminal side.
[0230] The docking of the model mainly refers to that the sender provides the model file and / or model parameter for the receiver for model deployment of the receiver, or for model development (such as model training) and model deployment of the receiver. In the embodiments of the present application, the sender can provide the model file and / or model parameter for the receiver for model training. FIG. 9 is a schematic diagram of model docking between the network side and the terminal side.
[0231] As shown in FIG. 9, the network side can first obtain the CSI compression model in a joint training manner. The specific procedure of joint training of the network side can refer to the description of FIG. 8, and will not be repeated here. The network side can distribute the CSI compression model or the CSI reconstruction model obtained by joint training to the terminal side. For example, the network side can send the model file and / or model parameters of the CSI compression model to the terminal side, or send the model file and / or model parameters of the CSI reconstruction model to the terminal side. The terminal side can perform model training according to the received model file and / or model parameters.
[0232] It should be understood that the CSI compression model shown in FIG. 8 and FIG. 9 above can be an encoder, and the CSI reconstruction model can be a decoder.
[0233] As described above, the dual-end AI model is usually jointly trained and can be used in matching with each other. In order to support the docking and model development or deployment of the dual-end AI model, a data set or a model needs to be provided for one or both of the dual-end AI models. For example, the network side can provide a data set for the terminal side for model development and deployment; for another example, the network side can provide a model for the terminal side for model development and / or deployment. As can be seen from the description of FIG. 8 and FIG. 9 above, the data contained in the data set is not fixed, and the data set can have multiple types, and different types of data sets contain different data. The model used for training can also have multiple types, such as the CSI compression model, or the CSI reconstruction model, or even other models. If the training of the terminal side does not match the received data set and / or model, the model obtained by training may not meet the requirements, or even be unusable. Therefore, how the terminal side performs model training according to the received data set or model becomes a technical problem to be solved.
[0234] Therefore, the present application provides a method of associating the data set and / or model used for model training with a performance indicator, and associating the performance indicator with an inference task, so that the terminal side can perform corresponding model training according to the received data set and / or model and the corresponding performance indicator, to obtain a model that meets the requirements, thereby improving the feedback performance of the CSI.
[0235] The method provided by the present application will be described in detail below in combination with the drawings.
[0236] FIG. 10 is a schematic flowchart of a communication method 1000 provided by an embodiment of the present application. FIG. 10 shows the flow of the method 1000 from the perspective of the interaction between the terminal side and the network side. The communication method 1000 shown in FIG. 10 can include steps 1010 to 1020. Optionally, the method 1000 further includes one or more of steps 1030 to 1060. Each step in the method 1000 will be described in detail below.
[0237] In step 1010, the terminal side acquires a data set and / or a model, which are used for model training.
[0238] In the embodiments of the present application, the device performing model training at the terminal side can be a terminal device or a server of an OTT system (hereinafter referred to as OTT system server), such as a host or a cloud server of the OTT system. The present application does not limit this. The device needs to acquire a data set and / or a model for model training.
[0239] In one example, the device performing model training at the terminal side is a terminal device, and step 1010 can specifically include: the terminal device receiving a data set and / or a model from a network device, or the terminal device acquiring a data set and / or a model from an OTT system server, which is received by the OTT system from an intelligent network element at the network side.
[0240] In another example, the device performing model training at the terminal side is an OTT system server, and step 1010 can specifically include: the OTT system server receiving a data set and / or a model from an intelligent network element, or the OTT system server acquiring a data set and / or a model from a terminal device, which is received by the terminal device from a network device.
[0241] As can be seen from the above, whether the terminal device acquires a data set and / or a model or the OTT system server acquires a data set and / or a model, the data set and / or the model are received from the network side. Therefore, one possible implementation of step 1010 is that the terminal side receives a data set and / or a model from the network side, and correspondingly, the network side sends a data set and / or a model to the terminal side.
[0242] The sending of a data set can mean sending a data set through a message bearer. The sending of a model can mean sending a model file and model parameters through a message bearer. The model file can be used to indicate a model structure, such as the aforementioned CNN, RNN, etc. The model parameters can be used to indicate the weights and other parameters of each neuron in the model. Based on the model file and the model parameters, the model can be determined.
[0243] The network side can send the dataset and / or the model to the terminal side through one or more messages, which is not limited in the present application. For example, the network device can send the dataset through one message and send the model through another message; for another example, the network side sends the dataset through one or more messages without sending the model; for another example, the network side sends the model through one or more messages without sending the dataset; for another example, the network side sends the dataset and the model through one message; and the like, which is not limited. In the figure, the process of sending the dataset and / or the model from the network side to the terminal side is shown through one step, which should not constitute any limitation to the present application.
[0244] In addition, for the convenience of distinguishing and description, the dataset received from the network side is recorded as the first dataset, and the dataset generated by the terminal side for training mentioned below is recorded as the second dataset.
[0245] In step 1020, the terminal side acquires a performance index, which corresponds to the acquired dataset and / or corresponds to the acquired model.
[0246] In the embodiment of the present application, the performance index corresponds to the inference task of the first model.
[0247] Since the performance index may be different due to different datasets and / or models, and the process of model training may be different due to different datasets and / or models, the terminal side may determine whether the performance index can be met in the process of training the first model, or may determine whether the performance index can be met in the process of training other models, which can be the model used to train the first model. In other words, the terminal side may have started the model training of the first model before determining whether the performance index can be met, or may not have started the model training of the first model.
[0248] Therefore, the performance index corresponding to the inference task of the first model can include that the performance index is a performance requirement to be met by performing the inference task of the first model, and / or the performance index is a performance requirement to be met by performing the inference task through the third model. The third model includes the first model. Illustratively, the inference task of the first model is the inference of compressing CSI. Correspondingly, the inference task of the third model includes the inference of compressing CSI and the inference of decompressing the compressed CSI.
[0249] Exemplarily, the performance indicator can be specifically a range of performance requirement, such as a range of accuracy, that a value of one or more loss functions should satisfy. The loss function may, for example, include but is not limited to a generalized cosine similarity (GCS), a square generalized cosine similarity (SGCS), a mean square error (MSE), a normalized mean square error (NMSE), a mean absolute error (MAE), etc. The MAE can also be referred to as L1 Loss.
[0250] The loss function may, for example, be predefined, such as predefined by a protocol, or indicated by the network side, without limitation in the present application.
[0251] Optionally, the performance indicator is indicated by the network side. The network side can indicate the performance indicator to the terminal side through information (e.g., the first information as shown below).
[0252] Correspondingly, the step 1020 includes 1020a: the terminal side receives first information, the first information being used to indicate the performance indicator. Correspondingly, the network side sends the first information.
[0253] Similar to the step 1010, in one example, the device performing the model training at the terminal side is a terminal device, and the step 1020a can include: the terminal device receives the first information from a network device, or the terminal device receives the first information from an intelligent network element. The information (e.g., the first information or other information) used to indicate the performance indicator can be received by an OTT system server from the intelligent network element. In another example, the device performing the model training at the terminal side is an OTT system server, and the step 1020a can include: the OTT system server receives the first information from an intelligent network element, or the OTT system server receives the first information from a terminal device. The information (e.g., the first information or other information) used to indicate the performance indicator can be received by the terminal device from a network device.
[0254] Optionally, the performance indicator can be predefined, such as predefined by a protocol. For example, the protocol can define a dataset and / or a model used for the model training, and a loss function and a performance indicator corresponding thereto. For another example, the protocol can define a plurality of combinations of the loss function and the performance indicator and a corresponding relationship with the dataset and / or the model.
[0255] Correspondingly, the step 1020 includes 1020b: the terminal side acquires the performance indicator from the terminal side.
[0256] Similar to step 1010, in one example, the device performing the model training at the terminal side is the terminal device, and step 1020b can include: the terminal device obtaining the performance indicator from the local (corresponding to the case that the performance indicator is saved in the terminal device), or the terminal device obtaining the performance indicator from the OTT system server (corresponding to the case that the performance indicator is saved in the OTT system server). In another example, the device performing the model training at the terminal side is the OTT system server, and step 1020b can include: the OTT system server obtaining the performance indicator from the local (corresponding to the case that the performance indicator is saved in the OTT system server), or the OTT system server obtaining the performance indicator from the terminal device (corresponding to the case that the performance indicator is saved in the terminal device).
[0257] The terminal device can obtain the performance indicator from the OTT system server, or the OTT system server can obtain the performance indicator from the terminal device.
[0258] Optionally, the method further includes step 1030, the terminal side performing the model training according to the obtained data set and / or model.
[0259] The data set that the terminal side can obtain can be of multiple types, and the model that the terminal side can obtain can also be of multiple types. Based on different types of data set and / or model, the process of model training is also different. Since the process of model training based on different types of data set and / or model will be described in detail below in combination with multiple drawings, no detailed description is given here.
[0260] It should be noted that since the process of model training can be different due to different data sets and / or models, the performance indicator can also be different due to different data sets and / or models. Therefore, the terminal side can determine whether the performance indicator can be satisfied in the process of training the first model, or can determine whether the performance indicator can be satisfied in the process of training other models, which can be the model used to train the first model. In other words, the terminal side can have started the model training of the first model before determining whether the performance indicator can be satisfied, or can have not started the model training of the first model.
[0261] During the process of model training at the terminal side, the performance indicator can be satisfied or can not be satisfied. The following will be described respectively for the two cases.
[0262] Optionally, the method further includes step 1040, the terminal side performing the inference task of the first model in the case that the performance indicator is satisfied.
[0263] The terminal side can perform the inference task of the first model in a case where the performance indicator is satisfied. Illustratively, the terminal side can obtain the first model through model training. The first model can be deployed to an actual scenario for performing the inference task of the first model in a case where the performance of the first model in performing the inference task of the first model satisfies the performance indicator.
[0264] Optionally, the method further includes a step 1050 of sending, by the terminal side, second information to the network side in a case where the performance indicator cannot be satisfied. Correspondingly, the network side receives the second information.
[0265] The terminal side can report to the network side through the second information in a case where the performance indicator cannot be satisfied.
[0266] Illustratively, the second information is used for one or more of the following:
[0267] a. the performance indicator cannot be satisfied;
[0268] b. requesting to replace the performance indicator;
[0269] c. requesting to replace the dataset;
[0270] d. requesting to replace the model;
[0271] e. requesting to replace the model parameter; or
[0272] f. requesting to stop the model training.
[0273] For example, since the terminal side has limited computing power, the performance indicator is difficult for the terminal side to satisfy, and thus the terminal side can request to replace the performance indicator, such as requesting to replace the performance indicator with a lower requirement, through the second information. The terminal side can also indicate that the performance indicator cannot be satisfied through the second information.
[0274] For another example, since the terminal side can perform model training according to the model sent by the network side, the terminal side can request to replace the model, such as requesting the network side to replace the model, such as replacing another type of model, through the second information. The terminal side can also indicate that the performance indicator cannot be satisfied through the second information.
[0275] For yet another example, since the terminal side can determine the model according to the model parameter sent by the network side and then perform model training, the terminal side can request to replace the model parameter through the second information. The terminal side can also indicate that the performance indicator cannot be satisfied through the second information.
[0276] For example, the terminal side can request to replace the data set by the second information, for example, request to replace the data set by another type of data set, since the terminal side can perform model training according to the data set sent by the network side. The terminal side can also indicate that the performance indicator cannot be met by the second information at the same time.
[0277] For another example, the terminal side can request to close the model training by the second information, since the terminal side has limited computing power and the performance indicator is difficult to meet for the terminal side. The terminal side can also indicate that the performance indicator cannot be met by the second information at the same time. Wherein, closing the model training can be understood as the terminal side no longer performs model training according to the data set and / or model issued by the network side. In other words, the terminal side does not deploy the first model. The network side also no longer configures the terminal side to perform inference tasks of the first model. It should be understood that the operation of the network side after closing the model training can be determined by the network side itself, and the present application does not limit it.
[0278] For another example, the terminal side can also indicate that the performance indicator cannot be met by the second information without indicating other information (such as any one of b to f listed above). What the network side should do after receiving the second information can be determined by the network side itself, or can also be predefined by the protocol. For example, the protocol can define that the model training is closed when the performance indicator is not met, or the performance indicator is replaced when the performance indicator is not met, and the like, which are not listed.
[0279] Similar to the steps above, in one example, the device of the terminal side that performs model training is a terminal device, and the terminal side sends the second information, which can be that the terminal device sends the second information to the network device, or the terminal device sends the second information to the OTT system server, so as to send information (for example, the second information or other information) that can be used to implement one or more of the above a to f to the intelligent network element through the OTT system server. In another example, the device of the terminal side that performs model training is an OTT system server, and the terminal side sends the second information, which can be that the OTT system server sends the second information to the intelligent network element, or the OTT system server sends the second information to the terminal device, so as to send information (for example, the second information or other information) that can be used to implement one or more of the above a to f to the network device through the terminal device.
[0280] It should be understood that steps 1030-1040 and steps 1050 are steps performed by the terminal side in the case that the performance indicator is met or cannot be met, and the terminal side can perform one of them according to the actual situation, and does not necessarily have to perform all of them.
[0281] Optionally, before step 1010, the method 1000 further comprises step 1060: the terminal side sends capability information to the network side, the capability information indicating the computing capability of the terminal side. Accordingly, the network side receives the capability information from the terminal side.
[0282] The computing capability of the terminal side can specifically refer to the computing capability of the device performing model training at the terminal side.
[0283] In one example, the device performing model training at the terminal side is the terminal device, and step 1060 can comprise: the terminal device sending the capability information to the network device, or the terminal device sending the capability information to the OTT system server, so as to send the capability information to the intelligent network element through the OTT system server. In another example, the device performing model training at the terminal side is the OTT system server, and step 1060 can comprise: the OTT system server sending the capability information to the intelligent network element, or the OTT system sending the capability information to the terminal device, so as to send the capability information to the network device through the terminal device.
[0284] The computing capability can be represented by parameters such as the computing speed of the device, the processing time of unit data volume, etc., which are not limited in the present application.
[0285] The network side can determine the data set and / or model with computing capability suitable for the terminal side according to the capability information of the terminal side, and can also determine the performance index of the data set and / or model to be sent to the terminal side according to the capability information of the terminal side and the business requirement. Thus, the data set and / or model for training sent can be avoided from being unsuitable for the capability of the terminal side, causing the model training of the terminal side to fail to meet the performance index, or causing the terminal side to re-request the data set and / or model from the network side, resulting in a lengthy process of model training.
[0286] Based on the above technical solution, the terminal side can obtain the data set and / or model for model training, and obtain the performance index corresponding to the data set and / or model obtained by the terminal side and the inference task of the first model. In other words, after obtaining the data set and / or model and the performance index, the terminal side can determine how to perform model training and the performance index to be met. Thus, the model training can be performed accordingly. In this way, the terminal side can obtain a model meeting the requirements through model training, thereby improving the feedback performance of CSI.
[0287] The training process will be described in detail below with reference to the accompanying drawings. In the following examples, examples one to three are model training according to the data set sent by the network side; and examples four to six are model training according to the model sent by the network side.
[0288] For convenience of distinguishing and description, the data set sent by the network side is referred to as a first data set, and the letters used to represent various parameters of the first data set are identified by a subscript "NW". The letters used to represent various parameters obtained by the terminal side for training in other ways are identified by a subscript "UE".
[0289] The first data set includes one or more of the following: a first target CSI (which can be represented as V NW ), first CSI feedback information, or first reconstructed CSI (which can be represented as ).
[0290] The data used for training on the terminal side includes a second target CSI (which can be represented as V UE ). The second target CSI can be obtained by measurement of the terminal device, can be obtained by simulation on the terminal side, can be predefined by a protocol, and the like, without limitation.
[0291] In addition, the terminal side can also generate a second data set for training by itself, and the second data set includes compressed CSI (which can be represented as C1 UE ) obtained by compressing the second target CSI by a model (such as a fourth model below) sent by the network side.
[0292] In addition, for convenience of distinguishing, the models involved in the training process are first described below.
[0293] The first model: a model that needs to be obtained by model training on the terminal side. The first model can be used to perform an inference task, and more specifically, the first model can be used to perform inference on compression of a target CSI, or in other words, compression of the target CSI. An example, the first model can be an encoder.
[0294] The second model: a model sent by the network side, and the second model is a model that can work with the first model. The second model can be used to decompress the compressed CSI. In a specific implementation, the compressed CSI can be compressed CSI with quantization loss obtained by dequantization of CSI feedback information, and therefore, the second model can also be used to decompress information (i.e., compressed CSI) obtained by dequantization of the CSI feedback information. An example, the second model can be a decoder.
[0295] The third model includes the first model and the second model and has the functions of the first model and the second model. The third model can be used to compress the target CSI to obtain compressed CSI (i.e., the function of the first model), and the third model can also be used to decompress the compressed CSI to obtain reconstructed CSI (i.e., the function of the second model). In a specific implementation, the decompression object (i.e., the compressed CSI) can be compressed CSI with quantization loss obtained by dequantizing the CSI feedback information. Therefore, the third model can also be used to compress the target CSI to obtain compressed CSI, and can also be used to decompress the information obtained by dequantizing the CSI feedback information to obtain reconstructed CSI, where the compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information. In one example, the third model is an AE model including an encoder and a decoder.
[0296] The fourth model is a model sent by the network side and is used for model training of the first model. The fourth model can be used to perform an inference task, and more specifically, the fourth model can be used to perform inference on compression of the target CSI, that is, to compress the target CSI.
[0297] In addition, in order to distinguish the models on the network side and the terminal side, the models on the network side are identified by "1" in the following description and the accompanying drawings, such as the encoder 1 and the decoder 1 on the network side; and the models on the terminal side are identified by "2", such as the encoder 2 and the decoder 2 on the terminal side.
[0298] Example one:
[0299] The network side sends a first data set, which includes a first target CSI (V NW ) and first CSI feedback information. The terminal side performs model training according to the first data set.
[0300] Exemplarily, the network side can first perform joint training of the encoder 1 and the decoder 1 to obtain the first data set. The process of joint training performed by the network side can refer to the exemplary description in the above examples in combination with FIG. 5 and FIG. 9, and will not be described herein again. The network side can send the first data set to the terminal side.
[0301] In one implementation manner, the terminal side can perform model training of the encoder 2 (i.e., an example of the first model) according to the first data set (which is shown as example one A below).
[0302] In another implementation manner, the terminal side can first perform model training of the decoder 2 (i.e., an example of the second model) according to the first data set, and then use the trained decoder to train the encoder 2 and the decoder 2 (i.e., an example of the third model) to obtain the encoder 2.
[0303] The processing flow at the terminal side will be described below with respect to Example 1A and Example 1B respectively.
[0304] The processing flow at the terminal side of Example 1A can refer to FIG. 11A, and can include the following steps:
[0305] ①, dequantize the first CSI feedback information in the first data set to obtain compressed CSI containing quantization loss, i.e., The dequantizer for dequantizing the first CSI feedback information can be predefined, such as protocol predefined, or indicated by the network side, which is not limited in the present application. Optionally, the first data set further includes the dequantizer. Since dequantization and quantization are relative, another possible implementation of the predefined dequantizer is to define a quantizer, and another possible implementation of the network side indicating the dequantizer is to indicate a quantizer.
[0306] ②, take the first target CSI (i.e., V NW ) in the first data set as the input of the encoder 2, and take the compressed CSI (i.e., ) obtained by ① as the label to train the encoder 2. The encoder 2 can compress the input first target CSI to obtain the compressed CSI (i.e., ). Through the training of the encoder 2, the output of the encoder 2 tends to be close to the label . Exemplarily, the model training of the first model can also be called supervised learning for .
[0307] Exemplarily, the performance metric function used for the model training can include one or more of the following:
[0308] The NMSE (i.e., NMSE ) between the information obtained by compressing the first target CSI by the encoder 2 and the information obtained by dequantizing the first CSI feedback information; The MSE (i.e., MSE ) between the information obtained by compressing the first target CSI by the encoder 2 and the information obtained by dequantizing the first CSI feedback information;
[0309] The MSE (i.e., MSE ) between the information obtained by compressing the first target CSI by the encoder 2 and the information obtained by dequantizing the first CSI feedback information; The MSE (i.e., MSE
[0310] ) between the information obtained by compressing the first target CSI by the encoder 2 and the information obtained by dequantizing the first CSI feedback information; the MAE between the information obtained by compressing the first target CSI by the encoder 2 and the information obtained by dequantizing the first CSI feedback information (i.e., the MAE );
[0311] the information obtained by compressing the first target CSI by the encoder 2 and the information obtained by dequantizing the first CSI feedback information a weighted sum of a plurality of the NMSE, the MSE or the MAE between the information obtained by compressing the first target CSI by the encoder 2 and the information obtained by dequantizing the first CSI feedback information (i.e., the NMSE MSE or the MAE ). The weighted coefficients of the plurality of terms can be predefined or indicated by the network side, which is not limited in the present application.
[0312] Correspondingly, the performance indicator for the model training includes a range that the value of one or more performance measurement functions should satisfy. That is, the performance indicator includes a range that one or more of the following should satisfy: the NMSE MSE MAE or a weighted sum of two or more of the above plurality of terms.
[0313] It can be seen that the several performance measurement functions listed above are loss functions. When the difference between the two values compared is greater, the corresponding NMSE, MSE and MAE are also greater; when the difference between the two values compared is smaller, the corresponding NMSE, MSE and MAE are also smaller. In the present example, since it is desired that the information output by the encoder 2 tends to be close to the label , therefore, the performance indicator for the model training can specifically be an upper bound of the value of one or more of the above performance measurement functions. That is, when the value of the performance measurement function is less than the upper bound of the corresponding performance measurement function, the performance indicator is satisfied; when the value of the performance measurement function is greater than the upper bound of the corresponding performance measurement function, the performance indicator is not satisfied.
[0314] For example, the performance measurement function is the MSE , and the performance indicator is the upper bound of the MSE . Then: the value of the MSE is less than the upper bound, and the performance indicator is satisfied; the value of the MSE is greater than the upper bound, and the performance indicator is not satisfied.
[0315] For another example, the performance measurement functions are the MAE and the NMSE , and the performance indicator is the upper bound of the MSE and the upper bound of the NMSE . Then: the value of the MAE is less than the upper bound of the MSE The upper bound of NMSE Less than NMSE When the upper bound is reached, the performance index is satisfied; MAE The value is greater than MSE The upper bound, and / or, NMSE The value is greater than NMSE When the upper bound is reached, the performance index is not met.
[0316] For example, the performance metric function is MAE. MSE and NMSE The performance index is the upper bound of the weighted sum of the three terms. Therefore, MAE is the weighted sum of the three terms. MSE and NMSE The weighted sum is less than the upper bound, satisfying the performance metric; add MAE MSE and NMSE If the sum of the weights is greater than the upper bound, the performance index is not met.
[0317] It should be noted that whether the upper bound of the loss function is considered to satisfy or not satisfy the performance metric can be predefined, such as by the protocol or by prior negotiation between the network and terminal sides. This application does not impose any limitations on this. Furthermore, the upper and lower bounds exemplified above are provided merely for ease of understanding of the performance metric; this application does not limit the specific form of the performance metric, nor does it specify the exact values of the upper and lower bounds. For example, the performance metric can be determined based on business requirements, or it can be determined in conjunction with the computing power of the terminal side, etc., without limitation.
[0318] When encoder 2 outputs With tags When the performance metrics are met, the training of the model can be stopped, thus obtaining the trained encoder 2.
[0319] ③ The trained encoder 2 can be deployed to a specific device to perform model inference. Since encoder 2 can be used to perform inference tasks after model deployment, it is no longer necessary to distinguish between datasets from the network side and the terminal side. Therefore, the target CSI input and output CSI feedback information of the deployed model are no longer distinguished by the first and second methods in the figure. For the same reason, the same processing is performed in the following flowcharts shown in Figures 11B, 12A, 12B, and 13 to 19. For the sake of brevity, it will not be repeated.
[0320] In the process described above and illustrated in Figure 11A, encoder 2 can be used to compress the first target CSI, but not the compressed CSI. Quantization is performed, thus, the terminal side needs to dequantize the first CSI feedback information in the first data set to obtain the compressed CSI as the label.
[0321] In another implementation, the encoder 2 can also be used to quantize the compressed CSI. For example, the encoder 2 also has a quantization function, such as including a quantizer. The quantizer can be predefined, such as protocol predefined, or also indicated by the network device. Optionally, the first data set also includes the quantizer. Correspondingly, the dequantization operation shown in ① above can be omitted. In the processing flow shown in ② above, the terminal side can take the first target CSI as the input of the encoder 2, the encoder 2 can compress the first target CSI and quantize the compressed CSI to obtain and output the CSI feedback information. The terminal side can take the first CSI feedback information in the first data set as the label to train the encoder 2. Through the training of the encoder 2, the CSI feedback information output by the encoder 2 tends to be close to the label first CSI feedback information. Exemplarily, the model training of the first model can be called supervised learning of the first CSI feedback information.
[0322] Since quantization and dequantization are relative, another possible implementation of predefining the quantizer is to predefine the dequantizer, and another possible implementation of the network side indicating the quantizer is to indicate the dequantizer by the network side. Therefore, the first data set optionally includes at least one of the quantizer or the dequantizer.
[0323] The performance metric function used for the model training can include one or more of the following:
[0324] NMSE between the CSI feedback information obtained by the encoder 2 compressing and quantizing the first target CSI and the first CSI feedback information;
[0325] MSE between the CSI feedback information obtained by the encoder 2 compressing and quantizing the first target CSI and the first CSI feedback information;
[0326] MAE between the CSI feedback information obtained by the encoder 2 compressing and quantizing the first target CSI and the first CSI feedback information; or
[0327] A weighted sum of multiple ones of the NMSE, the MSE or the MAE between the CSI feedback information obtained by the encoder 2 compressing and quantizing the first target CSI and the first CSI feedback information. The weighting coefficients of the respective ones can be predefined or indicated by the network side, which is not limited in the present application.
[0328] Correspondingly, the performance indicator for the model training includes a range that the value of one or more performance measurement functions should satisfy.
[0329] For more detailed description of the above-mentioned performance measurement functions and performance indicators, please refer to the relevant description in the above and The relevant description of the performance measurement functions and performance indicators is only a part of them is replaced by the first CSI feedback information, is replaced by the CSI feedback information output by the encoder 2, and will not be described again.
[0330] In addition, if the encoder 2 includes a quantizer, in the flow shown in the above ③, the encoder 2 after model deployment can compress and quantize the input target CSI to output the CSI feedback information.
[0331] Example IB: The processing flow on the terminal side can refer to FIG. 11B, which can include the following steps:
[0332] ①, dequantize the first CSI feedback information in the first data set to obtain compressed CSI containing quantization loss, that is, For more detailed description of ①, please refer to the relevant description in ① in Example IA, which will not be described again.
[0333] ②-1, use the compressed CSI (i.e., ) obtained by dequantizing the first CSI feedback information as the input of the decoder 2 (i.e., an example of the second model), and use the first reconstructed CSI (i.e., ) in the first data set as the label to train the decoder 2. The decoder 2 can decompress the input compressed CSI (i.e., ) to obtain the reconstructed CSI (i.e., ). To distinguish from the reconstructed CSI (i.e., ) in ②-2 below, the reconstructed CSI is denoted by in this place.
[0334] Through the training of the decoder 2, the output by the decoder 2 tends to be close to the label . Exemplarily, the model training of the second model can be called supervised learning for . The training of the decoder 2 is identified as training-1 in the figure.
[0335] ②-2, use the first target CSI (i.e., V NW) as the input of the encoder 2. The encoder 2 can compress the first target CSI to obtain a compressed CSI, which can be as the input of the decoder 2. The decoder 2 can decompress the compressed CSI to obtain the reconstructed CSI (i.e., )). The terminal side can train the encoder 2 and the decoder 2 with the first target CSI (i.e., V NW ) as the target. Exemplarily, the model training of the third model can be called self-supervised learning for V NW .
[0336] Exemplarily, the performance metric function for the third model can include one or more of the following, for example:
[0337] The MSE (i.e., MSE ) between the reconstructed CSI obtained by the decoder 2 compressing the information obtained by the dequantization of the first CSI feedback information and the first target CSI; The NMSE (i.e., NMSE
[0338] ) between the reconstructed CSI obtained by the decoder 2 decompressing the information obtained by the dequantization of the first CSI feedback information and the first target CSI; The MAE (i.e., MAE
[0339] ) between the reconstructed CSI obtained by the decoder 2 decompressing the information obtained by the dequantization of the first CSI feedback information and the first target CSI; The MSE (i.e., MSE
[0340] ) between the reconstructed CSI obtained by the encoder 2 and the decoder 2 processing the first target CSI and the first target CSI; The NMSE (i.e., NMSE
[0341] ) between the reconstructed CSI obtained by the encoder 2 and the decoder 2 processing the first target CSI and the first target CSI; The MAE (i.e., MAE
[0342] ) between the reconstructed CSI obtained by the encoder 2 and the decoder 2 processing the first target CSI and the first target CSI; The GCS (i.e., GCS
[0343] ) between the reconstructed CSI obtained by the encoder 2 and the decoder 2 processing the first target CSI and the first target CSI;
[0344] The GCS (i.e., SGCS) between the reconstructed CSI obtained by processing the first target CSI through encoder 2 and decoder 2 and the first target CSI is... );or
[0345] The MSE, NMSE, MAE of the reconstructed CSI obtained by decompressing the first CSI feedback information obtained by decoder 2 after dequantization, and the weighted sum of multiples of NMSE, MSE, MAE, GCS or SGCS between the reconstructed CSI and the first target CSI obtained by processing the first target CSI by encoder 2 and decoder 2 (i.e., MSE) NMSE MAE MSE NMSE MAE GCS or SGCS (The weighted sum of two or more terms in the equation). The weighting coefficients of each term can be predefined or indicated by the network side, which is not limited in this application.
[0346] Specifically, decoder 2 decompresses the information obtained by dequantizing the first CSI feedback information. This can mean that the information obtained by dequantizing the first CSI feedback information is used as input to decoder 2, and the dequantized information can be compressed CSI. Decoder 2 decompresses the dequantized information to obtain the reconstructed CSI.
[0347] The processing of the first target CSI by encoder 2 and decoder 2 may include: encoder 2 compressing the first target CSI to obtain compressed CSI, and decoder 2 decompressing the compressed CSI to obtain reconstructed CSI.
[0348] Accordingly, the performance metrics used for training this model include the range that the values of one or more performance metric functions should satisfy. That is, the performance metrics include the range that one or more of the following should satisfy: MSE NMSE MAE NMSE MSE MAE GCS SGCS Or a weighted sum of two or more of the above terms.
[0349] As can be seen from the performance metrics listed above, MSE NMSE MAE NMSE MSE and MAE is a loss function, and in this example, since it is desired that the output of the decoder 2 after processing tends to be close to the input V NW , and / or, the output of the encoder 2 and the output of the decoder 2 after processing tends to be close to the input V NW , therefore when the performance metric function includes MSE NMSE MAE NMSE MSE MAE or a weighted sum of two or more of the above, the performance indicator for the model training can be an upper bound of the value of each performance metric function. That is, when the value of the performance metric function is less than the upper bound of the corresponding performance metric function, the performance indicator is satisfied; when the value of the performance metric function is greater than the upper bound of the corresponding performance metric function, the performance indicator is not satisfied.
[0350] GCS SGCS or a weighted sum thereof is used to describe the degree of similarity between the two. In this example, since it is desired that the output of the encoder 2 and the output of the decoder 2 after processing tends to be close to the input V NW , therefore when the performance metric function includes GCS SGCS or a weighted sum thereof, the performance indicator for the model training can be a lower bound of the value of each performance metric function. That is, when the value of the performance metric function is greater than the lower bound of the corresponding performance metric function, the performance indicator is satisfied; when the value of the performance metric function is less than the lower bound of the corresponding performance metric function, the performance indicator is not satisfied.
[0351] For example, the performance metric function is MAE and the performance indicator is the upper bound of MAE , then the value of MAE is less than the upper bound, the performance indicator is satisfied; the value of MAE is greater than the upper bound, the performance indicator is not satisfied.
[0352] For another example, the performance metric function is MSE and GCS , the performance indicator is the upper bound of MSE and the lower bound of GCS lower bound of the MSE value of the MSE upper bound of the GCS value of the GCS lower bound of the MSE value of the MSE upper bound of the GCS value of the GCS lower bound of the MSE
[0353] For example, the performance metric function is the GCS and the SGCS , and the performance metric is a lower bound of the weighted sum of the GCS and the SGCS . Then, when the weighted sum of the GCS and the SGCS is greater than the lower bound, the performance metric is satisfied; and when the weighted sum of the GCS and the SGCS is less than the lower bound, the performance metric is not satisfied.
[0354] It should be noted that whether the value of the performance metric function is equal to the upper bound or the lower bound is determined as satisfying the performance metric or not satisfying the performance metric can be predefined, such as predefined by a protocol, or pre-agreed by the network side and the terminal side. The present application does not limit this. In addition, the upper bound and the lower bound shown in the above examples are only given for the convenience of understanding the performance metric, and the present application does not limit the specific form of the performance metric, nor the specific value of the upper bound and the lower bound. For example, the performance metric can be determined according to the service requirement, or can also be determined in combination with the computing capability of the terminal side, etc., which are not limited.
[0355] It should also be noted that if the GCS and the SGCS in the above are replaced by the difference between the GCS and 1, and the difference between the SGCS and 1 respectively, the performance metric can also be the upper bound of the corresponding performance metric function. The replacement of the GCS and the SGCS and the corresponding change of the performance metric here are applicable to the performance metric of the GCS and the SGCS in the following, and for the sake of brevity, will not be repeated in the following.
[0356] When the output of the decoder 2 satisfies the performance metric, the model training can be stopped, and thus the trained encoder 2 can be obtained. NW When the output of the decoder 2 satisfies the performance metric, the model training can be stopped, and thus the trained encoder 2 can be obtained.
[0357] ③, the trained encoder 2 can be deployed to a specific device to perform model inference.
[0358] In the process shown in FIG. 11B, the first CSI feedback information is dequantized before being input into the decoder 2, and the compressed CSI is input into the decoder 2.
[0359] In another implementation, the decoder 2 can also be used to dequantize the first CSI feedback information. For example, the decoder 2 also has a dequantization function, such as including a dequantizer. The dequantizer can be predefined, such as being predefined, or can also be indicated by the network side. Optionally, the first data set also includes the dequantizer. In this case, the dequantization operation shown in ① above can be omitted, and the terminal side can directly input the first CSI feedback information in the first data set into the decoder 2 in the process shown in ②, and obtain the reconstructed CSI by dequantization and decompression of the decoder 2.
[0360] Correspondingly, when the terminal side performs model training of the decoder 2, the first CSI feedback information can be directly input into the decoder 2. The decoder 2 can dequantize the first CSI feedback information to obtain the compressed CSI, and then decompress the compressed CSI to obtain the reconstructed CSI. When the terminal side performs model training of the encoder 2 and the decoder 2, the encoder 2 can include a quantizer, and the compressed CSI output by the encoder 2 is quantized by the quantizer to obtain the CSI feedback information, which is then input into the decoder 2.
[0361] In addition, if the encoder 2 includes the quantizer, in the process shown in ③ above, the encoder 2 after model deployment can compress and quantize the input target CSI to output the CSI feedback information.
[0362] It can be understood that dequantization and quantization are relative, and therefore another possible implementation of the predefined dequantizer is to define a quantizer, and another possible implementation of the network side indicating the dequantizer is to indicate a quantizer. Therefore, the first data set optionally includes at least one of the quantizer or the dequantizer.
[0363] Example II:
[0364] The network side sends the first data set, which includes the first CSI feedback information and the first reconstructed CSI The terminal side performs model training according to the first data set.
[0365] Exemplarily, the network side can first perform joint training of the encoder 1 and the decoder 1 to obtain the first data set. The process of the network side performing joint training can refer to the exemplary description in the above examples in combination with FIG. 5 and FIG. 9, and will not be described herein.
[0366] The network side can send the first data set to the terminal side. The terminal side can perform model training of the decoder 2 (i.e., an example of the second model) according to the first data set to obtain the decoder 2. Then, under the premise of parameter solidification of the decoder 2, the model training of the encoder 2 and the decoder 2 is performed by the second target CSI (i.e., V UE ) of the terminal side. Illustratively, this process can also be referred to as self-supervised learning for V UE . Alternatively, the terminal side can also perform the model training of the encoder 2 and the decoder 2 under the premise of parameter solidification of the decoder 2 by the first reconstruction CSI of the network side. Illustratively, this process can also be referred to as self-supervised learning for .
[0367] The processing flow of the terminal side can refer to FIGS. 12A and 12B, which are the same as other steps except that step ②-2 is different. For the sake of brevity, the same steps will not be described separately for FIGS. 12A and 12B. The flow shown in FIGS. 12A and 12B can include the following steps:
[0368] ①, dequantize the first CSI feedback information in the first data set to obtain compressed CSI containing quantization loss, i.e., For more detailed description of ①, please refer to the related description in ① of Example 1A, which will not be repeated.
[0369] ②-1, use the compressed CSI (i.e., ) obtained by dequantizing the first CSI feedback information as the input of the decoder 2 (i.e., an example of the second model), and use the first reconstruction CSI (i.e., ) in the first data set as the label to perform model training of the decoder 2. The decoder 2 can decompress the input compressed CSI (i.e., ) to obtain the reconstructed CSI (i.e., ). Through the training of the decoder 2, the output of the decoder 2 tends to be close to the label . Illustratively, the model training of the second model can be referred to as supervised learning for V . The training of the decoder 2 is identified as training-1 in the figure.
[0370] ②-2, solidify the parameters of the decoder 2, train the encoder 2 and the decoder 2. The encoder 2 and the decoder 2 can be regarded as a double-end model, which is an example of the third model. The model training of the encoder 2 and the decoder 2 can be regarded as the model training of the third model. Since the model training of the encoder 2 and the decoder 2 is the model training in the case of solidifying the parameters of the decoder 2, the parameters of the encoder 2 can be optimized, and thus the encoder 2 can be obtained by the model training of the encoder 2 and the decoder 2. Therefore, the model training of the encoder 2 and the decoder 2 can also be regarded as the model training of the encoder 2 (i.e., an example of the first model). The training of the encoder 2 and the decoder 2 is identified as training-2 in the figure.
[0371] In the flow shown in FIG. 12A, the second target CSI (i.e., V UE ) on the terminal side is taken as the input of the encoder 2. The encoder 2 can compress the second target CSI to obtain the compressed CSI, which can be taken as the input of the decoder 2. The decoder 2 can decompress the compressed CSI to obtain the reconstructed CSI (i.e., ) on the terminal side. The encoder 2 and the decoder 2 can be trained with the second target CSI (i.e., V UE ) on the terminal side as the target. Exemplarily, the model training of the third model can be called self-supervised learning for V UE .
[0372] Exemplarily, the performance measurement function used for the model training of the third model can include one or more of the following, for example:
[0373] MSE (i.e., MSE ) between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information by the decoder 2 and the first reconstructed CSI;
[0374] NMSE (i.e., NMSE ) between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information by the decoder 2 and the first reconstructed CSI;
[0375] MAE (i.e., MAE ) between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information by the decoder 2 and the first reconstructed CSI;
[0376] MSE (i.e., MSE ) between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI;
[0377] NMSE between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI (i.e., NMSE );
[0378] MAE between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI (i.e., MAE );
[0379] GCS between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI (i.e., GCS );
[0380] SGCS between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI (i.e., SGCS ); or
[0381] MSE, NMSE, MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information by the decoder 2 and the first reconstructed CSI, and the MSE, NMSE, MAE, GCS or SGCS of multiple items between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI (i.e., the weighted sum of two or more items of MSE NMSE MAE NMSE MSE MAE GCS or SGCS ). The weighting coefficients of each item can be predefined or indicated by the network side, which is not limited in the present application.
[0382] Wherein, the processing of the second target CSI by the encoder 2 and the decoder 2 can include: the encoder 2 compresses the second target CSI to obtain compressed CSI, and the decoder 2 decompresses the compressed CSI to obtain the reconstructed CSI.
[0383] Correspondingly, the performance indicators for training the model include the range that the value of one or more performance measurement functions should meet. That is, the performance indicators include the range that one or more of the following should meet: MSE NMSE MAE NMSE MSE MAE GCS SGCS or a weighted sum of two or more of the above.
[0384] It can be seen that among the several performance metric functions listed above, MSE NMSE MAE NMSE MSE and MAE are loss functions, and in the present example, since the output of the decoder 2 after processing tends to be close to the input , and / or, it is desired that the output of the encoder 2 and the decoder 2 after processing tends to be close to the input V UE , therefore when the performance metric function includes MSE NMSE MAE NMSE MSE MAE or a weighted sum of two or more of the above, the performance indicator for the model training can be an upper bound of the weighted sum of the two or more. That is, when the value of the performance metric function is less than the upper bound of the corresponding performance metric function, the performance indicator is satisfied; when the value of the performance metric function is greater than the upper bound of the corresponding performance metric function, the performance indicator is not satisfied.
[0385] GCS SGCS or a weighted sum thereof is used to describe the degree of similarity between. In the present example, since it is desired that the output of the encoder 2 and the decoder 2 after processing tends to be close to the input V UE , therefore when the performance metric function includes GCS SGCS or a weighted sum thereof, the performance indicator for the model training can be a lower bound of the value of each performance metric function. That is, when the value of the performance metric function is greater than the lower bound of the corresponding performance metric function, the performance indicator is satisfied; when the value of the performance metric function is less than the lower bound of the corresponding performance metric function, the performance indicator is not satisfied.
[0386] It should be noted that whether the value of the performance metric function is equal to the upper bound or the lower bound is determined as meeting the performance indicator or not meeting the performance indicator can be predefined, such as predefined by a protocol, or pre-agreed by the network side and the terminal side. The present application does not limit this. The correspondence between the performance metric function and the performance indicator can refer to the examples of Example 1A and Example 1B above, and will not be repeated here.
[0387] When the output of the decoder 2 meets the performance indicator, the model training can be stopped, and thus the trained encoder 2 can be obtained. UE When the output of the decoder 2 meets the performance indicator, the model training can be stopped, and thus the trained encoder 2 can be obtained.
[0388] In the flow shown in FIG. 12B, the first reconstructed CSI (i.e., ) of the network side is taken as the input of the encoder 2. The encoder 2 can compress the first reconstructed CSI to obtain the compressed CSI, which can be taken as the input of the decoder 2. The decoder 2 can decompress the compressed CSI to obtain the reconstructed CSI (for the convenience of distinguishing from the reconstructed CSI in the above examples, denoted as ). The terminal side can take the first reconstructed CSI (i.e., ) as the target to train the encoder 2 and the decoder 2. Exemplarily, the model training of the third model can be called self-supervised learning for.
[0389] Exemplarily, the performance metric function used for the model training of the third model can include one or more of the following, for example:
[0390] The MSE (i.e., MSE ) between the reconstructed CSI obtained by processing the first reconstructed CSI by the encoder 2 and the decoder 2 and the first reconstructed CSI;
[0391] The NMSE (i.e., NMSE ) between the reconstructed CSI obtained by processing the first reconstructed CSI by the encoder 2 and the decoder 2 and the first reconstructed CSI;
[0392] The MAE (i.e., MAE ) between the reconstructed CSI obtained by processing the first reconstructed CSI by the encoder 2 and the decoder 2 and the first reconstructed CSI;
[0393] The GCS (i.e., GCS ) between the reconstructed CSI obtained by processing the first reconstructed CSI by the encoder 2 and the decoder 2 and the first reconstructed CSI;
[0394] The GCS (i.e., SGCS) between the reconstructed CSI obtained by processing the first reconstructed CSI using encoder 2 and decoder 2 and the first reconstructed CSI is... );or
[0395] The weighted sum of multiples among NMSE, MSE, MAE, GCS, or SGCS between the reconstructed CSI obtained by processing the first reconstructed CSI through encoder 2 and decoder 2 (i.e., NMSE) and the first reconstructed CSI. MSE MAE GCS or SGCS (The weighted sum of two or more terms in the equation). The weighting coefficients of each term can be predefined or indicated by the network side, which is not limited in this application.
[0396] The processing of the first reconstructed CSI by encoder 2 and decoder 2 may include: encoder 2 compressing the first reconstructed CSI to obtain compressed CSI, and decoder 2 decompressing the compressed CSI to obtain reconstructed CSI.
[0397] Accordingly, the performance metrics used for training this model include the range that the values of one or more performance metric functions should satisfy. That is, the performance metrics include the range that one or more of the following should satisfy: NMSE MSE MAE GCS SGCS Or a weighted sum of two or more of the above terms.
[0398] For a more detailed explanation of performance metrics and indicators, please refer to the description of Figure 12A above. The difference is that V... UE It was replaced with It was replaced with No further explanation needed.
[0399] ③ The trained encoder 2 can be deployed to a specific device to perform model inference.
[0400] In the process described above in conjunction with Figures 12A and 12B, the first CSI feedback information needs to be dequantized before being input into decoder 2 to obtain compressed CSI, which is then input into decoder 2.
[0401] In another implementation, the decoder 2 can also be used to dequantize the first CSI feedback information. For example, the decoder 2 also has a dequantization function, such as including a dequantizer. The dequantizer can be predefined, such as predefined, or also indicated by the network side. Optionally, the first data set also includes the dequantizer. In this case, the dequantization operation shown in ① above can be omitted, and the terminal side can directly input the first CSI feedback information in the first data set into the decoder 2 in the process shown in ②, and obtain the reconstructed CSI by dequantization and decompression of the decoder 2.
[0402] Correspondingly, the terminal side can directly input the first CSI feedback information into the decoder 2 when performing model training of the decoder 2. The decoder 2 can dequantize the first CSI feedback information to obtain compressed CSI, and then decompress the compressed CSI to obtain the reconstructed CSI. When the terminal side performs model training of the encoder 2 and the decoder 2, the encoder 2 can include a quantizer to quantize the compressed CSI after output; or the encoder 2 can also not include the quantizer, and the compressed CSI output by the encoder 2 needs to be quantized by the quantizer to obtain the CSI feedback information before being input into the decoder 2.
[0403] In addition, if the encoder 2 includes the quantizer, in the process shown in ③ above, the encoder 2 after model deployment can compress and quantize the input target CSI or reconstructed CSI to output quantized information.
[0404] It can be understood that dequantization and quantization are relative, and therefore another possible implementation of the predefined dequantizer is to define a quantizer, and another possible implementation of the network side indicating the dequantizer is to indicate a quantizer. Therefore, the first data set optionally includes at least one of the quantizer or the dequantizer.
[0405] Example Three:
[0406] The network side sends the first data set, which includes the first target CSI (V NW ), the first CSI feedback information, and the first reconstructed The terminal side performs model training according to the first data set.
[0407] The network side can first perform joint training of the encoder 1 and the decoder 1 to obtain the first data set. The process of the network side performing joint training can refer to the exemplary description of the above in combination with FIG. 5 and FIG. 9, and will not be repeated here.
[0408] The network side can send the first data set to the terminal side. The terminal side can perform model training of the encoder 2 and the decoder 2 according to the first data set to obtain the encoder 2. It should be understood that the encoder 2 and the decoder 2 can be regarded as a double-end model, that is, an example of the third model. The model training of the encoder 2 and the decoder 2 can be regarded as the model training of the third model. In this way, the terminal side can obtain the trained encoder 2 (hereinafter referred to as Example IIIA as an example).
[0409] In another implementation manner, after the model training of the encoder 2 and the decoder 2, the terminal side can further train the encoder 2 (hereinafter referred to as Example IIIB as an example); or, further train the decoder 2, and then use the trained decoder 2 to train the encoder 1 (hereinafter referred to as Example III C as an example).
[0410] The processing flow of the terminal side will be described below with respect to Example IIIA, Example IIIB and Example III C respectively.
[0411] Example IIIA, the processing flow of the terminal side can refer to FIG. 13, and can include the following steps:
[0412] ①, the first target CSI (that is, V NW ) as the input of the encoder 2. The encoder 2 can compress the first target CSI to obtain the compressed CSI. Then, the compressed CSI is taken as the input of the decoder 2. The decoder 2 can decompress the compressed CSI to obtain the reconstructed CSI (that is, V ). The terminal side can take the first reconstructed CSI (that is, V ) as the label to train the encoder 2 and the decoder 2. For example, the model training of the third model can be supervised learning for V .
[0413] For example, the performance measurement function used for the model training of the third model can include one or more of the following:
[0414] MSE (that is, MSE ) between the reconstructed CSI obtained by processing the first target CSI through the third model and the first reconstructed CSI;
[0415] NMSE (that is, NMSE ) between the reconstructed CSI obtained by processing the first target CSI through the third model and the first reconstructed CSI;
[0416] MAE (that is, MAE ) between the reconstructed CSI obtained by processing the first target CSI through the third model and the first reconstructed CSI;
[0417] The GCS (Gross Component Relationship) between the reconstructed CSI obtained by processing the first target CSI using the third model and the first reconstructed CSI is... );
[0418] The SGCS (i.e., SGCS) between the reconstructed CSI obtained by processing the first target CSI using the third model and the first reconstructed CSI is... );or
[0419] The weighted sum of multiples among MSE, NMSE, MAE, GCS, or SGCS between the reconstructed CSI obtained by processing the first target CSI using the third model and the first reconstructed CSI (i.e., MSE) NMSE MAE GCS or SGCS (The weighted sum of two or more terms in the equation). The weighting coefficients of each term can be predefined or indicated by the network side, which is not limited in this application.
[0420] The third model's processing of the first target CSI may include: encoder 2 compressing the first target CSI to obtain compressed CSI, and decoder 2 decompressing the compressed CSI to obtain reconstructed CSI.
[0421] Accordingly, the performance metrics used for training this model include the range that the values of one or more performance metric functions should satisfy. That is, the performance metrics include the range that one or more of the following should satisfy: MSE NMSE MAE GCS SGCS Or a weighted sum of two or more of the above terms.
[0422] As can be seen from the performance metrics listed above, MSE NMSE MAE Or a weighted sum of two or more of them used to describe and The degree of difference between them, GCS SGCS Or its weighted sum is used to describe and The degree of similarity between them. In this example, since we want the outputs processed by encoder 2 and decoder 2 to be similar. With tags They are approaching each other, therefore, for MSE NMSE MAE or a weighted sum of two or more of them, the performance indicator can be an upper bound of the values of the individual functions; for GCS SGCS or a weighted sum of two or more of them, the performance indicator can be a lower bound of the values of the individual functions.
[0423] The cases of meeting and not meeting the performance indicator are already detailed in Example One and Example Two above in connection with multiple examples, and reference is made to the relevant description above, which will not be repeated.
[0424] When the performance indicator is met by the encoder output and the label the model training can be stopped, and thus the trained encoder 2 can be obtained.
[0425] ②, deploy the trained encoder 2 to a specific device to perform model inference.
[0426] In the flow shown in the above example in connection with FIG. 13, the encoder 2 does not include a quantizer. Therefore, in the model inference process, the encoder 2 is used to compress the input target CSI without quantization. Therefore, after the encoder outputs the compressed CSI, the quantizer can be used to quantize the compressed CSI to obtain the CSI feedback information.
[0427] In another implementation, the encoder 2 can also include a quantizer and have a quantization function. Correspondingly, the decoder 2 can include a dequantizer and have a dequantization function. Alternatively, the decoder 2 can not include a dequantizer, but the output of the encoder 2 can be input to the dequantizer for dequantization to obtain the compressed CSI before being input to the decoder 2. It can be understood that quantization and dequantization are relative, and therefore the quantizer and / or the dequantizer can be predefined, such as by a protocol, or can be indicated by the network side. Optionally, the first data set further includes at least one of the quantizer or the dequantizer.
[0428] In this case, the processing flow of ① above is as follows: the terminal side can input the first target CSI as the input of the encoder 2. The encoder 2 can compress and quantize the first target CSI to obtain the CSI feedback information. The terminal side can input the CSI feedback information to the decoder 2. The decoder 2 can dequantize and decompress the CSI feedback information to obtain the reconstructed CSI.
[0429] In addition, if the encoder 2 includes a quantizer, in the flow shown in ② above, the encoder 2 after model deployment can compress and quantize the input target CSI to output the CSI feedback information.
[0430] It should be understood that in Example 3A, the first dataset may include the first target CSI and the first reconstructed CSI, but may not include the first CSI feedback information, or may include the first CSI feedback information, without limitation.
[0431] Example 3B, the terminal-side processing flow can be seen in Figure 14, and may include the following steps:
[0432] ①-1, Set the first target CSI (i.e., V) NW The first target CSI is used as input to encoder 2. Encoder 2 can compress the first target CSI to obtain compressed CSI. The compressed CSI is then used as input to decoder 2. Decoder 2 can decompress the compressed CSI to obtain the reconstructed CSI (i.e., The terminal side can reconstruct the CSI as the first reconfiguration (i.e., Using labels () as the base, encoder 2 and decoder 2 are trained. Exemplarily, this process can also be referred to as training... Supervised learning. In the figure, the training of encoder 2 and decoder 2 is marked as training-1.
[0433] ①-2. The first target CSI (i.e., V) in the first dataset. NW The first CSI feedback information from the first dataset is dequantized to obtain the compressed CSI (i.e., ) as input to encoder 2. The first target CSI is used as a label to train the encoder 2 model. Encoder 2 can compress the input first target CSI to obtain the compressed CSI (i.e., By training encoder 2, the output of encoder 2 is made... With tags Approaching each other. For example, this process can also be called targeting... Supervised learning. In the figure, the training of encoder 2 is labeled as training-2.
[0434] Correspondingly, the performance metric function used for model training of encoder 2 may include one or more of the following:
[0435] Information obtained by compressing the first target CSI using encoder 2 Information obtained by dequantizing the first CSI feedback information The NMSE between (i.e., NMSE) );
[0436] Information obtained by compressing the first target CSI using encoder 2 Information obtained by dequantizing the first CSI feedback information The MSE between (i.e., MSE) );
[0437] Information obtained by compressing the first target CSI using encoder 2 Information obtained by dequantizing the first CSI feedback information The MAE between (i.e., MAE) );or
[0438] Information obtained by compressing the first target CSI using encoder 2 Information obtained by dequantizing the first CSI feedback information The weighted sum of multiple terms in NMSE, MSE, or MAE (i.e., NMSE) MSE or MAE (The weighted sum of two or more terms in the equation). The weighting coefficients of each term can be predefined or indicated by the network side, which is not limited in this application.
[0439] Accordingly, the performance metrics used for training this model include the range that the values of one or more performance metric functions should satisfy. That is, the performance metrics include the range that one or more of the following should satisfy: NMSE MSE MAE Or a weighted sum of two or more of the above terms.
[0440] When encoder 2 outputs With tags When the performance metrics are met, the training of the model can be stopped, thus obtaining the trained encoder 2.
[0441] For a more detailed explanation of the processing flow, performance metric function, and performance indicators for ①-2, please refer to the relevant description in ② of Example 1, which will not be repeated here. The situations where performance indicators are met and not met have been explained in detail with multiple examples in Examples 1 and 2 above, which can be referred to in the relevant descriptions above, and will not be repeated here.
[0442] ② Deploy the trained encoder 2 onto a specific device and execute model inference.
[0443] It should be understood that Example 3B can be regarded as a combination of Example 3A and Example 1. That is, the third model is trained first using the method provided by Example 3A, and then the first model is trained using the method provided by Example 1.
[0444] In the process shown in FIG. 14, the decoder 2 does not include a dequantizer. Therefore, after the terminal side obtains the first CSI feedback information, the terminal side can perform dequantization on the first CSI feedback information, obtain the compressed CSI, and then input the compressed CSI into the decoder 2 to obtain the compressed CSI with quantization loss.
[0445] In another implementation, the decoder 2 can also include a dequantizer and have a dequantization function. The dequantizer can be predefined, such as protocol predefined, or indicated by the network side. Optionally, the first data set also includes the dequantizer. In this case, the process in ①-2 above can be omitted.
[0446] Correspondingly, the encoder 2 can also include a quantizer and have a quantization function. The quantizer can be predefined, such as protocol predefined, or indicated by the network side. Optionally, the first data set also includes the quantizer. In this case, in the process in ①-2 above, the encoder 2 can compress and quantize the input first target CSI to obtain quantized information, i.e., the CSI feedback information. The terminal side can use the first CSI feedback information in the first data set as a label to train the model of the encoder 2.
[0447] In addition, if the encoder 2 includes the quantizer, in the process shown in ② above, the encoder 2 after model deployment can compress and quantize the input target CSI to output the CSI feedback information.
[0448] It can be understood that quantization and dequantization are relative, and therefore the quantizer and / or the dequantizer can be predefined, such as protocol predefined, or indicated by the network side. Optionally, the first data set also includes at least one of the quantizer or the dequantizer.
[0449] Example Three C, the processing flow of the terminal side can refer to FIG. 15 and can include the following steps:
[0450] ①-1, input the first target CSI (i.e., V NW ) into the encoder 2. The encoder 2 can compress the first target CSI to obtain the compressed CSI. Then input the compressed CSI into the decoder 2. The decoder 2 can decompress the compressed CSI to obtain the reconstructed CSI (i.e., ). The terminal side can use the first reconstructed CSI (i.e., ) as a label to train the encoder 2 and the decoder 2. Exemplarily, this process can also be referred to as supervised learning for . In the figure, the training of the encoder 2 and the decoder 2 is identified as training-1.
[0451]
[0452]
[0453]
[0454]
[0455]
[0456]
[0457]
[0458] The GCS (i.e., SGCS) between the reconstructed CSI obtained by processing the first target CSI using encoder 2 and decoder 2 and the first target reconstructed CSI );or
[0459] The weighted sum of multiples among NMSE, MSE, MAE, GCS, or SGCS between the reconstructed CSI obtained by processing the first target CSI through encoder 2 and decoder 2 and the first target reconstructed CSI (i.e., NMSE) MSE MAE GCS or SGCS (The weighted sum of two or more terms in the equation). The weighting coefficients of each term can be predefined or indicated by the network side, which is not limited in this application.
[0460] The encoder 2 and decoder 2 process the first target CSI by: the encoder 2 compressing the first CSI to obtain a compressed CSI; and the decoder 2 decompressing the compressed CSI to obtain a reconstructed CSI.
[0461] Accordingly, the performance metrics used for training this model include the range that the values of one or more performance metric functions should satisfy. That is, the performance metrics include the range that one or more of the following should satisfy: NMSE MSE MAE GCS SGCS Or a weighted sum of two or more of the above terms.
[0462] When decoder 2 outputs With tags When the performance metrics are met, the training of the model can be stopped, thus obtaining the trained encoder 2.
[0463] For a more detailed explanation of the processing flow, performance metric function, and performance indicators for ①-2 and ①-3, please refer to the relevant descriptions in ②-1 and ②-2 of Example 2. The difference is that in Example 3C, the training of encoder 2 in ①-3 is based on the data in the first dataset sent from the network side, and will not be elaborated further. The situations where performance indicators are met or not met have been explained in detail in Example 1 and Example 2 above with multiple examples; please refer to the relevant descriptions above, and will not be repeated here.
[0464] ② Deploy the trained encoder 2 onto a specific device and execute model inference.
[0465] It should be understood that example three C can be regarded as a combination of example three A and example two, i.e., the model training of the third model is performed by the method provided in example three A first, and then the model training is performed by the way provided in example two.
[0466] Similar to example three B above, the decoder 2 in example three C can also include a dequantizer, having a dequantization function. In this case, in the processing flow of ①-2 above, the dequantization process can be omitted.
[0467] The encoder 2 in example three C can also have a quantizer, having a quantization function. In this case, in the processing flow of ①-3 above, the output of the encoder 2 can be the CSI feedback information obtained by compressing and quantizing the first target CSI, and the input of the decoder 2 can be the CSI feedback information.
[0468] In addition, if the encoder 2 includes a quantizer, in the flow shown in ② above, the encoder 2 after model deployment can compress and quantize the input target CSI, and output the CSI feedback information.
[0469] Optionally, the first data set further includes at least one of a quantizer or a dequantizer. The dequantizer and the quantizer are described above and will not be repeated here.
[0470] It should be understood that in the above examples one to three, the first data set includes different data. In example one (including example one A and example one B), the first data set includes the first target CSI and the first CSI feedback information; in example two, the first data set includes the first CSI feedback information and the first reconstructed CSI; in example three, the first data set includes the first target CSI, the first CSI feedback information and the first reconstructed CSI. The first data set shown in the above examples can be regarded as different types of data sets. In other words, examples one to three show the processing flow at the terminal side when the first data set is of different types.
[0471] Example four:
[0472] The network side sends a fourth model. The terminal side performs model training of the first model according to the fourth model.
[0473] In a possible design, the fourth model can be a teacher model of the first model, used to train a student model of the first model to obtain the first model. In the present example, the fourth model can be the encoder 1 trained by the network side, which can be used for model training of the encoder 2 at the terminal side.
[0474] The network side can first perform joint training of the encoder 1 and the decoder 1 to obtain a teacher model of the first model. The process of joint training by the network side can be referred to the above example descriptions in combination with FIG. 5 and FIG. 9, and will not be repeated here.
[0475] The network side can send encoder 1 to the terminal side. For example, the network side can send the model file and / or model parameters corresponding to encoder 1 to the terminal side. The terminal side can then train the model for encoder 2 based on the received model file and / or model parameters.
[0476] In one implementation, the terminal side can employ knowledge distillation to train encoder 2 (i.e., the student model) based on encoder 1 (i.e., the teacher model) to obtain encoder 2. That is, the trained student model is used as the first model. This is illustrated below as Example 4A.
[0477] In another implementation, the terminal can train encoder 2 and decoder 2 through knowledge distillation, and then further train encoder 2 to obtain encoder 2. The trained encoder 2 is used as the first model. This is illustrated in Example 4B below.
[0478] The following sections will explain the processing flow on the terminal side, specifically for Examples 4A and 4B.
[0479] Example 4A:
[0480] The processing flow on the terminal side can be seen in Figure 16, and may include the following steps:
[0481] ① The second target CSI (i.e., V) is processed by encoder 1. UE The compressed CSI is obtained by compressing the CSI. The compressed CSI is used as a soft label. An encoder 2 is initialized as a student model. The second target CSI is used as input to encoder 2, which compresses the second target CSI to obtain the compressed CSI (i.e., By training encoder 2, the output of encoder 2 is made... With soft tags It tends to be closer, that is, to make the output of the student model more similar to the output of the teacher model.
[0482] Correspondingly, the performance metric function used for model training of encoder 2 may include one or more of the following:
[0483] Information obtained by compressing the second target CSI using encoder 2 Information obtained by compressing the second target CSI using encoder 1 The NMSE between (i.e., NMSE) );
[0484] Information obtained by compressing the second target CSI using encoder 2 the MSE between the information obtained by compressing the second target CSI through the encoder 1 and the information obtained by compressing the second target CSI through the encoder 2 (i.e., MSE );
[0485] the MAE between the information obtained by compressing the second target CSI through the encoder 1 and the information obtained by compressing the second target CSI through the encoder 2 (i.e., MAE );
[0486] the weighted sum of multiple terms in the NMSE, MSE or MAE between the information obtained by compressing the second target CSI through the encoder 1 and the information obtained by compressing the second target CSI through the encoder 2 (i.e., NMSE MSE or MAE ); Wherein the weighting coefficients of the terms can be predefined or indicated by the network side, which is not limited in the present application.
[0487] Correspondingly, the performance indicator for the model training includes the range that the value of one or more performance measurement functions should satisfy. That is, the performance indicator includes the range that one or more of the following should satisfy: NMSE MSE MAE or the weighted sum of two or more of the above multiple terms.
[0488] When the performance indicator is satisfied when the output by the encoder 2 and the soft label , the model training can be stopped, thereby obtaining the trained encoder 2.
[0489] For more detailed description of the performance measurement function and the performance indicator, please refer to the relevant description in ② of Example One, except that the training of the student model encoder 2 in Example Four A is the model training of the second target CSI at the terminal side, and the label for the model training is the information obtained by compressing the second target CSI through the teacher model encoder 1, which is not described again. The cases of satisfying and not satisfying the performance indicator have been described in detail in Examples One and Two above in combination with multiple examples, please refer to the relevant description above, which is not described again.
[0490] ②, deploy the trained encoder 2 to a specific device to perform model inference.
[0491] Similar to the above examples one to three, the encoder 1 and the encoder 2 can also include a quantizer with a quantization function. Since the encoder 1 in the present example is sent by the network side, the quantizer can be considered to be sent by the network side. In this case, in the process shown in ① above, the encoder 1 and the encoder 2 output compressed and quantized CSI feedback information.
[0492] In addition, if the encoder 2 includes a quantizer, in the process shown in ② above, the encoder 2 after model deployment can compress and quantize the input target CSI and output CSI feedback information.
[0493] Example four B:
[0494] The processing flow at the terminal side can refer to FIG. 17 and can include the following steps.
[0495] ①-1, compress the second target CSI (i.e., V UE ) through the encoder 1 to obtain compressed CSI (i.e., ). The compressed CSI is used as a soft label. An encoder 2 is initialized as a student model. The second target CSI is input into the encoder 2, which can compress the second target CSI to obtain compressed CSI (i.e., ). Through training of the encoder 2, the output of the encoder 2 tends to be close to the soft label , that is, the output of the student model tends to be closer to the output of the teacher model. The training of the encoder 2 is denoted as training-1 in the figure.
[0496] ①-2, simulate a decoder 2 to model train the encoder 2 and the decoder 2. The encoder 2 and the decoder 2 can be considered as a double-end model. The model training of the encoder 2 and the model training of the decoder 2 can be considered as model training of the double-end model. Through the model training, the encoder 2 can be obtained. The training of the encoder 2 and the decoder 2 is denoted as training-2 in the figure.
[0497] The second target CSI (i.e., V UE ) is input into the encoder 2. The encoder 2 can compress the second target CSI to obtain compressed CSI. The compressed CSI is further input into the decoder 2. The decoder 2 can decompress the compressed CSI to obtain the second reconstructed CSI (i.e., ). The terminal side can train the encoder 2 and the decoder 2 with the second target CSI (i.e., V UE ) as the target. Exemplarily, this process can be referred to as self-supervised learning for V UE .
[0498] Exemplarily, the performance metric function for the model training of the third model can include one or more of the following, for example:
[0499] MSE (i.e., MSE ) between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1;
[0500] NMSE (i.e., NMSE ) between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1;
[0501] MAE (i.e., MAE ) between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1;
[0502] GCS (i.e., GCS ) between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI;
[0503] SGCS (i.e., SGCS ) between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI; or
[0504] a weighted sum of two or more of the following: the MSE between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1, the NMSE between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1, the MAE between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1, the GCS between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI, or the SGCS between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 2 and the second target CSI (i.e., a weighted sum of two or more of the following: MSE NMSE MAE GCS or SGCS ). The weighting coefficients of each item can be predefined or indicated by the network side, which is not limited in the present application.
[0505] The processing of the second target CSI by the encoder 2 and the decoder 2 includes that the encoder 2 compresses the second target CSI to obtain compressed CSI, and the decoder 2 decompresses the compressed CSI to obtain reconstructed CSI. The information obtained by the encoder 2 compressing the second target CSI can be the compressed CSI, and the information obtained by the encoder 1 compressing the second target CSI can also be the compressed CSI.
[0506] Correspondingly, the performance indicator for the model training includes a range that the value of one or more performance measurement functions should satisfy. That is, the performance indicator includes a range that one or more of the following should satisfy: MSE NMSE MAE GCS SGCS or a weighted sum of two or more of the above.
[0507] When the performance indicator is satisfied, the model training can be stopped, and thus the trained encoder 2 can be obtained. UE
[0508] For more detailed descriptions of the performance measurement functions and the performance indicator, please refer to the related descriptions in ②-2 of Example Two, which will not be repeated here. The cases of satisfying and not satisfying the performance indicator have been described in detail in Example One and Example Two above in combination with multiple examples, please refer to the related descriptions above, which will not be repeated here.
[0509] ②, deploy the trained encoder 2 to a specific device to perform model inference.
[0510] Similar to Example Four A, the encoder 1 can include a quantizer with quantization function. Correspondingly, the encoder 2 also includes a quantizer with quantization function. Since the encoder 1 is sent by the network side, the quantizer can be regarded as being sent by the network side. In this case, in the process shown in ①-1 above, the CSI feedback information output by the encoder 1 and the encoder 2 is the CSI feedback information obtained by compressing and quantizing the second target CSI; in the process shown in ①-2, the CSI feedback information output by the encoder 2 is also the CSI feedback information obtained by compressing and quantizing the second target CSI.
[0511] Correspondingly, the decoder 2 can also include a dequantizer with dequantization function. The dequantizer can correspond to the quantizer in the encoder 1. In this case, in the process shown in ①-2 above, the decoder 2 can dequantize and decompress the CSI feedback information to obtain the reconstructed CSI.
[0512] Further, if the encoder 2 includes a quantizer, in the flow shown in the above ②, the encoder 2 after model deployment can compress and quantize the input target CSI, and output the CSI feedback information.
[0513] Example Five:
[0514] The network side sends a fourth model. The fourth model can be used to compress the input target CSI to obtain compressed CSI, which can be used as a data set for model training. In this example, the fourth model sent by the network side can be an encoder 1 trained by the network side. The second target CSI is compressed by the encoder 1 to obtain compressed CSI, which can be used as a data set to train the encoder 2. The data set is a data set generated by the terminal side, which is different from the first data set sent by the network side, and can be distinguished as a second data set. The terminal side can perform model training of the encoder 2 according to the second data set.
[0515] Exemplarily, the network side can first perform joint training of the encoder 1 and the decoder 1 to obtain the encoder 1 (i.e., an example of the fourth model). The process of joint training by the network side can refer to the exemplary description in the above in conjunction with FIG. 5 and FIG. 9, and will not be described in detail.
[0516] The network side can send the encoder 1 to the terminal side. Exemplarily, the network side can send the model file and / or model parameters corresponding to the encoder 1 to the terminal side.
[0517] The processing flow of the terminal side can refer to FIG. 18, which can include the following steps:
[0518] ①, according to the received model file and / or model parameter, the second target CSI (i.e., V UE ) is compressed by the encoder 1 to obtain the compressed CSI (i.e., ). The compressed CSI is an example of the second data set, which can be used for model training of the encoder 2 (i.e., an example of the first model).
[0519] ②, the second target CSI (i.e., V UE ) is used as the input of the encoder 2, and the compressed CSI (i.e., ) obtained by compressing the second target CSI by the encoder 1 is used as the label to train the model of the encoder 2. The encoder 2 can compress the input second target CSI to obtain the compressed CSI (i.e., ). Through the training of the encoder 2, the output of the encoder 2 tends to be close to the label . Exemplarily, the model training of the first model can also be called training for supervised learning.
[0520] Exemplarily, the performance metric function used for the model training may, for example, include one or more of the following:
[0521] the NMSE between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1 (i.e., NMSE );
[0522] the MSE between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1 (i.e., MSE );
[0523] the MAE between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1 (i.e., MAE ); or
[0524] a weighted sum of one or more of the NMSE, the MSE or the MAE between the information obtained by compressing the second target CSI by the encoder 2 and the information obtained by compressing the second target CSI by the encoder 1 (i.e., a weighted sum of two or more of NMSE MSE or MAE ). The weighting coefficients of the respective terms may, for example, be predefined or indicated by the network side, which is not limited in the present application.
[0525] Correspondingly, the performance indicator used for the model training includes a range that the value of one or more performance metric functions should satisfy. That is, the performance indicator includes a range that one or more of the following should satisfy: NMSE MSE MAE or a weighted sum of two or more of the above.
[0526] When the performance indicator is satisfied, the model training can be stopped, and thus the trained encoder 2 can be obtained.
[0527] More details about the process, performance metric function and performance indicator of ② can refer to the relevant description in ② of Example One, except that the input of the encoder 2 in Example Five is the second target CSI from the terminal side (i.e., V UE ), and the label for model training is the compressed CSI of the second dataset from the terminal side (i.e., ). Details are not repeated. The cases of meeting and not meeting the performance indicator have been described in detail in Example One and Example Two above in combination with multiple examples. Please refer to the relevant description above, and details are not repeated.
[0528] ③, deploy the trained encoder 2 to a specific device to perform model inference.
[0529] Similar to Examples One to Four above, the encoder 1 and the encoder 2 can also include a quantizer with quantization function. Since the encoder 1 in this example is sent from the network side, the quantizer can be considered to be sent from the network side. In this case, in the process shown in ① above, the encoder 1 outputs the second CSI feedback information after compressing and quantizing the second target CSI. Therefore, the second dataset includes the second CSI feedback information obtained by the encoder 1 compressing and quantizing the second target CSI. In the process shown in ② above, the encoder 2 outputs the CSI feedback information after compressing and quantizing the second target CSI. The model training of the encoder 2 can be exemplarily called supervised learning with the second CSI feedback information as the label.
[0530] In addition, if the encoder 2 includes a quantizer, in the process shown in ③ above, the encoder 2 after model deployment can compress and quantize the input target CSI and output the CSI feedback information.
[0531] Example Six:
[0532] The network side sends a second model, which is used to decompress the input compressed CSI to obtain the reconstructed CSI. Therefore, the second model can be used for training the terminal side third model to obtain the first model. In this example, the second model sent from the network side can be a decoder 1 trained by the network side, which can be used in conjunction with the encoder 2 to perform model training of the encoder 2.
[0533] Exemplarily, the network side can first perform joint training of the encoder 1 and the decoder 1 to obtain the decoder 1. The process of joint training by the network side can refer to the exemplary description above in combination with FIG. 5 and FIG. 9, and details are not repeated.
[0534] The network side can send the decoder 1 to the terminal side. Exemplarily, the network side can send the model file and / or model parameters corresponding to the decoder 1 to the terminal side.
[0535] The processing flow at the terminal side can refer to FIG. 19, and can include the following steps:
[0536] ①, solidify the decoder 1 according to the received model file and / or model parameters of the decoder 1. Take the second target CSI (i.e., V UE ) at the terminal side as the input of the encoder 2. The encoder 2 can compress the second target CSI to obtain compressed CSI, which can be taken as the input of the decoder 1. The decoder 1 can decompress the compressed CSI to obtain the reconstructed CSI (i.e., ). In the present example, it is not identified by 1 or 2, because the is output in the cooperation of the encoder 2 and the decoder, and it is not accurate enough to define it as the output of the terminal side model or the output of the network side model.
[0537] The terminal side can train the encoder 2 and the decoder 2 with the second target CSI (i.e., V UE ) as the target. Exemplarily, the model training of the third model can also be called self-supervised learning for V UE .
[0538] Exemplarily, the performance measurement function for the model training of the third model can include one or more of the following:
[0539] The SGCS (i.e., SGCS ) between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 1 and the second target CSI;
[0540] The GCS (i.e., GCS ) between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 1 and the second target CSI; or
[0541] The weighted sum of the SGCS and the GCS (i.e., the weighted sum of SGCS and GCS ) between the reconstructed CSI obtained by processing the second target CSI by the encoder 2 and the decoder 1 and the second target CSI. The weighting coefficients of each item can be predefined or indicated by the network side, which is not limited in the present application.
[0542] Wherein, the processing of the second target CSI by the encoder 2 and the decoder 1 includes: the encoder 2 compresses the second target CSI to obtain compressed CSI; the decoder 1 decompresses the compressed CSI to obtain the reconstructed CSI.
[0543] Correspondingly, the performance indicator for the model training includes a range that the value of one or more performance measurement functions should satisfy. That is, the performance indicator includes a range that one or more of the following should satisfy: SGCS GCS or a weighted sum of the above two.
[0544] When the output of the decoder 1 satisfies the performance indicator, the model training can be stopped, and thus the trained encoder 2 can be obtained. and the V UE When the performance indicator is satisfied, the model training can be stopped, and thus the trained encoder 2 can be obtained.
[0545] The processing flow related to ① can refer to the related description in ② of Example 2, and more detailed descriptions of the performance measurement function and the performance indicator can refer to the related description in ② of Example 1, which will not be repeated. The cases of satisfying and not satisfying the performance indicator have been described in detail in Example 1 and Example 2 above in combination with multiple examples, which can refer to the related description above, and will not be repeated.
[0546] ②, the trained encoder 2 can be deployed to a specific device to perform model inference.
[0547] Similar to Example 2 above, the decoder 1 can also include a dequantizer with a dequantization function. Since the decoder 1 in this example is sent by the network side, the dequantizer can be regarded as being sent by the network side.
[0548] Correspondingly, the encoder 2 can include a quantizer with a quantization function. The quantizer can correspond to the dequantizer in the decoder 1. In this case, in the flow shown in ① above, the encoder 2 can compress and quantize the second target CSI to obtain the CSI feedback information. The CSI feedback information is input into the decoder 1, and after dequantization and decompression, the reconstructed CSI is obtained.
[0549] In addition, if the encoder 2 includes a quantizer, in the flow shown in ② above, the encoder 2 after model deployment can compress and quantize the input target CSI to output the CSI feedback information.
[0550] The above describes the model training based on different data sets and the model training according to different models in combination with multiple figures. These examples are only shown for easy understanding, and should not constitute any limitation on the present application. Based on the examples provided above, those skilled in the art can also make simple combinations and / or transformations to obtain more possible embodiments.
[0551] For example, in Example Six above, the network side can send the first dataset and the second model (i.e., the decoder 1) to the terminal side, where the first dataset can include the first target CSI and the first reconstructed CSI. The second model can be the decoder 1. The terminal side can perform model training of the third model according to the second model and the first dataset. The third model includes the first model and the second model, which in this example can include the encoder 2 and the decoder 1. In the training process, the first target CSI can be used as the input of the third model (more specifically, the encoder 2 in the third model), and the first reconstructed CSI can be used as the label to perform model training of the third model. Exemplarily, this process can also be referred to as supervised learning for the first reconstructed CSI.
[0552] Correspondingly, the performance metric function for model training can include one or more of the following:
[0553] the NMSE (i.e., NMSE ) between the reconstructed CSI obtained by processing the first target CSI through the encoder 2 and the decoder 1 and the first reconstructed CSI;
[0554] the MSE (i.e., MSE ) between the reconstructed CSI obtained by processing the first target CSI through the encoder 2 and the decoder 1 and the first reconstructed CSI;
[0555] the MAE (i.e., MAE ) between the reconstructed CSI obtained by processing the first target CSI through the encoder 2 and the decoder 1 and the first reconstructed CSI;
[0556] the GCS (i.e., GCS ) between the reconstructed CSI obtained by processing the first target CSI through the encoder 2 and the decoder 1 and the first reconstructed CSI;
[0557] the SGCS (i.e., SGCS ) between the reconstructed CSI obtained by processing the first target CSI through the encoder 2 and the decoder 1 and the first reconstructed CSI; or
[0558] a weighted sum of two or more of the NMSE, the MSE, the MAE, the SGCS, or the GCS (i.e., NMSE MSE MAE SGCS or GCS a weighted sum of two or more of the above. The weighting coefficients of the respective terms can be predefined or indicated by the network side, which is not limited in the present application.
[0559] The performance indicator for the model training includes a range that the one or more performance measurement functions should satisfy. That is, the performance indicator includes a range that the one or more of the following should satisfy: NMSE MSE MAE GCS SGCS or a weighted sum of two or more of the above.
[0560] For more detailed description of the performance measurement function and the performance indicator, please refer to the relevant description of the respective examples above, which will not be repeated here.
[0561] It should be understood that in the above-mentioned Examples 4 to 6, the model sent by the network side is different. In Examples 4 and 5, the model sent by the network side is the fourth model, which is used to compress (or compress and quantize) the target CSI; in Example 6, the model sent by the network side is the second model, which is used to decompress the compressed CSI (or to dequantize and decompress the CSI feedback information). The models shown in the above examples can be regarded as different types of models. In other words, Examples 4 to 6 show the processing flow at the terminal side in the case where the network side sends different types of models.
[0562] It should also be understood that in the above-mentioned embodiments shown in FIGS. 10 to 19, the terminal side can be a terminal device, and the network side can also be a network device. In another design, the terminal side can also include a terminal device and a host or cloud server of an OTT system; the network side can include a network device and an intelligent network element. In this case, the devices at the terminal side can also communicate with each other, and the devices at the network side can also communicate with each other. The specific implementation process of the method provided in the present application will be described below by means of the method 2000 shown in FIG. 20 in an architecture in which a host or cloud server of an OTT system (hereinafter referred to as an OTT system server), a terminal device, a network device, and an intelligent network element are deployed.
[0563] FIG. 20 is another schematic flow chart of a communication method according to an embodiment of the present application. The method 2000 shown in FIG. 20 can include the following steps:
[0564] In step 2010, the terminal device sends capability information to the network device, the capability information indicating the computing capability supported by the terminal side. In the present example, the device performing the model training is the OTT system server, and the computing capability supported by the terminal side can refer to the computing capability supported by the OTT system server.
[0565] In step 2021, the intelligent network element sends the data set and / or the model to the network device, where the data set and / or the model are used for model training.
[0566] In step 2022, the network device sends the data set and / or the model to the terminal device.
[0567] In step 2023, the terminal device sends the data set and / or the model to the OTT system server.
[0568] Steps 2021-2023 show the process of the intelligent network element transmitting the data set and / or the model through the air interface between the network device and the terminal device. In another implementation manner, the intelligent network element can send the data set and / or the model to the OTT system server through wired transmission. As shown in step 2024 in the figure by a dashed line.
[0569] In step 2030, the network device sends first information to the terminal device, where the first information is used to indicate the performance indicator.
[0570] In step 2040, the terminal device sends information indicating the performance indicator to the OTT system server. The information indicating the performance indicator sent by the terminal device to the OTT system server can be the first information or other information, which is not limited.
[0571] In step 2050, the OTT system server performs model training according to the obtained data set and / or model.
[0572] In step 2060, the OTT system server performs an inference task of the first model in a case where the performance indicator is met.
[0573] In step 2070, the OTT system server sends information indicating that the performance indicator cannot be met to the terminal device in a case where the performance indicator cannot be met.
[0574] In step 2080, the terminal device sends second information to the network device, where the second information can be used to perform one or more of the following: indicating that the performance indicator cannot be met; requesting to replace the performance indicator; requesting to replace the data set; requesting to replace the model; or requesting to close the model training.
[0575] Optionally, the network device can also send the second information to the intelligent network element, or send indication information to the intelligent network element according to the second information, such as indicating the intelligent network element to replace the data set, or indicating the intelligent network element to replace the model, etc., which is not limited.
[0576] The detailed description of each step in the method 2000 can refer to the description of the related step in the method 1000 of FIG. 10, which will not be repeated. In addition, the technical solutions shown in the method 2000 correspond to the technical solutions shown in the method 1000, and thus the beneficial effects achieved are similar, which will not be repeated.
[0577] It should be understood that the flow listed above is only an example, and the present application includes but is not limited to this. For example, the terminal device and the OTT system server in FIG. 20 can also be replaced by a terminal device, in which case the communication between the terminal device and the OTT system is internal communication of the device; or the network device and the intelligent network element in FIG. 20 can also be replaced by a network device, in which case the communication between the network device and the intelligent network element is internal communication of the device. For brevity, no further description will be given in the drawings.
[0578] It should be understood that in each of the embodiments shown in the above multiple drawings, the size of the serial number of each step does not mean the order of execution, and the execution order of each step should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0579] In various embodiments of the present application, the terms and / or descriptions of different embodiments are consistent and can be referred to each other if there is no special description and logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0580] In some of the above embodiments, the devices in the current network architecture are mainly taken as examples for illustrative description (such as terminal devices, network devices, OTT system servers, intelligent network elements, etc.), and the specific form of the device is not limited in the embodiments of the present application. For example, devices that can achieve the same function in the future can also be applicable to the methods provided by the embodiments of the present application.
[0581] It can be understood that the methods and operations implemented by the devices (such as terminal devices, network devices, OTT system servers, intelligent network elements) in each of the above method embodiments can also be implemented by components (such as chips or circuits) of the devices.
[0582] The above provides a detailed description of the method provided by the embodiments of the present application in combination with multiple drawings. The following describes the apparatus provided by the embodiments of the present application in combination with the drawings.
[0583] FIGS. 21 and 22 are schematic block diagrams of possible apparatuses provided by the embodiments of the present application. These apparatuses can be used to implement the functions of the terminal side or the network side in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments.
[0584] FIG. 21 is a schematic block diagram of an apparatus provided in an embodiment of the present application. The apparatus 3000 shown in FIG. 21 can include a processing module 3010 and a communication module 3020.
[0585] In a possible design, the apparatus 3000 can be configured to implement the communication method performed by a terminal side in any of the embodiments shown in FIGS. 10 to 20. For example, the processing module 3010 can be configured to implement the steps of acquiring a data set and / or a model, acquiring a performance indicator, performing model training, determining whether the performance indicator can be satisfied, performing an inference task of a first model, and the like, which are performed by the terminal side in each method embodiment. The communication module 3020 can be configured to implement the steps of sending and / or receiving, and the like, performed by the terminal side in each method embodiment, such as one or more of receiving a data set and / or a model, receiving first information, sending second information, or sending capability information.
[0586] For example, the processing module 3010 can be configured to acquire a data set and / or a model, where the data set and / or the model are used for model training. The processing module 3010 can also be configured to acquire a performance indicator, where the performance indicator corresponds to the acquired data set and / or the acquired model, and the performance indicator corresponds to an inference task of a first model, where the first model is obtained based on model training of the data set and / or the model, and the inference task is an inference of compressing CSI.
[0587] Optionally, the communication module 3020 can be configured to receive first information, where the first information is used to indicate the performance indicator.
[0588] Optionally, the communication module 3020 can be configured to receive the data set and / or the model.
[0589] Optionally, when the communication module 3020 is configured to receive the data set and / or the model, the communication module 3020 can be specifically configured to receive a first data set. The first data set can include one or more of a first target CSI, first CSI feedback information, or a first reconstructed CSI.
[0590] Optionally, when the communication module 3020 is configured to receive the data set and / or the model, the communication module 3020 can be specifically configured to receive a fourth model, where the fourth model is used for model training of the first model.
[0591] Optionally, when the communication module 3020 is configured to receive the data set and / or the model, the communication module 3020 can be specifically configured to receive a second model, where the second model is used for model training of a third model to obtain the first model, and the third model includes the first model.
[0592] Optionally, the processing module 3010 can also be configured to perform the inference task of the first model when the performance indicator can be satisfied.
[0593] Optionally, the communication module 3020 can also be configured to send second information in a case where the performance indicator cannot be satisfied, the second information being used for one or more of the following: indicating that the performance indicator cannot be satisfied; requesting a change of the performance indicator; requesting a change of the data set; requesting a change of the model; or requesting a shutdown of the model training.
[0594] Optionally, the communication module 3020 can also be configured to send capability information, the capability information being used to indicate a computing capability supported by the terminal side.
[0595] More detailed descriptions of the processing module 3010 and the communication module 3020 can be directly obtained by referring to the related descriptions in the method embodiments shown in FIGS. 10-20, which will not be repeated here.
[0596] In another possible design, the apparatus 3000 can be configured to implement the communication method performed by the network side in any one of the embodiments shown in FIGS. 10-20. For example, the processing module 3010 can be configured to implement the steps of generating a data set and / or a model and the like related to processing performed by the network side in the method embodiments; and the communication module 3020 can be configured to implement the steps of sending and / or receiving and the like performed by the network side in the method embodiments, such as one or more of sending a data set and / or a model, sending first information, receiving second information, or receiving capability information.
[0597] For example, the communication module 3020 can be configured to send a data set and / or a model, the data set and / or the model being used for model training; and the communication module 3020 can also be configured to send a performance indicator corresponding to the obtained data set and / or the obtained model, the performance indicator corresponding to an inference task of a first model, the first model being obtained by model training based on the data set and / or the model, and the inference task being an inference of compressing CSI.
[0598] The first information is received, the first information being used to indicate a performance indicator.
[0599] Optionally, the communication module 3020 can be configured to send the data set and / or the model.
[0600] Optionally, when the communication module 3020 is configured to send a data set and / or a model, the communication module 3020 can be specifically configured to send a first data set. The first data set includes one or more of the following: a first target CSI, first CSI feedback information, or a first reconstructed CSI.
[0601] Optionally, when the communication module 3020 is configured to send a data set and / or a model, the communication module 3020 can be specifically configured to send a fourth model, the fourth model being used for model training of the first model.
[0602] Optionally, the communication module 3020 is configured to send the data set and / or the model, and specifically configured to send the second model, which is used for model training of the third model to obtain the first model, the third model comprising the first model.
[0603] Optionally, the communication module 3020 is further configured to receive second information in the case that the performance indicator cannot be met, the second information being used for one or more of the following: indicating that the performance indicator cannot be met; requesting to change the performance indicator; requesting to change the data set; requesting to change the model; or requesting to close the model training.
[0604] Optionally, the communication module 3020 is further configured to receive capability information, the capability information being used to indicate the computing capability supported by the terminal side.
[0605] Optionally, the processing module 3010 is further configured to determine the data set and / or the model according to the capability information.
[0606] Optionally, the processing module 3010 is further configured to determine the performance indicator according to the capability information.
[0607] For more details of the processing module 3010 and the communication module 3020, refer to the related description in the method embodiments shown in FIGS. 10-20, which will not be repeated here.
[0608] It should be noted that the communication module can also be referred to as a transceiver module, a transceiver unit, a transceiver, a transceiver device, or the like. The processing module can also be referred to as a processor, a processing unit, or the like. Optionally, the communication module is configured to perform the sending operation and the receiving operation of the first communication device or the second communication device in the above method. The device in the communication module for realizing the receiving function can be regarded as a receiving module, and the device in the communication module for realizing the sending function can be regarded as a sending module. That is, the communication module can include a receiving module and a sending module.
[0609] It should be further noted that in a possible design, the foregoing processing module and / or communication module can be implemented through a virtual module. For example, the processing module can be implemented through a software function unit or a virtual device, and the communication module can be implemented through a software function or a virtual device. In another possible design, the processing module or the communication module can also be implemented through an entity device. For example, if the device is implemented by using a chip / chip circuit, the communication module can be an input / output circuit and / or a communication interface, performing an input operation (corresponding to the foregoing receiving operation) and an output operation (corresponding to the foregoing sending operation). The processing module can be an integrated processor or a microprocessor or an integrated circuit.
[0610] The division of the modules in the embodiments of the present application is illustrative, and is merely a logical function division. In actual implementation, another division manner can be used. In addition, each function module in each example in the embodiments of the present application can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be realized in the form of hardware or in the form of a software function module.
[0611] Fig. 22 is a structural schematic diagram of a communication apparatus provided by another embodiment of the present application. As shown in Fig. 22, the apparatus 4000 includes a processing circuit 4010 and a communication circuit 4020. The processing circuit 4010 and the communication circuit 4020 are coupled with each other.
[0612] It can be understood that the processing circuit 4010 can be one or more processors, or can be all or part of the processing function circuit in the one or more processors.
[0613] It can be understood that the communication circuit 4020 can be a transceiver or an input / output interface.
[0614] Optionally, the apparatus 4000 can further include a memory 4030 for storing instructions executed by the processing circuit 4010 or storing input data required by the processing circuit 4010 for running instructions or storing data generated after the processing circuit 4010 runs instructions.
[0615] It can be understood that the memory 4030 can be located outside the processing circuit 4010, or located inside the processing circuit 4010.
[0616] As an example, the processing circuit 4010 is configured to implement the functions of the processing module 3010 described above, and the communication circuit 4020 is configured to implement the functions of the communication module 3020 described above.
[0617] As an example, the apparatus 4000 can be a communication device, or can be a chip applied to a communication device.
[0618] When the apparatus 4000 is a communication device, the communication circuit can be a transceiver; when the apparatus 4000 is a chip, the communication circuit can be an input / output circuit, a bus, a pin or other types of communication interfaces, wherein the input circuit in the input / output circuit can be used for receiving, and the output interface can be used for transmitting.
[0619] The present application also provides a computer program product, which, when running on a processor, can implement the communication method performed by a terminal side or the communication method performed by a network side in the method embodiment described above.
[0620] The application further provides a computer readable storage medium, which comprises computer instructions, and the computer instructions can implement the communication method executed by the terminal side or the communication method executed by the network side in the method embodiments when running on a processor.
[0621] The application further provides a communication system, which comprises the terminal side and the network side described above, the terminal side can be used to implement the communication method executed by the terminal side in the method embodiments, and the network side can be used to implement the communication method executed by the network side in the method embodiments.
[0622] It can be understood that the processor in the embodiments of the application can be all or part of the circuit of the following devices or the following devices for processing functions: a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), field programmable gate arrays (FPGAs) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.
[0623] The terms "unit", "module" and the like used in the specification can be used to represent computer-related entities, hardware, combinations of hardware and software, software, or software in execution.
[0624] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.
[0625] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0626] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the division of the above-described device embodiment is only a logical function division, and there can be another division manner for actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different functions can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0627] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0628] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.
[0629] In the above embodiments, the functions of the various functional units can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, the software can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the whole or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as digital video disc (DVD)), or semiconductor media (such as solid state disk (SSD)) and the like.
[0630] The functions, if implemented in the form of software functional units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art or the part of the technical solutions of the present application can be embodied in the form of software product, which is stored in a storage medium and includes a plurality of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and various media that can store program codes.
[0631] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A communication method characterized by comprising: The method comprises: obtaining a data set and / or a model, the data set and / or the model being used for model training; obtaining a performance indicator, the performance indicator corresponding to the obtained data set and / or the obtained model; the performance indicator corresponds to an inference task of a first model, the first model being obtained based on the model training based on the data set and / or the model, and the inference task being inference of compressing channel state information (CSI).
2. The method of claim 1, wherein, The performance indicator is obtained by: receiving first information, the first information being used to indicate the performance indicator.
3. The method of claim 1 or 2, wherein, The data set and / or the model are obtained by: obtaining a first data set, the first data set comprising one or more of the following data: first target CSI, first CSI feedback information, or first reconstructed CSI; wherein the first CSI feedback information is obtained by compressing and quantizing the first target CSI, and the first reconstructed CSI is obtained by dequantizing and decompressing the first CSI feedback information.
4. The method of claim 3, wherein, The first data set comprises the first target CSI and the first CSI feedback information.
5. The method of claim 4, wherein, The performance indicator corresponds to an inference task of a first model, comprising: the performance indicator is a performance indicator that should be met when performing the inference task of the first model.
6. The method of claim 4 or 5, wherein, The first target CSI and the first CSI feedback information are used for model training of the first model, and the performance indicator comprises one or more of the following ranges that should be met: a mean square error (MSE) between information obtained by compressing the first target CSI by the first model and information obtained by dequantizing the first CSI feedback information; a normalized mean square error (NMSE) between information obtained by compressing the first target CSI by the first model and information obtained by dequantizing the first CSI feedback information; a mean absolute error (MAE) between information obtained by compressing the first target CSI by the first model and information obtained by dequantizing the first CSI feedback information; or a weighted sum of a plurality of the following: the MSE, the NMSE, and the MAE, between information obtained by compressing the first target CSI by the first model and information obtained by dequantizing the first CSI feedback information. The first target CSI and the first CSI feedback information are used for model training of a second model, the second model being used for model training of a third model to obtain the first model, and the third model comprising the first model.
7. The method of claim 4, wherein, The third model is used to compress target CSI to obtain compressed CSI, and is also used to decompress information obtained by dequantizing CSI feedback information to obtain reconstructed CSI, and the compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information.
8. The method of claim 7, wherein, The performance indicator corresponds to an inference task of a first model, comprising: the performance indicator is a performance requirement that should be met when compressing and reconstructing second target CSI by the third model.
9. The method of claim 7 or 8, wherein, The performance indicator comprises one or more of the following ranges that should be met:
10. The method of any one of claims 7 to 9, wherein, MSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI; NMSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI; MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI; MSE between the reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; NMSE between the reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; MAE between the reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; GCS between the reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; SGCS between the reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; or a weighted sum of multiple items of the MSE, the NMSE, the MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI, the MSE, the NMSE, the MAE, the GCS or the SGCS between the reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI.
11. The method of claim 3, wherein, The first data set includes the first CSI feedback information and the first reconstructed CSI; the first CSI feedback information and the first reconstructed CSI are used for model training of a second model, the second model is used for model training of a third model to obtain the first model; and the third model includes the first model.
12. The method of claim 11, wherein, The second model is used for decompressing information obtained by dequantizing the first CSI feedback information; the third model is used for compressing a target CSI to obtain compressed CSI, and is used for decompressing information obtained by dequantizing CSI feedback information to obtain reconstructed CSI, the compressed CSI corresponding to the information obtained by dequantizing the CSI feedback information.
13. The method of claim 11 or 12, wherein, The performance indicator corresponds to an inference task of the first model, and includes: the performance indicator is a performance indicator that should be met by compression and reconstruction of a second target CSI through the third model.
14. The method of any one of claims 11 to 13, wherein, The performance indicator includes a range that should be met by one or more of the following: MSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; a NMSE between a reconstructed CSI obtained by decompressing information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; an MAE between a reconstructed CSI obtained by decompressing information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; an MSE between a reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; an NMSE between a reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; an MAE between a reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; a GCS between a reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; an SGCS between a reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; or a weighted sum of multiple ones of the MSE, the NMSE, the MAE between a reconstructed CSI obtained by decompressing information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI, the MSE, the NMSE, the MAE, the GCS or the SGCS between a reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI.
15. The method of any one of claims 11 to 13, wherein, the performance indicators include one or more ranges to be satisfied: an MSE between a reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; an NMSE between a reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; an MAE between a reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; a GCS between a reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; an SGCS between a reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; or a weighted sum of multiple ones of the MSE, the NMSE, the MAE, the GCS or the SGCS between a reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI.
16. The method of claim 3, wherein, the first data set includes the first target CSI, the first CSI feedback information and the first reconstructed CSI; wherein the first target CSI and the first reconstructed CSI are used for model training of a third model to obtain the first model, and the third model includes the first model.
17. The method of claim 16, wherein, The third model is used for compressing the first target CSI to obtain compressed CSI, and is used for decompressing information obtained by dequantizing CSI feedback information to obtain reconstructed CSI, the compressed CSI corresponding to the information obtained by dequantizing the CSI feedback information.
18. The method of claim 16 or 17, wherein, The performance index corresponds to an inference task of the first model, and includes: the performance index is a requirement to be met by compression and reconstruction of the first target CSI by the third model.
19. The method of any one of claims 16 to 18, wherein, The performance index includes one or more ranges to be met: MSE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; NMSE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; MAE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; GCS between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; SGCS between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; or A weighted sum of multiple items of MSE, NMSE, MAE, GCS or SGCS between reconstructed CSI obtained by inference of the first target CSI by the third model and the first reconstructed CSI.
20. The method of claim 1 or 2, wherein, The data set and / or model are obtained, including: obtaining a fourth model, the fourth model being used for model training of the first model.
21. The method of claim 20, wherein, The fourth model is a teacher model of the first model, used for model training of a student model of the first model to obtain the first model.
22. The method of claim 20, wherein, The fourth model is used to generate a second data set, the second data set being used for model training of the first model, and the second data set including compressed CSI obtained by compressing a second target CSI.
23. The method of claim 21 or 22, wherein, The performance index corresponds to an inference task of the first model, and includes: the performance index is a performance requirement to be met by performing the inference task of the first model.
24. The method of claim 1 or 2, wherein, The data set and / or model are obtained, including: obtaining a fourth model, the fourth model being used for model training of the third model to obtain the first model, the third model including the first model.
25. The method of claim 24, wherein, The fourth model is used for compressing a second target CSI; the third model is used for compressing the second target CSI to obtain compressed CSI, and is used for dequantizing information obtained by dequantizing CSI feedback information to obtain reconstructed CSI; the compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information.
26. The method of claim 24 or 25, wherein, The performance index corresponds to an inference task of the first model, and includes: the performance index includes: a requirement to be met by performing the inference task of the first model, and / or a performance requirement to be met by compression and reconstruction of the second target CSI by the third model.
27. The method of any one of claims 20 to 26, wherein, The performance index includes one or more ranges to be met: MSE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; NMSE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; MAE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; GCS between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; SGCS between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; or A weighted sum of multiple items of MSE, NMSE, MAE, GCS or SGCS between reconstructed CSI obtained by inference of the first target CSI by the third model and the first reconstructed CSI. MSE between information obtained by compressing a second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; NMSE between information obtained by compressing a second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; MAE between information obtained by compressing a second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; GCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI; SGCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI; or a weighted sum of multiple items of MSE, NMSE, MAE between information obtained by compressing a second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model, GCS or SGCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI.
28. The method of claim 1 or 2, wherein, The obtaining the dataset or the model comprises: obtaining a second model, the second model being used for model training of a third model to obtain the first model, the third model comprising the first model.
29. The method of claim 28, wherein, The second model is used for decompressing compressed CSI; the third model is used for compressing target CSI to obtain compressed CSI, and is used for decompressing information obtained by dequantizing CSI feedback information to obtain reconstructed CSI, the compressed CSI corresponding to the information obtained by dequantizing the CSI feedback information.
30. The method of claim 28 or 29, wherein, The performance indicator corresponds to an inference task of the first model, and the performance indicator is a performance requirement that should be met by compressing and reconstructing the second target CSI by the third model.
31. The method of claim 30, wherein, The performance indicator comprises a range that should be met by one or more of the following: GCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI; SGCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI; or a weighted sum of GCS and SGCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI.
32. The method of any one of claims 1 to 31, wherein, The obtaining the dataset and / or the model comprises: receiving the dataset and / or the model.
33. The method of any one of claims 1 to 32, wherein, The method further comprises: in a case where the performance indicator can be met, performing the inference task of the first model.
34. The method of any one of claims 1 to 33, wherein, The method further comprises: in a case where the performance indicator cannot be met, sending second information, the second information being used for one or more of the following: indicating that the performance indicator cannot be met; requesting to replace the performance indicator; requesting to replace the dataset; requesting to replace the model; or requesting to stop the model training.
35. The method of any one of claims 1 to 34, wherein, Before the receiving the data set and / or the model, the method further comprises: sending capability information, the capability information being used to indicate a computing capability supported at a terminal side.
36. A method of communication, comprising: comprising: sending a data set and / or a model, the data set and / or the model being used for model training; sending first information, the first information being used to indicate a performance indicator, the performance indicator corresponding to the sent data set and / or the sent model; and the performance indicator corresponding to an inference task of a first model, the first model being obtained based on the model training based on the data set and / or the model, the inference task being an inference of compressing channel state information (CSI).
37. The method of claim 36, wherein, The sending data set and / or model, comprising: sending a first data set, the first data set comprising one or more of the following data: a first target CSI, first CSI feedback information, or a first reconstructed CSI; wherein the first CSI feedback information is obtained by compressing and quantizing the first target CSI, and the first reconstructed CSI is obtained by dequantizing and decompressing the first CSI feedback information.
38. The method of claim 37, wherein, The first data set comprises the first target CSI and the first CSI feedback information.
39. The method of claim 38, wherein, The performance indicator corresponding to the inference task of the first model comprises: the performance indicator being a performance indicator that should be met when performing the inference task of the first model.
40. The method of claim 38 or 39, wherein, The first target CSI and the first CSI feedback information are used for model training of the first model, and the performance indicator comprises one or more of the following ranges that should be met: a mean square error (MSE) between information obtained by compressing the first target CSI by the first model and information obtained by dequantizing the first CSI feedback information; a normalized mean square error (NMSE) between information obtained by compressing the first target CSI by the first model and information obtained by dequantizing the first CSI feedback information; a mean absolute error (MAE) between information obtained by compressing the first target CSI by the first model and information obtained by dequantizing the first CSI feedback information; or a weighted sum of a plurality of the following: the MSE, the NMSE, or the MAE between information obtained by compressing the first target CSI by the first model and information obtained by dequantizing the first CSI feedback information. The first target CSI and the first CSI feedback information are used for model training of a second model, the second model being used for model training of a third model to obtain the first model, the third model comprising the first model.
41. The method of claim 38, wherein, The third model is used to compress a target CSI to obtain compressed CSI, and is also used to decompress information obtained by dequantizing CSI feedback information to obtain reconstructed CSI; the compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information.
42. The method of claim 41, wherein, 43. The method of claim 41 or 42, wherein, The performance indicator corresponds to an inference task of the first model, and includes: the performance indicator is a performance requirement to be met by the third model in compressing and reconstructing a second target CSI.
44. The method of any one of claims 41 to 43, wherein, The performance indicator includes one or more of the following ranges to be met: MSE between reconstructed CSI obtained by decompressing information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI; NMSE between reconstructed CSI obtained by decompressing information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI; MAE between reconstructed CSI obtained by decompressing information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI; MSE between reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; NMSE between reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; MAE between reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; GCS between reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; SGCS between reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI; or A weighted sum of multiple ones of the following: MSE between reconstructed CSI obtained by decompressing information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI, NMSE, MAE, MSE between reconstructed CSI obtained by processing the first target CSI through the third model and the first target CSI, NMSE, MAE, GCS, or SGCS.
45. The method of claim 37, wherein, The first data set includes: the first CSI feedback information and the first reconstructed CSI; the first CSI feedback information and the first reconstructed CSI are used for model training of a second model, the second model is used for model training of a third model to obtain the first model; and the third model includes the first model.
46. The method of claim 45, wherein, The second model is used to decompress information obtained by dequantizing the first CSI feedback information; and the third model is used to compress a target CSI to obtain compressed CSI, and is used to decompress information obtained by dequantizing CSI feedback information to obtain reconstructed CSI, the compressed CSI corresponding to the information obtained by dequantizing the CSI feedback information.
47. The method of claim 45 or 46, wherein, The performance indicator corresponds to an inference task of the first model, and includes: the performance indicator is a performance indicator to be met by the third model in compressing and reconstructing a second target CSI.
48. The method of any one of claims 45 to 47, wherein, The performance indicator includes one or more of the following ranges to be met: MSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; NMSE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first reconstructed CSI; MSE between the reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; NMSE between the reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; MAE between the reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; GCS between the reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; SGCS between the reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI; or a weighted sum of multiple ones of the MSE, the NMSE, the MAE between the reconstructed CSI obtained by decompressing the information obtained by dequantizing the first CSI feedback information through the second model and the first target CSI, the MSE, the NMSE, the MAE, the GCS or the SGCS between the reconstructed CSI obtained by processing the second target CSI through the third model and the second target CSI.
49. The method of any one of claims 45 to 47, wherein, The performance indicators include one or more ranges to be satisfied: MSE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; NMSE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; MAE between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; GCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; SGCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI; or a weighted sum of multiple ones of the MSE, the NMSE, the MAE, the GCS or the SGCS between the reconstructed CSI obtained by processing the first reconstructed CSI through the third model and the first reconstructed CSI.
50. The method of claim 37, wherein, The first data set includes the first target CSI, the first CSI feedback information and the first reconstructed CSI. The first target CSI and the first reconstructed CSI are used for model training of a third model to obtain the first model, and the third model includes the first model.
51. The method of claim 50, wherein, The third model is used for compressing the first target CSI to obtain compressed CSI, and is used for decompressing information obtained by dequantizing CSI feedback information to obtain reconstructed CSI, the compressed CSI corresponding to the information obtained by dequantizing the CSI feedback information.
52. The method of claim 50 or 51, wherein, The performance index corresponds to an inference task of the first model, and includes: the performance index is a requirement to be met by compression and reconstruction of the first target CSI by the third model.
53. The method of any one of claims 50 to 52, wherein, The performance index includes one or more ranges to be met: MSE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; NMSE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; MAE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; GCS between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; SGCS between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; or A weighted sum of multiple items of MSE, NMSE, MAE, GCS or SGCS between reconstructed CSI obtained by inference of the first target CSI by the third model and the first reconstructed CSI.
54. The method of claim 36, wherein, The transmitted data set and / or model includes: Transmitting a fourth model, the fourth model being used for model training of the first model.
55. The method of claim 54, wherein, The fourth model is a teacher model of the first model, used for model training of a student model of the first model to obtain the first model.
56. The method of claim 54, wherein, The fourth model is used to generate a second data set, the second data set being used for model training of the first model, and the second data set including compressed CSI obtained by compressing a second target CSI.
57. The method of claim 55 or 56, wherein, The performance index corresponds to an inference task of the first model, and includes: the performance index is a performance requirement to be met by performing the inference task of the first model.
58. The method of claim 36, wherein, The transmitted data set and / or model includes: Transmitting a fourth model, the fourth model being used for model training of the third model to obtain the first model, the third model including the first model.
59. The method of claim 58, wherein, The fourth model is used for compressing a second target CSI; the third model is used for compressing the second target CSI to obtain compressed CSI, and is used for dequantizing information obtained by dequantizing CSI feedback information to obtain reconstructed CSI; the compressed CSI corresponds to the information obtained by dequantizing the CSI feedback information.
60. The method of claim 58 or 59, wherein, The performance index corresponds to an inference task of the first model, and includes: the performance index includes: a requirement to be met by performing the inference task of the first model, and / or a performance requirement to be met by compression and reconstruction of the second target CSI by the third model.
61. The method of any one of claims 54 to 60, wherein, The performance index includes one or more ranges to be met: MSE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; NMSE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; MAE between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; GCS between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; SGCS between reconstructed CSI obtained by processing the first target CSI by the third model and the first reconstructed CSI; or A weighted sum of multiple items of MSE, NMSE, MAE, GCS or SGCS between reconstructed CSI obtained by inference of the first target CSI by the third model and the first reconstructed CSI. MSE between information obtained by compressing a second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; NMSE between information obtained by compressing a second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; MAE between information obtained by compressing a second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model; GCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI; SGCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI; or a weighted sum of multiple items of MSE, NMSE, MAE between information obtained by compressing a second target CSI by the first model and information obtained by compressing the second target CSI by the fourth model, GCS or SGCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI.
62. The method of claim 36, wherein, The sending data set or model comprises: sending a second model, the second model being used for model training of a third model to obtain the first model, the third model comprising the first model.
63. The method of claim 62, wherein, The second model is used for decompression of compressed CSI; the third model is used for compression of target CSI to obtain compressed CSI, and is used for decompression of information obtained by dequantization of CSI feedback information to obtain reconstructed CSI, the compressed CSI corresponding to the information obtained by dequantization of the CSI feedback information.
64. The method of claim 62 or 63, wherein, The performance indicator corresponds to an inference task of the first model, and the performance indicator comprises a performance requirement that should be met by compression and reconstruction of the second target CSI by the third model.
65. The method of claim 64, wherein, The performance indicator comprises a range that should be met by one or more of the following: GCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI; SGCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI; or a weighted sum of GCS and SGCS between reconstructed CSI obtained by processing a second target CSI by the third model and the second target CSI.
66. The method of any one of claims 36 to 65, wherein, The method further comprises: receiving second information, the second information being used for one or more of the following: indicating that the performance indicator cannot be met; requesting to replace the performance indicator; requesting to replace the data set; requesting to replace the model; or requesting to close the model training.
67. The method of any one of claims 36 to 66, wherein, Before the receiving data set and / or model, the method further comprises: sending capability information, the capability information being used to indicate a computing capability supported by a terminal side.
68. A communications device, characterized by comprise a module or unit for implementing the method according to any one of claims 1 to 35, or a module or unit for implementing the method according to any one of claims 36 to 67.
69. A communications device, characterized by comprise one or more processors and a communication circuit configured to at least one of input or output a signal for the communication apparatus; and the one or more processors are configured to implement the method according to any one of claims 1 to 35, or a module or unit for implementing the method according to any one of claims 36 to 67.
70. The device of claim 69, wherein, The communication apparatus is a network device or a terminal device, or a chip for the network device or the terminal device.
Citation Information
Patent Citations
Communication method and device
CN115802370A
Channel state information (CSI) determination method and apparatus, and readable storage medium
WO2024094177A1
Methods, devices and medium for communication
WO2024168517A1
Communication method and communication apparatus
WO2024169757A1