Communication method and apparatus

By acquiring datasets and models on the terminal side and training them based on multiple sets of metrics, the problem of AI model training mismatch between terminal devices and network devices is solved, improving CSI feedback performance and flexibility, and adapting to different business needs and device capabilities.

WO2026098312A1PCT designated stage Publication Date: 2026-05-15HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-10-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In wireless communication systems, there is a mismatch in AI model training between terminal devices and network devices, which leads to CSI feedback performance not meeting requirements.

Method used

Data sets and models are acquired through the terminal side, and the model is trained based on multiple sets of metrics, including performance metrics, training time and resource metrics. Flexible training methods are adopted to improve CSI feedback performance.

Benefits of technology

It enables compliant model training, improves CSI's feedback performance and flexibility, and adapts to different business needs and equipment capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025131112_15052026_PF_FP_ABST
    Figure CN2025131112_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a communication method and apparatus. In the method, a network side can not only send a data set and / or a model used for model training, but can also indicate a plurality of groups of indexes corresponding to the data set and / or the model, the plurality of groups of indexes correspond to an inference task of a first model, and the inference task can be compressing target CSI. A terminal side can perform model training on the basis of one or more groups of indexes among the plurality of groups of indexes, and the data set and / or the model, so as to obtain the first model. After acquiring the data set and / or the model, the terminal side can perform model training on the basis of the corresponding one or more groups of indexes. Therefore, the terminal side can obtain, by means of model training, a model meeting requirements, thereby improving the performance of CSI feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Communication methods and devices

[0001] This application claims priority to Chinese Patent Application No. 202411595237.9, filed with the China National Intellectual Property Administration on November 8, 2024, entitled "Communication Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of wireless communication, and more particularly to a communication method and apparatus. Background Technology

[0003] In wireless communication systems, artificial intelligence (AI) models can be used to compress and reconstruct (or recover) channel state information (CSI). The sender of the CSI report (such as a terminal device) can use the AI ​​model to compress the CSI and send the resulting CSI report. The receiver of the CSI report (such as a network device) can use the AI ​​model to reconstruct the channel measurement results based on the received CSI report.

[0004] To support the integration and development of dual-end AI models (such as models deployed on both the sender and receiver of a CSI report), datasets and / or models need to be provided to the sender and / or receiver for model training. One possible implementation is that the sender provides the dataset and / or model to the receiver for model training. However, the receiver's model training may not match the dataset or model provided by the sender, resulting in a trained model that is unsuitable or even unusable. Summary of the Invention

[0005] This application provides a communication method and apparatus that facilitates the training of models that meet the requirements and improves the feedback performance of CSI.

[0006] Firstly, a communication method is provided that can be applied to the terminal side, for example, it can be executed by a terminal device; or it can be executed by components deployed in the terminal device, such as circuits or chips inside the terminal device (e.g., modem chips, also known as baseband chips, or system-on-chip (SoC) chips or system-in-package (SIP) chips containing modem cores, etc.); or it can be executed by devices deployed outside the terminal device (e.g., the host of an over-the-top (OTT) system or a cloud server) or components within the device (e.g., chips, processors, or circuits inside the device); or it can be implemented by logic modules or software capable of realizing all or part of the functions of the terminal device, etc. Alternatively, the method can be executed by a first device, which can be a terminal device; it can also be a component within the terminal device, such as internal circuitry or a chip (e.g., a modem chip), or a SoC chip or SIP chip containing a modem core; it can also be a device outside the terminal device (e.g., the host of an OTT system or a cloud server) or a component within a device (e.g., internal chips, processors, or circuitry); or it can be a logic module or software capable of implementing some or all of the functionalities of the terminal side, etc. This application does not limit this. For ease of understanding and explanation, the method provided in the first aspect will be described below using the terminal side as an example.

[0007] For example, the method includes: acquiring a dataset and / or a model, the dataset and / or the model being used for model training; training the model based on a first metric and the dataset and / or the model to obtain a first model; the first metric is one or more sets of metrics, the sets of metrics including one or more of the following categories of metrics: performance metrics, metrics of training time required, or metrics of training resources required, the sets of metrics corresponding to the dataset and / or the model, and corresponding to the inference task of the first model, the inference task of the first model being to compress the target CSI.

[0008] The first model can be obtained by training the model on the terminal device or by training it on other devices besides the terminal device; this application does not limit this.

[0009] The dataset can be sent from the network side to the terminal side. The terminal side's acquisition of the dataset and / or model can include: the terminal device on the terminal side receiving the dataset and / or model from the network device on the network side; or, the host or cloud server of the OTT system on the terminal side receiving the dataset and / or model from the intelligent network element on the network side; or, the terminal device on the terminal side obtaining the dataset and / or model received from the intelligent network element from the host or cloud server of the OTT system; or, the host or cloud server of the OTT system on the terminal side obtaining the dataset and / or model received from the network device from the terminal device.

[0010] Based on the above scheme, the first metric not only corresponds to the dataset and / or model used for model training, but also to the inference task of the first model. While acquiring the dataset and / or model used for model training, the terminal can also acquire the corresponding first metric. Therefore, based on the dataset and / or model, and the associated first metric, it can use appropriate training methods to train the model to obtain a model that meets the requirements, thereby improving the feedback performance of CSI.

[0011] Furthermore, each dataset and / or each model can correspond to multiple sets of metrics, which means that the terminal can use various different training methods to train the model. The first metric is one or more of these sets of metrics; therefore, the terminal can use one or more corresponding training methods to train the model based on one or more of these sets of metrics. Thus, the terminal has a high degree of freedom; it can choose appropriate metrics and training methods for model training based on business needs, device capabilities, and other factors.

[0012] Furthermore, the multiple sets of metrics in this application may include one or more of the following categories: performance metrics, metrics of training time, or metrics of training resources. That is, the model training process is monitored from one or more dimensions of performance, time, or resources, so as to flexibly respond to different business needs.

[0013] In conjunction with the first aspect, in some possible implementations of the first aspect, the acquisition of the dataset and / or model includes: receiving the dataset and / or the model.

[0014] In conjunction with the first aspect, in some possible implementations of the first aspect, before training the model based on the first metric, the dataset, and / or the model, the method further includes: receiving first information, the first information being used to indicate the multiple sets of metrics.

[0015] This first piece of information can be sent from the network side to the terminal side. Based on this first piece of information, the terminal side can determine multiple sets of metrics corresponding to the acquired dataset and / or model.

[0016] Optionally, the method further includes: obtaining the multiple sets of indicators based on the first information.

[0017] These multiple sets of indicators can be carried within the first information. The terminal side can obtain these multiple sets of indicators by parsing the first information. By carrying these multiple sets of indicators through the first information, the network side can flexibly adjust the indicators, for example, according to business needs and network conditions.

[0018] In conjunction with the first aspect, in some possible implementations of the first aspect, the first indicator is a performance indicator, at least two of the multiple sets of indicators are the performance indicators, and the first indicator is one or more of the at least two sets of indicators.

[0019] The first metric is the one used for model training on the terminal side. Since this first metric is a performance metric, model training based on the first metric is equivalent to model training based on the performance metric. By training the model based on the performance metric, the resulting first model can achieve higher feedback accuracy, thereby improving the feedback performance of CSI.

[0020] In conjunction with the first aspect, in some possible implementations of the first aspect, the acquisition of the dataset and / or model includes: acquiring a first dataset, the first dataset including: target CSI and CSI feedback information from the network side, wherein the CSI feedback information is obtained by compressing the target CSI.

[0021] The CSI feedback information can be either compressed CSI obtained by compressing the target CSI, or quantized CSI obtained by compressing and quantizing the target CSI. This application does not limit this.

[0022] One possible design is that the first dataset is used to train the second model, and the second model is used to train the first model, wherein the inference task of the second model is to decompress the CSI feedback information.

[0023] Based on the above design for the first dataset, the terminal can first train the second model, and then train the first model based on the second model. This model training method allows performance to be monitored throughout the entire CSI compression and reconstruction process, which is beneficial for improving the performance of CSI feedback.

[0024] Optionally, the at least two sets of indicators correspond to at least two sets of functions, wherein the first set of functions in the at least two sets of functions includes one or more of the following functions: normalized mean square error (NMSE), mean square error (MSE), L1 loss (L1Loss), generalized cosine similarity (GCS), square generalized cosine similarity (SGCS) between the output of the second model and the label, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the label.

[0025] When training the second model, the labels can be target CSIs from the first dataset. When training the first model based on the second model, the input to the first model can be target CSIs from the first dataset or target CSIs from a third dataset obtained by the terminal; this application does not limit this. One possible implementation for training the first model based on the second model is to jointly train the first and second models. The labels for joint training can be the target CSIs input to the first model, i.e., target CSIs from the first dataset or target CSIs from the third dataset.

[0026] Further optionally, the output of the second model includes the output of the second model during training. Accordingly, the label is the target CSI in the first dataset. The first set of functions includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI in the first dataset, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the target CSI in the first dataset.

[0027] Monitoring can be performed on the terminal side while the second model is being trained separately, involving fewer steps and a simpler process. Since the second model is used to train the first model, a set of metrics corresponding to the first function set can be used to indirectly monitor the output of the first model. Therefore, this set of metrics can be said to correspond to the inference task of the first model.

[0028] Further optionally, the output of the second model includes the output of the second model when the first model is trained based on the second model. One possible implementation of training the first model based on the second model is to perform joint training of the first and second models, or end-to-end training. Accordingly, the label of the joint training is the target CSI in the first dataset or the target CSI in the third dataset. The first set of functions includes one or more of the following functions: a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI, where the target CSI is the target CSI in the first dataset or the target CSI in the third dataset.

[0029] On the terminal side, the performance of the first model can be indirectly monitored by monitoring the output of the second model when training the first model based on the second model. Therefore, the set of metrics corresponding to the first function set can be said to correspond to the inference task of the first model.

[0030] Alternatively, the terminal side can monitor both the separate training of the second model and the training of the first model based on the second model, thus enabling more comprehensive performance monitoring. Based on the explanation of the two stages above, it can be seen that the set of metrics corresponding to the first function set corresponds to the inference task of the first model.

[0031] The difference between the output of the second model and the target CSI (i.e., the label) is characterized by a weighted sum of NMSE, MSE, L1Loss, GCS, SGCS, or two or more of these metrics. The terminal can then optimize the parameters of the second model based on this difference to reduce the discrepancy between the second model's output and the target CSI in the first dataset. The first model is then trained using the second model, and through end-to-end training, the parameters of the first model are optimized to further reduce the discrepancy between the second model's output and the target CSI in either the first or third dataset. This results in a highly accurate first model for performing inference tasks. Furthermore, monitoring model training with more performance metrics allows for more comprehensive performance monitoring.

[0032] Another possible design is that the first dataset is used to train the first model.

[0033] The terminal can directly train the first model based on the first dataset. This model training involves fewer steps and has a simpler process, thus allowing the terminal to obtain the first model more quickly through training.

[0034] Optionally, the at least two sets of indicators correspond to at least two sets of functions, and the second set of functions in the at least two sets of functions includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the first model and the CSI feedback information, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS or SGCS between the output of the first model and the CSI feedback information.

[0035] The difference between the output of the first model and the CSI feedback information (i.e., labels) is characterized by a weighted sum of NMSE, MSE, L1Loss, GCS, SGCS, or two or more of the above. The terminal side can optimize the parameters of the first model based on this difference to reduce the discrepancy between the two, thereby obtaining a more accurate first model for performing inference tasks.

[0036] Since the set of metrics corresponding to the second set of functions directly monitors the output of the first model, that is, directly monitors the performance of the first model, it can be said that the set of metrics corresponding to the second set of functions corresponds to the inference task of the first model.

[0037] In conjunction with the first aspect, in some possible implementations of the first aspect, the acquisition of the dataset and / or model includes: acquiring a third model, wherein the inference task of the third model is to compress the target CSI.

[0038] The third model can be a model distributed from the network side. The inference task of the third model is the same as that of the first model. That is, the terminal side can train the model based on the third model distributed from the network side to obtain the first model of the local terminal (i.e., the terminal side).

[0039] The target CSI can be a target CSI from a first dataset on the network side, or a target CSI from a third dataset obtained by the terminal side itself. This application does not limit this.

[0040] One possible design is that the third model is used to train the second model, and the second model is used to train the first model, wherein the inference task of the second model is to decompress the CSI feedback information, which is obtained by the third model by compressing the input target CSI.

[0041] The target CSI input to the third model can come from the first dataset on the network side or from the third dataset obtained by the terminal side itself; this application does not limit this.

[0042] Because the network side has greater computing power, the third model trained on the network side usually has better performance. The terminal side trains the second model using the third model from the network side, and then trains the first model based on the second model. This is beneficial for obtaining a first model with similar performance to the third model.

[0043] The at least two function sets corresponding to the at least two sets of indicators include a third function set, which includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the label, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS or SGCS between the output of the second model and the label.

[0044] Further optionally, the output of the second model includes the output of the second model during training. Accordingly, the label is the target CSI. The third function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the target CSI, wherein the target CSI is the target CSI in the first dataset or the target CSI in the third dataset.

[0045] Monitoring can be performed on the terminal side while the second model is being trained separately, involving fewer steps and a simpler process. Since the second model is used to train the first model, the first metric can be used to indirectly monitor the output of the first model. Therefore, the first metric can be said to correspond to the inference task of the first model.

[0046] Further optionally, the output of the second model includes the output of the second model when training the first model based on the second model. One possible implementation of training the first model based on the second model is to jointly train the first and second models. Accordingly, the label of the joint training is the target CSI in the first dataset or the target CSI in the third dataset. The third function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the target CSI, wherein the target CSI is the target CSI in the first dataset or the target CSI in the third dataset.

[0047] The terminal side can also monitor the performance of the first model while training it based on the second model, allowing direct monitoring of the first model's performance. Therefore, the first metric can be said to correspond to the inference task of the first model.

[0048] Alternatively, the terminal side can monitor both the separate training of the second model and the training of the first model based on the second model, thus enabling more comprehensive performance monitoring. Based on the explanation of the two stages above, it is clear that the first metric corresponds to the inference task of the first model.

[0049] The difference between the output of the second model and the target CSI (i.e., the label) is characterized by a weighted sum of NMSE, MSE, L1Loss, GCS, SGCS, or two or more of these metrics. The terminal can then optimize the parameters of the second model based on this difference to reduce the discrepancy between the second model's output and the target CSI. The first model is then trained using the second model, and through end-to-end training, the parameters of the first model are optimized to further reduce the difference between the second model's output and the target CSI. This results in a highly accurate first model for performing inference tasks. Furthermore, monitoring model training with more performance metrics allows for more comprehensive performance monitoring.

[0050] Optionally, the third model is used to obtain a second dataset, which includes CSI feedback information, and the second dataset is used to train the first model.

[0051] The terminal generates a second dataset using a third model from the network side, and then trains the first model based on this second dataset. This training process involves fewer steps and is simpler, allowing the terminal to obtain the first model more quickly.

[0052] The at least two function sets corresponding to the at least two sets of indicators include a fourth function set, which includes one or more of the following functions: NMSE, MSE, L1Loss between the output of the first model and the CSI feedback information, or a weighted sum of multiple NMSE, MSE, or L1Loss between the output of the first model and the CSI feedback information.

[0053] The difference between the output of the first model and the CSI feedback information (i.e., labels) is characterized by a weighted sum of NMSE, MSE, L1Loss, or two or more of the above. The parameters of the first model are then optimized based on this difference to reduce the difference between the two, thereby obtaining a first model with higher accuracy for performing inference tasks.

[0054] The terminal side can monitor the output of the first model, which means it can directly monitor the performance of the first model. Therefore, the first metric corresponds to the inference task of the first model.

[0055] In conjunction with the first aspect, in some possible implementations of the first aspect, the method further includes: performing an inference task of the first model when the first metric is satisfied.

[0056] If the first metric is met, the inference task of the first model is executed, thereby ensuring the performance of the model on the terminal side, which is beneficial to business needs.

[0057] In conjunction with the first aspect, in some possible implementations of the first aspect, the method further includes: sending second information, the second information being used to indicate that the first indicator has been met.

[0058] The terminal can indicate to the network that the first metric has been met. On the one hand, the network can understand that the trained first model meets the first metric and can be deployed for use in real-world scenarios; on the other hand, the network can continue to monitor the first model deployed on the terminal based on the first metric, which is beneficial for ensuring the performance of CSI feedback in the long term.

[0059] In conjunction with the first aspect, in some possible implementations of the first aspect, the method further includes: sending third information when the first metric is not met, the third information being used for one or more of the following: indicating that the first metric is not met; requesting a change of the dataset; requesting a change of the model; or requesting to stop or pause the model training.

[0060] In this way, the network can respond when the first metric cannot be met, such as shutting down model training, changing performance metrics, changing the model, or changing the dataset, etc. Since the network's response can be made after understanding the situation on the terminal side, the models, datasets, metrics, etc., that are distributed are more compatible with the capabilities of the terminal side, thereby helping to improve the performance of the dual-end models.

[0061] Secondly, a communication method is provided, which can be applied to the network side. For example, it can be executed by a network device; or by a component deployed within the network device, such as internal circuits or chips (e.g., modem chips, also known as baseband chips, or SoC chips or SIP chips containing modem cores); or by a device outside the network device (e.g., intelligent network elements on the network side) or a component within the device (e.g., internal chips, processors, or circuits); or by a logic module or software capable of implementing all or part of the functions of the network device, etc. Alternatively, the method can be executed by a second device, which can be a network device; or a component within the network device, such as internal circuits or chips (e.g., modem chips), or SoC chips or SIP chips containing modem cores; or a device outside the network device (e.g., intelligent network elements) or a component within the device (e.g., internal chips, processors, or circuits); or a logic module or software capable of implementing some or all of the functions of the network side, etc. This application does not limit this.

[0062] For example, the method includes: sending a dataset and / or a model, the dataset and / or the model being used for model training; sending first information, the first information being used to indicate multiple sets of metrics, the multiple sets of metrics being used for model training; wherein the multiple sets of metrics include one or more of the following types of metrics: performance metrics, metrics of training time required, or metrics of training resources required, the multiple sets of metrics corresponding to the dataset and / or the model, and corresponding to the inference task of a first model, the inference task of the first model being to compress the target CSI.

[0063] In conjunction with the second aspect, in some possible implementations of the second aspect, at least two of the multiple sets of indicators are the performance indicators.

[0064] In conjunction with the second aspect, in some possible implementations of the second aspect, the sending of the dataset and / or model includes: sending a first dataset, the first dataset including: a target CSI and CSI feedback information, the CSI feedback information being obtained by compressing the target CSI.

[0065] One possible design is that the first dataset is used to train the second model, and the second model is used to train the first model, wherein the inference task of the second model is to decompress the CSI feedback information.

[0066] Optionally, the at least two function sets corresponding to the at least two sets of indicators include a first function set, which includes one or more of the following: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the label, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS or SGCS between the output of the second model and the label.

[0067] Further optionally, the output of the second model includes the output of the second model during training. Accordingly, the label is the target CSI in the first dataset. The first set of functions includes one or more of the following: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI in the first dataset; or, a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the target CSI in the first dataset.

[0068] Further optionally, the output of the second model includes: the output of the second model when training the first model based on the second model. One possible implementation of training the first model based on the second model is to jointly train the first and second models. Accordingly, the label of the joint training is a target CSI in the first dataset or a target CSI in the third dataset. The first set of functions includes one or more of the following: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the target CSI, wherein the target CSI is a target CSI in the first dataset or a target CSI in the third dataset.

[0069] Another possible design is that the first dataset is used to train the first model.

[0070] Optionally, the at least two function sets corresponding to the at least two sets of indicators include a second function set, the second function set including one or more of the following: NMSE, MSE, L1Loss, GCS, SGCS between the output of the first model and the CSI feedback information, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS or SGCS between the output of the first model and the CSI feedback information.

[0071] In conjunction with the second aspect, in some possible implementations of the second aspect, the sending of the dataset and / or model includes: sending a third model, the inference task of which is to compress the target CSI.

[0072] One possible design is that the third model is used to train the second model, and the second model is used to train the first model, wherein the inference task of the second model is to decompress the CSI feedback information, which is obtained by the third model by compressing the input target CSI.

[0073] Optionally, the at least two function sets corresponding to the at least two sets of indicators include a third function set, which includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the target CSI of the third model, wherein the target CSI is the target CSI in the first dataset or the target CSI in the third dataset.

[0074] Further optionally, the output of the second model includes the output of the second model during training. Accordingly, the label is the target CSI. The third function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the target CSI, wherein the target CSI is the target CSI in the first dataset or the target CSI in the third dataset.

[0075] Further optionally, the output of the second model includes the output of the second model when training the first model based on the second model. One possible implementation of training the first model based on the second model is to jointly train the first and second models. Accordingly, the label of the joint training is the target CSI in the first dataset or the target CSI in the third dataset. The third function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the target CSI, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the target CSI, wherein the target CSI is the target CSI in the first dataset or the target CSI in the third dataset.

[0076] Another possible design is that the third model is used to obtain a second dataset, which includes CSI feedback information, and the second dataset is used to train the first model.

[0077] Optionally, the at least two function sets corresponding to the at least two sets of indicators include a fourth function set, which includes one or more of the following functions: NMSE, MSE, L1Loss between the output of the first model and the CSI feedback information, or a weighted sum of multiple NMSE, MSE, or L1Loss between the output of the first model and the CSI feedback information.

[0078] In conjunction with the second aspect, in some possible implementations of the second aspect, the method further includes: receiving second information, the second information being used to indicate that a first indicator is satisfied, the first indicator being an indicator used for training the model, and the first indicator being one or more of the plurality of indicators.

[0079] In conjunction with the second aspect, in some possible implementations of the second aspect, the method further includes: receiving third information, the third information being used for one or more of the following: a first metric is not satisfied, the first metric being a metric used for model training, and the first metric being one or more of the multiple sets of metrics; requesting a change of the dataset; requesting a change of the model; or requesting to stop or pause the model training.

[0080] For more detailed explanations of the first, second, and third models in the various implementation methods of the second aspect, as well as the relevant explanations of the correspondence between multiple sets of indicators and the inference task of the first model, please refer to the relevant descriptions in the first aspect, which will not be repeated here.

[0081] It should be understood that the methods provided in the second aspect correspond to those in the first aspect. The descriptions and technical effects of the various implementation methods in the second aspect can be found in the relevant descriptions in the first aspect, and will not be repeated here.

[0082] Thirdly, an apparatus is provided. This apparatus may include functional modules corresponding to each of the methods / operations / steps / actions described in any possible implementation of the first aspect, or may include functional modules corresponding to each of the methods / operations / steps / actions described in any of the second aspects. The module may be a hardware circuit, software, or a combination of hardware circuitry and software implementation.

[0083] In one design, the device may include a processing module and a communication module. The communication module is used to perform the sending and receiving actions performed by the terminal side in the method described in the first aspect above, while the processing module is used to perform processing-related actions performed by the terminal side in the method described in the first aspect above.

[0084] In one design, the device can be a terminal device, or a device, module, circuit, or chip configured in the terminal device, or a device that can be used in conjunction with the terminal device, such as an OTT host or cloud server.

[0085] In one design, the device may include a processing module and a communication module. The communication module is used to perform the sending and receiving actions performed by the network side in the method described in the second aspect above, while the processing module is used to perform processing-related actions performed by the network side in the method described in the second aspect above.

[0086] In one design, the device can be a network device, or a device, module, circuit, or chip configured in the network device, or a device that can be used in conjunction with the network device, such as an intelligent network element with a radio access network (RAN) intelligent controller (RIC) deployed thereon.

[0087] Fourthly, an apparatus is provided, comprising a processor and a storage medium storing instructions that, when executed by the processor, cause a method as described in the first aspect or any possible implementation thereof to be implemented, or cause a method as described in the second aspect or any possible implementation thereof to be implemented.

[0088] Fifthly, an apparatus is provided, comprising a processing circuit for processing data and / or information such that a method as in the first aspect or any possible implementation thereof is implemented, or a method as in the second aspect or any possible implementation thereof is implemented.

[0089] The processing circuit may include one or more processors, or all or part of the circuitry in one or more processors used for control or processing functions.

[0090] Optionally, the apparatus may further include a memory for storing programs or instructions, and the processor for running the programs or instructions to implement the methods as described in the first aspect or any possible implementation thereof, or to implement the methods as described in the second aspect or any possible implementation thereof.

[0091] Optionally, the device may also include the transceiver circuit, or an input / output interface.

[0092] In a sixth aspect, a chip is provided, including processing circuitry for running a program or instructions to cause the method as described in the first aspect or any possible implementation thereof to be implemented, or to cause the method as described in the second aspect or any possible implementation thereof to be implemented.

[0093] Optionally, the chip may further include a memory for storing programs or instructions.

[0094] Optionally, the chip may also include transceiver circuitry, or input / output interfaces.

[0095] A seventh aspect provides a computer-readable storage medium comprising instructions that, when executed by a processor, cause the method as described in the first aspect or any possible implementation thereof to be implemented, or cause the method as described in the second aspect or any possible implementation thereof to be implemented.

[0096] Eighthly, a computer program product is provided, the computer program product comprising computer program code or instructions, which, when executed, cause the method as described in the first aspect and any possible implementation thereof to be implemented, or cause the method as described in the second aspect and any possible implementation thereof to be implemented.

[0097] Ninth aspect, a communication system is provided, the communication system including means for performing the first aspect and any possible implementation thereof, or including means for performing the second aspect and any possible implementation thereof.

[0098] It should be understood that the third to ninth aspects of this application correspond to the technical solutions of the first to second aspects of this application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description

[0099] Figure 1 is a schematic diagram of a communication system applicable to the communication method of this application embodiment;

[0100] Figure 2 is a schematic diagram of another communication system applicable to the communication method of this application embodiment;

[0101] Figure 3 is a schematic diagram of a possible application framework in a communication system;

[0102] Figure 4 is a schematic diagram of another possible application framework in a communication system;

[0103] Figure 5 is a schematic diagram of CSI feedback using an auto-encoder (AE) model provided in an embodiment of this application;

[0104] Figure 6 shows an example of a neuron structure;

[0105] Figure 7 is a schematic diagram of a deep neural network (DNN);

[0106] Figure 8 is a schematic diagram of the data set docking between the network side and the terminal side provided in an embodiment of this application;

[0107] Figure 9 is a schematic diagram of model interoperability between the network side and the terminal side provided in an embodiment of this application;

[0108] Figure 10 is a schematic flowchart of the communication method provided in an embodiment of this application;

[0109] Figures 11 to 14 are flowcharts of model training provided in the embodiments of this application;

[0110] Figure 15 is another schematic flowchart of the communication method provided in an embodiment of this application;

[0111] Figures 16 and 17 are schematic block diagrams of a communication device provided in an embodiment of this application. Detailed Implementation

[0112] The technical solution provided in this application will now be described with reference to the accompanying drawings.

[0113] To facilitate understanding of the embodiments of this application, the following points will be explained first:

[0114] First, in this application, the terminal side can also be referred to as the user equipment (UE) side, terminal-side equipment, etc., including: terminal equipment (or user equipment, terminal, etc.), components deployed in the terminal equipment (such as circuits or chips inside the terminal equipment), equipment deployed outside the terminal equipment (such as the host or cloud server of an OTT system, hereinafter referred to as the OTT system server), or components deployed in equipment outside the terminal equipment (such as circuits or chips inside the equipment). The network side (NW side) can also be referred to as network-side equipment, including: network equipment communicating with the terminal equipment, components deployed in the network equipment (such as circuits or chips inside the network equipment with near real-time radio access network (RAN) intelligent control functions), equipment deployed outside the network equipment (such as intelligent network elements, for example, intelligent network elements with near real-time RAN intelligent control functions), or components deployed in the intelligent network element (such as circuits or chips inside the intelligent network element). Among them, network equipment can include: access network equipment, core network equipment, or operation administration and maintenance (OAM).

[0115] Second, in this application, the indication includes direct indication (also known as explicit indication) and indirect indication (also known as implicit indication). Directly indicating information A means including information A; indirectly indicating information A can mean indicating information A through the correspondence between information A and information B and by directly indicating information B; or by indicating information A through a preset rule that can be used to determine A based on B and by directly indicating information B. The correspondence between information A and information B, and the preset rule, can be predefined, pre-stored, pre-burned, or pre-configured.

[0116] Third, in this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the preceding and following related objects, but it does not exclude the possibility of indicating an "and" relationship; the specific meaning can be understood in context. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c; a and b; a and c; b and c; or a and b and c. Here, a, b, and c can be single or multiple.

[0117] Fourth, the use of prefixes such as "first" and "second" in this application is merely for the purpose of distinguishing and describing different things belonging to the same name category, and does not constrain the order, size, or quantity of things. For example, "first dataset" and "second dataset" are simply different datasets, and do not limit the number, size, or priority of the datasets; similarly, "first model" and "second model" are simply different models, and do not limit the number, size, or priority of the models; furthermore, "first information" and "second information" are simply different indicative information, and do not limit the quantity, chronological order, size, or priority of the information.

[0118] Fifth, in this application, "send" and "receive" indicate the direction of signal transmission. For example, "send information to a terminal device" can be understood as the destination of the information being the terminal device, which may include direct transmission via the air interface or indirect transmission by other units or modules via the air interface. "Receive information from a network device" can be understood as the source of the information being the network device, which may include direct reception from the network device via the air interface or indirect reception from the network device by other units or modules via the air interface. "Send" can also be understood as the "output" of the chip interface, and "receive" can also be understood as the "input" of the chip interface. In other words, sending and receiving can occur between devices, such as between a terminal device and a computing node, or within a device, such as between components, modules, chips, software modules, or hardware modules within the device via a bus, wiring, or interface.

[0119] Sixth, in the embodiments of this application, "when," "if," and "if" all refer to the device making corresponding processing under certain objective circumstances, and are not limited to a time, nor do they require the device to make a judgment action when it is implemented, nor do they mean that there are other limitations.

[0120] Seventh, in this application, the words "example," "exemplarily," "for example," or "such as" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "example," "exemplarily," "for example," or "such as" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the words "example," "exemplarily," "for example," or "such as" is intended to present the relevant concepts in a specific manner.

[0121] Eighth, for ease of distinction and explanation, this paper introduces the terms encoder and decoder, quantizer and dequantizer. These names are given only to distinguish different functions and do not limit the structure of the device. For example, if the terminal side (such as a terminal device or OTT system server) has the inference function of the encoder, it can be said that the terminal side contains the encoder, which can be understood as a functional module of the terminal side; if the terminal side has the inference function of the decoder, it can be said that the terminal side contains the decoder, which can be understood as a functional module of the terminal side. If the encoder has the quantization function, it can be said that the encoder contains the quantizer. If the decoder has the dequantization function, it can be said that the decoder contains the dequantizer. In specific implementations, the encoder and decoder can be implemented through a neural network model. More specifically, the functions of the encoder and decoder, quantizer and dequantizer can be implemented separately in hardware, or in software, or in a combination of hardware and software; this application does not limit this.

[0122] The technical solutions provided in this application can be applied to various communication systems, such as 5th generation (5G) or new radio (NR) systems, frequency division duplex (FDD) systems, time division duplex (TDD) systems, wireless local area network (WLAN) systems, satellite communication systems, future communication systems, or integrated systems of multiple systems. The technical solutions provided in this application can also be applied to device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, machine-to-machine (M2M) communication, machine-type communication (MTC), and Internet of Things (IoT) communication systems or other communication systems.

[0123] In a communication system, one network element can send signals to or receive signals from another network element. These signals can include information, signaling, or data. The term "network element" can also be replaced by an entity, network entity, device, communication equipment, communication module, node, communication node, etc. This disclosure uses a network element as an example. For instance, a communication system can include at least one terminal device and at least one network device. The network device can send downlink signals to the terminal device, and / or the terminal device can send uplink signals to the network device. It is understood that the terminal device in this disclosure can be replaced by a first network element, and the network device can be replaced by a second network element, both performing the corresponding methods described in this disclosure.

[0124] Figure 1 is a schematic diagram of a communication system applicable to the communication method of this application embodiment. As shown in Figure 1, the communication system 100A may include at least one access network device, such as access network device 110 shown in Figure 1; the communication system 100A may also include at least one terminal device, such as terminal device 120 and terminal device 130 shown in Figure 1. Access network device 110 and terminal devices (such as terminal device 120 and terminal device 130) can communicate via a wireless link. The communication devices in this communication system, for example, access network device 110 and terminal device 120, can communicate via multi-antenna technology.

[0125] In wireless communication networks, such as mobile communication networks, the services supported by the network are becoming increasingly diverse, thus requiring increasingly diverse demands. For example, the network needs to support ultra-high speed, ultra-low latency, and / or massive connectivity. This characteristic makes network planning, network configuration, and / or resource scheduling increasingly complex. Furthermore, as network functions become more powerful, such as supporting higher spectrum, supporting higher-order multiple-input multiple-output (MIMO) technology, supporting beamforming (BF), and supporting beam management, network energy saving has become a hot research topic. These new demands, new scenarios, and new characteristics bring unprecedented challenges to network planning, operation, and efficient operation. To meet these challenges, artificial intelligence (AI) technology can be introduced into wireless communication networks to achieve network intelligence. To support AI technology in wireless networks, AI nodes may also be introduced. AI nodes can be AI network elements or AI modules. Figure 2 is a schematic diagram of another communication system applicable to the communication method of this application embodiment. Compared to the communication system 100A shown in Figure 1, the communication system 100B shown in Figure 2 also includes an AI network element 140. AI network element 140 is used to perform AI-related operations, such as building datasets or AI models. Here, AI network element can also be simply referred to as intelligent network element. In this disclosure, AI model can be simply referred to as model.

[0126] In one possible implementation, access network device 110 can send data related to the training of the AI ​​model to AI network element 140, whereby AI network element 140 constructs a dataset and trains the AI ​​model. For example, the data related to the training of the AI ​​model may include data reported by terminal devices. AI network element 140 can send the results of operations related to the AI ​​model to access network device 110, and then forward them to terminal devices via access network device 110. For example, the results of operations related to the AI ​​model may include at least one of the following: a trained AI model, model evaluation results, or test results, etc. Exemplarily, a portion of the trained AI model may be deployed on access network device 110, and another portion on terminal devices 120 and / or 130. Alternatively, the trained AI model may be deployed on access network device 110. Or, the trained AI model may be deployed on terminal devices 120 and / or 130.

[0127] It should be understood that Figure 2 is only used as an example of the AI ​​network element 140 being directly connected to the access network device 110. In other scenarios, the AI ​​network element 140 can also be connected to the terminal device. Alternatively, the AI ​​network element 140 can be connected to both the access network device 110 and the terminal device simultaneously. Alternatively, the AI ​​network element 140 can also be connected to the access network device 110 through a third-party network element. This application embodiment does not limit the connection relationship between the AI ​​network element and other network elements. For example, the AI ​​network element 140 can also be set as a module in the access network device and / or the terminal device, for example, in the access network device 110 or the terminal device shown in Figure 1.

[0128] It should be noted that Figures 1 and 2 are simplified schematic diagrams for ease of understanding. For example, the communication system may also include other devices, such as wireless relay devices and / or wireless backhaul devices, which are not shown in Figures 1 and 2. In practical applications, the communication system may include multiple access network devices and multiple terminal devices. The embodiments of this application do not limit the number of access network devices and terminal devices included in the communication system.

[0129] In the embodiments of this application, the terminal device may also be referred to as UE, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user equipment.

[0130] Terminal devices can be devices that provide voice / data, such as handheld devices with wireless connectivity, in-vehicle devices, etc. Currently, examples of terminals include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, wearable devices, terminal devices in 5G networks, or future public land mobile communication networks. Terminal devices in a network (PLMN), etc., are not limited to this in the embodiments of this application.

[0131] By way of example and not limitation, in this embodiment, the terminal device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0132] In this embodiment, the device for implementing the functions of the terminal device can be the terminal device itself, or it can be any device capable of supporting the terminal device in implementing those functions, such as a chip system. This device can be installed in or used in conjunction with the terminal device. In this embodiment, the chip system can be composed of chips or may include chips and other discrete components. This embodiment only uses the terminal device as an example to illustrate the device for implementing the functions of the terminal device, and does not constitute a limitation on the solution of this embodiment.

[0133] The network device in this application embodiment can be a device for communicating with a terminal device. This network device may include an access network device, a core network device, or other devices in the communication system. The access network device may be, for example, a base station. In this application embodiment, the access network device may refer to a radio access network (RAN) node (or device) that connects the terminal device to the wireless network. A base station can broadly encompass, or be replaced by, various names such as: NodeB, evolved NodeB (eNB), next-generation NodeB (gNB), relay station, access point, transmit / receive point (TRP), transmitting point (TP), master station, auxiliary station, motor slide retainer (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), radio unit (RU), positioning node, etc. A base station can be a macro base station, micro base station, relay node, donor node, or similar entities, or combinations thereof. A base station can also refer to a communication module, modem, or chip installed within the aforementioned equipment or apparatus. A base station can also be a mobile switching center, a device that performs base station functions in D2D, V2X, and M2M communications, or a device that performs base station functions in future communication systems. A base station can support networks using the same or different access technologies. Optionally, a RAN node can also be a server, wearable device, vehicle, or in-vehicle equipment. For example, the access network equipment in V2X technology can be a roadside unit (RSU). The embodiments of this application do not limit the specific technology or equipment form used in the access network equipment.

[0134] Base stations can be fixed or mobile. For example, a helicopter or drone can be configured to act as a mobile base station, and one or more cells can move depending on the location of the mobile base station. In other examples, a helicopter or drone can be configured as a device to communicate with another base station.

[0135] In some deployments, the access network equipment mentioned in the embodiments of this application may be a device including a CU, or a DU, or a device including both CU and DU, or a device with a control plane CU node (central unit-control plane (CU-CP)) and a user plane CU node (central unit-user plane (CU-UP)) and a DU node. For example, the access network equipment may include gNB-CU-CP, gNB-CU-UP, and gNB-DU.

[0136] In some deployments, multiple RAN nodes collaborate to assist terminals in achieving wireless access, with different RAN nodes each implementing some of the base station's functions. For example, RAN nodes can be CUs, DUs, CU-CPs, CU-UPs, or RUs. CUs and DUs can be configured separately or included in the same network element, such as a BBU. RUs can be included in radio frequency equipment or radio frequency units, such as RRUs, AAUs, or RRHs.

[0137] RAN nodes can support one or more types of fronthaul interfaces, each corresponding to a DU and RU with different functions. If the fronthaul interface between the DU and RU is a Common Public Radio Interface (CPRI), the DU is configured to implement one or more baseband functions, and the RU is configured to implement one or more radio frequency functions. If the fronthaul interface between the DU and RU is another type of interface, relative to CPRI, it moves some downlink and / or uplink baseband functions—for example, for downlink, precoding, digital beamforming, or one or more of inverse fast Fourier transform (IFFT) / adding a cyclic prefix (CP)—from the DU to the RU; and for uplink, digital beamforming, or one or more of fast Fourier transform (FFT) / removing CP—from the DU to the RU. In one possible implementation, this interface can be an enhanced common public radio interface (eCPRI). Under the eCPRI architecture, the splitting methods between DU and RU are different, corresponding to different types (category, Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, F.

[0138] Taking eCPRI Cat A as an example, for downlink transmission, layer mapping is used as the dividing line. The DU is configured to implement one or more functions preceding layer mapping (i.e., coding, rate matching, scrambling, modulation, and layer mapping itself), while other functions following layer mapping (e.g., resource element (RE) mapping, digital beamforming, or one or more of IFFT / CP addition) are implemented in the RU. For uplink transmission, de-RE mapping is used as the dividing line. The DU is configured to implement one or more functions preceding de-mapping (i.e., decoding, rate matching de-matching, descrambling, demodulation, inverse discrete Fourier transform (IDFT), channel equalization, and de-RE mapping itself), while other functions following de-mapping (e.g., digital BF or FFT / CP removal) are implemented in the RU. It is understood that descriptions of the functions of the DU and RU corresponding to various types of eCPRI can be found in the eCPRI protocol and will not be elaborated upon here.

[0139] In one possible design, the processing unit in the BBU used to implement baseband functions is called the baseband high (BBH) unit, and the processing unit in the RRU / AAU / RRH used to implement baseband functions is called the baseband low (BBL) unit.

[0140] In different systems, CU (or CU-CP and CU-UP), DU, or RU may have different names, but those skilled in the art will understand their meaning. For example, in an open RAN (ORAN) architecture, CU can also be called open CU (open-CU, O-CU), DU can also be called open DU (open-DU, O-DU), CU-CP can also be called open CU-CP (open-CU-CP) O-CU-CP, CU-UP can also be called open CU-UP (open-CU-UP, O-CU-UP), and RU can also be called open RU (open-RU, O-RU). Any of the units among CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through software modules, hardware modules, or a combination of software modules and hardware modules.

[0141] In this embodiment, the apparatus for implementing the functions of a network device can be a network device itself; it can also be an apparatus capable of supporting the network device in implementing those functions, such as a chip system, hardware circuit, software module, or a hardware circuit plus a software module. This apparatus can be installed in the network device or used in conjunction with the network device. In this embodiment, the example of a network device being used to implement the functions of a network device is provided only and does not constitute a limitation on the solutions described in this embodiment.

[0142] Network devices and / or terminal devices can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; they can also be deployed on water; and they can also be deployed in the air on airplanes, balloons, and satellites. This application does not limit the scenario in which the network devices and terminal devices are located. Furthermore, terminal devices and network devices can be hardware devices, or software functions running on dedicated hardware or general-purpose hardware, such as virtualization functions instantiated on a platform (e.g., a cloud platform), or entities that include dedicated or general-purpose hardware devices and software functions. This application does not limit the specific form of the terminal devices and network devices.

[0143] Optionally, the AI ​​node can be deployed in one or more of the following locations within the communication system: access network equipment, terminal equipment, or core network elements. Alternatively, the AI ​​node can also be deployed independently, for example, in a location other than any of the aforementioned devices, such as in the host or cloud server of an over-the-top (OTT) system. The AI ​​node can communicate with other devices in the communication system, which can be, for example, one or more of the following: access network equipment, terminal equipment, or core network elements.

[0144] It is understood that this application does not limit the number of AI nodes. For example, when there are multiple AI nodes, these nodes can be divided based on function, such as different AI nodes being responsible for different functions.

[0145] It can also be understood that AI nodes can be independent devices, or they can be integrated into the same device to achieve different functions. Alternatively, they can be network elements in hardware devices, software functions running on dedicated hardware, or virtualization functions instantiated on a platform (e.g., a cloud platform). This application does not limit the specific form of the aforementioned AI nodes.

[0146] Figure 3 illustrates a possible application framework in a communication system. As shown in Figure 3, network elements in the communication system are connected via interfaces (e.g., next-generation (NG) interfaces, Xn interfaces) or air interfaces. The NG interface is the interface between the radio access network and the 5G core network. The Xn interface is the interface between access network devices, and the air interface is the interface between access network devices and terminal devices. These network element nodes, such as core network devices, RAN nodes, terminal devices, or one or more devices in the OAM, are equipped with one or more AI modules (only one is shown in Figure 3 for clarity). The access network node can be a single RAN node or can include multiple RAN nodes, for example, including CU and DU. The CU and / or DU can also be equipped with one or more AI modules. Optionally, the CU can be further divided into CU-CP and CU-UP. One or more AI models are configured in the CU-CP and / or CU-UP.

[0147] The AI ​​module is used to implement corresponding AI functions. AI modules deployed in different network elements can be the same or different. Depending on the parameter configuration, the AI ​​module can implement different functions. The AI ​​module model can be configured based on one or more of the following parameters: structural parameters (e.g., at least one of the following: number of neural network layers, neural network width, inter-layer connections, neuron weights, neuron activation function, or bias in the activation function), input parameters (e.g., type and / or dimension of input parameters), or output parameters (e.g., type and / or dimension of output parameters). The bias in the activation function can also be referred to as the neural network bias.

[0148] An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or they can be deployed on the same node or device.

[0149] Network devices can be network devices equipped with one or more AI modules. These network devices can include one or more devices in the core network, RAN, or OAM as shown in Figure 3. For example, the AI ​​module can be a RAN intelligent controller (RIC) as shown in Figure 4, such as a near-real-time RIC (near-RT RIC) or a non-real-time RIC (non-RT RIC). For instance, a near-real-time RIC is located in a RAN node (e.g., in a CU or DU), while a non-real-time RIC is located in the OAM, a cloud server, a core network device, or other access network devices. The RIC can obtain subsets from multiple terminal devices from RAN nodes (e.g., CU, CU-CP, CU-UP, DU, and / or RU), reassemble them into a dataset, and train based on the dataset. Exemplarily, near-real-time RICs and non-real-time RICs can also be configured as separate network elements, and access network devices can be either near-real-time or non-real-time RICs.

[0150] Figure 4 illustrates another possible application framework in a communication system. In addition to access network nodes (CU, DU, and RU are shown in the figure) and terminals, the communication system shown in Figure 4 also includes an RIC (Regulator-Integrated Circuit). For example, the RIC could be the AI ​​module shown in Figure 3, which can be used to implement AI-related functions. The RIC includes near-real-time RICs and non-real-time RICs. Non-real-time RICs primarily process non-real-time information, such as data that is not sensitive to latency, with latency on the order of seconds. Real-time RICs primarily process near-real-time information, such as data that is relatively sensitive to latency, with latency on the order of tens of milliseconds.

[0151] The near real-time RIC is used for model training and inference. For example, it can be used to train an AI model and then use that AI model for inference. The near real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (e.g., CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or inference data. Optionally, the near real-time RIC can deliver inference results to RAN nodes and / or terminals. Optionally, inference results can be exchanged between CU and DU, and / or between DU and RU. For example, the near real-time RIC delivers the inference result to the DU, and the DU sends it to the RU.

[0152] The non-real-time RIC is also used for model training and inference. For example, it can be used to train an AI model and then use that model for inference. The non-real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (e.g., one or more of CU, CU-CP, CU-UP, DU, or RU) and / or terminals. This information can be used as training data or inference data, and the inference results can be delivered to the RAN nodes and / or terminals. Optionally, inference results can be exchanged between CU and DU, and / or between DU and RU; for example, the non-real-time RIC delivers the inference results to the DU, which then forwards them to the RU.

[0153] The near real-time RIC and non-real-time RIC can also be set up as separate network elements. Optionally, the near real-time RIC and non-real-time RIC can also be part of other devices. For example, the near real-time RIC can be set in the RAN node (e.g., in CU, DU), while the non-real-time RIC can be set in the OAM, cloud server, core network device, or other network device.

[0154] With the development of wireless communication technology and the increasing number of supported services, higher demands are being placed on communication systems in terms of system capacity, communication latency, and other indicators. Among these advancements, massive MIMO (Multi-User MIMO) systems, by configuring large-scale antenna arrays at the transceiver end, can achieve spatial diversity gain, thereby significantly increasing system capacity. For example, access network equipment can simultaneously transmit data to multiple terminal devices using the same time-frequency resources, i.e., multi-user MIMO (MU-MIMO); or, access network equipment can simultaneously transmit multiple data streams to the same terminal device, i.e., single-user MIMO (SU-MIMO). The data between these multiple terminal devices or the multiple data streams within the same terminal device are spatially multiplexed, thus becoming a key direction in the evolution of communication systems.

[0155] Access network equipment needs to obtain the CSI of the downlink channel to determine the resources, modulation and coding scheme (MCS), precoding and other configurations of the downlink data channel of the scheduling terminal equipment.

[0156] Taking precoding as an example, in massive MIMO, access network equipment needs to precode downlink data using a precoding matrix. By using precoding techniques, access network equipment can achieve spatial multiplexing between terminal devices or data streams. This means that data between different terminal devices or between different data streams of the same terminal device is spatially isolated, thereby reducing interference between different terminal devices or different data streams and improving the received signal-to-interference-plus-noise ratio (SINR) of the terminal devices. To calculate the precoding matrix, the access network equipment needs to obtain the CSI of the downlink channel and determine the precoding matrix based on the CSI.

[0157] In TDD systems, due to the reciprocity of uplink and downlink channels, access network devices can obtain the uplink CSI by measuring the uplink reference signal and then infer a relatively accurate downlink CSI, for example, using the uplink CSI as the downlink CSI. However, in FDD systems, uplink and downlink reciprocity cannot be guaranteed. The downlink CSI is obtained by the terminal device measuring the downlink reference signal, such as the channel state information reference signal (CSI-RS) or the synchronizing signal block (SSB). Therefore, the terminal device needs to generate a CSI report according to the protocol predefined method or the access network device configuration, and then feed the generated CSI report back to the access network device so that it can obtain the downlink CSI.

[0158] In FDD systems, a crucial part of CSI feedback is the precoding matrix indicator (PMI), which uses 0-1 bits to quantize the channel matrix or precoding matrix in CSI. PMI design (also known as codebook design) is a fundamental issue in mobile communication systems. Traditional codebook design methods predefine (or agree upon) a series of precoding matrices and their corresponding numbers in the protocol; these precoding matrices are called codewords. The channel matrix or precoding matrix can be approximated using predefined codewords or linear combinations of multiple predefined codewords. Therefore, the terminal device can feed back the corresponding codeword numbers and one or more weighting coefficients to the access network device via the PMI, which is used by the access network device to reconstruct the channel matrix or precoding matrix.

[0159] As the antenna array size of MIMO systems continues to increase, the number of supported antenna ports also increases, leading to a growth in the dimensionality of the corresponding channel matrix and precoding matrix. To enable terminal devices to estimate (or measure) the downlink channel, the overhead of the access network equipment transmitting reference signals increases. Simultaneously, the error of approximating large-scale channel matrices and precoding matrices with a finite number of predefined codewords increases. One method to improve channel reconstruction accuracy is to increase the number of codewords in the codebook, but this simultaneously increases the overhead of CSI feedback (including the corresponding codeword number and one or more weighting coefficients), thereby reducing the available resources for data transmission and causing system capacity loss. Therefore, it is necessary to investigate how to more effectively compress and represent channel information without increasing the overhead of transmitting reference signals and CSI feedback, and how to more effectively reconstruct the channel based on feedback information. Correlation exists between different elements in the downlink channel matrix between access network equipment and terminal equipment; furthermore, correlation exists between downlink channel matrices in different time slots. For example, the correlation between different elements in the channel matrix implies the existence of a basis (which can be represented by matrices U1 and U2). Projecting the channel matrix H onto this basis yields a sparse equivalent channel, i.e., H' = U1H. H U2 is a sparse matrix, where the superscript H denotes the conjugate transpose operation. Theoretically, the channel matrix H can be reconstructed simply by estimating the non-zero elements in H' using the reference signal and feeding them back. Therefore, the overhead of the reference signal transmission and CSI feedback has room for compression. However, traditional CSI feedback schemes, such as the codebook-based feedback method mentioned above, do not fully utilize the channel compression space, and the channel compression process may cause significant information loss. Machine learning methods (such as deep learning (DL)) have stronger nonlinear feature extraction capabilities, thus enabling more effective extraction of correlations between channel matrices. Consequently, compared to traditional schemes, they can more effectively compress and represent channel information, and more effectively reconstruct channel information based on the feedback information.

[0160] To facilitate understanding of the embodiments of this application, the terms involved in this application will be briefly explained below.

[0161] 1. Channel State Information (CSI): The meaning of CSI in this application is broader than that of traditional CSI, including but not limited to channel quality indication (CQI), PMI, rank indicator (RI), CSI-RS resource indicator (CRI), and may also include one or more of the following: channel response information (such as channel response matrix, frequency domain channel response information, time domain channel response information), weight information corresponding to the channel response, precoding matrix information corresponding to the channel response, reference signal receiving power (RSRP), reference signal receiving quality (RSRQ) or SINR, etc.

[0162] The following are some of the CSI-related terms used in this application:

[0163] Target CSI: Also known as the full CSI information, the uncompressed CSI information, the raw CSI, or the original CSI information. For ease of distinction and explanation, the target CSI will be represented by V in the following text.

[0164] CSI feedback information: also known as CSI feedback information, channel measurement result feedback information, channel information feedback information, compressed information, compressed channel information, compressed CSI information, compressed channel information, or compressed CSI, etc. For ease of distinction and explanation, CSI feedback information will be represented by "C" below.

[0165] It should be noted that in the embodiments of this application, during the training phase, CSI feedback information can refer to the information obtained after compressing the target CSI, which can be simply referred to as compressed CSI; or it can refer to the information obtained by compressing and quantizing the target CSI, which can be simply referred to as quantized CSI. This application does not limit this. During the model deployment phase, CSI feedback information can be information sent from the terminal side to the network side. To save feedback overhead, CSI feedback information can refer to the information obtained by compressing and quantizing the measured CSI (also called the target CSI), i.e., quantized CSI.

[0166] It should be understood that quantizing compressed CSI yields quantized CSI; dequantizing quantized CSI yields compressed CSI. Compressed CSI obtained by dequantizing quantized CSI is compressed CSI with quantization loss.

[0167] Reconstructed CSI: Also known as recovered CSI, reconstructed CSI, restored CSI, reconstructed channel information, decompressed CSI, decompressed channel information, etc. For ease of distinction and explanation, the term reconstructed CSI will be used in the following text. express.

[0168] 2. AI Model: A function model that maps an input of a certain dimension to an output of a certain dimension. Its parameters can be obtained through machine learning (ML). For example, f(x) = ax 2 +b is a quadratic function model, which can be viewed as an AI model. a and b correspond to the parameters of the model and can be obtained through machine learning training.

[0169] 3. AE Model: This can generally refer to a network structure consisting of two AI models, such as an encoder and a decoder. Each model can be an AI model. AE models are also called bilateral models, two-end models, collaborative models, etc. The encoder and decoder of an AE are usually trained together and can be used in a coordinated manner.

[0170] In this application, CSI feedback can be implemented based on the AI ​​model of AE. Figure 5 is a schematic diagram of CSI feedback using the AE model provided in an embodiment of this application.

[0171] As shown in the figure, the terminal side compresses the target CSI using an encoder, while the network side reconstructs the CSI using a decoder. For example, the terminal side can use the target CSI (i.e., V) as input to the encoder, which compresses the target CSI to obtain the compressed CSI. The terminal side can then quantize the compressed CSI to obtain the quantized CSI. The terminal side can then send the quantized CSI as CSI feedback information (i.e., C) to the network side, for example, through a CSI report.

[0172] The network side can first dequantize the CSI feedback information to obtain compressed CSI with quantization loss. The network side can then use this compressed CSI as input to the decoder, which can decompress it based on the compressed CSI to obtain the reconstructed CSI (i.e., ).

[0173] The quantizer used to perform quantization can be predefined, such as protocol predefined, or it can be indicated by the network side; this application does not limit this.

[0174] In another implementation, the encoder on the terminal side can also compress and quantize the target CSI, outputting CSI feedback information. The decoder on the network side can also dequantize and decompress the CSI feedback information to obtain the reconstructed CSI. In this case, the encoder output in Figure 5 is the CSI feedback information, and the decoder input is also the CSI feedback information.

[0175] It should be noted that the encoder on the terminal side can be deployed inside the terminal device or in other devices outside the terminal device, such as the aforementioned OTT server; the decoder on the network side can be deployed inside the network device or in other devices outside the network device, such as the aforementioned intelligent network element.

[0176] It should be understood that although the figure shows an encoder and a decoder, this is only a model division from a functional perspective. The encoder can also be called the first model, and the decoder can also be called the second model, etc. This application does not limit this. In addition, this application does not limit the number of models included in the AE model.

[0177] It should also be understood that the AE model is only one possible model for achieving the above functions and should not constitute any limitation on this application. The AE model can also be replaced by other AI models that can achieve the same or similar functions.

[0178] 4. Model training: By selecting an appropriate function (such as a loss function), the model parameters are trained using optimization algorithms to minimize the difference between the model's predicted value and the ground truth (or target value, label).

[0179] For example, model training methods include, but are not limited to, supervised learning, self-supervised learning, and knowledge distillation. These methods are briefly explained below.

[0180] Supervised learning, also known as supervised instruction, involves using machine learning algorithms to learn the mapping relationship between sample values ​​and labels based on collected sample values ​​and labels. This learned mapping relationship is then expressed using a machine learning model. The process of training the machine learning model is the process of learning this mapping relationship. For example, in signal detection, the noisy received signal is the sample, and the corresponding real constellation point is the label. Machine learning aims to learn the mapping relationship between samples and labels through training, that is, to enable the machine learning model to learn a signal detector. During training, the model parameters are optimized by calculating the error between the model's predicted values ​​and the real labels. Once the mapping relationship is learned, it can be used to predict the sample label of each new sample. The mapping relationship learned in supervised learning can include linear and nonlinear mappings. Based on the type of label, the learning task can be divided into classification tasks and regression tasks.

[0181] Self-supervised learning is a type of unsupervised learning. Unsupervised learning relies on collected sample values ​​to allow algorithms to discover inherent patterns within the samples themselves. Self-supervised learning uses the samples themselves as supervisory signals; that is, the model learns the mapping relationship from sample to sample. During training, model parameters are optimized by calculating the error between the model's predicted values ​​and the actual samples. Self-supervised learning can be used in signal compression and decompression recovery applications; common algorithms include autoencoders and generative adversarial networks.

[0182] Knowledge distillation: Generally, large models are often single complex networks or collections of networks, possessing excellent performance and generalization ability, while small models, due to their smaller network size, have limited expressive power. Therefore, the knowledge learned by the large model can be used to guide the training of the small model; this process is called knowledge distillation. Knowledge distillation can enable small models to achieve performance comparable to large models, but with fewer parameters and shorter inference latency, thus achieving model compression and acceleration. Furthermore, directly training small models with massive amounts of data often does not yield good performance, while training large models with massive amounts of data and then using the large model to perform knowledge distillation on the small model can achieve better continuation results. In addition, knowledge distillation can also be used to integrate and transfer datasets from different domains.

[0183] Knowledge distillation employs a teacher-student model, where a teacher model assists in training a student model. The teacher model is a complex, large model, while the student model is a simple, small model. Because the teacher model has strong learning capabilities, it can transfer the knowledge it learns to the relatively weaker student model, thereby enhancing the student model's generalization ability. The complex, cumbersome, but effective teacher model remains offline, simply acting as a mentor; the flexible and lightweight student model is the one actually deployed for prediction tasks.

[0184] The model training methods listed above are merely examples, and this application does not limit the methods used for model training.

[0185] Model training in this application embodiment may include initial training (or original training) or retraining of the model. Model training enables model enhancement. Model enhancement can refer to improving model functionality through model structure optimization and / or model training, such as adding other functions or neural networks to the existing model functionality, so that the enhanced model can be applied to more scenarios or richer requirements.

[0186] 5. Performance metric function: also simply referred to as a function. The performance metric function in this application is used to describe the performance of the model, specifically to describe the difference between the model's predicted values ​​and the true values. Model training can be terminated when the value of the performance metric function reaches the range required by the performance index.

[0187] The performance metric function can be a loss function, also known as an objective function or cost function, used to measure the difference between the predicted and actual values. A higher output value (loss) indicates a greater difference, and model training can be a process of minimizing this difference. Loss functions involved in the embodiments of this application include, for example, L1Loss (also known as mean absolute error (MAE), MSE, and NMSE).

[0188] The performance metric function can also be used to measure the similarity between predicted and actual values. In this application, examples of performance metric functions used to measure the similarity between predicted and actual values ​​include GCS and SGCS.

[0189] The calculation methods for each performance metric function can be found in the formulas below:

[0190] in, These are predicted values, including M values: H represents the true values, including M values: h0, h1, ..., h M-1 .

[0191] In another implementation, the above GCS and SGCS can also be replaced by the following loss functions: the difference between GCS and 1, and the difference between SGCS and 1. In this case, the model training process is the process of reducing this difference.

[0192] In the embodiments of this application, the performance metric function and the performance index are corresponding. For example, the performance metric function is NMSE, and the corresponding performance index is the upper bound of NMSE; another example is that the performance metric function is SGCS, and the corresponding performance index is the lower bound of SGCS, or the upper bound of the difference between SGCS and 1, and so on, and will not be listed further.

[0193] Furthermore, this application does not impose any restrictions on the specific values ​​of the performance metrics (such as upper and / or lower bounds) corresponding to each performance metric function.

[0194] 6. Model Files and Model Parameters: Model files and / or model parameters can be used to determine the model. Therefore, the model in this application can refer to the model itself, or it can refer to the model files and / or model parameters used to determine the model.

[0195] The model file can be used to indicate the model structure, which includes, but is not limited to, feedforward neural networks (FNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). The model file can have a fixed format, such as a standard predefined format or a format pre-negotiated by both ends of the connection. Model parameters can refer to parameters in the neural network model, such as, but not limited to, the number of layers in the neural network, the type and weights of neurons in each layer, etc. This application does not limit the method of distributing model parameters.

[0196] Take DNN as an example. The idea behind DNN comes from the neuronal structure of the brain. Each neuron can perform a weighted summation operation on its inputs and then use the result of the weighted summation operation to generate an output through a nonlinear function. Figure 6 shows an example of a neuron structure. The input of the neuron shown in Figure 6 is x = [x0 x1 … x N-1 The weights corresponding to the inputs are w = [w0 w1 … w] N-1 The bias of the weighted summation is b. The nonlinear function f() can take many forms; for example, the nonlinear function f() can be the maximum value function max{0, x}. Then the effect of a neuron's execution is... Where N is a positive integer; n is a positive integer greater than or equal to 0 and less than or equal to (N-1).

[0197] A DNN typically has multiple neural network layers, including an input layer, one or more hidden layers, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. Each layer contains multiple neurons. Layers are fully connected; that is, any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. The input layer processes the received numerical values ​​(i.e., the DNN's input) through neurons and then passes them to the hidden layers. Similarly, the hidden layers pass the computation results to the final output layer, producing the DNN's output. Figure 7 shows an example of a DNN. The DNN model shown in Figure 7 has three neural network layers: an input layer, a hidden layer, and an output layer.

[0198] It should be understood that the examples in conjunction with Figures 6 and 7 above are shown for ease of understanding only and should not constitute any limitation on this application. This application does not limit the structure and parameters used in the AI ​​model.

[0199] One of the model structure or model parameters can be predefined, while the other can be sent by the sender (e.g., the network side). Alternatively, both the model structure and model parameters can be sent by the sender (e.g., the network side). This application does not impose any restrictions on this.

[0200] In this embodiment of the application, sending a model may refer to sending a model file and / or model parameters, and receiving a model may refer to receiving a model file and / or model parameters.

[0201] 7. Two-way connection: This refers to the connection between the sender (e.g., the network side) and the receiver (e.g., the terminal side). It can be a connection between datasets or between models. A dataset is used in machine learning for model training, validation, and testing; the quantity and quality of the dataset affect the effectiveness of machine learning. In this embodiment, the dataset can be used for model training.

[0202] The following section uses CSI compressed inference tasks as an example to explain the connection between the dataset and the model.

[0203] Dataset integration mainly refers to the sender sending a dataset to the receiver for model training. Figure 8 is a schematic diagram of dataset integration between the network side and the terminal side.

[0204] As shown in Figure 8, the network can first acquire the dataset through joint training. For example, the network can use a virtual CSI compression model to compress the target CSI, and then use a CSI reconstruction model to decompress the target CSI. That is, the target CSI is the input to the CSI compression model, and the output of the CSI compression model can be the compressed CSI. This compressed CSI can be the input to the CSI reconstruction model, and the CSI reconstruction model can output the reconstructed CSI.

[0205] Optionally, the compressed CSI can be further quantized to obtain quantized CSI. During the model training phase, both the compressed CSI and the quantized CSI can be used as CSI feedback information.

[0206] For example, the dataset sent from the network side to the terminal side may include one or more of the following: target CSI, CSI feedback information, or reconstructed CSI. Optionally, the dataset may also include a quantization method. In other words, the quantization method is also indicated by the network side. Optionally, the quantization method is predefined, such as protocol predefined. The quantization method is used to determine how to quantize the compressed CSI and / or how to dequantize the quantized CSI. The quantization method can also be replaced by a dequantization method, a quantizer, or a dequantizer.

[0207] The network side can distribute this dataset to the terminal side. The terminal side can then train a model based on the received dataset to obtain a CSI compressed model for the terminal side.

[0208] Model interoperability mainly refers to the sender sending model files and / or model parameters to the receiver for model deployment, development (e.g., model training), and deployment. Figure 9 is a schematic diagram of model interoperability between the network side and the terminal side.

[0209] As shown in Figure 9, the network side can first obtain the CSI compressed model through joint training. The specific process of network-side joint training can be found in the explanation above combined with Figure 8, and will not be repeated here. The network side can then distribute the CSI compressed model or CSI reconstructed model obtained through joint training to the terminal side. For example, the network side can send the model file and / or model parameters of the CSI compressed model to the terminal side, or send the model file and / or model parameters of the CSI reconstructed model to the terminal side. The terminal side can then train the model based on the received model file and / or model parameters.

[0210] It should be understood that the CSI compression model shown above in conjunction with Figures 8 and 9 can be an encoder, and the CSI reconstruction model can be a decoder.

[0211] As mentioned earlier, dual-end AI models are typically co-trained and can be used interchangeably. To support the integration and development / deployment of dual-end AI models, a dataset or model needs to be provided to one or both of the models. For example, the network side can provide a dataset to the terminal side for model development and / or deployment; conversely, the network side can provide a model to the terminal side for model development and / or deployment. As can be seen from the description above combined with Figures 8 and 9, the data contained in the dataset is not fixed; datasets may be of various types, and different types of datasets contain different data. The models used for training may also be of various types, such as CSI compressed models, CSI reconstructed models, or even other models. Corresponding to different datasets or different models, there are also various ways for the terminal side to train the model, and the corresponding performance metrics and indicators can also vary. Even with the same dataset or the same model, there may be multiple ways for the terminal side to train the model, and the corresponding performance metrics and indicators may also vary. If the training method, performance metrics, and indicators used during training on the terminal side do not match the received dataset and / or model, the trained model may not meet the requirements or may even be unusable. Therefore, how to train a model on the terminal side based on the received dataset and / or model has become a technical problem that urgently needs to be solved.

[0212] In view of this, this application provides a method for associating datasets and / or models used for model training with metrics, and associating metrics with inference tasks. The network side can not only distribute datasets and / or models, but also distribute the associated metrics to the terminal side. This allows the terminal side to train the model using appropriate training methods based on the received datasets and / or models, as well as the associated metrics, to obtain a model that meets the requirements, thereby improving the feedback performance of CSI.

[0213] Furthermore, the network side can provide multiple sets of metrics for each dataset or model, which means that the terminal side can use various training methods for model training. The terminal side can train the model based on one or more of the multiple sets of metrics, using one or more corresponding training methods. Therefore, the terminal side has a high degree of freedom, and can choose appropriate metrics and training methods for model training according to business needs, device capabilities, etc.

[0214] Furthermore, the multiple sets of metrics in this application may include one or more of the following categories: performance metrics, metrics of training time, or metrics of training resources. That is, the model training process is monitored from one or more dimensions of performance, time, or resources, so as to flexibly respond to different business needs.

[0215] The method provided in this application will now be described in detail with reference to the accompanying drawings.

[0216] Figure 10 is a schematic flowchart of the communication method 1000 provided in an embodiment of this application. The method shown in Figure 10 illustrates the flow of the method 1000 from the perspective of interaction between the terminal side and the network side. More detailed descriptions of the terminal side and the network side can be found above and will not be repeated here.

[0217] The method 1000 shown in Figure 10 may include steps 1010 to 1020. Optionally, the method may also include one or more of steps 1030 to 1070. The various steps in method 1000 are described in detail below.

[0218] In step 1010, the terminal side acquires a dataset and / or a model, which is used for model training.

[0219] In the embodiments of this application, the device that performs model training on the terminal side can be referred to as a training device. The training device can be a terminal device, or it can be a device other than a terminal device, such as the aforementioned OTT system server. This application does not limit this.

[0220] In one example, the training device is a terminal device, and step 1010 may specifically include: the terminal device receiving a dataset and / or model from a network device, such as receiving it over an air interface; or, the terminal device obtaining a dataset and / or model from an OTT system server, which is received by the OTT system from a smart network element on the network side, such as receiving it through a wired network.

[0221] In another example, the training device is an OTT system server, and step 1010 may specifically include: the OTT system server receiving a dataset and / or model from a smart network element, such as receiving it through a wired network; or, the OTT system server obtaining a dataset and / or model from a terminal device, which is received by the terminal device from a network device, such as receiving it through an air interface.

[0222] In summary, whether the terminal device obtains the dataset and / or model or the OTT system server obtains the dataset and / or model, the dataset and / or model are received from the network side. Therefore, one possible implementation of step 1010 is that the terminal side receives the dataset and / or model from the network side, and correspondingly, the network side sends the dataset and / or model to the terminal side.

[0223] Sending the dataset can refer to sending the dataset via a message bearer. Sending the model can refer to sending the model file and / or model parameters via a message bearer. The model file can indicate the model structure, such as the aforementioned CNN or RNN; the model parameters can indicate parameters such as the weights of each neuron in the model. Based on the model file and / or model parameters, the model can be determined.

[0224] The network side can send datasets and / or models to the terminal side through one or more messages, and this application does not limit this. For example, the network device can send the dataset through one message and the model through another message; or, for another example, the network side can send the dataset through one or more messages without sending the model; or, for yet another example, the network side can send the model through one or more messages without sending the dataset; or, for yet another example, the network side can send both the dataset and the model through one message; and so on, without limitation. The figure is only an example, showing the process of the network side sending the dataset and / or model to the terminal side through one step, and should not constitute any limitation on this application.

[0225] In addition, for ease of distinction and explanation, the dataset received from the network side will be referred to as the first dataset, and the dataset obtained by the terminal side through other means mentioned below for training will be referred to as the second dataset. The other means include, but are not limited to, measurement, simulation, training generation, etc.

[0226] In step 1020, the terminal side performs model training based on the first metric, as well as the dataset and / or model, to obtain the first model.

[0227] The first model is the model that the terminal side hopes to obtain through model training. The inference task of the first model can be to compress the measured CSI. In this application, the target CSI can be used to represent the full information of CSI, and the CSI measured by the terminal side is the full information of CSI. Therefore, it can also be said that the inference task of the first model is to compress the target CSI.

[0228] It should be understood that the inference task of the first model is to compress the target CSI, but this does not mean that the first model is only used for compressing the target CSI. In some possible designs, the first model can also be used to quantize the compressed CSI, in which case the first model may also include a quantizer.

[0229] As mentioned earlier, there are multiple ways for the terminal to train the model based on the acquired dataset and / or model. For example, the terminal can directly train the first model to obtain the first model; or, the terminal can first train other models (such as the second model in the example below), and then use the other trained models to train the first model. Corresponding to different training methods, there can also be multiple sets of metrics used for model training. These multiple sets of metrics can correspond to the dataset and / or model acquired by the terminal and to the inference task of the first model. Therefore, the terminal can determine (or select) one or more sets of metrics from these multiple sets of metrics for model training based on the acquired dataset and / or model. In other words, the first metric is one or more sets of metrics among multiple sets of metrics.

[0230] In the embodiments of this application, the aforementioned multiple sets of indicators may include one or more of the following categories of indicators: performance indicators, indicators of training time required, or indicators of training resources required. Alternatively, the multiple sets of indicators may belong to one or more of the following categories of indicators: performance indicators, indicators of training time required, or indicators of training resources required. Or, the multiple sets of indicators may include indicators of one or more of the following dimensions: performance indicators, indicators of training time required, or indicators of training resources required.

[0231] The performance metrics are primarily defined from the perspective of the accuracy of the model's output. The terminal side can train the model based on these performance metrics, ensuring that the trained model's output meets performance requirements, thereby improving the performance of CSI feedback. Specific examples of these performance metrics will be provided later in conjunction with various training methods; they will not be detailed here.

[0232] The metric for training time required is primarily defined from the perspective of training duration. For example, the metric for training time required can include an upper bound on the duration. The terminal can monitor the training time required using the training method corresponding to that metric, or it can predict the training time required using the training method corresponding to that metric. If the actual training time or the predicted time does not exceed the upper bound, model training can be performed based on that metric and its corresponding training method. If the actual training time or the predicted time exceeds the upper bound, other training methods and metrics can be used for model training, or model training can be paused, and so on. This avoids the waste of resources caused by long training periods.

[0233] The metrics for training resources are primarily defined from the perspective of the resources required for model training. These resources may include, for example, computational resources. For instance, the metrics for training resources may include a lower bound on the resources. On the terminal side, if the resources available for model training from the training device are higher than this lower bound, model training can be performed based on this lower bound and the corresponding training method; if the resources available from the training device are lower than this lower bound, other training methods can be used, or model training can be temporarily suspended, etc., thereby helping to avoid resource waste caused by inappropriate training methods.

[0234] It should be noted that the descriptions of indicators above are distinguished by terms such as "individual," "group," and "class." Each indicator can refer to one or more specific thresholds (or limit values) corresponding to the same function, such as the maximum threshold (or upper bound) and / or the minimum threshold (or lower bound). An indicator can also be called a single indicator. Each group of indicators can refer to indicator values ​​corresponding to one or more functions within the same dimension. In other words, each group of indicators can include one or more indicators that are indicators within the same dimension. Or, each group of indicators belongs to a class of indicators. Each class of indicators can refer to multiple groups of indicators corresponding to indicators within the same dimension. Each class of indicators can include multiple groups of indicators, corresponding to various different training methods.

[0235] The model training process will be explained in detail below with several examples and accompanying figures. For ease of distinction and explanation, the models and datasets mentioned below will be briefly described.

[0236] First Model: The model that the terminal side hopes to obtain through model training. This first model can be used to perform inference for target CSI compression, or in other words, the inference task of this first model is to compress the target CSI. As an example, the first model is an encoder.

[0237] Second Model: A model jointly trained with the first model. This second model can be used to perform inference that decompresses CSI feedback information; in other words, the inference task of the second model is to decompress CSI feedback information. As an example, the second model is a decoder. It should be understood that the inference task of the second model is to decompress CSI feedback information, but this does not mean that the second model is only used for decompression. In some possible designs, the second model can also be used to dequantize quantized CSI (i.e., a possible form of CSI feedback information), in which case the second model may also include a dequantizer.

[0238] The third model: The model distributed from the network side. This third model can be used to perform inference for compressing the target CSI; in other words, the inference task of this third model is to compress the target CSI. In other words, the inference task of this third model is the same as that of the first model. For example, the third model is an encoder. It should be understood that, similar to the first model, the inference task of the third model is to compress the target CSI, but this does not mean that the third model is only used for compressing the target CSI. In some possible designs, the third model can also be used to quantize the compressed CSI; in this case, the third model may also include a quantizer.

[0239] In addition, to distinguish between the network-side and terminal-side models, the network-side models will be identified by "1" in the following text and figures, such as encoder 1 and decoder 1 on the network side; and the terminal-side models will be identified by "2", such as encoder 2 and decoder 2 on the terminal side.

[0240] The first dataset: the dataset distributed by the network side.

[0241] The second dataset: The dataset that the terminal side trains itself based on the model and / or the first dataset sent by the network side.

[0242] The third dataset: datasets obtained from the terminal side, such as those obtained through measurement, simulation, or generated by training a third model issued by the network side, etc. This application does not limit this.

[0243] Example 1: The network side sends a first dataset, which includes the target CSI (e.g., denoted as V1) and CSI feedback information (e.g., denoted as C1). The terminal side trains the model based on the first dataset. The CSI feedback information can be compressed CSI (i.e., unquantized CSI) or quantized CSI; there are no restrictions.

[0244] Based on the first dataset, the terminal side can use the following training methods to train the model.

[0245] Training Method 1: The terminal side first trains the second model (such as decoder 2), and then trains the first model (such as encoder 2) based on the second model.

[0246] In one possible design (hereinafter referred to as Design 1), the encoder 2 to be trained on the terminal side can be used for compression and quantization, or in other words, encoder 2 has the function of compression and quantization. The CSI feedback information in the first dataset can be quantized CSI. In this case, the inference task of the first model can also be said to be used for compression and quantization of the target CSI.

[0247] In another possible design (hereinafter referred to as Design 2), the encoder 2 to be trained on the terminal side can be used for compression but not for quantization; in other words, encoder 2 has compression functionality but not quantization functionality. The CSI feedback information in the first dataset can be either quantized CSI or compressed CSI.

[0248] It should be understood that whether an encoder has quantization functionality (or whether it includes a quantizer) can be predefined, such as by protocol predefinition, or indicated by the network side, or determined by the terminal side itself. This application does not limit this.

[0249] The following sections will explain these two possible designs.

[0250] In Design 1, the encoder can be used for compression and quantization, and correspondingly, the decoder can be used for dequantization and decompression. The CSI feedback information in the first dataset can be quantized CSI.

[0251] The process of model training on the terminal side can be seen in Figure 11, and may include the following steps:

[0252] Training 1-1: Train decoder 2 independently based on the first dataset. Decoder 2 can be a virtual decoder. On the terminal side, the CSI feedback information (i.e., C1) from the first dataset is used as input to decoder 2, and the target CSI (i.e., V1) from the first dataset is used as the label to train decoder 2. Decoder 2 can be used to dequantize and decompress the CSI feedback information (i.e., C1), outputting the reconstructed CSI (for ease of distinction and explanation, the reconstructed CSI output by decoder 2 in Training 1-1 is denoted as...). By training decoder 2, the output of decoder 2 (i.e., The training 1-1 tends to be close to the label (i.e., V1). For example, training 1-1 can be achieved through supervised learning for V1.

[0253] Training 1-2: Train encoder 2 based on decoder 2. Encoder 2 can be a model built by the terminal side, for example, the structure of which is designed by the terminal side. The terminal side can fix the parameters of decoder 2, that is, use frozen decoder 2, and use the target CSI (for easy distinction from the target CSI in training 1-1, for example, denoted as V) as the input of encoder 2 to perform end-to-end training on encoder 2 and decoder 2. It should be understood that the target CSI used in training 1-2 can be the target CSI from the first dataset on the network side (i.e., V can also be V1), or it can be the target CSI from the third dataset obtained by the terminal side. This application does not limit this.

[0254] For example, the end-to-end training process is as follows: The target CSI (i.e., V) is input to encoder 2, which can be used to compress and quantize the target CSI (i.e., V) and output CSI feedback information (i.e., the quantized CSI); the output of encoder 2 is used as the input of decoder 2, which can be used to dequantize and decompress the CSI feedback information output by encoder 2 and output the reconstructed CSI (for ease of distinction and explanation, the reconstructed CSI output by decoder 2 in training 1-2 is denoted as...). End-to-end training of encoder 2 and decoder 2 can also be called joint training of encoder 2 and decoder 2. The labels for this joint training can be the target CSI from the first dataset or the target CSI from the third dataset. Through end-to-end training, the output of decoder 2 (i.e., The label (i.e., V) tends to be close to the actual label. For example, training 1-2 can be achieved through self-supervised learning for V.

[0255] In Design 2, the encoder can be used for compression, not quantization. Correspondingly, the decoder can be used for decompression, not dequantization. The CSI feedback information can be either compressed CSI or quantized CSI. The corresponding processes differ depending on the type of CSI feedback information.

[0256] If the CSI feedback information is compressed CSI, the model training process on the terminal side is similar to the process shown in Figure 11 above. The difference is that decoder 2 is used to decompress the CSI feedback information, so the CSI feedback information input to decoder 2 is compressed CSI; encoder 2 is used to compress the target CSI, so the CSI feedback information output by encoder 2 is compressed CSI. Other processes can be referred to the relevant descriptions in Figure 11 above, and will not be repeated here.

[0257] If the CSI feedback information is quantized CSI, the process of model training on the terminal side can include the following steps:

[0258] Training 1-0: Dequantize the CSI feedback information in the first dataset to obtain compressed CSI including quantization loss.

[0259] Training 1-1': Based on the compressed CSI obtained from dequantization, including quantization loss, and the target CSI (i.e., V1) in the first dataset, decoder 2 is trained separately. On the terminal side, the compressed CSI with quantization loss can be used as input to decoder 2, and the target CSI (i.e., V1) in the first dataset can be used as the label to train decoder 2. Decoder 2 can be used to decompress the input compressed CSI and output the reconstructed CSI (i.e., V1). By training decoder 2, the output of decoder 2 (i.e., The training 1-1 tends to be close to the label (i.e., V1). For example, training 1-1 can be achieved through supervised learning for V.

[0260] Training 1-2': Train encoder 2 based on decoder 2. Encoder 2 can be a model built by the terminal itself, for example, the structure of which is designed by the terminal. For instance, the terminal can fix the parameters of decoder 2, i.e., use a fixed decoder 2, and use the target CSI (i.e., V) as the input and training label of encoder 2 to perform end-to-end training of encoder 2 and decoder 2. End-to-end training can reuse the process provided in Training 1-2 above; in this case, Training 1-2' is the same as Training 1-2. Alternatively, the end-to-end training process can differ from that provided in Training 1-2: The target CSI (i.e., V) is input to encoder 2, which compresses the target CSI (i.e., V) to obtain a compressed CSI; the compressed CSI is quantized to obtain a quantized CSI; the quantized CSI is dequantized to obtain a compressed CSI containing quantization loss; the compressed CSI containing quantization loss is used as input to decoder 2, which decompresses the compressed CSI containing quantization loss to obtain a reconstructed CSI (i.e., V). Through end-to-end training, the output of decoder 2 (i.e., The label (i.e., V) tends to be close to the label. For example, training 1-2 can be achieved through self-supervised learning for V.

[0261] Similar to Training 1-2, the target CSI used in Training 1-2' can be the target CSI from the first dataset on the network side (i.e., V can also be V1), or it can be the target CSI from the third dataset obtained by the terminal side itself. This application does not limit this.

[0262] In the training process provided by Design 2, quantization and dequantization operations can be implemented by modules other than the encoder and decoder. For example, quantization can be implemented by a quantizer, and dequantization can be implemented by a dequantizer. The quantizer and / or dequantizer can be predefined, such as protocol predefined, or indicated by the network side; this application does not limit this. Optionally, the first dataset also includes a quantizer and / or a dequantizer. Since dequantization and quantization are relative, one possible implementation of predefined quantizers and dequantizers is to predefine either a quantizer or a dequantizer, and another possible implementation of network-indicated quantizers and dequantizers is to indicate either a quantizer or a dequantizer on the network side.

[0263] Corresponding to training method one, the metrics used for model training can be one or more of the following categories: performance metrics, metrics for training time, or metrics for training resources. These categories will be explained in detail below.

[0264] 1) Performance indicators:

[0265] The set of functions used to monitor model performance (referred to as function set 1 for ease of distinction and explanation) may include one or more of the following functions: NMSE between the output of decoder 2 (i.e., an example of the second model) and the label, MSE between the output of decoder 2 and the label, L1Loss between the output of decoder 2 and the label, GCS between the output of decoder 2 and the label, SGCS between the output of decoder 2 and the label, or a weighted sum of two or more of the above. The weighting coefficients for each term may be predefined or indicated by the network side; this application does not impose any limitations on this.

[0266] The set of performance metrics corresponding to the aforementioned function set can be the range that the values ​​of each function in the function set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: NMSE between the output of decoder 2 and the label, MSE between the output of decoder 2 and the label, L1Loss between the output of decoder 2 and the label, GCS between the output of decoder 2 and the label, SGCS between the output of decoder 2 and the label, or a weighted sum of two or more of the above. The weighting coefficients for each item can be predefined or indicated by the network side; this application does not limit this.

[0267] As seen in Training 1-1 and Training 1-2 above, the output of Decoder 2 is not necessarily the same. Therefore, corresponding to Training 1-1, the set of functions used to monitor model performance (denoted as Function Set 1-1 for ease of distinction and explanation) may include one or more of the following functions:

[0268] The output of decoder 2 (i.e., The NMSE between the label (i.e., V1) and the tag (i.e., V1) is...

[0269] The output of decoder 2 (i.e., The MSE between the label (i.e., V1) and the tag (i.e., V1) is...

[0270] The output of decoder 2 (i.e., The L1 loss between the label (i.e., V1) and the tag (i.e., V1) is...

[0271] The output of decoder 2 (i.e., The GCS between the label (i.e., V1) and the tag (i.e., V1)

[0272] The output of decoder 2 (i.e., The SGCS between the label (i.e., V1) and the tag (i.e., V1) is... or,

[0273] The output of decoder 2 (i.e., The weighted sum of two or more of the following: NMSE, MSE, L1Loss, GCS, or SGCS between the label (i.e., V1) and the tag (i.e., V1). or The weighted sum of two or more terms.

[0274] The set of performance metrics corresponding to the above set of functions can be the range that the values ​​of each function in the set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: Or a weighted sum of two or more of the above. For ease of distinction and explanation, this set of performance indicators will be referred to as indicator group 1-1 below.

[0275] Corresponding to Training 1-2, the set of functions used to monitor model performance (denoted as Function Set 1-2 for ease of distinction and explanation) may include one or more of the following functions:

[0276] The output of decoder 2 (i.e., The NMSE between the label (i.e., V) and the tag (i.e., V) is...

[0277] The output of decoder 2 (i.e., The MSE between the label (i.e., V) and the tag (i.e., V) is...

[0278] The output of decoder 2 (i.e., The L1 loss between the label (i.e., V) and the tag (i.e., V) is...

[0279] The output of decoder 2 (i.e., The GCS between the label (i.e., V) and the tag (i.e., V)

[0280] The output of decoder 2 (i.e., The SGCS between the label (i.e., V) and the tag (i.e., V) or,

[0281] The output of decoder 2 (i.e., The weighted sum of two or more of the following: NMSE, MSE, L1Loss, GCS, or SGCS between the label (i.e., V) and the tag (V). or The sum of two or more terms. The weighting coefficients of each term can be predefined or indicated by the network side, and this application does not limit them.

[0282] The set of performance metrics corresponding to the above set of functions can be the range that the values ​​of each function in the set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: Or a weighted sum of two or more of the above. For ease of distinction and explanation, this set of performance indicators will be referred to as indicator group 1-2 below.

[0283] It should be understood that the functions listed above are merely examples and should not be construed as limiting this application. Those skilled in the art can perform simple mathematical transformations based on the functions listed above to obtain more functions; they can also use other functions to replace the functions listed above, based on the same concept. For example, Alternatively, they can be replaced with: 1 and The absolute value of the difference (i.e., 1 and The absolute value of the difference (i.e., ); Alternatively, they can be replaced with: 1 and The absolute value of the difference (i.e., ), 1 and The absolute value of the difference (i.e., This application includes, but is not limited to, etc.

[0284] As can be seen, the function sets corresponding to training method one include function set 1-1 and function set 1-2, and the corresponding performance metrics include metric set 1-1 and metric set 1-2. Function set 1-1 and metric set 1-1 correspond to training 1-1, and function set 1-2 and metric set 1-2 correspond to training 1-2. Therefore, the set of metrics corresponding to the first dataset can be: {metric set 1-1}, or {metric set 1-2}, or {metric set 1-1, metric set 1-2}, and the function set corresponding to the first dataset can be: {function set 1-1}, or {function set 1-2}, or {function set 1-1, function set 1-2}.

[0285] For example, if the terminal side wants to monitor training 1-1, it can train the model based on function set 1-1 and indicator group 1-1; if the terminal side wants to monitor training 1-2, it can train the model based on function set 1-2 and indicator group 1-2; if the terminal side wants to monitor the entire training process, it can train the model for training 1-1 based on function set 1-1 and indicator group 1-1, and train the model for training 1-2 based on function set 1-2 and indicator group 1-2.

[0286] 2) Metrics for training duration:

[0287] Since training method one includes two phases, training 1-1 and training 1-2, the required training time can specifically refer to the total time required for training 1-1 and training 1-2. The metrics for the required training time can include one or more of the following: the metrics for the required time of training 1-1, the metrics for the required time of training 1-2, or the metrics for the total time required for training 1-1 and training 1-2. That is, the metrics for the required training time can include one or more of the following: the metrics for the required time of training 1-1 (e.g., denoted as T). 1-1 The metric for the time required to train 1-2 (e.g., denoted as T) 1-2 Alternatively, a metric for the total time required to train 1-1 and 1-2 (e.g., denoted as T1). Therefore, a set of metrics corresponding to the first dataset could be: {T} 1-1}, or {T 1-2}, or {T 1-1 T 1-2}, or {T1}. For example, the terminal side can monitor the duration of training 1-1, or it can monitor the duration of training 1-2, or it can monitor both the duration of training 1-1 and the duration of training 1-2, or it can monitor the total duration of training 1-1 and training 1-2.

[0288] 3) Metrics of resources required for training:

[0289] Since training method one includes two phases, training 1-1 and training 1-2, the resources required for training can specifically refer to the resources needed for training 1-1 and training 1-2. The metrics for these resources can include one or more of the following: metrics for the resources required for training 1-1, metrics for the resources required for training 1-2, or metrics for the total resources required for training 1-1 and training 1-2. That is, the metrics for these resources can include one or more of the following: metrics for the resources required for training 1-1 (e.g., denoted as R). 1-1 The metrics for the resources required to train 1-2 (e.g., denoted as R). 1-2 Or, a metric for the total resources required for training 1-1 and training 1-2 (e.g., denoted as R1). Therefore, a set of metrics corresponding to the first dataset could be: {R...} 1-1}, or {R 1-2}, or {R 1-1 R 1-2}, or {R1}. For example, the resource size used by training 1-1 can be monitored, or the resource size used by training 1-2 can be monitored, or both the resource sizes used by training 1-1 and training 1-2 can be monitored, or the total resource size used by training 1-1 and training 1-2 can be monitored.

[0290] As can be seen, the aforementioned performance metrics, training time metrics, and training resource metrics are all directly or indirectly related to the inference task of the first model. These metrics can be used to monitor decoder-only training, end-to-end training, and both decoder-only and end-to-end training. Therefore, it can be said that the aforementioned performance metrics, training time metrics, and training resource metrics correspond to the inference task of the first model.

[0291] Training Method 2: Train the first model (such as encoder 2) on the terminal side, or in other words, train encoder 2 alone.

[0292] In one possible design (hereinafter referred to as Design 1), the encoder can be used for compression and quantization, or in other words, the encoder has the functions of compression and quantization. The CSI feedback information in the first dataset can be quantized CSI.

[0293] In another possible design (hereinafter referred to as Design Two), the encoder can be used for compression but not for quantization; or, the encoder has compression functionality but not quantization functionality. The CSI feedback information in the first dataset can be compressed CSI.

[0294] In another possible design (hereinafter referred to as Design 3), the encoder can be used for compression but not for quantization; or, in other words, the encoder has compression functionality but not quantization functionality. The CSI feedback information in the first dataset can be quantized CSI.

[0295] The following sections will explain these two possible designs.

[0296] In Design 1, the encoder can be used for compression and quantization, and the CSI feedback information in the first dataset can be quantized CSI.

[0297] The model training process on the terminal side is shown in Figure 12, as follows: The target CSI (i.e., V1) from the first dataset is used as the input to encoder 2, and the CSI feedback information (i.e., C1) from the first dataset is used as the label to train encoder 2. Encoder 2 can be used to compress and quantize the target CSI (i.e., V1), and output the CSI feedback information (i.e., C1). It should be understood that in Design 1, the CSI feedback information output by encoder 2 is quantized CSI. Through training encoder 2, the CSI feedback information output by encoder 2 (i.e., The value tends to be close to the label (i.e., C1). For example, this training can be achieved through supervised learning for C1.

[0298] In Design 2, the encoder can be used for compression but not for quantization, and the CSI feedback information in the first dataset can be compressed CSI.

[0299] The process of model training on the terminal side is similar to that shown in Figure 11 above. The difference is that encoder 2 is used to compress the target CSI, so the CSI feedback information output by encoder 2 is the compressed CSI. Other processes can be found in the relevant descriptions in Figure 12 above, and will not be repeated here.

[0300] In Design 3, the encoder can be used for compression but not for quantization, and the CSI feedback information in the first dataset can be quantized CSI.

[0301] The process of training a model on the terminal side can include the following steps:

[0302] Training 2-0: Dequantize the CSI feedback information in the first dataset to obtain compressed CSI with quantization loss. To distinguish it from the CSI feedback information in the first dataset (i.e., C1), the compressed CSI with quantization loss obtained by dequantizing the CSI feedback information is denoted as C1'.

[0303] Training 2-1: Using the target CSI (V1) from the first dataset as input to encoder 2, and the compressed CSI (C1') obtained from dequantization in Training 2-0 (including quantization loss) as the label, encoder 2 is trained. Encoder 2 can be used to compress the target CSI (V1) and output the compressed CSI. Through training encoder 2, the compressed CSI output by encoder 2 (e.g., denoted as...) becomes... The CSI and the label (i.e., C1') tend to be close. For example, this training can be achieved through supervised learning of compressed CSI with quantization loss, which is obtained by dequantizing the CSI feedback information in the first dataset, and therefore can also be called supervised learning of CSI feedback information.

[0304] In the training process provided in Design 2, the dequantization operation can be implemented by modules other than the encoder and decoder. For example, dequantization can be implemented by a dequantizer. The dequantizer can be predefined, such as protocol predefined, or it can be indicated by the network side; this application does not limit this. Optionally, the first dataset also includes a dequantizer. Since dequantization and quantization are relative, another possible implementation of the predefined dequantizer is a predefined quantizer, and another possible implementation of the network-indicated quantizer and dequantizer is a network-indicated quantizer.

[0305] Corresponding to training method two, the metrics used for model training can be one or more of the following categories: performance metrics, metrics for training time, or metrics for training resources. These categories of metrics will be explained in detail below.

[0306] 1) Performance indicators:

[0307] The set of functions used to monitor model performance may include one or more of the following functions: NMSE between the output of encoder 2 (i.e., an example of the first model) and the label, MSE between the output of encoder 2 and the label, L1Loss between the output of encoder 2 and the label, GCS between the output of encoder 2 and the label, SGCS between the output of encoder 2 and the label, or a weighted sum of two or more of the above. The weighting coefficients for each term may be predefined or indicated by the network side; this application does not limit this.

[0308] The set of performance metrics corresponding to the aforementioned function set can be the range that the values ​​of each function in the set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: NMSE between the output of encoder 2 and the tag, MSE between the output of encoder 2 and the tag, L1Loss between the output of encoder 2 and the tag, GCS between the output of encoder 2 and the tag, SGCS between the output of encoder 2 and the tag, or a weighted sum of two or more of the above. The weighting coefficients for each item can be predefined or indicated by the network side; this application does not limit this.

[0309] The performance of this model can be monitored by monitoring the difference between the output of encoder 2 and the label.

[0310] As seen in Design 1 and Design 2 above, the output of encoder 2 is not necessarily the same, and the labels used for training are also different. In Design 1, the label is the CSI feedback information C1, where C1 represents the quantized CSI. The CSI feedback information output by encoder 2 is different. This represents the quantized CSI; in Design 2, the label is CSI feedback information C1, where C1 represents compressed CSI, and the CSI feedback information output by encoder 2. The compressed CSI is represented. The set of functions used to monitor model performance (referred to as function set 2 for ease of distinction and explanation) may include one or more of the following functions:

[0311] CSI feedback information output by encoder 2 (i.e., The MSE between the CSI feedback information (i.e., C1) and the CSI feedback information (i.e., C1) is...

[0312] CSI feedback information of encoder 2 output (i.e., The NMSE between the CSI feedback information (i.e., C1) and the CSI feedback information (i.e., C1) is...

[0313] CSI feedback information output by encoder 2 (i.e., The L1 loss between the CSI feedback information (i.e., C1) and the CSI feedback information (i.e., C1) is...

[0314] CSI feedback information output by encoder 2 (i.e., The GCS between the CSI feedback information (i.e., C1) and the CSI feedback information (i.e., C1), i.e.

[0315] CSI feedback information output by encoder 2 (i.e., The SGCS between the CSI feedback information (i.e., C1) and the CSI feedback information (i.e., C1) or

[0316] CSI feedback information output by encoder 2 (i.e., The weighted sum of two or more of the following: NMSE, MSE, L1Loss, GCS, or SGCS between the CSI feedback information (i.e., C1) and the CSI feedback information (i.e., C1). or The sum of two or more terms. The weighting coefficients of each term can be predefined or indicated by the network side, and this application does not limit them.

[0317] The set of performance metrics corresponding to the above set of functions can be the range that the values ​​of each function in the set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: Or a weighted sum of two or more of the above. For ease of distinction and explanation, this set of performance indicators will be referred to as indicator group 2 below.

[0318] In Design 3, the label is the compressed CSI (i.e., C1') obtained by dequantizing the CSI feedback information, including quantization loss. This is the output of the second model. It is a compressed CSI, a set of functions used to monitor model performance (for ease of distinction and explanation, denoted as function set 2'), which may include one or more of the following functions:

[0319] The compressed CSI output by encoder 2 (i.e., The MSE between the compressed CSI (i.e., C1') and the CSI with quantization loss included is...

[0320] The compressed CSI output of encoder 2 (i.e., The NMSE between the compressed CSI (i.e., C1') and the CSI with quantization loss is...

[0321] The compressed CSI output by encoder 2 (i.e., The L1Loss between the compressed CSI (i.e., C1') and the CSI with quantization loss, i.e.

[0322] The compressed CSI output by encoder 2 (i.e., The GCS between the compressed CSI (i.e., C1') and the compressed CSI with quantization loss, i.e.

[0323] The compressed CSI output by encoder 2 (i.e., The SGCS between the compressed CSI (i.e., C1') and the CSI with quantization loss is... or

[0324] The compressed CSI output by encoder 2 (i.e., The weighted sum of two or more of the following: NMSE, MSE, L1Loss, GCS, or SGCS between the compressed CSI (i.e., C1') and the CSI containing quantization loss: or The sum of two or more terms. The weighting coefficients of each term can be predefined or indicated by the network side, and this application does not limit them.

[0325] The set of performance metrics corresponding to the above set of functions can be the range that the values ​​of each function in the set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: Or a weighted sum of two or more of the above. For ease of distinction and explanation, this set of performance indicators will be referred to as indicator group 2' below.

[0326] It should be understood that function set 2' is a variation of function set 2, and index group 2' is a variation of index group 2.

[0327] It should also be understood that the functions listed above are merely examples and should not be construed as limiting this application. Those skilled in the art can perform simple mathematical transformations based on the functions listed above to obtain more functions; they can also use other functions to replace the functions listed above, based on the same concept. For example, Alternatively, they can be replaced with: 1 and The absolute value of the difference (i.e., ), 1 and The absolute value of the difference (i.e., ); Alternatively, they can be replaced with: 1 and The absolute value of the difference (i.e., ), 1 and The absolute value of the difference (i.e., This application includes, but is not limited to, etc.

[0328] As can be seen, the function set corresponding to training method two includes function set 2 or function set 2', and the corresponding performance metrics include metric set 2 or metric set 2'. Therefore, the set of metrics corresponding to the first dataset can be {metric set 2} or {metric set 2'}, and the function set corresponding to the first dataset can be {function set 2} or {function set 2'}. For example, the terminal side can train the model based on one of the function sets {function set 2} or {function set 2'} and its corresponding set of metrics.

[0329] 2) Metrics for training duration:

[0330] The training time required can specifically refer to the time required to train encoder 2. Correspondingly, the metric for the training time required can be the metric for the time required to train encoder 2 (e.g., denoted as T2). Therefore, a set of metrics corresponding to the first dataset can be {T2}. For example, the terminal side can monitor the training time of encoder 2.

[0331] 3) Metrics of resources required for training:

[0332] The resources required for training can specifically refer to the size of the resources needed to train encoder 2. Correspondingly, the metric for the resources required for training can be the metric for the resources required to train encoder 2 (e.g., denoted as R²). Therefore, a set of metrics corresponding to the first dataset can be {R²}. For example, the terminal side can monitor the resources used to train encoder 2.

[0333] As can be seen, the above-mentioned performance metrics, training time metrics, and training resource metrics are all directly related to the inference task of the first model. Therefore, it can be said that the above-mentioned performance metrics, training time metrics, and training resource metrics correspond to the inference task of the first model.

[0334] Example 2: The network side sends a third model whose inference task is to compress the target CSI. The terminal side can then train a model based on this third model.

[0335] Based on the third model, the terminal side can use the following training methods to train the model.

[0336] Training method 3: The terminal first trains the second model (e.g., decoder 2) based on the third model (e.g., encoder 1), and then trains the first model (e.g., encoder 2) based on the second model.

[0337] In one possible design (hereinafter referred to as Design 1), encoder 1 can be used for compression and quantization, or in other words, encoder 1 has compression and quantization functions. Correspondingly, decoder 2 trained on the terminal side based on encoder 1 can be used for dequantization and decompression, or in other words, decoder 2 has dequantization and decompression functions. Therefore, encoder 2 trained on the terminal side based on decoder 2 can also have compression and quantization functions.

[0338] In another possible design (hereinafter referred to as Design 2), encoder 1 can be used for compression but not for quantization; in other words, encoder 1 has compression functionality but not quantization functionality. Correspondingly, decoder 2 trained on the terminal side based on encoder 1 can be used for decompression but not for dequantization; in other words, decoder 2 has decompression functionality but not dequantization functionality. Therefore, encoder 2 trained on the terminal side based on decoder 2 can also have compression functionality but not quantization functionality.

[0339] The process of model training on the terminal side can be seen in Figure 13, and may include the following steps:

[0340] Training 3-1: Train decoder 2 independently based on encoder 1. The terminal can train decoder 2 independently based on a dataset provided by encoder 1. This dataset can be the first dataset provided by the network side, or a third dataset obtained by the terminal itself, such as a dataset obtained through measurement or simulation, etc. This application does not impose any limitations on this. Decoder 2 can be a virtual decoder. The terminal can use the target CSI (e.g., denoted as V) from the dataset as the input and training labels of encoder 1 to train decoder 2.

[0341] Corresponding to Design 1, encoder 1 can be used to compress the input target CSI and output compressed CSI. The input of encoder 1 can be used as the input of decoder 2, which can be used to decompress the input compressed CSI to obtain the reconstructed CSI (e.g., denoted as ). ).

[0342] Corresponding to Design 2, encoder 1 can be used to compress and quantize the input target CSI (i.e., V), outputting the quantized CSI. The input of encoder 1 can be used as the input of decoder 2, which can be used to dequantize and decompress the input quantized CSI to obtain the reconstructed CSI (i.e., V). ).

[0343] By training decoder 2, the output of decoder 2 (i.e., The ) and the label (i.e., V) tend to be similar.

[0344] Training 3-2: Train encoder 2 based on decoder 2. Encoder 2 can be a model built by the terminal itself, for example, the structure of which is designed by the terminal. The terminal can fix the parameters of decoder 2, i.e., use a frozen decoder 2, and use the target CSI (i.e., V) in the dataset as the input of encoder 2 to perform end-to-end training, or joint training. For the end-to-end training process, please refer to the detailed explanation of combining design 1 and design 2 in training method 1 above, which will not be repeated here. The output of decoder 2 in training 3-2 is the reconstructed CSI (for convenience and explanation, the output of decoder 2 in training 3-2 is denoted as...). End-to-end training of encoder 2 and decoder 2 can also be called joint training of encoder 2 and decoder 2. The labels for this joint training can be the target CSI from the first dataset or the target CSI from the third dataset. Through end-to-end training, the output of decoder 2 (i.e., The ) and the label (i.e., V) tend to be similar.

[0345] For example, both training 3-1 and training 3-2 can be achieved through self-supervised learning for V.

[0346] As mentioned earlier, the dataset used for training can be a first dataset from the network side. Therefore, in training method three, the network side can distribute not only the third model but also the first dataset. Optionally, the terminal side obtains the dataset and / or model, including: the terminal side receives the first dataset and the third model, wherein the first dataset includes the target CSI, and the inference task of the third model is to compress the target CSI. Correspondingly, the network side sends the dataset and / or model, including: the network side sends the first dataset and the third model.

[0347] In Design 1, the quantization function can be implemented by a quantizer, and the dequantization function can be implemented by a dequantizer. The quantizer can be, for example, a module in the encoder capable of quantization, and the dequantizer can be, for example, a module in the decoder capable of dequantization. The quantizer and / or dequantizer can be predefined, such as protocol predefined, or indicated by the network side; this application does not limit this. Optionally, the terminal side acquires the dataset and / or model, including: the terminal side receives a first dataset and a third model, the first dataset further including a quantizer and / or a dequantizer, the inference task of which is to compress the target CSI. Correspondingly, the network side sends the dataset and / or model, including: the network side sends the first dataset and the third model. Since dequantization and quantization are relative, one possible implementation of predefined quantizers and dequantizers is to predefine either a quantizer or a dequantizer, and one possible implementation of the network side instructing the quantizer and dequantizer is to instruct either a quantizer or a dequantizer.

[0348] In one possible design, the first dataset mentioned above includes one or more of the following: target CSI, quantizer, or dequantizer.

[0349] Corresponding to training method one, the metrics used for model training can be one or more of the following categories: performance metrics, metrics for training time, or metrics for training resources. These categories will be explained in detail below.

[0350] 1) Performance indicators:

[0351] The set of functions used to monitor model performance (referred to as function set 3 for ease of distinction and explanation) may include one or more of the following functions: NMSE between the output of decoder 2 (i.e., an example of the second model) and the label, MSE between the output of decoder 2 and the label, L1Loss between the output of decoder 2 and the label, GCS between the output of decoder 2 and the label, SGCS between the output of decoder 2 and the label, or a weighted sum of two or more of the above. The weighting coefficients for each term may be predefined or indicated by the network side; this application does not impose any limitations on this.

[0352] The set of performance metrics corresponding to the aforementioned function set can be the range that the values ​​of each function in the function set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: NMSE between the output of decoder 2 and the label, MSE between the output of decoder 2 and the label, L1Loss between the output of decoder 2 and the label, GCS between the output of decoder 2 and the label, SGCS between the output of decoder 2 and the label, or a weighted sum of two or more of the above. The weighting coefficients for each item can be predefined or indicated by the network side; this application does not limit this.

[0353] As seen in Training 3-1 and Training 3-2 above, the output of Decoder 2 is not necessarily the same. Therefore, corresponding to Training 3-1, the set of functions used to monitor model performance (denoted as Function Set 3-1 for ease of distinction and explanation) may include one or more of the following functions:

[0354] The output of decoder 2 (i.e., The NMSE between the label (i.e., V) and the tag (i.e., V) is...

[0355] The output of decoder 2 (i.e., The MSE between the label (i.e., V) and the tag (i.e., V) is...

[0356] The output of decoder 2 (i.e. The L1 loss between the label (i.e., V) and the tag (i.e., V) is...

[0357] The output of decoder 2 (i.e., The GCS between the label (i.e., V) and the tag (i.e., V)

[0358] The output of decoder 2 (i.e., The SGCS between the label (i.e., V) and the tag (i.e., V) or,

[0359] The output of decoder 2 (i.e., The weighted sum of two or more of the following: NMSE, MSE, L1Loss, GCS, or SGCS between the label (i.e., V) and the tag (V). The weighted sum of two or more terms.

[0360] The set of performance metrics corresponding to the above set of functions can be the range that the values ​​of each function in the set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: Or a weighted sum of two or more of the above. For ease of distinction and explanation, this set of performance indicators will be referred to as indicator group 3-1 below.

[0361] Corresponding to Training 3-2, the set of functions used to monitor model performance (referred to as Function Set 3-2 for ease of distinction and explanation) may include one or more of the following functions:

[0362] The output of decoder 2 (i.e., The NMSE between the label (i.e., V) and the tag (i.e., V) is...

[0363] The output of decoder 2 (i.e., The MSE between the label (i.e., V) and the tag (i.e., V) is...

[0364] The output of decoder 2 (i.e., The L1 loss between the label (i.e., V) and the tag (i.e., V) is...

[0365] The output of decoder 2 (i.e., The GCS between the label (i.e., V) and the tag (i.e., V)

[0366] The output of decoder 2 (i.e., The SGCS between the label (i.e., V) and the tag (i.e., V) or,

[0367] The output of decoder 2 (i.e., The weighted sum of two or more of the following: NMSE, MSE, L1Loss, GCS, or SGCS between the label (i.e., V) and the tag (V). or The sum of two or more terms. The weighting coefficients of each term can be predefined or indicated by the network side, and this application does not limit them.

[0368] The set of performance metrics corresponding to the above set of functions can be the range that the values ​​of each function in the set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: Or a weighted sum of two or more of the above. For ease of distinction and explanation, this set of performance indicators will be referred to as indicator group 3-2 below.

[0369] It should be understood that the functions listed above are merely examples and should not be construed as limiting this application. Those skilled in the art can perform simple mathematical transformations based on the functions listed above to obtain more functions; they can also use other functions to replace the functions listed above, based on the same concept. For example, Alternatively, they can be replaced with: 1 and The absolute value of the difference (i.e., ), 1 and The absolute value of the difference (i.e., ); Alternatively, they can be replaced with: 1 and The absolute value of the difference (i.e., ), 1 and The absolute value of the difference (i.e., This application includes, but is not limited to, etc.

[0370] As can be seen, the function sets corresponding to training method three include function set 3-1 and function set 3-2, and the corresponding performance metrics include metric set 3-1 and metric set 3-2. Function set 3-1 and metric set 3-1 correspond to training 3-1, and function set 3-2 and metric set 3-2 correspond to training 3-2. Therefore, the set of metrics corresponding to the third model is: {metric set 3-1}, or {metric set 3-2}, or {metric set 3-1, metric set 3-2}, and the function sets corresponding to the third model are: {function set 3-1}, or {function set 3-2}, or {function set 3-1, function set 3-2}.

[0371] For example, if the terminal side wants to monitor training 3-1, it can train the model based on function set 1-1 and indicator group 1-1; if the terminal side wants to monitor training 3-2, it can train the model based on function set 3-2 and indicator group 3-2; if the terminal side wants to monitor the entire training process, it can train the model for training 3-1 based on function set 3-1 and indicator group 3-1, and train the model for training 3-2 based on function set 3-2 and indicator group 3-2.

[0372] 2) Metrics for training duration:

[0373] Since Training Method 3 includes two phases, Training 3-1 and Training 3-2, the required training time can specifically refer to the total time required for Training 3-1 and Training 3-2. The metrics for the required training time can include one or more of the following: the required time for Training 3-1, the required time for Training 3-2, or the total required time for Training 3-1 and Training 3-2. That is, the metrics for the required training time can include one or more of the following: the required time for Training 3-1 (e.g., denoted as T). 3-1 The metrics for the time required to train 3-2 (e.g., denoted as T) 3-2 Or, the metric for the total time required to train 3-1 and 3-2 (e.g., denoted as T3). Therefore, the first set of metrics corresponding to the third model is: {T 3-1}, or {T 3-2}, or {T 3-1 T 3-2}, or {T3}. For example, the terminal side can monitor the duration of training 3-1, or it can monitor the duration of training 3-2, or it can monitor both the duration of training 3-1 and the duration of training 3-2, or it can monitor the total duration of training 3-1 and training 3-2.

[0374] 3) Metrics of resources required for training:

[0375] Since training method three includes two phases, training 1-1 and training 3-2, the resources required for training can specifically refer to the resources needed for training 3-1 and training 3-2. The indicators for these resources can include one or more of the following: the indicators for resources required for training 3-1, the indicators for resources required for training 3-2, or the indicators for the total resources required for training 3-1 and training 3-2. That is, the indicators for these resources can include one or more of the following: the indicators for resources required for training 3-1 (e.g., denoted as R). 3-1 The metrics for the resources required to train 3-2 (e.g., denoted as R). 3-2 Or, the metric for the total resources required to train 3-1 and 3-2 (e.g., denoted as R3). Therefore, the first set of metrics corresponding to the third model is: {R 3-1}, or {R 3-2}, or {R 3-1 R 3-2}, or {R3}. For example, the resource size used for training 3-1 can be monitored, or the resource size used for training 3-2 can be monitored, or the resource size used for training 3-1 and training 3-2 can be monitored, or the total resource size used for training 3-1 and training 3-2 can be monitored.

[0376] As can be seen, the aforementioned performance metrics, training time metrics, and training resource metrics are all directly or indirectly related to the inference task of the first model. These metrics can be used to monitor decoder-only training, end-to-end training, and both decoder-only and end-to-end training. Therefore, it can be said that the aforementioned performance metrics, training time metrics, and training resource metrics correspond to the inference task of the first model.

[0377] Training Method 4: The terminal first generates a second dataset based on the third model (e.g., encoder 1), and then trains the first model (e.g., encoder 2) based on the second dataset.

[0378] In one possible design (hereinafter referred to as Design 1), encoder 1 can be used for compression and quantization, or in other words, encoder 1 has compression and quantization functions. Correspondingly, encoder 2 trained on the terminal side based on encoder 1 can also have compression and quantization functions.

[0379] In another possible design (hereinafter referred to as Design 2), encoder 1 can be used for compression but not for quantization; in other words, encoder 1 has compression functionality but not quantization functionality. Correspondingly, encoder 2 trained on the terminal side based on encoder 1 can also have compression functionality but not quantization functionality.

[0380] The process of model training on the terminal side can be seen in Figure 14, and may include the following steps:

[0381] Step 1: Dataset generation. The terminal can generate a second dataset based on the existing dataset and the received encoder 1. The existing dataset on the terminal can be, for example, the first dataset received from the network side, or a third dataset acquired by the terminal itself, such as a dataset obtained through measurement or simulation, etc. This application does not limit this. The existing dataset on the terminal can include the target CSI (e.g., denoted as V). The terminal can use this target CSI (i.e., V) as input to encoder 1.

[0382] Corresponding to Design 1, encoder 1 can be used to compress and quantize the target CSI to obtain a quantized CSI, which is an example of CSI feedback information.

[0383] Corresponding to Design 2, Encoder 1 can be used to compress the target CSI to obtain compressed CSI, which is another example of CSI feedback information.

[0384] The CSI feedback information obtained in this way can be used as the label for training encoder 2, that is, the second dataset includes CSI feedback information (e.g., denoted as C2).

[0385] Step 2: Training. The terminal can train encoder 2 based on the generated second dataset. The terminal can use the target CSI (i.e., V) as the input to encoder 2 and the CSI feedback information as the label to train encoder 2.

[0386] Corresponding to Design 1, encoder 2 can be used to compress and quantize the target CSI, outputting quantized CSI (i.e., a form of CSI feedback information, for example denoted as...). Corresponding to Design 2, encoder 2 can be used to compress the target CSI and output compressed CSI (i.e., another form of CSI feedback information, for example, also denoted as...). The terminal side can train encoder 2 to make the output of encoder 2 (i.e., The label (i.e., C2) tends to be similar to the label.

[0387] For example, step 2 can be achieved through supervised learning for C2.

[0388] As mentioned earlier, the dataset input to encoder 1 on the terminal side in step 1 can be the first dataset from the network side. Therefore, in training method four, the network side can send not only the third model but also the first dataset. Optionally, the terminal side obtaining the dataset and / or model includes: the terminal side receiving the first dataset and the third model, wherein the first dataset includes the target CSI, and the inference task of the third model is to compress the target CSI. Correspondingly, the network side sending the dataset and / or model includes: the network side sending the first dataset and the third model.

[0389] In Design 1, the quantization function can be implemented by a quantizer, and the dequantization function can be implemented by a dequantizer. The quantizer can be, for example, a module in the encoder capable of quantization, and the dequantizer can be, for example, a module in the decoder capable of dequantization. The quantizer and / or dequantizer can be predefined, such as protocol predefined, or indicated by the network side; this application does not limit this. Optionally, the terminal side acquires the dataset and / or model, including: the terminal side receives a first dataset and a third model, the first dataset further including a quantizer and / or a dequantizer, the inference task of which is to compress the target CSI. Correspondingly, the network side sends the dataset and / or model, including: the network side sends the first dataset and the third model. Since dequantization and quantization are relative, one possible implementation of predefined quantizers and dequantizers is to predefine either a quantizer or a dequantizer, and one possible implementation of the network side instructing the quantizer and dequantizer is to instruct either a quantizer or a dequantizer.

[0390] In one possible design, the first dataset mentioned above includes one or more of the following: target CSI, quantizer, or dequantizer.

[0391] Corresponding to training method four, the metrics used for model training can be one or more of the following categories: performance metrics, metrics for training time, or metrics for training resources. These categories of metrics will be explained in detail below.

[0392] 1) Performance indicators:

[0393] The set of functions used to monitor model performance (referred to as function set 4 for ease of distinction and explanation) may include one or more of the following functions:

[0394] The output of encoder 2 (i.e., The NMSE between the label (i.e., C2) and the tag (i.e., C2) is...

[0395] The output of encoder 2 (i.e., The MSE between the label (i.e., C2) and the tag (i.e., C2) is...

[0396] The output of encoder 2 (i.e., The L1 loss between the label (i.e., C2) and the tag (i.e., C2) is...

[0397] The output of encoder 2 (i.e., The GCS between the label (i.e., C2) and the tag (i.e., C2)

[0398] The output of encoder 2 (i.e., The SGCS between the label (i.e., C2) and the tag (i.e., C2) is... or

[0399] The output of encoder 2 (i.e., The NMSE between the label (i.e., C2) and the label (i.e., C2) is the weighted sum of two or more of the following: NMSE, MSE, L1Loss, GCS, or SGCS. or The sum of two or more terms. The weighting coefficients of each term can be predefined or indicated by the network side, and this application does not limit them.

[0400] The set of performance metrics corresponding to the above set of functions can be the range that the values ​​of each function in the set should satisfy. That is, a set of performance metrics can include the range that the values ​​of one or more of the following functions should satisfy: Or a weighted sum of two or more of the above terms. The weighting coefficients for each term can be predefined or indicated by the network side, and this application does not limit this. For ease of distinction and explanation, this set of performance indicators will be referred to as indicator group 4 below.

[0401] It should be understood that the functions listed above are merely examples and should not be construed as limiting this application. Those skilled in the art can perform simple mathematical transformations based on the functions listed above to obtain more functions; they can also use other functions to replace the functions listed above, based on the same concept. For example, Alternatively, they can be replaced with: 1 and The absolute value of the difference (i.e., ), 1 and The absolute value of the difference (i.e., This application includes, but is not limited to, etc.

[0402] As can be seen, the function set corresponding to training method four includes function set 4, and the corresponding performance metrics include metric group 4. Therefore, the set of metrics corresponding to the third model can be {metric group 4}, and the function set corresponding to the third model can be {function set 4}. For example, the terminal side can perform model training based on {function set 4} and its corresponding set of metrics {metric group 4}.

[0403] 2) Metrics for training duration:

[0404] The training time required can specifically refer to the time needed to train encoder 2, such as the time required for step 2 as shown in the example above, or the total time required for steps 1 and 2 as shown in the example above. Correspondingly, the metric for the training time required can be the metric for the time required to train encoder 2 (e.g., denoted as T4). Therefore, the set of metrics corresponding to the third model is {T4}. For example, the terminal side can monitor the training time of encoder 2.

[0405] 3) Metrics of resources required for training:

[0406] The resources required for training can specifically refer to the size of the resources needed to train encoder 2. For example, it could be the size of the resources required for step 2 as shown in the example above, or the size of the resources required for steps 1 and 2 as shown in the example above. Correspondingly, the metric for the resources required for training can be the metric for the resources required to train encoder 2 (e.g., denoted as R4). Therefore, the set of metrics corresponding to the third model is {R4}. For example, the terminal side can monitor the resources used to train encoder 2.

[0407] In summary, when the dataset is sent from the network side, the terminal side can use either training method one or training method two for model training. Corresponding to training method one, multiple sets of metrics are provided above, such as performance metrics: {metric group 1-1}, or {metric group 1-2}, or {metric group 1-1, metric group 1-2}, corresponding to {function set 1-1}, or {function set 1-2}, or {function set 1-1, function set 1-2} (i.e., multiple examples of the first function set); the metric for training time required: {T} 1-1}, or {T 1-2}, or {T 1-1 T 1-2}, or {T1}; metrics for resources required for training: {R} 1-1}, or {R 1-2}, or {R 1-1 R 1-2}, or {R1}. Corresponding to training method two, multiple sets of metrics are also provided above, such as performance metrics: {metric group 2}, or {metric group 2'}, corresponding to {function set 2}, or {function set 2'} (i.e., two examples of the second function set); the metric for training time {T2}; and the metric for training resources {R2}.

[0408] When the model is sent from the network side, the terminal side can use either training method three or training method four for model training. Corresponding to training method three, multiple sets of metrics are provided above, such as performance metrics: {metric group 3-1}, or {metric group 3-2}, or {metric group 3-1, metric group 3-2}, corresponding to {function set 3-1}, or {function set 3-2}, or {function set 3-1, function set 3-2} (i.e., multiple examples of the third function set); the metric for training time required: {T} 3-1}, or {T 3-2}, or {T 3-1 T 3-2}, or {T3}; metrics for resources required for training: {R} 3-1}, or {R 3-2}, or {R 3-1 R 3-2}, or {R3}. Corresponding to training method four, several sets of metrics are also provided above, such as performance metrics: {metric group 4}, corresponding to {function set 4} (i.e., an example of the fourth function set); the metric for training time required: {T4}; and the metric for training resources required: {R4}.

[0409] When the terminal receives the first dataset from the network side, including the target CSI and CSI feedback information, it can train the model based on training method one or training method two, and a set of metrics corresponding to the adopted training method. When the terminal receives the third model from the network side, it can train the model based on training method three or training method four, and a set of metrics corresponding to the adopted training method.

[0410] The foregoing, with the aid of several accompanying figures, provides a detailed explanation of various training methods, multiple sets of metrics, and sets of functions, as well as the corresponding relationships between them. It should be understood that these examples are provided for illustrative purposes only and should not be construed as limiting this application.

[0411] Optionally, the aforementioned multiple sets of indicators may be indicated by the network side. Before step 1020, the method further includes step 1030: the network side sends first information to the terminal side, the first information being used to indicate the multiple sets of indicators. Correspondingly, the terminal side receives the first information from the network side.

[0412] In this embodiment, the first information may be sent by a network device, such as to a terminal device or to an OTT system server on the terminal side; alternatively, the first information may also be sent by a smart network element on the network side, such as to a terminal device or to an OTT system server on the terminal side. This application does not limit this.

[0413] As an example, the first information can be sent from the network device to the terminal device. Step 1030 may specifically include: the network device sending the first information to the terminal device, for example, via the air interface. Accordingly, the terminal device receives the first information from the network side.

[0414] In another example, the first information can be sent from a network-side intelligent network element to an OTT system server on the terminal side. Step 1030 may specifically include: the intelligent network element sending the first information to the OTT system server, for example, via a wired network. Accordingly, the OTT system server receives the first information. Alternatively, step 1030 may specifically include: the intelligent network element sending indication information of the multiple sets of indicators to the network device; the network device sending the first information to the terminal device based on the received information of the multiple sets of indicators; and the terminal device then forwarding the first information to the OTT system server. The indication of the multiple sets of indicators sent by the intelligent network element to the network device may be, for example, the first information, or other information that can be used to indicate the multiple sets of indicators; this application does not limit this.

[0415] In another example, the first information can be sent to the terminal device by an intelligent network element on the network side. Step 1030 may specifically include: the intelligent network element sending indication information of the multiple sets of indicators to the network device; and the network device sending the first information to the terminal device based on the received indication information of the multiple sets of indicators. The indication information of the multiple sets of indicators sent by the intelligent network element to the network device may be, for example, the first information, or other information that can be used to indicate the multiple sets of indicators; this application does not limit this.

[0416] It should be understood that the examples above are provided for illustrative purposes only and should not be construed as limiting the scope of this application. Many more possible examples can be derived based on the same concept, but for the sake of brevity, they are not listed here.

[0417] The network side can, based on the dataset and / or model sent to the terminal side, indicate multiple sets of indicators corresponding to the dataset and / or model to the terminal side through first information, thereby associating the dataset and / or model with multiple sets of indicators.

[0418] One possible implementation is that the network side can include these multiple sets of indicators in the first information and send it to the terminal side. Accordingly, the terminal side can obtain these multiple sets of indicators based on the parsing of the first information. The terminal side can then determine the first indicator from the obtained multiple sets of indicators.

[0419] Another possible implementation is that the first information can carry identifiers or indexes for multiple sets of indicators. The network side and the terminal side can pre-store mapping information indicating the correspondence between multiple sets of indicators and multiple identifiers (or indexes), with each set of indicators indicated by one identifier (or index). The multiple sets of indicators indicated by the mapping information may include, but are not limited to, the multiple sets of indicators indicated by the network side to the terminal side through the first information; or, the multiple sets of indicators indicated by the network side to the terminal side through the first information are a subset of the multiple sets of indicators indicated by the mapping information. Based on this mapping information, the network side can carry the identifiers (or indexes) of the multiple sets of indicators corresponding to the dataset and / or model sent in step 1010 in the first information and send it to the terminal side. Correspondingly, the terminal side can determine the identifiers (or indexes) of the multiple sets of indicators based on the parsing of the first information, and thus obtain the multiple sets of indicators. The terminal side can determine the first indicator from the obtained multiple sets of indicators. Alternatively, after the terminal side determines the identifier (or index) of the multiple sets of indicators based on the parsing of the first information, it may not obtain the multiple sets of indicators, but instead directly select an identifier (or index) from the identifiers (or indexes) of the multiple sets of indicators and determine the corresponding set of indicators as the first indicator.

[0420] Optionally, the above mapping information can be predefined, such as protocol predefined; or it can be pre-negotiated between the network side and the terminal side. This application does not limit this.

[0421] The terminal can receive the dataset and / or model, and after determining the first metric, perform model training, and determine whether the first metric is satisfied based on the model training results; alternatively, the terminal can receive the dataset and / or model, determine whether the first metric is satisfied, and then perform model training if the first metric is satisfied. In other words, this application does not limit the execution order of the terminal determining whether the first metric is satisfied and the terminal performing model training.

[0422] The terminal side's prediction of whether the first metric is met can be based on the capabilities of its training device. Therefore, it can also be described as the terminal side predicting whether the first metric can be met. For example, if the training device is a terminal device, it can determine whether the first metric is met based on its own capabilities, such as computing power. This computing power can be characterized by parameters such as the device's computing speed and processing time per unit of data, which are not limited in this application.

[0423] For example, suppose the first metric is an upper bound on the time required for model training; that is, the training time should not exceed the first metric. The terminal device can predict the possible training time based on the received dataset and / or model, and the training method corresponding to the first metric, and then determine whether the training time exceeds the first metric. In other words, it determines whether the first metric is satisfied.

[0424] It should be understood that there can be multiple ways for the terminal side to determine whether the first indicator is met. The example above is only provided for ease of understanding and should not constitute any limitation on this application. This application does not limit the specific implementation method for the terminal side to determine whether the first indicator is met.

[0425] As previously stated, each metric may include an upper bound and / or a lower bound; therefore, the first metric also includes an upper bound and / or a lower bound. The first metric being satisfied may include: not higher than the upper bound and / or not lower than the lower bound; or, lower than or equal to the upper bound, and / or, greater than or equal to the lower bound. Alternatively, in another implementation, the first metric being satisfied may also include: lower than the upper bound and / or higher than the lower bound.

[0426] The following sections will explain the two scenarios: when the first indicator is met and when it is not.

[0427] Optionally, the method further includes step 1040: if the first metric is satisfied, the terminal side executes the inference task of the first model.

[0428] The terminal can train a model to obtain a first model. If the first metric is met, the terminal can deploy the first model to a real-world scenario to perform the inference task of the first model.

[0429] Furthermore, the method also includes step 1050: the terminal side sends second information to the network side, the second information being used to indicate that the first indicator has been met. Accordingly, the network side receives the second information from the terminal side.

[0430] As mentioned earlier, multiple training methods and multiple sets of indicators correspond to the first dataset, and multiple training methods and multiple sets of indicators correspond to the third model. The terminal can choose one of the training methods and the corresponding set of one or more sets of indicators for model training. Therefore, the terminal can also notify the network side through the second information that the first indicator has been satisfied. One possible implementation is that the terminal can carry the first indicator in the second information. Another possible implementation is that the terminal can carry the identifier (or index) of the first indicator in the second information. This application does not limit the specific implementation method by which the terminal indicates that the first indicator has been satisfied.

[0431] The network side can determine from the second information that the model training on the terminal side can meet the first indicator, and based on the first indicator, it can continue to monitor the first model during the model deployment phase.

[0432] Furthermore, the method also includes step 1060, whereby the network side monitors the first model.

[0433] One possible implementation is that the network side can continue to monitor the first model based on the first metric.

[0434] For example, the first indicator is indicator group 2 as shown above, and the network side can continue to monitor the first model based on indicator group 2.

[0435] Another possible implementation is that the network side monitors the first model based on other metrics related to the first metric.

[0436] For example, the first metric is metric group 1-1 as shown above. The network side can monitor the first model based on metric group 1. Since metric group 1-1 and metric group 1-2 are the metrics corresponding to training 1-1 respectively, and the first model has been deployed in the actual scenario, the network side can monitor the first model based on metric group 1-2.

[0437] The specific implementation of network-side monitoring of the first model is determined by the network side, and this application does not impose any restrictions on it.

[0438] Optionally, the method further includes step 1070: if the first indicator is not met, the terminal side sends third information to the network side, the third information being used for one or more of the following:

[0439] a. The first criterion is not met;

[0440] b. Request a change of dataset;

[0441] c. Request a model replacement; or

[0442] d. Request to stop or pause model training.

[0443] For example, due to the limited computing power of the terminal, the first indicator is difficult to meet for the terminal. Therefore, the first indicator can be determined by the second information that it is not met.

[0444] For example, since the terminal can train a model based on the model sent by the network, it can request a model change via the second information, such as requesting the network to change the model, or to change to a different type of model. Optionally, changing the model may include changing the model file and / or changing the model parameters. The terminal can also simultaneously indicate via the second information that the aforementioned first metric cannot be met. It can be understood that if the network changes to a different model, it means that the corresponding training method and metrics will also change. Therefore, requesting the network to change the model can also be understood as requesting the network to change the metrics.

[0445] For example, since the terminal can train the model based on the dataset sent by the network, it can request a change of dataset via the second information, such as requesting a different dataset type. The terminal can also simultaneously indicate via the second information that the first metric cannot be met. It's understandable that if the network changes to a different dataset type, the corresponding training method and metrics will also change. Therefore, requesting the network to change to a different dataset type can also be understood as requesting the network to change the metrics.

[0446] For example, due to limited computing power on the terminal side, the first metric may be difficult for the terminal side to meet. Therefore, the terminal side can also request to stop or pause model training through the second information. The terminal side can also simultaneously indicate that the above performance metric cannot be met through the second information. Stopping model training can be understood as the terminal side no longer training the model based on the dataset and / or model provided by the network side. Pausing model training can be understood as the terminal side temporarily not training the model based on the dataset and / or model provided by the network side until another event triggers, such as changes in terminal side resources or computing power, requesting the network side to provide the dataset and / or model to continue model training. Stopping or pausing model training on the terminal side also means that the first model is not deployed on the terminal side. The network side also no longer configures the terminal side for inference tasks of the first model. Therefore, it can also be called requesting the network side to shut down model training. It should be understood that the network side can decide its own actions after shutting down model training, and this application does not impose any restrictions on this.

[0447] Similar to the steps described above, in one example, the training device is a terminal device. The terminal sends the second information, which can be either sent from the terminal device to the network device or to the OTT system server. This information, through the OTT system server, is then sent to the intelligent network element to provide information that can be used to achieve one or more of the above-mentioned steps a to e (e.g., the second information itself, or other information). In another example, the training device is an OTT system server. The terminal sends the second information, which can be either sent from the OTT system server to the intelligent network element or to the terminal device. This information, through the terminal device, is then sent to the network device to provide information that can be used to achieve one or more of the above-mentioned steps a to e (e.g., the second information itself, or other information).

[0448] It should be understood that steps 1040-1060 and step 1070 are steps that the terminal side executes respectively depending on whether the first indicator is met or not. The terminal side can choose to execute one of them according to the actual situation, and does not necessarily have to execute all of them.

[0449] Based on the above scheme, the first metric not only corresponds to the dataset and / or model used for model training, but also to the inference task of the first model. While acquiring the dataset and / or model used for model training, the terminal can also acquire the corresponding first metric. Therefore, based on the dataset and / or model, and the associated first metric, it can use appropriate training methods to train the model to obtain a model that meets the requirements, thereby improving the feedback performance of CSI.

[0450] Furthermore, each dataset and / or each model can correspond to multiple sets of metrics, which means that the terminal can use various different training methods to train the model. The first metric is one or more of these sets of metrics; therefore, the terminal can use one or more corresponding training methods to train the model based on one or more of these sets of metrics. Thus, the terminal has a high degree of freedom; it can choose appropriate metrics and training methods for model training based on business needs, device capabilities, and other factors.

[0451] Furthermore, the multiple sets of metrics in this application may include one or more of the following categories: performance metrics, metrics of training time, or metrics of training resources. That is, the model training process is monitored from one or more dimensions of performance, time, or resources, so as to flexibly respond to different business needs.

[0452] It should also be understood that in the embodiments shown above in conjunction with Figures 10 to 14, the terminal side can be a terminal device, and the network side can be a network device. In another design, the terminal side can also include a terminal device and an OTT system host or cloud server; the network side can include network devices and intelligent network elements. In this case, devices on the terminal side can communicate with each other, and devices on the network side can also communicate with each other. The specific implementation flow of the method provided in this application in an architecture with an OTT system host or cloud server (hereinafter referred to as OTT system server), terminal device, network device, and intelligent network element will be described below using method 1500 shown in Figure 15. In this flow, it is assumed that the training device is an OTT system server.

[0453] Figure 15 is another schematic flowchart of the communication method provided in the embodiments of this application. The method 1500 shown in Figure 15 may include the following steps:

[0454] Step 1501: The intelligent network element sends a dataset and / or model to the network device, which is used for model training.

[0455] Step 1502: The network device sends the dataset and / or model to the terminal device.

[0456] Step 1503: The terminal device sends the dataset and / or model to the OTT system server.

[0457] Steps 1501-1503 illustrate the process by which an intelligent network element transmits datasets and / or models over the air interface between a network device and a terminal device. In another implementation, the intelligent network element can send datasets and / or models to an OTT system server via wired transmission, as shown in step 1504, indicated by the dashed line in the figure.

[0458] Step 1505: The network device sends first information to the terminal device, which is used to indicate multiple sets of indicators.

[0459] Step 1506: The terminal device sends information indicating the multiple sets of indicators to the OTT system server. The information indicating the multiple sets of indicators sent by the terminal device to the OTT system server can be the first set of information or other information, without limitation.

[0460] In another implementation, the terminal device can also determine the first indicator based on the first information, and then send information to the OTT system server to indicate the first indicator. That is, step 1506 can also be replaced by: the terminal device sending information to the OTT system server to indicate the first indicator.

[0461] Step 1507: The OTT system server trains the model based on the first indicator among multiple sets of indicators, as well as the received dataset and / or model.

[0462] Step 1508: The OTT system server determines whether the first metric has been met.

[0463] Since the OTT system server is a training device, it can determine whether the first metric is met on its own.

[0464] Step 1509: If the first metric is met, the OTT system server executes the inference task of the first model.

[0465] Step 1510: If the first indicator is met, the OTT system server sends information indicating that the first indicator has been met.

[0466] Step 1511: The terminal device sends a second message to the network device, which indicates that the first indicator has been met.

[0467] The information received by the terminal device from the OTT system server in step 1510 to indicate that the first indicator has been met can be the same as or different from the second information sent by the terminal device in step 1511. This application does not limit this.

[0468] Step 1512: The network device sends information to the intelligent network element to indicate that the first indicator has been met.

[0469] The second information received by the network device from the terminal device in step 1511 and the information sent by the network device in step 1512 to indicate that the first indicator has been met can be the same information or different information. This application does not limit this.

[0470] Step 1513: If the first indicator is not met, the OTT system server sends information to the terminal device indicating that the first indicator is not met.

[0471] In another implementation, step 1508 can also be performed by the terminal device. In this case, the terminal device can predict whether the first metric is met based on the capabilities of the OTT system server; alternatively, the OTT system server can send the output of the trained model during model training to the terminal device so that the terminal device can determine whether the first metric is met. Steps 1510 or 1513 can also be omitted.

[0472] Step 1514: The terminal device sends third information to the network device. This third information can be used for one or more of the following: indicating that the first metric is not met; requesting a change of dataset; requesting a change of model; or requesting to stop or pause model training.

[0473] If the first criterion is not met, the terminal device can send a third message.

[0474] The information received by the terminal device from the OTT system server in step 1513 regarding the failure to meet the first indicator can be the same as or different from the third information sent by the terminal device in step 1514. This application does not limit this.

[0475] Step 1515: The network device sends a fourth message to the intelligent network element. This fourth message can be used for one or more of the following: indicating that the first metric is not met; requesting a change of dataset; requesting a change of model; or requesting to stop or pause model training.

[0476] The fourth information sent by the network device in step 1515 and the third information received from the terminal device in step 1514 can be the same information or different information; this application does not limit this.

[0477] For detailed descriptions of each step in method 1500, please refer to the relevant descriptions in method 1000 above in conjunction with Figure 10, which will not be repeated here. Furthermore, the technical solution shown in method 1500 corresponds to the technical solution shown in method 1000 above, and therefore the beneficial effects obtained are similar, which will not be repeated here either.

[0478] It should be understood that the process shown in Figure 15 is merely one possible implementation of the method 1000 shown in Figure 10 in a system architecture deploying an OTT system server and intelligent network elements, and should not be construed as limiting this application in any way. Those skilled in the art can, based on the same concept, make simple substitutions or modifications to one or more steps of method 1000 to achieve the same effect. Such simple modifications or substitutions should fall within the protection scope of this application.

[0479] It should also be understood that the processes listed above are merely examples, and this application includes, but is not limited to, them. For example, the terminal device and OTT system server in Figure 15 can also be replaced by the terminal device itself. In this case, the communication between the terminal device and the OTT system server is internal communication within the terminal device. Alternatively, the network device and intelligent network element in Figure 15 can also be replaced by the network device itself. In this case, the communication between the network device and the intelligent network element is internal communication within the network device. For the sake of brevity, no further illustrations are provided in this document.

[0480] It should be understood that in the various embodiments shown above in conjunction with the accompanying drawings, the sequence number of each step does not imply the order of execution. The execution order of each step should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0481] It should also be understood that the various embodiments exemplified above can also be applied to scenarios where the terminal sends a dataset and / or model to the network side for model training via the network side. In this case, the steps on the terminal side in the above embodiments can be performed by the network side, and vice versa. Those skilled in the art can make simple modifications or substitutions based on the same concept to obtain more embodiments.

[0482] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of different embodiments are consistent and can be referenced by each other. The technical features of different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0483] In the above embodiments, exemplary descriptions are mainly based on devices in the current network architecture (such as terminal devices, network devices, OTT system servers, intelligent network elements, etc.). The specific form of the devices is not limited in the embodiments of this application. For example, devices that can achieve the same function in the future can also be applied to the methods provided in the embodiments of this application.

[0484] It is understood that the methods and operations implemented by devices (such as terminal devices, network devices, OTT system servers, and smart network elements) in the above-described method embodiments can also be implemented by components of the devices (such as chips or circuits).

[0485] The methods provided in the embodiments of this application have been described in detail above with reference to several accompanying drawings. The apparatus provided in the embodiments of this application will now be described with reference to the accompanying drawings.

[0486] Figures 16 and 17 are schematic block diagrams of possible apparatuses provided in embodiments of this application. These apparatuses can be used to implement the terminal-side or network-side functions in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments.

[0487] Figure 16 is a schematic block diagram of an apparatus provided in an embodiment of this application. The apparatus 1600 shown in Figure 16 may include a processing module 1610 and a communication module 1620.

[0488] In one possible design, device 1600 can be used to implement the communication method implemented by the terminal side in any of the embodiments shown in Figures 10 to 15. For example, processing module 1610 is used to implement model training performed by the terminal side in each method embodiment, determine whether the first metric is met, and perform inference tasks of the first model, etc., and other processing-related steps; communication module 1620 is used to implement sending and / or receiving steps performed by the terminal side in each method embodiment, such as receiving a dataset and / or model, receiving first information, sending second information, or sending third information, or one or more of these steps.

[0489] For example, the communication module 1620 can be used to: acquire a dataset and / or a model, the dataset and / or the model being used for model training; the processing module 1610 is used to perform model training based on a first metric, and the dataset and / or the model, to obtain a first model; the first metric is one or more sets of metrics, the sets of metrics including the following performance metrics corresponding to the acquired dataset, and / or corresponding to the acquired model; and the performance metrics correspond to the inference task of the first model, the first model being obtained by model training based on the dataset and / or the model, the inference task being inference of CSI compression.

[0490] Optionally, the communication module 1620 can be used to receive first information, which is used to indicate multiple sets of indicators.

[0491] Optionally, the processing module 1610 can be used to obtain the multiple sets of indicators based on the first information.

[0492] Optionally, when the communication module 1620 is used to receive a dataset and / or a model, it is specifically used to receive a first dataset. The first dataset includes: target CSI and CSI feedback information from the network side, wherein the CSI feedback information is obtained by compressing the target CSI.

[0493] Optionally, when the communication module 1620 is used to receive a dataset and / or a model, it is specifically used to receive a third model whose inference task is to compress the target CSI.

[0494] Optionally, the processing module 1610 is also used to perform the inference task of the first model if the first metric is satisfied.

[0495] Optionally, the communication module 1620 is further configured to send a second message when the first indicator is met, the second message indicating that the first indicator has been met.

[0496] Optionally, the communication module 1620 can also be used to send a third message when the performance metric is not met. The third message is used for one or more of the following: indicating that the first metric is not met; requesting a change of dataset; requesting a change of model; or requesting to stop or pause the training of the model.

[0497] A more detailed description of the processing module 1610 and the communication module 1620 can be obtained directly from the relevant descriptions in the method embodiments shown in Figures 10 to 15, and will not be repeated here.

[0498] In another possible design, device 1600 can be used to implement the communication method implemented by the network side in any of the embodiments shown in Figures 10 to 15. For example, processing module 1610 is used to implement processing-related steps such as generating datasets and / or models performed by the network side in each method embodiment; communication module 1620 is used to implement sending and / or receiving steps performed by the network side in each method embodiment, such as sending datasets and / or models, sending first information, receiving second information, or receiving third information, or one or more of these steps.

[0499] For example, the communication module 1620 can be used to send a dataset and / or a model for model training; the communication module 1620 is also used to send first information indicating multiple sets of metrics; the multiple sets of metrics include performance metrics corresponding to the acquired dataset and / or the acquired model; and the performance metrics correspond to the inference task of a first model, which is obtained by model training based on the dataset and / or the model, and the inference task is inference of CSI compression.

[0500] Optionally, when the communication module 1620 is used to send a dataset and / or a model, it is specifically used to send a first dataset. The first dataset includes a target CSI and CSI feedback information, wherein the CSI feedback information is obtained by compressing the target CSI.

[0501] Optionally, when the communication module 1620 is used to send a dataset and / or a model, it is specifically used to send a third model whose inference task is to compress the target CSI.

[0502] Optionally, the communication module 1620 can also be used to receive second information, which indicates that the first indicator has been met.

[0503] Optionally, the communication module 1620 can also be used to receive third information, which is used for one or more of the following: indicating that the first indicator cannot be met; requesting a change of dataset; requesting a change of model; or requesting to stop or pause the training of the model.

[0504] A more detailed description of the processing module 1610 and the communication module 1620 can be obtained directly from the relevant descriptions in the method embodiments shown in Figures 10 to 15, and will not be repeated here.

[0505] It should be noted that the communication module can also be called a transceiver module, transceiver unit, transceiver, transceiver device, or transceiver apparatus, etc. The processing module can also be called a processor, processing board, processing unit, or processing apparatus, etc. Optionally, the communication module is used to execute the sending and receiving operations of the first or second communication device in the above method. The device in the communication module that implements the receiving function can be considered as the receiving module, and the device in the communication module that implements the sending function can be considered as the sending module; that is, the communication module can include both a receiving module and a sending module.

[0506] It should also be noted that, in one possible design, the aforementioned processing module and / or communication module can be implemented through virtual modules. For example, the processing module can be implemented through software functional units or virtual devices, and the communication module can be implemented through software functions or virtual devices. In another possible design, the processing module or communication module can also be implemented through physical devices. For example, if the device is implemented using a chip / chip circuit, the communication module can be an input / output circuit and / or a communication interface, performing input operations (corresponding to the aforementioned receiving operation) and output operations (corresponding to the aforementioned sending operation); the processing module can be an integrated processor, a microprocessor, or an integrated circuit.

[0507] The module division in this embodiment is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional modules in the various examples of this embodiment can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0508] Figure 17 is a schematic diagram of the structure of a communication device provided in another embodiment of this application. As shown in Figure 17, the device 1700 includes a processing circuit 1710 and a communication circuit 1720. The processing circuit 1710 and the communication circuit 1720 are coupled to each other.

[0509] It can be understood that the processing circuit 1710 can be one or more processors, or it can be all or part of the processing functions of one or more processors.

[0510] Understandably, the communication circuit 1720 can be a transceiver or an input / output interface.

[0511] Optionally, the device 1700 may further include a memory 1730 for storing instructions executed by the processing circuit 1710, or storing input data required for the running instructions of the processing circuit 1710, or storing data generated after the running instructions of the processing circuit 1710.

[0512] It is understood that the memory 1730 may be located outside the processing circuit 1710, or inside the processing circuit 1710.

[0513] As an example, the processing circuit 1710 is used to implement the functions of the processing module 2010, and the communication circuit 1720 is used to implement the functions of the communication module 2020.

[0514] As an example, device 1700 can be a communication device or a chip used in a communication device.

[0515] When device 1700 is a communication device, the communication circuit can be a transceiver; when device 1700 is a chip, the communication circuit can be an input / output circuit, a bus, pins, or other types of communication interfaces. The input circuit in the input / output circuit can be used for receiving, and the output interface can be used for transmitting.

[0516] This application also provides a computer program product that, when run on a processor, can implement the communication method executed by the terminal side or the communication method executed by the network side in the above method embodiments.

[0517] This application also provides a computer-readable storage medium containing computer instructions that, when executed on a processor, can implement the communication method executed by the terminal side or the communication method executed by the network side in the above method embodiments.

[0518] This application also provides a communication system, including the aforementioned terminal side and network side. The terminal side can be used to implement the communication method implemented by the terminal side in the above method embodiments, and the network side can be used to implement the communication method implemented by the network side in the above method embodiments.

[0519] It is understood that the processor in the embodiments of this application may be any of the following devices or all or part of the circuitry used for processing functions: a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.

[0520] The terms “unit”, “module”, etc., used in this specification may be used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution.

[0521] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0522] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0523] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0524] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0525] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0526] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. This computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0527] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0528] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A communication method, characterized in that, include: Acquire a dataset and / or a model, which is used for model training; Based on the first metric, and the dataset and / or the model, the model is trained to obtain the first model; the first metric is one or more sets of metrics, which include one or more of the following categories of metrics: performance metrics, metrics of training time required, or metrics of training resources required, the multiple sets of metrics correspond to the dataset and / or the model, and correspond to the inference task of the first model, the inference task of the first model being to compress the target channel state information (CSI).

2. The method as described in claim 1, characterized in that, Before training the model based on the first metric, the dataset, and / or the model, the method further includes: Receive first information, which is used to indicate the multiple sets of indicators.

3. The method as described in claim 2, characterized in that, The method further includes: Based on the first information, obtain the multiple sets of indicators.

4. The method according to any one of claims 1 to 3, characterized in that, The first indicator is the performance indicator, and at least two of the multiple sets of indicators are the performance indicators, wherein the first indicator is one or more of the at least two sets of indicators.

5. The method as described in claim 4, characterized in that, The acquisition of the dataset and / or model includes: acquiring a first dataset, the first dataset including: target CSI and CSI feedback information from the network side, wherein the CSI feedback information is obtained by compressing the target CSI.

6. The method as described in claim 5, characterized in that, The first dataset is used to train the second model, and the second model is used to train the first model, wherein the inference task of the second model is to decompress the CSI feedback information; The at least two function sets corresponding to the at least two sets of indicators include a first function set, which includes one or more of the following functions: normalized mean squared error (NMSE), mean squared error (MSE), L1 loss (L1Loss), generalized cosine similarity (GCS), squared generalized cosine similarity (SGCS) between the output of the second model and the label; or, a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the label.

7. The method as described in claim 6, characterized in that, The output of the second model includes the output of the second model during training; the label is the target CSI in the first dataset; the first function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the label, or a weighted sum of multiple functions among NMSE, MSE, L1Loss, GCS and SGCS between the output of the second model and the label.

8. The method as described in claim 6, characterized in that, The output of the second model includes: the output of the second model when the first model is trained based on the second model; the label is the target CSI in the first dataset or the target CSI in the third dataset; the first function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the label, or the weighted sum of multiple functions among NMSE, MSE, L1Loss, GCS and SGCS between the output of the second model and the label; wherein, the third dataset is a dataset obtained by the terminal side through measurement or simulation.

9. The method as described in claim 5, characterized in that, The first dataset is used to train the first model; the at least two function sets corresponding to the at least two sets of indicators include a second function set, which includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the first model and the CSI feedback information, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS or SGCS between the output of the first model and the CSI feedback information.

10. The method as described in claim 4, characterized in that, The acquisition of the dataset and / or model includes: acquiring a third model, wherein the inference task of the third model is to compress the target CSI.

11. The method as described in claim 10, characterized in that, The third model is used to train the second model, and the second model is used to train the first model. The inference task of the second model is to decompress the CSI feedback information, which is obtained by the third model by compressing the target CSI. The at least two function sets corresponding to the at least two sets of indicators include a third function set, which includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the label, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the label.

12. The method as described in claim 11, characterized in that, The output of the second model includes: the output of the second model during training, wherein the label is the target CSI in the first dataset or the target CSI in the third dataset; the third function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SCGS between the output of the second model and the label, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS and SGCS between the output of the second model and the label; wherein the third dataset is a dataset obtained by the terminal side through measurement or simulation.

13. The method as described in claim 11, characterized in that, The output of the second model includes: the output of the second model when the first model is trained based on the second model; the label is the target CSI in the first dataset or the target CSI in the third dataset; the third function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SCGS between the output of the second model and the label, or the weighted sum of multiple functions among NMSE, MSE, L1Loss, GCS and SGCS between the output of the second model and the label; wherein, the third dataset is a dataset obtained by the terminal side through measurement or simulation.

14. The method as described in claim 10, characterized in that, The third model is used to obtain a second dataset, which includes CSI feedback information, and the second dataset is used to train the first model. The at least two function sets corresponding to the at least two sets of indicators include a fourth function set, which includes one or more of the following functions: NMSE, MSE, L1Loss between the output of the first model and the CSI feedback information, or a weighted sum of multiple NMSE, MSE, or L1Loss between the output of the first model and the CSI feedback information.

15. The method according to any one of claims 6 to 8, 11 to 13, characterized in that, The output of the second model includes: The output of the second model during training, and / or The output of the second model when the first model is trained based on the second model.

16. The method according to any one of claims 1 to 15, characterized in that, The method further includes: If the first metric is met, the inference task of the first model is executed.

17. The method according to any one of claims 1 to 16, characterized in that, The method further includes: Send a second message, which indicates that the first indicator has been met.

18. The method according to any one of claims 1 to 17, characterized in that, The method further includes: If the first indicator is not met, a third message is sent, the third message being used for one or more of the following: This indicates that the first indicator is not met; Request to replace the dataset; Request a replacement of the model; or Request to stop or pause the training of the model.

19. The method according to any one of claims 1 to 18, characterized in that, The acquisition of the dataset and / or model includes: Receive the dataset and / or the model.

20. A communication method, characterized in that, include: Send a dataset and / or a model, which is used for model training; Send a first message, which is used to indicate multiple sets of indicators, which are used for model training; wherein, the multiple sets of indicators include one or more of the following types of indicators: performance indicators, indicators of training time required, or indicators of training resources required, the multiple sets of indicators correspond to the dataset and / or the model, and correspond to the inference task of the first model, the inference task of the first model being to compress the target channel state information (CSI).

21. The method as described in claim 20, characterized in that, At least two of the multiple sets of indicators are the performance indicators.

22. The method as described in claim 21, characterized in that, The sending of the dataset and / or model includes: sending a first dataset, the first dataset including: target CSI and CSI feedback information, wherein the CSI feedback information is obtained by compressing the target CSI.

23. The method as described in claim 22, characterized in that, The first dataset is used to train the second model, and the second model is used to train the first model, wherein the inference task of the second model is to decompress the CSI feedback information; The at least two function sets corresponding to the at least two sets of indicators include a first function set, which includes one or more of the following functions: normalized mean squared error (NMSE), mean squared error (MSE), L1 loss (L1Loss), generalized cosine similarity (GCS), squared generalized cosine similarity (SGCS) between the output of the second model and the label; or, a weighted sum of multiples of NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the label.

24. The method as described in claim 23, characterized in that, The output of the second model includes the output of the second model during training; the label is the target CSI in the first dataset; the first function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the label, or a weighted sum of multiple functions among NMSE, MSE, L1Loss, GCS and SGCS between the output of the second model and the label.

25. The method as described in claim 23, characterized in that, The output of the second model includes: the output of the second model when the first model is trained based on the second model; the label is the target CSI in the first dataset or the target CSI in the third dataset; the first function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the label, or the weighted sum of multiple functions among NMSE, MSE, L1Loss, GCS and SGCS between the output of the second model and the label; wherein, the third dataset is a dataset obtained by the terminal side through measurement or simulation.

26. The method as described in claim 22, characterized in that, The first dataset is used to train the first model; The at least two function sets corresponding to the at least two sets of indicators include a second function set, which includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the first model and the CSI feedback information, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS or SGCS between the output of the first model and the CSI feedback information.

27. The method as described in claim 21, characterized in that, The sending of the dataset and / or model includes: sending a third model, wherein the inference task of the third model is to compress the target CSI.

28. The method as described in claim 27, characterized in that, The third model is used to train the second model, and the second model is used to train the first model. The inference task of the second model is to decompress the CSI feedback information, which is obtained by the third model by compressing the target CSI. The at least two function sets corresponding to the at least two sets of indicators include a third function set, which includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SGCS between the output of the second model and the label, or a weighted sum of multiple NMSE, MSE, L1Loss, GCS, or SGCS between the output of the second model and the label.

29. The method as described in claim 28, characterized in that, The output of the second model includes: the output of the second model during training, wherein the label is the target CSI in the first dataset or the target CSI in the third dataset; the third function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SCGS between the output of the second model and the label, or a weighted sum of multiples of NMSE, MSE, L1Loss, GCS and SGCS between the output of the second model and the label; wherein the third dataset is a dataset obtained by the terminal side through measurement or simulation.

30. The method as described in claim 28, characterized in that, The output of the second model includes: the output of the second model when the first model is trained based on the second model; the label is the target CSI in the first dataset or the target CSI in the third dataset; the third function set includes one or more of the following functions: NMSE, MSE, L1Loss, GCS, SCGS between the output of the second model and the label, or the weighted sum of multiple functions among NMSE, MSE, L1Loss, GCS and SGCS between the output of the second model and the label; wherein, the third dataset is a dataset obtained by the terminal side through measurement or simulation.

31. The method as described in claim 27, characterized in that, The third model is used to obtain a second dataset, which includes CSI feedback information. The second dataset is used to train the first model. The at least two function sets corresponding to the at least two sets of indicators include a fourth function set, which includes one or more of the following functions: NMSE, MSE, and L1Loss between the output of the first model and the CSI feedback information, or a weighted sum of multiple NMSE, MSE, or L1Loss between the output of the first model and the CSI feedback information.

32. The method according to any one of claims 23 to 25, 28 to 30, characterized in that, The output of the second model includes: The output of the second model during training, and / or The output of the second model when the first model is trained based on the second model.

33. The method according to any one of claims 20 to 32, characterized in that, The method further includes: Receive second information, the second information being used to indicate that a first indicator is satisfied, the first indicator being an indicator used for training the model, and the first indicator being one or more of the multiple sets of indicators.

34. The method according to any one of claims 20 to 33, characterized in that, The method further includes: Receive third information, which is used for one or more of the following: The first metric is not satisfied, where the first metric is a metric used for model training, and the first metric is one or more of the multiple sets of metrics. Request to replace the dataset; Request a replacement of the model; or Request to stop or pause the training of the model.

35. A communication device, characterized in that, Includes modules or units for performing the method according to any one of claims 1 to 34.

36. A communication device, characterized in that, Includes a processor configured to cause the communication device to perform the method of any one of claims 1 to 34.

37. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions that, when executed on a communication device, cause the communication device to perform the method as described in any one of claims 1 to 34.

38. A computer program product, characterized in that, The computer program product includes a computer program or instructions that, when executed on a communication device, cause the communication device to perform the method as described in any one of claims 1 to 34.