Method for acquiring downlink channel state information (CSI), model training method and device

By using a reinforcement learning model between the base station and the terminal to optimize signaling overhead, the problem of low efficiency in acquiring downlink CSI in massive MIMO systems is solved, and more efficient CSI acquisition is achieved.

CN116527215BActive Publication Date: 2026-02-10DATANG MOBILE COMM EQUIP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210068492.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-20
Publication Date
2026-02-10
Estimated Expiration
2042-01-20

AI Technical Summary

Technical Problem

As the number of base station antennas increases, the signaling overhead for acquiring downlink channel state information (CSI) becomes large, leading to low efficiency.

Method used

By employing a reinforcement learning model, the target action is determined by acquiring the state information of the terminal and the base station to obtain downlink CSI, including predicting or feeding back downlink CSI, thereby optimizing signaling overhead.

Benefits of technology

It effectively reduces signaling overhead, improves the efficiency and accuracy of obtaining downlink CSI, and reduces the resource consumption of the communication system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116527215B_ABST
    Figure CN116527215B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for acquiring downlink channel state information (CSI), and a model training method, and relates to the technical field of communication. The specific implementation scheme is as follows: acquiring state information of a terminal; determining a target action for acquiring downlink CSI according to the state information, wherein the target action is predicting downlink CSI or receiving downlink CSI fed back by the terminal; and executing the target action to acquire the downlink CSI. Thus, the target action for acquiring downlink CSI can be determined according to the state information of the terminal, the target action is predicting downlink CSI or receiving downlink CSI fed back by the terminal, and the target action is executed to acquire the downlink CSI. Downlink CSI can be acquired by selecting to predict downlink CSI or to receive downlink CSI fed back by the terminal according to the state information of the terminal, the signaling overhead consumed for acquiring downlink CSI is greatly reduced, and the accuracy and reliability of acquiring downlink CSI are higher.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, and in particular to a method for acquiring downlink channel state information (CSI), a model training method, a base station, a terminal, an apparatus, and a storage medium. BACKGROUND

[0002] At present, with the vigorous development of network technology, massive MIMO (massive MIMO) has been widely applied, and the number of antennas of a base station has gradually developed from the initial 8 antennas to 16, 32, 64, 256, 1024 antennas, etc. Massive MIMO has the advantages of large capacity and high spectrum utilization rate. However, with the increase in the number of antennas, the base station has the problem of large signaling overhead in acquiring downlink channel state information (CSI). SUMMARY

[0003] The present application provides a method for acquiring downlink channel state information (CSI), a model training method, a base station, a terminal, an apparatus, and a storage medium, to solve the technical problem of large signaling overhead in acquiring downlink channel state information by a base station in the related art.

[0004] According to a first aspect of the present application, a method for acquiring downlink channel state information (CSI) is provided, and the execution subject is a base station. The method comprises: acquiring state information of a terminal; determining a target action for acquiring downlink CSI according to the state information, wherein the target action is to predict downlink CSI or receive downlink CSI fed back by the terminal; and performing the target action to acquire the downlink CSI.

[0005] In an embodiment of the present application, the step of determining a target action for acquiring downlink CSI according to the state information comprises: inputting the state information into a trained reinforcement learning model, determining a cumulative reward of each candidate action under the state information by the reinforcement learning model, and determining the candidate action with the largest cumulative reward as the target action.

[0006] In an embodiment of the present application, the reinforcement learning model is configured to acquire a target parameter according to the state information, and acquire the cumulative reward according to the target parameter and signaling overhead corresponding to the candidate action for feeding back downlink CSI; wherein the target parameter comprises a block error rate of a physical downlink shared channel based on downlink CSI scheduling and / or beamforming, and / or an error between downlink CSI acquired by performing the candidate action and reference downlink CSI.

[0007] In an embodiment of the present application, in the case that the target action is the predicted downlink CSI, the performing the target action comprises: receiving a sounding reference signal (SRS) sent by the terminal, and predicting the downlink CSI according to the SRS.

[0008] In an embodiment of the present application, the predicting the downlink CSI according to the SRS comprises: acquiring uplink CSI according to the SRS; and predicting the downlink CSI according to the uplink CSI.

[0009] In an embodiment of the present application, before the receiving the SRS sent by the terminal, the method further comprises: sending first indication information to the terminal, wherein the first indication information is used to instruct the terminal to send the SRS.

[0010] In an embodiment of the present application, in the case that the target action is the predicted downlink CSI, the performing the target action comprises: acquiring historical downlink CSI fed back by the terminal, and predicting the downlink CSI according to the historical downlink CSI.

[0011] In an embodiment of the present application, in the case that the target action is the downlink CSI fed back by the terminal, the performing the target action comprises: sending a channel state information reference signal (CSI-RS) to the terminal, wherein the CSI-RS is used to instruct the terminal to acquire the downlink CSI based on the CSI-RS; and receiving the downlink CSI fed back by the terminal.

[0012] In an embodiment of the present application, before the sending the CSI-RS to the terminal, the method further comprises: sending second indication information to the terminal, wherein the second indication information is used to instruct the terminal to feed back the downlink CSI.

[0013] In an embodiment of the present application, the second indication information is further used to trigger setting of configuration information of downlink CSI feedback for the terminal, wherein the configuration information comprises at least one of the following: feedback quantity of downlink CSI, feedback period of downlink CSI, time domain resource of downlink CSI feedback, and frequency domain resource of downlink CSI feedback.

[0014] In an embodiment of the present application, the sending the CSI-RS to the terminal comprises: sending the CSI-RS to the terminal according to a sending period, wherein the sending period is equal to the feedback period.

[0015] In one embodiment of this application, obtaining the terminal's status information includes: receiving the status information sent by the terminal; and / or obtaining the pre-configured status information; and / or collecting the status information; and / or predicting the status information.

[0016] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0017] According to a second aspect of this application, another method for obtaining downlink channel state information (CSI) is provided, wherein the executing entity is a terminal, and the method includes: obtaining the terminal's own state information; determining a target action for obtaining the downlink CSI based on the state information, wherein the target action is a base station predicting the downlink CSI or feeding back the downlink CSI to the base station; and executing the target action to enable the base station to obtain the downlink CSI.

[0018] In one embodiment of this application, determining the target action for obtaining downlink CSI based on the state information includes: inputting the state information into a trained reinforcement learning model, having the reinforcement learning model determine the cumulative reward for each candidate action under the state information, and determining the candidate action with the largest cumulative reward as the target action.

[0019] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0020] In one embodiment of this application, when the target action is for the base station to predict the downlink CSI, the execution of the target action includes: sending first indication information to the base station, wherein the first indication information is used to instruct the base station to predict the downlink CSI.

[0021] In one embodiment of this application, when the target action is for the base station to predict the downlink CSI, the execution of the target action includes: sending a probe reference signal (SRS) to the base station, wherein the SRS is used to instruct the base station to predict the downlink CSI based on the SRS.

[0022] In one embodiment of this application, when the target action is to feed back downlink CSI to the base station, performing the target action includes: receiving a Channel State Information Reference Signal (CSI-RS) sent by the base station; obtaining the downlink CSI based on the CSI-RS; and feeding back the downlink CSI to the base station.

[0023] In one embodiment of this application, before receiving the Channel State Information Reference Signal (CSI-RS) sent by the base station, the method further includes: receiving second indication information sent by the base station, wherein the second indication information is used to trigger the setting of configuration information for downlink CSI feedback of the terminal itself, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0024] In one embodiment of this application, feeding back the downlink CSI to the base station includes: feeding back the downlink CSI to the base station according to the feedback period.

[0025] In one embodiment of this application, before receiving the Channel State Information Reference Signal (CSI-RS) sent by the base station, the method further includes: sending third indication information to the base station, wherein the third indication information is used to instruct the base station to send the CSI-RS.

[0026] In one embodiment of this application, obtaining the terminal's own status information includes: obtaining the pre-configured status information; and / or, collecting the status information.

[0027] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0028] According to a third aspect of this application, a model training method is provided, wherein the execution subject is a base station, the method comprising: acquiring training samples, wherein the training samples include sample state information of a terminal, a cumulative reward for each candidate action of a sample under the sample state information, and a target action of a sample, wherein the target action of a sample is the candidate action of the sample with the largest cumulative reward, and the target action of a sample is predicting downlink CSI or receiving downlink CSI feedback from the terminal; training a reinforcement learning model based on the training samples, and updating the model parameters of the reinforcement learning model; if the model training termination condition is not met, returning to use the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thereby generating a trained reinforcement learning model.

[0029] In one embodiment of this application, obtaining the cumulative sample reward for each candidate action under the sample state information includes: obtaining sample target parameters based on the sample state information, and obtaining the cumulative sample reward based on the sample target parameters and the sample signaling overhead for feeding back downlink CSI corresponding to the candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI obtained by executing the candidate action and the sample reference downlink CSI.

[0030] In one embodiment of this application, obtaining the sample target parameters based on the sample status information includes: when the status information of the terminal is the sample status information, receiving the Hybrid Automatic Repeat Request (HARQ) feedback information sent by the terminal, and obtaining the sample block error rate based on the HARQ feedback information.

[0031] In one embodiment of this application, the method further includes: when the state information of the terminal is the sample state information, performing the sample candidate action to obtain the sample downlink CSI.

[0032] In one embodiment of this application, the method further includes: when the state information of the terminal is the sample state information, receiving the sample reference downlink CSI sent by the terminal, wherein the sample reference downlink CSI is obtained based on the sample channel state information reference signal CSI-RS sent by the base station.

[0033] In one embodiment of this application, obtaining the sample status information of the terminal includes: receiving the sample status information sent by the terminal; and / or obtaining pre-configured sample status information; and / or collecting the sample status information; and / or predicting the sample status information.

[0034] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0035] According to a fourth aspect of this application, another model training method is provided, with the execution subject being a terminal. The method includes: acquiring training samples, wherein the training samples include sample state information of the terminal, a cumulative reward for each candidate action of the sample under the sample state information, and a target action of the sample, wherein the target action of the sample is the candidate action of the sample with the largest cumulative reward, and the target action of the sample is a base station predicting downlink CSI or feeding back downlink CSI to the base station; training a reinforcement learning model based on the training samples, and updating the model parameters of the reinforcement learning model; if the model training termination condition is not met, returning to use the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thereby generating a trained reinforcement learning model.

[0036] In one embodiment of this application, obtaining the cumulative sample reward for each candidate action under the sample state information includes: obtaining sample target parameters based on the sample state information, and obtaining the cumulative sample reward based on the sample target parameters and the sample signaling overhead for feeding back downlink CSI corresponding to the candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI obtained by executing the candidate action and the sample reference downlink CSI.

[0037] In one embodiment of this application, obtaining the sample target parameters based on the sample status information includes: when the status information of the terminal is the sample status information, collecting the sample error rate.

[0038] In one embodiment of this application, the method further includes: when the state information of the terminal is the sample state information, performing the sample candidate action to obtain the sample downlink CSI.

[0039] In one embodiment of this application, the method further includes: when the state information of the terminal is the sample state information, obtaining the sample reference downlink CSI based on the sample channel state information reference signal CSI-RS sent by the base station.

[0040] In one embodiment of this application, obtaining the sample status information of the terminal includes: obtaining pre-configured sample status information; and / or, collecting the sample status information.

[0041] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0042] According to a fifth aspect of this application, a base station is provided, including a memory, a transceiver, and a processor: the memory is used to store a computer program; the transceiver is used to transmit and receive data under the control of the processor; the processor is used to read the computer program in the memory and perform the following operations: acquiring terminal status information; determining a target action for acquiring downlink CSI based on the status information, wherein the target action is predicting downlink CSI or receiving downlink CSI fed back by the terminal; and executing the target action to acquire the downlink CSI.

[0043] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: inputting the state information into a trained reinforcement learning model, having the reinforcement learning model determine the cumulative reward for each candidate action under the state information, and determining the candidate action with the largest cumulative reward as the target action.

[0044] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0045] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: when the target action is the predicted downlink CSI, receive a probe reference signal (SRS) sent by the terminal, and predict the downlink CSI based on the SRS.

[0046] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: obtain an uplink CSI based on the SRS; and predict the downlink CSI based on the uplink CSI.

[0047] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: send a first instruction message to the terminal, wherein the first instruction message is used to instruct the terminal to send the SRS.

[0048] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: when the target action is the predicted downlink CSI, obtain the historical downlink CSI fed back by the terminal, and predict the downlink CSI based on the historical downlink CSI.

[0049] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: when the target action is to receive downlink CSI feedback from the terminal, sending a Channel State Information Reference Signal (CSI-RS) to the terminal, wherein the CSI-RS is used to instruct the terminal to obtain the downlink CSI based on the CSI-RS; and receiving the downlink CSI feedback from the terminal.

[0050] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: send a second indication message to the terminal, wherein the second indication message is used to instruct the terminal to provide feedback on the downlink CSI.

[0051] In one embodiment of this application, the second indication information is further used to trigger the setting of configuration information for downlink CSI feedback of the terminal, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0052] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: send the CSI-RS to the terminal according to a transmission period, wherein the transmission period is equal to the feedback period.

[0053] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: receive the status information sent by the terminal; and / or acquire the pre-configured status information; and / or collect the status information; and / or predict the status information.

[0054] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0055] According to a sixth aspect of this application, a terminal is provided, including a memory, a transceiver, and a processor: the memory is used to store a computer program; the transceiver is used to send and receive data under the control of the processor; the processor is used to read the computer program in the memory and perform the following operations: acquiring the terminal's own state information; determining a target action for acquiring downlink CSI based on the state information, wherein the target action is a base station predicting downlink CSI or feeding back downlink CSI to the base station; and executing the target action to enable the base station to acquire the downlink CSI.

[0056] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: inputting the state information into a trained reinforcement learning model, having the reinforcement learning model determine the cumulative reward for each candidate action under the state information, and determining the candidate action with the largest cumulative reward as the target action.

[0057] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0058] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: when the target action is the base station predicting downlink CSI, sending first indication information to the base station, wherein the first indication information is used to instruct the base station to predict the downlink CSI.

[0059] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: when the target action is the base station predicting downlink CSI, sending a probe reference signal (SRS) to the base station, wherein the SRS is used to instruct the base station to predict the downlink CSI based on the SRS.

[0060] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: when the target action is to feed back downlink CSI to the base station, receive the channel state information reference signal CSI-RS sent by the base station; obtain the downlink CSI according to the CSI-RS; and feed back the downlink CSI to the base station.

[0061] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: receiving second indication information sent by the base station, wherein the second indication information is used to trigger the setting of configuration information for downlink CSI feedback of the terminal itself, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0062] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: feed back the downlink CSI to the base station according to the feedback period.

[0063] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: send third indication information to the base station, wherein the third indication information is used to instruct the base station to send the CSI-RS.

[0064] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: obtain the pre-configured status information; and / or collect the status information.

[0065] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0066] According to a seventh aspect of this application, another base station is provided, comprising a memory, a transceiver, and a processor: the memory for storing a computer program; the transceiver for transmitting and receiving data under the control of the processor; and the processor for reading the computer program in the memory and performing the following operations: acquiring training samples, wherein the training samples include sample state information of a terminal, a cumulative reward for each candidate action of a sample under the sample state information, and a target action of a sample, wherein the target action of a sample is the candidate action of the sample with the largest cumulative reward, and the target action of a sample is predicting downlink CSI or receiving downlink CSI feedback from the terminal; training a reinforcement learning model based on the training samples, and updating the model parameters of the reinforcement learning model; and, if the model training termination condition is not met, returning to use the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thereby generating a trained reinforcement learning model.

[0067] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: obtain sample target parameters according to the sample state information, and obtain the sample cumulative reward according to the sample target parameters and the sample signaling overhead for feedback of downlink CSI corresponding to the sample candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI obtained by executing the sample candidate action and the sample reference downlink CSI.

[0068] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, receive the Hybrid Automatic Repeat Request (HARQ) feedback information sent by the terminal, and obtain the sample block error rate based on the HARQ feedback information.

[0069] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, execute the sample candidate action to obtain the sample downlink CSI.

[0070] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: when the state information of the terminal is the sample state information, receiving the sample reference downlink CSI sent by the terminal, wherein the sample reference downlink CSI is obtained based on the sample channel state information reference signal CSI-RS sent by the base station.

[0071] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: receive the sample status information sent by the terminal; and / or acquire the pre-configured sample status information; and / or collect the sample status information; and / or predict the sample status information.

[0072] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0073] According to an eighth aspect of this application, another terminal is provided, comprising a memory, a transceiver, and a processor: the memory for storing a computer program; the transceiver for transmitting and receiving data under the control of the processor; and the processor for reading the computer program in the memory and performing the following operations: acquiring training samples, wherein the training samples include sample state information of the terminal, a cumulative reward for each candidate action of the sample under the sample state information, and a target action of the sample, wherein the target action of the sample is the candidate action of the sample with the largest cumulative reward, and the target action of the sample is a base station predicting downlink CSI or feeding back downlink CSI to the base station; training a reinforcement learning model based on the training samples, and updating the model parameters of the reinforcement learning model; and, if the model training termination condition is not met, returning to use the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thereby generating a trained reinforcement learning model.

[0074] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: obtain sample target parameters according to the sample state information, and obtain the sample cumulative reward according to the sample target parameters and the sample signaling overhead for feedback of downlink CSI corresponding to the sample candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI obtained by executing the sample candidate action and the sample reference downlink CSI.

[0075] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, collect the sample error rate.

[0076] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, execute the sample candidate action to obtain the sample downlink CSI.

[0077] In one embodiment of this application, the processor is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, obtain the sample reference downlink CSI based on the sample channel state information reference signal CSI-RS sent by the base station.

[0078] In one embodiment of this application, the processor is further configured to read a computer program in the memory and perform the following operations: obtain the pre-configured sample status information; and / or collect the sample status information.

[0079] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0080] According to a ninth aspect of this application, an apparatus for acquiring downlink channel state information (CSI) is provided, comprising: an acquisition module for acquiring state information of a terminal; a determination module for determining a target action for acquiring downlink CSI based on the state information, wherein the target action is predicting downlink CSI or receiving downlink CSI feedback from the terminal; and an execution module for executing the target action to acquire the downlink CSI.

[0081] In one embodiment of this application, the determining module is further configured to: input the state information into a trained reinforcement learning model, have the reinforcement learning model determine the cumulative reward of each candidate action under the state information, and determine the candidate action with the largest cumulative reward as the target action.

[0082] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0083] In one embodiment of this application, when the target action is the predicted downlink CSI, the execution module is further configured to: receive a probe reference signal (SRS) sent by the terminal, and predict the downlink CSI based on the SRS.

[0084] In one embodiment of this application, the execution module is further configured to: obtain the uplink CSI based on the SRS; and predict the downlink CSI based on the uplink CSI.

[0085] In one embodiment of this application, the apparatus for obtaining downlink channel state information (CSI) further includes a sending module, which is configured to send first indication information to the terminal, wherein the first indication information is used to instruct the terminal to send the SRS.

[0086] In one embodiment of this application, when the target action is the predicted downlink CSI, the execution module is further configured to: obtain the historical downlink CSI fed back by the terminal, and predict the downlink CSI based on the historical downlink CSI.

[0087] In one embodiment of this application, when the target action is to receive downlink CSI feedback from the terminal, the execution module is further configured to: send a Channel State Information Reference Signal (CSI-RS) to the terminal, wherein the CSI-RS is used to instruct the terminal to obtain the downlink CSI based on the CSI-RS; and receive the downlink CSI feedback from the terminal.

[0088] In one embodiment of this application, the apparatus for obtaining downlink channel state information (CSI) further includes a sending module, which is further configured to send second indication information to the terminal, wherein the second indication information is used to instruct the terminal to provide feedback on the downlink CSI.

[0089] In one embodiment of this application, the second indication information is further used to trigger the setting of configuration information for downlink CSI feedback of the terminal, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0090] In one embodiment of this application, the execution module is further configured to: send the CSI-RS to the terminal according to the sending period, wherein the sending period is equal to the feedback period.

[0091] In one embodiment of this application, the acquisition module is further configured to: receive the status information sent by the terminal; and / or acquire the pre-configured status information; and / or collect the status information; and / or predict the status information.

[0092] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0093] According to a tenth aspect of this application, another apparatus for acquiring downlink channel state information (CSI) is provided, comprising: an acquisition module for acquiring state information of a terminal itself; a determination module for determining a target action for acquiring downlink CSI based on the state information, wherein the target action is a base station predicting downlink CSI or feeding back downlink CSI to the base station; and an execution module for executing the target action to enable the base station to acquire the downlink CSI.

[0094] In one embodiment of this application, the determining module is further configured to: input the state information into a trained reinforcement learning model, have the reinforcement learning model determine the cumulative reward of each candidate action under the state information, and determine the candidate action with the largest cumulative reward as the target action.

[0095] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0096] In one embodiment of this application, when the target action is for the base station to predict the downlink CSI, the execution module is further configured to: send first indication information to the base station, wherein the first indication information is used to instruct the base station to predict the downlink CSI.

[0097] In one embodiment of this application, when the target action is for the base station to predict the downlink CSI, the execution module is further configured to: send a probe reference signal (SRS) to the base station, wherein the SRS is used to instruct the base station to predict the downlink CSI based on the SRS.

[0098] In one embodiment of this application, when the target action is to feed back downlink CSI to the base station, the execution module is further configured to: receive a Channel State Information Reference Signal (CSI-RS) sent by the base station; obtain the downlink CSI based on the CSI-RS; and feed back the downlink CSI to the base station.

[0099] In one embodiment of this application, the apparatus for obtaining downlink channel state information (CSI) further includes a receiving module, which is configured to receive second indication information sent by the base station, wherein the second indication information is used to trigger the setting of configuration information for downlink CSI feedback of the terminal itself, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0100] In one embodiment of this application, the execution module is further configured to: feed back the downlink CSI to the base station according to the feedback period.

[0101] In one embodiment of this application, the apparatus for obtaining downlink channel state information (CSI) further includes a transmitting module, which is configured to transmit third indication information to the base station, wherein the third indication information is used to instruct the base station to transmit the CSI-RS.

[0102] In one embodiment of this application, the acquisition module is further configured to: acquire the pre-configured status information; and / or collect the status information.

[0103] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0104] According to the eleventh aspect of this application, a model training apparatus is provided, comprising: an acquisition module for acquiring training samples, wherein the training samples include sample state information of a terminal, a cumulative reward for each candidate action of a sample under the sample state information, and a target action of a sample, wherein the target action of a sample is the candidate action of the sample with the largest cumulative reward, and the target action of a sample is predicting downlink CSI or receiving downlink CSI feedback from the terminal; a training module for training a reinforcement learning model based on the training samples and updating the model parameters of the reinforcement learning model; the training module is further configured to, if the model training termination condition is not met, return to using the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thereby generating a trained reinforcement learning model.

[0105] In one embodiment of this application, the acquisition module is further configured to: acquire sample target parameters based on the sample state information, and acquire the sample cumulative reward based on the sample target parameters and the sample signaling overhead for feedback of downlink CSI corresponding to the sample candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI acquired by executing the sample candidate action and the sample reference downlink CSI.

[0106] In one embodiment of this application, the acquisition module is further configured to: when the status information of the terminal is the sample status information, receive the Hybrid Automatic Repeat Request (HARQ) feedback information sent by the terminal, and obtain the sample block error rate based on the HARQ feedback information.

[0107] In one embodiment of this application, the acquisition module is further configured to: when the status information of the terminal is the sample status information, perform the sample candidate action to obtain the sample downlink CSI.

[0108] In one embodiment of this application, the acquisition module is further configured to: receive the sample reference downlink CSI sent by the terminal when the terminal's status information is the sample status information, wherein the sample reference downlink CSI is obtained based on the sample channel state information reference signal CSI-RS sent by the base station.

[0109] In one embodiment of this application, the acquisition module is further configured to: receive the sample status information sent by the terminal; and / or acquire the pre-configured sample status information; and / or collect the sample status information; and / or predict the sample status information.

[0110] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0111] According to the twelfth aspect of this application, another model training apparatus is provided, comprising: an acquisition module for acquiring training samples, wherein the training samples include sample state information of a terminal, a cumulative reward for each candidate action of a sample under the sample state information, and a target action of a sample, wherein the target action of the sample is the candidate action of the sample with the largest cumulative reward, and the target action of the sample is a base station predicting downlink CSI or feeding back downlink CSI to the base station; a training module for training a reinforcement learning model based on the training samples and updating the model parameters of the reinforcement learning model; the training module is further configured to, if the model training termination condition is not met, return to using the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thereby generating a trained reinforcement learning model.

[0112] In one embodiment of this application, the acquisition module is further configured to: acquire sample target parameters based on the sample state information, and acquire the sample cumulative reward based on the sample target parameters and the sample signaling overhead for feedback of downlink CSI corresponding to the sample candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI acquired by executing the sample candidate action and the sample reference downlink CSI.

[0113] In one embodiment of this application, the acquisition module is further configured to: collect the sample error rate when the terminal's status information is the sample status information.

[0114] In one embodiment of this application, the acquisition module is further configured to: when the status information of the terminal is the sample status information, perform the sample candidate action to obtain the sample downlink CSI.

[0115] In one embodiment of this application, the acquisition module is further configured to: when the state information of the terminal is the sample state information, acquire the sample reference downlink CSI based on the sample channel state information reference signal CSI-RS sent by the base station.

[0116] In one embodiment of this application, the acquisition module is further configured to: acquire the pre-configured sample status information; and / or collect the sample status information.

[0117] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0118] According to a thirteenth aspect of this application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method for obtaining downlink channel state information (CSI) as described in the first aspect embodiment of this application, or to perform the method for obtaining downlink channel state information (CSI) as described in the second aspect embodiment of this application, or to perform the model training method as described in the third aspect embodiment of this application, or to perform the model training method as described in the fourth aspect embodiment of this application.

[0119] According to a fourteenth aspect of this application, a processor-readable storage medium is provided, the processor-readable storage medium storing a computer program for causing the processor to perform the method for acquiring downlink channel state information (CSI) as described in the first aspect embodiment.

[0120] According to the fifteenth aspect of this application, a processor-readable storage medium is provided, the processor-readable storage medium storing a computer program for causing the processor to perform the method for acquiring downlink channel state information (CSI) as described in the second aspect embodiment.

[0121] According to a sixteenth aspect of this application, a processor-readable storage medium is provided, the processor-readable storage medium storing a computer program for causing the processor to perform the model training method described in the third aspect embodiment.

[0122] According to the seventeenth aspect of this application, a processor-readable storage medium is provided, the processor-readable storage medium storing a computer program for causing the processor to perform the model training method described in the fourth aspect embodiment.

[0123] The technical solution provided by the embodiments of this application brings at least the following beneficial effects: Based on the terminal's state information, a target action for obtaining downlink CSI can be determined. The target action is either predicting downlink CSI or receiving downlink CSI feedback from the terminal, and the target action is executed to obtain the downlink CSI. Therefore, by selecting either predicted downlink CSI or receiving downlink CSI feedback from the terminal based on the terminal's state information to obtain the downlink CSI, the signaling overhead consumed in obtaining the downlink CSI is greatly reduced, and the accuracy and reliability of obtaining the downlink CSI are high.

[0124] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0125] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0126] Figure 1 This is a flowchart illustrating a method for obtaining downlink channel state information (CSI) according to an embodiment of this application.

[0127] Figure 2 This is a schematic diagram of a method for obtaining downlink channel state information (CSI) according to an embodiment of this application.

[0128] Figure 3 This is a schematic diagram of a method for obtaining downlink channel state information (CSI) according to another embodiment of this application.

[0129] Figure 4 This is a flowchart illustrating a method for obtaining downlink channel state information (CSI) according to another embodiment of this application;

[0130] Figure 5 This is a schematic flowchart of a model training method according to an embodiment of this application;

[0131] Figure 6 This is a block diagram of a base station and a terminal according to an embodiment of this application;

[0132] Figure 7 This is a schematic flowchart of a model training method according to another embodiment of this application;

[0133] Figure 8 This is a block diagram of a base station and a terminal according to another embodiment of this application;

[0134] Figure 9 This is a block diagram of a base station according to an embodiment of this application;

[0135] Figure 10 This is a block diagram of a terminal according to an embodiment of this application;

[0136] Figure 11 This is a block diagram of a base station according to another embodiment of this application;

[0137] Figure 12 This is a block diagram of a terminal according to another embodiment of this application;

[0138] Figure 13This is a block diagram of an apparatus for obtaining downlink channel state information (CSI) according to an embodiment of this application;

[0139] Figure 14 This is a block diagram of an apparatus for obtaining downlink channel state information (CSI) according to another embodiment of this application;

[0140] Figure 15 This is a block diagram of a model training apparatus according to an embodiment of this application;

[0141] Figure 16 This is a block diagram of a model training apparatus according to another embodiment of this application. Detailed Implementation

[0142] In the embodiments of this application, the term "and / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0143] In the embodiments of this application, the term "multiple" refers to two or more, and other quantifiers are similar.

[0144] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0145] This application provides a method, model training method, base station, terminal, device, electronic device, and storage medium for obtaining downlink channel state information (CSI), which solves the technical problem of large signaling overhead in base station acquisition of downlink channel state information (CSI) in related technologies.

[0146] The method and apparatus are based on the same concept of the application. Since the methods and apparatus solve problems in similar ways, the implementation of the apparatus and methods can refer to each other, and the repeated parts will not be described again.

[0147] Figure 1 This is a flowchart illustrating a method for obtaining downlink channel state information (CSI) according to an embodiment of this application.

[0148] like Figure 1 As shown in the embodiment of this application, the method for obtaining downlink channel state information (CSI) includes:

[0149] S101, Obtain the terminal's status information.

[0150] It should be noted that the execution subject of the method for obtaining downlink channel state information (CSI) in this application embodiment can be a base station.

[0151] In the embodiments of this application, the base station can obtain the terminal's status information. It should be noted that there are no excessive limitations on the method by which the base station obtains the terminal's status information, nor are there excessive limitations on the type of status information.

[0152] In one implementation, the terminal's status information includes, but is not limited to, at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency. For example, carrier frequency offset refers to the carrier frequency offset between the terminal's transmitter and receiver.

[0153] In one implementation, obtaining the terminal's status information may include at least one of the following implementation methods:

[0154] Method 1: Receive status information sent by the terminal.

[0155] In one implementation, the base station can obtain status information transmitted by the terminal on the uplink channel. It should be noted that the type of uplink channel is not strictly limited; for example, uplink channels include, but are not limited to, the Physical Uplink Shared Channel (PUSCH) and the Physical Uplink Control Channel (PUCCH). For instance, the base station can obtain the terminal's transmission speed, received signal-to-noise ratio, and carrier frequency offset on the PUSCH.

[0156] Method 2: Obtain pre-configured status information.

[0157] In one implementation, the base station can pre-configure the terminal's status information and store the pre-configured status information in the base station's storage space, and then retrieve the pre-configured status information from the base station's storage space. For example, the base station can retrieve the pre-configured uplink carrier frequency and downlink carrier frequency.

[0158] Method 3: Collect status information.

[0159] In one implementation, the base station can collect the terminal's status information.

[0160] For example, a base station can send and / or receive acquisition signals to collect terminal status information.

[0161] For example, a base station can collect terminal status information through a data acquisition device. For instance, a base station can use radar to collect the terminal's speed.

[0162] Method 4: Predict state information.

[0163] In one implementation, the base station can predict the status information of the terminal.

[0164] For example, a base station can predict the terminal's status information based on the terminal's historical status information. For instance, a base station can predict the terminal's current speed based on the speed between two minutes prior to the current time.

[0165] For example, a base station can predict the speed of a terminal based on its location. For instance, a base station can predict the speed of a terminal at its current moment based on its historical location two minutes prior and its current location.

[0166] S102, based on the status information, determine the target action for obtaining downlink CSI, wherein the target action is to predict downlink CSI or receive downlink CSI feedback from the terminal.

[0167] In the embodiments of this application, the base station can determine the target action for acquiring downlink CSI based on state information. The target action is either predicted downlink CSI or downlink CSI fed back by the receiving terminal. It should be noted that the base station itself can predict downlink CSI or receive downlink CSI fed back by the receiving terminal. The types of downlink CSI are not limited in detail; for example, downlink CSI includes, but is not limited to, information such as signal scattering, environmental attenuation, and distance attenuation.

[0168] In one implementation, determining the target action for obtaining downlink CSI based on the status information may include pre-establishing a mapping relationship or mapping table between the status information and the target action. After obtaining the status information, querying the aforementioned mapping relationship or mapping table can retrieve the target action mapped to the status information. It should be noted that the aforementioned mapping relationship or mapping table is not subject to excessive limitations.

[0169] In one implementation, determining the target action for acquiring downlink CSI based on status information may include identifying whether the current conditions for predicting downlink CSI are met based on the status information. If the current conditions for predicting downlink CSI are met, the target action for acquiring downlink CSI may be determined to be predicting downlink CSI; or, if the current conditions for predicting downlink CSI are not met, the target action for acquiring downlink CSI may be determined to be receiving downlink CSI feedback from the receiving terminal.

[0170] It should be noted that the setting conditions for predicted downlink CSI are not subject to excessive restrictions. For example, the setting conditions for predicted downlink CSI may include a carrier frequency offset less than a first set threshold, and / or a received signal-to-noise ratio greater than a second set threshold. It should be noted that neither the first nor the second set threshold is subject to excessive restrictions.

[0171] S103, Execute the target action to obtain downlink CSI.

[0172] In embodiments of this disclosure, the base station can perform a target action to obtain downlink CSI. For example, if the target action is to predict downlink CSI, the base station can predict the downlink CSI; or, if the target action is to receive downlink CSI fed back by the terminal, the base station can receive the downlink CSI fed back by the terminal.

[0173] In one implementation, the base station can send a reminder message carrying the target action to the terminal in a timely manner, so as to inform the terminal of the target action and facilitate the base station to execute the target action subsequently.

[0174] In one implementation, when the target action is to predict downlink CSI, executing the target action may include at least one of the following implementations:

[0175] Method 1: Receive the Sounding Reference Signal (SRS) sent by the receiving terminal, and predict the downlink CSI based on the SRS.

[0176] In one implementation, such as Figure 2 As shown, the base station can receive the SRS sent by the terminal and predict the downlink CSI based on the SRS. For example, the base station can obtain the uplink CSI based on the SRS and predict the downlink CSI based on the uplink CSI. For example, the base station can perform uplink channel estimation based on the SRS to obtain the uplink CSI.

[0177] In one embodiment, before receiving the SRS sent by the terminal, the method further includes sending first indication information to the terminal, wherein the first indication information is used to instruct the terminal to send the SRS.

[0178] Method 2: Obtain the historical downlink CSI from the terminal feedback, and predict the downlink CSI based on the historical downlink CSI.

[0179] In one implementation, the base station can obtain historical downlink CSIs fed back by the terminal and predict the downlink CSI based on these historical downlink CSIs. For example, the base station can store the historical downlink CSIs fed back by the terminal in its storage space and retrieve the historical downlink CSIs from the base station's storage space. Alternatively, the base station can sort the historical downlink CSIs according to their feedback time from earliest to latest, and predict the downlink CSI based on the N sorted historical downlink CSIs.

[0180] Method 3: Receive the SRS sent by the terminal, predict the first downlink CSI based on the SRS, obtain the historical downlink CSI fed back by the terminal, predict the second downlink CSI based on the historical downlink CSI, and obtain the downlink CSI based on the first downlink CSI and the second downlink CSI.

[0181] In one implementation, the base station can receive the SRS sent by the terminal, predict a first downlink CSI based on the SRS, obtain the historical downlink CSI fed back by the terminal, predict a second downlink CSI based on the historical downlink CSI, and obtain the downlink CSI based on the first and second downlink CSIs. For example, the average of the first and second downlink CSIs can be used as the downlink CSI.

[0182] It should be noted that the relevant content regarding the prediction of the first downlink CSI and the prediction of the second downlink CSI can be found in the above embodiments, and will not be repeated here.

[0183] In one implementation, such as Figure 3 As shown, when the target action is to receive downlink CSI feedback from the terminal, executing the target action may include sending a Channel State Information Reference Signal (CSI-RS) to the terminal, wherein the CSI-RS is used to instruct the terminal to obtain downlink CSI based on the CSI-RS and to receive downlink CSI feedback from the terminal.

[0184] In one embodiment, before sending CSI-RS to the terminal, a second indication message is sent to the terminal, wherein the second indication message is used to instruct the terminal to provide downlink CSI feedback.

[0185] In one embodiment, the second indication information is further used to trigger the setting of configuration information for downlink CSI feedback of the terminal, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0186] In one implementation, the base station can pre-configure the downlink CSI feedback configuration information of the terminal. For example, the base station sets the configuration information according to the target action, and different target actions can correspond to different configuration information. Accordingly, the second indication information is used to trigger the setting of the configuration information corresponding to the target action, that is, to set the downlink CSI feedback configuration information of the terminal to the configuration information corresponding to the target action.

[0187] In one implementation, sending CSI-RS to the terminal may include sending CSI-RS to the terminal according to a sending period, wherein the sending period is equal to the feedback period.

[0188] In summary, the method for obtaining downlink channel state information (CSI) according to the embodiments of this application can determine the target action for obtaining downlink CSI based on the terminal's state information. The target action is either predicting downlink CSI or receiving downlink CSI feedback from the terminal, and then executing the target action to obtain the downlink CSI. Therefore, by selecting either predicted downlink CSI or receiving downlink CSI feedback from the terminal based on the terminal's state information, the signaling overhead consumed in obtaining downlink CSI is greatly reduced, and the accuracy and reliability of obtaining downlink CSI are high.

[0189] Based on any of the above embodiments, step S102, which determines the target action for obtaining downlink CSI based on the state information, may include inputting the state information into a trained reinforcement learning model, having the reinforcement learning model determine the cumulative reward for each candidate action under the state information, and determining the candidate action with the largest cumulative reward as the target action. It should be noted that the reinforcement learning model is not subject to many limitations; it can be pre-set in the base station's storage space.

[0190] In one implementation, the reinforcement learning model is used to obtain target parameters based on state information, and to obtain a cumulative reward based on the target parameters and the signaling overhead for feeding back downlink CSI corresponding to the candidate actions. It should be noted that there are no strict limitations on the signaling overhead for feeding back downlink CSI corresponding to the candidate actions; different candidate actions can correspond to different signaling overheads. There are no strict limitations on the category of target parameters; for example, target parameters may include the block error rate (BLER) of the Physical Downlink Shared Channel (PDSCH) for scheduling and / or beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate actions and the reference downlink CSI.

[0191] In one implementation, the reinforcement learning model is used to predict the BLER of PDSCH based on downlink CSI scheduling and / or beamforming, the error between the downlink CSI obtained by executing candidate actions and the reference downlink CSI, and to obtain a cumulative reward based on the BLER, the error, and the signaling overhead for feeding back downlink CSI corresponding to the candidate actions.

[0192] For example, candidate action 'a' includes a0, a1 to a0. K Where a0 is the predicted downlink CSI, and a1 to a KTo receive downlink CSI feedback from the receiving terminal, the cumulative reward is obtained based on BLER, error, and the signaling overhead for feeding back downlink CSI corresponding to the candidate action. This can be achieved using the following formula:

[0193]

[0194]

[0195] Where Q(s,a) is the cumulative reward of candidate action a under state information s, and R is the reward of candidate action a under state information s, where R includes r0, r1 to r K , and r er r0, r1 to r K These are candidate actions a0, a1 to a2 under state information s. K The rewards, r0, r1 to r K Decreasing sequentially, and r0, r1 to r K All are greater than 0, r er ≤0. r0, r1 to r K , and r er All values ​​are obtained based on the error and the signaling overhead used for feedback of downlink CSI corresponding to the candidate action. η0, η1 to η K These are candidate actions a0, a1 to a2 under state information s. K The corresponding signaling overhead used for feedback of downlink CSI, f1 to f K Candidate actions a1 to a2 are respectively. K The corresponding maximum allowed signaling overhead for downlink CSI feedback, e0 is the maximum allowed BLER of PDSCH. It should be noted that for f1 to f... K Neither e0 nor e0 is subject to many restrictions.

[0196] In one implementation, the reinforcement learning model is used to predict the BLER of PDSCH based on downlink CSI scheduling and / or beamforming according to state information, and to obtain a cumulative reward based on the BLER.

[0197] For example, if candidate action 'a' includes a0 and a1, where a0 is the predicted downlink CSI and a1 is the downlink CSI fed back by the receiving terminal, then the cumulative reward based on BLER can be obtained through the following formula:

[0198]

[0199]

[0200] Where Q(s,a) is the cumulative reward of candidate action a under state information s, and R is the reward of candidate action a under state information s, where R includes r0, r1, r2, r3, r4, r5, r6, r7, r8, r9, r1, r1, r2, r3, r4 ...3, r4, r erLet r0 and r1 be the rewards for candidate actions a0 and a1 under state information s, respectively, where r0 > r1, and both r0 and r1 are greater than 0. er ≤0. η0 and η1 are the signaling overheads for feedback of downlink CSI corresponding to candidate actions a0 and a1 under state information s, respectively, and e0 is the maximum value allowed by the BLER of PDSCH. It should be noted that for r0, r1, r er Neither r0 nor e0 is subject to many restrictions. For example, a reward-based pruning method can be used, where r0 = 1, r1 = 0, and r... er =-1, at which point the reinforcement learning model can use the same hyperparameters, which helps to simplify the reinforcement learning model.

[0201] Figure 4 This is a flowchart illustrating a method for obtaining downlink channel state information (CSI) according to another embodiment of this application.

[0202] like Figure 4 As shown in the embodiment of this application, the method for obtaining downlink channel state information (CSI) includes:

[0203] S401, obtain the terminal's own status information.

[0204] It should be noted that the execution subject of the method for obtaining downlink channel state information (CSI) in this application embodiment can be a terminal.

[0205] In the embodiments of this application, the terminal can obtain its own status information. It should be noted that there are no excessive limitations on the way the terminal obtains its own status information, nor are there excessive limitations on the type of status information.

[0206] In one implementation, the terminal's status information includes, but is not limited to, at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency. For example, carrier frequency offset refers to the carrier frequency offset between the terminal's transmitter and receiver.

[0207] In one implementation, obtaining the terminal's own status information may include at least one of the following implementation methods:

[0208] Method 1: Obtain pre-configured status information.

[0209] In one implementation, the terminal's status information can be pre-configured by the base station and stored in the terminal's storage space, and then the pre-configured status information can be retrieved from the terminal's storage space. For example, the terminal can retrieve the pre-configured uplink carrier frequency and downlink carrier frequency.

[0210] Method 2: Collect status information.

[0211] In one implementation, the terminal can collect its own status information.

[0212] For example, a terminal can send and / or receive acquisition signals to collect its own status information.

[0213] For example, a terminal can collect its own status information through data acquisition devices. For instance, a terminal can collect its own speed through devices such as speed sensors or radar. It can also collect data such as the received signal-to-noise ratio and carrier frequency offset.

[0214] S402, based on the status information, determine the target action for obtaining downlink CSI, wherein the target action is the base station predicting downlink CSI or feeding back downlink CSI to the base station.

[0215] In the embodiments of this application, the terminal can determine the target action for acquiring downlink CSI based on the status information. The target action is either the base station predicting downlink CSI or feeding back downlink CSI to the base station. It should be noted that the terminal can feed back downlink CSI to the base station. The types of downlink CSI are not limited in detail; for example, downlink CSI includes, but is not limited to, information such as signal scattering, environmental attenuation, and distance attenuation.

[0216] It should be noted that the relevant content for determining the target action used to obtain downlink CSI based on the status information can be found in the above embodiments, and will not be repeated here.

[0217] In one implementation, determining the target action for acquiring downlink CSI based on state information may include inputting the state information into a trained reinforcement learning model, whereby the reinforcement learning model determines the cumulative reward for each candidate action under the given state information, and the candidate action with the highest cumulative reward is determined as the target action. It should be noted that the reinforcement learning model is not subject to many limitations and can be pre-set in the terminal's storage space.

[0218] In one implementation, the reinforcement learning model is used to obtain target parameters based on state information, and to obtain cumulative rewards based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate actions. The target parameters include the block error rate of the physical downlink shared channel for scheduling and beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate actions and the reference downlink CSI.

[0219] It should be noted that the relevant content of reinforcement learning models can be found in the above embodiments, and will not be repeated here.

[0220] S403, Execute the target action to enable the base station to obtain downlink CSI.

[0221] In embodiments of this disclosure, the terminal may perform a target action to enable the base station to acquire downlink CSI. For example, if the target action is for the base station to predict downlink CSI, the base station may predict the downlink CSI; or, if the target action is to feed back the downlink CSI to the base station, the base station may receive the downlink CSI fed back by the terminal.

[0222] In one implementation, the terminal can send a reminder message carrying the target action to the base station to inform the base station of the target action in a timely manner, so that the terminal can subsequently execute the target action.

[0223] In one implementation, when the target action is a base station predicted downlink CSI, executing the target action may include at least one of the following implementation methods:

[0224] Method 1: Send a first indication information to the base station, wherein the first indication information is used to instruct the base station to predict the downlink CSI.

[0225] In one implementation, the terminal may send a first indication message to the base station, the first indication message being used to instruct the base station to predict downlink CSI so as to promptly inform the base station to predict downlink CSI.

[0226] Method 2: Send SRS to the base station, where SRS is used to instruct the base station to predict downlink CSI based on SRS.

[0227] It should be noted that the relevant content of Method 2 can be found in the above embodiments, and will not be repeated here.

[0228] In one implementation, when the target action is to feed back downlink CSI to the base station, performing the target action may include receiving CSI-RS sent by the base station, obtaining the downlink CSI based on the CSI-RS, and feeding back the downlink CSI to the base station. For example, the terminal may perform downlink channel estimation based on the CSI-RS to obtain the downlink CSI.

[0229] In one embodiment, before receiving the CSI-RS sent by the base station, the system further includes receiving second indication information sent by the base station. The second indication information is used to trigger the setting of configuration information for the downlink CSI feedback of the terminal itself. The configuration information includes at least one of the following: the downlink CSI feedback amount, the downlink CSI feedback period, the downlink CSI feedback time domain resources, and the downlink CSI feedback frequency domain resources.

[0230] It should be noted that the relevant content regarding the second instruction information can be found in the above embodiments, and will not be repeated here.

[0231] In one implementation, feeding back downlink CSI to the base station may include feeding back downlink CSI to the base station according to a feedback period. Thus, the terminal can use multiple different feedback periods to feed back downlink CSI to the base station, resulting in higher accuracy of the downlink CSI fed back by the terminal to the base station.

[0232] In one embodiment, before receiving the CSI-RS sent by the base station, the method further includes sending third indication information to the base station, wherein the third indication information is used to instruct the base station to send CSI-RS.

[0233] In summary, the method for obtaining downlink channel state information (CSI) according to the embodiments of this application can determine the target action for obtaining downlink CSI based on the terminal's own state information. The target action is either for the base station to predict downlink CSI or to feed back downlink CSI to the base station, and the target action is executed to enable the base station to obtain downlink CSI. Therefore, by selecting either for the base station to predict or feed back downlink CSI based on the terminal's own state information, the signaling overhead incurred in obtaining downlink CSI is greatly reduced, and the accuracy and reliability of obtaining downlink CSI are high.

[0234] Figure 5 This is a schematic flowchart of a model training method according to an embodiment of this application.

[0235] like Figure 5 As shown, the model training method of this application embodiment includes:

[0236] S501, Obtain training samples, wherein the training samples include the terminal's sample state information, the cumulative reward of each sample candidate action under the sample state information, and the sample target action. The sample target action is the sample candidate action with the largest cumulative reward. The sample target action is the predicted downlink CSI or the downlink CSI fed back by the receiving terminal.

[0237] It should be noted that the execution subject of the model training method in this application embodiment can be a base station.

[0238] In the embodiments of this application, the base station can acquire a large number of training samples. Each training sample includes the terminal's sample state information, the cumulative reward for each candidate action under the sample state information, and the target action. The target action is the candidate action with the highest cumulative reward, and it is either a predicted downlink CSI or a received downlink CSI from the terminal. It should be noted that there are no excessive limitations on the method by which the base station acquires training samples, nor are there excessive limitations on the category of the sample state information.

[0239] In one implementation, the sample state information includes, but is not limited to, at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency. For example, sample carrier frequency offset refers to the carrier frequency offset between the transmitter and receiver of the terminal.

[0240] In one implementation, obtaining the sample status information of the terminal may include at least one of the following implementation methods:

[0241] Method 1: Receive sample status information sent by the receiving terminal.

[0242] Method 2: Obtain pre-configured sample status information.

[0243] Method 3: Collect sample status information.

[0244] Method 4: Predict sample state information.

[0245] It should be noted that the relevant content of methods 1 to 4 can be found in the above embodiments, and will not be repeated here.

[0246] In one implementation, obtaining the cumulative reward for each candidate action under sample state information may include obtaining sample target parameters based on the sample state information, and obtaining the cumulative reward based on the sample target parameters and the sample signaling overhead for feeding back downlink CSI corresponding to the candidate action. The sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI obtained by executing the candidate action and the sample reference downlink CSI.

[0247] In one implementation, obtaining sample target parameters based on sample status information may include, when the terminal's status information is sample status information, the base station may receive Hybrid Automatic Repeat Request (HARQ) feedback information sent by the terminal and obtain the sample block error rate based on the HARQ feedback information.

[0248] In one implementation, when the terminal's state information is sample state information, the base station can perform a sample candidate action to obtain sample downlink CSI. For example, when the sample candidate action is to predict downlink CSI, the base station can predict the sample downlink CSI; or, when the sample candidate action is to receive downlink CSI fed back by the terminal, the base station can receive the sample downlink CSI fed back by the terminal.

[0249] In one implementation, when the terminal's state information is sample state information, the base station can receive a sample reference downlink CSI sent by the terminal, wherein the sample reference downlink CSI is obtained based on the CSI-RS sent by the base station. It should be noted that the sample reference downlink CSI refers to the complete downlink CSI obtained by the terminal based on the CSI-RS.

[0250] S502, train the reinforcement learning model based on the training samples, and update the model parameters of the reinforcement learning model.

[0251] S503: If the model training termination condition is not met, return to the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, and a trained reinforcement learning model is generated.

[0252] In the embodiments of this application, a reinforcement learning model can be trained based on training samples, and the model parameters of the reinforcement learning model can be updated. If the model training termination condition is not met, the process returns to using the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thus generating a trained reinforcement learning model. It should be noted that there are no excessive limitations on the model training method or the model training termination condition. For example, the model training termination condition includes, but is not limited to, reaching a set number of training iterations or a set accuracy threshold.

[0253] In summary, the model training method according to the embodiments of this application can obtain training samples, wherein the training samples include the sample state information of the terminal, the sample cumulative reward of each sample candidate action under the sample state information, and the sample target action, and train a reinforcement learning model based on the training samples to generate a trained reinforcement learning model.

[0254] like Figure 6 As shown, the base station may include a reinforcement learning modeling module, a reinforcement learning training module, a reinforcement learning decision-making module, and a base station action execution module. The terminal may include a HARQ feedback module, a CSI estimation and reporting module, a status reporting module, and a terminal action execution module.

[0255] The modeling module is used to construct a reinforcement learning model, receive HARQ feedback information sent by the terminal's HARQ feedback module, and receive sample reference downlink CSI sent by the terminal's CSI estimation and reporting module. It also obtains the sample block error rate based on the HARQ feedback information, obtains the sample error between the sample downlink CSI obtained by executing the sample candidate action and the sample reference downlink CSI, and obtains the sample cumulative reward based on the sample block error rate, sample error, and the sample signaling overhead for the sample candidate action used to feed back the downlink CSI.

[0256] The training module receives sample state information from the terminal's state reporting module and sample cumulative reward and sample target action for each candidate action under the sample state information from the modeling module. Based on the sample state information, sample cumulative reward and sample target action, the training module trains a reinforcement learning model to generate the optimal policy, i.e., the trained reinforcement learning model.

[0257] The decision module receives status information from the terminal's status reporting module and inputs the status information into the trained reinforcement learning model to obtain the target action output by the reinforcement learning model.

[0258] The base station action execution module is used to receive the target action sent by the decision module and execute the target action to obtain downlink CSI.

[0259] The HARQ feedback module is used to send HARQ feedback information to the modeling module of the base station.

[0260] The CSI estimation and reporting module is used to obtain sample reference downlink CSI based on the CSI-RS sent by the base station and send the sample reference downlink CSI to the modeling module of the base station.

[0261] The status reporting module is used to send the terminal's sample status information to the base station's training module and to send the terminal's status information to the base station's decision module.

[0262] The terminal action execution module is used to receive the target action sent by the base station and execute the target action.

[0263] Figure 7 This is a schematic flowchart of a model training method according to another embodiment of this application.

[0264] like Figure 7 As shown, the model training method of this application embodiment includes:

[0265] S701, Obtain training samples, wherein the training samples include the terminal's sample state information, the cumulative reward of each candidate action under the sample state information, and the target action of the sample. The target action of the sample is the candidate action of the sample with the largest cumulative reward. The target action of the sample is the base station predicting downlink CSI or feeding back downlink CSI to the base station.

[0266] It should be noted that the execution subject of the model training method in this application embodiment can be a terminal.

[0267] In the embodiments of this application, the terminal can acquire a large number of training samples. Each training sample includes the terminal's sample state information, the cumulative reward of each candidate action under the sample state information, and the target action of the sample. The target action is the candidate action with the highest cumulative reward, and the target action is either the base station predicting downlink CSI or feeding back downlink CSI to the base station. It should be noted that there are no excessive restrictions on the method by which the terminal acquires training samples, nor are there excessive restrictions on the category of sample state information.

[0268] In one implementation, the sample state information includes, but is not limited to, at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency. For example, sample carrier frequency offset refers to the carrier frequency offset between the transmitter and receiver of the terminal.

[0269] In one implementation, obtaining the sample status information of the terminal may include at least one of the following implementation methods:

[0270] Method 1: Obtain pre-configured sample status information.

[0271] Method 2: Collect sample status information.

[0272] It should be noted that the relevant content of Method 1 to Method 2 can be found in the above embodiments, and will not be repeated here.

[0273] In one implementation, obtaining the cumulative reward for each candidate action under sample state information may include obtaining sample target parameters based on the sample state information, and obtaining the cumulative reward based on the sample target parameters and the sample signaling overhead for feeding back downlink CSI corresponding to the candidate action. The sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI obtained by executing the candidate action and the sample reference downlink CSI.

[0274] In one implementation, obtaining sample target parameters based on sample status information may include, when the terminal's status information is sample status information, the terminal may collect sample error rate.

[0275] In one implementation, when the terminal's state information is sample state information, the terminal performs a sample candidate action to obtain sample downlink CSI. For example, if the sample candidate action is for the base station to predict downlink CSI, the base station can predict the sample downlink CSI; or, if the sample candidate action is to feed back downlink CSI to the base station, the base station can receive the sample downlink CSI fed back by the terminal.

[0276] In one implementation, when the terminal's state information is sample state information, the terminal can obtain the sample reference downlink CSI based on the sample CSI-RS sent by the base station. For example, the terminal can perform downlink channel estimation based on the CSI-RS to obtain the sample reference downlink CSI.

[0277] S702, train the reinforcement learning model based on the training samples, and update the model parameters of the reinforcement learning model.

[0278] S703: If the model training termination condition is not met, return to the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, and a trained reinforcement learning model is generated.

[0279] It should be noted that the relevant content of steps S702-S703 can be found in the above embodiments, and will not be repeated here.

[0280] In summary, the model training method according to the embodiments of this application can obtain training samples, wherein the training samples include the sample state information of the terminal, the sample cumulative reward of each sample candidate action under the sample state information, and the sample target action, and train a reinforcement learning model based on the training samples to generate a trained reinforcement learning model.

[0281] like Figure 8 As shown, the base station may include a base station action execution module. The terminal may include a reinforcement learning modeling module, a reinforcement learning training module, a reinforcement learning decision-making module, and a terminal action execution module.

[0282] The base station execution action module is used to receive the target action sent by the terminal's decision module and execute the target action to obtain downlink CSI.

[0283] The modeling module is used to construct reinforcement learning models, obtain the terminal's own sample state information, collect sample block error rate, obtain the sample error between the sample downlink CSI obtained by executing sample candidate actions and the sample reference downlink CSI, and obtain sample cumulative reward based on sample block error rate, sample error, and sample signaling overhead for feedback downlink CSI corresponding to sample candidate actions.

[0284] The training module receives sample state information, cumulative reward for each candidate action under the sample state information, and target action of the sample from the modeling module. Based on the sample state information, cumulative reward and target action, the module trains a reinforcement learning model to generate the optimal policy, i.e., the trained reinforcement learning model.

[0285] The decision module is used to input the terminal's state information into the trained reinforcement learning model and obtain the target action output by the reinforcement learning model.

[0286] The terminal action module is used to receive the target action sent by the decision module and execute the target action to obtain downlink CSI.

[0287] The technical solutions provided in this application can be applied to various systems, especially 5G systems. For example, applicable systems include Global System for Mobile Communication (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA) General Packet Radio Service (GPRS), Long Term Evolution (LTE), LTE Frequency Division Duplex (FDD), LTE Time Division Duplex (TDD), Long Term Evolution Advanced (LTE-A), Universal Mobile Telecommunication System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX), and 5G New Radio (NR). All of these systems include terminal equipment and network equipment. The systems may also include a core network component, such as Evolved Packet System (EPS) and 5G system (5GS).

[0288] The terminal devices involved in the embodiments of this application can be devices that provide voice and / or data connectivity to users, handheld devices with wireless connectivity, or other processing devices connected to a wireless modem. The names of the terminal devices may differ in different systems; for example, in a 5G system, a terminal device can be called User Equipment (UE). Wireless terminal devices can communicate with one or more core networks (CNs) via a Radio Access Network (RAN). Wireless terminal devices can be mobile terminal devices, such as mobile phones (or "cellular" phones) and computers with mobile terminal devices, for example, portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile devices that exchange voice and / or data with the RAN. Examples include Personal Communication Service (PCS) phones, cordless phones, Session Initiated Protocol (SIP) phones, Wireless Local Loop (WLL) stations, and Personal Digital Assistants (PDAs). Wireless terminal equipment can also be referred to as a system, subscriber unit, subscriber station, mobile station, mobile station, remote station, access point, remote terminal, access terminal, user terminal, user agent, or user device, but is not limited to these terms in the embodiments of this application.

[0289] The base station involved in this application embodiment may include multiple cells providing services to terminals. Depending on the specific application, the base station may also be called an access point, or a device in the access network that communicates with the wireless terminal device through one or more sectors on the air interface, or other names. The network device can be used to exchange received air frames with Internet Protocol (IP) packets, acting as a router between the wireless terminal device and the rest of the access network, where the rest of the access network may include an Internet Protocol (IP) communication network. The network device can also coordinate the attribute management of the air interface. For example, the network equipment involved in the embodiments of this application can be a base transceiver station (BTS) in a Global System for Mobile communications (GSM) or Code Division Multiple Access (CDMA), a NodeB in a Wide-band Code Division Multiple Access (WCDMA) system, an evolved Node B (eNB or e-NodeB) in a long term evolution (LTE) system, a 5G base station (gNB) in a next generation system, a Home evolved Node B (HeNB), a relay node, a femto, a pico, etc., and is not limited in the embodiments of this application. In some network structures, the network equipment may include centralized unit (CU) nodes and distributed unit (DU) nodes, and the centralized unit and distributed unit may also be geographically separated.

[0290] Base stations and terminal devices can each use one or more antennas for multiple-input multiple-output (MIMO) transmission. MIMO transmission can be single-user MIMO (SU-MIMO) or multiple-user MIMO (MU-MIMO). Depending on the configuration and number of antenna combinations, MIMO transmission can be 2D-MIMO, 3D-MIMO, FD-MIMO, or massive-MIMO, and can also be diversity transmission, precoding transmission, or beamforming transmission, etc.

[0291] Figure 9 This is a block diagram of a base station according to an embodiment of this application.

[0292] like Figure 9 As shown, the base station 100 in this embodiment includes: a memory 110, a transceiver 120, and a processor 130.

[0293] The memory 110 is used to store computer programs; the transceiver 120 is used to send and receive data under the control of the processor 130; the processor 130 is used to read the computer program in the memory 110 and perform the following operations: obtain the terminal's status information; determine a target action for obtaining downlink CSI based on the status information, wherein the target action is to predict downlink CSI or receive downlink CSI feedback from the terminal; and execute the target action to obtain the downlink CSI.

[0294] Transceiver 120 is used to receive and send data under the control of processor 130.

[0295] Among them, Figure 9 In this context, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 130 and memory represented by memory 110 together. The bus architecture may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 120 may be multiple elements, including transmitters and receivers, providing units for communicating with various other devices over a transmission medium, including wireless channels, wired channels, optical fibers, and other transmission media.

[0296] The processor 130 is responsible for managing the bus architecture and general processing, and the memory 110 can store the data used by the processor 130 when performing operations.

[0297] The processor 130 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor can also adopt a multi-core architecture.

[0298] The processor 130 executes any of the methods described in the embodiments of this application by calling a computer program stored in the memory 110, according to the obtained executable instructions. The processor 130 and the memory 110 may also be physically separated.

[0299] In one embodiment of this application, the processor 130 is further configured to read the computer program in the memory and perform the following operations: inputting the state information into a trained reinforcement learning model, having the reinforcement learning model determine the cumulative reward for each candidate action under the state information, and determining the candidate action with the largest cumulative reward as the target action.

[0300] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0301] In one embodiment of this application, the processor 130 is further configured to read a computer program in the memory and perform the following operations: when the target action is the predicted downlink CSI, receive a probe reference signal (SRS) sent by the terminal, and predict the downlink CSI based on the SRS.

[0302] In one embodiment of this application, the processor 130 is further configured to read a computer program in the memory and perform the following operations: obtain the uplink CSI based on the SRS; and predict the downlink CSI based on the uplink CSI.

[0303] In one embodiment of this application, the processor 130 is further configured to read a computer program in the memory and perform the following operations: send a first instruction message to the terminal, wherein the first instruction message is used to instruct the terminal to send the SRS.

[0304] In one embodiment of this application, the processor 130 is further configured to read a computer program in the memory and perform the following operations: when the target action is the predicted downlink CSI, obtain the historical downlink CSI fed back by the terminal, and predict the downlink CSI based on the historical downlink CSI.

[0305] In one embodiment of this application, the processor 130 is further configured to read the computer program in the memory and perform the following operations: when the target action is to receive downlink CSI feedback from the terminal, send a Channel State Information Reference Signal (CSI-RS) to the terminal, wherein the CSI-RS is used to instruct the terminal to obtain the downlink CSI based on the CSI-RS; and receive the downlink CSI feedback from the terminal.

[0306] In one embodiment of this application, the processor 130 is further configured to read the computer program in the memory and perform the following operations: send a second indication message to the terminal, wherein the second indication message is used to instruct the terminal to provide feedback on the downlink CSI.

[0307] In one embodiment of this application, the second indication information is further used to trigger the setting of configuration information for downlink CSI feedback of the terminal, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0308] In one embodiment of this application, the processor 130 is further configured to read the computer program in the memory and perform the following operations: send the CSI-RS to the terminal according to the sending period, wherein the sending period is equal to the feedback period.

[0309] In one embodiment of this application, the processor 130 is further configured to read a computer program in the memory and perform the following operations: receive the status information sent by the terminal; and / or acquire the pre-configured status information; and / or collect the status information; and / or predict the status information.

[0310] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0311] In summary, the base station in this embodiment can determine the target action for acquiring downlink CSI based on the terminal's state information. The target action is either predicting downlink CSI or receiving downlink CSI feedback from the terminal, and then executing the target action to acquire the downlink CSI. Therefore, by selecting either predicted downlink CSI or received downlink CSI feedback from the terminal based on the terminal's state information, the signaling overhead incurred in acquiring downlink CSI is greatly reduced, and the accuracy and reliability of acquiring downlink CSI are high.

[0312] Figure 10 This is a block diagram of a terminal according to an embodiment of this application.

[0313] like Figure 10 As shown, the terminal 200 in this embodiment includes: a memory 210, a transceiver 220, and a processor 230.

[0314] The memory 210 is used to store computer programs; the transceiver 220 is used to send and receive data under the control of the processor 230; the processor 230 is used to read the computer program in the memory 210 and perform the following operations: obtain the terminal's own status information; determine a target action for obtaining downlink CSI based on the status information, wherein the target action is the base station predicting downlink CSI or feeding back downlink CSI to the base station; and execute the target action to enable the base station to obtain the downlink CSI.

[0315] Transceiver 220 is used to receive and send data under the control of processor 230.

[0316] Among them, Figure 10 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 230 and memory represented by memory 210 together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. Transceiver 220 can be multiple elements, including transmitters and receivers, providing units for communicating with various other devices over a transmission medium, including wireless channels, wired channels, optical fibers, and other transmission media.

[0317] The processor 230 is responsible for managing the bus architecture and general processing, and the memory 210 can store the data used by the processor 230 when performing operations.

[0318] The processor 230 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor can also adopt a multi-core architecture.

[0319] The processor 230 executes any of the methods described in the embodiments of this application by calling a computer program stored in the memory 210, according to the obtained executable instructions. The processor 230 and the memory 210 may also be physically separated.

[0320] In one embodiment of this application, the processor 230 is further configured to read the computer program in the memory and perform the following operations: inputting the state information into a trained reinforcement learning model, having the reinforcement learning model determine the cumulative reward for each candidate action under the state information, and determining the candidate action with the largest cumulative reward as the target action.

[0321] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0322] In one embodiment of this application, the processor 230 is further configured to read a computer program in the memory and perform the following operations: when the target action is the base station predicting downlink CSI, sending first indication information to the base station, wherein the first indication information is used to instruct the base station to predict the downlink CSI.

[0323] In one embodiment of this application, the processor 230 is further configured to read a computer program in the memory and perform the following operations: when the target action is the base station predicting downlink CSI, sending a probe reference signal (SRS) to the base station, wherein the SRS is used to instruct the base station to predict the downlink CSI based on the SRS.

[0324] In one embodiment of this application, the processor 230 is further configured to read the computer program in the memory and perform the following operations: when the target action is to feed back downlink CSI to the base station, receive the channel state information reference signal CSI-RS sent by the base station; obtain the downlink CSI according to the CSI-RS; and feed back the downlink CSI to the base station.

[0325] In one embodiment of this application, the processor 230 is further configured to read a computer program in the memory and perform the following operations: receiving second indication information sent by the base station, wherein the second indication information is used to trigger the setting of configuration information for downlink CSI feedback of the terminal itself, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0326] In one embodiment of this application, the processor 230 is further configured to read the computer program in the memory and perform the following operations: feed back the downlink CSI to the base station according to the feedback period.

[0327] In one embodiment of this application, the processor 230 is further configured to read a computer program in the memory and perform the following operations: send third indication information to the base station, wherein the third indication information is used to instruct the base station to send the CSI-RS.

[0328] In one embodiment of this application, the processor 230 is further configured to read the computer program in the memory and perform the following operations: obtain the pre-configured status information; and / or collect the status information.

[0329] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0330] In summary, the terminal in this embodiment can determine the target action for obtaining downlink CSI based on its own state information. The target action is either the base station predicting downlink CSI or feeding back downlink CSI to the base station, and then executing the target action to obtain the downlink CSI. Therefore, by selecting either base station prediction or feeding back downlink CSI to the base station based on the terminal's own state information, the signaling overhead incurred in obtaining downlink CSI is greatly reduced, and the accuracy and reliability of obtaining downlink CSI are high.

[0331] Figure 11 This is a block diagram of a base station according to another embodiment of this application.

[0332] like Figure 11 As shown, the base station 300 in this embodiment includes: a memory 310, a transceiver 320, and a processor 330.

[0333] The system includes a memory 310 for storing computer programs; a transceiver 320 for transmitting and receiving data under the control of a processor 330; and a processor 330 for reading the computer program from the memory 310 and performing the following operations: acquiring training samples, wherein the training samples include terminal sample state information, cumulative reward for each candidate action under the sample state information, and target action, wherein the target action is the candidate action with the largest cumulative reward, and the target action is either predicting downlink CSI or receiving downlink CSI feedback from the terminal; training a reinforcement learning model based on the training samples and updating the model parameters of the reinforcement learning model; and, if the model training termination condition is not met, returning to use the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thereby generating a trained reinforcement learning model.

[0334] Transceiver 320 is used to receive and send data under the control of processor 330.

[0335] Among them, Figure 11 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 330 and memory represented by memory 310 together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 320 can be multiple elements, including transmitters and receivers, providing units for communicating with various other devices over a transmission medium, including wireless channels, wired channels, optical fibers, and other transmission media.

[0336] The processor 330 is responsible for managing the bus architecture and general processing, while the memory 310 can store the data used by the processor 330 when performing operations.

[0337] The processor 330 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor can also adopt a multi-core architecture.

[0338] The processor 330 executes any of the methods described in the embodiments of this application by calling a computer program stored in the memory 310, according to the obtained executable instructions. The processor 330 and the memory 310 may also be physically separated.

[0339] In one embodiment of this application, the processor 330 is further configured to read the computer program in the memory and perform the following operations: obtain sample target parameters according to the sample state information, and obtain the sample cumulative reward according to the sample target parameters and the sample signaling overhead for feedback downlink CSI corresponding to the sample candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI obtained by executing the sample candidate action and the sample reference downlink CSI.

[0340] In one embodiment of this application, the processor 330 is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, receive the Hybrid Automatic Repeat Request (HARQ) feedback information sent by the terminal, and obtain the sample block error rate based on the HARQ feedback information.

[0341] In one embodiment of this application, the processor 330 is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, execute the sample candidate action to obtain the sample downlink CSI.

[0342] In one embodiment of this application, the processor 330 is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, receiving the sample reference downlink CSI sent by the terminal, wherein the sample reference downlink CSI is obtained based on the sample channel state information reference signal CSI-RS sent by the base station.

[0343] In one embodiment of this application, the processor 330 is further configured to read a computer program in the memory and perform the following operations: receive the sample status information sent by the terminal; and / or acquire the pre-configured sample status information; and / or collect the sample status information; and / or predict the sample status information.

[0344] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0345] In summary, the base station in this application embodiment can obtain training samples, wherein the training samples include the terminal's sample state information, the sample cumulative reward for each sample candidate action under the sample state information, and the sample target action, and train a reinforcement learning model based on the training samples to generate a trained reinforcement learning model.

[0346] Figure 12 This is a block diagram of a terminal according to another embodiment of this application.

[0347] like Figure 12 As shown, the terminal 400 in this embodiment includes: a memory 410, a transceiver 420, and a processor 430.

[0348] The system includes a memory 410 for storing computer programs; a transceiver 420 for transmitting and receiving data under the control of a processor 430; and a processor 430 for reading the computer program from the memory 410 and performing the following operations: acquiring training samples, wherein the training samples include terminal sample state information, cumulative reward for each candidate action under the sample state information, and target action, wherein the target action is the candidate action with the largest cumulative reward, and the target action is either a base station predicting downlink CSI or feeding back downlink CSI to the base station; training a reinforcement learning model based on the training samples and updating the model parameters of the reinforcement learning model; and if the model training termination condition is not met, returning to use the next training sample to continue training the updated reinforcement learning model until the model training termination condition is met, thereby generating a trained reinforcement learning model.

[0349] Transceiver 420 is used to receive and send data under the control of processor 430.

[0350] Among them, Figure 12 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits of one or more processors represented by processor 430 and memory represented by memory 410 together. The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. Transceiver 420 can be multiple elements, including transmitters and receivers, providing units for communicating with various other devices over a transmission medium, including wireless channels, wired channels, optical fibers, and other transmission media.

[0351] The processor 430 is responsible for managing the bus architecture and general processing, and the memory 410 can store the data used by the processor 430 when performing operations.

[0352] The processor 430 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor can also adopt a multi-core architecture.

[0353] The processor 430 executes any of the methods described in the embodiments of this application by calling a computer program stored in the memory 410, according to the obtained executable instructions. The processor 430 and the memory 410 may also be physically separated.

[0354] In one embodiment of this application, the processor 430 is further configured to read the computer program in the memory and perform the following operations: obtain sample target parameters according to the sample state information, and obtain the sample cumulative reward according to the sample target parameters and the sample signaling overhead for feedback downlink CSI corresponding to the sample candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI obtained by executing the sample candidate action and the sample reference downlink CSI.

[0355] In one embodiment of this application, the processor 430 is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, collect the sample error rate.

[0356] In one embodiment of this application, the processor 430 is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, execute the sample candidate action to obtain the sample downlink CSI.

[0357] In one embodiment of this application, the processor 430 is further configured to read the computer program in the memory and perform the following operations: when the status information of the terminal is the sample status information, obtain the sample reference downlink CSI according to the sample channel state information reference signal CSI-RS sent by the base station.

[0358] In one embodiment of this application, the processor 430 is further configured to read a computer program in the memory and perform the following operations: obtain the pre-configured sample status information; and / or collect the sample status information.

[0359] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0360] In summary, the terminal in this application embodiment can obtain training samples, wherein the training samples include the terminal's sample state information, the sample cumulative reward for each sample candidate action under the sample state information, and the sample target action, and train a reinforcement learning model based on the training samples to generate a trained reinforcement learning model.

[0361] Figure 13 This is a block diagram of an apparatus for obtaining downlink channel state information according to an embodiment of this application.

[0362] like Figure 13 As shown, the apparatus 500 for obtaining downlink channel state information according to an embodiment of this application includes: an acquisition module 510, a determination module 520, and an execution module 530.

[0363] The acquisition module 510 is used to acquire the terminal's status information;

[0364] The determining module 520 is used to determine the target action for obtaining downlink CSI based on the status information, wherein the target action is to predict downlink CSI or to receive downlink CSI fed back by the terminal;

[0365] The execution module 530 is used to perform the target action to obtain the downlink CSI.

[0366] In one embodiment of this application, the determining module 520 is further configured to: input the state information into a trained reinforcement learning model, have the reinforcement learning model determine the cumulative reward of each candidate action under the state information, and determine the candidate action with the largest cumulative reward as the target action.

[0367] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0368] In one embodiment of this application, when the target action is the predicted downlink CSI, the execution module 530 is further configured to: receive the probe reference signal (SRS) sent by the terminal, and predict the downlink CSI based on the SRS.

[0369] In one embodiment of this application, the execution module 530 is further configured to: obtain the uplink CSI based on the SRS; and predict the downlink CSI based on the uplink CSI.

[0370] In one embodiment of this application, the apparatus 500 for obtaining downlink channel state information (CSI) further includes a sending module, which is configured to send first indication information to the terminal, wherein the first indication information is used to instruct the terminal to send the SRS.

[0371] In one embodiment of this application, when the target action is the predicted downlink CSI, the execution module 530 is further configured to: obtain the historical downlink CSI fed back by the terminal, and predict the downlink CSI based on the historical downlink CSI.

[0372] In one embodiment of this application, when the target action is to receive downlink CSI feedback from the terminal, the execution module 530 is further configured to: send a channel state information reference signal CSI-RS to the terminal, wherein the CSI-RS is used to instruct the terminal to obtain the downlink CSI based on the CSI-RS; and receive the downlink CSI feedback from the terminal.

[0373] In one embodiment of this application, the apparatus 500 for obtaining downlink channel state information (CSI) further includes a sending module, which is further configured to send second indication information to the terminal, wherein the second indication information is used to instruct the terminal to provide feedback on the downlink CSI.

[0374] In one embodiment of this application, the second indication information is further used to trigger the setting of configuration information for downlink CSI feedback of the terminal, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0375] In one embodiment of this application, the execution module 530 is further configured to: send the CSI-RS to the terminal according to the sending period, wherein the sending period is equal to the feedback period.

[0376] In one embodiment of this application, the acquisition module 510 is further configured to: receive the status information sent by the terminal; and / or acquire the pre-configured status information; and / or collect the status information; and / or predict the status information.

[0377] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0378] In summary, the apparatus for obtaining downlink channel state information (CSI) according to the embodiments of this application can determine the target action for obtaining downlink CSI based on the terminal's state information. The target action is either predicting downlink CSI or receiving downlink CSI feedback from the terminal, and then executing the target action to obtain the downlink CSI. Therefore, by selecting either predicting downlink CSI or receiving downlink CSI feedback from the terminal based on the terminal's state information, the signaling overhead consumed in obtaining downlink CSI is greatly reduced, and the accuracy and reliability of obtaining downlink CSI are high.

[0379] Figure 14 This is a block diagram of an apparatus for obtaining downlink channel state information according to another embodiment of this application.

[0380] like Figure 14 As shown, the apparatus 600 for obtaining downlink channel state information according to an embodiment of this application includes: an acquisition module 610, a determination module 620, and an execution module 630.

[0381] The acquisition module 610 is used to acquire the terminal's own status information;

[0382] The determining module 620 is used to determine the target action for obtaining downlink CSI based on the status information, wherein the target action is the base station predicting downlink CSI or feeding back downlink CSI to the base station;

[0383] The execution module 630 is used to perform the target action so that the base station can acquire the downlink CSI.

[0384] In one embodiment of this application, the determining module 620 is further configured to: input the state information into a trained reinforcement learning model, have the reinforcement learning model determine the cumulative reward of each candidate action under the state information, and determine the candidate action with the largest cumulative reward as the target action.

[0385] In one embodiment of this application, the reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; wherein, the target parameters include the block error rate of the physical downlink shared channel for scheduling and beamforming based on downlink CSI, and / or the error between the downlink CSI obtained by executing the candidate action and the reference downlink CSI.

[0386] In one embodiment of this application, when the target action is for the base station to predict the downlink CSI, the execution module 630 is further configured to: send first indication information to the base station, wherein the first indication information is used to instruct the base station to predict the downlink CSI.

[0387] In one embodiment of this application, when the target action is for the base station to predict the downlink CSI, the execution module 630 is further configured to: send a probe reference signal (SRS) to the base station, wherein the SRS is used to instruct the base station to predict the downlink CSI based on the SRS.

[0388] In one embodiment of this application, when the target action is to feed back downlink CSI to the base station, the execution module 630 is further configured to: receive a channel state information reference signal (CSI-RS) sent by the base station; obtain the downlink CSI based on the CSI-RS; and feed back the downlink CSI to the base station.

[0389] In one embodiment of this application, the apparatus 600 for acquiring downlink channel state information (CSI) further includes a receiving module, which is configured to receive second indication information sent by the base station, wherein the second indication information is used to trigger the setting of configuration information for downlink CSI feedback of the terminal itself, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

[0390] In one embodiment of this application, the execution module 630 is further configured to: feed back the downlink CSI to the base station according to the feedback period.

[0391] In one embodiment of this application, the apparatus 600 for obtaining downlink channel state information (CSI) further includes a transmitting module, which is configured to transmit third indication information to the base station, wherein the third indication information is used to instruct the base station to transmit the CSI-RS.

[0392] In one embodiment of this application, the acquisition module 610 is further configured to: acquire the pre-configured status information; and / or collect the status information.

[0393] In one embodiment of this application, the status information includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency.

[0394] In summary, the apparatus for obtaining downlink channel state information (CSI) according to the embodiments of this application can determine the target action for obtaining downlink CSI based on the terminal's own state information. The target action is either the base station predicting downlink CSI or feeding back downlink CSI to the base station, and then executing the target action to obtain the downlink CSI. Therefore, by selecting either base station prediction of downlink CSI or feeding back downlink CSI to the base station based on the terminal's own state information, the signaling overhead consumed in obtaining downlink CSI is greatly reduced, and the accuracy and reliability of obtaining downlink CSI are high.

[0395] Figure 15 This is a block diagram of a model training apparatus according to an embodiment of this application.

[0396] like Figure 15 As shown, the model training device 700 of this application embodiment includes: an acquisition module 710 and a training module 720.

[0397] The acquisition module 710 is used to acquire training samples, wherein the training samples include the terminal's sample state information, the cumulative reward of each candidate action under the sample state information, and the target action of the sample. The target action of the sample is the candidate action of the sample with the largest cumulative reward, and the target action of the sample is to predict the downlink CSI or receive the downlink CSI fed back by the terminal.

[0398] The training module 720 is used to train a reinforcement learning model based on the training samples and update the model parameters of the reinforcement learning model.

[0399] The training module 720 is also used to return to the next training sample to continue training the updated reinforcement learning model if the model training termination condition is not met, until the model training termination condition is met and a trained reinforcement learning model is generated.

[0400] In one embodiment of this application, the acquisition module 710 is further configured to: acquire sample target parameters based on the sample state information, and acquire the sample cumulative reward based on the sample target parameters and the sample signaling overhead for feedback of downlink CSI corresponding to the sample candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI acquired by executing the sample candidate action and the sample reference downlink CSI.

[0401] In one embodiment of this application, the acquisition module 710 is further configured to: receive Hybrid Automatic Repeat Request (HARQ) feedback information sent by the terminal when the terminal's status information is the sample status information, and obtain the sample block error rate based on the HARQ feedback information.

[0402] In one embodiment of this application, the acquisition module 710 is further configured to: when the status information of the terminal is the sample status information, perform the sample candidate action to obtain the sample downlink CSI.

[0403] In one embodiment of this application, the acquisition module 710 is further configured to: receive the sample reference downlink CSI sent by the terminal when the terminal's status information is the sample status information, wherein the sample reference downlink CSI is obtained based on the sample channel state information reference signal CSI-RS sent by the base station.

[0404] In one embodiment of this application, the acquisition module 710 is further configured to: receive the sample status information sent by the terminal; and / or acquire the pre-configured sample status information; and / or collect the sample status information; and / or predict the sample status information.

[0405] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0406] In summary, the model training apparatus of this application embodiment can acquire training samples, wherein the training samples include sample state information of the terminal, sample cumulative reward for each sample candidate action under the sample state information, and sample target action, and train a reinforcement learning model based on the training samples to generate a trained reinforcement learning model.

[0407] Figure 16 This is a block diagram of a model training apparatus according to another embodiment of this application.

[0408] like Figure 16 As shown, the model training device 800 of this application embodiment includes: an acquisition module 810 and a training module 820.

[0409] The acquisition module 810 is used to acquire training samples, wherein the training samples include the terminal's sample state information, the cumulative reward of each candidate action under the sample state information, and the target action of the sample. The target action of the sample is the candidate action of the sample with the largest cumulative reward, and the target action of the sample is the base station predicting downlink CSI or feeding back downlink CSI to the base station.

[0410] The training module 820 is used to train a reinforcement learning model based on the training samples and update the model parameters of the reinforcement learning model.

[0411] The training module 820 is also used to return to the next training sample to continue training the updated reinforcement learning model if the model training termination condition is not met, until the model training termination condition is met and a trained reinforcement learning model is generated.

[0412] In one embodiment of this application, the acquisition module 810 is further configured to: acquire sample target parameters based on the sample state information, and acquire the sample cumulative reward based on the sample target parameters and the sample signaling overhead for feedback of downlink CSI corresponding to the sample candidate action; wherein, the sample target parameters include the sample block error rate of the physical downlink shared channel for scheduling and / or beamforming based on downlink CSI, and / or the sample error between the sample downlink CSI acquired by executing the sample candidate action and the sample reference downlink CSI.

[0413] In one embodiment of this application, the acquisition module 810 is further configured to: collect the sample error rate when the status information of the terminal is the sample status information.

[0414] In one embodiment of this application, the acquisition module 810 is further configured to: when the status information of the terminal is the sample status information, perform the sample candidate action to obtain the sample downlink CSI.

[0415] In one embodiment of this application, the acquisition module 810 is further configured to: when the status information of the terminal is the sample status information, acquire the sample reference downlink CSI based on the sample channel state information reference signal CSI-RS sent by the base station.

[0416] In one embodiment of this application, the acquisition module 810 is further configured to: acquire the pre-configured sample status information; and / or collect the sample status information.

[0417] In one embodiment of this application, the sample status information includes at least one of sample velocity, sample received signal-to-noise ratio (SNR), sample carrier frequency offset, sample uplink carrier frequency, and sample downlink carrier frequency.

[0418] In summary, the model training apparatus of this application embodiment can acquire training samples, wherein the training samples include sample state information of the terminal, sample cumulative reward for each sample candidate action under the sample state information, and sample target action, and train a reinforcement learning model based on the training samples to generate a trained reinforcement learning model.

[0419] It should be noted that the division of units in the embodiments of this application is illustrative and only represents one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.

[0420] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0421] According to embodiments of this application, this application also provides a processor-readable storage medium.

[0422] The processor-readable storage medium stores a computer program that is used to cause the processor to execute this application. Figures 1-3 The method for obtaining downlink channel state information (CSI) as described in the embodiment.

[0423] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0424] According to embodiments of this application, another processor-readable storage medium is also proposed.

[0425] The processor-readable storage medium stores a computer program that is used to cause the processor to execute this application. Figure 4 The method for obtaining downlink channel state information (CSI) as described in the embodiment.

[0426] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0427] According to embodiments of this application, another processor-readable storage medium is also proposed.

[0428] The processor-readable storage medium stores a computer program that is used to cause the processor to execute this application. Figure 5 The model training method described in the embodiments.

[0429] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0430] According to embodiments of this application, another processor-readable storage medium is also proposed.

[0431] The processor-readable storage medium stores a computer program that is used to cause the processor to execute this application. Figure 7 The model training method described in the embodiments.

[0432] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0433] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0434] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-executable instructions. These computer-executable instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0435] These processor-executable instructions may also be stored in a processor-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the processor-readable memory produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0436] These processors can execute instructions that can also be loaded onto a computer or other programmable data processing device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0437] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for obtaining downlink channel state information (CSI), characterized in that, The executing entity is a base station, and the method includes: Acquire the terminal's status information, which includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency. Based on the status information, a target action for obtaining downlink CSI is determined, wherein the target action is to predict downlink CSI or to receive downlink CSI feedback from the terminal; Perform the target action to obtain the downlink CSI; The step of determining the target action for obtaining downlink CSI based on the status information includes: The state information is input into the trained reinforcement learning model, which determines the cumulative reward for each candidate action under the state information. The candidate action with the largest cumulative reward is then determined as the target action.

2. The method according to claim 1, characterized in that, The reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; The target parameters include the block error rate of the physical downlink shared channel, which is scheduled and / or beamformed based on the downlink CSI, and / or the error between the downlink CSI obtained by performing the candidate action and the reference downlink CSI.

3. The method according to claim 1, characterized in that, When the target action is the predicted downlink CSI, executing the target action includes: The system receives the probe reference signal (SRS) sent by the terminal and predicts the downlink CSI based on the SRS.

4. The method according to claim 3, characterized in that, The step of predicting the downlink CSI based on the SRS includes: Based on the SRS, obtain the uplink CSI; Based on the uplink CSI, predict the downlink CSI.

5. The method according to claim 3, characterized in that, Before receiving the detection reference signal (SRS) sent by the terminal, the method further includes: Send a first indication message to the terminal, wherein the first indication message is used to instruct the terminal to send the SRS.

6. The method according to claim 1, characterized in that, When the target action is the predicted downlink CSI, executing the target action includes: Obtain the historical downlink CSI fed back by the terminal, and predict the downlink CSI based on the historical downlink CSI.

7. The method according to claim 1, characterized in that, When the target action is receiving downlink CSI feedback from the terminal, executing the target action includes: The terminal is sent a Channel State Information Reference Signal (CSI-RS), wherein the CSI-RS is used to instruct the terminal to obtain the downlink CSI based on the CSI-RS; Receive the downlink CSI fed back by the terminal.

8. The method according to claim 7, characterized in that, Before sending the Channel State Information Reference Signal (CSI-RS) to the terminal, the method further includes: Send a second indication message to the terminal, wherein the second indication message is used to instruct the terminal to provide feedback on the downlink CSI.

9. The method according to claim 8, characterized in that, The second indication information is also used to trigger the setting of configuration information for downlink CSI feedback of the terminal, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

10. The method according to claim 9, characterized in that, Sending the Channel State Information Reference Signal (CSI-RS) to the terminal includes: The CSI-RS is sent to the terminal according to the sending period, wherein the sending period is equal to the feedback period.

11. The method according to claim 1, characterized in that, The acquisition of terminal status information includes: Receive the status information sent by the terminal; and / or Obtain the pre-configured status information; and / or, Collect the status information; and / or, Predict the state information.

12. A method for obtaining downlink channel state information (CSI), characterized in that, The execution subject is a terminal, and the method includes: The terminal acquires its own status information, which includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency. Based on the status information, a target action for obtaining downlink CSI is determined, wherein the target action is for the base station to predict downlink CSI or to feed back downlink CSI to the base station; Perform the target action to enable the base station to acquire the downlink CSI; The step of determining the target action for obtaining downlink CSI based on the status information includes: The state information is input into the trained reinforcement learning model, which determines the cumulative reward for each candidate action under the state information. The candidate action with the largest cumulative reward is then determined as the target action.

13. The method according to claim 12, characterized in that, The reinforcement learning model is used to obtain target parameters based on the state information, and to obtain the cumulative reward based on the target parameters and the signaling overhead for feedback of downlink CSI corresponding to the candidate action; The target parameters include the block error rate of the physical downlink shared channel, which is scheduled and beamformed based on the downlink CSI, and / or the error between the downlink CSI obtained by performing the candidate action and the reference downlink CSI.

14. The method according to claim 12, characterized in that, When the target action is the base station's predicted downlink CSI, executing the target action includes: Send a first indication message to the base station, wherein the first indication message is used to instruct the base station to predict the downlink CSI.

15. The method according to claim 12, characterized in that, When the target action is the base station's predicted downlink CSI, executing the target action includes: A probe reference signal (SRS) is sent to the base station, wherein the SRS is used to instruct the base station to predict the downlink CSI based on the SRS.

16. The method according to claim 12, characterized in that, When the target action is to feed back downlink CSI to the base station, executing the target action includes: Receive the Channel State Information Reference Signal (CSI-RS) sent by the base station; The downlink CSI is obtained according to the CSI-RS; The downlink CSI is fed back to the base station.

17. The method according to claim 16, characterized in that, Before receiving the Channel State Information Reference Signal (CSI-RS) sent by the base station, the method further includes: The terminal receives a second indication message sent by the base station, wherein the second indication message is used to trigger the setting of configuration information for downlink CSI feedback of the terminal itself, wherein the configuration information includes at least one of the following: downlink CSI feedback amount, downlink CSI feedback period, downlink CSI feedback time domain resources, and downlink CSI feedback frequency domain resources.

18. The method according to claim 17, characterized in that, The step of feeding back the downlink CSI to the base station includes: The downlink CSI is fed back to the base station according to the feedback period.

19. The method according to claim 18, characterized in that, Before receiving the Channel State Information Reference Signal (CSI-RS) sent by the base station, the method further includes: Send a third indication message to the base station, wherein the third indication message is used to instruct the base station to send the CSI-RS.

20. The method according to claim 12, characterized in that, The acquisition of the terminal's own status information includes: Obtain the pre-configured status information; and / or, Collect the aforementioned status information.

21. A base station, characterized in that, Includes memory, transceiver, and processor: A memory for storing computer programs; a transceiver for sending and receiving data under the control of the processor; and a processor for reading the computer programs from the memory and performing the following operations: Acquire the terminal's status information, which includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency. Based on the status information, a target action for obtaining downlink CSI is determined, wherein the target action is to predict downlink CSI or to receive downlink CSI feedback from the terminal; Perform the target action to obtain the downlink CSI; The step of determining the target action for obtaining downlink CSI based on the status information includes: The state information is input into the trained reinforcement learning model, which determines the cumulative reward for each candidate action under the state information. The candidate action with the largest cumulative reward is then determined as the target action.

22. A terminal, characterized in that, Includes memory, transceiver, and processor: A memory for storing computer programs; a transceiver for sending and receiving data under the control of the processor; and a processor for reading the computer programs from the memory and performing the following operations: The terminal acquires its own status information, which includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency. Based on the status information, a target action for obtaining downlink CSI is determined, wherein the target action is for the base station to predict downlink CSI or to feed back downlink CSI to the base station; Perform the target action to enable the base station to acquire the downlink CSI; The step of determining the target action for obtaining downlink CSI based on the status information includes: The state information is input into the trained reinforcement learning model, which determines the cumulative reward for each candidate action under the state information. The candidate action with the largest cumulative reward is then determined as the target action.

23. An apparatus for acquiring downlink channel state information (CSI), characterized in that, include: The acquisition module is used to acquire the terminal's status information, which includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency. The determining module is configured to determine, based on the status information, a target action for obtaining downlink CSI, wherein the target action is either predicting downlink CSI or receiving downlink CSI feedback from the terminal; An execution module is used to perform the target action to obtain the downlink CSI; The determining module is specifically used for: The state information is input into the trained reinforcement learning model, which determines the cumulative reward for each candidate action under the state information. The candidate action with the largest cumulative reward is then determined as the target action.

24. An apparatus for acquiring downlink channel state information (CSI), characterized in that, include: The acquisition module is used to acquire the terminal's own status information, which includes at least one of speed, received signal-to-noise ratio (SNR), carrier frequency offset, uplink carrier frequency, and downlink carrier frequency. The determining module is configured to determine, based on the status information, a target action for obtaining downlink CSI, wherein the target action is either the base station predicting downlink CSI or feeding back downlink CSI to the base station; An execution module is configured to perform the target action to enable the base station to acquire the downlink CSI; The determining module is specifically used for: The state information is input into the trained reinforcement learning model, which determines the cumulative reward for each candidate action under the state information. The candidate action with the largest cumulative reward is then determined as the target action.

25. A processor-readable storage medium, characterized in that, The processor-readable storage medium stores a computer program for causing the processor to perform the method for acquiring downlink channel state information (CSI) as described in any one of claims 1-11.

26. A processor-readable storage medium, characterized in that, The processor-readable storage medium stores a computer program for causing the processor to perform the method for acquiring downlink channel state information (CSI) as described in any one of claims 12-20.

Citation Information

Patent Citations

  • Distributed antenna system, distributed antenna switching method, base station apparatus and antenna switching device

    CN102480316A