An irs-aided adaptive channel estimation and reflection optimization method

CN122513221APending Publication Date: 2026-08-04ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV OF TECH
Filing Date
2026-04-28
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

在实际应用中,不同的CE策略会显著影响系统的整体信道利用效率,尤其在信道相关性不稳定、环境动态变化较快的复杂场景下,简单的二元决策模式已难以满足系统的实时性与精度需求

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513221A_ABST
    Figure CN122513221A_ABST
Patent Text Reader

Abstract

This invention discloses an IRS-assisted adaptive channel estimation and reflection optimization method, comprising establishing an IRS-assisted downlink communication system. This IRS-assisted adaptive channel estimation and reflection optimization method constructs an objective function that maximizes the utility of the downlink communication system, and transforms the objective function into a two-layer partially observable Markov decision process. This two-layer partially observable Markov decision process includes a higher-layer decision to select the channel estimation mode for the current time slot and a lower-layer decision to select the phase reflection vector of the IRS for the current time slot, thereby improving the adaptability and stability of the downlink communication system under non-stationary channels. Through joint optimization using the discounted Thompson sampling algorithm and the deep deterministic policy gradient algorithm, the optimal channel estimation mode and the optimal IRS phase reflection vector are obtained, achieving channel estimation and reflection optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology, specifically relating to an IRS-assisted adaptive channel estimation and reflection optimization method. Background Technology

[0002] In recent years, Intelligent Reflecting Surfaces (IRS) have emerged as a promising technology and are widely recognized as an important means to improve the energy efficiency and spectral efficiency of next-generation sixth-generation (6G) wireless communication networks. IRS consists of a large number of low-cost passive reflective elements, which achieve intelligent control over the wireless propagation environment by dynamically adjusting the phase shift of the incident radio frequency signal. By precisely adjusting the phase of each reflective element, IRS can intelligently reconstruct channel propagation characteristics, thereby constructing a favorable communication link and significantly improving the overall system performance. In existing IRS-assisted communication systems, research on IRS phase reflection optimization mainly employs traditional optimization methods, such as convex optimization, to achieve optimal system performance. These methods typically rely on precise channel state information and are difficult to address real-time optimization problems in dynamic channel environments. With the rapid development of artificial intelligence technology, some research has begun to explore adaptive optimization strategies based on Deep Reinforcement Learning (DRL) to achieve intelligent control of IRS phase reflection under imperfect channel conditions, thereby effectively improving the long-term utility of the system.

[0003] Most existing research on IRS generally assumes that accurate channel state information (CSI) can be obtained through channel estimation (CE) before each decision when performing phase optimization. However, these methods often overlook the time and energy costs of CE. In practical IRS-assisted communication scenarios, such as terrestrial vehicle-to-everything (V2X) or airborne communication networks, these channels exhibit time correlation and memory, typically displaying rapid time-varying characteristics, making CSI feedback overhead non-negligible.

[0004] There is an inherent trade-off between channel estimation performance and system throughput. Existing research largely considers only two simple binary channel estimation decision modes: "estimate" or "don't estimate." In practical applications, different channel estimation (CE) strategies significantly impact the overall channel utilization efficiency, especially in complex scenarios with unstable channel correlation and rapidly changing environments. In such cases, the simple binary decision mode is insufficient to meet the system's real-time and accuracy requirements. Blindly performing frequent channel estimations in this situation leads to excessive energy and time overhead; conversely, reusing expired channel estimation (CSI) for extended periods results in accumulated estimation errors that degrade IRS beamforming performance, thereby reducing system throughput and stability. Therefore, existing technologies struggle to achieve an effective balance between channel dynamics and system energy efficiency. Summary of the Invention

[0005] The purpose of this invention is to address the problems raised in the background art by proposing an IRS-assisted adaptive channel estimation and reflection optimization method.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] This invention proposes an IRS-assisted adaptive channel estimation and reflection optimization method, comprising:

[0008] An IRS-assisted downlink communication system is established, and the downlink communication system includes a multi-antenna base station, a single-antenna user equipment, and an IRS. The base station first performs channel gain estimation and then transmits the information to the user equipment through the IRS.

[0009] Based on the transmission rate and energy consumption of the base station in the downlink communication system, an objective function for maximizing the utility of the downlink communication system is constructed.

[0010] The objective function is transformed into a two-layer partially observable Markov decision process, which includes a higher-layer decision to select the channel estimation mode for the current time slot and a lower-layer decision to select the phase reflection vector of the IRS for the current time slot.

[0011] In high-level decision-making, the discounted Thompson sampling algorithm is used to optimize the channel estimation mode for the current time slot. In low-level decision-making, the deep deterministic strategy gradient algorithm is used to optimize the phase reflection vector of the IRS for the current time slot. The optimal channel estimation mode and the phase reflection vector of the IRS for each time slot within a preset time period are applied to the downlink communication system to realize channel estimation and reflection optimization.

[0012] Preferably, the preset time period is Each time slot The number of antennas in the base station is One, IRS is represented as , The number of reflective elements in the IRS. The phase reflection vector of the time slot IRS is expressed as: , for The first time slot IRS The reflection amplitude coefficient of each reflecting element. for The first time slot IRS The phase reflection coefficient of each reflecting element, and the range of values ​​is 1. , For transpose;

[0013] The channel gain from the time-slot base station to the IRS is expressed as: , The channel gain from the time slot IRS to the user equipment is expressed as: , and They represent dimensions as follows: and For a complex matrix, then The channel gain of the time-slotted cascaded channel is expressed as: , This is the conjugate transpose.

[0014] Preferably, the channel estimation mode is represented as follows: ;

[0015] when When the time is used, it indicates that the mode is continued, and the channel gain of the concatenated channel estimated in the previous time slot is used as the channel gain of the concatenated channel estimated in the current time slot.

[0016] when The time indicates the test mode, which uses the channel gain of the cascaded channel estimated in history and the time series prediction model to estimate and predict the channel gain of the cascaded channel in the current time slot.

[0017] when When, it indicates the pilot mode, and the channel gain of the cascaded channel estimated in the current time slot is the channel gain of the cascaded channel in the current time slot.

[0018] Preferably, the objective function for maximizing the utility of the downlink communication system based on the transmission rate and energy consumption of the base station in the downlink communication system includes:

[0019] The objective function is expressed as follows:

[0020] ;

[0021] st C1: ;

[0022] C2: ;

[0023] in,

[0024] ;

[0025] in, for Time slots in channel estimation mode Phase reflection vector of IRS The function of downlink communication system utility. for Time-slot base stations in channel estimation mode Phase reflection vector of IRS The transmission rate at that time for Time slots in channel estimation mode Energy consumption of downlink communication systems This is a penalty factor.

[0026] Preferably, the step of converting the objective function into a two-level partially observable Markov decision process includes:

[0027] Define the role of high-level decision-making The observation of the time slot is ,in, This serves as the observation space corresponding to high-level decision-making, and Including the previous time slot Rewards, last time slot High-level decision-making actions and the previous time slot The channel gain is estimated, where the reward for each time slot is the utility of the downlink communication system in the corresponding time slot;

[0028] Define the role of high-level decision-making The action of the time slot is ,in, This provides the action space corresponding to high-level decision-making, and That is Channel estimation mode of time slot ;

[0029] Once the actions of the high-level decision-making process are completed, the observations from the low-level decision-making process in the current time slot are obtained, and the low-level decision-making process... The observation of the time slot is ,in, This is the observation space corresponding to low-level decision-making, and Include Channel gain estimation for time slots Actions in high-level decision-making within a time slot and the previous time slot Actions in low-level decision-making;

[0030] Define low-level decision-making The action of the time slot is ,in, This represents the action space corresponding to lower-level decisions, and That is Phase reflection vector of the IRS in the time slot ;

[0031] After the action in the lower-level decision-making is completed, the reward for the current time slot is obtained, and the reward for the current time slot is used as the input for the observation in the higher-level decision-making in the next time slot;

[0032] Define the joint strategy of high-level decision-making and low-level decision-making as ,in Strategies for high-level decision-making. Strategies for lower-level decision-making;

[0033] The belief state value function is calculated using a joint strategy, and the belief state value function is maximized. The strategy corresponding to maximizing the belief state value function is taken as the optimal strategy. The problem of solving the optimal strategy is transformed into solving the optimal channel estimation mode in high-level decision-making and solving the optimal phase reflection vector of the IRS in low-level decision-making.

[0034] Preferably, the step of using the discounted Thompson sampling algorithm in high-level decision-making to obtain the optimal channel estimation mode for the current time slot includes:

[0035] Channel estimation mode as an arm in the discounted Thompson sampling algorithm ;

[0036] At that time, initialize the discount factor respectively. Next Average discount experience for each arm In discount factor Next Discounts and cumulative rewards for each arm and in discount factor Next Valid sample size of each arm and initialization of the first The standard deviation of each arm is used. ;

[0037] Current time slot Next, samples are taken from each arm to obtain the sampled value corresponding to each arm. ,and Follows Gaussian distribution ,in, It follows a Gaussian distribution. For discount factor Down Time slot The average of the discount experience for each arm. for The next time slot Variance of each arm;

[0038] Select the largest sample value from the three sample values, and the arm corresponding to the largest sample value is the [arm value]. Slot-optimal channel estimation mode ;

[0039] For the current time slot The following rewards After normalization, we get The normalized reward was obtained by performing a logarithmic transformation. ;

[0040] Utilize the current time slot The result after the lower logarithmic transformation Discount factors respectively The average discount experience value, cumulative discount reward, and effective sample size for each arm are updated, along with the adoption standard deviation for each arm. These updated values ​​are then used as the values ​​for the next time slot. The update formulas for each value are as follows:

[0041] ;

[0042] ;

[0043] ;

[0044] ;

[0045] in, For discount factor Down Time slot Accumulated rewards for each arm's discount For discount factor Down Time slot The number of valid samples per arm For discount factor Down Time slot The average of the discount experience for each arm. for Time slot The standard deviation is used for each arm. For indicator functions, when When true, It is 1 if it is true, otherwise it is 0. This is the preset maximum sampling standard deviation.

[0046] Preferably, the step of using a deep deterministic strategy gradient algorithm to obtain the phase reflection vector of the current time slot's optimal IRS in low-level decision-making includes:

[0047] According to high-level decisions The optimal channel estimation mode under time slots is the action. After the high-level decision-making process is completed, the results are received. Observations in low-level decision-making under time slots ;

[0048] Current time slot Below, lower-level decisions are based on observations Output action And execute it to receive an immediate reward. ;

[0049] High-level decisions Actions under time slots and actions Observations of corresponding high-level decision-making ,as well as Observations of low-level decision-making in time slots ,action and rewards Forming a quintuple ;

[0050] If the experience replay pool is not full, the current quintuple is stored in the experience replay pool, and the interaction with the environment continues to form a new quintuple which is stored in the experience replay pool, until the experience replay pool is full.

[0051] If the experience replay pool is full, the current quintuple will be randomly replaced with a quintuple in the experience replay pool.

[0052] A predetermined number of quintuples are randomly selected from the experience replay pool to train and update the deep deterministic policy gradient algorithm;

[0053] For each update, the Q-value of the current policy is calculated using the Critic current network in the deep deterministic policy gradient algorithm. The parameters of the Critic current network are updated by minimizing the mean squared error loss function, and the parameters of the Actor current network are updated by the policy gradient.

[0054] Update the parameters of the Actor target network and the Critic target network;

[0055] The process is iterated and updated continuously until the preset number of training rounds is reached, resulting in the trained Actor network. The current time slot is then updated accordingly. Observations of low-level decision-making The input is fed into the trained Actor network to obtain the action of the low-level decision, which is the phase reflection vector of the optimal IRS.

[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0057] 1. This IRS-assisted adaptive channel estimation and reflection optimization method constructs an objective function that maximizes the utility of the downlink communication system. The objective function is then transformed into a two-layer partially observable Markov decision process, which includes a higher-layer decision to select the channel estimation mode for the current time slot and a lower-layer decision to select the phase reflection vector of the IRS for the current time slot. This improves the adaptability and stability of the downlink communication system under non-stationary channels. Joint optimization using the discounted Thompson sampling algorithm and the deep deterministic policy gradient algorithm yields the optimal channel estimation mode and the optimal IRS phase reflection vector, achieving channel estimation and reflection optimization.

[0058] 2. After the action in the low-level decision-making process is completed, the reward for the current time slot is obtained. The reward for the current time slot is used as the input for the observation in the high-level decision-making process of the next time slot, thus forming a closed-loop input-output dependency relationship between time slots. This achieves time-series coupling optimization, and has high adaptability, high energy efficiency and fast convergence performance, significantly improving the overall performance of the downlink communication system. Attached Figure Description

[0059] Figure 1 This is a flowchart illustrating the IRS-assisted adaptive channel estimation and reflection optimization method of the present invention.

[0060] Figure 2 This is a schematic diagram of the downlink communication system of the present invention;

[0061] Figure 3 This is a comparison chart showing the convergence of the cumulative reward values ​​of the method of the present invention with those of four other benchmark algorithms in the prior art. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention.

[0064] In one embodiment, such as Figures 1-3 As shown, an IRS-assisted adaptive channel estimation and reflection optimization method is provided, including:

[0065] Step 1: Establish an IRS-assisted downlink communication system, which includes a multi-antenna base station, a single-antenna user equipment, and an IRS. The base station first performs channel gain estimation, and then transmits the information to the user equipment through the IRS (in this embodiment, the direct link between the base station and the user equipment is blocked, such as by a building).

[0066] It should be noted that the channel estimation mode is expressed as ;

[0067] when When, it indicates that the mode is continued, and the channel gain of the concatenated channel estimated in the previous time slot is used as the channel gain of the concatenated channel estimated in the current time slot, that is... ,in, for Channel gain of cascaded channels estimated by time slots. for Channel gain of cascaded channels estimated by time slots.

[0068] when The time slot indicates the test mode, which uses historically estimated channel gain of the cascaded channel and a time series prediction model to estimate and predict the channel gain of the cascaded channel in the current time slot. ,in, For time series forecasting models (soon) (as input to time series forecasting models) For Channel gain of cascaded channels estimated by time slots. For the current time slot Previous A historical time slot.

[0069] when When, it indicates the pilot mode. The channel gain of the concatenated channel estimated in the current time slot is the channel gain of the concatenated channel in the current time slot, that is... ,in for Channel gain of time-slot concatenated channels.

[0070] The preset time period is Each time slot The number of antennas in the base station is One, base station The transmit power of a time slot is expressed as ,and , This is the maximum transmit power of the base station. For base stations The first time slot The transmit power of each antenna, IRS is expressed as , The number of reflective elements in the IRS (e.g., 32). The phase reflection vector of the time slot IRS is expressed as: IRS The reflection coefficient matrix of the time slot is , for The first time slot IRS The reflection amplitude coefficient of each reflective element (in total reflection mode, the reflection amplitude coefficient of all reflective elements is 1). for The first time slot IRS The phase reflection coefficient of each reflecting element, and the range of values ​​is 1. , For transpose;

[0071] The channel gain from the time-slot base station to the IRS is expressed as: , The channel gain from the time slot IRS to the user equipment is expressed as: , and They represent dimensions as follows: and For a complex matrix, then The channel gain of the time-slotted cascaded channel is expressed as: , This is the conjugate transpose. At each time slot t, the signal received by the user equipment... Represented as: ,in This represents Gaussian white noise. For base stations in time slots The transmitted data symbols For noise power, It follows a complex Gaussian distribution.

[0072] To describe the time-varying nature of the channel, let the channel gain of the cascaded channel be... Represented as , This represents the large-scale path loss coefficient. This represents the fast fading component at a small scale in time slot t, assuming... Satisfies the first-order Gaussian Markov property: , For time slots Gaussian random perturbation, The time-domain correlation coefficient between adjacent time slots (e.g., 0.99) indicates that the channel gain of the cascaded channel is time-dependent, and thus the entire channel follows a first-order Markov model.

[0073] Step 2: Based on the transmission rate and energy consumption of the base station in the downlink communication system, construct the objective function for maximizing the utility of the downlink communication system;

[0074] The objective function is expressed as follows:

[0075] ;

[0076] st C1: ;

[0077] C2: ;

[0078] in,

[0079] ;

[0080] ;

[0081] ;

[0082] in, for Time slots in channel estimation mode Phase reflection vector of IRS The function of downlink communication system utility. for Time-slot base stations in channel estimation mode Phase reflection vector of IRS The transmission rate at that time for Time slots in channel estimation mode The energy consumption of the downlink communication system is as follows: multiplexing mode has no additional energy consumption, prediction mode has moderate energy consumption, and pilot mode has the highest energy consumption. In addition, the channel estimation error variance... Closely related to the selected channel estimation mode, its magnitude satisfies: In other words, the multiplexing mode has the largest error, while the pilot mode is the most accurate. The estimation error will directly affect the design function of the IRS phase reflection vector. This can affect the performance of the downlink communication system. As a penalty factor, To estimate the delay for the normalized channel, and the delay varies depending on the mode, in the multiplexing mode. Prediction patterns Moderate latency, pilot mode Maximum delay, satisfying The delay will shorten the transmission time of each time slot to , Based on The maximum transmission precoding vector, for Channel gain of cascaded channels estimated by time slots. In channel estimation mode The corresponding overhead function represents the resource consumption incurred by adopting this mode.

[0083] In the objective function, C1 ensures that the reflection amplitude coefficient of the IRS reflection element is sparse and always 1 in the total reflection mode, and C2 restricts the channel estimation mode to be selected only among the three channel estimation modes.

[0084] Step 3: Transform the objective function into a two-layer partially observable Markov decision process. This process includes a higher-level decision to select the channel estimation mode for the current time slot and a lower-level decision to select the phase reflection vector of the IRS for the current time slot, including:

[0085] Transform the objective function into a two-level partially observable Markov decision process:

[0086] Step 3.1: Define the hidden state During time slot 𝑡, the actual channel state between the base station and user equipment via the IRS, i.e. ,in For state space;

[0087] Step 3.2: Define the role of high-level decision-making. The observation of the time slot is (Used to describe observational information in high-level decision-making), where, This serves as the observation space corresponding to high-level decision-making, and Including the previous time slot Rewards, last time slot High-level decision-making actions and the previous time slot The channel gain is estimated, where the reward for each time slot is the utility of the downlink communication system in the corresponding time slot, i.e. ;

[0088] Step 3.3: Define the role of high-level decision-making. The action of the time slot is ,in, This provides the action space corresponding to high-level decision-making, and That is Channel estimation mode of time slot ,Right now ;

[0089] Step 3.4: After the high-level decision-making action is completed, the observations in the current time slot of the low-level decision-making are obtained, and the low-level decision-making... The observation of the time slot is ,in, This is the observation space corresponding to low-level decision-making, and Include Channel gain estimation for time slots Actions in high-level decision-making within a time slot and the previous time slot Actions in low-level decision-making, namely ;

[0090] Step 3.5: Define the role of low-level decision-making. The action of the time slot is ,in, This represents the action space corresponding to lower-level decisions, and That is Phase reflection vector of the IRS in the time slot The reflection amplitude coefficient of all reflective elements is 1.

[0091] Step 3.6: After the actions in the lower-level decision-making are completed, the reward for the current time slot is obtained. Furthermore, the reward of the current time slot serves as the input for the observation in the high-level decision-making of the next time slot (the reward not only serves as the performance output in the low-level decision-making but also feeds back to the high-level decision-making, serving as the observation in the high-level decision-making of the next time slot). The input is used to form a closed-loop input-output dependency between time slots, thereby achieving timing coupling optimization.

[0092] Step 3.7, Define the state transition function: It is used to characterize the evolution of channel state over time;

[0093] Step 3.8: Define the joint strategy of high-level decision-making and low-level decision-making as follows: ,in This is a strategy for high-level decision-making (used to select the optimal channel estimation mode given observations). This is a strategy for low-level decision-making (used to select the phase reflection vector given an estimated channel).

[0094] Step 3.9: Calculate the belief state value function using the joint strategy, maximize the belief state value function, and take the strategy corresponding to maximizing the belief state value function as the optimal strategy. The problem of solving the optimal strategy is transformed into solving the optimal channel estimation mode in high-level decision-making and solving the optimal phase reflection vector of the IRS in low-level decision-making.

[0095] Among them, the state value function The formula is as follows:

[0096] ;

[0097] in, This is a discount factor used to balance short-term and long-term returns. A larger discount factor emphasizes long-term performance but converges more slowly, while a smaller discount factor emphasizes immediate performance but may lead to suboptimal long-term returns, due to the actual state of the market. Since it cannot be directly observed, the true state is represented by observation. and Instead, the belief state value function The formula is as follows:

[0098] ;

[0099] The optimal strategy is The problem of finding the optimal strategy is transformed into finding the optimal channel estimation mode in high-level decision-making and finding the optimal phase reflection vector of the IRS in low-level decision-making.

[0100] Step 4: In high-level decision-making, the discounted Thompson sampling algorithm is used to optimize and obtain the optimal channel estimation mode for the current time slot. In low-level decision-making, the deep deterministic policy gradient algorithm is used to optimize and obtain the optimal IRS phase reflection vector for the current time slot. The optimal channel estimation mode and IRS phase reflection vector for each time slot within a preset time period are applied to the downlink communication system to realize channel estimation and reflection optimization. (In this embodiment, the original non-convex coupled optimization problem is decomposed into two levels of reinforcement learning decision-making, which can effectively capture the dynamic characteristics and temporal dependencies of the channel, realize the synergistic optimization of channel estimation and IRS phase reflection vector, and ultimately maximize the long-term utility of the downlink communication system.)

[0101] Step 4.1: In high-level decision-making, the Discounted Thompson Sampling algorithm is used to obtain the optimal channel estimation mode for the current time slot, including:

[0102] Channel estimation mode As an arm in the discounted Thompson sampling algorithm ;

[0103] Step 4.1.1 At that time, initialize the discount factor respectively. Next Average discount experience for each arm In discount factor Next Discounts and cumulative rewards for each arm and in discount factor Next Valid sample size of each arm and initialization of the first The standard deviation of each arm is used. ;

[0104] Step 4.1.2, Current Time Slot Next, samples are taken from each arm to obtain the sampled value corresponding to each arm. ,and Follows Gaussian distribution ,in, It follows a Gaussian distribution. For discount factor Down Time slot The average of the discount experience for each arm. for The next time slot Variance of each arm;

[0105] Step 4.1.3: Select the largest sample value from the three sample values, and the arm corresponding to the largest sample value is the [arm name missing]. Slot-optimal channel estimation mode ( (For use in mode, test mode, or pilot mode).

[0106] Step 4.1.4: For the current time slot The following rewards After normalization, we get The normalized reward was obtained by performing a logarithmic transformation. ;

[0107] Step 4.1.5: Utilize the current time slot The result after the lower logarithmic transformation Discount factors respectively The average discount experience value, cumulative discount reward, and effective sample size for each arm are updated, along with the adoption standard deviation for each arm. These updated values ​​are then used as the values ​​for the next time slot. The update formulas for each value are as follows:

[0108] ;

[0109] ;

[0110] ;

[0111] ;

[0112] in, For discount factor Down Time slot Accumulated rewards for each arm's discount For discount factor Down Time slot The number of valid samples per arm For discount factor Down Time slot The average of the discount experience for each arm. for Time slot The standard deviation is used for each arm. For indicator functions, when When true, It is 1 if it is true, otherwise it is 0. This is the preset maximum sampling standard deviation.

[0113] Step 4.2: In the low-level decision-making process, the Deep Deterministic Policy Gradient (DDPG) algorithm is used to solve for the phase reflection vector of the IRS optimal for the current time slot, including:

[0114] Step 4.2.1: Based on high-level decisions The optimal channel estimation mode under time slots is the action. After the high-level decision-making process is completed, the results are received. Observations in low-level decision-making under time slots ;

[0115] Step 4.2.2, Current Time Slot Below, lower-level decisions are based on observations Output action And execute it to receive an immediate reward. ;

[0116] Step 4.2.3: Incorporate high-level decisions into... Actions under time slots and actions Observations of corresponding high-level decision-making ,as well as Observations of low-level decision-making in time slots ,action and rewards Forming a quintuple ;

[0117] Step 4.2.4: If the experience replay pool is not full, the current quintuple is stored in the experience replay pool, and the interaction with the environment continues to form a new quintuple and store it in the experience replay pool until the experience replay pool is full (this step is a conventional prior art in this field).

[0118] Step 4.2.5: If the experience replay pool is full, the current quintuple will be randomly replaced with a quintuple in the experience replay pool.

[0119] Step 4.2.6: Randomly select a predetermined number of 5-tuples from the experience replay pool to train and update the deep deterministic policy gradient algorithm (the specific training process is existing technology, as steps 4.2.7-4.2.9 are only briefly described. The parameters involved in the training process are: number of training rounds: 10000; number of steps per round: 100; discount factor γ: 0.99; learning rate: 0.001; batch size B: 256; experience replay pool capacity: 100000).

[0120] Step 4.2.7: For each update, use the Critic current network in the deep deterministic policy gradient algorithm to calculate the Q value of the current policy, update the parameters of the Critic current network by minimizing the mean squared error loss function, and update the parameters of the Actor current network by the policy gradient.

[0121] Step 4.2.8: Update the parameters of the Actor target network and the Critic target network;

[0122] Step 4.2.9: Iterate and update continuously until the preset number of training rounds is reached, obtain the trained Actor current network, and set the current time slot. Observations of low-level decision-making The input is fed into the trained Actor network to obtain the action of the low-level decision, which is the phase reflection vector of the optimal IRS;

[0123] Step 4.2.10: Proceed to the next time slot, iterating through steps 4.1.2 to 4.2.10 for each time slot. All operations are performed according to steps 4.1.2-4.2.10, thereby obtaining the optimal channel estimation mode and the optimal phase reflection vector of the IRS for each time slot. The optimal channel estimation mode and the phase reflection vector of the IRS for each time slot are applied to the downlink communication system to realize channel estimation and reflection optimization.

[0124] In another embodiment, the method is further verified by combining specific experimental results. The experiment is implemented in Python 3.7 and runs on a computer equipped with an Intel(R) Core(TM) i7-10510U CPU and 12GB of memory.

[0125] Figure 3 The cumulative reward comparison between our method and four existing benchmark optimization algorithms is shown.

[0126] DTS-Manifold: This method combines the Discounted Thompson Sampling (DTS) algorithm to optimize the optimal channel estimation mode with the Riemannian manifold optimization (Manifold) to optimize the phase reflection vector of the optimal IRS. The phase shift matrix of the IRS is parameterized on the Steifel manifold to satisfy the constant mode constraint.

[0127] DTS-PPO: This method combines the Discounted Thompson Sampling (DTS) algorithm to optimize the optimal channel estimation pattern with the Near-End Policy Optimization (PPO) algorithm to optimize the phase reflection vector of the optimal IRS.

[0128] TS-DDPG: This method uses standard Thompson sampling (TS) to optimize the optimal channel estimation pattern and uses the DDPG algorithm to optimize the phase reflection vector of the optimal IRS;

[0129] Stochastic strategy (i.e., both the optimization algorithms for high-level and low-level decisions are randomly selected): The channel estimation mode and IRS phase shift configuration are randomly selected in each time slot, serving as a stochastic baseline algorithm for the performance lower bound.

[0130] from Figure 3 It can be seen that the training rewards of both DDPG and PPO reinforcement learning algorithms steadily increase with each training round, indicating that they have good policy learning capabilities. In contrast, the reward curve of the Manifold algorithm remains almost constant, lacking performance improvement with training iterations, indicating that its adaptive learning ability in complex scenarios is limited.

[0131] Furthermore, both our proposed method and the existing DTS-PPO method exhibit a significant upward trend in rewards, with lower initial rewards but gradual convergence. Notably, our proposed method converges faster than the existing DTS-PPO method and achieves a higher average reward upon convergence. Specifically, our proposed method improves the average reward by approximately 4.1% and 5.6% compared to the existing DTS-PPO and DTS-Manifold methods, respectively, and achieves a performance improvement of approximately 71% compared to stochastic policies. These results clearly demonstrate the advantages of our proposed method in terms of learning efficiency and policy quality.

[0132] Further analysis reveals that in scenarios requiring rapid convergence to a deterministic optimal action, the stochasticity of the PPO strategy leads to a decrease in its exploration efficiency, resulting in slightly lower final performance compared to DDPG. For manifold optimization-based methods, due to their reliance on geometric modeling of the state space, they are generally more suitable for low-dimensional or structured problems. However, they lack continuous learning and adaptive capabilities in complex dynamic environments, thus their overall performance is inferior to DDPG-like algorithms with end-to-end deep learning characteristics.

[0133] Overall, this method demonstrates faster convergence speed, higher average reward, and more stable training performance in complex dynamic environments, verifying the effectiveness and superiority of the algorithm in balanced channel estimation decision-making and IRS beamforming optimization.

[0134] This IRS-assisted adaptive channel estimation and reflection optimization method constructs an objective function that maximizes the utility of the downlink communication system. This objective function is then transformed into a two-layer partially observable Markov decision process, which includes a higher-level decision to select the channel estimation mode for the current time slot and a lower-level decision to select the phase reflection vector of the IRS for the current time slot. This improves the adaptability and stability of the downlink communication system under non-stationary channels. Joint optimization using the discounted Thompson sampling algorithm and the deep deterministic policy gradient algorithm yields the optimal channel estimation mode and the optimal IRS phase reflection vector, achieving channel estimation and reflection optimization. After the actions in the lower-level decision are executed, the reward for the current time slot is obtained. This reward serves as the input for the observation in the higher-level decision of the next time slot, forming a closed-loop input-output dependency between time slots. This achieves time-coupling optimization, exhibiting high adaptability, high energy efficiency, and fast convergence performance, significantly improving the overall performance of the downlink communication system.

[0135] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0136] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0137] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A method of IRS-aided adaptive channel estimation and reflection optimization, characterized in that: The IRS-assisted adaptive channel estimation and reflection optimization method includes: An IRS-assisted downlink communication system is established, and the downlink communication system includes a multi-antenna base station, a single-antenna user equipment, and an IRS. The base station first performs channel gain estimation and then transmits the information to the user equipment through the IRS. Based on the transmission rate and energy consumption of the base station in the downlink communication system, an objective function for maximizing the utility of the downlink communication system is constructed. The objective function is transformed into a two-layer partially observable Markov decision process, which includes a higher-layer decision to select the channel estimation mode for the current time slot and a lower-layer decision to select the phase reflection vector of the IRS for the current time slot. In high-level decision-making, the discounted Thompson sampling algorithm is used to optimize the channel estimation mode for the current time slot. In low-level decision-making, the deep deterministic strategy gradient algorithm is used to optimize the phase reflection vector of the IRS for the current time slot. The optimal channel estimation mode and the phase reflection vector of the IRS for each time slot within a preset time period are applied to the downlink communication system to realize channel estimation and reflection optimization.

2. The IRS-assisted adaptive channel estimation and reflection optimization method as described in claim 1, characterized in that: The preset time period is: Each time slot The number of antennas in the base station is One, IRS is represented as , The number of reflective elements in the IRS. The phase reflection vector of the time slot IRS is expressed as: , for The first time slot IRS The reflection amplitude coefficient of each reflecting element. for The first time slot IRS The phase reflection coefficient of each reflecting element, and the range of values ​​is 1. , For transpose; The channel gain from the time-slot base station to the IRS is denoted by , The channel gain from the IRS to the user equipment in the time slot is denoted by , and denote the complex matrices of dimensions and respectively, then The channel gain of the cascaded channel in the time slot is denoted by , is the conjugate transpose.

3. The IRS-assisted adaptive channel estimation and reflection optimization method as described in claim 1, characterized in that: The channel estimation mode is represented as ; When Adopt the channel gain of the concatenated channel estimated in the last time slot as the channel gain of the concatenated channel estimated in the current time slot. When a test mode is indicated, the channel gain of the concatenated channel for the current time slot is estimated and predicted using the historically estimated channel gain of the concatenated channel and the time series prediction model. When , represents the pilot pattern, the channel gain of the concatenated channel estimated in the current time slot is the channel gain of the concatenated channel in the current time slot.

4. The IRS-assisted adaptive channel estimation and reflection optimization method of Claim 1, wherein: The objective function for maximizing the utility of the downlink communication system, based on the transmission rate and energy consumption of the base station in the downlink communication system, includes: The objective function is expressed as follows: ; s.t. C1: ; C2: ; in, ; wherein is the transmission rate of the time slot base station in the channel estimation mode and the phase reflection vector of the IRS, as a function of the utility of the downlink communication system in the channel estimation mode is the transmission rate of the time slot base station in the channel estimation mode and the phase reflection vector of the IRS, as a function of the utility of the downlink communication system in the channel estimation mode is the energy consumption of the time slot in the channel estimation mode as a function of the utility of the downlink communication system in the channel estimation mode is a penalty factor.

5. The IRS-assisted adaptive channel estimation and reflection optimization method as described in claim 4, characterized in that: The process of transforming the objective function into a two-level partially observable Markov decision process includes: Define the role of high-level decision-making The observation of the time slot is ,in, This serves as the observation space corresponding to high-level decision-making, and Including the previous time slot Rewards, last time slot High-level decision-making actions and the previous time slot The channel gain is estimated, where the reward for each time slot is the utility of the downlink communication system in the corresponding time slot; Define the role of high-level decision-making The action of the time slot is ,in, This provides the action space corresponding to high-level decision-making, and That is Channel estimation mode of time slot ; Once the actions of the high-level decision-making process are completed, the observations from the low-level decision-making process in the current time slot are obtained, and the low-level decision-making process... The observation of the time slot is ,in, This is the observation space corresponding to low-level decision-making, and Include Channel gain estimation for time slots Actions in high-level decision-making within a time slot and the previous time slot Actions in low-level decision-making; Define low-level decision-making The action of the time slot is ,in, This represents the action space corresponding to lower-level decisions, and That is Phase reflection vector of the IRS in the time slot ; After the action in the lower-level decision-making is completed, the reward for the current time slot is obtained, and the reward for the current time slot is used as the input for the observation in the higher-level decision-making in the next time slot; Define the joint strategy of high-level decision-making and low-level decision-making as ,in Strategies for high-level decision-making. Strategies for lower-level decision-making; The belief state value function is calculated using a joint strategy, and the belief state value function is maximized. The strategy corresponding to maximizing the belief state value function is taken as the optimal strategy. The problem of solving the optimal strategy is transformed into solving the optimal channel estimation mode in high-level decision-making and solving the optimal phase reflection vector of the IRS in low-level decision-making.

6. The IRS-assisted adaptive channel estimation and reflection optimization method as described in claim 5, characterized in that: The method of using the discounted Thompson sampling algorithm in high-level decision-making to obtain the optimal channel estimation mode for the current time slot includes: Channel estimation mode as an arm in the discounted Thompson sampling algorithm ; At that time, initialize the discount factor respectively. Next Average discount experience for each arm In discount factor Next Discounts and cumulative rewards for each arm and in discount factor Next Valid sample size of each arm and initialization of the first The standard deviation of each arm is used. ; Current time slot Next, samples are taken from each arm to obtain the sampled value corresponding to each arm. ,and Follows a Gaussian distribution ,in, It follows a Gaussian distribution. For discount factor Down Time slot The average of the discount experience for each arm for The next time slot Variance of each arm; Select the largest sample value from the three sample values, and the arm corresponding to the largest sample value is the [arm value]. Slot-optimal channel estimation mode ; For the current time slot The following rewards After normalization, we get The normalized reward was obtained by performing a logarithmic transformation. ; Utilize the current time slot The result after the lower logarithmic transformation Discount factors respectively The average discount experience value, cumulative discount reward, and effective sample size for each arm are updated, along with the adoption standard deviation for each arm. These updated values ​​are then used as the values ​​for the next time slot. The update formulas for each value are as follows: ; ; ; ; in, For discount factor Down Time slot Accumulated rewards for each arm's discount For discount factor Down Time slot The number of valid samples per arm For discount factor Down Time slot The average of the discount experience for each arm for Time slot The standard deviation is used for each arm. For indicator functions, when When true, It is 1 if it is true, otherwise it is 0. This is the preset maximum sampling standard deviation.

7. The IRS-assisted adaptive channel estimation and reflection optimization method as described in claim 6, characterized in that: The step of using a deep deterministic strategy gradient algorithm in low-level decision-making to obtain the phase reflection vector of the current time slot's optimal IRS includes: According to high-level decisions The optimal channel estimation mode under time slots is the action. After the high-level decision-making process is completed, the results are received. Observations in low-level decision-making under time slots ; Current time slot Below, lower-level decisions are based on observations Output action And execute it to receive an immediate reward. ; High-level decisions Actions under time slots and actions Observations of corresponding high-level decision-making ,as well as Observations of low-level decision-making in time slots ,action and rewards Forming a quintuple ; If the experience replay pool is not full, the current quintuple is stored in the experience replay pool, and the interaction with the environment continues to form a new quintuple which is stored in the experience replay pool, until the experience replay pool is full. If the experience replay pool is full, the current quintuple will be randomly replaced with a quintuple in the experience replay pool. A predetermined number of quintuples are randomly selected from the experience replay pool to train and update the deep deterministic policy gradient algorithm; For each update, the Q-value of the current policy is calculated using the Critic current network in the deep deterministic policy gradient algorithm. The parameters of the Critic current network are updated by minimizing the mean squared error loss function, and the parameters of the Actor current network are updated by the policy gradient. Update the parameters of the Actor target network and the Critic target network; The process is iterated and updated continuously until the preset number of training rounds is reached, resulting in the trained Actor network. The current time slot is then updated accordingly. Observations of low-level decision-making The input is fed into the trained Actor network to obtain the action of the low-level decision, which is the phase reflection vector of the optimal IRS.