A dual-time-scale beamforming method for intelligent reflecting surface-assisted systems

By optimizing the beamforming of base stations and smart reflective surfaces through deep reinforcement learning algorithms, the problems of high computational complexity and large training label requirements are solved, achieving fast and stable communication quality improvement, and is suitable for dual-time-scale transmission solutions.

CN116527093BActive Publication Date: 2025-10-03SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211613306.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2025-10-03
Estimated Expiration
2042-12-15

AI Technical Summary

Technical Problem

Under the dual-time-scale transmission scheme, the existing base station beamforming and intelligent reflecting surface beamforming optimization have problems with high computational complexity and large training label requirements, resulting in the inability to effectively improve the quality of communication services.

Method used

A deep reinforcement learning algorithm is used, combined with intelligent agents Θ and W, to optimize the base station beamforming matrix and the phase shift of the smart reflective surface through channel state information. An evaluation network, action network, experience pool and optimizer are constructed to train the reflection coefficient matrix of the base station and the smart reflective surface, achieving rapid optimization.

Benefits of technology

It reduces the algorithm execution time cost, improves the system spectrum efficiency, enhances the robustness and stability of the communication environment, and is suitable for various wireless communication environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116527093B_ABST
    Figure CN116527093B_ABST
Patent Text Reader

Abstract

This invention discloses a dual-time-scale intelligent reflecting surface-assisted system beamforming method, suitable for intelligent reflecting surface-assisted multi-user downlink transmission systems. Signals transmitted by the base station are first beamformed by the base station before being reflected by the intelligent reflecting surface to the user end, thereby enhancing the received signal at the user end. This method employs two agents, one for performing intelligent reflecting surface beamforming and the other for performing base station beamforming under statistical channel state information and the other for instantaneous channel state information, to maximize the system sum rate. This method features smooth and fast training, achieving near-optimal performance in a very short time when used online.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology, and in particular to a base station beamforming and intelligent reflection surface beamforming optimization method based on deep reinforcement learning under a dual-time-scale transmission scheme. Background Art

[0002] With the recent development of sixth-generation mobile communications (6G), wireless mobile communications are facing increasingly stringent requirements, such as ultra-high data rates and energy efficiency, global connectivity, high reliability, and low latency. Reconfigurable Intelligent Surfaces (RIS) have emerged as a green and cost-effective solution. They coordinate the reflections of all their components to modify the wireless channel between transmitters and receivers, improving the efficiency and reliability of wireless communications. RIS are also lightweight and easy to install, offering high flexibility and excellent compatibility in practical deployments.

[0003] In practice, base station beamforming and RIS beamforming optimization based entirely on instantaneous channel state information incur significant pilot overhead. Therefore, a dual-time-scale transmission scheme can be employed, designing the base station beamforming matrix and the long-term RIS phase shift matrix based on instantaneous and statistical channel state information, respectively. To this end, some researchers have proposed using the Stochastic Successive Convex Approximation (SSCA) algorithm for this purpose. However, this results in excessive computational complexity and is unsuitable for RIS deployments with large arrays. On the other hand, using deep learning to optimize beamforming requires acquiring a large number of training labels, which is practically impossible and therefore impractical.

[0004] Reinforcement learning, also known as enhanced learning, is based on two main optimization approaches: value-based and policy-based. Combined with neural networks, it forms deep reinforcement learning, which can handle high-dimensional state and action spaces, avoiding dimensionality explosion. Deep reinforcement learning is widely used in the communications field because it eliminates the need for large amounts of training labels and is widely applicable in complex scenarios. Summary of the Invention

[0005] Technical Problem: In view of this, the purpose of the present invention is to provide a dual-time-scale intelligent reflective surface-assisted system beamforming method to solve the technical problems mentioned in the background technology. The present invention configures a base station with a uniform linear array, deploys multiple single-antenna users, and places intelligent reflective surfaces to improve the quality of communication services. Under the dual-time-scale transmission scheme, a deep reinforcement learning algorithm is used to design the base station beamforming and the phase shift of the intelligent reflective surface based on channel state information to maximize the system spectrum efficiency. The model trained using deep reinforcement learning can greatly reduce the time cost of algorithm execution.

[0006] Technical solution: To achieve the above-mentioned purpose, the present invention provides a dual-time-scale intelligent reflective surface-assisted system beamforming method comprising the following steps:

[0007] Step 1: In a single-cell downlink transmission system, the base station is configured with a uniform linear antenna array, which includes M antenna elements; the smart reflective surface is configured with a uniform planar reflective unit, including N vertical directions. y Line reflection unit, N per line in horizontal direction x Reflection units, a total of N = N x N y reflection units, whose reflection coefficient matrix is ​​expressed as a diagonal matrix in is the nth diagonal element of Θ, q n is the known reflection amplitude of the nth reflection element, θ n is the reflection phase shift of the nth reflective element; the cell contains K single-antenna users;

[0008] The channel state information obtained by the base station and the smart reflective surface includes two parts: statistical channel state information and instantaneous channel state information; the statistical channel state information includes: the line-of-sight channel matrix from the base station to the smart reflective surface Line-of-sight channel vector from the smart reflecting surface to the kth user Line-of-sight channel vector from the base station to the kth user k = 1, ..., K; the instantaneous channel state information is the equivalent channel vector from the base station to each user k in is the instantaneous channel vector from the smart reflection surface to the kth user, is the instantaneous channel matrix from the base station to the smart reflective surface, is the instantaneous channel vector from the intelligent reflecting surface to the kth user; the base station uses the beamforming matrix For signal transmission, each column represents the beamforming vector for transmitting signals to user k;

[0009] Step 2: Construct agent Θ and agent W; the agent Θ takes the real and imaginary parts of the current channel state information as its state, which is expressed as Where Re(·) represents the real part, Im(·) represents the imaginary part, and the phase shift of the smart reflective surface at the current moment is taken as the action, which is expressed as The channel statistics at the current moment within the coherence time T H The average value of the users and rates in the time slot is the reward, which is expressed as in The base station adopts the time slot t′

[0010] The beamforming matrix W and the user and rate under the smart reflective surface reflection coefficient matrix corresponding to the current action, is the additive noise power of user k; the agent W takes the real and imaginary parts of the instantaneous equivalent channel vector of the current time slot as the state, which is expressed as Taking the real and imaginary parts of the base station beamforming matrix of the current time slot as the action, it is expressed as The current time slot user and rate are used as rewards, which can be expressed as

[0011] Construct the internal modules of the intelligent agent Θ, including: evaluation network, action network, experience pool U1, optimizer, where the evaluation network parameters are The action network parameter is θ Θ ; Construct the internal modules of the intelligent agent W, including: evaluation network, action network, experience pool U2, optimizer, where the evaluation network parameters are The action network parameter is θ W ;

[0012] The action network output mean u of the agent Θ Θ and standard deviation σ Θ , and then the mean is u Θ And the standard deviation is σ Θ The normal distribution of the sample is sampled, the sampled value is multiplied by π, and then clipped to the range [-π, π] as the action of the agent Θ; the action network output mean u of the agent W W and standard deviation σ W , and then the mean is u W And the standard deviation is σ W The normal distribution of is sampled, and the sampled value is used as the action of the intelligent agent W;

[0013] The evaluation network of the agent Θ outputs a value function based on the state to evaluate the quality of the action selected by the agent Θ; the evaluation network of the agent W outputs a value function based on the state to evaluate the quality of the action selected by the agent W;

[0014] The experience pool U1 of the agent Θ is responsible for storing samples generated during the learning process of the agent Θ, and the experience pool U2 of the agent W is responsible for storing samples generated during the learning process of the agent W;

[0015] The optimizer of the agent Θ is responsible for training the action network and evaluation network parameters of the agent Θ, and the optimizer of the agent W is responsible for training the action network and evaluation network parameters of the agent W;

[0016] Step 3: Train the evaluation network and action network of the two agents Θ and W constructed in step 2 to obtain the trained action network and evaluation network of the two agents Θ and W;

[0017] Step 4: The action networks of the two agents Θ and W trained in step 3 are used online. The agent Θ observes the current state. Input to its action network and output the mean value u of the network Θ Multiply by π and then clip to the range [-π,π] as the action represents the phase shift of the smart reflective surface element, and then the phase shift Mapped to smart reflective surface reflection coefficient matrix During the channel statistical coherence time from the current moment to the next moment, observe the state of the agent W in the current time slot Input to its action network and output the mean value u of the network W As an action Construct W by Re(W) and Im(W), and then form the base station beamforming matrix Finally, the optimal base station active beamforming matrix and smart reflective surface passive beamforming matrix are obtained.

[0018] In step 2, the evaluation network and action network of the two agents Θ and W each include an input layer, a hidden layer, and an output layer.

[0019] The training in step 3 specifically includes the following sub-steps:

[0020] a1) Randomly initialize the action network parameters θ of the agent Θ Θ and evaluate network parameters Their initial values ​​are recorded as and Randomly initialize the action network parameters θ of intelligent W W and evaluate network parameters Their initial values ​​are recorded as and Initialize the base station transmitter power to P, the number of statistical channel state information training is 2T, and the number of time slots within the channel statistical coherence time is T H, the maximum number of iterations is H, and the current number of iterations h is 1;

[0021] a2) When h≤H, set the initial time t=1, initialize the experience pool U1 and the experience pool U2, and go to step a3); otherwise, go to step a15);

[0022] a3) At time t, agent Θ observes the current statistical channel state information Set the state of the agent Θ at the current moment to Agent Θ will Input to its evaluation network, output Represents input When the parameter is The output of the evaluation network; Input to its action network, output mean u Θ and standard deviation σ Θ , the mean value is u Θ And the standard deviation is σ Θ The sampled value is multiplied by π and then clipped to the range [-π, π] as the action Record the agent Θ in state Downsampling Probability action As the phase shift of N reflection units of the smart reflection surface, the reflection coefficient matrix of the smart reflection surface can be further formed: And configure it;

[0023] a4) The channel statistical coherence time from time t to time t+1 contains T H Time slots, set the initialization time slot t'=1;

[0024] a5) At time slot t', agent W observes the instantaneous equivalent channel vector h from the base station to each user k in the current time slot. k , set the current time slot agent W state to Agent W will Input to its action network, output mean u W and standard deviation σ W , the mean u W As an action Construct W by Re(W) and Im(W), and then form the base station beamforming matrix of the current time slot as According to the instantaneous equivalent channel vector h of the current time slot k The instantaneous sum rate of the current time slot is calculated with the base station beamforming matrix W as

[0025] a6) If t'<T H, let t'=t'+1 and go to step a5); otherwise, for t'=1,...,T H T H time slots Calculate the average This is used as the reward value of the agent Θ at the current moment; the state of the agent Θ at the next moment is obtained Will It is stored in the experience pool U1 as an experience;

[0026] a7) Let t = t + 1. When t ≤ T, proceed to step a3) again. When t > T, proceed to step a8);

[0027] a8) At time t, the agent Θ observes the current statistical channel state information Set the state of the agent Θ at the current moment to Agent Θ will Input to its action network, output mean u Θ and standard deviation σ Θ , the mean value u of the network output Θ Multiply by π and then clip to the range [-π,π] as the action represents the phase shift of the smart reflective surface element, and then Mapped to smart reflective surface reflection coefficient matrix Initialization time slot t'=1;

[0028] a9) At time slot t', agent W observes the instantaneous equivalent channel vector h from the base station to each user k in the current time slot. k , set the current time slot agent W state to Agent W will Input to its evaluation network and output value function Represents input status When the parameter is The output of the evaluation network will be Input to its action network, output mean u W and standard deviation σ W , the mean value is u W The standard deviation is σ W Sampling from the normal distribution of And record the agent W in state Downsampling acquisition Probability Get Action After that, W is constructed by Re(W) and Im(W), and then the base station beamforming matrix of the current time slot is reconstructed as According to the instantaneous equivalent channel vector h of the current time slot k The instantaneous sum rate of the current time slot is calculated with the base station beamforming matrix W as Will As the reward value of the current time slot agent W; get the observation state of the next time slot agent W Will It is stored in the experience pool U2 as an experience;

[0029] a10) If t'<T H , set t'=t'+1 and go to step a9); otherwise, set t=t+1 and go to a11);

[0030] a11) If t≤2T, proceed to step a8); otherwise, proceed to step a12);

[0031] a12) Using the experience samples in the experience pool U1, calculate the advantage function of the agent Θ and target value Using the experience samples in the experience pool U2, calculate the advantage function of the agent W and target value Set the agent Θ mini-batch sampling size to B Θ , set the mini-batch sampling size of agent W to B W , set the total number of training times in one iteration to J, and the initial number of training times to j = 1;

[0032] a13) Randomly sample non-repeated experience from the experience pool U1 The batch size for each sampling is B Θ ;Will Input to the current agent Θ evaluation network to obtain the value function Randomly sample non-repeated experience from the experience pool U2 The batch size for each sampling is B W ;Will Input to the current agent W evaluation network to obtain the value function Use policy gradient to update the evaluation network and action network parameters of the two agents respectively;

[0033] a14) Set j = j + 1. When j ≤ J, proceed to a13). When j > J, set h = h + 1 and proceed to a2).

[0034] a15) Obtain the action network and evaluation network of the trained agents Θ and W.

[0035] In step a12), the advantage function of the agent Θ is calculated using the generalized advantage estimation method. and the advantage function of agent W The specific method is:

[0036]

[0037]

[0038] Where γ is the loss factor, λ is the generalized advantage estimation parameter, and r t is the agent’s reward value, Represents input s t When the parameter is The output of the evaluation network, Represents input s t+1 When the parameter is The output of the evaluation network;

[0039] The target value of the agent Θ in step a12) and the target value of agent W The calculation method is:

[0040]

[0041] Represents input s t When the parameter is The output of the evaluation network;

[0042] When the above advantage function and target value calculation method are used, when it is an intelligent agent Θ, When it is agent W,

[0043] In step a13), the evaluation network uses the gradient descent method to update the network parameters, and the update strategy gradient is:

[0044]

[0045] Among them, B is the agent mini-batch sampling size, Indicates the parameters Find partial derivatives and evaluate network parameters The update process is: α c is the learning rate; when evaluating network updates, when it is the agent Θ, B=B Θ , θ=θ Θ , When it is agent W, B=B W , θ=θ W ,

[0046] In step a13), the action network uses the gradient ascent method to update the network parameters, and the specific policy gradient is:

[0047]

[0048] in, is the ratio of strategy probabilities, π θ (a t |s t ) is an action network using parameter θ, state s t Downsampling to obtain action a t The probability of To use the parameter θ old When the action network is , the state s t Downsampling to obtain action a t The probability of clip(·) is the clipping function, which is to eliminate ρ t The excitation in (θ) is outside the interval [1-ε,1+ε], ε is a hyperparameter that controls the clipping range, H(π(·|s t )) is the policy entropy, c is a constant coefficient, ▽ θ Indicates the partial derivative of the variable θ, θ is the action network parameter, and its update process is: θ = θ + α a Δθ, α a is the learning rate;

[0049] When the action network is updated, when it is the agent Θ, B=B Θ , θ=θ Θ , When it is agent W, B=B W , θ=θ W ,

[0050] Beneficial effects: The dual-time-scale intelligent reflective surface-assisted system beamforming method of the present invention has the following advantages:

[0051] 1. The present invention has good robustness to channels under the dual-time-scale transmission scheme and is applicable to various typical wireless communication environments;

[0052] 2. The training process of the base station beamforming matrix and the smart reflective surface phase shift matrix of the present invention is stable, converges quickly, and is easy to implement;

[0053] 3. The online processing time of the present invention is extremely short, and higher system spectrum efficiency can be achieved at a lower time cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 Flowchart for offline training of two agents. DETAILED DESCRIPTION

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0056] Considering a multi-user MISO system with dual time scales, the base station beamforming matrix and the smart reflective surface beamforming matrix are designed based on the principle of maximizing spectrum efficiency. More specifically, the following steps are involved:

[0057] Step 1: In a single-cell downlink transmission system, the base station is configured with a uniform linear antenna array, which includes M = 4 antenna elements; the smart reflective surface is configured with a uniform planar reflective unit, including N vertical directions. y = 8 rows of reflective units, N per row in the horizontal direction x =8 reflective units, N=N x N y = 64 reflection units, and its reflection coefficient matrix is ​​expressed as a diagonal matrix The cell contains K = 2 single-antenna users;

[0058] The channel state information obtained by the base station and the smart reflective surface includes statistical channel state information and instantaneous channel state information. The statistical channel state information includes: the line-of-sight channel matrix from the base station to the smart reflective surface Line-of-sight channel vector from the smart reflecting surface to the kth user Line-of-sight channel vector from the base station to the kth user The instantaneous channel state information is the equivalent channel vector from the base station to each user k The base station uses a beamforming matrix For signal transmission, each column represents the beamforming vector for transmitting signals to user k.

[0059] Step 2: Construct agent Θ and agent W; the agent Θ takes the real and imaginary parts of the current channel state information as its state, which is expressed as Where Re(·) represents the real part, Im(·) represents the imaginary part, and the phase shift of the smart reflective surface at the current moment is taken as the action, which is expressed as The channel statistics at the current moment within the coherence time T H = The average value of the sum of users and rates in 10 time slots is the reward, expressed as The agent W takes the real and imaginary parts of the instantaneous equivalent channel vector of the current time slot as its state, which is expressed as Taking the real and imaginary parts of the base station beamforming matrix of the current time slot as the action, it is expressed as The current time slot user and rate are used as rewards, which can be expressed as Construct the internal modules of the intelligent agent Θ, including: evaluation network, action network, experience pool U1, optimizer, where the evaluation network parameters are The action network parameter is θ Θ ; Construct the internal modules of the intelligent agent W, including: evaluation network, action network, experience pool U2, optimizer, where the evaluation network parameters are The action network parameter is θ W ;

[0060] The action network output mean u of the agent Θ Θ and standard deviation σ Θ , and then the mean is u Θ And the standard deviation is σ Θ The normal distribution of the sample is sampled, the sampled value is multiplied by π, and then clipped to the range [-π, π] as the action of the agent Θ. The action network output mean u of the agent W is W and standard deviation σ W , and then the mean is u W And the standard deviation is σ W The normal distribution of is sampled, and the sampled value is used as the action of the intelligent agent W;

[0061] The evaluation network of the agent Θ outputs a value function based on the state, which is used to evaluate the quality of the action selected by the agent Θ. The evaluation network of the agent W outputs a value function based on the state, which is used to evaluate the quality of the action selected by the agent W;

[0062] The experience pool U1 of the agent Θ is responsible for storing samples generated during the learning process of the agent Θ, and the experience pool U2 of the agent W is responsible for storing samples generated during the learning process of the agent W;

[0063] The optimizer of the agent Θ is responsible for training the action network and evaluation network parameters of the agent Θ, and the optimizer of the agent W is responsible for training the action network and evaluation network parameters of the agent W;

[0064] Step 3: Train the evaluation network and action network of the two agents Θ and W constructed in step 2 to obtain the trained action network and evaluation network of the two agents Θ and W. The specific training includes the following sub-steps:

[0065] a1) Randomly initialize the action network parameters θ of the agent Θ Θ and evaluate network parameters Their initial values ​​are recorded as and Randomly initialize the action network parameters θ of intelligent W W and evaluate network parameters Their initial values ​​are recorded as and Initialize the base station transmitter power to P = 10dBm, the number of statistical channel state information training is 2T = 2000, and the number of time slots within the channel statistical coherence time is T H =10, the maximum number of iterations is H=8000, and the current number of iterations h is 1.

[0066] a2) When h≤H, set the initial time t=1, initialize the experience pools U1 and U2, and go to step a3); otherwise, go to step a15).

[0067] a3) At time t, agent Θ observes the current statistical channel state information Set the state of the agent Θ at the current moment to Agent Θ will Input to its evaluation network, output Represents input When the parameter is The output of the evaluation network; Input to its action network, output mean u Θ and standard deviation σ Θ , the mean value is u Θ And the standard deviation is σ Θ The sampled value is multiplied by π and then clipped to the range [-π, π] as the action Record the agent Θ in state Downsampling Probability Will action Mapped to the phase shift of N=64 reflective units on the smart reflective surface at the current moment Further form the intelligent reflection surface reflection coefficient matrix and configure it.

[0068] a4) The channel statistical coherence time from time t to time t+1 contains T H =10 time slots, set the initialization time slot t'=1.

[0069] a5) At time slot t', agent W observes the instantaneous equivalent channel vector h from the base station to each user k in the current time slot. k (k=1,2), set the state of the current time slot agent W to Agent W will Input to its action network, output mean u W and standard deviation σ W , the mean u W As an action Construct W by Re(W) and Im(W), and then form the base station beamforming matrix of the current time slot as According to the instantaneous equivalent channel vector h of the current time slot k (k=1,2) and the base station beamforming matrix W are used to calculate the instantaneous sum rate of the current time slot:

[0070] a6) If t'<T H , let t'=t'+1 and go to step a5); otherwise, for the 10 time slots of t'=1,...,10 Calculate the average This is used as the reward value of the agent Θ at the current moment; the state of the agent Θ at the next moment is obtained Will It is stored in the experience pool U1 as an experience;

[0071] a7) Let t = t + 1. When t ≤ T, go to step a3) again. When t > T, go to step a8).

[0072] a8) At time t, the agent Θ observes the current statistical channel state information Set the state of the agent Θ at the current moment to Agent Θ will Input to its action network, output mean u Θ and standard deviation σ Θ , the mean value u of the network output Θ Multiply by π and then clip to the range [-π,π] as the action represents the phase shift of the smart reflective surface element, and then Mapped to smart reflective surface reflection coefficient matrix And configure; initialize time slot t'=1;

[0073] a9) At time slot t', agent W observes the instantaneous equivalent channel vector h from the base station to each user k in the current time slot. k (k=1,2), set the state of the current time slot agent W to Agent W will Input to its evaluation network and output value function Represents input status When the parameter is The output of the evaluation network will be Input to its action network, output mean u W and standard deviation σ W , the mean value is u W The standard deviation is σ W Sampling from the normal distribution of And record the agent W in state Downsampling acquisition Probability Get Action After that, W is constructed by Re(W) and Im(W), and then the base station beamforming matrix of the current time slot is reconstructed as According to the instantaneous equivalent channel vector h of the current time slot k The instantaneous sum rate of the current time slot is calculated with the base station beamforming matrix W as Will As the reward value of the current time slot agent W; get the observation state of the next time slot agent W Will It is stored in the experience pool U2 as an experience;

[0074] a10) If t'<T H , let t'=t'+1 and go to step a9); otherwise, let t=t+1 and go to a11).

[0075] a11) If t≤2T, go to step a8); otherwise, go to step a12).

[0076] a12) Using the experience samples in the experience pool U1, calculate the advantage function of the agent Θ and target value Using the experience samples in the experience pool U2, calculate the advantage function of the agent W and target value Set the agent Θ mini-batch sampling size to B Θ =64, set the mini-batch sampling size of agent W to B W =128, set the total number of training times in one iteration to J=10, and the initial number of training times to j=1;

[0077] a13) Randomly sample non-repeated experience from the experience pool U1 The batch size for each sampling is B Θ =64; Input to the current agent Θ evaluation network to obtain the value function Randomly sample non-repeated experience from the experience pool U2 The batch size for each sampling is B W =128; Input to the current agent W evaluation network to obtain the value function Use policy gradient to update the evaluation network and action network parameters of the two agents respectively;

[0078] a14) Set j = j + 1. When j ≤ J, proceed to a13). When j > J, set h = h + 1 and proceed to a2).

[0079] a15) Obtain the action network and evaluation network of the trained agent Θ and agent W;

[0080] In the step a12), the advantage function of the agent Θ is calculated using the generalized advantage estimation method. and the advantage function of agent W The specific method is:

[0081]

[0082]

[0083] Where γ is the loss factor, λ is the generalized advantage estimation parameter, and r t is the agent’s reward value, Represents input s t When the parameter is The output of the evaluation network, Represents input s t+1 When the parameter is The output of the evaluation network; when it is the agent Θ, When it is agent W,

[0084] The target value of the agent Θ in step a12) and the target value of agent W The calculation method is:

[0085]

[0086] Represents input s t When the parameter is The output of the evaluation network; when it is the agent Θ, When it is agent W,

[0087] In step 3, sub-step a13), the evaluation network uses the gradient descent method to update the network parameters, and its update strategy gradient is:

[0088]

[0089] Among them, B is the agent mini-batch sampling size, Indicates the parameters Find partial derivatives and evaluate network parameters The update process is: α c is the learning rate;

[0090] The action network uses the gradient ascent method to update the network parameters. The specific strategy gradient is:

[0091]

[0092] in, is the ratio of strategy probabilities, π θ (a t |s t ) is an action network using parameter θ, state s t Downsampling to obtain action a t The probability of To use the parameter θ old When the action network is , the state s t Downsampling to obtain action a t The probability of clip(·) is the clipping function, which is to eliminate ρ t The excitation in (θ) is outside the interval [1-ε,1+ε], ε is a hyperparameter that controls the clipping range, H(π(·|s t )) is the policy entropy, c is a constant coefficient, Indicates the partial derivative of the variable θ, θ is the action network parameter, and its update process is: θ = θ + α a Δθ, α a is the learning rate; when the evaluation network and action network are updated, when it is the agent Θ, B=B Θ , θ=θ Θ , When it is agent W, B=B W , θ=θ W ,

[0093] Step 4: The action networks of the two agents Θ and W trained in step 3 are used online. The agent Θ observes the current state. Input to its action network and output the mean value u of the network Θ Multiply by π and then clip to the range [-π,π] as the action represents the phase shift of the smart reflective surface element, and then the phase shift Mapped to smart reflective surface reflection coefficient matrix During the channel statistical coherence time from the current moment to the next moment, observe the state of the agent W in the current time slot Input to its action network and output the mean value u of the network W As an action Construct W by Re(W) and Im(W), and then form the base station beamforming matrix Finally, the optimal base station active beamforming matrix and smart reflective surface passive beamforming matrix are obtained.

[0094] In summary, the present invention utilizes the powerful nonlinear modeling capabilities of deep neural networks to quickly learn the optimal base station beamforming matrix and intelligent metasurface beamforming matrix, ultimately achieving performance close to the optimal traditional algorithm with extremely short time loss.

[0095] Anything not described in detail in the present invention is well known to those skilled in the art.

[0096] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A dual-time-scale intelligent reflecting surface assisted system beamforming method, characterized in that: The following steps are involved: Step 1: In a single-cell downlink transmission system, the base station is configured with a uniform linear antenna array, which includes M antenna elements; the smart reflective surface is configured with a uniform planar reflective unit, including N vertical directions. y Line reflection unit, N per line in horizontal direction x Reflection units, a total of N = N x N y reflection units, whose reflection coefficient matrix is ​​expressed as a diagonal matrix in is the nth diagonal element of Θ, q n is the known reflection amplitude of the nth reflection element, θ n is the reflection phase shift of the nth reflective element; the cell contains K single-antenna users; The channel state information obtained by the base station and the smart reflective surface includes two parts: statistical channel state information and instantaneous channel state information; the statistical channel state information includes: the line-of-sight channel matrix from the base station to the smart reflective surface Line-of-sight channel vector from the smart reflecting surface to the kth user Line-of-sight channel vector from the base station to the kth user The instantaneous channel state information is the equivalent channel vector from the base station to each user k in is the instantaneous channel vector from the smart reflection surface to the kth user, is the instantaneous channel matrix from the base station to the smart reflective surface, is the instantaneous channel vector from the intelligent reflecting surface to the kth user; the base station uses the beamforming matrix For signal transmission, each column represents the beamforming vector for transmitting signals to user k; Step 2: Construct agent Θ and agent W; the agent Θ takes the real and imaginary parts of the current channel state information as its state, which is expressed as Where Re(·) represents the real part, Im(·) represents the imaginary part, and the phase shift of the smart reflective surface at the current moment is taken as the action, which is expressed as The channel statistics at the current moment within the coherence time T H The average value of the users and rates in the time slot is the reward, which is expressed as in is the user sum rate under the time slot t' when the base station adopts the beamforming matrix W and the reflection coefficient matrix of the smart reflection surface corresponding to the current action, is the additive noise power of user k; the agent W takes the real and imaginary parts of the instantaneous equivalent channel vector of the current time slot as the state, which is expressed as Taking the real and imaginary parts of the base station beamforming matrix of the current time slot as the action, it is expressed as The current time slot user and rate are used as rewards, which can be expressed as Construct the internal modules of the intelligent agent Θ, including: evaluation network, action network, experience pool U1, optimizer, where the evaluation network parameters are The action network parameter is θ Θ ; Construct the internal modules of the intelligent agent W, including: evaluation network, action network, experience pool U2, optimizer, where the evaluation network parameters are The action network parameter is θ W ; The action network output mean u of the agent Θ Θ and standard deviation σ Θ , and then the mean is u Θ And the standard deviation is σ Θ The normal distribution of the sample is sampled, the sampled value is multiplied by π, and then clipped to the range [-π, π] as the action of the agent Θ; the action network output mean u of the agent W W and standard deviation σ W , and then the mean is u W And the standard deviation is σ W The normal distribution of is sampled, and the sampled value is used as the action of the intelligent agent W; The evaluation network of the agent Θ outputs a value function based on the state to evaluate the quality of the action selected by the agent Θ; the evaluation network of the agent W outputs a value function based on the state to evaluate the quality of the action selected by the agent W; The experience pool U1 of the agent Θ is responsible for storing samples generated during the learning process of the agent Θ, and the experience pool U2 of the agent W is responsible for storing samples generated during the learning process of the agent W; The optimizer of the agent Θ is responsible for training the action network and evaluation network parameters of the agent Θ, and the optimizer of the agent W is responsible for training the action network and evaluation network parameters of the agent W; Step 3: Train the evaluation network and action network of the two agents Θ and W constructed in step 2 to obtain the trained action network and evaluation network of the two agents Θ and W; Step 4: The action networks of the two agents Θ and W trained in step 3 are used online. The agent Θ observes the current state s t Θ , input to its action network, and the mean value u output by the network Θ Multiply by π and then clip to the range [-π,π] as the action represents the phase shift of the smart reflective surface element, and then the phase shift Mapped to smart reflective surface reflection coefficient matrix During the channel statistical coherence time from the current moment to the next moment, observe the state of the agent W in the current time slot Input to its action network and output the mean value u of the network W As an action Construct W by Re(W) and Im(W), and then form the base station beamforming matrix Finally, the optimal base station active beamforming matrix and smart reflective surface passive beamforming matrix are obtained.

2. The dual-time-scale intelligent reflective surface-assisted system beamforming method according to claim 1, characterized in that: In step 2, the evaluation network and action network of the two agents Θ and W each include an input layer, a hidden layer, and an output layer.

3. The dual-time-scale intelligent reflective surface-assisted system beamforming method according to claim 1, characterized in that: The training in step 3 specifically includes the following sub-steps: a1) Randomly initialize the action network parameters θ of the agent Θ Θ and evaluate network parameters Their initial values ​​are recorded as and Randomly initialize the action network parameters θ of intelligent W W and evaluate network parameters Their initial values ​​are recorded as and Initialize the base station transmitter power to P, the number of statistical channel state information training is 2T, and the number of time slots within the channel statistical coherence time is T H , the maximum number of iterations is H, and the current number of iterations h is 1; a2) When h≤H, set the initial time t=1, initialize the experience pool U1 and the experience pool U2, and go to step a3); otherwise, go to step a15); a3) At time t, agent Θ observes the current statistical channel state information Set the state of the agent Θ at the current moment to Agent Θ will Input to its evaluation network, output Represents input When the parameter is The output of the evaluation network; Input to its action network, output mean u Θ and standard deviation σ Θ , the mean value is u Θ And the standard deviation is σ Θ The sampled value is multiplied by π and then clipped to the range [-π, π] as the action Record the agent Θ in state Downsampling Probability action As the phase shift of N reflection units of the smart reflection surface, the reflection coefficient matrix of the smart reflection surface can be further formed: And configure it; a4) The channel statistical coherence time from time t to time t+1 contains T H Time slots, set the initialization time slot t'=1; a5) At time slot t', agent W observes the instantaneous equivalent channel vector h from the base station to each user k in the current time slot. k , set the current time slot agent W state to Agent W will Input to its action network, output mean u W and standard deviation σ W , the mean u W As an action Construct W by Re(W) and Im(W), and then form the base station beamforming matrix of the current time slot as According to the instantaneous equivalent channel vector h of the current time slot k The instantaneous sum rate of the current time slot is calculated with the base station beamforming matrix W as a6) If t'<T H , let t'=t'+1 and go to step a5); otherwise, for t'=1,...,T H T H time slots Calculate the average This is used as the reward value of the agent Θ at the current moment; Get the state of the agent Θ at the next moment Will It is stored in the experience pool U1 as an experience; a7) Let t = t + 1. When t ≤ T, proceed to step a3) again. When t > T, proceed to step a8); a8) At time t, the agent Θ observes the current statistical channel state information Set the state of the agent Θ at the current moment to Agent Θ will Input to its action network, output mean u Θ and standard deviation σ Θ , the mean value u of the network output Θ Multiply by π and then clip to the range [-π,π] as the action represents the phase shift of the smart reflective surface element, and then Mapped to smart reflective surface reflection coefficient matrix Initialization time slot t'=1; a9) At time slot t', agent W observes the instantaneous equivalent channel vector h from the base station to each user k in the current time slot. k , set the current time slot agent W state to Agent W will Input to its evaluation network and output value function Represents input status When the parameter is The output of the evaluation network will be Input to its action network, output mean u W and standard deviation σ W , the mean value is u W The standard deviation is σ W Sampling from the normal distribution of And record the agent W in state Downsampling acquisition Probability Get Action After that, W is constructed by Re(W) and Im(W), and then the base station beamforming matrix of the current time slot is reconstructed as According to the instantaneous equivalent channel vector h of the current time slot k The instantaneous sum rate of the current time slot is calculated with the base station beamforming matrix W as Will As the reward value of the current time slot agent W; get the observation state of the next time slot agent W Will It is stored in the experience pool U2 as an experience; a10) If t'<T H , set t'=t'+1 and go to step a9); otherwise, set t=t+1 and go to a11); a11) If t≤2T, proceed to step a8); otherwise, proceed to step a12); a12) Using the experience samples in the experience pool U1, calculate the advantage function of the agent Θ and target value Using the experience samples in the experience pool U2, calculate the advantage function of the agent W and target value Set the agent Θ mini-batch sampling size to B Θ , set the mini-batch sampling size of agent W to B W , set the total number of training times in one iteration to J, and the initial number of training times to j = 1; a13) Randomly sample non-repeated experience from the experience pool U1 The batch size for each sampling is B Θ ;Will Input to the current agent Θ evaluation network to obtain the value function Randomly sample non-repeated experience from the experience pool U2 The batch size for each sampling is B W ;Will Input to the current agent W evaluation network to obtain the value function Use policy gradient to update the evaluation network and action network parameters of the two agents respectively; a14) Set j = j + 1. When j ≤ J, proceed to a13). When j > J, set h = h + 1 and proceed to a2). a15) Obtain the action network and evaluation network of the trained agents Θ and W.

4. The dual-time-scale intelligent reflective surface-assisted system beamforming method according to claim 3, characterized in that: In step a12), the advantage function of the agent Θ is calculated using the generalized advantage estimation method. and the advantage function of agent W The specific method is: Where γ is the loss factor, λ is the generalized advantage estimation parameter, and r t is the agent’s reward value, Represents input s t When the parameter is The output of the evaluation network, Represents input s t+1 When the parameter is The output of the evaluation network; The target value of the agent Θ in step a12) and the target value of agent W The calculation method is: Represents input s t When the parameter is The output of the evaluation network; When the above advantage function and target value calculation method are used, when it is an intelligent agent Θ, r t =r t Θ When it is agent W, r t =r t W .

5. The dual-time-scale intelligent reflective surface-assisted system beamforming method according to claim 3, characterized in that: In step a13), the evaluation network uses the gradient descent method to update the network parameters, and the update strategy gradient is: Among them, B is the agent mini-batch sampling size, Indicates the parameters Find partial derivatives and evaluate network parameters The update process is: α c is the learning rate; when evaluating network updates, when it is the agent Θ, B=B Θ , θ=θ Θ , r t =r t Θ , When it is agent W, B=B W , θ=θ W , r t =r t W , 6. The dual-time-scale intelligent reflective surface-assisted system beamforming method according to claim 3, characterized in that: In step a13), the action network uses the gradient ascent method to update the network parameters, and the specific policy gradient is: in, is the ratio of strategy probabilities, π θ (a t |s t ) is an action network using parameter θ, state s t Downsampling to obtain action a t The probability of To use the parameter θ old When the action network is , the state s t Downsampling to obtain action a t The probability of clip(·) is the clipping function, which is to eliminate ρ t The excitation in (θ) is outside the interval [1-ε,1+ε], ε is a hyperparameter that controls the clipping range, H(π(·|s t )) is the policy entropy, c is a constant coefficient, Indicates the partial derivative of the variable θ, θ is the action network parameter, and its update process is: θ = θ + α a Δθ, α a is the learning rate; When the action network is updated, when it is the agent Θ, B=B Θ , θ=θ Θ , r t =r t Θ , When it is agent W, B=B W , θ=θ W , r t =r t W ,

Citation Information

Patent Citations

  • Intelligent reflection surface phase optimization method based on deep reinforcement learning

    CN111181618A

  • Multi-RIS communication network rate improving method based on MADDPG

    CN114727318A