RIS auxiliary enhanced communication method based on evolution guidance strategy gradient
By introducing an evolutionary guidance strategy gradient method in the RIS auxiliary communication system, combined with TD3 and CEM algorithms, the joint optimization of the transmitter precoding matrix and RIS reflective phase shift matrix is realized, solving the problems of high computational complexity of traditional optimization algorithms and poor exploration of DRL algorithms, and improving the spectrum efficiency and convergence speed of the system.
Patent Information
- Application Number
- CN202510281547.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In RIS auxiliary communication systems, traditional optimization algorithms have high computational complexity, poor DRL algorithms have poor sample efficiency, and low sample efficiency of EL algorithms, making it difficult to effectively optimize the sending terminal precoding matrix and RIS reflective phase shift matrix, which in turn affects the efficiency of beamforming.
A method based on evolution-guided strategy gradient (EGPG) is proposed, combining the dual-delay depth deterministic strategy gradient algorithm (TD3) and the cross-entropy method (CEM) to realize the joint optimization of the sending terminal precoding matrix and the RIS reflective phase shift matrix in the RIS auxiliary communication system.
Through the EGPG algorithm, the joint design of active/passive beamforming of RIS-assisted communication system can be realized under low computing complexity, which improves the spectrum efficiency and convergence speed of the system.
Smart Images

Figure CN120223127A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a reflection phase shift optimization algorithm for a reconfigurable intelligent surface (RIS) in the field of wireless communication. Specifically, it is a beamforming method for jointly optimizing the transmit precoding matrix and the RIS reflection phase shift matrix based on the Evolution Guided Policy Gradient (EGPG). Background Art
[0002] The RIS is a two-dimensional plane composed of reconfigurable electromagnetic reflection units, which can control the wireless channel with relatively low hardware cost and system power consumption, significantly improving the communication system capacity, efficiency, and coverage. Different from traditional technical means that need to passively adapt to the wireless propagation channel, the communication link assisted by the RIS can change the wireless channel characteristics, making the wireless environment a part of the communication system design parameters, thus bringing an increase in the upper limit of the wireless channel capacity.
[0003] In the RIS-assisted enhanced communication scenario, due to the unit modulus constraint of the reflection unit and its coupling with the transmit precoding, the traditional convex optimization algorithm is not applicable, and the beamforming design is relatively difficult. The existing solutions mainly convert it into a convex optimization problem or use the method of alternating optimization for solution, such as semidefinite relaxation, Riemannian conjugate gradient, fractional programming, etc. However, this optimization method usually requires a high computational complexity, which is not feasible for endpoints with high timeliness requirements and insufficient hardware computing resources. In addition, in the case of multiple transmit antennas at the transmitter, it is necessary to jointly optimize the active transmit beamforming at the base station and the passive reflection beamforming at the RIS. The traditional optimization algorithm usually uses the method of alternating optimization to separately perform iterative optimization on the two, and the solution result depends on the selection of the initial value, and the computational complexity increases sharply with the complexity of the communication system, resulting in low efficiency for large-scale systems.
[0004] To solve the complex non - linear non - convex optimization problems in communication systems, machine learning (ML) algorithms are widely used. Based on deep learning (DL), a mapping relationship can be established between channel information and precoding design, so as to obtain the beamforming matrix of a multiple - input multiple - output (MIMO) system. However, DL - based methods require a large number of effective data sets to be obtained before training, and high - quality labels in the data sets are usually very difficult to obtain. Different from DL, deep reinforcement learning (DRL) is an online training algorithm that does not require labels to be obtained in advance. It mainly obtains empirical training data through the interaction between the agent and the environment, and continuously iteratively improves the agent during the interaction process. However, DRL needs to balance the exploration of the environment and the utilization of information. Researchers have proposed a series of improved algorithms based on state pseudo - counting, curiosity prediction models, noise networks, etc. to enhance the exploration ability of DRL. However, the problem - solving ability of these methods is limited, and some algorithms have large overheads. On the other hand, DRL algorithms are very sensitive to hyperparameters and require careful hyperparameter tuning in practical applications. Evolutionary learning (EL) is a heuristic algorithm that obtains the optimal solution by performing operations such as recombination and mutation on individuals through the idea of population iteration. However, EL is essentially an optimization algorithm based on Monte Carlo estimation, without a specific evolutionary direction guidance, and the sample efficiency of the population is low, only applicable to low - dimensional spaces. Summary of the Invention
[0005] Aiming at the problems of high computational complexity of traditional optimization algorithms, poor exploration ability of DRL algorithms, and low sample efficiency of EL algorithms in RIS - assisted communication systems, the present invention proposes a RIS - assisted enhanced communication method based on evolutionary guidance policy gradient (EGPG). The EGPG algorithm proposed by the present invention is based on the twin - delayed deep deterministic policy gradient algorithm (TD3) and the cross - entropy method (CEM) algorithm, combining the strong exploration ability of EA, the gradient characteristics of RL, and the non - linear fitting characteristics of DL, and can realize the joint design of active / passive beamforming in RIS - assisted communication systems.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A RIS - assisted enhanced communication method based on evolutionary guidance policy gradient, where the scenario of RIS - assisted enhanced communication includes RIS, base station, and user, and includes the following processes:
[0008] Step 1, initialize the TD3 value networks Q(s,a;ω1) and Q(s,a;ω2), and the TD3 policy network π(a|s;θ), where a is the action, s is the state, and ω1, ω2, and θ are the randomized neural network parameters; and initialize the target network parameters and θ - = θ, the experience replay pool mini-batch size M b the mean μ, variance σ, sampling size N s the elite set size E1, the number of interactive elite individuals E2, the number of iterations T, and the policy network update frequency d;
[0009] Step 2, input the state at time t Select an action through the TD3 policy network π(a|s;θ) and add noise ε. The noise ε follows a random Gaussian distribution with a mean of zero and a variance of σ, obtaining the action a at time t t ; and use the TD3 action a t to interact with the environment to obtain the reward r t and the channel gain s at the next moment t+1 , store the interaction trajectory <s t ,a t ,r t ,s t+1 > in the experience pool B; where the action a t is the precoding matrix of the transmitter and the RIS phase shift matrix a t = {Φ t ,W t}, ω k is the transmitter beamforming vector, the reward r t is the system spectral efficiency at time t γ k represents the system signal-to-interference-plus-noise ratio, h r,k is the channel from the RIS to user k, h d,k is the channel from the base station to user k, Φ is the reflection phase shift matrix of the RIS, G is the channel from the base station to the RIS, σ k is the Gaussian white noise, and K is the number of users;
[0010] Step 3, let the mean μ at time t t = μ, the variance σ t = σ, use CEM based on the random Gaussian distribution to sample and obtain N s action samples Send the N s action samples to interact with the environment respectively to obtain each action sample The corresponding reward and the channel gain at the next moment
[0011] Step 4: Based on the reward obtained in Step 3, sort the N s sampled actions to obtain the sorted action set F, and set Sample the E1 actions with the highest reward values from the action set F in sequence to obtain the elite action set F e1 = F(1:E1);
[0012] Step 5: Based on the elite action set F obtained in Step 4 e1 , update the mean and variance to obtain the mean at time t+1 and the variance at time t+1 where λ i is the weight evaluated based on the action sample a i , Z t is a constant vector related to t;
[0013] Step 6: If r max > r t , then sample the E2 action samples with the highest reward values in the elite action set in sequence to obtain the interactive elite action set F e2 = F(1:E2), and store the interactive elite action trajectory in the experience pool ; otherwise, directly go to Step 7;
[0014] Step 7: Randomly sample M trajectories from b and calculate where γ is the discounted return rate, ε follows a truncated Gaussian distribution with a mean of 0 and a variance of σ' and a truncation interval of [-c, c] c is the truncation interval bound;
[0015] Step 8: Update the value network parameters
[0016] Step 9: If t mod d == 0, then go to Step 10, otherwise go to Step 13;
[0017] Step 10: If r max > r t , then go to Step 11, otherwise go to Step 12;
[0018] Step 11: Update the policy network loss function:
[0019] Denotes the sum of the Euclidean distances between the actions generated by the TD3 policy network and the interaction elite actions obtained by CEM sampling, where a e2 Denotes the interaction elite action, β denotes the distance weight; update the policy network: θ new = θ now + α·▽ θ L'(θ, a e2 , β), where α is the policy network update learning rate; execute step 13;
[0020] Step 12: Update the policy network loss function: L(θ) = E S [-Q(s, π(a|s; θ))], update the policy network: θ new = θ now + α·▽ θ L(θ), where α is the policy network update learning rate;
[0021] Step 13: Update the value and policy target networks: where τ is the target network update learning rate;
[0022] Step 14: t = t + 1, if t = T, then output the policy network action a T = {Φ T , W T}, otherwise jump to step 2;
[0023] Step 15: Based on the action a T = {Φ T , W T} obtained at time T, obtain the transmitter precoding matrix and the RIS phase shift matrix, and calculate the system spectral efficiency
[0024] The present invention has the following advantages:
[0025] 1. The RIS-assisted enhanced communication method based on evolutionary guidance policy gradient proposed by the present invention combines the strong exploration of EL, the gradient characteristics of RL, and the non-linear fitting characteristics of DL. Without complex formula derivation and alternating iterative optimization, it can simultaneously achieve the joint optimization of the transmitter precoding and the RIS phase shift matrix in the RIS-assisted communication system.
[0026] 2. The RIS-assisted enhanced communication method based on evolutionary guidance policy gradient proposed by the present invention, aiming at the characteristic that EL cannot handle high-dimensional spaces, converts the optimization target from the high-dimensional policy network to the low-dimensional action space. EL can directly optimize the actions through population iteration, thus providing sample diversity for DRL.
[0027] 3. The RIS-assisted enhanced communication method based on the evolutionary guidance policy gradient proposed by the present invention uses an elite interaction strategy to positively optimize DRL in two aspects: experience pool filling and policy network-guided optimization, in order to prevent the poor experience generated during the evolution of EL from having an adverse impact on the DRL gradient optimization process. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 FIG. is a block diagram of the communication system of the present invention, a single RIS-assisted downlink single-input multi-output communication link, where the direct link is blocked and communication relies only on the RIS-based reflection link.
[0029] Figure 2 FIG. is a block diagram of the principle of the method of the present invention.
[0030] Figure 3 FIG. is a simulation convergence curve graph comparing the method of the present invention with a benchmark algorithm. DETAILED DESCRIPTION OF THE INVENTION
[0031] The present invention will be described in detail below with reference to the accompanying drawings.
[0032] The system model of the method of the present invention is as shown in Figure 1 FIG.. The RIS-assisted downlink communication consists of an RIS composed of N = 4 reflection units, the number of transmitting antennas at the base station is M = 4, and there are a total of K = 4 single-antenna users. G, and respectively represent the channels from the base station to the RIS, from the RIS to user k, and from the base station to user k. The complex baseband transmission signal transmitted at the base station can be expressed as where s k , k = 1,..., K are the transmission symbols after energy normalization for the base station to transmit to the k-th user, and s k obeys a Gaussian distribution with zero mean and unit variance, and ω k is the transmit beamforming vector. The received signal of user k can be expressed as:
[0033]
[0034] where Φ = diag(θ1,..., θ N ) represents the reflection phase shift matrix of the RIS, β n ∈[0,1] and respectively represent the reflection amplitude and reflection phase shift of the n-th electromagnetic reflection unit of the RIS, n k represents the noise of the k-th user, which obeys a Gaussian distribution with a mean of 0 and a variance of σ k . In the RIS-assisted downlink communication system, the state s is defined as the channel gain The action a is the transmit - end precoding matrix and the RIS phase - shift matrix a = {Φ, W}, where The reward r is the system spectral efficiency where γ k is the signal - to - interference - plus - noise ratio, and there is
[0035] As Figure 2 shown, a RIS - assisted enhanced communication method based on evolutionary - guidance policy gradient includes the following processes:
[0036] Step 1: Initialize the TD3 value networks Q(s, a; ω1) and Q(s, a; ω2), and the TD3 policy network π(a|s; θ), where a is the action, s is the state, and ω1, ω2, and θ are randomized neural - network parameters; and initialize the target - network parameters and θ - = θ, the experience replay pool B, the mini - batch size M b 、the mean μ, the variance σ, the sampling size N s 、the elite - set size E1, the number of interactive elite individuals E2, the number of iterations T, and the policy - network update frequency d.
[0037] Step 2: Input the state at time t Select an action through the TD3 policy network π(a|s; θ) and add noise ε, where the noise ε follows a random Gaussian distribution with a mean of zero and a variance of σ, to obtain the action at time t a t ; and use the TD3 action a t to interact with the environment to obtain the reward r t and the channel gain s at the next time t+1 , and store the interaction trajectory <s t , a t , r t , s t+1 > in the experience pool B; where the action a t is the transmit - end precoding matrix and the RIS phase - shift matrix a t = {Φ t , W t}, ω k is the transmit - end beamforming vector, the reward r t is the system spectral efficiency at time t γ k represents the system signal - to - interference - plus - noise ratio,
[0038] The channel from the base station to user k, Φ is the reflection phase - shift matrix of the RIS, G is the channel from the base station to the RIS, σ k is the Gaussian white noise, and K is the number of users.
[0039] Step 3: Let the mean at time t be μ t = μ, and the variance σ t = σ. Use Cross-Entropy Method (CEM) based on the random Gaussian distribution to sample and obtain N s action samples Interact with the environment for each of the N s action samples respectively, to obtain the corresponding reward r for each action sample t n s and the channel gain at the next time instant
[0040] Step 4: Based on the rewards obtained in Step 3, sort the N s sampled actions to obtain the sorted action set F, and set Sample the top E1 actions with the highest reward values from the action set F in sequence to obtain the elite action set F e1 = F(1:E1).
[0041] Step 5: Based on the elite action set F e1 obtained in Step 4, update the mean and variance to obtain the mean at time t+1 and the variance at time t+1 where λ i is the weight evaluated based on the action sample a i , and Z t is a constant vector related to t.
[0042] Step 6: If r max > r t , then sample the top E2 action samples with the highest reward values from the elite action set in sequence to obtain the interactive elite action set F e2 = F(1:E2), and store the interactive elite action trajectory in the experience pool B; otherwise, directly go to Step 7.
[0043] Step 7: Randomly sample M b trajectories from B, and calculate where γ is the discounted return rate, ε follows a truncated Gaussian distribution with mean 0 and variance σ' and truncation interval [-c, c] and c is the truncation interval bound.
[0044] Step 8: Update the value network parameters
[0045] Step 9: If t mod d == 0, go to Step 10; otherwise, go to Step 13.
[0046] Step 10: If rmax > r t , go to step 11, otherwise go to step 12.
[0047] Step 11: Update the policy network loss function: represents the sum of the Euclidean distances between the actions generated by the TD3 policy network and the interaction elite actions obtained by CEM sampling, where a e2 represents the interaction elite action, β represents the distance weight; update the policy network: θ new = θ now + α·▽ θ L'(θ, a e2 , β), where α is the policy network update learning rate; execute step 13.
[0048] Step 12: Update the policy network loss function: L(θ) = E S [-Q(s, π(a|s; θ))], update the policy network: θ new = θ now + α·▽ θ L(θ), where α is the policy network update learning rate.
[0049] Step 13: Update the value and policy target networks: where τ is the target network update learning rate.
[0050] Step 14: t = t + 1, if t = T, then output the policy network action a at time T T = {Φ T , W T}, otherwise jump to step 2.
[0051] Step 15: Based on the action a obtained at time T T = {Φ T , W T}, obtain the transmitter precoding matrix and the RIS phase shift matrix, and the system spectral efficiency can be obtained through calculation
[0052] Principle description:
[0053] The present invention utilizes the optimization framework of evolutionary reinforcement learning, embeds the evolutionary algorithm into the deep reinforcement learning, and optimizes the actions of the agent through population iteration. At the same time, the evaluated elite individual experience is injected into the deep reinforcement learning experience pool to increase sample diversity, and the interaction elite individuals are used to guide the update of the deep reinforcement learning policy network.
[0054] Figure 2It is the optimized flow block diagram of the proposed EGPG algorithm. It can be seen that TD3 is the DRL optimization part of the EGPG algorithm, and CEM is the EL part of the EGPG algorithm. CEM optimizes TD3 through two parts: Experience guide and Policy guide.
[0055] Figure 3 The convergence of the proposed EGPG algorithm is analyzed and compared with benchmark algorithms (DDPG, TD3, ERL). It can be seen that the convergence of gradient-based DRL algorithms (DDGP, TD3) is relatively gentle. When iterating to 50,000 time steps, the convergence value is 1.2 bps / Hz. The ERL algorithm that combines DDPG and the evolutionary algorithm (EA) has more exploration compared to gradient-based DRL algorithms, and the convergence value is 1.38 bps / Hz at 50,000 time steps. The proposed EGPG algorithm of the present invention reaches convergence at the 11,000th time step, with a faster convergence speed and a better convergence value compared to the benchmark algorithms (about 1.58 bps / Hz at the 50,000th time step).
[0056] In summary, the present invention jointly optimizes the transmitter precoding and the RIS reflection phase shift matrix in the RIS-assisted communication system through the proposed Evolutionary Guided Policy Gradient (EGPG) algorithm. Combining the strong fitting characteristics of deep learning, the gradient characteristics of reinforcement learning, and the strong exploration of evolutionary learning, and adopting the idea of evolutionary learning, it iteratively optimizes the agent's actions to increase the sample diversity for deep reinforcement learning and guides the optimization of the policy network of deep reinforcement learning. It has a simple and clear optimization framework and low computational complexity compared to traditional optimization algorithms, and at the same time has faster convergence and better system performance compared to gradient-based deep reinforcement learning algorithms and evolutionary reinforcement learning algorithms.
Claims
1. A RIS-assisted enhanced communication method based on evolutionary guided policy gradient, wherein the scenario of RIS-assisted enhanced communication includes RIS, base station and user, and is characterized in that: The process includes: Step 1: Initialize the TD3 value networks Q(s,a;ω1) and Q(s,a;ω2), and the TD3 policy network π(a|s;θ), where a is the action, s is the state, ω1, ω2 and θ are randomized neural network parameters; and initialize the target network parameters and θ - =θ, experience replay pool Mini-batch size M b , mean μ, variance σ, sampling size N s , elite set size E1, number of interacting elite individuals E2, number of iterations T, and strategy network update frequency d; Step 2: Enter the state at time t The action a at time t is obtained by selecting an action through the TD3 policy network π(a|s;θ) and superimposing noise ε, where the noise ε follows a random Gaussian distribution with a mean of zero and a variance of σ. t ; and use TD3 action a t Interact with the environment to get rewards r t and the channel gain s at the next moment t+1 , store the interaction trajectory t ,a t ,r t ,s t+1 >In the experience pool Among them, action a t is the transmitter precoding matrix and RIS phase shift matrix a at time t t ={Φ t ,W t }, ω k is the transmitting end beamforming vector, and the reward r t is the system spectrum efficiency at time t γ k represents the system signal-to-interference-noise ratio, h r,k is the channel from RIS to user k, h d,k is the channel from the base station to user k, Φ is the reflection phase shift matrix of RIS, G is the channel from the base station to RIS, σ k is Gaussian white noise, K is the number of users; Step 3: Let the mean μ at time t be t =μ, variance σ t =σ, using a random Gaussian distribution The CEM is sampled to obtain N s Action samples N s Action samples interact with the environment respectively, and each action sample is obtained Corresponding rewards and the channel gain at the next moment Step 4: Based on the reward obtained in step 3, s The sampled actions are sorted to obtain the sorted action set F, and set Sequentially sample the E1 actions with the highest reward value from the action set F to obtain the elite action set F e1 =F(1:E1); Step 5: Based on the elite action set F obtained in step 4 e1 , update the mean and variance to get the mean at time t+1 and the variance at time t+1 where λ i Based on action sample a i The weight after evaluation, Z t is a constant vector related to t; Step 6: If r max >r t , then sequentially sample the action samples E2 with the highest reward value in the elite action set to obtain the interactive elite action set F e2 =F(1:E2), and store the interactive elite action trajectory in the experience pool Otherwise, go directly to step 7; Step 7: From Random sampling M b Trajectory, calculation Where γ is the discounted rate of return, ε follows a truncated Gaussian distribution with mean 0, variance σ' and truncation interval [-c, c] c is the cutoff interval limit; Step 8: Update value network parameters Step 9: If tmod d == 0, go to step 10, otherwise go to step 13; Step 10: If r max >r t , then go to step 11, otherwise go to step 12; Step 11: Update the policy network loss function: represents the sum of the Euclidean distances between the actions generated by the TD3 policy network and the interactive elite actions sampled by CEM, where a e2 represents the interactive elite action, β represents the distance weight; update the strategy network: Where α is the learning rate for updating the policy network; execute step 13; Step 12: Update the policy network loss function: L(θ) = E S [-Q(s,π(a|s;θ))], update the policy network: Where α is the learning rate for updating the policy network; Step 13: Update the value and policy target network: Where τ is the target network update learning rate; Step 14: t = t + 1, if t = T, then output the policy network action a at time T T ={Φ T ,W T }, otherwise jump to step 2; Step 15: Action a obtained based on time T T ={Φ T ,W T }Get the transmitter precoding matrix and RIS phase shift matrix, and calculate the system spectrum efficiency
Citation Information
Cited By
6G terahertz communication optimization method based on machine learning
CN121567229A
A Machine Learning-Based Optimization Method for 6G Terahertz Communication
CN121567229B