Dual-intelligent reflecting surface dynamic reflection coefficient adjustment method based on deep reinforcement learning

By using a dynamic reflection coefficient adjustment method for dual intelligent reflective surfaces based on deep reinforcement learning, constructing a Markov process and adopting a distributed proximal strategy optimization algorithm, the difficulty of controlling intelligent reflective surfaces in dynamic environments is solved, and more efficient global optimization and rapid iteration are achieved.

CN119420384BActive Publication Date: 2025-10-17ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411539491.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-10-30
Filing Date
2024-10-31
Publication Date
2025-10-17
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively controlling the phase of smart reflective surfaces in dynamic and complex environments, causing traditional optimization methods to easily fall into local optimal solutions, and it is difficult to find a global optimal solution in high-dimensional decision spaces.

Method used

A dynamic reflection coefficient adjustment method for dual intelligent reflectors based on deep reinforcement learning is adopted. By constructing a Markov process and a distributed proximal strategy optimization algorithm, distributed computing is used to accelerate the iteration rate and dynamically adjust the reflection coefficient to optimize the performance of the communication system.

Benefits of technology

The communication performance of the system in a dynamic environment is improved, the adaptability is stronger, the iteration speed is faster, the local optimal solution is avoided, and a more efficient global optimization is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119420384B_ABST
    Figure CN119420384B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on depth reinforcement learning's double intelligent reflector dynamic reflection coefficient adjustment method.Current complexity and dynamic change of environment are important factors to influence intelligent reflector phase control, in dynamic and complex environment, long-term phase control is difficult thing.The method of the present application is based on depth reinforcement learning (DRL), constantly adjust through the feedback of environmental information, reach an optimal strategy, to carry out double intelligent reflector unit phase and amplitude adjustment.The method of the present application can provide a flexible, efficient, automatic method for double intelligent reflector wireless communication system to realize reflection coefficient optimization control, especially when facing changing environment and complex system dynamics.In this way, the energy efficiency of wireless network can be improved, the coverage range is expanded and the user experience is improved.In addition, the convergence rate is improved by distributing training for the method of the present application.Simulation results show that the method of the present application is effective and feasible, and the method of the present application has a faster convergence speed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of wireless communication, intelligent reflecting surface, deep reinforcement learning and DPPO, and particularly relates to a dual-intelligent reflecting surface cooperative transmission optimization method. BACKGROUND

[0002] The emergence of intelligent metasurface (intelligent reflecting surface) technology is a revolutionary supplement to modern wireless communication systems. It uses a series of controllable passive units to intelligently adjust the propagation characteristics of electromagnetic waves. The core advantage of intelligent reflecting surface is that it can finely control the reflection direction, phase and amplitude of signals, which enables it to effectively bypass obstacles and provide alternative propagation paths to achieve more reliable signal transmission even in complex urban environments.

[0003] Current research on intelligent reflecting surfaces is based on the assumption of known prior knowledge of the environment. However, with the development of communication technology, the Internet of Things has become a highly complex system, and it is difficult to obtain this prior knowledge in actual situations. The complexity and dynamic changes of the environment are important factors affecting the phase control of intelligent reflecting surfaces. In dynamic and complex environments, long-term phase control is a difficult task. Deep reinforcement learning, as a model-free algorithm, can continuously adjust itself through feedback from environmental information to achieve an optimal strategy. In addition, the phase control of intelligent reflecting surfaces involves a large number of reflecting units, each with multiple adjustable phase angles, resulting in a very high-dimensional decision space. Traditional optimization methods are prone to local optimal solutions in high-dimensional spaces, while deep reinforcement learning can effectively explore and utilize high-dimensional spaces through policy networks and value networks to find global optimal solutions. Therefore, applying deep reinforcement learning to intelligent reflecting surfaces will have good results. SUMMARY

[0004] The purpose of the present application is to solve the problems existing in the prior art and provide a dual-intelligent reflecting surface dynamic reflection coefficient adjustment method based on deep reinforcement learning.

[0005] In order to achieve the above-mentioned application purpose, the present application specifically adopts the following technical solutions:

[0006] A dual-intelligent reflecting surface dynamic reflection coefficient adjustment method based on deep reinforcement learning includes the following steps:

[0007] S1: Construct a dual-intelligent reflecting surface assisted wireless communication system model, all channel coefficients use the Rician channel model, and construct an optimization problem with the optimization objective of maximizing the system communication rate, with the constraint condition of the phase shift and amplitude reflection coefficient of the intelligent metasurface reflection module;

[0008] S2: Based on the transmission process of the intelligent metasurface smart reflector, the phase control process of the intelligent metasurface smart reflector is established as a Markov process;

[0009] S3: Based on the proximal policy optimization algorithm, the optimization problem is solved, and the distributed proximal policy optimization algorithm is used to accelerate the iteration rate through distributed computing, and the dynamic adjustment of the reflection coefficient is realized in the optimization solving process.

[0010] On the basis of the above scheme, each step can be implemented in the following preferred specific manner.

[0011] As a preferred, in step S1, the wireless communication system model comprises a base station, two intelligent metasurfaces, K single-antenna users and a controller for adjusting the phase of the reflector, the base station is equipped with M antennas, and the number of reflection units of the reflector is N.

[0012] As a preferred, in step S1, the optimization target is specifically:

[0013]

[0014] Wherein, Θ1 represents the reflection coefficient matrix of RIS1; Θ2 represents the reflection coefficient matrix of RIS2; P sum represents the sum of data rates; r k represents the data rate of user k; K represents the number of users; SINR k represents the signal-to-interference-and-noise ratio of user k; Θ fs represents the reflection coefficient matrix of the first or second intelligent metasurface; diag(·) represents a diagonal matrix, j represents an imaginary unit, respectively represent the phase shift of the 1st,..., Nth reflection module of the first or second intelligent metasurface; respectively represent the amplitude reflection coefficient of the 1st,..., Nth reflection module of the first or second intelligent metasurface; represents the phase shift of the nth reflection module of the first or second intelligent metasurface; represents the amplitude reflection coefficient of the nth reflection module of the first or second intelligent metasurface.

[0015] As a preferred, in step S3, the four-tuple represents a Markov process, represents a state space, represents an action space, represents a state transition function, represents a reward function; wherein the state space According to the elements that may affect the optimization target in the actual communication process, the state space is constructed, and the state set at time t is denoted as st = {H, Θ1, Θ2}, H = [h1, …, hK] is the set of all users' channels, h1, …, hK represent the channels of users 1, …, K respectively; the action space K is denoted as A = {a | a = {ΔΘ1, ΔΘ2}, ΔΘ1, ΔΘ2 represent the increments of the reflection coefficient matrices Θ1, Θ2 respectively}; the reward function K is denoted as R (a) = {P - Pth} if a = {0, 0}, R (a) = {P - Pth} if a = {0, ΔΘ2}, R (a) = {P - Pth} if a = {ΔΘ1, 0}, R (a) = {P - Pth} if a = {ΔΘ1, ΔΘ2}, where P is the rate threshold. The action set at time t is denoted as a t = {ΔΘ1, ΔΘ2}, ΔΘ1, ΔΘ2 represent the increments of the reflection coefficient matrices Θ1, Θ2 respectively; the reward function is denoted as R (a) = {P - Pth} if a = {0, 0}, R (a) = {P - Pth} if a = {0, ΔΘ2}, R (a) = {P - Pth} if a = {ΔΘ1, 0}, R (a) = {P - Pth} if a = {ΔΘ1, ΔΘ2}, where P is the rate threshold.

[0016]

[0017] where P t is the rate threshold.

[0018] As a preferred, in step S4, the process of solving the optimization problem using the proximal policy optimization algorithm is as follows:

[0019] S41. In the initialization stage, initialize the experience pool and configure the capacity C, initialize the related parameters of the proximal policy optimization algorithm, and set the environment and agent of the wireless communication system;

[0020] S42. In each iteration round of the loop stage, obtain the channel estimation number m and the state iteration number n, first judge whether the channel estimation number m is 0:

[0021] If m = 0, the proximal policy optimization algorithm ends;

[0022] If m ≠ 0, continue to judge whether the state iteration number n is 0: if n = 0, let m = m - 1, and then re-judge whether the channel estimation number is 0; if n ≠ 0, initialize the state s0, use the actor network to predict the probability of executing the action a j in the state s j at time j, use the critic network to predict the evaluation value in the state s j at time j, the agent executes the action a j , obtains the reward r j at time j and the state s j+1 at time j + 1, the state s j at time j, the action a j , the reward r j and the state s j+1 at time j + 1 constitute a four-tuple (s j , a j , r j , s j+1 ) stored in the experience pool, and the state s j at time j is updated to the state s j+1 at time j + 1.If the experience pool does not collect the four-tuple with the length of C and the loop does not end, let n = n-1, and then rejudge whether the state iteration number is 0; if the experience pool collects the four-tuple with the length of C or the loop ends, the following operations are performed: calculate the advantage function G, calculate the cumulative discounted reward corresponding to each state, update the actor network and the critic network;

[0023] S43. continuously iterate until the loop ends.

[0024] As preferred, the advantage function G t at time t is calculated as follows:

[0025] G t = δ t + (γλ)δ t+1 + … + (γλ) T-t+1 δ T-1

[0026] δ t = r t + γV(s t+1 ) - V(s t )

[0027] Wherein, γ is a discount factor; λ is a parameter in generalized advantage estimation; δ t , δ t+1 , δ T-1 respectively represent intermediate variables at times t, t+1, T-1; T represents a time parameter; V(s t+1 ), V(s t ) respectively represent the value of s t+1 , s t state.

[0028] As preferred, the proximal policy optimization algorithm adopts the PPO-Clip method, and the objective function is to maximize the following objective:

[0029]

[0030] Wherein, represents the value of the objective function when the policy parameter value is ; represents the policy parameter; represents taking expectation; represents the probability ratio; clip represents adopting the PPO-clip method; ∈ represents a hyperparameter; π represents the policy; is the policy parameter before updating;

[0031] The specific way of updating the policy parameter using gradient ascent is:

[0032]

[0033] wherein, denotes the updated policy parameter; a1 denotes a learning rate; denotes gradient.

[0034] The present application has the following beneficial effects relative to the prior art:

[0035] (1) More adaptable to dynamic communication environment: As a model-free algorithm, deep reinforcement learning can continuously adjust itself through feedback from environmental information to achieve an optimal strategy, avoiding the problem of easily falling into a local optimal solution in high-dimensional space with the unified optimization method.

[0036] (2) Faster algorithm iteration speed: The present application uses the DPPO algorithm to speed up the learning process by running multiple agent instances in parallel on multiple environment replicas. This parallelization allows the system to generate more data in the same time, enabling faster policy updates. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 is a step flowchart of the method of the present application;

[0038] Figure 2 is a structural diagram of the dual-intelligent reflector-assisted wireless communication system model of the present application;

[0039] Figure 3 is a structural diagram of the actor network and critic network of the present application;

[0040] Figure 4 is a flowchart of the wireless communication system of the present application;

[0041] Figure 5 is a workflow diagram of the distributed proximal policy optimization algorithm of the present application;

[0042] Figure 6 is a sequence diagram of the model-free intelligent reflector configuration of the present application;

[0043] Figure 7 is a comparative diagram of the performance of different algorithms in the embodiment of the present application;

[0044] Figure 8 is a comparative diagram of the performance of different algorithms under different numbers of reflection units in the embodiment of the present application;

[0045] Figure 9 is a bar chart of the average rate of different algorithms in the embodiment of the present application;

[0046] Figure 10 is a result graph of the bit error rate change with SNR for the method of the present application and the DDQN comparison method;

[0047] Figure 11 Figure 2 is a schematic diagram showing the performance comparison of different algorithms when the number of users is 2 in an embodiment of the present application;

[0048] Figure 12 Figure 3 is a schematic diagram showing the performance comparison of different algorithms when the number of users is 3 in an embodiment of the present application;

[0049] Figure 13 Figure 4 is a schematic diagram showing the performance comparison of different algorithms when the number of users is 4 in an embodiment of the present application;

[0050] Figure 14 Figure 5 is a schematic diagram showing the performance comparison of different algorithms when the number of users is 5 in an embodiment of the present application;

[0051] Figure 15 Figure 6 is a flowchart of the proximal policy optimization algorithm in an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the above objectives, features and advantages of the present application more apparent, specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced in a variety of ways other than those described herein without departing from the scope of the present application, and it will be apparent to those skilled in the art that the present application can be practiced with or without these specific details. Therefore, the specific embodiments disclosed below are not intended to limit the scope of the present application, and the technical features in each of the embodiments of the present application can be combined appropriately without conflict.

[0053] In the description of the present application, it should be understood that the terms "first", "second" are only used for distinguishing description purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features.

[0054] The present application mainly proposes a dual-intelligent reflecting surface dynamic reflection coefficient adjustment method based on deep reinforcement learning. The method establishes the entire intelligent reflecting surface communication process in the form of a Markov chain, adjusts and controls the phase of the intelligent reflecting surface through deep reinforcement learning, and improves the communication performance of the entire system in a dynamic environment through multiple rounds of iteration optimization. The specific implementation process is as follows:

[0055] As shown in Figure 1 In a preferred implementation manner of the present application, the dual-intelligent reflecting surface dynamic reflection coefficient adjustment method based on deep reinforcement learning includes the following S1-S4 steps. The specific implementation process will be described below.

[0056] S1: Construct a dual-intelligent reflecting surface assisted wireless communication system model.

[0057] In step S1, as shown in Figure 2 The wireless communication system model contains one base station, two intelligent metasurfaces (intelligent reflecting surfaces), K single-antenna users, and one controller that adjusts the phase of the reflecting surface. The base station is equipped with M antennas, and the reflecting surface has N reflecting units.

[0058] For the wireless communication system, let the first intelligent metasurface be RIS1 and the second intelligent metasurface be RIS2. The channel model between user k and the base station can be represented as:

[0059]

[0060] where h k represents the channel between user k and the base station; is the channel coefficient between the base station and user k; is the channel coefficient between the base station and RIS1; is the channel coefficient between RIS1 and RIS2; is the channel coefficient between RIS2 and user k; Θ1 represents the reflection coefficient matrix of RIS1, and Θ2 represents the reflection coefficient matrix of RIS2. The reflection coefficient matrices of RIS1 and RIS2 are unified as Θ fs represents a diagonal matrix, j represents the imaginary unit, respectively represent the phase shifts of the 1st,..., Nth reflecting modules of RIS1 or RIS2. respectively represent the amplitude reflection coefficients of the 1st,..., Nth reflecting modules of RIS1 or RIS2.

[0061] Then, the base station sends signals to the user, and the signal received by user k can be represented as:

[0062]

[0063] where y CEk is the part of the signal received by user k for channel estimation; y DTk is the part of the signal received by user k for information transmission; P s is the transmission power; a k is the power allocation coefficient of user k; α represents the proportion of the part of the signal used for channel estimation, and β represents the proportion of the part of the signal used for information transmission, and α + β = 1; x k is the signal sent to user k, and represents taking the expectation. is Gaussian noise that follows a Gaussian distribution with a mean of 0 and a variance of Therefore, the signal-to-interference-and-noise ratio (SINR) of user k can be represented as: ​

[0064]

[0065] wherein SINR k denotes the signal-to-interference-and-noise ratio SINR of user k; a t denotes the power allocation factor of user t.

[0066] For a communication system, the performance metric can be the SINR, data rate, frame error rate (FER), etc. Without loss of generality, the present application adopts the sum data rate P sum as the performance metric, and the calculation of P

[0067]

[0068] wherein r k denotes the data rate of user k.

[0069] In the wireless communication system model, all channel coefficients in the present application consider the Rician channel model, and h Sk is taken as an example, then the channel coefficient between the base station and user k can be expressed as:

[0070]

[0071] wherein h Sk,LoS denotes the line-of-sight (LoS) component of h Sk , and h Sk,NLoS represents the non-line-of-sight (NLoS) component of h Sk , and both components are independent and identically distributed (i.i.d.) and subject to circularly symmetric Gaussian random variables with zero mean and unit variance. In addition, the line-of-sight component is location-dependent and thus slowly time-varying; the non-line-of-sight component is caused by multipath effects and thus rapidly time-varying. is the ratio of the power in the line-of-sight path to the power in the non-line-of-sight path.

[0072] S2: Based on the wireless communication system model assisted by double intelligent reflecting surfaces, all channel coefficients adopt the Rician channel model, and an optimization problem is constructed with the optimization objective of maximizing the system communication rate, and the constraint condition is the phase shift and amplitude reflection coefficient of the intelligent metasurface reflection module.

[0073] In step S2, the optimization objective is specifically:

[0074]

[0075] S3: Based on the transmission process of the intelligent metasurface intelligent reflecting surface, the phase control process of the intelligent metasurface intelligent reflecting surface is established as a Markov process.

[0076] It should be noted that in the present application, generally speaking, some solutions about the optimal target usually adopt the problem of transforming the problem into an alternating iterative optimization problem and using related algorithms to solve. However, in this wireless communication system model, the environment is time-varying and very complex, and the traditional model-based optimization algorithm cannot handle such optimization problems. To solve this technical problem, the present application establishes a Markov process model for the phase control process of the intelligent reflecting surface based on the traditional intelligent reflecting surface transmission process. The whole regulation and control process of the controller for controlling the phase of the intelligent reflecting surface does not affect the wireless communication system, and the wireless communication system has no other connection with the controller except providing channel estimation information and phase information. The advantage of this design is that the controller can be deployed in other communication systems without modifying any content of the program. In deep reinforcement learning, an agent is an entity that can perceive the environment, make decisions, and take actions. It learns through interaction with the environment and gradually masters the optimal behavior strategy to achieve a specific goal. The agent of the present application is an intelligent reflecting surface control device, which can autonomously interact with the environment through the intelligent reflecting surface to meet the design goal.

[0077] For S3, the established Markov process is as follows:

[0078] A Markov process can be represented by a four-tuple , where represents the state space, represents the action space, represents the state transition function, represents the reward function, and the state transition function is continuously optimized in training. In order to apply reinforcement learning to intelligent reflecting surface configuration, the present application first models the intelligent reflecting surface assisted wireless communication as a Markov process. Specifically, the above state space is constructed according to the elements that may affect the optimization target in the actual communication process. In the present application, the state set at time t in the state space is denoted as s t = {H, Θ1, Θ2}, H = [h1, …, h K ] is the set of all user channels, and h1, …, h K represent the channels of users 1, …, K, respectively. Considering that the controller needs to regulate the phases of two intelligent reflecting surfaces in the wireless communication system, the action set at time t in the action space is denoted as a t= { ΔΘ1, ΔΘ2}, ΔΘ1, ΔΘ2 respectively represent the increments of the reflection coefficient matrices Θ1, Θ2. It should be noted that the reflection coefficient matrix has a range, therefore, the increment also has a range limit. When the state s t The action a t at time t is performed, the wireless communication system will get a reward r t at time t, and the wireless communication system moves to the state s t+1 .

[0079] The reward function has the following form:

[0080]

[0081] Where P t is the rate threshold.

[0082] When the sum data rate P sum is less than the rate threshold P t , the present application considers that this action has a greater impact on the performance of the wireless communication system, and the agent needs to be punished. The reason for setting it to 200 is to hope that the punishment will not be diluted by the daily reward. When the sum data rate P sum is greater than the rate threshold P t , the reward function is set to hope that the agent cannot always be still, nor can it always be moving to “brush points”, but strive to obtain a higher sum data rate P sum .

[0083] S4: The optimization problem is solved based on a proximal policy optimization algorithm (PPO), and a distributed proximal policy optimization algorithm (DPPO) is used to speed up the iteration rate through distributed computing, and the dynamic adjustment of the reflection coefficient is realized in the optimization solving process.

[0084] It should be noted that in the present application, the proximal policy optimization algorithm PPO is an algorithm widely used in the field of reinforcement learning, which was proposed by Schulman et al. in 2017. The proximal policy optimization algorithm PPO aims to solve some problems of the policy gradient method, such as sample efficiency and ease of implementation. Its core idea is to limit the step of policy change when updating the policy, in order to avoid performance collapse caused by too large update. The proximal policy optimization algorithm PPO tries to solve these problems by limiting the difference between the new policy and the old policy, and the algorithm introduces a clip function to prevent the new policy from deviating too far from the old policy. Therefore, the calculation of the advantage function G t at time t is as follows:

[0085] G t = δ t + (γλ) δ t+1 + … + (γλ) T-t+1 δ T-1

[0086] δ t = r t + γV(s t+1 ) - V(s t )

[0087] where γ is the discount factor; λ is the parameter in Generalized Advantage Estimation (GAE); δ t , δ t+1 , δ T-1 represent the intermediate variables at time t, t+1, T-1 respectively; T represents the time parameter; V(s t+1 ), V(s t ) represent the value of s t+1 , s t state respectively.

[0088] The advantage function tells the invention which actions are better than average, so they should be chosen more frequently when learning. By applying gradient ascent to those actions with positive advantage values, the policy can be biased towards better actions.

[0089] In the proximal policy optimization algorithm PPO, by limiting the step of policy update, it is ensured that the new policy will not deviate too far from the old policy. The use of the advantage function is consistent with this goal, as it promotes the stability of the update by emphasizing actions that are better than average, rather than all possible actions. The advantage function can encourage the proximal policy optimization algorithm to explore actions that may not be the current optimal but have potential. If the advantage value of an action is positive, even if it is not the optimal action, the proximal policy optimization algorithm will increase the probability of selecting it, thus achieving effective exploration.

[0090] The definition of the objective function L is also important, the proximal policy optimization algorithm PPO defines a special objective function, according to the difference of the optimization method adopted, it can be divided into two kinds: one is the objective function of PPO-Clip (Proximal Policy Optimization with Clip) method, and the other is the objective function of PPO-Penalty method. The invention adopts PPO-Clip method, for PPO-Clip, the objective function is to maximize the following objective:

[0091]

[0092] wherein, denotes the value of the objective function when the policy parameter takes the value ; denotes the policy parameter; denotes taking the expectation; denotes the probability ratio; clip denotes using the PPO-clip method; ∈ denotes a hyperparameter, in the present embodiment, the hyperparameter ∈ is set to be between 0.1 and 0.3; π denotes the policy; is the policy parameter before update.

[0093] PPO-Clip uses a clip function to prevent the policy update from deviating too far, which is achieved by limiting the policy ratio r t (θ) to the interval [1-∈, 1+∈], as described above. This clip function ensures that the policy update is not too large, even when the advantage function G t is very large. Finally, the setting of the policy update function, using gradient ascent to update the policy parameter:

[0094]

[0095] wherein, denotes the policy parameter after update; α1 denotes the learning rate; denotes taking the gradient.

[0096] The proximal policy optimization algorithm allows small-scale policy updates in this way, avoiding the problem of performance instability that may be caused by large-scale updates. In practice, the proximal policy optimization algorithm PPO has been proven to achieve good results on a variety of tasks, so it is one of the most popular reinforcement learning algorithms currently.

[0097] In addition, the proximal policy optimization algorithm is a reinforcement learning algorithm that uses an Actor-Critic architecture. In the Actor-Critic method, there are two main networks, the actor network Actor and the critic network Critic. The Actor network is responsible for learning the policy, i.e., given a state, deciding what action to take. In PPO, this is usually implemented through a neural network that outputs the probability of each possible action (in a discrete action space) or the parameters of the action (in a continuous action space) given a state. The Critic network evaluates the goodness of the current policy, usually by estimating the state value function V(s), i.e., the expected return that can be obtained from the current state by following the current policy.

[0098] The proximal policy optimization algorithm PPO optimizes the policy through an actor network, and uses a critic network to reduce the variance of the policy gradient estimate, which can accelerate the learning process and improve the stability of learning. The core of the proximal policy optimization algorithm is to limit the step size of policy update, to ensure that the updated policy does not deviate too far from the original policy, so as to avoid the performance of the training process from fluctuating sharply. For the dual-intelligent reflecting surface assisted wireless communication system, the actor network and the critic network are designed as shown in Figure 3 .

[0099] After introducing the proximal policy optimization algorithm and the actor network and the critic network, the present application will introduce the basic framework of the whole dual-intelligent reflecting surface assisted wireless communication system. As shown in Figure 4 , the optimization goal of the present application is to solve the optimization problem in step S2. First, at time t, the state of the environment is s t , in order to explore the optimal policy under the state s t , the actor network is used for prediction to obtain the mean and variance of the distribution of the reflection coefficient increment:

[0100]

[0101] Wherein, μ1(t), μ2(t) represent the mean of the RIS1, RIS2 reflection coefficient increment distribution at time t; σ1(t), σ2(t) represent the variance of the RIS1, RIS2 reflection coefficient increment distribution at time t; W represents the actor network, represents the actor network parameters.

[0102] After sampling the actor network, the increment a t ={ΔΘ1(t), ΔΘ2(t)} can be obtained. The intelligent reflecting surface modifies the reflection coefficient through the controller to obtain:

[0103] Θ1(t+1)=Θ1(t)⊙ΔΘ1(t)

[0104] Θ2(t+1)=Θ2(t)⊙ΔΘ2(t)

[0105] Wherein, Θ1(t+1) represents the reflection coefficient of RIS1 at time t+1; Θ1(t) represents the reflection coefficient of RIS1 at time t; ΔΘ1(t) represents the reflection coefficient increment of RIS1 at time t; Θ2(t+1) represents the reflection coefficient of RIS2 at time t+1; Θ2(t) represents the reflection coefficient of RIS2 at time t; ΔΘ2(t) represents the reflection coefficient increment of RIS2 at time t; and is the Hadamard product.

[0106] Subsequently, the wireless communication system performs the next round of signal transmission to obtain the channel estimation value H(t+1) and the immediate reward r t, H(t+1) and Θ1(t+1), Θ2(t+1) together constitute the state s at the next moment t+1 ,(s t ,a t ,r t ,s t+1 ) is stored as a four-tuple in the memory unit and used as an element for subsequent actor network and critic network updates. Then, s t+1 Substitute into the actor network to start the next cycle.

[0107] The actor network and the critic network are updated using the existing experience replay mechanism. When the quadruple in the memory unit reaches the capacity C, some quadruple is taken out to update the actor network and the critic network. l ,a l ,r l ,s l+1 ), first according to the actor network W of the current moment u u and network parameters Calculate the mean and variance of the output action, recorded as:

[0108]

[0109] in, Represents the incremental distribution of RIS1 and RIS2 reflection coefficients at time l, respectively, in the actor network W u and network parameters Calculate the new mean; Represents the incremental distribution of RIS1 and RIS2 reflection coefficients at time l, in the actor network W u and network parameters The new variance calculated; s l ,a l ,r l ,s l+1 They represent the state, action, reward at time l and the state at time l+1 respectively.

[0110] Correspondingly, the incremental reflection coefficients of RIS1 and RIS2 are distributed in the actor network W and the network parameters The distribution of is old, that is:

[0111]

[0112] in, Represents the incremental distribution of RIS1 and RIS2 reflection coefficients in the actor network W and network parameters respectively. The mean value before the update at time l under ; Represents the incremental distribution of RIS1 and RIS2 reflection coefficients in the actor network W and network parameters respectively. The variance before the update is calculated at time l under the following equation; μ1(l), μ2(l) represent the incremental distribution of RIS1 and RIS2 reflection coefficients in the actor network W and network parameters respectively. The value of the old mean at time l under σ1(l),σ2(l) represent the incremental distribution of RIS1 and RIS2 reflection coefficients in the actor network W and network parameters respectively. The value of the old variance calculated at time l under .

[0113] Then calculate the advantage function G t and the objective function Finally update the actor network and critic network.

[0114] Secondly, the critic network is adapted to generate an action-value function V(·) to evaluate whether the actor network’s strategy is correct. The critic network is updated by minimizing the loss function, which is defined as

[0115]

[0116] in, Indicates that the policy parameter value is The value of the objective function when Q u represents the actor network index; In the actor network Q u The following network parameters.

[0117] Finally, the policy update function is set up, using gradient ascent to update the policy parameters, namely:

[0118]

[0119] in, They represent the updated policy parameters under the actor network Q and the policy parameters before the update under the actor network Q respectively; Indicates that the policy parameter value under the actor network Q is The gradient value of the objective function when α2 is the learning rate.

[0120] According to the above discussion, the process of solving the optimization problem by the proximal strategy optimization algorithm proposed in step S4 of the present invention is as follows: Figure 15 As shown, the details are as follows:

[0121] S41. In the initialization phase, the experience pool is initialized and configured with a capacity of C, the relevant parameters of the proximal policy optimization algorithm are initialized, and the environment and agent of the wireless communication system are set;

[0122] S42. In each iteration of the loop phase, obtain the number of channel estimations m and the number of state iterations n. First, determine whether the number of channel estimations m is 0:

[0123] If m = 0, the proximal strategy optimization algorithm ends;

[0124] If m≠0, continue to determine whether the number of state iterations n is 0: if n=0, set m=m-1, and then re-determine whether the number of channel estimations is 0; if n≠0, initialize state s0 and use the actor network to predict the state s at time j j Next, perform action a j The probability of using the critic network to predict the state s at time j j Under the evaluation value, the agent performs action a j , get the reward r at time j j and the state s at time j+1 j+1 , the state s at time j j 、Action a j , reward r j And the state s at time j+1 j+1 Constitute a quadruple (s j ,a j ,r j ,s j+1 ) is stored in the experience pool, and the state s at time j is stored j Update to the state s at time +1 j+1 If the experience pool does not collect a four-tuple of length C and the loop has not ended, set n = n-1, and then re-determine whether the number of state iterations is 0; if the experience pool collects a four-tuple of length C or the loop ends, perform the following operations: calculate the advantage function G, calculate the cumulative discounted reward corresponding to each state, and update the actor network and critic network;

[0125] S43. Continue iterating until the loop ends.

[0126] It should be noted that in step S4 of the present invention, to accelerate the learning process of the entire wireless communication system, the present invention uses the Distributed Proximal Policy Optimization (DPPO) algorithm to accelerate the learning process by running multiple agent instances in parallel on multiple environment replicas. This parallelization allows the wireless communication system to generate more data in the same amount of time, thereby enabling faster policy updates.

[0127] Specifically, if Figure 5As shown, the DPPO algorithm has a main process running the reinforcement learning master model PPO. In this system, there are multiple sub-processes, sub-process 1, sub-process 2, and sub-process 3, which can be responsible for different tasks. Among them, each sub-process will collect experience data, which will be passed to the PPO model in the main process. The PPO model in the main process will update according to these experiences, and the updated model will continue to affect the running of the sub-processes. "net" and "GPU" represent the network structure and graphics processing unit, which play an important role in this system, such as model operation and acceleration. The entire architecture presents a distributed feature, with each sub-process cooperating with the main process to jointly promote the progress of reinforcement learning tasks.

[0128] The following will explain how parallelization helps to speed up the learning process and how the DPPO algorithm combines with the PPO algorithm.

[0129] In traditional PPO, an agent interacts with an environment to generate sequence data s t ,a t ,r t ,s t+1 . In DPPO, assume there are parallel workers, each of which can independently generate such sequences in its own environment copy. Therefore, if each worker runs for time steps, the system can generate a total of data points. This achieves significant acceleration compared to a single agent that can only produce data points in the same time.

[0130] In a distributed setting, after the policy parameters are updated, these updates need to be synchronized to all workers. This can be done through a central server that is responsible for collecting gradients and updating parameters, and then broadcasting the updated parameters back to each worker.

[0131] Therefore, by generating a large amount of data in parallel and more frequent policy updates, DPPO can achieve faster convergence speed in the learning process. This acceleration effect can be briefly summarized as follows:

[0132]

[0133]

[0134] where is the number of parallel workers, is the number of time steps collected by each worker.

[0135] It is important to note that the efficiency of a distributed system does not always scale linearly, as it can be affected by factors such as network bandwidth, communication latency, synchronization mechanisms, and more. Therefore, the actual efficiency of DPPO is not simply the product of the number of workers and the efficiency of a single worker. DPPO accelerates the learning process by utilizing multiple parallel workers, each independently collecting data in their own environment replica. This approach significantly increases data throughput and sample diversity relative to a single worker, and can effectively scale across multi-core or multi-machine distributed computing resources. The following sections will introduce how DPPO accelerates learning.

[0136] In DPPO, each worker performs the following operations:

[0137] 1) Independently interact with the environment

[0138] 2) Collect a set of transitions s, a, r, s'

[0139] 3) Compute the gradient of this data

[0140] If there are workers, then times more data can be collected in the same amount of time, thus speeding up the learning process. The policy gradient estimation and objective function setup are the same as previously introduced.

[0141] During synchronization and update, each worker independently computes the gradient, and then these gradients are sent to the central parameter server. The server updates the policy parameters and then broadcasts the updated parameters back to each worker. The mathematical representation of this process can be simplified as:

[0142]

[0143] where and represent the updated policy parameter values and the pre-updated policy parameter values, respectively; is the number of workers; and is the clipped objective function computed by the i-th worker.

[0144] In this way, DPPO is able to accumulate a large amount of data across multiple workers, iteratively and update the policy quickly, ultimately accelerating the learning process. This approach is particularly suitable for complex environments and tasks, where a single agent may require a large number of interactions to learn an effective policy.

[0145] It should be noted that, as Figure 6As shown, a cycle in the intelligent reflecting surface reflection coefficient optimization process is shown. After the intelligent reflecting surface updates the reflection coefficient according to the instructions of the agent, the wireless communication system performs signal transmission and channel estimation. Then, the base station BS and the user UE transmit reward information (Reward) and state information (State) to the controller, respectively. Finally, the controller updates the network according to the transmitted information.

[0146] In order to better show the specific implementation and technical effects of the present application, the deep reinforcement learning-based double intelligent reflecting surface dynamic reflection coefficient adjustment method shown in steps S1-S4 in the above preferred implementation mode will be applied to a specific example.

[0147] Embodiment

[0148] The deep reinforcement learning-based double intelligent reflecting surface dynamic reflection coefficient adjustment method used in this embodiment is as previously described, and will not be repeated here.

[0149] In this embodiment, some parameters are simulated to verify the effectiveness of the method of the present application. The settings of the experiments in this embodiment are as follows: the number of BS antennas is set to 2, the number of reflecting elements of the intelligent reflecting surface is set to 36, and the number of users is set to 2, each user is equipped with an antenna. The position of the intelligent reflecting surface is set to (-3, 5, 5), the position of the BS is set to (0, 0, 10), the users are randomly distributed in a square with a side length of 10, and the height is set to 0. The discount factor γ = 0.99, the learning rate α = 0.0003, the GAE parameter λ = 0.95, and the hyperparameter ε = 0.2. In the simulation experiment, three comparison methods are used for comparison, which are DDQN, MAB, and Random. Among them, the Random group adopts a completely random phase selection, MAB is to fixedly select the reflection module with the best average performance for reflection, and DDQN represents the use of the DDQN algorithm with a discrete action set to optimize the reflection coefficient.

[0150] Figure 7 are comparison graphs of different algorithms, where the x-axis is the number of iterations, and the y-axis is the P m moving average value, with a window length of 100. The values on the right of the line chart represent the average values after the four methods reach convergence. It can be seen that the method proposed in the present application has the best performance. Figure 7DPPO) is superior to the compared schemes in both iteration speed and final performance, the random phase method has no dependence on any external information and is undoubtedly the worst, while the MAB calculates the average performance of each block of reflection units and selects the best one, with slightly improved performance. The DDQN and the method of the present application, both using deep reinforcement learning methods, can far exceed the Random and MAB methods after a certain number of iterations, because the former two update the reflection coefficients from the feedback information of the environment for iterative optimization. Figure 7 It can be seen from the above that the method of the present application can reach a relatively stable value in about 100 rounds, while the DDQN needs about 200 rounds to reach a stable value. This may be influenced by many factors, but the biggest influencing factor is that the method of the present application adopts a distributed training process, which greatly reduces the number of rounds required to reach convergence. Secondly, it may also be related to the setting of the reward function or the selection of the algorithm. Finally, when the method of the present application and the DDQN reach a relatively stable value, the final convergence value of the method of the present application is higher than that of the DDQN. The reason is that the method of the present application is a continuous action set, while the DDQN is a discrete action set. The present application is more likely to reach the theoretical optimal value.

[0151] The above algorithms and the method of the present application are also compared under different numbers of reflection units (reflection modules), as shown in Figure 8 and Figure 9 The same, the x-axis of the figure is the number of iterations, and the y-axis is the moving average value of P m . Figure 8 The performance of the above four algorithms is shown when the number of reflection units is 36 and 16, respectively. It can be seen from Figure 8 that the performance of the algorithms is decreased when the number of reflection units is reduced, and the decrease is about 15% to 20%. Correspondingly, the gap between the algorithms will also increase with the increase of the number of reflection units, but the method of the present application is always superior to the other three comparison methods. Figure 9 The average rate comparison of the above four algorithms under different numbers of reflection units is shown. It can be seen that the average rate of the above four algorithms increases with the increase of the number of reflection units, but the rate of increase slows down and finally tends to a constant value, and the method of the present application is always superior to the other three comparison methods.

[0152] Figure 10 The bit error rate of the DDQN method and the method of the present application is compared with the change of SNR. The phase information after 2000 iterations is used as the test BER parameter, and the random user position is used. Figure 10 It can be seen from the above that the performance of the method of the present application is superior to that of the DDQN. Through simulation experiments, it is verified that the method of the present application based on DPPO can improve the communication performance under long time.

[0153] Figures 11-14 The performance of the above four different algorithms when the number of users is 2, 3, 4 and 5 respectively is shown, Figures 11-14 The x-axis is the number of iterations, and the y-axis is the moving average of P m It can be seen that no matter how many users, the method of the present application is always better than the other three comparative methods. In addition, it can be seen that as the number of users increases, the overall performance of each algorithm is declining, and the stability of the algorithm also decreases with the increase of the number of users. But the number of iterations to reach convergence does not change significantly, so it can be seen that the change of the number of users will not affect the speed of algorithm convergence.

[0154] In general, the present application proposes a deep reinforcement learning-based dynamic reflection coefficient adjustment method for a wireless communication system with dual-intelligent reflecting surfaces. In the method of the present application, how to establish an intelligent reflecting surface reflection coefficient optimization model suitable for dynamic environments is studied, which involves considering factors such as dynamic changes in the environment, multipath propagation, interference and noise, and incorporating them into the model. The present application will use deep reinforcement learning as the basis for the optimization algorithm, and study how to design a reinforcement learning framework suitable for intelligent reflecting surface reflection coefficient optimization, including state representation, action selection, reward design and value function estimation, etc. The present application designs a reasonable environment interaction mechanism to enable the intelligent reflecting surface reflection coefficient optimization algorithm to interact with the dynamic environment. Through interaction with the environment, iterative updates of the strategy will be realized to gradually optimize the configuration of the intelligent reflecting surface reflection coefficient. This will include steps such as selecting appropriate actions, evaluating feedback information and updating strategy parameters.

[0155] The above-described embodiments are only a preferred scheme of the present application, and are not intended to limit the present application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application. Therefore, any technical solutions obtained by equivalent replacement or equivalent transformation shall fall within the protection scope of the present application.

Claims

1. A method for adjusting the dynamic reflection coefficient of dual intelligent reflective surfaces based on deep reinforcement learning, characterized in that: The following steps are involved: S1: Construct a wireless communication system model assisted by dual smart reflectors. All channel coefficients adopt the Rice channel model. An optimization problem is constructed with the goal of maximizing the system communication rate. The constraints are the phase shift and amplitude reflection coefficient of the smart metasurface reflector module. S2: Based on the transmission process of the smart metasurface smart reflective surface, the phase control process of the smart metasurface smart reflective surface is established as a Markov process; S3: Solving the optimization problem based on the proximal strategy optimization algorithm, and accelerating the iteration rate through distributed computing using the distributed proximal strategy optimization algorithm, thereby achieving dynamic adjustment of the reflection coefficient during the optimization solution process; In step S1, the wireless communication system model includes a base station, two smart metasurfaces, A single antenna user and a controller that can adjust the phase of the reflector, the base station is equipped with antennas, and the number of reflecting units on the reflecting surface is ; In step S1, the optimization goal is specifically: ; ; ; ; in, express The reflection coefficient matrix; express The reflection coefficient matrix; represents the total data rate; represents the data rate of user k; Indicates the number of users; represents the signal-to-interference-and-noise ratio of user k; Represents the reflection coefficient matrix of the first or second smart metasurface; represents a diagonal matrix, represents the imaginary unit, Represents the first or second smart metasurface Phase shift of each reflection module; Represents the first or second smart metasurface The amplitude reflection coefficient of each reflection module; Indicates the first or second smart metasurface Phase shift of each reflection module; Indicates the first or second smart metasurface The amplitude reflection coefficient of each reflection module; In step S2, the four-tuple represents a Markov process, represents the state space, represents the action space, represents the state transition function, represents the reward function; where the state space According to the elements that may affect the optimization target in the actual communication process, the state space middle The state set at the moment is recorded as , is the set of all user channels, Represents users channel; the action space middle The action set at time , Represent the reflection coefficient matrix The increment of the reward function The functional form is: ; in, is the rate threshold.

2. A method for adjusting dynamic reflection coefficients of dual intelligent reflective surfaces based on deep reinforcement learning according to claim 1, characterized in that: In step S3, the process of solving the optimization problem using the proximal strategy optimization algorithm is as follows: S31. In the initialization phase, initialize the experience pool and configure the capacity to , initialize the relevant parameters of the proximal policy optimization algorithm and set the environment and intelligent agent of the wireless communication system; S32. In each iteration of the loop phase, obtain the number of channel estimations m and the number of state iterations n. First, determine whether the number of channel estimations m is 0. If m = 0, the proximal policy optimization algorithm ends. If m ≠ 0, continue to determine whether the number of state iterations n is 0. If n = 0, set m = m-1, and then re-determine whether the number of channel estimations is 0. If n≠0, initialize the state , using actor network prediction in State of the moment Next action The probability of using the critic network to predict State of the moment Under the evaluation value, the agent performs the action ,get Rewards of the moment and State of the moment ,Will State of the moment ,action ,award as well as State of the moment Forming a quadruple Stored in the experience pool, and State of the moment Update to State of the moment , if the experience pool does not collect a length of If the experience pool collects a quadruple of length The quadruple or loop ends, then the following operations are performed: Calculate the advantage function , calculate the cumulative discounted reward corresponding to each state, and update the actor network and critic network; S33. Continue iterating until the loop ends.

3. A method for adjusting dynamic reflection coefficients of dual intelligent reflective surfaces based on deep reinforcement learning according to claim 2, characterized in that: Time advantage function is calculated as follows: ; ; in, is the discount factor; is the parameter in the generalized advantage estimate; Respectively Intermediate variables at the moment; Represents time parameters; Respectively .

4. A method for adjusting dynamic reflection coefficients of dual intelligent reflective surfaces based on deep reinforcement learning according to claim 3, characterized in that: The proximal policy optimization algorithm adopts the PPO-Clip method, and the objective function is to maximize the following objectives: ; ; in, Indicates that the policy parameter value is The value of the objective function when ; represents the policy parameters; Expresses expectation; represents the probability ratio; Indicates the use of the PPO-clip method; represents a hyperparameter; express strategy; are the strategy parameters before updating; The specific way to use gradient ascent to update policy parameters is: ; in, represents the updated policy parameters; represents the learning rate; Indicates finding the gradient.

Citation Information

Patent Citations

  • MIMO full duplex power distribution method based on intelligent reflecting surface

    CN116318288A

  • RIS-assisted MU-MISO communication system intelligent beam forming method based on deep reinforcement learning

    CN118764055A