A Reconfigurable Intelligent Surface Beamforming Method Based on Model-Driven Reinforcement Learning
By combining model-driven and deep reinforcement learning methods, the complexity of beamforming design in multi-RIS assisted systems is solved, and efficient adaptive combined beamforming is achieved, improving system performance and training efficiency.
Patent Information
- Application Number
- CN202411324365.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2044-09-23
AI Technical Summary
The existing technology lacks effective joint beamforming design in multi-RIS auxiliary systems. The model-driven DL architecture has problems of parameter dependence and many iterations, and the DRL method faces the challenges of large training variables and slow convergence.
Combined with model-driven and deep reinforcement learning, the memory is played back, the actions are trained and subnet parameters are evaluated, and the beamforming vector is initialized using the zero-force algorithm, the agent state is iteratively updated, and the network is updated by minimizing the loss function to achieve adaptive beamforming.
It improves the reachable rate performance of RIS auxiliary system, reduces algorithm complexity and training time, and the network can adaptively adjust in a dynamic environment.
Smart Images

Figure CN119210528B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to communication technology, and in particular to a reconfigurable intelligent surface beamforming method based on model-driven reinforcement learning. Background Art
[0002] Advances in communications technology have triggered a demand for higher data rates and improved coverage. Reconfigurable Intelligent Surfaces (RIS) offer a potential solution for meeting these demands by manipulating the reflection, refraction, and absorption properties of electromagnetic waves, enabling customized communication environments. In this context, joint beamforming design in RIS-assisted systems has attracted significant attention. However, existing work has primarily focused on single-RIS-assisted systems, while research on multi-RIS-assisted systems remains limited.
[0003] In recent years, the application of deep learning (DL) technology in physical layer communications has increased. Existing DL-based joint beamforming work mainly focuses on single-RIS-assisted systems, relying on data-driven methods, performing poorly in dynamic environments and lacking interpretability. In contrast, model-driven DL methods improve resilience and interpretability by integrating expert knowledge. However, existing model-driven DL architectures suffer from layer-dependent parameters and non-adaptive iteration times, limiting their application in dynamic scenarios. Deep reinforcement learning (DRL) technology has also been applied to joint beamforming research. DRL autonomously collects data and trains itself through continuous interaction with the environment. However, DRL methods face the challenges of a large number of trainable variables and a slow convergence process. It is hoped that by combining model-driven methods with DRL, an innovative framework suitable for wireless communication applications can be proposed to alleviate network training complexity and improve performance by integrating expert knowledge. Summary of the Invention
[0004] Purpose of the invention: To address the problems existing in the prior art, the present invention provides a reconfigurable intelligent surface beamforming method based on model-driven reinforcement learning.
[0005] Technical solution: The RIS-assisted system joint beamforming method for model-driven reinforcement learning described in the present invention includes:
[0006] (1) Before deployment, the experience replay memory, training action subnet parameters, target action subnet parameters, training evaluation subnet parameters, and target evaluation subnet parameters of reinforcement learning are randomly initialized.
[0007] (2) After obtaining the channel estimation results, the passive beamforming vector of RIS is randomly initialized; the active beamforming vector is initialized using the ZeroForcing algorithm; and the state of the reinforcement learning network is initialized.
[0008] (3) During iteration, the training action subnet generates corresponding actions based on the current state, calculates rewards and updates the agent state; stores the experience in the replay memory; and outputs the active beamforming vector and the passive beamforming vector when the algorithm converges.
[0009] (4) Extract a batch of experiences from the experience replay memory, update the training action subnet and the training evaluation subnet by minimizing the loss function, and use soft update to update the target action subnet and the target evaluation subnet.
[0010] Furthermore, step (2) specifically includes:
[0011] (2.1) Randomly initialize the passive beamforming vector and
[0012] (2.2) Calculate the active beamforming vector W using the ZF algorithm (0) =H H (HH H ) -1 , where H=[h1,h2,…,h K ] H is the matrix composed of the combined channels of each user, K is the number of users, and hk is the combined channel of user k;
[0013] (2.3) Initialize the state of the reinforcement learning network a (0) ,H all They represent the achievable rate and corresponding to the initialization beamforming vector, the change in the achievable rate and between two adjacent iterations, the change in the active beamforming vector between two adjacent iterations, the action of the agent at initialization, and the matrix composed of all channels in the system. is the algorithm step size during initialization.
[0014] Furthermore, step (3) specifically includes:
[0015] (3.1) The action subnet generates the algorithm step size based on the observation and The observation is represented by t represents the time, s (t) [1:6] indicates state s (t) The first 6 elements in ;
[0016] (3.2) Calculate the auxiliary intermediate variables u and λ. The elements in u and λ are respectively recorded as:
[0017]
[0018] Among them, P T is the maximum transmit power, w i is the active beamforming vector of the i-th user, is the receiving noise variance of user k, u k is the receiving coefficient of user k, the * in the upper right corner represents the conjugate of the variable, λ k The intermediate variables introduced into the algorithm have no actual physical meaning;
[0019] (3.3) Update the active beamforming vector according to the auxiliary intermediate variables u and λ calculated in step (3.2);
[0020]
[0021] α k is the priority of user k;
[0022] (3.4) Calculate the auxiliary intermediate variables u and λ again;
[0023] (3.5) Using the step size output by the action subnet The passive beamforming vector of the first reconfigurable smart surface is updated as
[0024]
[0025] in is the value of the reflection phase φ1 of the first reconfigurable smart surface in the previous iteration,
[0026]
[0027] in represents the Hadamard product of the vectors, G1 and G2 are the channels from AP to the first reconfigurable smart surface and from AP to the second reconfigurable smart surface respectively; H 1,2 is the channel between two reconfigurable smart surfaces; h 1,k 、h 2,k are the channels from the first reconfigurable smart surface to user k and the channels from the second reconfigurable smart surface to user k respectively; represents the real part of the variable, J is the imaginary unit, φ2 is the reflection phase of the first reconfigurable smart surface;
[0028] (3.6) The third calculation of auxiliary intermediate variables u and λ;
[0029] (3.7) Using the step size output by the action subnet updating the passive beamforming vector of the second reconfigurable smart surface;
[0030] (3.8) Calculate the current reward
[0031]
[0032] where r con is the reward for algorithm convergence, r iter is the penalty for each iteration of the algorithm, η W They represent the change of the achievable rate and the active beamforming vector in the t-th iteration, the judgment threshold of the achievable rate and the convergence of the active beamforming vector, and the judgment threshold of the convergence of the active beamforming vector respectively;
[0033] (3.9) Update the agent status in Represents the change in the sum of the achievable rates of two adjacent rounds of iterations, represents the change of the active beamforming vector between two adjacent iterations, H all =[Vec(G1),Vec(G2),Vec(H1),Vec(H2),Vec(H 1,2 )] contains the real form of all channel matrices, where H1 = [h 1,1 ,…,h 1,K ] and H2=[h 2,1 ,…,h 2,K ],Vec(A)=[R(a 1,1 ,…,a 1,N ,…,a N,1 ,…,a N,N ),I(a 1,1 ,…,a 1,N ,…,a N,1 ,…,a N,N )] represents the operation of concatenating the real and imaginary parts of a matrix and converting them into a vector, a i,j Represents the element in the i-th row and j-th column of the matrix;
[0034] (3.10) Update agent experience {s (t) ,a (t) ,r (t) ,s (t+1)}store in playback memory;
[0035] (3.11) satisfies the convergence condition and When , the current active beamforming vector and passive beamforming vector are output.
[0036] Furthermore, step (4) specifically includes:
[0037] (4.1) Extract a batch of experience s from the experience replay memory (p) ,a (p) ,r (p) ,s (p+1 );
[0038] (4.2) Update the training evaluation subnet by minimizing the loss function
[0039]
[0040] (4.3) Update the training action subnet by minimizing the loss function
[0041]
[0042] (4.4) Use soft update to update the target action subnet and target evaluation subnet, denoted as
[0043]
[0044] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:
[0045] This invention rationally combines model-driven concepts with DRL. In RIS-assisted systems, the achievable rate performance is significantly improved while algorithmic complexity is greatly reduced. The invention also proposes a model-driven reinforcement learning network, which significantly reduces the number of network training parameters and training time. The number of network iterations can be dynamically adjusted based on different environments, and the network can adapt through online learning when the deployment environment changes. In summary, this invention provides a joint beamforming solution for RIS-assisted systems with practical deployment prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic flow chart of an embodiment of the present invention;
[0047] Figure 2 It is a structural diagram of the action subnet of the present invention. DETAILED DESCRIPTION
[0048] The present invention will be described in detail below with reference to the accompanying drawings and an example of a dual RIS-assisted downlink multi-user MISO system.
[0049] 1. Channel Model Applicable to This Embodiment
[0050] In a MU-MISO system with one AP and K single-antenna users, two RISs are introduced to assist the communication process of the system. In this embodiment, the number of users is K = 4, and both RIS1 and RIS2 have N1 = N2 = 40 reflection elements. The maximum transmit power of the AP is P T =9dBm, the carrier frequency is 3GHz, and the distance between adjacent antennas is half a wavelength. The channel from AP to RIS1, the channel from AP to RIS2, the channel from RIS1 to user k, the channel from RIS2 to user k, and the channel from RIS1 to RIS2 are represented by G1, G2, and h respectively. 1,k 、h 2,k and H 1,2 The phase shift introduced by the nth reflector unit of the lth RIS is expressed as φ l,n ∈[0,2π), record The corresponding reflection coefficient is expressed as where n∈1,2,…,N l , l∈1,2. The passive beamforming vector of RISl is denoted as N l is the number of reflection units of RIS1. The channel between AP and user can be expressed as:
[0051]
[0052] in, represents the phase shift matrix associated with RIS1.
[0053] The signal transmitted to user k is denoted as s k In this embodiment, it is assumed to be an independent random variable with zero mean and unit variance. The AP first uses the active beamforming vector w k Processes the transmission signal to user k. This vector is constrained The limit, where P T Represents the maximum transmission power. After transmission through the channel, the signal received by user k is expressed as
[0054]
[0055] in represents the received noise at user k. When the channel CSI is known, the received signal-to-noise ratio (SINR) can be expressed as:
[0056]
[0057] When the priority of user k is denoted as α k When , the user’s achievable rate is expressed as
[0058]
[0059] 2. Specific steps of this embodiment
[0060] like Figure 1 As shown in the figure, model-driven reinforcement learning is applied to wireless communication design and used for joint beamforming design of RIS system. The complete joint beamforming process includes three steps: initialization, algorithm iteration, and network update:
[0061] (1) Initialization
[0062] Before the algorithm is deployed, the experience replay memory of reinforcement learning is randomly initialized to train the action subnet parameters Target action subnet parameters Training and evaluation subnet parameters Target evaluation subnet parameters
[0063] During the deployment process, each time after obtaining CSI, the passive beamforming vector of RIS is randomly initialized. and Use ZF algorithm to calculate the active beamforming vector W (0) =H H (HH H ) -1 , where H=[h1,h2,…,h K ] H , K is the number of users, h k is the combined channel of user k. Initialize the state of the reinforcement learning network a (0) ,H all They represent the achievable rate and corresponding to the initialization beamforming vector, the change in the achievable rate and between two adjacent iterations, the change in the active beamforming vector between two adjacent iterations, the action of the agent at initialization, and the matrix composed of all channels in the system. is the algorithm step size during initialization.
[0064] (2) Algorithm iteration
[0065] In the tth iteration, the training action subnet is based on the current state Generate corresponding actions in Represents the change in the sum of the achievable rates of two adjacent rounds of iterations, represents the change of the active beamforming vector between two adjacent iterations, H all=[Vec(G1),Vec(G2),Vec(H1),Vec(H2),Vec(H 1,2 )] contains the real form of all CSI, where H1=[h 1,1 ,…,h 1,K ] and H2=[h 2,1 ,…,h 2,K ], Vec(·) represents the operation of concatenating the real and imaginary parts of the matrix and converting them into a vector. In order to enhance the exploration ability of DRL, the agent selects a random step size with a small preset probability ξ After getting the action, calculate the current reward
[0066]
[0067] where r con is the reward for algorithm convergence, r iter is the penalty for each iteration of the algorithm. Then, the agent updates the state and uses the experience {s (t) ,a (t) ,r (t) ,s (t+1)} is stored in the playback memory for subsequent training. Then, the convergence of the algorithm is judged. If the convergence condition is met, that is, and Then the current active beamforming vector and passive beamforming vector are output.
[0068] Specifically, the proposed model-driven reinforcement learning network structure is as follows Figure 2 The action subnet is composed of three layers of fully connected networks. The “observation” of the intelligent agent is selected as the input of the action subnet. (t) is defined as a part of the agent's state, represented as
[0069]
[0070] The action subnet generates the algorithm step size based on the observation and Then the WMMSE-SCA algorithm is used to update the active and passive beamforming vectors. Specifically, the auxiliary intermediate variables u and λ are first calculated and recorded as
[0071]
[0072] Among them, P T is the maximum transmit power, w i is the active beamforming vector of the i-th user, is the receiving noise variance of user k, u k is the receiving coefficient of user k, the * in the upper right corner represents the conjugate of the variable, λ kThe intermediate variables introduced into the algorithm have no actual physical meaning.
[0073] Then, update the active beamforming vector
[0074]
[0075] Next, update the auxiliary intermediate variables u and λ again, and use the step size output by the action subnet Update the passive beamforming vector of RIS1 and record it as
[0076]
[0077] in
[0078]
[0079] in Represents the Hadamard product of the vector, and then updates the auxiliary intermediate variables u and λ again, and uses the step size output by the action subnet Update the passive beamforming vector of RIS2 and record it as
[0080]
[0081] Thus, the agent can get the corresponding action
[0082] (3) Network Update
[0083] Extract a batch of experience s from the experience replay memory (p) ,a (p) ,r (p) ,s (p+1) , where p represents the index of experience. By minimizing the loss function
[0084]
[0085] Update the training review subnet. By minimizing the loss function
[0086]
[0087] Update the training action subnet. Finally, use soft update to update the target action subnet and target evaluation subnet to ensure training stability and convergence, denoted as
[0088]
[0089] where the variable τ c and τ a Determines the extent of soft updates.
[0090] The above disclosure is only a preferred embodiment of the present invention and cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A reconfigurable intelligent surface beamforming method based on model-driven reinforcement learning, characterized in that: Two reconfigurable smart surfaces are introduced to assist the communication process of the system. The method includes the following steps: (1) Before deployment, randomly initialize the reinforcement learning experience replay memory, training action subnet parameters, target action subnet parameters, training evaluation subnet parameters, and target evaluation subnet parameters; (2) After obtaining the channel estimation results, randomly initialize the passive beamforming vector of the reconfigurable smart surface; initialize the active beamforming vector; and initialize the state of the reinforcement learning network; (3) During iteration, the training action subnet is used to generate corresponding actions based on the current state, calculate rewards and update the agent state; the experience is stored in the replay memory; when the algorithm converges, the active beamforming vector and the passive beamforming vector are output; (4) Extract a batch of experiences from the experience replay memory, update the training action subnet and the training evaluation subnet by minimizing the loss function, and use soft update to update the target action subnet and the target evaluation subnet; Step (2) specifically includes: (2.1) Randomly initialize the passive beamforming vector and (2.2) Calculate the active beamforming vector W using the zero-forcing algorithm (0) =H H (HH H ) -1 , where H=[h1,h2,…,h k ] H is the matrix composed of each user combined with the channel, K is the number of users, h k is the combined channel of user k; (2.3) Initialize the state of the reinforcement learning network H all They represent the achievable rate and corresponding to the initialization beamforming vector, the change in the achievable rate and between two adjacent iterations, the change in the active beamforming vector between two adjacent iterations, the action of the agent at initialization, and the matrix composed of all channels in the system. is the algorithm step size during initialization.
2. The reconfigurable intelligent surface beamforming method based on model-driven reinforcement learning according to claim 1 is characterized in that: Step (3) specifically includes: (3.1) Training the action subnet based on the observation generation algorithm step size and The observation is represented by o (t) = t represents the time, s (t) [1:6] indicates state s (t) The first 6 elements in ; (3.2) Calculate the auxiliary intermediate variables u and λ. The elements in u and λ are respectively recorded as: Among them, P T is the maximum transmit power, w i is the active beamforming vector of the i-th user, is the receiving noise variance of user k, u k is the receiving coefficient of user k, the * in the upper right corner represents the conjugate of the variable, λ k The intermediate variables introduced into the algorithm have no actual physical meaning; (3.3) Update the active beamforming vector according to the auxiliary intermediate variables u and λ calculated in step (3.2); α k is the priority of user k; (3.4) Calculate the auxiliary intermediate variables u and λ again; (3.5) Using the step size output by the action subnet The passive beamforming vector of the first reconfigurable smart surface is updated as in is the value of the reflection phase φ1 of the first reconfigurable smart surface in the previous iteration, where ° represents the Hadamard product of the vectors, G1 and G2 are the channels from the AP to the first reconfigurable smart surface and from the AP to the second reconfigurable smart surface, respectively; H 1,2 is the channel between two reconfigurable smart surfaces; h 1,k 、h 2,k are the channels from the first reconfigurable smart surface to user k and the channels from the second reconfigurable smart surface to user k respectively; represents the real part of the variable, J is the imaginary unit, φ2 is the reflection phase of the first reconfigurable smart surface; (3.6) The third calculation of auxiliary intermediate variables u and λ; (3.7) Using the step size output by the action subnet updating the passive beamforming vector of the second reconfigurable smart surface; (3.8) Calculate the current reward where r con is the reward for algorithm convergence, r iter is the penalty for each iteration of the algorithm, They represent the change of the achievable rate and the active beamforming vector in the t-th iteration, the judgment threshold of the achievable rate and the convergence of the active beamforming vector, and the judgment threshold of the convergence of the active beamforming vector respectively; (3.9) Update the agent status in Represents the change in the sum of the achievable rates of two adjacent rounds of iterations, represents the change of the active beamforming vector between two adjacent iterations, H all =[Vec(G1),Vec(G2),Vec(H1),Vec(H2),Vec(H 1,2 )] contains the real form of all channel matrices, where H1 = [h 1,1 ,…,h 1,K ] and H2=[h 2,1 ,…,h 2,K ],Vec(A)=[R(a 1,1 ,…,a 1,N ,…,a N,1 ,…,a N,N ),I(a 1,1 ,…,a 1,N ,…,a N,1 ,…,a N,N )] represents the operation of concatenating the real and imaginary parts of a matrix and converting them into a vector, a i,j Represents the element in the i-th row and j-th column of the matrix; (3.10) Update agent experience {s (t) ,a (t) ,r (t) ,s (t+1) }store in playback memory; (3.11) satisfies the convergence condition and When , the current active beamforming vector and passive beamforming vector are output.
Citation Information
Patent Citations
RIS joint beamforming method based on deep reinforcement learning and communication system
CN117176218A
Beam forming design method for STAR-RIS-assisted multi-user MISO URLLC system
CN117375684A