Optimization method of IRS auxiliary short packet communication broadcasting system based on DDPG
By applying the DDPG algorithm to optimize the phase shift of beamforming vector and IRS at the base station in a multi-user broadcast scenario, the IRS optimization problem that is difficult to effectively solve in traditional methods is solved, and the system performance is maximized and adaptive optimization is achieved.
Patent Information
- Application Number
- CN202510338537.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In multi-user broadcast scenarios, traditional gradient or convex optimization methods are difficult to effectively solve IRS optimization problems, especially due to the nonlinear coupling characteristics of the actual phase shift model, which leads to problems such as difficult to solve and high computational complexity.
The deep deterministic policy gradient (DDPG) algorithm is used to optimize the phase shift of beamforming vectors and IRS at the base station, and maximize system performance by constructing end-to-end channel distribution and designing reward functions.
Through the optimization of the DDPG algorithm, system performance can be significantly improved in multi-user broadcast scenarios, overcome the limitations of traditional methods, and implement adaptive optimization strategies.
Smart Images

Figure CN120185652A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of new generation mobile communication technology, and is a new method for optimizing an intelligent reflecting surface assisted broadcast communication system by using deep reinforcement learning. Background Art
[0002] An intelligent reflecting surface (IRS) is a programmable reflecting surface composed of a large number of passive and low-cost reflecting units. Its working principle is to intelligently control the propagation direction and characteristics of wireless signals by adjusting the phase shift of the reflecting units, thereby significantly enhancing the performance of wireless communication systems. As one of the key technologies for the development of 6G mobile communication, IRS has become a research hotspot in intelligent communication systems due to its characteristics of low power consumption, low cost, and high flexibility.
[0003] In recent years, the research on IRS has mainly been based on two phase shift models: the ideal phase shift model and the practical phase shift model. Among them, the ideal phase shift model assumes that each reflection unit can independently and precisely adjust the phase shift without being restricted by hardware. However, under actual hardware conditions, this assumption is difficult to meet. To be closer to the real situation, the practical phase shift model (S. Abeywickrama, R. Zhang, Q. Wu, and C. Yuen, “Intelligent reflecting surface: Practical phase shift model and beamforming optimization,” IEEE Trans. Commun., vol. 68, no. 9, pp. 5849 - 5863, Sep. 2020.) has been proposed, which takes into account the coupling effect of the phase shift and amplitude of the IRS reflection units. Specifically, in the practical phase shift model, the amplitude of each reflection unit changes with the phase shift angle, and this coupling effect leads to additional complexity in design and optimization. In the broadcast scenario, the optimization problem of IRS usually involves the collaborative optimization of multi - user channels, and its core objective is to maximize the system performance by adjusting the phase shift and amplitude distribution of the IRS. However, limited by the non - linear coupling characteristics of the practical phase shift model, such optimization problems generally present as non - convex problems, with characteristics such as difficult to solve and high computational complexity. Therefore, traditional gradient - based or convex - optimization methods may face bottlenecks when solving these problems. To address these challenges, the Deep Deterministic Policy Gradient (DDPG) method can be adopted. The DDPG algorithm can effectively handle problems in high - dimensional state spaces and continuous action spaces. Through DDPG, an agent can be designed to dynamically adjust the phase shift and amplitude distribution of the IRS, thereby maximizing the system performance in the multi - user broadcast scenario. This method can not only overcome the limitations of traditional optimization methods but also achieve an adaptive optimization strategy in complex actual environments. Summary of the Invention
[0004] The present invention provides an optimization method for an IRS - assisted short - packet communication broadcast system based on DDPG, which optimizes the beamforming vector at the base station and the phase shift of the IRS by using the Deep Deterministic Policy Gradient (DDPG) algorithm to improve the performance of the multi - user broadcast communication system.
[0005] The object of the present invention is achieved through the following technical solutions:
[0006] An optimization method for an IRS - assisted short - packet communication broadcast system based on DDPG, and the specific implementation steps are as follows:
[0007] Step 1: Construct an IRS-assisted downlink broadcast system, including channel modeling among the Base Station (BS), IRS, and multiple User Equipments (UEs), and introduce an actual phase shift model to simulate the hardware limitations of the IRS.
[0008] Step 2: According to the system model established in Step 1, deduce the end-to-end channel distribution; establish a system optimization objective based on the statistical channel state information, that is, to maximize the minimum average signal-to-noise ratio among all users to ensure that each user can obtain good communication quality. The specific implementation is as follows:
[0009] The received signal at UEi is y i = H i x + w, where UEi represents the i-th user equipment, H i = h i Φg0f H , g0 and h i represent the channel coefficients between the BS and the IRS, and between the IRS and UEi respectively, Φ represents the reflection coefficient matrix of the IRS, is the beamforming vector at the BS, f H represents the conjugate transpose of f, x represents the short packet signal with power P s transmitted by the BS, w is the additive complex Gaussian white noise with mean 0 and variance σ 2 ; according to the received signal, the instantaneous signal-to-noise ratio at UEi is: γ i = ρ|h i Φg0f H | 2 , where |·| represents the modulus operation, ρ = P s / σ 2 is the transmit signal-to-noise ratio; to enable each user to receive a reliable short packet signal, when ρ is given, maximize the average signal-to-noise ratio E(γ i ) at UEi, which is equivalent to maximizing E(·) represents the mean operation;
[0010] Using the properties of the circularly symmetric complex Gaussian distribution, the distribution of the channel H i can be deduced: where represents the real part operation, represents the imaginary part operation, j represents the imaginary unit, obeys the normal distribution with mean μ i and variance σ i 2 , where
[0011] ξ0 and ξ i are the large-scale parameters of channel g0 and channel h i respectively, represents the array response at the BS, k0 and k i are the Rice factors of channel g0 and channel h i respectively, N represents the number of IRS reflection units, β(θ l ) ∈ [0,1] is the amplitude coefficient of the l-th reflection unit of the IRS, and are the line-of-sight components of channel g0 and channel h i respectively;
[0012] obeys a normal distribution with mean and variance , where
[0013]
[0014] Through the above derivation, it can be obtained that the distribution of E(|H i | 2 ) can be used for the design of the reward function, and the following optimization scheme is given accordingly:
[0015]
[0016] H i = h i Φg0f H , i ∈ {1,2,...,K}
[0017]
[0018] -π ≤ θ l < π
[0019] where min{·} represents taking the minimum element in the array.
[0020] Step 3: Use the DDPG algorithm to optimize the beamforming vector at the BS and the IRS phase shift. The DDPG model includes an actor network, a critic network, a target actor network, and a target critic network. The actor network outputs the beamforming vector and the IRS phase shift, and the critic network outputs the Q value, which is used to evaluate the quality of the action. The target actor network outputs the target action, and the target critic network outputs the target Q value. These outputs are used to train the actor network and the critic network respectively to reduce the fluctuations in the training process.
[0021] Step 4: Continuously iterate and train the DDPG model to optimize the beamforming vector at the BS and the IRS phase shift configuration, and save the optimized parameters.
[0022] Step 5: Finally, perform the minimum common block length analysis using the optimized beamforming vector at BS and the IRS phase shift.
[0023] Furthermore:
[0024] The specific implementation of Step 1 is as follows:
[0025] In an IRS-assisted downlink broadcast system, assume that there is a blockage between the BS and the UEs in the system and they cannot communicate directly. A BS equipped with M antennas broadcasts short-packet signals to K single-antenna UEs (denoted as UEi, i ∈ {1, 2,..., K}) with the assistance of an IRS with N reflection elements. The reflection coefficient matrix of the IRS is defined as: which represents a diagonal matrix with as the diagonal elements, θ l being the phase shift coefficient of the l-th reflection element of the IRS, θ l ∈[-π, π), and the actual phase shift model amplitude-phase shift coupling relation is:
[0026]
[0027] where β min ≥0, θ≥0, and τ≥0 are constants related to the specific circuit implementation. Assume that the antennas of the BS and the reflection elements of the IRS are both arranged in a uniform linear array, then the array responses at the BS and the IRS are where ζ ∈ {M, N}, d is the spacing between the reflection elements of the IRS or the spacing between the base station antennas, λ is the wavelength, θ is the Angle of Arrival, AOA or Angle of Departure, AOD of the signal. The channel coefficients between the BS and the IRS, and between the IRS and UEi are denoted as g0 and h i .
[0028] where α represents the path loss at a reference distance of 10 meters, α0 and α i are the path loss exponents, d0 is the distance from the BS to the center of the IRS, d i is the distance from the center of the IRS to UEi, and are the line-of-sight components, represents the complex domain, the AOD and AOA of the signal from the BS to the IRS are and the AOD from the IRS to UEi is It can be obtained that represents The conjugate transpose of and is the non-line-of-sight component, where each element follows a circularly symmetric complex Gaussian distribution with a mean of 0 and a variance of 1.
[0029] The specific implementation of step 3 is as follows:
[0030] Establish the Markov decision process of DDPG reinforcement learning. The beamforming vector at the BS and the phase shift of the IRS are the key parameters. Therefore, the beamforming vector at the BS and the phase shift of the IRS are used as the action vector, denoted as:
[0031] a t =(θ 1,t , θ 2,t ,..., θ N,t , |f t |, arg(f t ))
[0032] where arg(·) represents the operation of taking the argument, and the subscript t represents the t-th time step.
[0033] The state vector is mainly composed of the channel state, the beamforming vector at the BS, and the phase shift and amplitude of the IRS, denoted as:
[0034]
[0035] where
[0036] Under the statistical channel, the reward function can be expressed as:
[0037] r t = min{E(|H1| 2 ), E(|H2| 2 ),..., E(|H i | 2 ),..., E(|H K | 2 )}
[0038] The specific implementation of step 4 is as follows:
[0039] In the DDPG network framework, the Actor network and the target Actor network consist of a fully connected input layer, two hidden layers, a fully connected output layer, and a Tanh function; the Critic network and the target Critic network consist of a fully connected input layer, two hidden layers, and a fully connected output layer. Batch normalization layers and linear ReLU activation functions are used to connect between the layers of the network, and the Adam optimizer is used.
[0040] The update methods of each network are as follows:
[0041] Actor network parameters θ a Updated using gradient ascent, expressed as:
[0042]
[0043] Where α a is the learning rate for updating the Actor network, B is the sampling batch, represents the policy gradient, q(s t ,a t :θ c ) is the current action-value function Q generated after the input s c and a t to the Critic network with parameters θ t . π(s t :θ a ) is the execution policy that generates a t for the input s t to the Actor network. The target Critic network generates the target return:
[0044]
[0045] Where, λ r ∈(0,1] is the discount factor, is the target Q value generated after the input s ct and a t+1 to the target Critic network with parameters θ t+1 , and a t+1 is the action generated after the input s at to the target action network with parameters θ t+1 . To make the current Q value and the target Q value close, it is necessary to minimize their mean squared Bellman error loss, expressed as:
[0046]
[0047] And update θ c using gradient descent, expressed as:
[0048]
[0049] Where α c is the learning rate for updating the evaluation network. The target network weights and copies the parameters through soft updates to ensure the directionality and stability of the network training process. The soft updates of the target Actor network and the target Critic are respectively expressed as:
[0050] θ ct ←τ r θ c -(1-τ r )θ ct
[0051] θ at ←τ r θ a -(1 - τ r )θ at
[0052] where τ r ∈(0, 1] is the learning rate for updating the evaluation network.
[0053] After clarifying the action vector, state vector, and reward function, the specific process of the DDPG optimization algorithm is as follows:
[0054] 2 - 1. Initialize the Actor network, target Actor network, Critic network, and target Critic network;
[0055] 2 - 2. Initialize the training episode to 0;
[0056] 2 - 3. Initialize the channel state, phase shift, and amplitude of the IRS, etc.;
[0057] 2 - 4. Initialize the step size step to 0;
[0058] 2 - 5. According to the current state s t , the Actor network outputs the action a t ;
[0059] 2 - 6. Obtain the amplitude corresponding to each phase shift of the IRS unit according to the amplitude - phase shift coupling relationship;
[0060] 2 - 7. Calculate the reward of the current step min{E(|H1| 2 ), E(|H2| 2 ),..., E(|H i | 2 ),..., E(|H K | 2 )};
[0061] 2 - 8. Obtain the new state s t based on the current state s t and the action a t+1 ;
[0062] 2 - 9. Fill the obtained experience tuple (s t , a t , r t , s t+1 ) into the buffer;
[0063] 2-10. If the buffer is full, update the parameters of the Actor network, target Actor network, Critic network, and target Critic network;
[0064] 2-11. Update the current state s t = s t+1 ;
[0065] 2-12. Determine whether the time step step satisfies step < max_step (max_step represents the maximum step). If it is satisfied, return to step 2-5. If not, proceed to step 2-13;
[0066] 2-13. Determine whether the number of training episodes episode satisfies episode < max_episode (max_episode represents the maximum number of rounds). If it is satisfied, return to step 2-3. If not, the optimization process ends, and the optimized IRS phase shift and amplitude information, beamforming vector at the BS, and network weight parameters are obtained.
[0067] The specific implementation of step 5 is as follows: Finally, the minimum common block length analysis is performed using the optimized beamforming vector at the BS and IRS phase shift.
[0068] For a given short data packet block length T, signal-to-noise ratio γ, and short data packet block error rate Block Error Rate, BLER ε, the maximum achievable rate of short packet communication is approximately expressed as where C(γ) = log2(1 + γ) represents the Shannon capacity, V(γ) = (1 - (1 + γ) -2 )(log2e) 2 represents the channel divergence, Q -1 (·) represents the inverse of the Gaussian Q function. The physical layer information rate of the UE is where F represents the number of information bits, and ε can be approximated as The average BLER in a wireless fading channel can be approximated as where f γ (x) is the probability density function of γ. Assume is the target average BLER for all UEs. The minimum common block length optimization algorithm process is as follows:
[0069] (1) Given system parameters N, M, k0, k1,..., k K , d0, d1,..., d K , α, α1,..., α K , θ1,..., θ N , f, F, Initialize i to 1;
[0070] (2) Initialize the block length T, and limit the block length within [M min , M max . Let T - = M min , T + = M max ;
[0071] (3) Calculate the corresponding average - and T + respectively.
[0072] (4) If the calculated average BLER satisfies , let ( represents the minimum block length required for reliable communication of UEi), which means there is no solution;
[0073] (5) If the calculated average BLER satisfies , then let
[0074] (6) If neither of the above cases is satisfied, under the condition that T + - T - ≠ 1, let calculate the corresponding average represents the ceiling function;
[0075] (7) When is satisfied, then let ; otherwise let
[0076] (8) Update T - or T + , and continue steps (6) to (7) until T + - T - = 1 to stop the iteration, and output
[0077] (9) Determine whether i satisfies i < K. If it is satisfied, set i = i + 1 and enter step (2); if it is not satisfied, enter step (10);
[0078] (10) Traverse and select the maximum value among them and set it as the minimum common block length T min of all UEs.
[0079] The present invention proposes an optimization method for an IRS-assisted short-packet communication broadcast system based on DDPG. The core lies in the first derivation of the end-to-end channel distribution of the multi-user broadcast system under the actual phase shift model of the IRS, and the establishment of the DDPG system optimization goal from this, designing a DDPG reinforcement learning framework, jointly optimizing the beamforming vector at the BS and the phase shift of the IRS, and improving the communication quality of the system. Description of the Drawings
[0080] Figure 1 It is the system model of the present invention;
[0081] Figure 2 It is the implementation flowchart of the present invention. Detailed Embodiment
[0082] The present invention will be further described below in terms of theory and specific embodiments with reference to the accompanying drawings.
[0083] An optimization method for an IRS-assisted short-packet communication broadcast system based on DDPG is as follows:
[0084] 1. An optimization method for an IRS-assisted short-packet communication broadcast system based on DDPG, characterized in that the specific implementation steps are as follows:
[0085] Step 1: Construct a downlink broadcast system assisted by an Intelligent Reflecting Surface (IRS), including channel modeling between a Base Station (BS), an IRS, and multiple User Equipments (UEs), and introducing an actual phase shift model to simulate the hardware limitations of the IRS;
[0086] Step 2: Derive the end-to-end channel distribution according to the system model established in Step 1; establish a system optimization goal based on the statistical channel state information, that is, to maximize the minimum average signal-to-noise ratio among all users to ensure that each user can obtain good communication quality; specifically as follows:
[0087] The received signal at UEi is y i = H i x + w, where UEi represents the i-th user equipment, H i = h i Φg0f H , g0, h i respectively represent the channel coefficients between the BS and the IRS, and between the IRS and UEi, Φ represents the reflection coefficient matrix of the IRS, is the beamforming vector at the BS, f H represents the conjugate transpose of f, x represents a short-packet signal with power P s transmitted by the BS, and w is a Gaussian noise with a mean of 0 and a variance of σ2 Additive complex Gaussian white noise; According to the received signal, the instantaneous signal-to-noise ratio at the UEi side is: γ i = ρ|h i Φg0f H | 2 , where |·| represents the modulus operation, ρ = P s / σ 2 is the transmit signal-to-noise ratio; In order to enable each user to receive a reliable short packet signal, when ρ is given, maximizing the average signal-to-noise ratio E(γ i ) at the UEi side is equivalent to maximizing E(·) represents the mean operation;
[0088] Using the characteristics of circularly symmetric complex Gaussian distribution, the distribution of the channel H i can be deduced as follows: where represents the real part operation, represents the imaginary part operation, j represents the imaginary unit, obeys a normal distribution with mean μ i and variance σ i 2 , where
[0089]
[0090] ξ i are the large-scale parameters of the channels g0 and h i respectively, represents the array response at the BS, k0 and k i are the Rice factors of the channels g0 and h i respectively, N represents the number of IRS reflection elements, β(θ l ) ∈ [0, 1] is the amplitude coefficient of the l-th reflection element of the IRS, and are the line-of-sight components of the channels g0 and h i respectively;
[0091] obeys a normal distribution with mean and variance , where
[0092]
[0093] Through the above derivation, we can obtain The obtained E(|H i | 2) The distribution can be used for the design of the reward function, and the following optimization scheme is given: Using the idea of the Max-Min algorithm (C. Li, C. He, L. Jiang, and F. Liu, “Robust Beamforming Design for Max–Min SINR in MIMO Interference Channels,” IEEE Commun. Lett., vol. 20, no. 4, pp. 724–727, Apr. 2016.) to optimize the beamforming vector at the BS and the IRS phase shift to maximize the minimum average signal-to-noise ratio among all users, so that each user has a high communication quality. Therefore, the problem is modeled as:
[0094]
[0095] H i = h i Φg0f H , i ∈ {1, 2,..., K}
[0096]
[0097] -π ≤ θ l < π
[0098] where min{·} represents taking the minimum element in the array.
[0099] Step 3: Adopt the Deep Deterministic Policy Gradient (DDPG) algorithm to optimize the beamforming vector at the BS and the IRS phase shift; the DDPG model includes an actor network, a critic network, a target actor network, and a target critic network; the actor network outputs the beamforming vector and the IRS phase shift, and the critic network outputs the Q value to evaluate the quality of the action; the target actor network outputs the target action, and the target critic network outputs the target Q value, which are respectively used to train the actor network and the critic network to reduce the fluctuations during the training process;
[0100] Step 4: Continuously iteratively train the DDPG model to optimize the beamforming vector at the BS and the IRS phase shift configuration, and save the optimized parameters;
[0101] Step 5: Finally, use the optimized beamforming vector at the BS and the IRS phase shift to perform the minimum common block length analysis.
[0102] Furthermore:
[0103] The specific implementation of Step 1 is as follows:
[0104] In an IRS-assisted downlink broadcast system, assuming there is a blockage between the BS and the UEs in the system and they cannot communicate directly, a BS equipped with M antennas broadcasts short-packet signals to K single-antenna UEs (denoted as UEi, i ∈ {1, 2,..., K}) with the assistance of an IRS with N reflection units. The reflection coefficient matrix of the IRS is defined as: which represents a diagonal matrix with as the diagonal elements, θ l being the phase shift coefficient of the l-th reflection unit of the IRS, θ l ∈[-π, π), and the actual phase shift model amplitude-phase shift coupling relationship is:
[0105]
[0106] where β min ≥0, θ≥0, and τ≥0 are constants related to the specific circuit implementation. Assuming that the antennas of the BS and the reflection units of the IRS are both arranged in a uniform linear array, the array responses at the BS and the IRS are where ζ ∈ {M, N}, d is the spacing between the reflection units of the IRS or the spacing between the base station antennas, λ is the wavelength, θ is the angle of arrival (AOA) or angle of departure (AOD) of the signal. The channel coefficients between the BS and the IRS, and between the IRS and UEi are denoted as g0 and h i .
[0107] where α represents the path loss at a reference distance of 10 meters, α0 and α i are the path loss exponents, d0 is the distance from the BS to the center of the IRS, d i is the distance from the center of the IRS to UEi, and are the line-of-sight components, represents the complex domain, the AOD and AOA of the signal from the BS to the IRS are and respectively, the AOD from the IRS to UEi is θ AODi , and it can be obtained that represents the conjugate transpose of, and are the non-line-of-sight components, where each element follows a circularly symmetric complex Gaussian distribution with a mean of 0 and a variance of 1.
[0108] The specific implementation of step 3 is as follows:
[0109] Establish the Markov decision process of DDPG reinforcement learning. The beamforming vector at the BS and the IRS phase shift are the key parameters. Therefore, the beamforming vector at the BS and the IRS phase shift are used as the action vector, which is expressed as:
[0110] a t =(θ 1,t ,θ 2,t ,...,θ N,t ,|f t |,arg(f t ))
[0111] where arg(·) represents the operation of taking the argument, and the subscript t represents the time step t.
[0112] The state vector is mainly composed of the channel state, the beamforming vector at the BS, and the phase shift and amplitude of the IRS, which is expressed as:
[0113]
[0114] where
[0115] Under the statistical channel, the reward function is expressed as:
[0116] r t =min{E(|H1| 2 ),E(|H2| 2 ),...,E(|H i | 2 ),...,E(|H K | 2 )}
[0117] The specific implementation of Step 4 is as follows:
[0118] In the DDPG network framework, the Actor network and the target Actor network consist of a fully connected input layer, two hidden layers, a fully connected output layer, and the Tanh function; the Critic network and the target Critic network consist of a fully connected input layer, two hidden layers, and a fully connected output layer. Batch normalization layers and linear ReLU activation functions are used to connect between the layers of the network, and the Adam optimizer is used.
[0119] The update methods of each network are as follows:
[0120] The parameters θ of the Actor network a are updated using gradient ascent, which is expressed as:
[0121]
[0122] where α aTo update the learning rate of the Actor network, B is the sampling batch, denotes the policy gradient, q(s t ,a t :θ c ) is the input of the Critic network with parameter θ c for s t and a t to produce the current action-value function Q, π(s t :θ a ) is the input of the Actor network for s t to produce a t as the execution policy. The target Critic network generates the target return:
[0123]
[0124] where λ r ∈(0,1] is the discount factor, is the input of the target Critic network with parameter θ ct for s t+1 and a t+1 to produce the target Q value, a t+1 is the input of the target action network with parameter θ at for s t+1 to produce the action. To make the current Q value and the target Q value close, it is necessary to minimize their mean square Bellman error loss, expressed as:
[0125]
[0126] and update θ c by gradient descent, expressed as:
[0127] θ c ←θ c -α c ▽ θc L(θ c )
[0128] where α c is the learning rate for updating the evaluation network. The target network performs a soft update to weight-copy the parameters to ensure the directionality and stability of the network training process. The soft updates of the target Actor network and the target Critic are respectively expressed as:
[0129] θ ct ←τ r θ c -(1-τ r )θ ct
[0130] θ at ←τ r θa -(1 - τ r )θ at
[0131] where τ r ∈(0, 1] is the learning rate for updating the evaluation network.
[0132] After clarifying the action vector, state vector, and reward function, the specific process of the DDPG optimization algorithm is as follows:
[0133] 2 - 1. Initialize the Actor network, target Actor network, Critic network, and target Critic network;
[0134] 2 - 2. Initialize the training episode to 0;
[0135] 2 - 3. Initialize the channel state, phase shift, and amplitude of the IRS, etc.;
[0136] 2 - 4. Initialize the step size step to 0;
[0137] 2 - 5. According to the current state s t , the Actor network outputs the action a t ;
[0138] 2 - 6. Obtain the amplitude corresponding to each phase shift of the IRS unit according to the amplitude-phase shift coupling relationship;
[0139] 2 - 7. Calculate the reward for the current step as min{E(|H1| 2 ), E(|H2| 2 ),..., E(|H i | 2 ),..., E(|H K | 2 )};
[0140] 2 - 8. Obtain the new state s t based on the current state s t and the action a t+1 ;
[0141] 2 - 9. Fill the obtained experience tuple (s t , a t , r t , s t+1 ) into the buffer;
[0142] 2 - 10. If the buffer is full, update the parameters of the Actor network, target Actor network, Critic network, and target Critic network;
[0143] 2 - 11. Update the current state s t= s t+1 ;
[0144] 2 - 12. Determine whether the time step step satisfies step < max_step (where max_step represents the maximum step). If it is satisfied, return to step 2 - 5. If not, proceed to step 2 - 13;
[0145] 2 - 13. Determine whether the number of training episodes episode satisfies episode < max_episode (where max_episode represents the maximum number of rounds). If it is satisfied, return to step 2 - 3. If not, the optimization process ends, and the optimized IRS phase shift and amplitude information, beamforming vector at the BS, and network weight parameters are obtained.
[0146] Step 5: Finally, perform the minimum common block length analysis using the optimized beamforming vector at the BS and IRS phase shift.
[0147] For a given short data packet block length T, signal - to - noise ratio γ, and short data packet block error rate Block Error Rate, BLER ε, the maximum achievable rate of short - packet communication is approximately expressed as where C(γ) = log2(1 + γ) represents the Shannon capacity, V(γ)=(1-(1 + γ) -2 )(log2e) 2 represents the channel divergence, Q -1 (·) represents the inverse of the Gaussian Q - function. The physical - layer information rate of the UE is where F represents the number of information bits, and ε can be approximated as The average BLER in a wireless fading channel can be approximated as where f γ (x) is the probability density function of γ. Assume is the target average BLER for all UEs. The minimum common block length optimization algorithm flow is as follows:
[0148] (1) Given system parameters N, M, k0, k1,..., k K , d0, d1,..., d K , α, α1,..., α K , θ1,..., θ N , f, F, Initialize i to 1;
[0149] (2) Initialize the block length T, limit the block length within [M min , M max , let T - = M min , T + = M max ;
[0150] (3) Calculate the corresponding average when the block lengths are T - and T + respectively.
[0151] (4) If the calculated average BLER satisfies let ( represents the minimum block length required for reliable communication of UEi), that is, it means there is no solution;
[0152] (5) If the calculated average BLER satisfies then let
[0153] (6) If neither of the above situations is satisfied, under the condition that T + -T - ≠1, let calculate the corresponding average represents the ceiling function;
[0154] (7) When is satisfied, then let otherwise let
[0155] (8) Update T - or T + , and continue steps (6) to (7) until T + -T - =1 to stop the iteration, and output
[0156] (9) Judge whether i satisfies i < K. If it is satisfied, set i = i + 1 and enter step (2); if it is not satisfied, enter step (10);
[0157] (10) Traverse select the maximum value among them and set it as the minimum common block length T min for all UEs.
[0158] The present invention is further illustrated by the implementation process of the following specific embodiments:
[0159] An optimization method for an IRS-assisted short-packet communication broadcast system based on DDPG, the specific implementation steps are as follows:
[0160] Step 1: In an IRS-assisted downlink broadcast system, assuming there is a blockage between the BS and the UEs in the communication scenario and they cannot communicate directly, a BS equipped with M = 4 antennas broadcasts short-packet signals to K = 3 single-antenna UEs (denoted as UEi, i ∈ {1, 2,..., K}) with the assistance of an IRS with N = 64 reflection units. The reflection coefficient matrix of the IRS is defined as: which represents a diagonal matrix with as the diagonal elements, β(θ l ) is the amplitude coefficient of the l-th reflection unit of the IRS, β(θ l ) ∈ [0, 1], θ l is the phase shift coefficient of the l-th reflection unit of the IRS, θ l ∈ [-π, π), and the actual phase shift model amplitude-phase shift coupling relationship is: β min = 0.3, θ = 0.3π, τ = 1.6.
[0161] Assuming that the antennas of the BS and the reflection units of the IRS are both arranged in a uniform linear array, the responses at the BS and the IRS are where ζ ∈ {M, N}, d is the spacing between the reflection units of the IRS or the spacing between the base station antennas, λ is the wavelength, and let θ be the AOA or AOD of the signal. The channel coefficients between the BS and the IRS, and between the IRS and UEi are denoted as g0 and h i . Among them, ξ0 and ξ i are large-scale parameters, α = 0.01 represents the path loss at a reference distance of 10 meters, α0 and α i are path loss exponents, α0 = α1 = α2 = α3 = 2, d0 is the distance from the BS to the center of the IRS, d i is the distance from the center of the IRS to UEi, d0 = 40 meters, d1 = 30 meters, d2 = 40 meters, d3 = 70 meters, k0 and k i are the Rice factors of the channels g0 and h i respectively, k0 = 5, k1 = 3, k2 = 5, k3 = 7, and are the line-of-sight components, represents the complex domain, the AOD and AOA of the signal from the BS to the IRS are the AOD from the IRS to UEi is It can be obtained that represents the conjugate transpose of, and The non-line-of-sight component, where each element follows a circularly symmetric complex Gaussian distribution with a mean of 0 and a variance of 1.
[0162] Step 2: According to the system model established in Step 1, deduce the end-to-end channel distribution; establish a system optimization objective based on the statistical channel state information, that is, to maximize the minimum average signal-to-noise ratio among all users to ensure that each user can obtain good communication quality.
[0163] Step 3: Adopt the Deep Deterministic Policy Gradient (DDPG) algorithm to optimize the beamforming vector at the BS and the IRS phase shift. The DDPG model includes an action network, a critic network, a target action network, and a target critic network. The action network outputs the beamforming vector and the IRS phase shift, and the critic network outputs the Q value to evaluate the quality of the action. The target action network outputs the target action, and the target critic network outputs the target Q value, which are respectively used to train the action network and the critic network to reduce the fluctuations during the training process.
[0164] Step 4: Continuously iteratively train the DDPG model to optimize the beamforming vector at the BS and the IRS phase shift configuration, and save the optimized parameters.
[0165] Step 5: Finally, perform the minimum common block length analysis using the optimized beamforming vector at the BS and the IRS phase shift.
[0166] Under the ideal phase shift model, for UEi, the optimal phase shift value of the l-th reflection unit of the IRS is Table 1 gives the optimized IRS phase shift of the method of the present invention. Through the amplitude-phase shift coupling relationship of the actual phase shift model, the amplitude coefficients corresponding to the above two sets of phase shift values can be obtained, so as to obtain the IRS reflection coefficient matrix under the actual phase shift model. Table 2 gives the optimized beamforming vector at the BS of the method of the present invention. Assume that the number of information bits F = 300 and the target average BLER is And both use the beamforming vector at the BS given in Table 2. Through the minimum common block length optimization algorithm, it can be obtained that when using the IRS reflection coefficient matrix obtained by the method of the present invention, the minimum transmit signal-to-noise ratio ρ required for the minimum common block length T min = 200 is ρ = 55.7 dB, while when using the optimal phase shift values of UE1, UE2, and UE3 under the ideal phase shift model respectively, the minimum transmit signal-to-noise ratios ρ required for the minimum common block length T min = 200 are ρ = 99.2 dB, ρ = 99.3 dB, and ρ = 99.8 dB respectively. Through comparison, it can be seen that the IRS phase shift optimized by the method of the present invention has a significant effect on improving the performance of the IRS-assisted short packet communication broadcast system.
[0167] Table 1 shows the IRS phase shift values optimized for the specific implementation cases of the present invention
[0168]
[0169] Table 2 shows the beamforming vector values at the BS optimized for the specific embodiments of the present invention
[0170]
Claims
1. An optimization method for an IRS-assisted short packet communication broadcast system based on DDPG, characterized in that: The specific implementation steps are as follows: Step 1: Build an intelligent reflecting surface (IRS)-assisted downlink broadcast system, including channel modeling between the base station (BS), IRS, and multiple user equipment (UE), and introduce a practical phase shift model to simulate the hardware limitations of the IRS; Step 2: Based on the system model established in step 1, derive the end-to-end channel distribution; establish the system optimization goal based on the statistical channel state information, that is, maximize the minimum average signal-to-noise ratio among all users to ensure that each user can obtain good communication quality; the specific implementation is as follows: The signal received by UEi is y i =H i x+w, where UEi represents the i-th user equipment, H i =h i Φg0f H , g0, h i They represent the channel coefficients between BS and IRS, and between IRS and UEi, respectively. Φ represents the reflection coefficient matrix of IRS. is the beamforming vector at the BS, f H represents the conjugate transpose of f, and x represents the power transmitted by the BS as P s The short packet signal, w has a mean of 0 and a variance of σ 2 Additive complex Gaussian white noise; According to the received signal, the instantaneous signal-to-noise ratio at UEi is: γ i =ρ|h i Φg0f H | 2 , where |·| represents the modulus operation, ρ=P s / σ 2 is the transmission signal-to-noise ratio; in order for each user to receive a reliable short packet signal, when ρ is given, the average signal-to-noise ratio E(γ i ), which is equivalent to maximizing E(·) represents the mean operation; Using the cyclic symmetric complex Gaussian distribution characteristics, the channel H can be deduced i Distribution: in represents the real part operation, represents the operation of taking the imaginary part, j represents the imaginary unit, Subject to the mean μ i The variance is The normal distribution of ξ i They are channel g0 and channel h respectively i The large-scale parameters of represents the array response at BS, k0 and k i They are channel g0 and channel h i Rice factor, N represents the number of IRS reflection units, β(θ l )∈[0,1] is the amplitude coefficient of the lth reflection unit of IRS, and They are channel g0 and channel h respectively i The sight distance component of The mean is The variance is The normal distribution of Through the above derivation, we can get The obtained E(|H i | 2 ) can be used to design the reward function, which gives the following optimization scheme: H i =h i Φg0f H ,i∈{1,2,...,K} -π≤θ l <p Where min{·} means taking the minimum element in the array; Step 3: Use the Deep Deterministic Policy Gradient (DDPG) algorithm to optimize the beamforming vector and IRS phase shift at the BS. The DDPG model includes an action network, an evaluation network, a target action network, and a target evaluation network. The action network outputs the beamforming vector and IRS phase shift, and the evaluation network outputs the Q value, which is used to evaluate the quality of the action. The target action network outputs the target action, and the target evaluation network outputs the target Q value. These outputs are used to train the action network and the evaluation network, respectively, to reduce fluctuations during the training process. Step 4: Optimize the beamforming vector and IRS phase shift configuration at the BS by continuously iteratively training the DDPG model, and save the optimized parameters; Step 5: Finally, the minimum common block length analysis is performed using the optimized beamforming vector at the BS and the IRS phase shift.
2. According to claim 1, a DDPG-based IRS-assisted short packet communication broadcast system optimization method is characterized in that: The specific implementation of step 1 is as follows: In an IRS-assisted downlink broadcast system, assuming that there is a blockage between the BS and the UE in the system and they cannot communicate directly, a BS equipped with M antennas broadcasts short packet signals to K single-antenna UEs (denoted as UEi, i∈{1,2,...,K}) with the assistance of an IRS with N reflection units. The reflection coefficient matrix of IRS is defined as: It indicates that is a diagonal matrix with diagonal elements, θ l is the phase shift coefficient of the lth reflection unit of IRS, θ l ∈[-π,π), the actual phase shift model amplitude phase shift coupling relationship is: where β min ≥0, and τ ≥ 0 are constants related to the specific circuit implementation. Assuming that the BS antenna and the IRS reflector are arranged in a uniform linear array, the array response at the BS and IRS is Where ζ∈{M,N}, d is the spacing between IRS reflectors or the spacing between base station antennas, λ is the wavelength, and θ is the angle of arrival (AOA) or angle of departure (AOD). The channel coefficients between BS and IRS, and between IRS and UEi are represented by g0 and h i ;in α represents the path loss at a reference distance of 10 meters, α0 and α i is the path loss index, d0 is the distance from BS to IRS center, d i is the distance from the IRS center to UEi, and is the sight distance component, represents the complex domain, and the AOD and AOA of the signal from BS to IRS are and The AOD from IRS to UEi is Available express The conjugate transpose of and is the non-line-of-sight component, where each element obeys a cyclically symmetric complex Gaussian distribution with a mean of 0 and a variance of 1.
3. The optimization method of an IRS-assisted short packet communication broadcast system based on DDPG according to claim 1, characterized in that: The specific implementation of step 3 is as follows: The Markov decision process of DDPG reinforcement learning is established. The beamforming vector and IRS phase shift at the BS are key parameters. Therefore, the beamforming vector and IRS phase shift at the BS are used as action vectors and expressed as: a t =(θ 1,t ,i 2,t ,...,θ N,t ,|f t |,arg(f t )) Where arg(·) represents the angle calculation, and the subscript t represents the step size t. The state vector mainly consists of the channel state, the beamforming vector at the BS, and the phase shift and amplitude of the IRS, which can be expressed as: in Under the statistical channel, the reward function can be expressed as: r t =min{E(|H1| 2 ),E(|H2| 2 ),...,E(|H i | 2 ),...,E(|H K | 2 )} 4. The optimization method of an IRS-assisted short packet communication broadcast system based on DDPG according to claim 1, characterized in that: The specific implementation of step 4 is as follows: In the DDPG network framework, the Actor network and the target Actor network consist of a fully connected input layer, two hidden layers, a fully connected output layer, and a Tanh function; the Critic network and the target Critic network consist of a fully connected input layer, two hidden layers, and a fully connected output layer; each layer of the network is connected using a batch normalization layer and a linear ReLU activation function, and the Adam optimizer is used; The update method for each network is as follows: Actor network parameters θ a Using gradient ascent update, it is expressed as: where α a is the learning rate for updating the Actor network, B is the sampling batch, represents the policy gradient, q(s t ,a t :θ c ) is the parameter θ c The Critic network input s t and a t The current action-value function Q generated after, π(s t :θ a ) is the Actor network input s t Generate a t implementation strategy; The target critic network generates target returns: Among them, λ r ∈(0,1] is the discount factor, The parameter is θ ct The target Critic network input s t+1 and a t+1 The target Q value generated after t+1 The parameter is θ at The target action network input s t+1 To make the current Q value and the target Q value close, it is necessary to minimize their mean square Bellman error loss, which is expressed as: And update θ by gradient descent c , expressed as: where α c To update the learning rate of the evaluation network. The target network performs weighted replication of parameters through soft update to ensure the directionality and stability of the network training process. The soft updates of the target Actor network and the target Critic are expressed as: i ct ←t r i c -(1-t r )i ct i at ←t r i a -(1-t r )i at where τ r ∈(0,1] is the learning rate for updating the evaluation network.
5. The optimization method of the IRS-assisted short packet communication broadcast system based on DDPG according to claim 4 is characterized in that: After clarifying the action vector, state vector, and reward function, the specific process of the DDPG optimization algorithm is as follows: 2-1. Initialize the Actor network, target Actor network, Critic network and target Critic network; 2-2. Initialize the training episode to 0; 2-3. Initialize channel status, IRS phase shift and amplitude, etc.; 2-4. Initialize the step size to 0; 2-5. According to the current status t , the Actor network outputs action a t ; 2-6. According to the amplitude-phase-shift coupling relationship, the amplitude corresponding to each phase shift of the IRS unit is obtained; 2-7. Calculate the reward min{E(|H1| 2 ),E(|H2| 2 ),...,E(|H i | 2 ),...,E(|H K | 2 )}; 2-8. According to the current status s t and action a t Get the new state s t+1 ; 2-9. The experience tuple (s t ,a t ,r t ,s t+1 ) fill into buffer; 2-10. If the buffer is full, update the parameters of the Actor network, target Actor network, Critic network, and target Critic network. 2-11. Update the current status s t =s t+1 ; 2-12. Determine whether the time step step satisfies step<max_step (max_step represents the maximum step length). If so, return to step 2-5. If not, proceed to step 2-13. 2-13. Determine whether the number of training rounds episode satisfies episode<max_episode (max_episode represents the maximum number of rounds). If so, return to step 2-3. If not, the optimization process ends and the optimized IRS phase shift and amplitude information, beamforming vector at BS and network weight parameters are obtained.
6. The optimization method of the IRS-assisted short packet communication broadcast system based on DDPG according to claim 4 is characterized in that: The implementation of step 5 is as follows: For a given short data packet block length T, signal-to-noise ratio γ and short data packet block error rate Block Error Rate, BLERε, the maximum achievable rate of short packet communication is approximately expressed as Where C(γ) = log2(1+γ) represents Shannon capacity, V(γ) = (1-(1+γ) -2 )(log2 e) 2 represents the channel divergence, Q -1 (·) represents the inverse of the Gaussian Q function. The physical layer information rate of the UE is Where F represents the number of information bits, and ε can be approximated as The average BLER in a wireless fading channel can be approximated as where f γ (x) is the probability density function of γ; assuming For the target average BLER of all UEs, the minimum common block length optimization algorithm flow is as follows: (1) Given the system parameters N, M, k0, k1, ..., k K ,d0,d1,...,d K ,α,α1,...,α K ,θ1,...,θ N 、f、F、 Initialize i to 1; (2) Initialize the block length T and limit the block length to [M min ,M max ], let T - =M min , T + =M max ; (3) Calculate the block length as T - and T + The corresponding average (4) If the calculated average BLER satisfies make ( represents the minimum block length required for reliable communication of UEi), which means there is no solution; (5) If the calculated average BLER satisfies Then (6) If none of the above conditions are met, then + -T - ≠1, let Calculate the corresponding average represents the ceiling function; (7) When satisfied Then Otherwise, (8) Update T - or T + , continue steps (6) to (7) until T + -T - =1 to stop iteration and output (9) Determine whether i satisfies i<K. If so, set i=i+1 and go to step (2); if not, go to step (10); (10) Traversal Select the maximum value and set it as the minimum common block length T of all UEs min .
Citation Information
Patent Citations
Intelligent reflecting surface communication method based on distributed reinforcement learning
CN115802379A
Beam forming design method for STAR-RIS-assisted multi-user MISO URLLC system
CN117375684A
Deep reinforcement learning-based IRS assisted Internet of Vehicles system resource allocation method
CN117750527A
DDPG-based IRS-assisted cognitive radio system beam forming method
CN117767987A
IRS-SWIPT system resource allocation method based on deep reinforcement learning
CN118102327A
Cited By
TD3-based IRS auxiliary short packet communication system optimization method in interference environment
CN121055986A