Double-IRS auxiliary short packet communication system optimization method based on PPO
By jointly optimizing the beamforming of the base station and dual IRS using the PPO algorithm, the modeling and optimization complexity of the dual IRS-assisted communication system is solved, achieving high reliability and low latency short packet communication, and improving the average reachability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LANZHOU UNIV
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing dual intelligent reflector (IRS) assisted communication systems are difficult to model accurately under actual hardware constraints, have high optimization complexity, and have strict requirements for latency and reliability in short packet communication scenarios. Conventional deep reinforcement learning algorithms are difficult to meet the stability and efficiency requirements of high-dimensional and continuous control problems.
A near-end policy optimization (PPO) approach is adopted to jointly optimize the active beamforming of the base station and the passive beamforming of the dual IRS. By constructing an actual amplitude phase shift model, the beamforming parameters of the BS and IRS are optimized using the pruning strategy, importance sampling and entropy regularization of the PPO algorithm to achieve high reliability and low latency short packet communication.
It effectively improves the performance of the communication system, realizes highly reliable and low-latency short packet communication, solves the high-dimensional and non-convex optimization bottleneck in the dual IRS auxiliary system, and improves the average reachability of the system.
Smart Images

Figure CN121907294A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of next-generation mobile communication technology. It is a new method for optimizing intelligent reflective surface-assisted short packet communication systems using deep reinforcement learning. Specifically, it refers to an optimization method for dual intelligent reflective surface (IRS)-assisted short packet communication systems based on proximal policy optimization (PPO). Background Technology
[0002] IRS, as a key candidate technology for next-generation mobile communication, is a novel electromagnetic material composed of a large number of passive reflective elements. It can dynamically control parameters such as the phase and amplitude of incident electromagnetic waves in a programmable manner, thereby intelligently reconstructing the wireless propagation environment. This technology provides a new dimension for improving the spectral efficiency and energy efficiency of communication systems. Its low power consumption and easy deployment characteristics make it promising for future networks (X. Fang, "Joint channelestimation algorithm for IRS-Assisted multi-user MIMO systems," in IEEE Commun. Lett., vol. 28, no. 2, pp. 367-371, Feb. 2024).
[0003] The practical application of IRS technology faces two major challenges. The primary challenge lies in the problem of accurate modeling under actual hardware constraints. Existing research is mostly based on an ideal amplitude-phase shift model where the amplitude and phase shift of the reflecting unit can be independently controlled. This model assumes that the amplitude of each reflecting unit is unaffected by phase shift and is generally set to 1. However, S. Abeywickrama et al. pointed out in their research (S. Abeywickrama, R. Zhang, Q. Wu, and C. Yuen, “Intelligent reflecting surface: Practical phase shift model and beamforming optimization,” IEEE Trans. Commun., vol. 68, no. 9, pp. 5849–5863, Sep. 2020.) that the characteristics of physically realizable IRS reflecting units differ from the assumptions of the ideal amplitude-phase shift model. The amplitude and phase shift of their reflection coefficient are coupled, and the model describing this characteristic is called the “practical amplitude-phase shift model.” This coupling relationship leads to severe performance degradation of the IRS reflection coefficient optimized based on the ideal amplitude-phase shift model in physically realizable IRS-assisted wireless communication systems. Therefore, constructing an optimization framework that conforms to the actual amplitude and phase shift model has become crucial for advancing the implementation of this technology. Secondly, the performance limitations of a single IRS in complex wireless environments have prompted research to shift towards the more promising dual-IRS assisted communication architecture. In scenarios with severe obstruction or requiring wide-area coverage, a single IRS may be unable to establish an effective communication link due to its fixed location or limited field of view. Dual IRS, by coordinating the deployment of two IRSs, can form a longer cascaded reflection path, thereby significantly extending the coverage area.
[0004] The performance gains brought by dual IRS come at the cost of extremely high optimization complexity. The system requires the joint optimization of hundreds of reflection units on the dual IRS, with the scale of its adjustable parameters growing exponentially, forming an ultra-high-dimensional, strongly coupled continuous control problem. Furthermore, under the constraint of the actual amplitude-phase shift model, the optimization problem exhibits high non-convexity and nonlinearity. In applications such as short-packet communication, which are extremely sensitive to latency and reliability, these challenges are further amplified, posing severe tests to the optimization algorithm in terms of both computational complexity and real-time performance.
[0005] Facing the optimization challenges posed by dual IRS, Deep Reinforcement Learning (DRL) has shown potential in solving such high-dimensional and continuous control problems. However, communication systems have extremely high requirements for the stability of policy updates and sampling efficiency, which conventional DRL algorithms struggle to meet. Summary of the Invention
[0006] This invention provides an optimization method for a dual IRS-assisted short packet communication system based on PPO. The method utilizes the PPO algorithm to jointly optimize the active beamforming of the base station and the passive beamforming of the dual IRS, thereby improving the performance of the communication system and achieving highly reliable, low-latency communication.
[0007] The objective of this invention is achieved through the following technical solution:
[0008] An optimization method for a PPO-based dual IRS-assisted short packet communication system includes the following steps:
[0009] Step 1: Construct a dual-intelligent reflector IRS-assisted short packet communication system. This system includes a base station (BS) with N antennas, a single-antenna user equipment (UE), and two [unclear - possibly referring to devices with different antennas]. The IRS of each reflective unit; the two IRS are denoted as IRS1 and IRS2 respectively; due to obstruction, the BS can only communicate with the UE with the assistance of the two IRS; and a physically realizable amplitude and phase shift model is introduced to construct the amplitude and phase shift constraint relationship.
[0010] Step 2: Under Ricean fading channel conditions, a system optimization model is constructed using Statistical Channel State Information (CSI) to maximize the average reachability rate of the UE in short packet communication, ensuring that the UE achieves a high transmission rate under high reliability and low latency communication constraints; the specific implementation is as follows:
[0011] The BS transmits a signal that reaches the UE via a dual IRS reflection link; the UE receives the signal as follows: ,in The transmit power at BS, This is the power-normalized signal transmitted by the BS to the UE. The channel coefficients for the BS transmitted signal reaching the UE link after cascaded reflections through IRS1 and IRS2. To follow the pattern with a mean of 0 and a variance of Additive white Gaussian noise; instantaneous signal-to-noise ratio (SNR) of the UE in this scenario. Represented as:
[0012]
[0013] in This indicates the modulo operation.
[0014] Short packet communication under a given transport block length L and instantaneous received signal-to-noise ratio Block error rate BLER The achievable rate of a system in an additive white Gaussian noise channel is expressed as:
[0015]
[0016] in Indicates Shannon capacity, Represents channel divergence, under high signal-to-noise ratio. , express The inverse function; by jointly optimizing the normalized beamforming vector of the BS antenna array. IRS1 reflection coefficient matrix and IRS2 reflection coefficient matrix , Representing the complex field, we construct an optimization model for a dual IRS-assisted short packet communication system with the objective of maximizing the average reachable rate of the UE. The specific optimization problem is expressed as:
[0017]
[0018] in This represents the maximum average achievable rate of UE short packet communication under Ricean fading channels, expressed in bits per symbol (BPS). superscript Let be the conjugate transpose of the matrix, satisfying , Describes the 2-norm of a vector. These are the amplitude coefficient and phase shift coefficient of the BS antenna array factor, respectively. ,definition For IRS The ( The reflection coefficient of each reflecting unit. and They represent the first The IRS The amplitude coefficient and phase shift of each reflecting unit in the actual amplitude and phase shift model and The coupling relationship is , and These are constants related to the specific circuit implementation. It is the minimum amplitude. yes arrive Horizontal distance, This represents the mean calculation; according to existing technology, it can be obtained... for:
[0019] .
[0020] Step 3: The PPO algorithm, employing a near-end policy optimization, jointly optimizes the BS active beamforming vector and the IRS passive beamforming vector. Key design elements of PPO include state space, action space, and reward function. Its network model comprises a policy network and a value network. The policy network outputs the BS active beamforming vector and the IRS passive beamforming vector, while the value network outputs the state value function to estimate the advantage function. The network is trained by introducing pruning strategies, importance sampling, and entropy regularization. Through iterative training of the PPO model, the BS active beamforming vector and the IRS passive beamforming vector are jointly optimized. Using the optimized joint beamforming scheme, the maximum average reachable rate of the system is analyzed.
[0021] The specific implementation method of step 1 is as follows:
[0022] Assuming that BS and IRS are both uniform linear arrays (ULA), then the array responses at BS, IRS1, and IRS2 are: ,in , It is the spacing between adjacent units. It is the carrier wavelength. It refers to the signal's angle of arrival or departure angle; the departure angle and angle of arrival for the signal from BS to IRS1 are respectively... The departure angle and arrival angle from IRS1 to IRS2 are respectively The departure angle from IRS2 to UE is The channel gain of the BS transmitted signal reaching the UE link after cascaded reflections via IRS1 and IRS2 is: ,in The channel coefficients between BS and IRS1 The channel coefficients between IRS1 and IRS2 The channel coefficients between IRS2 and UE. It is the path loss exponent corresponding to the channel. This is the path loss at a unit reference distance of 1 meter. The distance from the center of BS to the center of IRS1. The distance from the center of IRS1 to the center of IRS2. The distance from the IRS2 center to the UE. , and They are respectively , and The corresponding Rice factor, The line-of-sight component is , The line-of-sight component is , The line-of-sight component is ; Non-line-of-sight components , Non-line-of-sight components and Non-line-of-sight components The elements in the set are independent and all follow a cyclic symmetric complex Gaussian random variable with mean 0 and variance 1.
[0023] The channel gain can be derived. Follow the mean
[0024]
[0025] variance is
[0026] The complex Gaussian distribution, in which , ,but .
[0027] The specific implementation of step 3 is as follows:
[0028] For the design of key elements of PPO, in order to fully consider the state space's description of environmental information, it is designed according to CSI as follows:
[0029]
[0030] in , , , , , Take the argument of the complex elements in the matrix. The current time step, superscript matrix express exist The time step value indicates that the agent's action space includes the amplitude coefficient, phase shift coefficient, and phase offset of the dual IRS of the BS beamforming vector. .
[0031] For active and passive beamforming designs based on statistical CSI, the reward function can be expressed as:
[0032]
[0033] The PPO algorithm's network model includes a policy network and a value network; its training process first initializes the policy network. and value network , Represents all learnable parameters of the policy network. Represents all learnable parameters of the value network, subscript or The policy network or value network is represented by the current parameter vector. or Perform parameterization; For input status, To output the action, For the strategy in the state Select action The probability, Input state The state-value function at time; creating an experience buffer; initializing the state at the start of training. In time step In the middle, using the current policy network Interacting with the environment Given the parameters of the action network in the previous trajectory, obtain the current state. Output the probability distribution of actions and obtain the actions by sampling according to the probability. Calculate the reflection coefficient of the IRS and the beamforming vector of the BS, and calculate the reward based on the reward function. And update status Meanwhile, the experience buffer collects complete interaction data for that time step, including state transition information and the actions taken by the current policy. The logarithmic probability of a tuple When the experience buffer reaches its preset capacity, multiple rounds of policy updates are performed on this batch of data. In each round of updates, the policy network... Based on the current state Recalculate the log probability of the new strategy And calculate the importance sampling ratio with the log probability of the old strategy. The pruning mechanism limits the scope of policy updates, while the value network estimates the value function based on empirical data. and dominance function Finally, update the policy network parameters. and value network parameters It also clears the experience buffer to prepare for the next data collection.
[0034] Generalized advantage estimation is used to obtain the advantage function.
[0035]
[0036] in , As a discount factor, In generalized dominance estimation, the parameter that balances variance and bias is... These are the parameters of the value network in the previous trajectory.
[0037] Importance sampling weights Constrain the update magnitude of the policy network
[0038]
[0039] The policy network calculates the loss by pruning the objective function.
[0040]
[0041] in This is a pruning function used to restrict policy updates to a specified range. Inside, This is the threshold for the cropping range.
[0042] Entropy regularization is used to avoid getting stuck in local optima during training. Based on the current strategy Calculated entropy
[0043]
[0044] The policy network uses a gradient algorithm in each After each trajectory is completed, the policy gradient is calculated based on the cumulative reward of the trajectory, and the parameters of the policy network are updated.
[0045]
[0046] in The learning rate of the policy network. The entropy coefficient, This represents gradient operation.
[0047] Value networks use the mean squared error function to calculate their loss function:
[0048]
[0049] The parameters of the value network are updated using backpropagation via gradient descent.
[0050]
[0051] in The learning rate of the value network.
[0052] The specific process of the PPO algorithm is as follows:
[0053] 4-1. Initialize policy network parameters and value network parameters Initialize relevant hyperparameters: learning rate and Discount Factor GAE parameters , clipping threshold And entropy coefficient .
[0054] 4-2. Initialize the training round index and set the maximum training round.
[0055] 4-3. Initialize the environment state, generate statistical CSI under Ricean fading channel, and initialize the system time step. And create an empty experience buffer.
[0056] 4-4. Obtain the system status at the current time step ,Will Input a policy network and output the action probability distribution under the current policy. .
[0057] 4-5. Sampling is performed based on the probability distribution to obtain the current action. Record the logarithmic probability at this time. .
[0058] 4-6. Perform the actions Mapping to the physical environment: Calculate the reflection amplitude corresponding to the phase shift of the IRS unit based on the actual amplitude phase shift model, and generate joint beamforming.
[0059] 4-7. Interacting with the environment: Calculate the average reachability of the current system, calculate the reward for the current step according to the reward function, and observe the state at the next time step. .
[0060] 4-8. The obtained experience tuples Fill the experience buffer.
[0061] 4-9. Update the current status Update time step .
[0062] 4-10. Determine if the experience buffer has reached the preset capacity: If not, return to step 3-4 to continue collecting data; if it has, proceed to step 3-11 to start updating network parameters.
[0063] 4-11. Estimating State Value Using Value Networks And calculate the advantage function according to the generalized advantage estimation formula. and the target return value.
[0064] 4-12. Perform multiple rounds of iterative updates on the data in the buffer: calculate the ratio of the old and new policies, and update the policy network parameters. and value network parameters .
[0065] 4-13. After the parameter update is complete, clear the experience buffer and assign the updated strategy parameters to the old strategy parameters. .
[0066] 4-14. Update the training rounds and determine whether the current round is less than the maximum round: if it is, return to step 3-3 to start a new round of training; if it is not, the optimization process ends and the optimization results and network weight parameters are output.
[0067] This invention utilizes the PPO algorithm to jointly optimize the active beamforming of the base station and the passive beamforming of the dual IRS, effectively ensuring the stability and convergence reliability of policy updates. It provides a new approach to addressing the high-dimensional, non-convex, and practically constrained optimization bottlenecks in dual IRS-assisted short packet communication systems. Furthermore, it effectively improves the performance of the communication system, achieving highly reliable, low-latency communication. Attached Figure Description
[0068] Figure 1 This is the system model of the present invention;
[0069] Figure 2 This is a schematic diagram of the PPO algorithm framework of the present invention;
[0070] Figure 3 This is a flowchart illustrating the implementation of the present invention;
[0071] Figure 4 This is to compare the maximum average reachable rate of the system optimized using the method described in this invention with the maximum average reachable rate of the system using the baseline scheme. , , , , dBm, dBm, , centimeter, centimeter, rice, rice, rice, , radian, radian, radian, radian, radian, , , radian, symbol, . Detailed Implementation
[0072] The present invention will be further described below with reference to the accompanying drawings, both in terms of theory and specific implementation.
[0073] An optimization method for a PPO-based dual IRS-assisted short packet communication system, the specific steps of which are as follows:
[0074] Step 1: Construct a dual-intelligent-reflecting-surface (IRS) assisted short packet communication system. This system includes a Base Station (BS) with N antennas, a User Equipment (UE) with a single antenna, and two [other components / equipment]... The IRS of each reflective unit; the two IRS are denoted as IRS1 and IRS2 respectively; due to obstruction, the BS can only communicate with the UE with the assistance of the two IRS; and a physically realizable amplitude and phase shift model is introduced to construct the amplitude and phase shift constraint relationship.
[0075] Step 2: Under Ricean fading channel conditions, utilize Statistical Channel State Information (CSI) to construct a system optimization model aimed at maximizing the average reachable rate of the UE in short packet communication, ensuring that the UE achieves a high transmission rate under high reliability and low latency communication constraints; the specific implementation is as follows:
[0076] The BS transmits a signal that reaches the UE via a dual IRS reflection link; the UE receives the signal as follows: ,in The transmit power at BS, This is the power-normalized signal transmitted by the BS to the UE. The channel coefficients for the BS transmitted signal reaching the UE link after cascaded reflections through IRS1 and IRS2. To follow the pattern with a mean of 0 and a variance of The additive white Gaussian noise. In this scenario, the instantaneous signal-to-noise ratio (SNR) received by the UE is... Represented as:
[0077]
[0078] in This indicates the modulo operation.
[0079] Short packet communication under a given transport block length L and instantaneous received signal-to-noise ratio Block Error Rate (BLER) The achievable rate of a system in an additive white Gaussian noise channel is expressed as:
[0080]
[0081] in Indicates Shannon capacity, Represents channel divergence, under high signal-to-noise ratio. , express The inverse function; by jointly optimizing the normalized beamforming vector of the BS antenna array. IRS1 reflection coefficient matrix and IRS2 reflection coefficient matrix , Representing the complex field, we construct an optimization model for a dual IRS-assisted short packet communication system with the objective of maximizing the average reachable rate of the UE. The specific optimization problem is expressed as:
[0082]
[0083] in This represents the maximum average achievable rate of UE short packet communication under Ricean fading channels, expressed in bits per symbol (BPS). superscript Let be the conjugate transpose of the matrix, satisfying , Describes the 2-norm of a vector. These are the amplitude coefficient and phase shift coefficient of the BS antenna array factor, respectively. ,definition For IRS The ( The reflection coefficient of each reflecting unit. and They represent the first The IRS The amplitude coefficient and phase shift of each reflecting unit in the actual amplitude and phase shift model and The coupling relationship is , and These are constants related to the specific circuit implementation. It is the minimum amplitude. yes arrive Horizontal distance, This represents the mean calculation; according to existing technology, it can be obtained... for:
[0084]
[0085] Step 3: Proximal Policy Optimization (PPO) is employed, where the PPO algorithm jointly optimizes the BS active beamforming vector and the IRS passive beamforming vector. Key design elements of PPO include the state space, action space, and reward function; its network model comprises a policy network and a value network. The policy network outputs the BS active beamforming vector and the IRS passive beamforming vector, while the value network outputs the state value function to estimate the advantage function. The network is trained using pruning strategies, importance sampling, and entropy regularization. Through iterative training of the PPO model, the BS active beamforming vector and the IRS passive beamforming vector are jointly optimized. Using the optimized joint beamforming scheme, the maximum average reachability of the system is analyzed.
[0086] Furthermore, the specific implementation method of step 1 is as follows:
[0087] Assuming that BS and IRS are both uniform linear arrays (ULA), then the array responses at BS, IRS1, and IRS2 are: ,in , It is the spacing between adjacent units. It is the carrier wavelength. This refers to the signal's angle of arrival or departure angle. The departure angle and angle of arrival for the signal from BS to IRS1 are respectively... The departure angle and arrival angle from IRS1 to IRS2 are respectively The departure angle from IRS2 to UE is The channel gain of the BS transmitted signal reaching the UE link after cascaded reflections via IRS1 and IRS2 is: ,in The channel coefficients between BS and IRS1 The channel coefficients between IRS1 and IRS2 The channel coefficients between IRS2 and UE. It is the path loss exponent corresponding to the channel. This is the path loss at a unit reference distance of 1 meter. The distance from the center of BS to the center of IRS1. The distance from the center of IRS1 to the center of IRS2. The distance from the IRS2 center to the UE. , and They are respectively , and The corresponding Rice factor, The line-of-sight component is , The line-of-sight component is , The line-of-sight component is ; Non-line-of-sight components , Non-line-of-sight components and Non-line-of-sight components The elements in the set are independent and all follow a cyclic symmetric complex Gaussian random variable with mean 0 and variance 1.
[0088] The channel gain can be derived. Follow the mean
[0089]
[0090] variance is
[0091] The complex Gaussian distribution, in which , ,but .
[0092] The specific implementation of step 3 is as follows:
[0093] For the design of key elements of PPO, in order to fully consider the state space's description of environmental information, it is designed according to CSI as follows:
[0094]
[0095] in , , , , , Take the argument of the complex elements in the matrix. The current time step, superscript matrix express exist The time step value indicates that the agent's action space includes the amplitude coefficient, phase shift coefficient, and phase offset of the dual IRS of the BS beamforming vector. .
[0096] For active and passive beamforming designs based on statistical CSI, the reward function can be expressed as:
[0097]
[0098] The PPO algorithm's network model consists of a policy network and a value network. Its training process begins by initializing the policy network. and value network , Represents all learnable parameters of the policy network. Represents all learnable parameters of the value network, subscript or The policy network or value network is represented by the current parameter vector. or Perform parameterization. For input status, To output the action, For the strategy in the state Select action The probability, Input state The state value function at time.
[0099] Create an experience buffer; initialize the state at the start of training. In time step In the middle, using the current policy network Interacting with the environment ( (The parameters of the action network in the previous trajectory) are used to obtain the current state. Output the probability distribution of actions and obtain the actions by sampling according to the probability. Calculate the reflection coefficient of the IRS and the beamforming vector of the BS, and calculate the reward based on the reward function. And update status Meanwhile, the experience buffer collects complete interaction data for that time step, including state transition information and the actions taken by the current policy. The logarithmic probability of a tuple When the experience buffer reaches its preset capacity, multiple rounds of policy updates are performed on this batch of data. In each round of updates, the policy network... Based on the current state Recalculate the log probability of the new strategy And calculate the importance sampling ratio with the log probability of the old strategy. The pruning mechanism limits the scope of policy updates, while the value network estimates the value function based on empirical data. and dominance function Finally, update the policy network parameters. and value network parameters It also clears the experience buffer to prepare for the next data collection.
[0100] Generalized advantage estimation is used to obtain the advantage function.
[0101]
[0102] in , As a discount factor, In generalized dominance estimation, the parameter that balances variance and bias is... These are the parameters of the value network in the previous trajectory.
[0103] Importance sampling weights Constrain the update magnitude of the policy network
[0104]
[0105] The policy network calculates the loss by pruning the objective function.
[0106]
[0107] in This is a pruning function used to restrict policy updates to a specified range. Inside, This is the threshold for the cropping range.
[0108] Entropy regularization is used to avoid getting stuck in local optima during training. Based on the current strategy Calculated entropy
[0109]
[0110] The policy network uses a gradient algorithm in each After each trajectory is completed, the policy gradient is calculated based on the cumulative reward of the trajectory, and the parameters of the policy network are updated.
[0111]
[0112] in The learning rate of the policy network. The entropy coefficient, This represents gradient operation.
[0113] Value networks use the mean squared error function to calculate their loss function:
[0114]
[0115] The parameters of the value network are updated using backpropagation via gradient descent.
[0116]
[0117] in The learning rate of the value network.
[0118] The specific process of the PPO algorithm is as follows:
[0119] 4-1. Initialize policy network parameters and value network parameters Initialize relevant hyperparameters: learning rate and Discount Factor GAE parameters , clipping threshold And entropy coefficient ;
[0120] 4-2. Initialize the training round index and set the maximum number of training rounds;
[0121] 4-3. Initialize the environment state, generate statistical CSI under Ricean fading channel, and initialize the system time step. and create an empty experience buffer;
[0122] 4-4. Obtain the system status at the current time step ,Will Input a policy network and output the action probability distribution under the current policy. ;
[0123] 4-5. Sampling is performed based on the probability distribution to obtain the current action. Record the logarithmic probability at this time. ;
[0124] 4-6. Perform the actions Mapping to the physical environment: Calculate the reflection amplitude corresponding to the phase shift of the IRS unit based on the actual amplitude phase shift model, and generate joint beamforming;
[0125] 4-7. Interacting with the environment: Calculate the average reachability of the current system, calculate the reward for the current step according to the reward function, and observe the state at the next time step. ;
[0126] 4-8. The obtained experience tuples Fill the experience buffer;
[0127] 4-9. Update the current status Update time step ;
[0128] 4-10. Determine if the experience buffer has reached the preset capacity: If not, return to step 3-4 to continue collecting data; if it has, proceed to step 3-11 to start updating network parameters.
[0129] 4-11. Estimating State Value Using Value Networks And calculate the advantage function according to the generalized advantage estimation formula. and target return value;
[0130] 4-12. Perform multiple rounds of iterative updates on the data in the buffer: calculate the ratio of the old and new policies, and update the policy network parameters. and value network parameters ;
[0131] 4-13. After the parameter update is complete, clear the experience buffer and assign the updated strategy parameters to the old strategy parameters. ;
[0132] 4-14. Update the training rounds and determine whether the current round is less than the maximum round: if it is, return to step 3-3 to start a new round of training; if it is not, the optimization process ends and the optimization results and network weight parameters are output.
[0133] The specific implementation process of this invention is as follows:
[0134] An optimization method for a PPO-based dual IRS-assisted short packet communication system is described below:
[0135] Step 1: Construct a dual-intelligent reflector IRS-assisted short packet communication system. This system includes a base station (BS) with N=4 antennas, a single-antenna user equipment (UE), and two antennas respectively equipped with... One and The IRS of each reflector element is represented as IRS1 and IRS2. Due to obstruction, the BS can only communicate with the UE with the assistance of two IRSs. A physically realizable amplitude and phase shift model of the IRS is introduced to construct amplitude and phase shift constraints. Assuming that both the BS and IRS are uniform linear arrays (ULA), the array responses at BS, IRS1, and IRS2 are... ,in , Centimeters is the distance between the centers of adjacent reflector elements or between BS antennas. Centimeters is the carrier wavelength. This refers to the signal arrival angle from the BS to the IRS or the signal departure angle from both IRSs to the UE. Therefore, the signal transmission angle and arrival angle from the BS to IRS1 are respectively... The emission angle and arrival angle from IRS1 to IRS2 are respectively The transmission angle from IRS2 to the UE is In radians; the channel gain of the BS-transmitted signal reaching the UE link after reflection from IRS1 and IRS2 is... ,in The channel coefficients between BS and IRS1 The channel coefficients between IRS1 and IRS2 The channel coefficients between IRS2 and UE. Denotes the field of complex numbers, where It is the path loss exponent corresponding to the channel. This is the path loss at a unit reference distance of 1 meter. Meters represent the distance from the center of the BS to the center of the IRS. Meters represent the distance from the center of IRS1 to the center of IRS2. Distance from the center of MiIRS2 to the UE , and They are respectively , and The corresponding Rice factor, The line-of-sight component is , The line-of-sight component is , The line-of-sight component is ; Non-line-of-sight components , Non-line-of-sight components and Non-line-of-sight components The elements in the vector are independent and all follow a cyclic symmetric complex Gaussian random variable with mean 0 and variance 1; the IRS reflection coefficient vector. ,definition For IRS No. The reflection coefficient of each IRS reflection unit in the actual amplitude phase shift model and The coupling relationship is , , , radian.
[0136] Step 2: Under Ricean fading channel conditions, use Statistical Channel State Information (CSI) to construct a system optimization model with the goal of maximizing the average reachability rate of the UE in short packet communication, so as to ensure that the UE achieves a high transmission rate under the constraints of high reliability and low latency communication.
[0137] Step 3: Proximal Policy Optimization (PPO) is employed, where the PPO algorithm jointly optimizes the BS active beamforming vector and the IRS passive beamforming vector. Key design elements of PPO include the state space, action space, and reward function; its network model comprises a policy network and a value network. The policy network outputs the BS active beamforming vector and the IRS passive beamforming vector, while the value network outputs the state value function to estimate the generalized advantage function. The network is trained by incorporating pruning policies, importance sampling, and entropy regularization.
[0138] In the proposed PPO framework, the training rounds are set to 2000, with each round containing 300 time steps, and the batch size for each training iteration is 256 trajectories. The number of neurons in each layer of the policy network and value network in PPO has been given above, as have the learning rates for the policy network and value network. The threshold of the cropping range Entropy coefficient setting Discount factor In generalized dominance estimation, the parameters that balance variance and bias are... Use the Adam optimizer.
[0139] Tables 1 and 2 present the optimized dual IRS phase shift values. First, the optimal reflection phase shift of the IRS unit is obtained through training. Then, the reflection amplitude is calculated using an actual amplitude phase shift model, leading to the optimized IRS reflection coefficient matrix. Table 3 presents the optimized beamforming vector at BS. Figure 4 This paper compares the average achievable rate of the system in this invention with that of a benchmark scheme. In the benchmark scheme, an ideal amplitude-phase shift model is used to optimize the IRS reflection coefficient. In the benchmark scheme, the amplitude of the reflection units of the two IRSs is set to 1, and the optimal phase shift value of the i-th reflection unit of IRS1 is... The optimal phase shift value of the i-th reflecting unit of IRS2 is The beamforming vector at BS is obtained based on the statistical CSI maximum ratio transmission. (Y. Jia, C. Ye and Y. Cui, “Analysis and Optimization of an Intelligent Reflecting Surface-Assisted System With Interference,” IEEE Trans. Wireless Commun, vol. 19, no. 12, pp. 8068-8082, Dec. 2020). Assuming a transport block length of... Symbols, block error rate The maximum average reachable rate of the system of this invention is significantly better than the benchmark scheme after optimization. For example, when the training rounds are 3000, the maximum average reachable rate of the system of this invention is 7.28 BPS, while that of the benchmark scheme is 6.45 BPS. The comparison results show that the BS active beamforming and dual IRS passive beamforming optimized by this invention can significantly improve the maximum average reachable rate of the dual IRS assisted short packet communication system.
[0140] Table 1
[0141]
[0142] Table 2
[0143]
[0144] Table 3
[0145]
Claims
1. An optimization method for a PPO-based dual IRS-assisted short packet communication system, characterized in that, Includes the following steps: Step 1: Construct a dual-intelligent reflector IRS-assisted short packet communication system. This system includes a base station (BS) with N antennas, a single-antenna user equipment (UE), and two [unclear - possibly referring to antennas or antennas]. IRS of each reflective unit; The two IRSs are designated as IRS1 and IRS2. Due to obstruction, the BS can only communicate with the UE with the assistance of the two IRSs. Furthermore, a physically realizable IRS actual amplitude and phase shift model is introduced to construct the amplitude and phase shift constraint relationship; Step 2: Under Ricean fading channel conditions, a system optimization model is constructed using Statistical Channel State Information (CSI) to maximize the average reachability rate of the UE in short packet communication, ensuring that the UE achieves a high transmission rate under high reliability and low latency communication constraints; the specific implementation is as follows: The BS transmits a signal that reaches the UE via a dual IRS reflection link; the UE receives the signal as follows: ,in The transmit power at BS, This is the power-normalized signal transmitted by the BS to the UE. The channel coefficients for the BS transmitted signal reaching the UE link after cascaded reflections through IRS1 and IRS2. To follow the pattern with a mean of 0 and a variance of Additive white Gaussian noise; instantaneous signal-to-noise ratio (SNR) of the UE in this scenario. Represented as: ; in This indicates the modulo operation; Short packet communication under a given transport block length L and instantaneous received signal-to-noise ratio Block error rate BLER The achievable rate of a system in an additive white Gaussian noise channel is expressed as: ; in Indicates Shannon capacity, Represents channel divergence at high signal-to-noise ratios. , express The inverse function; by jointly optimizing the normalized beamforming vector of the BS antenna array. IRS1 reflection coefficient matrix and IRS2 reflection coefficient matrix , Representing the complex field, we construct an optimization model for a dual IRS-assisted short packet communication system with the objective of maximizing the average reachable rate of the UE. The specific optimization problem is expressed as: ; in This represents the maximum average achievable rate of UE short packet communication under Ricean fading channels, expressed in bits per symbol (BPS). superscript Let be the conjugate transpose of the matrix, satisfying , Describes the 2-norm of a vector. These are the amplitude coefficient and phase shift coefficient of the BS antenna array factor, respectively. ,definition For IRS The ( The reflection coefficient of each reflecting unit. and They represent the first The IRS The amplitude coefficient and phase shift of each reflecting unit in the actual amplitude and phase shift model and The coupling relationship is , and These are constants related to the specific circuit implementation. It is the minimum amplitude. yes arrive Horizontal distance, This represents the mean calculation; according to existing technology, it can be obtained... for: ; Step 3: The PPO algorithm, employing a near-end policy optimization, jointly optimizes the BS active beamforming vector and the IRS passive beamforming vector. Key design elements of PPO include state space, action space, and reward function. Its network model comprises a policy network and a value network. The policy network outputs the BS active beamforming vector and the IRS passive beamforming vector, while the value network outputs the state value function to estimate the advantage function. The network is trained by introducing pruning strategies, importance sampling, and entropy regularization. Through iterative training of the PPO model, the BS active beamforming vector and the IRS passive beamforming vector are jointly optimized. Using the optimized joint beamforming scheme, the maximum average reachable rate of the system is analyzed.
2. The optimization method for a PPO-based dual IRS-assisted short packet communication system according to claim 1, characterized in that, The specific implementation method of step 1 is as follows: Assuming that BS and IRS are both uniform linear arrays ULA, then the array responses at BS, IRS1, and IRS2 are: ,in , It is the spacing between adjacent units. It is the carrier wavelength. It refers to the signal's angle of arrival or departure angle; the departure angle and angle of arrival for the signal from BS to IRS1 are respectively... The departure angle and arrival angle from IRS1 to IRS2 are respectively The departure angle from IRS2 to UE is The channel gain of the BS transmitted signal reaching the UE link after cascaded reflections via IRS1 and IRS2 is: ,in The channel coefficients between BS and IRS1 The channel coefficients between IRS1 and IRS2 The channel coefficients between IRS2 and UE. It is the path loss exponent corresponding to the channel. This is the path loss at a unit reference distance of 1 meter. The distance from the center of BS to the center of IRS1. The distance from the center of IRS1 to the center of IRS2. The distance from the IRS2 center to the UE. , and They are respectively , and The corresponding Rice factor, The line-of-sight component is , The line-of-sight component is , The line-of-sight component is ; Non-line-of-sight components , Non-line-of-sight components and Non-line-of-sight components The elements in the set are independent and all follow a cyclic symmetric complex Gaussian random variable with mean 0 and variance 1; The channel gain can be derived. Follow the mean ; variance is The complex Gaussian distribution, in which , ,but .
3. The optimization method for a short packet communication system assisted by a PPO-based dual IRS system according to claim 1, characterized in that, The specific implementation of step 3 is as follows: For the design of key elements of PPO, in order to fully consider the state space's description of environmental information, it is designed according to CSI as follows: ; in , , , , , Take the argument of the complex elements in the matrix. The current time step, superscript matrix express exist The time step value, the agent's action space includes the amplitude coefficient, phase shift coefficient, and phase offset of the dual IRS of the BS beamforming vector: ; For active and passive beamforming designs based on statistical CSI, the reward function can be expressed as: ; The network model of the PPO algorithm includes a policy network and a value network; Its training process first initializes the policy network. and value network , Represents all learnable parameters of the policy network. Represents all learnable parameters of the value network, subscript or The policy network or value network is represented by the current parameter vector. or Perform parameterization; For input status, To output the action, For the strategy in the state Select action The probability, Input state The state-value function at time; creating an experience buffer; initializing the state at the start of training. In time step In the middle, using the current policy network Interacting with the environment Given the parameters of the action network in the previous trajectory, obtain the current state. Output the probability distribution of actions and obtain the actions by sampling according to the probability. Calculate the reflection coefficient of the IRS and the beamforming vector of the BS, and calculate the reward based on the reward function. And update status Meanwhile, the experience buffer collects complete interaction data for that time step, including state transition information and the actions taken by the current policy. The logarithmic probability of a tuple When the experience buffer reaches its preset capacity, multiple rounds of policy updates are performed on this batch of data. In each round of updates, the policy network... Based on the current state Recalculate the log probability of the new strategy And calculate the importance sampling ratio with the log probability of the old strategy. The pruning mechanism limits the scope of policy updates, while the value network estimates the value function based on empirical data. and dominance function Finally, update the policy network parameters. and value network parameters And clear the experience buffer to prepare for the next data collection; The advantage function is obtained by using generalized advantage estimation: ; in , As a discount factor, In generalized dominance estimation, the parameter that balances variance and bias is... These are the parameters of the value network in the previous trajectory; Importance sampling weights Constrain the update magnitude of the policy network: ; The policy network calculates the loss using a pruning objective function: ; in This is a pruning function used to restrict policy updates to a specified range. Inside, The threshold for the cropping range; Entropy regularization is used to avoid getting stuck in local optima during training. Based on the current strategy The calculated entropy: ; The policy network uses a gradient algorithm in each After each trajectory is completed, the policy gradient is calculated based on the cumulative reward of the trajectory, and the parameters of the policy network are updated. ; ; in The learning rate of the policy network. The entropy coefficient, This represents gradient operation; Value networks use the mean squared error function to calculate their loss function: ; The parameters of the value network are updated using backpropagation via gradient descent. ; ; in The learning rate of the value network.
4. The optimization method for a PPO-based dual IRS system-assisted short packet communication system according to claim 3, characterized in that, The specific process of the PPO algorithm is as follows: 4-1. Initialize policy network parameters and value network parameters Initialize relevant hyperparameters: learning rate and Discount Factor GAE parameters , clipping threshold And entropy coefficient ; 4-2. Initialize the training round index and set the maximum number of training rounds; 4-3. Initialize the environment state, generate statistical CSI under Ricean fading channel, and initialize the system time step. and create an empty experience buffer; 4-4. Obtain the system status at the current time step ,Will Input a policy network and output the action probability distribution under the current policy. ; 4-5. Sampling is performed based on the probability distribution to obtain the current action. Record the logarithmic probability at this time. ; 4-6. Perform the actions Mapping to the physical environment: Calculate the reflection amplitude corresponding to the phase shift of the IRS unit based on the actual amplitude phase shift model, and generate joint beamforming; 4-7. Interacting with the environment: Calculate the average reachability of the current system, calculate the reward for the current step according to the reward function, and observe the state at the next time step. ; 4-8. The obtained experience tuples Fill the experience buffer; 4-9. Update the current status Update time step ; 4-10. Determine if the experience buffer has reached the preset capacity: If not, return to step 3-4 to continue collecting data; if it has, proceed to step 3-11 to start updating network parameters. 4-11. Estimating State Value Using Value Networks And calculate the advantage function according to the generalized advantage estimation formula. and target return value; 4-12. Perform multiple rounds of iterative updates on the data in the buffer: calculate the ratio of the old and new policies, and update the policy network parameters. and value network parameters ; 4-13. After the parameter update is complete, clear the experience buffer and assign the updated strategy parameters to the old strategy parameters. ; 4-14. Update the training rounds and determine whether the current round is less than the maximum round: if it is, return to step 3-3 to start a new round of training; if it is not, the optimization process ends and the optimization results and network weight parameters are output.
Citation Information
Patent Citations
Transmission and reflection type intelligent reflecting surface assisted non-orthogonal multiple access short packet communication method
CN116170856A
Short packet transmission method and system based on irregular IRS auxiliary backscatter communication
CN116980931A
Dynamic reflection coefficient adjusting method for double intelligent reflecting surfaces based on deep reinforcement learning
CN119420384A
Wireless communication transmission method based on intelligent reflecting surface
CN120264323A
Digital twin-based deduction and optimization method and system for intelligent reflecting surface communication system
US20250175216A1