Intelligent reflecting surface assisted resource allocation method based on dqn-ddpg

By introducing the DQN-DDPG algorithm into the OFDM communication system and optimizing the subcarrier allocation and IRS reflection surface phase shift, the complexity problem of improving system performance in multi-user scenarios is solved, and efficient resource allocation and system performance enhancement are achieved.

CN114040415BActive Publication Date: 2025-10-17NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111292938.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-10-17
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively solving the performance improvement of OFDM communication systems in multi-variable real-time joint optimization problems, especially in multi-user scenarios. Traditional methods are complex and cannot simultaneously optimize the discrete-continuous mixed action space.

Method used

A resource allocation method based on DQN-DDPG is adopted, combined with deep Q learning network and deep deterministic policy gradient to optimize subcarrier allocation, beamforming and IRS passive beam phase shift, and maximize the total system rate through Markov process and reward mechanism.

Benefits of technology

It significantly improves the receiving end signal strength and system performance of the OFDM communication system, simplifies the parameter optimization process, is applicable to various system scenarios, reduces complexity and achieves high achievable rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114040415B_ABST
    Figure CN114040415B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent reflecting surface (IRS) assisted OFDM communication system, which adjusts its phase shift by deploying an IRS to improve system throughput. By modeling a design problem of joint subcarrier allocation, base station transmit beamforming and IRS phase shift optimization, the application aims to maximize system throughput. In the method, multiple DQNs are used to solve the problem of excessive discrete action space, and DDPG is used to solve the problem of continuous action allocation. Simulation results show that, compared with other methods, the proposed DQN-DDPG based method can learn from the environment and continuously improve behavior, significantly improve the sum rate of the system, and has good convergence effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of wireless communication, in particular to a resource allocation method based on deep Q neural network (DQN) and deep deterministic policy gradient (DDPG) in intelligent reflecting surface (IRS) assisted OFDM communication system. The DQN-DDPG algorithm is introduced in the OFDM resource allocation system, which not only successfully solves the problem of simultaneous optimization of discrete and continuous variables, but also is easy to extend to various system scenarios. BACKGROUND

[0002] Orthogonal frequency division multiplexing (OFDM) is a widely used technology in many communications such as LTE and 5G, which can achieve high-speed and robust information transmission by using orthogonal subcarriers, and can effectively avoid inter-channel interference. At the same time, by optimizing the subcarrier and power allocation of the OFDM communication system, the system performance can be significantly improved. With the rapid development of mobile Internet and wireless services, the current mobile data volume is facing an explosive growth and the demand for higher data rate is a difficult problem. At the same time, the wireless channel fading environment will seriously weaken the performance of the OFDM communication system and weaken the user experience. Therefore, how to further improve the performance of the OFDM communication system to meet the growing user demand has become a pressing problem of general concern in the industry.

[0003] Recently, intelligent reflecting surface (IRS) aided wireless communication is considered as an ideal solution to the above problems. Specifically, IRS is a reconfigurable metasurface array consisting of a large number of passive reflecting elements that can independently induce phase shifts to the incident signals, thereby cooperatively altering the propagation of the reflected signals to achieve the desired channel response in wireless communication. By properly adjusting the phase shifts of the IRS elements, the reflected signals of different paths can be coherently combined at the receiver to maximize the link achievable rate. In addition, by changing the resistance load of each element, flexible control of the reflection amplitude of IRS can also be achieved. Among them, the paper "IRS-Enhanced OFDMA: Joint Resource Allocation and Passive Beamforming Optimization" is a study of IRS-aided downlink OFDM communication system, which uses alternating optimization algorithm and successive convex approximation technique to jointly optimize IRS reflection coefficient, time-frequency resource block and power allocation to maximize user downlink rate. Experiments show that IRS can significantly improve system performance, but the author only considers single antenna base station and does not consider the scenario of multi-antenna base station. The patent "A security beam forming method and device based on IRS" invented by Zhengzhou University is a study of IRS-aided secure communication, which also uses the alternating optimization method to maximize energy collection under the condition of meeting a certain safety rate.

[0004] But the above work uses the traditional technology of alternating optimization, adopts complex mathematical formula and numerical optimization method, and cannot meet the real-time processing requirements of large and complex communication systems. Therefore, inspired by the application of artificial intelligence technology, some works try to use deep learning algorithm to optimize the IRS phase shift matrix and beamforming matrix to maximize the system performance. The authors Chongwen H, Ronghong Mo and others in the published paper "Reconfigurable Intelligent Surface Assisted Multiuser MISO Systems Exploiting Deep Reinforcement Learning" use deep reinforcement learning (DRL) to solve the joint optimization problem of beamforming matrix and IRS phase shift matrix, but the authors only consider the ordinary scene. The authors Keming Feng and others in the published paper "Deep Reinforcement Learning Based Intelligent Reflecting Surface Optimization for MISO Communication Systems" use DRL to optimize the IRS phase shift matrix, and the simulation results show that the DRL algorithm can reach the upper bound of system performance at a lower time consumption compared with the semidefinite relaxation algorithm, but this work only considers the special single-user scenario, and the more general and complex multi-user scenario has not been analyzed and researched.

[0005] Most of the existing researches are based on the method of alternating optimization, and use complex mathematical formula and numerical optimization technology, which is difficult to truly solve the multi-variable real-time joint optimization problem, and only using DQN or DDPG method cannot solve the problem of discrete continuous mixed action space. Therefore, in view of the above problems, the resource allocation method based on DQN-DDPG is proposed, which jointly optimizes the subcarrier allocation, beamforming and IRS passive beam phase shift, has good convergence effect, and can be easily extended to various system scenarios. SUMMARY

[0006] The present application is aimed at the resource allocation problem in the IRS assisted OFDM communication scene, and proposes a reinforcement learning optimization method based on joint DQN and DDPG, which can ensure that the whole system can obtain the maximum total rate. Not only can it successfully solve the problem of simultaneous optimization of discrete continuous mixed variables, but also can be easily extended to different scenes.

[0007] To achieve the above purpose, the technical method of the present application comprises the following steps:

[0008] Comprise the following steps:

[0009] Step 1, set the positions of the base station, IRS and users, model the channels between the base station to the IRS, the IRS to the K users and the base station to the K users, and obtain the channel gains of the three;

[0010] Step 2, according to the channel gains of the three in step 1, obtain the optimization problem of the system sum rate;

[0011] Step 2.1, the achievable rate of the base station using subcarrier c to transmit data to user k can be expressed as

[0012] Step 2.2, the goal of the system is to jointly design subcarrier allocation, beamforming and IRS passive beamforming matrix to maximize the system sum rate, and this goal needs to meet the base station transmit power constraint condition, IRS unit reflection amplitude constraint condition, subcarrier usage constraint condition, user transmission mode constraint condition and user minimum rate requirement constraint condition;

[0013] Step 3, according to the user subcarrier allocation, base station beamforming, IRS passive beamforming phase shift, user minimum achievable rate requirement, system sum rate in the communication system, establish a Markov process;

[0014] Step 4, use the joint DQN and DDPG algorithm to optimize the reinforcement learning model;

[0015] Step 5, according to the optimized deep reinforcement learning model, obtain the optimized solution, and get the system sum rate;

[0016] Input the current system state s t , the deep reinforcement learning can learn the optimal action a t from the model, and obtain the solution of the optimization problem, the optimal subcarrier allocation, beamforming and passive beam phase shift.

[0017] Further, in step 1, the distribution of the IRS node, the base station node and the K users is defined as follows:

[0018] All communication nodes establish a three-dimensional Cartesian coordinate system, deploy K ground users, a base station with a fixed height is equipped with M antennas, and an IRS with a fixed height is equipped with N reflecting elements, and the phase of each reflecting element can be adjusted to receive a signal, then the positions of the base station, the IRS and the kth user are w B = [x B , y B , z B ] T , w R = [x R , y R , z R ] T , wk = [x k , y k , 0] T where the three numbers in each position represent the corresponding x, y, z-axis coordinates, respectively;

[0019] When the LoS path from the base station to the user is blocked, the channel from the base station to the user k can be modeled as a Rayleigh fading channel, and the channel gain can be expressed as: where is a complex Gaussian random vector with zero mean and unit variance, and PL B,k is the path loss between the base station and the user;

[0020] The channels between the base station and the IRS, and between the IRS and the user k are modeled as Rician fading, and the corresponding channel gains are:

[0021]

[0022]

[0023] where K1 and K2 are Rician-K factors, and are complex Gaussian random components with zero mean and unit variance, and and are deterministic components in the channels, and PL BR and PL Rk are the corresponding path losses;

[0024] The path loss can be modeled as: where PL0= 30 dB, D0= 1 m, ξ is the path loss exponent, and d is the distance between the links.

[0025] Further, in the step 2.2, the user achievable rate is:

[0026]

[0027] where B is the bandwidth, C is the number of subcarriers, represents that the base station transmits to the user k using subcarrier c, and vice versa, represents the beam used by the base station to transmit to the user k using subcarrier c, and represents the noise.

[0028] In the step 2.2, the system objective is expressed as The first constraint condition is that the total transmit power cannot exceed the maximum transmit power of the base station, i.e. The second constraint condition represents that the IRS reflection is full reflection, i.e. |φ n|=1, the third constraint condition is that each subcarrier can only be in two states, i.e. The fourth constraint condition is that one subcarrier can only be allocated to one user, and cannot be occupied by multiple users, i.e. The fifth constraint condition indicates that the minimum achievable rate requirement of the user must be met, i.e.

[0029] Further, in the step 3, the Markov process is specifically represented as:

[0030] Step 3.1 State space S: state s t The action, achievable rate and channel matrix at the t-1 time step are composed, and since the channel matrix has real and imaginary parts, the imaginary part and the real part can be regarded as independent inputs;

[0031] Step 3.2 Action space A: action a t The subcarrier allocation of the discrete action and the beamforming, IRS passive beam phase shift of the continuous action are composed, and are

[0032] Step 3.3 Instantaneous reward r: in order to ensure that the minimum achievable rate of each user is met while the sum rate is maximized, the reward function can be set as where w1 and w2 are constant coefficients, and δ k is expressed as;

[0033]

[0034] Further, in the step 4, the following steps are specifically included:

[0035] Step 4.1, the training round ep is initialized to 0;

[0036] Step 4.2, the time step t in the ep round is initialized to 0;

[0037] Step 4.3, according to the input state s t , the C DQN networks obtain the discrete action

[0038] Step 4.4, according to the input state s t , the Actor online network obtains the continuous action

[0039] Step 4.5, the total network instantaneous reward r t is obtained, and the next state s t+1 is converted, and the training set (s t , a t , r t , s t+1 ) is obtained;

[0040] Step 4.6, store the training data set into the experience replay pool D;

[0041] Step 4.7, determine whether t is satisfied <T,T为ep回合的总步数,若是则t=t+1,返回(4.3),若不是则进入(4.8);

[0042] Step 4.8: Randomly sample a dataset of N samples from the experience replay pool D and send it to the online DQN network, the target DQN network, the online actor network, the target actor network, the online critic network, and the target critic network.

[0043] Step 4.9, the number of DQN networks c is initialized to 0;

[0044] Step 4.10, according to the sampled data set, the cth DQN network is based on the state s i and a i Get the corresponding Q(s i ,a i ; w), according to s i+1 Get the optimal Q(s i+1 ; w) value, according to the return r i And Q value to get the network LOSS function is The online network updates the parameter w by minimizing the LOSS function;

[0045] Step 4.11: Determine whether c=C-1 is satisfied, where C is the total number of DQN networks. If not, then c=c+1 and return to step 4-10. If not, proceed to step (4.12).

[0046] Step 4.12, based on the sampled data set, the online Actor network is based on the state s i , get action a i =π(s i ; μ), the state s i and the resulting action a i =π(s i ; μ) input into online Critic network to obtain Q(s i ,π(s i ;μ);w), according to To update the online Actor network parameters θ, also use Update the online critic network parameters;

[0047] Step 4.13, every U rounds, use the online Actor network parameters μ to update the target Actor network parameters μ -updating the parameters theta in the target Critic network with the online Critic network parameters theta - ;

[0048] Step 4.14, judge whether the number of rounds ep meets ep < EP, EP is the total number of rounds, if yes, ep = ep + 1, return (4.2), if not, the optimization is ended, and the optimized reinforcement learning model is obtained.

[0049] The present application has the following advantages:

[0050] 1. The weakening characteristics of wireless channel propagation make it difficult to ensure the performance of the OFDM communication system, and the present application introduces IRS into the OFDM communication system, overcomes the adverse effects of channel fading, significantly enhances the signal strength at the receiving end, and guarantees the high achievable rate performance of the OFDM communication system.

[0051] 2. The present application first introduces the framework of DQN-DDPG in the IRS auxiliary resource allocation scene, compared with the traditional alternating optimization algorithm, solves the problem that numerous parameters are difficult to be optimized online at the same time, and the proposed DQN-DDPG method does not need to use complex mathematical formulas and numerical optimization techniques, and can be easily extended to various system scenes.

[0052] 3. Channel allocation is a combination problem, using a single DQN will make the discrete action space expand exponentially, significantly increasing the problem complexity, therefore, the present application uses multiple DQNs to solve the problem, changes the exponentially growing action space to the product size, and greatly reduces the complexity. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 It is a DQN-DDPG algorithm block diagram based on the present application.

[0054] Figure 2 It is a reward graph of the DQN-DDPG algorithm in the present application under the training step number.

[0055] Figure 3 It is a sum rate graph under different transmit power and IRS passive reflection unit number.

[0056] Figure 4 It is a user rate graph under the training step. DETAILED DESCRIPTION

[0057] The application will be further described below in combination with the drawings.

[0058] The technical method of the present application comprises the following steps.

[0059] Step 1, set the positions of base station, IRS and users. Model the channels between base station to IRS, IRS to K users and base station to K users, and get the channel gains of the three.

[0060] The distribution of IRS nodes, base station nodes and K users is defined as follows:

[0061] All communication nodes establish a three-dimensional Cartesian coordinate system, deploy K ground users, a base station with a fixed height is equipped with M antennas, and an IRS with a fixed height is equipped with N reflecting elements and the phase of each reflecting element can be adjusted to receive signals, then the coordinates of the base station, IRS and the kth user are w B = [x B ,y B ,z B ] T , w R = [x R ,y R ,z R ] T , w k = [x k ,y k ,0] T ;

[0062] When the LoS path of the base station to the user is blocked, the channel from the base station to the kth user can be modeled as a Rayleigh fading channel, and the channel gain can be represented as: where is a complex Gaussian random vector with zero mean and unit variance, PL B,k is the path loss between the base station and the user;

[0063] The channel between the base station and the IRS and the channel between the IRS and the kth user is modeled as Rician fading, so the corresponding channel gain is:

[0064]

[0065]

[0066] where K1 and K2 are Rician-K factors, and are complex Gaussian random components with zero mean and unit variance, and and are deterministic components in the channel, PL BR and PL Rk are the corresponding path losses.

[0067] The path loss can be modeled as: where PL0 = 30 dB, D0 = 1 m, ξ is the path loss exponent, and d is the distance between links.

[0068] Step 2, according to the channel gain obtained in step 1, model the optimization problem of maximizing the sum rate of the system.

[0069] The achievable rate of base station transmitting data to user k using subcarrier c is:

[0070]

[0071] where B is the bandwidth, C is the number of subcarriers, represents that the base station transmits to user k using subcarrier c, and vice versa. represents the beam used by the base station to transmit to user k using subcarrier c, and represents noise.

[0072] The goal of the system is to jointly design subcarrier allocation, beamforming, and IRS passive beam matrix to maximize the sum rate of the system, which can be expressed as The first constraint is that the total transmit power cannot exceed the maximum transmit power of the base station, that is, The second constraint represents that the IRS reflection is full reflection, that is, |φ n |=1. The third constraint is that each subcarrier can only be used or not used, that is, The fourth constraint is that a subcarrier can only be allocated to one user for use, and cannot be occupied by multiple users, that is, The fifth constraint represents that the minimum achievable rate requirement of the user must be met, that is,

[0073] Step 3, according to the subcarrier allocation of users in the communication system, the beamforming of the base station, the passive beam phase shift of the IRS, the minimum achievable rate requirement of the user, and the sum rate of the system, a deep reinforcement learning model is established.

[0074] The Markov process is established as:

[0075] Step 3-1, state space S: state s t From the action at the t-1 time step, the achievable rate and the channel matrix are composed, and since the channel matrix has real and imaginary parts, the imaginary part and the real part can be taken as independent inputs.

[0076] Step 3-2, action space A: action a t composed of discrete action subcarrier allocation and continuous action beamforming, IRS passive beam phase shift, that is,

[0077] Step 3-3, immediate reward r: to ensure that the minimum achievable rate of each user is met while maximizing the total rate, the reward function can be set as where w1 and w2 are constant coefficients, and δ k is expressed as:

[0078]

[0079] Step 4, using joint DQN and DDPG algorithm to optimize reinforcement learning model.

[0080] Specifically comprising the following steps:

[0081] Step 4-1, training round ep is initialized to 0;

[0082] Step 4-2, time step t in ep round is initialized to 0;

[0083] Step 4-3, according to the input state s t , C DQN networks obtain discrete action

[0084] Step 4-4, according to the input state s t , Actor online network obtains continuous action

[0085] Step 4-5, total network instant reward r t is obtained, and the next state s t+1 is converted, and the training set (s t , a t , r t , s t+1 ) is obtained;

[0086] Step 4-6, store the training data set into experience replay pool D;

[0087] Step 4-7, judge whether t < T, T is the total number of steps in ep round, if yes, t = t + 1, return to step 4-3, if not, go to step 4-8;

[0088] Step 4-8, randomly sample a data set composed of N number of samples from experience replay pool D, and send it to online DQN network, target DQN network, online Actor network, target Actor network, online Critic network and target Critic network;

[0089] Step 4-9, DQN network number c is initialized to 0;

[0090] Step 4-10, according to the sampled data set, the cth DQN network obtains the corresponding Q(s i , a i ; w) according to state s i and a i ; w), and s i+1get the optimal Q(s i+1 ;w) value, the loss function of the network is obtained according to the return r i and the Q value The online network updates the parameters w by minimizing the LOSS function.

[0091] Step 4-11, judge whether c=C-1 is met, C is the total number of DQN networks, if not met, c=c+1, return to step 4-10, if not, enter step 4-12.

[0092] Step 4-12, according to the sampled data set, the online Actor network obtains the action a i =π(s i ;μ) according to the state s i , input the state s i and the obtained action a i =π(s i ;μ) into the online Critic network, and obtains Q(s i ,π(s i ;μ);w), the online Actor network parameter θ is updated according to , and the online Critic network parameter is updated by ;

[0093] Step 4-13, every U rounds of online Actor network parameters μ are used to update the parameters μ - in the target Actor network, and the online Critic network parameters θ are used to update the parameters θ - in the target Critic network.

[0094] Step 4-14, judge whether the number of rounds ep<EP is met, EP is the total number of rounds, if yes, ep=ep+1, return to step 4-2, if not, the optimization is ended, and the optimized reinforcement learning model is obtained.

[0095] Step 5, the optimized solution is obtained according to the optimized deep reinforcement learning model, and the total rate of the system is obtained.

[0096] The current system state s t is input, and the deep reinforcement learning can learn the optimal action a t according to the model, and the solution of the optimization problem, the optimal subcarrier allocation, beamforming and passive beam phase shift can be obtained.

[0097] The performance effect of the application can be further illustrated by the following simulation:

[0098] 1. Simulation conditions

[0099] Assume that there are K = 3 downlink users in the communication system, the number of antennas of the base station is M = 8, the number of passive reflector units of the IRS is N = 16, and the number of subcarriers is C = 7. The position of the base station is [0,0,30] T , the position of IRS is [75,100,40] T , the position of user k is [x k ,y k ,0] T , of which 100 <x k <200,0 <y k <100. The path loss exponent of the direct link from the base station to the user is 3.75, while that of the reflection link is 2.2. The channel attenuation at a reference distance of 1m is 30dB, the maximum transmit power of the base station is 35dB, the channel bandwidth is 1MHz, and the noise power is -169dBm.

[0100] 2. Simulation content

[0101] Attachment Figure 2 The reward graphs for resource allocation using IRS-assisted DQN-DDPG, resource allocation without IRS-assisted DQN-DDPG, and random resource allocation are shown respectively. The curves using IRS and those not using IRS are compared, proving that the use of IRS can significantly improve the system's sum rate. Comparing the curves using DQN-DDPG with those using random allocation proves the effectiveness of the algorithm.

[0102] Attachment Figure 3 The system sum rate changes under different numbers of passive reflectors and different transmit powers are demonstrated. As the total base station power increases, more transmit power is allocated to users, and the proposed DQN-DDPG algorithm achieves a higher total system rate. The total rate increases with the number of passive reflectors, demonstrating the effectiveness of IRS in improving communication quality.

[0103] Attachment Figure 4 It shows that as the number of rounds increases, the final achievable rate of each user tends to stabilize. As the algorithm continues to interact with the environment, it can learn and adjust the optimization variables to approach the optimal solution, effectively meet the minimum transmission rate requirements of each user, make reasonable resource allocation, and at the same time achieve the maximum total system rate.

[0104] Based on the above simulation results and analysis, the DQN-DDPG-based intelligent reflector-assisted OFDM communication system resource allocation method proposed in the present invention can achieve the maximum sum rate for the entire system without requiring complex mathematical formula derivation and optimization techniques. The method achieves good real-time performance and is easily extended to various system settings, which makes the invention more applicable in practice.

Claims

1. The intelligent reflector-assisted DQN-DDPG-based resource allocation method is characterized by: It includes the following steps: Step 1: Set the positions of the base station, IRS, and users, model the channels between the base station and the IRS, between the IRS and the K users, and between the base station and the K users, and obtain the channel gains of the three. Among them, the distributions of the IRS nodes, base station nodes, and K users are defined as follows: All communication nodes establish a three-dimensional Cartesian coordinate system, deploy K ground users, a fixed-height base station equipped with M antennas and a fixed-height IRS equipped with N reflectors, and the phase of each reflector can adjust the received signal. Then the position of the base station, IRS, and the kth user is w B =[x B ,y B , z B ] T , w R =[x R ,y R , z R ] T , w k =[x k ,y k ,0] T , where the three numbers in each position represent the corresponding x, y, and z axis coordinates respectively; When the LoS path from the base station to the user is blocked, the channel from the base station to user k can be modeled as a Rayleigh fading channel, and the channel gain can be expressed as: in is a complex Gaussian random vector with zero mean and unit variance, PL B,k is the path loss between the base station and the user; The channels between the base station and the IRS and between the IRS and user k are modeled as Rice fading, so the corresponding channel gains are: where K1 and K2 are Rician-K factors, and is a complex Gaussian random component with zero mean and unit variance, and and is the deterministic component in the channel, PL BR and PL Rk is the corresponding path loss; the path loss can be modeled as: Where PL0 = 30dB, D0 = 1m, ξ is the path loss exponent, and d is the distance between links; Step 2: According to the channel gains of the three in Step 1, obtain the optimization problem of the system sum rate. In step 2.1, the achievable rate at which the base station transmits data to user k using subcarrier c can be expressed as Step 2.2: The goal of the system is to jointly design subcarrier allocation, beamforming, and IRS passive beamforming matrix to maximize the system sum rate, and this goal must satisfy the base station transmit power constraint, IRS unit reflection amplitude constraint, subcarrier usage constraint, user transmission mode constraint, and user minimum rate requirement constraint. Step 3: Establish a Markov process based on the user subcarrier allocation, base station beamforming, IRS passive beamforming phase shift, user minimum achievable rate requirement, and system sum rate in the communication system. Step 4: Use the joint DQN and DDPG algorithms to optimize the reinforcement learning model. Step 5: Obtain the optimized solution according to the optimized deep reinforcement learning model and get the system sum rate. Enter the current system status s t , deep reinforcement learning can learn the optimal action a based on the model t , the solution of the optimization problem, the optimal subcarrier allocation, beamforming and passive beam phase shifting can be obtained.

2. The intelligent reflective surface-assisted DQN-DDPG-based resource allocation method according to claim 1, characterized in that: In Step 2.2, the user achievable rate is: Where B is the bandwidth, C is the number of subcarriers, It means that the base station uses subcarrier c to transmit to user k, otherwise it means that subcarrier c is not used. represents the beam transmitted by the base station to user k using subcarrier c, and Indicates noise.

3. The intelligent reflective surface-assisted DQN-DDPG-based resource allocation method according to claim 1, characterized in that: In step 2.2, the system goal is expressed as The first constraint is that the total transmission power cannot exceed the maximum transmission power of the base station, that is, The second constraint states that the IRS reflection is total reflection, that is The third constraint is that each subcarrier has only two conditions: used and unused, that is, The fourth constraint is that a subcarrier can only be allocated to one user and cannot be occupied by multiple users. The fifth constraint indicates that the user's minimum achievable rate requirement must be met, that is, 4. The intelligent reflective surface-assisted DQN-DDPG-based resource allocation method according to claim 1, characterized in that: In Step 3, the Markov process is specifically expressed as: Step 3.1, state space S: state s t It is composed of the action at time step t-1, the achievable rate and the channel matrix. Since the channel matrix has imaginary and real parts, the imaginary and real parts can be used as independent inputs; Step 3.2, action space A: action a t It consists of discrete action subcarrier allocation and continuous action beamforming, IRS passive beam phase shift, and is Step 3.3, immediate reward r: To ensure that the minimum achievable rate for each user is met while maximizing the sum rate, the reward function can be set to Where w1 and w2 are constant coefficients, and δ k Expressed as; 5. The intelligent reflector-assisted DQN-DDPG-based resource allocation method according to claim 1, characterized in that: In Step 4, it specifically includes the following steps: Step 4.1: Initialize the training episode ep to 0. Step 4.2: Initialize the time step t in the ep episode to 0. Step 4.3, according to the input state s t , C DQN networks obtain discrete actions Step 4.4, according to the input state s t , Actor online network obtains continuous action Step 4.5, get the total network instant reward r t , and transition to the next state s t+1 , get the training set (s t , a t , r t , s t +1 ); Step 4.6: Store the training data set in the experience replay pool D. Step 4.7: Determine whether t < T is satisfied, where T is the total number of steps in the ep episode. If so, t = t + 1, and return to (4.3). If not, enter (4.8). Step 4.8: Randomly sample a data set composed of a batch of N samples from the experience replay pool D and send it to the online DQN network, target DQN network, online Actor network, target Actor network, online Critic network, and target Critic network. Step 4.9: Initialize the number c of DQN networks to 0. Step 4.10, according to the sampled data set, the cth DQN network is based on the state s i and a i Get the corresponding Q(s i , a i ; w), according to s i+1 Get the optimal Q(s i+1 ; w) value, according to the return r i And Q value to get the LO SS function of the network is The online network updates the parameter w by minimizing the LOSS function; Step 4.11: Determine whether c = C - 1 is satisfied, where C is the total number of DQN networks. If not, c = c + 1, and return to 4.

10. If not, enter (4.12). Step 4.12, based on the sampled data set, the online Actor network is based on the state s i , get action a i =π(s i ; μ), the state s i and the resulting action a i =π(s i ; μ) input into online Critic network to obtain Q(s i ,π(s i ;μ);w), according to To update the online Actor network parameters θ, also use Update the online critic network parameters; Step 4.13, every U rounds, use the online Actor network parameters μ to update the target Actor network parameters μ-, and use the online Critic network parameters θ to update the target Critic network parameters θ - ; Step 4.14: Determine whether the number of episodes ep < EP is satisfied, where EP is the total number of episodes. If so, ep = ep + 1, and return to (4.2). If not, the optimization ends, and the optimized reinforcement learning model is obtained.

Citation Information

Patent Citations

  • A D2D user resource allocation method based on a deep reinforcement learning DDPG algorithm

    CN109862610A

  • Service function chain low-cost intelligent deployment method based on environmental perception

    CN111093203A