An Optimization Method for Active and Passive Transmission of UAV-Assisted Intelligent Reflecting Surface Based on Deep Reinforcement Learning

By applying deep reinforcement learning to optimize the drone motion trajectory in the UAV-assisted Intelligent Reflective Surface (IRS) communication system, the problems of active passive signal bit error rate performance and system spectrum efficiency are solved, and efficient signal transmission and optimized energy efficiency performance are achieved.

CN115767581BActive Publication Date: 2025-06-27CHINA JILIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211118912.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2025-06-27
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problems of bit error performance and system spectrum efficiency of active passive signals in drone-assisted intelligent reflective surface (IRS) communication systems, especially when facing explosive data and complex mathematical formulas, the calculation time is long or cannot be solved.

Method used

Deep reinforcement learning (DRL)-based method is adopted, combined with deep Q network (DQN) to optimize the drone motion trajectory, integrate IRS and UAV to achieve total reflection modulation and index modulation, maximize the received signal-to-noise ratio of active signals and improve the spectrum efficiency of the system.

Benefits of technology

Optimizing the drone motion trajectory through deep reinforcement learning, solving the computational complexity of the non-convex optimization problem, achieving higher bit error rate performance of active passive signals and system spectrum efficiency, reducing the computing cost, and making it possible to deploy UAV in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115767581B_ABST
    Figure CN115767581B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for optimizing the active and passive transmission of an unmanned aerial vehicle (UAV)-assisted intelligent reflecting surface (IRS) based on deep reinforcement learning. In the active and passive transmission system of UAV-IRS, the UAV and IRS are integrated to improve the deployment flexibility of the IRS. The UAV is utilized to help the IRS reflect the active signal to the base station. Meanwhile, the IRS provides additional information bits for the proximal UAV and improves the spectral efficiency of the system. First, generalized orthogonal reflection modulation is adopted for the IRS to improve the transmission reliability of passive information. At the same time, the full reflection of the IRS can maximize the signal-to-noise ratio of the received active signal. Then, by jointly optimizing the UAV flight trajectory and the IRS scheduling of Internet of Things (IoT) devices, the objective function is to maximize the number of successfully transmitted bits and minimize the UAV energy consumption. The present invention uses the deep Q-network algorithm of DRL to solve the sub-optimal solution. Compared with the benchmark solution, the generalized orthogonal reflection modulation scheme based on IRS and the method optimized based on DRL can effectively improve the energy efficiency performance of the UAV-RIS system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to IRS communication systems, active and passive transmissions in the Internet of Things, index modulation, artificial intelligence, and the method of combining deep reinforcement learning to optimize the movement trajectory of drones, so as to assist the IRS in reflecting active signals in the Internet of Things system while transmitting passive signals for drones, and improve the system spectral efficiency and the bit error rate performance of active and passive signals. Specifically, it is a method for optimizing active and passive transmissions of drone-assisted intelligent reflecting surfaces based on deep reinforcement learning. Background Art

[0002] In the development of future wireless communication systems, it is necessary to not only meet the current situation of the continuous popularization of mobile intelligent terminals and the rapid growth of data volume, but also adapt to the development direction of large-scale Internet of Things (IoT). The Sixth Generation (6G) mobile communication will greatly enhance IoT communication to achieve the vision of "interconnection of all things" and integration of aerial, terrestrial, and space communications. However, the scarcity of radio spectrum, especially below 6 GHz, has become a key bottleneck in the development of wireless networks, and the requirement of low power consumption also brings huge challenges to system design. Therefore, revolutionary technologies are urgently needed to provide solutions with high spectral efficiency and high energy efficiency.

[0003] In recent years, the concept of Intelligent Reflecting Surface (IRS) has emerged. The IRS consists of a large number of low-cost, low-power, and nearly passive reflecting units. It reflects the incident signal with a specific reflection coefficient and changes the amplitude, phase, etc. of the incident signal to control the wireless network environment. Information bits can be mapped to the reflecting unit index to invisibly transmit messages, and the IRS has developed new dimensions and ideas for data transmission. Some new studies have begun to combine the IRS and drones to improve the performance of air-ground communication networks and IoT symbiotic radio communication systems. By combining the IRS and drones, the coverage of ground IoT communication can be expanded, thereby improving the various quality of service requirements of IoT devices. Specifically, some scholars have considered a downlink single-input single-output communication system composed of a mobile drone integrated with an IRS and ground devices. By jointly designing the trajectory of the drone and adjusting the phase shift of the IRS on the drone, the maximum average achievable rate is obtained, and the continuous convex approximation (SCA) method is used to obtain the optimal drone trajectory. However, it is very difficult to solve the joint optimization problem due to various constraint conditions. In existing studies, convex optimization theory is often used to solve it, but for explosive data, it usually causes problems such as long calculation time or even inability to solve due to complex mathematical formulas and numerical optimization.

[0004] In addition, since the IRS has natural advantages in implementing spatial index modulation technology, in a communication system assisted by a UAV-IRS, the on / off state of the IRS reflection elements is often used to modulate the passive information of the UAV or the proximal sensor device, and cooperate with the active transmission of the transmitting end to form a good active-passive transmission effect, so as to realize the active-passive transmission system composed of ground devices and UAVs. However, since there is always a part of the IRS elements in the off state in the activation mode of the on / off state, the inherent large-scale aperture gain of the IRS cannot be maximized, resulting in a decrease in the bit error rate performance of the active information of the ground device. In addition to using the IRS to achieve spatial modulation, existing research also realizes passive transmission based on IRS phase modulation by changing the phase of the incident signal, in which the active information of the ground device and the passive information of the UAV are jointly detected at the receiver. The modulation constellation points can be flexibly assigned to the active information and the passive information through the phase and amplitude to meet different transmission requirements in the active-passive transmission system. However, the design of the coexisting constellation points makes the active information and the passive information affect each other seriously, and the performance degradation of one party will cause the performance of the other party to decrease accordingly. Summary of the Invention

[0005] Aiming at the problems existing in the above-mentioned prior art, the present invention proposes an optimization method for active-passive transmission of a UAV (Unmanned Aerial Vehicle)-assisted intelligent reflecting surface (Intelligent Reflecting Surfaces, IRS) based on deep reinforcement learning (Deep Reinforcement Learning, DRL). The IRS uses generalized orthogonal reflection modulation to reactivate the off-state elements to maximize the received signal-to-noise ratio of the active signal. At the same time, the IRS reflection mode index modulation transmits additional information bits for the proximal UAV to improve the spectral efficiency of the system. To improve the deployment flexibility of the IRS, the IRS and the UAV are integrated into one body, and the deep reinforcement learning method is combined to optimize the UAV movement trajectory to maximize the energy efficiency performance of the UAV.

[0006] The present invention is realized by adopting the following technical solutions:

[0007] An optimization method for active-passive transmission of a UAV-assisted intelligent reflecting surface based on deep reinforcement learning, comprising the following steps:

[0008] 1) Construct an uplink communication system model for UAV-IRS-assisted Internet of Things devices, including a base station (BS) with N r receiving antennas, K ground Internet of Things (IoT) devices with single transmitting antennas, and a UAV integrated with L intelligent reflecting surface units. The active signals of the IoT devices in the system adopt M-QAM constellation modulation. Establish a Cartesian coordinate system for the range of the UAV's movable area. Denote the three-dimensional coordinate positions of the UAV and the IRS at time slot n. N is the total number of time slots. Among them, denotes the coordinate of the UAV in the horizontal direction, denotes the coordinate of the UAV in the vertical direction. Since the IoT and the base station are in a static state, so [x BS , y BS , z BS denotes the coordinate position where the base station is located. denotes the coordinate of the k-th IoT device. The IoT is stationary on the ground, and its height can be almost ignored, that is K denotes the total number of IoT devices.

[0009] 2) To ensure that the UAV assists IRS communication within a safe range, set the maximum and minimum heights that the UAV can move in the vertical direction, that is The duration of the UAV's movement in each time slot is where t min and t max respectively represent the minimum and maximum times that the UAV lasts in time slot n.

[0010] 3) According to steps 1) and 2), the speed of the UAV per unit time can be calculated by distance and time. The horizontal flight speed of the UAV in time slot n can be expressed as:

[0011]

[0012] Among them, is the maximum moving speed of the UAV in the horizontal direction. At the same time, the vertical flight speed of the UAV in time slot n can be expressed as:

[0013]

[0014] Among them, represents the maximum moving speed of the UAV in the vertical direction. If and are equal to 0, it means that in this time slot n, the UAV is stationary in the horizontal and vertical directions respectively.

[0015] 4) According to the speeds of the UAV in the horizontal and vertical directions described in step 3), the energy consumption required by the UAV in time slot n can be expressed as:

[0016]

[0017] Among them, P1, P2, P3, and P4 respectively represent the blade profile power of the UAV in the hovering state, the induced power of the UAV in the hovering state, the constant power of the UAV in the ascending or descending state, and the power required for the UAV to control the IRS. v tip represents the tip speed of the UAV rotor blade, d0 represents the self-drag ratio of the UAV, l represents the solidity of the UAV wind turbine, ρ represents the air density, G represents the rotor disk area, and v0 represents the average blade induction speed of the UAV in hovering.

[0018] 5) Model the channels for air-ground communication. The IoT-UAV and UAV-BS are modeled as Rice fading channels with deterministic line-of-sight components; the Rice fading channel models for the IoT-UAV link and the UAV-BS link are respectively expressed as:

[0019]

[0020]

[0021] Among them, κ1 and κ2 respectively represent the Rice factors of the two links, and are respectively the deterministic line-of-sight paths for the IoT-UAV link and the UAV-BS link; and respectively represent the non-Los paths with Rayleigh distribution for the IoT-UAV link and the UAV-BS link. Each element of the non-Los path follows the complex Gaussian distribution;

[0022] The calculation formula for the deterministic line-of-sight path is:

[0023]

[0024] Among them, λ represents the wavelength, d represents the spacing between adjacent two IRS elements, and it is assumed that and there is no coupling between IRS elements, ψ AoA is the angle of arrival; Considering the channel fading caused by the mobility of the UAV, on the basis of the Rice fading channel model, it is necessary to calculate the channel fading caused by the path distance. The IoT-UAV link and the UAV-BS link including path loss are respectively further expressed as:

[0025]

[0026] and

[0027]

[0028] where τ represents the path loss exponent, f c is the carrier frequency, c is the speed of light, and respectively represent the direct distance between the k-th IoT device and the UAV and the direct distance between the UAV and the BS at time slot n; is calculated from the coordinates of the scheduled IoT at time slot n and the coordinates of the UAV. The specific calculation formula is:

[0029]

[0030] is calculated from the coordinates of the UAV and the coordinates of the BS at time slot n and can be expressed as:

[0031]

[0032] 6) Assume that the input bit streams of the IoT device and the UAV device are grouped, and each group is divided into B1 = log2 M and B2 = log2 Q bits respectively. Both Q and M satisfy integer powers of 2. M represents the M-QAM constellation modulation order adopted by the IoT device, and Q represents the number of generalized orthogonal reflection patterns after IRS grouping.

[0033] 7) Divide the L reflection units of the IRS into adjacent L g groups, and each group consists of L / L g reflection units. Assume that L is divisible by L g . Activate g groups of IRS elements from L g groups to reflect the in-phase signal, where 1 ≤ g ≤ L g , and the remaining L g -g groups are re-activated to generate orthogonal signals to maximize the received signal-to-noise ratio of the active signal. There are combinations of generalized orthogonal reflection patterns, represents the binomial coefficient. Select the first Q generalized orthogonal reflection patterns from combinations for the transmission of UAV information bits; each generalized orthogonal reflection pattern is expressed as:

[0034]

[0035] where 1 indicates that this IRS element is used for in-phase reflection, and j indicates that the reflection phase of this IRS element rotates clockwise by

[0036] 8) Each constellation symbol s i mapped by the active signal is power-normalized and satisfies E[|s i | 2=1, i = 1, 2, ..., M; B1 bits are used for the active signal selection constellation symbol index of the IoT, and B2 bits are used for the passive signal selection IRS's generalized orthogonal reflection mode index to transmit information for the proximal UAV; when receiving, it is assumed that the channel state information is completely known, and the received signal is expressed as:

[0037]

[0038] where, diag(·) represents a diagonal matrix with elements on the main diagonal, P represents the fixed transmit power of the IoT device, is the additive white Gaussian noise AWGN, which follows the distribution CN(0, N0I Nr ), where, I Nr represents the identity matrix, and N0 is the complex noise variance;

[0039] 9) Use the maximum likelihood (ML) detector to jointly detect the active signal of the IoT and the passive signal of the UAV:

[0040]

[0041] where, and respectively represent the detected active signal index and passive signal index at the receiving end, and the corresponding bit information is restored through the index value;

[0042] 10) In time slot n, the total number of successfully transmitted bits of the system is expressed as:

[0043]

[0044] where, B w represents the bandwidth, represents the number of successfully transmitted bits detected in step 9), represents whether the k-th IoT device is scheduled to communicate with the base station in time slot n;

[0045] 11) In the given total time slots N, the optimization problem of energy consumption is defined as:

[0046]

[0047] where, the objective function is to minimize the total energy consumption of the UAV over all time slots;

[0048] Over the total time slots, the total number of successfully transmitted bits received by the receiving end is To comprehensively measure the energy efficiency index of the UAV-IRS system, its objective function can be further expressed as:

[0049]

[0050] As can be seen from Equation (16), the objective function is a non-convex problem. If traditional convex optimization methods are used to solve it, it will incur huge computational costs, making it difficult to be feasible in practical applications.

[0051] To solve the above problems, the present invention will use the Deep Q-Network (DQN) in Deep Reinforcement Learning (DRL) to find a sub-optimal solution. The construction process of the Deep Q-Network DQN is as shown in Steps 12)-16):

[0052] 12) Regarding the UAV as an agent object, the three-dimensional coordinate position of the UAV at the current time slot is used as the state, which is input into the neural network of the DQN based on the current state S(n). The discrete action behavior A(n) output by the network contains the index information of the UAV in the horizontal, vertical, scheduling, and duration directions respectively. The DQN uses Q value to evaluate the value of the action to determine whether to select this action A(n); the environment Env enters the next state S(n + 1) according to the UAV executing the action A(n) in the current state S(n), and scores and gives a reward R(n);

[0053] 13) Define the state action where a1(n) represents the movement index in the horizontal direction, and a2(n) represents the movement index in the vertical direction. is the scheduling for IoT, is the discrete variable of the duration, and its time interval is Δt; according to Equation (16), the reward function can be defined as:

[0054]

[0055] where ζ = {1, 200} is the penalty coefficient. When the flight area of the UAV integrated with IRS exceeds the controllable range of the specified moving area, Env gives a negative feedback ζ = 200;

[0056] 14) When Env updates the state S(n + 1) according to the executed action A(n), Δx = {0, +x u , -x u}, x u represents the distance between two adjacent coordinate points on the x-axis; similarly, Δy = {0, +y u , -y u}, y u represents the distance between two adjacent coordinate points on the y-axis; Δz = {0, +z u , -z u}, z u represents the spacing between two adjacent coordinates on the z-axis;

[0057] 15) Q in step 12) value is defined as:

[0058] Q value (S(n), A(n)) = E[U(n)|S(n), A(n)] (41)

[0059] where E represents expectation, and U(n) represents discounted return; U(n) is expressed as:

[0060] U(n) = R(n) + γR(n + 1) + γ 2 R(n + 2) + … (42)

[0061] where γ represents the discounted return factor; Q value (S(n), A(n)) is the conditional expectation of the return U(n), and its purpose is to eliminate the states and actions involved after the nth time slot in U(n) and score the quality of taking the action A(n) for the current state S(n); through formula (18), DQN adopts the strategy of traversing the maximum value during the learning process to find an action A(n) corresponding to the maximum value of Q value The strategy function for traversing the maximum value is expressed as:

[0062]

[0063] where π′(·) represents the policy function;

[0064] 16) In deep reinforcement learning, it is necessary to collect training set data, and each data set is expressed as:

[0065]

[0066] where S(n) is the current state, A(n) is the action to be executed, R(n) is the reward obtained by executing the action A(n) based on the current state S(n), and S(n + 1) is the next state entered after executing the action A(n); the UAV interacts with the environment to obtain training data After that, each piece of data is stored in the experience pool B buff When the number of data sets in the experience pool reaches the set threshold M size DQN starts to train the neural network, and during the training process, it dynamically obtains the latest data as the UAV interacts with the environment to replace the old data sets in the experience pool.

[0067] In the above technical solution, further, the training method of the deep Q network is specifically as follows:

[0068] 18) Optimization method for UAV-IRS based on DRL. The specific method of step 18) is as follows:

[0069] a) The number of input layers of the neural network model is 3, the number of neurons in the hidden layer is 20, and the number of output layers is related to the number of UAV behaviors A(n). The number of A(n) is There are 5 discrete actions in the horizontal direction and 3 discrete actions in the vertical direction. The discrete time variable is related to the accuracy of Δt. Assume that Δt can be divided by t max -t min divisible; The activation function of the neuron uses the rectified linear unit (ReLU) function:

[0070]

[0071] In the DQN algorithm, the Q in step 15) is fitted through a neural network value ; To alleviate the problem of slow model convergence caused by model errors, when selecting action A(n), the greediness is set to ε = 0.9, that is, there is a 10% probability of randomly selecting the UAV's action instead of always relying on the neural network to select; The learning rate lr = 0.005; The size of the experience pool is M size = 3200;

[0072] b) Set the initial coordinate position of the UAV as S(0) = [0, 0, H max , interact with Env and randomly obtain valid data and store it in the experience pool B buff ; When the amount of data in the experience pool reaches the set threshold M size , the neural network starts training;

[0073] c) Perform batch training on the data in the experience pool. The size of the batch training data is M bath = 25, M bath << M size ; Send the training data into the neural network in batches and perform forward calculations layer by layer until the output layer;

[0074] d) Calculate the loss using the mean squared error loss function:

[0075] L(w) = E[R(n) + γQ value (S(n + 1), A(n + 1)|w) - Q value (S(n), A(n)|w) 2 (46)

[0076] where, L(w) is the loss function, and w is the weight parameter of the neural network;

[0077] e) Through the chain rule, calculate the gradients of the loss function with respect to each layer layer by layer for backpropagation, and use the stochastic gradient descent algorithm to update the weight parameters of the neural network:

[0078]

[0079] where w′ is the updated weight parameter, and lr is the learning rate, denotes the derivative of w in the loss function;

[0080] f) Set the global variable T. After repeating the training of the neural network T times based on steps a)-e), observe whether the model converges. If it converges, save the model.

[0081] The invention principle of the present invention is:

[0082] In the present invention, by combining the advantages of IRS and index modulation, passive signals are transmitted while improving the communication quality of active signals by using IRS. First, the generalized orthogonal reflection mode of the fully activated IRS maximizes the inherent large-scale aperture gain of IRS to improve the bit error rate performance of active signals. Then, the UAV and IRS are integrated to improve the deployment flexibility of IRS. Combining the deep reinforcement learning method, taking the UAV as the object, discretize the actions, flight duration, and scheduling selection of the UAV in each time slot, and use a neural network to fit the Q value of each action to determine whether to select the current action. Solve the optimization problem with maximizing energy efficiency as the optimization goal.

[0083] The advantages and beneficial effects of the present invention are:

[0084] The present invention proposes a method for optimizing the active and passive transmission of a UAV-assisted intelligent reflecting surface based on deep reinforcement learning. IRS constructs passive signals using its own reflection mode index, achieving more with fewer resources without occupying additional frequency band resources. First, compared with the traditional reflection mode based on on / off states, the fully reflecting IRS can maximize the received signal-to-noise ratio of active signals. Then, compared with the active and passive reciprocal transmission scheme based on IRS phase index modulation, IRS can improve the bit error rate performance of active signals and reduce the performance impact brought by passive signals. The optimization method based on deep reinforcement learning can solve difficult non-convex form problems, and the training of its neural network can be regarded as offline learning, with a computational complexity much lower than that of traditional convex optimization methods, making the deployment of UAV in practice possible. Brief Description of the Drawings

[0085] Table 1 is a parameter table involved in the implementation of the method for optimizing the active and passive transmission of a UAV-assisted intelligent reflecting surface based on deep reinforcement learning proposed by the present invention;

[0086] Figure 1 It is an illustrative system diagram implemented according to the proposed method for optimizing the active and passive transmission of an intelligent reflecting surface assisted by a drone based on deep reinforcement learning of the present invention;

[0087] Figure 2 It is a comparison of the proposed deep reinforcement learning based optimization method and the benchmark optimization method in terms of the energy efficiency performance of the drone;

[0088] Figure 3 It is a performance comparison of the full reflection IRS index modulation proposed by the present invention and other IRS modulation schemes in terms of the overall system spectral efficiency;

[0089] Figure 4 It is a comparison of the active signal and the passive signal in terms of the bit error rate performance in the proposed active and passive modulation and optimization method. Detailed implementation mode

[0090] A method for optimizing the active and passive transmission of an intelligent reflecting surface assisted by a drone based on deep reinforcement learning, comprising the following steps:

[0091] 1) Construct an uplink communication system for UAV-IRS assisted Internet of Things devices, including a base station (BS) with N r receiving antennas, K ground Internet of Things (IoT) devices with single transmitting antennas, and a drone integrating L intelligent reflecting surface units. The active signals of the ground IoT devices in the system adopt M-QAM constellation modulation. Establish a Cartesian coordinate system for the movable area range of the drone, representing the three-dimensional coordinate positions of the drone and the IRS at time slot n, n = 1, 2,..., N, where N is the total number of time slots. Among them, represents the coordinate of the UAV in the horizontal direction, represents the coordinate of the UAV in the vertical direction; the IoT and the base station are in a stationary state, [x BS , y BS , z BS represents the coordinate position of the base station. The IoT is stationary on the ground, and its height can be almost ignored, that is, Therefore, represents the coordinate of the kth IoT device, k = 1, 2,..., K, where K represents the total number of Internet of Things devices.

[0092] 2) To ensure that the UAV assists IRS communication within a safe range, set the maximum height H max and the minimum height H min that the UAV can move in the vertical direction, that is, The duration of the UAV's movement in each time slot is where t minand t max respectively represent the minimum and maximum durations of the UAV in time slot n.

[0093] 3) According to steps 1) and 2), the speed of the UAV per unit time can be calculated based on distance and time. The horizontal flight speed of the UAV in time slot n can be expressed as:

[0094]

[0095] where, is the maximum moving speed of the UAV in the horizontal direction. At the same time, the vertical flight speed of the UAV in time slot n can be expressed as:

[0096]

[0097] where, represents the maximum moving speed of the UAV in the vertical direction. If and are equal to 0, it means that in this time slot n, the UAV is stationary in the horizontal and vertical directions respectively.

[0098] 4) Based on the speeds of the UAV in the horizontal and vertical directions described in step 3), the energy consumption required by the UAV in time slot n can be expressed as:

[0099]

[0100] where, P1, P2, P3, and P4 respectively represent the blade profile power of the UAV in the hover state, the induced power of the UAV in the hover state, the constant power of the UAV in the ascending or descending state, and the power required for the UAV to control the IRS. v tip represents the tip speed of the UAV rotor blade, d0 represents the self-drag ratio of the UAV, l represents the solidity of the UAV rotor, ρ represents the air density, G represents the rotor disk area, and v0 represents the average blade induction speed of the UAV in hover.

[0101] 5) Model the channel of the air-ground communication. The IoT-UAV and UAV-BS are modeled as Rice fading channels with deterministic line-of-sight components; the Rice fading channel model of the IoT-UAV link and the Rice fading channel model of the UAV-BS link are respectively expressed as:

[0102]

[0103]

[0104] where, κ1 and κ2 respectively represent the Rice factors of the two links, and The deterministic line-of-sight paths for the IoT-UAV link and the UAV-BS link, respectively; and respectively represent the non-Los paths with Rayleigh distribution for the IoT-UAV link and the UAV-BS link. Each element of the non-Los path follows the complex Gaussian distribution;

[0105] The calculation formula for the deterministic line-of-sight path is:

[0106]

[0107] where λ represents the wavelength, d represents the spacing between adjacent IRS elements, and it is assumed that and there is no coupling between IRS elements, ψ AoA is the angle of arrival; Considering the channel fading caused by the mobility of the UAV, based on the Rice fading channel model, it is necessary to calculate the channel fading caused by the path distance. The IoT-UAV link and the UAV-BS link including path loss can be further expressed as:

[0108]

[0109] and

[0110]

[0111] where τ represents the path loss exponent, f c is the carrier frequency, c is the speed of light, and respectively represent the direct distance between the k-th IoT device and the UAV and the direct distance between the UAV and the BS at time slot n. It is calculated from the coordinates of the scheduled IoT at time slot n and the coordinates of the UAV and can be expressed as:

[0112]

[0113] It is calculated from the coordinates of the UAV and the coordinates of the BS at time slot n and can be expressed as:

[0114]

[0115] 6) Assume that the input bit streams of the IoT device and the UAV device are grouped. Each group is divided into B1 = log2M and B2 = log2Q bits respectively. Both Q and M satisfy integer powers of 2. M represents the M-QAM constellation modulation order adopted by the IoT device, and Q represents the number of generalized orthogonal reflection modes after IRS grouping.

[0116] 7) Divide the L reflecting units of the IRS into adjacent L g groups, each group consists of L / L g reflecting units. Assume that L is divisible by L g . Activate g groups of IRS elements from the L g groups to reflect in-phase signals, where 1 ≤ g ≤ L g . Reactivate the remaining L g -g groups to generate orthogonal signals to maximize the received signal-to-noise ratio of the active signal. There are combinations of generalized orthogonal reflection patterns, denotes the binomial coefficient. Select the first Q generalized orthogonal reflection patterns from the combinations for the transmission of UAV information bits. Each generalized orthogonal reflection pattern can be expressed as:

[0117]

[0118] where 1 indicates that the IRS element is used for in-phase reflection, and j indicates that the reflection phase of the IRS element rotates clockwise

[0119] 8) Each constellation symbol s i mapped by the active signal is power-normalized and satisfies E[|s i | 2 = 1, where i = 1, 2,..., M. B1 bits are used for the constellation symbol index selection of the IoT active signal, and B2 bits are used for the generalized orthogonal reflection pattern index selection of the passive signal for the proximal UAV to transmit information. Assume that the channel state information is perfectly known at the receiver, and the received signal is expressed as:

[0120]

[0121] where diag(·) represents a diagonal matrix with elements on the main diagonal, P represents the fixed transmit power of the IoT device, is the additive white Gaussian noise AWGN, which follows the distribution CN(0, N0I Nr ), where I Nr represents the identity matrix, and N0 is the complex noise variance.

[0122] 9) Use the maximum likelihood (ML) detector to jointly detect the IoT active signal and the UAV passive signal:

[0123]

[0124] where, and They respectively represent the active signal index and the passive signal index detected by the receiving end, and the corresponding bit information is restored through the index value.

[0125] 10) In time slot n, the total number of successfully transmitted bits of the system can be expressed as:

[0126]

[0127] Among them, B w represents the bandwidth, represents the number of bits detected as successfully transmitted in step 9), represents whether the k-th IoT device is scheduled to communicate with the base station in time slot n;

[0128] 11) In the given total time slots N, the optimization problem of energy consumption can be defined as:

[0129]

[0130] Among them, the objective function is to minimize the total energy consumption of the UAV in all time slots;

[0131] To increase the total number of successfully transmitted bits of the system, the UAV adopts the generalized orthogonal reflection modulation of the full reflection of the IRS in step 7). In the total time slots, the total number of successfully transmitted bits received by the receiving end is To comprehensively measure the energy efficiency index of the UAV-IRS system, its objective function can be further expressed as:

[0132]

[0133] It can be seen from formula (16) that the objective function is a non-convex problem. If the traditional convex optimization method is used to solve it, it will cost huge computational costs and make it difficult to be possible in practical applications.

[0134] To solve the above problems, the present invention will use the deep Q-network (DQN) in the deep reinforcement learning (DRL) to solve the sub-optimal solution.

[0135] The construction process of the deep Q-network DQN is as in steps 12)-16):

[0136] 12) Regarding the UAV as an agent object, the three-dimensional coordinate position of the UAV in the current time slot is used as the state. Based on the current state S(n), it is input into the neural network of the DQN. The discrete action behavior A(n) output by the network contains the index information of the UAV in the horizontal, vertical, scheduling, and duration respectively. The DQN uses the Q of each action valueTo determine whether to select the A(n). The environment (Env) will enter the next state S(n+1) according to the UAV executing the action A(n) in the current state S(n), and perform scoring and give a reward R(n).

[0137] 13) Define the state Action Among them, a1(n) represents the movement index in the horizontal direction, and a2(n) represents the movement index in the vertical direction, For the scheduling of IoT, Is a discrete variable of the duration, and its time interval is Δt. According to formula (16), the reward function can be defined as:

[0138]

[0139] Among them, ζ={1,200} is the penalty coefficient. When the flight area of the UAV integrated with IRS exceeds the controllable range of the specified moving area, Env gives a negative feedback ζ=200.

[0140] 14) When Env updates the state S(n+1) according to the executed action A(n), Δx={0,+x u ,-x u}, x u Represents the spacing between two adjacent coordinate points on the x-axis. Similarly, Δy={0,+y u ,-y u}, y u Represents the spacing between two adjacent coordinate points on the y-axis. Δz={0,+z u ,-z u}, z u Represents the spacing between two adjacent coordinates on the z-axis.

[0141] 15) The Q in step 12) value Can be defined as:

[0142] Q value (S(n),A(n)) = E[U(n)|S(n),A(n)] (65)

[0143] Among them, E represents the expectation, and U(n) represents the discounted return; U(n) is represented as:

[0144] U(n) = R(n) + γR(n+1) + γ 2 R(n+2) + … (66)

[0145] Among them, γ represents the discounted return factor. Q value(S(n), A(n)) represents the conditional expectation of the return U(n), aiming to eliminate the states and actions involved after the n-th time slot in U(n), and score the quality of taking action A(n) for the current state S(n). Using formula (18), DQN adopts a strategy of traversing the maximum value to find an action A(n) corresponding to the maximum Q value The strategy function for traversing the maximum value can be expressed as:

[0146]

[0147] where π′(·) represents the policy function.

[0148] 16) In deep reinforcement learning, it is necessary to collect training set data, and each data set can be represented as:

[0149]

[0150] where S(n) is the current state, A(n) is the action to be executed, R(n) is the reward obtained by executing action A(n) based on the current state S(n), and S(n + 1) is the next state entered after executing action A(n); the UAV interacts with the environment to obtain training data After that, each piece of data is stored in the experience pool B buff When the number of data sets in the experience pool reaches the set threshold M size After that, DQN starts to train the neural network. During the training process, as the UAV and the environment continuously interact, the latest data is dynamically obtained to replace the old data sets in the experience pool.

[0151] The training method of the deep Q-network is specifically as follows:

[0152] a) The number of input layers of the neural network model is 3, the number of neurons in the hidden layer is 20, and the number of output layers is related to the number of behaviors A(n) of the UAV. The number of A(n) is There are 5 discrete actions in the horizontal direction and 3 discrete actions in the vertical direction. The discrete time variable is related to the accuracy of Δt. Assume that Δt can be divided by t max -t min divisible; the activation function of the neuron uses the rectified linear unit (ReLU) function:

[0153]

[0154] In the DQN algorithm, the neural network is used to fit Q in step 15) value; To alleviate the problem of slow model convergence caused by model errors, when selecting action A(n), the greediness is set to ε = 0.9, that is, there is a 10% probability of randomly selecting the action of the UAV instead of always relying on the neural network to select; the learning rate lr = 0.005, and the size of the experience pool is M size = 3200.

[0155] b) Set the initial coordinate position of the UAV as S(0) = [0, 0, H max , interact with the Env and randomly obtain valid data and store it in the experience pool B buff . When the amount of data in the experience pool reaches the set threshold M size , the neural network starts training.

[0156] c) Perform batch training on the data in the experience pool, and the size of the batch training data is M bath = 25, M bath << M size . Send the training data into the neural network in batches and perform forward calculations layer by layer until the output layer.

[0157] d) Use the mean square error loss function to calculate the loss:

[0158] L(w) = E[R(n) + γQ value (S(n + 1), A(n + 1)|w) - Q value (S(n), A(n)|w) 2 (70)

[0159] where L(w) is the loss function and w is the weight parameter of the neural network.

[0160] e) Through the chain rule, calculate the gradient of the loss function with respect to each layer layer by layer for backpropagation, and use the stochastic gradient descent algorithm to update the weight parameters of the neural network:

[0161]

[0162] where w′ is the updated weight parameter, lr is the learning rate in step 18.1), represents the derivative of w in the loss function.

[0163] f) Set the global variable T. After repeating the training of the neural network T times based on steps a) - e), observe whether the model converges. If it converges, save the model and compare the energy efficiency performance with the benchmark solution, and compare the bit error rate performance index of the traditional active and passive transmission schemes.

[0164] Next, specific embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0165] To demonstrate the superiority of the proposed UAV-assisted intelligent reflecting surface active and passive transmission optimization method based on deep reinforcement learning, the simulation parameters are configured as shown in Table 1. In the Cartesian coordinate system, the maximum horizontal movement range of the UAV is 0 - 500, the base station coordinates are [0, 0, 50], and the horizontal coordinates of K Internet of Things devices are randomly generated within the range of 450 - 500. Figure 2 The simulation results are used to evaluate the comparison of the DQN algorithm based on deep reinforcement learning and the benchmark greedy algorithm in terms of the energy efficiency performance of the UAV. At the same time, based on the DQN algorithm, the generalized orthogonal reflection modulation (GQRM) is compared with the existing active and passive transmission schemes in terms of the energy efficiency performance of the UAV. In the GQRM scheme, all Internet of Things devices use 16-QAM constellation symbols, and the number of reflection modes Q = 16. In the IRS modulation scheme based on ON / OFF, the Internet of Things devices use 16-QAM constellation symbols, and the number of reflection modes Q = 16. In the symbiotic space modulation (SSM) based on IRS phase index modulation, the Internet of Things devices use 16-PSK constellation symbols, and the IRS phase index is 16 to jointly form 256-APSK constellation points. Figure 3 The simulation results are used to evaluate the performance comparison between the overall spectral efficiency and the reflection power of the system under different modulation schemes in the method optimized based on deep reinforcement learning, where the transmit power range is from 0 dBm to 35 dBm. Figure 4 The simulation results are used to evaluate the influence of different modulation methods on the active and passive signals respectively in the method optimized based on deep reinforcement learning. The transmit power range is from 0 dBm to 30 dBm, and the bit error rate performance between the Internet of Things signal (active) and the UAV signal (passive) is compared respectively.

[0166] Table 1

[0167]

[0168] In Figure 2 , the cumulative distribution function (CDF) of the energy efficiency performance of the DQN algorithm and the benchmark greedy algorithm is compared. The DQN algorithm can obtain a performance gain of approximately 0.5 - 1 bits / J compared with the Greedy algorithm under the same modulation method. At the same time, under the DQN algorithm, the generalized orthogonal reflection modulation scheme based on the fully reflecting IRS can obtain better energy efficiency performance compared with the existing IRS modulation schemes.

[0169] From Figure 3 it can be seen that the generalized orthogonal reflection mode modulation method based on full reflection can obtain a higher successful transmission bit rate than the modulation methods based on ON / OFF states and IRS phase indices under the same transmit power.

[0170] In Figure 4Among them, B1 represents the number of bits of the IoT active signal, and B2 represents the number of bits of the UAV passive signal. In the active and passive transmission scheme of SSM, when B1 = 4 and B2 = 3, by observing Figure 4 (a) and Figure 4 (b), it can be found that the bit error rate performance of the active signal and the passive signal is very close, with a difference of about 0.1 dBm. However, when B1 = 4 and B2 = 4, as the spectral efficiency of the passive signal increases, while its bit error rate performance deteriorates, the bit error rate performance of the active signal is also affected accordingly. This is mainly because the phase modulation of the passive signal using IRS will change the phase of the active signal, resulting in a performance decline. In the active and passive transmission scheme based on the ON / OFF state, the bit error rate of the UAV passive signal almost coincides with that of the passive signal in the SSM scheme. However, for the active signal, at the same spectral efficiency, a better bit error rate performance can be obtained. This is mainly because the performance gain brought by modulating using the IRS spatial index resource is greater than that of the IRS phase index modulation. In addition, in addition to improving the bit error rate performance of the active signal, the generalized orthogonal reflection modulation scheme (GQRM) adopted in the present invention can maximize the signal-to-noise ratio of the active signal. In Figure 4 (a), at the same spectral efficiency, the bit error rate performance of GQRM is 3 dBm better than that of the ON / OFF scheme. This is mainly due to the performance gain brought by the total reflection of IRS. It should be noted that in the GQRM scheme, when changing the spectral efficiency of the passive signal, the bit error rate performance of the IoT active signal will not be affected like the phase index modulation.

Claims

1. A method for optimizing the active and passive transmission of an Unmanned Aerial Vehicle (UAV)-assisted Intelligent Reflecting Surface (IRS) based on deep reinforcement learning, characterized in that, This method is implemented based on the uplink communication system of UAV-IRS assisted Internet of Things (IoT) devices. The uplink communication system of UAV-IRS assisted IoT devices includes a base station BS with N r receiving antennas, K ground IoT devices each with a single transmitting antenna, and a drone UAV integrated with L intelligent reflecting surface units. The active signals of the ground IoT devices adopt M-QAM constellation modulation, and the passive signals of the UAV adopt generalized orthogonal reflection modulation. The optimization method includes the following steps: 1) Establish a Cartesian coordinate system for the movable area range of the UAV, represent the three-dimensional coordinate positions of the UAV and the IRS at time slot n, where n = 1, 2,..., N, and N is the total number of time slots. Among them, represents the coordinate of the UAV in the horizontal direction, represents the coordinate of the UAV in the vertical direction; [x BS , y BS , z BS represents the coordinate position of the base station, represents the coordinate of the k-th IoT device, where k = 1, 2,..., K, and K represents the total number of IoT devices; 2) Set the height range within which the UAV can move vertically: Among them, H min and H max respectively represent the maximum and minimum heights of the UAV in the vertical direction; the duration of the UAV's movement in each time slot The value range of: is Among them, t min and t max respectively represent the minimum and maximum times that the UAV lasts in time slot n; 3) According to steps 1) and 2), calculate the speed of the UAV per unit time through distance and time. The horizontal flight speed of the UAV in time slot n is expressed as: Among them, is the maximum moving speed of the UAV in the horizontal direction. At the same time, the flight speed of the UAV in the vertical direction at time slot n is expressed as: Among them, represents the maximum moving speed of the UAV in the vertical direction; if and are equal to 0, it means that in time slot n, the UAV is stationary in the horizontal and vertical directions respectively; 4) After calculating the speeds of the UAV in the horizontal and vertical directions according to step 3), the energy consumption required by the UAV in time slot n is expressed as: Among them, P1, P2, P3, and P4 respectively represent the blade profile power of the UAV in the hover state, the induced power of the UAV in the hover state, the constant power of the UAV in the ascending or descending state, and the power required for the UAV to control the IRS; v tip represents the tip speed of the UAV rotor blade, d0 represents the self-resistance ratio of the UAV, l represents the solidity of the UAV rotor, ρ represents the air density, G represents the rotor disk area, and v0 represents the average blade induction speed of the UAV in hover; 5) Modeling the air-to-ground communication channel, IoT-UAV and UAV-BS are modeled as Rician fading channels with deterministic line-of-sight components; Rician fading channel model for IoT-UAV link Rician fading channel model for UAV-BS link Respectively expressed as: where κ1 and κ2 represent the Rice factors of the two links, respectively, and are the deterministic line-of-sight paths of the IoT-UAV link and the UAV-BS link, respectively; and represent the non-Los paths with Rayleigh distribution of the IoT-UAV link and the UAV-BS link, respectively. Each element of the non-Los path follows the complex Gaussian distribution; The calculation formula for the deterministic line-of-sight path is: where λ represents the wavelength, d represents the spacing between two adjacent IRS elements, and it is assumed that and there is no coupling between IRS elements, ψ AoA is the angle of arrival; considering the channel fading caused by the mobility of the UAV, the channel fading caused by the path distance needs to be calculated based on the Rice fading channel model. The IoT-UAV link and the UAV-BS link including path loss are further expressed as follows: and where τ represents the path loss exponent, f c is the carrier frequency, c is the speed of light, and respectively represent the direct distance between the k-th IoT device and the UAV and the direct distance between the UAV and the BS at time slot n; is calculated from the coordinates of the scheduled IoT at time slot n and the coordinates of the UAV, and the specific calculation formula is: It can be calculated from the coordinates of the UAV and the coordinates of the BS at time slot n and can be expressed as: 6) Assume that the input bit streams of IoT devices and UAV devices are grouped. Each group is divided into B1 = log2M and B2 = log2Q bits respectively. Both Q and M satisfy integer powers of 2. M represents the M-QAM constellation modulation order adopted by IoT devices, and Q represents the number of generalized orthogonal reflection modes after IRS grouping; 7) Divide the L reflection units of the IRS into adjacent L g groups, with each group consisting of L / L g reflection units. Assume that L is divisible by L g . Activate g groups of IRS elements from the L g groups to reflect in-phase signals, where 1 ≤ g ≤ L g . Reactivate the remaining L g -g groups to generate orthogonal signals to maximize the received signal-to-noise ratio of the active signal. There are a total of combinations of generalized orthogonal reflection patterns, denoting the binomial coefficient. Select the first Q generalized orthogonal reflection patterns from the combinations for the transmission of UAV information bits; each generalized orthogonal reflection pattern is represented as: Among them, 1 indicates that the IRS element is used for co-phase reflection, and j indicates that the reflection phase of the IRS element rotates clockwise 8) Each constellation symbol s mapped by the active signal i After power normalization and satisfying E[|s i | 2 = 1, i = 1, 2,..., M; B1 bits are used for the constellation symbol index selection of the IoT active signal, and B2 bits are used for the generalized orthogonal reflection mode index selection of the passive signal for the IRS to transmit information to the proximal UAV; when receiving, it is assumed that the channel state information is completely known, and the received signal Is expressed as: where, diag(·) represents a diagonal matrix with elements on the main diagonal, P represents the fixed transmit power of the IoT device, is the additive white Gaussian noise AWGN, which follows the distribution CN(0, N0I Nr ), where, I Nr represents the identity matrix, and N0 is the complex noise variance; 9) Use the maximum likelihood (ML) detector to jointly detect the active signals of IoT and the passive signals of UAV: Among them, and respectively represent that the receiving end detects the active signal index and the passive signal index, and restores the corresponding bit information through the index value; 10) In time slot n, the total number of successfully transmitted bits of the system is expressed as: Among them, B w represents the bandwidth, represents the number of successfully transmitted bits detected in step 9), represents whether the k-th IoT device is scheduled to communicate with the base station in time slot n; 11) In the given total time slots N, the optimization problem of energy consumption is defined as: Among them, the objective function is to minimize the total energy consumption of the UAV in all time slots; Under the total time slots, the total number of successfully transmitted bits received by the receiver is To comprehensively measure the energy efficiency index of the UAV-IRS system, its objective function can be further expressed as: Use the deep Q-network DQN to solve the sub-optimal solution. The construction process of the deep Q-network DQN is as in steps 12)-16): 12) Taking the UAV as an agent object, the three-dimensional coordinate position of the UAV in the current time slot is used as the state, which is input into the neural network of the DQN based on the current state S(n). The discrete action behavior A(n) output by the network contains the index information of the UAV in the horizontal, vertical, scheduling, and duration aspects respectively. The DQN uses Q value to evaluate the value of the action to decide whether to select the action A(n); the environment Env enters the next state S(n+1) after the UAV executes the action A(n) in the current state S(n), and scores and gives a reward R(n); 13) Define the state Action where a1(n) represents the motion index in the horizontal direction, and a2(n) represents the motion index in the vertical direction For the scheduling of IoT is a discrete variable of the duration, with a time interval of Δt; according to formula (16), the reward function is defined as: Among them, ζ = {1, 200} is the penalty coefficient. When the flight area of the UAV integrated with IRS exceeds the controllable range of the specified moving area, Env gives negative feedback ζ = 200; 14) When Env updates the state S(n + 1) according to the executed action A(n), Δx = {0, +x u , -x u}, where x u represents the distance between two adjacent coordinate points on the x-axis; similarly, Δy = {0, +y u , -y u}, where y represents the distance between two adjacent coordinate points on the y-axis; Δz = {0, +z u , -z u}, where z u represents the distance between two adjacent coordinates on the z-axis; 15) The Q in step 12) value is defined as: Q value (S(n), A(n)) = Ε[U(n)|S(n), A(n)] (18) Among them, Ε represents expectation, and U(n) represents discounted return; U(n) is expressed as: U(n) = R(n) + γR(n + 1) + γ 2 R(n + 2) + … (19) where γ represents the discounted return factor; Q value (S(n), A(n)) is the conditional expectation of the return U(n), which aims to eliminate the states and actions involved after the n-th time slot in U(n) and score the quality of taking action A(n) for the current state S(n); through formula (18), DQN adopts a strategy of traversing the maximum value to find Q value an action A(n) corresponding to the maximum value, and the strategy function for traversing the maximum value is expressed as: Among them, π′(·) represents the policy function; 16) In deep reinforcement learning, it is necessary to collect training data sets. Each data set is expressed as: Among them, S(n) is the current state, A(n) is the action to be executed, R(n) is the reward obtained by executing action A(n) based on the current state S(n), and S(n+1) is the next state entered after executing action A(n); the UAV interacts with the environment to obtain training data After that, each piece of data is stored in the experience pool B buff In it, when the number of datasets in the experience pool reaches the set threshold M size After that, DQN starts to train the neural network. During the training process, as the UAV continuously interacts with the environment, the latest data is dynamically obtained to replace the old datasets in the experience pool.

2. The method for optimizing the active and passive transmission of an intelligent reflecting surface assisted by an unmanned aerial vehicle based on deep reinforcement learning according to claim 1, wherein The training method of the deep Q-network DQN is specifically as follows: a) The number of input layers of the neural network model is 3, the number of neurons in the hidden layer is 20, and the number of output layers is related to the number of UAV behaviors A(n). The number of A(n) is There are 5 discrete actions in the horizontal direction and 3 discrete actions in the vertical direction. The discrete time variable is related to the accuracy of Δt. Assume that Δt can be divided by t max -t min exactly; The activation function of the neuron uses the rectified linear unit (ReLU) function: When selecting action A(n), set the greediness to ε = 0.9, the learning rate lr = 0.005, and the size of the experience pool to M size = 3200; b) Set the initial coordinate position of the UAV as S(0) = [0, 0, H max , interact with the Env and randomly obtain valid data to store in the experience pool B buff . When the amount of data in the experience pool reaches the set threshold M size , the neural network starts training; c) Batch train the data in the experience pool, and the size of the batch training data volume is M bath = 25, M bath << M size ; Feed the training data into the neural network in batches and perform forward calculations layer by layer until the output layer; d) Calculate the loss using the mean square error loss function: L(w) = E[R(n) + γQ value (S(n + 1), A(n + 1)|w) - Q value (S(n), A(n)|w) 2 (23) Among them, L(w) is the loss function, and w is the weight parameter of the neural network; e) Through the chain rule, calculate the gradient of the loss function with respect to each layer layer by layer for backpropagation, and use the stochastic gradient descent algorithm to update the weight parameters of the neural network: where w′ is the updated weight parameter and lr is the learning rate, denotes the derivative of w in the loss function; f) Set the global variable T. Based on steps a)-e), repeat training the neural network T times and observe whether the model converges. If it converges, save the model.

Citation Information

Patent Citations

  • Unmanned aerial vehicle auxiliary ground communication method based on trajectory and phase joint optimization

    CN114980169A

  • Multi-user cooperative transmission method driven by intelligent reflecting surface

    CN115037337A