Method for joint beamforming and drone trajectory optimization for aerial ris-assisted mu-miso system
By using deep neural networks and deep reinforcement learning methods, the beamforming of the base station and RIS and the trajectory of the drone are jointly optimized, solving the problem of optimizing the energy consumption and transmission rate of the drone in the aerial RIS-assisted MU-MISO system, achieving improved system performance and reduced energy consumption.
Patent Information
- Application Number
- CN202510249178.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-03-04
AI Technical Summary
In the existing technology, the aerial RIS-assisted MU-MISO system is relatively mature in optimizing the transmission rate maximization, but there is less research on the energy consumption of UAVs, especially the technical research on how to jointly optimize the system transmission rate, the total number of flight time slots and the energy consumption of UAVs.
An aerial RIS-assisted MU-MISO system model was constructed. Through deep neural networks and deep reinforcement learning methods, the base station's active beamforming, the RIS's passive beamforming, and the drone's trajectory were jointly optimized to maximize the average total rate, minimize the number of flight slots, and reduce the drone's energy consumption. A deep neural network was used to train the base station and RIS's beamforming, and deep reinforcement learning was combined to optimize the drone's trajectory.
A good balance between complexity and performance is achieved, maximizing system transmission efficiency, reducing drone energy consumption, and optimizing the transmission performance of the aerial RIS-assisted MU-MISO system.
Smart Images

Figure CN120090677B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of unmanned aerial vehicle trajectory optimization, and particularly relates to a joint beamforming and unmanned aerial vehicle trajectory optimization method for an aerial RIS assisted MU-MISO system. BACKGROUND
[0002] Reconfigurable intelligent surface (RIS) has shown significant application potential in saving transmission power and improving communication performance. RIS is composed of a large number of low-cost passive reflecting units, which can intelligently adjust the amplitude and phase shift of the incident signal through a controller, thereby realizing flexible reconstruction of the wireless propagation environment. By creating an additional reflection channel between the sending end and the receiving end, RIS provides an effective means to improve the transmission performance of the communication system.
[0003] The application scenarios of RIS can be divided into fixed deployment and aerial deployment. Aerial deployment usually refers to integrating RIS on a UAV platform to form a so-called aerial RIS. The aerial RIS combines the high mobility of the UAV and the low power consumption and flexible configuration of the RIS, and can flexibly adjust the deployment position to provide better performance for the communication system. However, current researches mostly focus on the optimization problem of maximizing the transmission rate, and the research on the energy consumption of the UAV is relatively less, especially the research on how to jointly optimize the system transmission rate, the total time slot number of flight and the energy consumption of the UAV is relatively scarce. SUMMARY
[0004] The purpose of the present application is to provide a joint beamforming and unmanned aerial vehicle trajectory optimization method for an aerial RIS assisted MU-MISO system, which considers that the UAV can flexibly deploy the RIS by carrying the RIS, and the base station can provide better communication services for users.
[0005] The technical scheme of the present application is a joint beamforming and unmanned aerial vehicle trajectory optimization method for an aerial RIS assisted MU-MISO system, comprising the following steps:
[0006] The joint beamforming and unmanned aerial vehicle trajectory optimization method for an aerial RIS assisted MU-MISO system comprises the following steps:
[0007] An aerial RIS assisted MU-MISO system model is constructed, including a downlink transmission model, a motion model of the UAV, a communication model and a UAV energy consumption model;
[0008] A multi-objective optimization problem is constructed, and the objectives include maximizing the achievable average total rate, minimizing the total time slot number of flight, and reducing the energy consumption of the UAV;
[0009] training a deep neural network jointly optimizing active beamforming of the base station and passive beamforming of the RIS; including a beamforming network and a phase shift network, input of the phase shift network being channel information, passive beamforming of the RIS being obtained according to output of the phase shift network, input of the beamforming network being concatenated channel information after phase shift adjustment of the RIS, active beamforming of the base station being obtained according to output of the beamforming network;
[0010] constructing and optimizing a deep reinforcement learning model including a state function, an action space and a reward function, the state function including UAV position information, final coordinate information and user information, the action space including UAV flight speed, azimuth angle and elevation angle, the reward function being obtained according to the trained deep neural network;
[0011] obtaining optimal joint beamforming and UAV trajectory according to the optimized deep reinforcement learning model and the trained deep neural network.
[0012] Further, the multi-objective optimization problem includes:
[0013]
[0014] wherein, Λ represents the optimized multi-objective, N is the number of time slots, M is the number of users, is the achievable rate of user m in time slot n, E N is the energy consumption of the UAV in N time slots, W = [W1, W2,..., W N ] represents all active beamformings of the base station in N time slots, W n represents active beamforming of the base station in time slot n, Θ = [θ1, θ2,..., θ N ] is all passive beamformings of the RIS in N time slots, θ n represents passive beamforming of the RIS in time slot n, C is a three-dimensional trajectory of the UAV from the initial coordinate C 0 to the final coordinate C Ter , v n is the speed of the UAV in time slot n, v max is the maximum flight speed of the UAV, ξ n is the azimuth angle of the UAV in time slot n, φ n is the elevation angle of the UAV in time slot n, h n is the Z-direction height coordinate of the UAV in time slot n, h min and h max are the minimum and maximum flight heights of the UAV, x n and y n are the X-direction and Y-direction coordinates of the UAV in time slot n, D is the length of the rectangular flight area, C0 is the take-off point, C 0is the starting coordinate of the UAV, C N is the location of the UAV after N time slots, C Ter is the final coordinate of the drone, is the discrete phase shift of the kth RIS reflection unit in time slot n, F represents the number of quantization bits, and w n,m represents the active beamforming of the base station to user m in time slot n, P B is the maximum transmit power.
[0015] Furthermore, at each time slot n, given the current position of the UAV, the problem of jointly optimizing active and passive beamforming is formulated as:
[0016]
[0017] Among them, W n represents the active beamforming of the base station in time slot n, θ n represents the passive beamforming of RIS in time slot n, K is the number of reflection units of RIS, is the achievable rate of user m in time slot n, is the discrete phase shift of the kth RIS reflection unit in time slot n, F represents the number of quantization bits, and w n,m represents the active beamforming of the base station to user m in time slot n, P B is the maximum transmission power, M is the number of users;
[0018] use represents the output of the phase shift network at time slot n. Then, at time slot n, the complex reflection coefficient of the kth RIS reflection unit under continuous phase shift is expressed as:
[0019]
[0020] in, is the complex reflection coefficient of the kth RIS reflection unit in time slot n under continuous phase shift, and For reflection unit k, represents the amplitude of the kth reflection unit in time slot n, represents the continuous phase shift of the kth RIS reflection unit in time slot n, is the output of the phase shift network in time slot n The kth element of is the output of the phase shift network in time slot n The K+kth element of ; is the output of the phase shift network in time slot n The kth element of The output of the phase shift network for time slot n is The K+kth element of ; beamforming of RIS middle denotes the complex reflection coefficient of the k-th RIS reflecting unit under the discrete phase shift at time slot n, denotes the discrete phase shift of the k-th RIS reflecting unit at time slot n, which is selected by quantization medium distance The latest value obtains k = 1, 2, …, K; and the passive beamforming Θ of the RIS is obtained:
[0021] denotes the output of the beamforming network, then the active beamforming of the base station to the user m at time slot n is where each element is denoted as:
[0022]
[0023] where m ∈ {1, 2, …, M} represents the user index, n B ∈ {1, 2, …, N B} represents the base station antenna index; is the (m-1)N B +n B element of the output of the beamforming network at time slot n, is the n B +MN B +(m-1)N B element of the output of the beamforming network at time slot n, N B is the number of antennas contained by the base station, and finally the active beamforming W of the base station is obtained.
[0024] Further, the deep neural network model is trained offline in an unsupervised manner, and the loss function is:
[0025]
[0026] where, is the achievable rate of the user m, and M is the number of users; the parameters of the neural network are updated using the backpropagation gradient to minimize the loss function.
[0027] Further, the UAV position information includes the position information of the UAV, the azimuth angle, the elevation angle and the flight speed of the UAV, the final coordinate information includes the final coordinate of the UAV, the azimuth angle from the UAV to the final coordinate, the elevation angle and the distance to the final coordinate, the user information includes the transmission rate of all user equipment; the reward function is a weighted combination of the reward obtained by the UAV reaching the final coordinate, the system transmission rate reward, the energy consumption of the UAV and the motion reward of the UAV, wherein the system transmission rate reward is calculated by the achievable rate of the user, and the achievable rate is calculated by loading the trained deep neural network.
[0028] Further, the reward function r n is:
[0029]
[0030] where a, b and c are coefficients balancing the importance of reaching the final coordinates, improving the system rate and reducing energy consumption, is the reward for reaching the final coordinates by the UAV, defined as:
[0031]
[0032] where, and denote the distance of the UAV to the target point in the previous and current time slot, respectively, v max is the maximum speed of the UAV, δ n is the time length of the time slot, β penalty denotes the penalty of each time slot in the flight process;
[0033] is the reward of the system transmission rate in time slot n, defined as:
[0034]
[0035] where, denotes the achievable rate of user m in time slot n, which needs to load the trained deep neural network to calculate, taking the channel information as the input of the phase shift network, to obtain the beamforming of the RIS in time slot n, R H d,m denotes the direct link information from the base station to user m, H n,m denotes the cascaded link information from the base station to user m through the RIS in time slot n, T denotes the transpose of a vector; taking the cascaded channel information after adjusting the phase shift of the RIS as the input of the beamforming network, denotes the channel gain of the direct link from the base station to user m, denotes the channel gain of the cascaded link from the base station-RIS-user m, M is the number of users, to obtain the active beamforming of the base station to user m in time slot n, F N B is the number of antennas contained by the base station, the calculation formula is:
[0036]
[0037] where w n,i denotes the active beamforming of the base station to user i in time slot n, F denotes the noise variance of the received signal of user m;
[0038] for the energy consumption of the UAV in time slot n, for the movement reward of the UAV.
[0039] Further, the optimal joint beamforming and UAV trajectory acquisition method is:
[0040] The UAV inputs the channel state information and at each time slot n according to the current position and the trained deep neural network model n and the passive beamforming θ of the RIS n of the base station at time slot n; d,m denotes the direct link information from the base station to user m, H n,m denotes the concatenated link information from the base station to user m through the RIS at time slot n, and T denotes the transpose of a vector; denotes the channel gain of the direct link from the base station to user m, denotes the channel gain of the concatenated link from the base station-RIS-user m, and M is the number of users;
[0041] The UAV inputs the state information s n using the trained deep reinforcement learning model; and outputs the optimal action policy, i.e., the action of the UAV at the current time slot, including the flight speed v n , the flight azimuth angle ξ n , and the elevation angle φ n , which is represented as:
[0042] a n ={v n ,ξ n ,φ n}.
[0043] The system corresponding to the joint beamforming and UAV trajectory optimization method for the aerial RIS-aided MU-MISO system includes:
[0044] A system model construction unit for constructing an aerial RIS-aided MU-MISO system model, including: a downlink transmission model, a movement model of the UAV, a communication model, and a UAV energy consumption model;
[0045] A target construction unit for constructing a multi-objective optimization problem, whose objectives include: maximizing the achievable average total rate, minimizing the total time slot of flight, and reducing the energy consumption of the UAV;
[0046] The deep neural network design and training unit is configured to train a deep neural network that jointly optimizes active beamforming of the base station and passive beamforming of the RIS; the deep neural network comprises a beamforming network and a phase shift network, the input of the phase shift network is channel information, and passive beamforming of the RIS is obtained according to the output of the phase shift network, the input of the beamforming network is concatenated channel information after phase shift adjustment of the RIS, and active beamforming of the base station is obtained according to the output of the beamforming network;
[0047] The deep reinforcement learning model construction unit is configured to construct and optimize a deep reinforcement learning model comprising a state function, an action space and a reward function, the state function comprises UAV position information, final coordinate information and user information, the action space comprises UAV flight speed, azimuth angle and elevation angle, and the reward function is obtained according to the trained deep neural network;
[0048] The joint beamforming and UAV trajectory optimization unit is configured to obtain optimal joint beamforming and UAV trajectory according to the optimized deep reinforcement learning model and the trained deep neural network.
[0049] An electronic device for storing and executing the method comprises:
[0050] A memory storing executable program codes;
[0051] A processor coupled with the memory;
[0052] The processor invokes the executable program codes stored in the memory to execute the steps of the joint beamforming and UAV trajectory optimization method for the aerial RIS-aided MU-MISO system.
[0053] A computer readable storage medium for storing and executing the method, the computer readable storage medium stores computer instructions, and the computer instructions are configured to execute the steps of the joint beamforming and UAV trajectory optimization method for the aerial RIS-aided MU-MISO system when invoked.
[0054] Beneficial effects: Compared with the prior art, the significant technical effects of the present application are: (1) the unmanned aerial vehicle uses the deep reinforcement learning optimization strategy to obtain the optimal unmanned aerial vehicle flight angle and unmanned aerial vehicle flight speed, and the base station and RIS use the deep neural network to obtain the optimal active beamforming and passive beamforming; (2) by jointly optimizing the active beamforming of the base station, the passive beamforming of the RIS, the flight angle of the unmanned aerial vehicle, and the flight speed of the unmanned aerial vehicle, the system transmission efficiency is maximized; (3) using the SD3 algorithm and the deep neural network combination can effectively solve the joint optimization problem of the active beamforming of the base station, the passive beamforming of the RIS, the flight angle of the unmanned aerial vehicle, and the flight speed of the unmanned aerial vehicle in the aerial RIS assisted MU-MISO scene, and has relatively low complexity; (4) in the aerial RIS assisted MU-MISO scene, the method of the present application is superior in maximizing the average total rate that can be reached, minimizing the total time slot number of flight, and reducing the energy consumption of the unmanned aerial vehicle. Additional aspects and advantages of the present application will be in part given in the following description, some will become apparent from the following description, or will be learned by practice of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 A flowchart of a joint computation migration and unmanned aerial vehicle trajectory optimization method based on deep reinforcement learning provided according to an embodiment of the present application;
[0056] Figure 2 An aerial RIS assisted MU-MISO model schematic diagram provided according to an embodiment of the present application;
[0057] Figure 3 A simulation comparison diagram of a joint beamforming and unmanned aerial vehicle trajectory optimization method for an aerial RIS assisted MU-MISO system provided according to an embodiment of the present application and other algorithms. DETAILED DESCRIPTION
[0058] For scenarios where aerial RIS assists a Multi-User Multiple-Input Single-Output (MU-MISO) system, drones equipped with RIS can flexibly deploy RIS to assist base stations in providing better communication services to users. A multi-objective optimization problem is formulated to maximize the achievable average overall rate, minimize the total number of flight time slots, and reduce the drone's energy consumption. The goal is to jointly optimize the base station's active beamforming, the RIS's passive beamforming, and the drone's trajectory. A deep neural network is constructed and trained to jointly optimize the base station's active beamforming and the RIS's passive beamforming. The drone is treated as an intelligent agent, and a deep reinforcement learning model is established. An algorithm is used to optimize the deep reinforcement learning model. Based on the trained deep neural network and deep reinforcement learning model, the joint beamforming and drone trajectory are derived, achieving a good balance between complexity and performance, achieving multi-objective optimization.
[0059] The present invention is described in further detail below.
[0060] like Figure 1 As shown, the joint beamforming and UAV trajectory optimization method for the airborne RIS-assisted MU-MISO system of the present invention includes the following steps:
[0061] Step 1: Construct an aerial RIS-assisted MU-MISO system model, which includes an N B A base station with 1 antenna, M single-antenna users, and a drone. The M single-antenna users use a collection of It means that the UAV is equipped with a RIS consisting of K reflective units to enhance the communication from the base station to the user. The location of the base station is represented by C BS =(x BS ,y BS ,h BS ) T Indicates that x BS 、y BS 、h BS are the coordinates of the base station in the X, Y and Z directions respectively; the position of user m is represented by C m =(x m ,y m ,0) T Indicates that x m 、y m are the coordinates of user m in the X and Y directions respectively; specifically:
[0062] (1a) Construct the UAV motion model, considering that the UAV flight cycle consists of N time slots, Indicates that the position of the drone at time slot n is represented by C n =(x n ,yn h n ) T represent, x n , y n and h n are respectively the X direction coordinate, Y direction coordinate and Z direction height coordinate of the UAV in time slot n, which can be calculated from the position of the last time slot, that is:
[0063] x n = x n-1 + d n cos(ξ n )cos(φ n )(1)
[0064] y n = y n-1 + d n cos(ξ n )sin(φ n )(2)
[0065] h n = h n-1 + d n sin(ξ n )(3)
[0066] wherein x n-1 , y n-1 and h n-1 are respectively the X direction coordinate, Y direction coordinate and Z direction height coordinate of the UAV in time slot n-1, d n represents the distance of the UAV flight, which is determined by its speed v n and the time slot length δ n , that is d n = v n δ n , ξ n and φ n respectively represent the azimuth angle and the elevation angle of the UAV, in addition, the UAV is limited to fly in a rectangular area with length and width of D meters, and the height is limited between the minimum height h min and the maximum height h max , the starting coordinate and the final coordinate of the UAV flight are respectively C 0 and C Ter ;
[0067] (1b) construct a communication model, G n represents the channel gain of the base station to user m direct link, G n represents the channel gain of the base station to RIS link in time slot n, which is determined by the position C n,m of the UAV in time slot n, HnRIS,m (n) denotes the channel gain of the link from the RIS to user m at time slot n, and further the cascade link channel gain of the base station-RIS-user m may be expressed as:
[0068]
[0069] wherein, HnRIS (n) denotes the phase shift matrix of the RIS at time slot n, is the complex reflection coefficient of the kth RIS reflecting element under a discrete phase shift at time slot n, and For the reflecting element k, denotes the amplitude of the kth reflecting element at time slot n, denotes the discrete phase shift of the kth RIS reflecting element at time slot n, which is limited within a discrete value range, specifically:
[0070]
[0071] wherein, is the discrete value range of the phase shift, F denotes the number of quantization bits, and the phase shift of each RIS reflecting element can be adjusted to any one of the possible values, and the achievable rate of each user m at time slot n is expressed as:
[0072]
[0073] wherein, w n,m denotes the active beamforming of the base station to user m at time slot n, w n,i denotes the active beamforming of the base station to user i at time slot n, denotes the noise variance of the received signal of user m, and the noise follows a circularly symmetric complex Gaussian distribution with zero mean and variance ;
[0074] (1c) Constructing a UAV energy consumption model, for a rotor UAV, the energy consumption model E(V) in three-dimensional space can be approximately expressed in the following form:
[0075]
[0076] wherein, P(V(t)) is the instantaneous energy consumption of the UAV in the two-dimensional horizontal space at time t, V(t) represents the instantaneous UAV speed at time t, T is the total flight time, V(T) is the instantaneous UAV speed at time T, V(0) is the instantaneous UAV speed at time 0, M UAV is the mass of the UAV, g is the gravitational acceleration, and H(T)-H(0) represents the change in height of the UAV from the starting coordinates to the final coordinates.
[0077] In the two-dimensional horizontal space, the energy consumption model P(V(t)) of the UAV is given as:
[0078]
[0079] where P B is the basic power consumption, U tip is the rotor tip speed, P I is the induced power, and K0 is a constant related to the aerodynamic performance of the UAV, d1 is a constant related to aerodynamics, p, s, and A represent the air density, the wing span of the UAV, and the cross-sectional area of the UAV parallel to the flight direction, respectively.
[0080] Considering that the UAV maintains an average speed in each time slot, the speed at time slot n is v n , then the energy consumption E N of the rotor UAV in N time slots can be represented as:
[0081]
[0082] where P(v n ) is the energy consumption of the UAV in time slot n in the two-dimensional horizontal space, v N is the speed of the UAV at time slot N, and h ter -h0 represents the change in height of the UAV from the starting coordinate C 0 to the final coordinate C Ter .
[0083] Step 2: Construct a multi-objective optimization problem aimed at optimizing the three-dimensional trajectory C of the UAV from the starting coordinate C 0 to the final coordinate C Ter , while optimizing the active beamforming matrix W of the base station and the passive beamforming matrix Θ of the RIS, with objectives including: maximizing the achievable average total rate minimizing the total number of time slots N of the flight, and reducing the energy consumption E N of the UAV, i.e.:
[0084]
[0085] where W = [W1, W2,..., W N ] represents all active beamforming of the base station in N time slots, W n represents the active beamforming of the base station at time slot n, Θ = [θ1, θ2,..., θ N ] is the passive beamforming of the RIS in all N time slots, and θ n represents the passive beamforming of the RIS at time slot nPassive beamforming of RIS, Λ represents the multi-objective of optimization, C1-C6 restricts the movement of the UAV, and C1 constrains the speed v of the UAV in each time slot n n At 0 and maximum flight speed v max C2 and C3 constrain the azimuth angle ξ of the UAV flight. n and elevation angle φ n Between -π and π, C4 constrains the UAV’s flight height h in each time slot. n At the minimum height h min and the maximum height h max C5 and C6 constrain the flight range of the drone, that is, the horizontal and vertical coordinates are between 0 and D, D is the length of the rectangular flight area, C7 constrains the drone to fly from the starting coordinate to the final coordinate, and the take-off point C0 is the set position C 0 , the position C of the flying drone after N time slots N is the set end point C Ter , C8 specifies the discrete phase shift range of all RIS reflector units, and C9 limits the active beamforming of the base station to the maximum transmit power P B Within the range, w n,m represents the active beamforming of the base station to user m in time slot n.
[0086] Step 3: Train a deep neural network to jointly optimize the base station’s active beamforming W and the RIS’s passive beamforming Θ, specifically:
[0087] (3a) Order represents time slot n, passive beamforming of RIS, and channel gain It can be expressed as:
[0088]
[0089] in, represents the cascade channel, For time slot n, the passive beamforming of RIS is θ n The conjugate transpose of ; At each time slot n, given the current position of the UAV, the problem of jointly optimizing active and passive beamforming can be formulated as:
[0090]
[0091] in, represents the active beamforming of the base station in time slot n, subject to the maximum transmit power P B Restrictions and constraints, is the discrete phase shift of the k-th RIS reflection unit in time slot n, which is limited to a discrete value range, and F represents the number of quantization bits;
[0092] (3b) The deep neural network is designed to consist of two subnetworks, referred to as the beamforming network and the phase-shifting network, respectively. The phase-shifting network consists of two convolutional layers and one fully connected layer. The convolutional layers extract features from the input channel information H d,m represents the direct link information from the base station to user m, H n,m represents the cascaded link information from the base station to user m via the RIS at time slot n, T denotes the transpose of a vector, and employs the ReLU activation function and batch normalization. After the convolutional layers output is flattened, the fully connected layer performs batch normalization and ReLU on the result, which is then mapped to dimension 2K, representing the real and imaginary parts of the RIS phase shift, with denoting the output of the phase-shifting network at time slot n, then the passive beamforming of the RIS at time slot n is given by which can be calculated by
[0093]
[0094] where, is the complex reflection coefficient of the k-th RIS reflecting element under continuous phase shift at time slot n, and For reflecting element k, denotes the amplitude of the k-th reflecting element at time slot n, denotes the continuous phase shift of the k-th RIS reflecting element at time slot n, is the k-th element of the output of the phase-shifting network at time slot n, is the K+k-th element of the output of the phase-shifting network at time slot n; the discrete phase shift of the k-th RIS reflecting element at time slot n is given by is obtained by quantizing the selection of the closest value in the set , the complex reflection coefficient of the k-th RIS reflecting element under discrete phase shift at time slot n is given by
[0095] (3c) The beamforming network consists of three fully connected layers. This network takes the cascaded channel information adjusted by the RIS phase shift as input. The first fully connected layer maps the input to a 100-dimensional space, followed by batch normalization and ReLU activation. The second fully connected layer maintains the same dimension and also applies batch normalization and ReLU activation. The third fully connected layer maps the output to dimension 2MN B , representing the real and imaginary parts of the beamforming matrix generated by the base station for each user, with denoting the output of the beamforming network, then the active beamforming of the base station to user m at time slot n is given by where each element It can be expressed as:
[0096]
[0097] Among them, m∈{1,2,...,M} represents the user index, n B ∈{1,2,...,N B} represents the base station antenna index, is the (m-1)Nth output of the beamforming network in time slot n B +n B element, is the nth output of the beamforming network in time slot n B +MN B +(m-1)N B elements, the final beamforming matrix will be normalized to meet the transmit power constraint;
[0098] (3d) The proposed deep neural network model is trained offline in an unsupervised manner. The training samples are generated based on the actual building blocking conditions. The positions of the user equipment and the drone are randomly changed to generate diverse channel information for training. The loss function is defined as for:
[0099]
[0100] in, is the achievable rate of user m;
[0101] With the goal of minimizing the loss function, the parameters of the neural network are updated using reverse gradient.
[0102] Step 4: Define the state function s in time slot n n , action a n , reward function r n , build a deep reinforcement learning model, specifically:
[0103] (4a) Considering the UAV as an intelligent agent, a Markov decision process is constructed for the UAV trajectory optimization problem, and the state function s is defined in time slot n. n Contains drone location information Info self , Final coordinate information Info tar and user informationInfo ue , expressed as:
[0104] s n ={Info self ,Info tar ,Info ue} (16)
[0105] Among them, Infoself Position information C of the UAV n azimuth angle ξ of the UAV n elevation angle φ n and flight speed v n is denoted as:
[0106] Info self = {C n , ξ n , φ n , v n} (17)
[0107] Info tar Final coordinates C of the UAV Ter azimuth angle from the UAV to the final coordinates elevation angle and distance to the final coordinates
[0108]
[0109] Info ue contains the transmission rates of all user equipments, denoted as:
[0110]
[0111] where, is the achievable rate of user m at time slot n, m = 1, 2,... M;
[0112] (4b) defines the action a of the UAV at time slot n n contains flight speed v n azimuth angle ξ of the UAV n and elevation angle φ n is denoted as:
[0113] a n = {v n , ξ n , φ n} (20)
[0114] (4c) defines the reward function r at time slot n n is:
[0115]
[0116] where a, b and c are coefficients balancing the importance of the UAV reaching the final coordinates, improving the system rate and reducing energy consumption, is the reward measured for the UAV reaching the final coordinates, defined as:
[0117]
[0118] wherein, and denote the distance of the UAV to the target point in the previous and the current time slot, respectively, v max is the maximum speed of the UAV, δ n is the time slot length, and β penalty is a negative value, representing the penalty of each time slot in the flight process. The introduction of this penalty term aims to encourage the UAV to reach the target in as few time slots as possible, thereby optimizing flight efficiency; is the system transmission rate reward in time slot n, defined as:
[0119]
[0120] wherein, is the achievable rate of user m in time slot n
[0121] is the energy consumption of the UAV in time slot n, taking a negative value, is the movement reward of the UAV, encouraging the UAV to avoid buildings and remain within the specified boundary and height limits. If the UAV collides or exceeds the range, is negative; otherwise is zero.
[0122] Step 5: The deep reinforcement learning model is optimized using the Softmax Deep Double Deterministic Policy Gradients (SD3) algorithm, which contains two policy networks π1(s) and π2(s) with parameters and and two Q networks Q1(s, a) and Q2(s, a) with parameters φ1 and φ2, respectively. In each time slot n, the UAV observes a state s n . Using this state information, each policy network generates an action a n , expressed as:
[0123]
[0124] wherein, denotes a Gaussian noise, used to help the agent explore a wider action space. Subsequently, the UAV receives a reward r n , and the state transitions to s n+1 . The corresponding experience (s n , a n , s n+1 , r n , done n) stored in the experience replay buffer to perform network training, done n To complete the flag, if the UAV flies to the target point, then done n = True, otherwise done n = False.
[0125] By sampling a mini-batch B of experience set {(s, a, s', r, done)} from , the deep reinforcement learning model is trained, where s, a, s', r, done represent the batch set of s n , a n , s n+1 , r n , done n , the Q network is updated by gradient descent to minimize the loss between the target Q value and the predicted Q value, the target Q value is expressed as:
[0126] y j = r + (1 - done) γsoftmax β {Q(s', ·)}, j = 1, 2 (25)
[0127] where y j is the target Q value of the policy network j, r is the immediate reward, γ is the discount factor, which is used to weigh the importance of the immediate reward and the future reward, s' is the next state, the softmax function is applied, β is its parameter. The softmax function can be optimized smoothly and promote learning on experience, which is defined as:
[0128]
[0129] where, is the estimate of Q value, represents the expected cumulative reward of different actions in state s', a is the sampling probability of action , is the expectation operator, which represents the expected calculation of action sampled from Gaussian distribution p , and Q' i is the target Q network, whose parameters are φ' i , i = 1, 2. In addition is obtained by adding noise sampled from Gaussian distribution to the action of the target policy network, is the variance of the Gaussian distribution, the introduced noise is constrained to ensure that the target action is close to the original action, which is expressed as:
[0130]
[0131] Clip is a clipping operation used to limit the noise to the range of (-c, c), where c is a preset boundary value, and π j ′ is the target policy network, and its parameters are Therefore, the loss function of the Q network is expressed as:
[0132]
[0133] Where B is the batch size, the parameters of the policy network The update is performed using the policy gradient method, which is defined as follows:
[0134]
[0135] in, For the policy network parameters The gradient, For Q j (s,a;φ j ) is the jth Q network Q j (s,a;φ j ) with respect to the gradient of action a, Indicates that when calculating the gradient, action a takes the action π output by the policy network j (s).
[0136] The target policy network and target Q network are updated using soft updates. Every d time steps, the parameters of the target policy network and target Q network are gradually mixed with the parameters of the current policy network and Q network. The update formula is:
[0137] φ′ j =τφ j +(1-τ)φ′ j ,j=1,2 (31)
[0138] Where τ is the update rate. is the parameter of the target policy network j, φ′ j are the parameters of the target Q network j.
[0139] Specifically:
[0140] (5a) Start the environment simulator and initialize two policy networks π1(s) and 2(s), whose parameters are and and two Q networks Q1(s,a) and Q2(s,a), whose parameters are φ1 and φ2 respectively;
[0141] (5b) Initialize two target policy networks π1'(s) and π'2(s), whose parameters are and and two target Q networks Q′1(s,a) and Q′2(s,a), whose parameters are φ′1←φ1 and φ′2←φ2 respectively;
[0142] (5c) Initialize the experience replay buffer
[0143] (5d) Initialize the maximum number of rounds to N epoch , set epoch = 1, the maximum number of time slots in each round is N max ;
[0144] (5e) Initialize the current time slot n=1;
[0145] (5f) Initialize the position C of the drone and the position C of the base station BS and the locations C of M users m ,m=1,2,...,M;
[0146] (5g) UAV according to state s n And select action a based on two policy networks π1(s) and π2(s) n ;
[0147] (5h) The drone performs action a n And joint beamforming is output based on the trained deep neural network;
[0148] (5i) Drone Rewards n , next time slot state s n+1 And the completion flag done, if the drone flies to the target point, done = True, otherwise the opposite;
[0149] (5j) will (s n ,a n ,s n+1 ,r n ,done) is stored in middle;
[0150] (5k) from Sample a small batch B of experience set {(s,a,s′,r,done)};
[0151] (5l) Calculate the loss value Loss based on the two Q networks, adopt a small batch gradient descent strategy, and update the parameters of the two Q networks Q1(s,a) and Q2(s,a) through the back propagation of the neural network;
[0152] (5m) Use gradient descent to update the two policy networks π1(s) and π2(s);
[0153] (5n) The number of training times reaches the target policy network and target Q network update interval, and the target policy network parameters and target Q network parameters are updated according to the policy network parameters and Q network parameters;
[0154] (5o) If done = True, then epoch = epoch + 1, go to step (5q);
[0155] (5p) Determine whether n is satisfied <N max If n=n+1, go to step (5g), otherwise go to step (5q);
[0156] (5q) Determine whether epoch is satisfied <N epoch If so, epoch=epoch+1, go to step (5e), otherwise, the optimization ends and the optimized deep reinforcement learning model is obtained.
[0157] Step 6: Based on the optimized deep reinforcement learning model and the trained deep neural network, the optimal joint beamforming and drone trajectory are obtained. Specifically:
[0158] (6a) The drone inputs the channel state information according to the trained deep neural network model based on the current position of each time slot n and H d,m represents the direct link information from the base station to user m, H n,m represents the cascade link information from the base station to user m via RIS in time slot n, T represents the transpose of the vector, and the output is the active beamforming W of the base station in time slot n. n and RIS passive beamforming θ n ;
[0159] (6b) The drone uses the trained deep reinforcement learning model to input state information s n , and output the optimal action strategy, that is, the action of the drone in the current time slot includes the flight speed v n 、The azimuth angle of the UAVξ n and elevation angle φ n , expressed as:
[0160] a n ={v n ,ξ n ,φ n}(32)
[0161] exist Figure 1 In this paper, a flowchart of the joint beamforming and UAV trajectory optimization method for the aerial RIS-assisted MU-MISO system is described. First, a deep neural network for joint beamforming is trained, and the UAV selects the flight direction and speed based on the deep reinforcement learning model trained with SD3.
[0162] exist Figure 2 In [1], the system model of aerial RIS assisted MU-MISO is described. It can be seen that the UAV carries RIS as an aerial RIS-assisted base station for communication with users.
[0163] exist Figure 3 In the paper, a simulation comparison diagram of a joint beamforming and UAV trajectory optimization method for an aerial RIS-assisted MU-MISO system and other algorithms is described. It can be seen that the algorithm proposed in this invention has an advantage in maximizing the achievable average total rate.
[0164] Based on the description of the present invention, it should not be difficult for those skilled in the art to see that the present invention proposes a joint beamforming and drone trajectory optimization method for an aerial RIS-assisted MU-MISO system, which can achieve multi-objective optimization and achieve a good balance between complexity and performance.
[0165] The present invention also provides a system corresponding to the method, comprising:
[0166] System model building unit, used to build the aerial RIS-assisted MU-MISO system model, including: downlink transmission model, UAV motion model, communication model and UAV energy consumption model;
[0167] The goal construction unit is used to construct a multi-objective optimization problem, whose goals include maximizing the achievable average total rate, minimizing the total number of flight time slots, and reducing the energy consumption of the UAV;
[0168] A deep neural network design and training unit is used to train a deep neural network that jointly optimizes the base station's active beamforming and the RIS's passive beamforming. This unit includes a beamforming network and a phase shift network. The phase shift network takes channel information as input, and the RIS's passive beamforming is derived from its output. The beamforming network takes cascaded channel information after RIS phase shift adjustment as input, and the base station's active beamforming is derived from its output.
[0169] A deep reinforcement learning model construction unit is used to construct and optimize a deep reinforcement learning model including a state function, an action space, and a reward function. The state function includes the drone's position information, final coordinate information, and user information. The action space includes the drone's flight speed, azimuth, and elevation angle. The reward function is obtained based on the trained deep neural network.
[0170] The joint beamforming and UAV trajectory optimization unit is used to obtain the optimal joint beamforming and UAV trajectory based on the optimized deep reinforcement learning model and the trained deep neural network.
[0171] An electronic device for storing and executing the method, comprising:
[0172] a memory storing executable program code;
[0173] a processor coupled to the memory;
[0174] The processor calls the executable program code stored in the memory to execute the steps of the joint beamforming and UAV trajectory optimization method for the aerial RIS-assisted MU-MISO system.
[0175] A computer-readable storage medium for storing and executing the method, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are called, the steps of the joint beamforming and drone trajectory optimization method for the aerial RIS-assisted MU-MISO system are executed.
[0176] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.
[0177] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "N" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0178] Any process or method description described in a flowchart or otherwise herein may be understood to represent a module, segment, or portion of code comprising one or N executable instructions for implementing the steps of a custom logic function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which the embodiments of the present invention pertain. Any content not described in detail in this application belongs to the prior art known to those skilled in the art.
Claims
1. A joint beamforming and UAV trajectory optimization method for an aerial RIS-assisted MU-MISO system, characterized by: The following steps are involved: Construct an aerial RIS-assisted MU-MISO system model, including: downlink transmission model, UAV motion model, communication model, and UAV energy consumption model; For the aerial RIS-assisted MU-MISO system model, a multi-objective optimization problem is constructed, whose goals include maximizing the achievable average total rate, minimizing the total number of flight time slots, and reducing the energy consumption of the UAV; The multi-objective optimization problem is broken down into a joint optimization sub-problem of the base station's active beamforming and the RIS's passive beamforming. The optimization variables are the base station's active beamforming and the RIS's passive beamforming, and the optimization objective is to maximize the achievable total rate. This problem is solved by training a deep neural network that jointly optimizes the base station's active beamforming and the RIS's passive beamforming. The deep neural network includes a beamforming network and a phase shift network. The input of the phase shift network is the channel information of the airborne RIS-assisted MU-MISO system model. The output is the cascaded channel information of the RIS's passive beamforming adjusted by the RIS phase shift, which serves as the input of the beamforming network. The output of the beamforming network is the base station's active beamforming. The loss function used in the entire deep neural network training is the negative value of the sub-problem optimization objective, that is, minimizing the loss function is equivalent to maximizing the achievable total rate. Based on a trained deep neural network, the multi-objective optimization problem is transformed into a drone trajectory optimization problem. A deep reinforcement learning model is constructed and optimized, including a state function, an action space, and a reward function. The state function includes the drone's position information, final coordinate information, and user information, and the action space includes the drone's flight speed, azimuth, and elevation angle. The reward function is obtained based on the trained deep neural network, and the optimized deep reinforcement learning model is used to solve the drone trajectory optimization problem. Based on the optimized deep reinforcement learning model and the trained deep neural network, the optimal joint beamforming and drone trajectory are obtained.
2. The joint beamforming and UAV trajectory optimization method for an aerial RIS-assisted MU-MISO system according to claim 1 is characterized in that: Multi-objective optimization problems include: Where Λ represents the multi-objective of optimization, N is the number of time slots, M is the number of users, is the achievable rate of user m in time slot n, E N is the energy consumption of the UAV in N time slots, W=[W1,W2,...,W N ] represents all active beamforming of the base station in N time slots, W n represents the active beamforming of the base station in time slot n, Θ = [θ1,θ2,...,θ N ] is all the passive beamforming of RIS in N time slots, θ n represents the passive beamforming of RIS in time slot n, and C is the starting coordinate of the UAV from C 0 To the final coordinate C Ter The three-dimensional trajectory, v n is the speed of the UAV in time slot n, v max is the maximum flight speed of the UAV, ξ n is the azimuth of the UAV flying in time slot n, φ n is the elevation angle of the UAV at time slot n, h n is the Z-direction height coordinate of the UAV at time slot n, h min and h max are the minimum and maximum values of the UAV’s flight altitude, respectively, and x n and y n The X-direction coordinate and Y-direction coordinate of the drone in time slot n, respectively. D is the length of the rectangular flight area, C0 is the take-off point, and C 0 is the starting coordinate of the UAV, C N is the location of the UAV after N time slots, C Ter is the final coordinate of the drone, is the discrete phase shift of the kth RIS reflection unit in time slot n, F represents the number of quantization bits, and w n,m represents the active beamforming of the base station to user m in time slot n, P B is the maximum transmit power.
3. The joint beamforming and UAV trajectory optimization method for an aerial RIS-assisted MU-MISO system according to claim 1 is characterized in that: At each time slot n, given the current position of the UAV, the problem of jointly optimizing active and passive beamforming is formulated as: Among them, W n represents the active beamforming of the base station in time slot n, θ n represents the passive beamforming of RIS in time slot n, K is the number of reflection units of RIS, is the achievable rate of user m in time slot n, is the discrete phase shift of the kth RIS reflection unit in time slot n, F represents the number of quantization bits, and w n,m represents the active beamforming of the base station to user m in time slot n, P B is the maximum transmission power, M is the number of users; use represents the output of the phase shift network at time slot n. Then, at time slot n, the complex reflection coefficient of the kth RIS reflection unit under continuous phase shift is expressed as: in, is the complex reflection coefficient of the kth RIS reflection unit in time slot n under continuous phase shift, and represents the amplitude of the kth RIS reflection unit in time slot n, represents the continuous phase shift of the kth RIS reflection unit in time slot n, is the output of the phase shift network in time slot n The kth element of is the output of the phase shift network in time slot n The K+kth element of ; the passive beamforming of RIS in time slot n is expressed as represents the complex reflection coefficient of the kth RIS reflection unit at time slot n under discrete phase shifts, and represents the discrete phase shift of the kth RIS reflection unit in time slot n, which is selected by quantization Middle distance Get the most recent value, k = 1, 2, ..., K: use represents the output of the beamforming network, then the active beamforming of the base station to user m in time slot n is Each element Expressed as: Among them, m∈{1,2,...,M} represents the user index, n B ∈{1,2,...,N B } represents the base station antenna index; is the (m-1)Nth output of the beamforming network in time slot n B +n B element, is the nth output of the beamforming network in time slot n B +MN B +(m-1)N B Element, N B is the number of antennas included in the base station, and finally the active beamforming W of the base station is obtained.
4. The joint beamforming and UAV trajectory optimization method for an aerial RIS-assisted MU-MISO system according to claim 1 is characterized in that: The deep neural network model is trained offline in an unsupervised manner, and the loss function for: in, is the achievable rate of user m, M is the number of users; with the goal of minimizing the loss function, the reverse gradient is used to update the parameters of the neural network.
5. The joint beamforming and UAV trajectory optimization method for an aerial RIS-assisted MU-MISO system according to claim 1 is characterized in that: The drone's position information includes the drone's location information, azimuth, elevation, and flight speed. The final coordinate information includes the drone's final coordinates, the azimuth, elevation, and distance from the drone to the final coordinates. User information includes the transmission rate of all user devices. The reward function is a weighted combination of the reward obtained by the drone when it reaches the final coordinates, the system transmission rate reward, the drone's energy consumption, and the drone's motion reward. The system transmission rate reward is calculated based on the user's achievable rate, and the achievable rate is calculated by loading a trained deep neural network.
6. The joint beamforming and UAV trajectory optimization method for an aerial RIS-assisted MU-MISO system according to claim 5 is characterized in that: In time slot n, the reward function r n for: Among them, a, b, and c are coefficients that balance the importance of the UAV reaching the final coordinate, improving the system rate, and reducing energy consumption. To measure the reward obtained by the drone when it reaches the final coordinate, it is defined as: in, and Respectively represent the distances between the UAV and the target point in the two time slots before and after, v max is the maximum speed of the UAV, δ n is the time slot duration, β penalty represents the penalty for each time slot during the flight; is the system transmission rate bonus in time slot n, defined as: in, represents the achievable rate of user m in time slot n. Its calculation requires loading the trained deep neural network and transforming the channel information into As the input of the phase shift network, the passive beamforming of RIS in time slot n is obtained. H d,m represents the direct link information from the base station to user m, H n,m Indicates the cascade link information from the base station to user m via RIS in time slot n, T represents the transpose of the matrix, and H represents the conjugate transpose of the matrix; the cascade channel information after RIS phase shift adjustment As the input of the beamforming network, represents the channel gain of the direct link from the base station to user m, represents the cascade link channel gain of the base station-RIS-user m, M is the number of users, and the active beamforming of the base station to user m in time slot n is obtained. N B is the number of antennas included in the base station, The calculation formula is: Among them, w n,i represents the active beamforming of the base station to user i in time slot n, represents the noise variance of the signal received by user m; is the energy consumption of the UAV in time slot n, Bonus for drone movement.
7. The joint beamforming and UAV trajectory optimization method for an aerial RIS-assisted MU-MISO system according to claim 1 is characterized in that: The optimal joint beamforming and drone trajectory acquisition method is: The UAV inputs the channel state information in each time slot n according to the current position and the trained deep neural network model. and At time slot n, the base station’s active beamforming W is output n and RIS passive beamforming θ n ;H d,m represents the direct link information from the base station to user m, H n,m represents the cascade link information from the base station to user m via RIS in time slot n, and T represents the transpose of the matrix; represents the channel gain of the direct link from the base station to user m, represents the cascade link channel gain of the base station-RIS-user m, where M is the number of users; The drone uses the trained deep reinforcement learning model and inputs the state information s n Output the optimal action strategy, that is, the action of the drone in the current time slot includes the flight speed v n , flight azimuth ξ n and elevation angle φ n , expressed as: a n ={v n ,x n ,f n }。 8. Joint beamforming and UAV trajectory optimization system for aerial RIS-assisted MU-MISO system, characterized by: include: System model building unit, used to build the aerial RIS-assisted MU-MISO system model, including: downlink transmission model, UAV motion model, communication model and UAV energy consumption model; The goal construction unit is used to construct a multi-objective optimization problem for the aerial RIS-assisted MU-MISO system model, whose goals include maximizing the achievable average total rate, minimizing the total number of flight time slots, and reducing the energy consumption of the UAV; The deep neural network design and training unit is used to decompose and construct a multi-objective optimization problem, which is a joint optimization sub-problem of the base station's active beamforming and the RIS's passive beamforming. The optimization variables are the base station's active beamforming and the RIS's passive beamforming, and the optimization goal is to maximize the achievable total rate. The problem is solved by training a deep neural network that jointly optimizes the base station's active beamforming and the RIS's passive beamforming. The deep neural network includes a beamforming network and a phase shift network. The input of the phase shift network is the channel information of the airborne RIS-assisted MU-MISO system model. The output is the cascaded channel information of the RIS's passive beamforming after RIS phase shift adjustment, which serves as the input of the beamforming network. The output of the beamforming network is the base station's active beamforming. The loss function used in the entire deep neural network training is the negative value of the sub-problem optimization objective, that is, minimizing the loss function is equivalent to maximizing the achievable total rate. The deep reinforcement learning model construction unit is used to transform the constructed multi-objective optimization problem into a drone trajectory optimization problem based on the trained deep neural network. The deep reinforcement learning model including the state function, action space and reward function is constructed and optimized. The state function includes the drone's position information, final coordinate information and user information, and the action space includes the drone's flight speed, azimuth and elevation angle. The reward function is obtained based on the trained deep neural network, and the optimized deep reinforcement learning model is used to solve the drone trajectory optimization problem. The joint beamforming and UAV trajectory optimization unit is used to obtain the optimal joint beamforming and UAV trajectory based on the optimized deep reinforcement learning model and the trained deep neural network.
9. An electronic device, characterized in that: The device comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the steps of the joint beamforming and drone trajectory optimization method for the aerial RIS-assisted MU-MISO system as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which, when called, are used to execute the steps of the joint beamforming and drone trajectory optimization method for the aerial RIS-assisted MU-MISO system as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Capacity optimization method and device for RIS auxiliary communication system carried by unmanned aerial vehicle
CN114422363A
Unmanned aerial vehicle emergency communication method, system and equipment assisted by intelligent reflecting surface, and medium
CN119052830A