Relay communication network transmission method based on assistance of air intelligent reflecting surface
By optimizing the parameters of the intelligent reflector and base station, combined with simulated annealing and Markov decision process, the problems of ground user mobility and UAV flexibility in the aerial intelligent reflector relay communication network are solved, the system throughput is improved, and it is suitable for high-density mobile user communications in complex environments.
Patent Information
- Application Number
- CN202510789992.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, the relay communication network assisted by aerial intelligent reflectors is insufficient in considering the mobility of ground users and the flexibility of drones, and fails to effectively utilize the actual phase shift model, resulting in the optimal beamforming design no longer being optimal in practical applications.
By optimizing the reflection coefficient matrix of the intelligent reflective surface and the linear receiving beam vector of the base station, combined with the simulated annealing algorithm and Markov decision process, the flight trajectory of the UAV is optimized, the system throughput optimization problem is constructed, and the system throughput is maximized.
It improves the communication link quality between ground mobile users and base stations, enhances transmission performance, is suitable for high-capacity and high-density mobile user scenarios in complex environments, and maximizes system throughput.
Smart Images

Figure CN120639148A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of unmanned aerial vehicle (UAV) communication assisted by intelligent reflective surfaces, and relates to a relay communication network transmission method based on the assistance of aerial intelligent reflective surfaces. Background Art
[0002] As fifth-generation mobile communication technology evolves to sixth-generation, user demand for high speeds, low latency, and wide-area coverage continues to grow. Drones, with their high maneuverability and flexible deployment, have become a powerful tool for expanding network coverage and improving service quality. They act as aerial relay nodes, enabling data forwarding between ground users and base stations. Furthermore, smart reflective surfaces, as a new wireless communication technology, offer new avenues for improving spectrum efficiency and system capacity by efficiently manipulating the phase and amplitude of signal reflections. Mounting smart reflective surfaces on drones fully leverages the drone's three-dimensional mobility and its intelligent reflection capabilities, enabling the construction of an efficient aerial relay system and a seamless, integrated air-ground-air network. This system provides dynamically optimized communication support for ground mobile users and lays the foundation for a fully integrated 6G heterogeneous network.
[0003] Despite the numerous advantages of relay communication networks assisted by aerial smart reflectors, they still face difficulties in practical application. Current research on relay communication networks assisted by aerial smart reflectors primarily focuses on two-dimensional, static scenarios. These studies fail to fully consider the mobility of ground users and limit the flexibility of drones in complex environments. Furthermore, most employ an ideal phase-shift model, where the phase is adjustable within the range [0, 2π) and the reflection amplitude is constant. In practice, the reflection amplitude of smart reflector elements varies with phase, and phase adjustment is limited by hardware. This results in the fact that optimal beamforming designs under ideal phase-shift models are often no longer optimal under actual phase-shift models. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a relay communication network transmission method based on the assistance of an aerial intelligent reflecting surface, which maximizes the system throughput by optimizing the reflection coefficient matrix of the intelligent reflecting surface, the linear receiving beam vector of the base station and the flight trajectory of the drone, taking into account user mobility, regional constraints and flight restrictions of drones in a dynamic environment.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A method for relay communication network transmission based on the assistance of an aerial intelligent reflector comprises the following steps:
[0007] S1: Build a UAV-mounted intelligent reflective surface-assisted wireless relay communication system between the base station and ground mobile users;
[0008] S2: Establish a wireless channel model for the relay communication network and an actual phase shift model for the smart reflector;
[0009] S3: By adjusting the reflection coefficient matrix of the smart reflector and the linear receiving beam vector of the base station, the smart reflector phase shift optimization problem is constructed and simplified with the goal of maximizing the overall gain of multi-user synchronous transmission. A global perturbation strategy based on the simulated annealing algorithm is used to solve the problem and obtain the optimal phase shift of the smart reflector.
[0010] S4: After adjusting the reflection coefficient matrix of the smart reflector and the linear receiving beam vector of the base station, considering the user mobility in a dynamic environment and the flight constraints of the drone, a system throughput optimization problem is constructed. It is modeled as a Markov decision process and the TD3 algorithm is used to optimize the three-dimensional flight trajectory of the drone to maximize the system throughput.
[0011] Furthermore, in step S1, a UAV-equipped smart reflective surface-assisted wireless relay communication system is constructed, specifically including: a base station with N antennas, K ground mobile users, and a UAV equipped with a smart reflective surface; it is assumed that the UAV and the user are both equipped with a single omnidirectional antenna; the single-antenna users are randomly distributed in a given area and keep moving, k = {1, 2, ..., K} is the set of ground users; the smart reflective surface is composed of M reflective elements, denoted as the set m = {1, 2, ..., M}.
[0012] Furthermore, in step S2, a wireless channel model of the relay communication network is established, specifically including: the UAV carries an intelligent reflective surface, also known as an aerial intelligent reflective surface, denoted as ARIS; the system time is divided into discrete time slots, t = 0, 1, 2, ..., T; at time t = 0, ARIS starts from the starting point q s Departure, fly over the city to forward the data of ground mobile users to the base station, and arrive at the end point q e End; set the base station location information to The location information of the UAV and the kth user at time t are expressed as and Where [x(t), y(t)] represents the horizontal projection coordinates of the drone at time t, z(t) represents the height of the drone, [x k (t),y k (t)] represents the horizontal coordinate of user k at time t; spherical coordinates V = (v, λ, ρ) are defined to represent the flight action of the drone, where v = ||V||, v∈[0,v max ], v represents the flight speed of the drone, v max represents the maximum flight speed of the UAV, λ∈[0,π] represents the polar angle, and ρ∈[-π,π] represents the azimuth angle on the horizontal plane.
[0013] Furthermore, in step S2, the actual phase shift model of the intelligent reflective surface is established, specifically including: setting is the reflection coefficient of reflection unit m, then the reflection coefficient matrix can be described as θ m (t) and β m (t) represent the phase shift and reflection amplitude of the mth reflection unit in the smart reflection surface, respectively, where θ m (t)∈[0,2π) and β m (t)∈[0,1],m∈{1,...,M}; θ m (t) is selected from a finite number of discrete values from 0 to 2π; let the binary bit b represent the adjustable flexibility of the discrete phase shift, then the actual phase shift range is The reflection amplitude under the actual phase shift model can be calculated by the following formula:
[0014]
[0015] Among them, β min ≥0, φ≥0, and α≥0 are constants related to the specific circuit implementation; β min is the minimum amplitude, φ is -π / 2 and β min β min The horizontal distance between them, α controls the steepness of the function curve; it is worth noting that when β min When =1 or α=0, the above formula is equivalent to an ideal phase shift model with unit amplitude;
[0016] The channel gains of user-base station, user-ARIS and ARIS-base station at time t are: Therefore, the signal received at the base station can be expressed as:
[0017]
[0018] in, g k (t) is the linear receiving beam vector of the base station, s is the unit power complex baseband information symbol of the user, and represents the additive white Gaussian noise of the base station;
[0019] The signal-to-noise ratio γ received by the base station from user k k (t) is expressed as:
[0020]
[0021] Among them, P k (t) is the transmission power of user k, σ 2 is the noise power.
[0022] Furthermore, in step S3, by adjusting the reflection coefficient matrix μ of the smart reflective surface m (t) and the linear receiving beam vector g of the base station k (t), with the goal of maximizing the overall gain of multi-user synchronous transmission, the smart reflector phase shift optimization problem is constructed as:
[0023]
[0024] Multiple terrestrial users transmit synchronously based on frequency division multiple access technology, so the base station can generate a linear receiving beam vector for each transmitting user and optimize the linear receiving beam vector of the base station based on the maximum ratio transmission formula. where g k (t) * is the optimized linear receiving beam vector of the base station; therefore, problem P1 is transformed into the following:
[0025]
[0026] Due to the non-convex constraints of (4a), solving problem P2 is challenging. To address these difficulties, a penalty-based approach is used to simplify P2, which penalizes constraint violations by adding constraint-related penalty terms to the objective function of the optimization problem to eliminate intractable equality constraints; to this end, problem P2 can be formulated as
[0027]
[0028] stθ m (t)∈Ω,m=1,...,M(6a)
[0029] where ξ > 0 is a penalty parameter that imposes a penalty on violations of the constraints in (4a). In particular, as ξ → ∞, solving the above problem yields an approximate solution to problem P2. This results in a two-layer iterative algorithm, where the inner layer solves the penalized optimization problem P3, while the outer layer updates the penalty coefficient ξ until the solution converges.
[0030] Furthermore, in step S3, a global perturbation strategy based on a simulated annealing algorithm is used to solve and obtain the optimal phase shift of the smart reflective surface, specifically including:
[0031] a) For a given a, μ(t) in Problem P3 can be optimized by solving the following problem:
[0032]
[0033] in, It is not difficult to observe that the first and second terms in problem P4 are convex and concave respectively, so we can use the convex-concave process principle to approximate the solution in an iterative manner;
[0034] In each subsequent iteration, at a given point μ (i) (t) is used to perform a first-order Taylor expansion on the first term of the objective function in the above problem and ignore the constant term, thereby simplifying it to a linear function to form a convex approximate optimization problem, which is given by the following formula:
[0035]
[0036] It is not difficult to observe that P5 is an unconstrained convex optimization problem for which the closed-form optimal solution can be easily obtained (by setting the first-order derivative of the objective function with respect to μ(t) to zero);
[0037]
[0038] Among them, μ (i) (t) is the i-th order derivative of μ(t); Next, μ (i) (t) is updated to μ (i+1) (t), until the target value of P5 reaches convergence;
[0039] b) For a given μ(t), all a in Problem P4 can be optimized by solving the following problem:
[0040]
[0041] Since θ m (t) is completely separable in the objective function, so M independent sub-problems can be solved in parallel by expanding And ignoring the constant term, we can get:
[0042]
[0043] in, Since θ m (t) is chosen from a finite number of discrete values between 0 and 2π, so that the optimal The simulated annealing algorithm is used to perform global optimization based on the solution of formula (11) to escape from the local optimal solution.
[0044] Furthermore, in step S4, a system throughput optimization problem is constructed, specifically including: the transmission rate between the base station and user k is expressed as:
[0045] R k (t) = Blog2(1+γ k (t))(12)
[0046] Where B is the communication bandwidth of ARIS to users, γ k (t) is the channel ratio between the base station and user k, γ k (t) must be greater than or equal to the signal-to-noise ratio threshold γ th To accurately decode; the total transmission rate of the user can be expressed as:
[0047]
[0048] In order to maximize the system throughput in the ARIS-assisted wireless communication network within the specified time T, the following throughput maximization problem is constructed by optimizing the flight trajectory design of the UAV:
[0049]
[0050] γ(t)≥γ th (14a)
[0051] q u (0) = q s ,q u (T) = q e (14b)
[0052]
[0053] (μ m (t),g(t))=P1(14i)
[0054] Among them, q u (0) and q u (T) are the positions of the UAV at t = 0 and t = T, respectively, q s and q e are the starting point and end point of the drone, respectively, x min 、x max 、y min and y max is the constraint condition on the xy plane of the flight area, z min and z max is the restriction condition for the UAV’s flight altitude; Equation (14a) is the signal-to-noise ratio constraint; Equation (14b) ensures that the UAV must start from the starting point and reach the destination; Equations (14c) to (14e) restrict the UAV’s actions during flight; Equations (14f) to (14h) define the UAV’s feasible flight space.
[0055] Furthermore, in step S4, the system throughput optimization problem P7 is modeled as a Markov decision process, which specifically includes:
[0056] State design: The design of the state space includes the position, throughput, and remaining time and distance of the UAV, where the remaining time is tre =Tt, the remaining distance is the distance from the drone to the destination: The drone must arrive within the specified time. The agent knows its remaining time and the distance to the destination so that it can make a decision: whether to continue serving the user or arrive at the destination in time; therefore, the state space s(t) is defined as follows:
[0057]
[0058] Action design: The action of the agent corresponds to the flight control of the UAV, which is determined by the constraints (14c) to (14e). The flight action of the UAV includes the flight direction and flight speed. Therefore, the action a(t) = {v(t), λ(t), ρ(t)}, where v(t)∈[0,v max ] represents the flight speed of the UAV in the next time slot; λ(t)∈[0,π] represents the flight polar angle of the next time slot; ρ(t)∈[-π,π] represents the flight azimuth angle of the next time slot;
[0059] Reward function design: The design of the reward determines whether the agent can learn the desired strategy. In each time slot, throughput is generated between ARIS and the user. The optimization goal is to maximize the throughput of the system. Therefore, the throughput-related reward is set to a positive reward. The throughput part of the reward is defined as follows:
[0060] r1=c R R(t)(16)
[0061] Among them, c R is a positive constant used to reduce throughput, because too large a reward can lead to the “vanishing gradient” problem in the Tanh(·) activation function; Next is the reward for violating the flight area restriction, which is defined as:
[0062] r2=-c xy ·ξ xy -c h ·ξ h (17)
[0063] Among them, ξ xy is the horizontal plane constraint indicator, ξ h is a binary height constraint indicator; c xy and c h is a positive penalty constant associated with violating constraints (14f) to (14h). To ensure that the drone can reach the destination within the specified time, a reward related to the distance from the drone to the destination is set. The reward is defined as follows:
[0064] r3=c d w d(d max -d u,d (t))(18)
[0065] Among them, d max is the distance between the starting point and the end point, c d is a distance-dependent scaling factor; w d =e -ζ(T-t) is a time-related weight coefficient. By setting ζ reasonably, let w d As t increases nonlinearly, when the remaining time t re When it is insufficient, r3 becomes dominant, and the drone tends to fly to the destination to obtain a larger reward. The reward obtained by the agent when selecting action a(t) in a given state s(t) is defined as:
[0066] r(t)=r1+r2+r3(19)
[0067] Among them, r(t) is the reward function at time t.
[0068] The beneficial effects of the present invention are as follows: the present invention utilizes a drone equipped with an intelligent reflective surface to construct an additional signal propagation link through the aerial intelligent reflective surface, thereby enhancing the quality of the communication link between ground mobile users and base stations and improving transmission performance. By optimizing the reflection coefficient matrix of the intelligent reflective surface and the receiving beam of the base station, the channel gain of multi-user transmission is maximized; considering the user mobility and the flight constraints of the drone in a dynamic environment, a system throughput maximization problem is constructed. To solve this problem, it is modeled as a Markov decision process, and the TD3 algorithm is used to optimize the three-dimensional flight trajectory of the drone to obtain the maximum throughput of the system and improve the system transmission performance. The present invention is easy to implement, quick and efficient to deploy, and is suitable for complex wireless communication scenarios with high-capacity and high-density mobile users, and has high practical value.
[0069] In summary, the present invention can reconstruct the signal propagation environment by introducing drones equipped with intelligent reflective surfaces into the network to maximize the system throughput. At the same time, it takes into account the mobility of users in dynamic environments and adopts a practical phase shift model. It is easy to implement and has fast and efficient deployment. It is suitable for complex wireless communication scenarios with high-capacity and high-density mobile users, such as densely populated urban areas, emergency communication scenarios, and the Internet of Things, which are dynamic application environments with high requirements for communication capacity and low signal latency.
[0070] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0072] Figure 1 This is a model diagram of the UAV communication system assisted by the intelligent reflective surface in an embodiment of the present invention. DETAILED DESCRIPTION
[0073] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0074] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0075] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0076] See also Figure 1 This embodiment provides a relay communication network transmission method based on the assistance of an aerial intelligent reflector. Figure 1As shown in Figure 2, we consider a UAV-mounted smart reflector-assisted communication system consisting of a base station with N antennas, K ground mobile users, and a UAV equipped with a smart reflector. We assume that both the UAV and the users are equipped with a single omnidirectional antenna. The single-antenna users are randomly distributed and constantly mobile within a given area, with k = {1, 2, ..., K} being the set of ground users. The smart reflector consists of M reflective elements, denoted by the set m = {1, 2, ..., M}. In this network scenario, since urban buildings may block the direct link between the base station and users, UAVs equipped with smart reflectors are used to act as smart mobile relays to enhance communication.
[0077] The specific steps are as follows:
[0078] 1. In wireless communication networks between base stations and ground-based mobile users, various obstacles often exist, resulting in reduced communication transmission performance. To enhance the quality of the communication link between base stations and ground-based mobile users, drones equipped with intelligent reflective surfaces are introduced as aerial mobile relays, establishing additional signal transmission links and improving transmission performance.
[0079] 2. Establish a wireless channel model for the relay communication network and an actual phase shift model for the smart reflector. Specifically, this includes:
[0080] Establish a wireless channel model for the relay communication network: The UAV equipped with intelligent reflective surface is also called the aerial intelligent reflective surface, denoted as ARIS. The system time is divided into discrete time slots, t = 0, 1, 2, ..., T. At t = 0, ARIS starts from the starting point q s Departure, fly over the city to forward the data of ground mobile users to the base station, and arrive at the end point q e End. Set the base station location information to be expressed as The location information of the UAV and the kth user at time t are expressed as and Where [x(t), y(t)] represents the horizontal projection coordinates of the drone at time t, z(t) represents the height of the drone, [x k (t),y k (t)] represents the horizontal coordinate of user k at time t. Spherical coordinates V = (v, λ, ρ) are defined to represent the flight action of the drone, where v = ||V||, v∈[0,v max ],v max represents the maximum flight speed of the UAV, λ∈[0,π] represents the polar angle, and ρ∈[-π,π] represents the azimuth angle on the horizontal plane.
[0081] Establish the actual phase shift model of the smart reflector: Assume is the reflection coefficient of reflection unit m, then the reflection coefficient matrix can be described as θm (t) and β m (t) represent the phase shift and reflection amplitude of the mth reflection unit in the smart reflection surface, respectively, where θ m (t)∈[0,2π) and β m (t)∈[0,1],m∈{1,…,M}. θ m (t) is selected from a finite number of discrete values between 0 and 2π. Let the binary bit b represent the adjustable flexibility of the discrete phase shift, then the actual phase shift range is The reflection amplitude under the actual phase shift model can be calculated by the following formula:
[0082]
[0083] Among them, β min ≥0, φ≥0, and α≥0 are constants related to the specific circuit implementation. min is the minimum amplitude, φ is -π / 2 and β min β min The horizontal distance between them, α controls the steepness of the function curve. It is worth noting that when β min When α = 1 or (α = 0), the above formula is equivalent to an ideal phase shift model with unit amplitude.
[0084] The channel gains of user-base station, user-ARIS and ARIS-base station at time t are: Therefore, the signal received at the base station can be expressed as:
[0085]
[0086] in g k (t) is the linear receiving beam vector of the base station, s is the unit power complex baseband information symbol of the user, and represents the additive white Gaussian noise of the base station.
[0087] The signal-to-noise ratio received by the base station from user k is as follows:
[0088]
[0089] Among them, P k (t) is the transmission power of user k, σ 2 is the noise power.
[0090] 3. Establish and simplify the smart reflector phase shift optimization problem, and then obtain an exact solution through an improved penalty algorithm based on simulated annealing for global perturbation, thereby maximizing the overall gain of multi-user synchronous transmission. Specifically, it includes:
[0091] By adjusting the reflection coefficient matrix μ of the smart reflective surface m (t) and the linear receiving beam vector g of the base station k (t), with the goal of maximizing the overall gain of multi-user synchronous transmission, the smart reflector phase shift optimization problem is established:
[0092]
[0093] Multiple terrestrial users transmit synchronously based on frequency division multiple access technology, so the base station can generate a linear receiving beam vector for each transmitting user and optimize the linear receiving beam vector of the base station based on the maximum ratio transmission formula. Therefore, problem P1 is transformed into the following:
[0094]
[0095] Due to the non-convex constraints of (4a), solving problem P2 is challenging. To address these difficulties, a penalty-based approach is used to simplify P2, which penalizes constraint violations by adding constraint-related penalty terms in the objective function of the optimization problem to eliminate intractable equality constraints. To this end, problem P2 can be formulated as
[0096]
[0097] stθ m (t)∈Ω,m=1,...,M(6a)
[0098] Where ξ>0 is a penalty parameter that imposes a penalty on violations of the constraints in (4a). In particular, when ξ→∞, solving the above problem will produce an approximate solution to problem P2. This leads to a two-layer iterative algorithm, where the inner layer solves the penalized optimization problem P3, while the outer layer updates the penalty coefficient ξ until the solution converges. The specific solution steps are as follows:
[0099] a) For a given a, μ(t) in Problem P3 can be optimized by solving the following problem:
[0100]
[0101] in It is not difficult to observe that the first and second terms in problem P4 are convex and concave respectively, so we can use the convex-concave process principle to approximate the solution in an iterative manner.
[0102] In each subsequent iteration, at a given point μ (i)(t) is used to perform a first-order Taylor expansion on the first term of the objective function in the above problem and ignore the constant term, thereby simplifying it to a linear function to form a convex approximate optimization problem, which is given by the following formula:
[0103]
[0104] It is not difficult to observe that P5 is an unconstrained convex optimization problem for which the closed-form optimal solution can be easily obtained (by setting the first-order derivative of the objective function with respect to μ(t) to zero)
[0105]
[0106] Next, μ (i) (t) is updated to μ (i+1) (t), until the target value of P5 reaches convergence.
[0107] b) For a given μ(t), the a in Problem P4 can be optimized by solving the following problem
[0108]
[0109] stθ m (t)∈Ω,m=1,...,M(10a)
[0110] Since θ m (t) is completely separable in the objective function, so M independent sub-problems can be solved in parallel by expanding And ignoring the constant term, we can get:
[0111]
[0112] in Since θ m (t) is chosen from a finite number of discrete values between 0 and 2π, so that the optimal Although penalty-based algorithms can be used to solve P2, they can only guarantee a local optimal solution. To solve this problem, we propose a global perturbation strategy based on the simulated annealing algorithm. The simulated annealing algorithm performs global optimization based on the solution of Equation (11) to escape the local optimal solution.
[0113] 4. After adjusting the reflection coefficient matrix of the smart reflective surface and the linear receiving beam vector of the base station, considering the user mobility in a dynamic environment and the flight constraints of the drone, the drone trajectory optimization problem is constructed. It is modeled as a Markov decision process and the TD3 algorithm is used to optimize the drone's three-dimensional flight trajectory to maximize the system throughput. Specifically, it includes:
[0114] The transmission rate between the base station and user k is expressed as:
[0115] R k (t) = Blog2(1+γ k (t))(12)
[0116] Where B is the communication bandwidth of ARIS to users, γ k (t) is the channel ratio between the base station and user k, γ k (t) must be greater than or equal to the signal-to-noise ratio threshold γ th To accurately decode; the total transmission rate of the user can be expressed as:
[0117]
[0118] In order to maximize the system throughput in the ARIS-assisted wireless communication network within a specified time T, the following throughput maximization problem is constructed by optimizing the flight trajectory design of the UAV:
[0119]
[0120] γ(t)≥γ th (14a)
[0121] q u (0) = q s ,q u (T) = q e (14b)
[0122]
[0123] (μ m (t),g(t))=P1(14i)where q u (0) and q u (T) are the positions of the UAV at t = 0 and t = T, respectively, q s and q e are the starting point and end point of the drone, respectively, x min 、x max 、y min and y max is the constraint condition on the xy plane of the flight area, z min and z max is the restriction on the UAV’s flight altitude. Equation (14a) is the signal-to-noise ratio constraint; Equation (14b) ensures that the UAV must start from the starting point and reach the destination; Equations (14c) to (14e) restrict the UAV’s movements during flight; and Equations (14f) to (14h) define the UAV’s feasible flight space.
[0124] The above optimization problem is modeled as a Markov decision process. Next, the state space, action space, and reward function are given.
[0125] State design: The design of the state space includes the position, throughput, and remaining time and distance of the UAV, where the remaining time is t re =Tt, the remaining distance is the distance from the drone to the destination: The drone must arrive within the specified time. The agent knows its remaining time and the distance to the destination so that it can make a decision: whether to continue serving the user or arrive at the destination in time. Therefore, the state space s(t) is defined as follows:
[0126]
[0127] Action design: The action of the agent corresponds to the flight control of the UAV, which is determined by the constraints (14c) to (14e). The flight action of the UAV includes the flight direction and flight speed. Therefore, the action a(t) = {v(t), λ(t), ρ(t)}, where v(t)∈[0,v max ] represents the flight speed of the UAV in the next time slot; λ(t)∈[0,π] represents the flight polar angle of the next time slot; ρ(t)∈[-π,π] represents the flight azimuth angle of the next time slot.
[0128] Reward function design: The design of the reward determines whether the agent can learn the desired strategy. In each time slot, throughput is generated between ARIS and the user. The optimization goal is to maximize the throughput of the system. Therefore, the throughput-related reward is set to a positive reward. The throughput part of the reward is defined as follows:
[0129] r1=c R R(t)(16)
[0130] where c R is a positive constant used to reduce throughput, because too large a reward can lead to the "gradient vanishing" problem in the Tanh(·) activation function. Next is the reward for violating the flight area restriction. The reward for this part is defined as:
[0131] r2=-c xy ·ξ xy -c h ·ξ h (17)
[0132] Among them, ξ xy Used as a horizontal plane constraint indicator, ξ h is the binary height constraint indicator. c xy and c his a positive penalty constant associated with violating constraints (14f)–(14h). In order to ensure that the UAV can reach the destination within the specified time, a reward related to the distance from the UAV to the destination is set. The reward is defined as follows:
[0133] r3=c d w d (d max -d u,d (t))(18)
[0134] Among them, d max is the distance between the starting point and the end point, c d is a distance-dependent scaling factor. d =e -ζ(T-t) is a time-related weight coefficient. By setting ζ reasonably, let w d As t increases nonlinearly, when the remaining time t re When it is insufficient, r3 becomes dominant and the drone tends to fly to the destination to obtain a larger reward. The reward obtained by the agent when selecting action a(t) in a given state s(t) is defined as:
[0135] r(t)=r1+r2+r3(19)
[0136] Among them, r(t) is the reward function at time t.
[0137] Finally, according to the above Markov decision model, the TD3 algorithm is used to optimize the three-dimensional flight trajectory of the UAV, and the penalty algorithm based on simulated annealing for global perturbation is used to optimize the phase shift of the smart reflector to maximize the system throughput.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A relay communication network transmission method based on the assistance of an aerial intelligent reflector, characterized in that: The method specifically comprises the following steps: S1: Build a UAV-mounted intelligent reflective surface-assisted wireless relay communication system between the base station and ground mobile users; S2: Establish a wireless channel model for the relay communication network and an actual phase shift model for the smart reflector; S3: By adjusting the reflection coefficient matrix of the smart reflector and the linear receiving beam vector of the base station, the smart reflector phase shift optimization problem is constructed and simplified with the goal of maximizing the overall gain of multi-user synchronous transmission. A global perturbation strategy based on the simulated annealing algorithm is used to solve the problem and obtain the optimal phase shift of the smart reflector. S4: After adjusting the reflection coefficient matrix of the smart reflector and the linear receiving beam vector of the base station, considering the user mobility in a dynamic environment and the flight constraints of the drone, a system throughput optimization problem is constructed. It is modeled as a Markov decision process and the TD3 algorithm is used to optimize the three-dimensional flight trajectory of the drone to maximize the system throughput.
2. The relay communication network transmission method according to claim 1, characterized in that: In step S1, a UAV-mounted smart reflector-assisted wireless relay communication system is constructed, specifically including: a base station with N antennas, K ground mobile users, and a UAV equipped with a smart reflector. It is assumed that both the UAV and the user are equipped with a single omnidirectional antenna. The single-antenna users are randomly distributed and keep moving in a given area, k = {1, 2, ..., K} is the set of ground users. The smart reflector consists of M reflective elements, denoted as the set m = {1, 2, ..., M}.
3. The relay communication network transmission method according to claim 2, characterized in that: In step S2, a wireless channel model of the relay communication network is established, specifically including: the UAV carries an intelligent reflective surface, also known as an aerial intelligent reflective surface, denoted as ARIS; the system time is divided into discrete time slots, t = 0, 1, 2, ..., T; at time t = 0, ARIS starts from the starting point q s Departure, fly over the city to forward the data of ground mobile users to the base station, and arrive at the end point q e End; set the base station location information to The location information of the UAV and the kth user at time t are expressed as and 0≤t≤T, where [x(t),y(t)] represents the horizontal projection coordinates of the drone at time t, z(t) represents the altitude of the drone, [x k (t),y k (t)] represents the horizontal coordinate of user k at time t; spherical coordinates V = (v, λ, ρ) are defined to represent the flight action of the drone, where v = ||V||, v∈[0,v max ], v represents the flight speed of the drone, v max represents the maximum flight speed of the UAV, λ∈[0,π] represents the polar angle, and ρ∈[-π,π] represents the azimuth angle on the horizontal plane.
4. The relay communication network transmission method according to claim 3, characterized in that: In step S2, the actual phase shift model of the intelligent reflector is established, which specifically includes: setting is the reflection coefficient of reflection unit m, then the reflection coefficient matrix is described as θ m (t) and β m (t) represent the phase shift and reflection amplitude of the mth reflection unit in the smart reflection surface, respectively, where θ m (t)∈[0,2π) and β m (t)∈[0,1],m∈{1,…,M}; θ m (t) is selected from a finite number of discrete values from 0 to 2π; let the binary bit b represent the adjustable flexibility of the discrete phase shift, then the actual phase shift range is The reflection amplitude under the actual phase shift model is calculated by the following formula: Among them, β min ≥0, φ≥0 and α≥0 are constants related to circuit implementation; β min is the minimum amplitude, φ is -π / 2 and β min β min The horizontal distance between them, α controls the steepness of the function curve; when β min When =1 or α=0, the above formula is equivalent to an ideal phase shift model with unit amplitude; The channel gains of user-base station, user-ARIS and ARIS-base station at time t are: Therefore, the signal received at the base station is expressed as: in, g k (t) is the linear receiving beam vector of the base station, s is the unit power complex baseband information symbol of the user, and represents the additive white Gaussian noise of the base station; The signal-to-noise ratio γ received by the base station from user k k (t) is expressed as: Among them, P k (t) is the transmission power of user k, σ 2 is the noise power.
5. The relay communication network transmission method according to claim 4, characterized in that: In step S3, by adjusting the reflection coefficient matrix μ of the smart reflective surface m (t) and the linear receiving beam vector g of the base station k (t), with the goal of maximizing the overall gain of multi-user synchronous transmission, the smart reflector phase shift optimization problem is constructed as: Multiple terrestrial users transmit synchronously based on frequency division multiple access technology. Therefore, the base station generates a linear receiving beam vector for each transmitting user and optimizes the linear receiving beam vector of the base station based on the maximum ratio transmission formula. where g k (t) * is the optimized linear receiving beam vector of the base station; therefore, problem P1 is transformed into the following: st(4a)~(4b) Due to the non-convex constraints of (4a), a penalty-based approach is used to simplify P2. This approach penalizes constraint violations by adding constraint-related penalty terms to the objective function of the optimization problem to eliminate intractable equality constraints. To this end, problem P2 is formulated as s.t.θ m (t)∈Ω,m=1,...,M(6a) Among them, ξ>0 is the penalty parameter.
6. The relay communication network transmission method according to claim 5, characterized in that: In step S3, a global perturbation strategy based on a simulated annealing algorithm is used to solve and obtain the optimal phase shift of the smart reflective surface, which specifically includes: a) For a given a, μ(t) in Problem P3 is optimized by solving the following problem: in, It is not difficult to observe that the first and second terms in problem P4 are convex and concave respectively, so we use the convex-concave process principle to approximate the solution in an iterative manner; In each subsequent iteration, at a given point μ (i) (t) is used to perform a first-order Taylor expansion on the first term of the objective function in the above problem and ignore the constant term, thereby simplifying it to a linear function to form a convex approximate optimization problem, which is given by the following formula: It is not difficult to observe that P5 is an unconstrained convex optimization problem, for which the first-order derivative of the objective function with respect to μ(t) is set to zero; Among them, μ (i) (t) is the i-th order derivative of μ(t); Next, μ (i) (t) is updated to μ (i+1) (t), until the target value of P5 reaches convergence; b) For a given μ(t), problem a in problem P4 is optimized by solving the following problem: s.t.θ m (t)∈Ω,m=1,...,M(10a) Since θ m (t) is completely separable in the objective function, so we solve M independent subproblems in parallel by expanding And ignoring the constant term, we can get: in, Since θ m (t) is selected from a finite number of discrete values from 0 to 2π, so the optimal The simulated annealing algorithm is used to perform global optimization based on the solution of formula (11) to escape from the local optimal solution.
7. The relay communication network transmission method according to claim 6, characterized in that: In step S4, a system throughput optimization problem is constructed, specifically including: the transmission rate between the base station and user k is expressed as: R k (t)=Blog2(1+γ k (t)) (12) Where B is the communication bandwidth of ARIS to users, γ k (t) is the channel ratio between the base station and user k at time t, γ k (t) must be greater than or equal to the signal-to-noise ratio threshold γ th To accurately decode; the total transmission rate of the user can be expressed as: In order to maximize the system throughput in the ARIS-assisted wireless communication network within the specified time T, the following throughput maximization problem is constructed by optimizing the flight trajectory design of the UAV: γ(t)≥γ th (14a) what u (0)=q s ,q u (T)=q e (14b) (μ m (t),g(t))=P1 (14i)where q u (0) and q u (T) are the positions of the UAV at t = 0 and t = T, respectively, q s and q e are the starting point and end point of the drone, respectively, x min 、x max 、y min and y max is the constraint condition on the xy plane of the flight area, z min and z max is the restriction condition for the UAV’s flight altitude; Equation (14a) is the signal-to-noise ratio constraint; Equation (14b) ensures that the UAV must start from the starting point and reach the destination; Equations (14c) to (14e) restrict the UAV’s actions during flight; Equations (14f) to (14h) define the UAV’s feasible flight space.
8. The relay communication network transmission method according to claim 7, characterized in that: In step S4, the system throughput optimization problem P7 is modeled as a Markov decision process, which specifically includes: State design: The design of the state space includes the position, throughput, and remaining time and distance of the UAV, where the remaining time is t re =Tt, the remaining distance is the distance from the drone to the destination: The drone must arrive within the specified time. The agent knows its remaining time and the distance to the destination so that it can make a decision: whether to continue serving the user or arrive at the destination in time; therefore, the state space s(t) is defined as follows: Action design: The action of the agent corresponds to the flight control of the UAV, which is determined by the constraints (14c) to (14e). The flight action of the UAV includes the flight direction and flight speed. Therefore, the action a(t) = {v(t), λ(t), ρ(t)}, where v(t)∈[0,v max ] represents the flight speed of the UAV in the next time slot; λ(t)∈[0,π] represents the flight polar angle of the next time slot; ρ(t)∈[-π,π] represents the flight azimuth angle of the next time slot; Reward function design: The design of the reward determines whether the agent can learn the desired strategy. In each time slot, throughput is generated between ARIS and the user. The optimization goal is to maximize the throughput of the system. Therefore, the throughput-related reward is set to a positive reward. The throughput part of the reward is defined as follows: r1=c R R(t) (16) Among them, c R is a positive constant; the next is the reward for violating the flight area restriction, which is defined as: r2=-c xy ·x xy -c h ·x h (17) Among them, ξ xy is the horizontal plane constraint indicator, ξ h is a binary height constraint indicator; c xy and c h is a positive penalty constant associated with violating constraints (14f) to (14h). To ensure that the drone can reach the destination within the specified time, a reward related to the distance from the drone to the destination is set. The reward is defined as follows: r3=c d w d (d max -d u,d (t)) (18) Among them, d max is the distance between the starting point and the end point, c d is a distance-dependent scaling factor; w d =e -ζ(T-t) is a time-dependent weight coefficient, by setting ζ, let w d As t increases nonlinearly, when the remaining time t re When it is insufficient, r3 becomes dominant, and the drone tends to fly to the destination to obtain a larger reward. The reward obtained by the agent when selecting action a(t) in a given state s(t) is defined as: r(t)=r1+r2+r3 (19) Among them, r(t) is the reward function at time t.