A three-dimensional deployment and power allocation joint optimization method of a UAV

By optimizing the 3D deployment and power allocation of UAVs through deep deterministic policy gradient and water-filling algorithms, the problem of insufficient accuracy of UAV flight base stations in continuous states and action spaces is solved, thereby improving system throughput and network performance.

CN113206701BActive Publication Date: 2026-07-21CHONGQING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2021-04-30
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing methods for three-dimensional deployment and power allocation of UAV flight base stations suffer from insufficient accuracy when dealing with continuous states and action spaces, especially when UAVs have limited energy, making it difficult to optimize system throughput.

Method used

A Markov Decision Process (MDP) is constructed by combining a deep deterministic policy gradient algorithm with a water-filling algorithm to optimize the three-dimensional position and power allocation of the UAV. The action space is reduced in dimensionality by the deep deterministic policy gradient algorithm, and the power allocation is optimized by the water-filling algorithm.

Benefits of technology

It improved the throughput of the UAV system, enhanced the energy efficiency and network performance of the UAV, made full use of the distribution characteristics of ground users, and achieved optimal three-dimensional hovering position and power distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113206701B_ABST
    Figure CN113206701B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of unmanned aerial vehicle flight base stations, and particularly discloses a three-dimensional deployment and power distribution joint optimization method for dispatching unmanned aerial vehicles as flight base stations to serve ground user clusters. The influence of line-of-sight transmission and non-line-of-sight transmission on the air-ground channel from the unmanned aerial vehicle to each user is considered, and a maximum system throughput model for jointly optimizing the three-dimensional position of the unmanned aerial vehicle and power distribution is established. The model is solved in a continuous state and action space by using a deep reinforcement learning method, namely a deep deterministic policy gradient, and the action space is reduced in dimension by combining a water injection algorithm, so that the unmanned aerial vehicle successfully learns an optimal three-dimensional deployment position and power distribution strategy to provide maximum throughput for the served users, and the energy efficiency of the unmanned aerial vehicle is improved under the condition that the energy of the unmanned aerial vehicle is limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) flight base station technology, and in particular to a method for joint optimization of three-dimensional deployment and power allocation of UAVs. Background Technology

[0002] In the B5G era, drones offer a fast and cost-effective way to support temporary wireless connectivity needs, addressing issues such as terrestrial base station failures and network congestion. On one hand, compared to traditional terrestrial base stations, drone-based base stations can be rapidly deployed in remote areas (such as rural areas and mountainous regions) where extensive infrastructure deployment is difficult, and in areas hosting temporary hotspots (such as sporting events and concerts), significantly reducing the construction and maintenance costs of terrestrial infrastructure. On the other hand, flying drone base stations are more likely to establish line-of-sight links with ground users by adjusting their hovering position in three-dimensional space, thus providing higher data rates. Due to these advantages, researchers have extensively studied the optimal deployment of drone base stations. However, the three-dimensional deployment problem of drones is often a complex non-convex problem, and after incorporating resource allocation such as power, it involves optimizing continuous variables in higher dimensions. Current research is turning to machine learning methods for solving this problem. However, methods such as Q-learning and deep Q-networks, which have been widely used in previous studies, cannot handle continuous action spaces, leading to a loss of accuracy in the results. Therefore, employing a machine learning method capable of handling continuous state and action spaces to study the joint optimization of three-dimensional deployment and power allocation of UAV flight base stations with high-dimensional continuous variables can improve system throughput. This has significant practical implications for improving UAV energy efficiency and network performance, especially given the limited energy of UAVs. Summary of the Invention

[0003] This invention provides a method for joint optimization of three-dimensional deployment and power allocation of UAV flight base stations. The technical problem it solves is: how to determine the optimal hovering service position for UAVs to serve multiple ground users simultaneously, and how to allocate the optimal power to ground users in each location.

[0004] To address the above technical problems, this invention provides a method for joint optimization of three-dimensional deployment and power allocation of UAV flight base stations, comprising the following steps:

[0005] (1) Model of UAV base station system

[0006] S1: Establish a system model for UAV flight base station services to ground user clusters; the system model includes a UAV, and the UAV serves... A user cluster formed by ground users, and the air-to-ground channel from the UAV to the ground users;

[0007] (2) System throughput optimization model

[0008] S2: Simultaneously considering the impact of line-of-sight transmission and non-line-of-sight transmission on the air-to-ground channel, the path loss from the UAV to the ground user is obtained;

[0009] S3: With the goal of maximizing system throughput, the three-dimensional position and power allocation of the UAV are used as joint optimization variables to construct a system throughput optimization model for the UAV serving the ground user cluster; the system throughput optimization model includes an objective function that maximizes the sum of transmission rates of each ground user, as well as UAV altitude constraints, total transmit power constraints, user power non-negativity constraints, and reference signal received strength threshold constraints;

[0010] (3) Solving the system throughput optimization model

[0011] S4: Construct the system throughput optimization model as a Markov decision process (MDP), wherein the three-dimensional position of the UAV is set as the state space, the displacement of the UAV and the power allocated to each ground user are set as the action space, and a reward function is constructed based on the system throughput and the displacement of the UAV.

[0012] S5: During the state transition process of the MDP, the water-filling algorithm is used to determine the power allocation under the current three-dimensional position of the UAV, so that the action space of the MDP is reduced from UAV displacement and power allocation to UAV displacement. The depth deterministic strategy gradient algorithm is used to solve the dimensionality-reduced MDP to obtain the optimal three-dimensional deployment position of the UAV and the corresponding power allocation strategy.

[0013] Furthermore, the drone reaches a certain ground user The existence of line-of-sight transmission is represented as:

[0014]

[0015] in, and Indicates statistical parameters related to the geographical environment; This indicates that the drone is connected to the ground user. The angle of elevation, This represents the three-dimensional coordinates of the UAV. Indicates the ground user The three-dimensional coordinates This indicates that the drone is connected to the ground user. The straight-line distance.

[0016] Therefore, the corresponding non-line-of-sight transmission is represented as:

[0017]

[0018] Furthermore,

[0019]

[0020]

[0021] in, This represents the free-space propagation path loss. Indicates the carrier frequency. Represents the speed of light; This indicates that the drone is connected to the ground user. The total path loss is the mathematical expectation of the sum of the free-space propagation path loss and the additional path loss caused by line-of-sight and non-line-of-sight transmission. and These represent the additional path loss caused by line-of-sight transmission and non-line-of-sight transmission, respectively.

[0022] Furthermore, disregarding fast and slow fading in the channel, the UAV to the ground user Channel gain Represented as:

[0023]

[0024] in, It is based on equation (1) regarding , , and The function; except for the three-dimensional position of the UAV. In addition, the channel gain If all other parameters are known quantities or constants, then... It concerns the three-dimensional position of the drone. The function.

[0025] Furthermore, set For a ground user to successfully demodulate the UAV's transmitted signal at a reference signal received strength (RSRP) threshold, then the UAV to a certain ground user... transmission rate Represented as:

[0026]

[0027] in, Indicates the bandwidth of the system. This represents the total number of ground users. Bandwidth is orthogonally shared by each user. To avoid wireless interference, This represents the power spectral density of Gaussian white noise. Indicates the user RSRP value.

[0028] Therefore, based on equation (5), equation (6) is about the three-dimensional position of the UAV. and assigned to a certain ground user power The function.

[0029] Furthermore, in step S3, the established system throughput optimization model is specifically as follows:

[0030]

[0031]

[0032]

[0033]

[0034]

[0035] Wherein, the objective function (7) represents maximizing the system throughput, and the decision variable is the three-dimensional position of the UAV. and assigned to a certain ground user power , yes A set of ground users; constraint (8) represents the altitude limit of the UAV. and These represent the minimum and maximum allowed altitudes, respectively; constraint (9) represents the total transmit power limit of the UAV. Constraint (10) indicates that it is assigned to the user. The power is non-negative; constraint (11) indicates that the UAV only serves the RSRP value. Greater than RSRP threshold Users.

[0036] Furthermore, in step S4, the state space, action space, state transition probability, and reward function of the MDP are specifically as follows:

[0037] S41: Set the three-dimensional position of the UAV The state space of the MDP ;

[0038] S42: Set the displacement of the drone and the power allocated to the ground users The action space of the MDP ;

[0039] S43: Based on the aforementioned state space and action space, determine the next state of the UAV after performing the current action, wherein the current state is the current three-dimensional position of the UAV, the current action includes the displacement of the UAV and the power allocated to the ground user, and the next state is the sum of the current three-dimensional position and the displacement; the state transition probability of the MDP is expressed as:

[0040]

[0041] in, and These represent the next state and the current state, respectively. Indicates the current action.

[0042] S44: Based on the optimization objective of the system throughput optimization model and the actions of the UAV, set a certain state transition time. The reward value of the MDP is as follows:

[0043]

[0044] in, and These are the adjustment factors for the rewards. The first item in the rewards represents the reward for improving system throughput, and the second item represents the penalty for large-scale displacement of the drone.

[0045] Further, in step S5, the loss function for updating the parameters of the two estimation networks using the deep deterministic policy gradient is:

[0046]

[0047]

[0048] in, and These are Actor estimation networks. and Critic estimation network Parameters; Output actions based on the current state of the drone. The action is scored, and a Q value is given; the two estimation networks update their own parameters by minimizing the loss functions in equations (14) and (15), respectively.

[0049] Furthermore, in the loss function of equation (15) Represented as:

[0050]

[0051] in, It is the reward value of the MDP based on equation (13). Reward discount factor, and These are the target Actor network and the target Critic network for the deep deterministic policy gradient, respectively. The two target networks and the two estimation networks have the same structure, with parameter updates using a soft update method, where each update copies a portion of the parameters from the estimation network. The formula for the soft update is expressed as:

[0052]

[0053]

[0054] in, and These are the parameters of the target Actor network and the target Critic network, respectively. It is a soft update factor, satisfying .

[0055] This invention provides a joint optimization method for the 3D deployment and power allocation of UAV flight base stations. By employing a deep deterministic policy gradient, the UAV flight base station can fully utilize the distribution characteristics of ground users and learn the optimal 3D hovering position in a continuous state and action space. By combining a water-filling algorithm, the optimal power allocation for each state involved in training is obtained, thereby reducing the dimensionality of the action space. System throughput can be effectively improved through the optimal joint optimization of UAV 3D deployment and power allocation, which has significant practical implications. Attached Figure Description

[0056] Figure 1 This is a flowchart illustrating the steps of a method for joint optimization of three-dimensional deployment and power allocation of a UAV flight base station provided in an embodiment of the present invention.

[0057] Figure 2 This is a model diagram of the UAV base station system provided in an embodiment of the present invention;

[0058] Figure 3This is a schematic diagram of the gradient principle of a deep deterministic strategy provided in an embodiment of the present invention;

[0059] Figure 4 This is the gradient accumulation reward graph of the deep deterministic strategy provided in the embodiments of the present invention;

[0060] Figure 5 This is a comparison chart of system throughput provided in an embodiment of the present invention;

[0061] Figure 6 This is a 3D deployment diagram of a drone base station provided in an embodiment of the present invention; Detailed Implementation

[0062] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0063] To determine the optimal hovering service location for a drone flight base station to simultaneously serve multiple ground users and the optimal power allocation to each ground user, embodiments of this invention provide a joint optimization method for three-dimensional deployment and power allocation of drone flight base stations, such as... Figure 1 The steps are shown in the flowchart, specifically including the following steps:

[0064] (1) Model of UAV base station system

[0065] S1: Establish a system model for UAV flight base station services to ground user clusters; the system model includes a UAV, and the UAV serves... A user cluster formed by ground users, and the air-to-ground channel from the UAV to the ground users;

[0066] exist Figure 2 The system model shown considers Ground users at known locations (As shown by the dots in the diagram). Consider a drone serving this user cluster. The air-to-ground channel from the drone to the ground users includes two transmission modes: line-of-sight (LoS) and non-line-of-sight (NLoS).

[0067] (2) System throughput optimization model

[0068] The specific steps include:

[0069] S2: Simultaneously considering the impact of line-of-sight transmission and non-line-of-sight transmission on the air-to-ground channel, the path loss from the UAV to the ground user is obtained;

[0070] S3: With the goal of maximizing system throughput, the three-dimensional position and power allocation of the UAV are used as joint optimization variables to construct a system throughput optimization model for the UAV serving the ground user cluster; the system throughput optimization model includes an objective function that maximizes the sum of transmission rates of each ground user, as well as UAV altitude constraints, total transmit power constraints, user power non-negativity constraints, and reference signal received strength threshold constraints;

[0071] In step S2, we adopt a widely used air-to-ground channel model in the literature, which considers the possibility of both line-of-sight and non-line-of-sight transmission. The UAV then travels to a ground user... The existence of line-of-sight transmission is represented as:

[0072]

[0073] in, and Indicates statistical parameters related to the geographical environment; This indicates that the drone is connected to the ground user. The angle of elevation, This represents the three-dimensional coordinates of the UAV. Indicates the ground user The three-dimensional coordinates This indicates that the drone is connected to the ground user. The straight-line distance.

[0074] Therefore, the corresponding non-line-of-sight transmission is represented as:

[0075]

[0076] Then, the drone goes to the ground user. The total path loss can be expressed as the mathematical expectation of the free-space propagation path loss plus the additional path loss caused by line-of-sight and non-line-of-sight transmission, specifically:

[0077]

[0078]

[0079] in, This represents the free-space propagation path loss. Indicates the carrier frequency. Represents the speed of light; and These represent the additional path loss caused by line-of-sight transmission and non-line-of-sight transmission, respectively.

[0080] Next, we construct the system throughput optimization model described in step S3.

[0081] Ignoring fast and slow fading in the channel, the UAV to the ground user Channel gain Represented as:

[0082]

[0083] in, It is based on equation (1) regarding , , and The function; except for the three-dimensional position of the UAV. In addition, the channel gain If all other parameters are known quantities or constants, then... It concerns the three-dimensional position of the drone. The function.

[0084] definition The total transmission power of the UAV. To be allocated to a certain ground user The power. Then, set. For a ground user to successfully demodulate the UAV's transmitted signal at a reference signal received strength (RSRP) threshold, then the UAV to a certain ground user... transmission rate Represented as:

[0085]

[0086] in, Indicates the bandwidth of the system. This represents the total number of ground users. Bandwidth is orthogonally shared by each user. To avoid wireless interference, This represents the power spectral density of Gaussian white noise. Indicates the user RSRP value.

[0087] Therefore, based on equation (5), equation (6) is about the three-dimensional position of the UAV. and assigned to a certain ground user power The function.

[0088] The established system throughput optimization model is specifically as follows:

[0089]

[0090]

[0091]

[0092]

[0093]

[0094] Wherein, the objective function (7) represents maximizing the system throughput, and the decision variable is the three-dimensional position of the UAV. and assigned to a certain ground user power , yes A set of ground users; constraint (8) represents the altitude limit of the UAV. and These represent the minimum and maximum allowed altitudes, respectively; constraint (9) represents the total transmit power limit of the UAV. Constraint (10) indicates that it is assigned to the user. The power is non-negative; constraint (11) indicates that the UAV only serves the RSRP value. Greater than RSRP threshold Users.

[0095] (3) Solving the system throughput optimization model

[0096] The specific steps include:

[0097] S4: Construct the system throughput optimization model as a Markov decision process (MDP), wherein the three-dimensional position of the UAV is set as the state space, the displacement of the UAV and the power allocated to each ground user are set as the action space, and a reward function is constructed based on the system throughput and the displacement of the UAV.

[0098] S5: During the state transition process of the MDP, the water-filling algorithm is used to determine the power allocation under the current three-dimensional position of the UAV, so that the action space of the MDP is reduced from UAV displacement and power allocation to UAV displacement. The depth deterministic strategy gradient algorithm is used to solve the dimensionality-reduced MDP to obtain the optimal three-dimensional deployment position of the UAV and the corresponding power allocation strategy.

[0099] In step S4, the system throughput optimization model is established as a Markov Decision Process (MDP). The MDP is represented as a quadruple. This consists of the state space, action space, state transition probabilities, and reward. At each state transition moment, the drone transitions from the current state to the next state based on the current action and the state transition probability, and then receives the reward. This process is repeated until the maximum state transition moment is met.

[0100] The specific steps for constructing the MDP in this embodiment further include:

[0101] S41: Set the three-dimensional position of the UAV The state space of the MDP The state space has a dimension of 3;

[0102] S42: Set the displacement of the drone and the power allocated to the ground users The action space of the MDP The dimension of the action space is ;

[0103] S43: Based on the aforementioned state space and action space, determine the next state of the UAV after performing the current action, wherein the current state is the current three-dimensional position of the UAV, the current action includes the displacement of the UAV and the power allocated to the ground user, and the next state is the sum of the current three-dimensional position and the displacement; the state transition probability of the MDP is expressed as:

[0104]

[0105] in, and These represent the next state and the current state, respectively. Indicates the current action.

[0106] S44: For a certain state transition time According to the optimization objective of equation (7), the system throughput at this moment is taken as the reward value. However, at the moment of reaching the maximum state transition... Previously, drones would not stop transitioning between states. Therefore, if the drone was at a certain time... If the drone transitions to an optimal state, but the Actor network with a deep deterministic policy gradient outputs a large action (displacement) value, the drone will continue to transition states based on that action, thus entering a suboptimal state. Therefore, a punitive reward is needed to limit the action output by the network, i.e., the displacement of the drone. To improve convergence performance.

[0107] This embodiment will specify a state transition time. The reward value is set as follows:

[0108]

[0109] in, and These are the adjustment factors for the rewards. The first item in the rewards represents the reward for improving system throughput, and the second item represents the penalty for large-scale displacement of the drone.

[0110] In equation (13), by adjusting the factor and After readjusting the order of magnitude, the first term should be significantly larger than the second. Thus, in the initial stages of training a network with a deep deterministic policy gradient, the first term dominates the reward. After some training epochs, the increase in reward tends to level off. Then, the second displacement penalty begins to take effect, preventing the drone from conducting large-scale exploration, thereby allowing for a smoother convergence to the optimal position.

[0111] Next, the action space is reduced in dimensionality using the water-filling algorithm, and the MDP model is solved by gradient descent using a depth deterministic strategy.

[0112] The principle of the "water-filling" algorithm is to adaptively allocate the UAV's transmission power based on channel quality. Typically, more power is allocated to users with good channel quality, and less power is allocated to users with poor channel quality, thereby maximizing transmission power. The specific process of the water-filling algorithm can be described as follows:

[0113] 1) Based on the objective function and constraints of the original problem, construct the equation using the Lagrange multiplier method.

[0114] 2) Set the partial derivatives of the constructed equations to zero to obtain the power allocation expressions for each user with unknowns.

[0115] 3) Substitute the power allocation expressions for each user into the constraint conditions to obtain the unknowns.

[0116] 4) Substitute the obtained unknowns into the original expression to obtain the power allocation expression for each user without unknowns.

[0117] In step S5, the action space of the MDP is taken into account. In this case, if the dimension of power allocation is much larger than the dimension of UAV displacement, that is, if This will cause a dimensionality imbalance problem, making it difficult for the network training to converge to the optimal solution. Since the 3D position of the UAV is determined in any given state in a Multidimensional Power DP (Multidimensional Power DP), for a given state... According to equation (5), the path loss between the UAV and the ground user in state The following is also determined. Therefore, in state s, the problem... This is a convex power allocation problem, which can be easily solved using convex optimization methods. Therefore, to address the dimensionality imbalance problem, a water-filling algorithm is incorporated into the iterative process of the MDP to output the state. Optimal power distribution reduces the dimensionality of the motion space. .

[0118] The specific working principle of deep deterministic policy gradient is as follows: Figure 3 As shown, it stores the state transition iteration process of the MDP as experience in an experience replay buffer, and randomly selects experience samples from it to train two estimation networks: the Actor estimation network and the Critic estimation network, to fit the optimal action function and action-value function, respectively. The action function maps states to actions, and the action-value function scores actions and outputs a Q-value. To stabilize network training, the deep deterministic policy gradient uses a structurally identical sub-network in both the Actor and Critic networks, called the target network. The target network is not trained; instead, it is updated by copying a small set of parameters from the estimation network each time.

[0119] The loss function used in this embodiment to train and update the parameters of the two estimation networks is:

[0120]

[0121]

[0122] in, and These are Actor estimation networks. and Critic estimation network Parameters; Output actions based on the current state of the drone. The action is scored, and a Q value is given; the two estimation networks update their own parameters by minimizing the loss functions in equations (14) and (15), respectively. It is the size of the empirical sample.

[0123] The loss function in equation (15) Represented as:

[0124]

[0125] in, It is the reward value of the MDP based on equation (13). Reward discount factor, and These are the Actor target network and the Critic target network, respectively. The two target networks and the two estimation networks have the same structure, and the parameter updates are performed using soft updates, meaning that each update copies a portion of the parameters from the estimation network. The formula for soft updates is expressed as:

[0126]

[0127]

[0128] in, and These are the parameters of the target Actor network and the target Critic network, respectively. It is a soft update factor, satisfying .

[0129] The depth-deterministic strategy gradient algorithm combined with the water-filling algorithm in this embodiment can be described as follows:

[0130] 1: Initialize the Actor estimation network and the Critic estimation network , Initialize the Actor target network and the Critic target network. , ;have , 2: Initialize the experience replay buffer 3: for each training round do 4: Initialize the UAV's three-dimensional position 5: for each state transition time do 6: Drones observe their own three-dimensional position As a state 7: Actor Output Action 8: Run the water injection algorithm to obtain the system throughput in this state. And the drones receive reward value 9: Drone Operations And update To the next state 10: Storage Experience To the experience replay cache 11: if the experience replay cache is full then 12: Randomly select from the cache A sample of experiences 13: Calculate the loss according to equations (14) and (15), and update the two estimation networks. 14: Soft update the two target networks according to equations (17) and (18). 15: end if 16: end for 17: end for

[0131] In line 7 of the algorithm, during the training of the Actor network, its output actions are often accompanied by noise to prevent the drone from getting trapped in local optima. After the Actor network completes training, the noise in the output actions is removed.

[0132] Consider a specific implementation scenario, setting a 2 km... A rectangular geographical area of ​​2 km, with random distribution within the area. For each ground user, the other parameter settings are as follows:

[0133] parameter value 100 m 1000 m 2 GHz 100 MHz 9.61 0.16 1 20 20 dBm -174 dBm / Hz 25

[0134] In this embodiment, both the Actor and Critic networks consist of one input layer, two hidden layers, and one output layer. The number of neurons in the hidden layers is (200, 100) in the Actor network and (400, 200) in the Critic network. The activation function in the hidden layers is the ReLU function. The action noise follows a normal distribution with a mean of zero, and the bias decreases linearly from 0.3 to 0 after training epochs. The Adam optimizer is used to train the network with a learning rate of 0.0001. The remaining network parameters are set as shown in the table below:

[0135] parameter value 0.99 0.005 32 <![CDATA[1×10 -8 ]]> 0.001

[0136] This embodiment compares the performance of the proposed algorithm (JODP) with two other traditional methods (OA and OD) through experiments. In OA, the power of the UAV is evenly distributed among all ground users, and the UAV's planar position is fixed at the center of the user cluster (i.e., the origin of the coordinate system), optimizing only the UAV's altitude; in OD, the three-dimensional position of the UAV is optimized, and the power is evenly distributed among all ground users.

[0137] Figure 4 It is a cumulative reward graph of the gradient of a deep deterministic policy. From Figure 4 As can be seen, with the increase of training rounds, the JODP algorithm proposed in this embodiment can accumulate more rewards, and all three algorithms can converge stably. Figure 5 This is a comparison chart of system throughput. We use a Deep Q-Network (DQN) to illustrate the bias caused by discretizing the action space. From Figure 5 As can be seen, the JODP proposed in this embodiment outperforms OA and OD in terms of system throughput. Compared with the Deep Deterministic Policy Gradient (DDPG) in a continuous action space, the performance of the deep Q-network is worse, and the gap gradually widens. This is because the action space dimension of the three methods increases sequentially, and the bias caused by discretizing the action space also increases accordingly.

[0138] Figure 6 This is a 3D deployment diagram of the drone flight base station. From Figure 6 As can be seen, the drone altitude in OA (Automatic Access) is much higher than in other methods. This is because the drone's horizontal position is fixed in OA, so the drone must fly higher to establish more connections with ground users, at the expense of channel quality. In contrast, drones in OD (Operational Distributed Access) and JODP (Joint Operational Persistent Technology) can adjust their horizontal position to hover over hotspot areas with more users and establish better channels for these users. Furthermore, after considering optimal power allocation, JODP drones fly at a lower altitude than OD drones. This is because the water-filling algorithm allocates more power to users with better channels, prompting drones to approach hotspot areas. Therefore, when user distribution becomes more heterogeneous, JODP will outperform OD in terms of system throughput to a greater extent.

[0139] In summary, this invention provides a joint optimization method for the 3D deployment and power allocation of UAV flight base stations. By employing a deep deterministic policy gradient, the UAV flight base station can fully utilize the distribution characteristics of ground users and learn the optimal 3D hovering position in a continuous state and action space. Furthermore, by combining a water-filling algorithm, the optimal power allocation for each state involved in training is obtained, thereby reducing the dimensionality of the action space. System throughput can be effectively improved through the optimal joint optimization of UAV 3D deployment and power allocation, which has significant practical implications.

[0140] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A joint optimization method for three-dimensional deployment and power allocation of a UAV, characterized in that, Including the following steps: (1) Unmanned Aerial Vehicle System Model S1: Establish a system model for unmanned aerial vehicle (UAV) services to a cluster of ground users; the system model includes a UAV, and the UAV services... A user cluster formed by ground users, and the air-to-ground channel from the UAV to the ground users; (2) System throughput optimization model S2: Simultaneously considering the impact of line-of-sight transmission and non-line-of-sight transmission on the air-to-ground channel, the path loss from the UAV to the ground user is obtained; S3: With the goal of maximizing system throughput, the three-dimensional position and power allocation of the UAV are used as joint optimization variables to construct a system throughput optimization model for the UAV serving the ground user cluster; the system throughput optimization model includes an objective function that maximizes the sum of transmission rates of each ground user, as well as UAV altitude constraints, UAV total transmit power constraints, user power non-negativity constraints, and reference signal received strength threshold constraints; (3) Solving the system throughput optimization model S4: Construct the system throughput optimization model as a Markov decision process (MDP), wherein the three-dimensional position of the UAV is set as the state space, the displacement of the UAV and the power allocated to each ground user are set as the action space, and a reward function is constructed based on the system throughput and the displacement of the UAV. S5: During the state transition process of the MDP, the water-filling algorithm is used to determine the power allocation of the UAV at its current three-dimensional position, so that the action space of the MDP is reduced from UAV displacement and power allocation to UAV displacement. The depth deterministic strategy gradient algorithm is used to solve the dimensionality-reduced MDP to obtain the optimal three-dimensional position of the UAV and the corresponding power allocation.

2. The method for joint optimization of three-dimensional deployment and power allocation of a UAV as described in claim 1, characterized in that, In step S2, the drone reaches a ground user. The existence of line-of-sight transmission is represented as: in, and Indicates statistical parameters related to the geographical environment; This indicates that the drone is connected to the ground user. The angle of elevation, This represents the three-dimensional coordinates of the UAV. Indicates the ground user The three-dimensional coordinates This indicates that the drone is connected to the ground user. The straight-line distance; then, the corresponding non-line-of-sight transmission is represented as: 。 3. The method for joint optimization of three-dimensional deployment and power allocation of a UAV as described in claim 2, characterized in that: in, This represents the free-space propagation path loss. Indicates the carrier frequency. Represents the speed of light; This indicates that the drone is connected to the ground user. The total path loss is the mathematical expectation of the sum of the free-space propagation path loss and the additional path loss caused by line-of-sight and non-line-of-sight transmission. and These represent the additional path loss caused by line-of-sight transmission and non-line-of-sight transmission, respectively.

4. The method for joint optimization of three-dimensional deployment and power allocation of a UAV as described in claim 3, characterized in that, Ignoring fast and slow fading in the channel, the UAV to the ground user Channel gain Represented as: in, It is based on equation (1) regarding , , and The function; except for the three-dimensional position of the UAV. In addition, the channel gain If all other parameters are known quantities or constants, then... It concerns the three-dimensional position of the drone. The function.

5. The method for joint optimization of three-dimensional deployment and power allocation of a UAV as described in claim 4, characterized in that, set up For the ground user to successfully demodulate the UAV's transmitted signal, the reference signal received strength (RSRP) threshold is defined. Then, the UAV's signal to a certain ground user... transmission rate Represented as: in, Indicates the bandwidth of the system. This represents the total number of ground users. Bandwidth is orthogonally shared by each user. To avoid wireless interference, This represents the power spectral density of Gaussian white noise. Indicates the user The RSRP value; then, based on equation (5), equation (6) is about the three-dimensional position of the UAV. and assigned to a certain ground user power The function.

6. The method for joint optimization of three-dimensional deployment and power allocation of a UAV as described in claim 5, characterized in that, In step S3, the established system throughput optimization model is specifically as follows: Wherein, the objective function (7) represents maximizing the system throughput, and the decision variable is the three-dimensional position of the UAV. and assigned to a certain ground user power , yes A set of ground users; constraint (8) represents the altitude limit of the UAV. and These represent the minimum and maximum allowed altitudes, respectively; constraint (9) represents the total transmit power limit of the UAV. Constraint (10) indicates that it is assigned to the user. The power is non-negative; constraint (11) indicates that the UAV only serves the RSRP value. Greater than RSRP threshold Users.

7. The method for joint optimization of three-dimensional deployment and power allocation of a UAV as described in claim 6, characterized in that, In step S4, the state space, action space, state transition probabilities, and reward function of the MDP are specifically as follows: S41: Set the three-dimensional position of the UAV The state space of the MDP ; S42: Set the displacement of the drone and the power allocated to the ground users The action space of the MDP ; S43: Based on the aforementioned state space and action space, determine the next state of the UAV after performing the current action, wherein the current state is the current three-dimensional position of the UAV, the current action is the displacement of the UAV, and the next state is the sum of the current three-dimensional position and the displacement; the state transition probability of the MDP is expressed as: in, and These represent the next state and the current state, respectively. Indicates the current action; S44: Based on the optimization objective of the system throughput optimization model and the actions of the UAV, set a certain state transition time. The reward value of the MDP is as follows: in, and These are the adjustment factors for the rewards. The first item in the rewards represents the reward for improving system throughput, and the second item represents the penalty for large-scale displacement of the drone.