Unmanned aerial vehicle data collection trajectory optimization method based on customized channel
By combining ray tracing and machine learning to design customized channel models, and using Optimized PPO algorithm to optimize the drone trajectory, the problem of large channel modeling errors in urban areas is solved, and the efficiency of drone data collection is improved.
Patent Information
- Application Number
- CN202510356961.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
Smart Images

Figure CN120301540A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for optimizing the data collection trajectory of an unmanned aerial vehicle (UAV). Background Art
[0002] Unmanned aerial vehicles (UAVs) are widely used in various communication scenarios due to their high mobility, low cost, and line-of-sight propagation. UAVs have three typical applications in traditional wireless communication: 1) UAV-assisted full coverage; 2) UAV-assisted relaying; 3) UAV-assisted information dissemination and collection. Considering the third application, UAVs collect data. In existing UAV data collection scenarios, there are various UAV-to-ground channel models. The simplest one is to directly use the free space path loss model. Considering that there may be building blockages in urban areas, some studies adopt the Rayleigh channel. Some studies use measurement-based UAV-to-ground channel models. The measurement-based LOS probability model calculates the probabilities of line-of-sight (LOS) and non-line-of-sight (NLOS) propagation based on the elevation angle of the user relative to the UAV. The path loss model can be obtained by using the probability-weighted average of LOS and NLOS path losses. There is also the α-β path loss model based on measurement. The parameters α, β, and σ in the model change with the change of the UAV altitude. The measurement-based channel model is a statistical model. Due to the irregular shape of obstacles in a specific scenario, the UAV-to-ground channel mainly depends on environmental information. The statistical model will have a large error in a specific scenario. Summary of the Invention
[0003] The object of the present invention is to provide a reinforcement learning optimization method for UAV data collection based on a customized channel.
[0004] To achieve the above object, the technical solution of the present invention discloses a method for optimizing the data collection trajectory of an unmanned aerial vehicle based on a customized channel, which is characterized by including the following steps:
[0005] Step 1: After selecting the area to model the customized channel, obtain the three-dimensional topographic map of the area;
[0006] Step 2: Import the three-dimensional topographic map into ray tracing simulation software, and set the simulation parameters in the ray tracing simulation software, including the simulation frequency band, conductivity, permeability, transmitter position, and receiver position;
[0007] Step 3: Obtain the simulation data path loss and perform data cleaning, where each simulation data set contains path loss, TX coordinates and RX coordinates (x uav , y uav , h uav ) seven elements;
[0008] Step 4: Set the hyperparameters of the DNN model
[0009] Step 5: Use the simulated data path loss to train the DNN model to obtain a customized channel model, and apply this customized channel model to UAV data collection;
[0010] Step 6: Model the UAV data collection problem as an optimization problem, where:
[0011] For the uplink communication model where one UAV serves N user equipments, N≥1, the N user equipments are randomly distributed in the area selected in Step 1, and there is some data waiting to be collected. The UAV takes off from a fixed starting point, and collecting all the data of the user equipments within a specified time is regarded as successfully completing the task. Then the goal of the UAV data collection problem is to minimize the task completion time, and the optimization problem is expressed as:
[0012]
[0013] s.t. Eqs. (1)(2)(3)(4)(5)(6)(7)(8)(9)(10)
[0014] In the above optimization problem, Eq. (1) is expressed as:
[0015]
[0016] Eq. (2) is expressed as:
[0017]
[0018] Eq. (3) is expressed as:
[0019] 30m ≤ h uav [t] ≤ 120m
[0020] Eq. (4) is expressed as:
[0021]
[0022] Eq. (5) is expressed as:
[0023]
[0024] Eq. (6) is expressed as:
[0025]
[0026] Eq. (7) is expressed as:
[0027]
[0028] Eq. (8) is expressed as:
[0029]
[0030] Equation (9) is expressed as:
[0031]
[0032] Equation (10) is expressed as:
[0033]
[0034] Where: x uav [t] represents the X-axis coordinate of the UAV at time slot t;
[0035] represents the X-axis coordinate of user equipment n;
[0036] N is the set used to index user equipment, N = {1, 2, …, N};
[0037] T is the set used to index time slots, T = {1, 2, …, T};
[0038] y uav [t] represents the Y-axis coordinate of the UAV at time slot t;
[0039] represents the Y-axis coordinate of user equipment n;
[0040] h uav [t] represents the altitude of the UAV at time slot t;
[0041] l i [t] represents the index of the user equipment with the i-th largest path loss or the i-th smallest channel gain;
[0042] represents the user equipment l communicating with the UAV at time slot t i [t] is the transmission power used by user equipment l;
[0043] p max represents the maximum transmission power of user equipment;
[0044] a n [t] = 1 indicates that user equipment n has not completed data collection at time slot t and needs to communicate with the UAV; a n [t] = 0 indicates that user equipment n has completed data collection at time slot t and does not need to communicate with the UAV;
[0045] D n [t] represents the data transmission completed between the UAV and user equipment n at time slot t;
[0046] represents the data that user equipment n needs to collect;
[0047] t end The time for the drone to collect all user equipment data, t end ≤t max t max is the maximum working time of the drone. If the working time exceeds t max , the mission fails;
[0048] q[t] is the coordinate of the drone at time slot t;
[0049] v[t] is the speed of the drone at time slot t;
[0050] is the horizontal angle of the drone at time slot t;
[0051] ψ[t] is the elevation angle of the drone at time slot t;
[0052] v max is the maximum speed of the drone;
[0053] Step 7: Design the state, action, reward function, and hyperparameters of the Optimized PPO algorithm for reinforcement learning;
[0054] Step 8: Use the Optimized PPO algorithm for reinforcement learning to solve the optimization problem established in Step 6, and obtain the minimized t end .
[0055] Preferably, in Step 2: The transmitter is placed at a height of 1.5 m from the ground, and one transmitter is placed every 10 m until the area selected in Step 1 is covered: The horizontal position of the receiver is the same as that of the transmitter, and the height starts from 30 m and one layer is placed every 10 m; The simulated flight height of the drone is from 30 m to 120 m.
[0056] Preferably, in Step 4, the set DNN model hyperparameters include the learning rate, the number of neurons, the depth of the neural network, and the batch size.
[0057] Preferably, in Step 5, the customized channel model can predict the path loss, expressed as where PL(dB) is the path loss, and F DNN (·) represents the transfer function of the DNN.
[0058] Preferably, in Step 6, D n [t] is expressed as:
[0059]
[0060] In the formula, τ is the duration of each time period, and R n [t′] represents the transmission rate of user equipment n at time slot t′.
[0061] Preferably, user equipment l i [t] Transmission rate in time slot t Is expressed as:
[0062]
[0063] Where B is the bandwidth, N0 is the power spectral density of additive white Gaussian noise, Represents user l i [t] Coefficient indicating whether data collection is completed, Represents user equipment l i [t] And the path loss between the user equipment and the UAV in time slot t.
[0064] Preferably, in step 7, the state of time slot t is expressed as:
[0065]
[0066] Where N c [t] Is the number of UAVs that complete data collection in time slot t;
[0067] The action of time slot t is expressed as:
[0068]
[0069] The reward function of time slot t is expressed as:
[0070]
[0071] Where: Is the reward for the UAV to complete data acquisition,
[0072]
[0073] Represents the reward for each user equipment collected,
[0074] Is the total data collected in each time period,
[0075]
[0076] λ1, λ2, λ3, λ4 are constant parameters such that the reward Belongs to the same order of magnitude.
[0077] Preferably, in step 8, the following means are used to optimize the Reinforcement Learning Optimized PPO algorithm and accelerate the convergence of the Reinforcement Learning Optimized PPO algorithm:
[0078] Measure 1) Advantage function normalization
[0079] After calculating the advantage estimation function in a certain batch using Generalized Advantage Estimation, calculate the mean and standard deviation of all in the entire batch, and then subtract the mean and divide by the standard deviation;
[0080] Measure 2) Reward function normalization
[0081] Dynamically calculate the standard deviation of the cumulative discounted reward, and then divide the current reward by this standard deviation only;
[0082] Measure 3) Learning rate decay
[0083] Adopt the method of linearly decaying the learning rate, so that the learning rate linearly decreases from the initial value to 0 as the number of training steps increases, making the training more stable;
[0084] Measure 4) Gradient clipping
[0085] Introduce gradient clipping to prevent gradient explosion during the backpropagation of the critic network in the policy and Reinforcement Learning Optimized PPO algorithm.
[0086] Preferably, the advantage estimation function is expressed as:
[0087]
[0088] where: γ ∈ [0, 1], is the reward discount factor; λ ∈ [0, 1]; δ t = r t + γV μ (s t+1 ) - V μ (s t ), r t is the reward, V μ (s t ) is the state value output by the critic network, and μ is the parameter of the critic network.
[0089] The present invention combines ray tracing (RT) and machine learning (ML) to design a customized channel. RT combines the digital map of USC to generate path loss, and uses ML to fit the generated data to obtain a customized channel model. Different from the traditional channel model, the input of the customized channel model is the coordinates of the transmitter (TX) and the receiver (RX).
[0090] The present invention combines a customized channel and considers the scenario of UAV data collection. The UAV needs to collect data from user equipment (UE) randomly distributed on the ground. The objective of the present invention is to minimize the task completion time and find the optimal trajectory. To speed up the task completion, the present invention combines a new multiple access method, non-orthogonal multiple access (NOMA). NOMA has a higher communication rate than the traditional multiple access method, orthogonal multiple access (OMA). In addition, the present invention explores the optimal flight altitude of the UAV. Currently, the solutions to the optimization problem include convex optimization algorithms and deep reinforcement learning (DRL) algorithms. The DRL-based algorithm is superior to the traditional convex optimization and has a lower time complexity. The present invention designs an optimized proximal policy optimization (Optimized PPO) algorithm based on DRL to solve the optimization problem.
[0091] The present invention obtains a customized channel model in a fixed scenario through machine learning and ray tracing simulation, and applies this model to UAV data collection. Considering variables such as the UAV flight trajectory and signal transmission power, it optimizes the UAV data collection task completion time to minimize the time, and solves it through the reinforcement learning Optimized PPO algorithm. Compared with the prior art solutions, it has the following beneficial effects:
[0092] 1) The customized channel designed by combining ray tracing and machine learning can accurately model the channels in various different scenarios;
[0093] 2) A more reasonable and faster-converging reward function is designed;
[0094] 3) The convergence of the Optimized PPO algorithm is superior to that of the traditional PPO algorithm. Description of the Drawings
[0095] Figure 1 Schematically shows the customized channel modeling and simulation process of the method shown in the present invention;
[0096] Figure 2 Schematically shows the DNN channel model verification of the method shown in the present invention;
[0097] Figure 3 Schematically shows the UAV data collection optimization trajectory algorithm based on the PPO algorithm of the method shown in the present invention;
[0098] Figure 4 Schematically shows the convergence of the reward function in different cases of the method shown in the present invention, where (a) schematically shows the convergence in different multiple access methods and dimensions, and (b) schematically shows the convergence of the PPO algorithm using different techniques;
[0099] Figure 5 Schematically shows the UAV flight trajectories in different cases of the method shown in the present invention;
[0100] Figure 6 Illustrates the communication performance of the method shown in the present invention. Among them, (a) illustrates the number of completed UEs, and (b) illustrates the communication rate;
[0101] Figure 7 Illustrates the flowchart of a practical example of the method shown in the present invention. Detailed implementation manners
[0102] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0103] The following further elaborates on a method for optimizing the drone data collection trajectory based on a customized channel disclosed by the present invention in conjunction with the accompanying drawings and embodiments. This embodiment simulates the channels in the USC campus environment and conducts a drone collection task.
[0104] 1) Select the area to model the customized channel and export the three-dimensional topographic map of this area from OSM.
[0105] As Figure 1 shown, the selected area is the USC campus environment, and its size is 1 km × 1 km as shown in the satellite view. Set the drone flight area to 400 m × 200 m as shown by the red box in the digital map.
[0106] 2) Import the three-dimensional topographic map into the ray tracing simulation software WI.
[0107] As Figure 1 shown, export the electronic map of USC from OSM to WI, as shown in the digital map.
[0108] 3) Set the simulation parameters in WI: simulation frequency band, conductivity, permeability, transceiver position, etc.
[0109] The simulation frequency band is set to 3.5 GHz. The conductivity and permeability of different materials are shown in Table 1. The transmitter is placed 1.5 m above the ground, and a transmitter is placed every 10 m until the 400 m × 200 m area is covered. The horizontal position of the receiver is the same as that of the transmitter, and the height starts from 30 m, and a layer is placed every 10 m to simulate the drone flight height from 30 m to 120 m.
[0110] Table 1 Electromagnetic parameters of different materials
[0111] Material Permeability Conductivity (S / m) Concrete 7 0.015 Dry land 4 0.001 Grassland 2.4 0 Brick 4.44 0.001
[0112] 4) Obtain the simulation data path loss and perform data cleaning.
[0113] The number of TXs is 780, and there are 7800 RXs in total, distributed in layers every 10 m from 30 m to 120 m above the ground. Each simulation data set contains path loss, TX coordinates and RX coordinates (x uav , y uav , h uav ). There are seven elements. The path loss at the location where no signal is received in WI is 250 dB, which can be considered as abnormal data. After removing the abnormal data, there are a total of 5,445,285 groups of data.
[0114] 5) Set the hyperparameters of the DNN model: learning rate, number of neurons, depth of the neural network, batch size, etc.
[0115] The DNN neural network model we designed contains an input layer, three hidden layers with 128 neurons each, and an output layer. The activation function of the hidden layer is the rectified linear unit, the learning rate is 0.001, and the batch size is 256. The data is divided into a training set and a test set in a ratio of 8:2.
[0116] 6) Use the simulation data path loss to train the DNN model to obtain a customized channel.
[0117] Our goal is to use this data to train a model that can predict the path loss by inputting the TX and RX coordinates. The model can be expressed as where PL (dB) is the path loss, F DNN (·) represents the transfer function of the DNN, Figure 2 shows the prediction results of 1000 test sets. The red circles are the simulation values, and the green crosses are the prediction values of the DNN channel model. It can be seen that the accuracy is relatively high. We use the mean squared error (MSE) to measure the loss. The training loss is 27.40, and the test loss is 27.29.
[0118] 7) Model the UAV data collection problem as an optimization problem.
[0119] We consider an uplink communication model where a UAV serves N (N≥1) UEs. The UEs are randomly distributed in a given area and have some data waiting to be collected. The UAV takes off from a fixed starting point, and collecting all UE data within a specified time is regarded as successfully completing the task. Our goal is to minimize the task completion time. We use the sets N = {1, 2, …, N} and T = {1, 2, …, T} to index the UEs and time slots, and the UEs are considered stationary. The UAV can move freely in three-dimensional (3D) space. At time slot t, the positions of the UAV and the nth UE (denoted as UE n) are respectively represented by the vectors quav [t] = (x uav [t], y uav [t], h uav [t]), t ∈ T and description.
[0120] The UAV and the user equipment are restricted within a fixed area, such as Figure 1 shown. The coordinate boundaries of the UAV and the user equipment can be expressed as
[0121]
[0122] 30m ≤ h uav [t] ≤ 120m. (3)
[0123] UEn is randomly distributed on the ground and generates data The UAV takes off from a fixed starting point. At time slot t, the path loss between the UAV and UE n is NOMA is adopted for communication, which means that all UEs communicate at the same frequency. The UAV uses successive interference cancellation (SIC) to decode the signals from all UEs. For the uplink, the decoding order follows the descending order based on the channel gain of each UE, which means that the UAV preferentially decodes the stronger signals and treats the weaker signals as interference.
[0124] The UE list of the UAV at time slot t can be expressed as L[t] = {l1[t], l2[t], …, l ii [t]}, where l i [t] represents the index of the UE with the i-th largest path loss or the i-th smallest channel gain. Let the transmit power used for communication with the UAV at time slot t be expressed as which is subject to the following constraints
[0125]
[0126] where p max is the maximum transmit power of the UE. and the path loss (dB) between the UAV and at time slot t is expressed as The transmission rate at time slot t can be expressed as
[0127]
[0128] where B is the bandwidth, N0 is the power spectral density (PSD) of additive white Gaussian noise, is the coefficient indicating whether user l i [t] has completed data collection. Let Dn [t] represents the data transmission completed between the UAV and the UEn at time slot t, D n [t] can be expressed as
[0129]
[0130] where τ is the duration of each time period in the simulation, set to 1 second, R n [t′]. Then a n [t] can be expressed as
[0131]
[0132] where a n [t]=1 indicates that the UEn has not completed data collection at time slot t and needs to communicate with the UAV, a n [t]=0 means the opposite, represents the data that the user equipment n needs to collect. Completing the task can be expressed as
[0133]
[0134] where t end (t end ≤t max ) is the time for the UAV to collect all UE data, t max is the maximum working time of the UAV. If the working time exceeds t max , the task fails. The coordinates of the UAV can be expressed as
[0135]
[0136] where φ[t], ψ[t], v[t] represent the horizontal angle, elevation angle, and speed of the UAV respectively, and satisfy the following constraints
[0137]
[0138] where v max is the maximum speed of the UAV. Our goal is to minimize t end , which can be modeled as an optimization problem
[0139] min v[t],φ[t] t end
[0140] s.t. Eqs.(1)(2)(3)(4)(7)(8)(9)(10)(11)(12)(13)
[0141] According to Eq.(13), the UE power should take p max .
[0142] 8) Design the state, action, reward function, and hyperparameters of reinforcement learning.
[0143] a) State
[0144] The state is a vector containing environmental information. The state at time t can be represented as
[0145]
[0146] where N c [t] is the number of UEs that have completed data collection in time slot t. The elements in the state are normalized to [0, 1] before training. The state contains 5 + N elements.
[0147] b) Action
[0148] The action is what the agent might do after inputting the state into the environment. The action can be represented as
[0149]
[0150] where φ[t] ∈ (0, 2π], v[t] ∈ [0, v max are the horizontal angle, elevation angle, and speed of the UAV, respectively.
[0151] c) Reward function
[0152] The reward function consists of four parts. Our goal is to complete data collection, so the first part is designed as
[0153]
[0154] where r t 1 is the reward for the UAV to complete data collection, r t 1 is the sparse reward, and we design r t 2 to accelerate the convergence of r t 1 .
[0155]
[0156] where r t 2 represents the reward for each UE collected. To accelerate the convergence of r t 2 , we design r t 3 .
[0157]
[0158] where rt 3 is the total data collected in each time period. To prevent the drone from crossing the boundary, we design r t 4 as follows
[0159]
[0160] The reward function r t can be expressed as r t = r t 1 + r t 2 + r t 3 + r t 4 . λ1 to λ4 are constant parameters such that the rewards r t 1 to r t 4 are of the same order of magnitude.
[0161] 9) Use the Optimized PPO algorithm of reinforcement learning to solve.
[0162] PPO is an algorithm based on the policy gradient method. PPO is the same as the actor-critic (AC) algorithm and also includes two types of networks. The actor network is also called the policy network, which takes the input state s t , and outputs the action a t . We use θ and θ′ to represent the parameters of the trained and sampled policy networks respectively. We use π θ (a t |s t ) and π θ′ (a t |s t ) to represent the trained policy network and the sampled policy network respectively. The critic network takes the input state s t , and outputs the state value V μ (s t ), where μ is the parameter of the critic network. We use μ and μ′ to represent the parameters of the trained critic network and the sampled critic network respectively. We use V μ (s t ) and V μ′ (s t ) to represent the state values of the trained critic network and the sampled critic network respectively. The loss function of the policy network can be expressed as
[0163]
[0164] where, is the policy ratio, ε is the clipping factor, is the advantage estimation function, calculated by Generalized Advantage Estimation (GAE) as follows
[0165]
[0166] where γ ∈ [0, 1] is the reward discount factor, λ ∈ [0, 1], δ t = r t + γV μ (s t+1 ) - V μ (s t ), r t is the reward. The loss function of the critic network is We use some tricks to accelerate the convergence of PPO, which is called optimizing PPO. We describe these tricks as follows:
[0167] a) Advantage function normalization
[0168] After calculating in a certain batch using GAE, calculate the mean and standard deviation (STD) of all in the entire batch, and then subtract the mean and divide by the standard deviation (STD).
[0169] b) Reward function normalization
[0170] Reward function normalization is achieved by dynamically calculating the STD of the cumulative discounted sum of rewards and then dividing only the current reward by this standard deviation.
[0171] c) Learning rate decay
[0172] We adopt the method of linearly decaying the learning rate, making the learning rate linearly decrease from the initial value to 0 as the number of training steps increases, to make the training more stable.
[0173] d) Gradient clipping
[0174] Gradient clipping is introduced to prevent gradient explosion during the backpropagation of the policy and critic networks.
[0175] As Figure 3 shown, Algorithm 1 is the training process. We initialize the policy network π θ and the critic network V μ . Then we initialize the replay buffer B, and the training step e = 1. When e is less than the maximum number of training times e max , we initialize the positions of UAV and UE and set t = 1. We use π θ (a t |st ) Select a t . Then perform action a t , Move the UAV and calculate r t , Obtain the next state s t+1 , and store the transition (s t , a t , r t , s t+1 ) into B. If the number of transitions in B reaches the batch size, we sample a mini-batch of transitions from B and update θ and μ 10 times. If the UAV flies out of the boundary or has collected all UE data or t = t max , then this episode ends. Reset t to 1 and repeat the above process.
[0176] 10) Simulate and analyze the algorithm convergence degree and the UAV communication optimization performance.
[0177] a) Numerical simulation settings
[0178] Provide numerical results and analysis through simulation. We consider a system where 1 UAV serves 10 users. The users are uniformly distributed in a 200m×400m area. The take-off point of the UAV is (80m, 155m, 52.10m), above the roof of the tallest building in the area. t max is 100s; the carrier frequency is 3.5 GHz; the bandwidth is 10 MHz; the PSD of the noise N0 is 10 -17.4 W / Hz, the p of the user equipment max is 20 dBm; vmax is 15 m / s; the of all UEs is [10 Mb, 30 Mb]. The number of neurons in the policy and critic networks are (5 + N, 64, 64, 2) and (5 + N, 64, 64, 1) respectively. The batch size is 2024; the mini-batch is 64. The clipping factor ε, the discount factor γ and λ are 0.2, 0.99 and 0.95 respectively. The learning rates of the policy and critic networks are 0.001. (λ1, λ2, λ3, λ4) are (1, 2, 2, 5) respectively. The maximum training step e max is 3000000. Evaluate and observe the UAV trajectory and rewards every 5000 steps. The simulation is carried out using Python 3.8 and PyTorch 1.2.1. The hardware CPU and CPU use GTX4060 and R7-7945HX respectively.
[0179] b) Numerical simulation results and analysis
[0180] We compare the rewards of NOMA and OMA and the rewards of 3D and 2D trajectories. The height of the 2D trajectory is fixed at 52.10m. The transmission rate of OMA is Figure 4 (a) in it shows the relationship between the reward value and the number of training steps. The reward is set to From Figure 4 (a) in it, it can be seen that the reward of NOMA is greater than that of OMA. The reward of 3D is greater than that of 2D. The 3D deployment of the drone is more efficient than the 2D deployment.
[0181] To verify the effectiveness of the techniques, we conducted a comparative experiment with and without techniques in the NOMA 3D environment. As Figure 4 (b) in it shows that without using techniques a) and d), the training cannot proceed. Without using technique b), the reward will decrease. Without using technique c), the training is more unstable. We can conclude that all four techniques play a very crucial role.
[0182] In the final evaluation stage, we recorded the flight trajectory of the drone. Then, we compared the optimized PPO with the greedy solution. For the greedy solution, the drone flies to the nearest UE at a fixed altitude of 52.10m and v max with the maximum speed to collect data at this time. Figure 5 shows the flight trajectories in different situations. The 3D trajectory tends to fly lower to obtain a smaller path loss.
[0183] We also analyzed the variation of the number of completed UEs and the communication rate over time. Figure 6 (a) in it shows the variation of the number of completed UEs over time. NOMA 3D has the shortest completion time. 3D can always collect faster than 2D because in 3D, the drone tends to fly to a lower altitude to obtain a smaller path loss, and NOMA is more efficient than OMA. Finally, we can also see that the effect of PPO is much better than that of Greedy. Figure 6 (b) in it shows the communication rate over time. We can obtain the same conclusion as Figure 6 (a) in it.
[0184] c) Time complexity of the algorithm
[0185] For the optimized PPO, we use M and m i to represent the number of network layers and neurons in the i-th layer. The time complexity of the optimized PPO is The drone needs to fly to the nearest UE for the greedy solution, so the time complexity is O(N). For the DNN-based channel model, we use K and k j to represent the number of network layers and neurons in the j-th layer. The complexity of the common part is Therefore, the time complexities of the optimized PPO and Greedy are respectively and The training time of the optimized PPO is 5 hours. The test times of the optimized PPO and Greedy are 0.71s and 0.29s respectively.
[0186] The present invention proposes a customized channel model for UAV data collection and designs an optimized PPO algorithm to solve the UAV trajectory. The present invention proposes an optimization case based on a customized channel, and its flowchart is as Figure 7 shown. The present invention minimizes the task completion time, and the results are compared and analyzed using different multiple access methods and UAV trajectories (3D and 2D). The results show that NOMA is superior to OMA. The lower the UAV, the better the performance. In addition, the simulation results show that the techniques added in the PPO algorithm significantly improve the learning efficiency and convergence performance.
Claims
1. An optimization method for the UAV data collection trajectory based on a customized channel, characterized in that, It includes the following steps: Step 1: After selecting the area to model the customized channel, obtain the three-dimensional topographic map of this area; Step 2: Import the three-dimensional topographic map into the ray tracing simulation software, and set the simulation parameters in the ray tracing simulation software, including the simulation frequency band, conductivity, permeability, transmitter position, and receiver position; Step 3: Obtain the simulation data path loss and perform data cleaning, where each simulation data set contains path loss, TX coordinates and RX coordinates seven elements; Step 4: Set the hyperparameters of the DNN model; Step 5: Use the simulation data path loss to train the DNN model to obtain a customized channel model, and apply this customized channel model to UAV data collection; Step 6: Model the UAV data collection problem as an optimization problem, where: The uplink communication model of one UAV serving N user devices, N≥1, and the N user devices are randomly distributed in the area selected in Step 1, and there is some data waiting to be collected. The UAV takes off from a fixed starting point, and collecting all user device data within a specified time is regarded as successfully completing the task. Then the goal of the UAV data collection problem is to minimize the task completion time, and the optimization problem is expressed as: s.t. Eqs. (1)(2)(3)(4)(5)(6)(7)(8)(9)(10) In the above optimization problem, Eq. (1) is expressed as: Eq. (2) is expressed as: Eq. (3) is expressed as: 30m ≤ h uav [t] ≤ 120m Eq. (4) is expressed as: Eq. (5) is expressed as: Eq. (6) is expressed as: Eq. (7) is expressed as: Eq. (8) is expressed as: Eq. (9) is expressed as: Eq. (10) is expressed as: where: x uav [t] represents the X-axis coordinate of the UAV at time slot t; Represents the X-axis coordinate of the user device n; N is a set used to index user devices, N = {1, 2, …, N}; T is a set used to index time slots, T = {1, 2, …, T}; y uav [t] represents the Y-axis coordinate of the UAV at time slot t; Represents the Y-axis coordinate of the user equipment n; h uav [t] represents the altitude of the UAV at time slot t; l i [t] represents the index of the user equipment with the i-th largest path loss or the i-th smallest channel gain; User equipment l communicating with the UAV in time slot t i [t] Transmit power used; p max represents the maximum transmit power of the user equipment; a n [t]=1 indicates that user equipment n has not completed data collection in time slot t and needs to communicate with the UAV; a n [t]=0 indicates that user equipment n has completed data collection in time slot t and does not need to communicate with the UAV; D n [t] represents the data transmission completed between the UAV and user equipment n at time slot t; Indicates the data that the user equipment n needs to collect; t end The time for the drone to collect all user device data, t end ≤t max ,t max is the maximum working time of the drone. If the working time exceeds t max , the mission fails; q[t] is the coordinate of the UAV at time slot t; v[t] is the speed of the UAV at time slot t; is the horizontal angle of the UAV at time slot t; ψ[t] is the elevation angle of the UAV at time slot t; v max is the maximum speed of the UAV; Step 7: Design the state, action, reward function, and hyperparameters of the Reinforcement Learning Optimized PPO algorithm; Step 8: Use the Reinforcement Learning Optimized PPO algorithm to solve the optimization problem established in Step 6, and the minimized t is obtained through the solution end .
2. The method for optimizing the UAV data collection trajectory based on a customized channel according to claim 1, wherein, In Step 2: The transmitter is placed at a height of 1.5 m from the ground, and a transmitter is placed every 10 m until the area selected in Step 1 is covered: the horizontal position of the receiver is the same as that of the transmitter, and the height starts from 30 m and a layer is placed every 10 m; the simulated UAV flight height is from 30 m to 120 m.
3. The method for optimizing the UAV data collection trajectory based on a customized channel according to claim 1, characterized in that, In Step 4, the set DNN model hyperparameters include the learning rate, the number of neurons, the depth of the neural network, and the batch size.
4. The method for optimizing the UAV data collection trajectory based on a customized channel according to claim 1, wherein, In step 5, the customized channel model can predict the path loss, expressed as where PL(dB) is the path loss, and F DNN (·) represents the transfer function of the DNN.
5. The method for optimizing the drone data collection trajectory based on a customized channel according to claim 1, characterized in that, In step 6, D n [t] is expressed as: where τ is the duration of each time period, and R n [t′] represents the transmission rate of user equipment n at time slot t′.
6. The method for optimizing the drone data collection trajectory based on a customized channel according to claim 5, wherein User equipment l i [t] Transmission rate at time slot t Expressed as: where B is the bandwidth and N0 is the power spectral density of additive white Gaussian noise, denotes user l i is the coefficient indicating whether data collection is completed at [t], denotes user equipment l i is the path loss between [t] and the UAV at time slot t.
7. The method for optimizing the UAV data collection trajectory based on a customized channel according to claim 1, wherein In Step 7, the state at time slot t is expressed as: where N c [t] is the number of UAVs that complete data collection in time slot t; The action at time slot t is expressed as: The reward function at time slot t is expressed as: In the formula: is the reward for the UAV to complete data collection, Represents the rewards for each user device collected Total data collected for each time period λ1, λ2, λ3, λ4 are constant parameters such that the rewards are of the same order of magnitude.
8. The method for optimizing the UAV data collection trajectory based on a customized channel according to claim 1, wherein, In Step 8, use the following means to optimize the Reinforcement Learning Optimized PPO algorithm and accelerate the convergence of the Reinforcement Learning Optimized PPO algorithm: Mean 1) Advantage function normalization After calculating the advantage estimation function in a certain batch using generalized advantage estimation calculate the mean and standard deviation of all in the entire batch, and then subtract the mean and divide by the standard deviation; Mean 2) Reward function normalization Dynamically calculate the standard deviation of the cumulative discounted sum of rewards, and then divide the current reward only by this standard deviation; Mean 3) Learning rate decay Adopt the method of linearly decaying the learning rate, so that the learning rate linearly decreases from the initial value to 0 as the number of training steps increases, making the training more stable; Mean 4) Gradient clipping Gradient clipping is introduced to prevent gradient explosion during the backpropagation process of the critic network in the Optimized PPO algorithm for policy and reinforcement learning.
9. The method for optimizing the drone data collection trajectory based on a customized channel according to claim 1, wherein The advantage estimation function is expressed as: where: γ ∈ [0, 1], is the reward discount factor; λ ∈ [0, 1]; δ t = r t + γV μ (s t+1 ) - V μ (s t ), r t is the reward, V μ (s t ) is the output state value of the critic network, the parameter of the μ critic network.
Citation Information
Cited By
Unmanned aerial vehicle inspection method and equipment based on deep reinforcement learning, and medium
CN120512649A
A drone inspection method, equipment and medium based on deep reinforcement learning
CN120512649B