An Unmanned Aerial Vehicle (UAV) Trajectory Optimization Method Based on DDQN in a Data Acquisition Scenario
Through the DDQN-based drone trajectory optimization method, the problem of drone avoiding obstacles and interference in complex environments is solved, safe flight and efficient data acquisition are achieved, and the drone trajectory is optimized to improve data acquisition efficiency.
Patent Information
- Application Number
- CN202410175971.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-02-08
AI Technical Summary
The existing drone trajectory optimization methods fail to effectively consider the complexity of the three-dimensional environment, the impact of obstacle avoidance and external interference on signal quality in data acquisition scenarios, resulting in unsafe flight of the drone in complex environments and low data acquisition efficiency.
Using the DDQN-based drone trajectory optimization method, a communication scenario between the drone and ground equipment is constructed, an air-ground channel gain model is established, and the Markov decision-making process is standardized, and the drone trajectory is optimized through the DDQN algorithm, and a reward function is designed to achieve rapid convergence.
Implement safe flight and efficient data acquisition of drones in complex environments, shorten the mission completion time, improve data acquisition rate, adapt to environmental parameter changes, and reduce calculation complexity.
Smart Images

Figure CN117915375B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication countermeasure technologies, and particularly relates to a method for optimizing the trajectory of an unmanned aerial vehicle (UAV) based on Deep Deterministic Policy Gradient (DDQN) in a data collection scenario. Background Art
[0002] In recent years, using UAVs as airborne base stations to assist in offloading hotspots in existing terrestrial communication infrastructures and cellular networks has been considered a promising candidate technology. In some geographical regions, operators may be unable to build a complete cellular infrastructure, and UAVs can perform tasks such as collecting or transmitting data to ground Internet of Things (IoT) devices in specific areas, thereby reducing communication costs. However, with the continuous expansion of the scale and complexity of the IoT, UAVs still face some challenges while collecting data from IoT devices b, including limited flight time, limited mobility, and external malicious interference. Therefore, it is very important to study how to ensure that UAVs can efficiently perform data collection tasks in complex and realistic environments, while avoiding obstacles and interference and maintaining stable and excellent network performance.
[0003] With the rapid development of drone-assisted communication technology, traditional mathematical programming algorithms have achieved good results. S. Zhang achieved this goal by adopting convex optimization and graph theory techniques. The results show that the proposed technology shortens the task completion time and improves the signal-to-noise ratio (SNR) during the task. B. Li proposed an ant colony-based initial trajectory generation algorithm and an effective anti-collision scheme for drone flight trajectories. Q. Tang proposed a method to solve the non-convex problems of task allocation, power allocation, and drone flight trajectories in wireless communication services. However, it should be noted that the calculation time of these algorithms may increase exponentially with the increase in the scene scale and cannot fully adapt to the increasingly complex and scalable wireless network environment. The application of machine learning technology in drone communication has recently received renewed attention. Reinforcement learning is a model-free algorithm framework that has been proposed as an alternative to traditional algorithms. This method does not require modeling of specific environmental characteristic parameters and can train strategies through trial and error. It has practical significance for drone trajectory planning and wireless communication system optimization. Y. Wang optimized the drone trajectory and power allocation to maximize the fairness of throughput between sensor nodes. B. Zhang proposed a DRL-based framework that uses a convolutional neural network for feature extraction and the DQN algorithm for decision-making to design energy-efficient remote sensing routes for drones. H. Bayerlein studied a network model that utilizes the central layer and environmental information, processes the environmental layer through convolution, and the experimental results show that the efficiency of this algorithm is significantly higher. However, these studies only examined the flight trajectories of drones at a certain altitude, avoiding the complexity of the three-dimensional environment. Y. Zeng proposed a DRL method to minimize the task completion time of cellular-connected drones while maintaining good cellular network connectivity. B. Khamidehi proposed a double Q-learning method to solve the optimization problem involving drone trajectories under continuous time constraints. S. Yin coordinated between unmanned aerial vehicles to avoid collisions. The coordination was achieved through a sense-and-transmit protocol, and the main goal was to determine the optimal flight trajectory through a decentralized Q-learning algorithm, which shortened the convergence time and ensured the efficient transmission of sensor data. However, these studies did not consider the impact of obstacles and external interference attacks on signal quality and drone flight status.
[0004] Generally speaking, the existing drone trajectory optimization methods in data collection scenarios mainly have the following problems: 1) The model is not complete enough, avoiding the complexity of the action and state spaces of the drone's 3D trajectory; 2) Rarely considering how drones can accurately avoid obstacles to ensure safe flight in unknown urban environments; 3) Not considering the impact of malicious interference. Without prior wireless channel characteristics and the location of ground devices, drones must complete data collection from ground devices while considering external interference. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method for optimizing the trajectory of an unmanned aerial vehicle (UAV) based on Deep Double Q-Network (DDQN) in a data collection scenario, considering a more realistic environmental model in which multiple obstacles and jammers jointly affect the communication link between the UAV and ground Internet of Things (IoT) devices. By optimizing the three-dimensional flight trajectory of the UAV, the throughput is maximized, and data collection from the devices is completed while ensuring the flight safety of the UAV. Considering the limited computing power of the UAV, a trajectory optimization algorithm based on DDQN is proposed, and a corresponding reward function is set according to the scenario to achieve fast convergence of the algorithm. Compared with the traditional reinforcement learning algorithms Q-Learning and DQN, the present invention can significantly improve the algorithm convergence performance and data collection rate.
[0006] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for optimizing the trajectory of an unmanned aerial vehicle (UAV) based on DDQN in a data collection scenario, comprising:
[0008] Step S1, constructing a communication scenario between the UAV and ground devices;
[0009] Step S2, determining the UAV-airground channel gain based on the communication scenario and establishing a communication link model between the UAV and ground users;
[0010] Step S3, the UAV collects data and establishes a UAV trajectory optimization model in the data collection scenario;
[0011] Step S4, based on the communication link model, normalizing the optimization model into a Markov decision process;
[0012] Step S5, using a UAV trajectory optimization algorithm based on DDQN to solve the Markov decision process to obtain the optimal UAV trajectory.
[0013] To optimize the above technical solution, the specific measures taken further include:
[0014] The communication scenario between the UAV and ground devices constructed in the above step S1 includes a UAV, obstacles at different heights, U ground devices, and J directional jammers on the ground, where the UAV collects data from the U ground devices, with representing the set of ground devices, representing the set of jammers. The UAV mission completion time is discretized into N equal time intervals. At time n, the position of the UAV is the position of the ground device is the position of the jammer is is a vector space defined over the real number field.
[0015] The UAV-airground channel gain determined in the above step S2 is:
[0016]
[0017] where is the channel gain between the UAV and the ground device u within time n, is the distance between the UAV and the ground device u, ||·||2 is the L2 norm, q n , are the positions of the UAV and the ground device u respectively. Let z ∈ {LoS, NLoS} represent the line-of-sight link and non-line-of-sight link conditions, then β z represents the average channel gain at the reference distance d0 = 1m, and η z represents the shadow component modeled by the Gaussian distribution , and N is the number of equal time intervals into which the UAV mission completion time is discretized.
[0018] The communication link model between the UAV and the ground user established in the above step S2 under the channel gain model is:
[0019]
[0020]
[0021]
[0022]
[0023] where is the channel quality between the UAV and the ground device u, is the channel gain between the UAV and the ground device u within time n, P u is the transmit power of the ground device u, σ 2 is the white Gaussian noise power; INR n is the interference noise ratio of the jammer to the UAV within time n, P j is the transmit power of the jammer j, and J is the total number of jammers, is the signal-to-interference-plus-noise ratio between the ground device u and the UAV within time n, is the channel throughput between the ground device u and the UAV within time n, and B is the channel bandwidth; similar to , is the channel gain between the jammer and the UAV:
[0024]
[0025] In the above step S3, the communication between the UAV and the ground equipment adopts the TDMA strategy to collect data through the uplink channel, and according to the communication requirements of the UAV, a communication protocol between the UAV and the ground equipment is designed. The rule is that within each communication time n, the UAV only collects data from one ground equipment, and only the equipment with non-zero remaining data volume and the highest signal-to-noise ratio at the current time n can establish a communication link with the UAV. Specifically, is used to define the link state, means that the UAV has collected the data of device u at time n, means that the UAV has not collected the data of device u at time n.
[0026] The optimized UAV data collection model established in the above step S3 is:
[0027]
[0028] s.t.C1:
[0029] C2:b n >0,
[0030] C3: n∈[1,N]
[0031] C4: n∈[1,N]
[0032] where b n is the remaining power of the UAV, is the channel throughput between the ground equipment u and the UAV at time n, is the channel quality between the UAV and the ground equipment u, represents the obstacle area; γ th is the signal-to-noise ratio threshold; a n is the action selected by the UAV at time n; q n is the position of the UAV; N is the number of equal time intervals when the UAV task is completed and discretized; is the link state; represents the set of ground equipment, and U is the total number of ground equipment.
[0033] In the above step S4, the optimization model is specified as a Markov decision process represented by the quadruple ;
[0034] where represents the action space of the UAV, and the action selected by the UAV at time n where a x ,a y, a z ∈ {-1, 0, 1}; represents the state space of the UAV, and the state of the UAV at time n in this space is represented as s n = (s n,1 , s n,2 , s n,3 ), where s n,1 = {q n , b n , L n} represents the characteristics of the UAV, including the current instantaneous position q n of the unmanned aerial vehicle, the remaining battery power b n and the amount of data L n that has been collected; represents the characteristics of the UAV and ground equipment, represents the distance between the unmanned aerial vehicle and the equipment, is the channel quality between the UAV and the ground equipment u, is the link state, represents the remaining data volume of each device; s n,3 = {o n , o n+1} represents the observation space o n of the UAV at time n, and the observation space o n+1 predicted for the n + 1 time period in the case where the observable range of the UAV camera is relatively large; is the transition probability matrix, including the probability that when the UAV is in state s n , taking action a n transfers to the next state s n+1 ; is the reward information obtained by the UAV in the process of selecting action a n and reaching the next state s n+1 , and the reward function r n used is: r n = r n,1 - r n,2 - r n,3 , is the amount of data collected by the UAV in each time period, is the channel throughput between the ground equipment u and the UAV within time n, U is the total number of ground equipment, r n,2 is the power consumption penalty for the UAV's movement, and r n,3 is the collision penalty imposed when there are obstacles in the current observation space or the current position of the UAV exceeds the given area.
[0035] In the above step S5, the Markov decision process is solved using the DDQN-based UAV trajectory optimization algorithm to obtain the optimal UAV trajectory, including:
[0036] Step S51, initialize the position q of the drone n , randomly generate the position of the ground device and the position of the jammer Calculate the update the link state Determine the ground device selected by the drone in the current time slot;
[0037] Step S52, select an action a according to the ε-greedy strategy n , that is, randomly select an action a n with a probability of ε, and select an action according to the Q value with a probability of 1 - ε to ensure that the drone has a certain degree of exploration;
[0038] Step S53, input a n into the data collection environment to obtain the current state s n and the next position q of the drone n+1 , where q n+1 = q n+1 + a n ;
[0039] Step S54, the drone obtains r n according to the amount of data collected at the position q n,1 , obtains the power penalty r n,2 according to the movement amount of the drone, and obtains the collision penalty r n,3 according to whether the drone hits an obstacle in the current state and the next state, and finally obtains the total reward r n ;
[0040] Step S55, determine the next state s of the drone according to q n+1 ; n+1 ;
[0041] Step S56, save the transition result <s n+1 ,a n ,s n ,r n > to the experience pool . When the number of datasets in the experience pool reaches the threshold, select a sample set from the experience pool to train the neural network, and replace the old dataset in the experience pool with the newly obtained data during the training process;
[0042] Step S57, calculate the error function L(θ) = ||y n - Q(s n ,a n ; θ)|| 2 , where
[0043]
[0044] where y n is the target Q value; r n is the reward value for the n - time period; terminal represents the termination state; γ represents the discount factor, which mainly controls the influence of the new Q value on the neural network; represents finding the action corresponding to the maximum Q value in the current Q - network first;
[0045] Step S58, repeat steps S51 - S57. When the number of repetitions reaches the network update frequency N freq update the target network parameter θ' = θ, where θ is the current network parameter;
[0046] Step S59, when the maximum number of training times is reached, output the optimal UAV trajectory
[0047] The present invention has the following beneficial effects:
[0048] The present invention models the UAV trajectory optimization problem as an MDP in a three - dimensional environment. Different from the existing fixed - height two - dimensional scenarios, its action and state spaces are more complex, the solution space is larger, and it is more in line with the real scenario; the present invention considers the obstacle - avoidance problem of UAVs when performing data collection tasks in unknown complex urban environments. The UAV can sense jammers and obstacles in the environment in real time, make autonomous decisions to plan flight paths, resist interference, avoid obstacles, and ensure flight safety while completing data collection tasks; considering the limited computing power of UAVs, the present invention designs a trajectory optimization algorithm based on DDQN, sets corresponding reward values according to the scenario, and realizes the rapid convergence of the algorithm. The simulation results show that DDQN can well adapt to changes in environmental parameters, reduce the UAV task completion time, and has good convergence performance and short network training duration. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a scenario diagram of the UAV designed by the present invention for collecting data from ground equipment.
[0050] Figure 2 is the Markov process framework for UAV trajectory optimization in the data collection scenario designed by the present invention.
[0051] Figure 3 are the 3D diagram and top - view of the flight trajectory of the present invention using the proposed DDQN algorithm in different scenarios.
[0052] Figure 4 is the convergence comparison diagram of different reinforcement learning trajectory optimization methods of the present invention in different scenarios.
[0053] Figure 5This is a convergence comparison graph of the proposed DDQN trajectory optimization algorithm for the present invention in different scenarios.
[0054] Figure 6 This is a comparison graph of the convergence probability of the present invention under different algorithms.
[0055] Figure 7 This is a comparison graph of the data collection probability of the present invention under different algorithms. Detailed implementation manners
[0056] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0057] Although the steps in the present invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein relates to and encompasses any and all possible combinations of one or more of the associated listed items.
[0058] A DDQN-based UAV trajectory optimization method in a data collection scenario of the present invention includes:
[0059] Step S1, constructing a communication scenario between the UAV and ground devices.
[0060] Step S2, determining the UAV-airground channel gain based on the communication scenario and establishing a communication link model between the UAV and ground users;
[0061] Step S3, the UAV collects data and establishes a UAV trajectory optimization model in the data collection scenario;
[0062] Step S4, based on the communication link model, normalizing the optimization model into a Markov decision process;
[0063] Step S5, using a DDQN-based UAV trajectory optimization algorithm to solve the Markov decision process to obtain an optimal UAV trajectory.
[0064] In the embodiment, in step S1, the process of constructing a communication scenario between the UAV and ground devices includes:
[0065] It is assumed that a UAV is deployed in the environment to collect data from U ground devices within a specified area, and there are J directional jammers on the ground, where, represents the set of ground devices, represents the set of jammers.
[0066] Assume that the UAV mission completion time is discretized into N equal time intervals, and the position of the UAV at time n is The position of the ground device is The position of the jammer is is a vector space defined over the real number field. In addition, the urban environment also includes obstacles of different heights.
[0067] Figure 1 is a communication schematic diagram between the UAV and the ground device in the data acquisition scenario proposed by the present invention. Consider an urban scenario where a UAV is deployed in the environment, and high-rise buildings affect the communication link between the UAV and the ground device. There are U devices on the ground, and the UAV needs to communicate by accessing one of the devices. At the same time, there are J malicious jammers on the ground attempting to interfere with or block the communication between the UAV and the ground device.
[0068] In the embodiment, in step S2, to determine the UAV-airground channel gain and establish a communication link model between the UAV and the ground user, it includes:
[0069] Step S21, the channel gain model between the UAV and the device at time n is:
[0070]
[0071] where is the distance between the UAV and the ground device u, ||·||2 is the L2 norm, q n 、 are the positions of the UAV and the ground device u respectively. Let z ∈ {LoS, NLoS} represent the line-of-sight link and non-line-of-sight link conditions, then β z represents the average channel gain at the reference distance d0 = 1m, and η z represents the shadow component modeled by the Gaussian distribution
[0072] Step S22, under the channel gain model, establish a communication link model between the UAV and the ground user.
[0073]
[0074]
[0075]
[0076]
[0077] where, is the channel quality between the UAV and the ground device u, is the channel gain between the UAV and the ground device u within time n, P u is the transmit power of the ground device u, σ 2 is the white Gaussian noise power; INR n is the interference noise ratio of the jammer to the UAV within time n, P j is the transmit power of the jammer j, and J is the total number of jammers; is the signal-to-interference-plus-noise ratio between the ground device u and the UAV within time n, is the channel throughput between the ground device u and the UAV within time n, and B is the channel bandwidth; Similar to is the channel gain between the jammer and the UAV:
[0078]
[0079] In the embodiment, in step S3, the UAV uses the TDMA strategy to collect data through the uplink channel and establishes an optimization model for UAV data collection, including:
[0080] Step S31, it is set that the communication between the UAV and the ground device adopts the TDMA method. According to the communication requirements of the UAV, a communication protocol between the UAV and the ground device is designed, and its rules are:
[0081] Within each communication time n, the UAV can only collect data from one ground device. Only when the remaining data volume at the current time n is not zero and the signal-to-noise ratio is the highest can the device establish a communication link with the UAV. Specifically, it is defined by to define the link state, indicates that the UAV has collected the data of device u within time n, indicates that the UAV has not collected the data of device u within time n;
[0082] Step S32, the optimization goal is to optimize the flight trajectory of the UAV so as to collect data from the ground device to the maximum extent within the mission time. The UAV trajectory optimization problem (UAV data collection optimization model) in the data collection scenario can be expressed as:
[0083]
[0084] s.t.C1:
[0085] C2:b n >0,
[0086] C3: n∈[1,N]
[0087] C4: n∈[1,N]
[0088] where b n is the remaining power of the unmanned aerial vehicle, is the channel throughput between the ground device u and the UAV within time n, is the channel quality between the UAV and the ground device u, represents the obstacle area; γ th is the signal-to-noise ratio threshold; a n is the action selected by the UAV within time n; q n is the position of the UAV; N is the number of equal time intervals into which the UAV mission completion time is discretized; is the link state; represents the set of ground devices, and U is the total number of ground devices. In the embodiment, in step S4, based on the communication link model between the UAV and the ground device, with the UAV power (C2), the obstacle position (C1) and the given signal-to-noise ratio threshold (C3) as constraints, the optimization model is specified as a Markov decision process and transformed into an optimal decision problem. Specifically: the optimization model is specified as a Markov decision process represented by the quadruple ; where, represents the action space of the UAV, and the action selected by the UAV within time n where a x , a y , a z ∈{-1, 0, 1}; represents the state space of the UAV. The state of the UAV at time n in this space is represented as s n =(s n,1 , s n,2 , s n,3 ), where s n,1 ={q n , b n , L n} represents the characteristics of the UAV, including the current instantaneous position q n of the unmanned aerial vehicle, the remaining power b n and the amount of data collected L n ; represents the characteristics of the UAV and the ground device, represents the distance between the unmanned aerial vehicle and the device, is the channel quality between the UAV and the ground device u, is the link state, D n u represents the remaining data volume of each device; s n,3 ={o n , o n+1} represents the observation space o of the UAV at time n n, and the observation space o predicted for the n+1 time period when the observable range of the UAV camera is relatively large n+1 ; is the transition probability matrix, including the probability that when the UAV is in state s n , taking action a n transfers to the next state s n+1 ; is the reward information obtained by the UAV during the process of selecting action a n and reaching the next state s n+1 . The reward function r n is: r n = r n,1 - r n,2 - r n,3 , is the amount of data collected by the UAV in each time period, is the channel throughput between the ground device u and the UAV within time n. U is the total number of ground devices, r n,2 is the power consumption penalty for the UAV's movement, and r n,3 is the collision penalty imposed when there are obstacles in the current observation space or the current position of the UAV exceeds the given area.
[0089] Figure 2 is the Markov process framework for UAV trajectory optimization in the data collection scenario designed according to the present invention. The UAV, as an agent, interacts with the environment. The UAV inputs the current state s n information into the neural network. The network calculates the value of each action selected and guides the UAV to select action a n , causing the UAV to transfer to the next state s n+1 , and obtaining the reward value r n obtained by the UAV during this state transition process. The figure also details the generation process of the state space within any time slot [n, n+1], where processes ⑤ and ⑥ can be interchanged.
[0090] In the embodiment, in step S5, a UAV trajectory optimization algorithm framework based on DDQN is designed to solve the optimal decision-making problem, including:
[0091] Step S51, initialize the UAV position q n , randomly generate the ground device position and the jammer position Calculate the SINR between the UAV and the device Update the link state Determine the ground device selected by the UAV in the current time slot;
[0092] Step S52, select action a according to the ε-greedy policy n, that is, the probability of randomly selecting an action is ε, and the probability of selecting an action according to the Q value is 1 - ε, to ensure that the drone has a certain degree of exploration;
[0093] Step S53, input a n into the data collection environment to obtain the current state s n and the next position q of the drone n+1 , where q n+1 = q n+1 +a n ;
[0094] Step S54, the drone obtains r n according to the amount of data collected at position q n,1 , obtains the power penalty r n,2 according to the movement amount of the drone, and obtains the collision penalty r n,3 according to whether the drone hits an obstacle in the current state and the next state, and finally obtains the total reward r n .
[0095] Step S55, determine the next state s of the drone according to q n+1 . n+1 .
[0096] Step S56, save the transfer result <s n+1 , a n , s n , r n > into the experience pool . When the number of data sets in the experience pool reaches the threshold, select a sample set from the experience pool to train the neural network, and replace the old data set in the experience pool with the newly obtained data during the training process;
[0097] Step S57, calculate the error function L(θ) = ||y n - Q(s n , a n ; θ)|| 2 , where
[0098]
[0099] where, y n is the target Q value; r n is the reward value for the n time period; terminal represents the termination state; γ represents the discount factor, which mainly controls the influence of the new Q value on the neural network; represents finding the action corresponding to the maximum Q value in the current Q network first.
[0100] Step S58, repeat Steps S51 to S57. When the number of repetitions reaches the network update frequency N freq When it is, update the target network parameter θ' = θ, where θ is the current network parameter.
[0101] Step S59, when the maximum number of training times is reached, output the optimal UAV trajectory
[0102] Figure 3 The 3D map and top view of the UAV trajectory in different scenarios using the proposed algorithm are given. Among them, Figure (a) represents Scenario 1, that is, the number of ground devices is 3 and the number of obstacles is 5; Figure (b) represents Scenario 2, that is, the number of ground devices is 5 and the number of obstacles is 5; Figure (c) represents Scenario 3, that is, the number of ground devices is 3 and the number of obstacles is 8. It can be seen from the figure that the UAV will bypass the jammer and obstacles to minimize the interference during the data collection error and the possibility of the unmanned aerial vehicle crashing. In addition, the UAV flies towards the ground device to shorten the distance, thereby increasing the amount of data collected. The experimental results show that the overall trend of the UAV trajectory is to fly towards the optimization goal.
[0103] Figure 4 The convergence comparison diagrams of different reinforcement learning trajectory optimization methods in different scenarios are given. Among them, Figure (a) represents Scenario 1; Figure (b) represents Scenario 2; Figure (c) represents Scenario 3. The experimental results show that in terms of the final average reward, the Dueling DQN algorithm and the DDQN algorithm are not much different, but the Q-Learning algorithm only reaches 10% of the DDQN algorithm, and the DQN algorithm only reaches 50% of the DDQN, which shows that the proposed DDQN trajectory optimization method has obvious superiority in convergence performance.
[0104] Figure 5 The convergence comparison diagrams of the proposed DDQN trajectory optimization algorithm of the present invention in different scenarios are given. The simulation results are as follows: 1) Compared with Scenario 1, Scenario 2 increases the number of ground devices, the total data volume increases, and a higher average reward value is obtained; 2) Compared with Scenario 1, Scenario 2 increases the density of obstacles, which results in more penalties for the UAV, and the average reward value is lower than that of Scenario 3.
[0105] Figure 6 The convergence probability comparison diagrams of the present invention under different reinforcement learning algorithms are given. Taking the 90% confidence interval of the average reward value of each algorithm as the convergence interval, calculate the convergence rate of each algorithm in 3 scenarios, and then calculate the average convergence rate of each algorithm in the three scenarios respectively. The experimental results show that on the premise of maintaining the optimal average reward value, the convergence rates of DDQN and Dueling DQN are better, the Q-learning algorithm has a fast convergence rate, but the average reward value is the lowest and the convergence performance is poor.
[0106] Figure 7Figure 0 shows the comparison chart of data collection probabilities of the present invention under different algorithms. Taking Scenario 1 as an example, after the network is trained, the initial position of the UAV, the positions of the ground equipment and the jammer are fixed, and the data collection process of the UAV under different algorithms is observed. It can be seen from the figure that the Q-Learning algorithm is far less efficient than the Dueling DQN and DDQN algorithms in terms of data collection speed.
[0107] Table 1 shows the average training time of different reinforcement learning algorithms of the present invention in different scenarios. The experimental results show that the DDQN trajectory optimization algorithm proposed by the present invention has a shorter network training time while ensuring the convergence performance. Although the convergence performance of the Dueling DQN algorithm is not much different from that of the DDQN algorithm in the previous experiments, it is time-consuming in the network training stage. In addition, the execution time of the adopted algorithm on the system is much less than the training time, so the real-time requirement can be met.
[0108] Table 1 Average training time (seconds / iteration) of different reinforcement learning algorithms of the present invention in different scenarios
[0109] Algorithm Scenario 1 Scenario 2 Scenario 3 Q-Learning 1.34460 1.28929 1.39889 DQN 1.75778 4.26592 1.80001 Dueling DQN 2.05745 4.35702 2.13843 DDQN 1.63612 3.95510 1.71541
[0110] The experimental results show that the proposed DDQN trajectory optimization algorithm can better learn the unknown environment, collect data with the optimal reward value and the shortest number of steps to achieve the optimization goal.
[0111] In summary, the trajectory optimization method of the present invention has the following advantages:
[0112] 1) In the present invention, the UAV trajectory optimization problem is modeled as an MDP in a three-dimensional environment. Different from the existing two-dimensional scenarios with a fixed height, its action and state spaces are more complex, the solution space is larger, and it is closer to the real scenario.
[0113] 2) In an unknown environment, the positions of the jammer and the obstacles are randomly generated in a given area. The UAV can sense the jammer and the obstacles in the environment in real time, make autonomous decisions to plan the flight path, resist interference and avoid obstacles, and ensure flight safety while completing the data collection task.
[0114] 3) Considering the limited computing power of the UAV, the present invention designs a DDQN-based trajectory optimization algorithm, sets the corresponding reward value according to the scenario, and realizes fast convergence. The simulation results show that DDQN can well adapt to the changes of environmental parameters, has good convergence performance and a short training time.
[0115] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0116] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for optimizing the trajectory of an unmanned aerial vehicle based on DDQN in a data acquisition scenario, characterized in that, Including: Step S1, constructing a communication scenario between the drone and the ground device; Step S2, determining the drone-airground channel gain based on the communication scenario and establishing a communication link model between the drone and the ground user; Step S3, the drone collects data and establishes a drone trajectory optimization model in the data collection scenario; Step S4, based on the communication link model, normalizing the optimization model into a Markov decision process; Step S5, using a drone trajectory optimization algorithm based on DDQN to solve the Markov decision process to obtain the optimal drone trajectory; The drone-airground channel gain determined in Step S2 is: where, is the channel gain between the UAV and the ground device u within time n, is the distance between the UAV and the ground device u, ||·||2 is the L2 norm, q n 、 are the positions of the UAV and the ground device u; let z ∈ {LoS, NLoS} denote the line-of-sight link and non-line-of-sight link conditions, then β z represents the average channel gain at the reference distance d0 = 1m, η z represents the shadowing component of the model with a Gaussian distribution σ 2 is the white Gaussian noise power, and N is the number of equal time intervals into which the UAV mission completion time is discretized.
2. The method for optimizing the trajectory of an unmanned aerial vehicle based on DDQN in a data acquisition scenario according to claim 1, wherein The communication scenario between the drone and the ground devices constructed in step S1 includes a drone, obstacles at different heights, U ground devices, and J directional jammers on the ground. The drone collects data from the U ground devices to represent the set of ground devices, represent the set of jammers. The mission completion time of the drone is discretized into N equal time intervals. At time n, the position of the drone is the position of the ground device is the position of the jammer is is a vector space defined over the real number field.
3. A method for optimizing the trajectory of an unmanned aerial vehicle based on DDQN in a data acquisition scenario according to claim 1, wherein The communication link model between the drone and the ground user established in Step S2 under the channel gain model is: Among them, is the channel quality between the UAV and the ground device u, is the channel gain between the UAV and the ground device u within time n, P u is the transmit power of the ground device u, σ 2 is the white Gaussian noise power; INR n is the interference noise ratio of the jammer to the UAV within time n, P j is the transmit power of the jammer j, J is the total number of jammers, is the signal-to-interference-plus-noise ratio between the ground device u and the UAV within time n, is the channel throughput between the ground device u and the UAV within time n, B is the channel bandwidth; is the channel gain between the jammer and the UAV.
4. A method for optimizing the trajectory of an unmanned aerial vehicle based on DDQN in a data acquisition scenario according to claim 1, wherein In Step S3, the communication between the drone and the ground device uses the TDMA strategy to collect data through the uplink channel, and according to the communication requirements of the drone, a communication protocol between the drone and the ground device is designed, and its rules are: Within each communication time n, the UAV only collects data from one ground device. Only the device with non-zero remaining data volume and the highest signal-to-noise ratio at the current time n can establish a communication link with the UAV. Specifically, is used to define the link state. indicates that the UAV has collected data from device u at time n. indicates that the UAV has not collected data from device u at time n.
5. A method for optimizing the trajectory of an unmanned aerial vehicle based on DDQN in a data collection scenario according to claim 1, characterized in that, The drone data collection optimization model established in Step S3 is: where b n is the remaining battery power of the UAV; is the channel throughput between the ground device u and the UAV within time n; represents the obstacle area; γ th is the signal-to-noise ratio threshold; a n is the action selected by the UAV within time n; q n is the position of the UAV; N is the number of equal time intervals into which the UAV mission completion time is discretized; is the link state; represents the set of ground devices, and U is the total number of ground devices.
6. The method for optimizing the trajectory of an unmanned aerial vehicle based on DDQN in a data acquisition scenario according to claim 1, wherein In the step S4, the optimized model is specified as a Markov decision process represented by a quadruple ; where represents the action space of the UAV, and the action selected by the UAV at time n is where a x , a y , a z ∈{-1, 0, 1}; represents the state space of the UAV. The state of the UAV at time n in this space is represented as s n =(s n,1 , s n,2 , s n,3 ), where s n,1 ={q n , b n , L n} represents the characteristics of the UAV, including the current instantaneous position q n of the UAV, the remaining battery power b n and the amount of data collected L n ; represents the characteristics of the UAV and ground equipment, represents the distance between the UAV and the equipment, is the channel quality between the UAV and the ground equipment u, is the link state, represents the remaining data volume of each device; s n,3 ={o n , o n+1} represents the observation space o n of the UAV at time n, and the observation space o n+1 predicted in the case where the observable range of the UAV camera is relatively large for the n+1 time period; is the transition probability matrix, including the probability that when the UAV is in state s n , taking action a n transfers to the next state s n+1 ; is the reward information obtained by the UAV during the process of selecting action a n and reaching the next state s n+1 . The reward function r n used is: r n =r n,1 -r n,2 -r n,3 , is the amount of data collected by the UAV in each time period, is the channel throughput between the ground equipment u and the UAV at time n, U is the total number of ground equipment, r n,2 is the power consumption penalty for the UAV's movement, r n,3 The collision penalty applied when there are obstacles in the current observation space or the current position of the UAV exceeds the given area.
7. A method for optimizing the trajectory of an unmanned aerial vehicle based on DDQN in a data acquisition scenario according to claim 6, characterized in that, In Step S5, using a drone trajectory optimization algorithm based on DDQN to solve the Markov decision process to obtain the optimal drone trajectory, including: Step S51, initialize the position q of the UAV n , randomly generate the positions of the ground devices and the jammer Calculate the Update the link status Determine the ground device selected by the UAV in the current time slot; Step S52, select action a according to the ε-greedy policy n , that is, randomly select action a n with probability ε, and select according to the Q-value The probability of the action is 1 - ε to ensure that the drone has a certain degree of exploration; Step S53, input a n into the data collection environment of the input, to obtain the current state s n and the next position q of the drone n+1 , where q n+1 = q n+1 + a n ; Step S54, the drone obtains r according to the amount of data collected at position q n and obtains r according to the amount of movement of the drone n,1 . It obtains an electricity penalty r according to the amount of movement of the drone n,2 and obtains a collision penalty r based on whether the drone hits an obstacle in the current state and the next state n,3 . Finally, it obtains the total reward r n ; Step S55, determine the next state s of the UAV according to q n+1 ; n+1 Step S56, save the transfer result <s n+1 ,a n ,s n ,r n > into the experience pool . When the number of data sets in the experience pool reaches the threshold, select a sample set from the experience pool to train the neural network, and replace the old data set in the experience pool with the newly obtained data during the training process; Step S57, calculate the error function \(L(\theta)=\|y n -Q(s n ,a n ;\theta)\| 2 , where where y n is the target Q value; r n is the reward value for the n time period; terminal represents the termination state; γ represents the discount factor, which controls the influence of the new Q value on the neural network; represents finding the action corresponding to the maximum Q value in the current Q network first; Step S58, repeat steps S51 - S57. When the repetition count reaches the network update frequency N freq , update the target network parameter θ' = θ, where θ is the current network parameter; Step S59, when the maximum number of training times is reached, output the optimal UAV trajectory
Citation Information
Patent Citations
Unmanned aerial vehicle trajectory optimization method and system in Internet of Things data collection
CN113382060A
High-reliability short packet transmission method and device for unmanned aerial vehicle assisted backscatter communication
CN117082539A
Unmanned aerial vehicle cluster deployment and trajectory planning method based on reinforcement learning
CN117270559A