Unmanned aerial vehicle obstacle avoidance and navigation method based on lightweight breach dual-depth Q network
By defining the problems of obstacle avoidance and autonomous navigation as Markov decision-making process, deep neural networks with Dueling architecture and lightweight pruning technology are used to solve the problems of low learning efficiency and reduced endurance in complex environments, and efficient and stable obstacle avoidance and navigation performance are achieved.
Patent Information
- Application Number
- CN202510538649.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-18
AI Technical Summary
The existing drone obstacle avoidance and autonomous navigation technologies are incomplex environments with low learning efficiency and slow convergence speed, and high computing consumption of complex neural networks lead to a reduced battery life.
The problems of drone obstacle avoidance and autonomous navigation are defined as Markov's decision-making process, deep neural networks of Dueling architecture are adopted, and the model is optimized through lightweight pruning technology, combined with priority experience playback technology to improve learning efficiency and model lightweighting, design state space, action space and reward functions, and obtain drone status in real time for action selection.
It improves the learning efficiency and convergence speed of drone obstacle avoidance and autonomous navigation, reduces the computational complexity, improves battery life, and is suitable for practical applications in resource-constrained environments.
Smart Images

Figure CN120335488A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV navigation, and particularly to a UAV obstacle avoidance and navigation method based on a lightweight duel double deep Q-network. Background Art
[0002] The purpose of UAV autonomous navigation technology is to enable the UAV to autonomously reach the destination without human intervention. Since the UAV flies at a relatively high speed, strict guarantees are required for the response speed and accuracy of various decisions. For different scenarios, the decisions adopted are also different. With the development of technology, the existing UAV autonomous navigation methods can be divided into two categories: traditional path planning methods and deep learning-based path planning methods. Among them, the traditional method plans an optimal or sub-optimal path from the starting point to the target point according to the flight characteristics and flight environment of the UAV. In this process, the UAV needs to avoid obstacles, meet performance constraints, and optimize flight efficiency as much as possible, such as the artificial potential field method, the voronoi diagram method, etc. These methods are suitable for deployment in simple and structured environments. However, in complex real scenarios, due to the difficulty of accurately describing parameters such as the shape, position, and motion state of obstacles, these methods are prone to problems such as local optimal solutions, getting stuck in dead ends, or being unable to effectively plan safe paths in obstacle-dense areas. Currently, the methods to solve these problems include technical methods based on genetic algorithms, dynamic Bayesian networks, approximate dynamic programming, etc. However, the deployment of these methods requires complex modeling and a large amount of data sets as support, which makes the system calculation volume extremely large, greatly reducing the real-time decision-making efficiency, and then resulting in consequences such as slow decision-making during the high-speed flight of the UAV, unstable decision-making effects, and unacceptable training costs, bringing great difficulties to actual deployment.
[0003] By using the method of deep reinforcement learning, models such as DQN (Deep Q-network), SAC (Soft Actor-Critic), DDQN (Double Deep Q-Network), PPO (Proximal Policy Optimization), and DDGP (Deep Deterministic Policy Gradient) are used to respond to action requests with multiple basic actions. Both the DQN and DDGP methods use the Q-network to evaluate and select actions. When there is a lot of noise in the environment, more value estimations will be generated during the calculation, which is called the overestimation problem. This will have a greater impact on flight decisions. The existing methods are difficult to guarantee the efficiency of real-time decision-making while improving the overestimation problems of the two, resulting in situations such as uneven flight paths and path deadlocks of the UAV.
[0004] Previously, Xuefan Zhang et al. proposed an autonomous path planning method for drones based on the lightweight continuous SAC algorithm. By combining the state space, action space, and reward function of the drone, the smoothness of the drone's path was improved. A neural network was constructed to overcome the influence of noise on the flight state of the drone. Finally, model distillation was used to improve the response speed of the model. However, this method is still difficult to clearly model objects in complex environments. At the same time, the training process of this method is relatively complex and requires a large amount of data and computing resources to optimize model parameters. In addition, the stability of the algorithm is also an issue that needs to be considered. In practical applications, it may lead to a decline in algorithm performance.
[0005] Meng Ziyang from Tsinghua University disclosed a method of pre-training a deep reinforcement learning network using Lidar data and deploying it to a drone in the patent document "Lightweight Autonomous Navigation Method and Device Based on Neural Network Driving" (application number: 202410102883.0, application date: January 24, 2024, publication number: CN 117906614A). Although it reduces problems such as the computational complexity and memory occupancy of the neural network algorithm, the deep learning method still requires a large amount of data for calculation. When deployed on a drone platform with limited computing resources and energy, it still consumes a large amount of computing resources and energy, significantly reducing the endurance of the drone, while increasing the hardware cost and the difficulty of heat dissipation. Summary of the Invention
[0006] The object of the present invention is to solve the problems in the current technology, such as slow training speed, low completion rate of obstacle avoidance tasks, and reduced endurance of drones caused by high computational consumption of complex neural networks, and to propose a method for drone obstacle avoidance and navigation based on a lightweight dueling double deep Q network.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for drone obstacle avoidance and navigation based on a lightweight dueling double deep Q network, including the following steps; S1, Define the drone obstacle avoidance and autonomous navigation problem as a Markov decision process, and establish an MDP five-tuple MDP = (S, A, P, R, γ), where S is the state space, A is the action space, P is the state transition function, R is the reward function, and γ is the discount factor; S2, Construct an interaction framework between the environment and the drone in reinforcement learning, and design the state space, action space, and reward function; The state space S is composed of multiple states, and each state includes three consecutive depth maps, which are obtained by the drone in the Airsim simulation platform; The action space A includes multiple actions of the drone; The reward function R includes a result reward function, a distance reward function, and an action execution reward function; S3. Build a deep neural network based on the Dueling architecture, including a main network and a target network with the same structure; S4. Set up an experience pool for storing multiple pieces of experience data. The experience data is a five-tuple (s, a, r, s', f), where s is the current state, a is the current action, r is the reward for the current action, s' is the next state, and f indicates whether the current round of training is completed; S6. Use the D3QN-PER algorithm to train the deep neural network to obtain a trained main network and target network; S8. Use neural network pruning technology to lightweight the trained main network to obtain a lightweight model; S7. Real-time obtain the state of the drone during flight and select actions for execution using the lightweight model.
[0008] Preferably: In S1, the discount factor γ ∈ (0, 1), and the depth map is preprocessed.
[0009] Preferably: The action space includes 5 actions, namely; Action 0: Move along the positive x-axis at a speed of 1 m / s for 2 s; Action 1: Move along the negative x-axis at a speed of 1 m / s for 2 s; Action 2: Move along the positive y-axis at a speed of 1 m / s for 2 s; Action 3: Move along the negative y-axis at a speed of 1 m / s for 2 s; Action 4: Hover at the current position for 2 s.
[0010] Preferably: The reward function R is calculated according to the following formula; , , , , In the formula, r res , r d , r a respectively represent the result reward function, the distance reward function, and the action execution reward function, ω d , ω a respectively represent the distance reward function coefficient and the action execution reward function coefficient, D is the Euclidean distance in three-dimensional space between the current position of the drone and the target location, and threshold is a threshold used to specify that within this threshold range from the target point is considered to have reached the target point; D nextRepresents the Euclidean distance in three-dimensional space between the next position of the UAV and the target location; action = 0 to action = 4 are actions 0 to 4 in the action space respectively.
[0011] Preferably: The main network is based on the Dueling DQN network structure, including a feature extraction layer, whose output ends are respectively connected to a state value function layer and an advantage function layer; The feature extraction layer is a ResNet50 network, and its last fully connected layer is deleted; The state value function layer is used to output the state value V(s), and the advantage function layer is used to output the advantage function A(s,a); The input of the main network is three consecutive frames of depth maps captured by the UAV in the Airsim simulation platform, and the output is the Q value Q(s,a) of the UAV in state s and action a. Q(s,a)=V(s)+A(s,a)-mean(A(s,a)), where mean( ) is the mean function.
[0012] Preferably: S6 specifically includes S61~S63; S61, Use DepGraph to perform unstructured pruning on the Q network, with a pruning ratio of 10%; S62, Fine-tune the model obtained in S61; S63, Repeat steps S61~S62 five times to complete 50% pruning.
[0013] Compared with the prior art, the advantages of the present invention are as follows: (1) The present invention defines the UAV obstacle avoidance and autonomous navigation problem as a Markov decision process, designs a state space, a reward function, etc., and aims to improve the DQN model by adopting the Dueling architecture. In this method, the state is composed of three consecutive frames of preprocessed depth maps merged together, which can strengthen the dynamic perception of the environment. The reward function includes a result reward function, a distance reward function, and an action execution reward function. At the same time, the prioritized experience replay technique is used when extracting experiences, which effectively alleviates the sparse reward problem in reinforcement learning and improves the learning efficiency and convergence speed of UAV obstacle avoidance and autonomous navigation. This method can solve the problems existing in the current technology, such as low learning efficiency, slow convergence speed, and being prone to falling into local optimal solutions and being unable to plan safe routes in areas with dense obstacles.
[0014] (2)Prune the trained main network to achieve lightweight network architecture. The advantage of this method is that it can effectively reduce the computational complexity of the model, thus avoiding affecting the endurance performance of the drone due to excessive computational overhead. The pruned model has reduced computational complexity and storage overhead, while still being able to maintain a high obstacle avoidance task completion rate and improve the inference speed. This method improves the deployment efficiency of the drone obstacle avoidance algorithm in resource-constrained environments (such as embedded devices, edge computing platforms), making it more efficient and stable in practical applications. Description of the Drawings
[0015] Figure 1 Flowchart of the present invention; Figure 2 Schematic diagram of the main network structure; Figure 3 Schematic diagram of training a deep neural network using the D3QN-PER algorithm; Figure 4 A simulation environment built based on a forest scene in Embodiment 3; Figure 5 Experimental comparison chart of the reward function convergence curve; Figure 6 Experimental comparison chart of the task completion rate convergence curve. Detailed Implementation Manner
[0016] The present invention will be further described below in conjunction with the embodiments and the drawings.
[0017] Embodiment 1: Refer to Figures 1 to 3 , a drone obstacle avoidance and navigation method based on a lightweight dueling double deep Q-network, including the following steps; S1, Define the drone obstacle avoidance and autonomous navigation problem as a Markov decision process, and establish an MDP five-tuple MDP=(S, A, P, R, γ), where S is the state space, A is the action space, P is the state transition function, R is the reward function, and γ is the discount factor; S2, Construct the interaction framework between the environment and the drone in reinforcement learning, and design the state space, action space, and reward function; The state space S is composed of multiple states, and each state includes three consecutive frames of depth maps, which are obtained by the drone in the Airsim simulation platform; The action space A includes multiple actions of the drone; The reward function R includes a result reward function, a distance reward function, and an action execution reward function; S3, Build a deep neural network based on the Dueling architecture, including a main network and a target network with the same structure; S4. Set up an experience pool to store multiple pieces of experience data. The experience data is a five-tuple (s, a, r, s', f), where s is the current state, a is the current action, r is the reward for the current action, s' is the next state, and f indicates whether the current round of training is completed. S5. Train a deep neural network using the D3QN-PER algorithm to obtain a trained main network and target network. S6. Use neural network pruning technology to lightweight the trained main network to obtain a lightweight model. S7. Real-time obtain the state of the drone during flight and select actions for execution using the lightweight model.
[0018] In this embodiment, in S1, the discount factor γ ∈ (0, 1), and the depth map is preprocessed.
[0019] The action space includes 5 actions, namely: Action 0: Move along the positive x-axis at a speed of 1 m / s for 2 s. Action 1: Move along the negative x-axis at a speed of 1 m / s for 2 s. Action 2: Move along the positive y-axis at a speed of 1 m / s for 2 s. Action 3: Move along the negative y-axis at a speed of 1 m / s for 2 s. Action 4: Hover at the current position for 2 s.
[0020] The reward function R is calculated according to the following formula; , , , , In the formula, r res , r d , r a respectively represent the result reward function, distance reward function, and action execution reward function. ω d , ω a respectively represent the distance reward function coefficient and action execution reward function coefficient. D is the Euclidean distance in three-dimensional space between the current position of the drone and the target location. threshold is a threshold used to specify that within this threshold range from the target point is considered to have reached the target point; D next represents the Euclidean distance in three-dimensional space between the next position of the drone and the target location; action = 0 to action = 4 are actions 0 to 4 in the action space.
[0021] The main network is based on the Dueling DQN network structure and includes a feature extraction layer, whose output ends are respectively connected to a state value function layer and an advantage function layer; The feature extraction layer is a ResNet50 network, and its last fully connected layer is removed; The state value function layer is used to output the state value V(s), and the advantage function layer is used to output the advantage function A(s,a); The input of the main network is three consecutive depth maps captured by the drone in the Airsim simulation platform, and the output is the Q value Q(s,a) when the drone is in state s and action a. Q(s,a)=V(s)+A(s,a)-mean(A(s,a)), where mean( ) is the mean function.
[0022] S6 specifically includes S61 to S63; S61, use DepGraph to perform unstructured pruning on the Q network, with a pruning ratio of 10%; S62, fine-tune the model obtained in S61; S63, repeat steps S61 to S62 five times to complete 50% pruning.
[0023] Regarding the preprocessing of the depth map: The method for preprocessing the depth map obtained by the drone in the Airsim simulation platform is as follows: First, convert it into a depth map and perform inversion to enhance the recognition of nearby obstacles. Then, crop the image into a shape suitable for input to the network. Finally, merge three consecutive preprocessed depth maps as the current state to strengthen the dynamic perception of the environment.
[0024] Regarding the discount factor γ: γ ∈ (0,1), which is used to measure the influence degree of future rewards. The larger the discount factor, the more attention is paid to long-term benefits. The goal of D3QN is to maximize the cumulative discounted reward, that is .
[0025] Example 2: Refer to Figure 3 , based on Example 1, a specific implementation process of step S5 is given: S5, train the deep neural network with the D3QN-PER algorithm to obtain the trained main network and target network, including steps S51 to S55; S51, obtain the depth map of the drone in the external environment or the Airsim simulation platform, and generate states at different time steps after preprocessing; S52. Since the experience data in the experience pool is initially scarce and insufficient for batch training, the exploration rate epsilon is set to 1 at the beginning of training and gradually decreased to enhance the exploration of the UAV in different state spaces at the initial stage. Only experience data is collected in this stage, and no model optimization is performed. The experience data is stored in the form of a five-tuple (s, a, r, s', f), where: s is the current state, represented by three consecutive preprocessed depth maps; a is the executed action, r is the reward value calculated according to the set reward function; s' is the next state, also represented by three consecutive depth maps; f is used to indicate whether the current episode ends. When the UAV successfully reaches the target point or collides, f = 1, otherwise f = 0. S53. When the number of experiences in the experience pool reaches the threshold, model optimization begins: 64 pieces of experience data with higher weights are sampled from the experience pool using Prioritized Experience Replay (PER), and the next states in each piece of data are combined into a tensor as the input to the main network. S54. The combined tensor is input into the main network to obtain the Q-values of all actions, and the action with the largest Q-value is selected; the Q-value (target Q-value) of the selected action is evaluated using the target network, and the temporal difference error is calculated using this Q-value and the rewards in the 64 pieces of experience data. S55. The calculated temporal difference error and the output value (predicted Q-value) of the main network are used to calculate the loss function and perform gradient descent to optimize the main network.
[0026] Example 3: A simulation environment was built based on a forest scenario for this experiment, and its overall environment is as Figure 4 shown. Figure 4 As shown in the figure, the depth map being processed in real time is displayed in the green box at the lower left corner, and the real-time RGB scene map captured by the UAV's front camera is displayed in the purple box at the lower right corner. In this environment, a comparative experiment on the UAV obstacle avoidance flight task of the method of the present invention, the original D3QN algorithm, and the DQN algorithm was carried out, and the performance evaluation indicators were selected as the convergence of the reward function and the convergence of the task completion rate.
[0027] (1) Convergence of the reward function: The convergence of the reward function is an important basis for evaluating the stability of the reinforcement learning training process and the quality of the policy. Its final convergence value reflects the stable performance of the policy in the later stage of training, and the speed required for convergence reflects the learning efficiency and optimization ability of the model.
[0028] (2) Convergence of the task completion rate: It includes the convergence value of the task completion rate and the convergence speed. The task completion rate is the core indicator for measuring the obstacle avoidance performance of the drone, which is defined as the proportion of the drone successfully avoiding obstacles and reaching the target point smoothly after a certain number of training rounds. Its convergence value represents the effectiveness of the final strategy in the obstacle avoidance task, while the convergence speed reflects the efficiency of the strategy in achieving stable performance during the training process.
[0029] The experimental parameters are set as shown in Table 1: Table 1. Experimental Parameter Settings Parameter name Parameter value Learning rate 0.0001 Discount factor 0.99 Number of training rounds 20000 Experience pool capacity 10000 Number of sampled extractions 64 The experimental results are respectively as Figure 5 and Figure 6 shown, Figure 5 and Figure 6 Among them, DQN, D3QN (Original), and D3QN-PER (ours) refer to the DQN algorithm, the original D3QN algorithm, and the method of the present invention respectively.
[0030] From Figure 5 the experimental results, it can be seen that in terms of the convergence speed, the reward value of the DQN algorithm grows slowly within the first 5000 rounds, and its reward curve does not approach the convergence state until about 15000 rounds. Due to the introduction of the Dueling Network structure to enhance the state value estimation ability, the reward value of the D3QN algorithm shows a significant upward trend around 8000 rounds, and its convergence speed is about 46% higher than that of DQN. The method of the present invention further combines the Prioritized Experience Replay (PER) mechanism to accelerate the strategy optimization by resampling high TD-error samples, achieving a rapid increase in the reward value within 5000 rounds. Its convergence speed is 37% higher than that of D3QN and 67% higher than that of DQN, verifying the improvement effect of the PER mechanism on the training efficiency.
[0031] From Figure 6 the experimental results, it can be seen that the method of the present invention is superior to the traditional DQN and D3QN algorithms in both training efficiency and final performance. Specifically, the task completion rate of the method of the present invention begins to increase significantly at about 5000 rounds and reaches a relatively high level at about 10000 rounds, finally stabilizing above 90%. In contrast, the traditional D3QN algorithm does not start to improve significantly until after 8000 rounds, and its convergence speed is significantly slower than that of the method of the present invention, and the final task completion rate can only reach about 65%. The basic DQN algorithm performs the worst, with the task completion rate close to 0 during most of the training process (0 - 15000 rounds), and there is only a slight increase in the later stage of training. Even after 20000 rounds of training, its task completion rate still does not exceed 5%.
[0032] In summary, the learning efficiency of DQN in the UAV obstacle avoidance task is low, the convergence speed is slow, the task completion success rate is low, and the quality of the final strategy is poor. Compared with DQN, D3QN has a faster convergence speed, a higher task success rate, and a better obstacle avoidance strategy due to its dual valuation mechanism. The method of the present invention further improves the training efficiency and the final performance through prioritized experience replay, and can significantly improve the task success rate of the UAV while reducing the training time.
[0033] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for obstacle avoidance and navigation of an unmanned aerial vehicle based on a lightweight duel double deep Q-network, characterized in that: It includes the following steps; S1. Define the obstacle avoidance and autonomous navigation problem of the drone as a Markov decision process, and establish the MDP five-tuple MDP = (S, A, P, R, γ), where S is the state space, A is the action space, P is the state transition function, R is the reward function, and γ is the discount factor; S2. Construct the interaction framework between the environment and the drone in reinforcement learning, and design the state space, action space, and reward function; The state space S is composed of multiple states, and each state includes three consecutive depth maps, which are obtained by the drone in the Airsim simulation platform; The action space A includes multiple actions of the drone; The reward function R includes a result reward function, a distance reward function, and an action execution reward function; S3. Build a deep neural network based on the Dueling architecture, including a main network and a target network with the same structure; S4. Set up an experience pool for storing multiple pieces of experience data. The experience data is a five-tuple (s, a, r, s', f), where s is the current state, a is the current action, r is the reward of the current action, s' is the next state, and f is whether to complete the current round of training; S5. Train the deep neural network with the D3QN-PER algorithm to obtain the trained main network and target network; S6. Perform lightweight processing on the trained main network using neural network pruning technology to obtain a lightweight model; S7. Real-time obtain the state of the drone during flight, and select actions for execution using the lightweight model.
2. The method for obstacle avoidance and navigation of an unmanned aerial vehicle based on a lightweight duel double deep Q-network according to claim 1, wherein: In S1, the discount factor γ ∈ (0, 1), and the depth map is preprocessed.
3. The method for obstacle avoidance and navigation of an unmanned aerial vehicle based on a lightweight duel double deep Q-network according to claim 1, characterized in that: The action space includes 5 actions, which are respectively; Action 0: Move along the positive x-axis at a speed of 1 m / s for 2 s; Action 1: Move along the negative x-axis at a speed of 1 m / s for 2 s; Action 2: Move along the positive y-axis at a speed of 1 m / s for 2 s; Action 3: Move along the negative y-axis at a speed of 1 m / s for 2 s; Action 4: Hover at the current position for 2 s.
4. The method for obstacle avoidance and navigation of an unmanned aerial vehicle based on a lightweight duel double deep Q-network according to claim 1, wherein: The reward function R is calculated according to the following formula; , , , , where r res , r d , r a represent the result reward function, the distance reward function, and the action execution reward function respectively, ω d , ω a represent the distance reward function coefficient and the action execution reward function coefficient respectively, D is the Euclidean distance in three-dimensional space between the current position of the UAV and the target location, and threshold is a threshold used to specify that within this threshold range from the target point is considered to have reached the target point; D next represents the Euclidean distance in three-dimensional space between the next position of the UAV and the target location; action = 0 to action = 4 are respectively the actions 0 to 4 in the action space.
5. The method for obstacle avoidance and navigation of an unmanned aerial vehicle based on a lightweight duel double deep Q-network according to claim 1, characterized in that: The main network is based on the Dueling DQN network structure, including a feature extraction layer, whose output ends are respectively connected to a state value function layer and an advantage function layer; The feature extraction layer is a ResNet50 network, and its last fully connected layer is removed; The state value function layer is used to output the state value V(s), and the advantage function layer is used to output the advantage function A(s, a); The input of the main network is three consecutive depth maps captured by the drone in the Airsim simulation platform, and the output is the Q value of the drone in state s and action a, which is Q(s, a). Q(s, a) = V(s) + A(s, a) - mean(A(s, a)), where mean( ) is the mean function.
6. The method for obstacle avoidance and navigation of an unmanned aerial vehicle based on a lightweight duel double deep Q-network according to claim 1, wherein: S6 specifically includes S61~S63; S61. Use DepGraph to perform unstructured pruning on the Q network, with a pruning ratio of 10%; S62. Fine-tune the model obtained in S61; S63. Repeat steps S61~S62 five times to complete 50% pruning.
Citation Information
Patent Citations
Lightweight autonomous navigation method and device based on neural network driving
CN117906614A
Cited By
Low-altitude aircraft collision avoidance method based on intention prediction and deep reinforcement learning
CN121115517A