Multi-unmanned aerial vehicle obstacle avoidance method based on APF-Dueling DQN
By introducing the APF-Duelling DQN method into UAV obstacle avoidance technology, combined with the advantages of APF and DQN, the challenge of UAV obstacle avoidance in complex low-altitude environments has been solved, and more efficient and safer flight path planning has been achieved.
Patent Information
- Application Number
- CN202510126352.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-16
AI Technical Summary
Existing UAV obstacle avoidance technology is difficult to effectively avoid unknown or completely unknown obstacles in complex low-altitude environments, resulting in the threat of flight safety.
The obstacle avoidance method of multiple drones based on APF-Duelling DQN is adopted, combining the intuitiveness of the APF method and the learning ability of the DQN algorithm, and the stability and efficiency of the decision-making process are improved through the Dueling network structure.
It significantly improves the obstacle avoidance performance and adaptability of the drone in complex environments, and achieves beneficial effects of smoother flight trajectory, shorter flight distance, more effective obstacle avoidance and robustness.
Smart Images

Figure BDA0005259994070000031 
Figure BDA0005259994070000041 
Figure BDA0005259994070000042
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle control, and in particular relates to an obstacle avoidance method for multiple unmanned aerial vehicles based on APF-Duelling DQN. Background Art
[0002] With the intelligent development of UAV technology, this new type of aircraft has been widely used in military and civilian fields, with missions ranging from surveillance to search and rescue, escort protection, cargo delivery, and even for rescue situations or commercial business. UAVs have the advantages of high cost performance, flexible use, and no restrictions on the physiological conditions of pilots, so they are recognized by countries around the world. In recent decades, with the development of aviation, control and electronic information technologies, countries have continued to pay attention to and increase investment in the field of UAVs. UAV technology has made great progress and development, representing the development direction of today's high-tech. However, as the operating airspace of UAVs continues to expand from medium and high altitudes to low altitudes and even ultra-low altitudes, the complexity of the obstacle environment they face has gradually increased. The low-altitude obstacle environment has the characteristics of density, non-convexity, dynamics and uncertainty, and there may be unknown or completely unknown obstacle information. The complex obstacle environment brings great challenges to the flight safety of UAVs. Once a UAV collides with an obstacle, it will not only cause huge economic losses, but also threaten human life safety. Therefore, the obstacle avoidance and path planning technology of UAVs has received widespread attention and has broad research significance.
[0003] Obstacle avoidance path planning for unmanned aerial vehicles refers to the use of specific algorithms and technologies to determine an optimal flight path from the starting point to the target point when the unmanned aerial vehicle performs a flight mission, while avoiding obstacles on the flight path to ensure the flight safety of the unmanned aerial vehicle. At present, domestic and foreign scholars have conducted a lot of research on this issue from different angles and proposed various theories and methods. The current obstacle avoidance technology is mainly divided into obstacle avoidance methods based on optimization, potential field and machine learning. The obstacle avoidance method based on optimization can handle complex constraints and dynamic problems, but the mathematical optimization algorithm is complicated and difficult to understand. The heuristic algorithm has poor real-time performance and is not suitable for online obstacle avoidance planning scenarios. It is usually only suitable for offline planning or global initial planning; the obstacle avoidance method based on potential field can quickly generate obstacle avoidance paths, with good real-time and smooth paths, but it is impossible to add various constraints to the obstacle avoidance process and is prone to fall into local optimal solutions; the obstacle avoidance method based on machine learning performs well in planning real-time and global performance, and does not rely on environmental prior information. However, when the unmanned aerial vehicle is in a continuous state space and action space scenario, the offline learning and training of the model takes a long time, is not easy to converge, and is even difficult to complete the training. In addition, a large amount of training data and computing resources are required. Summary of the invention
[0004] In view of the above problems, the purpose of the present invention is to provide a multi-UAV obstacle avoidance method based on APF-Duelling DQN, which combines the intuitiveness of the APF method and the learning ability of the DQN algorithm, and improves the stability and efficiency of the decision-making process through the dueling network structure, thereby improving the obstacle avoidance performance and adaptability of UAVs in complex environments.
[0005] The present invention is achieved through the following technical solutions:
[0006] A multi-UAV obstacle avoidance method based on APF-Duelling DQN, comprising the following steps:
[0007] Step 1: Establish a 3D kinematic model of the UAV by comprehensively considering the flight dynamics of the UAV;
[0008] Step 2: Set the internal and external collision avoidance rules of the drone;
[0009] Step 3: Design Dueling-DQN method and APF method based on drone obstacle avoidance;
[0010] Step 4: Design the APF-Duelling DQN algorithm according to step 3 to form an obstacle avoidance strategy.
[0011] Furthermore, the 3D kinematic model of the UAV in step 1 includes the thrust generated by the engine, the lift generated by the relative movement of the wing and the air, and the steering control factor by adjusting the pitch angle and the heading angle.
[0012] Further, the step 2 specifically includes the following steps:
[0013] Step 2.1: Set the external collision avoidance rules of the drone, and set the warning area and safety area with the drone as the center;
[0014] Step 2.2: Set the internal collision avoidance rules of the drones to ensure that there is no internal collision between drones and that the drones maintain their formation.
[0015] Further, the step 3 specifically includes the following steps:
[0016] Step 3.1: The reinforcement learning method based on Dueling-DQN quickly converges to the optimal strategy;
[0017] Step 3.1.1: Obtain the optimal learning strategy for updating the action-value function through the Q-learning algorithm;
[0018] Step 3.1.2: Decompose the estimated value of the Q value in step 3.1.1 into a state value and an advantage value through Dueling DQN, effectively learn the relationship between the state and the action, and quickly converge to the optimal strategy;
[0019] Step 3.2: Use the APF method to establish the gravitational potential field and repulsive potential field to perform local path planning for the UAV.
[0020] Further, the step 4 specifically includes the following steps:
[0021] Step 4.1: Design a state set according to the kinetic equation in step 1;
[0022] Step 4.2: Design an action set based on the heading angle and pitch angle of the drone;
[0023] Step 4.3: Set up the Dueling-DQN algorithm for drone improvement rewards.
[0024] Further, the step 1 is specifically as follows:
[0025] The Y-direction UAV speed is calculated based on the pitch angle and heading angle. Taking N UAVs as an example, the motion equation of UAV (i = 1, 2…, N) is given by the following kinematic model:
[0026]
[0027] Where (x i ,y i ,z i ) is the position coordinate of the UAV, V i is the speed of drone i.
[0028] Further, the step 2 specifically includes the following steps:
[0029] Step 2.1: Set up the external collision avoidance rules for the drone
[0030] Assuming that the UAV has a limited sensing range, we model a radius d centered at the geometric center of the UAV. w This area is determined by its sensor capability and is set as a warning area. Then a circular area with a radius of d and a center of the UAV’s geometric center is modeled. s The circular area of the s <d w , set it as a safe area;
[0031] The safety area and warning area centered on the drone are divided into two spheres, the inner and outer spheres. io is the distance from UAV i to the center of the obstacle; d s is the safety distance, when d io <d s When the drone has no buffer distance to avoid obstacles; w is the warning distance, when ds <d io <d w When d s and d w , according to the actual situation, d io The calculation formula is as follows:
[0032]
[0033] Where (ob xo ,ob yo ,ob zo ) is the coordinate of the oth obstacle, assuming that the coordinate of the obstacle is known in the algorithm;
[0034] Step 2.2: Set up the internal collision avoidance rules of the drone
[0035] The two drones are arranged in a basically parallel manner, wherein d ij is the distance between UAV i and UAV j, d min is the minimum distance between two drones, when d ij >d min When >0, it ensures that there is no collision between drones and achieves the purpose of multiple drone missions.
[0036] Further, the step 3 specifically includes the following steps:
[0037] Step 3.1: Reinforcement learning method based on Dueling-DQN
[0038] Step 3.1.1: Q-Learning
[0039] Reinforcement learning allows the agent to learn an optimal strategy to maximize future cumulative rewards:
[0040]
[0041] Where γ is the discount factor 0<γ≤1, and r is the reward for the agent;
[0042] Action-value function Q π It can be expressed as:
[0043] Q π (s t ,a t )=Ε[U t |S t =s t ,A t =a t ]
[0044] where st is the state of the agent at time t, a t is the action performed by the agent at time t;
[0045] The action-value function can be converted to the optimal action-value function:
[0046]
[0047] The Q-learning algorithm learns the optimal strategy by updating the action value function. However, when selecting actions, the ε-greedy strategy is usually adopted, with random selection with ε probability and the optimal strategy selected with 1-ε probability. The update method is as follows:
[0048]
[0049] Where α is the learning rate;
[0050] Step 3.1.2: Dueling DQN
[0051] Dueling DQN is an improvement on the original DQN algorithm in terms of network structure. It improves learning efficiency and reduces estimation errors by outputting the state value (V) and the advantage value (A) independently. Dueling DQN decomposes the estimated value of the Q value into two parts: the state value (V) and the advantage value (A).
[0052] Q * (s,a;w)=V * (s; w V )+A * (s,a;w A )
[0053] where ω A is the network parameter of the advantage function neural network, ω V is the network parameter of the state-value function neural network. When only the current formula is used for updating, the challenge of "non-uniqueness" arises. For example, the addition and subtraction of values V and A may produce the same result Q, but the unique values V and A cannot be determined from Q. To solve this problem, the following formula is designed:
[0054]
[0055] Dueling DQN can learn the relationship between states and actions more effectively, thereby improving learning efficiency and performance. This decomposition helps reduce variance in the learning process because the estimation of the advantage value can offset the deviation in certain state values. Since Dueling DQN can learn the relationship between states and actions more effectively, it can usually converge to the optimal strategy at a faster speed.
[0056] Step 3.2: APF method
[0057] The drone is regarded as an object, and virtual gravitational potential fields and repulsive potential fields are added to the environment. Specifically, a virtual gravitational potential field is established at the target position to attract the drone, and a virtual repulsive potential field is established at the obstacle position to prevent the drone from moving to the obstacle. Therefore, the movement of the drone is affected by the combined force of the gravitational potential field and the repulsive potential field, which facilitates the planning of an optimal path so that the drone can avoid obstacles and approach the target. In the APF method, the gravitational potential field can be expressed as:
[0058]
[0059] Among them U att (p) is the gravitational potential field, ξ is the gravitational potential field coefficient, d ig is the distance between the target and the drone. The closer the drone is to the target, the smaller the gravitational potential field is, and its gradient is as follows:
[0060]
[0061] The direction of the drone's motion is generated by the negative gradient of the gravitational potential field function, so the gravitational force is:
[0062] F att (p)=-▽U att (p)=-ξd(p,p goal )
[0063] The repulsive potential field function is as follows:
[0064]
[0065] Where η is the gain coefficient of the repulsive potential field, d(p,p obs ) is the distance between the drone and the obstacle, d o is the influence radius of the obstacle, which is the threshold of the repulsive force of the obstacle. When the drone exceeds the threshold range, the drone is not affected by the repulsive force. When the drone is within the threshold range, the closer the obstacle is, the greater the repulsive force field will be, and its gradient is as follows:
[0066]
[0067] The direction of the drone's motion is generated by the negative gradient of the repulsive potential field function, so the repulsive force is:
[0068]
[0069] When subjected to the forces of gravity and repulsion, the drone is able to navigate to a safe destination while avoiding obstacles.
[0070] Further, the step 4 specifically includes the following steps:
[0071] Modeling is done according to the Markov decision process (MDP). The collision avoidance problem of multiple UAVs for the purpose of performance optimization and reaching the target point is expressed as a Markov decision process problem (S, A, P, R), where S is the state set, A is the set of finite actions, R is the finite set of rewards, and P represents the state transition probability. The following describes the set of states, actions, and rewards:
[0072] Step 4.1: Design Status
[0073] According to the dynamic equation in step 1, the state of the drone is composed of coordinate points. The state of the drone is:
[0074] s=(p1,p2,p3...p n )
[0075] p i =(x i ,y i ,z i )
[0076] where p i is the three-dimensional coordinate of UAV i at time t;
[0077] Step 4.2: Design Action
[0078] Under the assumption that the drone has a constant speed, the drone's action is determined by the heading angle and pitch angle. The combination of the finite discrete heading angle and pitch angle constitutes the drone's collision avoidance function. The action design of drone i is shown as follows:
[0079] a i =(θ i ,ψ i )
[0080] Since large angle changes are not possible, the following restrictions are added:
[0081]
[0082] The angular change of the pitch angle is smaller than that of the heading angle, which is beneficial to the stability of the drone. Then, the action a of all drones is expressed by the following formula:
[0083] a=(a1,a2,a3...a n )
[0084] a∈A;
[0085] Step 4.3: Setting the Reward Function
[0086] Set a reasonable reward so that the drones will try to figure out which actions will yield the most profitable reward and then complete the mission requirements. The total reward for all drones is set as follows:
[0087]
[0088] r i =r igoal +r idistance +r iaction +r iatt +r irep
[0089] There are five parts in the reward, which are responsible for reaching the target point, staying away from obstacles, energy consumption caused by changing the pitch angle and heading angle of the drone, the gravitational force of the target point on the drone, and the repulsion of the target point on the drone. The rewards for completing the goals are as follows:
[0090]
[0091] R ig It is a reward set according to the distance between the drone and the target point. This design allows the drone to receive a certain reward when it has not reached the target point, and the closer the drone is to the target point, the greater the reward, so that the drone can reach the target point as soon as possible. g is the reward when reaching the target point, R ig It is calculated as follows:
[0092]
[0093] Where (x g ,y g ,z g ) is the target point coordinate, k is the reward coefficient;
[0094] When changing the angle beyond a set angle, a slight penalty is given:
[0095]
[0096] The most important reward design is the reward for distance from obstacles. First of all, it is necessary to ensure that the drone cannot hit obstacles and cannot fly out of the set map, because this will directly lead to mission failure. At the same time, it is necessary to avoid the drone's action being too large, which will cause a greater action cost. In this technical solution, a coefficient λ is added according to the distance between the drone and the obstacle. i :
[0097]
[0098] Where, d iois the distance between the drone and the obstacle, and then sets the reward based on the distance from the obstacle:
[0099] r idistance =(λ i -1)R g
[0100] The target point's gravity bonus to the drone is as follows:
[0101]
[0102] The repulsion bonus of obstacles on the drone is as follows:
[0103]
[0104] The Dueling-DQN algorithm is as follows:
[0105] Step 1: Initialize the experience pool D to capacity N;
[0106] Initialize the action function with random weights;
[0107] Initialize Q with the same random weights _eval and Q _target ; where Q _eval is the evaluation network, Q _target is the target network;
[0108] Step 2: for episode = 1, M do; where M is the total number of episodes;
[0109] Initialize rewards and set the initial points of multiple drones;
[0110] for t=1,T do; where T is the number of steps in each episode;
[0111] Select a random action set a according to the exploration rate ε t ;
[0112] Otherwise, select the action set with the highest reward:
[0113] Execute the selected action t Get the reward r at the next moment and the state s at the next moment t+1 ;
[0114] calculate:
[0115] Sort Q(s,a) from largest to smallest to get rank(t);
[0116] Store the converted experience pool E into D in order;
[0117] Calculate the loss function loss = E((f(x)-y) 2 ), update the weights through a gradient descent procedure on the loss function;
[0118] Each step will Q _eval The parameters of Q are updated synchronously _target ;
[0119] Update status:s t ←s t+1 ;
[0120] end for;
[0121] end for.
[0122] Compared with the prior art, the present invention has the following beneficial effects: it combines the intuitiveness of artificial potential fields and the learning ability of deep Q networks, especially the DQN improved by the dueling network structure, the APF-Duelling DQN method significantly reduces the training time and improves the real-time performance of obstacle avoidance tasks. The APF-Duelling DQN method designed by combining the dueling deep Q network Dueling-DeepQ-Network (Dueling-DQN) and artificial potential field Artificial Potential Field (APF) technology solves the problem of multi-UAV obstacle avoidance in a multi-obstacle environment, and obtains the beneficial effects of smoother flight trajectory, shorter flight distance, more effective obstacle avoidance and robustness in complex environments. Specifically, it is manifested in:
[0123] 1. Enhanced learning ability: The adopted Dueling DQN improves learning efficiency and reduces estimation error by separating the estimation of state value (V) and action advantage (A). In APF-Dueling DQN, this learning ability enables the drone to better understand and predict its environment and make more accurate obstacle avoidance decisions;
[0124] 2. Improved decision-making stability: The dueling network structure used improves the stability of the DQN algorithm by reducing the variance in the learning process. This stability is reflected in APF-Dueling DQN, making the decision-making process of drones in complex environments more reliable;
[0125] 3. Optimized reward mechanism: In APF-Dueling DQN, the reward function is designed in combination with the APF concepts of gravity and repulsion, which helps guide the drone to reach the target more efficiently while avoiding obstacles. BRIEF DESCRIPTION OF THE DRAWINGS
[0126] Figure 1 It is a schematic diagram of the three-dimensional motion model of the UAV in the present invention;
[0127] Figure 2 This is a schematic diagram of the safety zone and warning zone of the drone in the present invention;
[0128] Figure 3 A schematic diagram of the distance between multiple drones in the present invention;
[0129] Figure 4 A schematic diagram showing the comparison of the obstacle avoidance of a single UAV using APF-Dueling DQN and Dueling DQN in the present invention;
[0130] Figure 5 A schematic diagram showing the comparison of using APF-Dueling DQN and Dueling DQN to perform obstacle avoidance for multiple drones in the present invention;
[0131] Figure 6 It is a schematic diagram of the algorithm flow in the present invention. DETAILED DESCRIPTION
[0132] The following is combined with Figure 1-6 The present invention is further illustrated by the following embodiments.
[0133] In an embodiment of the present invention, a method for avoiding obstacles for multiple drones based on APF-Duelling DQN specifically includes the following steps:
[0134] The step 1 is to establish a 3D kinematic model of a UAV.
[0135] The obstacle avoidance problem of fixed-wing UAVs requires comprehensive consideration of the flight dynamics of the UAV, including the thrust generated by the engine, the lift generated by the relative movement of the wing and the air, and steering control by adjusting the pitch and heading angles. Figure 1 is the Y-direction UAV velocity calculated based on the pitch angle and heading angle. Taking N UAVs as an example, the motion equation of UAV (i=1,2…,N) is given by the following kinematic model:
[0136]
[0137] Where (x i ,y i ,z i ) is the position coordinate of the UAV, V i is the speed of drone i.
[0138] Step 2: Setting internal and external collision avoidance rules for the drone.
[0139] Step 2.1: External collision avoidance of the drone
[0140] During the movement of the UAV, the speed is often very fast. This means that when there is an obstacle in front of it, no one can turn in time to avoid it due to inertia. Therefore, the UAV needs a certain buffer zone to react in advance. Assuming that the UAV has a limited sensing range, a radius d is modeled with the geometric center of the UAV as the center. w This area is determined by its sensor capability and is set as the warning area. Then model a circular area with a radius of d centered at the geometric center of the UAV. s The circular area of the s <d w , making it a safe area.
[0141] like Figure 2 As shown in the figure, the safety area and warning area are centered on the drone, which are divided into two spheres, the inner and outer spheres. io is the distance from UAV i to the center of the obstacle. s is the safety distance, when d io <d s When the drone is at a certain distance, it will have no buffer distance to avoid obstacles. w is the warning distance, when d s <d io <d w When d s and d w , according to the actual situation, d io The calculation formula is as follows:
[0142]
[0143] Where (ob xo ,ob yo ,ob zo ) is the coordinate of the oth obstacle. It is assumed that the coordinates of the obstacle are known in the algorithm.
[0144] Step 2.2: Internal Collision Avoidance of the Drone
[0145] When performing a mission, one drone often cannot complete the target requirements, and usually several drones form a formation to complete the mission together. Then the possible collision and distance maintenance between drones will be a problem worth discussing, so a simple setting is made below to ensure that there will be no internal collision between drones and to maintain the formation.
[0146] Figure 3 The two drones are basically arranged side by side, with d ij is the distance between UAV i and UAV j, d minis the minimum distance between two drones. ij >d min When >0, it ensures that there is no collision between drones and achieves the purpose of multiple drone missions.
[0147] Step 3: Introduce the reinforcement learning method and APF method of Dueling-DQN.
[0148] Step 3.1: Reinforcement Learning Method Based on Dueling-DQN
[0149] Step 3.1.1: Q-Learning
[0150] Reinforcement learning requires the subject to continuously participate in the environment and use the reward function obtained to enhance its understanding of the environment. Reinforcement learning aims to allow the agent to learn an optimal strategy to maximize the future cumulative rewards:
[0151]
[0152] Where γ is the discount factor 0<γ≤1, and r is the reward for the agent.
[0153] Action-value function Q π It can be expressed as:
[0154] Q π (s t ,a t )=Ε[U t |S t =s t ,A t =a t ]
[0155] where s t is the state of the agent at time t, a t is the action performed by the agent at time t.
[0156] The action-value function can be converted to the optimal action-value function:
[0157]
[0158] In practical reinforcement learning scenarios, determining the reward function and the probability of state transitions can be challenging. A reinforcement learning algorithm that does not rely on an environment model is called "model-free learning" and poses greater challenges than methods that include such models. Q-learning is a tabular model-free technique that is widely used to address the challenges of discrete state and action spaces in model-free environments. The Q-learning algorithm uses experience-based learning, which does not require an environment model, but instead lets the Q-table learn the value function directly from experience to obtain the optimal policy. The core idea of the Q-learning algorithm is to learn the optimal policy by updating the action-value function. However, when selecting actions, an ε-greedy strategy is usually adopted, with random selections made with ε probability and the optimal policy selected with 1-ε probability. The update method is as follows:
[0159]
[0160] Where α is the learning rate.
[0161] Step 3.1.2: Dueling DQN
[0162] Traditional DQN methods often have the characteristics of slow convergence and over-estimation. In order to solve these problems, the present invention introduces the Dueling DQN algorithm, which is a deep reinforcement learning algorithm. Dueling DQN is an improvement on the network structure of the original DQN algorithm. By outputting the state value (V) and the advantage value (A) independently, the learning efficiency is improved and the estimation error is reduced. The core idea of Dueling DQN is to decompose the estimated value of the Q value into two parts: the state value (V) and the advantage value (A).
[0163] Q * (s,a;w)=V * (s; w V )+A * (s,a;w A )
[0164] where ω A is the network parameter of the advantage function neural network, ω V is the network parameter of the state-value function neural network. When only the current formula is used for updating, the challenge of "non-uniqueness" arises. For example, the addition and subtraction of values V and A may produce the same result Q, but the unique values V and A cannot be determined from Q. To solve this problem, the following formula is designed:
[0165]
[0166] Dueling DQN is able to learn the relationship between states and actions more effectively, thereby improving learning efficiency and performance. This decomposition helps reduce variance in the learning process because the estimate of the advantage value can offset the bias in certain state values. Because Dueling DQN is able to learn the relationship between states and actions more effectively, it can usually converge to the optimal policy faster.
[0167] Step 3.2: APF Method
[0168] The APF algorithm is a commonly used algorithm for local path planning of UAVs. The basic idea of the algorithm is to regard the UAV as an object and add virtual gravitational potential fields and repulsive potential fields to the environment. Specifically, a virtual gravitational potential field is established at the target position to attract the UAV, and a virtual repulsive potential field is established at the obstacle position to prevent the UAV from moving onto the obstacle. Therefore, the movement of the UAV is affected by the combined force of the gravitational potential field and the repulsive potential field, which facilitates the planning of an optimal path so that the UAV can avoid obstacles and approach the target. In the APF method, the gravitational potential field can be expressed as:
[0169]
[0170] Among them U att (p) is the gravitational potential field, ξ is the gravitational potential field coefficient, d ig is the distance between the target and the drone. The closer the drone is to the target, the smaller the gravitational potential field is. Its gradient is as follows:
[0171] ▽U att (p) = ξd(p,p goal )
[0172] The direction of the drone's motion is generated by the negative gradient of the gravitational potential field function, so the gravitational force is:
[0173] F att (p)=-▽U att (p)=-ξd(p,p goal )
[0174] The repulsive potential field function is as follows:
[0175]
[0176] Where η is the gain coefficient of the repulsive potential field, d(p,p obs ) is the distance between the drone and the obstacle, d o is the influence radius of the obstacle, which is the threshold of the repulsive force of the obstacle. When the drone is beyond the threshold range, the drone is not affected by the repulsive force. When the drone is within the threshold range, the closer the obstacle is, the greater the repulsive force field will be. Its gradient is as follows:
[0177]
[0178] The direction of the drone's motion is generated by the negative gradient of the repulsive potential field function, so the repulsive force is:
[0179]
[0180] When subjected to the forces of gravity and repulsion, the drone is able to navigate to a safe target while avoiding obstacles. However, it is important to acknowledge that the APF technique does not ensure the discovery of a global optimum, as there is a risk that the drone will get stuck in a local optimum, potentially making the target unreachable.
[0181] The step 4: designing the APF-Duelling DQN algorithm.
[0182] Reinforcement learning is usually modeled using a Markov decision process (MDP). Markov decision process is one of the most important theoretical foundations of reinforcement learning and is also the core of many reinforcement learning algorithms. MDP is a five-tuple (S, A, P, R, S'). The collision avoidance problem of multiple drones for the purpose of performance optimization and reaching the target point can be expressed as a Markov decision process problem (S, A, P, R). Among them, S is the state set, A is the set of finite actions, R is the finite set of rewards, and P represents the state transition probability. The set of states, actions, and rewards will be described below.
[0183] Step 4.1: State
[0184] According to the dynamic equation in step 1, the state of the drone consists of coordinate points. Therefore, the state of the drone can be designed as
[0185] s=(p1,p2,p3...p n )
[0186] p i =(x i ,y i ,z i )
[0187] where p i is the three-dimensional coordinate of UAV i at time t.
[0188] Step 4.2: Action
[0189] Under the assumption that the drone has a constant speed, the drone's action is determined by the heading angle and pitch angle. The combination of finite discrete heading angle and pitch angle constitutes the drone's collision avoidance function. The action design of drone i is shown below:
[0190] a i =(θ i ,ψi )
[0191] Since large angle changes are not possible, the following restrictions are added:
[0192]
[0193] The pitch angle changes less than the heading angle, which is good for the stability of the drone. Then, the action a of all drones is expressed by the following formula:
[0194] a=(a1,a2,a3...a n )
[0195] a∈A
[0196] Step 4.3: Reward Function
[0197] Reinforcement learning is the process of maximizing digital benefit signals through actions. Therefore, after the drone takes action in the corresponding state, it needs to get a reward based on the next state, so the setting of this reward is very important. Unreasonable rewards may lead to failure to train the correct results. A reasonable reward needs to be set so that the drone will try to find out which actions will produce the most profitable rewards and then complete the task requirements. Set the total reward for all drones to:
[0198]
[0199] r i =r igoal +r idistance +r iaction +r iatt +r irep
[0200] There are five parts in the reward, which are responsible for reaching the target point, staying away from obstacles, energy consumption caused by changing the pitch angle and heading angle of the drone, the gravitational force of the target point on the drone, and the repulsion of the target point on the drone. The rewards for completing the goal are as follows:
[0201]
[0202] R ig It is a reward set according to the distance between the drone and the target point. This design allows the drone to receive a certain reward when it has not reached the target point, and the closer the drone is to the target point, the greater the reward, so that the drone can reach the target point as soon as possible. g is the reward when reaching the target point. ig It is calculated as follows:
[0203]
[0204] Where (x g ,y g ,z g ) is the target point coordinate. k is the reward coefficient.
[0205] When changing the angle beyond the set angle, a slight penalty is given:
[0206]
[0207] The most important reward design is the reward for distance from obstacles. First of all, it is necessary to ensure that the drone cannot hit obstacles and cannot fly out of the set map, because this will directly lead to mission failure. But at the same time, it is necessary to avoid the drone's movements being too large, which will cause greater movement costs. In this technical solution, a coefficient λ is added according to the distance between the drone and the obstacle. i :
[0208]
[0209] Where, d io The distance between the drone and the obstacle, and then set the reward based on the distance from the obstacle:
[0210] r idistance =(λ i -1)R g
[0211] The target point's gravity bonus to the drone is as follows:
[0212]
[0213] The repulsion bonus of obstacles on the drone is as follows:
[0214]
[0215] The Dueling-DQN algorithm is as follows:
[0216] Step 1: Initialize the experience pool D to capacity N;
[0217] Initialize the action function with random weights;
[0218] Initialize Q with the same random weights _eval and Q _target ; where Q _eval is the evaluation network, Q _target is the target network;
[0219] Step 2: for episode = 1, M do; where M is the total number of episodes;
[0220] Initialize rewards and set the initial points of multiple drones;
[0221] for t=1,T do; where T is the number of steps in each episode;
[0222] Select a random action set a according to the exploration rate ε t ;
[0223] Otherwise, select the action set with the highest reward:
[0224] Execute the selected action t Get the reward r at the next moment and the state s at the next moment t+1 ;
[0225] calculate:
[0226] Sort Q(s,a) from largest to smallest to get rank(t);
[0227] Store the converted experience pool E into D in order;
[0228] Calculate the loss function loss = E((f(x)-y) 2 ), update the weights through a gradient descent procedure on the loss function;
[0229] Each step will Q _eval The parameters of Q are updated synchronously _target ;
[0230] Update status:s t ←s t+1 ;
[0231] end for;
[0232] end for.
[0233] In the embodiment of the present invention, Figures 1 to 6 As shown, in order to more intuitively demonstrate and verify the stability and accuracy of the obstacle avoidance effect of the multi-UAV obstacle avoidance method based on APF-DuellingDQN proposed by the present invention, the present invention adopts the APF-Dueling DQN algorithm to solve the autonomous obstacle avoidance decision-making problem of UAVs, and improves the reward in reinforcement learning by introducing the concept of APF, and reasonably incorporates the concepts of gravitational potential field and repulsive potential field in APF into the reward setting; at the same time, the distance limit of UAVs is adopted to ensure that there is no collision between UAVs. Simulation verification is carried out based on the above.
[0234] This simulation example uses MATLAB to establish a three-dimensional environment map and a UAV motion model; a neural network is used to train one UAV and five UAVs at the same time, and the obstacle avoidance effect is analyzed through the flight trajectory of the UAV; finally, through simulation experiments, the feasibility of APF-Dueling DQN training UAVs and the effectiveness of the algorithm are verified, and this method can successfully achieve the obstacle avoidance effect.
[0235] like Figure 4 As shown in Figure 2, it is found that the drone trained with the DQN algorithm and the APF-Dueling DQN algorithm can reach the target point; comparison Figure 4 The two trajectories shown show that the trajectory of the APF-DuelingDQN algorithm is smoother, shorter, and has a good obstacle avoidance effect.
[0236] like Figure 5 As shown in the figure, the five drones trained by the DQN algorithm and the APF-DuelingDQN algorithm can reach the same target point; comparison Figure 5 The two trajectories shown show that the trajectories of the five drones using the APF-Dueling DQN algorithm are smoother, do not produce reverse trajectories, produce shorter paths, and have good obstacle avoidance effects.
[0237] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Other modifications or equivalent substitutions made to the technical solution of the present invention by ordinary technicians in the field should be included in the scope of the claims of the present invention as long as they do not depart from the spirit and scope of the technical solution of the present invention.
Claims
1. A multi-UAV obstacle avoidance method based on APF-Duelling DQN, characterized in that: The following steps are involved: Step 1: Establish a 3D kinematic model of the UAV by comprehensively considering the flight dynamics of the UAV; Step 2: Set the internal and external collision avoidance rules of the drone; Step 3: Design Dueling-DQN method and APF method based on drone obstacle avoidance; Step 4: Design the APF-Duelling DQN algorithm according to step 3 to form an obstacle avoidance strategy.
2. The obstacle avoidance method for multiple UAVs based on APF-Duelling DQN according to claim 1 is characterized in that: The 3D kinematic model of the UAV in step 1 includes the thrust generated by the engine, the lift generated by the relative motion between the wing and the air, and the steering control factor by adjusting the pitch angle and the heading angle.
3. The obstacle avoidance method for multiple UAVs based on APF-Duelling DQN according to claim 1 is characterized in that: The step 2 specifically includes the following steps: Step 2.1: Set the external collision avoidance rules of the drone, and set the warning area and safety area with the drone as the center; Step 2.2: Set the internal collision avoidance rules of the drones to ensure that there is no internal collision between drones and that the drones maintain their formation.
4. The obstacle avoidance method for multiple UAVs based on APF-Duelling DQN according to claim 1 is characterized in that: The step 3 specifically comprises the following steps: Step 3.1: The reinforcement learning method based on Dueling-DQN quickly converges to the optimal strategy; Step 3.1.1: Obtain the optimal learning strategy for updating the action-value function through the Q-learning algorithm; Step 3.1.2: Decompose the estimated value of the Q value in step 3.1.1 into a state value and an advantage value through Dueling DQN, effectively learn the relationship between the state and the action, and quickly converge to the optimal strategy; Step 3.2: Use the APF method to establish the gravitational potential field and repulsive potential field to perform local path planning for the UAV.
5. The obstacle avoidance method for multiple UAVs based on APF-Duelling DQN according to claim 1 is characterized in that: The step 4 specifically comprises the following steps: Step 4.1: Design a state set according to the kinetic equation in step 1; Step 4.2: Design an action set based on the heading angle and pitch angle of the drone; Step 4.3: Set up the Dueling-DQN algorithm for drone improvement rewards.
6. The obstacle avoidance method for multiple UAVs based on APF-Duelling DQN according to claim 2 is characterized in that: The step 1 is specifically as follows: The Y-direction UAV speed is calculated based on the pitch angle and heading angle. Taking N UAVs as an example, the motion equation of UAV (i = 1, 2…, N) is given by the following kinematic model: Where (x i ,y i ,z i ) is the position coordinate of the UAV, V i is the speed of drone i.
7. The obstacle avoidance method for multiple UAVs based on APF-Duelling DQN according to claim 3 is characterized in that: The step 2 specifically includes the following steps: Step 2.1: Set up the external collision avoidance rules for the drone Assuming that the UAV has a limited sensing range, we model a radius d centered at the geometric center of the UAV. w This area is determined by its sensor capability and is set as a warning area. Then a circular area with a radius of d and a center of the UAV’s geometric center is modeled. s The circular area of the s <d w , set it as a safe area; The safety area and warning area centered on the drone are divided into two spheres, the inner and outer spheres. io is the distance from UAV i to the center of the obstacle; d s is the safety distance, when d io <d s When the drone has no buffer distance to avoid obstacles; w is the warning distance, when d s <d io <d w When d s and d w , according to the actual situation, d io The calculation formula is as follows: Where (ob xo ,ob yo ,ob zo ) is the coordinate of the oth obstacle, assuming that the coordinate of the obstacle is known in the algorithm; Step 2.2: Set up the internal collision avoidance rules of the drone The two drones are arranged in a basically parallel manner, wherein d ij is the distance between UAV i and UAV j, d min is the minimum distance between two drones, when d ij >d min When >0, it ensures that there is no collision between drones and achieves the purpose of multiple drone missions.
8. The obstacle avoidance method for multiple UAVs based on APF-Duelling DQN according to claim 4 is characterized in that: The step 3 specifically comprises the following steps: Step 3.1: Reinforcement learning method based on Dueling-DQN Step 3.1.1: Q-Learning Reinforcement learning allows the agent to learn an optimal strategy to maximize future cumulative rewards: Where γ is the discount factor 0<γ≤1, and r is the reward for the agent; Action-value function Q π It can be expressed as: Q π (s t ,a t )=Ε[U t |S t =s t ,A t =a t ] where s t is the state of the agent at time t, a t is the action performed by the agent at time t; The action-value function can be converted to the optimal action-value function: The Q-learning algorithm learns the optimal strategy by updating the action value function. However, when selecting actions, the ε-greedy strategy is usually adopted, with random selection with ε probability and the optimal strategy selected with 1-ε probability. The update method is as follows: Where α is the learning rate; Step 3.1.2: Dueling DQN Dueling DQN is an improvement on the original DQN algorithm in terms of network structure. It improves learning efficiency and reduces estimation errors by outputting the state value (V) and the advantage value (A) independently. Dueling DQN decomposes the estimated value of the Q value into two parts: the state value (V) and the advantage value (A). Q * (s,a;w)=V * (s;w V )+A * (s,a;w A ) where ω A is the network parameter of the advantage function neural network, ω V is the network parameter of the state-value function neural network. When only the current formula is used for updating, the challenge of "non-uniqueness" arises. For example, the addition and subtraction of values V and A may produce the same result Q, but the unique values V and A cannot be determined from Q. To solve this problem, the following formula is designed: Dueling DQN can learn the relationship between states and actions more effectively, thereby improving learning efficiency and performance. This decomposition helps reduce variance in the learning process because the estimation of the advantage value can offset the deviation in certain state values. Since Dueling DQN can learn the relationship between states and actions more effectively, it can usually converge to the optimal strategy at a faster speed. Step 3.2: APF method The drone is regarded as an object, and virtual gravitational potential fields and repulsive potential fields are added to the environment. Specifically, a virtual gravitational potential field is established at the target position to attract the drone, and a virtual repulsive potential field is established at the obstacle position to prevent the drone from moving to the obstacle. Therefore, the movement of the drone is affected by the combined force of the gravitational potential field and the repulsive potential field, which facilitates the planning of an optimal path so that the drone can avoid obstacles and approach the target. In the APF method, the gravitational potential field can be expressed as: Among them U att (p) is the gravitational potential field, ξ is the gravitational potential field coefficient, d ig is the distance between the target and the drone. The closer the drone is to the target, the smaller the gravitational potential field is, and its gradient is as follows: The direction of the drone's motion is generated by the negative gradient of the gravitational potential field function, so the gravitational force is: The repulsive potential field function is as follows: Where η is the gain coefficient of the repulsive potential field, d(p,p obs ) is the distance between the drone and the obstacle, d o is the influence radius of the obstacle, which is the threshold of the repulsive force of the obstacle. When the drone exceeds the threshold range, the drone is not affected by the repulsive force. When the drone is within the threshold range, the closer the obstacle is, the greater the repulsive force field will be, and its gradient is as follows: The direction of the drone's motion is generated by the negative gradient of the repulsive potential field function, so the repulsive force is: When subjected to the forces of gravity and repulsion, the drone is able to navigate to a safe destination while avoiding obstacles.
9. The obstacle avoidance method for multiple UAVs based on APF-Duelling DQN according to claim 5, characterized in that: The step 4 specifically comprises the following steps: Modeling is done according to the Markov decision process (MDP). The collision avoidance problem of multiple UAVs for the purpose of performance optimization and reaching the target point is expressed as a Markov decision process problem (S, A, P, R), where S is the state set, A is the set of finite actions, R is the finite set of rewards, and P represents the state transition probability. The following describes the set of states, actions, and rewards: Step 4.1: Design Status According to the dynamic equation in step 1, the state of the drone is composed of coordinate points. The state of the drone is: s=(p1,p2,p3...p n ) p i =(x i ,y i ,z i ) where p i is the three-dimensional coordinate of UAV i at time t; Step 4.2: Design Action Under the assumption that the drone has a constant speed, the drone's action is determined by the heading angle and pitch angle. The combination of the finite discrete heading angle and pitch angle constitutes the drone's collision avoidance function. The action design of drone i is shown as follows: a i =(θ i ,ψ i ) Since large angle changes are not possible, the following restrictions are added: The angular change of the pitch angle is smaller than that of the heading angle, which is beneficial to the stability of the drone. Then, the action a of all drones is expressed by the following formula: <h2 style=";text-align:left;direction:ltr">a = (a1, a2, a3...a)<h2 style=";text-align:left;direction:ltr"> n <h2 style=";text-align:left;direction:ltr"> ) a∈A; Step 4.3: Setting the Reward Function Set a reasonable reward so that the drones will try to figure out which actions will yield the most profitable reward and then complete the mission requirements. The total reward for all drones is set as follows: r i =r igoal +r idistance +r iaction +r iatt +r irep There are five parts in the reward, which are responsible for reaching the target point, staying away from obstacles, energy consumption caused by changing the pitch angle and heading angle of the drone, the gravitational force of the target point on the drone, and the repulsion of the target point on the drone. The rewards for completing the goals are as follows: R ig It is a reward set according to the distance between the drone and the target point. This design allows the drone to receive a certain reward when it has not reached the target point, and the closer the drone is to the target point, the greater the reward, so that the drone can reach the target point as soon as possible. g is the reward when reaching the target point, R ig It is calculated as follows: Where (x g ,y g ,z g ) is the target point coordinate, k is the reward coefficient; When changing the angle beyond a set angle, a slight penalty is given: The most important reward design is the reward for distance from obstacles. First of all, it is necessary to ensure that the drone cannot hit obstacles and cannot fly out of the set map, because this will directly lead to mission failure. At the same time, it is necessary to avoid the drone's action being too large, which will cause a greater action cost. In this technical solution, a coefficient λ is added according to the distance between the drone and the obstacle. i : Where, d io is the distance between the drone and the obstacle, and then sets the reward based on the distance from the obstacle: r idistance =(λ i -1)R g The target point's gravity bonus to the drone is as follows: The repulsion bonus of obstacles on the drone is as follows: The Dueling-DQN algorithm is as follows: Step 1: Initialize the experience pool D to capacity N; Initialize the action function with random weights; Initialize Q with the same random weights _eval and Q _target ; where Q _eval is the evaluation network, Q _target is the target network; Step 2: for episode = 1, M do; where M is the total number of episodes; Initialize rewards and set the initial points of multiple drones; for t=1,T do; where T is the number of steps in each episode; Select a random action set a according to the exploration rate ε t ; Otherwise, select the action set with the highest reward: Execute the selected action t Get the reward r at the next moment and the state s at the next moment t+1 ; calculate: Sort Q(s,a) from largest to smallest to get rank(t); Store the converted experience pool E into D in order; Calculate the loss function loss = E((f(x)-y) 2 ), update the weights through a gradient descent procedure on the loss function; Each step will Q _eval The parameters of Q are updated synchronously _target ; Update status:s t ←s t+1 ; end for; end for.