Unmanned aerial vehicle dynamic path planning method based on improved DDQN network

By combining artificial potential field method and improved DDQN network, the drone path is generated and optimized, and the local optimal and real-time problems of path planning in dynamic environments are solved, and efficient obstacle avoidance for drones in complex low-altitude scenarios are achieved.

CN120445200APending Publication Date: 2025-08-08SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510514013.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art drone path planning has problems with insufficient path quality and local optimality in dynamic environments, which is difficult to meet the real-time requirements.

Method used

Combining the artificial potential field method and the improved dual-depth Q network (DDQN), the environment information and drone state are described through state vectors, and the artificial potential field method is used to calculate and generate heuristic paths, and the DDQN neural network is optimized for real-time path planning and adjustment.

Benefits of technology

It improves the path planning accuracy and real-time performance of the drone in dynamic and complex environments, reduces the random exploration time in the early stage of training, and ensures the global optimality of the path and real-time obstacle avoidance capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120445200A_ABST
    Figure CN120445200A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle dynamic path planning method based on an improved DDQN network, and relates to the technical field of unmanned aerial vehicle path planning. The method comprises the following steps: firstly, initializing the environment and state of a task execution area of the unmanned aerial vehicle, and describing environment information and the state of the unmanned aerial vehicle by using a state vector; then calculating resultant force borne by the unmanned aerial vehicle under the combined action of the target point and the obstacle in the movement process by using an artificial potential field method, calculating the movement direction of the unmanned aerial vehicle, and generating a heuristic path; optimizing a heuristic path generated by an artificial potential field method by using a DDQN neural network, selecting a flight action, and ensuring that the unmanned aerial vehicle avoids obstacles and advances towards a target; and finally, planning and adjusting the path of the unmanned aerial vehicle in real time until the unmanned aerial vehicle completes a flight task. According to the method, an artificial potential field method and a double depth Q network DDQN are combined, and the problems of path optimization and real-time obstacle avoidance of the unmanned aerial vehicle in a low-altitude dynamic complex environment are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) path planning, and in particular to a UAV dynamic path planning method based on an improved DDQN network. Background Art

[0002] The low-altitude economy, as an emerging industrial concept, holds great promise. Within this sector, drones, as the primary carriers, play a vital role. Their applications are increasingly widespread in logistics, disaster relief, agricultural monitoring, and environmental protection, and improving autonomous navigation capabilities has become a research hotspot. Drone path planning involves planning a safe, efficient, and mission-compliant path in known or unknown environments, guiding the drone from its starting point to its destination. The core challenge lies in rapidly generating a path and adjusting it in real time to accommodate changes in obstacles and targets in a dynamic environment. Therefore, path planning technology is a key area of drone autonomy and is attracting increasing attention.

[0003] Traditional path planning methods can effectively handle path planning problems in static environments. However, these algorithms require modeling the entire environment and cannot update the path in a timely manner when the environment changes dynamically, making them difficult to meet real-time requirements. In recent years, with the development of deep learning technology, deep reinforcement learning (DRL) has gradually become a research hotspot. Deep reinforcement learning learns optimal strategies by simulating the interaction between drones and the environment, making it particularly suitable for solving path planning problems in high-dimensional and dynamic environments. Artificial potential fields (APF), a classic path planning method, guide drones to the target point and avoid obstacles by constructing attractive and repulsive potential fields. APF is a common method used in early research on dynamic path planning due to its simple calculations and fast planning speed. However, APF is prone to falling into local optimality and uneven paths, and exhibits significant limitations in complex scenarios.

[0004] Therefore, to address the limitations of existing technologies in dynamic environments, a hybrid path planning algorithm combining an artificial potential field method with a dual deep Q-network (APF-DDQN) has been proposed, which has become a research hotspot in the field of UAV path planning. By deeply integrating the advantages of APF and DQN, this invention solves the path planning problem of UAVs in dynamic and complex scenarios, providing technical support for their widespread application in low-altitude economic and highly complex mission environments. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned existing technologies and provide a UAV dynamic path planning method based on an improved DDQN network to solve the path quality and local optimality problems faced by UAVs when performing path planning in low-altitude complex scenes.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is: a method for dynamic path planning of unmanned aerial vehicles based on an improved DDQN network, comprising:

[0007] Initialize the environment and state of the UAV's mission area, and use the state vector to describe the environment information and UAV state;

[0008] The artificial potential field method is used to calculate the combined force of the target point and obstacles on the UAV during its movement, calculate the UAV's movement direction, and generate a heuristic path.

[0009] The DDQN neural network is used to optimize the heuristic path generated by the artificial potential field method and select flight actions to ensure that the drone avoids obstacles and moves towards the target;

[0010] Perform real-time planning and adjustment of the drone's path until the drone completes its flight mission.

[0011] Furthermore, the initialization operation of the environment and state of the drone's mission area and the use of a state vector to describe the environment information and the drone's state include:

[0012] At the beginning of the path planning task, the environment information and the UAV status are initialized, and a static map of the area where the UAV is performing the planned task, as well as the starting point and end point information, is obtained; an environment model is constructed based on the UAV's initial position, target position, and the distribution of obstacles in the environment; and information including the UAV's current position, target position, and obstacle distribution is obtained;

[0013] The state vector is used to describe the environmental information and the drone state, as shown in the following formula:

[0014] s={p,v,p goal ,{p obs},F total}

[0015] Among them, p, v, p goal 、{p obs} respectively represent the initial position, initial speed, target point position and obstacle position set obtained by the UAV; F total It represents the resultant force on the UAV caused by the target point and obstacles during its motion, calculated by the artificial potential field method. It is used to guide the preliminary path planning.

[0016] Furthermore, the method of using the artificial potential field method to calculate the resultant force exerted on the UAV by the target point and obstacles during its movement, and to calculate the movement direction of the UAV to generate a heuristic path includes:

[0017] Calculate the target attraction of the drone, which is used to guide the drone to move towards the target point;

[0018] Calculate the total repulsive force on the drone to ensure it stays away from obstacles;

[0019] Calculate the resultant force on the drone from the target point and obstacles during its motion and determine the motion state of the drone;

[0020] F total =λ att F att +λ rep F rep

[0021]

[0022] Among them, F total F is the resultant force exerted on the UAV by the target point and obstacles during its movement, and is used to control the flight direction of the UAV; att F is the attraction of the drone; rep is the total repulsive force on the drone; att is the weight coefficient of attraction, which is used to control the intensity of attraction; rep is the weight coefficient of the repulsive force, which is used to control the strength of the repulsive force; the two weight coefficients λ att ,λ rep Changes with the motion state of the drone;

[0023] The movement direction of the drone is calculated based on the resultant force acting on the drone and a heuristic action selection is performed.

[0024] Furthermore, the calculating of the movement direction of the UAV and performing heuristic action selection includes:

[0025] Determine the direction of motion of the drone by normalizing the resultant force vector acting on the drone;

[0026] Map the calculated motion direction to a set of discrete actions and select the action that is closest to the calculated motion direction:

[0027]

[0028] Among them, a APF is the heuristic path action recommended by the artificial potential field; is a set of discrete actions of the drone, including combinations of different directions and speeds; a j Indicates the current drone's flight action; d move is the direction of movement of the drone.

[0029] Furthermore, the optimization of the heuristic path generated by the artificial potential field method using the DDQN neural network includes:

[0030] Step S1: Obtain the state vector s describing the environment information and the state of the drone;

[0031] Step S2: Design the DDQN neural network structure; the DDQN neural network structure includes an online network and a target network. The DDQN neural network structure solves the problem of Q value overestimation by separating the selection and updating of Q values;

[0032] Step S3: Drone action selection; the DDQN neural network selects the optimal action using the ε-greedy strategy, striking a balance between exploring new paths and utilizing the current optimal path;

[0033] Step S4: Design a reward function to encourage the drone to move towards the target, avoid obstacles, and take a smooth path;

[0034] Step S5: Designing a Q value function; the Q value function includes the Q value of the online network and the Q value of the target network;

[0035] Step S6: Update the Q value; the Q value of the online network is updated by the target Q value, and the mean square error is used to minimize the difference between the Q value of the online network and the target Q value;

[0036] Step S7: Repeat steps S1-S6 to train the DDQN neural network; use the flight maneuvers selected by the DDQN neural network to ensure that the drone avoids obstacles and moves towards the target.

[0037] Furthermore, the formula for the DDQN neural network to select the optimal action through the ε-greedy strategy is:

[0038]

[0039] Among them, π(a|s enhanced ) means that in a given state s enhanced Next, the strategy probability of selecting action a; Q online (s enhanced ,a) represents the online network for state s enhanced and the Q-value estimate of action a; represents the action with the largest Q value; p APF represents the probability of the action guided by the artificial potential field method; the action selection of the UAV is based on pAPF The probability of selecting heuristic action a APF ; 1-ε-p APF The optimal action predicted by the online network is selected with probability ε; random exploration is performed with probability ε.

[0040] Furthermore, the reward function is designed based on the target attraction, the obstacle repulsion, the target proximity reward and the collision penalty, as shown in the following formula:

[0041]

[0042] △d g =||p goal -p||-||p goal -p new ||

[0043] Among them, R is the immediate reward obtained by the drone after performing the current action; U att is the target gravitational potential field function; k att is the attraction coefficient; ||pp goal || is the distance between the UAV and the target; ||p goal -p new ||Mark value indicates the new position p of the drone after executing the action new and the target position p goal the distance between them; represents the repulsive potential field function of the obstacle; k rep is the repulsive force coefficient; △d g represents the target proximity reward; C c represents the collision penalty; α, β, γ, and η represent the weights of the reward values, and the action selection during the flight of the drone is controlled by adjusting the weights.

[0044] Furthermore, the Q value Q of the online network o (s,a APF ) Heuristic path action a recommended based on the current state vector s of the UAV and the artificial potential field APF It is obtained by approximation calculation through neural network; the calculation formula of the Q value y of the target network is as follows:

[0045]

[0046] Where γ is the discount factor; s′ is the next state vector of the drone after executing the current action; a′ is all possible actions under the next state vector s′; a′ APF Indicates the heuristic path action selected for the next state based on the artificial potential field calculation; is the Q value of the optimal action a′ selected by the target network for the next state vector s′.

[0047] Furthermore, the update formula of the Q value of the online network is:

[0048]

[0049] in, It is the loss function, which represents the loss of the online network and is used to measure the Q value Q of the current online network. o The difference between (s,a) and the Q value y of the target network; θ is the parameter of the online network, which is updated by the backpropagation algorithm to minimize this loss function; α is the learning rate, which represents the step size of the parameter update; is the gradient of the loss function with respect to the online network parameters.

[0050] Furthermore, the real-time planning and adjustment of the drone path includes:

[0051] Real-time status update, using the onboard drone sensors to detect the surrounding environment in real time, update the location of obstacles, and update the state vector that describes the current environmental information and drone status;

[0052] The updated state vector is input into the DDQN neural network to make decisions on the drone’s path;

[0053] Dynamically adjust the path. When the obstacle and target positions change, the path is adjusted in real time using the actions generated by the DDQN neural network; the new resultant force F is calculated in real time. total , and combined with the updated state information to optimize the path decision;

[0054] Adjust the obstacle avoidance strategy in real time, and adjust the weight coefficients λ of attraction and repulsion in real time according to the obstacle position detected by the sensor att and λ rep , avoid collision with obstacles;

[0055] Repeat the status update, dynamically adjust the path and adjust the obstacle avoidance strategy until the drone completes the flight mission.

[0056] The beneficial effects of the above technical solution are as follows: The present invention provides a dynamic UAV path planning method based on an improved DDQN network, combining the artificial potential field method (APF) and the Double Deep Q-Network (DDQN), solving the problems of path optimization and real-time obstacle avoidance for UAVs in low-altitude dynamic and complex environments. The APF method is also introduced to generate heuristic paths, leveraging the APF's computational simplicity and efficiency to provide initial path guidance for the DDQN. This significantly reduces the random exploration time in the initial stages of training, effectively improving the algorithm's convergence efficiency and ensuring the method's real-time and applicability.

[0057] This paper utilizes heuristic path information generated by the APF to expand the DDQN state representation. Integrating dynamic obstacle sensor data for real-time state updates solves the issues of insufficient path planning accuracy and real-time performance in dynamic and complex environments. The direction of the resultant force calculated by the APF is introduced into the state vector as an auxiliary feature, helping the DDQN find the target path more quickly in sparse reward environments. Furthermore, by dynamically adjusting the weights of the heuristic action selection, this paper ensures that the DDQN learning strategy is fully dominant in the later stages of training, further improving the global optimality of the planned path.

[0058] This paper, based on a heuristic action selection strategy, combines the path suggestions provided by an artificial potential field with the reinforcement learning strategy of DDQN. This approach promotes rapid network convergence through reward function design. Under the same environment, the APF-DDQN combination achieves faster convergence and higher path planning performance than either DDQN or traditional artificial potential field methods alone. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is an overall flow chart of a dynamic UAV path planning method based on an improved DDQN network provided by an embodiment of the present invention;

[0060] Figure 2 is a schematic diagram of the interaction between a drone and an environment provided by an embodiment of the present invention;

[0061] Figure 3 Schematic diagram of obstacle avoidance using an artificial potential field method according to an embodiment of the present invention;

[0062] Figure 4 This is a schematic diagram of the APF-DDQN neural network training provided by an embodiment of the present invention;

[0063] Figure 5 It is a schematic diagram of the online network and target network structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0065] In this embodiment, a dynamic UAV path planning method based on an improved DDQN network is as follows: Figure 1 As shown, the following steps are included:

[0066] Step 1: Initialize the environment and status of the UAV's mission area;

[0067] Step 1.1: At the beginning of the path planning task, initialize the environment information and the UAV status, obtain the static map of the area where the UAV is performing the planning task, as well as the starting point, end point and other information; obtain the current position of the UAV, the target position and the obstacle distribution information, and build the environment model based on the current position of the UAV, the target position and the distribution information of the obstacles in the environment, such as Figure 2 As shown:

[0068] p=(x,y,z) (1)

[0069] v=(v x ,v y ,v z ) (2)

[0070] p goal =(x goal ,y goal ,z goal ) (3)

[0071]

[0072] Among them, p, v, p goal 、{p obs} respectively represent the current position, current speed, target point position and obstacle position set obtained by the drone; (x, y, z) are the current position coordinates obtained by the drone; (v x ,v y ,v z ) is the current velocity component obtained by the drone; (x goal ,y goal ,z goal ) is the target point position coordinate obtained by the UAV, (x obs,n ,y obs,n ,z obs,n ) is the position coordinate of the nth obstacle obtained by the drone;

[0073] Step 1.2: Use the state vector to describe the environment information and the drone state:

[0074] s={p,v,p goal ,{p obs},F total} (5)

[0075] Among them, F total Represents the resultant force on the UAV from the target point and obstacles during its motion, calculated by the artificial potential field method. It is used to guide preliminary path planning.

[0076] Step 2: Use the artificial potential field method to calculate the combined force on the drone caused by the target point and obstacles during its movement, calculate the direction of movement of the drone, and generate a heuristic path to ensure that the drone moves towards the target and avoids obstacles, such as Figure 3 shown.

[0077] Step 2.1: Calculate the target attraction of the drone, which is used to guide the drone to move towards the target point;

[0078] The target attraction of a drone is given by the following formula:

[0079] F att =-k att (p c -p goal ) (6)

[0080] Among them, F att is the attraction vector; k att is the attraction gain coefficient, which controls the intensity of attraction; p c Represents the current position of the drone; the direction of attraction points to the target point, and the magnitude is proportional to the distance between the drone and the target point; when the drone approaches the target point, the attraction gradually decreases; when it reaches the target point, the attraction is zero, ensuring that the drone stays stable.

[0081] Step 2.2: Calculate the total repulsive force on the drone to ensure it stays away from obstacles.

[0082] The repulsive force of the i-th obstacle on the drone is shown in the following formula:

[0083]

[0084] Among them, F rep,i is the repulsive force of the i-th obstacle, k rep is the repulsive force gain coefficient, d obs is the impact range of the obstacle; p obs,i is the position of the i-th obstacle; the repulsive force will take effect only when the distance between the obstacle and the drone is within the influence range.

[0085] When there are multiple obstacles in the environment, the total repulsive force on the drone is the vector sum of the repulsive forces of all obstacles, as shown in the following formula:

[0086]

[0087] Among them, F rep is the total repulsive force on the drone;

[0088] Step 2.3: Calculate the resultant force on the drone from the target point and obstacles during its motion and determine the drone's motion state;

[0089]

[0090] Among them, F total λ is the resultant force exerted on the drone by the target point and obstacles during its movement, which is used to control the flight direction of the drone; att is the weight coefficient of attraction, which is used to control the intensity of attraction; rep is the weight coefficient of the repulsive force, which is used to control the strength of the repulsive force;

[0091] Among them, as the UAV's motion state changes, the two weight coefficients are also adjusted accordingly; when the UAV approaches an obstacle, the weight of the repulsive force λ rep Increase, give priority to avoiding obstacles; when λ rep >λ att When , the direction of the resultant force will be more affected by the repulsive force, ensuring that the drone stays away from obstacles; when the drone is away from obstacles, the weight of the attractive force λ att Increase to ensure that the UAV moves towards the target point; at this time, λ rep <λ att ,The UAV is mainly guided by the target attraction;

[0092] Step 2.4: Calculate the UAV's movement direction and perform heuristic action selection;

[0093] By normalizing the resultant force vector on the drone, the direction of motion d of the drone is determined move , as shown in the following formula:

[0094]

[0095] The calculated motion direction d move Mapping to a discrete set of actions and select the action that is closest to the calculated motion direction:

[0096]

[0097] Among them, a APF is the heuristic path action recommended by the artificial potential field; is a set of discrete actions of the drone, including combinations of different directions and speeds; a j Indicates the current drone's flight actions;

[0098] Step 3: Use the DDQN neural network to optimize the heuristic path generated by the artificial potential field method;

[0099] Step 3.1: Obtain the state vector s that describes the environment information and the drone state obtained in step 1.2;

[0100] Step 3.2: Design the DDQN neural network structure; Figure 4 As shown in Figure 2, the DDQN neural network structure includes an online network and a target network. The DDQN neural network structure solves the problem of overestimation of Q values by separating the selection and updating of Q values. The online network and the target network have the same structure, including an input layer, a hidden layer, and an output layer. Figure 5 shown.

[0101] Step 3.3: Drone action selection; the DDQN neural network selects the optimal action using the ε-greedy strategy, striking a balance between exploring new paths and utilizing the current optimal path;

[0102] The formula for the DDQN neural network to select the optimal action through the ε-greedy strategy is:

[0103]

[0104] Among them, π(a|s enhanced ) means that in a given state s enhanced Next, the strategy probability of selecting action a; Q online (s enhanced ,a) represents the online network for state s enhanced and the Q-value estimate of action a; represents the action with the largest Q value; p APF represents the probability of the action guided by the artificial potential field method; the action selection of the UAV is based on p APF The probability of selecting heuristic action a APF ; 1-ε-p APF The probability of selecting the optimal action predicted by the online network is ε; random exploration is performed with the probability of ε; p APF It is high in the initial stage and gradually decreases to ensure that the neural network structure can cover a wider state space.

[0105] Step 3.4: Design a reward function. Design a reward function that encourages the drone to move toward the goal, avoid obstacles, and take a smooth path.

[0106] The reward function is designed based on the target attraction, obstacle repulsion, target proximity reward and collision penalty, as shown in the following formula:

[0107]

[0108] △d g =||p goal -p||-||p goal -pnew || (18)

[0109] Among them, R is the immediate reward obtained by the drone after performing the current action; U att is the target gravitational potential field function; k att is the attraction coefficient; ||pp goal || is the distance between the UAV and the target; ||p goal -p new ||Mark value indicates the new position p of the drone after executing the action new and the target position p goal the distance between them; represents the repulsive potential field function of the obstacle; k rep is the repulsive force coefficient; △d g represents the target proximity reward; C c represents the collision penalty; α, β, γ, and η represent the weights of the reward values, and the action selection during the flight of the drone is controlled by adjusting the weights;

[0110] Step 3.5: Design a Q-value function; the Q-value function includes the Q-value of the online network and the Q-value of the target network;

[0111] The Q value of the online network is Q o (s,a APF ) The heuristic path action a recommended based on the current state vector s of the UAV and the artificial potential field calculated in step 2.4 APF It is obtained by approximation calculation through neural network; the calculation formula of the Q value y of the target network is as follows:

[0112]

[0113] Where γ is the discount factor; s′ is the next state vector of the drone after executing the current action; a′ is all possible actions under the next state vector s′; a′ APF Indicates the heuristic path action selected for the next state based on the artificial potential field calculation; is the Q value of the optimal action a′ selected by the target network for the next state vector s′;

[0114] Step 3.6: Update the Q value; the Q value of the online network is updated by the target Q value, and the mean square error is used to minimize the difference between the Q value of the online network and the target Q value;

[0115] The update formula of the Q value of the online network is:

[0116]

[0117]

[0118] in, It is the loss function, which represents the loss of the online network and is used to measure the Q value Q of the current online network. o The difference between (s,a) and the Q value y of the target network; θ is the parameter of the online network, which is updated by the back-propagation algorithm to minimize this loss function;

[0119] Formula (19) is the update formula of the back propagation algorithm, where α is the learning rate, which represents the step size of parameter update; is the gradient of the loss function with respect to the online network parameters; by minimizing the loss, the parameters θ of the online network will be updated so that the Q value of the online network is closer to the Q value y of the target network, thereby improving the accuracy and convergence speed of path planning.

[0120] Step 3.7: Repeat steps 3.1-3.6 to train the DDQN neural network. Use the flight maneuvers selected by the DDQN neural network to ensure that the drone efficiently avoids obstacles and moves toward the target.

[0121] Step 4: Real-time planning and adjustment of the drone’s path;

[0122] Step 4.1: Real-time status update, using the onboard drone sensor to detect the surrounding environment in real time and update the obstacle position {p obs}, and update the state vector s that currently describes the environment information and the drone state c :

[0123] s c ={p,v,p goal ,{p obs},F total} (twenty two)

[0124] The updated state vector s c Input into the DDQN neural network to make decisions on the drone’s path;

[0125] Step 4.2: Dynamically adjust the path. When the obstacle and target positions change, use the actions generated by the DDQN neural network to adjust the path in real time; calculate the new resultant force F in real time using formula (9) total , and combined with the updated state information to optimize the path decision;

[0126] Step 4.3: Adjust the obstacle avoidance strategy in real time. According to the obstacle position detected by the sensor, adjust the weight coefficients λ of the attraction and repulsion forces in real time according to formulas (10) and (11). att and λ rep , avoid collision with obstacles;

[0127] Step 5: Repeat steps 2 to 4 until the drone completes its flight mission.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for dynamic path planning of unmanned aerial vehicles based on an improved DDQN network, characterized by: include: Initialize the environment and state of the UAV's mission area, and use the state vector to describe the environment information and UAV state; The artificial potential field method is used to calculate the combined force of the target point and obstacles on the UAV during its movement, calculate the UAV's movement direction, and generate a heuristic path. The DDQN neural network is used to optimize the heuristic path generated by the artificial potential field method and select flight actions to ensure that the drone avoids obstacles and moves towards the target; Perform real-time planning and adjustment of the drone's path until the drone completes its flight mission.

2. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 1, characterized in that: The initialization operation of the environment and state of the UAV's mission area and the use of a state vector to describe the environment information and the UAV state include: At the beginning of the path planning task, the environment information and the UAV status are initialized, and a static map of the area where the UAV is performing the planned task, as well as the starting point and end point information, is obtained; an environment model is constructed based on the UAV's initial position, target position, and the distribution of obstacles in the environment; and information including the UAV's current position, target position, and obstacle distribution is obtained; The state vector is used to describe the environmental information and the drone state, as shown in the following formula: s={p,v,p goal ,{p obs },F total } Among them, p, v, p goal 、{p obs } respectively represent the initial position, initial speed, target point position and obstacle position set obtained by the UAV; F total It represents the resultant force on the UAV caused by the target point and obstacles during its motion, calculated by the artificial potential field method. It is used to guide the preliminary path planning.

3. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 2, characterized in that: The method of using the artificial potential field to calculate the resultant force exerted on the UAV by the target point and obstacles during its movement, and to calculate the movement direction of the UAV and generate a heuristic path includes: Calculate the target attraction of the drone, which is used to guide the drone to move towards the target point; Calculate the total repulsive force on the drone to ensure it stays away from obstacles; Calculate the resultant force on the drone from the target point and obstacles during its motion and determine the motion state of the drone; F total =λ att F att +λ rep F rep Among them, F total F is the resultant force exerted on the UAV by the target point and obstacles during its movement, and is used to control the flight direction of the UAV; att F is the attraction of the drone; rep is the total repulsive force on the drone; att is the weight coefficient of attraction, which is used to control the intensity of attraction; rep is the weight coefficient of the repulsive force, which is used to control the strength of the repulsive force; the two weight coefficients λ att ,λ rep Changes with the motion state of the drone; The movement direction of the drone is calculated based on the resultant force acting on the drone and a heuristic action selection is performed.

4. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 3, characterized in that: The calculation of the UAV's motion direction and the heuristic action selection include: Determine the direction of motion of the drone by normalizing the resultant force vector acting on the drone; Map the calculated motion direction to a set of discrete actions and select the action that is closest to the calculated motion direction: Among them, a APF is the heuristic path action recommended by the artificial potential field; is a set of discrete actions of the drone, including combinations of different directions and speeds; a j Indicates the current drone's flight action; d move is the direction of movement of the drone.

5. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 4, characterized in that: The method of optimizing the heuristic path generated by the artificial potential field method using the DDQN neural network includes: Step S1: Obtain the state vector s describing the environment information and the state of the drone; Step S2: Design the DDQN neural network structure; the DDQN neural network structure includes an online network and a target network. The DDQN neural network structure solves the problem of Q value overestimation by separating the selection and updating of Q values; Step S3: Drone action selection; the DDQN neural network selects the optimal action using the ε-greedy strategy, striking a balance between exploring new paths and utilizing the current optimal path; Step S4: Design a reward function to encourage the drone to move towards the target, avoid obstacles, and take a smooth path; Step S5: Designing a Q value function; the Q value function includes the Q value of the online network and the Q value of the target network; Step S6: Update the Q value; the Q value of the online network is updated by the target Q value, and the mean square error is used to minimize the difference between the Q value of the online network and the target Q value; Step S7: Repeat steps S1-S6 to train the DDQN neural network; use the flight maneuvers selected by the DDQN neural network to ensure that the drone avoids obstacles and moves towards the target.

6. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 5, characterized in that: The formula for the DDQN neural network to select the optimal action through the ε-greedy strategy is: Among them, π(a|s enhanced ) means that in a given state s enhanced Next, the strategy probability of selecting action a; Q online (s enhanced ,a) represents the online network for state s enhanced and the Q-value estimate of action a; represents the action with the largest Q value; p APF represents the probability of the action guided by the artificial potential field method; the action selection of the UAV is based on p APF The probability of selecting heuristic action a APF ; 1-ε-p APF The optimal action predicted by the online network is selected with probability ε; random exploration is performed with probability ε.

7. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 6, characterized in that: The reward function is designed based on the target attraction, obstacle repulsion, target proximity reward and collision penalty, as shown in the following formula: △d g =||p goal -p||-||p goal -p new || Among them, R is the immediate reward obtained by the drone after performing the current action; U att is the target gravitational potential field function; k att is the attraction coefficient; ||pp goal || is the distance between the UAV and the target; ||p goal -p new ||Mark value indicates the new position p of the drone after executing the action new and the target position p goal the distance between them; represents the repulsive potential field function of the obstacle; k rep is the repulsive force coefficient; △d g represents the target proximity reward; C c represents the collision penalty; α, β, γ, and η represent the weights of the reward values, and the action selection during the flight of the drone is controlled by adjusting the weights.

8. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 7, characterized in that: The Q value Q of the online network o (s,a APF ) Heuristic path action a recommended based on the current state vector s of the UAV and the artificial potential field APF It is obtained by approximation calculation through neural network; the calculation formula of the Q value y of the target network is as follows: Where γ is the discount factor; s′ is the next state vector of the drone after executing the current action; a′ is all possible actions under the next state vector s′; a′ APF Indicates the heuristic path action selected for the next state based on the artificial potential field calculation; is the Q value of the optimal action a′ selected by the target network for the next state vector s′.

9. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 8, characterized in that: The update formula of the Q value of the online network is: in, It is the loss function, which represents the loss of the online network and is used to measure the Q value Q of the current online network. o The difference between (s,a) and the Q value y of the target network; θ is the parameter of the online network, which is updated by the backpropagation algorithm to minimize this loss function; α is the learning rate, which represents the step size of the parameter update; is the gradient of the loss function with respect to the online network parameters.

10. The method for dynamic path planning of an unmanned aerial vehicle based on an improved DDQN network according to claim 9, characterized in that: The real-time planning and adjustment of the drone path includes: Real-time status update, using the onboard drone sensors to detect the surrounding environment in real time, update the location of obstacles, and update the state vector that describes the current environmental information and drone status; The updated state vector is input into the DDQN neural network to make decisions on the drone’s path; Dynamically adjust the path. When the obstacle and target positions change, the path is adjusted in real time using the actions generated by the DDQN neural network; the new resultant force F is calculated in real time. total , and combined with the updated state information to optimize the path decision; Adjust the obstacle avoidance strategy in real time, and adjust the weight coefficients λ of attraction and repulsion in real time according to the obstacle position detected by the sensor att and λ rep , avoid collision with obstacles; Repeat the status update, dynamically adjust the path and adjust the obstacle avoidance strategy until the drone completes the flight mission.