A three-dimensional dynamic path planning method for unmanned aerial vehicles based on adaptive dynamic programming
By adopting an adaptive dynamic planning method in drone path planning, the cost function and utility function are redefined, and the BP neural network is used for approximate solution, the dynamic path planning problem of multi-UAV in complex three-dimensional environments is solved, and a safer and more efficient path planning is achieved.
Patent Information
- Application Number
- CN202210789966.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-07-05
AI Technical Summary
The existing drone path planning methods are difficult to effectively solve the dynamic path planning problem of multiple drones in three-dimensional space in complex environments, especially when obstacles exist, it is difficult to ensure the safety and efficiency of the path.
Adaptive dynamic programming method is adopted to redefine the cost function and utility function, and design the dynamic path planning method of multi-UAV in three-dimensional space. By establishing a mathematical model of the outer ring of the drone position and linearizing the model, the approximate solution of the execution network and the evaluation network is achieved in combination with the BP neural network.
Dynamic path planning of multiple drones in complex three-dimensional environments is realized, which improves the safety and efficiency of the path, can effectively avoid obstacles and optimize the path.
Smart Images

Figure CN115328190B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle path planning, and in particular to a three-dimensional dynamic path planning method for an unmanned aerial vehicle based on adaptive dynamic planning. Background Art
[0002] Compared with manned aircraft, UAVs have significant features such as low cost, small size, high flexibility and good adaptability, which has expanded their applications in various fields and occupied an irreplaceable position. Path planning is regarded as an important part of UAV research and is of great significance to the safety and mission efficiency of UAVs.
[0003] The essence of path planning is to find a collision-free route that can safely perform tasks. When a drone performs an area coverage mission, the planned path needs to cover the mission area as much as possible in the shortest time and at the lowest cost. When a drone searches for a target in a complex environment and is required to reach a specified location, it needs to find a safer and shorter route. Therefore, different tasks performed by drones have different requirements for path planning. Many methods have been proposed for the problem of drone path planning, including the artificial potential field method using "virtual force", the A* algorithm, and intelligent algorithms such as the ant colony algorithm, particle swarm algorithm, genetic algorithm, and pigeon colony algorithm based on the bionic behavior of biological groups in nature.
[0004] In addition, the dynamic programming (DP) algorithm has also achieved certain results in the research of path planning due to its optimality and good adaptability. However, the solution of the Hamilton-Jacobi-Bellman (HJB) equation is very difficult, and when the scale of the problem to be solved increases, the "curse of dimensionality" problem will appear. Therefore, Werbos proposed the adaptive dynamic programming (ADP) algorithm, which has been applied in the control of UAVs. However, the application research of the adaptive dynamic programming algorithm in path planning and UAVs is relatively small and not comprehensive, especially in the path planning of UAVs. In the existing research results, the research on ADP in path planning is mostly in the two-dimensional plane. Summary of the invention
[0005] Purpose of the invention: In view of the problems existing in the above-mentioned background technology, the present invention provides a three-dimensional dynamic path planning method for UAVs based on adaptive dynamic programming, redefines the cost function and utility function, and proposes a design method for dynamic path planning of multiple UAVs in three-dimensional space based on an adaptive dynamic programming algorithm.
[0006] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is:
[0007] A three-dimensional dynamic path planning method for an unmanned aerial vehicle based on adaptive dynamic planning comprises the following steps:
[0008] Step S1, establishing a mathematical model of the outer ring of the drone position, and performing linearization processing on the outer ring model of the drone position; performing regularization processing on obstacles with irregular shapes detected by the drone;
[0009] Step S2: designing an adaptive dynamic programming method based on the linearized model; the adaptive dynamic programming method includes designing a utility function, a cost function, an execution network, and an evaluation network;
[0010] Step S3: Plan the path of the UAV in three-dimensional space based on the adaptive dynamic planning method designed in step S2, obtain the planning scheme and test it.
[0011] Furthermore, the mathematical model of the outer loop of the drone position in step S1 is established as follows:
[0012]
[0013]
[0014]
[0015] where x i ,y i and z i represents the position of the i-th UAV in the inertial coordinate system, V i , γ i and χ i represents the speed, climb angle and heading angle of the ith UAV respectively; there is no requirement for the attitude angle, and γ i and χ i As a constant, the drone position is only related to the acceleration; the outer loop of the drone position is linearized as follows:
[0016]
[0017]
[0018] Where P i (x i ,y i ,z i ) is the position coordinate of the i-th UAV, V i (v ix ,v iy ,v iz) represents the speed of the ith UAV, u i (u ix ,u iy ,u iz ) represents the acceleration of the i-th UAV in the inertial coordinate system.
[0019] Furthermore, in step S1, the irregularly shaped obstacle is specially regularized and the obstacle is regarded as a sphere.
[0020] Furthermore, the utility function designed in step S2 is as follows:
[0021] U*x(t),u(t))=U obj (k)+U obs (k)+U ij (k)+u′(k)u(k)
[0022] Among them, U obj (k) and U obs (k) represents the utility function between the UAV and the target point, and between the UAV and the obstacle, respectively. ij (k) is the collision avoidance utility function between UAVs, and u(k) is the control input;
[0023] The target utility function U reflects the relative distance between the UAV and the target point obj (k) The expression is as follows:
[0024]
[0025]
[0026] Among them, (x i ,y i ,z i ) represents the current position of UAV i, (x obj ,y obj ,z obj ) represents the position of the current target point, ξ1 is the target utility coefficient;
[0027] The target utility function U reflects the relative distance between the UAV and the obstacle obs (k) The expression is as follows:
[0028]
[0029]
[0030]
[0031] in, and They represent the distance between the geometric center of the multi-UAVs and the obstacle center and the distance between UAV i and the geometric center respectively; R threat is the threat radius of the regularized obstacle; (x obs ,y obs ,z obs ) and (x vir ,y vir ,z vir ) represent the center position of the obstacle sphere and the geometric center position of the multi-UAV formation respectively; ξ2 is the obstacle avoidance utility coefficient;
[0032] Detect unknown obstacles in the environment and obtain the position of the point on the obstacle closest to the drone; set the maximum detection distance L of the onboard sensor max ; then when When ξ1=0 and ξ2=1; when When ξ1=1 and ξ2=0;
[0033] The collision avoidance utility function U that reflects the relative distance between UAVs ij (k) The expression is as follows:
[0034]
[0035] U ij (k) = ξ3(d ij -d min )
[0036] Among them, d min is the minimum distance allowed between UAVs, ξ3 is the collision avoidance utility coefficient; when d ij <d min , ξ3=1, otherwise ξ3=0.
[0037] Furthermore, the cost function designed in step S2 is specifically as follows:
[0038]
[0039] Where γ is the discount factor.
[0040] Furthermore, in step S2, both the execution network and the evaluation network are implemented using BP neural network;
[0041] The outputs of the execution network and the evaluation network are as follows:
[0042] u(k)=Ah out (k)·Wa2(k)
[0043]
[0044] Among them, u(k) represents the output of the execution network, Represents the estimated cost function obtained by the evaluation network; Ah out (k) and Ch out (k) are the outputs of the hidden layers of the execution network and the evaluation network respectively; Wa2(k) and Wc2 are the weight matrices between the hidden layers and the output layers of the execution network and the evaluation network respectively.
[0045] Furthermore, the execution network adopts a 3-layer BP neural network, in which the number of input neurons is the same as the number of system states, and the number of output neurons is the same as the number of system control inputs; the execution network also includes a number of hidden layer neurons; the input of the execution network is expressed as follows:
[0046] In a (k) = x(k)
[0047] The input and output of the hidden layer are calculated as follows:
[0048] Ah in (k)=In a (k)·Wa1(k)
[0049]
[0050] The output of executing the network is represented as follows:
[0051] u(k)=Ah out (k)·Wa2(k)
[0052] Where Wa1 and Wa2 are the weights between the input layer and the hidden layer and the weights between the hidden layer and the output layer respectively; the execution network error is defined as follows:
[0053]
[0054]
[0055] Among them U c is the expected utility function; let U c =0, the execution network is trained using the gradient descent method, and the network weights are updated using the following formula:
[0056] w a (k+1)=w a (k)+Δw a (k)
[0057]
[0058] The evaluation network also uses a 3-layer BP network, where the number of input neurons is the number of system states and control inputs, the number of output neurons is one, and the number of hidden layer neurons is several. The input of the evaluation network at time k is expressed as follows:
[0059] in c (k)=[x(k),u(k)]
[0060] The input of the hidden layer is Ch in (k) and output Ch out (k) are calculated by the following formula:
[0061] Ch in (k)=In c (k)·Wc1(k)
[0062]
[0063] The output of the evaluation network is expressed as follows:
[0064]
[0065] Among them, Wc1 and Wc2 are the weights between the input layer and the hidden layer and the weights between the hidden layer and the output layer respectively; the gradient descent method is used to train the neural network and the evaluation network error is defined as follows:
[0066]
[0067]
[0068] Use the following formula to update the weights:
[0069] w c (k+1)=w c (k)+Δw c (k)
[0070]
[0071] Beneficial effects:
[0072] The present invention linearizes the outer loop of the position of the UAV, and on this basis designs a dynamic path planning method for multiple UAVs in three-dimensional space based on an adaptive dynamic programming algorithm. Among them, new cost functions and utility functions are defined, and neural networks are used for approximate solutions. On the one hand, the path planning problem of the UAV flying to the specified location is solved by the redefined target utility function. On the other hand, the obstacle avoidance utility function realizes the obstacle avoidance problem when the UAV detects the presence of obstacles in the flight process. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 This is a schematic diagram of the special treatment of obstacles by the present invention;
[0074] Figure 2 A design block diagram of the three-dimensional dynamic path planning method for unmanned aerial vehicles provided by the present invention;
[0075] Figure 3 A network structure diagram for executing the three-dimensional dynamic path planning method for unmanned aerial vehicles provided by the present invention;
[0076] Figure 4 A network structure diagram for evaluating the three-dimensional dynamic path planning method for unmanned aerial vehicles provided by the present invention;
[0077] Figure 5 The three-dimensional dynamic path planning result diagram of the UAV provided by the present invention;
[0078] Figure 6 The control input for the drone. DETAILED DESCRIPTION
[0079] The present invention is further described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0080] The present invention provides a method for three-dimensional dynamic path planning of an unmanned aerial vehicle based on adaptive dynamic planning, which specifically comprises the following steps:
[0081] Step S1: Establish a mathematical model of the outer ring of the drone position, and perform linearization on the outer ring model of the drone position. Perform regularization on irregularly shaped obstacles detected by the drone, and after processing, treat the obstacles as a sphere.
[0082] First, the mathematical model of the outer ring of the drone position is established as follows:
[0083]
[0084]
[0085]
[0086] where x i ,y i and z i represents the position of the i-th UAV in the inertial coordinate system, V i , γ i and χ i represents the speed, climbing angle and heading angle of the ith UAV respectively. There is no requirement for the attitude angle, so γi and χ i If we regard it as a constant, the drone position is only related to the speed. The outer loop of the drone position is linearized as follows:
[0087]
[0088]
[0089] Where P i (x i ,y i ,z i ) is the position coordinate of the i-th UAV, V i (v ix ,v iy ,v iz ) represents the speed of the ith UAV, u i (u ix ,u iy ,u iz ) represents the acceleration of the i-th UAV in the inertial coordinate system.
[0090] During the flight of the drone, unexpected obstacles may be encountered. The present invention makes special treatment for these obstacles, that is, the obstacles are regarded as spheres, and the central cross-section diagram is as follows: Figure 1 As shown. Among them, O obs represents the center of the obstacle sphere, R obs is the radius of the obstacle sphere, R threat Indicates the radius of the threat area created by the obstacle.
[0091] Step S2: Based on the linearized model, an adaptive dynamic programming method is designed. The adaptive dynamic programming method includes designing a utility function, a cost function, an execution network, and an evaluation network. The utility function includes a target utility function, an obstacle avoidance utility function, and a collision avoidance utility function, which respectively solve the path planning, obstacle avoidance, and collision avoidance problems of multiple UAVs.
[0092] According to the linearized model in step 1, the specific details of the adaptive dynamic programming algorithm are designed. Its structure diagram is as follows: Figure 2 shown.
[0093] First, design the utility function as follows:
[0094] U(x(t),u(t))=U obj (k)+U obs (k)+Uij(k)+u′(k)u(k)
[0095] Among them, U obj (k) and U obs(k) represents the utility function between the UAV and the target point, and between the UAV and the obstacle, respectively. ij (k) is the collision avoidance utility function between UAVs, and u(k) is the control input.
[0096] The target utility function U reflects the relative distance between the UAV and the target point obj (k) The expression is as follows:
[0097]
[0098]
[0099] Among them, (x i ,y i ,z i ) represents the current position of UAV i, (x obj ,y obj ,z obj ) represents the current target point position, and ξ1 is the target utility coefficient.
[0100] The target utility function U reflects the relative distance between the UAV and the obstacle obs (k) The expression is as follows:
[0101]
[0102]
[0103]
[0104] in, and They represent the distance between the geometric center of the multi-UAVs and the obstacle center and the distance between UAV i and the geometric center. threat is the threat radius of the regularized obstacle. (x obs ,y obs ,z obs ) and (x vir ,y vir ,z vir ) represent the center position of the obstacle sphere and the geometric center position of the multi-UAV formation respectively. ξ2 is the obstacle avoidance utility coefficient.
[0105] Detect unknown obstacles in the environment and obtain the position of the point on the obstacle closest to the drone. Set the maximum detection distance L of the onboard sensor max . Then when When ξ1=0 and ξ2=1. , ξ1=1 and ξ2=0.
[0106] The collision avoidance utility function U that reflects the relative distance between UAVs ij (k) The expression is as follows:
[0107]
[0108] U ij (k) = ξ3(d ij -d min )
[0109] Among them, d min is the minimum distance allowed between UAVs, and ξ3 is the collision avoidance utility coefficient. ij <d min , ξ3=1, otherwise ξ3=0.
[0110] Then design the cost function as follows:
[0111]
[0112] Where γ is the discount factor.
[0113] In the present invention, the execution network and the evaluation network use BP neural network to approximately solve the HJB equation and obtain the optimal control law. Specifically, it is as follows:
[0114] For the execution network, the present invention adopts a three-layer BP neural network, in which the number of input neurons is the same as the number of system states, the number of output neurons is the same as the number of system control inputs, and a number of hidden layer neurons. The structure diagram is as follows: Figure 3 As shown. Figure 2 It can be obtained that the input of the execution network is
[0115] In a (k) = x(k)
[0116] The input and output of the hidden layer are calculated as follows
[0117] Ah in (k)=In a (k)·Wa1(k)
[0118]
[0119] The output of the execution network is
[0120] u(k)=Ah out (k)·Wa2(k)
[0121] Among them, Wa1 and Wa2 are the weights between the input layer and the hidden layer and the weights between the hidden layer and the output layer respectively. The error of executing the network is defined as
[0122]
[0123]
[0124] Among them, U c is the expected utility function, and the present invention sets U c = 0. The training of the execution network is also performed using the gradient descent method. The network weights are updated using the following formula:
[0125] w a (k+1)=w a (k)+Δw a (k)
[0126]
[0127] The evaluation network also uses a three-layer BP neural network, in which the number of input neurons is the number of system states and control inputs, the number of output neurons is one and several hidden layer neurons. The structure diagram is as follows: Figure 4 As shown. Figure 2 It can be obtained that the input of the evaluation network at time k is:
[0128] In c (k)=[x(k),u(k)]
[0129] The input and output of the hidden layer are calculated as follows:
[0130] Ch in (k)=In c (k)·Wc1(k)
[0131]
[0132] Finally, the output of the evaluation network is
[0133]
[0134] Among them, Wc1 and Wc2 are the weights between the input layer and the hidden layer and the weights between the hidden layer and the output layer respectively. In order to achieve the control requirements, the gradient descent method is used to train the neural network. The following definition evaluates the network error:
[0135]
[0136]
[0137] Use the following formula to update the weights
[0138] w c (k+1)=w c (k)+Δw c (k)
[0139]
[0140] Step 3: Apply the designed adaptive dynamic programming algorithm to the path planning of multiple UAVs in three-dimensional space and conduct tests.
[0141] Based on the MATLAB / SIMULINK simulation platform, the planned paths and control inputs of multiple UAVs in three-dimensional space are obtained, such as Figure 5 and Figure 6 As shown. This verifies the design method of dynamic path planning of multiple UAVs in three-dimensional space based on the adaptive dynamic programming algorithm.
[0142] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A three-dimensional dynamic path planning method for unmanned aerial vehicles based on adaptive dynamic programming, characterized in that: The following steps are involved: Step S1, establishing a mathematical model of the outer ring of the drone position, and performing linearization processing on the outer ring model of the drone position; performing regularization processing on obstacles with irregular shapes detected by the drone; Step S2: designing an adaptive dynamic programming method based on the linearized model; the adaptive dynamic programming method includes designing a utility function, a cost function, an execution network and an evaluation network; Step S3, planning the path of the UAV in three-dimensional space based on the adaptive dynamic planning method designed in step S2, obtaining the planning scheme and testing it; The mathematical model of the outer ring of the drone position in step S1 is established as follows: where x i ,y i and z i represents the position of the i-th UAV in the inertial coordinate system, V i , γ i and χ i represents the speed, climb angle and heading angle of the i-th UAV respectively; There is no requirement for the attitude angle. i and χ i As a constant, the drone position is only related to the speed; the outer loop of the drone position is linearized as follows: Where P i (x i ,y i , z i ) is the position coordinate of the ith UAV, V i (v ix , v iy , v iz ) represents the speed of the ith UAV, u i (u ix ,u iy ,u iz ) represents the acceleration of the i-th UAV in the inertial coordinate system; The utility function designed in step S2 is as follows: U(x(t),u(t))=U obj (k)+U obs (k)+U ij (k)+u′(k)u(k) Among them, U obj (k) and U obs (k) represents the utility function between the UAV and the target point, and between the UAV and the obstacle, respectively. ij (k) is the collision avoidance utility function between UAVs, and u(k) is the control input; The target utility function U reflects the relative distance between the UAV and the target point obj (k) The expression is as follows: Among them, (x i ,y i , z i ) represents the current position of UAV i, (x obj ,y obj , z obj ) represents the position of the current target point, ξ1 is the target utility coefficient; The target utility function U reflects the relative distance between the UAV and the obstacle obs (k) The expression is as follows: in, and They represent the distance between the geometric center of the multi-UAVs and the obstacle center and the distance between UAV i and the geometric center respectively; R threat is the threat radius of the regularized obstacle; (x obs ,y obs , z obs ) and (x vir ,y vir , z vir ) represent the center position of the obstacle sphere and the geometric center position of the multi-UAV formation respectively; ξ2 is the obstacle avoidance utility coefficient; Detect unknown obstacles in the environment and obtain the position of the point on the obstacle closest to the drone; set the maximum detection distance L of the onboard sensor max ; then when When ξ1=0 and ξ2=1; when When ξ1=1 and ξ2=0; The collision avoidance utility function U that reflects the relative distance between UAVs ij (k) The expression is as follows: U ij (k)=ξ3(d ij -d min ) Among them, d min is the minimum distance allowed between UAVs, ξ3 is the collision avoidance utility coefficient; when d ij <d min , ξ3=1, otherwise ξ3=0; The cost function designed in step S2 is specifically as follows: Where γ is the discount factor; In step S2, both the execution network and the evaluation network are implemented using BP neural network; The outputs of the execution network and the evaluation network are as follows: u(k)=Ah out (k)·Wa2(k) Among them, u(k) represents the output of the execution network, Represents the estimated cost function obtained by the evaluation network; Ah out (k) and Ch out (k) are the outputs of the hidden layers of the execution network and the evaluation network, respectively; Wa2(k) and Wc2 are the weight matrices between the hidden layers and the output layers of the execution network and the evaluation network, respectively; The execution network adopts a 3-layer BP neural network, in which the number of input neurons is the same as the number of system states, and the number of output neurons is the same as the number of system control inputs; the execution network also includes a number of hidden layer neurons; the input of the execution network is expressed as follows: In a (k)=x(k) The input and output of the hidden layer are calculated as follows: Ah in (k)=In a (k)·Wa1(k) The output of the execution network is represented as follows: u(k)=Ah out (k)·Wa2(k) Where Wa1 and Wa2 are the weights between the input layer and the hidden layer and the weights between the hidden layer and the output layer respectively; the execution network error is defined as follows: Among them U c is the expected utility function; let U c =0, the execution network is trained using the gradient descent method, and the network weights are updated using the following formula: In a (k+1)=in a (k)+Δw a (k) The evaluation network also uses a 3-layer BP network, in which the number of input neurons is the number of system states and control inputs, the number of output neurons is one, and the number of hidden layer neurons is several; the input of the evaluation network at time k is expressed as follows: In c (k)=[x(k),u(k)] The input of the hidden layer is Ch in (k) and output Ch out (k) are calculated by the following formula: Ch in (k)=In c (k)·Wc1(k) The output of the evaluation network is expressed as follows: Among them, Wc1 and Wc2 are the weights between the input layer and the hidden layer and the weights between the hidden layer and the output layer respectively; the gradient descent method is used to train the neural network and the evaluation network error is defined as follows: Use the following formula to update the weights: In c (k+1)=in c (k)+Δw c (k) 2. The method for three-dimensional dynamic path planning of an unmanned aerial vehicle based on adaptive dynamic programming according to claim 1, characterized in that: In the step S1, the irregularly shaped obstacle is specially regularized and the obstacle is regarded as a sphere.
Citation Information
Patent Citations
Cooperative path planning method of multiple maneuvering objects of tracking of drone
CN106873628A
Method for planning paths of unmanned aerial vehicles on basis of Q(lambda) algorithms
CN109655066A