Unmanned Aerial Vehicle Path Planning Method Based on Inverse Reinforcement Learning

By adopting the reverse reinforcement learning method in drone path planning, combining expert experience loss and maximum entropy inverse reinforcement learning algorithm, the problem of poor path planning in traditional methods in complex environments is solved, and more efficient and stable path planning is achieved.

CN115826601BActive Publication Date: 2025-06-24NAVAL AVIATION UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211437557.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2025-06-24
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

Traditional drone path planning algorithms are difficult to deal with dynamic obstacles in complex environments, and there are subjectivity and sparse reward problems in reward function design, resulting in poor path planning effects.

Method used

Using an inverse reinforcement learning method, the DDPG algorithm with expert experience loss and the maximum entropy inverse reinforcement learning algorithm is used to build an experience pool, optimize the strategy network and solve the optimal reward function to improve the efficiency and effect of path planning.

Benefits of technology

This method can effectively plan the UAV path in complex environments, reduce exploration costs, improve convergence speed and obstacle avoidance effects, and overcome the difficulties in designing reward function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115826601B_ABST
    Figure CN115826601B_ABST
Patent Text Reader

Abstract

To solve the problems such as slow convergence speed and difficult setting of reward function when the deep deterministic policy gradient algorithm is used to plan the safe collision avoidance path of unmanned aerial vehicles (UAVs), the present invention proposes a UAV path planning method based on inverse reinforcement learning. First, a demonstration trajectory dataset of an expert operating the UAV to avoid obstacles is collected based on simulator software. Secondly, a hybrid sampling mechanism is adopted to update the network parameters by integrating high-quality expert demonstration trajectory data into the self-exploration data, so as to reduce the exploration cost of the algorithm. Finally, the optimal reward function implicit in the expert experience is solved according to the maximum entropy inverse reinforcement learning algorithm, solving the problem of difficult setting of the reward function in complex tasks. The comparative experimental results show that the present invention can effectively improve the training efficiency of the algorithm and has better collision avoidance performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of UAV path planning, and particularly relates to a UAV path planning method based on inverse reinforcement learning. Background Technique

[0002] With the further opening of the UAV (Unmanned Aerial Vehicle) field, the dense dynamic obstacles in complex environments such as cities and mountains pose a great threat to the flight safety of UAVs. Traditional path planning algorithms, such as heuristic algorithms like A* and D*, and graph theory-based visibility graph method, Voronoi diagram method, etc., can only cope with simple environments where obstacle information is known in advance. However, due to the complex and variable terrain of cities and mountains, and the difficulty in obtaining specific parameters of obstacles, the application scope of traditional obstacle avoidance algorithms is limited.

[0003] Different from the above traditional path planning methods, the navigation method based on reinforcement learning draws on the learning method of biological acquired perception development, and continuously optimizes the obstacle avoidance strategy through interaction with the environment. It not only avoids the dependence on obstacle modeling and supervised learning, but also has strong generalization ability and robustness. Especially in recent years, deep reinforcement learning has utilized the powerful perception and function fitting ability of deep learning to effectively alleviate the "exponential explosion" problem of high-dimensional environmental state space and decision space, providing new ideas for the UAV path planning problem in dense dynamic obstacle environments. The Sliver, Google DeepMind team, Dr. John Schulman of the University of California, Berkeley, and OpenAI have successively proposed deep reinforcement learning algorithms such as DDPG (Deep deterministic policy gradient), asynchronous advantage AC (Asynchronous Advantage Actor Critic, abbreviated as: A3C), trust region policy optimization (abbreviated as: TRPO), and proximal policy optimization (abbreviated as: PPO).

[0004] Although these methods have obvious advantages in UAV path planning, they often need to explore a large number of random obstacle environment samples to try new strategies and are prone to falling into local optima. Therefore, the present invention proposes a DDPG algorithm that fuses expert experience loss, introduces expert demonstration trajectory samples with high reward values on the basis of self-exploration samples to save the exploration space; at the same time, introduces an expert experience loss gradient function to optimize network parameters and obtain the optimal strategy.

[0005] The above DDPG algorithm that incorporates the expert experience loss solves the problem of policy iteration optimization. However, the design of the reward function still has strong subjectivity, and the rewards obtained through interaction with the environment are usually sparse, resulting in extremely difficult convergence during algorithm training and poor path planning effects. When experts complete obstacle avoidance tasks, their strategies are often optimal. Therefore, learning expert experience from expert demonstration trajectories and then constructing a reward function is, to a large extent, more in line with real-world needs than the form of a reward function designed manually. An algorithm that, given expert trajectories, reverse-derives the reward function implicit in expert experience is inverse reinforcement learning (IRL). IRL can be divided into two major categories: maximum margin and maximum entropy. However, methods based on maximum margin often produce ambiguities, that is, different reward functions with random preferences can be derived from the same expert policy. The maximum entropy model is completely constructed based on known data (i.e., expert trajectories) without making any subjective assumptions about the distribution of unknown information, effectively avoiding the problem of ambiguity. Therefore, the present invention uses the IRL algorithm based on maximum entropy to solve the optimal reward function implicit in expert demonstration trajectories. Summary of the Invention

[0006] To overcome the problems in the prior art, the present invention proposes a UAV path planning method based on inverse reinforcement learning.

[0007] The technical solution of the present invention to solve the above technical problems is as follows:

[0008] A UAV path planning method based on inverse reinforcement learning, comprising the following steps:

[0009] Step 1. Collect an expert demonstration trajectory dataset and a self-exploration trajectory dataset of an expert operating a UAV to avoid obstacles;

[0010] Step 2. Construct an experience pool, which is jointly composed of the expert demonstration trajectory dataset and the self-exploration trajectory dataset, and adopt a hybrid sampling mechanism to sample from the two datasets respectively to form the final training samples;

[0011] Step 3. Based on DDPG, introduce an expert experience loss function to guide the iterative update of DDPG parameters and accelerate the solution of the optimal policy;

[0012] Step 4. Construct a reward function, and solve the reward function based on the maximum entropy inverse reinforcement learning algorithm, that is, given the known expert demonstration trajectories, solve the implicit probability model that generates these trajectories;

[0013] Step 5. Train DDPG until DDPG completes the flight task with the optimal policy under the optimal reward function implicit in the expert trajectories.

[0014] Furthermore, the specific steps for constructing the experience pool in step 2 are as follows:

[0015] The experience pool consists of the expert demonstration trajectory dataset T expert and the self-exploration trajectory dataset T discover and they jointly form the final training sample T by using a hybrid sampling mechanism to sample from the two datasets respectively:

[0016] T = α·T expert + β·T discover (1)

[0017] In the formula, α is the sampling ratio from the training set T expert and β is the sampling ratio from the training set T discover .

[0018] Furthermore, the DDPG algorithm guided by the expert experience loss function in step 3 includes the online policy network μ(s|θ μ ), the online value function network Q(s,a|θ Q ), the target policy network μ'(s|θ μ ') and the target value function network Q'(s,a|θ Q ').

[0019] Furthermore, the optimization of the parameters of the online value function network Q(s,a|θ Q ) specifically includes the following steps:

[0020] According to the Bellman equation, at the i-th training time step, the action target value y i of the online value function network is:

[0021] y i = r i + γQ'(s i+1 ,μ'(s i+1 |θ μ' )|θ Q' ) (2)

[0022] Then the error δ i between the action target value of the online value function network and the actual output Q(s i ,a Q |θ i is:

[0023] δ i = y i - Q(s i ,a i |θ Q ) (3)

[0024] Substitute Equation (3) into Equation (2) to obtain the loss function of the online value function network:

[0025]

[0026] Minimize the loss function J(θ Q ) with respect to the parameters θ Q of the online value function network through gradient descent, and let J(θ Q ) be differentiated with respect to the network parameters θ Q . It can be seen that its gradient value is:

[0027]

[0028] The update of the parameters of the online value function network is carried out according to Equation (5).

[0029] Furthermore, the optimization of the parameters of the online policy network specifically includes the following steps:

[0030] The optimization of the parameters of the online policy network is divided into two parts: the expert demonstration trajectory samples and the self-exploration samples;

[0031] For the expert demonstration trajectory data, the mean square error J between the immediate policy a i predicted by the online policy network based on the current expert state and the true expert policy exp (θ μ ) is introduced as the expert experience loss, so that the predicted output policy of the network continuously approaches the expert policy:

[0032]

[0033] where is the immediate policy predicted by the online policy network based on the current expert state ;

[0034] Let the expert experience loss J exp (θ μ ) be differentiated with respect to the policy network parameters θ μ . The gradient value can be obtained as

[0035]

[0036] Update the parameters θ according to the online policy gradient value μ of the original DDPG algorithm:

[0037]

[0038] Adopt the fusion gradient Update the parameters of the online policy network:

[0039]

[0040] Where λ is the fusion gradient adjustment factor.

[0041] Furthermore, the update of the target network parameters is based on the online network parameters in a soft update manner:

[0042]

[0043] Where τ < 1.

[0044] Furthermore, constructing the reward function in step 4 includes the following steps:

[0045] Given the trajectory ζ generated by the expert's manipulation of the UAV to avoid obstacles:

[0046] ζ = {(s1,a1),(s2,a2),…(s n ,a n )} (11)

[0047] Then the reward value r(ζ) of this trajectory is:

[0048]

[0049] Use a linear combination of a finite number of important feature functions f(·) to fit the reward function, then

[0050]

[0051] Where f i is the i-th feature component of the reward function, and θ i is the i-th component of the weight vector of the reward function; n is the number of feature vectors in the reward function;

[0052] If the Euclidean distance d of the UAV relative to the obstacle, the relative distance heading angle ψ d , the relative distance climb angle The motion speed v of the UAV relative to the obstacle, the relative motion speed heading angle ψ v , the relative motion speed climb angle and other information belong to the important features in the UAV obstacle avoidance process, so

[0053]

[0054] Define F(ζ) as the sum of the feature components of each state in Equation (14), and the specific form is:

[0055]

[0056] Substitute Equation (15) into Equation (13), and the reward value of each trajectory is expressed as:

[0057] r(ζ) = θ T F(ζ) (16).

[0058] Furthermore, the specific steps for solving the reward function based on the maximum entropy inverse reinforcement learning algorithm in Step 4 are as follows:

[0059] Given m expert trajectories, the feature expectation of the experts is:

[0060]

[0061] In the case of known expert trajectories, assume the potential probability distribution is p(ζ i |θ), then the feature expectation of the expert trajectories is:

[0062]

[0063] The maximum entropy model is constructed based on the expert trajectories in the above formula, and the problem of solving the maximum entropy is transformed into an optimization problem:

[0064]

[0065] In the formula, p = p(ζ i |θ);

[0066] Transform the above optimization problem into the dual form:

[0067]

[0068] In the formula, λ j , λ0 are Lagrange multipliers;

[0069] Let the loss function L(p) be differentiated with respect to the probability p of the expert demonstration trajectory distribution, and we can get:

[0070]

[0071] Let the above formula be equal to 0, then the maximum entropy probability model of the expert demonstration trajectory is obtained:

[0072]

[0073] In the formula, λ j corresponds to the weight vector θ of the feature function in the reward function;

[0074]

[0075] In the formula, Z(θ) is the partition term, that is, the sum of the probabilities of all possible expert trajectories;

[0076] In the probability model shown above, the greater the probability of the expert trajectory appearing, that is, the greater \(Z(\theta)\), the closer the setting of the reward function is to the optimal policy implicit in the expert examples. The solution of the optimal reward function can be transformed into optimizing by maximizing the entropy of the expert trajectory distribution:

[0077]

[0078] Transform the above formula into solving the loss by minimizing the negative log-likelihood function of the weights \(\theta\) of the reward function feature components:

[0079]

[0080] By calculating the expert trajectory prediction partition function \(Z(\theta)\) under the current policy:

[0081]

[0082] In the formula, \(T\) samp represents the expert trajectory under the current policy, and \(n\) represents the number of expert trajectories under the current policy;

[0083] For the continuous expert states and the corresponding true expert policies in the sampled expert trajectories, perform discretization processing, and randomly batch sample from them. Transform formula (25) into:

[0084]

[0085] In the above formula, the loss function \(J(\theta)\) is:

[0086]

[0087] Let the loss function \(J(\theta)\) be differentiated with respect to the weights \(\theta\) of the reward function, and solve the optimal reward function through the gradient descent method, we can get:

[0088]

[0089] In summary, through formula (29), the global optimal solution \(r\) of the reward function is finally learned * (s i ,a i ).

[0090] Furthermore, training the DDPG in step 5 until the DDPG completes the flight mission with the optimal policy under the optimal reward function implicit in the expert trajectory includes the following steps:

[0091] Randomly initialize the network parameters \(\theta\) of the online policy network \(Q(s,a|\theta\) Q ) and the online value function network \(\mu(s|\theta\) μ )μ and θ Q , initialize the target networks μ′ and Q', their weights, and the weights θ of the reward function i ; initialize the experience pool, and store the collected dataset T of expert demonstration trajectories expert in the experience pool;

[0092] a) The on-policy network obtains an action a = πθ(s) + η based on the current state s t , where η t is random noise, and the action selection policy π depends on the design of the reward function;

[0093] b) Interact with the environment to execute the action a, obtain the new state s', and the immediate reward value r;

[0094] c) Store the self-exploration sample data (s, a, r, s′) generated by interacting with the environment, i.e., T discover in the experience pool;

[0095] d) s = s';

[0096] e) Randomly sample N sample data from the experience pool for training, estimate the partition function Z(θ) according to Equation (26), minimize the objective value shown in Equation (28), and optimize the weights θ of the reward function i to obtain the optimal reward function;

[0097] f) Update the value function network parameters θ according to Equation (5) Q , if the training data T ∈ T expert , then update the policy network parameters according to Equation (7), if the training data T ∈ T discover , then update the policy network parameters θ according to Equation (8) μ ;

[0098] g) Update the target network parameters θ according to Equation (10) Q′ and θ μ′ ;

[0099] h) When s' is the termination state, the current iteration ends, otherwise go to step a).

[0100] Compared with the prior art, the present invention has the following technical effects:

[0101] In view of the UAV path planning problem, the present invention proposes a UAV path planning method based on inverse reinforcement learning. First, a demonstration trajectory data set of an expert manipulating the UAV to avoid obstacles is collected based on simulator software; secondly, a hybrid sampling mechanism is adopted to fuse high-quality expert demonstration trajectory data in the self-exploration data to update network parameters, so as to reduce the exploration cost; finally, the optimal reward function is solved according to the maximum entropy inverse reinforcement learning algorithm, solving the problems that the reward function is difficult to design and the invisible expert experience cannot be fully mined. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] Figure 1 It is a schematic diagram of the simulator component and data display of the present invention;

[0103] Figure 2 It is a block diagram of the simulator training of the present invention;

[0104] Figure 3 It is a schematic diagram of the expert demonstration trajectory data set of the present invention;

[0105] Figure 4 It is a schematic diagram of the test environment of the present invention;

[0106] Figure 5 It is a block diagram of the hybrid sampling mechanism of the present invention;

[0107] Figure 6 It is a training framework diagram of the improved DDPG algorithm of the present invention;

[0108] Figure 7 It is a curve diagram of the reward value integrating expert experience loss of the present invention;

[0109] Figure 8 It is a curve diagram of the reward value based on the maximum entropy inverse reinforcement learning of the present invention;

[0110] Figure 9 It is a curve diagram of the reward value of the comprehensively improved DDPG algorithm of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0111] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0112] To solve the problems of slow convergence speed and difficult reward function setting in the deep deterministic policy gradient algorithm for planning the safe collision avoidance path of UAVs, the present invention proposes a UAV path planning method based on inverse reinforcement learning, including the following steps: Step 1. Collect the expert demonstration trajectory dataset and self-exploration trajectory dataset of the expert manipulating the UAV to avoid obstacles; Step 2. Construct an experience pool, which is jointly composed of the expert demonstration trajectory dataset and the self-exploration trajectory dataset, and adopt a hybrid sampling mechanism to sample from the two datasets respectively to form the final training samples; Step 3. Construct DDPG, introduce the expert experience loss function to guide the iterative update of the DDPG parameters, and accelerate the solution of the optimal policy; Step 4. Construct a reward function, solve the reward function based on the maximum entropy inverse reinforcement learning algorithm, that is, given the expert demonstration trajectory, solve the implicit probability model that generates this trajectory; Step 5. Train DDPG until DDPG completes the flight task with the optimal policy under the optimal reward function implied by the expert trajectory.

[0113] The above steps are described in detail as follows:

[0114] Step 1. Collect the expert demonstration trajectory dataset and self-exploration trajectory dataset of the expert manipulating the UAV to avoid obstacles;

[0115] Based on the complex obstacle scenarios provided in the simulator Phoenix RC software, obtain the expert demonstration trajectories, and build a simple obstacle scenario based on Python for the UAV to interact with the environment to generate self-exploration data samples.

[0116] The collection of expert demonstration trajectories is based on the professional radio control flight simulation software Phoenix RC developed by a game company. When collecting expert demonstration trajectories in the Phoenix RC simulator, the following three components are mainly used: 1) Remote control module: Make the channels of the remote control match the functions of the UAV model, including the stick movement strokes of the rudder, elevator, throttle, and aileron to control the movement of the UAV; 2) UAV model module: Includes various aircraft models such as fixed-wing and rotary-wing; 3) Scenario module: Provides hundreds of three-dimensional simulation obstacle environments, and variables such as wind and light can be customized to simulate reality. In the simulation environment, data such as the UAV flight speed, heading angle, GPS positioning, gyroscope, and barometer can be obtained through interface functions and displayed in real time. The components and data display are as Figure 1 shown.

[0117] The framework of manually manipulating the UAV model to avoid obstacles in the Phoenix RC simulator is as Figure 2As shown in the figure. After obtaining the environmental state information in the three-dimensional simulation obstacle environment, the environmental state information includes obstacle positions, distances between the UAV and obstacles, etc. The expert manually operates the rudder, elevator, throttle, and aileron lever strokes of the remote control, and continuously adjusts the heading angle, pitch angle, flight speed, etc. of the UAV model to avoid obstacles.

[0118] Partial expert demonstration trajectories collected from the Phoenix RC simulator are as Figure 3 shown. The advantages of collecting the obstacle environment dataset in the simulator are as follows: 1) The obstacle settings and scene types in the simulator are complex and variable, and the degree of closeness to the real world is high; 2) The training is completely carried out in the simulated scene, and the UAV can be manually operated to simulate a variety of different maneuvering actions to determine the optimal flight strategy; 3) The simulator intuitively displays the monocular RGB images at each moment during the obstacle avoidance process and parameters such as the heading angle and flight speed of the UAV without using complex sensors for perception and measurement; 4) There is no need to consider collision damage and safety issues.

[0119] To test the algorithm performance, a three-dimensional obstacle environment Q as Figure 4 shown is built, enabling the UAV to generate a self-exploration dataset during the interaction with the environment. The obstacle environment has a length of L, a width of W, and a height of H. There are dynamic and static obstacles with different threat levels in the environment, and the spatial positions, movement speeds, and influence ranges of the obstacles are all unknown.

[0120] Step 2. Construct an experience pool, which is jointly composed of the expert demonstration trajectory dataset and the self-exploration trajectory dataset, and a hybrid sampling mechanism is used to sample from the two datasets respectively to form the final training samples.

[0121] In the UAV obstacle avoidance training, in order to avoid the waste of resources caused by the random and inefficient exploration in the initial training stage of the UAV, and at the same time to achieve the diversification of samples as much as possible, and then break through the upper limit implied by the expert strategy, as Figure 5 shown, the experience pool of this implementation is jointly composed of the expert demonstration trajectory dataset T expert and the self-exploration trajectory dataset T discover and a hybrid sampling mechanism is used to sample from the two datasets respectively to form the final training sample T:

[0122] T = α·T expert + β·T discover (1)

[0123] In the formula, α is the sampling ratio from the training set T expert , and β is the sampling ratio from the training set T discover .

[0124] Step 3. Based on DDPG, introduce an expert experience loss function to guide the iterative update of DDPG parameters and accelerate the solution of the optimal strategy.

[0125] Aiming at the disadvantages of the original DDPG algorithm, such as large exploration space and low sample reward value in the initial stage, an improved DDPG algorithm optimization strategy iteration that integrates expert experience loss is proposed. The training samples of the original DDPG algorithm only contain the data set generated by self-exploration through interaction with the same environment. In this embodiment, a hybrid sampling mechanism is adopted, and some expert demonstration trajectory samples are introduced on the basis of self-exploration samples. For the expert trajectory data set, an expert experience loss function is introduced to guide the iterative update of DDPG parameters, accelerating the solution of the optimal strategy; the self-exploration data samples are still updated according to the original DDPG algorithm.

[0126] The DDPG algorithm guided by the introduction of the expert experience loss function includes an online policy network μ(s|θ μ ), an online value function network Q(s,a|θ Q ), a target policy network μ'(s|θ μ ) and a target value function network Q'(s,a|θ Q' ) in four parts.

[0127] According to the Bellman equation, at the i-th training time step, the action target value y i of the online value function network is:

[0128] y i =r i +γQ'(s i+1 ,μ'(s i+1 |θ μ' )|θ Q' ) (2)

[0129] Then the error δ i between the action target value of the online value function network and the actual output Q(s i ,a Q |θ i is:

[0130] δ i =y i -Q(s i ,a i |θ Q ) (3)

[0131] Substituting Equation (3) into Equation (2), the loss function of the online value function network can be obtained:

[0132]

[0133] Minimize the loss function J(θ Q ) with respect to the online value function network parameters θ Q through gradient descent method for optimization update, and let J(θ Q ) with respect to the network parameters θQ Taking the derivative, its gradient value can be obtained is:

[0134]

[0135] The update of the online value function network parameters is carried out according to Equation (5).

[0136] The optimization of the online policy network parameters is divided into two parts: expert demonstration trajectory samples and self-exploration samples. For the expert demonstration trajectory data, the online policy network can be based on the current expert state The predicted immediate policy a i and the true expert policy The mean square error J exp (θ μ ) is introduced into the policy network as the expert experience loss, making the predicted output policy of the network continuously tend to the expert policy:

[0137]

[0138] In the formula, is the immediate policy predicted by the online policy network based on the current expert state predicted.

[0139] Let the expert experience loss J exp (θ μ ) take the derivative with respect to the policy network parameter θ μ Taking the derivative, its gradient value can be obtained is

[0140]

[0141] Since the expert policy trajectory is limited and cannot cover the entire state and action space, while the UAV can explore a larger space in the interaction with the environment, thereby breaking through the upper limit implied by the expert policy and improving the algorithm stability. Therefore, while introducing the expert experience loss gradient to optimize the online policy iteration process, the self-exploration trajectory dataset is also retained and the parameters θ are updated according to the online policy gradient value of the original DDPG algorithm μ :

[0142]

[0143] An expert experience loss function method that includes the online policy gradient of the expert experience loss and the original online policy gradient is adopted. On the one hand, introducing high-quality expert policies saves the exploration space in the initial stage and improves the algorithm convergence efficiency. On the other hand, it continuously learns in self-exploration to try to obtain better policies not covered in the expert trajectory. Finally, the fusion gradient Update the parameters of the online policy network:

[0144]

[0145] where λ is the fusion gradient adjustment factor.

[0146] The update of the target network parameters is based on the online network parameters and adopts a soft update method:

[0147]

[0148] where τ < 1.

[0149] Step 4. Construct a reward function and solve the reward function based on the maximum entropy inverse reinforcement learning algorithm, that is, given the trajectory ζ of the expert demonstration, solve the implicit probability model that generates this trajectory.

[0150] Given the trajectory ζ generated by the expert's manipulation of the UAV to avoid obstacles:

[0151] ζ = {(s1,a1),(s2,a2),…(s n ,a n )} (11)

[0152] Then the reward value r(ζ) of this trajectory is:

[0153]

[0154] Use the linear combination of a finite number of important feature functions f(·) to fit the reward function, then

[0155]

[0156] where f i is the i-th feature component of the reward function, and θ i is the i-th component of the weight vector of the reward function; n is the number of feature vectors in the reward function.

[0157] During the process of the expert manipulating the UAV to avoid obstacles, the expert operator often makes decisions based on the current UAV flight speed, the azimuth distance between the UAV and the obstacle, etc. Therefore, the Euclidean distance d of the UAV relative to the obstacle, the relative distance course angle ψ d , the relative distance climb angle the relative motion speed v of the UAV relative to the obstacle, the relative motion speed course angle ψ v , the relative motion speed climb angle and other information belong to the important features in the process of the UAV avoiding obstacles. Therefore

[0158]

[0159] Define \(F(\zeta)\) as the sum of the characteristic components of each state in Equation (14), and its specific form is:

[0160]

[0161] Substitute Equation (15) into Equation (13), then the reward value of each trajectory can be expressed as:

[0162] \(r(\zeta)=\theta\) T \(F(\zeta)\) (16)

[0163] Given \(m\) expert trajectories, the characteristic expectation of the expert is:

[0164]

[0165] In the case of known expert trajectories, assuming the potential probability distribution is \(p(\zeta\) i |\(\theta\)), then the characteristic expectation of the expert trajectory is:

[0166]

[0167] The maximum entropy model is completely constructed based on the known data (i.e., expert trajectories) in the above equation, without making any subjective assumptions about unknown situations. Therefore, it can effectively avoid the ambiguity problem existing in the custom reward function. Convert the problem of solving the maximum entropy problem into an optimization problem:

[0168]

[0169] where \(p = p(\zeta\) i |\(\theta\)).

[0170] Convert the above optimization problem into its dual form:

[0171]

[0172] where \(\lambda\) j and \(\lambda_0\) are Lagrange multipliers.

[0173] Take the derivative of the loss function \(L(p)\) with respect to the probability \(p\) of the expert demonstration trajectory distribution, and we can get:

[0174]

[0175] Let the above equation be equal to 0, then we get the maximum entropy probability model of the expert demonstration trajectory:

[0176]

[0177] where \(\lambda\) j corresponds to the weight vector \(\theta\) of the characteristic function in the reward function.

[0178]

[0179] In the formula, \(Z(\theta)\) is the partition item, that is, the sum of the probabilities of all possible expert trajectories.

[0180] In the probability model shown above, the greater the probability of the expert trajectory appearing, that is, the greater \(Z(\theta)\), the closer the setting of the reward function is to the optimal policy implied in the expert examples. The solution of the optimal reward function can be transformed into optimizing by maximizing the entropy of the expert trajectory distribution:

[0181]

[0182] Transform the above formula into solving the loss by minimizing the negative log-likelihood function of the weight \(\theta\) of the reward function feature component:

[0183]

[0184] By calculating the expert trajectory prediction partition function \(Z(\theta)\) under the current policy:

[0185]

[0186] In the formula, \(T\) samp represents the expert trajectory under the current policy, and \(n\) represents the number of expert trajectories under the current policy.

[0187] Due to certain differences in expert cognition, in order to reduce the fitting variance of the weight \(\theta\), the continuous expert states in the sampled expert trajectories and the corresponding true expert policies are discretized, and random batch sampling is performed from them, and the formula (25) is transformed into:

[0188]

[0189] In the above formula, the loss function \(J(\theta)\) is:

[0190]

[0191] Let the loss function \(J(\theta)\) be differentiated with respect to the weight \(\theta\) of the reward function, and the optimal reward function is solved by the gradient descent method, and we can get:

[0192]

[0193] In summary, through the formula (29), the global optimal solution \(r\) of the reward function can finally be learned * (s i ,a i ).

[0194] Step 5. Train DDPG until DDPG completes the flight mission with the optimal policy under the optimal reward function implied by the expert trajectory.

[0195] The problem of obstacle avoidance using the improved DDPG algorithm can be described as follows: at a series of consecutive decision-making moments, the network makes a decision based on the current state s of the UAV; after the decision is implemented, the network obtains an immediate reward value according to the reward function designed by inverse reinforcement learning, and this reward corresponds to the network decision and the environmental state. Then the network will enter the state at the next moment corresponding to the decision and update the network parameters forward through the DDPG algorithm that integrates the expert supervision loss; at a new training time, the network will execute a new decision based on the new state it is in and obtain a new reward value, and so on in a loop until the network completes the flight mission with the optimal policy under the optimal reward function implied by the expert trajectory.

[0196] The specific training includes the following steps:

[0197] 1. Randomly initialize the network parameters θ μ of the online policy network μ(s|θ Q ) and the online value function network Q(s,a|θ μ ), and initialize the target networks μ′ and Q' and their weights; Q

[0198] 2. Construct the reward function and initialize the reward function weights;

[0199] 3. Initialize the experience pool and store the collected expert demonstration trajectory dataset T expert in the experience pool;

[0200] 4. Conduct network training with the number of iterations being M;

[0201] a) The online policy network obtains the action a = πθ(s)+η t based on the current state s, where η t is random noise, and the action selection strategy π depends on the design of the reward function;

[0202] b) Interact with the environment to execute the action a, obtain the new state s' and the immediate reward value r;

[0203] c) Store the self-exploration sample data (s,a,r,s′) generated by interacting with the environment, that is, T discover in the experience pool;

[0204] d) s = s';

[0205] e) Randomly sample N sample data from the experience pool for training, estimate the partition function Z(θ), minimize the objective value shown in Equation (28), optimize the reward function weights θ i to obtain the optimal reward function;

[0206] f) Update the value function network parameters θ Q according to Equation (5), if the training data \(T\in\mathcal{T}\) expert , then update the policy network parameters according to Equation (7). If the training data \(T\in\mathcal{T}\) discover , then update the policy network parameters \(\theta\) according to Equation (8) μ ;

[0207] g) Update the target network parameters \(\theta\) and \(\hat{\theta}\) according to Equation (10) Q′ and \(\hat{\theta}\) μ′ ;

[0208] h) When \(s'\) is the terminal state, the current iteration ends. Otherwise, go to step a).

[0209] Test the obstacle avoidance performance of the algorithm of the present invention in a simulation environment. The simulation experimental environment is Python 3.7, an Inter Core i5 processor with a main frequency of 2.42 GHz, and a Windows 10 operating system. The test scenario is a dense dynamic and static obstacle environment as shown in Figure 4 . Set the simulation environment as a three-dimensional area of \(200\times250\times400\), and there are several dynamic obstacles moving in different directions of heading angle and climbing angle at different moving speeds in the area.

[0210] To test the policy learning performance of the DDPG algorithm with integrated expert experience loss, the present invention compares and tests the influence of the improved algorithm and the original algorithm of the present invention on the network convergence speed while keeping the reward function consistent. Design the reward function as shown in Equation (30), where \(R\) is the reward function, \(d\) min is the minimum value of \(d\) described in Equation (14) i , that is, the distance between the UAV and the nearest obstacle, so as to drive the UAV away from the obstacle; \(d\) tar is the distance between the UAV and the end point, so as to prompt the UAV to fly towards the end point; \(k_1 = 2\), \(k_2 = 0.5\).

[0211]

[0212] The reward value during the UAV training process is as shown in Figure 7As can be seen from the figure, for the UAV using the original DDPG algorithm, the policy learning speed almost linearly increases within 100 to 200 iterations; however, within 200 to 600 iterations, the model falls into a local optimal solution, resulting in difficulty for the network to converge; within 600 to 1600 iterations, the policy learning speed continues to increase, but the growth rate slows down; after 1600 iterations, the reward value gradually converges and is roughly stable at 110, but with large fluctuations. For the DDPG algorithm with integrated expert experience loss, the policy learning speed of the UAV continues to increase rapidly within 100 to 300 iterations; after 300 iterations, the growth rate slows down, but there is no stagnation; after 1300 iterations, the network converges, the reward value is roughly stable at 125, which is higher than the original DDPG algorithm, and the fluctuation range is smaller. In summary, the DDPG algorithm with integrated experience loss has a faster convergence speed, better stability, and better obstacle avoidance effect.

[0213] Regarding the analysis of the influence of the maximum entropy inverse reinforcement learning algorithm on the obstacle avoidance effect, the reward value curve of the UAV using the optimal reward function solved by the maximum entropy inverse reinforcement learning is as Figure 8 shown. The UAV trained with the optimal reward function has a high-speed growth in the policy learning speed within 100 to 500 iterations; after 500 iterations, the growth rate slightly slows down; after 1300 iterations, the reward value roughly converges to 170. Compared with the original DDPG algorithm, the UAV trained with the optimal reward function has a significantly higher reward value, a faster policy learning speed, and better stability.

[0214] The improved DDPG algorithm integrating expert experience loss and inverse reinforcement learning. Figure 9 The reward value curve of the improved DDPG algorithm integrating expert experience loss and inverse reinforcement learning is given. As can be seen from the figure, the DDPG algorithm improved by combining the above two points combines the respective advantages of both. It not only has a faster policy learning speed in the initial stage, but also has a better obstacle avoidance effect. The reward value can converge to 175 around 1300 iterations.

[0215] The present invention aims at the UAV path planning problem and improves the DDPG algorithm. By introducing an expert experience loss function to optimize the iterative process of the policy network, the exploration cost in the initial stage of the original DDPG algorithm is saved, and the network convergence speed is accelerated; at the same time, based on the maximum entropy inverse reinforcement learning algorithm, the optimal reward function hidden in the expert demonstration trajectory is solved, overcoming the problem of difficulty in artificially setting the reward function in complex tasks. Comparative experiments show that the improved DDPG algorithm of the present invention has a faster convergence speed and better obstacle avoidance effect.

[0216] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A UAV path planning method based on inverse reinforcement learning, characterized in that, It includes the following steps: Step 1. Collect the expert demonstration trajectory dataset and the self-exploration trajectory dataset of the expert's manipulation of UAV obstacle avoidance; Step 2. Construct an experience pool, which is jointly composed of the expert demonstration trajectory dataset and the self-exploration trajectory dataset, and a hybrid sampling mechanism is used to sample from the two datasets respectively to form the final training samples; Step 3. Based on DDPG, introduce an expert experience loss function to guide the iterative update of DDPG parameters and accelerate the solution of the optimal policy; the DDPG algorithm guided by the introduced expert experience loss function includes an online policy network μ(s|θ μ ), an online value function network Q(s,a|θ Q ), a target policy network μ'(s|θ μ '), and a target value function network Q'(s,a|θ Q '); The optimization of the online policy network parameters specifically includes the following steps: The optimization of the online policy network parameters is divided into two parts: the expert demonstration trajectory samples and the self-exploration samples; For the expert demonstration trajectory data, the online policy network is based on the current expert state The predicted immediate policy a i And the true expert policy The mean squared error J exp (θ μ ) is introduced as the expert experience loss, making the predicted output policy of the network continuously tend to the expert policy: In the formula, is the immediate policy predicted by the online policy network based on the current expert state ; Let the expert experience loss be J exp (θ μ ) with respect to the policy network parameters θ μ Take the derivative to obtain its gradient value which is Online policy gradient value according to the original DDPG algorithm Update parameter θ μ : Adopt the fusion gradient Update the parameters of the online policy network: In the formula, λ is the fusion gradient adjustment factor; Step 4. Construct a reward function, and solve the reward function based on the maximum entropy inverse reinforcement learning algorithm, that is, given the expert demonstration trajectory, solve the implicit probability model that generates this trajectory; Step 5. Train DDPG until DDPG completes the flight task with the optimal policy under the optimal reward function implied by the expert trajectory.

2. The method for path planning of an unmanned aerial vehicle based on inverse reinforcement learning according to claim 1, wherein, The specific steps of constructing the experience pool in Step 2 include the following: The experience pool consists of the expert demonstration trajectory dataset T expert and the self-exploration trajectory dataset T discover which together form the final training sample T by using a hybrid sampling mechanism to sample from the two datasets respectively: T = α·T expert + β·T discover (1) where α is the sampling proportion from the training set T expert and β is the sampling proportion from the training set T discover ​ 3. A method for path planning of an unmanned aerial vehicle based on inverse reinforcement learning according to claim 1, characterized in that, Online value function network Q(s,a|θ Q ) The optimization of the parameters specifically includes the following steps: According to the Bellman equation, at the i-th training time step, the action target value y of the online value function network i is as follows: y i = r i + γQ'(s i+1 , μ'(s i+1 |θ μ' )|θ Q' ) (2) Then the error δ between the action target value of the online value function network and the actual output Q(s i , a i |θ Q ) is as follows: i is: δ i = y i - Q(s i , a i | θ Q ) (3) Substitute Equation (3) into Equation (2) to obtain the loss function of the online value function network: Minimize the loss function \(J(\theta)\) through gradient descent to optimize and update the online value function network parameters \(\theta\). Let \(J(\theta)\) take the derivative with respect to the network parameters \(\theta\). It can be seen that its gradient value is: Q Q Q Q ​​​​​ The update of the online value function network parameters is carried out according to Equation (5).

4. A method for UAV path planning based on inverse reinforcement learning according to claim 1, characterized in that, The update of the target network parameters is based on the online network parameters and adopts a soft update method: In the formula, τ < 1.

5. A method for UAV path planning based on inverse reinforcement learning according to claim 4, characterized in that, The steps of constructing the reward function in Step 4 include the following: Given the trajectory ζ generated by the expert's manipulation of UAV obstacle avoidance: ζ = {(s1, a1), (s2, a2), …(s n , a n )} (11) Then the reward value r(ζ) of this trajectory is: Use the linear combination of a finite number of important feature functions f(·) to fit the reward function, then where, f i is the i-th feature component of the reward function, and θ i is the i-th component of the weight vector of the reward function; n is the number of feature vectors in the reward function; If the Euclidean distance d of the UAV relative to the obstacle, the relative distance heading angle ψ d , the relative distance climb angle The motion speed v of the UAV relative to the obstacle, the relative motion speed heading angle ψ v , the relative motion speed climb angle The information belongs to the important features in the UAV obstacle avoidance process, so Define F(ζ) as the sum of the feature components of each state in Equation (14), and the specific form is: Substitute Equation (15) into Equation (13), then the reward value of each trajectory is expressed as: r(ζ) = θ T F(ζ) (16).

6. A method for path planning of an unmanned aerial vehicle based on inverse reinforcement learning according to claim 5, characterized in that, The specific steps of solving the reward function based on the maximum entropy inverse reinforcement learning algorithm in Step 4 include the following: Given m expert trajectories, the feature expectation of the expert is: Given the known expert trajectory, assuming the potential probability distribution is p(ζ i |θ), the characteristic expectation of the expert trajectory is as follows: The maximum entropy model is constructed based on the expert trajectories in the above formula, and the problem of solving the maximum entropy is converted into an optimization problem: where p = p(ζ i |θ); Convert the above optimization problem into a dual form: where λ j and λ0 are Lagrange multipliers; Let the loss function L(p) be differentiated with respect to the distribution probability p of the expert demonstration trajectory, and we get: Let the above formula be equal to 0, then the maximum entropy probability model of the expert demonstration trajectory is obtained: where λ j corresponds to the weight vector θ of the feature function in the reward function; In the formula, Z(θ) is the partition item, that is, the sum of the probabilities of all possible expert trajectories; In the probability model shown above, the greater the probability of the expert trajectory appearing, that is, the greater Z(θ), the closer the reward function setting is to the optimal policy implied in the expert example; the problem of solving the optimal reward function is optimized by maximizing the entropy of the expert trajectory distribution: Convert the above formula into the minimization of the negative log-likelihood function of the reward function feature component weight θ to solve the loss amount: By calculating the expert trajectory prediction partition function Z(θ) under the current policy: where, T samp represents the expert trajectory under the current policy, and n represents the number of expert trajectories under the current policy; For consecutive expert states in the sampled expert trajectories and the corresponding true expert policies perform discretization processing, randomly batch sample from them, and transform Equation (25) into: In the above formula, the loss function J(θ) is: Let the loss function J(θ) be differentiated with respect to the weight θ of the reward function, and solve the optimal reward function through the gradient descent method, we can get: In summary, the global optimal solution r of the reward function is finally learned through Equation (29). * (s i ,a i ).

7. A method for path planning of an unmanned aerial vehicle based on inverse reinforcement learning according to claim 6, characterized in that, The steps of training DDPG in Step 5 until DDPG completes the flight task with the optimal policy under the optimal reward function implied by the expert trajectory include the following: Randomly initialize the network parameters θ of the online policy network Q(s,a|θ Q ) and the online value function network μ(s|θ μ ), and θ μ and θ Q , initialize the target networks μ′ and Q' and their weights, and initialize the reward function weight θ i ; Initialize the experience pool and store the collected expert demonstration trajectory dataset T expert in the experience pool; a) The online policy network obtains the action a = πθ(s) + η based on the current state s t , where η t is random noise, and the action selection policy π depends on the design of the reward function; b) Interact with the environment to execute the action a, obtain the new state s', and the immediate reward value r; c) Store the self-exploration sample data (s, a, r, s′) generated by interacting with the environment, i.e., T discover into the experience pool; d) s = s′; e) Randomly sample N sample data from the experience pool for training, estimate the partition function Z(θ) according to Equation (26), minimize the objective value shown in Equation (28), and optimize the reward function weight θ i to obtain the optimal reward function; f) Update the value function network parameter θ according to Equation (5). Q If the training data T ∈ T expert then update the policy network parameter according to Equation (7), if the training data T ∈ T discover then update the policy network parameter θ according to Equation (8). μ ; g) Update the target network parameters θ according to Equation (10) Q′ and θ μ′ ; h) When s' is the termination state, the current iteration ends; otherwise, go to step a).