Unmanned aerial vehicle trajectory planning and tracking method, system and device based on deep reinforcement learning and adaptive nonlinear model predictive control, and medium

By combining deep reinforcement learning and adaptive nonlinear model prediction control, the drone trajectory planning and tracking methods are optimized, and the problems of high computational complexity and low learning efficiency in complex environments are solved, and the efficient and stable flight of the drone in dynamic environments is achieved.

CN120276462APending Publication Date: 2025-07-08XIDIAN UNIV
View PDF 0 Cites 17 Cited by

Patent Information

Application Number
CN202510419204.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing UAV trajectory planning and tracking methods have high computational complexity and poor real-time performance in complex environments, making it difficult to cope with dynamic perturbations and parameter changes. The reward feedback delay in deep reinforcement learning leads to low learning efficiency, and the nonlinear model prediction control is insufficient in changing environments.

Method used

Combining deep reinforcement learning and adaptive nonlinear model prediction control, Agent, Environment and Reward are optimized, target position cost function, stage cost function and smooth nonlinear control law function are introduced, SoftActor-Critic algorithm is integrated, optimization problem parameters are dynamically adjusted, and the optimal reference trajectory is generated.

Benefits of technology

It improves the trajectory planning and tracking performance of the drone in complex environments, enhances the robustness and stability of the system, and can quickly respond to dynamic obstacles and wind noise interference, achieving efficient and safe flight control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276462A_ABST
    Figure CN120276462A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle trajectory planning and tracking method, system and device based on deep reinforcement learning and adaptive nonlinear model predictive control, and a medium. The method comprises the following steps: constructing various static multi-obstacle and dynamic multi-obstacle simulation environments; constructing a kinetic model of the unmanned aerial vehicle; constructing an adaptive nonlinear model predictive control (ANMPC) algorithm; constructing a reward function of the tracking performance of the unmanned aerial vehicle to the reference trajectory generated by the adaptive nonlinear model predictive control algorithm; constructing a network framework based on deep reinforcement learning and an adaptive nonlinear model predictive control algorithm; setting network parameters; training a network framework, and selecting an optimal weight file; outputting a test result; the system, the device and the medium are used for realizing the unmanned aerial vehicle trajectory planning and tracking method. The method can effectively cope with changes of targets and environments, shows strong obstacle avoidance capability and anti-interference performance when facing dynamic obstacles and wind noise interference, and embodies a high intelligent decision-making level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of the integration of deep reinforcement learning and UAV control technology, and particularly relates to a UAV trajectory planning and tracking method, system, device and medium based on deep reinforcement learning and adaptive non-linear model predictive control. Background Art

[0002] With the reduction of UAV costs and the rapid development of sensing technology, automation and artificial intelligence, UAVs have been equipped with cameras, global positioning systems (GPS), a variety of hardware sensors and mission payload systems, and are widely used in fields such as reconnaissance, monitoring, countermeasure and communication, performing tasks that are difficult to complete by traditional manpower. In these applications, the path planning ability of UAVs is crucial. To improve the path planning speed, optimize the efficiency and reduce the loss risk, the research on path planning algorithms has become the focus in the field of UAVs. However, traditional remote control and preset trajectory control methods are difficult to cope with complex environments, so more intelligent and autonomous trajectory planning and tracking control methods are needed. At present, the combination of deep reinforcement learning and control algorithms has become one of the latest development trends in UAV flight trajectory planning and predictive tracking.

[0003] Deep reinforcement learning has shown great potential in many key fields such as robot control, autonomous driving, aircraft control and power system scheduling, and can effectively handle high-dimensional state and complex policy optimization problems. Among them, the SoftActor-Critic algorithm is particularly excellent and is good at solving problems in continuous action spaces, providing accurate and efficient technical support for the trajectory planning and predictive tracking of UAVs. At the same time, as an advanced model-based control strategy, adaptive non-linear model predictive control combines adaptive control and non-linear model predictive control, and uses a prediction model to optimize future control inputs, and can adjust parameters in real time when facing dynamic environments and uncertainties, significantly improving the robustness and performance of the system. The combination of deep reinforcement learning and adaptive non-linear model predictive control is gradually becoming a research hotspot, which can not only improve the control effect in high-dimensional non-linear problems, but also be extended to more complex application scenarios. In the future, the integration of the two is expected to play a more important role in UAV intelligent control and related fields, accelerating the innovation and development of intelligent control technology.

[0004] At present, the application of the method of UAV trajectory planning and tracking combining deep reinforcement learning and model predictive control has made substantial progress in UAV trajectory planning and tracking. The UAV can explore in a complex dynamic environment and gradually learn to plan a path with high safety and short distance to the target. In (Wang Y, Gao Z, Zhang J, Cao X, Zheng D, Gao Y, Ng DW, Di Renzo M. Trajectory design for UAV-based Internet of Things data collection: A deep reinforcement learning approach. IEEE Internet of Things Journal. 2021 Aug 3;9(5):3899-912.), inspired by the state-of-the-art deep reinforcement learning method, the twin-delayed deep deterministic policy gradient (TD3) is used to design the UAV trajectory, and a trajectory design of the time-to-destination minimization (TD3-TDCTM) algorithm based on TD3 is proposed. In (Tong GU, Jiang N, Biyue LI, Xi ZH, Ya WA, Wenbo DU. UAV navigation in high dynamic environments: A deep reinforcement learning approach. Chinese Journal of Aeronautics. 2021 Feb 1;34(2):479-89.), an improved deep reinforcement learning-based method is proposed to solve the problem of non-convergence.

[0005] The defects and deficiencies of the existing related inventions are as follows:

[0006] 1. Nonlinear model predictive control (NMPC) generates the optimal control input by solving an optimization problem. However, due to the need to perform multi-step predictions of future behavior in each control cycle and the dependence on complex system models and control strategies, the computational complexity is relatively high, resulting in poor real-time performance. In complex nonlinear and coupled systems, the computational amount increases sharply. Especially in dynamic and uncertain environments, traditional nonlinear model predictive control lacks adaptability and is difficult to cope with external disturbances and parameter changes. Therefore, combining algorithms such as adaptive mechanisms or deep reinforcement learning can improve the real-time performance and robustness of nonlinear model predictive control and enhance its adaptability in changing environments.

[0007] 2. In deep reinforcement learning, the reward function usually gives positive rewards only after the UAV reaches the target, which leads to delayed reward feedback and reduces the exploration efficiency. Since deep neural networks rely on a large amount of training data, the update process is slow and the convergence speed is relatively slow, making it difficult to quickly learn the optimal strategy. To improve the learning efficiency, the design of the reward function can be improved, an immediate feedback mechanism can be adopted, or a model-based reinforcement learning method can be introduced to accelerate training and enhance task performance.

[0008] 3. Deep reinforcement learning may overfit the training data during the training process, especially when the amount of data is limited, which may lead to insufficient generalization ability in practical applications. When fused with nonlinear model predictive control, the nonlinear dynamics model of the UAV may perform well in some specific situations, but it is difficult to maintain stability in an environment with large actual changes. Summary of the Invention

[0009] To overcome the above deficiencies of the prior art, the purpose of the present invention is to provide a UAV trajectory planning and tracking method, system, device and medium based on deep reinforcement learning and adaptive nonlinear model predictive control. First, the Agent (intelligent agent strategy), Environment (environmental modeling) and Reward (reward function) are optimized. The trajectory planning intelligence is improved through deep reinforcement learning, the model generalization ability is enhanced by using a high-fidelity dynamics environment, and the reward function is redesigned to improve the trajectory tracking accuracy and energy efficiency, enabling the UAV to have better autonomous flight capabilities in complex environments. Secondly, by introducing the target position cost function E, the stage cost function L and the smooth nonlinear control law function g, the cost function in the adaptive nonlinear model predictive control optimization problem is simplified, thereby reducing the computational amount and complexity and significantly improving the learning efficiency. In addition, the SoftActor-Critic algorithm is integrated under the adaptive nonlinear model predictive control algorithm. By dynamically adjusting the parameters of the underlying optimization problem, the optimal reference trajectory of the UAV is generated and it is controlled to track this trajectory to reach the target. Finally, the algorithm is verified in the designed dynamic multi-obstacle environment. The results show that the present invention has excellent performance in terms of cumulative reward r, flight path length L, path global distortion GS, whether there is a collision and threat index, etc. At the same time, the UAV trajectory planning and tracking method, system, device and medium based on deep reinforcement learning and adaptive nonlinear model predictive control are outstanding in terms of robustness, stability and safety. This method can effectively cope with the changes of the target and the environment, and shows strong obstacle avoidance ability and anti-interference performance when facing dynamic obstacles and wind noise interference, reflecting a high level of intelligent decision-making.

[0010] To achieve the above purpose, the technical solution adopted by the present invention is:

[0011] A method for UAV trajectory planning and tracking based on deep reinforcement learning and adaptive nonlinear model predictive control, comprising the following steps:

[0012] Step 1: Construct various static multi-obstacle and dynamic multi-obstacle simulation environments;

[0013] Step 2: Construct the dynamic model of the UAV, including its motion equation, attitude control model, and external disturbance modeling;

[0014] Step 3: Construct an adaptive nonlinear model predictive control (ANMPC) algorithm;

[0015] Step 4: Construct a reward function for the tracking performance of the UAV for the reference trajectory generated by the adaptive nonlinear model predictive control algorithm;

[0016] Step 5: Based on the adaptive nonlinear model predictive control (ANMPC) algorithm constructed in Step 3, construct a network framework based on the deep reinforcement learning and adaptive nonlinear model predictive control algorithm;

[0017] Step 6: Set network parameters;

[0018] Step 7: According to the network parameters set in Step 6, input the reward function constructed in Step 4 into the network framework based on the deep reinforcement learning and adaptive nonlinear model predictive control algorithm constructed in Step 5; use the dynamic model of the UAV constructed in Step 2 as the control object, and perform training in the various static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1. After each round of training, output the training weight file; by comparing with the training results generated in all previous training rounds before this training, select the weight file corresponding to the training result with the highest reward value and the shortest flight path length as the optimal weight file;

[0019] Step 8: Input the optimal weight file obtained in Step 7 and the dynamic model of the UAV after applying it to the algorithm network framework based on the deep reinforcement learning and adaptive nonlinear model predictive control constructed in Step 5 into the various static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1 for testing, and obtain the UAV flight trajectory map, reward value r, flight path length L, path maximum distortion LS, path global distortion GS, whether a collision occurs, and threat index results.

[0020] Compared with the prior art, the beneficial effects of the present invention are:

[0021] 1. The present invention develops an Adaptive Nonlinear Model Predictive Control (ANMPC) algorithm to solve the trajectory planning and tracking problems of unmanned aerial vehicles in a nonlinear, multivariable, and strongly coupled dynamic environment. This algorithm introduces a deep reinforcement learning algorithm to adjust the optimization problem parameters, generate an optimal reference trajectory, enhance the system's adaptability to environmental changes, and perform real-time optimization under complex flight conditions. By comprehensively optimizing the trajectory error, energy consumption, and control input, the network framework based on the deep reinforcement learning and adaptive nonlinear model predictive control algorithm achieves efficient, stable, and safe flight performance, and has a strong adaptive ability to maintain high-precision trajectory tracking under dynamic obstacles and interference noise.

[0022] 2. By introducing the target position cost function E (Equation (10)), the stage cost function L (Equation (11)), and the smooth nonlinear control law function g, the present invention simplifies the cost function (Equation (9)) of the adaptive nonlinear model predictive control (ANMPC) optimization problem, reduces the computational amount and complexity, significantly improves the system's computational speed, and achieves efficient real-time control. The target position cost function (Equation (10)) ensures that the unmanned aerial vehicle can accurately reach the target position, the stage cost function L (Equation (11)) ensures the trajectory tracking accuracy, and the smooth nonlinear control law function g reduces the drastic fluctuations of the control signal, improving the system's stability and robustness. These optimizations enable the system to quickly respond in a complex environment while ensuring flight performance and control accuracy.

[0023] 3. By introducing a reward value function into the network framework of the adaptive nonlinear model predictive control (ANMPC) algorithm based on deep reinforcement learning to measure the tracking performance of the unmanned aerial vehicle with respect to the reference trajectory, the convergence speed of network updates can be accelerated. This enables the unmanned aerial vehicle to learn the optimal strategy from exploration more quickly and accurately, enhancing its adaptability and real-time control efficiency in complex environments.

[0024] In summary, the present invention first redesigned the Agent, Environment, and Reward parts of deep reinforcement learning to improve the performance and efficiency of UAV trajectory planning and tracking. Secondly, by introducing the target position cost function E (Equation (10)), the stage cost function L (Equation (11)), and the smooth nonlinear control law function g, the cost function in the adaptive nonlinear model predictive control (ANMPC) optimization problem was simplified, thereby reducing the computational amount and complexity and significantly improving the learning efficiency. In addition, the SoftActor-Critic algorithm was integrated under the adaptive nonlinear model predictive control algorithm. By dynamically adjusting the parameters of the underlying optimization problem, the optimal reference trajectory of the UAV was generated, and it was controlled to accurately track this trajectory to reach the target. The combination of deep reinforcement learning and the adaptive nonlinear model predictive control algorithm unifies the advantages of the model-based adaptive nonlinear model predictive control method and the data-based SoftActor-Critic algorithm, retaining both the stability and optimization ability of model predictive control and improving the adaptability and exploration ability of the algorithm. This innovative method performs excellently in complex environments such as dynamic obstacles and wind noise interference, significantly improving the trajectory tracking accuracy, real-time performance, and system robustness of the UAV. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is the network framework diagram of the present invention's embodiment based on deep reinforcement learning and adaptive nonlinear model predictive control algorithm.

[0026] Figure 2 is the network framework diagram of the present invention's embodiment based on SoftActor-Critic algorithm and adaptive nonlinear model predictive control algorithm.

[0027] Figure 3 is the coordinate definition of the UAV dynamics modeling in the present invention's embodiment.

[0028] Figure 4 is the effect comparison diagram of the flight trajectory and the reference trajectory in a certain local trajectory planning and tracking in the present invention's embodiment.

[0029] Figure 5 is the cumulative reward curve graph of the policy network of the present invention (ANMPC-SAC algorithm) and the double-delayed deep deterministic policy gradient algorithm (TD3), proximal policy optimization algorithm (PPO), and deep deterministic policy gradient algorithm (DDPG) in 4 dynamic multi-obstacle environments. Among them, Figure 5 (a) is the comparison curve graph of the cumulative rewards of the UAVs trained in the No. 1 dynamic multi-obstacle environment using the double-delayed deep deterministic policy gradient algorithm (TD3), proximal policy optimization algorithm (PPO), deep deterministic policy gradient algorithm (DDPG), and the present invention (ANMPC-SAC algorithm);Figure 5 (b) is a comparison curve graph of the cumulative rewards of drones using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG), and the present invention (ANMPC-SAC algorithm) during training in the No. 2 dynamic multi-obstacle environment; Figure 5 (c) is a comparison curve graph of the cumulative rewards of drones using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG), and the present invention (ANMPC-SAC algorithm) during training in the No. 3 dynamic multi-obstacle environment; Figure 5 (d) is a comparison curve graph of the cumulative rewards of drones using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG), and the present invention (ANMPC-SAC algorithm) during training in the No. 4 dynamic multi-obstacle environment.

[0030] Figure 6 are simulation diagrams of drones using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG), and the present invention (ANMPC-SAC algorithm) in a dynamic multi-obstacle environment. The bright blue dots are the starting points, and the magenta dots are the target points. Among them, Figure 6 (a) is a simulation diagram of a drone using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3) flying in the No. 3 dynamic multi-obstacle environment; Figure 6 (b) is a simulation diagram of a drone using the Proximal Policy Optimization algorithm (PPO) flying in the No. 3 dynamic multi-obstacle environment; Figure 6 (c) is a simulation diagram of a drone using the Deep Deterministic Policy Gradient algorithm (DDPG) flying in the No. 3 dynamic multi-obstacle environment; Figure 6 (d) is a simulation diagram of a drone using the present invention (ANMPC-SAC algorithm) flying in the No. 3 dynamic multi-obstacle environment.

[0031] Figure 7 are comparison graphs of the training durations of drones using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG), and the present invention (ANMPC-SAC algorithm) in a dynamic multi-obstacle simulation environment. Among them, Figure 7 (a) is a comparison graph of the total training durations of drones using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG), and the present invention (ANMPC-SAC algorithm) in a dynamic multi-obstacle simulation environment; Figure 7(b) is a comparison graph of the training duration per time of the unmanned aerial vehicle using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG) and the present invention (ANMPC-SAC algorithm) in a dynamic multi-obstacle simulation environment.

[0032] Figure 8 is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) in a dynamic multi-obstacle environment with different target point positions. Among them, Figure 8 (a) is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) flying in the No. 3 dynamic multi-obstacle environment with the target point of (5, 10, 10); Figure 8 (b) is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) flying in the No. 3 dynamic multi-obstacle environment with the target point of (10, 5, 10); Figure 8 (c) is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) flying in the No. 3 dynamic multi-obstacle environment with the target point of (10, 10, 5); Figure 8 (d) is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) flying in the No. 3 dynamic multi-obstacle environment with the target point of (10, 10, 10).

[0033] Figure 9 is a simulation graph of the unmanned aerial vehicle using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG) and the present invention (ANMPC-SAC algorithm) flying in different dynamic multi-obstacle environments. Among them, Figure 9 (a) is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) flying in the No. 1 dynamic multi-obstacle environment, Figure 9 (b) is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) flying in the No. 2 dynamic multi-obstacle environment, Figure 9 (c) is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) flying in the No. 3 dynamic multi-obstacle environment, Figure 9 (d) is a simulation graph of the unmanned aerial vehicle using the present invention (ANMPC-SAC algorithm) flying in the No. 4 dynamic multi-obstacle environment.

[0034] Figure 10 is a simulation graph of the unmanned aerial vehicle using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG) and the present invention (ANMPC-SAC algorithm) flying in a dynamic multi-obstacle environment with the influence of wind noise. Among them, Figure 10(a) is a simulation diagram of a drone using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3) flying in a No. 3 dynamic multi-obstacle environment with a 0.05 wind noise impact. Figure 10 (b) is a simulation diagram of a drone using the Proximal Policy Optimization algorithm (PPO) flying in a No. 3 dynamic multi-obstacle environment with a 0.05 wind noise impact. Figure 10 (c) is a simulation diagram of a drone using the Deep Deterministic Policy Gradient algorithm (DDPG) flying in a No. 3 dynamic multi-obstacle environment with a 0.05 wind noise impact. Figure 10 (d) is a simulation diagram of a drone using the present invention (ANMPC-SAC algorithm) flying in a No. 3 dynamic multi-obstacle environment with a 0.05 wind noise impact. Detailed implementation manners

[0035] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0036] A UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control includes the following steps:

[0037] Step 1: Construct a simulation environment of various static multi-obstacles and dynamic multi-obstacles of 10×10×10 (km) in MatlabR2018b software;

[0038] Step 2: Construct the dynamic model of the UAV, including its motion equation, attitude control model, and external disturbance modeling;

[0039] Step 3: Construct an Adaptive Nonlinear Model Predictive Control (ANMPC) algorithm;

[0040] Step 4: Construct a reward function for the tracking performance of the UAV for the reference trajectory generated by the adaptive nonlinear model predictive control algorithm;

[0041] Step 5: Based on the Adaptive Nonlinear Model Predictive Control (ANMPC) algorithm constructed in Step 3, construct a network framework based on the deep reinforcement learning and adaptive nonlinear model predictive control algorithm;

[0042] Step 6: Set network parameters;

[0043] Step 7: According to the network parameters set in Step 6, input the reward function constructed in Step 4 into the network framework of the algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control constructed in Step 5; take the dynamic model of the UAV constructed in Step 2 as the control object, and perform training in the multiple static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1. After each round of training, output the training weight file; by comparing with the training results generated in all previous training rounds before this training, select the weight file corresponding to the training result with the highest reward value and the shortest flight path length as the optimal weight file;

[0044] Step 8: Input the optimal weight file obtained in Step 7 and the dynamic model of the UAV after applying it to the algorithm network framework based on deep reinforcement learning and adaptive nonlinear model predictive control constructed in Step 5 into the multiple static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1 for testing, and obtain the UAV flight trajectory diagram, reward value r, flight path length L, path maximum distortion LS, path global distortion GS, whether a collision occurs, and threat index results.

[0045] The implementation method of Step 1 includes:

[0046] In the present invention, flight experiments are respectively carried out on the UAV in the static and dynamic multi-obstacle simulation environments constructed in Matlab R2018b. In the static multi-obstacle environment, the positions and shapes of the obstacles are fixed, mainly used to test the navigation and obstacle avoidance capabilities of the UAV in a complex space, and focus on evaluating the path planning efficiency and the effectiveness of the obstacle avoidance strategy. In the dynamic multi-obstacle environment, the positions and states of the obstacles change with time, simulating a more complex scenario closer to actual applications. In this case, the UAV needs to have the ability to perceive the environment in real time and dynamically adjust the flight path to cope with the movement of the obstacles. The following are the specific setting methods for the static and dynamic multi-obstacle environments:

[0047] Step 101: Set multiple static multi-obstacle environments: Set the center position of the sphere (X I , Y I , Z I ), radius R, and the number of spheres n sphere ; the center position of the bottom circle of the cylinder (X I , Y I , Z I ), radius R, height H, and the number of cylinders n cylinder ; the center position of the bottom circle of the cone (X I , Y I , Z I ), radius R, height H, and the number of cones n cone ;

[0048] Step 102: Set up multiple dynamic multi-obstacle environments: Set the center position (X I , Y I , Z I ) of the sphere, the radius R, the motion trajectory r(t) of the sphere, the moving speed v, and the number n of spheres sphere .

[0049] The specific implementation method of the second step is as follows:

[0050] Based on the Newton-Euler method, analyze the dynamic characteristics of the UAV, establish an accurate dynamic model of the UAV, so as to understand and analyze the system characteristics of the UAV and use it as the simulation model for the simulation experiment.

[0051] As Figure 3 shown, define the geographical coordinate system O I (X I , Y I , Z I ) and the UAV body coordinate system O B (X B , Y B , Z B ). The position vector of the UAV in the geographical coordinate system O I is defined as ζ:

[0052]

[0053] The velocity vector ν of the UAV is defined as:

[0054]

[0055] The attitude vector of the UAV in space is defined as η:

[0056]

[0057] The roll angle φ in Equation (3) represents the rotation angle of the UAV along the X I axis of the geographical coordinate system O I ; the pitch angle θ represents the rotation angle of the UAV along the Y I axis of the geographical coordinate system O I ; the heading angle ψ represents the rotation angle of the UAV along the Z I axis of the geographical coordinate system O I . The Euler angle rate vector ω is:

[0058]

[0059] In Equation (4), p, q, and r are the UAV's along the X B , Y B , Z B in the UAV body coordinate system O B under the condition ofB The angular velocity of the axis rotation, and the control quantity state space is defined as:

[0060] U = [U1 U2 U3 U4] T (5)

[0061] In Equation (5), is the sum of the thrusts generated by the four rotors, representing the displacement movement of the UAV along the Z-axis of the UAV body coordinate system O B ; B ; is the roll moment, representing the rolling movement of the UAV rotating along the X-axis of the UAV body coordinate system O B ; B ; is the pitch moment, representing the pitching movement of the UAV rotating along the Y-axis of the UAV body coordinate system O B ; B ; is the yaw moment, representing the yaw movement of the UAV rotating along the Z-axis of the UAV body coordinate system O B ; B ; K is the lift coefficient; K d is the anti-torque coefficient; l is the distance from the center of each rotor of the quadrotor to its center of mass; ω1, ω2, ω3, ω4 are the rotational speeds of the four rotors respectively;

[0062] The state space and state vector of the dynamic model of the UAV are:

[0063]

[0064] In Equation (6), the state vector of the UAV is Therefore, the state X(k) of the UAV at time k belongs to X; k x , k y , k z is obtained by decomposing the air resistance coefficient in the geodetic coordinate system O I ; k φ , k θ , k ψ are the components of the total air resistance torque coefficient on the X I , Y I , Z I axes respectively; I x , I y , I z are the moments of inertia of the UAV along the X B , Y B , Z B , Z B axes of the UAV body coordinate system O; J RP\(J\) is the moment of inertia of the propeller along the rotation axis of the drone; \(\Omega = -\omega_1+\omega_2-\omega_3+\omega_4\) is the algebraic sum of the motor speeds.

[0065] The specific implementation method of the third step is as follows:

[0066] The adaptive nonlinear model predictive control (ANMPC) adopted in the control technology of the drone in the present invention is an optimization-based control method, aiming to generate a reference trajectory and control input for the closed-loop negative feedback control system of the drone. In its optimization framework, the adaptive nonlinear model prediction uses a nonlinear system model, a nonlinear control law, a set of state vectors, control vectors, and the maximum and minimum values of the output vector control quantities, as well as the constraints of static and dynamic obstacles in the environment, to connect the nonlinear object with the control law, aiming to reduce the nonlinearity of the entire closed-loop negative feedback system, thereby reducing the non-convexity of the related optimization problem. This greatly improves the efficiency of the optimization calculation, enabling the real-time generated optimization trajectory to be run on the drone.

[0067] The drone is a nonlinear system, where the state vector, control vector, and output vector are respectively and It is assumed that the output vector is a subset of the state vector Furthermore, let \(f(x(k),u(k)):X\times U\rightarrow X\) be a known smooth nonlinear function based on online parameter estimation, and \(g(x(k),y ref (k)):X\times Y\rightarrow U\) be the mapping of the smooth nonlinear control law; given a finite prediction horizon \(T\), the finite sequences of the predicted control vector \(u(t k+T-1 )\) and the predicted state vector \(x(t k+T )\) are:

[0068] \(u(t k+T-1 )=(u_0,u_1,\cdots,u N-1 )\) and \(x(t k+T )=(x_0,x_1,\cdots,x N )\ (7)

[0069] For the convenience of computer processing, the adaptive nonlinear model predictive control (ANMPC) is a digital control method. Therefore, it is necessary to discretize the continuous state space model. Through discretization, the continuous-time control system is transformed into a discrete-time model suitable for digital computers, enabling the control problem to be optimized and solved at discrete time steps. Therefore, the general form of the discrete expression of the state space of the drone is:

[0070] \(x(t k+1 ) = f(x(t k ),u(t k )) = A\cdot x(tk ) + B·u(t k ) (8)

[0071] The adaptive nonlinear model predictive control algorithm generates the reference trajectory y for the UAV flight. ref , and then, this trajectory is tracked by the UAV closed-loop negative feedback control system composed of the controlled object and the control law for the reference trajectory y ref to fly to the target position of the current local trajectory planning and tracking; the cost function J of the adaptive nonlinear model predictive control optimization problem is:

[0072]

[0073] In Equation (9), the target position cost function E and the stage cost function L are respectively:

[0074]

[0075] Within the prediction interval T of one local trajectory planning and tracking, x′(τ) is the predicted UAV system state vector, and y ref (τ) is the predicted UAV flight trajectory. The smaller the target position cost function E (Equation (10)), the closer the actual flight end point of the UAV is to the predicted trajectory end point; the smaller the stage cost function L, the closer the flight trajectory of the UAV is to the predicted trajectory; the smaller the smooth nonlinear control law function g, the less energy the UAV consumes;

[0076] The errors in Equations (10) and (11) are weighted by the matrices and and will be adaptively adjusted using the deep reinforcement learning algorithm in this work. k represents the kth trajectory planning and tracking during the flight; within the prediction interval T of one local trajectory planning and tracking, the UAV closed-loop negative feedback control system reaches the target stable setpoint x tar from the current state x, and the adaptive nonlinear model predictive optimization problem is shown in Equation (12); let t k+i , i = 0, 1,... T represent the continuous sampling times. At each sampling moment, for x(t k+i ) and y ref (t k+i ), when ||x tar - x(t k+i )|| ≤ Δ, it means that the UAV has reached the target position of the current local trajectory planning and tracking, where is the tolerance specified by the user; to better control the UAV, it is necessary to minimize the cost function J(x′(τ), y ref (τ)), and the adaptive nonlinear model predictive control optimization problem is:

[0077]

[0078] In Equation (12), A represents the weight coefficient matrix of the system state; B represents the weight coefficient matrix of the controller output; X, U, and Y are respectively the data sets of the state, input, and output trajectories; u min (τ) and u max (τ) are respectively the minimum value of the control quantity and the maximum value of the control quantity;

[0079] The calculation process of the adaptive nonlinear model predictive control is as follows: First, detect the state of the UAV at time t n . Then, under the constraints of x′(τ) = f(x(τ), u(τ)), u(τ) = g(x′(τ), y ref (τ)), and u min (τ) ≤ u(τ) ≤ u max (τ), minimize the cost function J(x′(τ), y k , t k+T in the prediction interval [t ref (τ)) of one local trajectory planning and tracking to provide the prediction of the values of x′(t k+i ) and y ref (t k+i ). Finally, input the predicted UAV system state vector x′(t k+i ) and the reference trajectory y ref (t k+i ) into the UAV closed-loop negative feedback control system, and output the control vector u(t k+i ) through the smooth nonlinear control law function g for trajectory tracking; this process is repeatedly calculated in real time at the iteration rate set by the user until the required target state value is reached.

[0080] In the constructed complex dynamic environment, the adaptive nonlinear model predictive control algorithm pre-obtains the target position, and the modular global motion planner gradually generates stage target positions successively in the unexplored areas of the established complex dynamic environment. These stage target positions, the current state vector of the UAV, the UAV nonlinear model, the constraints of static and dynamic obstacles in the environment, the maximum / minimum values of the control quantity, and the numerical values of the weighting matrix are sent to the adaptive nonlinear model predictive control optimization problem to calculate the best reference trajectory between the current position of the UAV and the current stage target position.

[0081] The specific implementation method of Step 4 is as follows:

[0082] The reward of the deep reinforcement learning in the present invention is the reward function of the tracking performance of the UAV for the reference trajectory generated by the adaptive nonlinear model predictive control algorithm, which is respectively defined as the trajectory tracking reward r traj and the target position reward rss 、Completion reward r c and smooth reward r smo :

[0083] (1) Trajectory tracking reward r traj , r traj represents the degree of fit between the real-time flight trajectory of the UAV and the predicted reference trajectory generated by the adaptive nonlinear model predictive control algorithm; the trajectory tracking reward r traj is:

[0084]

[0085] In formula (13), is the root mean square error between the generated predicted reference trajectory and the flight trajectory, r traj,max and r traj,th are the maximum value and threshold of the trajectory tracking reward r traj respectively;

[0086] (2) Target position reward r tar , r tar is the distance difference between the target position x(t k+T ) of the real-time flight trajectory of the UAV and the target position p tar of the predicted reference trajectory generated by the adaptive nonlinear model predictive control algorithm; the target position reward r tar is:

[0087]

[0088] In formula (14), e tar =||x(t k+T ) - p tar || is the distance difference between the target position x(t k+T ) of the real-time flight trajectory of the UAV and the target position p tar of the trajectory generated by the adaptive nonlinear model predictive control algorithm, r tar,max and r tar,th are the maximum value and threshold of the target position reward r tar respectively;

[0089] (3) Completion reward r c , r c represents the distance difference between the position x(t k ) of the real-time flight trajectory of the UAV and the target position p tar of the trajectory generated by the adaptive nonlinear model predictive control algorithm during a local trajectory planning and tracking process; the completion reward is r c :

[0090]

[0091] In Equation (15), e tar = ||x(t k ) - p tar || represents the distance difference between the position of the UAV at time t k and the trajectory target position p generated by the adaptive nonlinear model predictive control algorithm. r tar and r c,max are respectively the maximum value and the threshold of the completion reward r c,th ; whenever the error e c exceeds the specified threshold r c (e c,th > r c ), this factor is more emphasized by reducing the total reward (r < 0); therefore, the algorithm designed in the present invention will give priority to ensuring that the UAV reaches the target position within the specified time. c,th )

[0092] (4) Smooth the reward r smo , r smo represents the smoothness of the flight trajectory of the UAV during a local trajectory planning and tracking process; the smooth reward r smo is:

[0093]

[0094] In Equation (16), τ(p) is the distortion degree of the UAV flight trajectory at position p, and the trajectory distortion degree of a local trajectory planning and tracking process r smo,max and r smo,th are respectively the maximum value and the threshold of the smooth reward r smo ;

[0095] In summary, the reward designed in the present invention is:

[0096] r = λ1·r traj + λ2·r tar + λ3·r c + λ4·r smo (17)

[0097] In Equation (17), λ1, λ2, λ3, and λ4 respectively represent the weights of the four rewards, and λ1 + λ2 + λ3 + λ4 = 1.

[0098] The specific implementation method of the fifth step is as follows:

[0099] The Soft Actor-Critic algorithm in deep reinforcement learning adopted by the present invention is a deep reinforcement learning method based on the maximum entropy theory. By maximizing the entropy of the policy, this algorithm promotes full exploration of the environment, thus achieving a better balance between exploration and exploitation. Entropy is a measure of the randomness of the policy. Increasing entropy means enhanced randomness of the policy, which in turn prompts the algorithm to obtain more information during exploration, accelerate the learning speed, and prevent the policy from prematurely converging to a local optimal solution.

[0100] The Soft Actor-Critic algorithm belongs to the probabilistic framework, is built on the basis of the Actor-Critic model, and combines the idea of Q-learning. In SAC, the optimization of the policy not only depends on the traditional reward signal but also enhances the exploration ability by maximizing entropy. Specifically, the Soft Actor-Critic algorithm encourages more policy variations during the optimization process by introducing an entropy term, thereby better exploring the environment in the initial stage and avoiding getting stuck in a local optimum.

[0101] The present invention uses the Soft Actor-Critic algorithm in deep reinforcement learning to optimize the trajectory planning and control of drones. By leveraging its advantage of balancing exploration and exploitation, significant performance improvements are achieved in the trajectory planning and control of drones. By introducing the maximum entropy theory, the Soft Actor-Critic algorithm can effectively accelerate the learning speed and ensure that drones can make more flexible and robust decisions when facing complex dynamic environments.

[0102] The Soft Actor-Critic (SAC) algorithm consists of two action-value networks (Q networks), two target Q networks, and one policy network (Actor). The Soft Actor-Critic (SAC) algorithm trains the two action-value networks (Q networks) simultaneously, adopts a stochastic policy, and uses the modified mean-square Bellman equation to update the two action-value networks (Q networks). The Soft Actor-Critic (SAC) algorithm improves the exploration ability and stability of the policy by training the two action-value networks (Q networks) simultaneously and adopting a stochastic policy. Both Q networks are updated using the modified mean-square Bellman equation to reduce overestimation bias. In addition, SAC enhances the stability of learning by introducing target Q networks (for smoothing Q-value updates) and performs weighted average updates on network parameters during training. Since SAC adopts a stochastic policy and directly obtains the next state-action value from the current policy, there is no need for a target policy network, thus simplifying the training process and improving the sampling efficiency. The Bellman equation is as follows:

[0103]

[0104] Let \(P\) denote the distribution that the state-action pairs encountered by the Agent follow under the control of policy \(\pi\). The goal of the SoftActor-Critic algorithm is to maximize the entropy of the policy and the expected return to achieve efficient exploration and exploitation, thereby improving the performance of the Agent in reinforcement learning problems. Therefore, the Bellman equation of the Q-function for reinforcement learning based on entropy regularization is redesigned as follows:

[0105]

[0106] In Equation (19), \(\alpha\) is the coefficient that adjusts the balance between the expected entropy and the return, \(a'\) is the next action, \(s'\) is the next state, and \(\gamma\) is the discount factor used in deep reinforcement learning to adjust the influence of the near and far future;

[0107] The expectation is calculated for the next state (from the replay buffer) and the next action (from the current policy rather than the replay buffer). Therefore, samples can be used to approximately estimate the Q value:

[0108] Q π (s,a)\(\approx r(s)+\gamma(Q π (s',a')-\alpha\log π (a'|s')),a'\sim\pi(\cdot|s')(20)

[0109] The loss function of the Q-function network in the Soft Actor-Critic algorithm is:

[0110]

[0111] In Equation (21), the form of the Bellman equation is:

[0112]

[0113] In the SoftActor-Critic algorithm, the action-value network is updated using the gradient descent method, while the policy network is updated using the gradient ascent method; the goal of the Agent's policy is to maximize the state-value function \(V π (s)\), and the state-value function is:

[0114]

[0115] Equation (23) represents the expected return of the Agent starting from state \(s\) and following policy \(\pi\); the policy is optimized using the reparameterization method. By calculating the deterministic function of the state, policy parameters, and independent noise, samples are drawn from it. The reparameterization technique allows the expectation of the action to be rewritten as the expectation of the noise; this calculation process is expressed as:

[0116] a θ(s, ξ) = tanh(μ θ (s) + σ θ (s) ⊙ ξ), ξ ~ N(0, I) (24)

[0117] Policy optimization can be achieved by maximizing the Q - function, which implicitly maximizes the entropy of the trajectory. The loss function of the policy network in the SoftActor - Critic algorithm is as follows:

[0118]

[0119] As Figure 4 shown, it is a comparison diagram of the flight trajectory and the predicted reference trajectory in a local trajectory planning and tracking. From the perspective of deep reinforcement learning, in each flight mission, the whole flight mission is divided into k segments, and the UAV is controlled to fly over k target points; therefore, each flight mission consists of k local trajectory planning and tracking; after the UAV completes the previous trajectory planning and tracking each time, three observation values are sent to the UAV, which are respectively: the initial velocity of the UAV in this local trajectory planning and tracking Initial velocity vector between the vector and the ss angle Figure 4 is the comparison effect diagram of the flight trajectory and the predicted reference trajectory in a local trajectory planning and tracking. In addition, in order to verify the robustness and stability of the present invention, static obstacles, dynamic obstacles and wind noise are also added.

[0120] The network framework based on deep reinforcement learning (DRL) and adaptive nonlinear model predictive control (ANMPC) algorithm is as Figure 1 shown, and its core includes three processes: interaction, update and iteration. In the interaction process, the UAV, as an agent, explores and executes tasks in a three - dimensional space, continuously exchanges experiences with the environment, and stores the collected experiences in the replay buffer. The update process randomly extracts experiences from the buffer to update the neural network, including the Actor network and multiple Critic networks (Critic#1, Critic#2 and their target networks). In the iteration process, the policy network generates the weight matrix of the adaptive nonlinear model predictive control cost function in real - time, constructs an optimization model in combination with the UAV state vector, and obtains the optimal control vector by solving, so as to achieve precise control of the UAV.

[0121] In a local trajectory planning process, the UAV uses the adaptive nonlinear model predictive control algorithm based on deep reinforcement learning to plan the optimal reference trajectory y to the target point from the current state ref, and predict its future state vector x'. Subsequently, a control vector u is generated through the designed smooth non-linear control law function g to accurately track the reference trajectory y ref . Figure 1 The cyan box (ANMPC algorithm module) in Figure 1 contains the non-linear dynamics model of the UAV, the non-linear flight control law, the solution scheme for the optimization problem, the cost function, and the constraint conditions characterizing the drive and environmental obstacles. Through this framework, the UAV can achieve autonomous flight and trajectory optimization in complex environments.

[0122] As Figure 2 shown, based on the network framework of the SoftActor-Critic algorithm and the adaptive non-linear model predictive control algorithm, the actions generated by the policy network are the weight coefficient matrix A of the system state and the weight coefficient matrix B of the control input in the adaptive non-linear model predictive control optimization problem. Execute the generated actions, and the adaptive non-linear model predictive control calculates the predicted system state x'(τ) and the generated reference trajectory y that optimize the target position cost function E (Equation (10)), the stage cost function L (Equation (11)), and the smooth non-linear control law function g ref (τ). These results are used by the control system of the UAV, combined with the current state of the UAV, to generate the control quantity u(τ) and calculate the observed value and the reward r stored in the replay buffer.

[0123] The specific implementation method of the sixth step is as follows:

[0124] The simulation experiment was carried out in a constructed three-dimensional environment with static and dynamic obstacles. The type, size, position, and number of obstacles vary randomly, and wind noise interference is introduced to simulate the real situation. The present invention compares several popular deep reinforcement learning algorithms - Twin Delayed Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO), and Deep Deterministic Policy Gradient (DDPG) with a UAV trajectory planning and tracking algorithm based on deep reinforcement learning and adaptive non-linear model predictive control proposed by the present invention. All algorithms are configured with the same parameters: 10 3 training times, a 5-layer MLP architecture, 256 neurons in each layer, and the learning rate of the Adam policy optimizer is 3×10 -4 . Table 1 summarizes the main hyperparameters of each algorithm.

[0125] Table 1 Hyperparameter setting table of Twin Delayed Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO), and Deep Deterministic Policy Gradient (DDPG) and a UAV trajectory planning and tracking algorithm based on deep reinforcement learning and adaptive non-linear model predictive control proposed by the present invention

[0126]

[0127]

[0128] The specific implementation method of Step 7 is as follows:

[0129] According to the network parameters set in Step 6, input the reward function constructed in Step 4 into the network framework of the algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control constructed in Step 5; use the dynamic model of the drone constructed in Step 2 as the control object, and perform training in the simulation environments of various static multi-obstacles and dynamic multi-obstacles constructed in Step 1. After each round of training, output the training weight file; by comparing with the training results generated in all previous training rounds before this training, select the weight file corresponding to the training result with the highest reward value and the shortest flight path length as the optimal weight file.

[0130] The specific implementation method of Step 8 is as follows:

[0131] The calculation of experimental data such as flight path length L, path maximum distortion LS, path global distortion GS, whether a collision occurs, and threat index is as follows:

[0132] (1) Flight path length L

[0133]

[0134] In Equation (26), S k represents the distance that the drone moves from time k to time k + 1, which is equal to the distance between the drone and the target at time k and the distance between the drone and the target at time k + 1 the absolute value of the difference;

[0135]

[0136] In Equation (27), the flight path length L is the iterative sum of S k and N is the number of algorithm iterations;

[0137] (2) Path maximum distortion LS

[0138] The distortion τ of the flight curve of the drone at time k k is:

[0139]

[0140] In Equation (28), T k is the tangent vector of the flight curve of the drone at time k, N k is the normal vector of the flight curve of the drone at time k, B kis the cross product of the normal vectors of the UAV flight curve at time k (also known as the binormal vector of the hyperbola), B′ k is the derivative of the binormal vector of the UAV flight curve at time k;

[0141] LS = max{τ 0 , τ 1 ,..., τ N} (29)

[0142] (3) Path global distortion GS

[0143]

[0144] Integrate the distortion on the path to obtain the path global distortion measure GS. In Equation (36), τ(s) is the distortion of the point s on the path, and a and b are the starting and ending points of the UAV flight path.

[0145] (4) Whether a collision occurs

[0146]

[0147] In Equation (31), when the distance between the UAV and the obstacle at time k is less than or equal to the radius R of the obstacle, a collision occurs between the UAV and the obstacle; when the distance between the UAV and the obstacle at time k is greater than the radius R of the obstacle, no collision occurs between the UAV and the obstacle;

[0148] (5) Threat index

[0149]

[0150] In Equation (32), when , when a collision occurs between the UAV and the obstacle, the threat index is equal to ∞; when , when the UAV is within the obstacle repulsion influence range but no collision occurs, the threat index is equal to ν k is the speed of the UAV at time k, is the distance between the UAV and the center of the obstacle at time k; when , when the UAV has not entered the obstacle repulsion influence range, the threat index is 0.

[0151] Embodiment

[0152] As Figure 1 shown, an unmanned aerial vehicle trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control provided by an embodiment of the present invention includes the following steps:

[0153] Step 1: Construct a simulation environment of various static multi-obstacles and dynamic multi-obstacles with a size of 10×10×10 (km) in MatlabR2018b software;

[0154] Step 2: Build the dynamic model of the UAV in the simulation environment of various static multi-obstacles and dynamic multi-obstacles; Define the geographic coordinate system O I (X I ,Y I ,Z I ), the body coordinate system O B (X B ,Y B ,Z B ) of the UAV, set the position vector ζ, velocity vector ν, attitude vector η, Euler angle rate vector ω, state quantity X, control quantity U and output quantity Y of the UAV in the geographic coordinate system O I (X I ,Y I ,Z I );

[0155] Step 3: Construct the Adaptive Nonlinear Model Predictive Control (ANMPC) algorithm;

[0156] Define the cost function of the adaptive nonlinear model predictive control based on the UAV tracking performance:

[0157]

[0158] In Equation (33), within the prediction interval T of one local trajectory planning and tracking, x′(τ) is the predicted UAV system state vector, and y ref (τ) is the predicted UAV flight trajectory; the target position cost function E and the stage cost function L are respectively:

[0159]

[0160] Let t k+i , i = 0, 1,... T represent the continuous sampling times. At each sampling moment, for x(t k+i ) and y ref (t k+i ), when ||x tar - x(t k+i )|| ≤ Δ, it means that the UAV has reached the target position of this local trajectory planning and tracking, where is the tolerance specified by the user; in order to better control the UAV, it is necessary to minimize the cost function J(x′(τ), y ref (τ)), and the adaptive nonlinear model predictive control optimization problem is:

[0161]

[0162] In Equation (36), A represents the weight coefficient matrix of the system state; B represents the weight coefficient matrix of the controller output; X, U, and Y are the data sets of the state, input, and output trajectories of the UAV, respectively. u min (τ) and u max (τ) are the minimum and maximum control amounts, respectively.

[0163] Step 4: Construct a reward function for the tracking performance of the UAV for the reference trajectory generated by the adaptive nonlinear model predictive control algorithm;

[0164] The reward function of the present invention is set considering the following four rewards: (1) trajectory tracking reward r traj ; (2) target position reward r tar ; (3) completion reward r c ; (4) smoothness reward r smo ;

[0165] Step 5: Construct a network framework based on the deep reinforcement learning and the adaptive nonlinear model predictive control algorithm. The main components of the network framework of the adaptive nonlinear model predictive control based on the deep reinforcement learning are as Figure 1 shown. An autonomously flying UAV explores in the constructed three-dimensional space and interacts with the environment. In a local trajectory planning, the UAV uses the adaptive nonlinear model predictive control algorithm based on the deep reinforcement learning to plan an optimal reference trajectory y from the current state to the target point ref and predicts the state vector x' of the UAV. Subsequently, the control vector u is output through the smooth nonlinear control law function g to track the optimal reference trajectory.

[0166] The actions generated by the policy network are the weight coefficient matrix A of the system state and the weight coefficient matrix B of the control input in the optimization problem of the adaptive nonlinear model predictive control. By executing the generated actions, the adaptive nonlinear model predictive control calculates the predicted system state x'(τ) and the generated reference trajectory y ref (τ) that optimize the target position cost function E (Equation (10)), the stage cost function L (Equation (11)), and the smooth nonlinear control law function g. These results are used by the control system of the UAV and combined with the current state of the UAV to generate the control amount u(τ) and calculate the observation value and the reward r stored in the replay buffer.

[0167] Step 6: Set the network parameters;

[0168] The simulation experiments were conducted in a three-dimensional environment with static and dynamic obstacles. The types, sizes, positions, and quantities of the obstacles were randomly varied, and wind noise interference was introduced to simulate real situations. The present invention compared several popular deep reinforcement learning algorithms - Twin Delayed Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO), and Deep Deterministic Policy Gradient (DDPG) with a UAV trajectory planning and tracking algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control proposed by the present invention. All algorithms were configured with the same parameters: 10 3 training epochs, a 5-layer MLP architecture with 256 neurons in each layer, and a learning rate of 3×10 -4 for the Adam policy optimizer. Table 1 summarizes the main hyperparameters of each algorithm.

[0169] Table 1 Hyperparameter settings of Twin Delayed Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO), Deep Deterministic Policy Gradient (DDPG), and a UAV trajectory planning and tracking algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control proposed by the present invention

[0170]

[0171] Step 7: According to the network parameters set in Step 6, input the reward function constructed in Step 4 into the network framework of the algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control constructed in Step 5; use the dynamic model of the UAV constructed in Step 2 as the control object, and train it in the multiple static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1. After each training round, output the training weight file; by comparing with the training results generated in all previous training rounds before this training, select the weight file corresponding to the training result with the highest reward value and the shortest flight path length as the optimal weight file;

[0172] Step 8: Input the optimal weight file obtained in Step 7 and the dynamic model of the UAV after applying it to the algorithm network framework of the deep reinforcement learning and adaptive nonlinear model predictive control constructed in Step 5 into the multiple static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1 for testing, and obtain the UAV flight trajectory diagram, reward value r, flight path length L, maximum path distortion LS, global path distortion GS, whether a collision occurs, and threat index results.

[0173] As Figure 2 shown, the present invention provides a network framework based on the SoftActor-Critic algorithm and the adaptive nonlinear model predictive control algorithm.

[0174] In the present invention, the environment module of the Soft Actor-Critic (SAC) algorithm combines the adaptive nonlinear model predictive control (ANMPC) optimization problem with the unmanned aerial vehicle (UAV) system. The adaptive nonlinear model predictive control algorithm is responsible for generating the predicted reference flight trajectory y ref and the UAV state x′, providing dynamic reference support for the flight of the UAV. The network structure of this framework includes the Actor and Critic modules. Among them, the Actor module consists of a policy network, aiming to find and select actions that maximize the total reward; the Critic module improves the policy network based on the feedback temporal difference error (TD error). The SAC algorithm includes a policy network, two Q-function networks, two target Q-function networks, and an optimization problem, constituting a complete learning and decision-making framework. As an off-policy type of deep reinforcement learning algorithm, SAC stores previous interaction data through a replay buffer pool. During the update iteration process, the algorithm randomly extracts the stored data to improve the Actor and Critic modules, thereby improving the training efficiency and policy performance.

[0175] As Figure 3 shown, the present invention provides a method for defining the coordinate system of a UAV.

[0176] The UAV is a nonlinear, multivariable, strongly coupled underactuated control system, and directly performing dynamic modeling on it is relatively complex. To simplify the complexity of the model, the following assumptions are made for the research object:

[0177] (1) The quadrotor is a uniformly symmetric rigid body;

[0178] (2) The mass and moment of inertia of the quadrotor do not change;

[0179] (3) The geometric center of the quadrotor coincides with its center of gravity;

[0180] (4) The quadrotor is only affected by gravity and propeller thrust.

[0181] Based on the Newton-Euler method, the dynamic characteristics of the UAV are analyzed, and an accurate UAV model is established to understand and analyze the system characteristics of the UAV and serve as the simulation model for simulation experiments. Secondly, the dynamic model is reasonably simplified as the prediction model for the model predictive control method.

[0182] As Figure 4 shown, the present invention provides a comparison effect diagram of the flight trajectory and the predicted reference trajectory in one-time local trajectory planning and tracking.

[0183] From the perspective of deep reinforcement learning, in each flight mission, the entire flight mission is divided into k segments, and the UAV is controlled to fly over k target points. Therefore, each flight mission consists of k local trajectory planning and tracking. After the UAV completes the previous trajectory planning and tracking each time, three observation values are sent to the UAV, namely: the initial velocity of the UAV's current local trajectory planning and tracking Initial velocity vector And the vector The included angle between End point p ss The distance between the end point p and the initial point p0 Figure 4 Figure is a comparison effect diagram of the flight trajectory and the predicted reference trajectory in a local trajectory planning and tracking. In addition, in order to verify the robustness and stability of the algorithm, static obstacles, dynamic obstacles and wind noise are added

[0184] The effects of the present invention are further described below in conjunction with simulation experiments

[0185] 1. Simulation experiment conditions

[0186] The hardware platform of the simulation experiment of the present invention is: CPU: Intel I5-13490F, main frequency 2.5GHz, 16G running memory; GPU: RTX4060Ti

[0187] The software platform of the simulation experiment platform of the present invention is: Windows 11 operating system, PyCharm Community Edition 2022.1.4, PyTorch 2.01, CUDA 12.4 and MATLAB R2018b

[0188] The simulation experiment is carried out in a constructed three-dimensional environment with static and dynamic obstacles. The types, sizes, positions and numbers of obstacles change randomly, and wind noise interference is introduced to simulate the real situation. Several popular deep reinforcement learning algorithms - Twin Delayed Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO) and Deep Deterministic Policy Gradient (DDPG) are compared with an algorithm for UAV trajectory planning and tracking based on deep reinforcement learning and adaptive nonlinear model predictive control proposed by the present invention. All algorithms are configured with the same parameters: 10 3 Training times, 5-layer MLP architecture, 256 neurons in each layer, and the learning rate of the Adam policy optimizer is 3×10 -4。Table 1 summarizes the main hyperparameters of each algorithm. Table 1 Hyperparameter settings of Twin Delayed Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO), Deep Deterministic Policy Gradient (DDPG), and an unmanned aerial vehicle trajectory planning and tracking algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control proposed by the present invention

[0189]

[0190] 2. Simulation steps

[0191] Step 1: According to the set network parameters, input the reward function of the tracking performance of the constructed unmanned aerial vehicle for the reference trajectory generated by the adaptive nonlinear model predictive control algorithm into the network framework of the algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control. Then, take the dynamic model of the unmanned aerial vehicle as the control object and conduct training in the constructed simulation environments with multiple static multi-obstacles and dynamic multi-obstacles. After each round of training, output the training weight file. By comparing with the previous training results, select the weight file corresponding to the training result with the highest reward value and the shortest flight path length as the optimal weight file;

[0192] Step 2: Input the optimal weight file obtained in Step 1 and the dynamic model of the unmanned aerial vehicle applying the algorithm network framework based on deep reinforcement learning and adaptive nonlinear model predictive control into the constructed simulation environments with multiple static multi-obstacles and dynamic multi-obstacles for testing, and obtain the unmanned aerial vehicle flight trajectory diagram, reward value r, flight path length L, path maximum distortion LS, path global distortion GS, whether a collision occurs, and threat index results.

[0193] 3. Simulation content and its result analysis

[0194] The simulation experiment of the present invention is verified by using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG), and an unmanned aerial vehicle based on the unmanned aerial vehicle trajectory planning and tracking algorithm proposed by the present invention which combines deep reinforcement learning and adaptive nonlinear model predictive control in four dynamic multi-obstacle environments. By comparing the performances of different algorithms in these environments, the effectiveness of the unmanned aerial vehicle trajectory planning and tracking method combining deep reinforcement learning and adaptive nonlinear model predictive control is verified.

[0195] In order to verify the performance of an unmanned aerial vehicle trajectory planning and tracking method proposed by the present invention from all aspects and meet the actual needs, five groups of comparative experiments are conducted in different static or dynamic multi-obstacle environments. The main experimental data for comparison are:

[0196] (1) Reward value r: During the training process, the drone receives feedback from the environmental model, which reflects its performance. The higher the value, the better the performance of the drone.

[0197] (2) Flight path length L: The flight distance of the drone from the starting point to the target point, indicating the efficiency of the drone. The smaller the value, the higher the efficiency.

[0198] (3) Maximum path distortion GS (°): This parameter represents the maximum change in the heading angle between two consecutive samples during the flight of the drone, serving as a measure of flight stability. The smaller the value, the better the stability.

[0199] (4) Path global distortion LS (°): The total change in the heading angle of the drone during the entire flight is measured by this parameter, indicating the flatness of the drone's trajectory. The smaller the value, the higher the smoothness.

[0200] (5) Threat index: The degree of threat of obstacles to the drone during flight, reflecting the obstacle avoidance performance of the drone. The smaller the value, the stronger the obstacle avoidance ability of the drone.

[0201] (6) Whether a collision occurs: Whether the drone collides with an obstacle during flight, reflecting the safety of the drone. "No" means the drone does not collide with an obstacle; "Yes" means the drone collides with an obstacle.

[0202] In order to verify the performance of a UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control proposed in the present invention in all aspects and meet the actual requirements, five groups of comparative experiments were carried out in a variety of constructed static multi-obstacle and dynamic multi-obstacle simulation environments.

[0203] 1. Validity verification experiment

[0204] In different dynamic multi-obstacle environments, the UAV conducts optimal path planning and tracking in a globally known environment from the starting point (0, 0, 0) to the target point (10, 10, 10). The experimental results are shown in Table 2 and Figure 5 、 6 as shown. Table 2 shows the training results of four algorithms in different dynamic multi-obstacle environments; Figure 5 is the cumulative reward curve graph of the policy network of the present invention (ANMPC-SAC algorithm) and the twin-delayed deep deterministic policy gradient algorithm (TD3), proximal policy optimization algorithm (PPO), and deep deterministic policy gradient algorithm (DDPG) in 4 dynamic multi-obstacle environments; Figure 6Simulation diagrams of drones using the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3), Proximal Policy Optimization algorithm (PPO), Deep Deterministic Policy Gradient algorithm (DDPG), and the present invention (ANMPC-SAC algorithm) in a dynamic multi-obstacle environment.

[0205] Table 2 Experimental data table for effectiveness verification

[0206]

[0207]

[0208] Among these four algorithms, the ANMPC-SAC algorithm generally shows better superiority, while the TD3 algorithm shows the worst performance in handling this task. The experimental results are analyzed from the following aspects:

[0209] (1) Generally speaking, the reward curve of the present invention converges faster. It converges after about 200 training rounds and stabilizes near the highest reward value, indicating that it has found the optimal optimization path. This performance is significantly better than other algorithms. In addition, the present invention shows good stability during the entire training process, and the fluctuation range of the reward curve is small, indicating that it can maintain relatively stable performance during training.

[0210] (2) As shown in Table 2, in four dynamic multi-obstacle environments, the experimental data of the present invention such as reward, length, threat index, and collision are better than other algorithms. The rewards of the present invention are all better than other algorithms, indicating that it has better training effects and convergence speed; the lengths of the present invention are all better than other algorithms, indicating that it has better path planning and tracking capabilities; in Environments 1, 2, and 3, the GS(°) and LS(°) of PPO are slightly better than those of the present invention, but the present invention is better than other algorithms (TD3, DDPG). This indicates that compared with the present invention, PPO has better trajectory stability and smoothness; in Environments 1, 3, and 4, TD3 collided with obstacles, resulting in an infinite threat index, while the present invention did not collide with obstacles and the threat coefficient is small. Only in Environment 1, the threat index is slightly larger than that of DDPG. This shows that the present invention has better safety and obstacle avoidance capabilities.

[0211] (3) As Figure 5As shown in the figure, in four dynamic multi-obstacle environments, the reward return curves of the four algorithms all have convergence characteristics, indicating the effectiveness of the four algorithms, namely TD3, PPO, DDPG, and the present invention (ANMPC-SAC algorithm). The overall reward of the present invention in 4 different dynamic multi-obstacle environments is greater than that of other algorithms and has higher stability. Among them, the present invention (ANMPC-SAC algorithm) is the red curve, the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3) is the blue curve, the Proximal Policy Optimization algorithm (PPO) is the orange curve, and the Deep Deterministic Policy Gradient algorithm (DDPG) is the green curve.

[0212] (4) Figure 6 It is a visualization result graph of the training of the four algorithms in Environment 3. TD3 collided with the obstacle, while PPO, DDPG, and the present invention did not collide with the obstacle. Compared with other algorithms, the flight trajectory of the present invention is more stable, smoother, and the path is shorter, without colliding with dynamic obstacles. Among them, the flight trajectory of the present invention (ANMPC-SAC algorithm) is the red curve, the flight trajectory of the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3) is the blue curve, the flight trajectory of the Proximal Policy Optimization algorithm (PPO) is the orange curve, and the flight trajectory of the Deep Deterministic Policy Gradient algorithm (DDPG) is the green curve.

[0213] 2. Training Time Comparison Experiment

[0214] Through simulation in a dynamic multi-obstacle environment, the total training time and time per iteration of the four algorithms (TD3, PPO, DDPG, and ANMPC-SAC) in solving the path planning and tracking problem were evaluated and compared. The experimental results are as Figure 7 shown.

[0215] As Figure 7 shown, the present invention is significantly superior to the other three algorithms in terms of both total training time and time per iteration. Its main advantage stems from the optimization in the adaptive non-linear model predictive controller, especially in the design of the cost function. By introducing the target position cost function, stage cost function, and smooth non-linear control law function, the optimization problem is simplified, and the computational amount and complexity are reduced. These optimizations effectively improve the computational speed of the system, enabling the algorithm to complete each iteration in a shorter time, while reducing the overhead of each calculation during the training process, thereby enhancing the overall training efficiency. In contrast, the other three algorithms are more efficient in terms of computational complexity and computational resource consumption, but due to their inability to fully utilize the model information in the control strategy, they result in longer training times and slower convergence speeds.

[0216] These results indicate that through the combination of deep reinforcement learning and adaptive non - linear model predictive control, the present invention can perform trajectory planning and tracking more efficiently in a dynamic environment, with faster computing speed and lower computational overhead, providing better real - time support for practical applications.

[0217] 3. Target Change Experiment

[0218] In practical applications, the target point may change its position due to various unexpected situations. Therefore, further study the trajectory planning and tracking performance of the UAV when the target changes. In this set of comparative experiments, the starting point of the UAV is fixed, and the target point moves randomly. Next, experiments are carried out in a dynamic multi - obstacle environment. The experimental data are shown in Table 3, and the visualized experimental results (trained in Environment 3) are as Figure 8 shown.

[0219] Table 3 Data Table of Target Change Experiment

[0220]

[0221] Compared with the effectiveness verification experiment, due to the change of the target point in each task, the difficulty of model training increases greatly, resulting in the failure of TD3. Analyze the experimental results from the following aspects:

[0222] Generally speaking, considering the random distribution characteristics of the target points, the training experimental results shown in Table 3 fail to reach the level of the training results in Table 2.

[0223] As shown in Table 3, facing the challenges brought by the randomness of the target points, the present invention demonstrates remarkable robustness. Its performance not only does not show an obvious decline, but on the contrary, compared with other algorithms, it performs outstandingly in maintaining high efficiency and high accuracy. This indicates that the present invention has unique advantages and adaptability in dealing with complex training tasks with random target points.

[0224] As shown in Table 3 and Figure 8 shown, while maintaining high performance, the flight trajectory of the present invention can still maintain high stability and smoothness, and does not collide with dynamic obstacles, having high safety and obstacle avoidance ability. Among them, the light blue point is the starting point, the magenta point is the target point, and the red trajectory is the flight trajectory of the UAV. The flight trajectory of the present invention (ANMPC - SAC algorithm) is the red curve.

[0225] 1. Environment Change Experiment

[0226] In more challenging practical applications, when the UAV has not explored the global environmental information, it can only perform path planning based on limited environmental information and target point information. The ability of the UAV to cope with environmental changes has more substantial practical benefits, which means that the UAV can work in any environment. Next, experiments are carried out in multiple dynamic multi-obstacle environments to verify the adaptability and efficiency of a UAV trajectory planning and tracking method, system, device and medium based on deep reinforcement learning and adaptive non-linear model predictive control proposed by the present invention in dealing with environments of different complexities. The experimental data is shown in Table 4, and the visualized experimental results are as Figure 9 shown.

[0227] Table 4 Experimental data table of environmental changes

[0228]

[0229] In this group of comparative experiments, each obstacle in the dynamic multi-obstacle environment is changing. As the task complexity increases, the requirement for the stability of the algorithm becomes higher and higher. At this time, the other three algorithms may collide due to difficulty in adapting to the changes of dynamic obstacles and ultimately fail to achieve the goal of reaching the target point. In contrast, the method of the present invention can successfully meet this challenge. Figure 9 The meaning of Figure 8 is similar, and it intuitively demonstrates the ability of the present invention to perform trajectory planning and tracking in a variety of complex and changeable dynamic multi-obstacle environments.

[0230] 2. Experiments on the influence of wind noise interference

[0231] In practical applications, the flight of the UAV will inevitably be disturbed by wind noise, and this kind of disturbance has significant uncertainty. Therefore, a UAV trajectory planning and tracking method, system, device and medium based on deep reinforcement learning and adaptive non-linear model predictive control proposed by the present invention must have sufficient robustness and stability to ensure that the UAV can effectively resist the interference of wind noise during flight. To verify this, a group of experiments are designed to compare the performance of the present invention in dealing with wind noise of different intensities.

[0232] Table 5 Experimental data table of the influence of wind noise interference

[0233]

[0234] Considering that the experimental results often show similarity when the wind force fluctuates within a small range, the present invention decides to expand the intensity range of wind noise from 0.01 to 0.5 and use multiplicative growth to simulate more significant wind force changes. The specific data of the experiment are shown in Table 5, and the visualized experimental results are as Figure 10As shown. Through in-depth analysis of the experimental results, it is found that wind noise has the following effects on the flight of the UAV:

[0235] (1) Under the environmental conditions with wind noise introduced, TD3 cannot avoid collisions with dynamic obstacles during each flight mission.

[0236] (2) According to the data shown in Table 4, as the environmental wind noise intensity increases to 0.1, DDPG begins to collide with dynamic obstacles; when the wind noise further increases to 0.5, PPO also shows the phenomenon of colliding with dynamic obstacles.

[0237] (3) Generally speaking, the present invention demonstrates many remarkable advantages, including the highest reward value, the best anti-wind noise interference ability, the shortest path length, and good smoothness and stability performance.

[0238] In summary, from the above five groups of comparative experiments, it can be concluded that: The performance of the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3) is the worst. It will collide with obstacles in the set dynamic multi-obstacle environment and cannot complete the trajectory planning task. The trajectory planned by the Proximal Policy Optimization algorithm (PPO) has better smoothness and has a certain degree of generality and anti-interference ability. However, it is sensitive to hyperparameters and the tuning is complex, resulting in too long training time. The Deep Deterministic Policy Gradient algorithm (DDPG) has poor adaptability when facing different dynamic multi-obstacle environments and cannot be applied to unknown random environments. The UAV trajectory planning and tracking algorithm based on deep reinforcement learning and adaptive non-linear model predictive control proposed in the present invention shows the best performance in terms of robustness, stability and safety. It can cope with the changes of targets and environments, shows strong obstacle avoidance and anti-interference abilities against the interference of dynamic obstacles and wind noise, and reflects a high level of intelligent decision-making.

[0239] This invention proposes an Unmanned Aerial Vehicle (UAV) Trajectory Planning and Tracking (ANMPC-SAC) algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control. This algorithm has been innovated in many aspects, optimizing the training process, network architecture, and algorithm model, significantly improving the UAV's trajectory planning and tracking performance in complex environments. First, for the actual requirements of this study, the core components of the reinforcement learning system are redesigned, especially the Agent, Environment, and Reward. After optimization, these components are more adapted to the characteristics of UAV trajectory planning and tracking, enabling the algorithm to make more efficient and accurate decisions when facing dynamic environments. Through this customized design, reinforcement learning can better understand and adapt to the environmental changes and task requirements in actual flight missions, thus accelerating the learning process and improving the accuracy of trajectory tracking. Second, the cost function in the adaptive nonlinear model predictive control optimization problem is simplified, reducing the computational amount and complexity, and significantly improving the learning efficiency. By introducing the target position cost function, stage cost function, and smooth nonlinear control law function, not only the calculation process is optimized, but also the smoothness and feasibility of trajectory planning are ensured. This optimized design effectively reduces the computational burden of the system, enabling the UAV to complete trajectory planning and control calculations in a shorter time, thus improving the efficiency of real-time control. In addition, under the adaptive nonlinear model predictive control algorithm, the SoftActor-Critic (SAC) algorithm is integrated. The SAC algorithm promotes the balance between exploration and exploitation by introducing the principle of maximum entropy, enabling the system to quickly adjust the strategy and generate the optimal reference trajectory in a dynamic environment. The SAC algorithm can also dynamically adjust the optimization parameters according to environmental changes, ensuring that the UAV can stably track the reference trajectory and accurately reach the target in complex flight missions. To verify the effectiveness and application performance of this invention, a number of experiments are designed and carried out, including effectiveness verification, training time analysis, target change adaptability, environmental change adaptability, and the impact of wind noise interference, etc. The experimental results fully prove the superiority of this invention in many aspects, especially outstanding in terms of robustness, stability, and safety. Specifically, this algorithm can effectively handle dynamic obstacles, target changes, and external disturbances, demonstrating strong adaptability and real-time response capabilities. The experimental results also show that the UAV using the ANMPC-SAC algorithm of this invention has excellent performance in multiple key performance indicators, including cumulative reward, flight path length, path global distortion, collision avoidance ability, and threat index, etc. Especially when facing dynamic obstacles, this algorithm can quickly sense the presence of obstacles and timely adjust the flight path to avoid collisions, ensuring the safety of flight missions.In summary, the UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control proposed by the present invention has good online planning ability, excellent performance and high safety, can achieve efficient, stable and safe UAV flight in complex dynamic environments, and has broad application prospects.

[0240] The present invention also provides a UAV trajectory planning and tracking system based on deep reinforcement learning and adaptive nonlinear model predictive control, including:

[0241] A simulation environment construction module: used to implement the construction of various static multi-obstacle and dynamic multi-obstacle simulation environments in Step 1;

[0242] A dynamic model construction module of the UAV: used to implement the construction of the dynamic model of the UAV in Step 2, including its motion equation, attitude control model and external disturbance modeling;

[0243] An adaptive nonlinear model predictive control (ANMPC) algorithm construction module, used to implement the construction of the adaptive nonlinear model predictive control (ANMPC) algorithm in Step 3;

[0244] A deep reinforcement learning reward function construction module: used to implement the construction of the reward function for the UAV to track the performance of the reference trajectory generated by the adaptive nonlinear model predictive control algorithm in Step 4;

[0245] A network construction module: used to implement the construction of a network framework based on deep reinforcement learning and adaptive nonlinear model predictive control algorithms in Step 5 based on the adaptive nonlinear model predictive control (ANMPC) algorithm constructed in Step 3;

[0246] A network training module: used to implement, in Step 7, according to the network parameters set in Step 6, input the reward function constructed in Step 4 into the network framework of the algorithm based on deep reinforcement learning and adaptive nonlinear model predictive control constructed in Step 5; use the dynamic model of the UAV constructed in Step 2 as the control object, and perform training in the various static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1. After each round of training ends, output the training weight file; by comparing with the training results generated in all previous training rounds before this training, select the weight file corresponding to the training result with the highest reward value and the shortest flight path length as the optimal weight file;

[0247] A test result output module: used to implement, in Step 8, input the optimal weight file obtained in Step 7 and the dynamic model of the UAV after applying it to the algorithm network framework of the deep reinforcement learning and adaptive nonlinear model predictive control constructed in Step 5 into the various static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1 for testing to obtain the UAV flight trajectory map、 Reward value r 、 Flight path length L 、 Path maximum distortion LS, path global distortion GS, whether a collision occurs, and threat index results.

[0248] The present invention also provides a UAV trajectory planning and tracking device based on deep reinforcement learning and adaptive nonlinear model predictive control, including:

[0249] Memory: storing a computer program of the above-mentioned UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control, which is a computer-readable device;

[0250] Processor: used to implement the above-mentioned UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control when executing the computer program.

[0251] The present invention also provides a computer-readable storage medium, which stores a computer program that can implement the above-mentioned UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control when executed by a processor.

Claims

1. A method for UAV trajectory planning and tracking based on deep reinforcement learning and adaptive nonlinear model predictive control, characterized in that, It includes the following steps: Step 1: Construct multiple static multi-obstacle and dynamic multi-obstacle simulation environments; Step 2: Construct the dynamic model of the UAV, including its motion equation, attitude control model, and external disturbance modeling; Step 3: Construct an Adaptive Nonlinear Model Predictive Control (ANMPC) algorithm; Step 4: Construct a reward function for the tracking performance of the UAV for the reference trajectory generated by the adaptive nonlinear model predictive control algorithm; Step 5: Based on the Adaptive Nonlinear Model Predictive Control (ANMPC) algorithm constructed in Step 3, construct a network framework based on the deep reinforcement learning and the adaptive nonlinear model predictive control algorithm; Step 6: Set network parameters; Step 7: According to the network parameters set in Step 6, input the reward function constructed in Step 4 into the network framework based on the deep reinforcement learning and the adaptive nonlinear model predictive control algorithm constructed in Step 5; use the dynamic model of the UAV constructed in Step 2 as the control object, and perform training in the multiple static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1. After each round of training, output the training weight file; by comparing with the training results generated in all previous training rounds before this training, select the weight file corresponding to the training result with the highest reward value and the shortest flight path length as the optimal weight file; Step 8: Input the optimal weight file obtained in Step 7 and the dynamic model of the UAV after applying it to the algorithm network framework based on the deep reinforcement learning and the adaptive nonlinear model predictive control constructed in Step 5 into the multiple static multi-obstacle and dynamic multi-obstacle simulation environments constructed in Step 1 for testing, and obtain the UAV flight trajectory diagram, reward value r, flight path length L, maximum path distortion LS, global path distortion GS, whether a collision occurs, and threat index results.

2. A method for UAV trajectory planning and tracking based on deep reinforcement learning and adaptive nonlinear model predictive control according to claim 1, characterized in that, The implementation method of the said Step 1 includes: Step 101: Set multiple static multi-obstacle environments: Set the center position (X I , Y I , Z I ), radius R, and the number of spheres n sphere ; the center position (X I , Y I , Z I ) of the bottom circle of the cylinder, radius R, height H, and the number of cylinders n cylinder ; the center position (X I , Y I , Z I ) of the bottom circle of the cone, radius R, height H, and the number of cones n cone ; Step 102: Set up multiple dynamic multi-obstacle environments: Set the center position (X I , Y I , Z I ) of the sphere, the radius R, the motion trajectory r(t) of the sphere, the moving speed v, and the number n of spheres sphere .

3. A UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control according to claim 1, characterized in that, The specific implementation method of the said Step 2 is: Based on the Newton-Euler method, analyze the dynamic characteristics of the UAV and establish the dynamic model of the UAV; Define the geographic coordinate system O I (X I , Y I , Z I ) and the UAV body coordinate system O B (X B , Y B , Z B ). The position vector of the UAV in the geographic coordinate system O I is defined as ζ: The velocity vector ν of the UAV is defined as: The attitude vector of the UAV in space is defined as η: The roll angle φ in Equation (3) represents the rotation angle of the UAV along the X-axis of the geographic coordinate system O I ; the pitch angle θ represents the rotation angle of the UAV along the Y-axis of the geographic coordinate system O I ; the heading angle ψ represents the rotation angle of the UAV along the Z-axis of the geographic coordinate system O I ; and the Euler angle rate vector ω is: I The roll angle φ in Equation (3) represents the rotation angle of the UAV along the X-axis of the geographic coordinate system O I ; the pitch angle θ represents the rotation angle of the UAV along the Y-axis of the geographic coordinate system O I ; the heading angle ψ represents the rotation angle of the UAV along the Z-axis of the geographic coordinate system O In Equation (4), p, q, and r are the angular velocities of the drone rotating along the X B , Y B , and Z B axes in the drone body coordinate system O B . The control quantity state space is defined as: U = [U1 U2 U3 U4] T (5) In formula (5), It is the sum of the thrusts generated by the four rotors, representing the UAV along the UAV body coordinate system O B The coordinate axis Z B displacement movement; is the rolling moment, which represents the UAV along the UAV body coordinate system O B The coordinate axis X B a rotational tumbling motion; is the pitch moment, representing the UAV along the UAV body coordinate system O B The Y axis B Rotational pitch motion; is the yaw moment, representing the yaw moment along the UAV body coordinate system O B Coordinate axis Z B Rotational yaw motion; K is the lift coefficient; K d is the anti-torque coefficient; l is the distance from the center of each rotor of the quadrotor to its center of mass; ω1, ω2, ω3, ω4 are the rotation speeds of the four rotors respectively; The state space and state vector of the dynamic model of the UAV are: In Equation (6), the state vector of the UAV is Therefore, the state of the UAV at time k, X(k) ∈ X; k x , k y , k z is obtained by decomposing the air resistance coefficient in the geographic coordinate system O I ; k φ , k θ , k ψ are the components of the total air resistance torque coefficient on the X I , Y I , Z I axes respectively; I x , I y , I z are the moments of inertia of the UAV along the coordinate axes X B , Y B , Z B of the UAV body coordinate system O B ; J RP is the moment of inertia of the propeller along the rotation axis of the UAV; Ω = -ω1 + ω2 - ω3 + ω4 is the algebraic sum of the motor speeds.

4. A method for UAV trajectory planning and tracking based on deep reinforcement learning and adaptive nonlinear model predictive control according to claim 1, characterized in that, The specific implementation method of the said Step 3 is: Adaptive Nonlinear Model Predictive Control (ANMPC) generates a reference trajectory and control input for the UAV closed-loop negative feedback control system. Adaptive nonlinear model prediction uses a nonlinear system model, a nonlinear control law, a set of state vectors, control vectors, and the maximum and minimum values of the output vector control quantities, as well as the constraints of static and dynamic obstacles in the environment, and connects the nonlinear object with the control law; The unmanned aerial vehicle is a non - linear system, where the state vector, control vector, and output vector are respectively and Assume that the output vector is a subset of the state vector Let \(f(x(k),u(k)):X\times U\rightarrow X\) be a known smooth non - linear function based on online parameter estimation, and \(g(x(k),y ref (k)):X\times Y\rightarrow U\) be the mapping of a smooth non - linear control law; given a finite prediction interval \(T\), the finite sequence of predicted control vectors \(u(t_{k} +T-1 )\) and predicted state vectors \(x(t k+T )\) is as follows: u(t k+T-1 ) = (u0, u1, …, u N-1 ) and x(t k+T ) = (x0, x1, …, x N ) (7) For the continuous state space model, first discretize it. The general form of the discrete expression of the UAV state space is: x(t k+1 ) = f(x(t k ), u(t k )) = A·x(t k ) + B·u(t k ) (8) The adaptive nonlinear model predictive control algorithm generates the reference trajectory y for the UAV flight ref , and then, the UAV closed-loop negative feedback control system composed of the controlled object and the control law tracks the reference trajectory y ref to fly to the target position of the current local trajectory planning and tracking; the cost function J of the adaptive nonlinear model predictive control optimization problem is as follows: In Equation (9), the target position cost function E and the stage cost function L are respectively: During the prediction interval T of a local trajectory planning and tracking, x′(τ) is the predicted UAV system state vector, and y ref (τ) is the predicted UAV flight reference trajectory. The smaller the target position cost function E (Equation (10)), the closer the actual flight end point of the UAV is to the predicted trajectory end point; the smaller the stage cost function L, the closer the flight trajectory of the UAV is to the predicted reference trajectory; the smaller the smooth nonlinear control law function g, the less energy the UAV consumes; The errors in equations (10) and (11) are weighted by the matrices and and will be adaptively adjusted using a deep reinforcement learning algorithm in this work. k represents the k-th trajectory planning and tracking during flight. Within the prediction interval T of a local trajectory planning and tracking, the UAV closed-loop negative feedback control system reaches the target stable setpoint x tar from the current state x. The adaptive nonlinear model predictive optimization problem is shown in equation (12). Let t k+i , i = 0, 1,... T represent the continuous sampling times. At each sampling time, for x(t k+i ) and y ref (t k+i ), when ||x tar - x(t k+i )|| ≤ Δ, it means that the UAV has reached the target position of this local trajectory planning and tracking, where is the tolerance specified by the user. Minimize the cost function J(x′(τ), y ref (τ)) to control the UAV. The adaptive nonlinear model predictive control optimization problem is as follows: In Equation (12), A represents the weight coefficient matrix of the system state; B represents the weight coefficient matrix of the controller output; X, U, and Y are the data sets of the state, input, and output trajectories, respectively; u min (τ) and u max (τ) are the minimum value and the maximum value of the control quantity, respectively; The calculation process of adaptive nonlinear model predictive control is as follows: First, detect the UAV at t n The state at the moment, then x′(τ)=f(x(τ),u(τ)),u(τ=g(x′(τ),y ref (τ)) and u min (τ)≤u(τ)≤u max (τ) constraint, in the prediction interval [t k ,t k+T ] minimize the cost function J(x′(τ),y ref (τ)) to provide x′(t k+i ) and y ref (t k+i ) value prediction; finally, the predicted UAV system state vector x′(t k+i ) and the reference trajectory y ref (t k+i ) is input into the UAV closed-loop negative feedback control system, and the control vector u(t k+i ) for trajectory tracking; the process is repeated in real time at an iteration rate set by the user until the desired target state value is reached.

5. A method for UAV trajectory planning and tracking based on deep reinforcement learning and adaptive nonlinear model predictive control according to claim 1, characterized in that The specific implementation method of the said Step 4 is: The reward of deep reinforcement learning is a reward function for the tracking performance of the drone with respect to the reference trajectory generated by the adaptive non - linear model predictive control algorithm, which is respectively defined as the trajectory tracking reward r traj , the target position reward r ss , the completion reward r c and the smoothness reward r smo : (1) Trajectory tracking reward r traj , r traj represents the degree of fit between the real-time flight trajectory of the UAV and the predicted reference trajectory generated by the adaptive nonlinear model predictive control algorithm; the trajectory tracking reward r traj is as follows: In Equation (13), is the root mean square error for generating the predicted reference trajectory and the flight trajectory, r traj,max and r traj,th are the maximum value and the threshold of the trajectory tracking reward r traj respectively; (2) Target position reward r tar , r tar is the distance difference between the target position x(t k+T ) of the real-time flight trajectory of the UAV and the target position p of the predicted reference trajectory generated by the adaptive nonlinear model predictive control algorithm; the target position reward r tar is as follows: tar is: In Equation (14), e tar = ||x(t k+T ) - p tar || is the distance difference between the target position x(t k+T ) of the real-time flight trajectory of the UAV and the target position p tar of the trajectory generated by the adaptive nonlinear model predictive control algorithm. r tar,max and r tar,th are the maximum value and threshold of the target position reward r tar respectively; (3) Completion reward r c , r c represents the distance difference between the position x(t k ) of the real-time flight trajectory of the UAV and the trajectory target position p tar generated by the adaptive nonlinear model predictive control algorithm during a local trajectory planning and tracking process; the completion reward is r c : In Equation (15), e tar = ||x(t k ) - p tar || represents the distance difference between the position of the UAV at time t k and the trajectory target position p tar generated by the adaptive nonlinear model predictive control algorithm. r c,max and r c,th are the maximum value and threshold of the completion reward r c , respectively. Whenever the error e c exceeds the specified threshold r c,th (e c > r c,th ), this factor is more emphasized by reducing the total reward (r < 0); (4) Smooth reward r smo , r smo represents the smoothness of the flight trajectory of the UAV during a local trajectory planning and tracking process; the smooth reward r smo is as follows: In Equation (16), τ(p) is the distortion degree of the UAV flight trajectory at position p, and the distortion degree of the trajectory during one local trajectory planning and tracking process r smo,max and r smo,th are the maximum value and threshold of the smoothing reward r smo respectively; In summary, the designed reward is: r = λ1·r traj + λ2·r tar + λ3·r c + λ4·r smo (17) In Equation (17), λ1, λ2, λ3, and λ4 respectively represent the weights of four rewards.

6. A method for UAV trajectory planning and tracking based on deep reinforcement learning and adaptive nonlinear model predictive control according to claim 1, characterized in that The specific implementation method of the said Step 5 is: Optimize the trajectory planning and control of the UAV using the SoftActor-Critic algorithm in deep reinforcement learning; The SoftActor-Critic (SAC) algorithm consists of two action-value networks (Q networks), two target Q networks, and a policy network (Actor). The two action-value networks (Q networks) are trained simultaneously using the SoftActor-Critic (SAC) algorithm. A stochastic policy is adopted, and the two action-value networks (Q networks) are updated using the modified mean-square Bellman equation; The Bellman equation is: P represents the distribution followed by the state-action pairs that the Agent will encounter under the control of the policy π. The goal of the SoftActor-Critic algorithm is to maximize the entropy of the policy and the expected return to achieve efficient exploration and exploitation. The Q-function Bellman equation for reinforcement learning based on entropy regularization is redesigned as: In Equation (19), α is the coefficient that regulates the balance between the expected entropy and the reward, a′ is the next action, s′ is the next state, and γ is the discount factor used in deep reinforcement learning to regulate the influence of the near and far future; The solution of the expectation is for the next state, which comes from the replay buffer, and the next action, which comes from the current policy, and the Q value is approximately estimated using samples: Q π (s,a) ≈ r(s) + γ(Q π (s′,a′) - α log π (a′|s′)), a′ ∼ π(·|s′) (20) The loss function of the Q-function network in the SoftActor-Critic algorithm is: In Equation (21), the form of the Bellman equation is: The action-value network of the SoftActor-Critic algorithm is updated using gradient descent, and the policy network is updated using gradient ascent; the goal of the Agent's policy is to maximize the state value function V π (s), and the state value function is as follows: Equation (23) represents the expected return of the Agent starting from state s and following policy π; The optimized policy uses the reparameterization method to calculate the deterministic function of the state, policy parameters, and independent noise, and samples are drawn from it. Reparameterization enables the expectation of the action to be rewritten as the expectation of the noise; This calculation process is expressed as: a θ (s, ξ) = tanh(μ θ (s) + σ θ (s) ⊙ ξ), ξ ~ N(0, I) (24) Policy optimization is achieved by maximizing the Q function, and the Q function implicitly maximizes the entropy of the trajectory; The loss function of the policy network in the SoftActor-Critic algorithm is: From the perspective of deep reinforcement learning, in each flight mission, the entire flight mission is divided into k segments, and the UAV is controlled to fly over k target points; Therefore, each flight mission consists of k times of local trajectory planning and tracking; after the UAV finishes the previous trajectory planning and tracking each time, three observation values are sent to the UAV, which are respectively: the initial velocity of the UAV's local trajectory planning and tracking this time Initial velocity vector And the vector The included angle between Endpoint p tar The distance between the starting point p0 7. A method for UAV trajectory planning and tracking based on deep reinforcement learning and adaptive nonlinear model predictive control according to claim 1, characterized in that, The specific implementation method of Step 8 is: The calculation of the flight path length L, the maximum path distortion LS, the global path distortion GS, whether a collision occurs, and the threat index is as follows: (1) Flight path length L In formula (26), S k represents the distance that the UAV moves from time k to time k + 1, which is equal to the distance between the UAV and the target at time k and the distance between the UAV and the target at time k + 1 of the absolute value of the difference; In Equation (27), the flight path length L is equal to S k which is the iterative sum of, where N is the number of algorithm iterations; (2) Maximum path distortion LS The distortion degree τ of the UAV flight curve at time k k is as follows: In Equation (28), T k is the tangent vector of the UAV flight curve at time k, N k is the normal vector of the UAV flight curve at time k, B k is the cross product of the normal vectors of the UAV flight curve at time k (also known as the binormal vector of the hyperbola), B′ k is the derivative of the binormal vector of the UAV flight curve at time k; LS = max{τ 0 , τ 1 ,..., τ N} (29) (3) Global path distortion GS Integrate the distortion on the path to obtain the global path distortion measure GS. In Equation (36), τ(s) is the distortion of the point s on the path, and a and b are the starting and ending points of the UAV flight path; (4) Whether a collision occurs In formula (31), when the distance between the UAV and the obstacle at time k is less than or equal to the radius R of the obstacle, the UAV collides with the obstacle; when the distance between the UAV and the obstacle at time k is greater than the radius R of the obstacle, the UAV does not collide with the obstacle; (5) Threat index In formula (32), when , when the drone collides with an obstacle, the threat index is equal to ∞; when , when the drone is within the obstacle repulsion influence range but no collision occurs, the threat index is equal to ν k is the speed of the drone at time k, is the distance between the drone and the center of the obstacle at time k; when , when the drone does not enter the obstacle repulsion influence range, the threat index is 0.

8. A UAV trajectory planning and tracking system based on deep reinforcement learning and adaptive nonlinear model predictive control according to the method of any one of claims 1 to 7, characterized in that, Include: Simulation environment construction module: Used to construct various simulation environments of static multi-obstacles and dynamic multi-obstacles; Dynamic model construction module of the UAV: Used to construct the dynamic model of the UAV, including its motion equation, attitude control model, and external disturbance modeling; Adaptive Nonlinear Model Predictive Control (ANMPC) Algorithm Construction Module: used to construct the Adaptive Nonlinear Model Predictive Control (ANMPC) algorithm; Deep Reinforcement Learning Reward Function Construction Module: used to construct the reward function for the tracking performance of the UAV with respect to the reference trajectory generated by the adaptive nonlinear model predictive control algorithm; Network Construction Module: used to implement the construction of a network framework based on deep reinforcement learning and adaptive nonlinear model predictive control algorithm according to the adaptive nonlinear model predictive control (ANMPC) algorithm; Network Training Module: used to input the reward function into the network framework based on deep reinforcement learning and adaptive nonlinear model predictive control algorithm according to the set network parameters; Taking the dynamic model of the UAV as the control object, training is carried out in various static multi-obstacle and dynamic multi-obstacle simulation environments. After each round of training, a training weight file is output; by comparing with the training results generated in all previous training rounds before this training, the weight file corresponding to the training result with the highest reward value and the shortest flight path length is selected as the optimal weight file; Test Result Output Module: used to input the optimal weight file and the dynamic model of the UAV after applying the algorithm network framework based on deep reinforcement learning and adaptive nonlinear model predictive control into various static multi-obstacle and dynamic multi-obstacle simulation environments for testing, and obtain the UAV flight trajectory map, reward value r, flight path length L, path maximum distortion LS, path global distortion GS, whether a collision occurs, and threat index results.

9. An unmanned aerial vehicle trajectory planning and tracking device based on deep reinforcement learning and adaptive nonlinear model predictive control, characterized in that, Including: Memory: storing a computer program of a UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control according to any one of claims 1-7, being a computer-readable device; Processor: used to implement a UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control according to any one of claims 1-7 when executing the computer program.

10. A computer-readable storage medium, characterized in that A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a UAV trajectory planning and tracking method based on deep reinforcement learning and adaptive nonlinear model predictive control according to any one of claims 1-7.

Citation Information

Cited By

  • Method and system for predicting and analyzing energy-saving effect in climbing process of unmanned aerial vehicle

    CN120449719A

  • Intelligent unmanned aerial vehicle navigation method and system based on bimodal obstacle feature extraction

    CN120506958A

  • Unmanned aerial vehicle network system phase point trajectory simulation modeling method and system, medium and software product

    CN120802670A

  • An unmanned aerial vehicle network system phase point trajectory simulation modeling method, system, medium and software product

    CN120802670B

  • Dynamic defense model construction method for unmanned aerial vehicle countering

    CN121028568A