Unmanned aerial vehicle agile flight control system and method based on differentiable physical simulator
By building a lightweight strategic neural network based on differentiable physics simulator, and optimizing the UAV controller with backpropagation loss gradient, the problem of insufficient error accumulation and generalization capabilities of traditional UAV navigation in high-speed and dynamic environments is solved, and efficient autonomous flight control and multi-machine self-organized navigation are achieved.
Patent Information
- Application Number
- CN202311851947.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional UAV navigation methods have problems such as accumulation of errors, delays, large data demands, insufficient generalization capabilities and poor flexibility in high-speed flights and dynamic environments. The existing reinforcement learning and imitation learning methods are insufficient in the adaptability of new tasks.
By building a lightweight policy neural network based on differentiable physics simulator, the neural network controller is optimized by using the backpropagation loss gradient, and combining the convolutional recurrent neural network and the differentiable physics simulator, the UAV policy network is trained to achieve autonomous flight control.
It significantly improves the efficiency of drone training and optimization of navigation strategies, and can achieve high-speed obstacle avoidance and autonomous flight in dynamic environments in complex environments, without relying on sensors with high computing complexity, and is suitable for single-machine and multi-machine self-organized navigation tasks.
Smart Images

Figure CN120233794A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of unmanned aerial vehicle (UAV) flight control, and specifically to an agile flight control system and method for UAVs based on a differentiable physical simulator. Background Art
[0002] Traditional autonomous UAV methods decompose navigation into localization and mapping, planning, and control. For example, Fast-Planner. Affected by error accumulation and latency, traditional methods cannot achieve high-speed robust and flexible flight. Reinforcement learning methods rely on a large amount of random exploration and do not utilize the differential of the physical model to obtain the direction of policy optimization, so they have slow convergence and high data requirements. Imitation learning frameworks rely on the quality and comprehensiveness of expert demonstrations. They have insufficient generalization ability outside the training data, and the agent's imitation of the expert may be inaccurate. In addition, the design of the expert model is complex and it is difficult to flexibly solve new tasks. Summary of the Invention
[0003] Aiming at the deficiencies of the prior art in terms of latency, error accumulation, data efficiency, generalization, and flexibility, the present invention proposes an agile flight control system and method for UAVs based on a differentiable physical simulator. By utilizing the differentiability of the physical system and backpropagating the loss gradient, the neural network controller can be directly optimized, which can significantly improve the training efficiency and the optimization effect of the neural network. Moreover, it can be flexibly extended to solve high-speed obstacle avoidance at 20 m / s in a single-machine forest, dynamic environment obstacle avoidance, and cluster self-organizing navigation tasks without communication.
[0004] The present invention is realized through the following technical solutions:
[0005] The present invention relates to an agile flight control method for UAVs based on a differentiable physical simulator. By defining the robot navigation task as an optimization problem in the offline stage, modeling the robot as a discrete-time dynamic system with a continuous state and control input space, constructing a lightweight policy neural network and a differentiable physical simulator using a convolutional recurrent neural network (CRNN) structure, and backpropagating the flight trajectory loss to train the policy network; in the online stage, the trained lightweight policy neural network is deployed on the UAV to achieve autonomous flight control.
[0006] The present invention relates to an agile flight control system for an unmanned aerial vehicle (UAV) that implements the above method, including: a lightweight policy neural network unit, a motion simulation unit, a network training unit, and a policy generation unit, where: the lightweight policy neural network unit uses a max pooling layer to obtain a low-resolution nearest depth map, then extracts spatial features of the depth map sequence through a convolutional neural network, and uses a gated recurrent unit (GRU) to receive UAV attitude and target flight speed inputs to extract temporal features; a full connection layer is used to obtain a thrust vector sequence planned by the UAV policy network; the motion simulation unit first uses the UAV particle model and the thrust vector sequence of the lightweight policy neural network unit to perform UAV motion simulation to obtain the position and attitude of the UAV at the next time step; the motion simulation unit renders a depth map of the UAV's perspective based on pre-generated obstacles and the ground to obtain a simulated UAV depth map at the next time step; the network training unit sets the lightweight policy neural network unit based on vision to perform iterative simulations for multiple time steps in the motion simulation unit to obtain the flight trajectory of the agent in the simulator and its computational graph, and uses the backpropagation algorithm on the computational graph to calculate the loss function for the gradient of the network weights, and then obtains the optimized neural network weights through the neural network optimizer AdamW; the policy generation unit first infers the optimized UAV policy network based on the depth map collected by the depth camera and the attitude measured by the UAV inertial sensor, combined with the flight speed target given by the user, and calculates the current required thrust vector; the UAV flight controller adjusts the attitude and throttle of the UAV according to the thrust vector planned by the policy network, so as to control the real UAV to avoid obstacles and achieve obstacle avoidance flight. Technical effects
[0007] The present invention uses the first-order differential of the UAV physical model to obtain an accurate UAV policy network update gradient through backpropagation trajectory loss. Compared with the prior art, the present invention updates the UAV policy network more effectively and improves the optimization effect of the navigation strategy. The described training method based on a differentiable simulator (network training unit) can optimize the lightweight neural network to achieve end-to-end navigation. The present invention not only significantly reduces the amount of calculation, calculation delay, and error accumulation, but also can be flexibly extended to solve high-speed obstacle avoidance at 20 m / s in a single-machine forest, dynamic environment obstacle avoidance, and cluster self-organizing navigation tasks without communication. By efficiently implementing agile flight in complex environments, the present invention has great potential in future autonomous UAV deployment. Description of the drawings
[0008] Figure 1 is a flowchart of the present invention;
[0009] Figure 2 is a schematic diagram of the lightweight policy neural network unit;
[0010] Figure 3 Brief description of the technical solution for optimizing the neural network for the differentiable drone simulator and theoretical schematic diagram;
[0011] Figure 4 Schematic diagram of the implementation scenario of the embodiment of the present invention;
[0012] Figure 5 Effect examples and data charts of the present invention;
[0013] In the figure: The present invention can reach a high speed of 10 m / s in woods, urban environments, and when avoiding dynamic obstacles, and has a high success rate. The experiments shown do not rely on visual inertial odometry or GPS positioning;
[0014] Figure 6 Schematic diagram of the quantitative comparison between the flexible flight method of the present invention and the existing complex environment in the simulation environment;
[0015] In the figure: A - C are three different test scenarios; on the left is the success rate of various methods at different flight speeds. On the right is the schematic diagram of the test environment and the flight trajectory under the present invention;
[0016] Figure 7 Schematic diagram of the present invention enabling multiple drones to have the ability to self - coordinate and swap positions through narrow spaces without communication;
[0017] Figure 8 Schematic diagram of the quantitative comparison between the flexible flight method of the present invention and the existing complex environment in the simulation environment;
[0018] Figure 9 Schematic diagram of the training scenario of the present invention. Detailed implementation method
[0019] As Figure 1 shown, this embodiment relates to a method for optimizing the flight control of a drone based on the above - mentioned system. By defining the robot navigation task as an optimization problem in the offline stage, the robot is modeled as a discrete - time dynamic system with a continuous state and control input space. At each time step k, the system state is x k , and the corresponding control input is u k . Based on the current state, an observation value o is generated through the sensor model h k . The dynamics of the system are described by a function that defines the discrete - time evolution of the system over time. At each time step k, the robot receives a cost signal which is a function of the current state and control input, specifically including:
[0020] The first step is to construct as Figure 2The lightweight policy neural network adopting a convolutional recurrent neural network (CRNN) structure as shown is used for depth map feature extraction, where: the depth map preprocessing unit performs maximum pooling processing according to the depth map information obtained from the depth camera to obtain a low-resolution depth map of 16×12; the spatial feature extraction unit performs convolutional neural network (CNN) feature extraction processing according to the low-resolution depth map feature information to obtain a depth map feature sequence; the temporal feature extraction unit performs gated recurrent unit (GRU) processing according to the depth map feature sequence information to obtain temporal features; the planning and control unit performs fully connected layer processing according to the temporal feature information to obtain the policy planned by the network, that is, the thrust vector expected for the drone to reach.
[0021] As Figure 2 shown, the number of convolutional layer channels (filter sizes) of each convolutional layer of the convolutional neural network (CNN) are respectively: 32(2)-64(3)-128(3), and the features output by it are flattened and projected onto a 192-dimensional feature through a fully connected layer, where: the stride of all convolutional layers is 1 and is followed by a LeakyReLU activation function.
[0022] The parameters are shared among the convolutional neural networks at all time steps in the spatial feature extraction unit.
[0023] The fully connected layer projects the target speed, attitude estimation, and speed estimation (optional) onto a 192-dimensional feature, fuses them with the image features and inputs them into the GRU unit, and uses the GRU state to predict the required thrust acceleration and the current speed estimation through another fully connected layer.
[0024] Step 2: Construct a differentiable drone physical simulator including a differentiable physical simulation sub-module and a depth map rendering sub-module, randomly generate geometric obstacles and virtual depth cameras in the simulation environment, and construct a training environment as Figure 9 shown.
[0025] The so-called differentiable physical simulation means: establishing a differentiable physical simulation for a quadrotor drone using a particle dynamics model, and the specific particle dynamics model is: Where: represents the three-dimensional spatial position of the drone, represents the three-dimensional velocity of the drone, represents the three-dimensional acceleration of the drone. Δt represents the step time. The subscript k represents the k-th time step.
[0026] In this embodiment, the integration time step Δt is 1 / 15 second.
[0027] The so-called depth map rendering means: placing geometric obstacles and a virtual depth camera at random positions, calculating a loss function using the states of the obstacles and the UAV, and implementing fast depth image rendering using the ray casting method.
[0028] The virtual depth camera is located at the centroid of the UAV and has an elevation angle of 10 degrees upward relative to the UAV body. Its field of view angle and aspect ratio are the same as those of the camera used for depth map information acquisition in the first step.
[0029] The so-called ray casting means: taking the center of the virtual UAV as the optical center, determining the plane size and pixel coordinates of the virtual depth camera according to the field of view angle, making rays from the optical center to the pixels in the plane of the virtual depth camera, finding the intersection points of the rays and all obstacles, and taking the projection distance value of the point with the closest distance as the value of the pixel to complete the depth map rendering.
[0030] The geometric obstacles include cuboids, spheres, cylinders, rings, and planes.
[0031] In the third step, as Figure 1 and Figure 2 shown, optimize the lightweight policy neural network obtained in the first step in the differentiable physical simulator obtained in the second step. Specifically: use the vision-based lightweight policy neural network to perform iterative simulations for N = 150 time steps in the differentiable UAV physical simulator to obtain the flight trajectory of the agent in the simulator and its computational graph as shown in Figure 3 Calculate the loss function on the obtained flight trajectory Use the backpropagation algorithm on the computational graph to calculate the policy gradient of the network weights θ, and then obtain the neural network weights after policy training through the neural network optimizer AdamW.
[0032] The so-called flight trajectory means: in the differentiable physical simulator, taking the output of the lightweight policy neural network as the input of the UAV dynamics model, introducing a first-order step response to simulate the control response of the real UAV to the output of the neural network, and recursively calculating the state of the UAV in the next time step according to discrete time, including: the three-dimensional position, three-dimensional velocity, three-dimensional acceleration, and three-dimensional jerk of the UAV.
[0033] The flight trajectory consists of the states of the UAV at N = 150 time steps.
[0034] The loss function where: λ represents the weight of the loss function, the collision loss function the velocity tracking loss function the acceleration smoothing loss function the jerk smoothing loss function where: dk denotes the minimum distance of the UAV from the obstacle at the k-th time step, r q denotes the radius of the UAV, denotes the speed of the UAV flying towards the obstacle at the k-th time step, denotes the desired speed of the UAV at the k-th time step, denotes the average speed of the UAV at nearby times. SmoothL1 refers to the default SmoothL1 loss function in pytorch( https: / / pytorch.org / docs / stable / generated / torch.nn.SmoothL1Loss.html );a k denotes the acceleration of the UAV at the k-th moment. Δt is the integration time step in the particle dynamics model.
[0035] The so-called policy gradient refers to: where: l i denotes the loss function component at the i-th time step, x i denotes the state of the UAV at the i-th moment, including the three-dimensional position, speed, acceleration and jerk of the UAV, where: e -αΔt is the time gradient decay introduced in the training process, and α is the gradient decay coefficient. The time gradient decay, as Figure 3 shown in the upper right, exponentially decays the long-term influence, alleviates gradient explosion and stabilizes the optimization process.
[0036] In this embodiment, the coefficient α of the decay term is taken as 0.4. This loss function uses a differentiable physical simulator to directly calculate the policy gradient through backpropagation, and efficiently trains the neural network to complete the vision-based navigation task. As Figure 1 shown, it is the gradient decay term that alleviates the gradient explosion problem in backpropagation.
[0037] As Figure 3 shown, it is the computational graph composed of the iterative operation of the policy network and the differentiable simulator. As Figure 3 shown, it is the method of obtaining the gradient for training the policy network by backpropagation in time, and finally training the policy network using the gradient descent method.
[0038] The optimization process is carried out after forward simulation of N = 150 time steps in 64 parallel differentiable physical environments.
[0039] In the training environment, the training parameters are the AdamW optimizer, the learning rate is 0.001, the learning rate uses Cosine decay, and the number of training iterations M = 50000. It can be obtained that in the training environment of a single RTX4090 and CUDA 11.8, the training duration is 70 minutes, the loss converges to 4.46, and the optimized policy network weights are obtained.
[0040] Step 4: Deploy the weights of the lightweight policy neural network optimized in Steps 2 and 3 on a real-world drone: Use the output of the lightweight policy neural network as the thrust vector, and track the required thrust through the inner-loop angular velocity controller during deployment.
[0041] The inner-loop angular velocity controller mentioned above refers to: calculating the attitude error based on the attitude represented by the quaternion of the drone to obtain the angular velocity control target ω fb = 2·K att ·q e , where: q is the current attitude of the drone, q des is the direction of the thrust vector output by the network, q e is the attitude error, ω fb is the angular velocity control target, and K att is the proportional control coefficient.
[0042] In this embodiment,
[0043] After specific actual experiments, deploy the network weights and running code on a quadrotor drone. And conduct agile obstacle avoidance flight tests on the quadrotor drone in the real world shown in Figure 4 . Figure 4 The flight trajectories shown in Figure 5 show that the present invention can have excellent flight speed and robustness in various complex environments with dynamic obstacles. The quadrotor drone mentioned above is a small, agile, autonomous and low-cost drone. In the actual experiment mentioned above, the drone uses a Roma 3-inch frame, GEMFAN 3-inch propellers and 16063750KV motors. Use a RealSense D435i camera as the depth perception module, and the frequency of the depth image is 15Hz. In addition, use a Mango Pi as the on-board computer, equipped with a quad-core Cortex-A53 processor and 1G of memory. This patent repeatedly tests the obstacle avoidance results of the drone at different desired speeds under static and dynamic obstacles, and can obtain
[0044] Compared with the prior art, the present method uses the first derivative of the drone physical model to obtain an accurate update gradient of the drone policy network through backpropagation of the trajectory loss. The training method based on the differentiable simulator can more effectively optimize the drone policy network and improve the effect of the navigation strategy. This makes the training process do not require reinforcement learning techniques such as random exploration, policy gradient, and Q-network, nor artificial or expert model demonstrations. Through differentiable physical training, vision-based agile navigation is achieved.
[0045] As shown inFigure 6 As shown, compared with the existing methods Integrated perception and control athighspeed:Evaluating collision avoidance maneuvers withoutmaps(Reactive), Fast-planner, Learning High-Speed Flight inthe Wild(Agile), Swarm ofmicro flyingrobots inthe wild(Ego v2), the present invention conducts comprehensive experiments in a simulation environment, benchmarking its method against existing methods, including the traditional map-based method Ego v2 and the learning-based method Agile.
[0046] To ensure a fair comparison, the present invention uses the Flightmare simulator and conducts experiments according to the experimental settings outlined in Agile. In addition, the present invention also uses the AirSim simulator for additional evaluation. All methods are tested at different target speeds (1, 4, 7, 10, and 13 m / s). Each experiment is repeated 10 times, and the average flight speed, maximum flight speed, and success rate are calculated. If the drone reaches the target position within a 5-meter radius without a collision, the task is considered successful. The average speed is calculated based on these successful trials.
[0047] The results are summarized as Figure 6 shown. The traditional map-based method Ego v2 performs relatively well at very low speeds, such as below 3 m / s; however, as the target speed increases, its performance significantly degrades. This is consistent with the real experiments shown in the original paper, where the average flight speed in the position scenario is approximately 2 m / s.
[0048] Although the learning-based method (called Agile) performs well in a forest environment, its training is also in a forest-specific environment. When tested in an unknown environment, such as in an urban environment ( Figure 6 B-C), the success rate of Agile is very low. The expert-based method has difficulty generalizing to unknown environments, resulting in a significant degradation in performance.
[0049] In addition to the flight effect of the drone, the present invention verifies that it exceeds the Agile method in terms of the convergence speed during training. In the experiment, we extract the point cloud from the training scenario used by Agile and evaluate the success rate and average speed under various optimization steps. Figure 6Demonstrating the faster convergence speed of the present invention: At the 1000th iteration, the present invention has achieved a 100% success rate at an average speed of 5 m / s, while Agile has only just started to show success with a success rate of only 20%. These findings highlight the extremely high efficiency of the present invention.
[0050] Because the present invention adopts the first-order differential of the physical model and obtains the accurate update gradient of the UAV policy network through the backpropagation trajectory loss, the network can converge more quickly and has stronger generalization ability. Eventually, it achieves a faster speed and a higher success rate.
[0051] In addition, compared with the above two methods, the policy network of the present invention does not need to rely on the current position information of the UAV as input. In the experiment shown in Figure 4 , Figure 5, in the case where the position and speed of the UAV cannot be obtained, the present invention outputs control commands based on the UAV attitude and visual input, maintaining effectiveness in the extreme case without GPS and odometer. In the real aircraft deployment, the present invention adopts a lightweight network and does not require an inertial vision odometer module with high computational complexity, so it has significant advantages in the deployment of low-computing-power UAV platforms. In the experiment shown in Figure 4 , Figure 5, the present invention runs on the Mango Pi micro-computing board sold for only 140 yuan, demonstrating the ability to navigate a small UAV to fly agilely under extremely limited computing resources.
[0052] In summary, the application effect of the present invention in a multi-UAV system. After training in a multi-UAV simulation environment (as shown in Figure 7 ), it can enable multiple UAVs to self-coordinately exchange positions through a narrow door. In addition, the UAVs also exhibit decentralized self-organizing behaviors such as waiting, following, conflict resolution, and mutual obstacle avoidance. The results of the simulation environment experiment (as shown in Figure 8 ) show that, different from the prior art, the present invention can achieve self-coordinated obstacle avoidance without mutual communication between multiple UAVs, and has a high speed and success rate.
[0053] The above specific embodiments can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific embodiments, and all implementation solutions within its scope are subject to the present invention.
Claims
1. A method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator, characterized in that, By defining the robot navigation task as an optimization problem in the offline stage, modeling the robot as a discrete-time dynamic system with a continuous state and control input space, constructing a lightweight policy neural network and a differentiable physical simulator using a convolutional recurrent neural network (CRNN) structure, backpropagating the flight trajectory loss to train the lightweight policy neural network; in the online stage, deploying the trained lightweight policy neural network on the unmanned aerial vehicle to achieve autonomous flight control.
2. The method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator according to claim 1, wherein The lightweight policy neural network includes: a depth map preprocessing unit, a spatial feature extraction unit, a temporal feature extraction unit, and a planning and control unit, where: the depth map preprocessing unit performs max pooling processing on the depth map information obtained from the depth camera to obtain a low-resolution depth map of 16×12; the spatial feature extraction unit performs convolutional neural network (CNN) feature extraction processing on the low-resolution depth map feature information to obtain a depth map feature sequence; the temporal feature extraction unit performs gated recurrent unit (GRU) processing on the depth map feature sequence information to obtain temporal features; the planning and control unit performs fully connected layer processing on the temporal feature information to obtain the policy planned by the network, that is, the thrust vector expected for the unmanned aerial vehicle to reach.
3. The method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator according to claim 2, wherein The number of convolutional layer channels (filter sizes) of each layer of the convolutional neural network (CNN) is: 32(2)-64(3)-128(3), and the output features are flattened and projected onto a 192-dimensional feature through a fully connected layer, where: the stride of all convolutional layers is 1 and is followed by a LeakyReLU activation function; The parameters are shared among the convolutional neural networks at all time steps in the spatial feature extraction unit; The fully connected layer projects the target speed, attitude estimation, and speed estimation onto a 192-dimensional feature, fuses them with the image features and inputs them into the GRU unit, and uses the GRU state to predict the required thrust acceleration and the current speed estimation through another fully connected layer.
4. The method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator according to claim 1, characterized in that, The differentiable physical simulator includes: a differentiable physical simulation sub-module and a depth map rendering sub-module; The differentiable physical simulation mentioned above refers to: establishing a differentiable physical simulation for a quadrotor UAV using a particle dynamics model, and the specific particle dynamics model is as follows: Where: represents the three-dimensional spatial position of the UAV, represents the three-dimensional velocity of the UAV, represents the three-dimensional acceleration of the UAV, Δt represents the step time, and the subscript k represents the k-th time step; The depth map rendering refers to: placing geometric obstacles and a virtual depth camera at random positions, calculating the loss function using the states of the obstacles and the unmanned aerial vehicle, and implementing fast depth image rendering using the ray casting method.
5. The method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator according to claim 4, characterized in that, The virtual depth camera is located at the centroid of the unmanned aerial vehicle and has an elevation angle of 10 degrees upward relative to the body of the unmanned aerial vehicle, and its field of view angle and screen ratio are the same as those of the camera used for depth map information acquisition in the first step.
6. The method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator according to claim 4 or 5, characterized in that, The ray casting refers to: taking the center of the virtual unmanned aerial vehicle as the optical center, determining the plane size and pixel coordinates of the virtual depth camera according to the field of view angle, making rays from the optical center to the pixels in the plane of the virtual depth camera, finding the intersection points of the rays and all obstacles, and taking the projection distance value of the point with the closest distance as the value of the pixel to complete the depth map rendering.
7. The method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator according to claim 2, wherein The training mentioned above refers to: optimizing the lightweight policy neural network obtained in the first step in a differentiable physics simulator. Specifically, use the vision-based lightweight policy neural network to perform several iterative simulations in the differentiable drone physics simulator to obtain the flight trajectory of the agent in the simulator and its computational graph, and calculate the loss function on the obtained flight trajectory. Use the backpropagation algorithm on the computational graph to calculate the policy gradient of the network weights θ, and then obtain the neural network weights after policy training through the neural network optimizer AdamW. The flight trajectory mentioned above refers to: in a differentiable physics simulator, taking the output of a lightweight policy neural network as the input of the UAV dynamics model, introducing a first-order step response to simulate the control response of a real UAV to the output of the neural network, and recursively calculating the state of the UAV at the next time step according to discrete time, including: the three-dimensional position, three-dimensional velocity, three-dimensional acceleration, and three-dimensional jerk of the UAV.
8. The method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator according to claim 7, characterized in that, The described loss function where: λ represents the weight of the loss function, and the collision loss function The velocity tracking loss function The acceleration smoothing loss function The jerk smoothing loss function where: d k represents the minimum distance of the UAV from the obstacle at the k-th time step, r q represents the radius of the UAV, represents the velocity of the UAV flying towards the obstacle at the k-th time step, represents the desired velocity of the UAV at the k-th time step, represents the average velocity at the nearby time of the UAV, and SmoothL1 refers to the default SmoothL1 loss function in pytorch; a k represents the acceleration of the UAV at the k-th moment, and Δt is the integration time step in the particle dynamics model; The so-called policy gradient refers to: Where: l i represents the loss function component at the i-th time step, and x i represents the state of the UAV at the i-th moment, including the three-dimensional position, velocity, acceleration, and jerk of the UAV, where: e -αΔt is the time gradient attenuation introduced during the training process, and α is the gradient attenuation coefficient; The time gradient decay exponentially decays long-term effects to mitigate gradient explosion and stabilize the optimization process.
9. The method for agile flight control of an unmanned aerial vehicle based on a differentiable physical simulator according to claim 1, wherein The deployment mentioned above refers to: deploying the weights of the optimized lightweight policy neural network on a real-world UAV: taking the output of the lightweight policy neural network as the thrust vector, and tracking the required thrust through an inner-loop angular velocity controller during deployment; The inner-loop angular velocity controller refers to: calculating the attitude error based on the attitude represented by the quaternion of the UAV to obtain the angular velocity control target where: q is the current attitude of the UAV, q des is the direction of the thrust vector output by the network, q e is the attitude error, ω fb is the angular velocity control target, K att is the proportional control coefficient.
10. An agile flight control system for an unmanned aerial vehicle implementing the method according to any one of claims 1-9, characterized in that, including: A lightweight policy neural network unit, a motion simulation unit, a network training unit, and a policy generation unit, where: The lightweight policy neural network unit uses a max pooling layer to obtain a low-resolution nearest depth map, then extracts spatial features of the depth map sequence through a convolutional neural network, and uses a gated recurrent unit (GRU) to receive the input of the UAV attitude and the target flight speed to extract temporal features; uses a fully connected layer to obtain the thrust vector sequence planned by the UAV policy network; The motion simulation unit first uses the UAV particle model and the thrust vector sequence of the lightweight policy neural network unit to perform UAV motion simulation to obtain the position and attitude of the UAV at the next time step; The motion simulation unit renders the depth map from the perspective of the UAV based on the pre-generated obstacles and the ground to obtain the simulated UAV depth map at the next time step; The network training unit sets the lightweight policy neural network unit based on vision to perform iterative simulations for multiple time steps in the motion simulation unit to obtain the flight trajectory of the agent in the simulator and its computational graph, and uses the backpropagation algorithm on the computational graph to calculate the loss function For the gradient of the network weights, the optimized neural network weights are obtained through the neural network optimizer AdamW; The policy generation unit first infers the optimized UAV policy network according to the depth map collected by the depth camera and the attitude measured by the UAV inertial sensor, combined with the flight speed target given by the user, and calculates the thrust vector required at present; The UAV flight controller adjusts the attitude and throttle size of the UAV according to the thrust vector planned by the policy network, so as to control the real UAV to avoid obstacles and realize obstacle avoidance flight.