Method for path planning of eVTOL unmanned aerial vehicle in dynamic environment
Through the reinforcement learning framework TDDPG and LSTM model combined with the composite reward function, the problem of path planning of eVTOL drones in dynamic environments is solved, and efficient and safe path planning and obstacle avoidance capabilities are achieved to ensure the accuracy of the flight trajectory.
Patent Information
- Application Number
- CN202510489469.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to efficiently and safely carry out path planning of eVTOL drones in dynamic environments, especially in complex dynamic obstacle environments, which cannot effectively avoid collisions and ensure the accuracy of flight trajectory.
The reinforcement learning framework TDDPG combined with the LSTM model is adopted to design a composite reward function, and an initial exploration mechanism and a priority experience playback mechanism are introduced. Through continuous action space and kinematics and aerodynamics-based models, the movement trajectory of obstacles is predicted to achieve dynamic obstacle avoidance and endpoint planning.
It improves the efficiency and learning stability of the drone in dynamic environments, ensures the accuracy and safety of the flight trajectory, and can avoid collisions in advance.
Smart Images

Figure CN120371001A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of eVTOL drones, and particularly relates to a method for path planning of eVTOL drones in a dynamic environment. Background Art
[0002] In recent years, the development of eVTOL aircraft (pure electric vertical takeoff and landing aircraft) has attracted extensive attention from aerospace enterprises, the automotive industry, the transportation industry, and academia. Potential future applications of eVTOL involve various scenario modes such as urban passenger transportation, freight transportation, personal aircraft, and emergency medical services. Currently in China, eVTOL aircraft are also called flying cars, mainly used for urban air traffic, and are a current popular investment field. In the task of complex dynamic environment scenarios, drone path planning is one of the most important links in completing the entire task, which determines the spatio-temporal information of the drone. The purpose of path planning is to find a series of points from the starting point to the ending point to ensure that the drone does not collide with obstacles. Therefore, efficient, reliable, and safe path planning algorithms are worthy of research.
[0003] For this reason, the present invention proposes a method for path planning of eVTOL drones in a dynamic environment. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for path planning of eVTOL drones in a dynamic environment to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for path planning of eVTOL drones in a dynamic environment, the specific steps include: S1: Design a reinforcement learning framework TDDPG (Temporal Difference Deep Deterministic Policy Gradient), and use the LSTM (Long Short-Term Memory) model as the policy model;
[0006] S2: Combine the continuous action space to solve the exploration and path planning problems in the dynamic environment;
[0007] S3: To accurately simulate the motion trajectory of the drone in a complex three-dimensional environment, construct a model based on the principles of kinematics and aerodynamics;
[0008] S4: Design a composite reward function that comprehensively considers various obstacles and environmental factors, enabling the eVTOL to successfully plan a flight path in a dynamic and unknown environment;
[0009] S5: Introduce an initialization exploration mechanism and a prioritized experience replay mechanism to accelerate the convergence speed of the TDDPG algorithm and improve the learning efficiency of the algorithm in a complex environment.
[0010] Preferably, in S1, a deep reinforcement learning algorithm TDDPG that combines the Deep Deterministic Policy Gradient algorithm DDPG with the Long Short-Term Memory network LSTM enables the UAV to perform dynamic obstacle avoidance and dynamic end-point planning based on historical state and action information by leveraging the continuous action optimization ability of DDPG and the time series modeling ability of LSTM, thereby solving the problem of three-dimensional path planning for UAVs in a dynamic obstacle environment.
[0011] Preferably, in S2, a continuous action space is adopted, allowing the UAV to select any direction within the entire 360-degree range and adjust its speed in real time according to the environment. The continuous action space enables the UAV to make more precise and flexible adjustments and respond promptly to environmental changes during flight.
[0012] Preferably, in S3, the model takes into account the three-dimensional position (X, Y, Z), speed (v0, v1, v2) of the UAV, as well as the dynamic changes in its attitude angles; the motion state of the UAV is gradually updated through a recursive formula, combining control inputs and dynamic environmental factors to dynamically adjust the position, speed, and attitude angle variables; through this model, the state variables not only include the distance between the UAV and the target point and the relative position to the obstacle, but also incorporate the position changes of dynamic obstacles, ensuring that the simulation results truly reflect the flight trajectory of the UAV in a multi-obstacle environment and can effectively simulate the state changes of the UAV at different flight stages to ensure the accuracy of its motion trajectory.
[0013] Preferably, the position (x, y, z) of the UAV, as the first part of the state variable, represents the position of the UAV in three-dimensional space, and follows the following dynamic formula:
[0014] x t+1 =x t +v t cos(θ t )cos(φ t )Δt
[0015] y t+1 =y t +v t cos(θ t )sin(φ t )Δt
[0016] z t+1 =z t -v t sin(θ t )Δt
[0017] where x t , y t , z t are the current three-dimensional coordinates of the UAV, and v tis the current speed of the drone, and θ t and φ t are the pitch angle and yaw angle respectively;
[0018] The speed of the eVTOL drone is decomposed into components (v0, v1, v2) in three directions, where v0, v1, and v2 represent the speed components along the X, Y, and Z axes respectively. The formulas are as follows:
[0019] v0 = v·cos(θ)·cos(φ), v1 = v·cos(θ)·sin(φ), v2 = -v·sin(θ)
[0020] where v represents the speed of the eVTOL, θ is the pitch angle, and φ is the yaw angle;
[0021] The distance d between the eVTOL and the target is used to measure the proximity of the drone to the target; this state variable is calculated using the Euclidean distance formula, as follows:
[0022]
[0023] where X opp , Y opp , Z opp represent the coordinates of the target position;
[0024] The heading error reflects the angular deviation between the current heading of the eVOTL and the target direction; this error is obtained by calculating the angle between the direction vector between the eVTOL and the target and the eVTOL speed vector:
[0025]
[0026] where D0, D1, and D2 are the direction vectors between the eVTOL and the target, v is the speed of the eVTOL, and d is the distance between the eVTOL and the target;
[0027] The state variables include not only the distance to static obstacles but also the states of dynamic obstacles; the distance between a static obstacle and the eVTOL is calculated using the following formula:
[0028]
[0029] where x obs , y obs , z obs are the coordinates of the obstacle, and r obs is the radius of the obstacle; the position of each dynamic obstacle is obtained through dynamic x, dynamic_y and dynamic_z indicate that these positions will be updated as time progresses; the dynamic state information can help the agent predict the future movement trajectory of the obstacle, so as to plan the path in advance and avoid collisions; at the same time, in order to capture the movement trend of the dynamic obstacle, the historical dynamic obstacle movement information is saved as history_dynamic, which is used as a component of the state to improve the agent's prediction ability of the movement trend of the dynamic obstacle.
[0030] Preferably, in S4, the composite reward function is as follows:
[0031] R = R collision + R align + R terminal
[0032] In the formula, R collision represents the obstacle collision penalty, R align represents the orientation penalty, R terminal represents the terminal state reward;
[0033] R collision The function is as follows:
[0034]
[0035] In the formula, s min is the minimum safety distance. For each obstacle, the distance between the UAV and the obstacle is calculated and then the radius of the obstacle is subtracted, and the result is used as the safety distance index; when s min is less than 1, it is considered that the distance between the UAV and the obstacle is too close, and then a penalty is imposed; when the distance is exactly 1.0, the penalty is 0; if the distance further decreases, the penalty will increase exponentially, so that the UAV and the obstacle can maintain a sufficient safety distance;
[0036] R align The function is as follows:
[0037]
[0038] In the formula, θ e is the pitch angle error, which reflects the deviation between the UAV and the target height direction in the vertical plane; ψ e is the yaw angle error, which reflects the deviation between the actual flight direction of the UAV on the horizontal plane and the ideal target direction;
[0039] R terminal The function is as follows:
[0040]
[0041] The termination conditions of the task are determined according to the state of the drone, specifically divided into reaching the target, colliding, or flying out of the boundary; when the eVTOL successfully explores the target position, the reward will increase; if the eVTOL collides with an obstacle, a penalty will be imposed; if the eVTOL flies out of the flight boundary, a corresponding penalty will also be triggered.
[0042] Preferably, in S5, in the initial stage of reinforcement learning, the agent accumulates experience through random exploration, and uses the prioritized experience replay mechanism to accelerate the learning process, gradually shifting from random exploration to policy selection based on the Actor network. At the same time, Ornstein-Uhlenbeck noise is added to maintain exploration, and a linear scheduling strategy is adopted to optimize the sampling weights of important experiences, thereby improving the path planning efficiency and learning stability of the agent in complex dynamic environments.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] The method designed by the present invention enables the drone to perform dynamic obstacle avoidance and dynamic end-point planning based on historical state and action information by utilizing the continuous action optimization ability of DDPG and the time series modeling ability of LSTM, thereby solving the problem of three-dimensional path planning of drones in a dynamic obstacle environment;
[0045] The designed model based on kinematic and aerodynamic principles ensures that the simulation results truly reflect the flight trajectory of the drone in a multi-obstacle environment, and can also effectively simulate the state changes of the drone at different flight stages, ensuring the accuracy of its motion trajectory;
[0046] The present invention can predict the future motion trajectory of obstacles, thereby planning the path in advance to avoid collisions;
[0047] The present invention introduces an initialization exploration mechanism and a prioritized experience replay mechanism to improve the path planning efficiency and learning stability of the agent in complex dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] Embodiment 1
[0051] Please refer to Figure 1 The present invention provides a method technical solution: a method for path planning of an eVTOL drone in a dynamic environment, and the specific steps include: S1: Design a reinforcement learning framework TDDPG (Temporal-Dependent Deep Deterministic Policy Gradient), and adopt an LSTM (Long Short-Term Memory) model as the policy model;
[0052] S2: Combine the continuous action space to solve the exploration and path planning problems in the dynamic environment;
[0053] S3: To accurately simulate the motion trajectory of the drone in a complex three-dimensional environment, construct a model based on kinematic and aerodynamic principles;
[0054] S4: Design a composite reward function that comprehensively considers various obstacles and environmental factors, enabling the eVTOL to successfully plan a flight path in a dynamic and unknown environment;
[0055] S5: Introduce an initialization exploration mechanism and a prioritized experience replay mechanism to accelerate the convergence speed of the TDDPG algorithm and improve the learning efficiency of the algorithm in a complex environment.
[0056] In this embodiment, preferably, in S1, the deep reinforcement learning algorithm TDDPG that combines the deep deterministic policy gradient algorithm DDPG and the long short-term memory network LSTM enables the drone to perform dynamic obstacle avoidance and dynamic end-point planning based on historical state and action information by utilizing the continuous action optimization ability of DDPG and the time series modeling ability of LSTM, thereby solving the problem of three-dimensional path planning of the drone in a dynamic obstacle environment.
[0057] In this embodiment, preferably, in S2, the continuous action space is adopted, allowing the drone to select any direction within the entire 360-degree range and adjust the speed in real time according to the environment. Through the continuous action space, the drone can make more accurate and flexible adjustments and can respond to environmental changes in a timely manner during flight.
[0058] In this embodiment, preferably, in S3, the model considers the three-dimensional position (X, Y, Z), speed (v0, v1, v2) of the drone, and the dynamic changes of its attitude angles; the motion state of the drone is gradually updated through a recursive formula, and combines the control input and dynamic environmental factors to dynamically adjust the position, speed, and attitude angle variables; through this model, the state variables not only include the distance between the drone and the target point and the relative position with the obstacle, but also add the position changes of the dynamic obstacles, ensuring that the simulation results truly reflect the flight trajectory of the drone in a multi-obstacle environment, and at the same time can effectively simulate the state changes of the drone at different flight stages to ensure the accuracy of its motion trajectory.
[0059] In this embodiment, preferably, the position (x, y, z) of the drone is used as the first part of the state variable, representing the position of the drone in three-dimensional space. The following dynamic formula is followed:
[0060] x t+1 = x t + v t cos(θ t )cos(φ t )Δt
[0061] y t+1 = y t + v t cos(θ t )sin(φ t )Δt
[0062] z t+1 = z t - v t sin(θ t )Δt
[0063] Where x t , y t , z t are the current three-dimensional coordinates of the drone, v t is the current speed of the drone, θ t and φ t are the pitch angle and yaw angle respectively;
[0064] The speed of the eVTOL drone is decomposed into three components (v0, v1, v2) in three directions. Among them, v0, v1, and v2 respectively represent the speed components along the X, Y, and Z axes. The formula is as follows:
[0065] v0 = v·cos(θ)·cos(φ), v1 = v·cos(θ)·sin(φ), v2 = -v·sin(θ)
[0066] Where v represents the speed of the eVTOL, θ is the pitch angle, and φ is the yaw angle;
[0067] The distance d between the eVTOL and the target is used to measure the proximity of the drone to the target; this state variable is calculated by the Euclidean distance formula. The formula is as follows:
[0068]
[0069] Where X opp , Y opp , Z opp represent the coordinates of the target position;
[0070] The heading error reflects the angular deviation between the current heading of the eVOTL and the target direction; this error is obtained by calculating the angle between the direction vector between the eVOTL and the target and the eVOTL velocity vector:
[0071]
[0072] where D0, D1, D2 are the direction vectors between the eVOTL and the target, v is the velocity of the eVOTL, and d is the distance between the eVOTL and the target;
[0073] The state variables include not only the distance from static obstacles but also the states of dynamic obstacles; the distance between a static obstacle and the eVOTL is calculated by the following formula:
[0074]
[0075] where, x obs , y obs , z obs are the coordinates of the obstacle, and r obs is the radius of the obstacle; the position of each dynamic obstacle is represented by dynamic x , dynamic_y, dynamic_z, and these positions are updated as time progresses; the dynamic state information can help the agent predict the future movement trajectory of the obstacle, so as to plan the path in advance and avoid collisions; at the same time, in order to capture the movement trend of the dynamic obstacle, the historical dynamic obstacle movement information is saved as history_dynamic and used as a component of the state to improve the agent's prediction ability of the movement trend of the dynamic obstacle.
[0076] In this embodiment, preferably, in S4, the composite reward function is as follows:
[0077] R = R collision + R align + R terminal
[0078] In the formula, R collision represents the obstacle collision penalty, R align represents the heading penalty, and R terminal represents the terminal state reward;
[0079] The R collision function is as follows:
[0080]
[0081] In the formula, s minis the minimum safety distance. For each obstacle, calculate the distance between the drone and the obstacle and subtract the radius of the obstacle. The resulting value is used as the safety distance metric. When s min is less than 1, it is considered that the distance between the drone and the obstacle is too close, and a penalty is imposed. When the distance is exactly 1.0, the penalty is 0. If the distance further decreases, the penalty will increase exponentially, so as to keep a sufficient safety distance between the drone and the obstacle;
[0082] R align The function is as follows:
[0083]
[0084] In the formula, θ e is the pitch angle error, which reflects the deviation between the drone and the target height direction in the vertical plane; ψ e is the yaw angle error, which reflects the deviation between the actual flight direction of the drone and the ideal target direction on the horizontal plane;
[0085] R terminal The function is as follows:
[0086]
[0087] The termination condition of the task is determined according to the state of the drone, specifically divided into reaching the target, colliding, or flying out of the boundary; when the eVTOL successfully explores the target position, the reward will increase; if the eVTOL collides with an obstacle, a penalty will be imposed; if the eVTOL flies out of the flight boundary, a corresponding penalty will also be triggered.
[0088] In this embodiment, preferably, in S5, in the initial stage of reinforcement learning, the agent accumulates experience through random exploration, and uses the prioritized experience replay mechanism to accelerate the learning process, gradually shifting from random exploration to policy selection based on the Actor network. At the same time, Ornstein-Uhlenbeck noise is added to maintain exploration, and a linear scheduling strategy is adopted to optimize the sampling weights of important experiences, so as to improve the path planning efficiency and learning stability of the agent in complex dynamic environments.
[0089] Although the embodiments of the present invention have been shown and described, see the above detailed description. For those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for path planning of an eVTOL drone in a dynamic environment, characterized in that: The specific steps include: S1: Design the deep reinforcement learning framework TDDPG and adopt the LSTM model as the policy model; S2: Combine the continuous action space to solve the exploration and path planning problems in the dynamic environment; S3: To accurately simulate the flight trajectory of the UAV in the complex three-dimensional environment, construct a model based on the principles of kinematics and aerodynamics; S4: Design a composite reward function by comprehensively considering various obstacles and environmental factors, enabling the eVTOL to successfully plan the flight path in the dynamic and unknown environment; S5: Introduce the initialization exploration mechanism and the prioritized experience replay mechanism to accelerate the convergence speed of the TDDPG algorithm and improve the learning efficiency of the algorithm in the complex environment.
2. A method for path planning of an eVTOL drone in a dynamic environment according to claim 1, characterized in that: In the above S1, the deep reinforcement learning algorithm TDDPG that combines the deep deterministic policy gradient algorithm DDPG and the long short-term memory network LSTM enables the UAV to perform dynamic obstacle avoidance and dynamic end-point planning based on historical state and action information by utilizing the continuous action optimization ability of DDPG and the time series modeling ability of LSTM.
3. A method for path planning of an eVTOL drone in a dynamic environment according to claim 1, characterized in that: In the above S2, adopt the continuous action space, allowing the UAV to select any direction within the entire 360-degree range and adjust the speed in real time according to the environment. Through the continuous action space, the UAV can make more accurate and flexible adjustments and can respond to environmental changes in a timely manner during flight.
4. A method for path planning of an eVTOL UAV in a dynamic environment according to claim 1, characterized in that: In the above S3, this model takes into account the three-dimensional position (X, Y, Z), speed (v0, v1, v2) of the UAV, and the dynamic changes in its attitude angles; the motion state of the UAV is gradually updated through recursive formulas, combining control inputs and dynamic environmental factors to dynamically adjust the position, speed, and attitude angle variables; through this model, the state variables not only include the distance between the UAV and the target point and the relative position with respect to the obstacle, but also incorporate the position changes of dynamic obstacles, ensuring that the simulation results truly reflect the flight trajectory of the UAV in the multi-obstacle environment and can also effectively simulate the state changes of the UAV at different flight stages to ensure the accuracy of its motion trajectory.
5. A method for path planning of an eVTOL drone in a dynamic environment according to claim 1, characterized in that: The position (x, y, z) of the UAV, as the first part of the state variable, represents the position of the UAV in the three-dimensional space, and follows the following dynamic formula: x t+1 = x t + v t cos(θ t ) cos(φ t ) Δt y t+1 = y t + v t cos(θ t ) sin(φ t ) Δt z t+1 = z t - v t sin(θ t )Δt Among them, x t , y t , z t are the current three-dimensional coordinates of the UAV, v t is the current speed of the UAV, θ t and φ t are the pitch angle and yaw angle respectively; The speed of the eVTOL UAV is decomposed into components (v0, v1, v2) in three directions, where v0, v1, v2 respectively represent the speed components along the X, Y, and Z axes, and the formula is as follows: v0 = v·cos(θ)·cos(φ), v1 = v·cos(θ)·sin(φ), v2 = -v·sin(θ) where, v represents the speed of the eVTOL, θ is the pitch angle, and φ is the yaw angle; The distance d between the eVTOL and the target is used to measure the closeness of the UAV to the target; this state variable is calculated by the Euclidean distance formula, and the formula is as follows: where X opp , Y opp , Z opp represent the coordinates of the target position; The heading error reflects the angular deviation between the current heading of the eVOTL and the target direction; this error is obtained by calculating the included angle between the direction vector between the eVTOL and the target and the eVTOL speed vector: Where D0, D1, and D2 are the direction vectors between the eVTOL and the target, v is the speed of the eVTOL, and d is the distance between the eVTOL and the target; The state variables include not only the distance to static obstacles but also the states of dynamic obstacles; the distance between a static obstacle and the eVTOL is calculated by the following formula: where x obs , y obs , z obs are the coordinates of the obstacle, and r obs is the radius of the obstacle; the position of each dynamic obstacle is represented by dynamic x , dynamic_y, dynamic_z, and these positions will be updated as time progresses; the dynamic state information can help the agent predict the future movement trajectory of the obstacle, so as to plan the path in advance and avoid collisions; at the same time, in order to capture the movement trend of the dynamic obstacle, the historical dynamic obstacle movement information is saved as history_dynamic and used as a component of the state to improve the agent's ability to predict the movement trend of the dynamic obstacle.
6. A method for path planning of an eVTOL drone in a dynamic environment according to claim 1, characterized in that: In S4, the composite reward function is as follows: R = R collision + R align + R terminal where R collision represents the obstacle collision penalty, R align represents the orientation penalty, R terminal represents the terminal state reward; R collision The function is as follows: Where s min is the minimum safety distance. For each obstacle, the distance between the UAV and the obstacle is calculated and then the radius of the obstacle is subtracted, and the resulting value is used as the safety distance index. When s min is less than 1, it is considered that the distance between the UAV and the obstacle is too close, and then a penalty is imposed. When the distance is exactly 1.0, the penalty is 0. If the distance further decreases, the penalty will increase exponentially, so as to keep a sufficient safety distance between the UAV and the obstacle; R align The function is as follows: where θ e is the pitch angle error, reflecting the deviation between the UAV and the target height direction in the vertical plane; ψ e is the yaw angle error, reflecting the deviation between the actual flight direction of the UAV and the ideal target direction on the horizontal plane. R terminal The function is as follows: The termination conditions of the task are determined according to the state of the drone, specifically divided into reaching the target, colliding, or flying out of the boundary; when the eVTOL successfully explores the target location, the reward increases; if the eVTOL collides with an obstacle, a penalty is imposed; if the eVTOL flies out of the flight boundary, a corresponding penalty is also triggered.
7. A method for path planning of an eVTOL drone in a dynamic environment according to claim 1, characterized in that: In S5, in the initial stage of reinforcement learning, the agent accumulates experience through random exploration and uses the prioritized experience replay mechanism to accelerate the learning process, gradually shifting from random exploration to policy selection based on the Actor network. At the same time, Ornstein-Uhlenbeck noise is added to maintain exploration, and a linear scheduling strategy is adopted to optimize the sampling weights of important experiences, thereby improving the path planning efficiency and learning stability of the agent in complex dynamic environments.
Citation Information
Cited By
Multi-quad-rotor unmanned aerial vehicle centralized task path planning method based on reinforcement learning
CN116301007A
Multi-quadrotor unmanned aerial vehicle rendezvous task path planning method based on reinforcement learning
CN116301007B
Collision-free trajectory dynamic planning method and system for machining unit robot
CN122253227A