Real-time obstacle avoidance unmanned aerial vehicle flight path planning method
By combining methods such as multi-agent deep reinforcement learning and rapid exploration of random trees, sensor data is used to predict dynamic obstacle trajectories and make collaborative decisions, the challenges of UAV clusters to avoid obstacles and manage energy in complex environments are solved, and safe and efficient navigation is achieved.
Patent Information
- Application Number
- CN202411918387.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-25
AI Technical Summary
In complex environments, dynamic obstacles pose a major challenge to the safe flight of drones, and it is difficult for the existing technology to achieve collaborative cooperation and real-time obstacle avoidance of drones under energy constraints.
Multi-agent deep reinforcement learning is combined with rapid exploration of random trees and other methods, and environmental information is captured in real time through airborne multi-sensors, convolutional neural networks and recurrent neural networks are used to predict the future motion trajectory of dynamic obstacles, dynamically update the environment model, and generate preliminary flight trajectories from a global perspective, and collaborative decision-making is made in combination with the maximum reciprocity reward mechanism.
It realizes safe navigation and efficient energy management of drone groups in complex dynamic environments, ensuring the safety of drone groups and the efficiency of energy utilization.
Smart Images

Figure CN119937626A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of unmanned aerial vehicles, and in particular relates to a real-time obstacle avoidance unmanned aerial vehicle track planning method. Background Art
[0002] In the trajectory planning of drone swarms under energy constraints, obstacles (including static and dynamic obstacles) in complex environments pose a major challenge to the safe flight of drones. During the execution of tasks, drone swarms need to face the ever-changing environment and obstacles while maintaining efficient energy use and achieving mission goals. Especially in multi-task scenarios, drones must not only maintain flight accuracy and robustness, but also perform real-time obstacle detection and avoidance.
[0003] To achieve this goal, the technical solution of the trajectory planning method needs to be autonomous and collaborative. By utilizing multiple onboard sensors (such as lidar, cameras, etc.), drones can perceive the surrounding environment in real time and identify static obstacles (such as buildings, terrain, etc.) as well as dynamic obstacles (such as moving vehicles, other drones, etc.). This environmental perception information must be quickly and accurately integrated and combined with advanced path planning algorithms for dynamic obstacle avoidance.
[0004] The technical solution of obstacle avoidance trajectory planning not only considers the obstacle avoidance needs of a single drone, but also focuses on the collaboration of drone clusters under energy-constrained conditions to maximize global efficiency. This collaborative obstacle avoidance and energy-aware planning strategy can significantly improve the safety and execution efficiency of drone groups in complex and dynamic environments, thereby ensuring stable flight and effective energy management in diverse missions.
[0005] Therefore, it is a very practical task to study the real-time obstacle avoidance UAV trajectory planning method to ensure the safety of the UAV swarm when facing a complex and changing environment. Summary of the invention
[0006] In view of this, the object of the present invention is to propose a real-time obstacle avoidance UAV trajectory planning method, comprising the following steps:
[0007] Step 1: Use the onboard multi-sensor to detect the environment and capture the position, speed and direction of obstacles in the environment in real time;
[0008] Step 2: Use convolutional neural networks and recurrent neural networks to process the data collected by sensors and predict the future movement trajectory of dynamic obstacles;
[0009] Step 3: Build an environmental model to quantitatively assess the risks of the drone’s surroundings and dynamically update the model when new obstacles are detected;
[0010] Step 4: Generate a preliminary flight trajectory from a global perspective using deep reinforcement learning;
[0011] Step 5: Replan the trajectory to avoid dynamic obstacles based on the real-time environment information provided by the obstacle detection module;
[0012] Step 6: The drone group makes collaborative decisions collectively based on the experience sharing of multiple agents and the maximum reciprocal reward mechanism to ensure the safety of the entire group.
[0013] Specifically, the method of using a convolutional neural network and a recurrent neural network to process data collected by sensors and predict the future motion trajectory of a dynamic obstacle includes the following steps:
[0014] Acquire environmental data from multiple sensors, including image data from cameras and point cloud data from lidar;
[0015] The normalized image data and point cloud data are input into the convolutional neural network to extract spatial features;
[0016] The spatial features extracted from the convolutional neural network are input into the recurrent neural network for time series modeling to learn the movement patterns of obstacles;
[0017] Through recursive operations over multiple time steps, the recursive neural network predicts the trajectory of dynamic obstacles.
[0018] Specifically, the use of deep reinforcement learning to generate a preliminary flight trajectory from a global perspective includes the following steps:
[0019] State definition: The state of each drone includes its position, speed, remaining power, and information about surrounding obstacles. Let the state of the i-th drone be s i , whose state vector includes: s i =[x i ,y i ,z i ,v i ,E i ,O i ], where: (x i ,y i ,z i ) represents the position coordinates of the drone, v i represents the speed of the drone, E i Represents the remaining energy of the drone, O i Represents the information of surrounding obstacles, including distance, direction, and speed. In order to achieve collaboration among multiple agents, the global state S is used to represent the state of all drones and the environment: S = [s 1 ,s 2 ,…,s N], where N is the number of drones;
[0020] Action definition, the action of each drone includes its next displacement or speed adjustment. Let the action of the i-th drone be a i , whose action vector is: a i =[Δx i ,Δy i ,Δz i ,Δv i ], where Δx i ,Δy i ,Δz i represents the displacement increment in three dimensions, Δv i represents the velocity increment, and the global action A represents the joint action of all drones, A=[a 1 ,a 2 ,…,a N ];
[0021] Reward function definition: The reward function is used to measure the quality of the trajectory and encourage energy-efficient and obstacle-avoiding trajectory planning. Suppose the reward function of the i-th drone is r i , which is defined as follows: i =-αd i -βE i +γC i -δR i , where d i Indicates the distance between the drone and the target point. The shorter the distance, the higher the reward. i Represents energy consumption. The lower the energy consumption, the higher the reward. i represents the cooperation utility with other drones. The higher the utility, the higher the reward. i Represents the risk distance to the obstacle. The closer to the obstacle, the greater the penalty. α, β, γ, δ: coefficients for weighing various rewards;
[0022] During the training phase, the actions of each agent are controlled by a decentralized actor network, while the critic network is centralized and can obtain the global state S and global action A. The parameters of each actor network are determined by the strategy μ i (s i |θ i ) indicates that the critic network uses Q(S,A|φ) to evaluate the overall reward. Q(S,A|φ) is the critic network, which is used to evaluate the value of the global action.
[0023] The actor network uses the data in the experience replay buffer to update the policy parameters θ of the i-th drone. i , Represents the experience replay buffer, which stores past state-action-reward-next state samples; the critic network update uses the Bellman equation to update the critic network parameters φ;
[0024] During execution, each UAV uses only its own local observation s i And the trained actor network to make action decisions and achieve autonomous flight.
[0025] Preferably, before trajectory planning, UAVs with more energy are assigned to tasks that require higher energy, thereby balancing energy consumption. In the deployment phase, the initial energy E of each UAV is evaluated. i , assign tasks to them to ensure the optimal overall energy utilization. In order to plan energy-efficient paths, the critic network is used to evaluate the energy consumption and reward of each path, and select the path that maximizes the global reward. The path optimization function is defined as:
[0026]
[0027] in, The optimal action is to minimize energy consumption and maximize coordination with other drones; according to the evaluation results of the critic network, the flight action with the lowest energy consumption and that can meet the mission requirements is selected, and the position of the drone is updated at each time step:
[0028] s i,t+1 =s i,t +a i
[0029] Among them, s i,t+1 is the state of the i-th UAV at time step t+1, a i The currently selected action.
[0030] Specifically, the replanning of the trajectory to avoid dynamic obstacles according to the real-time environmental information provided by the obstacle detection module includes the following steps:
[0031] Obtain surrounding environment information through the obstacle detection module, including the real-time position, speed, and direction of obstacles;
[0032] Define the risk function R of each obstacle relative to the drone ij , used to quantify the threat level of obstacles to drones: Among them, d ij is the distance between UAV j and obstacle i, v ij is the relative speed of the obstacle to the UAV. The larger the value of the risk function, the higher the threat level of the obstacle.
[0033] For highly dependent drones, each other's actions are considered when making obstacle avoidance decisions to avoid conflicts. A policy function is used to determine the optimal action a of drone i. i : Among them, C(s i ,a) is the energy consumption cost, λ is the adjustment coefficient, s i represents the environmental state of the i-th UAV, and a represents the optional action of the i-th UAV, which is used to balance the relationship between obstacle avoidance risk and energy consumption;
[0034] Plan the optimal path for a drone to avoid obstacles using a fast-exploring random tree algorithm or a neural network-based approach.
[0035] Furthermore, the method of using a fast exploration random tree algorithm or a neural network-based method to plan the optimal path for a drone to avoid obstacles includes the following steps:
[0036] Initialize the fast exploration random tree and set the current position of the drone s start As the root node of the fast exploration random tree, the target position is s goal , using the information provided by the obstacle detection module, define the obstacle area And set it as the unreachable area in path planning;
[0037] Fast exploration of random tree path expansion: randomly sample a new position s in the planning space rand , find the distance s rand The nearest tree node s near , and along the s near To s rand Direction expansion, get new node s new : Where ∈ is the expansion step size.
[0038] Obstacle avoidance judgment: judge the newly expanded node s new Is it in the obstacle area? If Then add the node to the fast exploration random tree;
[0039] Repeat the above expansion process until the path is extended from the root node to the target location s goal ;
[0040] Obstacle trajectory predicted using a neural network Future collision checking of paths generated by the fast-exploring random tree algorithm;
[0041] For each path point s j , calculate the path point s j With obstacles The distance at time step t+l
[0042] If the distance is less than the safety distance threshold d safe , it is considered that there is a collision risk and the path needs to be replanned;
[0043] For path segments that may collide, a fast exploration random tree algorithm is used to resample and expand from the current node to find a new path to avoid obstacles;
[0044] The generated path is optimized using path smoothing techniques to reduce energy consumption and improve path feasibility;
[0045] The optimized path P after path smoothing smooth It is expressed as: Among them, P i represents the i-th point on the path. The goal is to minimize the distance between path points to obtain a smoother path. n represents the number of path points.
[0046] Specifically, the drone group collectively makes collaborative decisions based on the experience sharing of multiple agents and the maximum reciprocal reward mechanism, including the following steps:
[0047] During the mission, each drone stores its observations, actions, rewards, and state change information as experience. The experience of the i-th drone at time step t is:
[0048]
[0049] in, represents the state at time step t, Indicates the action to be performed. Indicates immediate reward, Indicates the new state after executing the action;
[0050] All drones store their experience into a shared experience pool A shared experience pool is used for learning and optimization in decision making;
[0051] When a dynamic obstacle is detected, A batch of experiences B is randomly sampled for learning to ensure the diversity and robustness of the learning process;
[0052] Define the reciprocal reward R rec , the reciprocal reward measures the contribution of each drone to other drones:
[0053]
[0054] in, is the point mutual information, which is used to quantify the degree of information sharing between UAV i and UAV j. is the contribution or impact of the action of drone i on drone j;
[0055] The total reward for each drone includes its own immediate reward and reciprocal rewards
[0056]
[0057] Each drone updates its policy network based on shared experience and maximum reciprocal reward mechanism. To determine the best action in: is the state-action value function, which is used to estimate the possible reward after taking action a;
[0058] Trajectory generation and adjustment: Based on the local actions of all drones, a collective trajectory is generated to ensure that the entire drone group can pass safely;
[0059] Path optimization objective,The goal of path optimization is to maximize the overall reward of the group while minimizing energy consumption and avoiding obstacles:
[0060]
[0061] Among them, P is the candidate path; is the total energy consumption of the ith UAV in T time steps, α is the weight coefficient of energy consumption,
[0062] The trajectory P generated by the drone according to the collective decision i Perform flights while maintaining communication with other drones to obtain mutual status information;
[0063] When a new obstacle or environmental change is detected, the drone swarm reevaluates the current trajectory and updates the action, adjusting in real time by sampling again from the shared experience pool and combining it with the current maximum reciprocal reward.
[0064] The beneficial effects of the present invention are as follows: In order to effectively respond to dynamic environmental changes, the technical solution combines multi-agent deep reinforcement learning with methods such as rapid exploration random trees to ensure decentralized collaborative decision-making within the drone cluster. Each drone relies on local information for path planning based on a partial observation Markov decision process, and improves the response speed and decision-making quality of the entire cluster through a shared experience pool. In order to ensure the safety of the drone cluster and efficient use of energy, the solution also combines the maximum reciprocal reward mechanism to evaluate the cooperation effect between drones to achieve global optimal task execution and energy allocation. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 An overall flow chart of a real-time obstacle avoidance UAV trajectory planning method is shown. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0067] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0068] like Figure 1 As shown, this embodiment proposes a real-time obstacle avoidance UAV trajectory planning method, comprising the following steps:
[0069] Step 1: Use the onboard multi-sensor to detect the environment and capture the position, speed and direction of obstacles in the environment in real time;
[0070] Step 2: Use convolutional neural networks and recurrent neural networks to process the data collected by sensors and predict the future movement trajectory of dynamic obstacles;
[0071] Step 3: Build an environmental model to quantitatively assess the risks of the drone’s surroundings and dynamically update the model when new obstacles are detected;
[0072] Step 4: Generate a preliminary flight trajectory from a global perspective using deep reinforcement learning;
[0073] Step 5: Replan the trajectory to avoid dynamic obstacles based on the real-time environment information provided by the obstacle detection module;
[0074] Step 6: The drone group makes collaborative decisions collectively based on the experience sharing of multiple agents and the maximum reciprocal reward mechanism to ensure the safety of the entire group.
[0075] The data collected by multiple sensors (such as lidar, visual sensors, infrared sensors, etc.) are used to process and predict the future movement trajectory of dynamic obstacles through the combination of CNN and RNN. CNN is used to extract spatial features from input images or point cloud data, and RNN is used to capture the time series features of dynamic obstacles, thereby completing the prediction of the future movement trajectory of obstacles.
[0076] Specifically, the method of using a convolutional neural network and a recurrent neural network to process data collected by sensors and predict the future motion trajectory of a dynamic obstacle includes the following steps:
[0077] Acquire environmental data from multiple sensors, including image data from cameras and point cloud data from lidar;
[0078] The normalized image data and point cloud data are input into the convolutional neural network to extract spatial features;
[0079] The spatial features extracted from the convolutional neural network are input into the recurrent neural network for time series modeling to learn the movement patterns of obstacles;
[0080] Through recursive operations over multiple time steps, the recursive neural network predicts the trajectory of dynamic obstacles.
[0081] By combining CNN and RNN methods, we can extract spatial and temporal features from data collected by multiple sensors to effectively predict the future trajectory of dynamic obstacles. Convolutional neural networks (CNN) are used for feature extraction, while recurrent neural networks (RNN), especially LSTM or GRU, are used to capture time series characteristics. This combined method provides an intelligent obstacle detection and avoidance solution for energy-constrained drones, thereby ensuring safe navigation in complex dynamic environments.
[0082] Perform multi-agent trajectory planning from a global perspective and combine it with an energy-efficient deployment strategy to generate preliminary flight trajectories. This approach ensures that each drone completes its mission using the optimal path while coordinating with other drones. The steps include defining states, actions and rewards, training, and applying energy-efficient path planning.
[0083] Specifically, the use of deep reinforcement learning to generate a preliminary flight trajectory from a global perspective includes the following steps:
[0084] State definition: The state of each drone includes its position, speed, remaining power, and information about surrounding obstacles. Let the state of the i-th drone be s i , whose state vector includes: s i =[x i ,y i ,z i ,v i ,E i ,O i ], where: (x i ,y i ,z i ) represents the position coordinates of the drone, v i represents the speed of the drone, Ei Represents the remaining energy of the drone, O i Represents the information of surrounding obstacles, including distance, direction, and speed. In order to achieve collaboration among multiple agents, the global state S is used to represent the state of all drones and the environment: S = [s 1 ,s 2 ,…,s N ], where N is the number of drones;
[0085] Action definition, the action of each drone includes its next displacement or speed adjustment. Let the action of the i-th drone be a i , whose action vector is: a i =[Δx i ,Δy i ,Δz i ,Δv i ], where Δx i ,Δy i ,Δz i represents the displacement increment in three dimensions, Δv i represents the velocity increment, and the global action A represents the joint action of all drones, A=[a 1 ,a 2 ,…,a N ];
[0086] Reward function definition: The reward function is used to measure the quality of the trajectory and encourage energy-efficient and obstacle-avoiding trajectory planning. Suppose the reward function of the i-th drone is r i , which is defined as follows: i =-αd i -βE i +γC i -δR i , where d i Indicates the distance between the drone and the target point. The shorter the distance, the higher the reward. i Represents energy consumption. The lower the energy consumption, the higher the reward. i represents the cooperation utility with other drones. The higher the utility, the higher the reward. i Represents the risk distance to the obstacle. The closer to the obstacle, the greater the penalty. α, β, γ, δ: coefficients for weighing various rewards;
[0087] During the training phase, the actions of each agent are controlled by a decentralized actor network, while the critic network is centralized and can obtain the global state S and global action A. The parameters of each actor network are determined by the strategy μ i (s i |θ i) indicates that the critic network uses Q(S,A|φ) to evaluate the overall reward. Q(S,A|φ) is the critic network, which is used to evaluate the value of the global action.
[0088] The actor network uses the data in the experience replay buffer to update the policy parameters θ of the i-th drone. i , Represents the experience replay buffer, which stores past state-action-reward-next state samples; the critic network update uses the Bellman equation to update the critic network parameters φ;
[0089] During execution, each UAV uses only its own local observation s i And the trained actor network to make action decisions and achieve autonomous flight.
[0090] By defining global states, actions, and reward functions, collaboration and path optimization among multiple agents are achieved. Through centralized training and decentralized execution, preliminary flight trajectories can be generated from a global perspective, while energy-efficient deployment strategies are combined to ensure that the drone swarm can complete the task with optimal energy efficiency. In complex dynamic environments, this solution can achieve collaborative obstacle avoidance and efficient flight of drones through intelligent path planning.
[0091] Preferably, before trajectory planning, UAVs with more energy are assigned to tasks that require higher energy, thereby balancing energy consumption. In the deployment phase, the initial energy E of each UAV is evaluated. i , assign tasks to them to ensure the optimal overall energy utilization. In order to plan energy-efficient paths, the critic network is used to evaluate the energy consumption and reward of each path, and select the path that maximizes the global reward. The path optimization function is defined as:
[0092]
[0093] in, The optimal action is to minimize energy consumption and maximize coordination with other drones; according to the evaluation results of the critic network, the flight action with the lowest energy consumption and that can meet the mission requirements is selected, and the position of the drone is updated at each time step:
[0094] s i,t+1 =s i,t +a i
[0095] Among them, s i,t+1 is the state of the i-th UAV at time step t+1, a i The currently selected action.
[0096] In the trajectory planning of drone swarms under energy constraints, dynamic obstacle avoidance is an important part to ensure the safe operation of drones in complex environments. Combined with the real-time environmental information provided by the obstacle detection module and the multi-agent point mutual information mechanism, the drone swarm can efficiently achieve dynamic obstacle avoidance.
[0097] Specifically, the replanning of the trajectory to avoid dynamic obstacles according to the real-time environmental information provided by the obstacle detection module includes the following steps:
[0098] Obtain surrounding environment information through the obstacle detection module, including the real-time position, speed, and direction of obstacles;
[0099] Define the risk function R of each obstacle relative to the drone ij , used to quantify the threat level of obstacles to drones: Among them, d ij is the distance between UAV j and obstacle i, v ij is the relative speed of the obstacle to the UAV. The larger the value of the risk function, the higher the threat level of the obstacle.
[0100] For highly dependent drones, each other's actions are considered when making obstacle avoidance decisions to avoid conflicts. A policy function is used to determine the optimal action a of drone i. i : Among them, C(s i ,a) is the energy consumption cost, λ is the adjustment coefficient, s i represents the environmental state of the i-th UAV, and a represents the optional action of the i-th UAV, which is used to balance the relationship between obstacle avoidance risk and energy consumption;
[0101] Plan the optimal path for a drone to avoid obstacles using a fast-exploring random tree algorithm or a neural network-based approach.
[0102] In the trajectory planning of drone swarms under energy constraints, the combination of fast exploration random trees and trajectory prediction methods based on neural networks can achieve real-time planning for dynamic obstacle avoidance and ensure the safety of drones. Fast exploration random trees are an efficient path search algorithm that can quickly find the path from the starting point to the end point, while neural networks are used to predict the future trajectory of dynamic obstacles, thereby optimizing the obstacle avoidance path.
[0103] Furthermore, the method of using a fast exploration random tree algorithm or a neural network-based method to plan the optimal path for a drone to avoid obstacles includes the following steps:
[0104] Initialize the fast exploration random tree and set the current position of the drone s startAs the root node of the fast exploration random tree, the target position is s goal , using the information provided by the obstacle detection module, define the obstacle area And set it as the unreachable area in path planning;
[0105] Fast exploration of random tree path expansion: randomly sample a new position s in the planning space rand , find the distance s rand The nearest tree node s near , and along the s near To s rand Direction expansion, get new node s new : Where ∈ is the expansion step size.
[0106] Obstacle avoidance judgment: judge the newly expanded node s new Is it in the obstacle area? If Then add the node to the fast exploration random tree;
[0107] Repeat the above expansion process until the path is extended from the root node to the target location s goal ;
[0108] Obstacle trajectory predicted using a neural network Future collision checking of paths generated by the fast-exploring random tree algorithm;
[0109] For each path point s j , calculate the path point s j With obstacles The distance at time step t+l
[0110] If the distance is less than the safety distance threshold d safe , it is considered that there is a collision risk and the path needs to be replanned;
[0111] For path segments that may collide, a fast exploration random tree algorithm is used to resample and expand from the current node to find a new path to avoid obstacles;
[0112] The generated path is optimized using path smoothing techniques to reduce energy consumption and improve path feasibility;
[0113] The optimized path P after path smoothing smooth It is expressed as: Among them, P i represents the i-th point on the path. The goal is to minimize the distance between path points to obtain a smoother path. n represents the number of path points.
[0114] The main purpose of the experience replay mechanism is to improve the learning ability of drones in complex environments by storing and reusing past experience, especially when facing new obstacles, countermeasures can be obtained from experience. This embodiment combines the experience replay mechanism and the maximum reciprocal reward to re-optimize the trajectory of the drone. Experience replay is used to store and reuse past experience, and the maximum reciprocal reward algorithm is used to ensure cooperation between drone groups, so that trajectory optimization not only considers individual goals, but also improves overall performance.
[0115] Specifically, the drone group collectively makes collaborative decisions based on the experience sharing of multiple agents and the maximum reciprocal reward mechanism, including the following steps:
[0116] During the mission, each drone stores its observations, actions, rewards, and state change information as experience. The experience of the i-th drone at time step t is:
[0117]
[0118] in, represents the state at time step t, Indicates the action to be performed. Indicates immediate reward, Indicates the new state after executing the action;
[0119] All drones store their experience into a shared experience pool A shared experience pool is used for learning and optimization in decision making;
[0120] When a dynamic obstacle is detected, Randomly sample a batch B of experience for learning: Among them, M is the batch size, which is used to ensure the diversity and robustness of the learning process;
[0121] Define the reciprocal reward R rec , the reciprocal reward measures the contribution of each drone to other drones:
[0122]
[0123] in, is the point mutual information, which is used to quantify the degree of information sharing between UAV i and UAV j. is the contribution or impact of the action of drone i on drone j;
[0124] The total reward for each drone includes its own immediate reward and reciprocal rewards
[0125]
[0126] Each drone updates its policy network based on shared experience and maximum reciprocal reward mechanism. To determine the best action in: is the state-action value function, which is used to estimate the possible reward after taking action a;
[0127] Trajectory generation and adjustment, based on the local actions of all drones, a collective trajectory is generated to ensure that the entire drone group can pass safely. Let the trajectory of the i-th drone be P i , which consists of a series of location points:
[0128]
[0129] Where T is the number of time steps for planning;
[0130] Path optimization objective,The goal of path optimization is to maximize the overall reward of the group while minimizing energy consumption and avoiding obstacles:
[0131]
[0132] Among them, P is the candidate path; is the total energy consumption of the ith UAV in T time steps, α is the weight coefficient of energy consumption,
[0133] The trajectory P generated by the drone according to the collective decision i Perform flights while maintaining communication with other drones to obtain mutual status information;
[0134] When a new obstacle or environmental change is detected, the drone swarm reevaluates the current trajectory and updates the action, adjusting in real time by sampling again from the shared experience pool and combining it with the current maximum reciprocal reward.
[0135] In the trajectory planning of drone swarms under energy constraints, when dynamic obstacles are detected, multi-agent experience sharing and maximum reciprocal reward mechanisms are used to ensure that the drone swarm can make collective collaborative decisions. The experience sharing mechanism helps the drone swarm improve its learning ability to cope with complex environments, while the maximum reciprocal reward mechanism ensures cooperation between drones and maximizes the overall mission benefits. In trajectory planning, by maximizing rewards and minimizing energy consumption, the drone swarm can achieve efficient and safe dynamic obstacle avoidance.
[0136] As used herein, the word "preferred" is intended to be used as an example, instance, or illustration. Any aspect or design described herein as "preferred" is not necessarily to be construed as being more advantageous than other aspects or designs. On the contrary, the use of the word "preferred" is intended to present concepts in a specific way. The term "or" as used in this application is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" means any one of the naturally included permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.
[0137] Moreover, although the present disclosure has been shown and described with respect to one or implementations, those skilled in the art will think of equivalent variations and modifications based on the reading and understanding of this specification and the accompanying drawings. The present disclosure includes all such modifications and variations, and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-mentioned components (such as elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (such as it is functionally equivalent), even if the structure is not equivalent to the disclosed structure of the function in the exemplary implementation of the present disclosure shown herein. In addition, although the specific features of the present disclosure have been disclosed with respect to only one of several implementations, such features can be combined with one or other features of other implementations that may be desired and advantageous for a given or specific application. Moreover, insofar as the terms "including", "having", "containing" or their variations are used in specific embodiments or claims, such terms are intended to be included in a manner similar to the term "comprising".
[0138] The functional units in the embodiments of the present invention may be integrated into a processing module, or each unit may exist physically separately, or multiple or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc. The above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.
[0139] To sum up, the above embodiment is an implementation mode of the present invention, but the implementation mode of the present invention is not limited by the embodiment. Any other changes, modifications, substitutions, combinations, and simplifications that deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A real-time obstacle avoidance UAV trajectory planning method, characterized in that: The steps include: Step 1: Use the onboard multi-sensor to detect the environment and capture the position, speed and direction of obstacles in the environment in real time; Step 2: Use convolutional neural networks and recurrent neural networks to process the data collected by sensors and predict the future movement trajectory of dynamic obstacles; Step 3: Build an environmental model to quantitatively assess the risks of the drone’s surroundings and dynamically update the model when new obstacles are detected; Step 4: Generate a preliminary flight trajectory from a global perspective using deep reinforcement learning; Step 5: Replan the trajectory to avoid dynamic obstacles based on the real-time environment information provided by the obstacle detection module; Step 6: The drone group makes collaborative decisions collectively based on the experience sharing of multiple agents and the maximum reciprocal reward mechanism to ensure the safety of the entire group.
2. A real-time obstacle avoidance UAV trajectory planning method according to claim 1, characterized in that: The method of using a convolutional neural network and a recurrent neural network to process the data collected by the sensor and predict the future motion trajectory of the dynamic obstacle includes the following steps: Acquire environmental data from multiple sensors, including image data from cameras and point cloud data from lidar; The normalized image data and point cloud data are input into the convolutional neural network to extract spatial features; The spatial features extracted from the convolutional neural network are input into the recurrent neural network for time series modeling to learn the movement patterns of obstacles; Through recursive operations over multiple time steps, the recursive neural network predicts the trajectory of dynamic obstacles.
3. The real-time obstacle avoidance UAV trajectory planning method according to claim 1 is characterized in that: The method of generating a preliminary flight trajectory from a global perspective based on deep reinforcement learning includes the following steps: State definition: The state of each drone includes its position, speed, remaining power, and information about surrounding obstacles. Let the state of the i-th drone be s i , whose state vector includes: s i =[x i ,y i ,z i ,v i ,E i ,O i ], where (x i ,y i ,z i ) represents the position coordinates of the drone, v i represents the speed of the drone, E i Represents the remaining energy of the drone, O i Represents the information of surrounding obstacles, including distance, direction, and speed. In order to achieve collaboration among multiple agents, the global state S is used to represent the state of all drones and the environment: S = [s1, s2, …, s N ], where N is the number of drones; Action definition, the action of each drone includes its next displacement or speed adjustment. Let the action of the i-th drone be a i , whose action vector is: a i =[Δx i ,Δy i ,Δz i ,Δv i ], where Δx i ,Δy i ,Δz i represents the displacement increment in three dimensions, Δv i represents the velocity increment, and the global action A represents the joint action of all drones, A=[a1,a2,…,a N ]; Reward function definition: The reward function is used to measure the quality of the trajectory and encourage energy-efficient and obstacle-avoiding trajectory planning. Suppose the reward function of the i-th drone is r i , defined as follows: i =-αd i -βE i +γC i -δR i , where d i Indicates the distance between the drone and the target point. The shorter the distance, the higher the reward. i Represents energy consumption. The lower the energy consumption, the higher the reward. i represents the cooperation utility with other drones. The higher the utility, the higher the reward. i Represents the risk distance to the obstacle. The closer to the obstacle, the greater the penalty. α, β, γ, and δ are used to weigh the coefficients of various rewards. During the training phase, the actions of each agent are controlled by a decentralized actor network, while the critic network is centralized and can obtain the global state S and global action A. The parameters of each actor network are determined by the strategy μ i (s i |θ i ) indicates that the critic network uses Q(S,A|φ) to evaluate the overall reward. Q(S,A|φ) is the critic network, which is used to evaluate the value of the global action. The actor network uses the data in the experience replay buffer to update the policy parameters θ of the i-th drone. i , Represents the experience replay buffer, which stores past state-action-reward-next state samples; the critic network update uses the Bellman equation to update the critic network parameters φ; During execution, each UAV uses only its own local observation s i And the trained actor network to make action decisions and achieve autonomous flight.
4. A real-time obstacle avoidance UAV trajectory planning method according to claim 3, characterized in that: Before trajectory planning, drones with more energy are assigned to tasks that require higher energy to balance energy consumption. In the deployment phase, the initial energy of each drone is evaluated and tasks are assigned to ensure optimal overall energy utilization. In order to plan energy-efficient paths, the critic network is used to evaluate the energy consumption and reward of each path and select the path that maximizes the global reward. The path optimization function is defined as: in, The optimal action is to minimize energy consumption and maximize coordination with other drones. According to the evaluation results of the critic network, the flight action with the lowest energy consumption and meeting the mission requirements is selected, and the position of the drone is updated at each time step: s i,t+1 =s i,t +a i Among them, s i,t+1 is the state of the i-th UAV at time step t+1, a i The currently selected action.
5. The real-time obstacle avoidance UAV trajectory planning method according to claim 1 is characterized in that: The replanning of the trajectory to avoid dynamic obstacles according to the real-time environmental information provided by the obstacle detection module comprises the following steps: Obtain surrounding environment information through the obstacle detection module, including the real-time position, speed, and direction of obstacles; Define the risk function R of each obstacle relative to the drone ij , used to quantify the threat level of obstacles to drones: Among them, d ij is the distance between UAV j and obstacle i, v ij is the relative speed of the obstacle to the UAV. The larger the value of the risk function, the higher the threat level of the obstacle. For highly dependent drones, each other's actions are considered when making obstacle avoidance decisions to avoid conflicts. A policy function is used to determine the optimal action a of drone i. i : Among them, C(s i ,a) is the energy consumption cost, λ is the adjustment coefficient, s i represents the environmental state of the i-th UAV, and a represents the optional action of the i-th UAV, which is used to balance the relationship between obstacle avoidance risk and energy consumption; Plan the optimal path for a drone to avoid obstacles using a fast-exploring random tree algorithm or a neural network-based approach.
6. A real-time obstacle avoidance UAV trajectory planning method according to claim 5, characterized in that: The method of using a fast exploration random tree algorithm or a neural network-based method to plan the optimal path for a drone to avoid obstacles includes the following steps: Initialize the fast exploration random tree and set the current position of the drone s start As the root node of the fast exploration random tree, the target position is s goal , using the information provided by the obstacle detection module, define the obstacle area And set it as the unreachable area in path planning; Fast exploration of random tree path expansion: randomly sample a new position s in the planning space rand , find the distance s rand The nearest tree node s near , and along the s near To s rand Direction expansion, get new node s new : Among them, ∈ is the expansion step size. Obstacle avoidance judgment: judge the newly expanded node s new Is it in the obstacle area? If Then add the node to the fast exploration random tree; Repeat the above expansion process until the path is extended from the root node to the target location s goal ; Obstacle trajectory predicted using a neural network Future collision checking of paths generated by the fast-exploring random tree algorithm; For each path point s j , calculate the path point s j With obstacles The distance at time step t+l If the distance is less than the safety distance threshold d safe , it is considered that there is a collision risk and the path needs to be replanned; For path segments that may collide, a fast exploration random tree algorithm is used to resample and expand from the current node to find a new path to avoid obstacles; The generated path is optimized using path smoothing techniques to reduce energy consumption and improve path feasibility; The optimized path P after path smoothing smooth It is expressed as: Among them, P i represents the i-th point on the path. The goal is to minimize the distance between path points to obtain a smoother path. n represents the number of path points.
7. The real-time obstacle avoidance UAV trajectory planning method according to claim 1, characterized in that: The drone group collectively makes collaborative decisions based on the experience sharing of multiple agents and the maximum reciprocal reward mechanism, including the following steps: During the mission, each drone stores its observations, actions, rewards, and state change information as experience. The experience of the i-th drone at time step t is: in, represents the state at time step t, Indicates the action to be performed. Indicates immediate reward, Indicates the new state after executing the action; All drones store their experience into a shared experience pool A shared experience pool is used for learning and optimization in decision making; When a dynamic obstacle is detected, A batch of experiences B is randomly sampled for learning to ensure the diversity and robustness of the learning process; Define the reciprocal reward R rec , the reciprocal reward measures the contribution of each drone to other drones: in, is the point mutual information, which is used to quantify the degree of information sharing between UAV i and UAV j. is the contribution or impact of the action of drone i on drone j; The total reward for each drone includes its own immediate reward and reciprocal rewards Each drone updates its policy network based on shared experience and maximum reciprocal reward mechanism. To determine the best action in: is the state-action value function, which is used to estimate the possible reward after taking action a; Trajectory generation and adjustment: Based on the local actions of all drones, a collective trajectory is generated to ensure that the entire drone group can pass safely; Path optimization objective,The goal of path optimization is to maximize the overall reward of the group while minimizing energy consumption and avoiding obstacles: Among them, P is the candidate path; is the total energy consumption of the ith UAV in T time steps, α is the weight coefficient of energy consumption, The trajectory P generated by the drone according to the collective decision i Perform flights while maintaining communication with other drones to obtain mutual status information; When a new obstacle or environmental change is detected, the drone swarm reevaluates the current trajectory and updates the action, adjusting in real time by sampling again from the shared experience pool and combining it with the current maximum reciprocal reward.
Citation Information
Patent Citations
Target tracking control method for autonomous underwater vehicle based on trajectory prediction
CN115657689A
Unmanned aerial vehicle obstacle avoidance method and device, storage medium and equipment
CN117666605A
Multi-unmanned aerial vehicle trajectory planning and data collection method based on collaborative reinforcement learning
CN118859987A
Formation path planning method fusing experience sharing and balance award Actor-Critic network
CN118915772A
Multi-robot trajectory planning method
WO2022241808A1
Cited By
Multi-unmanned aerial vehicle adaptive collaborative planning method based on game theory
CN120560343A
Unmanned aerial vehicle group intelligent path planning and obstacle avoidance method and system based on artificial intelligence
CN120848543A
Unmanned aerial vehicle route intelligent planning and obstacle avoidance method based on deep learning
CN121070026A
Unmanned aerial vehicle real-time flight path planning method based on improved A* algorithm
CN122429819A