Unmanned aerial vehicle path planning method under energy constraint condition
By using probability maps and reinforcement learning technology in drone path planning, dynamically adjusting the paths to cope with energy constraints and environmental changes, solving the problem that traditional methods are difficult to achieve efficient and safe task planning in complex environments, and achieving efficient and safe task completion of drones in complex environments.
Patent Information
- Application Number
- CN202411918386.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-25
AI Technical Summary
In complex environments, traditional drone path planning methods are difficult to effectively deal with energy constraints, environmental uncertainty and dynamically changing task requirements, making it difficult to achieve efficient and safe task planning under limited energy.
A method of path planning under energy constraints is proposed. Through the probability map, real-time update is carried out, multiple paths are generated and energy consumption is estimated, the optimal initial path is selected, and the path is dynamically adjusted through reinforcement learning to ensure that the drone completes tasks efficiently and safely in complex environments.
It realizes the ability to complete tasks efficiently and safely under limited energy, improves the navigation capabilities of the drone in complex environments and the intelligence level of the system, and is suitable for complex and uncertain environments.
Smart Images

Figure CN119937616A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicles, and in particular relates to a path planning method for unmanned aerial vehicles under energy constraint conditions. Background Art
[0002] As an emerging intelligent flight platform, drones are increasingly used in military, logistics, monitoring, agriculture and other fields. With the development of drone technology, its autonomous navigation and mission execution capabilities in complex environments have gradually become the focus of research. Traditional drone path planning methods mostly rely on pre-set rules or path optimization techniques based on graph search algorithms. Although these methods can cope with relatively simple static environments, they show great limitations when facing dynamically changing, uncertain and complex environments.
[0003] UAV flight missions in complex environments are usually accompanied by many challenges, including limited energy resources, complex terrain obstacles, unstable weather conditions, and dynamically changing mission requirements. In order to solve these problems, path planning not only needs to consider the shortest path or the lowest energy consumption, but also needs to deal with environmental uncertainties, avoid obstacles, and ensure the safe flight of the UAV. At the same time, UAVs need to take into account both energy efficiency and the real-time nature of the mission when performing missions. How to complete efficient and safe mission planning under the premise of limited energy has become an important topic in current UAV path planning research.
[0004] Therefore, we study the efficient, safe and stable planning method of UAV paths under energy constraints, so as to ensure that the UAV can optimize energy usage while completing the mission when facing complex environments. Summary of the invention
[0005] In view of this, the object of the present invention is to propose a UAV path planning method under energy constraints, comprising the following steps:
[0006] Step 1: Obtain environment and drone information and use probability maps for real-time updates;
[0007] Step 2: Considering the task requirements and energy constraints, generate multiple possible paths, estimate the energy consumption of each path, and select the optimal initial path;
[0008] Step 3: Based on the optimal initial path, learn and dynamically adjust the flight path according to the feedback from the probability map;
[0009] Step 4: Continuously monitor the drone’s energy level and adjust the drone’s mission and return path in real time.
[0010] Specifically, the real-time updating using the probability map includes the following steps:
[0011] Initialize the map. The map is represented by a grid map or a continuous space, where each grid or area contains an initial probability indicating whether the location is occupied;
[0012] Collect surrounding environment data in real time through sensors as the basis for environmental perception;
[0013] Based on the sensor's measurements, the probability of whether there is an obstacle at each location is updated using the Bayesian update rule;
[0014] As the drone moves and environmental information is updated, the probability map is constantly adjusted and updated.
[0015] Specifically, the Bayesian update rule is used to update the probability of whether there is an obstacle at each position, including the following steps: when the drone is flying, the sensor collects the information corresponding to a certain grid m i The relevant measurement value z t , the measured value z t Represents the grid m at time t i Whether it is occupied or not, at this time, the probability of each grid is updated using the Bayesian formula, and the Bayesian update formula is:
[0016]
[0017] Among them, m i represents the position of the i-th grid, s i Indicates the state of the i-th grid, whether it is occupied, P(s i |z t ) indicates that when the sensor observes z t After that, grid m i The posterior probability of being occupied, P(z t |s i ) represents a given grid m i The sensor observes z t The probability of i ) represents the grid m i Prior probability of being occupied, P(z t ) indicates that the sensor observes z t The edge probability is calculated by weighting all possible grid states.
[0018] Furthermore, the method of considering the task requirements and energy constraints, generating multiple possible paths, estimating the energy consumption of each path, and selecting the optimal initial path includes the following steps:
[0019] The probabilistic path search algorithm is used to define the state space as a set of nodes in a grid map. The state transition is the process of moving from one grid to the next grid. The goal is to find the path from the initial position x0 to the target position x g Multiple paths of
[0020] The cumulative risk of a path is estimated by the grid occupancy probability in the probability map. The risk R(π) of a path π is expressed as the sum of the occupancy probabilities of all grids passed on the path:
[0021]
[0022] Among them, P(s i =1) indicates grid m i The probability of being occupied by an obstacle, i represents a grid on the path π;
[0023] Calculate the energy consumption of flight distance and speed. The energy consumption per unit time of the drone on the path is expressed as flight power P, which is proportional to the flight speed v and the flight distance d. The energy consumption E(π) of path π is calculated by the following formula:
[0024]
[0025] Among them, P(v i ) represents the flight power of the UAV in the i-th segment of the path, which depends on the flight speed v i , represents the flight time of the i-th segment, d i is the flight distance between two locations on the path, and n represents the number of segments of the path;
[0026] Calculate the energy consumption of terrain and wind speed. When considering the influence of terrain, the energy consumption is calculated by adding the vertical height change h caused by the terrain. i To make corrections:
[0027]
[0028] Where γ is a constant related to the weight of the drone, h i represents the height difference on the i-th path, and the wind speed w i The effect of is taken into account by correcting the flight power:
[0029] P(v i ,w i )=P(v i )·(1+αw i )
[0030] Among them, α is the coefficient of wind speed on flight power, w i represents the wind speed of the i-th path;
[0031] When the energy consumption E(π) and risk R(π) of multiple paths are calculated, an optimization algorithm is used to select the optimal path, and a comprehensive objective function J(π) is defined to evaluate the pros and cons of each path:
[0032] J(π)=λ1E(π)+λ2R(π)
[0033] Among them, λ1 and λ2 are the weight coefficients of energy and risk, respectively, reflecting the priority of the task for different goals. By adjusting the weights, the system finds a balance between energy saving and risk avoidance.
[0034] Preferably, the method of considering the task requirements and energy constraints, generating multiple possible paths, estimating the energy consumption of each path, and selecting the optimal initial path includes the following steps:
[0035] Step 201, initialize a tree structure T = (V, E), where V is a node set, initially including the starting point V = {x start}, E is the set of edges, initially empty; set the initial position of the drone to x start and the target position is x goal , construct a probability map Each position m in the map i The probability of including obstacles P(s i ), which is used in the path evaluation process to determine whether a certain area can be safely passed through and to initialize the energy constraint condition E max , represents the maximum available energy of the UAV;
[0036] Step 202, randomly sample a point x from the space rand , the point is within the range of the configuration space, which is defined as the flight area of the UAV;
[0037] Step 203, find the distance x in the current tree T rand The nearest node x nearest , the distance calculation formula is as follows: ‖·‖ represents the Euclidean distance;
[0038] Step 204, from x nearest To x rand Expand and generate a new node x new , the new node is at x nearest to x rand In the direction of , and expand with a step size of Δd:
[0039]
[0040] The new node x after expansion new Add the node set V of the tree and add the edge (xnearest ,x new )Add edge set E;
[0041] Step 205, when expanding the tree, check nearest to x new Whether the path passes through the obstacle area in the probability map, calculate the risk of the path, the risk R(x nearest ,x new ) is estimated by the occupancy probability of the grid on the path:
[0042]
[0043] Among them, π(x nearest ,x new ) indicates that from x nearest to x new All grids passed by on the path, P(s i =1) is the grid s i The probability of being occupied by obstacles; if the cumulative risk of the path exceeds the preset risk threshold, the extension is discarded and x is not new Add to the tree;
[0044] Step 206: for each newly generated path segment, consider the flight distance and speed and evaluate the energy consumption. The energy consumption E(x nearest ,x new ) is:
[0045]
[0046] Where P(v) is the flight power corresponding to the flight speed v, ‖x new -x nearest ‖ is the distance of the path segment, v is the flight speed, if the cumulative energy consumption of the path exceeds the maximum energy E of the drone max , then stop expanding the path;
[0047] Step 207, repeat steps 202 to 206, and continue to expand the tree until the connection starting point x is found start and the target point x goal A feasible path of x, or the preset maximum number of iterations is reached; multiple potential paths are generated, each path corresponds to start to x goal The energy cost and risk of each path on a tree are calculated and recorded during the expansion process.
[0048] Step 208, among the generated multiple paths, select the best path that meets the energy constraint, and select the path through an optimization function, which comprehensively considers energy consumption E(π) and risk R(π):
[0049] J(π)=λ1E(π)+λ2R(π)
[0050] Among them, E(π) represents the total energy consumption of path π, R(π) represents the risk of the path, λ1,λ2 are the weight coefficients of energy consumption and risk, which are adjusted according to task requirements.
[0051] Preferably, in step 206, considering the influence of terrain height change h and wind speed w, the energy consumption of the path segment is corrected to:
[0052]
[0053] Where: γ is the coefficient of altitude change on energy consumption; α is the coefficient of wind speed on energy consumption.
[0054] Furthermore, when the energy consumption is considered in terms of flight distance and speed, the flight power P(v) varies with speed, and the calculation formula is: P(v) = c1v 3 +c2v 2 +c3v+c4, where c1, c2, c3, c4 are constants related to the drone model, c1v 3 represents aerodynamic drag, c2v 2 represents propulsion power, c3v represents mechanical loss, and c4 represents fixed power loss.
[0055] Furthermore, the energy consumption also takes into account the energy consumption of attitude adjustment, including the energy consumption of steering and the energy consumption of acceleration and deceleration. During the steering process of the UAV in the path, due to the effect of inertia, the additional energy consumption is related to the steering angle θ: E turn (θ) = k turn ·θ, where E turn (θ) represents the energy consumption when the steering angle is θ, k turn is a constant related to the UAV model and flight dynamics, which represents the energy consumption coefficient of steering; if the UAV needs to accelerate or decelerate on the path, additional power is required to overcome the change in speed. Assuming that the speed of the UAV changes from v1 to v2 between the path segment x1 and x2, the energy consumption is expressed as: Among them, E acc / dec is the additional energy consumption during the speed change, m is the mass of the UAV, and v1 and v2 are the velocities at the beginning and end of the path segment, respectively.
[0056] Specifically, learning and dynamically adjusting the flight path according to the feedback of the probability map includes the following steps:
[0057] Define the state and action space. The state space S includes: the current position information of the drone (x, y, z), the current flight speed v, the current flight direction θ, and the remaining energy E. remaining , environmental information includes obstacle location, wind speed, terrain height, environmental characteristics Env, and the state is represented by a vector s t =x t ,y t ,z t ,v t ,θ t ,E remaining ,Env t ); the action space A includes speed adjustment, direction adjustment, and height adjustment, and each action is represented by a vector a t =(a v ,a θ ,a h ), a v Indicates adjusting the flight speed, a θ Indicates adjusting the flight direction, a h Indicates adjusting the flight altitude;
[0058] Define the reward function, which includes energy efficiency reward, safety reward, task completion reward, and energy efficiency reward R E (s, a), assuming that the energy consumed by executing action a in state s is E(s, a), then the energy reward can be designed to be a negative value, indicating that the greater the energy consumption, the lower the reward, R E (s,a)=-α E E(s,a), where α E is the weight coefficient of energy consumption, E(s,a) is the energy consumption required to perform action a; safety reward R S (s,a), based on the distance d between the drone and the obstacle obs (s) Design, the closer the distance, the lower the reward. Among them, α S is the safety weight coefficient, ∈ is a small positive number used to prevent division by zero errors; the task completion reward R T (s), when the drone successfully reaches the target point, it is given a high reward R T (s) = β T , indicating that the task is completed; comprehensive reward function, the total reward R(s,a) for the drone to perform action a in state s is expressed as the weighted sum of the above rewards, R(s,a)=R E (s,a)+R S (s,a)+R T (s);
[0059] State transfer and update, the drone is in state s t Next, perform action at After that, based on the state transition equation of the dynamic model and the environment, it enters the next state s t+1 , the state transition model is expressed as: s t+1 =f(s t ,a t )+ω t , where f(s t ,a t ) is the state transfer function of the drone, which describes the state change after executing the action in the current state, ω t is the process noise, which represents the influence of environmental uncertainty;
[0060] Energy consumption calculation, in each state s t Next, perform action a t Calculate the energy consumption E(s t ,a t );
[0061] Strategy updating and optimization, using reinforcement learning to maximize cumulative rewards through strategy updating;
[0062] Through reinforcement learning, the speed, direction and altitude of the drone are autonomously adjusted to achieve the optimal path planning goal.
[0063] Furthermore, the method of using reinforcement learning to maximize the cumulative reward by updating the strategy includes the following steps:
[0064] Initialize the parameters θ, set the learning rate α, discount factor γ, entropy regularization coefficient β, and trust region limit step size δ;
[0065] Sampling generation paths, from strategy π θ (a|s) Sampling generates multiple flight paths, recording states, actions, rewards, and state transitions;
[0066] Calculate the advantage function A(s t ,a t ), using a generalized dominance-based estimation method to reduce variance.
[0067] Update the policy gradient and update the policy parameters under the trust region constraint: J(θ) is the expected cumulative reward;
[0068] Adding entropy regularization term to encourage the exploration of strategies and prevent premature convergence to suboptimal solutions;
[0069] Strategy update and iteration, repeat the above steps until the strategy converges, and finally obtain the optimized path planning strategy;
[0070] The advantage function A(s t ,a t) represents the advantage of the current action compared with the expected action in the strategy, A(s t ,a t )=Q(s t ,a t )-V(s t ), where Q(s t ,a t ) is the state action value function, which means that in state s t Next, perform action a t The cumulative reward, V(s t ) is the state value function, indicating that in state s t The expected cumulative reward;
[0071] The generalized advantage estimate The calculation formula is: Among them, δ t =R(s t ,a t )+γV(s t+1 )-V(s t ) is the temporal difference error, and λ is a parameter that weighs the bias and variance, which is between [0,1].
[0072] Specifically, the continuous monitoring of the drone's energy level and real-time adjustment of the drone's mission and return path include the following steps: During the flight, the drone's energy level is continuously monitored and the path planning is adjusted in real time. When the energy is close to the warning threshold, the system will automatically calculate the shortest path back to the charging station and trigger the drone to return to the charging station when necessary. If the drone's energy is insufficient to complete the current mission, other drones are notified to take over the mission, or the formation is coordinated to adapt to the current energy conditions.
[0073] The beneficial effects of the present invention are as follows: The present invention takes into account the path planning of UAVs under energy constraints, provides full-time and full-domain energy consumption calculation and monitoring, and also adopts a path planning method that combines pre-path planning and real-time path planning, which not only improves the energy utilization efficiency of UAVs when performing tasks, but also enhances their navigation capabilities in complex and uncertain environments. Such a comprehensive solution enables UAVs to complete multiple tasks efficiently and safely, while improving the intelligence level and adaptability of the system, providing a good foundation for more complex UAV application scenarios in the future. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 The overall flow chart of the UAV path planning method under energy constraints is shown. DETAILED DESCRIPTION
[0075] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0076] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0077] like Figure 1 As shown, this embodiment proposes a UAV path planning method under energy constraint conditions, comprising the following steps:
[0078] Step 1: Obtain environment and drone information and use probability maps for real-time updates;
[0079] Step 2: Considering the task requirements and energy constraints, generate multiple possible paths, estimate the energy consumption of each path, and select the optimal initial path;
[0080] Step 3: Based on the optimal initial path, learn and dynamically adjust the flight path according to the feedback from the probability map;
[0081] Step 4: Continuously monitor the drone’s energy level and adjust the drone’s mission or return path in real time.
[0082] The basic idea of this embodiment is: first, assign tasks to each drone, and determine the starting position, target position and related constraints (such as energy, obstacles, etc.). After the drone is started, the initial energy will be recorded and the maximum range allowed by the current energy will be calculated. After the drone is started, the environmental information is obtained through sensors, and the map is updated using the probabilistic map construction method, including the location of obstacles, targets and charging stations. In a dynamic environment, the environmental information is corrected in real time to ensure that the drone can make decisions based on the latest data when planning the path. Considering the task requirements and energy constraints, multiple possible paths are generated, and the energy consumption of each path is estimated. By evaluating the task completion rate and energy consumption of each path, the optimal initial path is selected. Based on the path planning, the reinforcement learning module will dynamically adjust the path according to real-time feedback. Through the reinforcement learning model, the drone can autonomously adjust the flight speed, direction and altitude to improve the energy efficiency and safety of the path. Reinforcement learning gradually improves the path selection through simulation and historical data to ensure better performance in subsequent flight tasks. During the flight, the drone energy level is continuously monitored and the path planning is adjusted in real time. When the energy is close to the warning threshold, the system automatically calculates the shortest path back to the charging station and triggers the drone to return to the charging station if necessary. If the drone does not have enough energy to complete the current mission, it will notify other drones to take over the mission or coordinate the formation to adapt to the current energy conditions.
[0083] The probabilistic map construction method is used to deal with the uncertainty in the environment, and describes the location of obstacles, targets and other important landmarks through a continuously updated probability distribution. Specifically, the real-time update of the probabilistic map includes the following steps:
[0084] Initialize the map. The map is represented by a grid map or a continuous space, where each grid or area contains an initial probability indicating whether the location is occupied;
[0085] Collect surrounding environment data in real time through sensors as the basis for environmental perception;
[0086] Based on the sensor's measurements, the probability of whether there is an obstacle at each location is updated using the Bayesian update rule;
[0087] As the drone moves and environmental information is updated, the probability map is constantly adjusted and updated.
[0088] Specifically, the Bayesian update rule is used to update the probability of whether there is an obstacle at each position, including the following steps: when the drone is flying, the sensor collects the information corresponding to a certain grid m i The relevant measurement value z t , the measured value z t Represents the grid m at time t iWhether it is occupied or not, at this time, the probability of each grid is updated using the Bayesian formula, and the Bayesian update formula is:
[0089]
[0090] Among them, m i represents the position of the i-th grid, s i Indicates the state of the i-th grid, whether it is occupied, P(s i |z t ) indicates that when the sensor observes z t After that, grid m i The posterior probability of being occupied, P(z t |s i ) represents a given grid m i The sensor observes z t The probability of i ) represents the grid m i Prior probability of being occupied, P(z t ) indicates that the sensor observes z t The edge probability is calculated by weighting all possible grid states.
[0091] The probability map provides a clear probability estimate of the occupancy of each location, and uses the updated probability map to plan the optimal path based on the drone’s current location, target location, and energy status, ensuring that the drone can complete its mission safely and efficiently.
[0092] Furthermore, the method of considering the task requirements and energy constraints, generating multiple possible paths, estimating the energy consumption of each path, and selecting the optimal initial path includes the following steps:
[0093] The probabilistic path search algorithm is used to define the state space as a set of nodes in a grid map. The state transition is the process of moving from one grid to the next grid. The goal is to find the path from the initial position x0 to the target position x g Multiple paths of
[0094] The cumulative risk of a path is estimated by the grid occupancy probability in the probability map. The risk R(π) of a path π is expressed as the sum of the occupancy probabilities of all grids passed on the path:
[0095]
[0096] Among them, P(s i =1) indicates grid m i The probability of being occupied by an obstacle, i represents a grid on the path π;
[0097] Calculate the energy consumption of flight distance and speed. The energy consumption per unit time of the drone on the path is expressed as flight power P, which is proportional to the flight speed v and the flight distance d. The energy consumption E(π) of path π is calculated by the following formula:
[0098]
[0099] Among them, P(v i ) represents the flight power of the UAV in the i-th segment of the path, which depends on the flight speed v i , represents the flight time of the i-th segment, d i is the flight distance between two locations on the path, and n represents the number of segments of the path;
[0100] Calculate the energy consumption of terrain and wind speed. When considering the influence of terrain, the energy consumption is calculated by adding the vertical height change h caused by the terrain. i To make corrections:
[0101]
[0102] Where γ is a constant related to the weight of the drone, h i represents the height difference on the i-th path, and the wind speed w i The effect of is taken into account by correcting the flight power:
[0103] P(v i ,w i )=P(v i )·(1+αw i )
[0104] Among them, α is the coefficient of wind speed on flight power, w i represents the wind speed of the i-th path;
[0105] When the energy consumption E(π) and risk R(π) of multiple paths are calculated, an optimization algorithm is used to select the optimal path, and a comprehensive objective function J(π) is defined to evaluate the pros and cons of each path:
[0106] J(π)=λ1E(π)+λ2R(π)
[0107] Among them, λ1 and λ2 are the weight coefficients of energy and risk, respectively, reflecting the priority of the task for different goals. By adjusting the weights, the system finds a balance between energy saving and risk avoidance.
[0108] By generating multiple potential paths and estimating the energy consumption of each path, the UAV can effectively select the optimal path in a complex environment by considering the probability map, mission requirements and energy constraints. Combining path risk assessment with multi-objective optimization ensures that the UAV can complete the mission efficiently while avoiding unnecessary energy consumption and risks.
[0109] This embodiment proposes a new idea to generate multiple potential paths for the drone and estimate the energy consumption and risk of each path. Then, through a multi-objective optimization method, an optimal path with minimal energy consumption and path safety is selected. This method is suitable for complex and uncertain environments and can ensure that the drone successfully completes its mission under limited energy conditions.
[0110] Preferably, the method of considering the task requirements and energy constraints, generating multiple possible paths, estimating the energy consumption of each path, and selecting the optimal initial path includes the following steps:
[0111] Step 201, initialize a tree structure T = (V, E), where V is a node set, initially including the starting point V = {x start}, E is the set of edges, initially empty; set the initial position of the drone to x start and the target position is x goal , construct a probability map Each position m in the map i The probability of including obstacles P(s i ), which is used in the path evaluation process to determine whether a certain area can be safely passed through and to initialize the energy constraint condition E max , represents the maximum available energy of the UAV;
[0112] Step 202, randomly sample a point x from the space rand , the point is within the range of the configuration space, which is defined as the flight area of the UAV;
[0113] Step 203, find the distance x in the current tree T rand The nearest node x nearest , the distance calculation formula is as follows: ‖·‖ represents the Euclidean distance;
[0114] Step 204, from x nearest To x rand Expand and generate a new node x new , the new node is at x nearest to x rand In the direction of , and expand with a step size of Δd:
[0115]
[0116] The new node x after expansion new Add the node set V of the tree and add the edge (x nearest ,x new )Add edge set E;
[0117] Step 205, when expanding the tree, check nearest to x new Whether the path passes through the obstacle area in the probability map, calculate the risk of the path, the risk R(x nearest ,x new ) is estimated by the occupancy probability of the grid on the path:
[0118]
[0119] Among them, π(x nearest ,x new ) indicates that from x nearest to x new All grids passed by on the path, P(s i =1) is the grid s i The probability of being occupied by obstacles; if the cumulative risk of the path exceeds the preset risk threshold, the extension is discarded and x is not new Add to the tree;
[0120] Step 206: for each newly generated path segment, consider the flight distance and speed and evaluate the energy consumption. The energy consumption E(x nearest ,x new ) is:
[0121]
[0122] Where P(v) is the flight power corresponding to the flight speed v, ‖x new -x nearest ‖ is the distance of the path segment, v is the flight speed, if the cumulative energy consumption of the path exceeds the maximum energy E of the drone max , then stop expanding the path;
[0123] Step 207, repeat steps 202 to 206, and continue to expand the tree until the connection starting point x is found start and the target point x goal A feasible path of x, or the preset maximum number of iterations is reached; multiple potential paths are generated, each path corresponds to start to x goal The energy cost and risk of each path on a tree are calculated and recorded during the expansion process.
[0124] Step 208, among the generated multiple paths, select the best path that meets the energy constraint, and select the path through an optimization function, which comprehensively considers energy consumption E(π) and risk R(π):
[0125] J(π)=λ1E(π)+λ2R(π)
[0126] Among them, E(π) represents the total energy consumption of path π, R(π) represents the risk of the path, λ1,λ2 are the weight coefficients of energy consumption and risk, which are adjusted according to task requirements.
[0127] Preferably, in step 206, considering the influence of terrain height change h and wind speed w, the energy consumption of the path segment is corrected to:
[0128]
[0129] Among them: γ1 is the coefficient of height change on energy consumption; α is the coefficient of wind speed on energy consumption.
[0130] Furthermore, when the energy consumption is considered in terms of flight distance and speed, the flight power P(v) varies with speed, and the calculation formula is: P(v) = c1v 3 +c2v 2 +c3v+c4, where c1, c2, c3, c4 are constants related to the drone model, c1v 3 represents aerodynamic drag, c2v 2 represents propulsion power, c3v represents mechanical loss, and c4 represents fixed power loss.
[0131] Furthermore, the energy consumption also takes into account the energy consumption of attitude adjustment, including the energy consumption of steering and the energy consumption of acceleration and deceleration. During the steering process of the UAV in the path, due to the effect of inertia, the additional energy consumption is related to the steering angle θ: E turn (θ) = k turn ·θ, where E turn (θ) represents the energy consumption when the steering angle is θ, k turn is a constant related to the UAV model and flight dynamics, which represents the energy consumption coefficient of steering; if the UAV needs to accelerate or decelerate on the path, additional power is required to overcome the change in speed. Assuming that the speed of the UAV changes from v1 to v2 between the path segment x1 and x2, the energy consumption is expressed as: Among them, E acc / dec is the additional energy consumption during the speed change, m is the mass of the UAV, and v1 and v2 are the velocities at the beginning and end of the path segment, respectively.
[0132] In this embodiment, reinforcement learning-driven path optimization is an intelligent method that enables drones to autonomously learn the optimal path in a dynamic and uncertain environment. Through reinforcement learning, drones can adjust their flight speed, direction, and altitude based on environmental feedback to improve the energy efficiency and safety of the path.
[0133] Specifically, learning and dynamically adjusting the flight path according to the feedback of the probability map includes the following steps:
[0134] Define the state and action space. The state space S includes: the current position information of the drone (x, y, z), the current flight speed v, the current flight direction θ, and the remaining energy E. remaining , environmental information includes obstacle location, wind speed, terrain height, environmental characteristics Env, and the state is represented by a vector s t =x t ,y t ,z t ,v t ,θ t ,E remaining ,Env t ); the action space A includes speed adjustment, direction adjustment, and height adjustment, and each action is represented by a vector a t =(a v ,a θ ,a h ), a v Indicates adjusting the flight speed, a θ Indicates adjusting the flight direction, a h Indicates adjusting the flight altitude;
[0135] Define the reward function, which includes energy efficiency reward, safety reward, task completion reward, and energy efficiency reward R E (s, a), assuming that the energy consumed by executing action a in state s is E(s, a), then the energy reward can be designed to be a negative value, indicating that the greater the energy consumption, the lower the reward, R E (s,a)=-α E E(s,a), where α E is the weight coefficient of energy consumption, E(s,a) is the energy consumption required to perform action a; safety reward R S (s,a), based on the distance d between the drone and the obstacle obs (s) Design, the closer the distance, the lower the reward. Among them, α S is the safety weight coefficient, ∈ is a small positive number used to prevent division by zero errors; the task completion reward R T (s), when the drone successfully reaches the target point, it is given a high reward R T (s) = βT , indicating that the task is completed; comprehensive reward function, the total reward R(s,a) for the drone to perform action a in state s is expressed as the weighted sum of the above rewards, R(s,a)=R E (s,a)+R S (s,a)+R T (s);
[0136] State transfer and update, the drone is in state s t Next, perform action a t After that, based on the state transition equation of the dynamic model and the environment, it enters the next state s t+1 , the state transition model is expressed as: s t+1 =f(s t ,a t )+ω t , where f(s t ,a t ) is the state transfer function of the drone, which describes the state change after executing the action in the current state, ω t is the process noise, which represents the influence of environmental uncertainty;
[0137] Energy consumption calculation, in each state s t Next, perform action a t Calculate the energy consumption E(s t ,a t );
[0138] Strategy updating and optimization, using reinforcement learning to maximize cumulative rewards through strategy updating;
[0139] Through reinforcement learning, the speed, direction and altitude of the drone are autonomously adjusted to achieve the optimal path planning goal.
[0140] In reinforcement learning-driven path optimization, the improved policy gradient method (policy-based reinforcement learning) can improve computational efficiency and reduce energy consumption, including the balance between reducing sample complexity, improving learning efficiency, and optimizing the exploration and exploitation of strategies.
[0141] Furthermore, the method of using reinforcement learning to maximize the cumulative reward by updating the strategy includes the following steps:
[0142] Initialize the parameters θ, set the learning rate α, discount factor γ, entropy regularization coefficient β, and trust region limit step size δ;
[0143] Sampling generation paths, from strategy π θ (a|s) Sampling generates multiple flight paths, recording states, actions, rewards, and state transitions;
[0144] Calculate the advantage function A(s t ,a t ), using a generalized dominance-based estimation method to reduce variance.
[0145] Update the policy gradient and update the policy parameters under the trust region constraint: J(θ) is the expected cumulative reward;
[0146] Adding entropy regularization term to encourage the exploration of strategies and prevent premature convergence to suboptimal solutions;
[0147] Strategy update and iteration, repeat the above steps until the strategy converges, and finally obtain the optimized path planning strategy;
[0148] The advantage function A(s t ,a t ) represents the advantage of the current action compared with the expected action in the strategy, A(s t ,a t )=Q(s t ,a t )-V(s t ), where Q(s t ,a t ) is the state action value function, which means that in state s t Next, perform action a t The cumulative reward, V(s t ) is the state value function, indicating that in state s t The expected cumulative reward;
[0149] The generalized advantage estimate The calculation formula is: Among them, δ t =R(s t ,a t )+γV(s t+1 )-V(s t ) is the temporal difference error, and λ is a parameter that weighs the bias and variance, which is between [0,1].
[0150] The traditional policy gradient method directly uses reward signals for optimization, but this method has high sample complexity and is easily affected by noise. The improved method introduces an advantage function to replace the direct use of cumulative reward signals, thereby reducing the computational burden and improving learning efficiency.
[0151] Specifically, the continuous monitoring of the drone's energy level and real-time adjustment of the drone's mission and return path include the following steps: During the flight, the drone's energy level is continuously monitored and the path planning is adjusted in real time. When the energy is close to the warning threshold, the system will automatically calculate the shortest path back to the charging station and trigger the drone to return to the charging station when necessary. If the drone's energy is insufficient to complete the current mission, other drones are notified to take over the mission, or the formation is coordinated to adapt to the current energy conditions.
[0152] The advantages and beneficial effects of the present invention are: real-time monitoring of the remaining power of the drone, and optimizing the flight path according to energy consumption and mission requirements, thereby maximizing the possibility of mission completion, which ensures that the drone can effectively utilize every bit of energy when performing the mission and avoid unnecessary energy waste; dynamic energy evaluation: by evaluating the energy state of the drone in real time, its endurance can be predicted to ensure that the flight will not be interrupted due to insufficient energy during mission execution. Path planning can make effective decisions in the face of environmental uncertainties, such as dynamic obstacles, target position changes, etc., which enables drones to navigate safely and effectively in complex and unstable environments; by considering multiple factors such as energy consumption, mission completion probability and collision risk, the path with the highest expected return can be calculated, thereby improving the efficiency and success rate of mission completion; according to real-time environmental feedback, the planned path is dynamically adjusted to ensure that the drone can quickly respond to changes and maintain mission objectives during flight. The deep reinforcement learning algorithm enables the drone to autonomously optimize its flight strategy and improve the intelligence level of path planning by continuously learning historical data and real-time feedback. This capability enables drones to make more efficient decisions when facing complex environments and changing tasks.
[0153] As used herein, the word "preferred" is intended to be used as an example, instance, or illustration. Any aspect or design described herein as "preferred" is not necessarily to be construed as being more advantageous than other aspects or designs. On the contrary, the use of the word "preferred" is intended to present concepts in a specific way. The term "or" as used in this application is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" means any one of the naturally included permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.
[0154] Moreover, although the present disclosure has been shown and described with respect to one or implementations, those skilled in the art will think of equivalent variations and modifications based on the reading and understanding of this specification and the accompanying drawings. The present disclosure includes all such modifications and variations, and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-mentioned components (such as elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (such as it is functionally equivalent), even if the structure is not equivalent to the disclosed structure of the function in the exemplary implementation of the present disclosure shown herein. In addition, although the specific features of the present disclosure have been disclosed with respect to only one of several implementations, such features can be combined with one or other features of other implementations that may be desired and advantageous for a given or specific application. Moreover, insofar as the terms "including", "having", "containing" or their variations are used in specific embodiments or claims, such terms are intended to be included in a manner similar to the term "comprising".
[0155] The functional units in the embodiments of the present invention may be integrated into a processing module, or each unit may exist physically separately, or multiple or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc. The above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.
[0156] To sum up, the above embodiment is an implementation mode of the present invention, but the implementation mode of the present invention is not limited by the embodiment, and any other changes, modifications, substitutions, combinations, and simplifications that deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A method for UAV path planning under energy constraints, characterized in that: The steps include: Step 1: Obtain environment and drone information and use probability maps for real-time updates; Step 2: Considering the task requirements and energy constraints, generate multiple possible paths, estimate the energy consumption of each path, and select the optimal initial path; Step 3: Based on the optimal initial path, learn and dynamically adjust the flight path according to the feedback from the probability map; Step 4: Continuously monitor the drone’s energy level and adjust the drone’s mission or return path in real time; The real-time updating using the probability map comprises the following steps: Initialize the map. The map is represented by a grid map or a continuous space, where each grid or area contains an initial probability indicating whether the location is occupied; Collect surrounding environment data in real time through sensors as the basis for environmental perception; Based on the sensor's measurements, the probability of whether there is an obstacle at each location is updated using the Bayesian update rule; As the drone moves and environmental information is updated, the probability map is constantly adjusted and updated.
2. The method for UAV path planning under energy constraints according to claim 1 is characterized in that: The Bayesian update rule is used to update the probability of whether there is an obstacle at each position, including the following steps: when the drone is flying, the sensor collects the information corresponding to a certain grid m i The relevant measurement value z t , the measured value z t Represents the grid m at time t i Whether it is occupied or not, at this time, the probability of each grid is updated using the Bayesian formula, and the Bayesian update formula is: Among them, m i represents the position of the i-th grid, s i Indicates the state of the i-th grid, whether it is occupied, P(s i |z t ) indicates that when the sensor observes z t After that, grid m i The posterior probability of being occupied, P(z t |s i ) represents a given grid m i The sensor observes z t The probability of i ) represents the grid m i Prior probability of being occupied, P(z t ) indicates that the sensor observes z t The edge probability is calculated by weighting all possible grid states.
3. The method for UAV path planning under energy constraints according to claim 1 is characterized in that: The method of considering the task requirements and energy constraints, generating multiple possible paths, estimating the energy consumption of each path, and selecting the optimal initial path includes the following steps: The probabilistic path search algorithm is used to define the state space as a set of nodes in a grid map. The state transition is the process of moving from one grid to the next grid. The goal is to find the path from the initial position x0 to the target position x g Multiple paths of The cumulative risk of a path is estimated by the grid occupancy probability in the probability map. The risk R(π) of a path π is expressed as the sum of the occupancy probabilities of all grids passed on the path: Among them, P(s i =1) indicates grid m i The probability of being occupied by an obstacle, i represents a grid on the path π; Calculate the energy consumption of flight distance and speed. The energy consumption per unit time of the drone on the path is expressed as flight power P, which is proportional to the flight speed v and the flight distance d. The energy consumption E(π) of path π is calculated by the following formula: Among them, P(v i ) represents the flight power of the UAV in the i-th segment of the path, which depends on the flight speed v i , represents the flight time of the i-th segment, d i is the flight distance between two locations on the path, and n represents the number of segments of the path; Calculate the energy consumption of terrain and wind speed. When considering the influence of terrain, the energy consumption is calculated by adding the vertical height change h caused by the terrain. i To make corrections: Where γ is a constant related to the weight of the drone, h i represents the height difference on the i-th path, and the wind speed w i The effect of is taken into account by correcting the flight power: P(v i ,w i )=P(v i )·(1+αw i ) Among them, α is the coefficient of wind speed on flight power, w i represents the wind speed of the i-th path; When the energy consumption E(π) and risk R(π) of multiple paths are calculated, an optimization algorithm is used to select the optimal path, and a comprehensive objective function J(π) is defined to evaluate the pros and cons of each path: J(π)=λ1E(π)+λ2R(π) Among them, λ1 and λ2 are the weight coefficients of energy and risk, respectively, reflecting the priority of the task for different goals. By adjusting the weights, the system finds a balance between energy saving and risk avoidance.
4. The method for UAV path planning under energy constraints according to claim 1, characterized in that: The method of considering the task requirements and energy constraints, generating multiple possible paths, estimating the energy consumption of each path, and selecting the optimal initial path includes the following steps: Step 201, initialize a tree structure T = (V, E), where V is a node set, initially including the starting point V = {x start }, E is the set of edges, initially empty; set the initial position of the drone to x start and the target position is x goal , construct a probability map Each position m in the map i The probability of including obstacles P(s i ), which is used in the path evaluation process to determine whether a certain area can be safely passed through and to initialize the energy constraint condition E max , represents the maximum available energy of the UAV; Step 202, randomly sample a point x from the space rand , the point is within the range of the configuration space, which is defined as the flight area of the UAV; Step 203, find the distance x in the current tree T rand The nearest node x nearest , the distance calculation formula is as follows: ‖·‖ represents the Euclidean distance; Step 204, from x nearest To x rand Expand and generate a new node x new , the new node is at x nearest to x rand In the direction of , and expand with a step size of Δd: The new node x after expansion new Add the node set V of the tree and add the edge (x nearest ,x new )Add edge set E; Step 205, when expanding the tree, check nearest to x new Whether the path passes through the obstacle area in the probability map, calculate the risk of the path, the risk R(x nearest ,x new ) is estimated by the occupancy probability of the grid on the path: Among them, π(x nearest ,x new ) indicates that from x nearest to x new All grids passed by on the path, P(s i =1) is the grid s i The probability of being occupied by obstacles; if the cumulative risk of the path exceeds the preset risk threshold, the extension is discarded and x is not new Add to the tree; Step 206: for each newly generated path segment, consider the flight distance and speed and evaluate the energy consumption. The energy consumption E(x nearest ,x new ) is: Where P(v) is the flight power corresponding to the flight speed v, ‖x new -x nearest ‖ is the distance of the path segment, v is the flight speed, if the cumulative energy consumption of the path exceeds the maximum energy E of the drone max , then stop expanding the path; Step 207, repeat steps 202 to 206, and continue to expand the tree until the connection starting point x is found start and the target point x goal A feasible path of x, or the preset maximum number of iterations is reached; multiple potential paths are generated, each path corresponds to start to x goal The energy cost and risk of each path on a tree are calculated and recorded during the expansion process. Step 208, among the generated multiple paths, select the best path that meets the energy constraint, and select the path through an optimization function, which comprehensively considers energy consumption E(π) and risk R(π): J(π)=λ1E(π)+λ2R(π) Among them, E(π) represents the total energy consumption of path π, R(π) represents the risk of the path, λ1,λ2 are the weight coefficients of energy consumption and risk, which are adjusted according to task requirements.
5. The method for UAV path planning under energy constraints according to claim 4 is characterized in that: In step 206, considering the influence of terrain height change h and wind speed w, the energy consumption of the path segment is corrected to: Among them: γ1 is the coefficient of height change on energy consumption; α is the coefficient of wind speed on energy consumption.
6. A method for UAV path planning under energy constraints according to claim 3 or 4, characterized in that: When considering the flight distance and speed, the energy consumption is calculated as follows: P(v) = c1v 3 +c2v 2 +c3v+c4, where c1, c2, c3, c4 are constants related to the drone model, c1v 3 represents aerodynamic drag, c2v 2 represents propulsion power, c3v represents mechanical loss, and c4 represents fixed power loss.
7. The method for UAV path planning under energy constraints according to claim 6 is characterized in that: The energy consumption also takes into account the energy consumption of attitude adjustment, including steering energy consumption and energy consumption of acceleration and deceleration. During the steering process of the UAV in the path, due to the effect of inertia, the additional energy consumption is related to the steering angle θ: E turn (θ) = k turn ·θ, where E turn (θ) represents the energy consumption when the steering angle is θ, k turn is a constant related to the UAV model and flight dynamics, which represents the energy consumption coefficient of steering; if the UAV needs to accelerate or decelerate on the path, additional power is required to overcome the change in speed. Assuming that the speed of the UAV changes from v1 to v2 between the path segment x1 and x2, the energy consumption is expressed as: Among them, E acc / dec is the additional energy consumption during the speed change, m is the mass of the UAV, and v1 and v2 are the velocities at the beginning and end of the path segment, respectively.
8. The method for UAV path planning under energy constraints according to claim 7, characterized in that: The method of learning and dynamically adjusting the flight path based on the feedback of the probability map includes the following steps: Define the state and action space. The state space S includes: the current position information of the drone (x, y, z), the current flight speed v, the current flight direction θ, and the remaining energy E. remaining , environmental information includes obstacle location, wind speed, terrain height, environmental characteristics Env, and the state is represented by a vector s t =(x t ,y t ,z t ,v t ,θ t ,E remaining ,Env t ); the action space A includes speed adjustment, direction adjustment, and height adjustment, and each action is represented by a vector a t =(a v ,a θ ,a h ), a v Indicates adjusting the flight speed, a θ Indicates adjusting the flight direction, a h Indicates adjusting the flight altitude; Define the reward function, which includes energy efficiency reward, safety reward, task completion reward, and energy efficiency reward R E (s, a), assuming that the energy consumed by executing action a in state s is E(s, a), then the energy reward can be designed to be a negative value, indicating that the greater the energy consumption, the lower the reward, R E (s,a)=-α E E(s,a), where α E is the weight coefficient of energy consumption, E(s,a) is the energy consumption required to perform action a; safety reward R S (s,a), based on the distance d between the drone and the obstacle obs (s) Design, the closer the distance, the lower the reward. Among them, α S is the safety weight coefficient, ∈ is a small positive number used to prevent division by zero errors; the task completion reward R T (s), when the drone successfully reaches the target point, it is given a high reward R T (s) = β T , indicating that the task is completed; comprehensive reward function, the total reward R(s,a) for the drone to perform action a in state s is expressed as the weighted sum of the above rewards, R(s,a)=R E (s,a)+R S (s,a)+R T (s); State transfer and update, the drone is in state s t Next, perform action a t After that, based on the state transition equation of the dynamic model and the environment, it enters the next state s t+1 , the state transition model is expressed as: s t+1 =f(s t ,a t )+ω t , where f(s t ,a t ) is the state transfer function of the drone, which describes the state change after executing the action in the current state, ω t is the process noise, which represents the influence of environmental uncertainty; Energy consumption calculation, in each state s t Next, perform action a t Calculate the energy consumption E(s t ,a t ); Strategy updating and optimization, using reinforcement learning to maximize cumulative rewards through strategy updating; Through reinforcement learning, the speed, direction and altitude of the drone are autonomously adjusted to achieve the optimal path planning goal.
9. The method for UAV path planning under energy constraints according to claim 8, characterized in that: The method of using reinforcement learning to maximize the cumulative reward through strategy updating includes the following steps: Initialize the parameters θ, set the learning rate α, discount factor γ, entropy regularization coefficient β, and trust region limit step size δ; Sampling generation paths, from strategy π θ (a|s) Sampling generates multiple flight paths, recording states, actions, rewards, and state transitions; Calculate the advantage function A(s t ,a t ), using a generalized dominance-based estimation method to reduce variance. Update the policy gradient and update the policy parameters under the trust region constraint: J(θ) is the expected cumulative reward; Adding entropy regularization term to encourage the exploration of strategies and prevent premature convergence to suboptimal solutions; Strategy update and iteration, repeat the above steps until the strategy converges, and finally obtain the optimized path planning strategy; The advantage function A(s t ,a t ) represents the advantage of the current action compared with the expected action in the strategy, A(s t ,a t )=Q(s t ,a t )-V(s t ), where Q(s t ,a t ) is the state action value function, which means that in state s t Next, perform action a t The cumulative reward, V(s t ) is the state value function, indicating that in state s t The expected cumulative reward; The generalized advantage estimate The calculation formula is: Among them, δ t is the temporal difference error, and λ is a parameter that weighs the bias and variance, between [0,1].
10. The method for UAV path planning under energy constraints according to claim 1, characterized in that: The continuous monitoring of the drone energy level and real-time adjustment of the drone's mission and return path include the following steps: during flight, the drone energy level is continuously monitored and the path planning is adjusted in real time. When the energy is close to the warning threshold, the shortest path back to the charging station is calculated, and the drone is triggered to return to the charging station when necessary. If the drone energy is insufficient to complete the current mission, other drones are notified to take over the mission, or the formation is coordinated to adapt to the current energy conditions.
Citation Information
Patent Citations
Multi-UAV (unmanned aerial vehicle) cooperative searching method and system based on path planning and information fusion
CN107844129A
Energy consumption optimal underwater area coverage method based on bilevel programming framework under ocean current influence
CN115655274A
Multi-agent unmanned aerial vehicle search task energy optimization method based on entropy maximization strategy
CN119151062A
System for marking lesion location
KR1020240166655A
Cited By
Autonomous navigation and obstacle avoidance system of fire-fighting robot
CN120949764A
Path planning method and system for target airframe
CN120972957A