Sectional type unmanned aerial vehicle swarm path planning method and device

Through the segmented path planning method, combined with the A* algorithm, virtual rigid body algorithm and dynamic obstacle avoidance mechanism, the problem of low path planning efficiency of drone swarms in dynamic environments is solved, and efficient dynamic adaptation and stable flight in multi-objective tasks are achieved.

CN120293138APending Publication Date: 2025-07-11SUN YAT SEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510381548.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing drone swarm path planning algorithm has low computational efficiency in dynamic environments, is difficult to cope with environmental changes in real time, and is difficult to achieve performance balance in multi-objective tasks.

Method used

The segmented path planning method is adopted, combined with the A* algorithm and the virtual rigid body algorithm to generate the initial global path, determine the path segments that need to be optimized through path re-planning, use the flexible action-evaluation algorithm to perform local optimization, and adjust the path through the dynamic obstacle avoidance mechanism to improve dynamic adaptability.

Benefits of technology

The dynamic adaptability and computing efficiency of drone swarm path planning are improved, and comprehensive optimization of path quality, formation stability, energy efficiency and task completion in multi-objective tasks can be achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120293138A_ABST
    Figure CN120293138A_ABST
Patent Text Reader

Abstract

The invention discloses a segmented unmanned aerial vehicle swarm path planning method and device, and the method comprises the steps: obtaining a first global path of an unmanned aerial vehicle swarm through an A * algorithm and a virtual rigid body algorithm; according to a path re-planning judgment index, judging whether a path section needing to be re-planned exists in the first global path or not; if the path segment needing to be replanned does not exist, taking the first global path as a final global path of the unmanned aerial vehicle swarm; if the path segment needing to be replanned exists, marking the path segment needing to be replanned; optimizing the path segment needing to be replanned through a flexible action-evaluation algorithm to obtain a second global path of the unmanned aerial vehicle swarm; and adjusting the second global path according to a dynamic obstacle avoidance mechanism, and taking the adjusted second global path as a final global path of the unmanned aerial vehicle swarm. Compared with the prior art, the method can improve the dynamic adaptability and calculation efficiency of unmanned aerial vehicle swarm path planning in a multi-target task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drone swarms, and particularly to a segmented drone swarm path planning method and device. Background Art

[0002] In the technical field of drone swarms, path planning is the core issue for realizing the flight control of drone swarms. Traditional global path planning methods (such as the A* algorithm and the Dijkstra algorithm) perform well in static environments, but have low computational efficiency in dynamic environments and are difficult to respond to environmental changes in real time. At the same time, local path planning algorithms based on deep reinforcement learning (such as the DDPG algorithm) although improve the dynamic adaptability, it is difficult to achieve performance balance in the face of complex multi-objective tasks (such as path quality, energy efficiency, team stability, and task completion), and the computational efficiency is low, making it difficult to meet the real-time requirements. In addition, the virtual rigid body algorithm (Virtual Rigid Formation, VRF), as an effective drone swarm control method, can maintain the formation stability through geometric constraints, but its adaptability in a dynamic obstacle environment is weak, and it is difficult to cope with the obstacles and path adjustment requirements in dynamic scenarios. Therefore, it is particularly important to improve the dynamic adaptability and computational efficiency of the drone swarm path planning algorithm in multi-objective tasks. Summary of the Invention

[0003] The present invention provides a segmented drone swarm path planning method and device, which can improve the dynamic adaptability and computational efficiency of the drone swarm path planning in multi-objective tasks.

[0004] In a first aspect, an embodiment of the present invention provides a segmented drone swarm path planning method, including:

[0005] Obtaining a first global path of the drone swarm through the A* algorithm and the virtual rigid body algorithm;

[0006] Judging whether there is a path segment that needs to be replanned in the first global path according to the path replanning judgment index; if there is no path segment that needs to be replanned, taking the first global path as the final global path of the drone swarm; if there is a path segment that needs to be replanned, marking the path segment that needs to be replanned;

[0007] Optimizing the path segment that needs to be replanned through a flexible action-evaluation algorithm to obtain a second global path of the drone swarm;

[0008] Adjusting the second global path according to the dynamic obstacle avoidance mechanism, and taking the adjusted second global path as the final global path of the drone swarm.

[0009] In the embodiments of the present invention, the A* algorithm and the virtual rigid body algorithm are used to obtain the first global path of the UAV swarm, providing a stable overall planning direction for the UAV swarm and maintaining the formation stability of the UAV swarm; according to the path replanning determination index, it is determined whether there is a path segment that needs to be replanned in the first global path, providing a judgment basis and data foundation for whether local optimization is needed and which path segment needs to be optimized in the follow-up; through the flexible action-evaluation algorithm, the path segment that needs to be replanned is optimized to obtain the second global path of the UAV swarm, improving the dynamic adaptability of path planning while comprehensively considering the multi-objective optimization requirements of path quality, formation stability, energy efficiency, and task completion; according to the dynamic obstacle avoidance mechanism, the second global path is adjusted, and the adjusted second global path is used as the final global path of the UAV swarm, improving the adaptability of path planning in a dynamic obstacle environment. Compared with the prior art, the present application segments the result of global path planning, dynamically optimizes the local path, and adjusts the optimization threshold according to different task requirements, capable of improving the dynamic adaptation ability and computational efficiency of the UAV swarm path planning in multi-objective tasks.

[0010] Further, the obtaining of the first global path of the UAV swarm through the A* algorithm and the virtual rigid body algorithm is specifically as follows:

[0011] Through the A* algorithm, the global path of the UAV leader is obtained;

[0012] Through the virtual rigid body algorithm, the relative positions between the UAV leader and all UAV followers are obtained;

[0013] According to the global path of the UAV leader and the relative positions between the UAV leader and all UAV followers, the first global path of the UAV swarm is obtained.

[0014] In the embodiments of the present invention, the A* algorithm is combined with the virtual rigid body algorithm to provide a stable overall planning direction for the UAV swarm and maintain the formation stability of the UAV swarm.

[0015] Further, in the described segmented UAV swarm path planning method, the obtaining of the global path of the UAV leader through the A* algorithm is specifically as follows:

[0016] The task area is modeled as a grid map;

[0017] According to the preset step size and movement cost, the UAV leader is controlled to perform grid expansion;

[0018] According to the cost function and B-spline curve interpolation, the global path of the UAV leader is obtained; where the cost function is defined as:

[0019] f(n) = g(n) + h(n);

[0020] Wherein, f(n) is the total path cost of the UAV leader from the starting point to grid n; g(n) is the actual path cost of the UAV leader from the starting point to grid n; h(n) is the heuristic estimated path cost of the UAV leader from the starting point to grid n.

[0021] In the embodiment of the present invention, the global path of the UAV leader is generated by the A* algorithm to provide an overall reference path for the UAV swarm.

[0022] Further, the relative positions between the UAV leader and all UAV followers are obtained by the virtual rigid body algorithm, specifically:

[0023] The relative positions between the UAV leader and all UAV followers are maintained by geometric constraints; wherein, the geometric constraints are defined as:

[0024]

[0025] Wherein, (x l , y l ) is the position of the UAV leader; θ l is the heading angle of the UAV leader; (x i , y i ) is the position of the i-th UAV follower; d i is the relative distance of the i-th UAV follower; β i is the angle of the i-th UAV follower; R(θ l ) is the rotation matrix;

[0026]

[0027] Wherein, v l is the speed of the UAV leader; ω l is the angular velocity of the UAV leader;

[0028]

[0029] In the embodiment of the present invention, the relative positions between UAVs are maintained by the virtual rigid body algorithm to ensure the formation stability of the UAV swarm.

[0030] Further, it is determined whether there is a path segment to be replanned in the first global path according to the path replanning determination index, specifically:

[0031] Calculate the perpendicular distance from the current path segment to the shortest path. If the perpendicular distance exceeds a preset threshold, mark the current path segment as a path segment to be replanned; wherein, the calculation formula of the perpendicular distance is as follows:

[0032]

[0033] Among them, D ⊥ is the perpendicular distance; (x1, y1) and (x2, y2) are two points on the shortest path; (x0, y0) is a point on the current path segment;

[0034] Calculate the energy consumption of the UAV swarm in the current path segment. If the energy consumption exceeds the preset threshold, mark the current path segment as a path segment that needs to be replanned; among them, the calculation formula of the energy consumption is as follows:

[0035]

[0036] Among them, E is the energy consumption; v i is the speed of the UAV swarm in the i-th path segment; a i is the acceleration of the UAV swarm in the i-th path segment; α and β are the weights of the speed and acceleration of the UAV swarm in the i-th path segment respectively;

[0037] Calculate the formation circle radius of the UAV swarm in the current path segment. If the formation circle radius exceeds the preset threshold, mark the current path segment as a path segment that needs to be replanned; among them, the calculation formula of the formation circle radius is as follows:

[0038]

[0039] Among them, R formation is the formation circle radius; N is the total number of UAVs; D i is the distance between the i-th UAV and the formation center.

[0040] By comprehensively considering the influence of the perpendicular length of the shortest path, energy consumption and formation circle radius indicators, the embodiment of the present invention segments the result of the global path planning, providing a judgment basis and data basis for whether local optimization is needed and which path segments need to be optimized in the follow-up.

[0041] Further, through the flexible action-evaluation algorithm, the path segment that needs to be replanned is optimized to obtain the second global path of the UAV swarm, specifically:

[0042] Determine the state space and action space of the UAV swarm;

[0043] According to the state space, action space and the preset reward function, optimize the path segment that needs to be replanned; among them, the reward function is defined according to the multi-objective evaluation index, and the multi-objective evaluation index includes path quality, formation stability, energy efficiency and task completion degree;

[0044] Replace the corresponding path segment in the first global path with the optimized path segment to obtain the second global path of the UAV swarm.

[0045] In the embodiment of the present invention, through the flexible action-evaluation algorithm, the path segment that needs to be replanned is optimized and adjusted to improve the dynamic adaptability of path planning.

[0046] Further, optimizing the path segment that needs to be replanned according to the state space, action space, and a preset reward function specifically includes:

[0047] Obtain the motion strategy range of the UAV swarm through the state space and action space;

[0048] The flexible action-evaluation algorithm optimizes the path segment that needs to be replanned by maximizing the weighted sum of the expected reward and the policy entropy; wherein, the flexible action-evaluation algorithm includes prioritized experience replay and soft update of the target network.

[0049] In the embodiment of the present invention, by maximizing the weighted sum of the expected reward and the policy entropy, the multi-objective optimization requirements of path quality, formation stability, energy efficiency, and task completion are comprehensively considered.

[0050] Further, adjusting the second global path according to the dynamic obstacle avoidance mechanism and using the adjusted second global path as the final global path of the UAV swarm specifically includes:

[0051] Obtain the synthetic control force between the UAV swarm and the obstacle through the artificial potential field method;

[0052] Adjust the second global path of the UAV swarm through the synthetic control force to obtain the final global path of the UAV swarm.

[0053] In the embodiment of the present invention, obstacle avoidance is carried out by combining the artificial potential field method with path adjustment to improve the adaptability of path planning in a dynamic obstacle environment.

[0054] Further, obtaining the synthetic control force between the UAV swarm and the obstacle through the artificial potential field method specifically includes:

[0055] F att =-ζ(q - q next );

[0056] Wherein, F att is the gravitational potential field; ζ is the gravitational strength coefficient; q is the current position of the UAV swarm; q next is the next target point on the second global path;

[0057]

[0058] Wherein, Frep is a repulsive potential field; η is the potential field strength coefficient; ρ is the distance between the UAV swarm and the obstacle; ρ0 is the influence range;

[0059] F = F att + F rep ;

[0060] where F is the synthetic control force.

[0061] In the embodiment of the present invention, by generating the synthetic control force of the gravitational potential field and the repulsive potential field, the UAV swarm is controlled to avoid obstacles.

[0062] In a second aspect, the embodiment of the present invention provides a segmented UAV swarm path planning device, including a global path planning module, a path replanning determination module, a local path optimization module, and a dynamic obstacle avoidance module:

[0063] The global path planning module is used to obtain the first global path of the UAV swarm through the A* algorithm and the virtual rigid body algorithm;

[0064] The path replanning determination module is used to determine whether there is a path segment that needs to be replanned in the first global path according to the path replanning determination index; if there is no path segment that needs to be replanned, the first global path is used as the final global path of the UAV swarm; if there is a path segment that needs to be replanned, the path segment that needs to be replanned is marked;

[0065] The local path optimization module is used to optimize the path segment that needs to be replanned through the flexible action-evaluation algorithm to obtain the second global path of the UAV swarm;

[0066] The dynamic obstacle avoidance module is used to adjust the second global path according to the dynamic obstacle avoidance mechanism, and use the adjusted second global path as the final global path of the UAV swarm.

[0067] In the embodiment of the present invention, through the global path planning module, the first global path of the UAV swarm is obtained, providing a stable overall planning direction for the UAV swarm and maintaining the formation stability of the UAV swarm; according to the path replanning determination module, it is determined whether there is a path segment that needs to be replanned in the first global path, providing a judgment basis and data basis for whether local optimization is needed and which path segment needs to be optimized in the follow-up; through the local path optimization module, the path segment that needs to be replanned is optimized to obtain the second global path of the UAV swarm, while improving the dynamic adaptability of path planning, comprehensively considering the multi-objective optimization requirements of path quality, formation stability, energy efficiency, and task completion; according to the dynamic obstacle avoidance module, the second global path is adjusted to obtain the final global path of the UAV swarm, improving the adaptability of path planning in a dynamic obstacle environment. Description of the Drawings

[0068] Figure 1 It is a schematic flow chart of the segmented UAV swarm path planning method provided by an embodiment of the present invention;

[0069] Figure 2 It is a specific flow chart of the segmented UAV swarm path planning method provided by an embodiment of the present invention;

[0070] Figure 3 It is an overall structure diagram of the segmented UAV swarm path planning method provided by an embodiment of the present invention;

[0071] Figure 4 It is a local path optimization architecture diagram based on a flexible action-evaluation algorithm provided by an embodiment of the present invention;

[0072] Figure 5 It is a schematic structural diagram of the segmented UAV swarm path planning device provided by an embodiment of the present invention. Detailed Embodiments

[0073] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.

[0074] Please refer to Figure 1 , a segmented UAV swarm path planning method provided by an embodiment of the present invention, including steps S101 to S104, which are described in detail as follows:

[0075] Step S101, obtaining a first global path of the UAV swarm through the A* algorithm and the virtual rigid body algorithm.

[0076] In this step, the obtaining of the first global path of the UAV swarm through the A* algorithm and the virtual rigid body algorithm is specifically as follows:

[0077] Obtaining a global path of the UAV leader through the A* algorithm;

[0078] Obtaining the relative positions between the UAV leader and all UAV followers through the virtual rigid body algorithm;

[0079] Obtaining the first global path of the UAV swarm according to the global path of the UAV leader and the relative positions between the UAV leader and all UAV followers.

[0080] In the embodiments of the present invention, the A* algorithm is combined with the virtual rigid body algorithm to provide a stable overall planning direction for the UAV swarm and maintain the formation stability of the UAV swarm.

[0081] Further, by using the A* algorithm, the global path of the UAV leader is obtained, specifically as follows:

[0082] Model the mission area as a grid map; optionally, the grid value of the grid map represents the obstacle probability (0 for free area, 1 for obstacle);

[0083] According to the preset step size and movement cost, control the UAV leader to perform grid expansion; optionally, the preset step size and movement cost are specifically: the UAV leader starts from the starting point and expands 8-neighborhood grids each time (diagonal movement is allowed); the horizontal / vertical movement cost is 1, and the diagonal movement cost is √2;

[0084] Obtain the global path of the UAV leader according to the cost function and B-spline curve interpolation; wherein, the cost function is defined as:

[0085] f(n) = g(n) + h(n);

[0086] Among them, f(n) is the total path cost of the UAV leader from the starting point to grid n; g(n) is the actual path cost of the UAV leader from the starting point to grid n; h(n) is the heuristic estimated path cost of the UAV leader from the starting point to grid n; optionally, h(n) adopts the Euclidean distance, specifically as follows:

[0087]

[0088] Among them, h(n) is the Euclidean distance; (x n , y n ) is the coordinate value of the current grid; (x goal , y goal ) is the coordinate value of the target grid;

[0089] Optionally, the A* algorithm backtracks from the end point to the starting point to generate a global path point sequence P = {p1, p2,..., p N}; perform B-spline curve interpolation on the global path point sequence to generate a smooth global path of the UAV leader; calculate the heading angle of the UAV leader according to the smooth global path of the UAV leader According to the kinematic model v L = ∥p k+1 - p k ∥ / Δt to calculate the speed command of the UAV leader.

[0090] In the embodiments of the present invention, the global path of the UAV leader is generated through the A* algorithm, providing an overall reference path for the UAV swarm.

[0091] Further, through the virtual rigid body algorithm, the relative positions between the UAV leader and all UAV followers are obtained, specifically:

[0092] The relative positions between the UAV leader and all UAV followers are maintained through geometric constraints; wherein, the geometric constraints are defined as:

[0093]

[0094] wherein, (x l , y l ) is the position of the UAV leader; θ l is the heading angle of the UAV leader; (x i , y i ) is the position of the i-th UAV follower; d i is the relative distance of the i-th UAV follower; β i is the angle of the i-th UAV follower; R(θ l ) is the rotation matrix;

[0095]

[0096] wherein, v l is the speed of the UAV leader; ω l is the angular velocity of the UAV leader;

[0097]

[0098] Optionally, in this embodiment, the control law is designed through the Lyapunov function to ensure the asymptotic convergence of the system error and guarantee that the formation error of the UAV swarm remains within an acceptable range in the dynamic environment, specifically:

[0099] This embodiment considers the dynamic model of the multi-agent system, and the formula is as follows:

[0100]

[0101] wherein, x i (t) ∈ R n is the state variable; u i (t) ∈ R p is the control input; ψ(x i (t), t) is the non-linear dynamic term; A, B, D are the known matrices of the system;

[0102] By designing the following control law, the formation stability of the UAV swarm is guaranteed:

[0103] u i (t)=K1e i +K2v l ;

[0104] where K1 and K2 are control gain matrices; e i is the system error, specifically:

[0105] e i =[x i -x d ,y i -y d T ;

[0106] where (x d ,y d ) is the desired position; the control objective of the system is to ensure that e i converges asymptotically, and can ensure control stability in the face of the existence of non - linear dynamic terms ψ(x i ,t) and communication topology changes;

[0107] Substitute the system error dynamic equation into the control law, and the following formula is obtained:

[0108]

[0109] Select the following Lyapunov function:

[0110]

[0111] where K=-B T P -1 ; Take the derivative of the Lyapunov function to get:

[0112]

[0113] Substitute the control law into the derivative of the Lyapunov function to get:

[0114]

[0115] To ensure system stability, it is required that That is:

[0116] (A + BK1) T P+P(A + BK1)<0;

[0117] Define the matrix Q = P -1 , then the Lyapunov stability condition can be transformed into the following linear matrix inequality problem:

[0118] ​

[0119] Among them, Q is a positive definite matrix; γ>0 is a positive scalar used to adjust the performance index; the gain matrix of the controller is optimized by solving the linear matrix inequality problem, and the linear matrix inequality problem ensures that the system is asymptotically stable.

[0120] In the embodiment of the present invention, the relative positions between the unmanned aerial vehicles are maintained by the virtual rigid body algorithm, ensuring the formation stability of the unmanned aerial vehicle swarm.

[0121] Step S102, according to the path replanning determination index, determine whether there is a path segment in the first global path that needs to be replanned; if there is no path segment that needs to be replanned, use the first global path as the final global path of the unmanned aerial vehicle swarm; if there is a path segment that needs to be replanned, mark the path segment that needs to be replanned.

[0122] In this step, the determination of whether there is a path segment in the first global path that needs to be replanned according to the path replanning determination index is specifically:

[0123] Calculate the perpendicular distance from the current path segment to the shortest path. If the perpendicular distance exceeds the preset threshold, mark the current path segment as a path segment that needs to be replanned; where the calculation formula for the perpendicular distance is as follows:

[0124]

[0125] where D ⊥ is the perpendicular distance; (x1, y1) and (x2, y2) are two points on the shortest path; (x0, y0) is a point on the current path segment;

[0126] Calculate the energy consumption of the unmanned aerial vehicle swarm in the current path segment. If the energy consumption exceeds the preset threshold, mark the current path segment as a path segment that needs to be replanned; where the calculation formula for the energy consumption is as follows:

[0127]

[0128] where E is the energy consumption; v i is the speed of the unmanned aerial vehicle swarm in the i-th path segment; a i is the acceleration of the unmanned aerial vehicle swarm in the i-th path segment; α and β are the weights of the speed and acceleration of the unmanned aerial vehicle swarm in the i-th path segment respectively;

[0129] Calculate the formation circle radius of the unmanned aerial vehicle swarm in the current path segment. If the formation circle radius exceeds the preset threshold, mark the current path segment as a path segment that needs to be replanned; where the calculation formula for the formation circle radius is as follows:

[0130]

[0131] Among them, R formation is the formation circle radius; N is the total number of UAVs; D i is the distance between the i-th UAV and the formation center.

[0132] By comprehensively considering the influence of the perpendicular line length of the shortest path, energy consumption, and formation circle radius index, the embodiment of the present invention segments the result of the global path planning, providing a judgment basis and data foundation for whether local optimization is needed and which path segment needs to be optimized subsequently.

[0133] Step S103, optimize the path segment to be replanned through a flexible action-evaluation algorithm to obtain the second global path of the UAV swarm.

[0134] In this step, the process of optimizing the path segment to be replanned through a flexible action-evaluation algorithm to obtain the second global path of the UAV swarm is specifically as follows:

[0135] Determine the state space and action space of the UAV swarm; optionally, the state space includes information such as the current position, current speed of the UAV, and relative positions with respect to the target point and obstacles, and the action space includes linear velocity, angular velocity, and acceleration adjustment of the UAV;

[0136] Optimize the path segment to be replanned according to the state space, action space, and a preset reward function; wherein, the reward function is defined according to multi-objective evaluation indexes, and the multi-objective evaluation indexes include path quality, formation stability, energy efficiency, and task completion degree;

[0137] Replace the corresponding path segment in the first global path with the optimized path segment to obtain the second global path of the UAV swarm.

[0138] The embodiment of the present invention optimizes and adjusts the path segment to be replanned through a flexible action-evaluation algorithm, improving the dynamic adaptability of path planning.

[0139] Optionally, the reward function is defined according to multi-objective evaluation indexes, specifically as follows:

[0140] R t = w1·R goal + w2·R path + w3·R obs + w4·R energy + w5·R time + w6·

[0141] R formation ;

[0142] Among them, R t is the expected reward; w1, w2, w3, w4, w5, and w6 are the weights of each index item respectively; R goal , R path , R obs , R energy , R time , and R formation are the reward item for reaching the target point, the dynamic reward item for approaching the target, the obstacle distance penalty item, the energy efficiency penalty item, the time step penalty item, and the formation stability penalty item respectively; Classify each index into four categories: path quality, formation stability, energy efficiency, and task completion. Specifically:

[0143] Path quality includes the dynamic reward item for approaching the target and the obstacle distance penalty item, which are respectively:

[0144]

[0145] Among them, d max = 150m is the maximum diagonal distance of the grid map (the grid map range is set to [150, 150]); The role of R path is that the closer the UAV is to the target, the higher the reward value, and the reward growth rate increases non-linearly with the decrease of d t .

[0146]

[0147] Among them, d obs = min(|p t - p obs |) is the distance between the UAV and the nearest obstacle at the current moment; The role of R obs is that the penalty value increases hyperbolically with the decrease of d obs , forcing the UAV to stay away from the obstacle;

[0148] Formation stability includes the formation stability penalty item:

[0149]

[0150] Among them, p i is the actual position of the i-th UAV; is the expected position generated by the virtual rigid body algorithm; β is the UAV formation error penalty coefficient;

[0151] Energy efficiency includes the energy efficiency penalty item:

[0152] R energy = -0.05·|a t | 2 ;

[0153] Among them, a t is the action vector of the UAV at the current moment; |a t | 2 is the square of the Euclidean norm of the action, representing the magnitude of the acceleration; R energy is used to suppress drastic changes in speed or direction and reduce energy consumption;

[0154] The task completion degree includes a target point arrival reward item and a time step penalty item, which are respectively:

[0155] R goal = 200·I(d t <r th );

[0156] Among them, d t = |p t -p goal | is the Euclidean distance between the UAV and the target point at the current moment; r th = 1m is the UAV arrival judgment threshold (about 1.5 times the diameter of the UAV); I(·) is the indicator function (takes 1 when the condition is satisfied, otherwise takes 0); R goal is used to give a high-intensity positive reward when the UAV enters the target point neighborhood, guiding rapid convergence to the target;

[0157]

[0158] Among them, T max = 500 is the maximum number of time steps for single-round training (which can be designed according to specific tasks); R time is used to linearly increase the penalty over time, prompting the UAV to complete the task quickly.

[0159] Furthermore, optimizing the path segment that needs to be replanned according to the state space, action space, and the preset reward function is specifically:

[0160] Obtain the motion strategy range of the UAV swarm through the state space and action space;

[0161] The flexible action-evaluation algorithm optimizes the path segment that needs to be replanned by maximizing the weighted sum of the expected reward and the policy entropy; among them, the flexible action-evaluation algorithm includes prioritized experience replay and soft update of the target network.

[0162] Exemplarily, the flexible action-evaluation algorithm includes the following steps:

[0163] State input:

[0164] s t = [x t ,y t ,v x,t, v y,t , d obs , cosθ t , sinθ t ;

[0165] Among them, x t , y t is the current position of the UAV; v x,t , v y,t is the velocity component of the UAV; d obs is the distance to the nearest obstacle; θ t is the target azimuth angle (decomposed into cos / sin to avoid angle jumps);

[0166] Actor network output:

[0167] μ t = Actor(s t ) is the action mean; a t ~ N(μ t , σ t ) is the exploration action with added Gaussian noise;

[0168] Critic network output:

[0169] Q1(s t , a t ) = Critic1(s t , a t );

[0170] Q2(s t , a t ) = Critic2(s t , a t );

[0171] Target Q-value calculation:

[0172] Q target = min(Q1, Q2) + αH(π(·|s t ));

[0173] Among them, H(π) = -logπ(a t |s t ) is the policy entropy; α is the entropy temperature coefficient;

[0174] Actor loss function:

[0175] L Actor = -Ea t ~ π[Q(s t , a t ) - αlogπ(a t |s t )];

[0176] Critic loss function:

[0177] L Critic = E[(Q(s t , a t ) - y t ) 2 ;

[0178] where y t = r t + γ.Q target (s t+1 , a t+1 ) is the target value; γ = 0.99 is the discount factor;

[0179] Soft update of the target network:

[0180] Q target ← τθ + (1 - τ)Q target ;

[0181] where τ = 0.005 is the hyperparameter controlling the update rate;

[0182] Experience replay and prioritized sampling:

[0183]

[0184] where, is to prevent zero priority;

[0185] Sampling probability:

[0186]

[0187] where α = 0.6 is to control the sampling bias towards high-error samples;

[0188] State space expansion:

[0189] is the UAV formation constraint information (such as formation center offset);

[0190] Output normalized velocity command:

[0191] a t = [Δv x , Δv y ∈ [-1, 1];

[0192] Actual velocity update:

[0193] v t+1 = v t + a t · v max ;

[0194] Among them, v max = 5 m / s (this value can be set according to the specific model).

[0195] In the embodiment of the present invention, by maximizing the weighted sum of the expected reward and the policy entropy, the multi-objective optimization requirements of path quality, formation stability, energy efficiency, and task completion are comprehensively considered.

[0196] Step S104, according to the dynamic obstacle avoidance mechanism, adjust the second global path, and use the adjusted second global path as the final global path of the UAV swarm.

[0197] In this step, the adjusting the second global path according to the dynamic obstacle avoidance mechanism and using the adjusted second global path as the final global path of the UAV swarm specifically includes:

[0198] Obtain the synthetic control force between the UAV swarm and the obstacle through the artificial potential field method;

[0199] Adjust the second global path of the UAV swarm through the synthetic control force to obtain the final global path of the UAV swarm.

[0200] In the embodiment of the present invention, obstacle avoidance is performed by combining the artificial potential field method with path adjustment, improving the adaptability of path planning in a dynamic obstacle environment.

[0201] Further, the obtaining the synthetic control force between the UAV swarm and the obstacle through the artificial potential field method specifically includes:

[0202] F att = -ζ(q - q next );

[0203] Among them, F att is the gravitational potential field; ζ is the gravitational strength coefficient; q is the current position of the UAV swarm; q next is the next target point on the second global path;

[0204]

[0205] Among them, F rep is the repulsive potential field; η is the potential field strength coefficient; ρ is the distance between the UAV swarm and the obstacle; ρ0 is the influence range;

[0206] F = F att + F rep ;

[0207] Among them, F is the synthetic control force.

[0208] In the embodiment of the present invention, the synthetic control force of the gravitational potential field and the repulsive potential field is generated to control the UAV swarm to avoid obstacles.

[0209] Please refer to Figure 2 , which is the specific flowchart of a segmented UAV swarm path planning method provided by an embodiment of the present invention. In this embodiment, through the A* algorithm and the Virtual Rigid Formation (VRF) algorithm, the first global path of the UAV swarm is obtained, providing a stable overall planning direction for the UAV swarm and maintaining the formation stability of the UAV swarm; according to the path replanning determination index, it is determined whether there is a path segment that needs to be replanned in the first global path, providing a judgment basis and data foundation for whether local optimization is needed and which path segment needs to be optimized in the follow-up; through the Soft Actor-Critic (SAC) algorithm, the path segment that needs to be replanned is optimized to obtain the second global path of the UAV swarm, while improving the dynamic adaptability of path planning, comprehensively considering the multi-objective optimization requirements of path quality, formation stability, energy efficiency, and task completion; according to the dynamic obstacle avoidance mechanism, the second global path is adjusted, and the adjusted second global path is used as the final global path of the UAV swarm, improving the adaptability of path planning in a dynamic obstacle environment. Please refer to Figure 3 , which is the overall structure diagram of a segmented UAV swarm path planning method provided by an embodiment of the present invention. Compared with the prior art, the present application segments the result of global path planning, dynamically optimizes the local path, and adjusts the optimization threshold according to different task requirements, capable of improving the dynamic adaptability and calculation efficiency of UAV swarm path planning in multi-objective tasks.

[0210] Please refer to Figure 4 , which is the local path optimization architecture diagram based on the Soft Actor-Critic (SAC) algorithm provided by an embodiment of the present invention. This architecture adopts a hierarchical strategy to optimize the trajectory of the UAV through an intelligent decision-making network, enabling it to have stronger adaptive adjustment capabilities in complex environments, including the core algorithm process and the dynamic optimization mechanism, which are described in detail as follows:

[0211] The core algorithm process includes action generation and execution, reward calculation, and multi-objective reward synthesis;

[0212] Further, the action generation and execution include policy output, action mapping, and path update, specifically:

[0213] Policy output:

[0214]

[0215] Action mapping refers to converting the normalized instruction into the actual speed:

[0216] v t+1 = vt +a t ·v max ;

[0217] Path update means updating the position according to the speed:

[0218] x t+1 =x t +v x,t ·Δt, y t+1 =y t +v y,t ·Δt;

[0219] Furthermore, the reward calculation is specifically as follows:

[0220]

[0221] Furthermore, the multi-objective reward synthesis includes Critic network update, Actor network update, and target network soft update, specifically as follows:

[0222] The Critic network is updated according to minimizing the mean squared error:

[0223] L Critic =E0(Q i (s t , a t )-(r t +γminQ target (s t+1 , a t+1 ))) 2 1, i = 1, 2;

[0224] The Actor network is updated according to maximizing the entropy-regularized return:

[0225] L Actor =-E[minQ i (s t , a t ) - αlogπ(a t |s t )];

[0226] Target network soft update:

[0227] θ target ←0.005·θ + 0.995·θ target ;

[0228] The dynamic optimization mechanism includes path replanning trigger and obstacle avoidance strategy;

[0229] Furthermore, the path replanning trigger includes deviation threshold and soft switching transition, specifically as follows:

[0230] If the perpendicular distance between the path point and the global reference path is greater than 3m, local optimization is triggered; through the Sigmoid function Smoothly switch between the global path and the locally optimized path;

[0231] Furthermore, the obstacle avoidance strategy is specifically as follows:

[0232] When d obs < 2m, the forced adjustment action is:

[0233]

[0234] Please refer to Figure 5 , a segmented UAV swarm path planning device provided by an embodiment of the present invention, including a global path planning module 501, a path replanning determination module 502, a local path optimization module 503, and a dynamic obstacle avoidance module 504, which are described in detail as follows:

[0235] The global path planning module 501 is used to obtain the first global path of the UAV swarm through the A* algorithm and the virtual rigid body algorithm;

[0236] The path replanning determination module 502 is used to determine whether there is a path segment that needs to be replanned in the first global path according to the path replanning determination index; if there is no path segment that needs to be replanned, the first global path is used as the final global path of the UAV swarm; if there is a path segment that needs to be replanned, the path segment that needs to be replanned is marked;

[0237] The local path optimization module 503 is used to optimize the path segment that needs to be replanned through the flexible action-evaluation algorithm to obtain the second global path of the UAV swarm;

[0238] The dynamic obstacle avoidance module 504 is used to adjust the second global path according to the dynamic obstacle avoidance mechanism, and use the adjusted second global path as the final global path of the UAV swarm.

[0239] In this embodiment, the global path planning module 501 includes a leader path sub-module, a formation stability sub-module, and a global path sub-module, specifically:

[0240] The leader path sub-module is used to obtain the global path of the UAV leader through the A* algorithm;

[0241] The formation stability sub-module is used to obtain the relative positions between the UAV leader and all UAV followers through the virtual rigid body algorithm;

[0242] The global path sub-module is used to obtain the first global path of the UAV swarm according to the global path of the UAV leader and the relative positions between the UAV leader and all UAV followers.

[0243] In the embodiment of the present invention, the A* algorithm is combined with the virtual rigid body algorithm to provide a stable overall planning direction for the UAV swarm and maintain the stable formation of the UAV swarm.

[0244] In this embodiment, the leader path sub-module includes a map modeling unit, a grid expansion unit, and a path generation unit, specifically:

[0245] The map modeling unit is used to model the task area as a grid map;

[0246] The grid expansion unit is used to control the UAV leader to perform grid expansion according to a preset step size and movement cost;

[0247] The path generation unit is used to obtain the global path of the UAV leader according to the cost function and B-spline curve interpolation; wherein, the cost function is defined as:

[0248] f(n) = g(n) + h(n);

[0249] Wherein, f(n) is the total path cost of the UAV leader from the starting point to grid n; g(n) is the actual path cost of the UAV leader from the starting point to grid n; h(n) is the heuristic estimated path cost of the UAV leader from the starting point to grid n.

[0250] In the embodiment of the present invention, the global path of the UAV leader is generated by the A* algorithm, providing an overall reference path for the UAV swarm.

[0251] In this embodiment, the formation stability sub-module includes a geometric constraint unit, specifically:

[0252] The geometric constraint unit is used to maintain the relative positions between the UAV leader and all UAV followers through geometric constraints; wherein, the geometric constraints are defined as:

[0253]

[0254] Wherein, (x l , y l ) is the position of the UAV leader; θ l is the heading angle of the UAV leader; (x i , y i ) is the position of the i-th UAV follower; d i is the relative distance of the i-th UAV follower; β iis the angle of the i-th UAV follower; R(θ l ) is the rotation matrix;

[0255]

[0256] where v l is the speed of the UAV leader; ω l is the angular velocity of the UAV leader;

[0257]

[0258] In the embodiment of the present invention, the relative positions between UAVs are maintained by the virtual rigid body algorithm, ensuring the formation stability of the UAV swarm.

[0259] In this embodiment, the path replanning determination module 502 includes a perpendicular distance determination sub-module, an energy consumption determination sub-module, and a formation circle radius determination sub-module, specifically:

[0260] The perpendicular distance determination sub-module is used to calculate the perpendicular distance from the current path segment to the shortest path. If the perpendicular distance exceeds a preset threshold, the current path segment is marked as a path segment that needs to be replanned; where the calculation formula for the perpendicular distance is as follows:

[0261]

[0262] where D ⊥ is the perpendicular distance; (x1, y1) and (x2, y2) are two points on the shortest path; (x0, y0) is a point on the current path segment;

[0263] The energy consumption determination sub-module is used to calculate the energy consumption of the UAV swarm in the current path segment. If the energy consumption exceeds a preset threshold, the current path segment is marked as a path segment that needs to be replanned; where the calculation formula for the energy consumption is as follows:

[0264]

[0265] where E is the energy consumption; v i is the speed of the UAV swarm in the i-th path segment; a i is the acceleration of the UAV swarm in the i-th path segment; α and β are the weights of the speed and acceleration of the UAV swarm in the i-th path segment respectively;

[0266] The formation circle radius determination sub-module is used to calculate the formation circle radius of the UAV swarm in the current path segment. If the formation circle radius exceeds a preset threshold, the current path segment is marked as a path segment that needs to be replanned; where the calculation formula for the formation circle radius is as follows:

[0267]

[0268] wherein, R formation is the formation circle radius; N is the total number of UAVs; D i is the distance between the i-th UAV and the formation center.

[0269] In the embodiment of the present invention, by comprehensively considering the influence of the perpendicular length of the shortest path, energy consumption, and formation circle radius index, the result of the global path planning is segmented, providing a judgment basis and data foundation for whether local optimization is required and which path segments need to be optimized subsequently.

[0270] In this embodiment, the local path optimization module 503 includes a space determination sub-module, a path optimization sub-module, and a path replacement sub-module, specifically as follows:

[0271] The space determination sub-module is used to determine the state space and action space of the UAV swarm;

[0272] The path optimization sub-module is used to optimize the path segment that needs to be replanned according to the state space, action space, and a preset reward function; wherein, the reward function is defined according to multi-objective evaluation indicators, and the multi-objective evaluation indicators include path quality, formation stability, energy efficiency, and task completion degree;

[0273] The path replacement sub-module is used to replace the corresponding path segment in the first global path with the optimized path segment to obtain the second global path of the UAV swarm.

[0274] In the embodiment of the present invention, through the flexible action-evaluation algorithm, the path segment that needs to be replanned is optimized and adjusted to improve the dynamic adaptability of the path planning.

[0275] In this embodiment, the path optimization sub-module includes a strategy range unit and a path optimization unit, specifically as follows:

[0276] The strategy range unit is used to obtain the motion strategy range of the UAV swarm through the state space and action space;

[0277] The strategy range unit is used to optimize the path segment that needs to be replanned through the flexible action-evaluation algorithm by maximizing the weighted sum of the expected reward and the policy entropy; wherein, the flexible action-evaluation algorithm includes prioritized experience replay and soft update of the target network.

[0278] In the embodiment of the present invention, by maximizing the weighted sum of the expected reward and the policy entropy, the multi-objective optimization requirements of path quality, formation stability, energy efficiency, and task completion degree are comprehensively considered.

[0279] In this embodiment, the dynamic obstacle avoidance module 504 includes a synthetic control force sub-module and a path adjustment sub-module, specifically:

[0280] The synthetic control force sub-module is used to obtain the synthetic control force between the UAV swarm and the obstacle through the artificial potential field method;

[0281] The path adjustment sub-module is used to adjust the second global path of the UAV swarm through the synthetic control force to obtain the final global path of the UAV swarm.

[0282] In the embodiment of the present invention, obstacle avoidance is performed by combining the artificial potential field method with path adjustment, which improves the adaptability of path planning in a dynamic obstacle environment.

[0283] In this embodiment, the synthetic control force sub-module is specifically:

[0284] F att =-ζ(q - q next );

[0285] Among them, F att is the gravitational potential field; ζ is the gravitational strength coefficient; q is the current position of the UAV swarm; q next is the next target point on the second global path;

[0286]

[0287] Among them, F rep is the repulsive potential field; η is the potential field strength coefficient; ρ is the distance between the UAV swarm and the obstacle; ρ0 is the influence range;

[0288] F = F att + F rep ;

[0289] Among them, F is the synthetic control force.

[0290] In the embodiment of the present invention, the synthetic control force of the gravitational potential field and the repulsive potential field is generated to control the UAV swarm to avoid obstacles.

[0291] The above-mentioned segmented UAV swarm path planning device can implement the segmented UAV swarm path planning method in the above method embodiment. The optional items in the above method embodiment are also applicable to this embodiment, which will not be elaborated here. The remaining content of the embodiment of the present application can refer to the content of the above method embodiment, and will not be repeated in this embodiment.

[0292] In this embodiment, through the global path planning module, the first global path of the UAV swarm is obtained, providing a stable overall planning direction for the UAV swarm and maintaining the formation stability of the UAV swarm; according to the path replanning determination module, it is determined whether there is a path segment that needs to be replanned in the first global path, providing a judgment basis and data foundation for whether local optimization is needed and which path segment needs to be optimized in the follow-up; through the local path optimization module, the path segment that needs to be replanned is optimized to obtain the second global path of the UAV swarm, while improving the dynamic adaptability of path planning, comprehensively considering the multi-objective optimization requirements of path quality, formation stability, energy efficiency and task completion; according to the dynamic obstacle avoidance module, the second global path is adjusted to obtain the final global path of the UAV swarm, improving the adaptability of path planning in a dynamic obstacle environment.

[0293] The specific embodiments described above further elaborate on the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only specific embodiments of the present invention and is not used to limit the protection scope of the present invention. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A segmented UAV swarm path planning method, characterized in that, Including: Obtain the first global path of the UAV swarm through the A* algorithm and the virtual rigid body algorithm; According to the path replanning determination index, determine whether there is a path segment that needs to be replanned in the first global path; if there is no path segment that needs to be replanned, use the first global path as the final global path of the UAV swarm; if there is a path segment that needs to be replanned, mark the path segment that needs to be replanned; Optimize the path segment that needs to be replanned through the flexible action-evaluation algorithm to obtain the second global path of the UAV swarm; Adjust the second global path according to the dynamic obstacle avoidance mechanism, and use the adjusted second global path as the final global path of the UAV swarm.

2. The segmented UAV swarm path planning method according to claim 1, characterized in that, The specific method for obtaining the first global path of the UAV swarm through the A* algorithm and the virtual rigid body algorithm is as follows: Obtain the global path of the UAV leader through the A* algorithm; Obtain the relative positions between the UAV leader and all UAV followers through the virtual rigid body algorithm; According to the global path of the UAV leader and the relative positions between the UAV leader and all UAV followers, obtain the first global path of the UAV swarm.

3. The segmented UAV swarm path planning method according to claim 2, characterized in that The specific method for obtaining the global path of the UAV leader through the A* algorithm is as follows: Model the task area as a grid map; Control the UAV leader to perform grid expansion according to the preset step size and movement cost; Obtain the global path of the UAV leader according to the cost function and B-spline curve interpolation; where, the cost function is defined as: f(n) = g(n) + h(n); Where, f(n) is the total path cost of the UAV leader from the starting point to grid n; g(n) is the actual path cost of the UAV leader from the starting point to grid n; h(n) is the heuristic estimated path cost of the UAV leader from the starting point to grid n.

4. The segmented UAV swarm path planning method according to claim 2, wherein, The specific method for obtaining the relative positions between the UAV leader and all UAV followers through the virtual rigid body algorithm is as follows: Maintain the relative positions between the UAV leader and all UAV followers through geometric constraints; where, the geometric constraints are defined as: Among them, (x l , y l ) is the position of the UAV leader; θ l is the heading angle of the UAV leader; (x i , y i ) is the position of the i-th UAV follower; d i is the relative distance of the i-th UAV follower; β i is the angle of the i-th UAV follower; R(θ l ) is the rotation matrix; where v l is the speed of the UAV leader; ω l is the angular velocity of the UAV leader; 5. The segmented UAV swarm path planning method according to claim 1, wherein, The specific method for determining whether there is a path segment that needs to be replanned in the first global path according to the path replanning determination index is as follows: Calculate the perpendicular distance from the current path segment to the shortest path. If the perpendicular distance exceeds the preset threshold, mark the current path segment as a path segment that needs to be replanned; where, the calculation formula for the perpendicular distance is as follows: where D ⊥ is the perpendicular distance; (x1, y1) and (x2, y2) are two points on the shortest path; (x0, y0) is a point on the current path segment; Calculate the energy consumption of the UAV swarm in the current path segment. If the energy consumption exceeds the preset threshold, mark the current path segment as a path segment that needs to be replanned; where, the calculation formula for the energy consumption is as follows: Among them, E is the energy consumption; v i is the speed of the UAV swarm in the i-th path segment; a i is the acceleration of the UAV swarm in the i-th path segment; α and β are the weights of the speed and acceleration of the UAV swarm in the i-th path segment, respectively; Calculate the formation circle radius of the UAV swarm in the current path segment. If the formation circle radius exceeds the preset threshold, mark the current path segment as a path segment that needs to be replanned; where, the calculation formula for the formation circle radius is as follows: where R formation is the formation circle radius; N is the total number of UAVs; D i is the distance between the i-th UAV and the formation center.

6. The segmented UAV swarm path planning method according to claim 1, wherein, The specific method for optimizing the path segment that needs to be replanned through the flexible action-evaluation algorithm to obtain the second global path of the UAV swarm is as follows: Determine the state space and action space of the UAV swarm; Optimize the path segment to be replanned according to the state space, action space, and a preset reward function; wherein, the reward function is defined according to multi-objective evaluation metrics, and the multi-objective evaluation metrics include path quality, formation stability, energy efficiency, and task completion degree; Replace the corresponding path segment in the first global path with the optimized path segment to obtain the second global path of the UAV swarm.

7. The segmented UAV swarm path planning method according to claim 6, wherein, The optimizing the path segment to be replanned according to the state space, action space, and a preset reward function is specifically as follows: Obtain the motion strategy range of the UAV swarm through the state space and action space; The flexible actor-critic algorithm optimizes the path segment to be replanned by maximizing the weighted sum of the expected reward and the policy entropy; wherein, the flexible actor-critic algorithm includes prioritized experience replay and soft update of the target network.

8. A segmented UAV swarm path planning method according to claim 1, wherein, The adjusting the second global path according to the dynamic obstacle avoidance mechanism and taking the adjusted second global path as the final global path of the UAV swarm is specifically as follows: Obtain the synthetic control force between the UAV swarm and the obstacles through the artificial potential field method; Adjust the second global path of the UAV swarm through the synthetic control force to obtain the final global path of the UAV swarm.

9. The segmented UAV swarm path planning method according to claim 8, wherein The obtaining the synthetic control force between the UAV swarm and the obstacles through the artificial potential field method is specifically as follows: F att = -ζ(q - q next ); Among them, F att is the gravitational potential field; ζ is the gravitational strength coefficient; q is the current position of the UAV swarm; q next is the next target point on the second global path; Among them, F rep is the repulsive potential field; η is the potential field strength coefficient; ρ is the distance between the UAV swarm and the obstacle; ρ0 is the influence range; F = F att + F rep ; Wherein, F is the synthetic control force.

10. A segmented UAV swarm path planning device, characterized in that, It includes a global path planning module, a path replanning determination module, a local path optimization module, and a dynamic obstacle avoidance module: The global path planning module is used to obtain the first global path of the UAV swarm through the A* algorithm and the virtual rigid body algorithm; The path replanning determination module is used to determine whether there is a path segment to be replanned in the first global path according to the path replanning determination index; If there is no path segment to be replanned, take the first global path as the final global path of the UAV swarm; if there is a path segment to be replanned, mark the path segment to be replanned; The local path optimization module is used to optimize the path segment to be replanned through the flexible actor-critic algorithm to obtain the second global path of the UAV swarm; The dynamic obstacle avoidance module is used to adjust the second global path according to the dynamic obstacle avoidance mechanism and take the adjusted second global path as the final global path of the UAV swarm.

Citation Information

Patent Citations

  • Distributed cluster unmanned aerial vehicle formation flight path generation method in complex unknown environment

    CN114610066A

  • AUV path planning control method based on SAC algorithm

    CN115493597A

  • Multi-unmanned aerial vehicle global and local path intelligent planning method and system

    CN115494866A

  • Unmanned ship cluster hunting method based on improved LSTM network trajectory prediction

    CN116466726A