Unmanned aerial vehicle swarm collaborative path planning method and system based on deep reinforcement learning
Through deep reinforcement learning and multi-potential field coupling model, combined with distributed control and fault tolerance mechanism, the problem of path planning in the collaborative task of drone swarms is solved, efficient and safe path planning and collaborative movement are achieved, and the robustness of the system and the reliability of task execution are improved.
Patent Information
- Application Number
- CN202510583693.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional path planning algorithms have problems such as global search and local adjustment separation, limitations in obstacle avoidance strategies, insufficient colony coordinated control and limited flexibility in intelligent decision-making in drone swarm collaborative tasks, which are difficult to meet the requirements of high-dimensional state space, dynamic environment and real-time tasks.
A deep reinforcement learning method is adopted, combined with sampling algorithms and multi-potential field coupling model, path planning and action decisions are carried out through deep deterministic strategy gradient algorithms and Actor-Critic framework, distributed collaborative control is used to achieve swarm collaborative movement, and fault tolerance mechanisms and security strategies are introduced.
It improves the accuracy of path planning and the coordination of drone operations, and can achieve safe and smooth path planning in complex environments, reduce communication burden and improve system robustness, prevent software attacks, and ensure the safety and reliability of task execution.
Smart Images

Figure CN120447617A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of drone swarm collaboration technology, and in particular to a drone swarm collaborative path planning method and system based on deep reinforcement learning. Background Art
[0002] With the rapid development of drone technology, drone swarms have shown great potential in military operations, disaster relief, environmental monitoring, and other fields. Traditional path planning algorithms (such as Dijkstra, A*, RRT, and its improved RRT* and Informed-RRT*) have achieved some success in solving shortest or globally suboptimal paths. However, they suffer from the following shortcomings when faced with high-dimensional state spaces, dynamic environments, and swarm collaborative tasks:
[0003] 1. Separation of global search and local adjustment: Traditional algorithms often separate path planning into front-end global path search and back-end trajectory optimization, making it difficult to simultaneously meet real-time and dynamic adjustment requirements.
[0004] 2. Limitations of obstacle avoidance strategies: Although methods based on artificial potential fields or collision cones have low computational complexity, they are prone to falling into local optimality and have difficulty dealing with complex obstacles.
[0005] 3. Insufficient swarm collaborative control: Traditional centralized or simple distributed control faces problems such as communication delay, information redundancy and reduced system robustness when coordinating large-scale drones.
[0006] 4. Lack of intelligent decision-making: Faced with ever-changing battlefield environments or mission scenarios, rule-based decision-making algorithms have limited flexibility and are difficult to adapt to real-time mission requirements. Summary of the Invention
[0007] Based on the technical problems existing in the background technology, the present invention proposes a drone swarm collaborative path planning method and system based on deep reinforcement learning, which improves the accuracy of path planning and the coordination of drone movements.
[0008] The UAV swarm collaborative path planning method based on deep reinforcement learning proposed in this invention includes:
[0009] Obtain environmental state information, determine a preliminary path in a pre-built state space based on a sampling algorithm, and perform local smoothing and obstacle avoidance on the preliminary path;
[0010] Use deep reinforcement learning to determine the drone's action instructions at each time step, and correct the initial path after smoothing and obstacle avoidance in real time;
[0011] The drones regularly exchange local information to achieve the overall coordinated movement of the swarm.
[0012] Furthermore, in determining a preliminary path in a pre-constructed state space based on a sampling algorithm, an improved sampling algorithm is used to determine a global path from a starting point to a target point in the pre-constructed sampling space, and an elliptical sampling area is used to narrow the search range.
[0013] Furthermore, the elliptical sampling area formula is as follows:
[0014]
[0015] Among them, q is the sampling point, c max is the current optimal path length, q start is the starting point of path planning, q goal The target point for path planning.
[0016] Furthermore, in the local smoothing and obstacle avoidance of the preliminary path, specifically:
[0017] Based on real-time sensor data, an artificial potential field method is used to perform local path adjustments;
[0018] When an obstacle is detected, the local path is replanned using a local optimal algorithm to ensure collision avoidance.
[0019] Furthermore, the local path is replanned through a local optimal algorithm to ensure collision avoidance, specifically:
[0020] A multi-potential field coupling model is constructed using the target gravitational field, the repulsive field of the neighboring drones, the repulsive field of the obstacle, the formation maintenance force, and the safety redundancy control field;
[0021] Calculating the total potential field force vector on the UAV through the multi-potential field coupling model to obtain the current state of the UAV;
[0022] Starting from the current state, sample a set number of alternative trajectories τ j And select the optimal trajectory through the cost function to obtain the path planning result;
[0023] According to the position of the UAV in the swarm formation, the weight coefficient of each potential field in the multi-potential field coupling model is dynamically adjusted to balance obstacle avoidance, formation stability and path smoothness.
[0024] Furthermore, the total potential field force vector is calculated as follows:
[0025]
[0026] Among them, F i is the total potential force vector of the i-th UAV, F att,i is the target gravitational field of the i-th UAV, F rep,j is the repulsive force generated by the neighboring drone, Fobs,k is the repulsive force from the repellent k, F form,i is the formation maintenance force of the i-th UAV, j is the index of the neighboring UAVs of the i-th UAV, N i is the set of neighboring drones of the i-th drone, and O is the set of repellers.
[0027] Furthermore, the cost function formula is as follows:
[0028]
[0029] Among them, F rep is the obstacle avoidance repulsion strength, F form is the degree of deviation from the formation structure, is the velocity change, and α, β, and γ are weight factors respectively.
[0030] Furthermore, in using deep reinforcement learning to determine the drone's action instructions at each time step and correct the initial path in real time, a deep deterministic policy gradient algorithm is adopted to realize continuous action space decision-making based on the Actor-Critic framework.
[0031] Furthermore, in the continuous action space decision-making based on the Actor-Critic framework, the critic network update formula is:
[0032] L(θ Q )=Ε (s,a,r,s') [(Q(s,a|θ Q )-(r+γQ'(s',μ'(s')|θ Q' ))) 2 ];
[0033] Among them, θ Q is the parameter of the main critic network, which needs to be updated by optimizing the loss function, Q(s,a|θ Q ) is the output value of the current Critic network. θ Q' are the parameters of the target Critic network, u' is the target Actor network, μ'(s') is the action output by the target Actor network, Q'(s',μ'(s')|θ Q' ) is the target Q value estimate, r is the immediate reward, γ is the discount factor, L is the updated Critic network, E is the expected value, s is the current state, a is the action performed in the current state s, and s' is the next state transferred to after performing action a.
[0034] Furthermore, in the process of realizing the overall coordinated movement of the swarm by regularly exchanging local information among drones, the multi-drone consistency control formula is:
[0035]
[0036] Among them, u i is the control input of the i-th UAV, N i is the neighbor set of the i-th drone, x i , v i are the position and velocity of the i-th UAV, x j , v j are the position and velocity of the jth UAV, a ij is the communication weight between the i-th UAV and the j-th UAV.
[0037] Furthermore, in the collaborative path planning of drone swarms, the software integrity and operating status are continuously monitored, and fault tolerance mechanisms or safe return strategies are enabled when anomalies are encountered.
[0038] A UAV swarm collaborative path planning system based on deep reinforcement learning, including an environmental perception module, a global planning module, a local optimization module, a deep reinforcement learning decision module, and a distributed collaborative control module;
[0039] The environment perception module is used to obtain environment status information;
[0040] The global planning module is used to determine the preliminary path in the pre-constructed state space based on the sampling algorithm.
[0041] The local optimization module is used to perform local smoothing and obstacle avoidance on the preliminary path;
[0042] The deep reinforcement learning decision module is used to use deep reinforcement learning to determine the action instructions of the drone at each time step, and to correct the initial path after smoothing and obstacle avoidance in real time;
[0043] The distributed collaborative control module is used to periodically exchange local information through drones to achieve overall collaborative movement of the swarm.
[0044] Furthermore, it also includes a security and fault tolerance module;
[0045] The safety and fault-tolerance module is used to continuously monitor the integrity and operating status of the software, and to enable a fault-tolerance mechanism or a safe return strategy when an anomaly is encountered.
[0046] The advantages of the UAV swarm collaborative path planning method and system based on deep reinforcement learning provided by the present invention are: using multi-level planning and online decision-making to ensure safe and smooth path planning in complex environments; distributed control and consistency protocols enable each node to work collaboratively, reducing the communication burden and improving system robustness; local planning based on artificial potential field and online optimization can quickly respond to sudden obstacles; integrating security mechanisms such as TPM and control flow authentication to effectively prevent software and dynamic attacks, ensuring safe and reliable mission execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the process of the present invention;
[0048] Figure 2 This is a schematic diagram of the drone swarm navigation loop;
[0049] Figure 3 This is a schematic diagram of the deep reinforcement learning and distributed consistency control module structure, showing the deep reinforcement learning network (Actor-Critic framework) built into each drone node, the local decision-making module, and the architecture diagram of information interaction with neighboring drones. DETAILED DESCRIPTION
[0050] The technical solutions of the present invention are described in detail below through specific embodiments. Numerous specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0051] like Figures 1 to 3 As shown, the UAV swarm collaborative path planning method based on deep reinforcement learning proposed in the present invention includes:
[0052] Step 1: Obtain environmental state information, determine a preliminary path in a pre-constructed state space based on a sampling algorithm, and perform local smoothing and obstacle avoidance on the preliminary path;
[0053] Collect real-time data through multiple sensors to model obstacles, terrain, weather, etc. in the environment;
[0054] Global path search layer: In a pre-constructed grid or sampling space, an improved sampling algorithm (Informed-RRT*) is used to determine a global path from the starting point to the target point, and an elliptical sampling area is used to narrow the search range.
[0055] That is, the heuristic function of the state space is used to quickly find the preliminary path. The elliptical sampling area formula of Informed-RRT* is:
[0056]
[0057] Among them, c max is the current optimal path length, q start is the starting point of path planning, q goal is the target point of path planning, q is the sampling point, and the search efficiency is improved by limiting the sampling area.
[0058] Local path optimization layer: A local trajectory optimization algorithm is used to smooth the preliminary path so that the path meets the kinematic requirements of the drone. That is, based on real-time sensor data, an artificial potential field method is used to adjust the local path. When an obstacle is detected, the local path is replanned using a local optimal algorithm to ensure collision avoidance.
[0059] Hierarchical planning coordination mechanism: global search determines the approximate route, and local real-time adjustments ensure planning efficiency and obstacle avoidance safety.
[0060] When an obstacle is detected, the local path is replanned using a local optimal algorithm to ensure collision avoidance, specifically (1) to (4):
[0061] (1) Construct a multi-potential field coupling model using the target gravity field, the neighboring UAV repulsion field, the obstacle repulsion field, the formation maintenance force, and the safety redundancy control field;
[0062] (2) calculating the total potential field force vector acting on the UAV through the multi-potential field coupling model to obtain the current state of the UAV;
[0063] The total potential field force vector is calculated as follows:
[0064]
[0065] Among them, F i is the total potential force vector of the i-th UAV, F att,i is the target gravity of the i-th UAV, F rep,j is the repulsive force generated by the neighboring drone, F obs,k is the repulsive force from the repellent k, F form,i is the formation maintenance force of the i-th UAV, j is the index of the neighboring UAVs of the i-th UAV, N i is the set of neighboring drones of the i-th drone, and O is the set of repellers.
[0066] Among them, the target gravity is used to pull the drone towards the target point and is the main direction of the flight path:
[0067] F att,i =-k att (p i -p goal );(3)
[0068] Among them, p i is the current position (three-dimensional coordinates) of the i-th UAV, p goal is the location of the task target point, k att is the attraction coefficient. The larger the value, the stronger the attraction.
[0069] Neighboring drone repulsion field (prevents collisions within the swarm):
[0070]
[0071] Among them, p j is the position of the neighboring UAV j, d min is the minimum safe distance between drones, k rep is the repulsive force coefficient, ||p i -p j || is the Euclidean distance between drones i and j, p i and p j are the i-th and j-th drones respectively. When the distance between two drones is less than the set minimum distance d min As the two objects move closer, a stronger and stronger repulsive force is generated, which pushes them away and prevents them from colliding.
[0072] Obstacle repulsion field (prevent collision with walls, avoid dynamic obstacles):
[0073]
[0074] Among them, p obs,k is the position of the kth obstacle, d obs is the interaction threshold between the drone and the obstacle. If the distance exceeds this, no repulsion will be generated. k obs is the obstacle repulsion coefficient. Similar to the repulsion field of a neighboring drone, when approaching an obstacle, the repulsion increases dramatically, causing the drone to quickly change direction to avoid it.
[0075] The formation-maintaining force field (which maintains the formation structure) ensures that the drones do not stray too far from the formation due to obstacle avoidance maneuvers. It is a bit like the magnetic force in a magnetic field, pulling them back to their original position:
[0076]
[0077] in, is the reference position of the i-th UAV in the swarm formation, k form It is the formation pull-back coefficient, which is used to maintain the stability of the swarm structure.
[0078] The total potential force vector is the comprehensive driving force for drone movement in complex environments. Its core function lies in: multi-factor integration: By coupling the attraction of the target, the repulsion of neighboring drones, the repulsion of obstacles, and the formation maintenance force, it provides directional guidance for drones, enabling them to move toward the target while avoiding obstacles and neighboring individuals, while maintaining the swarm formation structure. Real-time dynamic input: The total potential force vector directly serves as the basis for subsequent path planning and weight adjustment. Its calculation reflects the current environmental state (such as obstacle location and neighboring drone distribution).
[0079] (3) Starting from the current state, sample a set number of alternative trajectories τj And select the optimal trajectory through the cost function to obtain the path planning result;
[0080] If a drone detects an obstacle or receives an avoidance instruction from a neighboring individual, it will plan its local path in real time based on a dynamic adjustment mechanism, including (a1) to (a4):
[0081] (a1) Starting from the current state, a certain number of candidate trajectories τ are sampled. j , according to the different values of j, a certain number of trajectories are obtained;
[0082] (a2) Define the trajectory cost function to select the best obstacle avoidance path:
[0083] (a3) Select the trajectory with the minimum cost to execute.
[0084] (a4) If all candidate trajectories cannot meet the obstacle avoidance requirements, the local topology fallback mechanism is triggered and the system returns to the previous node for resampling.
[0085] The trajectory cost function is as follows:
[0086]
[0087] Among them, J(τ j ) is the trajectory τ j The cost function value, F rep To avoid the repulsive force (it will be very strong if it is too close to the neighboring machine / obstacle), F form is the degree of deviation from the formation structure, is the speed change, α, β, and γ are weight factors, which are flexibly adjusted according to the actual task to balance the effects of obstacle avoidance, formation, and speed stability.
[0088] The relationship between the total potential field vector and path planning is as follows:
[0089] Path Generation Basics: The total potential field force vector determines the current force state of the drone, directly influencing the generation of alternative trajectories during path planning. For example, when the repulsive field is strong, alternative trajectories should prioritize avoiding areas with high repulsive forces.
[0090] Cost function evaluation: The cost function of path selection (such as obstacle avoidance cost, formation deviation cost) directly depends on the contribution value of each component force in the total potential field force (such as || F rep || 2 、||F form || 2 ). The magnitude and direction of the total potential field force determine the feasibility of the trajectory.
[0091] Dynamic feedback adjustment: If all alternative trajectories cannot meet the obstacle avoidance requirements (e.g., potential field conflict leads to a local optimal trap), the path planning module triggers the weight coefficient in the weight adjustment step (4) through the fallback mechanism, indirectly correcting the calculation logic of the total potential field force.
[0092] (4) According to the position of the UAV in the swarm formation, the weight coefficient of each potential field in the multi-potential field coupling model is dynamically adjusted to balance obstacle avoidance, formation stability and path smoothness.
[0093] The relationship between the total potential field vector and the weight coefficient adjustment:
[0094] Dynamic weight adjustment is based on the following: The position of a drone in the swarm (e.g., edge or core) directly affects the formation-maintaining force weight (β), while obstacles or the distance to neighboring drones determine the repulsion weight (α). Adjustments to these weight coefficients directly affect the calculation of the total potential field force vector, altering the relative strength of each component force.
[0095] Avoid local optimality: Through safe redundant control fields (such as dynamic threshold correction), weight adjustment can alleviate the force direction conflict problem caused by the superposition of multiple potential fields, ensuring that the total potential field force always guides the drone to approach the global optimal path.
[0096] The relationship between path planning and weight coefficient adjustment:
[0097] Path planning feedback-driven weight adjustment: When the trajectory cost function in path planning is too high (for example, the formation deviates too much or obstacle avoidance is difficult), the formation maintenance weight (β) or repulsion weight (α) is increased to optimize the distribution of the total potential field force, thereby generating a more reasonable alternative trajectory.
[0098] Weight adjustment optimizes path selection: Dynamic weight coefficients make the total potential field force vector more adaptable to complex scenarios (such as dense obstacle areas or formation structure adjustments), ultimately reducing the cost function value of the path planning module and improving the feasibility and smoothness of the path.
[0099] Therefore, for steps (2) to (4), the logical framework is briefly described as follows:
[0100] (b1) Input layer: Environmental information (obstacles, neighboring aircraft positions, formation structure) is converted into a total potential field force vector through step (2);
[0101] (b2) Decision layer: Step (3) generates alternative paths based on the total potential field force and selects the optimal trajectory through the cost function.
[0102] (b3) Optimization layer: Step (4) dynamically adjusts the weight coefficient according to the path planning results and formation status to optimize the total potential field force calculation.
[0103] (b4) Closed-loop feedback: The three form a dynamic cycle of "calculation (b1) → planning (b2) → optimization (b3) → recalculation (b1)", ensuring that the bee colony can operate efficiently and collaboratively in a complex environment.
[0104] The local optimal algorithm of this embodiment has the following advantages:
[0105] (c1) Multi-potential field joint driving mechanism: introducing compound potential fields such as obstacle repulsion, neighbor repulsion, formation maintenance, and target attraction to achieve dynamic coordination between individuals and groups in the swarm.
[0106] (c2) Adaptive weight mechanism: Each drone in the swarm adaptively adjusts the weights of various potential fields according to its position at the edge or core of the formation to improve global path coordination. Specifically, when a drone is at the edge of the swarm formation, the weight coefficient β of the formation maintenance force is increased; when a drone is close to an obstacle or a neighboring individual, the weight coefficient α of the obstacle repulsion or the neighboring drone repulsion is increased; when a smooth path is required, the weight coefficient γ of the speed change term is increased.
[0107] (c3) Local sampling optimization strategy: Combining trajectory sampling with cost function minimization method, the obstacle avoidance path is not only safe but also meets the dynamics and formation requirements.
[0108] (c4) The safety redundancy control field introduces dynamic threshold correction to prevent the local optimal trap caused by the superposition of potential fields. Specifically, when the force directions of multiple potential fields conflict, the weight of non-critical potential fields is reduced first; safety distance redundancy is added in path planning to ensure the robustness of obstacle avoidance actions.
[0109] Step 2: Use deep reinforcement learning to determine the drone’s action instructions at each time step, and correct the initial path after smoothing and obstacle avoidance in real time;
[0110] Based on the global path, the Actor-Critic model is used to determine the action instructions for each time step and correct the path in real time. Each agent performs actions based on its own state and the state of its neighbors to form a coordinated movement. Specifically:
[0111] Environmental perception and state modeling: using multiple sensors to collect environmental information and build an actor-critic model;
[0112] The Deep Deterministic Policy Gradient (DDPG) algorithm is used to implement continuous action space decision-making based on the Actor-Critic framework. The update formula of the Critic network is:
[0113] L(θ Q )=Ε (s,a,r,s') [(Q(s,a|θ Q)-(r+γQ′(s′,μ′(s′)|θ Q' ))) 2 ];(8)
[0114] Among them, θ Q is the parameter of the main critic network, which needs to be updated by optimizing the loss function, Q(s,a|θ Q ) is the output value of the current Critic network, that is, the Q value estimate of the current state-action pair (s, a) by the main Critic network; θ Q′ are the parameters of the target Critic network, which are usually obtained by soft updating or periodically copying the main network parameters for stable training. u′ is the target Actor network, which is used to generate the action a′=μ′(s′) for the next state s′. μ′(s′) is the action output by the target Actor network. Q′(s′,μ′(s′)|θ Q' ) is the Q-value estimate of the target Critic network in state s' and target Actor action μ'(s'), which constitutes the target value, r is the immediate reward obtained after executing action a, γ is the discount factor, which weighs the importance of current rewards and future rewards (the value range is usually [0,1)), L is the loss function of the Critic network, which is based on the mean square error of the temporal difference (TD) error, E is the expectation, s is the current state, a is the action performed in the current state s, and s' is the next state transferred to after executing action a.
[0115] The core of formula (2) is to minimize the difference between the current Q value and the target Q value, and to improve the training stability by separating the target network (Critic and Actor). That is, the current Critic network's prediction of the Q value of state s and action a should be close to the target Q value calculated by the target network. The target Q value is composed of the immediate reward r and the Q value after taking the action in the next state (estimated by the target Critic network). Finally, by minimizing this loss, θ Q Perform gradient descent updates to improve training stability.
[0116] Online policy update: Adjust path planning strategies through real-time interaction to cope with unexpected obstacles and task changes.
[0117] Step 3: Regularly exchange local information through drones to achieve overall coordinated movement of the swarm;
[0118] Each drone regularly exchanges status information and adjusts speed and direction through a consistency control protocol to ensure the overall motion consistency of the swarm. Specifically:
[0119] Distributed communication network design: reduce communication burden and use directed topology to reduce system coupling;
[0120] Consistency control protocol: Using the multi-UAV (agent) consistency method, the group can achieve a consistent motion state and improve the collaborative operation effect. Multi-agent consistency control formula:
[0121]
[0122] Among them, u I is the control input of the i-th UAV, N I is the neighbor set of the i-th drone, x I , v I are the position and velocity of the i-th UAV, x j , v j are the position and velocity of the jth UAV, a Ij is the communication weight between the i-th UAV and the j-th UAV.
[0123] Dynamic collision avoidance: Combining artificial potential fields with deep reinforcement learning for online decision-making to avoid local extrema problems.
[0124] Step 4: Continuously monitor software integrity and operating status, and enable fault tolerance mechanisms or safe return strategies when anomalies are encountered.
[0125] Software reliability design: modular design and strict testing process to ensure safe return in the event of a failure;
[0126] Secure encryption and authentication: Using TPM and real-time monitoring to effectively prevent malicious attacks.
[0127] Steps 1 to 4 have the following effects:
[0128] 1. Improve autonomous navigation capabilities: Utilize multi-level planning and online decision-making to ensure safe and smooth path planning in complex environments;
[0129] 2. Enhanced swarm collaboration: Distributed control and consistency protocols enable nodes to work together, reducing communication burdens and improving system robustness.
[0130] 3. Real-time response and obstacle avoidance: Local planning based on artificial potential fields and online optimization can quickly respond to sudden obstacles;
[0131] 4. Security assurance: Integrated security mechanisms such as TPM and control flow authentication effectively prevent software and dynamic attacks, ensuring safe and reliable task execution.
[0132] As an embodiment
[0133] (1) The system is first initialized. Each drone collects surrounding environment data and completes initialization information registration through the distributed network.
[0134] (2) The global planning module uses the improved Informed-RRT* algorithm to generate a preliminary path and plan the overall route of the swarm.
[0135] (3) The local optimization module uses an artificial potential field algorithm to avoid obstacles in front and enemy defensive fire areas in real time, while the deep reinforcement learning module adjusts the drone's movements.
[0136] (4) The distributed consistency control module ensures that each UAV maintains its formation after performing evasive maneuvers, thus achieving flexible and coordinated combat.
[0137] (5) The system security module uses TPM and real-time monitoring technology to prevent possible dynamic attacks and software intrusions to ensure the smooth completion of the task.
[0138] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A UAV swarm collaborative path planning method based on deep reinforcement learning, characterized by: include: Obtain environmental state information, determine a preliminary path in a pre-built state space based on a sampling algorithm, and perform local smoothing and obstacle avoidance on the preliminary path; Use deep reinforcement learning to determine the drone's action instructions at each time step, and correct the initial path after smoothing and obstacle avoidance in real time; The drones regularly exchange local information to achieve the overall coordinated movement of the swarm.
2. The UAV swarm collaborative path planning method according to claim 1, characterized in that: In determining a preliminary path in a pre-constructed state space based on a sampling algorithm, an improved sampling algorithm is used to determine a global path from a starting point to a target point in the pre-constructed sampling space, and an elliptical sampling area is used to narrow the search range. The elliptical sampling area formula is as follows: Among them, q is the sampling point, c max is the current optimal path length, q start is the starting point of path planning, q goal The target point for path planning.
3. The UAV swarm collaborative path planning method according to claim 1, characterized in that: In the local smoothing and obstacle avoidance of the preliminary path, specifically: Based on real-time sensor data, an artificial potential field method is used to perform local path adjustments; When an obstacle is detected, the local path is replanned using a local optimal algorithm to ensure collision avoidance.
4. The UAV swarm collaborative path planning method according to claim 3, characterized in that: In replanning the local path through the local optimal algorithm to ensure collision avoidance, specifically: A multi-potential field coupling model is constructed using the target gravitational field, the repulsive field of the neighboring drones, the repulsive field of the obstacle, the formation maintenance force, and the safety redundancy control field; Calculating the total potential field force vector on the UAV through the multi-potential field coupling model to obtain the current state of the UAV; Starting from the current state, sample a set number of alternative trajectories τ j And select the optimal trajectory through the cost function to obtain the path planning result; According to the position of the UAV in the swarm formation, the weight coefficient of each potential field in the multi-potential field coupling model is dynamically adjusted to balance obstacle avoidance, formation stability and path smoothness.
5. The UAV swarm collaborative path planning method according to claim 4, characterized in that: The total potential field force vector is calculated as follows: Among them, F i is the total potential force vector of the i-th UAV, F att,i is the target gravitational field of the i-th UAV, F rep,j is the repulsive force generated by the neighboring drone, F obs,k is the repulsive force from the repellent k, F form,i is the formation maintenance force of the i-th UAV, j is the index of the neighboring UAVs of the i-th UAV, N i is the set of neighboring drones of the i-th drone, and O is the set of repellers.
6. The UAV swarm collaborative path planning method according to claim 4, characterized in that: The cost function formula is as follows: Among them, F rep is the obstacle avoidance repulsion strength, F form is the degree of deviation from the formation structure, is the velocity change, and α, β, and γ are weight factors respectively.
7. The UAV swarm collaborative path planning method according to claim 1, characterized in that: In using deep reinforcement learning to determine the drone's action instructions at each time step and correct the initial path in real time, a deep deterministic policy gradient algorithm is adopted to achieve continuous action space decision-making based on the Actor-Critic framework.
8. The UAV swarm collaborative path planning method according to claim 1, characterized in that: In the continuous action space decision-making based on the Actor-Critic framework, the critic network update formula is: L(θ Q )=E (s,a,r,s′) [(Q(s,a|θ Q )-(r+γQ′(s′,μ′(s′)|θ Q' ))) 2 ] Among them, θ Q is the parameter of the main critic network, which needs to be updated by optimizing the loss function, Q(s,a|θ Q ) is the output value of the current Critic network, θ Q′ are the parameters of the target Critic network, u′ is the target Actor network, μ′(s′) is the action output by the target Actor network, Q′(s′,μ′(s′)|θ Q′ ) is the target Q value estimate, r is the immediate reward, γ is the discount factor, L is the updated Critic network, E is the expected value, s is the current state, a is the action performed in the current state s, and s′ is the next state transferred to after performing action a.
9. The UAV swarm collaborative path planning method according to claim 1, characterized in that: In the process of realizing the coordinated movement of the swarm by periodically exchanging local information between drones, the multi-drone consistency control formula is: Among them, u i is the control input of the i-th UAV, N i is the neighbor set of the i-th drone, x i , v i are the position and velocity of the i-th UAV, x j , v j are the position and velocity of the jth UAV, a ij is the communication weight between the i-th UAV and the j-th UAV.
10. UAV swarm collaborative path planning system based on deep reinforcement learning, characterized by: It includes environmental perception module, global planning module, local optimization module, deep reinforcement learning decision module and distributed collaborative control module; The environment perception module is used to obtain environment status information; The global planning module is used to determine the preliminary path in the pre-constructed state space based on the sampling algorithm. The local optimization module is used to perform local smoothing and obstacle avoidance on the preliminary path; The deep reinforcement learning decision module is used to use deep reinforcement learning to determine the action instructions of the drone at each time step, and to correct the initial path after smoothing and obstacle avoidance in real time; The distributed collaborative control module is used to periodically exchange local information through drones to achieve overall collaborative movement of the swarm.
Citation Information
Cited By
Cluster robot collaborative operation method and system based on hierarchical path planning
CN121165793A
Unmanned cluster confrontation combat decision-making method, device and program product
CN121348781A
Anti-unmanned ship unmanned aerial vehicle swarm dynamic clustering algorithm and saturation attack path planning method
CN121558029A