Method for path planning of unmanned aerial vehicle swarm in complex environment based on reinforcement learning

By employing a reinforcement learning-based drone swarm path planning method, combined with virtual leadership and formation strategy optimization, the problem of drone swarms rapidly reaching targets and reducing radar detection and failure risks in complex environments is solved, achieving robust path planning and mission success in complex environments.

CN121977578BActive Publication Date: 2026-06-09NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2026-04-03
Publication Date
2026-06-09

Smart Images

  • Figure CN121977578B_ABST
    Figure CN121977578B_ABST
Patent Text Reader

Abstract

The application discloses a path planning method for a UAV swarm in a complex environment based on reinforcement learning, comprising the following steps: step 1, obtaining a converged proximal policy network; step 2, obtaining an expected path of a virtual leader composed of multiple waypoints; step 3, calculating the probability of the UAV swarm being detected by a radar cluster system and the number of failed UAVs newly added to the UAV swarm at each waypoint, and determining whether the UAV swarm is evenly flown at each waypoint; and step 4, obtaining an expected path of each UAV. The application adopts a proximal policy optimization algorithm to construct a reinforcement learning path planning framework, simultaneously considers a composite reward function of path length, the probability of being detected by a radar and the probability of single-machine failure, and can obtain a stable and high-quality strategy in a complex three-dimensional environment and a high-dimensional state space, thereby reducing the risk of falling into a local optimum based on an artificial rule or a single cost function method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of UAV path planning technology, and particularly relates to a path planning method for UAV swarms in complex environments based on reinforcement learning. Background Technology

[0002] Unmanned aerial vehicles (UAVs) demonstrate broad application potential due to their superior maneuverability, high degree of autonomy, and multi-functionality. However, a single UAV has significant limitations in mission execution. In contrast, UAV swarms can significantly improve mission efficiency. By flying in formation, swarms can enhance mission efficiency, improve system survivability, and, to some extent, mitigate the risks associated with a single UAV.

[0003] In complex environments, such as mountainous terrain or networked systems composed of multiple ground radars, drone swarms need to minimize the risk of drone failure due to radar detection and interference while ensuring overall mission completion, thereby improving swarm uptime and mission success rate. Traditional rule-based or cost function-based path planning methods typically rely on manually designed heuristic strategies, which struggle to simultaneously consider multiple factors such as terrain masking, radar detection probability, and individual drone failure probability, and are prone to getting trapped in local optima.

[0004] Existing technologies offer various drone swarming and obstacle avoidance control methods that can maintain swarm formation and achieve local obstacle avoidance while ensuring communication topology connectivity. Some methods also consider avoidance of threat areas. However, these methods are mostly designed for static obstacles or local conflicts, typically simplifying the radar threat area into an insurmountable obstacle. They fail to incorporate the radar swarm system detection probability and the drone failure probability into a unified path optimization framework, and they do not adaptively split and merge the swarm based on the characteristics of the radar resolution unit. Therefore, it is difficult to simultaneously achieve rapid target arrival and reduce the number of drone swarm failures in complex threat environments. Summary of the Invention

[0005] The purpose of this invention is to provide a path planning method for UAV swarms based on reinforcement learning in complex environments, in order to solve the problem that existing methods are difficult to incorporate the risks of detection and failure into unified planning in complex terrain and radar cluster system threat environments, thus making it difficult to simultaneously take into account both rapid target arrival and overall survivability.

[0006] This invention employs the following technical solution: a path planning method for UAV swarms in complex environments based on reinforcement learning, comprising:

[0007] Step 1: Place a device with... The centroid of the drone swarm in a drone formation is defined as the virtual leader, and the current state of the virtual leader is used as the state input for reinforcement learning. The policy network and value network are trained using the proximal policy optimization algorithm to obtain the converged proximal policy network.

[0008] Step 2: Based on the control quantity of the virtual leader output by the converged proximal policy network, perform forward extrapolation based on the dynamic model to obtain the expected path of the virtual leader consisting of multiple waypoints.

[0009] Step 3: Calculate the probability that the drone swarm will be detected by the radar cluster system at each waypoint on the virtual leader's expected path, and the number of newly added drones that fail in the drone swarm.

[0010] When the probability of a drone swarm being detected by a radar cluster system is greater than the first preset threshold or the number of drone failures is greater than the second preset threshold, the drone formations of the drone swarm at the previous waypoint are divided into two formation subgroups for flight, and the positions of each formation subgroup are adjusted so that the distance between the centroids of the two formation subgroups is greater than or equal to the maximum distance between two adjacent radar resolution units.

[0011] Step 4: Based on the expected path of the virtual leader, the deviation vector between the centroid of each drone formation and the virtual leader, and the formation deviation vector between the centroid of each drone formation and the multiple drones it leads, obtain the expected path of each drone.

[0012] The beneficial effects of this invention are:

[0013] This invention equates the multi-target echo effect of a radar swarm to a single-target threat assessment, thereby reducing the computational complexity of radar threat assessment and facilitating training and optimization within a reinforcement learning framework. Simultaneously, it introduces the proportion of non-failed drones in the swarm, enabling path planning to optimize range while mitigating the risks of detection and failure. Furthermore, by combining a mechanism of equally divided flight, spatial dispersion is achieved by increasing the spacing between drone formations in high-risk segments, while restoring the spacing between drone formations in low-risk segments to improve coordination efficiency. This results in a robust and executable desired path for the drone swarm in complex threat environments.

[0014] This invention employs a proximal policy optimization algorithm to construct a reinforcement learning path planning framework. It considers a composite reward function that takes into account path length, probability of radar detection, and single-machine failure probability. This enables the acquisition of stable and high-quality policies in complex three-dimensional environments and high-dimensional state spaces, reducing the risk of getting trapped in local optima by methods based on manual rules or single cost functions.

[0015] This invention limits the distance between the centroids of two formation subgroups to be greater than or equal to the maximum distance between two adjacent radar resolution cells, thereby incorporating radar detection, single-unit failure, the proportion of non-failed UAVs in the swarm, and equal flight into the same decision-making closed loop. This significantly enhances robustness and overall survivability in complex threat environments while ensuring optimized range. Attached Figure Description

[0016] Figure 1 This is a three-dimensional schematic diagram of the simulated path planning of the drone swarm with simulation number 1 from the black starting point to the white target position in the embodiment.

[0017] Figure 2 This is a three-dimensional schematic diagram of the simulated path planning of the drone swarm (serial number 2) from the black starting point to the white target position in the embodiment.

[0018] Figure 3 This is a top view of the drone swarm with simulation number 1 in the embodiment, from the black starting point to the white target position.

[0019] Figure 4 This is a top view of the drone swarm with simulation number 2 in the embodiment, from the black starting point to the white target position;

[0020] Figure 5 The curve showing the change in the radar detection area of ​​the command indicating whether the drone swarm with simulation number 1 in the embodiment is flying in equal segments;

[0021] Figure 6 The curve of the change command for whether the drone swarm with simulation number 2 in the embodiment flies in equal segments within the radar detection area is shown.

[0022] Figure 7 The figure shows the change curve of the number of non-failed drones in the drone swarm with simulation number 1 in the embodiment.

[0023] Figure 8 The figure shows the change curve of the number of non-failed drones in the drone swarm with simulation number 2 in the embodiment. Detailed Implementation

[0024] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0025] This invention discloses a path planning method for drone swarms in complex environments based on reinforcement learning, including steps 1-4.

[0026] Step 1: Place a device with... The centroid of the drone swarm in a drone formation is defined as the virtual leader. The current state of the virtual leader is used as the state input for reinforcement learning. The policy network and value network are trained using the proximal policy optimization algorithm to obtain the converged proximal policy network.

[0027] In step 1, the current state of the virtual leader includes the flight state and the environmental state. The flight state includes position, speed, tilt angle, azimuth angle, and roll angle. The environmental state includes the current terrain obstacle space, the relative distance and azimuth between the virtual leader and the target point, and the relative distance between the virtual leader and each radar.

[0028] Specifically: Define the current state as ;in, This is a virtual leadership position; For the speed of virtual leadership, ; The minimum speed for virtual leadership; The maximum speed for virtual leadership; The angle of inclination. ; It is the azimuth angle. ; For the virtual leadership's rolling corner, ; For virtual leadership and target position The relative position, distance, and direction. ; , From the perspective of the objective relative to virtual leadership, ; For the target location; The relative distance between the virtual leader and the radar cluster system. ,in, The relative distance between the virtual leader and the first radar in a radar cluster system. The relative distance between the virtual leader and the second radar in a radar cluster system. For the virtual leader and the first in the radar cluster system The relative distance between the radars This represents the total number of radars included in the radar cluster system.

[0029] The control output of the reinforcement learning network is tangentially overloaded by the virtual leader. Virtual leadership's normal overload The Rolling Angle of Virtual Leadership ,in, , , .

[0030] During the training of the near-end strategy optimization algorithm, the algorithm's reward function is a composite reward function that takes into account path length, radar detection probability, and single-drone failure probability. By penalizing events such as approaching high-threat areas and a decrease in the proportion of non-failed drones in the drone swarm, the strategy is guided to complete rapid crossing while ensuring the swarm's survivability.

[0031] The constraints in step 1 when training the policy network and the value network are as follows:

[0032] ;

[0033] in, The probability that a swarm of drones will be detected by a radar cluster system. For the total number of radars, For the first radar, For radar Performance parameters, For the UAV center and radar in the resolution unit The distance between them For each radar in a radar cluster system Number of drones within the corresponding resolution unit The maximum value, For radar The number of drones in the corresponding resolution unit; The radar cross-section of a single UAV. Used to characterize the worst-case detection risk of a radar cluster system, and used in training and planning to characterize the upper bound of the detection risk caused by joint echo superposition by the number of clustered units of the same resolution under the worst-case scenario.

[0034] The reward function for training the policy network and value network in step 1 is:

[0035] ;

[0036] in, For the reward function; The reward for the drone swarm reaching the target location; and All are weighting coefficients, through and The swarm of drones tends to either reach the target directly or evade radar detection. Rewards for drone swarms approaching the target location; A reward for drone swarms flying at a preset altitude; The penalty for being detected by a radar cluster system; The penalty is the total number of failed drones in the drone swarm.

[0037] in, ;

[0038] in, For virtual leadership Distance to the target location , The location of the target; For virtual leadership Location; The percentage of failed drones in a drone swarm. ; This represents the total number of failed drones. The total number of drones in the drone swarm;

[0039] in, ;

[0040] in, and All are constants. This represents the distance difference between the current time and the target position at the previous time.

[0041] in, ;

[0042] in, The coefficient representing the impact of height deviation on system rewards. The height of virtual leadership; The desired preset flight altitude for the drone;

[0043] in, ;

[0044] in, The penalty coefficient for a swarm of drones entering the detection range of a radar cluster system. The penalty coefficient for a swarm of drones entering the dangerous zone of a radar cluster system; The preset danger zone for drones; This is the radius corresponding to the radar's maximum detection range; Position and radar for virtual leadership The distance;

[0045] in, ;

[0046] in, This represents the number of drones that are not malfunctioning in the drone swarm. This represents the number of drones that failed within the drone swarm. It is a constant greater than zero, used to adjust the degree of impact of failure on the entire drone swarm.

[0047] Of course, it is also necessary to avoid collisions between drone swarms and terrain, i.e. It can be guaranteed at all times, among which, This is a terrain obstacle space constructed from a digital elevation model using two-dimensional cubic convolution interpolation technology. This represents the outer envelope space of the drone swarm. When the two spaces intersect, it indicates that the drone swarm has collided with a terrain obstacle. In this case, the current round is stopped, and a new round is started for training.

[0048] ,

[0049] in, For the position of virtual leader, The collision radius of the drone swarm. To form points in the outer envelope space of a drone swarm, It is a three-dimensional real number space.

[0050] Step 2: Based on the control quantity of the virtual leader output by the converged proximal policy network, perform forward extrapolation based on the dynamic model to obtain the desired path of the virtual leader consisting of multiple waypoints.

[0051] The dynamic model of each drone in the drone swarm is as follows:

[0052] ;

[0053] in, , For drones speed, For the horizontal plane to Angle, Indicates drone The angle of inclination; For drones Location, , Indicates drone The azimuth angle; For drones Around The roll angle, It is the acceleration due to gravity. For drones Tangential overload; For drones Normal overload.

[0054] Step 3: Calculate the probability that the drone swarm will be detected by the radar cluster system at each waypoint on the virtual leader's expected path, and the number of newly added drone failures in the drone swarm.

[0055] When the probability of a drone swarm being detected by a radar cluster system is greater than a first preset threshold or the number of drone failures is greater than a second preset threshold, the drone formations at the previous waypoint are divided into two equal formation subgroups for flight, and the positions of each formation subgroup are adjusted so that the distance between the centroids of the two formation subgroups is greater than or equal to the maximum distance between two adjacent radar resolution cells.

[0056] When the probability of the drone swarm being detected by the radar cluster system is less than or equal to the first preset threshold, or the number of drone failures is less than or equal to the second preset threshold, the drone swarm will maintain its flight status at the previous waypoint.

[0057] Preferably, the first preset threshold is 0.5 and the second preset threshold is 10.

[0058] The method for calculating the number of newly added drone failures in step 3 is as follows:

[0059] First, determine whether the failure probability of each UAV is greater than a third preset threshold. If it is greater than the third preset threshold, the UAV is determined to have failed. The failure probability of a UAV refers to the probability that a UAV loses its ability to continue performing a mission due to being within the suppression range caused by continuous radar tracking. Preferably, the third preset threshold is 0.7.

[0060] The method for calculating the failure probability of a drone is as follows:

[0061] ,

[0062] in, The failure probability of the drone. The radius of the area that causes the failure. For drones Location, The duration for which a drone swarm is continuously tracked by radar; This refers to the current moment.

[0063] Since the relative positions of multiple drones within each drone formation are fixed, and the position of each drone formation is represented by its centroid, the overall formation geometry of the drone swarm can be changed by adjusting the distance between the centroids of different drone formations.

[0064] Step 3, dividing the drone swarm at the previous waypoint into two equal sub-swarms for flight, refers to equal-segment flight. The conditions for equal-segment flight are that the probability of the drone swarm being detected by the radar cluster system is greater than a first preset threshold, or the number of newly added failed drones in the swarm is greater than a second preset threshold. During equal-segment flight, the positions of each sub-swarm are adjusted so that the distance between the centroids of the two sub-swarms is greater than or equal to the maximum distance between two adjacent radar resolution cells. The method for calculating the maximum distance between two adjacent resolution cells of a single radar is as follows:

[0065] ,

[0066] in, For each radar, the elevation resolution unit is... This is the azimuth resolution unit for radar. This is the ranging resolution unit for radar.

[0067] and In comparison, the distance between two adjacent drones in a formation subgroup is much smaller. Therefore, if the distance between the centroids of any two formation subgroups is less than... Then they are within the same radar resolution unit. Conversely, if the distance between the centroids of the two formation subgroups is not less than... Therefore, they will be dispersed across multiple radar resolution cells. Thus, by dividing the flight into equal segments, the number of drones in each radar resolution cell is reduced, thereby decreasing the probability of the entire drone swarm being detected by radar.

[0068] If in step 3 there exists When there are several consecutive waypoints that do not meet the condition for equal flight division, the flight shall proceed according to the flight status of the waypoint preceding the one that previously met the condition for equal flight division. Preferably, .

[0069] Step 4: Based on the expected path of the virtual leader, the deviation vector between the centroid of each drone formation and the virtual leader, and the formation deviation vector between the centroid of each drone formation and the multiple drones it leads, obtain the expected path of each drone.

[0070] Among them, the expected location of the drone The calculation method is as follows:

[0071] ;

[0072] in, The rotation matrix is ​​constructed from the virtual leadership posture of the drone swarm; For the first The expected position of the centroid of a drone formation; For the first The first drone formation group The formation deviation vector of the individual drones relative to the centroid of the drone formation;

[0073] in, , For virtual leadership in the The expected position at each moment; The deviation vector between the centroid of each drone formation and the virtual leader.

[0074] in, It is determined based on the expected path of the virtual leader, specifically: determining the unit's direction based on the expected path of the virtual leader. , Used to map one-dimensional cumulative coordinates to a three-dimensional deviation vector. The desired displacement direction between adjacent waypoints is determined by the virtual leader, and can be orthogonalized by combining with a preset reference direction to ensure direction continuity; when the displacement of adjacent waypoints is too small or the direction degenerates, the direction of the previous waypoint or a preset fixed direction is used instead.

[0075] Example: In two simulation instances, five radars are randomly distributed in the environment, and the initial positions of the UAV swarm are randomly generated coordinates, as shown in Table 1. The reinforcement learning neural network structure is a fully connected neural network, with the output layer using the tanh activation function and the others using the ReLU activation function. The learning rate is 0.000125, the discount factor is 0.9, and the batch size during training is 64 samples. The flight time between two adjacent waypoints of the virtual leader is 1 second, and the maximum number of steps in a single training round is capped at 550. The training round ends if any of the following conditions are met: the maximum number of steps in a single round, the UAV swarm collides with the terrain, or reaches the target.

[0076] Table 1 Initialization Location Information

[0077]

[0078] Number of drones in a drone swarm in the simulation The radius of the area where the failure occurred The duration of continuous radar tracking of drone swarms Seconds, the radius corresponding to the radar's maximum detection range Preset drone danger zone Collision radius of drone swarm The expected preset flight altitude of the drone Drone speed constraints , , , , , , , , , The discount factor in reinforcement learning is 0.9, the first preset threshold is 0.5, the second preset threshold is 10, and the third preset threshold is 0.7.

[0079] from Figure 1 and Figure 3 As can be seen, the paths generated by each UAV in this embodiment remain feasible under complex terrain and radar swarm system threat distribution, and achieve planning from random initial position to target position. The UAV swarm avoids terrain obstacles in its direction of flight and traverses two radar zones to reach the target position.

[0080] from Figure 2 and Figure 4 It can be seen that the drone swarm passed through the terrain valley during its flight and crossed a radar area along the way to reach the target location.

[0081] Figure 5 and Figure 6 The horizontal axis represents time. The vertical axis represents the change commands for whether the drone formation flies in equal segments. . Used to characterize the degree of fragmentation of the drone swarm at the current moment; when An increase indicates that the drone swarm is flying in equal segments.

[0082] from Figure 5 As can be seen, the drone swarm with simulation number 1 performed equal-segment flight when its probability of being detected increased as it approached the radar center; when the drone swarm moved away from the radar center and left the radar detection area, and there were three consecutive waypoints that did not meet the equal-segment flight condition, it flew according to the flight state of the waypoint preceding the one that met the equal-segment flight condition. Figure 6 As can be seen from the simulation, the drone swarm with simulation number 2 only traversed one radar detection area along the way.

[0083] Figure 7 and Figure 8 The horizontal axis represents time. The vertical axis represents the number of drones that have not failed in the drone swarm. The step-like decline of the curve in the graph indicates that new failure events occurred within the corresponding time period, resulting in a segmented decrease in the number of non-failed events.

[0084] from Figure 7 and Figure 8As can be seen, failure events are mainly concentrated within the radar detection area, indicating a strong correlation between drone swarm survival and radar threat. Within the threat zone, dispersing drones reduces the number of drones clustered in the same resolution cell, thereby lowering the detection probability and the intensity of continuous tracking suppression, thus suppressing the frequency and cumulative number of failure events. Restoring formation while moving away from the threat zone ensures coordination efficiency and reduces unnecessary dispersal costs.

[0085] The simulation results of this embodiment show that the generated drone swarm trajectory can avoid risks by combining terrain and passable channels when approaching the radar threat area, or choose a relatively reasonable crossing method when crossing is unavoidable. This ensures that the target is reached while suppressing the probability of radar detection of the drone swarm and the number of drones that fail. It reduces the risk of failure caused by detection probability and continuous tracking, while improving the overall survivability of the swarm and the mission success rate.

Claims

1. A path planning method for UAV swarms in complex environments based on reinforcement learning, characterized in that, include: Step 1: Place a device with... The centroid of the drone swarm in a drone formation is defined as the virtual leader, and the current state of the virtual leader is used as the state input for reinforcement learning. The policy network and value network are trained using the proximal policy optimization algorithm to obtain the converged proximal policy network. Step 2: Based on the control quantity of the virtual leader output by the converged proximal policy network, perform forward extrapolation based on the dynamic model to obtain the desired path of the virtual leader consisting of multiple waypoints; Step 3: Calculate the probability that the drone swarm will be detected by the radar cluster system at each waypoint on the expected path of the virtual leader, and the number of newly added drones that fail in the drone swarm. When the probability of the drone swarm being detected by the radar cluster system is greater than a first preset threshold or the number of drone failures is greater than a second preset threshold, the drone formations of the drone swarm at the previous waypoint are divided into two formation subgroups for flight, and the positions of each formation subgroup are adjusted so that the distance between the centroids of the two formation subgroups is greater than or equal to the maximum distance between two adjacent radar resolution units. Step 4: Based on the expected path of the virtual leader, the deviation vector between the centroid of each drone formation and the virtual leader, and the formation deviation vector between the centroid of each drone formation and the multiple drones it leads, obtain the expected path of each drone. In step 1, the current state of the virtual leader includes the flight state and the environmental state. The flight state includes position, speed, tilt angle, azimuth angle, and roll angle. The environmental state includes the current terrain obstacle space, the relative distance and azimuth between the virtual leader and the target point, and the relative distance between the virtual leader and each radar. The constraints in step 1 when training the policy network and the value network are as follows: ; in, The probability that a swarm of drones will be detected by a radar cluster system. For the total number of radars, For the first radar, For radar Performance parameters, For the UAV center and radar in the resolution unit The distance between them For each radar in a radar cluster system Number of drones within the corresponding resolution unit The maximum value, For radar The number of drones in the corresponding resolution unit; This is the radar cross-section of a single UAV.

2. The path planning method for UAV swarms in complex environments based on reinforcement learning according to claim 1, characterized in that, The reward function for training the policy network and value network in step 1 is: ; in, For the reward function; The reward for the drone swarm reaching the target location; and All are weighting coefficients, through and The swarm of drones tends to either reach the target directly or evade radar detection. Rewards for drone swarms approaching the target location; A reward for drone swarms flying at a preset altitude; The penalty for being detected by a radar cluster system; Penalty for the total number of failed drones in the drone swarm; in, ; in, For virtual leadership Distance to the target location , The location of the target; For virtual leadership Location; The percentage of failed drones in a drone swarm. ; This represents the total number of failed drones. The total number of drones in the drone swarm; in, ; in, and All are constants. This represents the distance difference between the current time and the target position at the previous time. in, ; in, The coefficient representing the impact of height deviation on system rewards. The height of virtual leadership; The desired preset flight altitude for the drone; in, ; in, The penalty coefficient for a swarm of drones entering the detection range of a radar cluster system. The penalty coefficient for a swarm of drones entering the dangerous zone of a radar cluster system; The preset danger zone for drones; This is the radius corresponding to the radar's maximum detection range; Location and radar for virtual leadership The distance; in, ; in, This represents the number of drones that are not malfunctioning in the drone swarm. This represents the number of drones that failed within the drone swarm. It is a constant greater than zero, used to adjust the degree of impact of failure on the entire drone swarm.

3. The path planning method for UAV swarms in complex environments based on reinforcement learning according to claim 2, characterized in that, The method for calculating the number of newly added failed drones in the drone swarm in step 3 is as follows: First, determine whether the failure probability of each drone is greater than a third preset threshold. If it is greater than the third preset threshold, the drone is determined to have failed. The method for calculating the failure probability of the drone is as follows: ; in, The failure probability of the drone. The radius of the area that causes the failure. For drones Location, This refers to the duration during which the drone swarm is continuously tracked by radar. This refers to the current moment.

Citation Information

Patent Citations

  • Quadrotor attitude trajectory control method using PPO algorithm based on dimension cutting

    CN113885549A

  • Multi-unmanned-aerial-vehicle-to-multi-target cooperative air combat maneuver decision-making method based on reinforcement learning

    CN117032300A