Unmanned aerial vehicle cluster attack task planning method based on hierarchical fusion framework
By employing clustering, genetic, and reinforcement learning methods within a hierarchical fusion framework, the problem of task allocation in complex combat scenarios for UAV swarms was solved, enabling efficient task planning and real-time decision-making, and improving the mission efficiency of UAV swarms.
Patent Information
- Application Number
- CN202511009472.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-18
AI Technical Summary
Existing UAV swarm task allocation methods suffer from high computational complexity, insufficient dynamic adaptability, low efficiency in multi-objective optimization, and weak real-time decision-making capabilities in complex combat scenarios, and lack a hybrid algorithm framework that combines global optimization and dynamic decision-making.
A hierarchical fusion framework is adopted. The top-level task planning uses clustering and genetic algorithms to divide the battlefield into multiple task groups, while the bottom-level real-time decision-making is carried out by reinforcement learning. The multi-agent reinforcement learning model is combined to optimize the drone swarm strike mission.
It significantly reduces the complexity of reinforcement learning steps, improves the robustness and real-time decision-making capabilities of the algorithm, adapts to complex battlefield environments, and enhances the mission efficiency of UAV swarms.
Smart Images

Figure CN120975443A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of large-scale unmanned aerial vehicle cluster intelligence attack, and particularly relates to an unmanned aerial vehicle cluster attack task planning method based on a hierarchical fusion framework. BACKGROUND
[0002] In the early development of unmanned systems, various fields and departments developed a large number of special unmanned equipment according to their own needs, but as the complexity of modern war increases, it has become an important development trend to form a "combat group" composed of multiple types of heterogeneous unmanned systems to realize cross-domain collaborative operations. It is difficult to meet the needs of agile and efficient combat design under the conditions of joint operations by relying on traditional manual task planning methods and lacking effective auxiliary decision-making means support. Therefore, it is necessary to study a targeted cluster collaborative task planning method to uniformly plan and schedule the cooperative combat of unmanned systems. In recent years, unmanned aerial vehicle cluster combat has become an important development direction of modern military technology, and the core challenge lies in how to efficiently complete the scheduling of large unmanned aerial vehicle clusters and multi-task oriented decision-making.
[0003] In future battlefields, unmanned aerial vehicle clusters will fully exert their advantages of low cost, large quantity, and high intelligence, and will be an effective asymmetric means to attack high-value targets. In the research of unmanned aerial vehicle cluster coordination technology, scholars and institutions in Europe and the United States have launched a large number of unmanned cluster application research projects. DRAPA proposed the automatic hybrid control (MICA) project, aiming to improve the autonomy and coordination ability of unmanned aerial vehicle systems. It also includes multi-unmanned aerial vehicle control architecture, multi-unmanned aerial vehicle autonomous formation control method, unmanned aerial vehicle cooperative combat modeling and simulation technology, etc., and its goal is to realize the control of large-scale unmanned combat aircraft formation by fewer operators; the wide area search munitions (WASM) project aims at the key technologies of unmanned aerial vehicle reconnaissance / attack integration, and proposes a four-level cooperative control architecture of unmanned aerial vehicle cluster cooperative control distributed system based on formation-based task allocation, task coordination, path planning and trajectory control, to realize hierarchical regulation and control for large-scale complex task scenarios.
[0004] Traditional task allocation methods mainly include optimization methods and heuristic methods. The optimization methods mostly belong to deterministic search algorithms, which have been strictly theoretically proved, and the same algorithm can obtain the same solution at different times. However, due to the fact that the solving complexity of the multi-unmanned aerial vehicle task allocation problem increases significantly with the increase of the number of unmanned aerial vehicles and targets in the scene, the calculation cost and time consumption are greatly increased. The heuristic methods mostly belong to random search methods, which do not try to traverse the entire search space to obtain the optimal solution, but compromise between the solution quality and the calculation time to obtain a feasible solution of the optimization problem within an acceptable time. In order to solve the problems caused by the above two methods, some scholars propose to use reinforcement learning method to solve the multi-unmanned aerial vehicle decision problem. Reinforcement learning divides the complex decision problem into multiple stages for decision in sequence. Under the premise that the transition equation between the current stage decision and the next stage state is determined, the optimal decision chain is solved by using recursive method. The reinforcement learning algorithm based on this has good performance in discrete environment, but has problems such as insufficient dynamic adaptability, low multi-objective optimization efficiency and weak real-time decision ability when facing complex unmanned aerial vehicle combat scenes. Although the existing reinforcement learning algorithm can realize real-time decision, it lacks a global task allocation framework. Therefore, a hybrid algorithm framework combining global optimization and dynamic decision is urgently needed to improve the task efficiency of the unmanned aerial vehicle group in complex combat scenes.
[0005] The present application proposes a "clustering algorithm, genetic algorithm and reinforcement learning hierarchical fusion" unmanned aerial vehicle task allocation and decision framework to solve the problem of intelligent attack of unmanned aerial vehicle cluster, and the innovation is embodied in that: the clustering algorithm and the genetic algorithm are used in the top task planning, the battlefield with large scale is divided into multiple task groups with small scale through multi-objective optimization, and a global task allocation scheme is generated; each unmanned aerial vehicle makes decision based on the reinforcement learning and the top allocation scheme in the bottom real-time decision, and the specific action of each simulation step of the unmanned aerial vehicle cluster is formed. Compared with the single-layer decision method, the complexity of the reinforcement learning step is reduced, the dimension of the observation space and the action space is significantly reduced, the convergence is easier, and the allocation scheme of the top layer also guides the bottom layer, thereby improving the robustness of the algorithm. SUMMARY
[0006] In order to solve the above problems, the present application proposes a task planning method for attacking unmanned aerial vehicle cluster based on hierarchical fusion framework, which comprises the following steps:
[0007] Step 1, importing the unmanned aerial vehicle attack task scene information, including the initial position, speed and target threat degree of the unmanned aerial vehicle and the task target; confirming the unmanned aerial vehicle attack task target;
[0008] Step 1.1, in the form of red and blue confrontation, establish a UAV cluster attack task scene, red party is a UAV cluster, blue party is a plurality of attack targets with air defense capability. In the UAV cluster attack task described in the application, the UAV attack and target air defense model adopts a simplified probability model, that is, when the red UAV enters the air defense range of the blue target, the target will launch air defense weapons to attack it, and will destroy it according to a certain probability; when the blue target enters the red target's bomb range, the UAV will bomb it. After the red and blue parties launch weapons, they will enter a certain cooling time, and cannot attack any target during the cooling time. When the total threat degree of the red UAV cluster destroying the blue target exceeds a given threshold threat max , the task is considered to be completed, and the red party wins; when all the red UAVs in the combat area are destroyed, the task fails; in addition, when the red UAV cluster completely leaves the battlefield range or the combat time exceeds the specified maximum combat time, the task is forced to fail.
[0009] Step 2, according to the position information of the attack target, the task target is divided into a plurality of task groups based on the K-means++ algorithm;
[0010] Step 2.1, first construct a K-means++ clustering algorithm model. In the K-means++ algorithm, each task target group obtained by clustering is called a task cluster, and it is assumed that the cluster population is marked as (c1, c2,..., c k ), where k is the number of task clusters obtained by clustering. Each attack target to be clustered is marked as (x1, x2,..., x m ), where m is the number of blue attack targets. The goal of the algorithm is to minimize the sum of squared errors (SSE) of the distance from each attack target to the cluster center of the task cluster to which it is assigned , and the calculation method of SSE is as follows:
[0011]
[0012] Where x i ∈c i is the target in the task cluster c i , corresponding to the geographic position coordinates of each blue attack target. Since the number of cluster centers, that is, the number of task clusters k, is uncertain at this time, set k=2 when clustering for the first time.
[0013] Step 2.2, after determining k, randomly select a position of a blue attack target as the first initial cluster center u0, and calculate the distance D(x) from other attack targets to u0. The remaining k-1 initial cluster centers u1, u 2, …u k-1 positions are selected from the remaining n t- 1 roulette wheel method is used to select the target position, and the probability of each target position being selected as the next cluster center is:
[0014]
[0015] Step 2.3, after determining all the initial cluster centers, for each blue target position, calculate the distance to all cluster centers and classify it into the cluster corresponding to the nearest cluster center. Then for each cluster, use the average of all sample position coordinates in the cluster as the new cluster center. After that, repeat this step until the cluster center no longer changes or the set number of iterations is reached.
[0016] Step 2.4, after using the K-means++ clustering algorithm to get the clustering result when k = 2, it is necessary to judge whether the clustering effect of k = 2 is good enough through two conditions. Whether the selection of k is appropriate needs to consider whether the number of each group of attack targets is reasonable and whether the cluster center position is appropriate. The former mainly affects the observation space dimension of the later reinforcement learning algorithm, and the judgment condition and value are determined according to the efficiency of the selected reinforcement learning algorithm and the computing power of the program running device. The latter uses the elbow method to determine the standard, that is, calculate the sum of squared errors (SSE) of the distance from each attack target to the cluster center of the task cluster to which it is assigned. When the number of clusters does not reach the optimal number k best , with the increase of the number of clusters, SSE decreases rapidly. After reaching the optimal number of clusters k best , SSE will become slow. In order to determine the value of k best , it is necessary to increase the value of k and repeat step 2.3 and step 2.4, record each k value and the corresponding SSE value and draw a graph. In this image, the place where the slope changes the most and the number of attack targets in each group is reasonable is the optimal k value.
[0017] Step 3, based on the distributed parallel genetic algorithm, assign tasks to each UAV, and divide each UAV corresponding to each task target into the task group to which the task target belongs;
[0018] Step 3.1, build a distributed parallel genetic algorithm model, the core of the distributed genetic algorithm is the parallel processing of the basic genetic algorithm, and multiple basic genetic algorithms are dispersed to each red unmanned aerial vehicle. The fitness function construction and selection process are consistent with the basic genetic algorithm. In this genetic algorithm model, each chromosome in the population is a two-dimensional table of h rows and n columns. Among them, n is the total number of red unmanned aerial vehicles, and h is the maximum number of targets that each red unmanned aerial vehicle can destroy + 1, which depends on the weapon load of the red unmanned aerial vehicle. The first row of each column is the number of the corresponding red unmanned aerial vehicle, and the remaining rows represent the targets that need to be attacked by the unmanned aerial vehicle. When initializing the population, it is necessary to ensure that each target is attacked by an unmanned aerial vehicle, and each unmanned aerial vehicle participates in attacking the target. If the actual number of targets that a certain unmanned aerial vehicle participates in attacking is less than m-1, then the corresponding row of the column is assigned a value of-1, representing that the unmanned aerial vehicle has no task at this time.
[0019] When designing the fitness function, it is considered that the task planning layer as the upper layer design should more comprehensively consider the combat effectiveness of the unmanned aerial vehicle cluster attacking multiple targets, therefore, the weighted sum of the reciprocal of the total range of each unmanned aerial vehicle completing the attack task, the total time of reaching the target point, and the total threat degree of the enemy target is selected as the final fitness function.
[0020]
[0021] Among them, s i represents the total range required by each red unmanned aerial vehicle to complete all tasks, v i represents the speed of each unmanned aerial vehicle, Threat i is the total threat value of the task target assigned to each unmanned aerial vehicle, and λ1, λ2 are correction coefficients.
[0022] Step 3.2, initialize the population of the genetic algorithm, randomly generate the initial population, each individual represents a task allocation scheme, ensure that each blue target is assigned to at least one red unmanned aerial vehicle, and each red unmanned aerial vehicle is assigned to at least one blue target for attack. When initializing the population, first assign the first row of each column to the unmanned aerial vehicle number, then select a task target for each unmanned aerial vehicle and assign it to the second row of each column to ensure that each unmanned aerial vehicle has a target to attack; then, for other task targets of each unmanned aerial vehicle, first randomly sort the remaining task targets, then sort the positions of the remaining unfilled task targets, and select the positions equal to the number of remaining task targets to arrange the task targets. Finally, assign the positions of the unfilled task targets to-1. Finally, verify whether each individual meets the basic requirements of task allocation, such as the weapon load quantity limit of the unmanned aerial vehicle and the uniqueness of task target allocation. If not, regenerate the individual. After initialization, calculate the fitness F(i) of all chromosomes in the current population.
[0023] Step 3.3, for the selection of good chromosomes, a proportional selection strategy is adopted, that is, the probability of each chromosome being selected for genetic algorithm is the ratio of the fitness value of the chromosome to the sum of the fitness values of all individuals in the population. Let the fitness of chromosome c be F(c), and the chromosome population size be N, then the probability of this chromosome being selected is:
[0024]
[0025] From the above formula, the greater the fitness value, the greater the selection probability. Then, according to the above probability, N times are drawn from the chromosome population, and the selected results are the offspring chromosomes.
[0026] Step 3.4, the offspring chromosomes are subjected to crossover and mutation operations. For the chromosome in the form of a two-dimensional table of h rows and n columns described in the present application, the crossover operation adopts two-point crossover and crossover operation based on specified target. Two-point crossover is a standard genetic algorithm crossover operation, specifically, it randomly selects two crossover points in two chromosomes, and then exchanges the gene fragments between the two crossover points; the crossover based on the specified target means that a task target is randomly selected in two chromosomes, and then the corresponding red unmanned aerial vehicle is crossed; the mutation operation adopts task ordering mutation and task block mutation. Among them, the task ordering mutation is to randomly disturb the order of the task targets attacked by one or more unmanned aerial vehicles in the chromosome, and the task block mutation is to select a number of unmanned aerial vehicles corresponding to the columns and the corresponding attack targets in these columns, and then reassign these attack targets in these columns. After completing the crossover and mutation, finally verify whether each individual satisfies the unmanned aerial vehicle weapon load quantity limit, task target allocation uniqueness, etc. If not, re-execute the crossover or mutation operation that does not meet the requirements.
[0027] Step 3.5, repeat steps 3.3 and 3.4 until the number of iteration rounds is reached or the fitness value remains stable in continuous generations, or the preset threshold is reached, which can be considered as the iteration of the genetic algorithm meeting the stopping criterion, and the iteration stops. Select the individual with the highest fitness value from the final population as the task allocation scheme of the unmanned aerial vehicle. Then, according to the task allocation scheme of the optimal individual, each task target is allocated to the corresponding unmanned aerial vehicle, and the unmanned aerial vehicle is divided into the task group to which the task target belongs.
[0028] Step 4, construct a multi-agent reinforcement learning model, regard each unmanned aerial vehicle as an agent, construct a simulation environment model, an observation space dimension, an action space dimension, and a reward function;
[0029] Step 4.1: Construction of the UAV swarm attack simulation environment includes the motion model of the red team's UAVs, the attack model of the red team's UAVs, and the air defense model of the blue team's target. This invention uses a three-degree-of-freedom dynamics and kinematics model with a relatively fast solution speed to model the motion of the red team's UAVs, establishing a model with the UAV's center of gravity O as the reference point. b Body coordinate system O with origin b x b y b z b Among them, O b x b The axis is located within the plane of symmetry of the UAV, parallel to the fuselage axis, and points towards the nose; O b z b The axis is also located within the plane of symmetry of the UAV and is perpendicular to O. b x b Axis and downward; O b y b The axis is perpendicular to the plane of symmetry of the UAV and points to the right according to the right-hand rule. The kinematic and dynamic model of the three-degree-of-freedom UAV is as follows:
[0030]
[0031]
[0032] Where (x,y,z) represents the spatial coordinates of the UAV, based on the UAV's center of gravity O. b Body coordinate system O with origin b x b y b z b 。 ;n x and n z These represent the longitudinal and normal overloads of the drone, respectively; V represents the drone's velocity; and g is the acceleration due to gravity. This indicates the attitude angle of the drone.
[0033] Step 4.2: First, construct the state space as the input parameters for the decision network. To simplify the state space and prevent invalid output parameters, this invention selects only battlefield information that influences the UAV swarm's action selection to construct the state space. The local observations S of each UAV are combined into joint observation information O after data deduplication. The specific form of the local observations S is as follows:
[0034]
[0035] The state space is divided into two parts: the first part is the current state of the UAV, and the second part is the current state of the target. Wherein, (ID, x, y, z, h, ID) t) is the state information of a UAV, which is the number, three-dimensional position coordinates, survival state and the first target number assigned in step 3 in turn; ) is the state information of a target, which is the number, three-dimensional position coordinates, survival state and threat degree in turn. max ) is the number of UAVs in the task group with the largest number of UAVs of the red side obtained in step 3; m max ) is the number of targets in the task group with the largest number of targets obtained in step 3. If the number of UAVs in a task group is less than n max or the number of targets is less than m max , the position of the observation space vacated is assigned as F NULL = 10 9 to ensure that invalid data is suppressed in gradient calculation.
[0036] Step 4.3, then, the action space is established as the output of the decision network. The action space of each red UAV is designed as A = {θ desire , ψ desire}, wherein θ desire is the desired pitch angle; and ψ desire is the desired yaw angle.
[0037] Step 4.4, in the process of building the reinforcement learning model, a corresponding reward function needs to be designed to evaluate the pros and cons of the generated action strategy scheme. At the same time, the reward function also needs to reflect the guiding role of the top-level task allocation result on the bottom-level real-time decision, that is, to guide the UAV to attack the task target assigned in the task allocation layer. The specific reward function R is in the form of:
[0038]
[0039] is the reward obtained by a single UAV in this step action. Wherein, is the threat degree value of the target destroyed by the UAV in the current step; bonus is the upper-level decision reward factor, which rewards if the UAV is closer to the top-level task allocation target than the previous step, or successfully destroys the top-level task allocation target in the current step; penalty is the punishment for the failure of the UAV, which should be deducted from the reward if the UAV is destroyed or exceeds the predetermined battlefield range. γ is a time correction factor, the goal of which is to prevent the UAV from completing the task too late, and its calculation method is:
[0040]
[0041] , wherein T current is the time experienced by the UAV from the start of the task to the completion of the target attack, and T maxThe maximum permissible mission time for a drone swarm is determined by the battlefield size, drone speed, and maximum flight time.
[0042] Step 5: Train the multi-agent reinforcement learning network. If the mission target is attacked or the drone malfunctions during the training process, return to Step 3 to re-divide the task groups;
[0043] Step 5.1: Before starting the training of the multi-agent reinforcement learning algorithm, initialization is performed. Specifically, all drones are treated as independent agents, and their environmental states are initialized. Each drone's local observations S are deduplicated to form joint observation information O. Initialization also requires obtaining the initial state's joint observation information O0. Simultaneously, replay pools D and D' are constructed and initialized. These pools will store the experiential data generated during the interaction between the drones and the environment for subsequent experience replay and learning. Furthermore, the multi-agent reinforcement learning network also needs to be initialized. After initialization, training begins.
[0044] Step 5.2: The period from the start of the task to its success or failure is considered a training cycle. Within each training cycle, an iteration is performed every 0.1 seconds as a training step. In the t-th training step, the i-th red team UAV, based on the current joint observation O... t Use its own strategy π i :O i ×A i →[0,1] Select the action to be performed in this step, a i The individual actions of all agents constitute a joint action u∈A.
[0045] Step 5.3: Then, all agents interact with the environment through joint actions to obtain corresponding rewards and the next joint observation state. R(O,u) represents the reward function that all agents can obtain after taking joint action u under joint observation information O. The reward value r can be obtained according to this function. t Then execute the joint action u t To obtain the next step, local observations S for each UAV. t+1 and joint observation information O t+1 The state transitions generated during this interaction process will be applied to {O}. t ,u t ,r t O t+1 Store it in playback pool D.
[0046] Step 5.4, after completing the reward function calculation, if the current training step has a red UAV destroyed or out of bounds resulting in failure, or a blue strike target successfully destroyed by a UAV, it is considered that the current environment has changed, and step 3 is returned to run the genetic algorithm to reassign the task target for the current UAV that is not failed, and redivide the task group according to the new task target.
[0047] Step 5.5, if the task succeeds or fails at this time, the trajectory generated in the entire training period (containing a series of state transition pairs) is stored in the replay pool D'. Then, from the replay pools D and D', the parameters ψ, θ, are updated. In this way, the network can learn the rules of obtaining rewards under different states and actions, thereby continuously optimizing the policy function and value function of the agent and improving the decision-making ability of the agent in a complex environment. If the training has not ended at this time, return to step 5.1 to reinitialize the environment and perform the next training cycle of training.
[0048] Step 6, repeat step 5 to perform multiple rounds of red UAVs and blue strike targets, continuously interact with the agent and the environment, collect data and update the network, and observe the reward function at each step of training. If the number of training steps reaches the set upper limit, or the reward function converges and exceeds a set threshold, the training is completed, and the trained reinforcement learning network model is extracted.
[0049] The advantages and beneficial effects of the present application are that the present application aims at the problem of UAV cluster strike task planning in complex battlefield environment, combines the constraint conditions of UAV combat, constructs an efficient strike task planning model, and finally outputs the pitch angle and yaw angle of each step of UAV, realizes full-process planning simulation. This method comprehensively considers the combat efficiency and survivability of UAV, and effectively deals with the complexity and uncertainty of the battlefield environment through flexible adjustment of task allocation and real-time decision mechanism. This method significantly reduces the dimension of the observation space and the action space, making it easier to converge, and the top-level allocation scheme will also guide the bottom layer. In addition, the application of this method in the field of UAV cluster strike task planning provides a new idea and technical support for task allocation problems in other complex environments. This method is also applicable to other fields that require task allocation, such as intelligent transportation and logistics distribution, and has wide applicability and promotional value. BRIEF DESCRIPTION OF DRAWINGS
[0050] The above-mentioned features of the present application can be more clearly understood by referring to the embodiments described in detail with reference to the accompanying drawings.
[0051] Figure 1 A flowchart of a UAV cluster strike task planning method based on a hierarchical fusion framework.
[0052] Figure 2 Flow chart for clustering targets of attack based on K-means++ clustering algorithm.
[0053] Figure 3 Target grouping result of attack based on K-means++ clustering.
[0054] Figure 4 Flow chart for running distributed parallel genetic algorithm.
[0055] Figure 5 Initial distributed parallel genetic algorithm task planning result.
[0056] Figure 6 Flow chart for reinforcement learning training of multi-agent reinforcement learning algorithm. DETAILED DESCRIPTION
[0057] The present application will be explained in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are used to explain and illustrate the present application, but are not used to limit the scope of the present application.
[0058] As shown in Figure 1 , the present embodiment takes the attack of enemy ground target group by unmanned aerial vehicle (UAV) cluster as an example, and proposes a UAV cluster attack task planning method based on hierarchical fusion framework. The specific method is shown below.
[0059] Step 1, import the UAV attack task scene information, including the initial position, speed and target threat degree of UAV and task target; confirm the UAV attack task target.
[0060] In the present embodiment, the UAV cluster attack task scene is established in the form of red and blue confrontation. The red side is the UAV cluster, and the blue side is a plurality of attack targets with air defense capability. The red side contains 5 fixed-wing UAVs with attack capability, which start from the same position; the blue side contains 9 dispersed ground attack targets. The red UAV approaches and attacks the target to make it invalid, which is regarded as completing the attack task.
[0061] In this process, the blue attack target will carry out air defense counterattack. In the present embodiment, the UAV attack and target air defense model adopts a simplified probability model. When the red UAV enters the air defense range of the blue target, the target will launch air defense weapons to attack it, and will destroy it with a certain probability; when the blue target enters the bomb range of the red target, the UAV will launch bombs to destroy it. The hit rate of each blue target air defense counterattack is multiplied by 100 as the threat degree of the target. After the red and blue sides launch weapons, they will enter a certain cooling time, and cannot attack any target during the cooling time. When the total threat degree of the blue target destroyed by the red UAV cluster exceeds a given threshold threat maxAfter, as the task is completed, the red side wins; when all the red side UAVs in the combat area are destroyed, the task fails; in addition, when the red side UAV cluster completely leaves the battlefield range or the combat time exceeds the specified maximum combat time, the task is forced to fail.
[0062] Step 2, according to the position information of the blue side attack targets, the K-means++ algorithm is used to divide them into multiple task groups. In this embodiment, the process of clustering the above-mentioned 9 ground targets using the K-means++ clustering algorithm is as shown in Figure 2
[0063] In the first clustering, set k = 2, randomly select a position of an attack target as the first initial clustering center u0, and calculate the distance D(x) of other attack targets and u0. The position of another initial clustering center u1 is extracted from the remaining 36 attack targets using the roulette method, and the probability of each target position being selected as the next clustering center is:
[0064]
[0065] After determining all the initial clustering centers, for each attack target position, the distance to all clustering center points is calculated, and it is classified into the cluster corresponding to the nearest clustering center. Then, for each cluster, the average value of all position coordinates of all samples in the cluster is used as the new clustering center. Then, the step is repeatedly executed until the clustering center no longer changes or the set iteration number is reached.
[0066] After obtaining the clustering results of k = 2, 3, 4, and 5 respectively according to the above method, the sum of squared errors (SSE) of the distance of each blue side attack target to the clustering center of the task cluster to which it is assigned is calculated, and each k value and the corresponding SSE value is recorded. It is found that the optimal k value is k = 3 when the slope change is the largest and the number of attack targets in each group is reasonable, and three task target groups each containing 3 attack targets are obtained as shown in Figure 3
[0067] Step 3, based on the distributed parallel genetic algorithm, each UAV is assigned a task, and the UAV corresponding to each task target is divided into the task group to which the task target belongs. In this embodiment, the process of using the distributed parallel genetic algorithm to assign tasks to the above-mentioned 5 UAVs is as shown in Figure 4
[0068] In the genetic algorithm model, each individual in the population is a 4-row 5-column two-dimensional table. The column number represents the total number of macro unmanned aerial vehicles, and the row number represents the number of each unmanned aerial vehicle itself and the maximum number of three blue strike targets that can be attacked. When initializing the population, first, a task target is selected for each unmanned aerial vehicle to ensure that each unmanned aerial vehicle has a target that can be attacked; then, for the second and third task targets of each unmanned aerial vehicle, first, the remaining task targets are randomly sorted, and then the positions of the remaining unfilled task targets are randomly selected and arranged. Finally, the positions of the unfilled task targets are assigned to -1, representing that the unmanned aerial vehicle has no task at this time. After initialization, the fitness F(i) of all chromosomes in the current population is calculated, and the calculation method is as follows:
[0069]
[0070] where s i represents the total distance required for each unmanned aerial vehicle to complete all tasks, v i represents the speed of each unmanned aerial vehicle, Threat i is the total threat value of each unmanned aerial vehicle assigned to the task target, λ1 and λ2 are correction coefficients, in this embodiment, λ1 = λ2 = 0.3, and the solution of the distance s i is expressed by Dubins curve. The Dubins library in Python programming is used to generate the path from the starting point to the task target and between the task targets, and the length of the path is output.
[0071] Then, selection operation is performed according to the fitness of all chromosomes. Let the fitness of chromosome i be F(i), and the chromosome population size be N, then the probability of being selected is:
[0072]
[0073] As can be seen from the above formula, the greater the fitness value, the greater the selection probability. Then, N times are selected in the chromosome population according to the above probability, and the selected result is the offspring chromosome. Then, the offspring chromosome is subjected to crossover and mutation operations. The crossover operation includes two-point crossover and crossover operation based on specified targets, and the mutation operation includes task sorting mutation and task block mutation operation. When performing the crossover and mutation operations of the genetic algorithm, the probabilities of selecting the above two crossover operations and two mutation operations are equal. After completing the crossover and mutation, finally, it is verified whether each individual satisfies the weapon load quantity limit of the unmanned aerial vehicle, the uniqueness of task target allocation, etc. If not, re-execute the crossover or mutation operation that does not meet the requirements.
[0074] In this embodiment, through repeated selection, crossover, and mutation operations, the fitness value of the population will eventually stabilize over multiple generations, satisfying the stopping criterion, and the iteration will stop. The individual with the highest fitness is selected from the final population as the task allocation scheme for the UAVs. After the distributed genetic algorithm completes the task allocation for each UAV, the UAVs are assigned to each task group according to the task group of the first target each UAV attacks. The distributed genetic algorithm task planning result obtained in this embodiment is as follows: Figure 5 As shown:
[0075] like Figure 5 As shown, the task groups for targets 5 and 9 were assigned to drone 3; the task groups for targets 1, 3, 7, and 8 were assigned to drones 2 and 4; and the task groups for targets 2, 4, and 6 were assigned to drones 1 and 5. The three diagrams also include the path and order in which each Red Team drone attacked its targets, with drones closer to their launch point being prioritized for attack. If a drone attacks targets from multiple task groups during this phase, the task group of the first target it attacks will be considered its assigned task group.
[0076] Step 4: Construct a multi-agent reinforcement learning model, treating each drone as an agent. Construct the observation space dimension based on the task group with the most objectives; construct the action space dimension based on the task group with the most drones. Construct the reward function.
[0077] Step 4.1: In this embodiment, a three-degree-of-freedom dynamics and kinematics model with faster calculation speed is used to model the motion of the fixed-wing UAV. The three-degree-of-freedom UAV kinematics and dynamics model is as follows:
[0078]
[0079] Where (x,y,z) represents the spatial coordinates of the UAV, based on the UAV's center of gravity O. b Body coordinate system O with origin b x b y b z b 。 ;n x and n z These represent the longitudinal and normal overloads of the drone, respectively; V represents the drone's velocity; and g is the acceleration due to gravity. This indicates the attitude angle of the drone.
[0080] Then, a state space is constructed as the input parameters for the decision network. To simplify the state space and prevent invalid output parameters from the network, this invention selects only battlefield information that influences the UAV swarm's action selection to construct the state space. The specific form of the state space S is as follows:
[0081]
[0082] The state space is divided into two parts, the first part is the current state of the UAV, and the second part is the current state of the attack target. Among them, (ID, x, y, z, h, ID t ) is the state information of a UAV, in turn, the number, three-dimensional position coordinates, survival state and the first task target number allocated in step 3; is the state information of an attack target, in turn, the number, three-dimensional position coordinates, survival state and threat degree. n max is the number of UAVs contained in the task group with the largest number of UAVs, which is 13 in this embodiment; m max is the number of targets contained in the task group with the largest number of targets, which is 17 in this embodiment. In this embodiment, the number of UAVs of the first and second task groups is less than n max or the number of targets is less than m max , then the positions vacated by the observation space of the two groups are uniformly assigned as F NULL = 10 9 to ensure that invalid data is suppressed in gradient calculation.
[0083] Then, the action space is established as the output of the decision network. The action space of each intelligent agent is designed as A = {θ desire , ψ desire}, wherein θ desire is the expected pitch angle; and ψ desire is the expected yaw angle.
[0084] In this embodiment, the reward function R is in the form of:
[0085]
[0086] is the reward obtained by a single UAV in this step action. Among them, is the threat degree value of the target destroyed by the UAV in the current step; bonus is the upper-level decision reward factor, if the UAV is closer to the top-level task allocation target than the previous step, or successfully destroys the top-level task allocation target in the current step, the UAV is rewarded; penalty is the punishment of the UAV failure, if the UAV is destroyed or exceeds the predetermined battlefield range, the reward should be deducted. γ is a time correction factor, the goal is to prevent the UAV from completing the task too late, and its calculation method is:
[0087]
[0088] Among them, T current is the time experienced by the UAV from the beginning of the task to the completion of the target attack, Tmax The maximum task allowable time for the UAV cluster is determined by the battlefield size, the UAV speed, and the maximum endurance.
[0089] Step 5, multi-agent reinforcement learning network training is performed. If the task target is hit or the UAV fails in the training step, return to step 3 to redivide the task group;
[0090] In this embodiment, the MAPPO algorithm is selected as the multi-agent reinforcement learning algorithm model, and the training process is as shown in Figure 6 To further speed up the training, the multi-threading is used in the sampling process of the interaction between the UAV cluster and the environment, that is, all three groups simultaneously interact. In this way, the computing potential of the computer in the training process can be fully utilized, the training efficiency is improved, and the rewards and states of all combat groups can be observed in the training process, which helps to evaluate the training effectiveness.
[0091] Before the multi-agent reinforcement learning algorithm training starts, initialization is first performed. Specifically, all UAVs are regarded as independent agents, and the environment state in which they are located is initialized. The local observation S of each UAV is combined into joint observation information O after data deduplication, and the joint observation information O0 of the initial state also needs to be obtained during initialization. At the same time, the replay pool D and D' are constructed and initialized, which are used to store the experience data generated during the interaction between the UAV and the environment, so as to perform experience replay and learning subsequently. In addition, the multi-agent reinforcement learning network also needs to be initialized. After the initialization is completed, the training is started.
[0092] A training cycle is regarded as starting from the beginning of the task to the success or failure of the task, and iteration is performed according to every 0.1 second as a training step in each training cycle. In the tth training step, each agent i selects the action a t to be executed in this step according to the current joint observation O i using the self-strategy π i :O i ×A i , and the individual actions of all agents form a joint action u∈A. Then the joint action of all agents is interacted with the environment, so as to obtain the corresponding reward and the joint observation state of the next step. R(O,u) represents the reward function that can be obtained by all agents after taking the joint action u under the joint observation information O, and the reward value r t can be obtained according to the function; then the joint action u t is executed to obtain the local observation S t+1 of each UAV in the next step and the joint observation information O t+1 , and the state transition pair {O t , ut ,r t ,O t+1} into the replay pool D.
[0093] After the reward function calculation is completed, if no drone is destroyed or out of bounds, resulting in failure, or the target is successfully destroyed by the drone, it is considered that the current environment has changed, and step 3 is returned to run the genetic algorithm to reassign the task target for the current non-failed drone, and redivide the task group according to the new task target.
[0094] If the task has succeeded or failed at this time, the trajectory generated in the entire training period (containing a series of state transition pairs) is stored in the replay pool D'. Then, random sampling is performed from the replay pools D and D', and the multi-agent reinforcement learning network is updated. In this way, the network can learn the rules of obtaining rewards under different states and actions, thereby continuously optimizing the policy function and value function of the agent and improving the decision-making ability of the agent in a complex environment. If the training has not ended at this time, the environment is reinitialized, and the training of the next training period is performed.
[0095] In summary, the reinforcement learning training method used in the present embodiment is shown in the following pseudo code.
[0096]
[0097] Step 6, continue to interact between the agent and the environment, collect data and update the network, observe the reward function of each step of training, if the reward function converges and exceeds a certain threshold, the training is completed, and the intelligent decision-making network model is obtained. When the algorithm is deployed and used, only the action that can bring the most expected return is selected according to the learned policy function and value function, that is, the drone swarm as the agent can perform the action considered best by the decision-making model.
Claims
1. A method for planning UAV swarm strike missions based on a hierarchical fusion framework, characterized in that: Includes the following steps: Step 1: Import drone strike mission scenario information, including the initial position, speed, and target threat level of the drone and the mission target; confirm the drone strike mission target; Step 2: Based on the location information of the targets, divide the mission targets into multiple task groups using the K-means++ algorithm; Step 3: Assign tasks to each UAV based on a distributed parallel genetic algorithm, and assign the UAV corresponding to each task objective to the task group to which the task objective belongs; Step 4: Construct a multi-agent reinforcement learning model, treating each UAV as an agent, and build a simulation environment model, observation space dimension, action space dimension, and reward function; Step 5: Train the multi-agent reinforcement learning network; If the mission target is hit or the drone malfunctions during the training process, return to step 3 to regroup the mission groups. Step 6: Repeat Step 5 to conduct multiple rounds of confrontation between the red team's drones and the blue team's target, continuously carry out the interaction between the agent and the environment, data collection and network updates, and observe the reward function of each training step; if the number of training steps reaches the set upper limit, or the reward function converges and exceeds a set threshold, then the training is completed and the trained reinforcement learning network model is extracted.
2. The method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 1, characterized in that: In step 1, specifically: a drone swarm attack scenario is established using a red-blue adversarial approach. The red team represents a drone swarm, and the blue team represents multiple targets with air defense capabilities. In this drone swarm attack mission, a simplified probabilistic model is used for both drone attack and target air defense. Specifically, when a red team drone enters the air defense range of a blue team target, the target will launch air defense weapons to attack it, destroying it according to probability. When a blue team target enters the bombing range of a red team target, the drone will drop bombs to destroy it. Both the red and blue teams will enter a cooldown period after firing their weapons, during which they cannot attack any targets. When the total threat level of the red team's drone swarm destroying blue team targets exceeds a given threshold, the attack is initiated. max After that, the mission is considered complete and the Red Team wins; when all Red Team drones in the combat area are destroyed, the mission fails; in addition, when the entire Red Team drone swarm leaves the battlefield or the combat time exceeds the specified maximum combat time, the mission is forcibly failed.
3. The method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 1, characterized in that: In step 2, specifically: Step 2.1: First, construct the K-means++ clustering algorithm model. In the K-means++ algorithm, each group of task objectives obtained by clustering is called a task cluster, and the cluster group is labeled as (c1, c2, ..., c...). k ), where k is the number of task clusters that should be obtained from clustering; each target to be clustered is labeled as (x1, x2, ..., x). m ), where m is the number of targets attacked by the Blue Team; Step 2.2: After determining k, randomly select the location of a target attacked by the Blue Team as the first initial cluster center u0, and calculate the distance D(x) from other targets to u0; the remaining k-1 initial cluster centers u1, u2, u3, u4, u5, u6, u7, u8, u9, u1, u9, u1, u9, u1, u9, u1, u2 ... 2, …u k-1 Position from the remaining n t The target is selected using a roulette wheel method from -1 targets. The probability that each target location will be selected as the next cluster center is: Step 2.3: After determining all the initial cluster centers, for each blue team target location, calculate its distance to all cluster center points and classify it into the cluster corresponding to the nearest cluster center; then, for each cluster, use the average of all position coordinates of all samples in the same cluster as the new cluster center; then, repeat this step until the cluster centers no longer change or the set number of iterations is reached. Step 2.4: After obtaining the clustering results when k=2 using the K-means++ clustering algorithm, it is necessary to judge whether the clustering effect of k=2 is good enough by two conditions: whether the choice of k is appropriate should be considered in terms of whether the number of targets in each group is reasonable and whether the location of the cluster center is appropriate.
4. The method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 3, characterized in that: In step 2.1, the goal of the K-means++ clustering algorithm is to find the cluster center of each target to the task cluster to which it is assigned. The sum of squared errors (SSE) of the distances is minimized. The SSE is calculated as follows: Where, x i ∈c i For task cluster c i The targets in the map correspond to the geographical coordinates of each target that the Blue Team attacks.
5. The method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 1, characterized in that: In step 3, specifically: Step 3.1: Construct a distributed parallel genetic algorithm model. In this genetic algorithm model, each chromosome in the population is a two-dimensional table with h rows and n columns. Here, n is the total number of red team drones, and h is the maximum number of targets that each red team drone can destroy + 1, which depends on the weapon payload of the red team drone. The first row of each column is the red team drone number corresponding to that column, and the remaining rows represent the targets that the drone needs to attack. Step 3.2: Initialize the population for the genetic algorithm. Randomly generate an initial population, with each individual representing a task allocation scheme. Ensure that each blue team target is assigned to at least one red team drone, and each red team drone is assigned at least one blue team target to attack. During population initialization, first assign the drone number to the first row of each column. Then, select a target for each drone and assign it to the second row of each column, ensuring that each drone has a target to attack. Next, for the other targets of each drone, first randomly sort the remaining targets, then number the positions of the remaining unfilled targets, and randomly select positions with the same number of remaining targets to assign targets. Finally, assign -1 to the positions of the unfilled targets. Finally, verify whether each individual meets the task allocation requirements. Step 3.3: A proportional selection strategy is used for chromosome selection, meaning the probability of each chromosome being selected for the genetic algorithm is the ratio of its fitness to the sum of the fitness values of all individuals in the population. Let the fitness of chromosome c be F(c), and the chromosome population size be N. Then the probability of this chromosome being selected is: Following the above probability, N selections are made from the chromosome population, and the selected results are the offspring chromosomes. Step 3.4: Perform crossover and mutation operations on the offspring chromosomes; for the chromosomes in the form of a two-dimensional table with h rows and n columns, the crossover operation adopts two-point crossover and crossover operation based on a specified target; the mutation operation adopts two types: task sorting mutation and task block mutation; after completing the crossover and mutation, verify whether each individual meets the weapon payload quantity limit of the UAV and the uniqueness of the task target allocation; if not, re-execute the crossover or mutation operation that does not meet the requirements. Step 3.5: Repeat steps 3.3 and 3.4 until the number of iteration rounds or the fitness value remains stable over multiple generations, or a preset threshold is reached. At this point, the iteration of the genetic algorithm is considered to have met the stopping criterion, and the iteration stops. Select the individual with the highest fitness from the final population as the task allocation scheme for the UAV. Then, according to the task allocation scheme of the optimal individual, assign each task objective to the corresponding UAV and assign the UAV to the task group to which the task objective belongs.
6. The method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 5, characterized in that: In step 3.1, when designing the fitness function, considering the task planning layer as the upper layer design, the fitness function is selected as the reciprocal of the weighted sum of the total flight distance of each UAV to complete the strike mission, the total time to reach the target point, and the total threat level of the enemy target. Among them, s i ν represents the total distance required for each Red Team drone to complete all tasks. i Threat represents the speed of each drone. i The total threat value of the mission target assigned to each UAV, where λ1 and λ2 are correction coefficients.
7. The method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 1, characterized in that: In step 4, specifically: Step 4.1: The construction of the UAV swarm attack simulation environment includes the motion model of the red team's UAVs, the attack model of the red team's UAVs, and the air defense model of the blue team's targets; a three-degree-of-freedom dynamics and kinematic model is used to model the motion of the red team's UAVs; Step 4.2: Construct the state space as the input parameters of the decision network; In order to simplify the state space and prevent invalid output parameters of the network, only battlefield information that affects the UAV swarm's action selection is selected to construct the state space. Step 4.3: Establish the action space as the output of the decision network; design the action space of each red team UAV as A = {θ} desire ,ψ desire }, where θ desire ψ is the desired pitch angle. desire The desired yaw angle; Step 4.4: In the process of building a reinforcement learning model, it is necessary to design a corresponding reward function to evaluate the merits of the generated action strategy. At the same time, the reward function should reflect the guiding role of the top-level task allocation result on the bottom-level real-time decision-making, that is, guide the UAV to strike the task target assigned at the task allocation layer.
8. The method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 7, characterized in that: In step 4.2, the local observations S of each UAV are combined into joint observation information O after data deduplication; the specific form of the local observations S is as follows: The state space is divided into two parts: the first part is the current state of the drone, and the second part is the current state of the target. Among them, (ID,x,y,z,h,ID) t The information consists of the drone's ID, three-dimensional position coordinates, survival status, and the sequence number of the first assigned mission target. The status information for a target includes, in order: ID, three-dimensional location coordinates, survival status, and threat level; n max The number of drones in the task group with the largest number of Red Team drones; m max The number of blue team objectives contained in the task group that yields the most objectives.
9. A method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 7, characterized in that: In step 4.4, the reward function R is in the form of: The reward for a single drone's contribution to this step of the operation; whereby, The threat level of the target destroyed by the drone in the current step is represented by: `bonus` is the reward factor for higher-level decision-making; a reward is given if the drone is closer to the target assigned by the top-level task than in the previous step, or if it successfully destroys the target assigned by the top-level task in the current step; `penalty` is the penalty for drone failure; the reward is deducted if the drone is destroyed or exceeds the predetermined battlefield range; `γ` is the time correction factor, which aims to prevent the drone from completing its mission too late, and its calculation method is as follows: Among them, T current T represents the time elapsed from the start of the mission to the completion of the target strike by the drone. max The maximum permissible mission time for a drone swarm is determined by the battlefield size, drone speed, and maximum flight time.
10. The method for planning UAV swarm strike missions based on a hierarchical fusion framework according to claim 1, characterized in that: In step 5, specifically: Step 5.1 Before the training of the multi-agent reinforcement learning algorithm begins, an initialization operation is performed. Specifically, all drones are treated as independent agents, and their environmental states are initialized. The local observations S of each drone are combined into joint observation information O after data deduplication. During initialization, the joint observation information O0 of the initial state needs to be obtained. At the same time, the replay pools D and D' are constructed and initialized. Step 5.2: Consider the period from the start of the task to its success or failure as one training cycle. Within each training cycle, iterate every 0.1 seconds as one training step. In the t-th training step, the i-th red team UAV, based on the current joint observation O... t Use its own strategy π i :O i ×A i →[0,1] Select the action to be performed in this step, a i The individual actions of all agents constitute a joint action u∈A; Step 5.3: All agents interact with the environment through joint actions to obtain corresponding rewards and the next joint observation state; R(O,u) represents the reward function obtained by all agents after taking joint action u under joint observation information O, and the reward value r is obtained according to this reward function. t Then execute the joint action u t To obtain the next step, local observation S for each UAV. t+1 and joint observation information O t+1 The state transitions generated during this interaction process will be applied to {O}. t ,u t ,r t O t+1 Stored in playback pool D; Step 5.4: After completing the reward function calculation, if a red team drone is destroyed or goes out of bounds and fails in the current training step, or a blue team drone successfully attacks a target and is destroyed, it is considered that the current environment has changed. Return to step 3 to run the genetic algorithm to reassign task objectives to the drones that have not failed, and re-divide the task groups according to the new task objectives. Step 5.5: If the task succeeds or fails at this point, the trajectories generated throughout the entire training cycle are stored in the replay pool D'; subsequently, random sampling is performed from replay pools D and D' to process parameters ψ. θ、 Update; If the training has not ended at this point, return to step 5.1 to reinitialize the environment and begin training for the next training cycle.
Citation Information
Cited By
Anti-unmanned aerial vehicle cluster dynamic target allocation simulation method based on MAPPO laser defense system
CN121638001A
Anti- drone cluster dynamic target assignment simulation method based on MAPPO laser defense system
CN121638001B