A method for drone swarm task allocation based on collective intelligence-driven alliance game theory
By establishing a collaborative game model based on collective intelligence and an improved spatial game adaptive learning algorithm, the problem of information interaction and coordination scheduling in UAV swarm task allocation was solved, achieving efficient, balanced, and rapid response in UAV swarm task allocation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-11
- Publication Date
- 2026-03-13
AI Technical Summary
Existing drone swarm task allocation methods are difficult to achieve information interaction and coordinated scheduling between individuals in multi-drone collaborative control, cannot be effectively applied in complex scenarios, and traditional methods are difficult to meet the task requirements of speed and real-time.
A coalition game model based on collective intelligence is established. By designing time cost, resource consumption, and individual reputation functions, the task allocation problem is transformed into an optimal coalition structure problem. An improved spatial game adaptive learning algorithm is then used to optimize individual strategies to achieve task allocation for UAV swarms.
Within a limited timeframe, a balanced distribution of tasks and a high degree of benefit were achieved in the allocation of tasks among drone swarms, thereby improving the collaborative execution efficiency and task coverage of drone swarms.
Smart Images

Figure CN115963724B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for allocating tasks in a drone swarm based on collective intelligence-driven alliance game theory, belonging to the field of drone autonomous control. Background Technology
[0002] Unmanned aerial vehicle (UAV) swarms are complex systems composed of multiple UAVs. With the continuous improvement of UAV autonomy and artificial intelligence, swarm cooperative control methods have become an important research direction in the future UAV field. Compared with single-UAV missions, UAV swarms offer higher mission execution efficiency and a wider mission coverage, and can ensure the robustness and self-organization of the swarm system through the design of cooperative control methods. Efficient allocation methods play a crucial role in improving the collaborative mission execution capabilities of UAVs. By rationally allocating the tasks performed by UAVs, it is possible to complete tasks with minimal global cost or maximum global benefit. Facing the complex application requirements of future UAVs involving multiple tasks and objectives, developing UAV swarm cooperative control based on swarm intelligence has become a major research direction for the future.
[0003] The task allocation problem for unmanned aerial vehicles (UAVs) is a key issue in the cooperative control of UAV swarms, with its core lying in the research on target allocation. Currently, traditional task allocation methods include mathematical programming, auction algorithms, reinforcement learning, and biomimetic intelligent algorithms. However, traditional programming methods primarily address target allocation for single UAVs, rather than multi-UAV cooperative control. Auction algorithms can effectively solve the allocation problem for multiple UAVs and multiple targets, but they struggle to achieve information exchange and coordinated scheduling among individuals, limiting their application in complex scenarios. Reinforcement learning can yield the optimal solution to the task allocation problem, demonstrating high application value in dynamic and complex scenarios. However, it requires a suitable training set for model training, making the process difficult and time-consuming, thus failing to meet the requirements of speed and real-time performance. Biomimetic intelligent algorithms, including wolf pack and bee swarm algorithms, mimic the intrinsic mechanisms of typical biological group movement and decision-making to solve key problems in UAV swarm cooperative control, but further in-depth research combining them with swarm intelligence methods is needed. Treating UAVs as intelligent individuals and considering the collective intelligence-generating properties of UAV swarms, an effective approach is to apply game theory to the UAV task allocation problem. By establishing a game theory model for task allocation, individual interests can be aligned with global interests. As intelligent individuals, drones do not require complex decision-making systems, but rather the focus is on the study of swarm intelligence characteristics.
[0004] Game theory, as a common method for resolving conflicts between individual behavior and group interests in the field of swarm intelligence, is increasingly being applied to the cooperative control of UAV swarms. By stimulating cooperative behavior, it reflects the attributes of swarm intelligence. As a typical model in cooperative game theory, coalition game theory, by forming coalitions composed of several participants and dividing the group according to individual attributes and preferences, maximizes the interests of both individuals and the group, transforming the task allocation problem into solving the problem of finding the optimal coalition structure. Addressing the practical needs of UAV task allocation, UAVs form coalitions based on factors such as location, speed, and resources, with each coalition completing a corresponding objective task. Therefore, this invention establishes a coalition game model from the perspectives of time cost, resource consumption, individual reputation, and individual strategy update rules, and proposes a UAV swarm task allocation method based on swarm intelligence-stimulated coalition game theory to solve the UAV swarm task allocation problem. Summary of the Invention
[0005] This invention proposes a drone task allocation method based on swarm intelligence alliance game theory. It establishes the relationship between the alliance game model and the drone dynamics model, and designs functions for time cost, resource consumption, and individual reputation, transforming the task allocation problem into an optimal alliance structure problem, ensuring a one-to-one correspondence between drone alliances and swarm tasks. Finally, an optimization algorithm is used to optimize the individual update strategy, thereby obtaining the optimal solution for swarm task allocation and solving problems such as simple structure and poor adaptability in drone swarm task allocation.
[0006] To achieve the above objectives, the main steps of the UAV task allocation method based on swarm intelligence alliance game theory of the present invention are as follows:
[0007] Step 1: UAV Dynamics Modeling
[0008] like Figure 1 As shown, in the inertial coordinate system X g Y g Z g Below, a dynamic model of the UAV is established. Assuming the UAV is equipped with an autopilot for speed, heading, and altitude control, the 3-DOF nonlinear motion model of the UAV is simplified, resulting in: To control input variables,
[0009] With [x i ,y i ,h i ,v i ,ψ i ,λ i The simplified nonlinear model of the UAV with state variables is as follows:
[0010]
[0011] Where N represents the number of drones, i = 1, 2, ..., N. The horizontal position and altitude of the drones are represented by (x...). i ,y i ) and h i To represent. v i ψ represents the horizontal speed of the drone. i and λ i These represent the heading angle and altitude change rate of the UAV, respectively. These represent the control inputs of the drone's autopilot, τ V , τ ψ , and (τ λ ,τ h The numbers ) represent the control parameters of the drone's autopilot. In addition, the drone is also subjected to a thrust T. i Resistance D i Lift L i and gravity m i The role of g.
[0012] Considering the motion characteristics of the UAV, its speed, heading angle, and rate of change of altitude are subject to the following constraints:
[0013]
[0014] Among them, v min ,v max Let n represent the minimum and maximum speeds of the drone, respectively. max Indicates the normal overload of the drone, λ min ,λ max These represent the minimum and maximum values of the drone's altitude change rate, respectively.
[0015] Step 2: Modeling the Coalition Game Theory
[0016] The task allocation problem is transformed into a coalition grouping problem in coalition games, where the set of drones is M = {M1, M2, ..., M}. nM The target set is T = {T1, T2, ..., T}. nT The drones have different initial speeds, positions, and carried resources. The alliance divisions correspond one-to-one with the target sets, defined as S = {S1, S2, ..., S}. nT Each drone is assigned only one task at a time.
[0017]
[0018] Drone i selects a from the set i As the target it chooses, use the vector group a = (a1, a2, ..., a...) nMLet represent a set of results for task assignment. Let represent a set of UAVs that select target j under assignment solution a.
[0019]
[0020] Let K(i) further represent the allocation target of UAV i, S K(i) Indicates the alliance to which drone i belongs.
[0021] S K(i) ={S j ∈K|M i ∈S j} (5)
[0022] Global reward is defined as the total reward when all tasks are completed. The ultimate goal of task allocation is to maximize global reward.
[0023]
[0024] in Let represent the payoff function for objective j under the allocation of solution 'a'. It is defined as the task reward minus the task cost. When a task cannot be assigned, the payoff function is the penalty for ignoring that objective.
[0025]
[0026] Where, r j c represents the reward for completing task j, but also the penalty for ignoring that goal. ij This represents the cost function for drone i to complete target j, which includes both time cost and resource consumption.
[0027] Step 3: Design of UAV Cost Function
[0028] In the coalition game model, the formation of a coalition primarily considers the initial position, speed, carried resources, and individual reputation of the drones. The goal of this combinatorial optimization problem is to complete the task in a timely and efficient manner. This goal requires that the resources possessed by the coalition members are sufficient to guarantee the completion of the corresponding task, and the task completion time depends on the arrival time of all drones at the target; drones closer to the target need to wait for drones farther away. Therefore, the cost function of the drones consists of two aspects: time cost and resource consumption.
[0029]
[0030] Where d ij Let ω1, ω2, ε represent the distance between UAV i and target j. t ,ε e These represent the weighting coefficients.
[0031] Thus, the individual revenue function of the drone is obtained.
[0032]
[0033] Where |S j | indicates alliance S j The number of drones in a consortium. As the number of drones within the same consortium increases, the probability of drone collisions and communication load also increase significantly. Therefore, simply increasing the size of the consortium may harm the benefits of individual drones, and the size of the consortium needs to be reasonably controlled.
[0034] When the individual payoff function reaches its optimal solution, all individuals have achieved optimal allocation. Given that the allocation outcomes for other individuals in this combination remain unchanged, no individual has an incentive to unilaterally change their own goal; that is, Nash equilibrium is reached.
[0035]
[0036] Step 4: Design of Individual Reputation Functions for UAVs
[0037] To regulate the cooperative behavior of drones, a cumulative cooperative credit is defined for each drone based on the amount of resources it contributes to the mission. It is assumed that all drones have equal initial credit.
[0038]
[0039] Update the drone's accumulated reputation every moment.
[0040]
[0041] The change in reputation is defined as
[0042]
[0043] Where r j a represents the reward of task j. i Indicates the relative resource contribution of drones
[0044]
[0045] Normalize the cumulative reputation of drones.
[0046]
[0047] in These represent the maximum and minimum individual reputation values at the current moment. The individual reputation ranges from [0,1] at each moment. An individual's current reputation will influence future alliance formation. When the reputation falls below a certain threshold η... c At that time, drones were considered low-value partners and would be difficult to include in the formation of alliances.
[0048] Step 5: Design of Individual Strategy Update Rules
[0049] In coalition game models, an individual's strategy is the chosen target. Traditional strategy update rules include unconditional imitation, the Fermi rule, and pairwise comparison rules. These rules are relatively simple, but they mainly rely on neighbor information or global information and cannot reflect the characteristics of swarm intelligence. The Spatial Game Adaptive Learning (SAP) algorithm, however, features the characteristic of randomly selecting the target individual for updating with equal probability in each iteration. The selected drone M... i Calculate the task selection probability according to the following formula.
[0050]
[0051] Where σ(·) is the logit probability function
[0052]
[0053] However, random selection by individuals in a distributed environment is difficult, so a periodic adaptive selection mechanism is introduced. Each drone updates its policy in numerical order within a period. Although some randomness is sacrificed, the requirements of the SAP algorithm are still met.
[0054] Meanwhile, to accelerate the algorithm's convergence speed, a reference based on neighbor and historical information was added during the individual update task. If the updated individual reward is lower than the neighbor's optimal solution or the historical optimal solution, the drone will abandon the task update and randomly select a new task update.
[0055]
[0056] in This indicates the highest or historical highest profit among drone neighbors.
[0057] Step Six: Drone Swarm Task Allocation Model and Output
[0058] like Figure 3As shown, the task allocation process for a drone swarm based on collective intelligence-driven alliance game theory includes the following steps: First, the drones initialize target allocation, forming alliances based on different task objectives, with the individual first assigned to that objective becoming the leader. Then, a group payoff function and individual payoff functions are constructed based on the drone cost function and task value. Next, alliance members are screened based on individual reputation, eliminating individuals with lower reputations. Finally, an improved SAP algorithm is used to update the target allocation results for individuals in the alliance until the optimal allocation solution is reached. After obtaining the target allocation results, drone flight paths need to be planned to complete the task allocation process. The final output includes the task allocation result and the flight path planning result.
[0059] The proportional guidance method is used as the control law for route planning, such as Figure 2 As shown, the equations of relative motion
[0060]
[0061] Where r represents the relative distance between the UAV and the target, and the line connecting the UAV and the target is called the target line of sight; q represents the angle between the target line of sight and a baseline in the attack plane, called the target line of sight azimuth angle; V represents the speed of the UAV; σ represents the angle between the speed and the baseline; η represents the angle between the speed and the target line of sight, i.e., the UAV's forward angle.
[0062] Route planning primarily considers time consistency constraints, namely, coordinated flight time. The arrival time is greater than the expected arrival time of all drones within the same alliance, ensuring that all drones reach the target within the same timeframe.
[0063]
[0064] This invention presents a task allocation method for UAV swarms based on collective intelligence-inspired alliance game theory. Addressing the requirements of distributed coordination control in UAV task allocation, it proposes a task allocation method based on improved potential strategies. First, the task allocation problem is described as a joint formation game model, and individual and global payoff functions for each UAV are designed. Second, an improved SAP algorithm is used to optimize and solve the game model. Simulation results demonstrate that the proposed method exhibits superior allocation equilibrium and task benefit advantages within a finite time frame, validating its effectiveness. Attached Figure Description
[0065] Figure 1 This is a schematic diagram of the dynamics model of a drone.
[0066] Figure 2 This is a schematic diagram of the proportional guidance method.
[0067] Figure 3This is a flowchart of the alliance game algorithm of the present invention.
[0068] Figures 4(a) and (b) are location trajectory diagrams of the UAV cluster.
[0069] Figures 5(a) and (b) are curves showing the changes in revenue from drone swarms.
[0070] X g — x-axis of inertial coordinate system
[0071] O g —Z-axis of the inertial coordinate system
[0072] Y g —Y-axis of inertial coordinate system
[0073] Z g —Z-axis of the inertial coordinate system
[0074] x i — x-coordinate of the UAV in the inertial coordinate system
[0075] y i — The y-coordinate of the UAV in the inertial coordinate system
[0076] h i — The z-coordinate of the UAV in the inertial coordinate system
[0077] v i —The current speed of the drone
[0078] ψ i —The current heading angle of the drone
[0079] L i —Lift force experienced by the drone
[0080] D i —Resistance encountered by drones
[0081] m i g – the force of gravity acting on the drone
[0082] -- Rate of change of drone altitude
[0083] —UAV altitude control input
[0084] —UAV speed control input
[0085] --Unmanned Aerial Vehicle (UAV) heading angle control input
[0086] r—relative distance between the drone and the target
[0087] q — Angle between the target's line of sight and the baseline
[0088] σ—The angle between the velocity and the baseline
[0089] η—The angle between the velocity and the target's line of sight Detailed Implementation
[0090] The effectiveness of the proposed method is verified below through a specific example of drone swarm task allocation. The experimental computer was configured with an Intel Core i7-7700HQ processor, 2.8GHz clock speed, 8GB RAM, and MATLAB 2018a software.
[0091] The present invention provides a method for drone swarm task allocation based on collective intelligence-inspired alliance game theory, the process of which is as follows: Figure 3 As shown.
[0092] Step 1: UAV Dynamics Modeling
[0093] A dynamic model of the UAV is established, assuming that the UAV is equipped with an autopilot for speed, heading, and altitude control. The 3-DOF nonlinear motion model of the UAV is simplified to obtain the following: To control input variables, with [x i ,y i ,h i ,v i ,ψ i ,λ i The simplified nonlinear model of the UAV with state variables is as follows:
[0094]
[0095] Where N represents the number of drones, N = 15. The horizontal position and altitude of the drones are represented by (x... i ,y i ) and h i To represent. v i ψ represents the horizontal speed of the drone. i and λ i These represent the heading angle and altitude change rate of the UAV, respectively. These represent the control inputs of the drone's autopilot, τ V , τ ψ , and (τ λ ,τ h The numbers ) represent the control parameters of the drone's autopilot. In addition, the drone is also subjected to a thrust T. i Resistance D i Lift L i and gravity m i The role of g.
[0096] Considering the motion characteristics of the UAV, its speed, heading angle, and rate of change of altitude are subject to the following constraints:
[0097]
[0098] Among them, v min =15m / s,v max =120m / s represents the minimum and maximum speed of the drone, respectively.
[0099] n max =5 indicates the normal overload of the drone, λ min = -5m / s,λ max =5m / s represents the minimum and maximum values of the drone's altitude change rate, respectively.
[0100] Step 2: Modeling the Coalition Game Theory
[0101] The task allocation problem is transformed into a coalition grouping problem in coalition games, where the set of drones is M = {M1, M2, ..., M}. nM The target set is T = {T1, T2, ..., T}. nT}, where n M =15,n T =5. The alliance divisions correspond one-to-one with the target sets, defined as S = {S1, S2, ..., S...} nT Each drone is assigned only one task at a time.
[0102]
[0103] Drone i selects a from the set i As the target it chooses, use the vector group a = (a1, a2, ..., a...) nM Let represent a set of results for task assignment. Let represent a set of UAVs that select target j under assignment solution a.
[0104]
[0105] Let K(i) further represent the allocation target of UAV i, S K(i) Indicates the alliance to which drone i belongs.
[0106] S K(i) ={S j ∈K|M i ∈S j} (5)
[0107] Global reward is defined as the total reward when all tasks are completed. The ultimate goal of task allocation is to maximize global reward.
[0108]
[0109] in Let represent the payoff function for objective j under the allocation of solution 'a'. It is defined as the task reward minus the task cost. When a task cannot be assigned, the payoff function is the penalty for ignoring that objective.
[0110]
[0111] Where, r j =1.5 represents the reward for completing task j, but also the penalty for ignoring the objective. ij This represents the cost function for drone i to complete target j, which includes both time cost and resource consumption.
[0112] Step 3: Design of UAV Cost Function
[0113] The cost function of a drone consists of two parts: time cost and resource consumption.
[0114]
[0115] Where d ij Let ω1 = 0.05, ω2 = 0.01, and ε represent the distance between drone i and target j. t =1,ε e =1 represents the weighting coefficient.
[0116] Thus, the individual revenue function of the drone is obtained.
[0117]
[0118] Where |S j | indicates alliance S j The number of drones in the country.
[0119] When the individual payoff function reaches its optimal solution, all individuals have achieved optimal allocation. Given that the allocation outcomes for other individuals in this combination remain unchanged, no individual has an incentive to unilaterally change their own goal; that is, Nash equilibrium is reached.
[0120]
[0121] Step 4: Individual Reputation Design for Drones
[0122] To regulate the cooperative behavior of drones, a cumulative cooperative credit is defined for each drone based on the amount of resources it contributes to the mission. It is assumed that all drones have equal initial credit.
[0123]
[0124] Update the drone's accumulated reputation every moment.
[0125]
[0126] The change in reputation is defined as
[0127]
[0128] Where r j a represents the reward of task j. i Indicates the relative resource contribution of drones
[0129]
[0130] Normalize the cumulative reputation of drones.
[0131]
[0132] in These represent the maximum and minimum individual reputation values at the current moment. The individual reputation ranges from [0,1] at each moment. An individual's current reputation will influence future alliance formation. When the reputation falls below a certain threshold η... c When the coefficient is 0.6, drones are considered low-value partners and will be difficult to include in the formation of alliances.
[0133] Step 5: Design of Individual Strategy Update Rules
[0134] The Spatial Game Adaptive Learning (SAP) algorithm is used as the update rule for individual policies. The selected drone M i Calculate the task selection probability according to the following formula.
[0135]
[0136] Where σ(·) is the logit probability function
[0137]
[0138] A periodic adaptive selection mechanism is introduced to facilitate sequential updates of individual UAV policies. Furthermore, to accelerate algorithm convergence, references based on neighbor and historical information are added during individual update tasks. If the updated individual reward is lower than the neighbor's optimal solution or the historical optimal solution, the UAV will abandon the task update and randomly select a new task for update.
[0139]
[0140] in This indicates the highest or historical highest profit among drone neighbors.
[0141] Step Six: Drone Swarm Task Allocation Model and Output
[0142] The proportional guidance method is used as the control law for route planning, such as Figure 2 As shown, the equations of relative motion
[0143]
[0144] Where r represents the relative distance between the UAV and the target, and the line connecting the UAV and the target is called the target line of sight; q represents the angle between the target line of sight and a baseline in the attack plane, called the target line of sight azimuth angle; V represents the speed of the UAV; σ represents the angle between the speed and the baseline; η represents the angle between the speed and the target line of sight, i.e., the UAV's forward angle.
[0145] Route planning primarily considers time consistency constraints, namely, coordinated flight time. The arrival time is greater than the expected arrival time of all drones within the same alliance, to ensure that all drones reach the target within the same timeframe:
[0146]
[0147] Assuming the simulation scenario is a rectangular space of 10km×10km×1km, the simulation results show the task allocation results of the UAV swarm, including flight trajectories, as shown in Figures 4(a) and (b), where Figure 4(a) is the three-dimensional trajectory and Figure 4(b) is the two-dimensional trajectory; and the UAV swarm benefit change curves are shown in Figures 5(a) and (b), where Figure 5(a) is the individual benefit change curve and Figure 5(b) is the group benefit change curve.
Claims
1.A method for task allocation of a UAV swarm based on swarm intelligence inspired coalition game, characterized in that: The method steps are as follows: Step one: unmanned aerial vehicle dynamics modeling In the inertial coordinate system X g Y g Z g Next, the dynamic model of UAV is established; assuming that the UAV has been equipped with an autopilot for speed, heading and altitude control, the 3-DOF nonlinear motion model of the UAV is simplified to take as the control input variable, with [x i , y i , h i , v i , ψ i , λ i ] as state variables, is as follows: where N denotes the number of UAVs, i = 1, 2,..., N; the horizontal position and height of the UAV are denoted by (x i ,y i ) and h i , respectively; v i denotes the horizontal velocity of the UAV, ψ i and λ i denote the heading angle and the rate of change of height of the UAV, respectively; denote the control inputs of the autopilot of the UAV, τ V , τ ψ , and (τ λ , τ h ) denote the control parameters of the autopilot of the UAV, respectively; in addition, the UAV is also subject to the thrust T i , the drag D i , the lift L i , and the gravity m i g. Step two: coalition game model modeling The task allocation problem is converted into a coalition grouping problem in a coalition game, where the set of UAVs is The target set is The UAVs have different initial speeds, positions, and carry resources; the coalition division is one-to-one corresponding to the target set, defined as Each UAV is only assigned to one task at the same time, that is Drones i select a from the set i As self-selected targets, a set of vectors A set of results representing task assignment; a set of drones representing the selected targets j under assignment solution a Further, K(i) represents the allocation target of the UAV i, S K(i) represents the alliance to which the UAV i belongs S K(i) = {S j ∈ K | M i ∈ S j} (5) The global benefit is defined as the total benefit when all tasks are completed, and the final purpose of task allocation is to maximize the global benefit where, represents the payment function for target j under assignment solution a; it is defined as the task reward minus the task cost; when the task fails to be assigned, the payment function is a penalty of ignoring the target where r j represents the reward of completing task j, and is also the punishment of ignoring the objective, c ij represents the cost function of UAV i completing task j, including time cost and resource consumption Step three: unmanned aerial vehicle cost function design In the coalition game model, the cost function of the unmanned aerial vehicle consists of two parts, namely time cost and resource consumption Wherein, d ij represents the distance between the unmanned aerial vehicle i and the target j, ω1, ω2, ε t , ε e respectively represent the weight coefficient; Thus, the individual benefit function of the unmanned aerial vehicle is obtained where |S j | denotes the number of drones in the coalition S j When the individual benefit function reaches the optimal solution, all individuals achieve optimal allocation, given the allocation results of other individuals in the allocation combination, any individual has no incentive to change his own target unilaterally, that is, to achieve Nash equilibrium Step four: design of individual reputation function of unmanned aerial vehicle In order to regulate the cooperative behavior of unmanned aerial vehicle, according to the amount of resources contributed by unmanned aerial vehicle in the task, the cumulative cooperation credit of each unmanned aerial vehicle is defined; it is assumed that all unmanned aerial vehicles have equal initial credit At each moment, the cumulative reputation of the unmanned aerial vehicle is updated Where the reputation change is defined as where r j represents the benefit of task j, a i represents the relative resource contribution of the UAV The cumulative reputation of the unmanned aerial vehicle is normalized wherein, are the maximum and minimum values of individual reputation at the current time; the range of individual reputation at each time is [0, 1], and the current individual reputation will have an impact on the future formation of alliance, when the reputation is lower than a certain threshold η c When the reputation is lower than a certain threshold η, the UAV is regarded as a low-value cooperative object and will be difficult to participate in the formation of alliance. The step one further comprises: considering the motion characteristics of the unmanned aerial vehicle, the speed, heading angle and height change rate are constrained accordingly, as follows: wherein v min , v max represent the minimum and maximum values of the speed of the UAV, respectively, n max represents the normal overload of the UAV, λ min , λ max represent the minimum and maximum values of the rate of change of the height of the UAV, respectively; Step five: design of individual strategy update rule In the coalition game model, the strategy of individual is the selected target, and the space game adaptive learning algorithm has the characteristics of randomly selecting the updated target individual with equal probability in each iteration process. The selected UAV M i The task selection probability of the UAV M is calculated according to the following formula Where σ(·) is the logit probability function Periodic adaptive selection mechanism is introduced, and each unmanned aerial vehicle updates the strategy in the order of number within a period; At the same time, in order to speed up the convergence speed of the algorithm, the reference based on neighbor information and historical information is increased when the individual updates the task; if the updated individual benefit is lower than the neighbor optimal solution or the historical optimal solution, the unmanned aerial vehicle will give up task update, and randomly select a new task update wherein, represents the highest drone neighbor gain or the highest historical gain; Step six: unmanned aerial vehicle cluster task allocation model and output; the proportional guidance method is used as the control law of the route planning, and the relative motion equation is: wherein r represents the relative distance between the UAV and the target, the line connecting the UAV and the target is called the target line of sight; q represents the angle between the target line of sight and a reference line in the attack plane, which is called the target line of sight azimuth; V represents the speed of the UAV; σ represents the angle between the speed and the reference line; η represents the angle between the speed and the target line of sight, i.e. the forward angle of the UAV; in the path planning, the constraint of time consistency, i.e. the cooperative sailing time is mainly considered greater than the expected arrival time of all UAVs within the same alliance to ensure that all UAVs arrive at the target at the same time: