Unmanned aerial vehicle cluster collaborative confrontation system and method

By decomposing the UAV swarm collaboration problem into the decision-making layer and the action layer, and using intent recognition and QMIX algorithm to optimize UAV actions, the problems of dynamic environment response and communication limitations in UAV swarm confrontation are solved, achieving efficient defense effects.

CN120631052APending Publication Date: 2025-09-12XI AN JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510763350.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

During drone swarm confrontations, existing technologies are unable to effectively respond to dynamic environmental changes, and there are problems such as difficulty in modeling and verification, limited communication capabilities, high testing costs, and low efficiency.

Method used

The UAV swarm collaboration problem is decomposed into the decision-making layer and the action layer. The decision-making layer uses intention recognition and scheduling algorithms to group UAVs and decompose tasks. The action layer uses the QMIX algorithm to optimize motion and is verified in a simulation environment.

Benefits of technology

It achieves flexible response and efficient defense against drone swarm confrontations, improves the defense win rate, and enhances the efficiency of confrontation and the foresight of strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631052A_ABST
    Figure CN120631052A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle cluster collaborative confrontation system and method. The system comprises a decision-making layer, an action layer and a simulation environment. The decision-making layer is used for intention recognition and performing unmanned aerial vehicle dynamic marshalling and task decomposition through a related algorithm; the action layer is used for executing the decomposed marshalling-level tasks; the simulation environment is used for verifying related algorithm performance in the decision-making layer. The method has the characteristics of layering and flexibility. An unmanned aerial vehicle cluster cooperation problem is decomposed into two layers, namely a decision-making layer and an action layer, so that opponents with different intentions and strategies can be flexibly coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drone swarm collaborative confrontation, and in particular to a drone swarm collaborative confrontation system and method. Background Art

[0002] Currently, in drone swarm combat, situational drones must flexibly adjust their strategies based on the dynamically changing environment to achieve highly coordinated operations. From a decision-making and control perspective, there are three main approaches. The first is based on expert systems, which can leverage prior knowledge and rules for decision-making. While technically mature, these approaches perform poorly in large-scale and complex scenarios.

[0003] The other type is based on swarm intelligence methods, such as wolf packs, ant colonies, and artificial bee colonies. Each drone has the ability to make autonomous decisions, and there will be simple communication and interaction between drones, so that the cluster can exhibit complex behaviors and stronger capabilities.

[0004] At the same time, research on drone swarm countermeasures faces a number of challenges. First, modeling and analyzing drone swarm countermeasures is challenging, requiring consideration of multiple factors, including drone communication, perception, and coordination. Furthermore, the research needs to adapt to future larger drone swarms and more complex and unknown scenarios. Second, verifying swarm countermeasures is challenging. Using physical drones for verification presents challenges such as limited testing sites, high testing costs, high risks, and low efficiency. Furthermore, drone communication capabilities are limited by hardware requirements, communication protocols, and observation range and accuracy, hindering the efficiency and effectiveness of drone swarm countermeasures. Summary of the Invention

[0005] To overcome the shortcomings of the existing technologies, the present invention provides a UAV swarm collaborative countermeasure system and method. This system and method features a hierarchical and flexible approach. By decomposing the UAV swarm collaboration problem into two levels: the decision-making layer and the action layer, it can flexibly respond to adversaries with different intentions and strategies.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is:

[0007] A UAV swarm collaborative confrontation system, including a decision layer, an action layer and a simulation environment;

[0008] The decision layer is used for intention recognition and dynamic grouping and task decomposition of drones through relevant algorithms;

[0009] The action layer is used to execute the decomposed group-level tasks;

[0010] The simulation environment is used to verify the performance of relevant algorithms in the decision-making layer.

[0011] The decision-making layer includes demarcating defense zones, analyzing zone patterns, and dispatching drones;

[0012] The above-mentioned demarcation of defense zones takes into account the shape of the zones, and the zones should ensure that the entire target area is fully covered; the number of zones should be moderate, as too few zones may result in defense loopholes, while too many zones will lead to complex management or waste of resources.

[0013] The partitioning forms include two methods:

[0014] Method 1 is rule-based: it calculates the overall threat index of each zone based on the size of the opponent and the proximity of the Red team's opponent to the target area. Method 2 is intention-based: it uses HMM to analyze the opponent's intention, providing both group-based and individual-based analysis methods.

[0015] In group-based intention recognition, the average distance and average speed statistics of the red team to the other team are used to construct observation variables, and then the intention of the entire group is inferred based on the observation variable sequence;

[0016] Individual-based intention recognition starts with a single drone. Observation variables are designed based on the speed and other characteristic quantities of the individual drone. Then, the intention of each drone of the red team is recognized separately. Finally, a voting method is used to obtain the probability distribution of the intention of the entire group.

[0017] The dispatching of drones includes dispatching timing and dispatching rules;

[0018] The specific method and steps for scheduling timing are: considering the communication cost and maneuverability of UAV scheduling, scheduling is carried out at an appropriate frequency;

[0019] The specific rules of the scheduling rules are based on the principle of neighbor priority scheduling. The scheduling priority of adjacent areas is higher than that of intervening areas. For human-machines transferred from the same area to another area, the UAV with the shortest path is prioritized.

[0020] A method for using a UAV swarm collaborative countermeasure system comprises the following steps:

[0021] Step 1: Set a circular target area and divide it into several defensive zones. Analyze the situation of each defensive zone to assess the relative importance of each zone and determine the number of drones required for each zone. Consider the length of the dispatch path, the urgency of the task, and the dispatch time consumption to rationally plan the dispatch timing and the dispatch path of the drones.

[0022] Step 2: A distributed execution approach is adopted at the action layer. Individual drones select actions based on local observations, and the QMIX algorithm is used to optimize and coordinate the drones' actions to complete the local collaborative defense task.

[0023] The step 1 is specifically as follows:

[0024] 1.1: Delineate defensive zones;

[0025] 1.2: Introduce the threat index T to measure the threat level of the Red Team to each area, and use it as the basis for drone scheduling, and perform scheduling based on the rule-based scheduling algorithm;

[0026] 1.3: Scheduling based on the scheduling algorithm of intent recognition;

[0027] 1.4: Schedule according to the scheduling timing and the scheduling rules based on steps 1.2 and 1.3.

[0028] The step 1.1 is specifically as follows: in the decision layer, the target area is divided into four identical 90° sector areas around the center of the target area in the experiment, corresponding to the four directions respectively;

[0029] Each defensive sector needs to determine a defensive base point as the gathering center for drones guarding the area, and set the base point at the midpoint of the arc section of each sector.

[0030] Step 1.2 specifically includes: introducing a threat index T to measure the threat level of the Red Army to each area, and using it as the basis for UAV dispatch;

[0031] In rule-based scheduling algorithms,

[0032] First, consider the scale of enemy drones in the defensive zone. The more enemy drones observed in the zone, the greater the potential threat to the Red Team. More drones are needed for effective interception and attack. Therefore, the threat index of this zone is proportional to the number of enemy drones. The normalized threat index is:

[0033]

[0034] Where N div is the number of defensive zones, n i Indicates the number of enemy drones observed in area i, threat index T num is a relative value, satisfying:

[0035]

[0036] Secondly, the proximity between the red team and the target area can be analyzed by the average distance between the red team group and the target center. To indicate that here The smaller the size, the greater the threat to the target area, so the threat index of this part is the same as It is negatively correlated and is calculated as:

[0037]

[0038] where d max The maximum distance allowed by the venue must meet

[0039] Combining the threat indices of the above two aspects, the overall threat index of each zone is calculated in a linear weighted form:

[0040] T(i)=α num T num (i)+α dis T dis (i),i=1,2,3,4(4)

[0041] where α num With α dis is the weight of the two threat indices, satisfying α num +α num =1. In the experiment, the two parts are considered to be equally important, so both are set to 0.5.

[0042] The step 1.3 is specifically as follows:

[0043] Situation 1) Group-based intention recognition:

[0044] Based on the local patterns that the Red Army may adopt, its intentions are designed into five types: strong attack, feint attack, retreat, transfer, and no threat;

[0045] When using HMM modeling, it is treated as a hidden state and represented by the set Q = {q1,q2,q3,q4,q5};

[0046] A strong attack means that the Red side concentrates a large number of troops to launch a direct attack, quickly approaching the target area and forcibly breaking through the Blue side's defense line;

[0047] A feint attack is when the Red side sends out some drones as bait to conduct a false attack, attempting to draw the Blue side's firepower to tie down its defenses and create opportunities for other drones to attack.

[0048] Retreat is the Red Army's rapid withdrawal from the target area to a safe zone to avoid excessive losses, readjust its strategy and prepare for the next attack;

[0049] The red team circles around the target area, leaving the current battle zone and turning to other battle zones, trying to find a gap in the blue team's defense before attacking.

[0050] No Threat means that the Blue Team has not observed any enemy drones in the defensive zone, but it does not mean that there are actually no enemy drones in the zone.

[0051] The situation selection statistics are the number of red groups, average speed, and average speed direction, and each statistic is divided into multiple levels according to different value ranges;

[0052] Among the three types of statistics, different levels of average speed v distinguish different attack intensities of the Red side; the direction of average speed is represented by the angle α between the moving direction and the target area direction. The three ranges of direction angles represent approach, circling, and moving away, respectively, and can be used to distinguish between intentions such as strong attack, transfer, and retreat;

[0053] The three types of statistics are further combined into observation variables p1~p 12 , and additionally increase p 13 The variable indicates that no enemy drones were observed in this partition;

[0054] The time window size of intent recognition is set to 5 time steps, that is, only a limited length of observation sequence {o t-4 ,o t-3 ,…,o t Infer the Red Team's intentions; Before using HMM for intention recognition, it is necessary to learn the model parameters λ = (A, B, π). A supervised learning approach is used, that is, given an observation sequence O and a state sequence I, the model parameters are calculated using the maximum likelihood estimation method.

[0055] First, we need to create training data, conduct a certain simulated confrontation process in the simulation environment, collect observation data O from the environment and label the intention I, combine the characteristics of each intention and the properties of the observation variables, and for the observation sequence O obtained at time t t =(o t-4 ,o t-3 ,…,o t ), whose intention t The marking method is:

[0056] After obtaining the training data, the maximum likelihood estimation method is used to calculate the initial state probability π, the state transition probability A, and the observation probability B:

[0057]

[0058]

[0059] Where s is the total number of samples, n is i (π) is the initial state q i The frequency, n ij (A) is the frequency of state i transitioning to state j, n ik (B) is the frequency of observation k in state j;

[0060] After obtaining the HMM model parameters λ = (A, B, π) from the above steps, the red team's intention is identified in the actual confrontation.

[0061] Furthermore, during the game, the collected original information sequence needs to be processed into an observation sequence at each time step. When the decision layer performs scheduling, the finite observation sequence corresponding to that moment is extracted;

[0062] The forward algorithm is used to calculate the probability distribution of the red team's intentions as follows:

[0063] Assume that the observation sequence obtained at time T is O T =(o1,o2,…,o T ), in state q at time t i The forward probability is α t (i), first calculate the initial value:

[0064] α1(i)=π i b i (o i ),i=1,2,…,N (8)

[0065] Next, recursively perform the following for t=1, 2, …, T-1:

[0066]

[0067] Finally, the forward probability of each state at time T is normalized to obtain the probability distribution of each intention:

[0068]

[0069] For each intention q j Set an importance coefficient m j , intention Q = {force attack, feint attack, retreat, transfer, no threat} corresponding to the value of m = [0.35, 0.25, 0.15, 0.20, 0.05], and then the intention probability distribution γ obtained by the above process j The intention threat index of partition i is calculated as:

[0070]

[0071] Consider the average distance in the rule section The purpose is to obtain the distance threat index T dis , the calculation method is the same as formula (3);

[0072] Based on the above process, the threat index finally calculated by the group-based intention recognition method is:

[0073] T(i)=α int T int (i)+αdis T dis (i),i=1,2,3,4(12)

[0074] where α int With α dis are the weights of the intention threat index and the distance threat index respectively.

[0075] 2) Individual-based intention recognition: First, this method classifies intentions into four types: strong attack, feint attack, retreat, and transfer;

[0076] Secondly, we select the individual characteristic quantities of the speed and movement direction of a single enemy drone. At the same time, we cannot use quantitative characteristics to judge individual intentions, which makes the combination of observed variables more simplified. The quantitative characteristics will be considered in the rule part.

[0077] Use supervised learning and use the maximum likelihood estimation method to calculate λ = (A, B, π);

[0078] According to the model parameters, the intention of a single drone is first identified. After the forward algorithm is used to obtain the probability distribution of the intention of a single enemy drone at time T, the intention with the highest probability is taken as the intention of the drone. Then, based on the individual intention identification results, the probability distribution of the intention of the entire group is calculated by voting. Let the total number of enemy drones in the partition be n, and the number of enemy drones corresponding to intention i be n i , then the probability distribution of group intention is:

[0079]

[0080] The threat index still includes the intention part and the rule part. In the intention part, the importance coefficients of the four intentions Q = {force attack, feint attack, retreat, transfer} are taken as m = [0.40, 0.25, 0.15, 0.20], and the intention threat index T is obtained by the calculation method of formula (11) int ; In the rule part, it is necessary to consider both the number of groups n and the average distance The influence of the two aspects is calculated according to the method of formula (1) and formula (3) respectively. num With T dis Finally, based on the individual intention recognition method, the overall threat index is calculated as:

[0081] T(i)=α int T int (i)+α num T num (i)+α dis T dis (i),i=1,2,3,4 (14).

[0082] 1.4 Scheduling Timing and Rules

[0083] The step 1.4 is specifically as follows:

[0084] 1) Scheduling timing

[0085] Schedule at an appropriate frequency, every 2 time steps;

[0086] 2) Scheduling rules

[0087] In the rule-based and intent-based scheduling algorithms, the operational situation of each defensive zone has been analyzed, and by calculating the threat index T, a quantitative assessment of the relative defensive strength required for each zone is achieved. Next, based on the threat index T, the number of drones deployed to each zone and the defensive area each drone is responsible for are determined;

[0088] The scheduling process requires rational planning to minimize the length and time of the drone's dispatch path while ensuring the completion of the assigned tasks. Therefore, the algorithm follows the principle of neighbor-first scheduling, with scheduling priority given to adjacent areas over those in between. For drones dispatched from one area to another, the drone with the shortest path is prioritized.

[0089] The algorithm is mainly divided into two steps: first, determine the scheduling direction and quantity between partitions; second, determine the defense area to which each drone is dispatched.

[0090] The step 2 is specifically as follows:

[0091] The action layer is responsible for outputting the specific actions that the drone needs to perform, which is implemented using the QMIX algorithm. The QMIX algorithm in the action layer adopts a distributed execution method, and each drone selects the best action based on local observation information.

[0092] Among the designed observation values, those related to the position are the coordinates (x, y) of the drone and the distance d from the target center, where the coordinates of the target center are taken as the origin;

[0093] The drone is dispatched to a specific defense zone, and its observation values ​​are taken as the relative coordinates (x-x0, y-y0) and the relative distance d(x-x0, y-y0) relative to the defense zone base point (x0, y0). In this way, the drone can defend a specific zone.

[0094] Beneficial effects of the present invention:

[0095] The technical solution of the present invention can effectively solve the problem of drone cluster confrontation.

[0096] The present invention is hierarchical and flexible. By decomposing the UAV swarm collaboration problem into two levels: the decision-making level and the action level, it can flexibly respond to opponents with different intentions and strategies.

[0097] The present invention processes and analyzes battlefield data to infer the possible intentions and behavior patterns of the Red Army, thereby better predicting potential threats and making defense strategies more forward-looking.

[0098] The present invention schedules drones through the decision layer, and the QMIX algorithm in the action layer outputs the drone's actions. In the decision layer, scheduling algorithms are designed based on rules and HMM intent recognition, with intent recognition further divided into group intent and individual intent-based methods.

[0099] The decision-making algorithm based on HMM individual intention recognition in the present invention has the best effect and provides a new method and idea for achieving efficient drone cluster defense operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 Schematic diagram of the system framework of the present invention.

[0101] Figure 2 This is a schematic diagram of the defense partition of the present invention.

[0102] Figure 3 Schematic diagram of the test results of four defensive strategies under different attack strategies of the present invention.

[0103] Figure 4 This is a schematic diagram of the changes in various data of the four defensive strategies over time under different attack strategies. DETAILED DESCRIPTION

[0104] The present invention will be described in further detail below with reference to the accompanying drawings.

[0105] Hierarchical decision-making algorithm architecture

[0106] In the drone swarm confrontation scenario set by the present invention, the target area defended by the blue team is circular, which requires the blue team to reasonably deploy drones in space according to the combat situation in different areas on the battlefield, so as to achieve all-round collaborative defense.

[0107] Specifically, if a certain area is under the focus of the Red Army's attack or the Blue Army suffers too much casualties, it is necessary to dispatch drones in a timely manner to strengthen the deployment of troops in that area; conversely, for areas where the Red Army's attack intensity is relatively low, there is no need to focus on defense, but rather dispatch them to other areas with weaker defenses.

[0108] In order to achieve a high degree of spatial collaboration among the blue team's drones, the present invention hierarchically organizes the blue team's defense strategy, from high-level overall deployment to low-level action execution, and gradually refines it from the overall to the local, making decision-making more flexible and efficient.

[0109] The adversarial system mainly consists of two layers, the decision layer and the action layer. The overall architecture of the adversarial system is as follows: Figure 1 shown.

[0110] The decision-making layer plays a macro-control role in the entire defense mission, realizes the coordinated scheduling of drones, and ultimately determines the defense area that each drone is responsible for.

[0111] First, the area needs to be divided into several defense zones, which is the basis for cross-regional drone dispatching;

[0112] Next, the situation in each defensive zone is analyzed. This paper employs both rule-based and intent-recognition analysis methods to assess the relative importance of each defensive zone, thereby determining the number of drones required for each. Finally, a specific scheduling plan is developed by comprehensively considering factors such as the length of the dispatch path, the urgency of the task, and the time required for dispatching. This plan rationally plans the dispatch timing and the drone dispatch paths, maximizing dispatch efficiency.

[0113] The action layer is responsible for determining the specific actions that the drone needs to perform.

[0114] At the decision-making level, drones have already learned their respective defense zones through information exchange and negotiation. To enhance defensive flexibility, the action layer employs a distributed execution approach, with individual drones selecting actions based on local observations. This ensures that even if some drones malfunction or lose connection, the overall system's operation is not impacted. The QMIX algorithm optimizes and coordinates drone actions, enabling coordinated local defense.

[0115] It's worth noting that coordinated scheduling at the decision-making level requires the blue drones to share observation information so they can aggregate and analyze situational information across each defensive zone and make collective scheduling decisions. Given the limited communication capabilities of drone swarms, information aggregation at the decision-making level requires complex communication and data transmission, incurring significant communication costs. The action-level, on the other hand, only requires local information, which carries a relatively low communication cost. Therefore, the decision-making level's scheduling frequency is lower than the action-level's output frequency.

[0116] The decision-making layer is specifically:

[0117] 1.1 Delineation of defensive zones:

[0118] At the decision-making level, the division of the defense area is the premise of drone scheduling. Considering the shape of the partition, since the blue team needs to defend a circular area, the partition should ensure that the entire target area is fully covered. At the same time, the number of partitions should be moderate. Too few may lead to defense loopholes, while too many will lead to complex management or waste of resources. Combining the cluster game scenario with the scale of drone confrontation, the target area is divided into four identical 90° fan-shaped areas around the center of the target area, corresponding to the four directions, such as Figure 2 shown.

[0119] Each defensive zone needs to determine a defensive base point, which serves as the gathering center for drones guarding the area. The base point is set at the midpoint of the arc of each sector. This ensures that the dispatch distance of drones between areas will not be too long, and also expands the overall observation range of the drone cluster as much as possible, so as to promptly detect and intercept enemy drones in the periphery.

[0120] 1.2 Rule-based scheduling algorithm:

[0121] In order to determine the number of drones required to be allocated to each defensive zone, it is necessary to fully collect battlefield information and comprehensively analyze the situation from multiple angles, so as to reasonably assess the defensive strength required for each zone.

[0122] To achieve quantitative analysis, the present invention introduces the threat index T to measure the threat level of the red team to each area, and uses it as the basis for UAV scheduling.

[0123] The rule-based scheduling algorithm primarily analyzes enemy drones based on their size and proximity to the target area. For each defensive sector, only enemy drone information within that sector is counted. This refers to the 90° azimuth interval corresponding to the defensive sector.

[0124] First, consider the scale of enemy drones in the defensive zone. The more enemy drones observed in the zone, the greater the potential threat to the Red Team. More drones are needed for effective interception and attack. Therefore, the threat index of this zone is proportional to the number of enemy drones. The normalized threat index is:

[0125]

[0126] Where N div is the number of defensive zones, n i Indicates the number of enemy drones observed in area i, threat index T num is a relative value, satisfying:

[0127]

[0128] Secondly, the proximity between the red team and the target area can be analyzed by the average distance between the red team group and the target center. To indicate that here The smaller the size, the greater the threat to the target area, so the threat index of this part is the same as It is negatively correlated and is calculated as:

[0129]

[0130] where d max The maximum distance allowed by the venue must meet

[0131] Combining the threat indices of the above two aspects, the overall threat index of each zone is calculated in a linear weighted form:

[0132] T(i)=α num T num (i)+α dis T dis (i),i=1,2,3,4(4)

[0133] where α num With α dis is the weight of the two threat indices, satisfying α num +α num =1. In the experiment, the two parts are considered to be equally important, so both are set to 0.5.

[0134] 1.3 Scheduling algorithm based on intent recognition:

[0135] 1) Algorithm Description

[0136] In drone swarm confrontations, for the Blue team to make more efficient scheduling decisions, it is necessary to mine deeper information from limited battlefield data. To address this, this paper designs a scheduling method based on intent recognition. By processing and analyzing battlefield data, it infers the Red team's possible intentions and behavior patterns, thereby better predicting potential threats and making defensive strategies more proactive.

[0137] This paper uses a hidden Markov model (HMM) to identify opponent intent, using the Red team's intent as the hidden state and processed battlefield data as the observed quantity. The Red team's intent is generally uncertain and will, to a certain extent, influence the changing patterns of situational information. On-field data changes dynamically over time, with certain logical relationships between them. The HMM is well suited for modeling such time-series data, effectively handling the uncertainty of the hidden state and describing the opponent's intent through a probability distribution.

[0138] When using HMM to analyze opponent intentions, this invention provides two analysis methods: group-based and individual-based. In group-based intention identification, observation variables are constructed using statistics such as the average distance and average speed of the Red team's group. The intention of the entire group is then inferred based on the sequence of observed variables. Individual-based intention identification begins with a single drone, designing observation variables based on individual drone characteristics such as speed. Intent is then identified for each Red team drone individually, and finally, a voting method is used to determine the probability distribution of the entire group's intentions.

[0139] Overall, the algorithm needs to identify intent for each defensive zone separately. Based on the intent analysis results and other auxiliary information, it calculates the threat index for each defensive zone. The relative size of this index serves as the basis for drone dispatch. The following will explain two intent analysis methods in detail.

[0140] 2) Group-based intention recognition (1) Intent design:

[0141] Based on the local patterns that the Red Army may adopt, its intentions are designed into five types: strong attack, feint attack, retreat, transfer, and no threat;

[0142] When using HMM modeling, it is treated as a hidden state and represented by the set Q = {q1,q2,q3,q4,q5};

[0143] A strong attack means that the Red side concentrates a large number of troops to launch a direct attack, quickly approaching the target area and forcibly breaking through the Blue side's defense line;

[0144] A feint attack is when the Red side sends out some drones as bait to conduct a false attack, attempting to draw the Blue side's firepower to tie down its defenses and create opportunities for other drones to attack.

[0145] Retreat is the Red Army's rapid withdrawal from the target area to a safe zone to avoid excessive losses, readjust its strategy and prepare for the next attack;

[0146] The red team circles around the target area, leaving the current battle zone and turning to other battle zones, trying to find a gap in the blue team's defense before attacking.

[0147] No threat means that the blue team has not observed any enemy drones in the defensive zone, but it does not mean that there are actually no enemy drones in the zone.

[0148] (2) Observation variable design:

[0149] The observed variables are processed situational information used to analyze the Red Team's intentions. To analyze group intentions, we need to select representative statistics that can distinguish each intention based on statistical characteristics. The statistics selected in this paper are the number of Red Team groups, average speed, and average speed direction. Each statistic is divided into multiple levels based on its value range, as shown in Table 1:

[0150] Table 1 Classification of various statistics (based on groups)

[0151]

[0152] Among the three types of statistics, different levels of average speed v distinguish the different attack intensities of the Red side; the direction of the average speed is represented by the angle α between the moving direction and the direction of the target area. The three ranges of direction angles in Table 1 represent approach, circle, and distance, respectively, and can be used to distinguish intentions such as force attack, transfer, and retreat. For the quantitative feature n, since the number of drones on both sides will continue to decrease as the confrontation progresses, the absolute number cannot be used as a feature. Instead, the proportion of the number is counted. The 25% division is used here to measure the degree of emphasis of the Red side in deploying troops to this area (a total of four divisions).

[0153] The three types of statistics are further combined into observation variables p1~p 12 , and additionally increase p 13 The variable indicates that no enemy drones were observed in this partition, as shown in Table 2:

[0154] Table 2 Observation variable design (based on population)

[0155]

[0156]

[0157] (3) HMM model parameter learning

[0158] In the drone game, the action strategy of the red team is constantly changing, which makes the intention recognition of the red team group time-sensitive and needs to be completed in a relatively short time. The time window size of the intention recognition is set to 5 time steps, that is, at time t, only a limited length of observation sequence {o t-4 ,o t-3 ,…,o t Before using HMM for intention recognition, it is necessary to learn the model parameters λ = (A, B, π). This paper adopts a supervised learning approach, that is, given an observation sequence O and a state sequence I, the model parameters are calculated using the maximum likelihood estimation method.

[0159] First, we need to create training data. This requires a certain simulation confrontation process in the simulation environment, collecting observation data O from the environment and labeling the intention I. Combining the characteristics of each intention and the properties of the observation variables, for the observation sequence O obtained at time t t =(o t-4 ,o t-3 ,…,o t ), whose intention t The marking method is:

[0160] Table 3 Intent labeling rules (based on groups)

[0161]

[0162] After obtaining the training data, the maximum likelihood estimation method is used to calculate the initial state probability π, the state transition probability A, and the observation probability B:

[0163]

[0164] Where s is the total number of samples, n is i (π) is the initial state q i The frequency, n ij (A) is the frequency of state i transitioning to state j, n ik (B) is the frequency of observation k in state j.

[0165] (4) Calculate the probability distribution of intention:

[0166] After obtaining the HMM model parameters λ = (A, B, π) from the above steps, the Red team's intentions can be identified in actual confrontations. During the game, the collected raw information sequence must be processed into an observation sequence at each time step. When the decision layer executes the scheduling, the corresponding finite observation sequence is extracted.

[0167] The present invention uses a forward algorithm to calculate the probability distribution of the red team's intentions, and the method is as follows.

[0168] Assume that the observation sequence obtained at time T is O T =(o1,o2,…,o T ), in state q at time t i The forward probability is α t (i), first calculate the initial value:

[0169] α1(i)=π i b i (o i ),i=1,2,…,N(8)

[0170] Next, recursively perform the following for t=1, 2, …, T-1:

[0171]

[0172] Finally, the forward probability of each state at time T is normalized to obtain the probability distribution of each intention:

[0173]

[0174] (5) Calculate the partition threat index:

[0175] The threat index of the partition includes the intention recognition part and the rule part. For the intention recognition part, the potential threat of different red team intentions is different, so this paper uses each intention q j Set an importance coefficient m j , intention Q = {force attack, feint attack, retreat, transfer, no threat} corresponding to the value of m = [0.35, 0.25, 0.15, 0.20, 0.05], and then the intention probability distribution γ obtained by the above process j The intention threat index of partition i is calculated as:

[0176]

[0177] It can be found that the average distance between the red group and the target center is not considered in the intention recognition. This is because the statistic is not easy to distinguish between different types of intentions. The impact brought about cannot be ignored, so it is considered in the rules part, the purpose is to obtain the distance threat index T dis , the calculation method is the same as formula (3).

[0178] Based on the above process, the threat index finally calculated by the group-based intention recognition method is:

[0179] T(i)=α int T int (i)+α dis T dis (i),i=1,2,3,4(12)

[0180] where α int With α dis are the weights of the intention threat index and the distance threat index, respectively. The experiment focuses on the role of intention recognition, so they are taken as 0.8 and 0.2 respectively.

[0181] 3) Individual-based intention identification (1) Design of intention and observation variables:

[0182] The individual-based intent recognition method starts with identifying the intent of a single drone and then uses voting to calculate the probability distribution of the intent of the entire group. This method is very similar to the group-based method in terms of process, so this section only describes the differences in details.

[0183] First, this method categorizes intent into four types: assault, feint, retreat, and relocation. Compared to group intent, it eliminates the "non-threatening" category. Second, the individual-based approach does not select group statistics, but instead selects individual characteristic quantities such as the speed and direction of movement of a single opposing drone. Furthermore, quantitative characteristics cannot be used to determine individual intent, simplifying the combination of observed variables. Quantitative characteristics are then considered in the rules. The classification of characteristic quantities and the design of observed variables are shown in Table 4:

[0184] Table 4 Characteristic quantity classification and observation variable design (based on individuals)

[0185]

[0186] (2) HMM model parameter learning and intent recognition:

[0187] For the learning of HMM model parameters, the supervised learning method is still used as in the population-based method, and λ = (A, B, π) is calculated using the maximum likelihood estimation method.

[0188] Based on the model parameters, the intention of a single drone is first identified. Unlike the group method, after using the forward algorithm to obtain the probability distribution of the intention of a single enemy drone at time T, the intention with the highest probability is taken as the intention of the drone. Then, based on the individual intention identification results, the probability distribution of the intention of the entire group is calculated by voting. Let the total number of enemy drones in the partition be n, and the number of enemy drones corresponding to intention i be n i , then the probability distribution of group intention is:

[0189]

[0190] (3) Calculate the partition threat index:

[0191] The threat index still includes the intention part and the rule part. In the intention part, the importance coefficients of the four intentions Q = {force attack, feint attack, retreat, transfer} are taken as m = [0.40, 0.25, 0.15, 0.20], and the intention threat index T is obtained by the calculation method of formula (11) int In the rule part, we need to consider both the number of groups n and the average distance The influence of the two aspects is calculated according to the method of formula (1) and formula (3) respectively. num With T disFinally, based on the individual intent recognition method, the overall threat index is calculated as:

[0192] T(i)=α int T int (i)+α num T num (i)+α dis T dis (i),i=1,2,3,4(14)

[0193] In the above formula, the weight α int , α num With α dis Take them as 0.6, 0.2 and 0.2 respectively.

[0194] 1.4 Scheduling timing and rules:

[0195] 1) Scheduling timing:

[0196] The decision-making layer's scheduling timing is influenced by a variety of factors. For one thing, scheduling is subject to certain constraints. As mentioned earlier, before the decision-making layer can dispatch, all drones must share observation information for aggregated analysis and centralized decision-making. This process is limited by drones' communication and data transmission capabilities and consumes significant energy. Therefore, scheduling should be limited. Furthermore, the dynamic nature of the situation requires drones to possess sufficient maneuverability and flexibility to quickly adapt to changing mission requirements and environmental conditions, and to promptly adjust their deployment and operations.

[0197] This experiment balances the communication cost and maneuverability of UAV scheduling and decides to schedule at an appropriate frequency, scheduling every 2 time steps.

[0198] 2) Scheduling rules:

[0199] The rule-based and intent-recognition-based scheduling algorithms analyze the operational situation in each defensive zone and, by calculating the threat index T, quantitatively assess the relative defensive strength required for each zone. Next, based on the threat index T, the number of drones to be deployed to each zone and the defensive area each drone is responsible for are determined.

[0200] The scheduling process requires rational planning to minimize the length and time of the drone's dispatch path while ensuring the completion of the assigned tasks. Therefore, the algorithm follows the principle of neighbor-first scheduling, with scheduling priority given to adjacent areas over those in between. For drones dispatched from one area to another, the drone with the shortest path is prioritized.

[0201] The algorithm consists of two steps: first, determining the dispatch direction and number between zones; second, determining the defense zone to which each drone is dispatched. Algorithm 5-1 describes the specific implementation process of the dispatch:

[0202]

[0203]

[0204] The action layer is specifically:

[0205] The action layer is responsible for outputting the specific actions that the drone needs to perform, and is implemented using the QMIX algorithm.

[0206] This invention focuses on how the blue team's drones can defend different partitions at the action layer after being dispatched by the decision layer.

[0207] It should be noted that the scheduling of the decision-making layer only determines the defense zone to which each drone is deployed, but does not control the movement of the drone. The actual action is output through the action layer.

[0208] The QMIX algorithm at the action layer adopts a distributed execution mode, and each UAV selects the best action based on local observation information.

[0209] Among the designed observation values, those related to the position are the coordinates (x, y) of the drone and the distance d from the target center. Since the coordinates of the target center are taken as the origin, these position information are relative to the target center.

[0210] Therefore, in order to dispatch the drone to a specific defense partition, its observation values ​​are taken as the relative coordinates (x-x0, y-y0) and the relative distance d(x-x0, y-y0) relative to the defense partition base point (x0, y0), so that the drone can defend a specific partition.

[0211] Experimental results and analysis

[0212] Comparative analysis of final indicators

[0213] The QMIX-based decision-making algorithm and the three hierarchical decision-making algorithms designed in this chapter were tested under the five attacking strategies Rule 1 to Rule 5, and the statistical results of various end-game indicators were obtained:

[0214] In the experiment, 200 rounds were tested under each attacking strategy, and the average values ​​of various end-game indicators were calculated, resulting in the statistical data in Table 5:

[0215] Table 5 Mean values ​​of various indicators of the four defensive strategies under the attacking strategy of Rule 1 to Rule 5

[0216]

[0217] for Figure 3 The statistical results of the indicators in Table 5 can be analyzed from the following two perspectives.

[0218] On the one hand, the hierarchical decision-making algorithm demonstrates significant advantages over the single-layer QMIX decision-making algorithm in all collaborative defense metrics. Looking at the average defensive win rate under the Red team's five offensive strategies, all three algorithms based on hierarchical decision-making show at least a 33% improvement over the single-layer QMIX algorithm. In particular, the decision-making algorithm based on HMM individual intent recognition achieved a 45.50% improvement. This demonstrates that hierarchical decision-making enables drone swarms to better accomplish collaborative defense tasks and demonstrate stronger combat capabilities. Hierarchical decision-making offers significant improvements over single-layer decision-making, demonstrating advantages in collaborative defense efficiency and resource utilization. Regarding time steps, while the single-layer QMIX algorithm consumes fewer time steps than the hierarchical algorithm, seemingly achieving faster combat speed, it actually loses the target area sooner.

[0219] Combining the above indicators, it can be found that the hierarchical decision-making model has obvious advantages. The reason is that it hierarchizes the decision-making process, from the overall scheduling of the decision-making layer to the specific action output of the action layer, which enables drone clusters to better cope with complex and changing situations, while improving the interpretability of the decision-making process.

[0220] On the other hand, among the three algorithms based on hierarchical decision-making models, the HMM-based individual intention recognition algorithm performed best overall, followed by the rule-based algorithm, while the HMM-based group intention recognition algorithm performed relatively poorly. First, both the rule-based and HMM-based individual intention recognition algorithms demonstrated strong defensive effectiveness, each with its own advantages. The rule-based and HMM-based individual intention recognition algorithms only require scheduling based on the number of drones in the defensive zone and the approach distance, resulting in a simple and easy-to-implement algorithm. The HMM-based individual intention recognition strategy, however, outperformed the rule-based strategy because intention recognition can extract deeper information from raw battlefield data. By modeling the Red Army's behavioral patterns through the HMM, the HMM-based individual intention recognition strategy can better understand its operational motivations. Secondly, while the HMM-based group intention recognition decision algorithm significantly improved over the single-layer QMIX decision algorithm, its advantage was not as clear in the hierarchical algorithm. This is because group intention is analyzed based on statistics such as average speed and average direction of movement, focusing on reflecting the overall behavioral trends of the Red team while ignoring the differences in individual behavior. As a result, the behavioral characteristics of some individuals are easily obscured by the group, making it difficult to accurately infer the Red team's true intentions. Finally, in terms of generalization, both the rule-based approach and the HMM-based individual intention recognition approach achieved relatively high defensive win rates under each attacking strategy. In Rule 4, the attackers adopted a massed formation, making defense more difficult, but still achieved a win rate of over 45%, demonstrating the adaptability of both defensive strategies to different environments.

[0221] Evaluation and Analysis of the Confrontation Process

[0222] In addition to evaluating various defensive strategies based on final indicators such as winning rate, the characteristics of the four types of defense can also be analyzed from the specific process of drone confrontation. Figure 4 The following graph shows the curves of the survival rate, enemy annihilation rate, and the shortest distance from the attacker to the target center over time for the four defensive strategies of the blue side under the offensive strategies of Rule 1 to Rule 5:

[0223] according to Figure 4The test results compare the three metrics of the QMIX-based single-layer decision-making algorithm with the three hierarchical decision-making algorithms, each of which changes over time. A notable characteristic is that, while the survival rate of the single-layer decision-making algorithm is high in the first half of the confrontation, it begins to drop sharply at a certain time step, reaching a final stable value much lower than that of the hierarchical decision-making algorithm. In the simulator, the confrontation process reveals that under the single-layer decision-making algorithm, the Blue team's drones are mostly concentrated near the target center, waiting until the Red team approaches sufficiently before moving outward to launch an active attack. However, by this time, it is too late; the Red team has already launched a large-scale offensive, resulting in a sharp drop in the Blue team's survival rate. In contrast, under the hierarchical decision-making algorithm, the Blue team's drones are dispatched to various defensive zones, allowing them to observe the attacking enemy drones in advance. This allows them to enter the confrontation earlier, fully prepare for the subsequent defense, and achieve a higher defense success rate. In terms of enemy kill rate, the single-layer decision-making algorithm's kill rate is almost always lower than that of the hierarchical decision-making algorithm, demonstrating the defensive lag problem inherent in the single-layer decision-making algorithm and highlighting the advantages of the hierarchical decision-making algorithm. Analyzing the attacker's shortest approach distance, in the first half of the round, due to the attacker's long distance, the drones under various strategies did not enter the confrontation phase, resulting in almost overlapping curves. However, a significant difference emerged in the second half, with the curve under the single-layer decision-making algorithm dropping the fastest, indicating low defensive efficiency and the Red side's easy entry. However, the layered decision-making algorithm was able to promptly block the attacking opponent's drones in the later stages, preventing the Red side from ever approaching the target area.

[0224] Furthermore, the HMM-based individual intention recognition strategy continues to perform the best among the four defensive strategies, according to the curve trends. This strategy boasts significant advantages in survival rate, enemy kill rate, and Red Team's approach distance, demonstrating the superior overall performance of drone collaboration and combat efficiency under this strategy, enabling more efficient completion of defensive missions.

[0225] In response to the problems of poor collaboration and defensive loopholes in the single-layer defense strategy based on QMIX, the present invention refines the defense strategy and proposes a hierarchical decision-making model based on the fusion of rules and reinforcement learning. The decision-making layer schedules drones, and the QMIX algorithm in the action layer outputs the drone's actions. In the decision-making layer, scheduling algorithms based on rules and HMM intent recognition are designed, respectively. Intent recognition is further divided into two methods: group intent-based and individual intent-based. The single-layer QMIX decision algorithm and three hierarchical decision algorithms were tested in a simulation environment. The results show that the hierarchical decision model has better collaborative defense effects than the single-layer decision model. Among them, the decision algorithm based on HMM individual intent recognition has the best effect, providing a new method and idea for achieving efficient drone cluster defense operations.

Claims

1. A UAV swarm collaborative confrontation system, characterized by: Including decision-making layer, action layer and simulation environment; The decision layer is used for intention recognition and dynamic grouping and task decomposition of drones through relevant algorithms; The action layer is used to execute the decomposed group-level tasks; The simulation environment is used to verify the performance of relevant algorithms in the decision-making layer.

2. The UAV swarm collaborative confrontation system according to claim 1, characterized in that: The decision-making layer includes demarcating defense zones, analyzing zone patterns, and dispatching drones; The above mentioned delineation of defensive zones takes into account the shape of the zones, which should ensure that the entire target area is fully covered; The partitioning forms include two methods: Method 1 is rule-based: the overall threat index of each zone is calculated based on the size of the opponent and the proximity of the Red team's opponent to the target area; The second method is based on intention: when using HMM to analyze the opponent's intention, two analysis methods are provided: group-based and individual-based; The dispatching of drones includes dispatching timing and dispatching rules; The specific method and steps for scheduling timing are: considering the communication cost and maneuverability of UAV scheduling, scheduling is carried out at an appropriate frequency; The specific rules of the scheduling rules are based on the principle of neighbor priority scheduling. The scheduling priority of adjacent areas is higher than that of intervening areas. For human-machines transferred from the same area to another area, the UAV with the shortest path is prioritized.

3. The UAV swarm collaborative confrontation system according to claim 2, characterized in that: In group-based intention recognition, the average distance and average speed statistics of the red team to the other team are used to construct observation variables, and then the intention of the entire group is inferred based on the observation variable sequence; Individual-based intention recognition starts with a single drone. Observation variables are designed based on the speed and other characteristic quantities of the individual drone. Then, the intention of each drone of the red team is recognized separately. Finally, a voting method is used to obtain the probability distribution of the intention of the entire group.

4. A method for using a drone swarm collaborative countermeasure system according to any one of claims 1 to 3, characterized in that: The following steps are included: Step 1: Set a circular target area and divide it into several defensive zones. Analyze the situation in each zone to assess the relative importance of each zone and determine the number of drones required to allocate to each zone. Comprehensively consider the dispatch path length, task urgency, and dispatch time consumption to rationally plan the dispatch timing and the dispatch path of the UAV; Step 2: A distributed execution approach is adopted at the action layer. Individual drones select actions based on local observations, and the QMIX algorithm is used to optimize and coordinate the drones' actions to complete the local collaborative defense task.

5. The method for using the UAV swarm collaborative confrontation system according to claim 4, characterized in that: The step 1 is specifically as follows: 1.1: Delineate defensive zones; 1.2: Introduce the threat index T to measure the threat level of the Red Team to each area, and use it as the basis for drone scheduling, and perform scheduling based on the rule-based scheduling algorithm; 1.3: Scheduling based on the scheduling algorithm of intent recognition; 1.4: Schedule according to the scheduling timing and the scheduling rules based on steps 1.2 and 1.

3.

6. The method for using the UAV swarm collaborative confrontation system according to claim 5, characterized in that: The step 1.1 is specifically as follows: in the decision layer, the target area is divided into four identical 90° sector areas around the center of the target area in the experiment, corresponding to the four directions respectively; Each defensive sector needs to determine a defensive base point as the gathering center for drones guarding the area, and set the base point at the midpoint of the arc section of each sector.

7. The method for using the UAV swarm collaborative confrontation system according to claim 5, characterized in that: The step 1.2 is specifically as follows: In the rule-based scheduling algorithm, the size of the enemy drones in the defensive zone is first considered. The threat index is proportional to the number of enemy drones. The normalized threat index is: Where N div is the number of defensive zones, n i Indicates the number of enemy drones observed in area i, threat index T num is a relative value, satisfying: Secondly, the proximity between the red team and the target area can be analyzed by the average distance between the red team group and the target center. To indicate that here The smaller the target area is, the greater the threat it faces. It is negatively correlated and is calculated as: where d max The maximum distance allowed by the venue must meet Combining the threat indices of the above two aspects, the overall threat index of each zone is calculated in a linear weighted form: T(i)=α num T num (i)+α dis T dis (i), i=1,2,3,4 (4) where α num With α dis is the weight of the two threat indices, satisfying α num +α num =1.

8. The method for using the UAV swarm collaborative confrontation system according to claim 5, characterized in that: The step 1.3 is specifically as follows: Situation 1) Group-based intention recognition: Based on the local patterns that the Red Army may adopt, its intentions are designed into five types: strong attack, feint attack, retreat, transfer, and no threat; When using HMM modeling, it is treated as a hidden state and represented by the set Q = {q1,q2,q3,q4,q5}; A strong attack means that the Red side concentrates a large number of troops to launch a direct attack, quickly approaching the target area and forcibly breaking through the Blue side's defense line; A feint attack is when the Red side sends out some drones as bait to conduct a false attack, attempting to draw the Blue side's firepower to tie down its defenses and create opportunities for other drones to attack. Retreat is the Red Army's rapid withdrawal from the target area to a safe zone to avoid excessive losses, readjust its strategy and prepare for the next attack; The red team circles around the target area, leaving the current battle zone and turning to other battle zones, trying to find a gap in the blue team's defense before attacking. No Threat means that the Blue Team has not observed any enemy drones in the defensive zone, but it does not mean that there are actually no enemy drones in the zone. The situation selection statistics are the number of red groups, average speed, and average speed direction, and each statistic is divided into multiple levels according to different value ranges; Among the three types of statistics, different levels of average speed v distinguish different attack intensities of the Red side; the direction of average speed is represented by the angle α between the moving direction and the target area direction. The three ranges of direction angles represent approach, circling, and moving away, respectively, and can be used to distinguish between intentions such as strong attack, transfer, and retreat; The three types of statistics are further combined into observation variables p1~p 12 , and additionally increase p 13 The variable indicates that no enemy drones were observed in this partition; The time window size of intent recognition is set to 5 time steps, that is, only a limited length of observation sequence {o t-4 ,o t-3 ,…,o t Infer the Red Team's intentions; Before using HMM for intention recognition, it is necessary to learn the model parameters λ = (A, B, π). A supervised learning approach is used, that is, given an observation sequence O and a state sequence I, the model parameters are calculated using the maximum likelihood estimation method. First, we need to create training data, conduct a certain simulated confrontation process in the simulation environment, collect observation data O from the environment and label the intention I, combine the characteristics of each intention and the properties of the observation variables, and for the observation sequence O obtained at time t t =(o t-4 ,o t-3 ,…,o t ), whose intention t The marking method is: After obtaining the training data, the maximum likelihood estimation method is used to calculate the initial state probability π, the state transition probability A, and the observation probability B: Where s is the total number of samples, n is i (π) is the initial state q i The frequency, n ij (A) is the frequency of state i transitioning to state j, n ik (B) is the frequency of observation k in state j; After obtaining the HMM model parameters λ = (A, B, π) from the above steps, the intention of the red team is identified in the actual confrontation; During the game, the collected original information sequence needs to be processed into an observation sequence at each time step. When the decision layer executes the scheduling, the finite observation sequence corresponding to that moment is extracted. The forward algorithm is used to calculate the probability distribution of the red team's intentions as follows: Assume that the observation sequence obtained at time T is O T =(o1,o2,…,o T ), in state q at time t i The forward probability is α t (i), first calculate the initial value: α1(i)=π i b i (about i ),i=1,2,…,N (8) Next, recursively perform the following for t=1, 2, …, T-1: Finally, the forward probability of each state at time T is normalized to obtain the probability distribution of each intention: For each intention q j Set an importance coefficient m j , intention Q = {force attack, feint attack, retreat, transfer, no threat}, and then the intention probability distribution γ obtained by the above process j The intention threat index of partition i is calculated as: Consider the average distance in the rule section The purpose is to obtain the distance threat index T dis , the calculation method is the same as formula (3); Based on the above process, the threat index finally calculated by the group-based intention recognition method is: T(i)=α int T int (i)+α dis T dis (i), i=1,2,3,4 (12) where α int With α dis are the weights of the intention threat index and the distance threat index respectively; 2) Individual-based intention recognition: First, this method classifies intentions into four types: strong attack, feint attack, retreat, and transfer; Secondly, we select the individual characteristic quantities of the speed and movement direction of a single enemy drone. At the same time, we cannot use quantitative characteristics to judge individual intentions, which makes the combination of observed variables more simplified. The quantitative characteristics will be considered in the rule part. Use supervised learning and use the maximum likelihood estimation method to calculate λ = (A, B, π); According to the model parameters, the intention of a single drone is first identified. After the forward algorithm is used to obtain the probability distribution of the intention of a single enemy drone at time T, the intention with the highest probability is taken as the intention of the drone. Then, based on the individual intention identification results, the probability distribution of the intention of the entire group is calculated by voting. Let the total number of enemy drones in the partition be n, and the number of enemy drones corresponding to intention i be n i , then the probability distribution of group intention is: The threat index still includes the intention part and the rule part. In the intention part, the importance coefficients of the four intentions Q = {force attack, feint attack, retreat, transfer} are taken as m = [0.40, 0.25, 0.15, 0.20], and the intention threat index T is obtained by the calculation method of formula (11) int ; In the rule part, it is necessary to consider both the number of groups n and the average distance The influence of the two aspects is calculated according to the method of formula (1) and formula (3) respectively. num With T dis Finally, based on the individual intention recognition method, the overall threat index is calculated as: T(i)=α int T int (i)+α num T num (i)+α dis T dis (i), i=1,2,3,4 (14)。 9. The method for using the UAV swarm collaborative confrontation system according to claim 8, characterized in that: The step 1.4 is specifically as follows: Schedule at an appropriate frequency, every 2 time steps; In the rule-based and intent-based scheduling algorithms, the operational situation of each defensive zone has been analyzed, and by calculating the threat index T, a quantitative assessment of the relative defensive strength required for each zone is achieved. Next, based on the threat index T, the number of drones deployed to each zone and the defensive area each drone is responsible for are determined; The scheduling process needs to be reasonably planned to minimize the length of the UAV's scheduling path and the scheduling time while ensuring the completion of the assigned tasks; Following the principle of neighbor-first scheduling, the scheduling priority of adjacent areas is higher than that of intervening areas. For drones transferred from the same area to another area, the drone with the shortest path is prioritized.

10. The method for using the UAV swarm collaborative confrontation system according to claim 1, characterized in that: The step 2 is specifically as follows: The action layer is responsible for outputting the specific actions that the drone needs to perform, which is implemented using the QMIX algorithm. The QMIX algorithm in the action layer adopts a distributed execution method, and each drone selects the best action based on local observation information. Among the designed observation values, those related to the position are the coordinates (x, y) of the drone and the distance d from the target center, where the coordinates of the target center are taken as the origin; The drone is dispatched to a specific defense zone, and its observation values ​​are taken as the relative coordinates (x-x0, y-y0) and the relative distance d(x-x0, y-y0) relative to the defense zone base point (x0, y0). In this way, the drone can defend a specific zone.