Method and system for autonomous coordinated counter-unmanned aerial vehicle swarm by unmanned aerial vehicle swarm

By employing an autonomous and collaborative countermeasure method for drone swarms, and utilizing neural networks and multi-agent reinforcement learning combined with a large language model, efficient collaborative countermeasures against drone swarms are achieved. This solves the adaptability and real-time issues of traditional anti-drone systems in dynamic environments, thereby enhancing countermeasure capabilities.

CN121613948BActive Publication Date: 2026-05-05INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF AUTOMATION CHINESE ACAD OF SCI
Filing Date
2026-01-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing counter-drone systems are ill-suited to the dynamic and changing adversarial environment, lacking adaptability, intelligence, and coordination. Traditional methods are insufficient in their generalization ability when facing large-scale drone swarms, and suffer from high decision-making complexity and poor real-time performance.

Method used

A method for autonomous and collaborative countermeasures by drone swarms is adopted. Task allocation is performed through a neural network model, and collaborative control is achieved by combining multi-agent reinforcement learning and a large language model. Action masks are used to filter invalid or collision actions, thereby realizing efficient collaborative countermeasures by drone swarms.

Benefits of technology

It improves the decision-making efficiency and safety of drone swarms in dynamic environments, and realizes autonomous, collaborative, safe and efficient countermeasure capabilities, meeting real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121613948B_ABST
    Figure CN121613948B_ABST
Patent Text Reader

Abstract

This invention provides a method for autonomous and collaborative counter-drone swarming of drones, applied in the field of anti-drone technology. The method includes: acquiring environmental information; inputting the environmental information into a neural network model for task allocation, obtaining the task allocation result output by the neural network model; inputting the task allocation result into a multi-agent reinforcement learning model for cooperative control strategy learning, obtaining a target cooperative control strategy; inputting the target cooperative control strategy into a trained large language model for knowledge enhancement processing, obtaining action control commands output by the large language model; converting the action control commands into drone flight control commands, and controlling each drone in the drone swarm based on the drone flight control commands to counter the target. This invention improves the decision-making efficiency, generalization ability, and security of anti-drone systems in dynamic environments, forming a new type of anti-drone capability of "countering drones with drones" and "using swarms to control swarms."
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of anti-drone technology, and in particular to a method for autonomous and collaborative countermeasures against drone swarms. Background Technology

[0002] With the rapid development and widespread application of drone technology, drones have shown great potential in fields such as aerial photography, search and reconnaissance, area monitoring, and precision strikes, but they have also brought new security challenges.

[0003] When drones are used for infiltration, reconnaissance, or countermeasures, traditional single-defense methods struggle to cope with their numerical superiority, coordinated mobility, and distributed nature. Especially in defensive scenarios, the short reaction time and high decision-making complexity place extremely high demands on the real-time performance, intelligence, and coordination of countermeasure systems. Most existing anti-drone systems are based on preset rules or centralized control architectures, making them ill-suited to dynamically changing adversarial environments.

[0004] This shows that the anti-drone methods in related technologies have technical problems such as low adaptability. Summary of the Invention

[0005] This invention provides a method for autonomous and collaborative counter-drone swarming, which addresses the shortcomings of existing anti-drone methods in terms of low adaptability and improves the decision-making efficiency, generalization ability, and security of anti-drone systems in dynamic environments.

[0006] This invention provides a method for autonomous and cooperative counter-drone swarming of drones, comprising the following steps: Acquiring environmental information, including the drone swarm, the target object to be countered, and obstacle information; inputting the environmental information into a neural network model for task allocation, obtaining the task allocation result output by the neural network model, wherein the task allocation result is used to specify the target object corresponding to each drone in the drone swarm; inputting the task allocation result into a multi-agent reinforcement learning model for cooperative control strategy learning, obtaining the target cooperative control strategy output by the multi-agent reinforcement learning model; inputting the target cooperative control strategy into a trained large language model for knowledge augmentation processing, obtaining the action control commands output by the large language model, wherein the knowledge augmentation processing filters invalid or collision actions through action masking; converting the action control commands into drone flight control commands, and controlling each drone in the drone swarm based on the drone flight control commands to counter the target object.

[0007] According to the present invention, a method for autonomous and collaborative countermeasures against drone swarms by drone swarms is provided, wherein the neural network model adopts a dual-tower Transformer structure.

[0008] The step of inputting the environmental information into a neural network model for task allocation and obtaining the task allocation result output by the neural network model includes:

[0009] The positions of each UAV in the UAV cluster and the position of the target object in the environmental information are respectively embedded to obtain the UAV embedding matrix and the target object embedding matrix:

[0010] The assigned logical value matrix is ​​determined by scaling the dot product cross attention based on the UAV embedding matrix and the target object embedding matrix:

[0011] ;

[0012] in, Represents the allocation of logical values ​​matrix. Represents the UAV embedding matrix. Represents the embedding matrix of the target object. These represent the number of drones and the number of target objects, respectively. Indicates transpose. Represents the vector dimension. Represents the space of real numbers;

[0013] The allocation logic value matrix is ​​normalized to obtain the allocation probability matrix;

[0014] ;

[0015] in, Represents the distribution probability matrix. Represents the normalization function;

[0016] Based on the allocation probability matrix, the task allocation result is determined:

[0017] ;

[0018] in, Indicates the first Task allocation results for each drone This indicates taking the parameter corresponding to the maximum value. Indicates the first The drone was assigned to the first The probability value of each target object.

[0019] According to the present invention, a method for autonomous and cooperative countermeasures against drone swarms, prior to inputting the environmental information into a neural network model for task allocation and obtaining the task allocation result output by the neural network model, the method further includes:

[0020] The preset environmental information samples are processed using an integer linear programming algorithm to generate an expert label set offline.

[0021] Based on the expert label set, the neural network parameters of the neural network model are optimized by minimizing the cross-entropy loss function, which is expressed by the following formula:

[0022] ;

[0023] in, Represents the cross-entropy loss function. Represents the parameters of the neural network. Indicates the number of drones, For drone indexing, Indicating in neural network parameters The following model prediction results With expert tags The probability values ​​are consistent.

[0024] According to the present invention, a method for autonomous cooperative countermeasures against unmanned aerial vehicle (UAV) swarms includes the step of inputting the task allocation result into a multi-agent reinforcement learning model for cooperative control policy learning, and obtaining the target cooperative control policy output by the multi-agent reinforcement learning model, comprising:

[0025] Based on the task allocation results, a sub-team corresponding to each target object is determined, and an encirclement relationship is established, wherein the sub-team includes at least one drone;

[0026] Based on the encirclement relationship, local observation information of each UAV in the sub-team is obtained, wherein the local observation information includes the status of all UAVs in the sub-team and the status of the target object;

[0027] Based on the local observation information, a cooperative control strategy is learned through a multi-agent proximal policy optimization algorithm to obtain the target cooperative control strategy.

[0028] According to the present invention, a method for autonomous and cooperative countermeasures against drone swarms is provided, the method further comprising:

[0029] Based on the aforementioned collaborative control strategy, the success of capturing the target object is determined according to preset capture success conditions, wherein the capture success conditions include:

[0030] Quantity condition: The number of drones in the sub-team is greater than a preset quantity threshold;

[0031] Distance condition: The error between the distance from each drone in the sub-team to the target object and the preset capture radius is less than the distance error range;

[0032] Angle uniformity condition: The distribution angle uniformity of the UAVs in the sub-team around the target object is less than the angle error range.

[0033] According to the present invention, a method for autonomous cooperative countermeasures against drone swarms, before inputting the target cooperative control strategy into a trained large language model for knowledge augmentation processing to obtain the action control commands output by the large language model, the method further includes:

[0034] Obtain a dataset of state observations of the interactions between each drone in the drone cluster and the target object;

[0035] The state observation dataset is input into the trained large language model to generate a corresponding binary action mask set as a label, wherein the binary action mask set is used to identify actions that are invalid or have a collision risk under the current observation.

[0036] Based on the state observation dataset and the binary action mask set, a pre-defined lightweight action mask generation network is trained using supervised learning to obtain a trained action mask generation network.

[0037] The supervised learning uses a binary cross-entropy loss function to optimize network parameters. The trained action mask generation network is used to replace the large language model to generate action masks, and the action masks are applied to the output of the large language model.

[0038] This invention also provides an autonomous and collaborative counter-drone swarm system, comprising the following modules:

[0039] The acquisition module is used to acquire environmental information, which includes information on the drone cluster, the target object to be countered, and obstacles.

[0040] The task allocation module is used to input the environmental information into the neural network model to allocate tasks and obtain the task allocation result output by the neural network model. The task allocation result is used to specify the target object corresponding to each drone in the drone cluster.

[0041] The coordination and control module is used to input the task allocation result into the multi-agent reinforcement learning model to learn the cooperative control strategy, and obtain the target cooperative control strategy output by the multi-agent reinforcement learning model.

[0042] The knowledge enhancement module is used to input the target cooperative control strategy into the trained large language model for knowledge enhancement processing to obtain the action control instructions output by the large language model. The knowledge enhancement processing filters invalid actions or collision actions through action masking.

[0043] The control module is used to convert the motion control commands into UAV flight control commands, and control each UAV in the UAV cluster based on the UAV flight control commands to counter the target object.

[0044] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for autonomous and cooperative countermeasures against drone swarms.

[0045] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for autonomous and cooperative counter-drone swarming as described above.

[0046] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for autonomous and cooperative counter-drone swarming of drones.

[0047] The present invention provides an autonomous and collaborative method for countering drone swarms. By acquiring environmental information, it provides a comprehensive situational awareness foundation for the system, ensuring real-time monitoring of the drone swarm, target objects, and obstacles. Next, a neural network model allocates tasks, optimizing the resource configuration of the drone swarm and clearly defining the countermeasure targets for each drone, thus improving the orderliness of task execution. Then, a multi-agent reinforcement learning model learns collaborative control strategies, enabling the drone swarm to dynamically adapt to the environment and achieve efficient collaborative action. Furthermore, the trained large language model undergoes knowledge augmentation processing, using action masks to filter invalid or colliding actions, enhancing the safety and accuracy of decision-making. Finally, the action control commands are converted into drone flight control commands, driving the drone swarm to accurately execute countermeasures, thereby achieving an overall autonomous, collaborative, safe, and efficient countermeasure capability for the drone swarm. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0049] Figure 1 This is a flowchart illustrating the method for autonomous and collaborative counter-drone swarming provided by the present invention.

[0050] Figure 2This is a schematic diagram of our drone surrounding and capturing a target drone, provided by the present invention.

[0051] Figure 3 This is a schematic diagram of the module of the UAV swarm autonomous collaborative counter-UAV swarm system provided by the present invention.

[0052] Figure 4 This is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0054] Traditional anti-munitions methods mainly include physical interception, electronic jamming, and laser destruction. These methods show significant shortcomings in multi-target, highly mobile swarm warfare: rule-based systems lack learning and adaptability, making them unable to cope with unknown tactics; centralized control architectures suffer from communication delays and single-point-of-failure risks, making it difficult to support distributed decision-making; optimization algorithms such as integer linear programming, while providing theoretically optimal solutions, have high computational complexity and cannot meet the requirements for millisecond-level real-time response. Furthermore, existing methods are mostly designed for single-target or small-scale scenarios, and their generalization ability is severely insufficient when facing large-scale swarm attacks.

[0055] In recent years, multi-agent reinforcement learning (MAL) has provided new insights into autonomous collaborative decision-making. For example, some studies have achieved local collaboration through distributed policy learning or introduced trapping incentives into the reward function to guide agent behavior. However, these methods still face significant bottlenecks: First, MRL training is inefficient, especially in large-scale scenarios requiring extensive interactive data and exhibiting slow convergence; second, the generalization ability of learned policies is poor, with performance plummeting when the number, speed, or environmental layout of targets changes; furthermore, the joint action space grows exponentially with the number of agents, making exploration difficult and hindering the learning of effective collaborative strategies. Although some studies have attempted to introduce prior knowledge to improve learning efficiency, the knowledge sources are singular, and the integration methods are crude, failing to systematically combine domain knowledge, expert experience, and learning capabilities. Especially in real-world adversarial situations, the complex and ever-changing landscape means that purely data-driven methods lack an understanding of physical constraints, tactical rules, and security boundaries, potentially leading to high decision-making risks and insufficient reliability. Therefore, how to organically integrate knowledge enhancement mechanisms with hierarchical learning architectures to achieve efficient, secure, and scalable collaborative anti-drone technology remains a key challenge to overcome.

[0056] To address the aforementioned issues in existing anti-drone technologies and meet the practical requirements of modern drone countermeasures, this invention proposes an autonomous and collaborative method for countering drone swarms. This method achieves efficient target allocation at the upper layer and integrates graph attention mechanisms and action masks at the lower layer. It effectively combines prior knowledge from large models with data-driven learning, significantly improving the system's decision-making efficiency, generalization ability, and security in dynamic environments.

[0057] refer to Figure 1 , Figure 1 This is a flowchart illustrating the method for autonomous and collaborative counter-drone swarming provided by the present invention, as shown below. Figure 1 As shown, the method includes the following steps.

[0058] Step 101: Obtain environmental information, which includes information on the drone swarm, the target object to be countered, and obstacles.

[0059] In this embodiment of the invention, complete environmental information is acquired in real time through a variety of pre-deployed sensors.

[0060] For example, visual sensors, lidar, and millimeter-wave radar deployed on each drone work together to perceive the surrounding airspace. Visual sensors identify and initially track targets to be countered, acquiring their appearance features and movement trends; lidar precisely measures the relative distance, orientation, and contours of the target, other drones, and various obstacles, generating high-precision point cloud data; and millimeter-wave radar provides target velocity information and compensates for perception limitations in adverse weather conditions. Simultaneously, each drone reports its own position, velocity, and attitude information in real time through its onboard satellite positioning module and inertial measurement unit.

[0061] In some embodiments, the target to be countered is at least one target drone that our drone swarm needs to counter.

[0062] Step 102: Input the environmental information into the neural network model for task allocation, and obtain the task allocation result output by the neural network model.

[0063] The task allocation result is used to specify the target object corresponding to each drone in the drone cluster.

[0064] In this embodiment of the invention, a neural network model based on a graph neural network is constructed. This model treats the UAV and each target object as nodes in a graph. Their relative positions, velocity vectors, and mission suitability (such as remaining battery power and the type of countermeasures equipment mounted) are encoded as node features, while the relationships between the UAV and the target object, and between target objects themselves, are modeled as edges in the graph. The graph neural network model learns complex cluster-target interaction relationships through multi-layer message passing and aggregation, ultimately outputting an allocation probability matrix.

[0065] The assignment probability matrix is ​​passed through a differentiable optimal matching layer to generate deterministic task assignment results, explicitly specifying the target object corresponding to each drone in the drone swarm. This neural network model is pre-trained using a large amount of historical task data and simulated adversarial data, enabling near-optimal real-time computation.

[0066] Step 103: Input the task allocation result into the multi-agent reinforcement learning model to learn the cooperative control strategy, and obtain the target cooperative control strategy output by the multi-agent reinforcement learning model.

[0067] In this embodiment of the invention, a centralized training and distributed execution architecture is adopted.

[0068] During the training phase, a drone swarm is run in a high-fidelity simulation environment. Each drone acts as an agent, and its observation information includes its own state, the state of its assigned target, the states of neighboring drones, and local obstacle information. All agents share a centralized policy evaluation network, but each has its own policy execution network. By designing a comprehensive reward function that includes rewards for successful counterattacks, rewards for approaching the target, penalties for avoiding collisions, and penalties for energy consumption, and using multi-agent reinforcement learning algorithms (such as MAPPO) for long-term training, the agents learn complex tactics such as coordinated encirclement, interception, and expulsion.

[0069] In the actual execution phase, each UAV can generate its own flight control strategy in a distributed manner by relying only on the local policy execution network and real-time observation. The collection of these strategies constitutes the target cooperative control strategy.

[0070] Step 104: Input the target collaborative control strategy into the trained large language model for knowledge augmentation processing to obtain the action control instructions output by the large language model.

[0071] Among them, knowledge augmentation processing filters out invalid or colliding actions through action masks.

[0072] In this embodiment of the invention, an action mask training dataset is constructed, and labeled samples are generated based on historical flight data and a simulation environment using a large language model. Each sample contains UAV state information, environmental situation data, and a corresponding set of valid actions. A lightweight action mask generation network is trained using supervised learning. This network employs a binary classifier structure and outputs a binary decision to retain or mask each possible action.

[0073] During the real-time inference phase, the trained action mask generation network receives current UAV state observations, including its own position, velocity, heading angle, and the relative positions of teammate UAVs, target objects, and obstacles. Based on learned safety rules and constraints, the action mask generation network quickly calculates the feasibility score for each action and generates a corresponding binary action mask vector.

[0074] The generated action mask is applied to the output distribution of the original cooperative control strategy. This is achieved by modifying the action sampling probability, masking action options identified as invalid or high-risk, while increasing the probability of selecting safe actions. A mask-based softmax function is used to recalculate the action distribution, ensuring that the output commands conform to both tactical objectives and safety constraints.

[0075] It should be noted that, to meet real-time requirements, this invention employs knowledge distillation technology to transfer the decision-making capabilities of a large language model to a lightweight network. By designing a specialized loss function to optimize network parameters, inference speed is significantly improved while maintaining decision quality.

[0076] Step 105: Convert the motion control commands into UAV flight control commands, and control each UAV in the UAV cluster based on the UAV flight control commands to counter the target object.

[0077] In this embodiment of the invention, the motion control commands include specific motion parameters and action sequences. The conversion process first parses the command format, decomposing it into UAV flight control commands that the UAV can recognize. Each UAV flight control command includes basic flight parameters such as speed commands, heading angle commands, and altitude commands, as well as triggering conditions for specific countermeasure actions. The conversion algorithm maps high-level motion commands to specific low-level control signals such as control surface deflection and motor speed, based on the UAV's dynamic characteristics and current flight state.

[0078] The converted flight control commands are distributed to each drone. The drone's onboard flight control system independently receives and executes the commands while maintaining coordination with other members of the swarm. During execution, the drones adjust their flight attitude and trajectory according to the command requirements to perform countermeasures such as encirclement and interception. Each drone provides real-time feedback on its execution status through sensors.

[0079] During the command execution phase, information synchronization among drones is ensured. When a drone detects an environmental change or execution anomaly, it immediately broadcasts a status update to the cluster via the communication network. Control commands are dynamically adjusted based on the new environmental information to maintain coordinated cluster actions. This mechanism, combining distributed decision-making with centralized coordination, guarantees both the autonomy of individual drones and the unity of overall action.

[0080] In some embodiments, the drone swarm adopts differentiated countermeasures for different types of targets; for high-speed maneuvering targets, it implements multi-directional encirclement and capture; for swarm targets, it adopts group encirclement; all countermeasures are carried out under the premise of ensuring safety and avoiding secondary damage.

[0081] Through this embodiment of the invention, environmental information is acquired to provide a comprehensive situational awareness foundation for the system, ensuring real-time monitoring of the drone swarm, target objects, and obstacles. Next, a neural network model allocates tasks, optimizing the resource configuration of the drone swarm, clearly specifying the countermeasure targets for each drone, and improving the orderliness of task execution. Then, a multi-agent reinforcement learning model learns cooperative control strategies, enabling the drone swarm to dynamically adapt to the environment and achieve efficient cooperative action. Furthermore, the trained large language model undergoes knowledge augmentation processing, using action masks to filter invalid or collision actions, enhancing the safety and accuracy of decision-making. Finally, action control commands are converted into drone flight control commands, driving the drone swarm to accurately execute countermeasures, thereby achieving an overall autonomous, cooperative, safe, and efficient countermeasure capability for the drone swarm.

[0082] According to the present invention, a method for autonomous and collaborative countermeasures against drone swarms is provided, wherein the neural network model adopts a dual-tower Transformer structure.

[0083] Environmental information is input into a neural network model for task allocation, and the task allocation results output by the neural network model are obtained, including:

[0084] The positions of each UAV in the UAV swarm and the target object in the environmental information are respectively embedded to obtain the UAV embedding matrix and the target object embedding matrix:

[0085] By scaling the dot product cross attention, the assigned logical value matrix is ​​determined based on the UAV embedding matrix and the target object embedding matrix:

[0086] ;

[0087] in, Represents the allocation of logical values ​​matrix. Represents the UAV embedding matrix. Represents the embedding matrix of the target object. These represent the number of drones and the number of target objects, respectively. Indicates transpose. Represents the vector dimension. Represents the space of real numbers;

[0088] Normalize the allocation logic value matrix to obtain the allocation probability matrix;

[0089] ;

[0090] in, Represents the distribution probability matrix. Represents the normalization function;

[0091] Based on the assignment probability matrix, the task assignment result is determined:

[0092] ;

[0093] in, Indicates the first Task allocation results for each drone This indicates taking the parameter corresponding to the maximum value. Indicates the first The drone was assigned to the first The probability value of each target object.

[0094] In this embodiment of the invention, integer linear programming is used to generate an allocation sample set. The allocation sample set is then used to train a neural network model.

[0095] The neural network model uses a dual-tower Transformer, first embedding the positions of the drone and the target object:

[0096] ;

[0097] in, Indicates the first Embedded representation of a drone Indicates parameterization A three-layer perceptron is used as an embedding function; Indicates the first The location information of a drone is usually a two-dimensional or three-dimensional coordinate vector (such as latitude and longitude or Cartesian coordinates). express 3D real space. Indicates the first Embedded representation of a target object Indicates parameterization The three-layer perceptron, as an embedding function, Indicates the first The location information of a target object is usually a two-dimensional or three-dimensional coordinate vector.

[0098] Calculate the assigned logits by scaling the dot product cross attention:

[0099] ;

[0100] in, , , and Let represent the number of drones and the number of target objects, respectively. Thus, we obtain the allocation probability matrix:

[0101]

[0102] The final decision is:

[0103] ;

[0104] In this embodiment of the invention, the position information of each UAV in the UAV swarm and the position information of the target object are respectively processed by embedding representation. Two independent three-layer perceptron neural networks are used to map the position coordinates of the UAVs and the target object to a high-dimensional vector space, generating corresponding embedding vectors. The embedding vectors of all UAVs are combined to form a UAV embedding matrix, and the embedding vectors of all target objects are combined to form a target object embedding matrix.

[0105] The UAV embedding matrix and the target object embedding matrix are multiplied together, and the results are scaled. The scaling factor is the reciprocal of the square root of the embedding vector dimension. The final assigned logical value matrix reflects the degree of matching between each UAV and each target object.

[0106] The allocation logic value matrix is ​​subjected to softmax normalization to convert it into a probability distribution form. This process is performed row-wise to ensure that the sum of the probability values ​​assigned to each target object for each drone is 1. Through exponential operations and normalization calculations, the logic values ​​are converted into probability values ​​between 0 and 1, forming the allocation probability matrix.

[0107] For each drone, the index of the location containing the maximum value in its corresponding probability distribution is selected; this index is the target object number assigned to that drone. By traversing all drones, the task allocation decision for the entire cluster is completed, ensuring that each drone clearly understands the target object it needs to counter.

[0108] In practical deployments, this neural network model achieves efficient task allocation through a combination of offline training and online inference. During the training phase, expert samples generated using integer linear programming are used for knowledge distillation to optimize model parameters. During the inference phase, the trained model is directly used for forward computation, enabling task allocation for large-scale clusters to be completed within milliseconds, meeting real-time countermeasure requirements.

[0109] This invention utilizes a dual-tower Transformer structure and a scaling dot product cross-attention mechanism to efficiently model the interaction between UAVs and target objects, generating a reasonable task allocation strategy. Compared to traditional methods, it can more accurately and dynamically achieve many-to-many collaborative countermeasures, improving cluster response speed and adversarial effectiveness.

[0110] According to the present invention, a method for autonomous and cooperative countermeasures against drone swarms, before inputting environmental information into a neural network model for task allocation and obtaining the task allocation result output by the neural network model, the method further includes:

[0111] The preset environmental information samples are processed using an integer linear programming algorithm to generate an expert label set offline.

[0112] Based on the expert label set, the neural network parameters of the neural network model are optimized by minimizing the cross-entropy loss function, which is expressed by the following formula:

[0113] ;

[0114] in, Represents the cross-entropy loss function. Represents the parameters of the neural network. Indicates the number of drones, For drone indexing, Indicating in neural network parameters The following model prediction results With expert tags The probability values ​​are consistent.

[0115] In this embodiment of the invention, during the training phase, an integer linear programming algorithm is used to generate expert labels offline. Network parameters Optimize by minimizing the cross-entropy loss:

[0116] ;

[0117] An integer linear programming algorithm is used to process pre-defined environmental information samples to generate a high-quality training dataset. These environmental information samples include the swarm status of UAVs, target object distribution, and obstacle layout under various typical combat scenarios. Based on optimal assignment theory, the integer linear programming algorithm calculates the theoretically optimal task assignment scheme for each sample scenario. These schemes constitute an expert label set. Each label in the expert label set explicitly specifies the optimal target object that each UAV should correspond to under a specific environmental situation.

[0118] Based on the generated set of expert labels, a supervised training process is performed on the neural network model. During training, environmental information samples are used as input to the neural network model, and the corresponding expert labels are used as supervision signals. The neural network parameters are optimized by minimizing the cross-entropy loss function, so that the model's predicted output gradually approximates the expert label assignment results. The cross-entropy loss function is specifically calculated by summing the differences between the prediction result of each UAV and the expert label, and a logarithmic probability measure is used to ensure the differentiability and optimization stability of the loss function.

[0119] The training process is essentially a knowledge transfer process that distills the solution capabilities of integer linear programming algorithms into a neural network model. By continuously adjusting its internal parameters, the neural network model learns the thought process of making optimal decisions under complex constraints using integer linear programming algorithms. This distillation training allows the neural network model to retain the theoretical optimality of integer linear programming while possessing the rapid reasoning capabilities that the former lacks, enabling it to complete task allocation decisions for large-scale clusters within milliseconds.

[0120] Throughout the training process, a quality control mechanism is implemented. By dividing the training and validation sets, the model's generalization performance on unseen data is monitored to prevent overfitting. An early stopping strategy is employed to terminate training promptly when the model's performance tends to stabilize, ensuring that the final neural network model possesses both good assignment accuracy and efficient inference speed.

[0121] Through the embodiments of the present invention, expert labels are generated by integer linear programming, and the neural network is trained under supervision using cross-entropy loss, enabling the model to learn the optimal task allocation strategy, thereby improving allocation accuracy and generalization ability, and balancing performance optimization and real-time performance.

[0122] According to the present invention, a method for autonomous cooperative countermeasures against unmanned aerial vehicle (UAV) swarms involves inputting task allocation results into a multi-agent reinforcement learning model to learn a cooperative control strategy, thereby obtaining a target cooperative control strategy output by the multi-agent reinforcement learning model, including:

[0123] Based on the task allocation results, determine the sub-team corresponding to each target object and establish an encirclement relationship. The sub-team includes at least one drone.

[0124] Based on the encirclement relationship, local observation information of each UAV in the sub-team is obtained, including the status of all UAVs in the sub-team and the status of the target object.

[0125] Based on local observation information, a cooperative control strategy is learned through a multi-agent proximal policy optimization algorithm to obtain the target cooperative control strategy.

[0126] In this embodiment of the invention, a graph attention mechanism is used to model the state in order to characterize the interaction relationship between drone clusters.

[0127] At every moment, the sub-team The drones inside ) Receive local observation information ,in Indicates drone state, position ,speed ( (representing three-dimensional real space), yaw angle The state of the target object is denoted as , respectively representing sub-teams The corresponding target object's position, velocity, and yaw angle. This policy model is trained using a centralized training and distributed execution framework, and the optimization algorithm is a multi-agent near-field policy optimization.

[0128] During the training phase, the centralized value function utilizes global information to improve policy learning performance; during the execution phase, the distributed agent relies only on local observations to make decisions, thereby achieving efficient distributed control.

[0129] Once the encirclement relationship is established, local observation information is acquired for each UAV within a sub-team. This local observation information comprises two main parts: the status data of all friendly UAVs within the sub-team, and the real-time status of the target UAV. Specific status information includes the UAV's position coordinates, three-dimensional velocity vector, yaw angle, attitude, and other motion parameters. This observation data is collected in real time through a sensor system and shared within the sub-team, providing a data foundation for collaborative decision-making.

[0130] A multi-agent proximal policy optimization algorithm is employed for cooperative control policy learning. This algorithm is based on a framework of centralized training and distributed execution. During the training phase, it utilizes global information to optimize the policy, while during the execution phase, each UAV makes independent decisions based solely on local observations. Each UAV acts as an independent agent, with its policy network receiving local observation information and outputting specific control actions.

[0131] During training, policy parameters are optimized using extensive interactive data from simulated environments. Each agent updates its policy network using a policy gradient method, focusing on learning key skills required for encirclement, such as cooperative maneuvers and obstacle avoidance. The training objective is to minimize task completion time while satisfying safety constraints. In the actual execution phase, each UAV autonomously generates control actions based on the trained policy network and real-time local observations, achieving distributed cooperative decision-making.

[0132] The multi-agent reinforcement learning model outputs a target cooperative control strategy. This strategy, manifested as the optimal action selection rules for each UAV under specific observation conditions, guides the UAV swarm to achieve efficient encirclement operations.

[0133] Through the embodiments of the present invention, the target cooperative control strategy output ensures the coordination of actions among UAVs within the sub-team, while avoiding collisions with obstacles or other UAVs, thus ensuring the safety and effectiveness of countermeasures.

[0134] According to the present invention, a method for autonomous and cooperative countermeasures against drone swarms is provided, the method further includes:

[0135] Based on a collaborative control strategy, the system determines whether the target has been successfully captured according to preset capture success conditions. These capture success conditions include:

[0136] Quantity condition: The number of drones in the sub-team is greater than the preset quantity threshold;

[0137] Distance condition: The error between the distance from each drone in the sub-team to the target object and the preset capture radius is less than the distance error range;

[0138] Angle uniformity condition: The distribution angle uniformity of UAVs around the target object in the sub-team is less than the angle error range.

[0139] In this embodiment of the invention, the movement of the drone swarm is guided to gradually form an ideal encirclement formation, while avoiding collisions with obstacles or other drones. For any target object... ( (representing the set of target objects to be countered), if there exists a sub-team ( If our drone swarm meets the following three conditions, the target is considered to have been successfully captured:

[0140] Quantity condition: The number of drones in the sub-team Not less than the preset quantity threshold B :

[0141] ;

[0142] Distance condition: Each drone in the sub-team To the target object The distance is approximately equal to the encirclement radius. =10:

[0143] ;

[0144] Angular uniformity condition: The capture drones are approximately uniformly distributed around the target. Indicates the angle value. Indicates the first One drone (relative to the target object) ) angle value, Indicates the first One drone (relative to the target object) (angle value)

[0145] ;

[0146] in, Indicates the range of distance error. This indicates the angular error range. When the cumulative time steps to the target (e.g., 10) meet the above conditions, the target object is considered to have been successfully captured. The mission is completed when all target objects are successfully captured.

[0147] The goal of drone swarms is to minimize mission completion time while avoiding collisions.

[0148] ;

[0149] in, This indicates the time required to complete the encirclement under the joint strategy π. Let π represent the mathematical expectation of the joint strategy.

[0150] refer to Figure 2 , Figure 2 This is a schematic diagram of our drone surrounding and capturing a target drone provided by the present invention, which includes a drone and a target drone (i.e., the target object).

[0151] According to the present invention, a method for autonomous cooperative countermeasures against unmanned aerial vehicle (UAV) swarms, before inputting the target cooperative control strategy into a trained large language model for knowledge augmentation processing to obtain the action control commands output by the large language model, the method further includes:

[0152] Obtain a dataset of state observations of interactions between each drone and the target object in a drone swarm.

[0153] The state observation dataset is input into the trained large language model to generate a corresponding set of binary action masks as annotations. The set of binary action masks is used to identify actions that are invalid or have a risk of collision under the current observation.

[0154] Supervised learning training is performed on a pre-defined lightweight action mask generation network based on a state observation dataset and a binary action mask set to obtain a trained action mask generation network.

[0155] In this process, supervised learning uses a binary cross-entropy loss function to optimize network parameters. The trained action mask generation network is used to replace the large language model to generate action masks, and the action masks are applied to the output of the large language model.

[0156] In this embodiment of the invention, at each time step, the large language model receives the current UAV position and yaw angle, the position and velocity of teammate UAVs, and the position and velocity of surrounding target objects, and outputs a binary action mask. This mask reduces the exploration space and avoids collisions. Specifically, Time Drone The observation vector is The motion control command is The mask model is defined as follows:

[0157] ;

[0158] in, It is a drone The state space, It is a discrete action space. This represents the action mask vector. If... ( Indicates in Time of the first For actions using drones (The actual mask label value), then the action It will be blocked and cannot be executed; otherwise, it can be executed.

[0159] Since calling large models for online inference at every time step is extremely slow, the prior knowledge of large models can be distilled into a lightweight, deployable action mask generation network.

[0160] Specifically, a dataset (i.e., a state observation dataset) is built through interaction. And use the real mask (i.e., the set of binary action masks) generated by the large model. Each observation vector is labeled. The action mask generation network is then trained using supervised learning. Since the action mask is binary, a Bernoulli distribution is used for each action, and the binary cross-entropy loss function is optimized. :

[0161] ;

[0162] in, These represent the drone index and the time step index, respectively. Indicates in Time of the first For actions using drones The actual mask label values ​​are generated by the large language model; This indicates that the lightweight motion mask generation network is in The first time generated For actions using drones The predicted mask value, Actions that indicate prediction The probability that it should be retained.

[0163] During the policy learning phase, the action mask predicted by the action mask generation network is applied to the policy model output. Let... Indicates an action that is not normalized. This represents the predicted action mask. The formula for calculating the action distribution after the action mask is:

[0164]

[0165] Then from Sampling action and execution, Indicates in the observation vector Next, the drone selects actions. The probability distribution, Represents the normalization function. This represents the unnormalized action logits output by the multi-agent reinforcement learning model. The action mask is a binary vector (each element takes the value 0 or 1).

[0166] In summary, to address the issues of poor generalization ability and low decision-making efficiency of existing anti-drone systems in dynamic adversarial environments, this invention proposes an autonomous and collaborative method for countering drone swarms. The upper layer employs a target allocation network based on neural networks to decompose complex multi-target adversarial tasks into multiple single-target encirclement sub-tasks. The lower layer combines graph attention networks and knowledge-enhanced action masking mechanisms to achieve efficient collaborative motion control of the drone swarm. Through prior knowledge distillation of a large language model and expert-guided training, the method significantly improves adaptability in scenarios with varying target numbers, speeds, and maneuverability. It possesses strong generalization and high real-time performance, making it suitable for autonomous and collaborative countering drones in complex environments.

[0167] The following describes the UAV swarm autonomous collaborative counter-UAV swarm system provided by the present invention. The UAV swarm autonomous collaborative counter-UAV swarm system described below can be referred to in correspondence with the UAV swarm autonomous collaborative counter-UAV swarm method described above.

[0168] refer to Figure 3 , Figure 3 This is a schematic diagram of the module of the UAV swarm autonomous collaborative counter-UAV swarm system provided by the present invention.

[0169] The acquisition module 301 is used to acquire environmental information, including the drone cluster, the target object to be countered, and obstacle information.

[0170] The task allocation module 302 is used to input environmental information into the neural network model for task allocation and obtain the task allocation result output by the neural network model. The task allocation result is used to specify the target object corresponding to each drone in the drone cluster.

[0171] The coordination and control module 303 is used to input the task allocation result into the multi-agent reinforcement learning model to learn the cooperative control strategy and obtain the target cooperative control strategy output by the multi-agent reinforcement learning model.

[0172] The knowledge enhancement module 304 is used to input the target cooperative control strategy into the trained large language model for knowledge enhancement processing, and obtain the action control instructions output by the large language model. The knowledge enhancement processing filters invalid actions or collision actions through action masking.

[0173] The control module 305 is used to convert motion control commands into UAV flight control commands, and control each UAV in the UAV cluster based on the UAV flight control commands to counter the target object.

[0174] Specifically, the UAV swarm autonomous collaborative counter-UAV swarm system provided by the present invention can realize all the method steps implemented in the above-mentioned UAV swarm autonomous collaborative counter-UAV swarm method embodiment, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.

[0175] Figure 4 This is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communications bus 440. The processor 410 can call logical instructions in the memory 430 to execute a method for autonomous and cooperative counter-drone swarming of drones. This method includes: acquiring environmental information, including the drone swarm, the target object to be countered, and obstacle information; inputting the environmental information into a neural network model for task allocation, obtaining the task allocation result output by the neural network model, wherein the task allocation result is used to specify the target object corresponding to each drone in the drone swarm; inputting the task allocation result into a multi-agent reinforcement learning model for cooperative control strategy learning, obtaining the target cooperative control strategy output by the multi-agent reinforcement learning model; inputting the target cooperative control strategy into a trained large language model for knowledge augmentation processing, obtaining the action control instructions output by the large language model, wherein the knowledge augmentation processing filters invalid or collision actions through action masks; converting the action control instructions into drone flight control instructions, and controlling each drone in the drone swarm based on the drone flight control instructions to counter the target object.

[0176] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0177] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the autonomous cooperative counter-drone swarm method provided by the above methods. The method includes: acquiring environmental information, wherein the environmental information includes a drone swarm, a target object to be countered, and obstacle information; inputting the environmental information into a neural network model for task allocation, and obtaining the task allocation result output by the neural network model, wherein the task allocation result is used to specify the target object corresponding to each drone in the drone swarm; inputting the task allocation result into a multi-agent reinforcement learning model for cooperative control strategy learning, and obtaining the target cooperative control strategy output by the multi-agent reinforcement learning model; inputting the target cooperative control strategy into a trained large language model for knowledge enhancement processing, and obtaining the action control command output by the large language model, wherein the knowledge enhancement processing filters invalid actions or collision actions through action masking; converting the action control command into drone flight control command, and controlling each drone in the drone swarm based on the drone flight control command to counter the target object.

[0178] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the method for autonomously and collaboratively countering a drone swarm provided by the above methods. The method includes: acquiring environmental information, wherein the environmental information includes a drone swarm, a target object to be countered, and obstacle information; inputting the environmental information into a neural network model for task allocation, obtaining a task allocation result output by the neural network model, wherein the task allocation result is used to specify the target object corresponding to each drone in the drone swarm; inputting the task allocation result into a multi-agent reinforcement learning model for cooperative control strategy learning, obtaining a target cooperative control strategy output by the multi-agent reinforcement learning model; inputting the target cooperative control strategy into a trained large language model for knowledge augmentation processing, obtaining action control instructions output by the large language model, wherein the knowledge augmentation processing filters invalid actions or collision actions through action masks; converting the action control instructions into drone flight control instructions, and controlling each drone in the drone swarm based on the drone flight control instructions to counter the target object.

[0179] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0180] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for autonomous and collaborative countermeasures against unmanned aerial vehicle (UAV) swarms, characterized in that, include: Acquire environmental information, including information on the drone swarm, the target object to be countered, and obstacles; The environmental information is input into a neural network model for task allocation, and the task allocation result output by the neural network model is obtained. The task allocation result is used to specify the target object corresponding to each drone in the drone cluster. The task allocation result is input into a multi-agent reinforcement learning model to learn a cooperative control strategy, and the target cooperative control strategy output by the multi-agent reinforcement learning model is obtained. The target collaborative control strategy is input into a trained large language model for knowledge augmentation processing to obtain the action control instructions output by the large language model. The knowledge augmentation processing filters invalid or collision actions through action masking. The motion control commands are converted into UAV flight control commands, and each UAV in the UAV cluster is controlled based on the UAV flight control commands to counter the target object; The neural network model adopts a dual-tower Transformer structure. The step of inputting the environmental information into a neural network model for task allocation and obtaining the task allocation result output by the neural network model includes: The positions of each UAV in the UAV cluster and the position of the target object in the environmental information are respectively embedded to obtain the UAV embedding matrix and the target object embedding matrix: The assigned logical value matrix is ​​determined by scaling the dot product cross attention based on the UAV embedding matrix and the target object embedding matrix: ; in, Represents the allocation of logical values ​​matrix. Represents the UAV embedding matrix. Represents the embedding matrix of the target object. These represent the number of drones and the number of target objects, respectively. Indicates transpose. Represents the vector dimension. Represents the space of real numbers; The allocation logic value matrix is ​​normalized to obtain the allocation probability matrix; ; in, Represents the distribution probability matrix. Represents the normalization function; Based on the allocation probability matrix, the task allocation result is determined: ; in, Indicates the first Task allocation results for each drone This indicates taking the parameter corresponding to the maximum value. Indicates the first The drone was assigned to the first The probability value of each target object.

2. The method for autonomous and collaborative counter-drone swarming of drones according to claim 1, characterized in that, Before inputting the environmental information into the neural network model for task allocation and obtaining the task allocation result output by the neural network model, the method further includes: The preset environmental information samples are processed using an integer linear programming algorithm to generate an expert label set offline. Based on the expert label set, the neural network parameters of the neural network model are optimized by minimizing the cross-entropy loss function, which is expressed by the following formula: ; in, Represents the cross-entropy loss function. Represents the parameters of the neural network. Indicates the number of drones, For drone indexing, Indicating in neural network parameters The following model prediction results With expert tags The probability value is consistent.

3. The method for autonomous and collaborative counter-drone swarming of drones according to claim 1, characterized in that, The step of inputting the task allocation result into a multi-agent reinforcement learning model for cooperative control policy learning, and obtaining the target cooperative control policy output by the multi-agent reinforcement learning model, includes: Based on the task allocation results, a sub-team corresponding to each target object is determined, and an encirclement relationship is established, wherein the sub-team includes at least one drone; Based on the encirclement relationship, local observation information of each UAV in the sub-team is obtained, wherein the local observation information includes the status of all UAVs in the sub-team and the status of the target object; Based on the local observation information, a cooperative control strategy is learned through a multi-agent proximal policy optimization algorithm to obtain the target cooperative control strategy.

4. The method for autonomous and collaborative counter-drone swarming of drones according to claim 3, characterized in that, The method further includes: Based on the aforementioned collaborative control strategy, the success of capturing the target object is determined according to preset capture success conditions, wherein the capture success conditions include: Quantity condition: The number of drones in the sub-team is greater than a preset quantity threshold; Distance condition: The error between the distance from each drone in the sub-team to the target object and the preset capture radius is less than the distance error range; Angle uniformity condition: The distribution angle uniformity of the UAVs in the sub-team around the target object is less than the angle error range.

5. The method for autonomous and collaborative counter-drone swarming of drones according to claim 1, characterized in that, Before inputting the target collaborative control strategy into the trained large language model for knowledge augmentation processing to obtain the action control command output by the large language model, the method further includes: Obtain a dataset of state observations of the interactions between each drone in the drone cluster and the target object; The state observation dataset is input into the trained large language model to generate a corresponding binary action mask set as a label, wherein the binary action mask set is used to identify actions that are invalid or have a collision risk under the current observation. Based on the state observation dataset and the binary action mask set, a pre-defined lightweight action mask generation network is trained using supervised learning to obtain a trained action mask generation network. The supervised learning uses a binary cross-entropy loss function to optimize network parameters. The trained action mask generation network is used to replace the large language model to generate action masks, and the action masks are applied to the output of the large language model.

6. A system for autonomous and collaborative counter-drone swarming of unmanned aerial vehicles (UAVs), characterized in that, include: The acquisition module is used to acquire environmental information, which includes information on the drone cluster, the target object to be countered, and obstacles. The task allocation module is used to input the environmental information into the neural network model to allocate tasks and obtain the task allocation result output by the neural network model. The task allocation result is used to specify the target object corresponding to each drone in the drone cluster. The coordination and control module is used to input the task allocation result into the multi-agent reinforcement learning model to learn the cooperative control strategy, and obtain the target cooperative control strategy output by the multi-agent reinforcement learning model. The knowledge enhancement module is used to input the target cooperative control strategy into the trained large language model for knowledge enhancement processing to obtain the action control instructions output by the large language model. The knowledge enhancement processing filters invalid actions or collision actions through action masking. The control module is used to convert the motion control commands into UAV flight control commands, and control each UAV in the UAV cluster based on the UAV flight control commands to counter the target object; The neural network model adopts a dual-tower Transformer structure. The step of inputting the environmental information into a neural network model for task allocation and obtaining the task allocation result output by the neural network model includes: The positions of each UAV in the UAV cluster and the position of the target object in the environmental information are respectively embedded to obtain the UAV embedding matrix and the target object embedding matrix: The assigned logical value matrix is ​​determined by scaling the dot product cross attention based on the UAV embedding matrix and the target object embedding matrix: ; in, Represents the allocation of logical values ​​matrix. Represents the UAV embedding matrix. Represents the embedding matrix of the target object. These represent the number of drones and the number of target objects, respectively. Indicates transpose. Represents the vector dimension. Represents the space of real numbers; The allocation logic value matrix is ​​normalized to obtain the allocation probability matrix; ; in, Represents the distribution probability matrix. Represents the normalization function; Based on the allocation probability matrix, the task allocation result is determined: ; in, Indicates the first Task allocation results for each drone This indicates taking the parameter corresponding to the maximum value. Indicates the first The drone was assigned to the first The probability value of each target object.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for autonomous and cooperative counter-drone swarming as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for autonomous and cooperative counter-drone swarming as described in any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for autonomous and cooperative counter-drone swarming as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-subject pursuit optimal strategy method based on threat degree reinforcement learning algorithm

    CN118276438A