Multi-rotor unmanned aerial vehicle cooperative anti-unmanned aerial vehicle method and system

Through the multi-agent reinforcement learning training of collaborative anti-noise strategy, the advantages of multi-machine collaboration are utilized to solve the problem of insufficient application scope and flexibility of existing counter-drone technology, and efficient multi-objective anti-noise task and regional coverage are achieved.

CN120194566APending Publication Date: 2025-06-24POLIXIR TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204032.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing counter-drone technology relies on manual control or autonomous navigation, and its actual application scope and flexibility are still limited, making it difficult to effectively deal with complex drone management problems.

Method used

Multi-agent reinforcement learning training collaborative anti-noise strategy is adopted to improve the anti-task success rate through the advantages of multi-machine collaboration in situational awareness and decision-making efficiency. The specific solution includes setting up multiple types of low-altitude targets and rotor UAVs in a simulation environment, using neural network structures as a strategy model, training the collaborative anti-strategy-free model through multi-agent reinforcement learning and human-machine collaborative training framework, and deploying it on an edge computing board.

Benefits of technology

Multiple drones have achieved autonomous completion of multi-target anti-no tasks, improved anti-ineffectiveness and regional coverage, and solved the problem of insufficient flexibility and adaptability of traditional methods in dealing with complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120194566A_ABST
    Figure CN120194566A_ABST
Patent Text Reader

Abstract

The invention provides a multi-rotor unmanned aerial vehicle cooperative anti-unmanned aerial vehicle method and system, and the method comprises the following steps: S1, setting various types of low-altitude targets and the motion strategies of the targets in a simulation environment, and setting a plurality of rotor unmanned aerial vehicles; s2, setting a reverse non-task function and a corresponding reward function in a simulation environment; s3, selecting a neural network structure supporting input of any number of teammates as a strategy model; s4, randomly generating a batch of training tasks, and training a strategy model in each task by using a multi-agent reinforcement learning and man-machine cooperative training framework to obtain a cooperative anti-strategy-free model; and S5, migrating the trained collaborative anti-policy-free model, and deploying the collaborative anti-policy-free model to an edge computing board card. According to the method, countering can be completed on a multi-machine black flight target, or during single-machine black flight, the countering efficiency is improved, and the area coverage rate is increased. Meanwhile, the problem that a traditional planning method or an artificial rule cannot deal with tasks in time when executing the tasks can be solved, and the adaptability and the success rate of different types of anti-unmanned-aerial-vehicle-free cooperative tasks are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to a method and system for multi-aircraft collaborative anti-unmanned aerial vehicle of rotor unmanned aerial vehicles. Background Art

[0002] The rapid development of unmanned aerial vehicle technology in recent years has enabled it to be widely used in fields such as aerial photography, logistics, surveying and mapping, and agricultural plant protection, promoting the booming rise of the low-altitude economy. However, the problem of "black flying" (unauthorized flight) has intensified accordingly, not only interfering with aviation order, but also posing potential risks to safety and privacy protection. To address this challenge, the research and development of technical means to counter unmanned aerial vehicles has become an important topic.

[0003] Countermeasure technologies are generally divided into two links: detection and intervention. After target positioning through wireless signals, radar, optical or acoustic means, interception can be implemented by means such as radio interference, laser strikes, or using anti-unmanned aerial vehicles. Among them, radio and laser technologies are likely to have an impact on legitimate unmanned aerial vehicles or the environment, while anti-unmanned aerial vehicles have become a more efficient option due to their flexibility and can accurately respond to highly autonomous targets through means such as tracking and interference. Currently, the operation of anti-unmanned aerial vehicles relies on manual control or autonomous navigation technology, but the actual application scope and flexibility still have limitations and need to be further improved in the future to better address the increasingly complex unmanned aerial vehicle management problems.

[0004] This application proposes a collaborative anti-unmanned aerial vehicle defense method and system using multiple aircraft, trains collaborative anti-unmanned aerial vehicle strategies based on multi-agent reinforcement learning, and utilizes the advantages of multi-aircraft collaboration in situation awareness and decision-making efficiency to improve the success rate of anti-unmanned aerial vehicle tasks. Summary of the Invention

[0005] To solve the above problems, the present invention discloses a method and system for multi-aircraft collaborative anti-unmanned aerial vehicle of rotor unmanned aerial vehicles.

[0006] The specific solutions are as follows: A method for multi-aircraft collaborative anti-unmanned aerial vehicle of rotor unmanned aerial vehicles, characterized by comprising the following steps: S1. Set multiple types of low-altitude targets and the movement strategies of the targets in a simulation environment, and set multiple rotor unmanned aerial vehicles; S2. Set anti-unmanned aerial vehicle tasks and corresponding reward functions in the simulation environment; S3. Select a neural network structure that supports input of any number of teammates as a policy model; S4. Randomly generate a batch of training tasks, use a multi-agent reinforcement learning and human-machine collaborative training framework, and train the policy model in each task to obtain a collaborative anti-unmanned aerial vehicle policy model; S5. Transfer the trained collaborative anti-unmanned aerial vehicle policy model and deploy it on an edge computing board.

[0007] Furthermore, the rotary-wing UAV has a high maximum operating speed, and the actions of each UAV can be autonomously controlled; during the anti-drone mission process, humans can intervene at any time, including aborting the current mission, selecting a new mission for execution, and confirming whether to execute critical missions.

[0008] Furthermore, the training process of the collaborative anti-drone strategy model specifically includes the following steps: (1) Simulate the area information where the initial anti-drone targets appear and provide it to the simulated ground station or humans, and issue the initial target area information, which is initially perceived through various communication means; at the same time, give the maximum target UAV and the maximum number of anti-drone UAVs that can be called. (2) Set corresponding task completion rewards according to different anti-drone missions. (21) Use the Population Invariant Network (PIN) to solve the problem of changes in input and output dimensions in multi-UAV agent transfer learning, which consists of a policy network and a value evaluation network of a single agent; use the attention mechanism to process different inputs; the policy network can receive information and observations from nearby teammates to synchronize task information. (22) Train the collaborative anti-drone strategy model that supports dynamic collaborative tasks, use population-based training and evolutionary algorithms to generate diverse and incompatible teammates, introduce the Jensen-Shannon Divergence (JSD) as a diversity metric to ensure differences between teammates; at the same time, by reducing the compatibility between teammates and the learnable agent's strategy, ensure the uniqueness of the generated teammates, where high compatibility means high cooperation benefits between the learnable agent and the teammate, and low compatibility means low cooperation benefits; the optimization objectives include maximizing the team reward, increasing diversity, that is, making the JSD between the new learnable agent and the agents in the population larger, and reducing the compatibility with the learnable agent, that is, making the cooperation benefits between the new learnable agent and the existing agents in the population lower, and adjust the corresponding hyperparameters to balance the three. Usually, the weight of maximizing the team reward should be greater than the sum of the weights of the two objectives of diversity and compatibility. For example, set the weights for maximizing the team reward, increasing diversity, and reducing compatibility to 2:1:1 respectively. (23) During the training process, randomly reassign the current mission objective to simulate the switching of mission objective phases, where the mission objectives include patrol and reconnaissance, warning, and attack. (3) Transfer the trained collaborative anti-drone strategy model, perform domain adaptation transfer using a small amount of flight data in the real environment, and deploy it to the board of the anti-drone UAV.

[0009] Furthermore, in the step (2), the anti-drone missions and the corresponding task completion rewards include: a) When encountering false targets, if these targets are correctly identified as non - dangerous and the aircraft can quickly return, it is regarded as the task completed, and a completion reward of 100 is given upon completion. b) When encountering potentially threatening targets, take actions to persuade them to return. If the threat level of the target UAV does not decrease, consider taking further measures and output the measures in advance for manual confirmation. If it times out, it will be automatically confirmed. If the task is successfully completed according to the process, a completion reward of 100 is given. c) When it is confirmed that there are threatening targets, if the threat level of the target UAV does not decrease and the warning is ineffective, consider taking further measures and output the measures in advance for manual confirmation. If it times out, it will be automatically confirmed. If the task is successfully completed according to the process, a completion reward of 100 is given. d) To encourage the task to be completed as soon as possible, a penalty inversely proportional to the completion time is set. This penalty is cumulatively controlled within the scale of - 10. If the task is not completed according to the process or times out, an additional penalty of - 100 is uniformly given.

[0010] A multi - rotor UAV multi - aircraft collaborative anti - UAV system, characterized by comprising: A simulation environment construction module, used to construct a simulation environment for collaborative anti - UAV strategy training, and set various types of low - altitude targets, the movement strategies of the targets, and multiple rotor UAVs. A collaborative anti - UAV strategy model training module, used to set anti - UAV tasks and corresponding reward functions in the simulation environment; select a neural network structure that supports input of any number of teammates as the strategy model; randomly generate a batch of training tasks, and use a multi - agent reinforcement learning and human - machine collaborative training framework to train the strategy model in each task to obtain a collaborative anti - UAV strategy model. A migration and deployment module, used to migrate the trained collaborative anti - UAV strategy model and deploy it on an edge computing board.

[0011] The beneficial effects of the present invention are as follows: 1. The present invention uses multiple high - speed UAVs to collaboratively complete anti - UAV tasks. Through the training method and system proposed by the present invention, multiple UAVs can autonomously complete multi - target anti - UAV tasks, counter multi - aircraft black - flying targets, or when there is a single - aircraft black - flying situation, improve the anti - UAV efficiency and expand the regional coverage rate, etc. At the same time, by using the methods of multi - agent reinforcement learning and human - machine collaboration, it can solve the difficulties of traditional planning methods or manual rules in coping with emergencies during task execution, and improve the adaptability and success rate of collaborative tasks with different types of anti - UAV UAVs.

[0012] 2. The present invention uses a population-based training and evolutionary algorithm, combined with diversity metrics, to enhance the diversity of the population during training, enabling our drones to cooperate with the anti-drone strategies of various other drones of our side. The "population-based training and evolutionary algorithm for generating diverse and incompatible teammates" adopted during the training process of the present invention can significantly improve the success rate of multi-aircraft in collaborative tasks, especially in the process of human-unmanned collaboration. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flowchart of the present invention.

[0014] Figure 2 It is a flowchart of training the collaborative anti-drone strategy in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0015] The following further clarifies the present invention in conjunction with the specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0016] The present invention proposes a method and system for using multiple high-speed rotor drones to cooperate in countering target drones. The rotor drones therein need to have a relatively high maximum operating speed. For example, the maximum speed can reach 200 km / h. The actions of each drone in the system can be autonomously controlled. During the anti-drone mission process, a person can intervene at any time, including aborting the current mission, selecting a new mission to execute after issuance, or confirming whether to execute certain key missions. The present invention needs to complete the strategy training in a simulation environment and finally migrate to the real machine. As Figure 1 shown, the method steps of the present invention are as follows: 100. Set multiple types of low-altitude targets and the strategies of the targets, and set multiple drones in the simulation environment; 200. Set the anti-drone mission and the corresponding reward function in the simulation environment; 210. Select a neural network structure that supports the input of any number of teammates as the strategy model; 220. Randomly generate a batch of training tasks, and use a multi-agent reinforcement learning and human-machine collaborative training framework to train the strategy model in each task to obtain a collaborative anti-drone strategy; 300. Migrate the trained model and deploy it on the edge computing board.

[0017] In this embodiment, the present invention needs to complete basic strategy training with the help of a high-fidelity simulation environment. In the simulation environment, various anti-drone targets are designed in advance according to the possible low-altitude targets encountered in anti-drone tasks, such as balloons, birds, and various low, slow, and small unmanned aerial vehicle targets, and corresponding motion strategies are designed for them to simulate various situations that may be encountered in the actual anti-drone process. During the training process of collaborative anti-drone, human guidance or confirmation is required. At the same time, in each collaborative anti-drone task, the number of targets faced and the number of teammates who can cooperate will also change. It is necessary to adopt a training paradigm that supports human-computer interaction and a neural network strategy model that can handle quantity changes. At the same time, the maximum available number of anti-drone drones and the maximum possible number of target drones are specified. After initially setting the possible low-altitude targets and corresponding strategies in the anti-drone task, specifically, as Figure 2 shown, the collaborative anti-drone strategy can be trained according to the following process: A100. Simulate the area information where the initial anti-drone targets appear and provide it to the simulated ground station or person, and issue the initial target area information. This information can be initially sensed through various communication means such as radar and wireless and input into the collaborative anti-drone system of the present invention. The maximum number of target drones that can appear simultaneously in the system is 8, and the maximum number of anti-drone drones that can be called is 8; A200. Set corresponding task completion rewards according to different anti-drone tasks. Specifically, the following methods can be adopted: a) When encountering false targets, such as birds, non-threatening balloons, and drones that comply with flight regulations, if these targets are correctly identified as non-dangerous and can return quickly, it is regarded as the task completed, and a large completion reward of 100 is given when completed; b) When encountering potentially threatening targets, take actions such as surrounding or warning to persuade them to return. If the threat level of the target drone does not decrease, further measures such as ramming are considered, and the measures are output in advance for manual confirmation. If the time limit is exceeded, it is automatically confirmed. If the task is successfully completed according to the process, a large completion reward of 100 is given; c) When it is confirmed that a threatening target is encountered, if the threat level of the target drone does not decrease and the warning is ineffective, further measures such as ramming are considered, and the measures are output in advance for manual confirmation. If the time limit is exceeded, it is automatically confirmed. If the task is successfully completed according to the process, a large completion reward of 100 is given; d) At the same time, in order to encourage the task to be completed as soon as possible, a penalty inversely proportional to the completion time is set. This penalty is accumulated and controlled within the scale of -10. If the task is not completed according to the process or is completed overtime, an additional penalty of -100 is uniformly given.

[0018] A210. Use Population Invariant Networks (PIN) to solve the problem of input and output dimension changes in multi-agent transfer learning. It can be composed of an individual's policy network and value evaluation network, and uses an attention mechanism to process different inputs. The policy network can receive information and observations from nearby teammates to synchronize task information; A220. Train an anti-policy that supports dynamic collaborative tasks. Use population-based training and evolutionary algorithms to generate diverse and incompatible teammates, and introduce Jensen-Shannon divergence (JSD) as a diversity metric to ensure differences among teammates. At the same time, by reducing the compatibility between teammates and the learnable agent's strategy, ensure the uniqueness of the generated teammates. Here, high compatibility means that the learnable agent has a high return when cooperating with teammates, and low compatibility means a low return. The optimization objectives include maximizing rewards, increasing diversity, and reducing compatibility with the learnable agent, and adjust the corresponding hyperparameters to balance the three; A230. During the training process, randomly reassign the current task objective to simulate the switching of task objective phases, where the task objectives include patrol and reconnaissance, warning, and attack; A300. Transfer the model obtained from the above training, use a small amount of flight data in the real environment for domain adaptation transfer, and deploy it to the anti-drone board.

[0019] The collaborative anti-policy obtained through the above training process can be transferred and deployed to the anti-drone board, enabling each anti-drone to cooperate with multiple anti-drones within a certain range to complete anti-drone target confirmation, warning, and strike tasks. Anti-drones can share observations with each other, accept tasks provided by humans through the video transmission control interface or the ground station, and are set to give priority to tasks issued by the operator in the system.

[0020] The present invention uses population-based training and evolutionary algorithms, combined with diversity metrics, to enhance the diversity of the population during training, so that our drones can cooperate with the anti-drone strategies of various other drones of ours. The "population-based training and evolutionary algorithms to generate diverse and incompatible teammates" adopted in the training process of the present invention can significantly improve the success rate of multi-aircraft in collaborative tasks, especially in the process of human-unmanned collaboration.

[0021] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for multi-rotor UAV cooperative anti-UAV, characterized in that: The following steps are involved: S1. Set up various types of low-altitude targets and target movement strategies in the simulation environment, and set up multiple rotor drones; S2, setting the anti-nothing task and the corresponding reward function in the simulation environment; S3. Select a neural network structure that supports input of any number of teammates as the strategy model; S4. Randomly generate a batch of training tasks, use multi-agent reinforcement learning and human-machine collaborative training framework, train the policy model in each task, and obtain a collaborative anti-no-policy model; S5. Migrate the trained collaborative anti-strategy model and deploy it to the edge computing board.

2. The method for multi-rotor UAV cooperative anti-UAV according to claim 1 is characterized in that: The rotor UAV has a high maximum operating speed, and the action of each UAV can be controlled autonomously; in the anti-task process, people can intervene at any time, including terminating the current task, selecting a new task to be executed, and confirming whether to execute a key task.

3. The method for multi-rotor UAV cooperative anti-UAV according to claim 1 is characterized in that: The training process of the collaborative anti-off-policy model specifically includes the following steps: (1) Simulate the initial area information where the anti-unmanned aerial vehicle appears and provide it to the simulated ground station or person, which sends the initial target area information. This information is initially perceived through multiple communication means; at the same time, the maximum target drone and the maximum number of anti-unmanned aerial vehicles that can be called are given; (2) Set corresponding task completion rewards according to different anti-nothing tasks; (21) The population invariant network PIN is used to solve the problem of input and output dimension changes in multi-UAV agent transfer learning. It consists of a single agent's policy network and a value evaluation network. The attention mechanism is used to process different inputs. The policy network can receive information and observations from nearby teammates to synchronize task information. (22) Train a collaborative anti-strategy model that supports dynamic collaborative tasks. Use population-based training and evolutionary algorithms to generate diverse and incompatible teammates. Introduce JSD divergence as a diversity metric to ensure differences between teammates. At the same time, the uniqueness of generated teammates is ensured by reducing the compatibility of teammates with the strategies of the learnable agents. High compatibility means that the benefits of cooperation between the learnable agents and teammates are high, and low compatibility means that the benefits of cooperation are low. The optimization objectives include maximizing team returns, increasing diversity, even if the JSD of new learnable agents with agents in the population becomes larger, and reducing compatibility with learnable agents, even if the benefits of cooperation between new learnable agents and existing agents in the population become lower. Adjust the corresponding hyperparameters to balance the three. (23) During the training process, the current mission objective is randomly reassigned to simulate the switching of mission objective stages, where the mission objectives include patrol reconnaissance, warning, and attack; (3) Migrate the trained collaborative anti-drone strategy model, use a small amount of flight data in the real environment to perform domain adaptive migration, and deploy it on the anti-drone board.

4. The method for multi-rotor UAV cooperative anti-UAV according to claim 3 is characterized in that: In step (2), the anti-no task and the corresponding task completion reward include: a) When encountering false targets, if you correctly identify these targets as non-dangerous and can return quickly, the mission is considered completed and a completion reward of 100 will be given upon completion; b) When encountering a potential threat target, take action to persuade it to return. If the threat level of the target drone has not been reduced, consider taking further measures and output the measures in advance for manual confirmation. If it exceeds the time limit, it will be automatically confirmed. If the task is successfully completed according to the process, a completion reward of 100 will be given; c) When a threat target is confirmed, if the threat level of the target drone has not been reduced and the warning is ineffective, further measures will be considered and output in advance for manual confirmation. If the timeout is exceeded, it will be automatically confirmed. If the task is successfully completed according to the process, a completion reward of 100 will be given; d) In order to encourage the task to be completed as soon as possible, a penalty inversely proportional to the completion time is set. The cumulative penalty is controlled at a scale of -10. If the task is not completed according to the process or is completed overtime, an additional penalty of -100 will be given.

5. A multi-rotor UAV coordinated anti-UAV system, characterized in that: include: The simulation environment construction module is used to build a simulation environment for collaborative anti-strategy training and set up various types of low-altitude targets, target movement strategies, and multiple rotor drones; The collaborative anti-off-policy model training module is used to set anti-off-policy tasks and corresponding reward functions in the simulation environment; and a neural network structure that supports the input of any number of teammates is selected as the policy model; a batch of training tasks are randomly generated, and the multi-agent reinforcement learning and human-computer collaborative training framework are used to train the policy model in each task to obtain the collaborative anti-off-policy model; The migration and deployment module is used to migrate the trained collaborative anti-strategy-free model and deploy it to the edge computing board.