Rotor unmanned aerial vehicle autonomous anti-unmanned aerial vehicle method and system
Through the autonomous anti-UAV method based on reinforcement learning, the neural network strategy model and reinforcement learning algorithm are used to train anti-unless strategies, the problem of difficulty in effectively countering high-autonomous drones in the existing technology is solved, and efficient and economical autonomous countermeasures are achieved.
Patent Information
- Application Number
- CN202510217205.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively counter highly autonomous target drones with visual perception, especially in close-range defense and multi-environmental scenarios. Traditional radio interference and laser interception methods have limited effects and are costly, making it difficult to meet the needs of real-time adjustment and high-speed control.
The autonomous anti-UAV method based on reinforcement learning is adopted to provide the initial positioning of the target drone through radar or wireless signals, and quickly dispatch high-speed rotor drone for confirmation, warning and physical interception. The neural network strategy model and reinforcement learning algorithm are used to train anti-unless strategies to achieve autonomous countermeasures.
It improves the success rate and practicality of counter-drones, reduces dependence on high-precision positioning and navigation systems, reduces equipment costs, enhances the counter-response capabilities of fully autonomous medium and low-speed drones, and can cope with closer target drones, and is suitable for more production environments.
Smart Images

Figure CN120103869A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of using a rotary-wing UAV for autonomous anti-UAV operations, and in particular to a rotary-wing UAV autonomous anti-UAV method and system. Background Art
[0002] In recent years, with the rapid development of drone technology, drones have been widely used in aerial photography, logistics, surveying and mapping, agricultural plant protection and other fields, becoming an important force in promoting the development of the low-altitude economy. The rise of the low-altitude economy has injected new vitality into industrial development and economic growth. However, along with this, the phenomenon of "illegal flying" (flying without permission) of drones has increased, posing a severe challenge to aviation safety, public privacy and the protection of critical infrastructure. Many drone users lack knowledge of flight regulations or actively evade supervision, resulting in frequent illegal flights, which not only disrupts normal aviation order, but may also threaten civil aviation flight safety or bring safety hazards to key production departments. Therefore, it is necessary to develop technical means to counter illegal drones, namely "anti-drones".
[0003] At present, anti-drone means are mainly divided into two stages: detection and interception. In the detection stage, wireless signal monitoring, radar detection, optical and acoustic detection and other means are usually used to detect the approximate location information of the target, and then interception measures are taken, such as radio interference interception, laser interception, drone interception and other means. Among them, radio interference requires the speculation of the channel of the target drone, and will interfere with drones flying legally on the same channel. Laser interception means generally consumes a lot of energy, has high costs, and has certain requirements for the environment. The defense effect at close range is usually poor, and radio interference and lasers cannot directly counter highly autonomous target drones with visual perception. UAVs are needed for anti-drone, that is, special anti-drone drones are dispatched to track, collide or interfere with target drones.
[0004] As the last line of defense against drones, drones are highly flexible and can be used in complex environments. Usually, drones are manually controlled by professional pilots, or use high-precision navigation and positioning systems and very precise control technology to achieve autonomous anti-drone, which limits their use scenarios. At present, autonomous flight control methods are difficult to adjust plans in real time and difficult to cope with high-speed control and other issues. Summary of the invention
[0005] To solve the above problems, the present invention discloses a method and system for autonomous anti-drone of a rotorcraft drone in the technical field of physical collision countermeasure against a target drone based on reinforcement learning and training under the human-computer interaction paradigm. Based on the initial positioning of the target drone provided by radar or wireless signals (for example, 5G-A signals), a high-speed drone is quickly dispatched for confirmation, warning and physical interception, wherein the strategy used for anti-drone drones is obtained based on reinforcement learning training. The method proposed in this application can improve the success rate of anti-drone drones, thereby improving the practicality and reliability of using anti-drone drones.
[0006] A method for using a high-speed rotor drone to collide and counter a target drone comprises the following steps: Step 100: construct a variety of low-altitude targets and countermeasure tasks in a simulation environment; Step 210: construct a neural network strategy model and corresponding input, output and reward functions; Step 220: Use a reinforcement learning algorithm to perform anti-strategy training for the drone in a simulation environment; Step 300: Perform real-machine migration deployment on the trained strategy.
[0007] Furthermore, in step 100, for the various low-altitude targets, the countermeasure tasks include warning, following, and impacting; wherein the countermeasure task is randomly initialized each time.
[0008] Furthermore, the step 210 defines the situation input and the task target input according to the fusion situation data of the drone used, defines the control output and the task currently performed by the drone according to the control command interface, namely, warning, tracking, collision, etc., and designs the task reward function, and reserves the corresponding status display interface in the drone's image transmission interface.
[0009] Furthermore, the reward function of the corresponding task, specifically, a: When encountering a false target, you need to identify it as a false target or a non-threatening target and return home quickly. If the mission is successfully completed, you will be given a larger completion reward of 100, as well as a reward inversely proportional to the completion time. The cumulative result of this reward does not exceed 10; b: When encountering a potential threat target, take circling or warning actions. If the threat level of the target drone has not decreased, take further measures such as collision and output the measures in advance, waiting for manual confirmation. If it exceeds the time limit, it will be automatically confirmed. If the task is successfully completed according to the process, a completion reward of 100 will be given, as well as a reward inversely proportional to the completion time. The cumulative result of this reward shall not exceed 10; c: When it is confirmed that a threatening target is encountered, if the threat level of the target drone has not decreased and the warning is ineffective, take further measures such as collision and output the measures in advance, waiting for manual confirmation. If it exceeds the time limit, it will be automatically confirmed. If the task is successfully completed according to the process, a completion reward of 100 will be given, as well as a reward inversely proportional to the completion time. The cumulative result of this reward shall not exceed 10; If the task is not completed or exceeds the time limit, a uniform score of -100 will be given.
[0010] Furthermore, when the step 300 strategy is executed, the control amount is first executed according to the approximate target area information given by the ground station or people, and the input mission objectives, according to the trained strategy, and the mission currently executed by the UAV is output, and the UAV information and mission information are displayed on the image transmission interface. The controller can select a new task from the task list and send it to the UAV or confirm the execution of the current task. Among them, the collision task forces a confirmation prompt to pop up, and sets a default confirmation time. If the confirmation timeout is exceeded and the collision is not manually terminated, the collision task will continue to be executed.
[0011] A system for using a rotary wing UAV to collide and counter a target UAV, comprising a low-altitude target task construction module: constructing a variety of low-altitude targets and counter-tasks in a simulation environment; Neural network strategy module: builds the neural network strategy model and the corresponding input, output and reward functions; Drone anti-strategy training module: Use reinforcement learning algorithm to train drone anti-strategy training in a simulation environment; Real-machine migration and deployment module: performs real-machine migration and deployment of trained strategies.
[0012] Beneficial effects of the present invention: The anti-drone method of the present invention is based on preliminary target location intelligence information, and dispatches high-speed rotor drones to perform anti-drone tasks. During the execution of the task, the task can be switched at any time according to human guidance. The anti-drone execution process designed in the present invention can identify false reports and return, and issue an alarm when it is confirmed to be a threatening target drone. Finally, after human confirmation, anti-drone collision is performed. The anti-drone method of the present invention can counter fully autonomous medium and low-speed drones, solve the problem of failure of radar and electronic interference in the face of fully autonomous target drones, use the autonomous anti-drone strategy obtained through reinforcement learning training, reduce the dependence on high-precision positioning and navigation systems, reduce equipment costs, greatly improve the availability of drone anti-drone, and be able to deal with target drones at closer distances, and be deployed in more production environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 , flow chart of the present invention. DETAILED DESCRIPTION
[0014] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. It should be noted that the words "front", "rear", "left", "right", "upper" and "lower" used in the following description refer to directions in the accompanying drawings, and the words "inner" and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.
[0015] This embodiment proposes a method and system for using a high-speed rotor drone to collide and counter a target drone. The drone's actions can be fully autonomously controlled, and it gives priority to executing tasks given by humans or executing key actions such as confirming warnings and colliding with target drones. The rotor drone in the present invention needs to have a high maximum operating speed, for example, a maximum speed of 160-200km / h, and be equipped with at least a binocular camera to complete vision-based recognition and detection. In addition, the present invention needs to perform anti-no-strategy training in a simulation environment with high fidelity, and the simulation needs to include a variety of targets that may be encountered in the low altitude, including false targets such as birds, kites, balloons, and various types of drones. At the same time, according to the drone's situation and control interface and mission objectives, a reinforcement learning algorithm is used to train the anti-no-strategy in a simulation environment, and then the trained strategy is migrated. As shown in Figure 1, the operating steps and units of the present invention are as follows: 100. Construct a variety of low-altitude targets and countermeasure tasks in the simulation environment. The countermeasure tasks include warning, following, and collision. The countermeasure tasks are randomly initialized each time. 210. Construct a neural network strategy model, define the situation input and task target input according to the UAV fusion situation data used, define the control output and the task currently performed by the UAV according to the control command interface, i.e. warning, tracking, collision, etc., and design the task reward function, and reserve the corresponding status display interface in the UAV's image transmission interface; 220. Using reinforcement learning algorithms to train drones to counteract no strategy in a simulation environment; 300. Perform real-machine migration of trained strategies.
[0016] When the strategy is executed, it first executes the control volume according to the trained strategy based on the approximate target area information given by the ground station or people, as well as the input mission objectives. At the same time, it outputs the mission currently executed by the drone and displays the drone information and mission information on the image transmission interface. The controller can select a new task from the task list and send it to the drone or confirm the execution of the current task. Among them, the collision task is forced to pop up a confirmation prompt and set a default confirmation time. If the confirmation timeout is exceeded and the collision is not manually terminated, the collision task will continue to be executed.
[0017] In order to explain the technical solutions related to the present invention in more detail, the present invention will be further described below in conjunction with specific implementation methods and drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Example
[0018] The drone in the present invention needs to take preliminary target drone location information as input, which can be provided by radar or wireless equipment after sensing. After obtaining the preliminary location, the drone is quickly dispatched to confirm the category and location of the target. If a false alarm is encountered, it will return, otherwise it will issue an alarm or collision. This process requires human-machine cooperation. People select high-level tasks and drones execute them. Alternatively, when the drone performs key actions, it will give people prompts for human confirmation before execution. People can interrupt the task process at any time. This process is completed in the image transmission interface. The anti-unmanned strategy of the drone is trained in a simulation environment. Specifically, it can be implemented according to the following process: A100, set up a variety of low-altitude targets in the simulation environment, and set the target flight strategy or track. The low-altitude targets include birds, balloons, kites, and various drones. Simulate several different types of threat tasks, such as false alarm missions, drone countermeasures against non-threatening targets, drone countermeasures against potential threatening targets, and drone countermeasures against threatening targets. Use one-hot encoding to encode different types of tasks. A200, choose deep neural network as the policy model, usually use a network composed of fully connected and residual modules, and combine LSTM or GRU memory units to construct the basic policy network; A210. According to the situation and control interface of the anti-UAV to be trained, the input of the strategy model is set to the situation and the corresponding task code, where the original task code needs to be embedded in the embedding layer to obtain low-dimensional continuous features, and the output control amount is set for the strategy model; A220. Set a training task, i.e., a simulation reset function, which can randomly generate a type of target (false target, or randomly select a target with different threat levels), set a corresponding flight strategy for the target, and set a reward function for the corresponding task. Specifically, when encountering a false target, it is necessary to identify it as a false target or a non-threatening target and return home quickly. If the task is successfully completed, a large completion reward of 100 will be given, as well as a reward inversely proportional to the completion time. The cumulative result of this reward generally does not exceed 10. When encountering a potential threat target, actions such as circling or warning can be taken. If the threat level of the target drone does not decrease, further measures such as collision will be taken. If the task is successfully completed according to the process, a large completion reward of 100 will be given, as well as a reward inversely proportional to the completion time. The cumulative result of this reward generally does not exceed 10. When it is confirmed that there is a threatening target, if the threat level of the target drone has not been reduced and the warning is invalid, further measures such as collision will be taken, and measures will be output in advance, waiting for manual confirmation. If the timeout is exceeded, it will be automatically confirmed. If the task is successfully completed according to the process, a large completion reward of 100 will be given, as well as a reward inversely proportional to the completion time. The cumulative result of this reward generally does not exceed 10. If the task is not completed or exceeds the time, a uniform score of -100 will be given. A230. Use a reinforcement learning algorithm, such as PPO or SAC, to train the policy model. Specifically, reset the simulation environment to generate training tasks each time, and continue training until the cumulative reward function on multiple tasks no longer increases. A300, verify the trained strategy and prepare for strategy migration. It is necessary to collect a small amount of data from actual task scenarios in advance and use methods such as domain randomization and domain adaptation to align the data feature distribution of the simulation environment and the actual environment, so as to correct the inconsistency between the strategy trained in the simulation environment and the real environment. After the above process, a policy model for practical anti-nothing tasks can be obtained. Example
[0019] The strategy obtained in Example 1 can be migrated and deployed to the actual machine to directly execute the task. However, in the actual anti-task process, since recognition and perception are not completely reliable, at this stage, people are still needed to confirm the current task in the loop, or complete the task under human supervision. Specifically, when people participate, it can be implemented according to the following process: B100, send the received preliminary target area information to the anti-unmanned aerial vehicle. After receiving the mission signal, the anti-unmanned aerial vehicle calls the strategy model to start the mission, quickly flies to the target area for patrol and reconnaissance, and displays the current status and mission execution status on the image transmission interface, and updates in real time; B200, in patrol reconnaissance, anti-drone uses pre-trained visual detection models to identify target objects, confirm target types, and decide to perform return or further countermeasure tasks (warning, and collision after warning) according to the loaded strategy model, and update the task status in real time on the image transmission interface; B210, the image transmission interface provides a task list, which can perform tasks such as tracking, warning, and collision on the discovered targets. When directly assigning tasks, a second confirmation is required on the image transmission interface; B300, when the warning is invalid and the collision mission is confirmed, or when the collision mission is directly specified by a person, a prompt box for confirming the collision mission will pop up on the image transmission interface. The person can choose to execute the mission or terminate the mission, and a countdown will be prompted. The countdown is usually set to 10-15 seconds. If there is no confirmation within the timeout, the default behavior (execution of the collision mission) will be taken. The image transmission interface has an additional abort button, which can abort the mission at any time; B400, Anti-UAV, when the mission target is lost, disappears, or leaves the danger zone, the collision mission is terminated, and when it is confirmed that the target is no longer within the range, it returns to the site (take-off point).
[0020] When the rotor UAV in the present invention performs an impact mission, further confirmation by a human is required to reduce the risk of misjudgment.
[0021] This embodiment is based on reinforcement learning and mission goal-oriented anti-UAV strategy training, multiple low-altitude target constructions and multiple anti-UAV task settings, and adopts anti-UAV strategy execution process and human-computer interaction logic design. , reducing the dependence on high-precision positioning and navigation systems, reducing equipment costs, greatly improving the availability of UAV anti-UAV, and being able to deal with target UAVs at closer distances, and deployed in more production environments.
[0022] The technical means disclosed in the scheme of the present invention are not limited to the technical means disclosed in the above-mentioned implementation mode, but also include technical schemes composed of any combination of the above technical features.
Claims
1. A method for a rotary wing UAV to collide with a target UAV, characterized in that: The following steps are involved: Step 100: construct a variety of low-altitude targets and countermeasure tasks in a simulation environment; Step 210: construct a neural network strategy model and corresponding input, output and reward functions; Step 220: Use a reinforcement learning algorithm to perform anti-strategy training for the drone in a simulation environment; Step 300: Perform real-machine migration deployment on the trained strategy.
2. The method for a rotary-wing UAV to collide and counter a target UAV according to claim 1, characterized in that: In step 100, there are multiple low-altitude targets, and the countermeasure tasks include warning, following, and impacting; wherein the countermeasure task is randomly initialized each time.
3. The method for a rotary-wing UAV to collide and counter a target UAV according to claim 1, characterized in that: The step 210 defines the situation input and the task target input according to the fusion situation data of the drone used, defines the control output and the task currently performed by the drone according to the control instruction interface, namely, warning, tracking, and collision, and designs the task reward function, and reserves the corresponding status display interface in the drone's image transmission interface.
4. The method for a rotary-wing UAV to collide and counter a target UAV according to claim 3, characterized in that: The reward function for the corresponding task, specifically, a: When encountering a false target, you need to identify it as a false target or a non-threatening target and return home quickly. If the mission is successfully completed, you will be given a larger completion reward of 100, as well as a reward inversely proportional to the completion time. The cumulative result of this reward does not exceed 10; b: When encountering a potential threat target, take circling or warning actions. If the threat level of the target drone does not decrease, take further measures of collision and output the measures in advance for manual confirmation. If it exceeds the time limit, it will be automatically confirmed. If the task is successfully completed according to the process, a completion reward of 100 will be given, as well as a reward inversely proportional to the completion time. The cumulative result of this reward shall not exceed 10; c: When it is confirmed that there is a threatening target, if the threat level of the target drone has not been reduced and the warning is ineffective, further measures such as collision will be taken, and the measures will be output in advance for manual confirmation. If it times out, it will be automatically confirmed. If the task is successfully completed according to the process, a completion reward of 100 will be given, as well as a reward inversely proportional to the completion time. The cumulative result of this reward shall not exceed 10; if the task is not completed or times out, a uniform score of -100 will be given.
5. The method for a rotary-wing UAV to collide and counter a target UAV according to claim 1, characterized in that: When executing the step 300 strategy, first, according to the approximate target area information given by the ground station or people, and the input mission objectives, the control volume is executed according to the trained strategy, and the mission currently executed by the UAV is output, and the UAV information and mission information are displayed on the image transmission interface. The controller can select a new task from the task list and send it to the UAV or confirm the execution of the current task. Among them, the collision task is forced to pop up a confirmation prompt and set a default confirmation time. If the confirmation timeout is not manually terminated, the collision task will continue to be executed.
6. A system for countering a target drone by using a rotary-wing drone, characterized in that: Includes low-altitude target task construction module: constructs a variety of low-altitude targets and countermeasure tasks in a simulation environment; Neural network strategy module: builds the neural network strategy model and the corresponding input, output and reward functions; Drone anti-strategy training module: Use reinforcement learning algorithm to train drone anti-strategy training in a simulation environment; Real-machine migration and deployment module: performs real-machine migration and deployment of trained strategies.
Citation Information
Patent Citations
Unmanned aerial vehicle countering system test method and platform
CN111929650A
Unmanned aerial vehicle autonomous flight control method based on offline reinforcement learning
CN113110546A
Intelligent autonomous unmanned aerial vehicle countering system
CN114911267A
Multi-unmanned-aerial-vehicle-to-multi-target cooperative air combat maneuver decision-making method based on reinforcement learning
CN117032300A
Cluster collaborative target search method based on digital twinning and deep reinforcement learning
CN117930863A