Unmanned aerial vehicle swarm countering method based on deep learning

By adopting deep learning and multi-sensor data fusion methods in the drone swarm counter technology, the problems of low multi-objective tracking accuracy, poor adaptability of counter strategies, and insufficient system stability in complex environments are solved, and efficient, flexible and stable counter task execution is achieved.

CN120074734AInactive Publication Date: 2025-05-30BEIJING CENT POLICE JINSHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510205297.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex environments, the existing drone swarm countermeasure technology has problems such as low multi-target tracking accuracy, poor adaptability of countermeasure strategies, and insufficient system stability.

Method used

Using a deep learning-based method, by establishing a dynamic system model of drones and targets, using Li Qun and Li algebra to model the relative motion of drones and targets, defining optimal control problems, and optimizing counter strategies through deep reinforcement learning, combining multi-sensor data fusion and topological optimization algorithms, Li Yapunov stability theory is used to analyze system stability.

Benefits of technology

It improves the multi-objective tracking accuracy and adaptability of counter strategies in complex environments, enhances the stability and robustness of the system, and significantly improves the execution efficiency and success rate of the drone swarm counter tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074734A_ABST
    Figure CN120074734A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicles, and discloses an unmanned aerial vehicle swarm countering method based on deep learning, comprising the following steps: step 1, establishing dynamic system models of unmanned aerial vehicles and targets, the state of each unmanned aerial vehicle being represented by the position and speed vector of the unmanned aerial vehicle, and the state of each target being represented by the position and speed vector of the target; and step 2, the relative motion of the unmanned aerial vehicle and the target is modeled by adopting the Lie group and the Lie algebra, the Lie group is used for representing the rigid motion of the unmanned aerial vehicle and the target, and the transformation of the unmanned aerial vehicle and the target in a three-dimensional space is described through the combination of a rotation matrix and a translation matrix. The relative motion of the unmanned aerial vehicle and the target is modeled by adopting the Lie group and the Lie algebra, so that the problem of errors in multi-target tracking in a complex environment in the existing method is solved, the dynamic relationship between the unmanned aerial vehicle and the target is accurately captured, effective countering capability and target tracking are obtained, and the task execution efficiency in the multi-target environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles, and specifically to an anti-swarm method for unmanned aerial vehicles based on deep learning. Background Art

[0002] With the rapid development of unmanned aerial vehicle (UAV) swarm technology, its potential in the fields of military, security, and disaster response is becoming increasingly apparent. However, UAV swarms face many technical challenges when performing tasks, especially when carrying out anti-swarm tasks in complex environments. Existing technologies have limitations in target tracking, adaptability of anti-swarm strategies, and system stability.

[0003] Currently, traditional multi-target tracking methods usually rely on classical control algorithms. When dealing with large-scale and dynamically changing targets, they face problems such as low tracking accuracy and ineffective response to mutual interference between targets. Especially in complex environments, the rapid changes of targets and the coordination problem among multiple UAVs exacerbate the complexity of control strategies, resulting in low efficiency in performing anti-swarm tasks.

[0004] In addition, most existing anti-swarm strategies are based on static models and preset paths, making the system unable to cope with sudden changes or new tactics. In a rapidly changing combat environment, traditional static anti-swarm strategies are easily affected by enemy interference and changes, and cannot be effectively adjusted according to environmental changes.

[0005] In terms of the stability of UAV swarm anti-swarm tasks, traditional methods lack effective solutions to problems such as system dynamic instability and accumulation of target tracking errors, resulting in risks of unstable control and inability to complete tasks when the system performs long-term tasks.

[0006] Therefore, how to solve the problems of low multi-target tracking accuracy, poor adaptability of anti-swarm strategies, and insufficient system stability in complex environments has become the key to improving the anti-swarm ability of UAV swarms.

[0007] Therefore, those skilled in the art provide an anti-swarm method for UAV swarms based on deep learning to solve the above-mentioned problems. Summary of the Invention

[0008] In view of the deficiencies of the prior art, the present invention provides an anti-swarm method for UAV swarms based on deep learning to solve the problems raised in the above background art.

[0009] To achieve the above objectives, the present invention is realized through the following technical solutions: An anti-swarm method for UAV swarms based on deep learning, comprising:

[0010] Step 1: Establish a dynamic system model of UAVs and targets. The state of each UAV is represented by the position and velocity vectors of the UAV, and the state of each target is represented by the position and velocity vectors of the target;

[0011] Step 2: Use Lie groups and Lie algebras to model the relative motion between the UAV and the target. The Lie group is used to represent the rigid body motion of the UAV and the target, and through the combination of rotation and translation matrices, it describes the transformation of the UAV and the target in three-dimensional space. The Lie algebra is used to represent the relative motion between the UAV and the target;

[0012] Step 3: Define the optimal control problem, and optimize the countermeasure process by designing a cost function. The cost function comprehensively considers the trajectory deviation of the target and the cost of the UAV control input, and solves the optimal control problem through the optimal control theory;

[0013] Step 4: Based on the results of optimal control, introduce deep reinforcement learning to optimize the UAV's countermeasure strategy, and by training the deep Q network, enable the UAV to adaptively adjust its countermeasure strategy according to the feedback reward signal in a dynamically changing environment;

[0014] Step 5: Use multi-sensor data fusion technology, combine Kalman filtering and extended Kalman filtering to estimate the state of the target in real time, and by fusing data from different sensors;

[0015] Step 6: Through the path planning algorithm, formulate an optimal countermeasure path for the UAV. During this process, consider the position of the target, the current position of the UAV, and obstacles, and calculate a path that can quickly and effectively approach the target, avoiding collisions with obstacles during flight;

[0016] Step 7: Through the topology optimization algorithm, adjust the path of the UAV. Topology optimization considers the real-time position changes between the target and the UAV, and dynamically optimizes the path of the UAV;

[0017] Step 8: Apply Lyapunov stability theory to analyze the stability of the countermeasure system. By constructing a Lyapunov function, analyze the stability conditions of the Lyapunov function to successfully complete the task.

[0018] Preferably, the design of the cost function is as follows:

[0019]

[0020] where, x i (t) is the state of target i at time t, x i,ref (t) is the reference trajectory position of target i at time t, is the difference between the target trajectory deviation and the reference trajectory,

[0021] u j (t) is the control input of UAV j at time t,

[0022] is the control input cost of UAV j, and α is a constant coefficient. is the velocity of UAV j at time t, T is the total time for task execution, Q is a weight matrix, R is a weight matrix, m is the number of state variables, n is the number of control inputs, and dt represents integration with respect to time t.

[0023] Preferably, the application of the optimal control theory is based on a Hopfield neural network, and the Hopfield neural network is used to optimize the generation of the countermeasure path. The energy function is defined as:

[0024]

[0025] where E is the energy function, x j (t) is the state of UAV j at time t, x j,target (t) is the target position of UAV j at time t.

[0026] ||x j (t) - x j,target (t)|| 2 is the deviation between UAV j and the target position, x i (t) is the state of target i at time t, x i,ref (t) is the reference trajectory position of target i at time t.

[0027] ||x i (t) - x i,ref (t)|| 2 is the deviation between target i and the reference trajectory, m is the number of state variables, and n is the number of control inputs.

[0028] Preferably, the reward function in the deep Q network is designed as:

[0029] R t = -β 1 ||x i (t) - x i,ref (t)|| 2 - β 2 ||x j (t) - x j,target (t)|| 2 + γ·||Δu j (t)|| 2 ,

[0030] where x i (t) is the state of target i at time t, x i,ref (t) is the reference trajectory of target i at time t, ||x i (t) - x i,ref (t)||2 is the deviation between the target i and the reference trajectory,

[0031] β 1 is the target deviation weighting coefficient, x j (t) is the state of UAV j at time t, x j,target (t) is the target position of UAV j at time t,

[0032] ||x j (t)-x j,target (t)|| 2 is the deviation between UAV j and the target position,

[0033] β 2 is the UAV deviation weighting coefficient, γ is the weighting coefficient of the control input change,

[0034] Δu j (t) is the control input change of UAV j at time t, ||Δu j (t)|| 2 is the cost of the UAV control input change.

[0035] Preferably, the distributed policy gradient method is adopted in the training process of the deep Q network to enhance the reaction ability of the system in multi-UAV cooperative combat, and the gradient update formula is:

[0036]

[0037] where, θ t+1 is the policy parameter at time t+1, θ t is the current policy parameter, η is the learning rate, is the policy gradient.

[0038] Preferably, the multi-sensor data fusion technology includes fusing radar, infrared and visual sensor data to improve the accuracy of target state estimation, and the estimation formula of the sensor fusion is:

[0039]

[0040] where, represents the estimated value of the target state at time k, represents the state estimation calculated in the previous time step, K k represents the fusion ratio of the current time measurement information and the prediction information, z k represents the target state observation data obtained by the sensor at time k, H k is the observation matrix.

[0041] Preferably, the process of the Kalman filter includes using the extended Kalman filter for state estimation of a non-linear system. The state update formula of the extended Kalman filter is:

[0042]

[0043] where, represents the estimated value of the target state at time k, z k represents the observed data of the target state obtained by the sensor at time k, H k is the observation matrix, is the predicted value of the target state at time k - 1, P k represents the uncertainty of the target state estimation.

[0044] Preferably, the path planning algorithm includes using the A* algorithm to plan the optimal countermeasure path of the UAV. The path calculation process is optimized by the following formula:

[0045] f(x) = g(x) + h(x),

[0046] where, g(x) is the actual cost from the starting point to the current node x, h(x) is the estimated cost from the node x to the target, and f(x) is the total cost of the current path.

[0047] Preferably, the topology optimization algorithm includes introducing environmental constraints and dynamic adjustment factors during the path planning process. The objective function of the topology optimization is:

[0048]

[0049] where, λ is the weight coefficient for adjusting the influence of obstacles, obstacle(p j (t)) represents the influence of obstacles in the path, p j (t) represents the path of the UAV, p i (t) is the path of the target,

[0050] ||p j (t) - p i (t)|| 2 is the distance deviation between UAV j and target i at time t.

[0051] Preferably, the Lyapunov stability analysis is based on the stability theory of generalized linear systems to analyze the stability of the UAV swarm system. The Lyapunov function is:

[0052]

[0053] where, x j (t) is the UAV state, x j,ref(t) is the desired trajectory, V(x) is the Lyapunov function, ||x j (t) - x j,ref (t)|| 2 is the error between the state of UAV j at time t and the position of the reference trajectory, and n is the number of control inputs.

[0054] The present invention provides a method for countering UAV swarms based on deep learning. It has the following beneficial effects:

[0055] 1. By using Lie groups and Lie algebras to model the relative motion of UAVs and targets, the present invention solves the error problem in multi-target tracking of existing methods in complex environments, realizes accurate capture of the dynamic relationship between UAVs and targets, obtains effective countermeasure capabilities and target tracking, and improves the task execution efficiency in multi-target environments.

[0056] 2. By training a deep Q-network for deep reinforcement learning, the UAVs can adaptively adjust their countermeasure strategies according to feedback signals in a dynamically changing environment, solve the limitation that traditional static countermeasure strategies cannot cope with sudden changes, obtain flexible countermeasure capabilities, and significantly improve the adaptability and real-time performance of the system in complex scenarios.

[0057] 3. By applying the Lyapunov stability theory, the present invention analyzes and ensures the stability of the countermeasure system, solves the instability problem that occurs during the execution of the system, obtains continuous and stable task execution capabilities, and enhances the robustness and reliability of the system in complex dynamic environments. Description of the Drawings

[0058] Figure 1 is the flowchart of the present invention. Detailed Embodiments

[0059] To enable those skilled in the art to understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0060] The present invention will be described in detail below with reference to the accompanying drawings:

[0061] Embodiment:

[0062] Please refer to the attached Figure 1 , the embodiment of the present invention provides a method for countering UAV swarms based on deep learning, including:

[0063] Step 1: Establish a dynamic system model for the drones and the targets. The state of each drone is represented by the position and velocity vectors of the drone, and the state of each target is represented by the position and velocity vectors of the target;

[0064] Step 2: Use Lie groups and Lie algebras to model the relative motion of the drones and the targets. Lie groups are used to represent the rigid body motion of the drones and the targets, and through the combination of rotation and translation matrices, the transformation of the drones and the targets in three-dimensional space is described. Lie algebras are used to represent the relative motion between the drones and the targets;

[0065] Step 3: Define an optimal control problem, and optimize the countermeasure process by designing a cost function. The cost function comprehensively considers the trajectory deviation of the target and the cost of the drone control input, and the optimal control problem is solved through optimal control theory;

[0066] Step 4: Based on the results of optimal control, introduce deep reinforcement learning to optimize the countermeasure strategy of the drones, and by training a deep Q-network, enable the drones to adaptively adjust the countermeasure strategy of the drones according to the feedback reward signal in a dynamically changing environment;

[0067] Step 5: Use multi-sensor data fusion technology, combine Kalman filtering and extended Kalman filtering to estimate the state of the target in real time, and by fusing data from different sensors;

[0068] Step 6: Through a path planning algorithm, formulate an optimal countermeasure path for the drones. In this process, consider the position of the target, the current position of the drones, and the obstacles, and calculate a path that can quickly and effectively approach the target, avoiding collisions with obstacles during flight;

[0069] Step 7: Through a topology optimization algorithm, adjust the path of the drones. Topology optimization considers the real-time position changes between the target and the drones and dynamically optimizes the path of the drones;

[0070] Step 8: Apply Lyapunov stability theory to analyze the stability of the countermeasure system. By constructing a Lyapunov function, analyze the stability conditions of the Lyapunov function to successfully complete the task.

[0071] The benefit of Step 1 is to provide a clear mathematical basis for subsequent target tracking, path planning, and control decision-making, making the analysis and control of the countermeasure system accurate in a multi-target environment.

[0072] The benefit of Step 2 is to accurately represent the relative changes between the drones and the targets in three-dimensional space through the combination of rotation and translation matrices, providing strong support for accurate target tracking and dynamic control, especially in a complex battlefield environment.

[0073] The benefit-cost function in Step 3 comprehensively considers the deviation of the target trajectory and the cost of the UAV control input, effectively guiding the UAV to maintain an efficient and precise countermeasure strategy in a complex environment.

[0074] The benefit of Step 4 solves the problem that traditional static control strategies cannot cope with the dynamic changes of the environment, enhancing the adaptability, flexibility, and real-time response ability of the system in different scenarios, when facing uncertain actions of the enemy.

[0075] The benefit of Step 5 combines the Kalman filter and the extended Kalman filter to fuse the data obtained from multiple groups of sensors, significantly improving the accuracy of target state estimation. Multi-sensor fusion makes up for the limitations of a single sensor, effectively reducing the influence caused by measurement noise and occlusion, and providing reliable information support for the accurate execution of the countermeasure strategy.

[0076] The benefit of Step 6 ensures that the UAV can quickly and effectively approach the target and execute the countermeasure task, especially avoiding obstacles in a complex environment and improving the execution efficiency of the countermeasure task.

[0077] The benefit of Step 7 improves the path optimization ability of the countermeasure task, ensuring that the UAV has an efficient path when executing tasks in a changing environment, and avoiding unnecessary path conflicts and energy waste.

[0078] The benefit of Step 8 avoids the problem of system instability caused by dynamic changes or error accumulation, enhancing the reliability and robustness of the countermeasure task, and ensuring that the UAV swarm can stably execute the task until completion.

[0079] The design of the cost function is as follows:

[0080]

[0081] where, \(x\) i \((t)\) is the state of target \(i\) at time \(t\), \(x\) i,ref \((t)\) is the reference trajectory position of target \(i\) at time \(t\), is the difference between the target trajectory deviation and the reference trajectory,

[0082] \(u\) j \((t)\) is the control input of UAV \(j\) at time \(t\),

[0083] is the cost of the control input of UAV \(j\), \(\alpha\) is a constant coefficient, is the speed of UAV \(j\) at time \(t\), \(T\) is the total time of task execution, \(Q\) is the weight matrix, \(R\) is the weight matrix, \(m\) is the number of state variables, \(n\) is the number of control inputs, and \(dt\) represents the integration with respect to time \(t\).

[0084] The cost function design comprehensively considers the target trajectory deviation, the cost of control input, and the impact of speed changes, enabling the optimization process to balance various factors in multiple aspects. By minimizing the cost function, the system can accurately track the target and avoid unnecessary energy consumption.

[0085] By adjusting the constant coefficients α and the weighting matrices Q and R, the weights of the system for target trajectory deviation, control input, and speed changes can be flexibly adjusted. It can perform adaptive optimization according to different task requirements and environmental changes.

[0086] By introducing the speed change term and controlling its weight, the UAV can smoothly adjust its flight speed, avoid sharp acceleration and deceleration, enhance the stability of the system, and reduce the risk of system instability.

[0087] The total time term T in the cost function enables the method to focus on the accuracy of target tracking and consider the time optimization of task execution, ensuring that the task can be efficiently completed within a reasonable time.

[0088] By adding weights to the target and control input in the cost function, it can ensure the coordinated operation and coordinated actions among multiple targets and multiple UAVs, and avoid the efficiency decline caused by target conflicts among individual UAVs.

[0089] The application of the optimal control theory is based on the Hopfield neural network, which is used to optimize the generation of the countermeasure path. The energy function is defined as:

[0090]

[0091] where E is the energy function, x j (t) is the state of UAV j at time t, x j,target (t0 is the target position of UAV j at time t,

[0092] ||x j (t) - x j,target (t)|| 2 is the deviation between UAV j and the target position, x i (t) is the state of target i at time t, x i,ref (t) is the reference trajectory position of target i at time t,

[0093] ||x i (t) - x i,ref (t)|| 2 is the deviation between target i and the reference trajectory, m is the number of state variables, and n is the number of control inputs.

[0094] By introducing the Hopfield neural network and the energy function, the anti - interference path can be dynamically optimized to ensure that the UAV can reduce the deviation between the tracking target and the reference trajectory, and then approach the target efficiently and successfully execute the anti - interference task.

[0095] By minimizing the deviation between the target position and the UAV control input, the energy function can adaptively adjust the path in actual tasks, effectively coping with the dynamic changes of the target and the environment. The system can maintain flexible adaptability in the face of complex tactics and emergencies.

[0096] By simultaneously optimizing the errors between multiple targets, it is ensured that the UAV can effectively coordinate its behavior when performing the anti - interference task, avoid path conflicts between multiple UAVs, and improve the efficiency of the anti - interference task.

[0097] Through the optimization of the Hopfield neural network, an accurate match between the UAV path and the target position can be achieved, reducing invalid control inputs, and then realizing smoother flight and avoiding unnecessary energy consumption.

[0098] By minimizing the energy function, the system can quickly and accurately complete target anti - interference within the given task time, improving the overall anti - interference efficiency, especially in the case of multiple targets and complex environments.

[0099] The reward function in the deep Q - network is designed as:

[0100] R t =-β 1 ||x i (t)-x i,ref (t)|| 2 -β 2 ||x j (t)-x j,target (t)|| 2 +γ·||Δu j (t)|| 2 ,

[0101] where x i (t) is the state of target i at time t, x i,ref (t) is the reference trajectory of target i at time t, ||x i (t)-x i,ref (t)|| 2 is the deviation between target i and the reference trajectory,

[0102] β 1 is the target deviation weighting coefficient, x j (t) is the state of UAV j at time t, x j,target (t) is the target position of UAV j at time t,

[0103] ||xj (t)-x j,target (t)|| 2 is the deviation between the drone j and the target position,

[0104] β 2 is the deviation weighting coefficient of the drone, and γ is the weighting coefficient of the control input change,

[0105] Δu j (t) is the control input change of the drone j at time t, ||Δu j (t)|| 2 is the cost of the control input change of the drone.

[0106] By designing a reward function that includes the deviation between the target and the reference trajectory, the deviation between the drone and the target, and the cost of the control input change, the anti-countermeasure strategy can be precisely guided to optimize through deep reinforcement learning. The reward function enables the drone to minimize the target deviation and the control input change while improving the reaction speed and accuracy of the system.

[0107] The reward function considers multiple factors and can achieve comprehensive optimization in the scenario of multi-target tracking. By adjusting the weight coefficients of the target and the drone, these factors can be flexibly balanced, thereby adapting to the requirements of different anti-countermeasure tasks.

[0108] By introducing the cost of the control input change and the corresponding weighting coefficient γ, the design can reduce the situation of excessive adjustment of the control input, avoid unnecessary energy consumption and unstable control strategies, and ensure the smoothness and efficiency of the anti-countermeasure process.

[0109] The design of the reward function can adaptively adjust the weight coefficients and dynamically optimize the control strategy based on the feedback signal. The design enables the system to cope with changes in complex environments. Especially when dealing with the dynamic adjustment of the enemy's strategy, the deep Q-network can quickly respond according to the reward signal and maintain the efficient execution of the task.

[0110] By comprehensively considering the error of the target trajectory, the tracking accuracy of the drone, and the cost of the control input through the reward function, the system can efficiently execute anti-countermeasure tasks in complex environments, and improve the anti-countermeasure accuracy and task completion efficiency of the system by optimizing the learning process.

[0111] During the training process of the deep Q-network, the distributed policy gradient method is adopted to enhance the reaction ability of the system in multi-drone cooperative operations. The gradient update formula is:

[0112]

[0113] where θ t+1 is the policy parameter at time t+1, θ t is the current policy parameter, and η is the learning rate, is the policy gradient.

[0114] Adopt the distributed policy gradient method to enable multiple UAVs to share experiences during collaborative operations, and improve the overall response speed and accuracy by jointly updating the policy. In multi-UAV tasks, each UAV can independently learn and share gradient information, avoiding the limitations in the UAV learning process, and thus improving the efficiency and flexibility of overall task completion.

[0115] The distributed policy gradient method can accelerate the training process, especially in complex environments and multi-target scenarios. Through parallel computing and sharing gradient information, the training time is significantly shortened, enabling the system to adapt to environmental changes in a shorter time and adjust the policy in real time to cope with the dynamic changes of enemy targets.

[0116] In multi-UAV collaborative operations, the policies of UAVs reflect local decisions, while the distributed policy gradient method can unify the behaviors of multiple UAVs into a globally optimized policy. Through gradient sharing, the system can optimize the countermeasure paths and behaviors globally, avoiding conflicts and resource waste when multiple UAVs execute tasks.

[0117] The multi-sensor data fusion technology includes fusing radar, infrared, and visual sensor data to improve the accuracy of target state estimation. The estimation formula for sensor fusion is:

[0118]

[0119] where represents the estimated value of the target state at time k, represents the state estimation calculated in the previous time step, K k represents the fusion ratio of the measurement information and the prediction information at the current moment, z k represents the observed data of the target state obtained by the sensor at time k, H k is the observation matrix.

[0120] The multi-sensor data fusion technology combines the data of radar, infrared, and visual sensors, which can effectively make up for the deficiencies of a single sensor and improve the accuracy of target state estimation. By fusing redundant information from different sensors, the system can eliminate the noise and errors of a single sensor and obtain accurate estimates of the target position and speed.

[0121] By fusing data from different types of sensors, the system can maintain good performance under various environmental conditions, especially in low visibility and complex backgrounds. For example, radar works without being affected by light, the infrared sensor can provide effective information at night and in low light conditions, while the visual sensor is suitable for providing high-resolution target recognition data. Fusing the data enables the system to maintain robustness and reliability in a changing environment.

[0122] Different sensors can capture target information from different angles. For example, radar mainly provides distance and speed information, infrared sensors can provide temperature distribution data, and vision sensors can provide detailed images of the target. By fusing the information, the system can obtain comprehensive and rich target information, which helps to more accurately judge the state and behavior of the target and improve the quality of target tracking.

[0123] The state estimation after multi-sensor fusion is accurate, which helps to optimize the countermeasure strategy. Through accurate target positioning and state estimation, the UAV can make more reasonable countermeasure decisions in real time. The system can quickly identify the target based on the data of multi-sensor fusion and formulate an effective countermeasure path in the shortest time, improving the response speed and execution efficiency of the task.

[0124] The process of Kalman filtering includes using the Extended Kalman Filter (EKF) for state estimation of a non-linear system. The state update formula of the EKF is as follows:

[0125]

[0126] where represents the estimated value of the target state at time k, z k represents the observed data of the target state obtained by the sensor at time k, H k is the observation matrix, is the predicted value of the target state at time k - 1, and P k represents the uncertainty of the target state estimation.

[0127] As a non-linear system state estimation technique, the Extended Kalman Filter (EKF) significantly improves the accuracy of target state estimation in the present invention, especially in dynamic and complex environments. By real-time fusing sensor data, processing non-linear dynamics, and evaluating estimation uncertainty, the EKF improves the accuracy of the countermeasure task, enhances the stability and robustness of the system, ensures that the UAV can accurately and reliably perform multi-target tracking during the countermeasure task, adapt to environmental changes and make rapid responses, and improve the countermeasure efficiency and the success rate of task execution.

[0128] The path planning algorithm includes using the A* algorithm to plan the optimal countermeasure path of the UAV. The path calculation process is optimized by the following formula:

[0129] f(x) = g(x) + h(x),

[0130] where g(x) is the actual cost from the starting point to the current node x, h(x) is the estimated cost from the node x to the target, and f(x) is the total cost of the current path.

[0131] The A* algorithm in the present invention, as the core technology for path planning, provides efficient and accurate path planning capabilities for the UAV swarm when performing countermeasure tasks. By combining the actual cost and the estimated cost, the A* algorithm can ensure the optimality of the path and can adjust the path planning according to the dynamic environment to cope with real-time changes. The characteristics of high efficiency, flexibility, and optimal path generation enable the UAVs to quickly and effectively perform countermeasure tasks in complex environments, improving the task completion efficiency and saving energy.

[0132] The topology optimization algorithm includes introducing environmental constraints and dynamic adjustment factors in the path planning process. The objective function of topology optimization is:

[0133]

[0134] where λ is the weight coefficient for adjusting the influence of obstacles, obstacle(p j (t)) represents the influence of obstacles in the path, p j (t) represents the path of the UAV, p i (t) is the path of the target,

[0135] ||p j (t)-p i (t)|| 2 is the distance deviation between UAV j and target i at time t.

[0136] The topology optimization algorithm can introduce environmental constraints in the UAV path planning process. The UAV can flexibly respond to changes in obstacles in complex environments and continuously optimize the path, avoiding path conflicts or unsafe situations during flight.

[0137] By introducing the weight coefficient λ of the obstacle influence and the obstacle influence function, the system can effectively avoid obstacles and dynamically adjust the path according to the density and position of the obstacles, minimizing the collision risk with obstacles to the greatest extent and ensuring the safe flight of the UAV.

[0138] The topology optimization algorithm can simultaneously consider the position changes of multiple groups of targets and the path planning of the UAVs, enabling multiple UAVs to collaboratively optimize their respective paths, avoiding conflicts and duplicate tasks among each other, and improving the overall execution efficiency of the countermeasure tasks.

[0139] The Lyapunov stability analysis is based on the stability theory of generalized linear systems to analyze the stability of the UAV swarm system. The Lyapunov function is:

[0140]

[0141] where, x j (t) is the UAV state, x j,ref(t) is the desired trajectory, V(x) is the Lyapunov function, ∥x j (t) - x j,ref (t)∥ 2 is the error between the state of UAV j at time t and the reference trajectory position, and n is the number of control inputs.

[0142] Lyapunov stability analysis provides theoretical support for the system stability in the UAV swarm countermeasure mission. By using the Lyapunov function, the system can evaluate the error between the UAV swarm state and the reference trajectory in real time, ensuring that the UAVs can maintain a stable flight state and avoid the instability caused by dynamic changes. Lyapunov stability analysis improves the reliability and robustness of the system, ensuring efficient cooperation and stability during the mission execution, and is one of the key technologies for the successful UAV swarm countermeasure mission.

[0143] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A drone swarm countermeasure method based on deep learning, characterized in that: include: Step 1: Establish a dynamic system model of UAVs and targets. The state of each UAV is represented by the position and velocity vector of the UAV, and the state of each target is represented by the position and velocity vector of the target. Step 2: Lie group and Lie algebra are used to model the relative motion between the UAV and the target. The Lie group is used to represent the rigid body motion of the UAV and the target. The transformation of the UAV and the target in three-dimensional space is described by combining rotation and translation matrices. The Lie algebra is used to represent the relative motion between the UAV and the target. Step 3: Define the optimal control problem and optimize the countermeasure process by designing a cost function, which comprehensively considers the trajectory deviation of the target and the cost of the UAV control input, and solves the optimal control problem through optimal control theory; Step 4: Based on the results of optimal control, deep reinforcement learning is introduced to optimize the drone's countermeasure strategy. By training the deep Q network, the drone can adaptively adjust its countermeasure strategy according to the feedback reward signal in a dynamically changing environment. Step 5: Use multi-sensor data fusion technology, combined with Kalman filtering and extended Kalman filtering to estimate the state of the target in real time, and fuse the data from different sensors; Step 6: Use the path planning algorithm to develop the optimal countermeasure path for the drone. In this process, the location of the target, the current location of the drone, and obstacles are considered to calculate a path that can quickly and effectively approach the target and avoid conflicts with obstacles during flight. Step 7: Use the topology optimization algorithm to adjust the path of the UAV. The topology optimization takes into account the real-time position changes between the target and the UAV and dynamically optimizes the path of the UAV. Step 8. Apply Lyapunov stability theory to analyze the stability of the countermeasure system. By constructing the Lyapunov function, analyze the stability conditions of the Lyapunov function to successfully complete the task.

2. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The cost function is designed as: Among them, x i (t) is the state of target i at time t, x i,ref (t) is the reference trajectory position of target i at time t, is the difference between the target trajectory deviation and the reference trajectory, u j (t) is the control input of UAV j at time t, is the control input cost of UAV j, α is a constant coefficient, is the speed of UAV j at time t, T is the total time of task execution, Q is the weight matrix, R is the weight matrix, m is the number of state variables, n is the number of control inputs, and dt represents the integration over time t.

3. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The application of the optimal control theory is based on the Hopfield neural network, which is used to optimize the generation of the countermeasure path. The energy function is defined as: Where E is the energy function, x j (t) is the state of UAV j at time t, x j,target (t) is the target position of UAV j at time t, ||x j (t)-x j,target (t)|| 2 is the deviation between UAV j and the target position, x i (t) is the state of target i at time t, x i,ref (t) is the reference trajectory position of target i at time t, ||x i (t)-x i,ref (t)|| 2 is the deviation between target i and the reference trajectory, m is the number of state variables, and n is the number of control inputs.

4. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The reward function in the deep Q network is designed as: R t =-β1||x i (t)-x i,ref (t)|| 2 -β2||x j (t)-x j,target (t)|| 2 +γ·||Δu j (t)|| 2 , Among them, x i (t) is the state of target i at time t, x i,ref (t) is the reference trajectory of target i at time t, ||x i (t)-x i,ref (t)|| 2 is the deviation of target i from the reference trajectory, β1 is the target deviation weighting coefficient, x j (t) is the state of UAV j at time t, x j,target (t) is the target position of UAV j at time t, ||x j (t)-x j,target (t)|| 2 is the deviation between UAV j and the target position, β2 is the weighting coefficient of the UAV deviation, γ is the weighting coefficient of the control input change, Δu j (t) is the change in control input of UAV j at time t, ||Δu j (t)|| 2 is the cost of the change in the control input of the drone.

5. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The distributed policy gradient method is used in the deep Q network training process to enhance the system's responsiveness in multi-UAV collaborative combat. The gradient update formula is: Among them, θ t+1 is the strategy parameter at time t+1, θ t is the current strategy parameter, η is the learning rate, is the policy gradient.

6. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The multi-sensor data fusion technology includes fusing radar, infrared and visual sensor data to improve the accuracy of target state estimation. The estimation formula of the sensor fusion is: in, represents the estimated value of the target state at time k, represents the state estimate calculated at the previous time step, K k Indicates the fusion ratio of the current measurement information and the predicted information, z k represents the target state observation data obtained by the sensor at time k, H k is the observation matrix.

7. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The Kalman filtering process includes using an extended Kalman filter to perform state estimation on a nonlinear system. The state update formula of the extended Kalman filter is: in, represents the estimated value of the target state at time k, z k represents the target state observation data obtained by the sensor at time k, H k is the observation matrix, is the predicted value of the target state at time k-1, P k Represents the uncertainty in the target state estimate.

8. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The path planning algorithm includes using the A* algorithm to plan the optimal countermeasure path of the drone, and the path calculation process is optimized by the following formula: f(x)=g(x)+h(x), Among them, g(x) is the actual cost from the starting point to the current node x, h(x) is the estimated cost from node x to the target, and f(x) is the total cost of the current path.

9. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The topology optimization algorithm includes introducing environmental constraints and dynamic adjustment factors in the path planning process. The objective function of topology optimization is: Among them, λ is the weight coefficient for adjusting the impact of obstacles, obstacle(p j (t)) represents the influence of obstacles in the path, p j (t) represents the path of the UAV, p i (t) is the path to the target, ||p j (t)-p i (t)|| 2 The distance deviation between UAV j and target i at time t.

10. The method for countering drone swarms based on deep learning according to claim 1, characterized in that: The Lyapunov stability analysis is based on the generalized linear system stability theory to analyze the stability of the UAV group system. The Lyapunov function is: Among them, x j (t) is the state of the drone, x j,ref (t) is the expected trajectory, V(x) is the Lyapunov function, ||x j (t)-x j,ref (t)|| 2 The error between the state of drone j at time t and the reference trajectory position, and n is the number of control inputs.