Unmanned aerial vehicle countering model training method, unmanned aerial vehicle countering method, control equipment and system
Through deep learning and reinforcement learning technology, the problem of low intelligence of the existing drone countermeasure system is solved, and automated intelligent decision-making and efficient countermeasures are realized in complex scenarios.
Patent Information
- Application Number
- CN202510621405.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
AI Technical Summary
The decision-making part of the existing drone counter system is low in intelligence and cannot meet the needs of efficient countermeasures in complex scenarios. Especially in high-demand scenarios such as important locations and key personnel protection, the existing countermeasures cannot provide the optimal decision-making plan.
By obtaining the status information of the drone in the target area, using deep learning and reinforcement learning technology to generate counter strategies, evaluate the counter effect and adjust the model, and realize automated intelligent decision-making and countermeasures.
The counter accuracy and intelligence level of the drone counter model are improved, and automated intelligent decisions can be made in complex scenarios, optimal countermeasures are given and real-time optimization is optimized, reducing manual intervention.
Smart Images

Figure CN120488877A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of drone countermeasure technology, and specifically to a drone countermeasure model training method, a drone countermeasure method, a control device, and a system. Background Art
[0002] In recent years, drone technology has experienced tremendous development. Technological advancements in miniaturization, low cost, adaptability, and portability have led to the widespread use of small drones in a variety of applications. In the field of electronic countermeasures, low-cost small drones have rapidly gained popularity worldwide, with their applications increasing dramatically. These drones are capable of carrying out illegal reconnaissance and attacks against fixed locations, personnel, and vehicles, among other threats. To address this issue, demand for counter-drone products is increasing, and counter-drone technologies are rapidly developing.
[0003] The inventors of this application discovered during their research that the decision-making part of existing anti-UAV systems is usually completed by the human brain or manually formulated logic. Only by manually judging the overall situation based on various environmental and target information obtained by the detection equipment can a countermeasure plan be made. This countermeasure method against drones has a low degree of intelligence and is relatively inefficient. Moreover, for high-demand scenarios, such as protection of key locations and important personnel, the existing drone countermeasure method can no longer fully meet actual needs. There is an urgent need for a drone-attended anti-UAV system with intelligent capabilities. Summary of the Invention
[0004] In view of the above problems, the embodiments of the present application provide a drone countermeasure model training method, a drone countermeasure method, a control device and a system to solve the above technical problems existing in the prior art.
[0005] According to one aspect of an embodiment of the present application, a method for training a drone countermeasure model is proposed, the method comprising:
[0006] Obtaining the first status information of each UAV in the target area;
[0007] Generating a first overall situation score corresponding to the target area according to the first state information and a preset overall situation assessment algorithm;
[0008] generating a countermeasure strategy according to the first state information of each of the drones to counter the drones in the target area, wherein the countermeasure strategy includes an action space vector;
[0009] Acquire second status information of each of the UAVs in the target area after the countermeasure strategy is implemented;
[0010] Generating a second overall situation score corresponding to the target area according to the second state information of each of the drones and a preset overall situation assessment algorithm;
[0011] generating a reward function according to the first overall situation score and the second overall situation score;
[0012] Evaluate the countermeasure strategy according to the reward function and the action space vector to generate an action value function;
[0013] The drone countermeasure model is adjusted according to the action-value function.
[0014] Preferably, in some embodiments, obtaining the first status information of each drone in the target area includes:
[0015] Obtain the distance information s between each drone and the drone countermeasure device, the speed information v of each drone, and the speed component v of each drone toward the drone countermeasure device. n ;
[0016] According to the preset distance quantization unit ρ s and the speed quantization unit ρ v The distance information s, speed information v and speed component v of each drone are respectively n Perform discretization processing;
[0017] According to the distance information s, speed information v and speed component v after the discretization process n First state information is generated.
[0018] Preferably, in some embodiments, generating a first overall situation score corresponding to the target area based on the first state information and a preset overall situation assessment algorithm includes:
[0019] Get the direct attack threat coefficient α of each drone a , UAV potential attack threat coefficient α pa and drone reconnaissance threat coefficient α o ;
[0020] Generate a single drone situation score D based on the overall situation assessment algorithm:
[0021]
[0022] Generate a first overall situation score L(t) corresponding to the target area according to the situation score D of each UAV:
[0023]
[0024] Where t is the current time, and N is the number of drones in the target area.
[0025] Preferably, in some embodiments, generating a countermeasure strategy based on the first state information of each drone to counter the drones in the target area includes:
[0026] Get the type information m and the quantity information M corresponding to each type of drone countermeasure equipment m ;
[0027] The action space vector A is determined according to the number N of drones in the target area:
[0028] A=[a(1),a(2),…,a(i)];
[0029] Where a(i) is the countermeasure mission of the i-th UAV countermeasure device against the a(i)-th UAV, 0≤a(i)≤N(t)≤N, 1≤i≤M; M is the total number of UAV countermeasure devices,
[0030] Generate a countermeasure strategy based on the first state information of each drone and the action space vector A through a first deep learning network;
[0031] The UAVs in the target area are countered according to the countermeasure strategy.
[0032] Preferably, in some embodiments, the first deep learning network includes: a convolutional layer, a fully connected layer and a Softmax layer.
[0033] Preferably, in some embodiments, generating a reward function according to the first overall situation score and the second overall situation score comprises:
[0034] Determining an immediate reward based on the first overall situation score and the second overall situation score;
[0035] A cumulative reward is determined based on the instant reward.
[0036] Preferably, in some embodiments, the drone countermeasure model includes a second deep learning network, and the second deep learning network includes a first convolutional layer, a second convolutional layer, a first fully connected layer, a second fully connected layer, a third fully connected layer and a first connection layer;
[0037] The step of evaluating the countermeasure strategy according to the reward function and the action space vector to generate an action-value function includes:
[0038] The first convolutional layer performs convolution processing on the second state information to extract state space features, and sends the state space features to the first fully connected layer;
[0039] The first fully connected layer is used to classify the state space features according to the reward function and send the processing results to the first connected layer;
[0040] The second convolutional layer performs convolution processing on the action space vector to extract action space features;
[0041] The second fully connected layer performs classification processing on the action space features and sends the processing results to the first connected layer;
[0042] The first connection layer is used to perform superposition processing on the state space feature and the action space feature;
[0043] The third fully connected layer is used to perform regression processing on the state space features and the action space features after superposition processing, and output an action value function.
[0044] According to another aspect of the embodiments of the present application, a method for countering a drone is also proposed, the method comprising:
[0045] Obtain status information of each drone in the target area;
[0046] Inputting the status information of each drone into a preset drone countermeasure model, wherein the drone countermeasure model is trained using the drone countermeasure model training method proposed in the above embodiment;
[0047] The UAV countermeasure model outputs a countermeasure strategy to counter each UAV in the target area.
[0048] According to a third aspect of an embodiment of the present application, a drone countermeasure control device is further provided, the device comprising a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;
[0049] The memory is used to store at least one program, and the program enables the processor to execute the operations of the drone countermeasure method proposed in the above embodiment.
[0050] According to a fourth aspect of the embodiments of the present application, a drone countermeasure system is further proposed, comprising a drone detection device, a drone countermeasure device, and the drone countermeasure control device proposed in the above embodiments;
[0051] The drone detection device is used to detect drones in the target area and obtain status information of the drones;
[0052] The UAV countermeasure control device is used to generate a countermeasure strategy according to the status information of the UAV and send the countermeasure strategy to the UAV countermeasure device;
[0053] The drone countermeasure device is used to perform countermeasure operations on drones in the target area according to the countermeasure strategy.
[0054] In summary, the embodiments of the present application obtain first state information of drones within a target area, evaluate the overall situation of drones within the target area, generate a countermeasure strategy, and control the drone countermeasure device to counter the drone. After the countermeasure is completed, second state information of the drones within the target area is obtained to evaluate the countermeasure effect. A reward function is determined based on the situation scores before and after. The countermeasure strategy is timely evaluated based on the reward function and the action space vector to generate an action value function. The drone countermeasure model is adjusted based on the action value function. This training method improves the countermeasure accuracy of the drone countermeasure model. When countering drones using the trained drone countermeasure model, drone state information in the target area can be obtained in real time, and a countermeasure strategy can be generated in a timely manner. Countermeasure operations against drones are performed according to the countermeasure strategy. This greatly improves the intelligent level of drone countermeasure system operation, does not require excessive human intervention, and can comprehensively evaluate detection capabilities, countermeasure resources, and effectiveness analysis in complex counter-drone scenarios. It makes automated intelligent decisions for different drone or drone swarm intrusion behaviors, provides and executes the optimal countermeasure plan, and intelligently optimizes the countermeasure strategy in real time according to situation changes to achieve the optimal overall countermeasure effectiveness.
[0055] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to more clearly understand the technical means of the embodiments of the present application, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present application. In addition, the same reference symbols are used to represent the same components throughout the drawings. In the drawings:
[0057] Figure 1 The following is a schematic diagram showing the structure of the drone countermeasure system proposed in an embodiment of the present application;
[0058] Figure 2 A flowchart of the drone countermeasure model training method proposed in an embodiment of the present application is shown;
[0059] Figure 3 A schematic diagram of the structure of the first deep learning network proposed in an embodiment of the present application is shown;
[0060] Figure 4 A schematic diagram of the structure of the second deep learning network proposed in the embodiment of the present application is shown;
[0061] Figure 5 A schematic diagram of the flow of the drone countermeasure method proposed in an embodiment of the present application is shown;
[0062] Figure 6 The figure shows a schematic structural diagram of the UAV countermeasure control device proposed in an embodiment of the present application. DETAILED DESCRIPTION
[0063] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein.
[0064] With the advancement of drone technology, drones are increasingly being used in various scenarios. Consequently, the demand for counter-drone products is increasing, and counter-drone technologies are rapidly developing. The decision-making process in existing counter-drone systems is typically performed by the human brain or manually formulated logic. Based on various environmental and target information acquired by detection equipment, humans perceive, assess, and judge the overall situation, formulate a countermeasure plan, and then use drone countermeasure control equipment to counter drones and evaluate the countermeasure effectiveness in real time. Such decision-making systems are suitable for simple counter-drone scenarios. However, for complex counter-drone scenarios, particularly those requiring unmanned operation, rapid response, and countering swarm attacks, as well as those with high drone countermeasure requirements, such as protecting key locations and personnel, the system requires fully automated detection, perception, decision-making, and countermeasure execution throughout the entire chain. This requires the drone countermeasure system to be able to automatically recommend a decision plan.
[0065] With the development of artificial intelligence (AI), deep reinforcement learning (DL) technology has become increasingly mature. Deep reinforcement learning is an AI technology that combines reinforcement learning and deep learning. Reinforcement learning uses trial-and-error interactions within an environment to learn optimal behavioral strategies to maximize long-term cumulative rewards. Reinforcement learning algorithms perform actions within an environment, which then responds with rewards or penalties based on these actions. Based on these rewards and penalties, the reinforcement learning algorithm continuously adjusts its behavioral strategies to maximize cumulative rewards. Deep learning is a machine learning method based on deep neural networks. It possesses powerful function fitting capabilities and can automatically extract features from large amounts of data. Deep learning algorithms construct complex functions to represent policies (policy functions) or value functions, enabling them to learn and make decisions more effectively in high-dimensional state spaces and complex environments.
[0066] In response to the problem that decisions made by the human brain or manually formulated logic in complex anti-UAV scenarios cannot provide optimal decision-making solutions, the inventors of this application, based on the characteristics of anti-UAV scenarios and the advantages of artificial intelligence technology, have proposed a UAV counter-measure model training method, a UAV counter-measure method, a UAV counter-measure control device, and a UAV counter-measure system. The UAV counter-measure model training method proposed in this application obtains the first state information of the UAVs in the target area, evaluates the overall situation of the UAVs in the target area, and generates a counter-measure strategy to control the UAV counter-measure device to counter the UAVs. After the counter-measure is completed, the second state information of the UAVs in the target area is obtained to evaluate the counter-measure effect. The reward function is determined based on the situation scores before and after, and the counter-measure strategy is timely evaluated based on the reward function and the action space vector to generate an action value function. The UAV counter-measure model is adjusted through the action value function. This training method improves the counter-measure accuracy of the UAV counter-measure model. When using a trained drone countermeasure model to counter drones, it is possible to obtain drone status information in the target area in real time, generate countermeasure strategies in a timely manner, and perform countermeasure operations on drones based on the countermeasure strategies, greatly improving the intelligence level of drone countermeasure system operations. It does not require excessive human intervention and can conduct a comprehensive assessment of detection capabilities, countermeasure resources, and effectiveness analysis in complex anti-drone scenarios. It can make automated intelligent decisions for the intrusion behaviors of different drones or drone swarms, provide and execute the optimal countermeasure plan, and intelligently optimize the countermeasure strategy in real time according to changes in the situation to achieve the best overall countermeasure effectiveness.
[0067] Figure 1 A structural schematic diagram of the drone countermeasure system proposed in an embodiment of the present application is shown. In the drone countermeasure system, a drone detection device 10, a drone countermeasure device 20 and a drone countermeasure control device 30 are included. The drone detection device 10 and the drone countermeasure device 20 are respectively communicated with the drone countermeasure control device 30, and the drone countermeasure control device 30 is used to control the drone detection device 10 and the drone countermeasure device 20.
[0068] Drone detection equipment 10 includes various sensors, including radar, visual, thermal imaging, radio, and optoelectronics, for detecting and tracking multiple different types of drone targets and measuring their various parameters to support overall situational awareness. These detected drone parameters, including the drone's three-dimensional position, velocity vector, communication frequency band, and drone type, are transmitted to drone countermeasure control equipment 30. Certain drone detection equipment 10 may also combine multiple detection technologies, such as radar, optoelectronics, acoustic waves, and radio frequency. Through data fusion and sharing, the advantages of each detection method can be fully utilized to overcome the shortcomings of a single technology, thereby achieving more accurate, comprehensive, and reliable drone detection.
[0069] The drone countermeasure device 20 is mainly used to deal with the threat of drones. After the drone countermeasure control device 30 receives the drone information detected by the drone detection device 10, when the drone detection device 10 detects a drone threat, the drone countermeasure device 10 is used to counterattack. Common countermeasures used by the drone countermeasure device 20 include radio interference, navigation deception, laser damage, physical net capture, etc., which can effectively form different levels of countermeasure capabilities against drones, including forced return, forced landing, deception capture, physical damage, etc. The drone countermeasure device 20 can include fixed integrated countermeasure equipment, vehicle-mounted integrated countermeasure equipment, and portable integrated countermeasure equipment. Various countermeasures can use radio interference, navigation deception, laser damage, physical net capture and other technologies separately, or they can integrate multiple detection methods and countermeasure technologies into one, which can achieve all-round and all-weather drone detection and warning, identity recognition, target direction finding, interference disposal, etc.
[0070] The drone countermeasure control device 30 is the core of the drone countermeasure system. Serving as a user interaction platform, the drone countermeasure control device 30 can be a dedicated terminal device or an integrated device integrated with the drone detection device 10 or the drone countermeasure device 20. The drone countermeasure control device 30 is primarily responsible for analyzing and displaying drone information data detected by the drone detection device 10 and, based on the analysis results, transmitting countermeasure strategies to the drone countermeasure device 20, enabling the drone countermeasure device 20 to perform countermeasure operations against the drone.
[0071] In order to realize the intelligent countermeasure operation of the drone countermeasure control device 30, the drone countermeasure control device 30 needs to have the ability to call the drone countermeasure model, realize automatic analysis of the drone status information in the target area through the drone countermeasure model, generate a countermeasure strategy, and send the countermeasure strategy to the drone countermeasure device 20 to counter the drone. The drone countermeasure model can realize automated intelligent decision-making for different drones and provide the optimal countermeasure solution. It can also optimize the countermeasure strategy in real time according to the situation changes to achieve the optimal overall countermeasure effectiveness. Among them, the drone countermeasure model can be set on the drone countermeasure control device 30 or on a remote server, and the drone countermeasure control device 30 can remotely call the drone countermeasure model through remote operation. There is no specific limitation in the embodiment of this application.
[0072] In order to realize intelligent countermeasures against drones through the drone countermeasure model, it is necessary to first train the drone countermeasure model. Figure 2 A drone countermeasure model training method proposed in an embodiment of the present application is shown, specifically including:
[0073] Step S110: Acquire first status information of each UAV in the target area;
[0074] In this embodiment, a specific anti-drone scenario is hypothesized. Multiple drone detection devices and drone countermeasures are deployed within a target drone defense area. The drone countermeasure system needs to obtain status information for multiple drones being countered. During the actual training of the drone countermeasure model, the first status information for each drone is data obtained by the drone detection device or previously acquired historical data.
[0075] The drone status information obtained by the drone detection device may generally include multiple types of information, including the drone's three-dimensional position, velocity vector, communication frequency band, and drone type. In the embodiment of the present application, in order to simplify the operation of the drone countermeasure model, only drone information related to the security situation of the target area may be used, such as: the distance information s between the drone and the drone countermeasure device, the velocity information v of each drone, and the velocity component v of each drone toward the drone countermeasure device. n Other drone information does not participate in decision-making and may not be processed in the embodiment of the present application.
[0076] When acquiring the first state information of a drone, in order to address the issue of drone state data continuity and avoid the introduction of complex models into the continuous state space, and because anti-drone decision-making is based on long-term reward optimization, it is insensitive to the numerical accuracy of single-shot data and can be quantized to form a discretization of the state. Therefore, in the embodiments of this application, it is proposed to first discretize the drone data, wherein the discretization process can be performed through the following process.
[0077] When performing discretization processing, the distance quantization unit ρ is first pre-set s and the speed quantization unit ρ v , for the distance information s, velocity information v and velocity component v n The following conversion is formed:
[0078] s=n s ρ s ;
[0079] v=n v ρ v ;
[0080] v n =n vn ρ v ;
[0081] Among them, n s 、nv 、n vn are all integers and satisfy the following conditions:
[0082]
[0083] |n vn |≤n v ;
[0084] Among them, s max is the maximum distance between the UAV and the UAV countermeasure device, v max The maximum possible speed of the drone.
[0085] The first state information after discretization is as follows:
[0086]
[0087] Among them, N(t) represents the Nth drone at time t, and the upper limit of N(t) is N max , where N max is the maximum number of drones in the target area. The above state information matrix consists of three different parameters (n s 、n v 、n vn ) forms different state space data, and multiple drone targets constitute the 3*N matrix.
[0088] Step S120: generating a first overall situation score corresponding to the target area according to the first state information and a preset overall situation assessment algorithm;
[0089] In order to be able to evaluate the overall situation in the target area, in an embodiment of the present application, the target area is scored using a preset overall situation evaluation algorithm, wherein the overall situation evaluation algorithm is used to evaluate the threat level of the drone target and the safety level of the overall situation. The overall situation evaluation algorithm has a direct correspondence with subsequent intelligent decision-making and reward functions.
[0090] In this embodiment of the present application, the situation score D of a single drone is determined by the following formula:
[0091]
[0092] Among them, α a is the direct attack threat coefficient of the drone, α pa is the potential attack threat coefficient of the drone, α o is the UAV detection threat coefficient; s is the distance between the UAV and the UAV countermeasure equipment, v is the speed of the UAV, and v n is the velocity component of the UAV toward the UAV countermeasure device.
[0093] After determining the situation score D of a single UAV, the first overall situation score L(t) of the target area is determined based on the situation scores of all individual UAVs:
[0094]
[0095] Where t is the current time, and N is the number of drones in the target area.
[0096] The first overall situation score of the target area reflects the overall security situation of the target area at the current moment, and is used to compare with the overall security situation of the target area after the countermeasure strategy is implemented to evaluate the effectiveness of the implementation of the countermeasure strategy.
[0097] Step S130: generating a countermeasure strategy based on the first state information of each drone to counter the drones in the target area, wherein the countermeasure strategy includes an action space vector;
[0098] In the embodiment of the present application, a counter-action space vector is designed for the counter-action strategy of the drone counter-action device. In order to simplify the types and inputs of the drone counter-action device, the embodiment of the present application discretizes the continuous actions of the drone counter-action device in the time dimension. It is assumed that the drone counter-action device includes m types, and each type of drone counter-action device includes M m UAV countermeasures equipment, that is, the number of each type is M1, M2, ..., M m , then the total number of drone countermeasure devices is:
[0099]
[0100] Among them, it is set that any one of the M drone countermeasure devices can only perform effective countermeasure actions on a certain drone at the same time. Based on the model-free training idea, the differences in the types of countermeasure devices themselves are ignored at the decision-making level, and they are uniformly considered as M countermeasure devices. Only the different countermeasure device attributes are distinguished when the action is actually executed.
[0101] The action space vector is based on all possible situations in which M countermeasure devices counter N(t) drone targets. An M-dimensional vector is used to represent the action space vector A, as follows:
[0102] A=[a(1),a(2),…a(i)…,a(M)];
[0103] Among them, a(i) is the countermeasure task of the i-th UAV countermeasure device against the a(i) UAV, 0≤a(i)≤N(t)≤N, 1≤i≤M; when a(i) is 0, it means that the countermeasure task is not executed. The highest dimension in the action space vector is (N max +1) M .
[0104] After setting the above-mentioned action space vector A, a countermeasure strategy is generated according to the first state information of the UAV through a first deep learning network, wherein the countermeasure action included in the countermeasure strategy is included in the action space vector A.
[0105] The first deep learning network may be an Actor network, and its structural diagram is as follows: Figure 3 As shown, the first deep learning network includes a convolutional layer, a fully connected layer and a Softmax layer.
[0106] The convolutional layer is used to extract features from the state-space data in step S110. It also extracts features from the input state information matrix (e.g., a matrix consisting of information such as the drone's position and velocity). The convolution kernel is used to slide over the input data to perform a convolution operation, capturing local features, such as potential threat patterns reflected by different position and velocity combinations. The convolutional layer also reduces the dimensionality of the input data, reducing the data volume while retaining key features, reducing the complexity of subsequent processing, and improving network computing efficiency.
[0107] The fully connected layer integrates the features extracted by the convolutional layer. Each neuron is connected to all neurons in the previous layer, enabling global analysis and correlation of input features. It integrates the different local features extracted by the convolutional layer and determines the significance of the combined features. For example, it can assess the threat level of a drone by combining features such as its distance and speed, and determine the required countermeasure strategy based on the aforementioned action space vector.
[0108] The Softmax layer receives the output of the fully connected layer and converts it into a probability distribution. In the anti-drone intelligent decision-making scenario, the drone countermeasure system's preference for different countermeasure actions is converted into the probability of executing each countermeasure action. For example, the network might output the probabilities of different countermeasure device selections and different countermeasure action combinations, such as the probability of selecting drone countermeasure device 1 to interfere with drone A is 0.7, while the probability of selecting drone countermeasure device 2 to capture drone B is 0.3, and so on.
[0109] According to different combination probabilities, the countermeasure strategy π is output, where π = π(a|s), a is the action and s is the state.
[0110] Step S140: Acquire second status information of each of the UAVs in the target area after the countermeasure strategy is implemented;
[0111] The drone countermeasure control device outputs the determined countermeasure strategy to the drone countermeasure device. The drone countermeasure device then performs a countermeasure operation against drones in the target area according to the countermeasure strategy. After the countermeasure operation is completed, the drone detection device continues to detect drones in the target area to evaluate the effectiveness of the countermeasure operation and obtain second state information of each drone in the target area. During the drone countermeasure model training process, the second state information can be trained using historical data.
[0112] Step S150: generating a second overall situation score corresponding to the target area according to the second state information of each of the drones and a preset overall situation assessment algorithm;
[0113] The drone countermeasure model recalculates the second overall situation score corresponding to the target area after the countermeasure strategy is implemented. The method for determining the second overall situation score is the same as step S120 and will not be repeated here. The second overall situation score is used to evaluate the effectiveness of the countermeasure strategy.
[0114] Step S160: generating a reward function according to the first overall situation score and the second overall situation score;
[0115] In order to reflect the change in the overall situation score after the countermeasure strategy is implemented, in this embodiment of the application, a reward function R is defined: c .
[0116] The first overall situation score is L(t), and the second overall situation score is L(t+1). For the convenience of calculation, the second overall situation score L(t+1) is based on the number of drones at a certain time node t. At the next time t+1, the score is not considered when there is a new drone intrusion.
[0117] According to the first overall situation score L(t) and the second overall situation score L(t+1), determine the immediate reward R i (t+1)=L(t+1)-L(t);
[0118] Then determine the reward function R according to the immediate reward c for:
[0119]
[0120] γ is a discount factor, ranging from 0 to 1, that balances the weight of immediate and long-term rewards. The discount factor determines the importance of long-term rewards in current decision-making. If γ is close to 1, long-term rewards are prioritized; if γ is close to 0, immediate rewards are prioritized.
[0121] Step S170: Evaluate the countermeasure strategy according to the reward function and the action space vector to generate an action value function;
[0122] In the embodiment of the present application, the drone countermeasure model further includes a second deep learning network. In the embodiment of the present application, the second deep learning network may be a Critic value network, such as Figure 4 As shown, the second deep learning network includes a first convolutional layer, a second convolutional layer, a first fully connected layer, a second fully connected layer, a third fully connected layer and a first connection layer.
[0123] The first convolutional layer performs convolution processing on the state space data in the second state information to extract state space features, and sends the state space features to the first fully connected layer; the first convolutional layer is also used to perform dimensionality reduction processing on the input data, reducing the amount of data while retaining key features, reducing the complexity of subsequent processing, and improving network computing efficiency.
[0124] The first fully connected layer is used to classify the state space features extracted by the first convolutional layer according to the reward function. Each neuron is connected to all neurons in the previous layer, and the input features are globally analyzed and associated to determine the significance of the feature combination. The first fully connected layer is also used to introduce nonlinear factors, enabling the network to model complex relationships and learn the complex nonlinear mapping relationship between state, action, and value. The processing results of the first fully connected layer are sent to the first connection layer;
[0125] The second convolutional layer performs convolution processing on the action space vector A to extract action space features. It is also used to perform dimensionality reduction processing on the input data, reducing the amount of data while retaining key features, reducing the complexity of subsequent processing, and improving network computing efficiency.
[0126] The second fully connected layer performs classification processing on the action space features and sends the processing results to the first connected layer;
[0127] The first connection layer is used to perform superposition processing on the state space feature and the action space feature;
[0128] The third fully connected layer is used to further fuse and process feature information based on the first and second fully connected layers, conducting a deeper correlation analysis between state-space features and corresponding counteractions, and more accurately assessing the value of each action. It is also used to regress the superimposed state-space features and action-space features and output the action-value function q = q(s, a), where a is the action and s is the state.
[0129] The embodiment of the present application utilizes nonlinear mapping capabilities through the third fully connected layer to refine the value assessment of each counter-action, so that the network can more accurately learn the value differences of different counter-actions in different states, providing a more accurate basis for the final action value output.
[0130] Step S180: Adjusting the drone countermeasure model according to the action value function.
[0131] In the embodiment of the present application, the action value function q(s,a) represents the expected cumulative reward that the drone countermeasure model can obtain after taking action a in state s. Through the size of the q value, the drone countermeasure model can evaluate the pros and cons of different actions in a specific state. An action with a higher q value means that the countermeasure strategy is more likely to successfully counter the drone in the future, thereby obtaining a higher cumulative reward.
[0132] The action-value function q provides a clear direction for policy optimization. The drone countermeasure model's goal is to select actions that maximize long-term cumulative rewards. Based on q-values, the drone countermeasure model optimizes its policy by selecting actions with the highest q-values. During training, the drone countermeasure model continuously updates and adjusts q-values, gradually improving its policy and enhancing the quality of its decisions.
[0133] In summary, the drone countermeasure model training method proposed in the embodiments of the present application combines artificial intelligence technologies such as reinforcement learning and deep learning. By acquiring first state information of drones within a target area, the overall situation of drones within the target area is evaluated and a countermeasure strategy is formed. After implementing the countermeasure strategy, second state information of drones within the target area is acquired to evaluate the countermeasure effect. A reward function is determined based on the situation scores before and after. The countermeasure strategy is evaluated in real time based on the reward function and the action space vector, generating an action-value function. The drone countermeasure model is adjusted based on the action-value function, thereby improving the countermeasure accuracy of the drone countermeasure model. The trained drone countermeasure model significantly improves the intelligent operation level of the drone countermeasure system, eliminating the need for excessive human intervention. It can comprehensively evaluate detection capabilities, countermeasure resources, and effectiveness analysis in complex countermeasure scenarios, make automated intelligent decisions for different drone or drone swarm intrusion behaviors, generate and execute the optimal countermeasure plan, and intelligently optimize the countermeasure strategy in real time based on situation changes to achieve the optimal overall countermeasure effectiveness.
[0134] After completing the training of the above-mentioned drone countermeasure model, the trained drone countermeasure model can be set on the drone countermeasure control device. Alternatively, the trained drone countermeasure model can be set on a cloud server, and the drone countermeasure control device can call it through the network to perform intelligent countermeasures on drones in the target area. Based on the above embodiment, the embodiment of the present application also proposes a drone countermeasure method, such as Figure 5 and Figure 1 As shown, the drone countermeasure method is applied to the drone countermeasure control device 30, and the drone countermeasure control device 30 performs the following steps:
[0135] Step S210: Obtaining status information of each UAV in the target area;
[0136] The drone countermeasure control device 30 receives the status information of each drone in the target area obtained by the drone detection device 10. The status information of the drone obtained by the drone detection device 10 can generally include a variety of information, including the three-dimensional position, velocity vector, communication frequency band and drone type of the drone.
[0137] Step S220: inputting the status information of each drone into a preset drone countermeasure model;
[0138] The drone countermeasure model is trained using the aforementioned drone countermeasure model training method, which will not be further described here. Optionally, in an embodiment of the present application, the drone countermeasure model may include an actor network and a critic network. The actor network and critic network parameters are trained using a discrete Markov decision model. The actor network is used to formulate an anti-drone execution strategy and select different drone countermeasure devices 20 for comprehensive countermeasures against different drones. The critic network is used to evaluate the value of anti-drone actions.
[0139] Step S230: The drone countermeasure model outputs a countermeasure strategy, so that the drone countermeasure device counters each drone in the target area according to the countermeasure strategy.
[0140] Based on the status information of the above-mentioned drone, the status information is input into the Actor network. The drone countermeasure model gives a targeted countermeasure strategy according to the trained Actor network, and sends the countermeasure strategy to the drone countermeasure device 20, so that the drone countermeasure device 20 can counter the specific drone according to the countermeasure strategy.
[0141] After the countermeasure strategy is implemented, the UAV countermeasure control device 30 can also continuously adjust the countermeasure strategy according to changes in the UAV target state to achieve long-term optimal countermeasure effectiveness.
[0142] To sum up, in the solution of countering drones through the above-mentioned trained drone countering model, the deep reinforcement learning technology in artificial intelligence is used to build an intelligent decision-making system to achieve the ability to intelligently select the optimal strategy and automatically execute actions; the anti-drone intelligent decision-making system using reinforcement learning technology has the ability to learn independently, and the continuously running network can use real-time data and historical experience data to intelligently train and learn the model, continuously improving the system capabilities and greatly reducing the decision-making process’s dependence on human experience; further using the Actor and Critic network to upgrade reinforcement learning to deep reinforcement learning, solve the problem of high-dimensional state space learning difficulties, and achieve more accurate situational awareness and counter-effectiveness evaluation for complex anti-drone scenarios, greatly improving the upper limit of technical capabilities compared with manual evaluation methods.
[0143] On the basis of the above embodiments, the present application also proposes a drone countermeasure control device 30, such as Figure 6 As shown, the UAV countermeasure control device 30 is applied to Figure 1In the illustrated drone countermeasure system, the drone countermeasure control device 30 may include a processor 302, a communications interface 304, a memory 306, and a communications bus 308. The processor 302, communications interface 304, and memory 306 communicate with each other via the communications bus 308. The communications interface 304 is used to communicate with other devices, such as client devices or other server network elements. The processor 302 is used to execute a program 310, which may specifically execute the drone countermeasure method proposed in the embodiments of the present application.
[0144] Specifically, program 310 may include program code, which includes computer-executable instructions. Processor 302 may be a central processing unit (CPU), an ASIC, or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the computing device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0145] Memory 306 is used to store program 310. Memory 306 may include high-speed RAM memory or non-volatile memory, such as at least one disk drive. Memory 306 stores code for executing the relevant steps of the drone countermeasure method embodiment proposed in the present application.
[0146] An embodiment of the present application further provides a computer-readable storage medium, in which executable instructions are stored. When the executable instructions are executed on a drone countermeasure control device, the drone countermeasure control device executes the operations of the drone countermeasure method provided in any of the above embodiments.
[0147] An embodiment of the present application also provides a drone countermeasure program, which is used to execute the drone countermeasure method provided in the above embodiment.
[0148] The algorithm or demonstration provided here are not inherently relevant to any particular computer, virtual system or other equipment. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing this type of system. In addition, the present application embodiment is not directed to any specific programming language yet. It should be understood that various programming languages can be utilized to realize the content of the present application described here, and the above description of specific languages is for the purpose of disclosing the best mode of implementation of the present application.
[0149] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0150] Similarly, it should be understood that in order to streamline the present application and assist in understanding one or more of the various aspects of the invention, in the above description of the exemplary embodiments of the present application, the various features of the embodiments of the present application are sometimes grouped together into a single embodiment, figure, or description thereof.
[0151] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and can be divided into multiple submodules or subunits or subassemblies. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying abstract and drawings) and all processes or units of any method or device disclosed in this manner can be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying abstract and drawings) can be replaced by alternative features providing the same, equivalent or similar purpose.
[0152] It should be noted that the above embodiments are illustrative rather than limiting of the present invention, and that those skilled in the art may design alternative embodiments without departing from the scope of the present invention. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order in which they are to be performed.
Claims
1. A UAV countermeasure model training method, characterized in that: The method comprises: Obtaining the first status information of each UAV in the target area; Generating a first overall situation score corresponding to the target area according to the first state information and a preset overall situation assessment algorithm; generating a countermeasure strategy according to the first state information of each of the drones to counter the drones in the target area, wherein the countermeasure strategy includes an action space vector; Acquire second status information of each of the UAVs in the target area after the countermeasure strategy is implemented; Generating a second overall situation score corresponding to the target area according to the second state information of each of the drones and a preset overall situation assessment algorithm; generating a reward function according to the first overall situation score and the second overall situation score; Evaluate the countermeasure strategy according to the reward function and the action space vector to generate an action value function; The drone countermeasure model is adjusted according to the action-value function.
2. The method according to claim 1, characterized in that The obtaining of first status information of each UAV in the target area includes: Obtain the distance information s between each drone and the drone countermeasure device, the speed information v of each drone, and the speed component v of each drone toward the drone countermeasure device. n ; According to the preset distance quantization unit ρ s and the speed quantization unit ρ v The distance information s, speed information v and speed component v of each drone are respectively n Perform discretization processing; According to the distance information s, speed information v and speed component v after the discretization process n First state information is generated.
3. The method according to claim 2, characterized in that Generating a first overall situation score corresponding to the target area according to the first state information and a preset overall situation assessment algorithm includes: Get the direct attack threat coefficient α of each drone a , UAV potential attack threat coefficient α pa and drone reconnaissance threat coefficient α o ; Generate a single drone situation score D based on the overall situation assessment algorithm: Generate a first overall situation score L(t) corresponding to the target area according to the situation score D of each UAV: Where t is the current time, and N is the number of drones in the target area.
4. The method according to claim 3, characterized in that Generating a countermeasure strategy according to the first state information of each drone to counter the drones in the target area includes: Get the type information m and the quantity information M corresponding to each type of drone countermeasure equipment m ; The action space vector A is determined according to the number N of drones in the target area: A=[a(1),a(2),…,a(i)]; Where a(i) is the countermeasure mission of the i-th UAV countermeasure device against the a(i)-th UAV, 0≤a(i)≤N(t)≤N, 1≤i≤M; M is the total number of UAV countermeasure devices, Generate a countermeasure strategy based on the first state information of each drone and the action space vector A through a first deep learning network; The UAVs in the target area are countered according to the countermeasure strategy.
5. The method according to claim 4, characterized in that The first deep learning network includes: a convolutional layer, a fully connected layer and a Softmax layer.
6. The method according to claim 4, characterized in that Generating a reward function according to the first overall situation score and the second overall situation score includes: Determining an immediate reward based on the first overall situation score and the second overall situation score; A cumulative reward is determined based on the instant reward.
7. The method according to claim 4, characterized in that The drone countermeasure model includes a second deep learning network, which includes a first convolutional layer, a second convolutional layer, a first fully connected layer, a second fully connected layer, a third fully connected layer and a first connection layer; The step of evaluating the countermeasure strategy according to the reward function and the action space vector to generate an action-value function includes: The first convolutional layer performs convolution processing on the second state information to extract state space features, and sends the state space features to the first fully connected layer; The first fully connected layer is used to classify the state space features according to the reward function and send the processing results to the first connected layer; The second convolutional layer performs convolution processing on the action space vector to extract action space features; The second fully connected layer performs classification processing on the action space features and sends the processing results to the first connected layer; The first connection layer is used to perform superposition processing on the state space feature and the action space feature; The third fully connected layer is used to perform regression processing on the state space features and the action space features after superposition processing, and output an action value function.
8. A method for countering drones, characterized in that: include: Obtain status information of each drone in the target area; Inputting the status information of each drone into a preset drone countermeasure model, wherein the drone countermeasure model is trained by the drone countermeasure model training method according to any one of claims 1 to 7; The UAV countermeasure model outputs a countermeasure strategy to counter each UAV in the target area.
9. A UAV countermeasure control device, characterized in that: comprising a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one program, and the program enables the processor to perform the operations of the drone countermeasure method according to claim 8.
10. A drone countermeasure system, characterized in that: The system comprises: a drone detection device, a drone countermeasure device and a drone countermeasure control device as claimed in claim 9; The drone detection device is used to detect drones in the target area and obtain status information of the drones; The UAV countermeasure control device is used to generate a countermeasure strategy according to the state information of the UAV and send the countermeasure strategy to the UAV countermeasure device; The drone countermeasure device is used to perform countermeasure operations on drones in the target area according to the countermeasure strategy.
Citation Information
Cited By
Unmanned aerial vehicle countering plan generation method based on artificial intelligence technology
CN121279746A