An anti-drone intelligent interference system and method for low-altitude safety
By combining multimodal detection and deep reinforcement learning, high-precision detection and accurate interference of UAVs are achieved, solving the problems of low detection accuracy, inaccurate interference and electromagnetic pollution in existing technologies, and improving the system's adaptability and energy efficiency.
Patent Information
- Application Number
- CN202610782520.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-04
AI Technical Summary
Existing anti-drone systems suffer from low detection accuracy, inaccurate jamming, easy generation of electromagnetic environmental pollution, and lack of adaptive learning capabilities, making them unable to effectively deal with new types of drones and frequency hopping communication.
It employs a multimodal detection front-end, an intelligent fusion tracking module, a threat assessment and decision-making module, a cognitive radio jamming module, and an adaptive feedback learning module, combined with multi-hypothesis tracking algorithms, deep reinforcement learning algorithms, and cognitive radio technology, to achieve multimodal data fusion, real-time jamming strategy optimization, and adaptive adjustment.
It achieves high-precision detection and precise interference of UAVs, reduces electromagnetic pollution, improves the system's adaptability, effectively copes with new types of UAVs, reduces energy consumption, and minimizes the impact on legitimate communications.
Smart Images

Figure CN122506499A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-altitude security defense technology, and in particular to an intelligent anti-drone jamming system and method for low-altitude security. Background Technology
[0002] With the popularization of civilian drone technology, "unauthorized flights" have become frequent, posing a serious threat to the safety of sensitive areas such as airports, nuclear power plants, large event venues, and government agencies.
[0003] Traditional counter-drone methods suffer from the following shortcomings: single detection methods are susceptible to environmental interference, resulting in a high false alarm rate; jamming methods are mostly full-band blocking, affecting legitimate communications in the vicinity (such as Wi-Fi, Bluetooth, 4G / 5G); they lack adaptive learning capabilities and cannot cope with new types of drones and frequency-hopping communications; and their jamming strategies are fixed, resulting in high energy consumption and low efficiency. Therefore, there is an urgent need for an intelligent, precise, and low-collateral-damage counter-drone system. Summary of the Invention
[0004] This invention provides an intelligent anti-drone jamming system and method for low-altitude safety, aiming to solve the problems of low detection accuracy, inaccurate jamming, and easy generation of electromagnetic environmental pollution in the prior art.
[0005] To address the problems of existing technologies, this invention discloses an intelligent anti-drone jamming system and method for low-altitude security.
[0006] On one hand, the present invention provides an anti-drone intelligent jamming system for low-altitude security, comprising: a multi-modal detection front end for simultaneously receiving radio frequency signals, radar echo signals and acoustic signals from drones; The intelligent fusion tracking module, based on extended Kalman filtering and multi-hypothesis tracking algorithm, performs spatiotemporal alignment and fusion of heterogeneous sensor data output by the multimodal detection front end to form a continuous and stable target trajectory, and outputs the real-time status information of the UAV, including position, speed, heading and radio frequency characteristic parameters. The threat assessment and decision-making module is used to calculate the threat level and determine the optimal interference strategy and interference waveform parameters based on the real-time status information and the preset no-fly zone and important facility location data. The cognitive radio jamming module includes a software-defined radio platform and an adjustable gain power amplifier, used to generate and transmit cognitive jamming signals that match the uplink / downlink of the UAV according to the instructions output by the threat assessment and decision module; The adaptive feedback learning module is used to collect data on changes in the flight status of the UAV and the response of the communication link after the jamming is implemented, calculate the jamming effectiveness evaluation index, and dynamically correct the decision parameters in the threat assessment and decision-making module and the waveform parameters in the cognitive radio jamming module through an online deep reinforcement learning algorithm.
[0007] Furthermore, the multimodal detection front end includes: A passive radio frequency detection unit is used to monitor the image and data transmission frequency bands between the UAV and the remote controller. By using time difference of arrival and angle of arrival estimation technology, the initial location and radio frequency fingerprint of the UAV can be obtained. A low probability of intercept continuous wave radar unit is used to actively detect UAV targets and obtain their precise range, radial velocity and micro-Doppler characteristics; A microphone array acoustic detection unit is used to collect specific frequency band acoustic signatures generated by the drone rotor and serve as an auxiliary verification signal during radio frequency silence or radar blockage.
[0008] Furthermore, the intelligent fusion tracking module has an embedded feature-level fusion submodule. This submodule uses a pre-trained convolutional neural network to extract depth feature vectors from the radio frequency signal time-frequency map, radar point cloud data and acoustic spectrum map, respectively. The depth feature vectors are then concatenated and input into a gated recurrent unit network to output the UAV type confidence and fine-grained state information.
[0009] Furthermore, the threat assessment and decision-making module includes: The threat quantification unit is used to calculate the comprehensive threat index based on the minimum approach distance, arrival time, and payload risk coefficient of the UAV's current speed vector to the center of the critical asset. The strategy generation unit includes an interference decision agent based on a deep Q-network. The agent takes the current electromagnetic environment state, system energy margin, and threat index as inputs, and outputs the action space including interference signal type, interference duty cycle, transmission power level, and interference direction.
[0010] Furthermore, the cognitive radio jamming module includes a frequency hopping prediction submodule. This submodule employs a predictor that combines a long short-term memory network with a hidden Markov model to receive and analyze the frequency hopping sequence of the UAV's uplink in real time, predict the frequency point, dwell time, and modulation scheme of the next hop, and control the jamming generator to transmit a tracking jamming signal 0.5 to 2 milliseconds in advance of the predicted frequency point.
[0011] Furthermore, the cognitive radio jamming module also includes a waveform adaptive synthesizer, which can dynamically generate one or more of the following jamming waveforms according to the instructions of the threat assessment and decision-making module: Gaussian amplitude modulation jamming for single-carrier signals, comb spectrum jamming for OFDM signals, partial band noise jamming for spread spectrum signals, and satellite navigation spoofing signals for deceiving UAV navigation systems.
[0012] On the other hand, the present invention provides a method for countering intelligent interference from unmanned aerial vehicles (UAVs) for low-altitude security, characterized by comprising the following steps: Step S1: Multimodal cooperative detection, simultaneously acquiring radio, radar and acoustic data, and normalizing and aligning them with timestamps; Step S2: Perform asynchronous hierarchical data fusion. First, perform single-sensor tracking locally on each sensor, and then perform global fusion through a joint probabilistic data association algorithm to generate a comprehensive state estimate of the UAV. Step S3: Based on the comprehensive state estimation and the preset security strategy knowledge base, the current threat level is dynamically calculated using the fuzzy hierarchical analysis method; Step S4: When the threat level exceeds a preset threshold, the cognitive interference decision engine is triggered. The engine selects the optimal interference action online through a deep reinforcement learning model. Step S5: Based on the selected interference action, generate the corresponding cognitive interference signal and transmit it directionally to the target UAV via the phased array antenna; Step S6: Monitor the changes in the UAV's transmitted signals or flight behavior after interference in real time, and calculate the interference effectiveness. Step S7: The interference effectiveness rate is used as a reward signal and input into the deep reinforcement learning model to update the Q-value network parameters and form a closed-loop self-optimization.
[0013] Furthermore, the deep reinforcement learning model in step S4 adopts a dual deep Q-network architecture, whose state space includes: target UAV distance, speed, signal-to-noise ratio, frequency hopping rate, frequency band occupied by surrounding legal base stations, and system remaining power; its reward function is designed as follows: a positive reward is given when the UAV is successfully forced to land or driven away without interfering with legal communication, and a negative reward is given when interference fails or the authorized frequency band is interfered with.
[0014] Furthermore, the algorithm for generating the cognitive interference signal in step S5 includes the following mathematical expression: Let the predicted next hop frequency of the drone be... bandwidth is B The generated interference signal J ( t The expression in the frequency domain is:
[0015] in A The power amplitude is adaptively adjusted according to distance. s The spectrum shaping factor, Rect, is a rectangular function that concentrates interference energy around the predicted frequency point. At the same time, the coherent signal generated by the digital radio frequency memory is synchronized with the duty cycle of the UAV signal and is transmitted only during the active period of the signal.
[0016] Furthermore, the adaptive feedback learning module in step S7 also performs a transfer learning step: when the system encounters a new communication protocol of the UAV that it cannot recognize, the module uploads the time-frequency features encountered to the cloud database for matching or clustering, downloads the initial parameters of the matching interference strategy, and uses them as the initial network weights of the deep reinforcement learning model to accelerate the adaptation process to the new target.
[0017] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: 1. This invention employs a multimodal detection front-end, achieving redundant and complementary detection through feature-level fusion of heterogeneous sensor data and joint probabilistic data association algorithms. Even if one sensor is affected by environmental obstruction, electromagnetic silence, or weather interference, other modes can still maintain stable tracking. This system significantly improves the overall detection success rate for low, slow, and small UAVs, effectively solving the problem of frequent false alarms caused by flocks of birds, fallen leaves, and rain clutter in sensitive areas such as airports and nuclear power plants.
[0018] 2. This invention employs a cognitive radio jamming module and phased array directional transmission technology. Through a frequency hopping prediction submodule, it accurately locks the next hopping frequency of the UAV, transmitting a matched jamming waveform only at the predicted frequency with a narrow beam and low duty cycle. Furthermore, it reduces sidelobes through frequency-domain Gaussian shaping. Experiments have shown that while successfully blocking the UAV's image / data transmission links, no perceptible increase in packet loss rate was observed in Wi-Fi access points, Bluetooth devices, and mobile communication users within a radius of approximately 50 meters, effectively overcoming the electromagnetic pollution challenges of traditional jammers.
[0019] 3. This invention introduces a decision-making agent based on deep reinforcement learning and an adaptive feedback learning module. The system can analyze the frequency hopping sequence, modulation method, and communication protocol characteristics of UAVs online, predict the frequency hopping pattern in real time using an LSTM-HMM model, and dynamically adjust the jamming strategy. For novel UAVs using unknown frequency hopping protocols, the system further adapts to the new target within tens of seconds through a transfer learning mechanism, without the need for manual reprogramming or hardware replacement. It can automatically switch to the optimal jamming strategy when encountering various types of UAVs, achieving true cognitive adversarial capabilities.
[0020] 4. This invention employs a deep reinforcement learning agent to dynamically decide the jamming power level and duty cycle, combined with precise frequency hopping prediction and directional beam transmission, to perform pulse jamming only when necessary and with the minimum required power. Compared to traditional full-band jamming systems, this invention has a lower average transmit power and a duty cycle typically below 50%, effectively reducing total system energy consumption. For vehicle-mounted or battery-powered mobile deployment platforms, this characteristic allows for longer continuous defense time, greatly enhancing its practicality in environments with limited power supply or in the field.
[0021] 5. This invention, through a waveform adaptive synthesizer, can flexibly select from multiple response strategies based on the threat level, ranging from navigation deception-induced return to partial frequency band suppression and expulsion, and finally to full-power blocking and forced landing. For amateur drones that have strayed into restricted areas, low-power deception signals can be used to gently guide them away, avoiding the legal risks and public controversy associated with the traditional system's immediate and forceful suppression. For malicious drones carrying dangerous payloads, a rapid switch to high-power tracking and jamming mode can force them to land. This tiered response capability fills the gap in existing anti-drone systems' extreme modes of either ignoring them or shooting them down / strongly jamming them, better meeting the requirements of refined management in civilian security scenarios. Attached Figure Description
[0022] Figure 1 This is a modular schematic diagram of the anti-drone intelligent jamming system of the present invention; Figure 2 This is a schematic diagram of the overall steps of the anti-drone intelligent interference method of the present invention. Detailed Implementation
[0023] like Figure 1-2 As shown, an intelligent anti-drone jamming system and method for low-altitude security are described: In this embodiment, an anti-drone intelligent jamming system for low-altitude security is deployed at a high point surrounding a major nuclear power plant. The system consists of an integrated cabinet and a small phased array radar / jamming integrated antenna tower. Internally, it includes four core physical modules: a multimodal detection front-end, an intelligent fusion tracking module, a threat assessment and decision-making module, a cognitive radio jamming module, and an adaptive feedback learning module that runs as a software algorithm on a high-speed FPGA+GPU heterogeneous computing platform.
[0024] Specifically, after the system powers on, the multimodal detection front-end continuously scans the airspace. When an unauthorized quadcopter drone intrudes into the 3-kilometer warning zone around the nuclear power plant: The passive radio frequency detection unit first intercepted the 2.4GHz band image transmission signal between the drone and the remote controller. Using three-station time difference of arrival (TDOA) positioning, the drone's azimuth angle was initially calculated to be approximately 30 degrees north of east, with a distance of about 1200 meters.
[0025] This initial information triggered a low probability of intercept (LPI) continuous wave radar unit to perform a staring scan of the area. The radar emitted a set of linear frequency modulated continuous waves, and based on the Doppler frequency shift and time delay of the echo, the distance to the UAV was accurately measured to be 1205 meters, and the radial velocity was 15 m / s. At the same time, the micro-Doppler modulation sidebands generated by the rotor rotation were extracted, and it was initially identified as a quadcopter UAV.
[0026] Meanwhile, the microphone array acoustic detection unit captured the acoustic signature signal of the UAV rotor at a characteristic frequency of 210Hz. The beamforming algorithm confirmed that the direction was consistent with the radio frequency / radar signal, further verifying the authenticity of the target.
[0027] The aforementioned multimodal data streams (RF IQ data, radar spot data, and acoustic MFCC feature vectors) all carry precise nanosecond-level timestamps and are transmitted to the intelligent fusion tracking module via 10 Gigabit Ethernet. This module first performs time interpolation and spatial coordinate transformation (unifying radar spherical coordinates and RF angular coordinates to a geodetic rectangular coordinate system). Then, it uses an extended Kalman filter (EKF) to independently track each sensor, and finally fuses the tracks using a joint probabilistic data association (JPDA) algorithm. This process effectively suppresses multipath reflections and false alarms, forming a stable target track and continuously outputting the UAV's real-time status: coordinates (X, Y, Z), velocity vector (Vx, Vy, Vz), heading, and RF characteristic parameters (center frequency, frequency hopping rate, and modulation type estimation).
[0028] Furthermore, upon receiving the aforementioned status information, the threat assessment and decision-making module invokes a pre-set no-fly zone database. It calculates the minimum approach distance between the extrapolated trajectory of the UAV's current velocity vector and the nuclear power plant reactor building model to be 800 meters, with an estimated arrival time of 40 seconds. The embedded threat quantification unit, based on the formula Threat = w1 * (1 / approach distance) + w2 * velocity + w3 * payload risk (here, high weights are set according to the sensitivity level: w1=0.6, w2=0.2, w3=0.2), calculates a comprehensive threat index of 0.85 (out of 1; threshold 0.6 is a warning, 0.8 is mandatory interception). The decision-making module then invokes its internal Deep Q-Network (DQN)-based agent, inputting the current electromagnetic environment (a nearby 2.4GHz Wi-Fi hotspot), system energy (80% battery), and threat index 0.85. The output is the optimal jamming action: employing "intelligent tracking jamming" mode, using partial-band noise as the jamming signal type, a duty cycle of 50%, a transmission power of 2W, and precisely pointing the jamming direction at the target UAV.
[0029] Furthermore, the aforementioned decision-making instructions are transmitted in real time to the cognitive radio jamming module. This module's software-defined radio (SDR) platform utilizes the AD9361 RF transceiver chip and a Zynq FPGA series. Internally, the FPGA runs a frequency hopping prediction submodule employing a hybrid LSTM-HMM model. This model has continuously analyzed the UAV's frequency hopping sequence over the past 5 seconds (hopping every 1ms, with pseudo-random frequency changes between 2401MHz and 2479MHz). At time T, the model predicts the most likely frequency for the next hop at time T+1 to be 2445MHz with a confidence level of 92%. The FPGA controls the local oscillator (LO) to quickly lock to 2445MHz and generates a Gaussian amplitude-modulated jamming signal with a center frequency of 2445MHz and a bandwidth of 2MHz. The power is amplified to 2W by an adjustable gain power amplifier (HMC641) and finally transmitted with precise beamforming towards the UAV via the directional elements in the antenna array.
[0030] Furthermore, simultaneously, the adaptive feedback learning module runs in the background. It continuously monitors the interference effect: observing whether the drone exhibits flight wobbling, abnormal hovering, interrupted image transmission signal, or begins to return to base via a multimodal detection front-end. This module calculates interference effectiveness metrics, such as the decrease in uplink signal-to-noise ratio and the drone's speed change rate. Assuming that the drone's image transmission signal is completely interrupted after the first second of interference, and the drone begins to land in place, a high reward value of +10 is obtained. This reward value is fed back to the DQN agent in the threat assessment and decision-making module to update the Q-network parameters. Conversely, if the interference is ineffective (the drone continues to fly along its original route), a negative reward of -5 is obtained, and the system will attempt to adjust the interference action in the next decision cycle (e.g., increasing power or changing the interference waveform).
[0031] Through the aforementioned multimodal collaboration and intelligent closed-loop control, this system achieved a detection rate of over 95% for consumer-grade drones in testing, a false alarm rate of less than 0.1 times per hour, and an interference success rate of over 90%. Moreover, it did not affect the normal communication of surrounding 2.4GHz Wi-Fi users throughout the entire process (due to the extremely narrow interference beam and transmission only at the frequency hopping point), demonstrating the characteristics of high precision and low collateral damage.
[0032] In this embodiment, based on the above embodiments, the hardware and algorithm configuration of the multimodal detection front-end is defined in detail. This front-end comprises three independent but collaboratively working units: First, the passive radio frequency detection unit consists of three high-gain omnidirectional antennas distributed at different locations and three time-synchronized broadband receivers (each covering 70MHz to 6GHz). Employing the TDOA algorithm, the UAV's position is determined by calculating the time difference of the same signal arriving at different receivers and establishing a hyperbolic equation system. Simultaneously, the angle of arrival (AOA) is obtained using interferometric direction finding. By fusing TDOA and AOA information, initial positioning with an accuracy of approximately 50 meters can be obtained without active radiation.
[0033] Second, the low probability of intercept (LPI) continuous wave radar unit employs a sawtooth frequency modulated continuous wave (FMCW) system with a center frequency of 24 GHz (ISM band), a transmit bandwidth of 250 MHz, and a range resolution of up to 0.6 meters. The transmitted waveform is a pseudo-random phase-coded continuous wave with a peak power of only 0.1 W, making it difficult for enemy reconnaissance receivers to detect in noise. The receiver uses de-ramp processing to convert the echo delay into beat frequency, and extracts range and Doppler signals using FFT. The micro-Doppler analysis module performs a short-time Fourier transform on the time-frequency diagram of the rotor echo to extract blade rotation speed and number characteristics.
[0034] Third, the microphone array acoustic detection unit consists of a uniform circular array of 32 MEMS microphones with a diameter of 0.5 meters. The MUSIC algorithm is used for two-dimensional direction estimation. The digital signal processor (DSP) pre-emphasizes, frames, and windowes the acquired signal, then extracts the Mel-frequency cepstral coefficients (MFCCs) as the voiceprint feature. The real-time voiceprint is then dynamically time-warped and matched against templates in the database, such as "DJI Phantom," "DJI Mavic," and "Racing Machine."
[0035] Therefore, in the default "silent surveillance" mode, only the passive radio frequency (RF) unit and acoustic unit are activated to avoid any electromagnetic exposure. When the passive RF unit first detects a suspicious signal (e.g., a frequency-hopping signal from a non-Wi-Fi protocol in the 2.4 GHz band), the system confidence level is set to 0.6. At this point, the LPI radar unit is activated for precise confirmation. The radar unit emits a probe beam; if a point with micro-Doppler characteristics is simultaneously obtained in the same airspace, the system confidence level rises to 0.95. If the acoustic unit also extracts the UAV's acoustic signature in the corresponding direction, final confirmation is achieved, and full-scale tracking is initiated. If the RF unit fails in certain environments with strong electromagnetic interference, the system can operate solely using the radar + acoustic dual-mode; if the radar is required to remain RF silent, the system can rely solely on RF TDOA + acoustic joint positioning. These three modes are redundant, significantly improving the system's survivability and reliability in complex battlefield environments.
[0036] In summary, this multimodal front-end solves the inherent problem of single sensors being susceptible to deception or environmental interference. For example, a pure radar system can easily mistake flocks of birds for drones, but by combining a radio frequency unit (which identifies birds if there is no radio frequency signal) and an acoustic unit (bird voiceprints are different from drone voiceprints), false alarms caused by birds are successfully eliminated. In testing, the system achieved a detection range of 5 kilometers (against a DJI Mavic 3) for low-altitude, slow-moving, small targets, with a positioning accuracy better than 3 meters (RSS), fully meeting the requirements for low-altitude security.
[0037] This embodiment details an innovative submodule within the intelligent fusion tracking module: the feature-level fusion submodule. Traditional sensor fusion is mostly data-level or decision-level, while this invention introduces feature-level fusion, utilizing deep learning to automatically extract high-level features for each modality. The specific algorithm implementation is as follows: Radio frequency (RF) feature extraction is performed by converting the 10ms long IQ complex signal data acquired by the passive RF unit into instantaneous amplitude and phase data using the CORDIC algorithm, generating a 256x256 pixel time-frequency map (using short-time Fourier transform and Hamming window). This time-frequency map is then input into a pre-trained convolutional neural network (CNN) with a 5-layer convolutional layer and a 2-layer fully connected layer architecture, outputting a 128-dimensional deep feature vector. .
[0038] Radar feature extraction is performed by organizing the radar point cloud data (range, azimuth, elevation, signal-to-noise ratio, and micro-Doppler spectrum) into a sparse tensor. An improved PointNet++-based network is then used to process this point cloud, extracting the latent features of the UAV's geometry (number of rotors, arm length) and outputting a 128-dimensional feature vector. .
[0039] Acoustic feature extraction is performed by calculating a Mel spectrogram (128 Mel filters, 10ms temporal resolution) for a 1-second audio frame. This Mel spectrogram is input into a lightweight CNN with a MobileNetV2 architecture, and outputs a 128-dimensional feature vector. .
[0040] Perform feature concatenation and GRU prediction, and , , The features are concatenated along the dimensions to form a 384-dimensional joint feature vector. This vector is then fed into a gated recurrent unit (GRU) network with two layers, each containing 256 hidden units. The GRU network processes the joint features from consecutive time frames in a time-series manner. The hidden state at the last moment is used by a Softmax classifier to output the drone type confidence (e.g., DJI M300: 0.9, DIY racing drone: 0.05, others: 0.05). Simultaneously, a fully connected regression layer outputs fine-grained state information (such as drone attitude angle and payload weight estimation).
[0041] In summary, when the system first detects the target, it acquires only a small amount of coarse information. After accumulating 0.5 seconds of continuous multimodal data, the feature-level fusion submodule is activated. It processes the data streams from each sensor in parallel to generate the aforementioned deep features. Notably, even if a sensor temporarily fails (e.g., the acoustic unit is interfered with by strong winds), the GRU network can still infer the current target state using historical feature sequences and features from the remaining modalities, ensuring the continuity of tracking. When training this GRU network, a large dataset combining real-world flying drones and simulation data was used. Labels included drone model, attitude, distance, etc., and the loss function was a weighted sum of cross-entropy and mean squared error.
[0042] Compared to traditional Kalman filtering or simple weighted fusion, this method improved the accuracy of UAV model identification from 78% to 96% in experiments, and reduced the standard deviation of the error in UAV instantaneous speed estimation from 0.5 m / s to 0.12 m / s. Furthermore, when two different UAV models appear simultaneously, this module can successfully distinguish them and establish separate tracks, effectively preventing identity swapping errors when tracks intersect.
[0043] This embodiment describes how to utilize deep reinforcement learning to achieve intelligent threat assessment and interference strategy generation, replacing the traditional rigid logic based on rule tables.
[0044] The comprehensive threat index is calculated in real time through the threat quantification unit. IT = α ⋅(1 / ( d + d 0))+ β ⋅ v + c ⋅ P , where d is the minimum approach distance between the target and the center of the protected asset (the shortest distance between the predicted trajectory and the protected polygon, extrapolated by linear Kalman prediction), v is the radial velocity, and P is the payload risk coefficient (estimated by UAV type and carried heat source characteristics). α , β , cThe weights are determined using fuzzy hierarchical analysis (FAHP) and dynamically adjusted based on the type of protected target (e.g., tank farm vs. administration building). For example, for a nuclear power plant, α Take 0.7, β Take 0.2, c Take 0.1; for gatherings of people, c The value is increased to 0.4.
[0045] The policy generation unit is a perturbation decision-making agent based on a deep Q-network (DQN). Its specific configuration is as follows: State space S: 8-dimensional continuous vector [target distance (normalized 0-1), radial velocity (normalized), received signal-to-noise ratio (normalized), frequency hopping rate (Hz), occupancy rate of surrounding legal base stations (0-1), system remaining energy (0-1), historical interference success rate (0-1), comprehensive threat index (0-1)].
[0046] Action Space A: Discrete action combinations. This includes interference mode selection (0: no interference, 1: deception interference, 2: suppression interference, 3: tracking interference), interference power level (0: low, 1: medium, 2: high), and waveform selection (0: white noise, 1: comb spectrum, 2: partial frequency band noise). The total number of actions is 4*3*3=36.
[0047] Q-network architecture: a three-layer fully connected neural network, with 8 nodes in the input layer, 128 nodes in the first hidden layer (ReLU activation), 64 nodes in the second hidden layer (ReLU activation), and 36 nodes in the output layer (linear output). The optimizer used is Adam, with a learning rate of 0.001.
[0048] Experience replay pool: capacity 10,000 transition tuples (S, A, R, S'). Mini-batch training is performed every 100 decision steps (batch size=32).
[0049] The decision-making process is as follows: when the threat index threshold exceeds 0.6, the strategy generation unit is activated. The current state S is input into the DQN, and the network outputs the Q values of all actions, selecting the action corresponding to the highest Q value. For example, when the target distance is >2000 meters and the frequency hopping rate is low, the action might be "tracking jamming + medium power + comb spectrum" to attempt to block image transmission and induce return; when the target distance is <500 meters and the threat index is high, the action becomes "suppression jamming + high power + partial frequency band noise" to force it to land.
[0050] Through online learning, after 50 jamming missions, the system achieved a 35% improvement in average interception success rate compared to a fixed strategy, while reducing system energy consumption by 27%. This was because it learned an energy-saving strategy of using low power at long range and targeted high power at close range. Compared to static rules, the DQN intelligent agent can adapt to different drone behavior patterns.
[0051] In this embodiment, the frequency hopping prediction submodule is described in detail. This is the key to achieving accurate tracking of interference, especially for UAVs using frequency hopping communication.
[0052] LSTM-HMM predictor architecture: Input: A sequence of frequency points recording the past N transition cycles. Time interval .
[0053] Long Short-Term Memory (LSTM) layer: The frequency point sequence is one-hot encoded (assuming the frequency point set size is M=200) and input into the LSTM network (containing 3 layers, each with 256 LSTM units). The LSTM layer extracts the dynamic patterns of the time series and outputs a hidden state. .
[0054] Hidden Markov Model (HMM) Layers: The hidden states of the HMM correspond to the internal register states of the UAV's pseudo-random code generator (unknown, assumed to have K possibilities). The observation probability matrix of the HMM. B Output features of LSTM Parameterization is performed, meaning the probability of the current observation frequency point is determined by the sum of the previous hidden state and... A joint decision.
[0055] Predicted Output: Decode the most probable next hidden state using the Viterbi algorithm, and then output the predicted next frequency hopping point based on the observation distribution of that hidden state. and the predicted length of stay. .
[0056] Implementation details: In the FPGA, the predictor operates in a sliding window manner. Each time a new frequency hopping signal is received, the system demodulates the current frequency point in real time and updates the LSTM sequence window. The matrix operation unit within the FPGA (using a Xilinx deep learning processor unit) can complete forward inference within 0.1ms. When the next frequency hopping point is predicted... The length of stay is Then, the FPGA immediately tunes the frequency synthesizer (such as the LMX2594) to... To address prediction uncertainties, the system generates a frequency sweeping interference to cover... And its adjacent ±2.5MHz range. The jammer begins transmitting 0.5ms in advance to "capture" the frequency-hopping signal.
[0057] In actual testing of a DJI drone model, the drone's frequency hopping rate was 1000 hops per second. Traditional unpredictable jamming requires coverage of the entire 2400-2483MHz frequency band and consumes 20W of power. In contrast, the predictive jamming of this invention only needs to accurately cover the predicted frequency point before the hopping moment, requiring an average transmission power of only 2W, while still achieving a jamming success rate of 92%. For military drones with fixed pseudo-random sequences, the LSTM-HMM model can even achieve 100% prediction accuracy after several learning cycles, thus enabling "agile" tracking jamming.
[0058] In this embodiment, we illustrate the waveform adaptive synthesizer in the cognitive radio interference module, which can quickly generate various interference waveforms according to commands.
[0059] The hardware foundation is based on a direct digital frequency synthesis (DDS) architecture using FPGA+DAC (AD9176, sampling rate 12GSPS).
[0060] Supported waveform generation algorithms include Gaussian amplitude modulation interference, used to interfere with single-carrier analog image transmission. The mathematical expression is as follows: ,in A ( t ) is the envelope of the Gaussian white noise sequence after low-pass filtering, and the bandwidth can be set by the filter coefficients.
[0061] Comb spectrum interference: Used to interfere with OFDM signals. It generates multiple equally spaced narrowband pulses in the frequency domain, corresponding to OFDM subcarriers. In the FPGA, this is achieved by placing high amplitude values at specified subcarrier indices and setting other subcarriers to zero using inverse Fourier transform (IFFT). The notch depth can be flexibly set.
[0062] Partial-band noise interference is used to disrupt spread spectrum signals. The generation method involves bandpass filtering wide-bandwidth Gaussian white noise, retaining only a specific frequency band (e.g., 1MHz) with the same bandwidth as the spread spectrum signal, and maintaining a flat power spectral density within this band. The FPGA implementation uses a multiplicative windowing sequence to perform frequency domain shaping on the noise sequence.
[0063] Satellite navigation spoofing signals are used to deceive UAV satellite navigation receivers. Baseband signal samples of GPS L1 C / A code or BeiDou B1I code are pre-stored. Based on false position and time parameters provided by the decision module, the FPGA synthesizes navigation messages in real time, adjusting the code phase and Doppler shift to generate a navigation simulation signal with the same format as the real satellite signal but with false content.
[0064] When the threat assessment module determines that the drone is engaged in "potential intelligence gathering" rather than "malicious attack," it commands the synthesizer to generate precise satellite navigation false signals, inducing the drone's flight control system to believe that its position has deviated, thus causing it to slowly drift to the designated safe landing area (deception strategy). When the drone is determined to be approaching at high speed and carrying explosives, it immediately switches to "partial frequency band noise jamming" mode to suppress its control and image transmission links with maximum power (hard kill strategy).
[0065] The waveform adaptive capability enables this system to complete multi-level responses ranging from "gentle expulsion" to "violent suppression," avoiding excessive use of force or ineffective interference. For example, in one exercise, it successfully used false signals to deceive a commercial drone, causing it to land smoothly in a safe zone, while preserving the stored data within the drone for evidence collection.
[0066] In this embodiment, a complete method for countering intelligent interference from unmanned aerial vehicles using the aforementioned system is described, strictly following the seven steps of claim 7.
[0067] Step S1: Multimodal cooperative detection The system simultaneously initiates RF, radar, and acoustic scans, with each sensor outputting timestamped data packets at a frequency of 20Hz. Nanosecond-level synchronization is achieved through the IEEE 1588 precision time protocol on the FPGA. Before being fed into the fusion process, the data undergoes uniform normalization (removing the influence of dimensions) and outlier removal (based on median filtering).
[0068] Step S2: Asynchronous hierarchical data fusion Due to the different data rates of the various sensors (radar 50Hz, RF 10Hz, acoustic 20Hz), asynchronous fusion is employed. First, single-sensor tracking is performed locally on each sensor (using its respective motion model, such as CV or CT). Then, a global Joint Probabilistic Data Association (JPDA) module performs a fusion cycle every 50ms. It collects the latest state estimates from each sensor, calculates the association probability of each measurement belonging to each existing track, and updates the global state estimates. The final generated target state includes position (3σ ellipse), velocity vector, and a "fusion confidence" label.
[0069] Step S3: Fuzzy Hierarchical Analysis Threat Assessment The "distance," "approach speed," "target type confidence level," and "historical intrusion count" output from step S2 are used as input. A three-layer hierarchical structure is constructed (target layer: comprehensive threat; criterion layer: behavioral threat, attribute threat; indicator layer: distance, speed, etc.). A fuzzy judgment matrix is established, and the weights of each indicator are calculated using triangular fuzzy number operations. Finally, a weighted summation is used to obtain the precise threat level value (0-1). This method is more consistent with expert experience than the traditional weighted method and is robust to minor judgment conflicts.
[0070] Step S4: Trigger the cognitive interference decision engine Once the threat level exceeds a threshold (e.g., 0.65), the deep reinforcement learning model is immediately activated. This model has undergone at least 5000 epochs of offline pre-training and employs the dual deep Q-network architecture described in claim 8. During online decision-making, the model executes an epsilon-greedy policy (epsilon=0.1), selecting the action with the highest Q-value with a 90% probability and randomly exploring new actions with a 10% probability. The decision time is less than 5ms.
[0071] Step S5: Generate and directionally transmit jamming signals Based on the action command (e.g., action index 15 corresponds to "tracking interference + medium power + partial band noise"), the FPGA executes the frequency hopping prediction algorithm of claim 5 to generate the corresponding waveform. Simultaneously, it drives the phased array antenna (e.g., based on a 4x4 T / R module, operating in dual-band 2.4GHz / 5.8GHz). By calculating the target's current azimuth and elevation angles, a corresponding phase shift value is configured for each antenna element, forming a beam pointing towards the target. The transmitted beamwidth is approximately 15 degrees, with sidelobes below -20dB.
[0072] Step S6: Real-time monitoring of interference effectiveness During the jamming transmission intervals, a multimodal detection front-end continuously monitors the uplink / downlink signals of the target UAV using the jamming front-end (or using a separate listening channel). Key metrics are calculated: the decrease in Received Signal Strength Indicator (RSSI), the decrease in Signal-to-Noise Ratio (SNR), and whether the UAV loses lock (no signal feedback). Simultaneously, the radar continuously tracks the UAV's trajectory, calculating its velocity direction change rate and altitude descent rate. These metrics are then combined into a jamming effectiveness η (0-100%).
[0073] Step S7: Closed-loop self-optimization η is mapped to a reward value R = η - 0.5 * (number of false positives to valid communication) - 0.1 * (power consumption ratio). This reward value (S, A, R, S') is stored in the experience replay pool. Every 32 new experiences are accumulated, model training is initiated, and the DQN parameters are updated. The training process is performed asynchronously on the GPU and does not affect real-time decision-making.
[0074] In summary, this method achieves a complete OODA loop (Observe-Judgment-Decision-Action), typically taking only 0.5 to 2 seconds from detection to successful interference. In a field test at an airport, the system successfully intercepted 30 intruding drones of different types, with an average interference time (from detection to forced landing) of 15 seconds, causing no flight delays.
[0075] In this embodiment, more detailed parameter definitions are provided for the deep reinforcement learning model in the above-described method to ensure its convergence and effectiveness.
[0076] State-space quantization: Target UAV distance (d): 0-5000m, normalized and divided by 5000.
[0077] Velocity (v): 0-50 m / s, normalized by dividing by 50.
[0078] Signal-to-noise ratio (SNR): -10dB to 30dB, normalized linear mapping to 0-1.
[0079] Frequency hopping rate (hop_rate): 0-2000 hops / s, normalized and divided by 2000, then compressed using the logarithm.
[0080] Occupied frequency bands by surrounding legal base stations: Through real-time spectrum sensing, the proportion of bandwidth occupied in the 2.4G and 5.8G frequency bands is counted, ranging from 0 to 1.
[0081] System remaining power (power_left): Battery capacity or power reserve, 0-1.
[0082] Reward function design details: In the early stages of training, dense rewards are used to guide exploration: At each time step, if the target distance decreases (i.e., it is driven away), then add 0.01 * (distance decreased) / maximum distance; if the distance increases (approaches), then subtract 0.02.
[0083] A Dazheng reward of +100 is given when a drone is successfully forced to "land" (height < 1 meter and speed < 1 m / s) or "return" (distance > 2 times the initial intrusion distance).
[0084] If a drone flies out of the defense zone without causing damage, a moderate positive reward of +50 is given.
[0085] If interference causes the drone to crash out of control but not land in a safe zone, a negative reward of -30 will be given (due to the potential for secondary disasters).
[0086] Each time interference is transmitted, if packet loss or bit error is detected in a legitimate frequency band (such as a pre-recorded Wi-Fi beacon frame), an immediate negative reward of -5 is given.
[0087] Each disturbance action consumes energy, and a small negative reward (-0.1 / -0.2 / -0.3) is given according to the power level to promote energy conservation.
[0088] Network structure and training parameters: Dual DQN (DDQN) is employed to overcome the overestimation problem. An online network is used to select actions, and a target network is used to evaluate the value of the actions. Parameters are synchronized every 1000 steps. The Adam optimizer is used with a learning rate of 0.0001 and a discount factor γ = 0.95.
[0089] In summary, after 500 training rounds (each round simulating for 3 minutes), the agent learned to prioritize using "low-power comb spectrum" instead of "full-power noise" when at close range and with unknown drone types, thus avoiding alerting the drone. At the same time, it learned to precisely shift the interference frequency band to the channel used by the drone, rather than the entire ISM band, when multiple Wi-Fi devices are operating nearby, thereby achieving a balance between high interference effectiveness and low electromagnetic pollution.
[0090] In this embodiment, the preferred algorithm for generating cognitive interference signals in step S5 of the above method is described in detail, highlighting its specificity and accuracy.
[0091] Mathematical principle: Suppose that the next hop frequency is obtained through the LSTM-HMM predictor of claim 5. and frequency hopping bandwidth B Ideally, the tracking interference should coincide with the target signal in both time and frequency. Due to prediction errors, this invention designs a robust interference waveform, whose time-frequency energy distribution expression is:
[0092] in: rect( x () is a rectangular function, and the disturbance is defined only during the predicted dwell time. Internal emission (listen-emit alternation strategy).
[0093] A To achieve adaptive power amplitude, distance is measured via radar. R and the system's set desired received interference power Calculate: A = ,in For antenna gain, l λ is the wavelength.
[0094] This is the spectral shaping factor, whose value is based on the prediction confidence level. conf Dynamic adjustment: When the confidence level is 90%, Small, with highly concentrated interference energy; at a confidence level of 50%. Increase the interference bandwidth to cover more possibilities.
[0095] Interference bandwidth, usually set to B (Same as signal bandwidth) or 2 B (When the prediction is inaccurate).
[0096] The Gaussian function exp(...) ensures that the interference energy has a bell-shaped distribution in the frequency domain with extremely low sidelobes, further reducing adjacent-channel interference.
[0097] The algorithm is implemented on an FPGA, which uses Digital Radio Frequency Memory (DRFM) technology to capture a copy of the received UAV signal. This copy is then digitally modulated: copied, delayed, and shifted to a predicted frequency. The signal is then convolved with the generated noise sequence to produce the final output. This method generates an interference signal that is highly coherent with the target signal in terms of time-domain structure and modulation, making it more deceptive and even capable of inserting false data packets to interfere with its communication protocol.
[0098] In summary, the power spectral density of the generated interference signal perfectly matches the target signal using the algorithm described above. In actual testing, for a UAV with a frequency hopping rate of 500 hops / second, a conventional jammer needs to continuously transmit 20W of power across the entire frequency band to suppress it; this system only needs to pulse-transmit 2W of power at the predicted frequency point to achieve the same effect, and reduces interference to adjacent channels (such as Wi-Fi channel 1) by more than 15dB, demonstrating excellent "agility" and electromagnetic compatibility.
[0099] In this embodiment, we describe how to use transfer learning to quickly adapt when the system encounters a new type of drone with an unknown protocol, thus avoiding the long process of learning from scratch.
[0100] Problem Scenario: Suppose that one day, the system detects a drone whose frequency hopping pattern and modulation method (such as a new type of LoRa spread spectrum) have never appeared in the local training data. Initially, the cognitive interference module tries to use existing strategies, but the interference effect is very poor (interference effectiveness η < 20%).
[0101] The steps of transfer learning are as follows: Feature extraction and encapsulation are performed. The adaptive feedback learning module automatically extracts the time-frequency diagram, frequency hopping sequence, and IQ signal constellation diagram features accumulated in the past 10 seconds and encapsulates them into an "unknown target feature vector".
[0102] The system performs a cloud-based query and uploads the feature vector to a cloud-based expert system database via a secure, encrypted link. This database stores the fingerprint features of thousands of drone protocols worldwide and their corresponding optimal interference strategy parameters (e.g., optimal interference waveform, frequency sweep range, power curve).
[0103] For matching and clustering, the cloud server runs a cosine similarity-based clustering algorithm. If a known target with a similarity greater than 90% is found, the corresponding interference policy parameters are returned directly. If there is no perfect match, but a category with a similarity greater than 70% is found (e.g., "unknown CSS-modulated drone"), the initial parameters of the general policy for that category are returned. If all matches are low, an "exploratory policy" (e.g., wideband frequency sweep) is returned.
[0104] After downloading and initializing parameters, the local cognitive interference module receives the returned policy parameters (e.g., optimal interference waveform is "wideband noise", center frequency 905MHz, sweep range ±10MHz, power ramp-up). The system does not directly copy these parameters but uses them as the initial network weights for the DQN agent. initial Instead of starting with random weights or previous weights.
[0105] Rapid online fine-tuning is possible because the initial weights already possess a certain level of knowledge (i.e., "know how to interfere with similar targets"). The system only needs a small number of explorations (e.g., 10-20 interference actions) to increase the interference success rate to over 80%, whereas learning from scratch would require hundreds of attempts.
[0106] The local system periodically (e.g., weekly) packages and uploads accumulated successful interference cases and corresponding raw IQ data (after anonymization) to the cloud for continuous updates to the cloud-based general model, forming "swarm intelligence." Meanwhile, for sensitive scenarios involving national security, the system can choose to disable cloud migration and use only the locally fixed model to ensure absolute security.
[0107] In a simulated adversarial test, a novel "racing drone" was introduced, employing a non-public frequency-hopping spread spectrum protocol. Without transfer learning, the system achieved an 80% success rate after 132 interference attempts (approximately 7 minutes). However, using the transfer learning mechanism described in this invention, the system matched a similar protocol family in the cloud, downloaded initialization parameters, and achieved an 85% success rate with only 12 interference attempts (40 seconds), significantly reducing response time. This is crucial for defending against suicide drone swarm attacks.
[0108] In summary, this invention constructs a complete intelligent anti-drone system through multimodal cognitive detection, deep reinforcement learning decision-making, accurate frequency hopping prediction, and clever interference waveform synthesis, solving the key pain points of existing technologies and meeting the requirements of patent law for inventions.
[0109] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A smart anti-drone jamming system for low-altitude security, characterized in that, include: A multimodal detection front end, used to simultaneously receive radio frequency signals, radar echo signals, and acoustic signals from UAVs; The intelligent fusion tracking module, based on extended Kalman filtering and multi-hypothesis tracking algorithm, performs spatiotemporal alignment and fusion of heterogeneous sensor data output by the multimodal detection front end to form a continuous and stable target trajectory, and outputs the real-time status information of the UAV, including position, speed, heading and radio frequency characteristic parameters. The threat assessment and decision-making module is used to calculate the threat level and determine the optimal interference strategy and interference waveform parameters based on the real-time status information and the preset no-fly zone and important facility location data. The cognitive radio jamming module includes a software-defined radio platform and an adjustable gain power amplifier, used to generate and transmit cognitive jamming signals that match the uplink / downlink of the UAV according to the instructions output by the threat assessment and decision module; The adaptive feedback learning module is used to collect data on changes in the flight status of the UAV and the response of the communication link after the jamming is implemented, calculate the jamming effectiveness evaluation index, and dynamically correct the decision parameters in the threat assessment and decision-making module and the waveform parameters in the cognitive radio jamming module through an online deep reinforcement learning algorithm.
2. The anti-drone intelligent jamming system for low-altitude security according to claim 2, characterized in that, The multimodal detection front end includes: A passive radio frequency detection unit is used to monitor the image and data transmission frequency bands between the UAV and the remote controller. By using time difference of arrival and angle of arrival estimation technology, the initial location and radio frequency fingerprint of the UAV can be obtained. A low probability of intercept continuous wave radar unit is used to actively detect UAV targets and obtain their precise range, radial velocity and micro-Doppler characteristics; A microphone array acoustic detection unit is used to collect specific frequency band acoustic signatures generated by the drone rotor and serve as an auxiliary verification signal during radio frequency silence or radar blockage.
3. The anti-drone intelligent jamming system for low-altitude security according to claim 1, characterized in that, The intelligent fusion tracking module has an embedded feature-level fusion submodule. This submodule uses a pre-trained convolutional neural network to extract depth feature vectors from the radio frequency signal time-frequency map, radar point cloud data and acoustic spectrum map, respectively. The depth feature vectors are then concatenated and input into a gated recurrent unit network to output the UAV type confidence and fine-grained state information.
4. The anti-drone intelligent jamming system for low-altitude security according to claim 1, characterized in that, The threat assessment and decision-making module includes: The threat quantification unit is used to calculate the comprehensive threat index based on the minimum approach distance, arrival time, and payload risk coefficient of the UAV's current speed vector to the center of the critical asset. The strategy generation unit includes an interference decision agent based on a deep Q-network. The agent takes the current electromagnetic environment state, system energy margin, and threat index as inputs, and outputs the action space including interference signal type, interference duty cycle, transmission power level, and interference direction.
5. The anti-drone intelligent jamming system for low-altitude security according to claim 1, characterized in that, The cognitive radio jamming module includes a frequency hopping prediction submodule. This submodule uses a predictor that combines a long short-term memory network with a hidden Markov model to receive and analyze the frequency hopping sequence of the UAV uplink in real time, predict the frequency point, dwell time and modulation mode of the next hop, and control the jamming generator to transmit a tracking jamming signal 0.5 to 2 milliseconds in advance of the predicted frequency point.
6. The anti-drone intelligent jamming system and method for low-altitude security according to claim 1, characterized in that, The cognitive radio jamming module also includes a waveform adaptive synthesizer, which can dynamically generate one or more of the following jamming waveforms according to the instructions of the threat assessment and decision-making module: Gaussian amplitude modulation jamming for single-carrier signals, comb spectrum jamming for OFDM signals, partial band noise jamming for spread spectrum signals, and satellite navigation spoofing signals for deceiving UAV navigation systems.
7. A method for countering intelligent interference from unmanned aerial vehicles (UAVs) applied to the system described in any one of claims 1 to 6, characterized in that, Includes the following steps: Step S1: Multimodal cooperative detection, simultaneously acquiring radio, radar and acoustic data, and normalizing and aligning them with timestamps; Step S2: Perform asynchronous hierarchical data fusion. First, perform single-sensor tracking locally on each sensor, and then perform global fusion through a joint probabilistic data association algorithm to generate a comprehensive state estimate of the UAV. Step S3: Based on the comprehensive state estimation and the preset security strategy knowledge base, the current threat level is dynamically calculated using the fuzzy hierarchical analysis method; Step S4: When the threat level exceeds a preset threshold, the cognitive interference decision engine is triggered. The engine selects the optimal interference action online through a deep reinforcement learning model. Step S5: Based on the selected interference action, generate the corresponding cognitive interference signal and transmit it directionally to the target UAV via the phased array antenna; Step S6: Monitor the changes in the UAV's transmitted signals or flight behavior after interference in real time, and calculate the interference effectiveness. Step S7: The interference effectiveness rate is used as a reward signal and input into the deep reinforcement learning model to update the Q-value network parameters and form a closed-loop self-optimization.
8. The method according to claim 7, characterized in that, The deep reinforcement learning model in step S4 adopts a dual deep Q-network architecture. Its state space includes: target UAV distance, speed, signal-to-noise ratio, frequency hopping rate, frequency band occupied by surrounding legal base stations, and system remaining power. Its reward function is designed as follows: a positive reward is given when the UAV is successfully forced to land or driven away without interfering with legal communication, and a negative reward is given when interference fails or the authorized frequency band is interfered with.
9. The method according to claim 7, characterized in that, The algorithm for generating the cognitive interference signal in step S5 includes the following mathematical expression: Let the predicted next hop frequency of the drone be... bandwidth is B The generated interference signal J ( t The expression in the frequency domain is: , in A The power amplitude is adaptively adjusted according to distance. σ The spectrum shaping factor, Rect, is a rectangular function that concentrates interference energy around the predicted frequency point. At the same time, the coherent signal generated by the digital radio frequency memory is synchronized with the duty cycle of the UAV signal and is transmitted only during the active period of the signal.
10. The method according to claim 7, characterized in that, The adaptive feedback learning module in step S7 also performs a transfer learning step: when the system encounters a new communication protocol of the UAV that it cannot recognize, the module uploads the current time-frequency features to the cloud database for matching or clustering, downloads the initial parameters of the matching interference strategy, and uses them as the initial network weights of the deep reinforcement learning model to accelerate the adaptation process to the new target.