Adaptive tracking unmanned aerial vehicle spectrum detection module

By using an adaptive tracking UAV spectrum detection module, and leveraging technologies such as dynamic fingerprint database generation, deep reinforcement learning, and federated learning, the problem of signal acquisition and tracking in complex environments under traditional UAV spectrum detection is solved, achieving high-precision, real-time UAV signal monitoring.

CN120934660AActive Publication Date: 2025-11-11SHENZHEN YANUOXUN TECH CO LTD

Patent Information

Application Number
CN202511294744.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-11-11
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

Traditional UAV spectrum detection technology struggles to cope with dynamic changes in complex electromagnetic environments, suffers from insufficient signal acquisition capabilities, lags in parameter adjustment, poor environmental adaptability, and its centralized processing architecture cannot support swarm intelligence optimization under distributed deployment.

Method used

An adaptive tracking UAV spectrum detection module is adopted, including dynamic fingerprint database generation, deep reinforcement learning decision-making, intent prediction and compensation, extended tracking control and federated learning evolutionary unit. Through polarimetric reconfigurable broadband antenna array, time-frequency-space three-dimensional joint sampling, lightweight deep Q network, hidden Markov model and federated learning, dynamic acquisition and adaptive tracking of signals are achieved.

Benefits of technology

It achieves high-precision and high-reliability tracking capability for UAV signals, improves signal acquisition rate and tracking real-time performance, actively compensates for signal fading, enhances the robustness of cross-modal data fusion, and realizes group knowledge sharing and continuous evolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120934660A_ABST
    Figure CN120934660A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicles, in particular to a self-adaptive tracking unmanned aerial vehicle spectrum detection module which comprises a dynamic fingerprint database generation unit, a deep reinforcement learning decision unit, an intention prediction and compensation unit, an extended tracking control unit and a federal learning evolution unit. The dynamic fingerprint database accurately depicts the spatial-temporal characteristics of the signals; the deep reinforcement learning optimizes detection parameters in real time; the intention prediction unit actively compensates signal fading; the cross-modal data fusion improves the tracking robustness; therefore, the system has high-precision and high-reliability tracking capability in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to an adaptive tracking UAV spectrum detection module. Background Technology

[0002] With the rapid development and widespread application of UAV technology, its role in security, logistics, surveying and mapping is becoming increasingly prominent, but it also brings security risks such as illegal intrusion and privacy leaks. Traditional UAV spectrum detection technology mainly relies on fixed frequency band scanning and static parameter configuration, which is difficult to cope with the dynamic changes of UAV signals in complex electromagnetic environments. Existing technologies have the following limitations: (1) Insufficient signal acquisition capability, narrowband antennas and fixed sampling rates cannot adapt to new communication protocols such as frequency hopping and spread spectrum; (2) Lagging parameter adjustment, rule-based control strategies are difficult to track the rapid frequency shift and Doppler effect caused by UAV maneuvers in real time; (3) Poor environmental adaptability, single sensor data lacks the ability to compensate for signal fading and multipath interference. In addition, existing systems mostly adopt a centralized processing architecture, which cannot support the collective intelligent optimization under distributed deployment. Although some studies have attempted to introduce machine learning algorithms, key issues such as dynamic signal fingerprint modeling, cross-modal data fusion and collaborative learning under privacy protection have not yet been solved. Therefore, there is an urgent need for an adaptive spectrum detection module with environmental perception, intent prediction and collective evolution capabilities to improve the monitoring efficiency of UAVs in complex scenarios. Summary of the Invention

[0003] This invention overcomes the shortcomings of the prior art and provides an adaptive tracking UAV spectrum detection module.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: This invention discloses an adaptive tracking drone spectrum detection module, comprising a dynamic fingerprint database generation unit, a deep reinforcement learning decision-making unit, an intent prediction and compensation unit, an extended tracking control unit, and a federated learning evolution unit; The dynamic fingerprint database generation unit receives UAV signals through a polarization reconfigurable broadband antenna array, and extracts and fuses the transient frequency offset characteristics, Doppler spread characteristics and spatial angle of arrival fluctuation parameters of the signal by combining time-frequency-space three-dimensional joint sampling to construct a dynamic signal fingerprint database that includes the time-varying characteristics and spatiotemporal correlation of the signal. The deep reinforcement learning decision unit utilizes a lightweight deep Q-network, with the dynamic signal fingerprint database as the state input, to play the optimal frequency domain resolution, noise suppression threshold, and adaptive filter order in real time, thereby achieving nonlinear dynamic optimization of the detection parameters. The intent prediction and compensation unit models the UAV's motion intent using a hidden Markov model, combines the spatiotemporal correlation of the dynamic signal fingerprint database, predicts the signal fading range and frequency offset direction, and triggers the anti-fading tracking mode and frequency domain compensation strategy in advance. The extended tracking control unit fuses the radio frequency signal features optimized by the deep reinforcement learning decision unit with cross-modal data from external radar sensors, filters high-confidence features through an attention mechanism, distills and generates an environment-adaptive tracking strategy for the UAV signal, and outputs it to the distributed edge computing nodes. The federated learning evolutionary unit aggregates local adjustment experience from multiple detection modules, including parameter optimization experience from deep reinforcement learning decision-making units, prediction results from intent prediction and compensation units, and fusion strategies from extended tracking control units. Through federated learning, it iterative updates of the global parameter optimization model form a swarm intelligence-driven spectrum detection self-evolution capability.

[0005] Specifically, by receiving UAV signals through a polarization-reconfigurable broadband antenna array and combining time-frequency-space three-dimensional joint sampling, the transient frequency offset characteristics, Doppler spread characteristics, and spatial angle of arrival fluctuation parameters of the signal are extracted and fused to construct a dynamic signal fingerprint database containing the time-varying characteristics and spatiotemporal correlation of the signal. By synchronously receiving the raw radio frequency signals of the UAV through the horizontal and vertical polarization channels using a polarization-reconfigurable broadband antenna array, a dual-polarization time-domain signal stream is generated. The dual-polarized time-domain signal stream is subjected to time-frequency-space three-dimensional synchronous sampling: the signal stream is segmented with millisecond-level time windows to extract transient signal segments; short-time Fourier transform is performed on each transient signal segment to generate a time-frequency matrix; the spatial angle of arrival of the signal is calculated based on the phase difference of the antenna array, and the angle of arrival time sequence is output. In the time-frequency matrix, carrier frequency jump points are detected, and the frequency offset and offset rate between jump points are calculated to generate a transient frequency offset feature vector. The Doppler spread characteristics of the transient frequency offset feature vector are analyzed, and the Doppler spread coefficient of the frequency offset vector is determined by combining the signal propagation delay change rate of the time-domain slice, and the Doppler-frequency offset joint feature matrix is ​​output. Based on the Doppler-frequency offset joint feature matrix and the angle of arrival time series, the spatiotemporal correlation feature tensor is generated by tensor fusion analysis of the covariance relationship between the angle of arrival fluctuation and the Doppler spread coefficient. Tensor dimensionality reduction and feature decoupling are performed on the spatiotemporal correlation feature tensor to extract time-varying characteristic codes and spatial correlation codes. A dynamic fingerprint database of the signal environment is constructed by associating historical states through a rolling time window.

[0006] Specifically, a lightweight deep Q-network is used, with the dynamic signal fingerprint database as the state input, to achieve nonlinear dynamic optimization of the detection parameters by real-time game theory, optimizing the frequency domain resolution, noise suppression threshold, and adaptive filter order. The time-varying characteristic code and spatial correlation code of the current moment are extracted from the dynamic signal fingerprint database, and an environmental state feature vector is generated after feature concatenation. The environmental state feature vector is aligned with the parameter adjustment action record of the previous decision cycle in time series, and the historical state transition relationship is fused through a gated loop unit to generate an enhanced state tensor. The enhanced state tensor is input into a lightweight deep Q-network, and action value function approximation is performed in the network hidden layer: the output layer generates three discrete action vectors: frequency domain resolution adjustment action, noise suppression threshold action, and filter order action; the action combination corresponding to the maximum Q value is selected through an ε-greedy strategy to optimize the parameter combination. Based on the parameter optimization action combination, the detection parameters are configured in real time: the frequency domain resolution adjustment action is mapped to the FFT point setting of the spectrum analyzer; the noise suppression threshold action is converted into the stopband attenuation depth of the digital filter; and the filter order action is associated with the tap coefficient update of the adaptive filter. The signal-to-noise ratio improvement rate and feature false detection rate after monitoring parameter adjustment are combined with the newly generated spatiotemporal correlation feature tensor in the dynamic fingerprint database to construct an environmental feedback tuple; The environmental feedback tuples are stored in the priority experience replay pool, and the sampling weights are calculated through temporal differential error. The weight parameters of the deep Q network are iteratively updated to form a closed-loop optimization mechanism.

[0007] Specifically, by modeling the UAV's motion intention using a Hidden Markov Model and combining the spatiotemporal correlation of a dynamic signal fingerprint database, the signal fading interval and frequency offset direction are predicted, triggering anti-fading tracking mode and frequency domain compensation strategy in advance. Obtain the spatiotemporal correlation feature tensor within the current rolling time window from the dynamic signal fingerprint database, and separate the time-varying characteristic encoding and spatial correlation encoding therein; The time-varying characteristics are encoded and input into the state layer of the Hidden Markov Model to analyze the signal Doppler frequency shift change law caused by the UAV motion and generate a state transition probability matrix containing the acceleration mutation probability. The spatial correlation encoding is input into the observation layer of the hidden Markov model, and combined with historical signal fading interval data, the statistical correlation between spatial angle of arrival fluctuation and frequency offset direction is calculated, and the observation probability matrix is ​​output. By fusing the state transition probability matrix and the observation probability matrix using the Viterbi algorithm, the UAV intention state sequence for the next three decision cycles is decoded and generated, including three types of motion intention labels: climb, dive, and sharp turn. Based on the UAV intention state sequence, and combined with the historical attenuation patterns of the Doppler-frequency offset joint feature matrix in the dynamic signal fingerprint database, the signal fading start time and duration interval are deduced, and a fading prediction vector is generated. Based on the fading prediction vector and the frequency offset direction label in the intention state sequence, the anti-fading tracking mode is activated in advance: when the prediction is a climb intention, Doppler fading margin compensation is initiated, and when the prediction is a sharp turn intention, spatial beamforming compensation strategy is activated.

[0008] Specifically, when an ascent intention is predicted, Doppler fading margin compensation is initiated; when a sharp turn intention is predicted, a spatial beamforming compensation strategy is activated. Based on the frequency offset direction label in the intent state sequence, extract the positive Doppler offset corresponding to the climbing intent or the negative Doppler offset corresponding to the sharp turn intent, and generate an intent mode identifier. When the intent mode identifier is climb intent, the maximum fading depth record with the same positive Doppler offset in the historical Doppler-frequency offset joint feature matrix is ​​called to generate the Doppler margin compensation value; and the frequency compensation operation is triggered before the fading prediction start time point (such as the first 200 milliseconds) to increase the receiver local oscillator frequency by the Doppler margin compensation value. When the intent mode identifier is a sudden turn intent, the azimuth angle change rate is calculated based on the angle of arrival time sequence to generate the beam steering angle acceleration; and the phase control unit of the antenna array is driven to generate the spatial beam scanning trajectory according to the beam steering angle acceleration. The spatiotemporal correlation feature tensor after compensation is collected in real time to verify the volatility of the Doppler spread coefficient or the stability of the spatial angle of arrival, generate the anti-fading effectiveness index, associate the anti-fading effectiveness index with the current intention mode identifier and feed it back to the anti-fading strategy library to update the corresponding scenario compensation parameters.

[0009] Specifically, the radio frequency signal features optimized by the deep reinforcement learning decision unit are fused with cross-modal data from external radar sensors. High-confidence features are then selected through an attention mechanism, and a drone signal environment-adaptive tracking strategy is generated through distillation. This strategy is then output to distributed edge computing nodes. The system receives the radio frequency signal feature vector output by the deep reinforcement learning decision unit and simultaneously acquires point cloud trajectory data provided by an external radar sensor; it also matches the radio frequency feature timestamp with the radar trajectory timestamp to generate a spatiotemporally synchronized multimodal feature tensor; the radio frequency signal includes frequency domain resolution and filtering order; Cross-domain confidence assessment is performed on the multimodal feature tensor: the time-frequency stability index of the radio frequency features is extracted as the spectrum confidence factor; the spatial continuity score of the radar point cloud trajectory is calculated as the trajectory confidence factor; the two confidence factors are input into the gated attention layer to generate a multimodal confidence attention vector. The multimodal feature tensor is weighted and filtered based on the multimodal confidence attention vector: the first type of radio frequency feature subset with a spectral confidence factor higher than the first threshold is retained; the second type of radar trajectory feature subset with a trajectory confidence factor higher than the second threshold is selected; the two feature subsets are fused through feature crossover to generate a high-confidence fused feature vector; The high-confidence fused feature vector is input into the policy distillation network: spatial-spectral joint feature patterns are extracted in the convolutional embedding layer; and a tracking policy parameter set containing beam pointing angle, scan rate, and frequency compensation is generated through the fully connected policy compilation layer. The tracking strategy parameter group is encoded into binary instructions, and after adding timestamps and location tags, an environment-adaptive tracking strategy instruction set is generated. This set is then distributed to distributed edge computing nodes via a low-latency communication protocol to perform beam reconfiguration operations.

[0010] Specifically, it aggregates local adjustment experience from multiple detection modules, including deep reinforcement learning decision unit parameter optimization experience, intent prediction and compensation unit prediction results, and extended tracking control unit fusion strategy. Through federated learning, it iteratively updates the global parameter optimization model, forming a swarm intelligence-driven spectrum detection self-evolution capability. Collect local tuning experience from each detection module, including parameter optimization action records of the deep reinforcement learning decision unit, fading prediction vector of the intent prediction and compensation unit, and high-confidence fusion feature vector of the extended tracking control unit, to generate a multi-source local experience dataset. The multi-source local experience dataset is timestamped and its feature dimensions are normalized. Common feature parameters are extracted, including frequency domain resolution adjustment, Doppler compensation value, and beam pointing angle deviation, to generate a standardized local experience matrix. The standardized local experience matrix is ​​input into the local parameter optimization model of each detection module, and the model parameter increment is calculated by gradient descent to generate an encrypted local gradient update vector. It receives encrypted local gradient update vectors from multiple probe modules, performs weighted aggregation through a federated averaging algorithm, and generates a global gradient update vector, where the weights are dynamically allocated based on the amount of local empirical data of each module. Add Gaussian-distributed noise perturbation to the global gradient update vector to generate a privacy-preserving global gradient update vector, ensuring data security. The privacy-preserving global gradient update vector is distributed to each detection module to update the local parameter optimization model, forming a globally consistent spectrum detection strategy optimization model. Monitor the tracking accuracy and response latency of the updated model to generate model performance evaluation metrics, which are used to dynamically adjust the aggregation frequency and learning rate of federated learning.

[0011] This invention addresses the technical deficiencies in the prior art and has the following beneficial effects: Through a multi-unit collaborative intelligent spectrum detection architecture, it achieves dynamic acquisition and adaptive tracking of UAV signals; a dynamic fingerprint database accurately characterizes the spatiotemporal features of signals; deep reinforcement learning optimizes detection parameters in real time; an intent prediction unit actively compensates for signal fading; cross-modal data fusion enhances tracking robustness; and federated learning enables group knowledge sharing and continuous evolution, giving the system high-precision and high-reliability tracking capabilities in complex environments. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained from these drawings without creative effort.

[0013] Figure 1 This is a system framework diagram of the adaptive tracking UAV spectrum detection module; Figure 2 This is a flowchart of the workflow of the adaptive tracking UAV spectrum detection module. Detailed Implementation

[0014] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0015] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0016] like Figure 1 , 2 As shown, the present invention discloses an adaptive tracking UAV spectrum detection module, including a dynamic fingerprint database generation unit, a deep reinforcement learning decision unit, an intent prediction and compensation unit, an extended tracking control unit, and a federated learning evolution unit; S1. The dynamic fingerprint database generation unit receives UAV signals through a polarization reconfigurable broadband antenna array, and extracts and fuses the transient frequency offset characteristics, Doppler spread characteristics and spatial angle of arrival fluctuation parameters of the signal by combining time-frequency-space three-dimensional joint sampling to construct a dynamic signal fingerprint database containing the time-varying characteristics and spatiotemporal correlation of the signal. S2. The deep reinforcement learning decision unit utilizes a lightweight deep Q-network, with the dynamic signal fingerprint database as the state input, to play the optimal frequency domain resolution, noise suppression threshold, and adaptive filter order in real time, thereby achieving nonlinear dynamic optimization of the detection parameters. S3. The intent prediction and compensation unit models the UAV's motion intent using a hidden Markov model, and combines the spatiotemporal correlation of the dynamic signal fingerprint database to predict the signal fading range and frequency offset direction, thereby triggering the anti-fading tracking mode and frequency domain compensation strategy in advance. S4. The extended tracking control unit fuses the radio frequency signal features optimized by the deep reinforcement learning decision unit with the cross-modal data of the external radar sensor, filters high-confidence features through an attention mechanism, distills and generates an environment-adaptive tracking strategy for the UAV signal, and outputs it to the distributed edge computing node. S5. The federated learning evolution unit aggregates the local adjustment experience of multiple detection modules, including the parameter optimization experience of the deep reinforcement learning decision unit, the prediction results of the intent prediction and compensation unit, and the fusion strategy of the extended tracking control unit. Through federated learning, the global parameter optimization model is iteratively updated to form a spectrum detection self-evolution capability driven by swarm intelligence.

[0017] Specifically, by receiving UAV signals through a polarization-reconfigurable broadband antenna array and combining time-frequency-space three-dimensional joint sampling, the transient frequency offset characteristics, Doppler spread characteristics, and spatial angle of arrival fluctuation parameters of the signal are extracted and fused to construct a dynamic signal fingerprint database containing the time-varying characteristics and spatiotemporal correlation of the signal. By synchronously receiving the raw radio frequency signals of the UAV through the horizontal and vertical polarization channels using a polarization-reconfigurable broadband antenna array, a dual-polarization time-domain signal stream is generated. The dual-polarized time-domain signal stream is subjected to time-frequency-space three-dimensional synchronous sampling: the signal stream is segmented with millisecond-level time windows to extract transient signal segments; short-time Fourier transform is performed on each transient signal segment to generate a time-frequency matrix; the spatial angle of arrival of the signal is calculated based on the phase difference of the antenna array, and the angle of arrival time sequence is output. In the time-frequency matrix, carrier frequency jump points are detected, and the frequency offset and offset rate between jump points are calculated to generate a transient frequency offset feature vector. The Doppler spread characteristics of the transient frequency offset feature vector are analyzed, and the Doppler spread coefficient of the frequency offset vector is determined by combining the signal propagation delay change rate of the time-domain slice, and the Doppler-frequency offset joint feature matrix is ​​output. Based on the Doppler-frequency offset joint feature matrix and the angle of arrival time series, the spatiotemporal correlation feature tensor is generated by tensor fusion analysis of the covariance relationship between the angle of arrival fluctuation and the Doppler spread coefficient. Tensor dimensionality reduction and feature decoupling are performed on the spatiotemporal correlation feature tensor to extract time-varying characteristic codes and spatial correlation codes. A dynamic fingerprint database of the signal environment is constructed by associating historical states through a rolling time window.

[0018] It should be noted that since most UAVs employ frequency hopping and spread spectrum techniques during flight, and are affected by factors such as the Doppler effect and environmental obstruction, the signal frequency, polarization state, and spatial angle of arrival can change rapidly. This invention uses a polarization-reconfigurable broadband antenna array (such as a 4×4 UCA array) to simultaneously receive UAV RF signals from horizontal (H) and vertical (V) polarized channels, generating a dual-polarized time-domain signal stream. This antenna supports dynamic switching of polarization modes to adapt to the communication polarization characteristics of different UAVs (such as linear / circular polarization). The dual-polarized signal stream is segmented in 10ms time windows to extract transient signal segments (e.g., each segment contains 1024 sampling points). A short-time Fourier transform is performed on each transient segment to generate a time-frequency matrix (time-frequency-energy distribution), and carrier frequency jump points are detected. The spatial angle of arrival of the signal is calculated using the phase difference of the antenna array, and the angle of arrival time sequence is output (e.g., 100 sampling points per second). In the time-frequency matrix, carrier frequency jump points are detected (such as frequency abrupt changes caused by frequency hopping communication), and the frequency offset (Δf) and offset rate (df / dt) of adjacent jump points are calculated to generate a transient frequency offset feature vector. For example, if a UAV signal jumps from 2.4GHz to 2.42GHz within 50ms, then Δf = 20MHz and df / dt = 0.4MHz / ms.

[0019] By combining the signal propagation delay change rate of time-domain slices (e.g., calculated based on correlation peak displacement), the Doppler spread characteristics in the frequency offset vector (e.g., frequency broadening caused by high-speed motion) are analyzed to determine the Doppler spread coefficients, and the Doppler-frequency offset joint feature matrix is ​​output (dimension: time × frequency × Doppler coefficients). The Doppler-frequency offset joint feature matrix is ​​then fused with the angle-of-arrival time series using tensor fusion (e.g., using Tucker decomposition) to analyze the covariance relationship between the angle-of-arrival fluctuation and the Doppler spread coefficients, generating a spatiotemporal correlation feature tensor (dimension: time × frequency × space × Doppler). Tensor dimensionality reduction (PCA) and feature decoupling are performed on the spatiotemporal correlation feature tensor to extract time-varying characteristic codes (e.g., LSTM time series features) and spatial correlation codes (e.g., convolutional features as shown in the figure). Historical states are then associated through a rolling time window (e.g., a 30-second window) to construct a dynamically updated signal fingerprint database.

[0020] In summary, this invention constructs a dynamic signal fingerprint database in real time by using a polarization-reconfigurable antenna array and time-frequency-space three-dimensional joint sampling, accurately capturing the transient frequency offset, Doppler characteristics and spatial fluctuations of UAV signals, thereby improving the signal acquisition rate and tracking real-time performance in complex environments.

[0021] Specifically, a lightweight deep Q-network is used, with the dynamic signal fingerprint database as the state input, to achieve nonlinear dynamic optimization of the detection parameters by real-time game theory, optimizing the frequency domain resolution, noise suppression threshold, and adaptive filter order. The time-varying characteristic code and spatial correlation code of the current moment are extracted from the dynamic signal fingerprint database, and an environmental state feature vector is generated after feature concatenation. It should be noted that the time-varying characteristic code (such as Doppler frequency shift trend) and spatial correlation code (such as angle of arrival fluctuation) of the current moment are extracted from the dynamic signal fingerprint database, and the two are concatenated to generate an environmental state feature vector.

[0022] The environmental state feature vector is aligned with the parameter adjustment action record of the previous decision cycle in time series, and the historical state transition relationship is fused through a gated loop unit to generate an enhanced state tensor. It should be noted that the current state features are time-aligned with the parameter adjustment actions of the previous decision cycle (such as the previous FFT point setting), input into the gated recurrent unit, and historical state transition relationships are fused to output an enhanced state tensor. The hidden layer of the gated recurrent unit can memorize long-term dependencies, such as the asymptotic frequency offset caused by the drone's continuous left turns.

[0023] The enhanced state tensor is input into a lightweight deep Q-network, and action value function approximation is performed in the network hidden layer: the output layer generates three discrete action vectors: frequency domain resolution adjustment action, noise suppression threshold action, and filter order action; the action combination corresponding to the maximum Q value is selected through an ε-greedy strategy to optimize the parameter combination. It should be noted that the enhanced state tensor is input into a lightweight DQN (such as a 4-layer fully connected network with parameter count controlled within 1MB), and the output layer generates three discrete action vectors: frequency domain resolution adjustment action (e.g., FFT point count can be selected as 256 / 512 / 1024); noise suppression threshold action (e.g., stopband attenuation depth can be selected as 20dB / 30dB / 40dB); and filter order action (e.g., adaptive filter tap count can be selected as 16 / 32 / 64). Then, an ε-greedy strategy (ε=0.1) is used to select the action combination with the largest Q value, balancing exploration and exploitation.

[0024] Based on the parameter optimization action combination, the detection parameters are configured in real time: the frequency domain resolution adjustment action is mapped to the FFT point setting of the spectrum analyzer; the noise suppression threshold action is converted into the stopband attenuation depth of the digital filter; and the filter order action is associated with the tap coefficient update of the adaptive filter. The signal-to-noise ratio improvement rate and feature false detection rate after monitoring parameter adjustment are combined with the newly generated spatiotemporal correlation feature tensor in the dynamic fingerprint database to construct an environmental feedback tuple; It should be noted that the FFT points are mapped to the spectrum analyzer; for example, a 512-point FFT is selected to balance frequency resolution and computational delay. The noise suppression threshold is converted into the stopband attenuation depth of the digital filter (e.g., 30dB) to suppress adjacent channel interference. The tap coefficients of the adaptive filter are updated according to the filter order to optimize signal separation.

[0025] Next, the improved signal-to-noise ratio and false detection rate of the adjusted signal-to-noise ratio are monitored, and the new spatiotemporal correlation feature tensor of the dynamic fingerprint database is combined to construct a feedback tuple (state, action, reward, new state).

[0026] The environmental feedback tuples are stored in the priority experience replay pool, and the sampling weights are calculated through temporal differential error. The weight parameters of the deep Q network are iteratively updated to form a closed-loop optimization mechanism.

[0027] It should be noted that the feedback tuples are stored in a priority experience replay pool, and sampling weights are assigned according to the temporal differential error (TD-error). For example, samples in frequency offset abrupt change scenarios have higher priority. The DQN weights are iteratively updated by minimizing the Bellman error, forming an adaptive optimization closed loop.

[0028] In summary, this step enables UAVs to meet the real-time tracking requirements of rapid maneuvering scenarios, while effectively improving the signal-to-noise ratio and reducing the bit error rate. Furthermore, the closed-loop feedback mechanism allows the system to maintain stable tracking even under complex electromagnetic interference.

[0029] Specifically, by modeling the UAV's motion intention using a Hidden Markov Model and combining the spatiotemporal correlation of a dynamic signal fingerprint database, the signal fading interval and frequency offset direction are predicted, triggering anti-fading tracking mode and frequency domain compensation strategy in advance. Obtain the spatiotemporal correlation feature tensor within the current rolling time window from the dynamic signal fingerprint database, and separate the time-varying characteristic encoding and spatial correlation encoding therein; It should be noted that the spatiotemporal correlation feature tensor within the current rolling time window (e.g., 500ms) is extracted from the dynamic signal fingerprint database, and the time-varying characteristic code (including Doppler frequency shift trend, signal attenuation rate, etc.) and spatial correlation code (including angle of arrival fluctuation mode, polarization state change, etc.) are separated through feature decoupling.

[0030] The time-varying characteristics are encoded and input into the state layer of the Hidden Markov Model to analyze the signal Doppler frequency shift change law caused by the UAV motion and generate a state transition probability matrix containing the acceleration mutation probability. The spatial correlation encoding is input into the observation layer of the hidden Markov model, and combined with historical signal fading interval data, the statistical correlation between spatial angle of arrival fluctuation and frequency offset direction is calculated, and the observation probability matrix is ​​output. By fusing the state transition probability matrix and the observation probability matrix using the Viterbi algorithm, the UAV intention state sequence for the next three decision cycles is decoded and generated, including three types of motion intention labels: climb, dive, and sharp turn. Based on the UAV intention state sequence, and combined with the historical attenuation patterns of the Doppler-frequency offset joint feature matrix in the dynamic signal fingerprint database, the signal fading start time and duration interval are deduced, and a fading prediction vector is generated. It should be noted that the Viterbi algorithm is used to decode the optimal state sequence, outputting the intention prediction results for the next three decision cycles (e.g., 300ms). Simultaneously, by combining a historical attenuation pattern library (recording signal fading characteristics under different motion states), the expected fading start time and duration are calculated.

[0031] Based on the fading prediction vector and the frequency offset direction label in the intention state sequence, the anti-fading tracking mode is activated in advance: when the prediction is a climb intention, Doppler fading margin compensation is initiated, and when the prediction is a sharp turn intention, spatial beamforming compensation strategy is activated.

[0032] It should be noted that when a UAV performs maneuvers such as climbing, diving, or sharp turns, its communication signal experiences rapid frequency shift due to the Doppler effect, accompanied by a sharp decline in signal strength. Existing compensation mechanisms, typically employing fixed thresholds, cannot predict the UAV's movement intentions, resulting in compensation actions that are always "half a beat late," severely impacting the continuity and stability of tracking. Especially in complex electromagnetic environments, multipath effects further exacerbate signal fading, making it difficult to maintain reliable UAV tracking. Therefore, this invention uses a Hidden Markov Model to intelligently predict the UAV's movement intentions, combined with spatiotemporal feature analysis of a signal fingerprint database, to achieve early prediction of signal fading intervals and frequency shift directions. This proactively triggers targeted anti-fading compensation strategies, thereby improving the continuity and stability of UAV signal tracking and effectively overcoming the lag problem of traditional passive compensation methods, enabling the system to maintain reliable signal acquisition capabilities even during rapid UAV maneuvers.

[0033] Specifically, when an ascent intention is predicted, Doppler fading margin compensation is initiated; when a sharp turn intention is predicted, a spatial beamforming compensation strategy is activated. Based on the frequency offset direction label in the intent state sequence, extract the positive Doppler offset corresponding to the climbing intent or the negative Doppler offset corresponding to the sharp turn intent, and generate an intent mode identifier. When the intent mode identifier is climb intent, the maximum fading depth record with the same positive Doppler offset in the historical Doppler-frequency offset joint feature matrix is ​​called to generate the Doppler margin compensation value; and the frequency compensation operation is triggered before the fading prediction start time point (such as the first 200 milliseconds) to increase the receiver local oscillator frequency by the Doppler margin compensation value. It should be noted that when the intent identifier is climb (the frequency offset direction label is "positive"), the maximum fading depth record under the same positive Doppler offset is retrieved from the history database (for example, the maximum fading depth of a certain model is 12dB at a 20° climb angle). 200ms before the predicted fading start time, the receiver local oscillator frequency is increased by the compensation value (e.g., increased by +45kHz) to cancel the Doppler frequency shift.

[0034] When the intent mode identifier is a sudden turn intent, the azimuth angle change rate is calculated based on the angle of arrival time sequence to generate the beam steering angle acceleration; and the phase control unit of the antenna array is driven to generate the spatial beam scanning trajectory according to the beam steering angle acceleration. It should be noted that by linearly fitting the continuous angle-of-arrival sampling values ​​through a sliding time window (e.g., 100ms), the ratio of the azimuth angle change within the window to the time interval is calculated to obtain the instantaneous angular velocity. Then, the angular acceleration is obtained by differentiating the difference in angular velocity between adjacent time windows with the window interval. This angular acceleration is used to drive the beam control unit, so that the antenna beam turning rate is dynamically adjusted according to the angular acceleration value, achieving smooth tracking synchronized with the UAV's turning actions and avoiding signal loss caused by beam jumps.

[0035] The spatiotemporal correlation feature tensor after compensation is collected in real time to verify the volatility of the Doppler spread coefficient or the stability of the spatial angle of arrival, generate the anti-fading effectiveness index, associate the anti-fading effectiveness index with the current intention mode identifier and feed it back to the anti-fading strategy library to update the corresponding scenario compensation parameters.

[0036] It should be noted that the spatiotemporal correlation feature tensor after compensation is collected, the Doppler expansion coefficient volatility (greater than or equal to 5% after compensation) or the angle of arrival stability (deviation greater than or equal to 2°) is calculated, and the anti-fading effectiveness index is generated (if the stability meets the standard, the index is 1, otherwise it is 0). After associating with the current intent identifier, it is fed back to the policy library.

[0037] Specifically, the radio frequency signal features optimized by the deep reinforcement learning decision unit are fused with cross-modal data from external radar sensors. High-confidence features are then selected through an attention mechanism, and a drone signal environment-adaptive tracking strategy is generated through distillation. This strategy is then output to distributed edge computing nodes. It receives the radio frequency signal feature vector output by the deep reinforcement learning decision unit, and simultaneously acquires the point cloud trajectory data provided by the external radar sensor; and matches the radio frequency feature timestamp with the radar trajectory timestamp to generate a spatiotemporally synchronized multimodal feature tensor. It should be noted that the system receives the radio frequency feature vector (including parameters such as frequency domain resolution and filter order, with a sampling rate of 1kHz) output by the deep reinforcement learning decision unit; it simultaneously acquires the point cloud trajectory data of the millimeter-wave radar (sampling rate of 100Hz), upsamples the radar data to 1kHz through bilinear interpolation, and then uses hardware timestamp alignment (PTP protocol) to generate a spatiotemporally synchronized multimodal feature tensor (dimension: time × spectral features × spatial coordinates).

[0038] Cross-domain confidence assessment is performed on the multimodal feature tensor: the time-frequency stability index of the radio frequency features is extracted as the spectrum confidence factor; the spatial continuity score of the radar point cloud trajectory is calculated as the trajectory confidence factor; the two confidence factors are input into the gated attention layer to generate a multimodal confidence attention vector. For the spectral confidence factor: calculate the standard deviation of the carrier frequency of the radio frequency signal within a sliding window (100ms) (e.g., 1 point for less than 50kHz, 0.5 points for 50-100kHz).

[0039] For the trajectory confidence factor: evaluate the continuity and cluster density of the radar point cloud (e.g., 1 point is awarded when the displacement difference between adjacent frames is less than 0.5m and the number of point clouds is greater than 50).

[0040] The two factors mentioned above are input into a gated attention layer (containing a sigmoid activation function), and the output attention vector is [0.8, 0.2], which indicates that radio frequency features are preferred.

[0041] The multimodal feature tensor is weighted and filtered based on the multimodal confidence attention vector: the first type of radio frequency feature subset with a spectral confidence factor higher than the first threshold is retained; the second type of radar trajectory feature subset with a trajectory confidence factor higher than the second threshold is selected; the two feature subsets are fused through feature crossover to generate a high-confidence fused feature vector; It should be noted that the multimodal feature tensor is masked: frequency band features with a spectral confidence greater than 0.7 (such as the 2.4-2.485GHz sub-band) and spatial regions with a trajectory confidence greater than 0.6 are preserved. For the selected high-confidence RF feature subsets and radar trajectory feature subsets, a cross-modal correlation matrix is ​​constructed using tensor outer product operations, and then a fused feature vector is generated through attention-weighted pooling. Specifically, the RF spectral feature vector (dimension M) and the radar spatial coordinate vector (dimension N) are subjected to a Kronecker product to obtain an M×N joint representation matrix. Then, a pre-trained attention weight matrix is ​​used to perform row-column bidirectional compression on this joint representation, finally outputting a fused feature vector of dimension K, where each element represents the energy distribution of the spectral-spatial coupling features.

[0042] The high-confidence fused feature vector is input into the policy distillation network: spatial-spectral joint feature patterns are extracted in the convolutional embedding layer; and a tracking policy parameter set containing beam pointing angle, scan rate, and frequency compensation is generated through the fully connected policy compilation layer. The tracking strategy parameter group is encoded into binary instructions, and after adding timestamps and location tags, an environment-adaptive tracking strategy instruction set is generated. This set is then distributed to distributed edge computing nodes via a low-latency communication protocol to perform beam reconfiguration operations.

[0043] It should be noted that the strategy parameters are encoded into 16-bit binary instructions (such as 0x5A3F representing an azimuth angle of 35.2°), and GPS timestamps (UTC format) and base station location tags (ECEF coordinates) are added. These instructions are then distributed to edge nodes through the TSN time-sensitive network to drive the phased array antenna to complete beam reconfiguration.

[0044] As can be seen, this invention achieves intelligent fusion of radio frequency signals and radar trajectories, thereby improving the accuracy and real-time performance of UAV tracking in complex environments, ensuring that the system can quickly adapt to environmental changes, and effectively solving the problems of insufficient multi-source data fusion and delayed decision-making in traditional methods.

[0045] Specifically, it aggregates local adjustment experience from multiple detection modules, including deep reinforcement learning decision unit parameter optimization experience, intent prediction and compensation unit prediction results, and extended tracking control unit fusion strategy. Through federated learning, it iteratively updates the global parameter optimization model, forming a swarm intelligence-driven spectrum detection self-evolution capability. Collect local tuning experience from each detection module, including parameter optimization action records of the deep reinforcement learning decision unit, fading prediction vector of the intent prediction and compensation unit, and high-confidence fusion feature vector of the extended tracking control unit, to generate a multi-source local experience dataset. The multi-source local experience dataset is timestamped and its feature dimensions are normalized. Common feature parameters are extracted, including frequency domain resolution adjustment, Doppler compensation value, and beam pointing angle deviation, to generate a standardized local experience matrix. The standardized local experience matrix is ​​input into the local parameter optimization model of each detection module, and the model parameter increment is calculated by gradient descent to generate an encrypted local gradient update vector. It should be noted that after inputting the standardized local experience matrix into the local parameter optimization model of each detection module, the current predicted output is calculated through model forward propagation. Then, based on the loss function (such as the mean squared loss of policy error), the backpropagation algorithm is used to calculate the update gradient of the model parameters. Specifically, for each trainable parameter in the model, its partial derivative is calculated to form a parameter gradient vector. To ensure data privacy, differential privacy encryption technology is used to process this gradient vector: first, the gradient values ​​are clipped, and then random noise that follows a Gaussian distribution is added to generate an encrypted local gradient update vector. This vector retains the validity of the parameter optimization direction while ensuring that the original training data cannot be restored.

[0046] It receives encrypted local gradient update vectors from multiple probe modules, performs weighted aggregation through a federated averaging algorithm, and generates a global gradient update vector, where the weights are dynamically allocated based on the amount of local empirical data of each module. It should be noted that the system receives encrypted gradient vectors uploaded by each detection module, dynamically assigns aggregation weights based on the proportion of data in each module's local experience dataset (e.g., if a module contributes 1000 data points, accounting for 20% of the total data, its weight is 0.2), and then performs a weighted summation of all encrypted gradient vectors (each vector is multiplied by its corresponding weight and then summed) to generate a global gradient update vector. In this process, the weight allocation is proportional to the amount of data, ensuring that nodes with larger data contributions have a greater impact on the global model, while encryption protects the privacy of the original data of each node.

[0047] Add Gaussian-distributed noise perturbation to the global gradient update vector to generate a privacy-preserving global gradient update vector, ensuring data security. The privacy-preserving global gradient update vector is distributed to each detection module to update the local parameter optimization model, forming a globally consistent spectrum detection strategy optimization model. Monitor the tracking accuracy and response latency of the updated model to generate model performance evaluation metrics, which are used to dynamically adjust the aggregation frequency and learning rate of federated learning.

[0048] In summary, this invention aggregates the local optimization experience of multiple detection modules through a federated learning framework, realizing the collaborative evolution and knowledge sharing of a distributed spectrum detection system. While ensuring the data privacy of each node, it uses weighted gradient aggregation to construct a global optimization model, enabling the system to continuously adapt to environmental changes. At the same time, a dynamic adjustment mechanism ensures a balance between model update efficiency and tracking performance, improving the accuracy and robustness of swarm intelligence decision-making.

[0049] In practical applications, this spectrum detection module may also include the following steps: The sudden gradient of the Doppler spread coefficient in the spatiotemporal correlation feature tensor generated by the dynamic fingerprint database is monitored in real time. When the signal-to-noise ratio decrease rate exceeds the environmental adaptation threshold within three consecutive decision cycles, a signal-to-noise ratio abnormal fluctuation vector is generated. The signal-to-noise ratio anomaly fluctuation vector is time-aligned with the parameter optimization action combination currently output by the deep Q network. The frequency domain resolution adjustment action and noise suppression threshold action are analyzed by correlation convolution kernel to analyze the suppression failure characteristics of the anomaly fluctuation and generate action failure identifier. Based on the action failure identifier, the network structure is adjusted: a dedicated interference suppression action branch is added to the output layer of the deep Q network. This branch includes three types of discrete actions: notch filter center frequency shift, adaptive stopband width expansion, and polarization anti-interference weight switching, forming an expanded action space. The signal-to-noise ratio abnormal fluctuation vector is input into the extended action space. The Q value of the newly added action is initialized through transfer learning. The interference suppression action strategy is iteratively optimized using the timing difference error. The anti-interference action combination command is output to the radio frequency front end for real-time reconstruction. The reconstructed spatiotemporal correlation feature tensor is collected, and the Doppler spread coefficient recovery rate and signal-to-noise ratio recovery slope are extracted to generate an anti-interference performance tensor and feed it back to the federated learning evolution unit to drive the global model to update the weight distribution of the interference suppression action strategy.

[0050] It should be noted that the system monitors the abrupt gradient of the Doppler expansion coefficient in the dynamic fingerprint database in real time (e.g., the gradient value increases from 0.2 to 1.5 within 1 second); when the signal-to-noise ratio (SNR) decrease rate exceeds the environmental adaptation threshold within 3 consecutive decision cycles (30ms) (e.g., a decrease of more than 5dB per cycle), an abnormal SNR fluctuation vector (including the decrease rate, frequency distribution, etc.) is generated; the abnormal vector is temporally aligned with the current action combination of the deep Q network, and the action failure features are analyzed through the correlation convolution kernel (e.g., if the SNR continues to decrease after the noise suppression threshold is adjusted to 40dB, an action failure identifier "01" is generated).

[0051] It should be noted that in environments with strong electromagnetic interference (such as sudden signals from nearby base stations), the deep reinforcement learning decision unit of traditional spectrum detection modules may fail. However, conventional parameter adjustments (such as frequency domain resolution or filter order) are usually insufficient to suppress abrupt drops in signal-to-noise ratio (SNR), leading to tracking interruptions in unknown interference scenarios. Therefore, in this embodiment, when abnormal SNR fluctuations are detected, the system dynamically expands the action space of the decision network, adds a dedicated interference suppression strategy, and rapidly generates the optimal anti-interference scheme through transfer learning and online optimization. This enables the RF front-end to reconstruct parameters in real time, effectively restoring signal quality, while simultaneously feeding optimization experience back to the global model, continuously improving the system's robustness in complex interference environments.

[0052] In practical applications, this spectrum detection module may also include the following steps: Tensor rank decomposition is performed on the spatiotemporal correlation feature tensor to separate the arrival angle fluctuation vector, Doppler spread coefficient sequence and trajectory continuity factor representing different signal sources, and an initial topological map with signal sources as nodes is constructed. The shared population separation strategy of the federated learning evolutionary unit is invoked to extract the spectral isolation weight and spatial conflict probability for the same multi-objective scenario, and mapped to the edge weight correction coefficient of the topological graph. The frequency isolation weights are multiplied by the Hadamard product with the angle of arrival fluctuation vector to generate a frequency domain isolation feature vector; at the same time, the spatial conflict probability and the trajectory continuity factor are convolved and fused to output a spatial domain scheduling priority matrix. Based on the frequency domain isolation feature vector, available frequency band resource blocks are divided, and the beam pointing sequence is sorted in combination with the spatial domain scheduling priority matrix to generate a joint resource allocation matrix containing frequency band-beam binding relationships; The signal-to-noise ratio balance and feature capture completeness of each signal source are monitored after allocation. When trajectory breakage caused by resource competition is detected, a topology reconstruction instruction is generated and fed back to the federated learning evolutionary unit to update the population separation strategy.

[0053] It should be noted that in complex electromagnetic environments, multi-UAV target tracking often faces the problems of frequency domain resource competition and airspace beam pointing conflict. Traditional methods rely on preset fixed allocation strategies, which are difficult to adapt to dynamic changes in signal sources, resulting in a persistently high trajectory breakage rate. Therefore, this embodiment achieves dynamic resource scheduling through the following steps: Tensor rank decomposition is performed using the spatiotemporal correlation feature tensor (containing features of mixed signal sources) generated by the dynamic fingerprint database. Specifically, it is decomposed into three physical layer feature vectors: (1) angle of arrival fluctuation vector (representing changes in signal orientation); (2) Doppler spread coefficient sequence (reflecting motion velocity characteristics); and (3) trajectory continuity factor (calculating the temporal integrity of the signal). A topological map is constructed with each independent signal source as a node to form a basic framework for resource allocation.

[0054] Next, the historical optimization data shared by the federated learning evolutionary units is used to extract key parameters of the swarm separation strategy for multi-objective scenarios: spectral isolation weight (an empirical value for optimizing frequency band spacing); and spatial conflict probability (a statistic of beam pointing collisions). These two parameters are mapped to weight correction coefficients of the connecting edges between nodes in the topology graph, enabling the graph to inherit the anti-conflict capability of swarm intelligence optimization.

[0055] Subsequently, the spectral isolation weights and the angle-of-arrival fluctuation vector are subjected to a Hadamard product operation to output a frequency domain isolation feature vector (quantifying the required frequency band isolation for each signal source). Simultaneously, the spatial conflict probability and the trajectory continuity factor are convolved and fused to generate a spatial scheduling priority matrix (determining the beam service order). The two are then jointly optimized using a resource conflict resolution operator, ultimately outputting a joint resource allocation matrix representing the frequency band-beam binding relationship.

[0056] Finally, the signal-to-noise ratio (SNR) balance (frequency domain resource competition index) and feature capture completeness (spatial domain tracking continuity index) of each signal source are monitored in real time after resource allocation. When a trajectory break event is detected (completeness is lower than the threshold and accompanied by a sudden change in SNR), a topology reconstruction instruction is generated and fed back to the federated learning evolutionary unit to drive the iterative update of the population separation strategy.

[0057] In summary, this embodiment can effectively eliminate signal trajectory breakage caused by resource contention, improve multi-target parallel tracking capability, and achieve system adaptive evolution.

[0058] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. An adaptive tracking unmanned aerial vehicle (UAV) spectrum detection module, characterized in that, It includes a dynamic fingerprint database generation unit, a deep reinforcement learning decision-making unit, an intent prediction and compensation unit, an extended tracking control unit, and a federated learning evolution unit; The dynamic fingerprint database generation unit receives UAV signals through a polarization reconfigurable broadband antenna array, and extracts and fuses the transient frequency offset characteristics, Doppler spread characteristics and spatial angle of arrival fluctuation parameters of the signal by combining time-frequency-space three-dimensional joint sampling to construct a dynamic signal fingerprint database that includes the time-varying characteristics and spatiotemporal correlation of the signal. The deep reinforcement learning decision unit utilizes a lightweight deep Q-network, with the dynamic signal fingerprint database as the state input, to play the optimal frequency domain resolution, noise suppression threshold, and adaptive filter order in real time, thereby achieving nonlinear dynamic optimization of the detection parameters. The intent prediction and compensation unit models the UAV's motion intent using a hidden Markov model, combines the spatiotemporal correlation of the dynamic signal fingerprint database, predicts the signal fading range and frequency offset direction, and triggers the anti-fading tracking mode and frequency domain compensation strategy in advance. The extended tracking control unit fuses the radio frequency signal features optimized by the deep reinforcement learning decision unit with cross-modal data from external radar sensors, filters high-confidence features through an attention mechanism, distills and generates an environment-adaptive tracking strategy for the UAV signal, and outputs it to the distributed edge computing nodes. The federated learning evolutionary unit aggregates local adjustment experience from multiple detection modules, including parameter optimization experience from deep reinforcement learning decision-making units, prediction results from intent prediction and compensation units, and fusion strategies from extended tracking control units. Through federated learning, it iterative updates of the global parameter optimization model form a swarm intelligence-driven spectrum detection self-evolution capability.

2. The adaptive tracking UAV spectrum detection module according to claim 1, characterized in that, By receiving UAV signals through a polarization-reconfigurable broadband antenna array and combining time-frequency-space three-dimensional joint sampling, transient frequency offset characteristics, Doppler spread characteristics, and spatial angle of arrival fluctuation parameters of the signal are extracted and fused to construct a dynamic signal fingerprint database containing the signal's time-varying characteristics and spatiotemporal correlation. Specifically: By synchronously receiving the raw radio frequency signals of the UAV through the horizontal and vertical polarization channels using a polarization-reconfigurable broadband antenna array, a dual-polarization time-domain signal stream is generated. The dual-polarized time-domain signal stream is subjected to time-frequency-space three-dimensional synchronous sampling: the signal stream is segmented with millisecond-level time windows to extract transient signal segments; short-time Fourier transform is performed on each transient signal segment to generate a time-frequency matrix; the spatial angle of arrival of the signal is calculated based on the phase difference of the antenna array, and the angle of arrival time sequence is output. In the time-frequency matrix, carrier frequency jump points are detected, and the frequency offset and offset rate between jump points are calculated to generate a transient frequency offset feature vector. The Doppler spread characteristics of the transient frequency offset feature vector are analyzed, and the Doppler spread coefficient of the frequency offset vector is determined by combining the signal propagation delay change rate of the time-domain slice, and the Doppler-frequency offset joint feature matrix is ​​output. Based on the Doppler-frequency offset joint feature matrix and the angle of arrival time series, the spatiotemporal correlation feature tensor is generated by tensor fusion analysis of the covariance relationship between the angle of arrival fluctuation and the Doppler spread coefficient. Tensor dimensionality reduction and feature decoupling are performed on the spatiotemporal correlation feature tensor to extract time-varying characteristic codes and spatial correlation codes. A dynamic fingerprint database of the signal environment is constructed by associating historical states through a rolling time window.

3. The adaptive tracking UAV spectrum detection module according to claim 1, characterized in that, By utilizing a lightweight deep Q-network and taking the dynamic signal fingerprint database as the state input, the optimal frequency domain resolution, noise suppression threshold, and adaptive filter order are dynamically optimized in real time to achieve nonlinear optimization of the detection parameters. Specifically: The time-varying characteristic code and spatial correlation code of the current moment are extracted from the dynamic signal fingerprint database, and an environmental state feature vector is generated after feature concatenation. The environmental state feature vector is aligned with the parameter adjustment action record of the previous decision cycle in time series, and the historical state transition relationship is fused through a gated loop unit to generate an enhanced state tensor. The enhanced state tensor is input into a lightweight deep Q-network, and action value function approximation is performed in the network hidden layer: the output layer generates three discrete action vectors: frequency domain resolution adjustment action, noise suppression threshold action, and filter order action; the action combination corresponding to the maximum Q value is selected through an ε-greedy strategy to optimize the parameter combination. Based on the parameter optimization action combination, the detection parameters are configured in real time: the frequency domain resolution adjustment action is mapped to the FFT point setting of the spectrum analyzer; the noise suppression threshold action is converted into the stopband attenuation depth of the digital filter; and the filter order action is associated with the tap coefficient update of the adaptive filter. The signal-to-noise ratio improvement rate and feature false detection rate after monitoring parameter adjustment are combined with the newly generated spatiotemporal correlation feature tensor in the dynamic fingerprint database to construct an environmental feedback tuple; The environmental feedback tuples are stored in the priority experience replay pool, and the sampling weights are calculated through temporal differential error. The weight parameters of the deep Q network are iteratively updated to form a closed-loop optimization mechanism.

4. The adaptive tracking UAV spectrum detection module according to claim 1, characterized in that, By modeling the drone's motion intent using a Hidden Markov Model and combining the spatiotemporal correlation of a dynamic signal fingerprint database, the signal fading interval and frequency offset direction are predicted, triggering anti-fading tracking mode and frequency domain compensation strategy in advance. Specifically: Obtain the spatiotemporal correlation feature tensor within the current rolling time window from the dynamic signal fingerprint database, and separate the time-varying characteristic encoding and spatial correlation encoding therein; The time-varying characteristics are encoded and input into the state layer of the Hidden Markov Model to analyze the signal Doppler frequency shift change law caused by the UAV motion and generate a state transition probability matrix containing the acceleration mutation probability. The spatial correlation encoding is input into the observation layer of the hidden Markov model, and combined with historical signal fading interval data, the statistical correlation between spatial angle of arrival fluctuation and frequency offset direction is calculated, and the observation probability matrix is ​​output. By fusing the state transition probability matrix and the observation probability matrix using the Viterbi algorithm, the UAV intention state sequence for the next three decision cycles is decoded and generated, including three types of motion intention labels: climb, dive, and sharp turn. Based on the UAV intention state sequence, and combined with the historical attenuation patterns of the Doppler-frequency offset joint feature matrix in the dynamic signal fingerprint database, the signal fading start time and duration interval are deduced, and a fading prediction vector is generated. Based on the fading prediction vector and the frequency offset direction label in the intention state sequence, the anti-fading tracking mode is activated in advance: when the prediction is a climb intention, Doppler fading margin compensation is initiated, and when the prediction is a sharp turn intention, spatial beamforming compensation strategy is activated.

5. The adaptive tracking UAV spectrum detection module according to claim 4, characterized in that, When an intention to climb is anticipated, Doppler fading margin compensation is initiated; when an intention to make a sharp turn is anticipated, a spatial beamforming compensation strategy is activated. Specifically: Based on the frequency offset direction label in the intent state sequence, extract the positive Doppler offset corresponding to the climbing intent or the negative Doppler offset corresponding to the sharp turn intent, and generate an intent mode identifier. When the intent mode identifier is a climb intent, the maximum fading depth record with the same positive Doppler offset in the historical Doppler-frequency offset joint feature matrix is ​​called to generate a Doppler margin compensation value; and a frequency compensation operation is triggered before the fading prediction start time point to increase the receiver local oscillator frequency by the Doppler margin compensation value. When the intent mode identifier is a sudden turn intent, the azimuth angle change rate is calculated based on the angle of arrival time sequence, and the beam steering angle acceleration is generated. It also drives the phase control unit of the antenna array to generate a spatial beam scanning trajectory according to the beam steering angle acceleration; The spatiotemporal correlation feature tensor after compensation is collected in real time to verify the volatility of the Doppler spread coefficient or the stability of the spatial angle of arrival, generate the anti-fading effectiveness index, associate the anti-fading effectiveness index with the current intention mode identifier and feed it back to the anti-fading strategy library to update the corresponding scenario compensation parameters.

6. The adaptive tracking UAV spectrum detection module according to claim 1, characterized in that, The radio frequency signal features optimized by the deep reinforcement learning decision unit are fused with cross-modal data from external radar sensors. High-confidence features are then selected through an attention mechanism, and the resulting distillation process generates an environment-adaptive tracking strategy for the UAV signal. This strategy is then output to distributed edge computing nodes. Specifically: It receives the radio frequency signal feature vector output by the deep reinforcement learning decision unit, and simultaneously acquires the point cloud trajectory data provided by the external radar sensor; and matches the radio frequency feature timestamp with the radar trajectory timestamp to generate a spatiotemporally synchronized multimodal feature tensor. Cross-domain confidence assessment is performed on the multimodal feature tensor: the time-frequency stability index of the radio frequency features is extracted as the spectral confidence factor; The spatial continuity score of the radar point cloud trajectory is calculated as the trajectory confidence factor; the two confidence factors are input into the gated attention layer to generate a multimodal confidence attention vector; The multimodal feature tensor is weighted and filtered based on the multimodal confidence attention vector: the first type of radio frequency feature subset with a spectral confidence factor higher than the first threshold is retained; the second type of radar trajectory feature subset with a trajectory confidence factor higher than the second threshold is selected. By fusing two feature subsets through feature crossover, a high-confidence fused feature vector is generated. The high-confidence fused feature vector is input into the policy distillation network: spatial-spectral joint feature patterns are extracted in the convolutional embedding layer; and a tracking policy parameter set containing beam pointing angle, scan rate, and frequency compensation is generated through the fully connected policy compilation layer. The tracking strategy parameter group is encoded into binary instructions, and after adding timestamps and location tags, an environment-adaptive tracking strategy instruction set is generated. This set is then distributed to distributed edge computing nodes via a low-latency communication protocol to perform beam reconfiguration operations.

7. The adaptive tracking UAV spectrum detection module according to claim 6, characterized in that: The radio frequency signal includes frequency domain resolution and filter order.

8. The adaptive tracking UAV spectrum detection module according to claim 1, characterized in that, By aggregating the local adjustment experience of multiple detection modules, including the parameter optimization experience of deep reinforcement learning decision-making units, the prediction results of intent prediction and compensation units, and the fusion strategy of extended tracking control units, and iteratively updating the global parameter optimization model through federated learning, a swarm intelligence-driven spectrum detection self-evolution capability is formed, specifically: Collect local tuning experience from each detection module, including parameter optimization action records of the deep reinforcement learning decision unit, fading prediction vector of the intent prediction and compensation unit, and high-confidence fusion feature vector of the extended tracking control unit, to generate a multi-source local experience dataset. The multi-source local experience dataset is timestamped and its feature dimensions are normalized. Common feature parameters are extracted, including frequency domain resolution adjustment, Doppler compensation value, and beam pointing angle deviation, to generate a standardized local experience matrix. The standardized local experience matrix is ​​input into the local parameter optimization model of each detection module, and the model parameter increment is calculated by gradient descent to generate an encrypted local gradient update vector. It receives encrypted local gradient update vectors from multiple probe modules, performs weighted aggregation through a federated averaging algorithm, and generates a global gradient update vector, where the weights are dynamically allocated based on the amount of local empirical data of each module. Add Gaussian-distributed noise perturbation to the global gradient update vector to generate a privacy-preserving global gradient update vector and ensure data security. The privacy-preserving global gradient update vector is distributed to each detection module to update the local parameter optimization model, forming a globally consistent spectrum detection strategy optimization model. Monitor the tracking accuracy and response latency of the updated model to generate model performance evaluation metrics, which are used to dynamically adjust the aggregation frequency and learning rate of federated learning.

Citation Information

Patent Citations

  • Low-power-consumption star flash control module applied to unmanned aerial vehicle and dummy pilot control method

    CN120183166A

  • Power load prediction method and system based on association rule analysis

    CN120280912A

  • Unmanned aerial vehicle full-band countering method and system based on acousto-optic-electric composite detection

    CN120320900A

  • Intelligent photoelectric theodolite aerial target positioning and tracking system

    CN120538494A

  • Constant-information ranging for dynamic spectrum access in a joint positioning-communications system

    US20220256496A1

Cited By

  • Unmanned aerial vehicle identification method and system based on linkage of radar and wireless signal

    CN122310240A