An adaptive tracking unmanned aerial vehicle spectrum detection module

By using an adaptive tracking UAV spectrum detection module, and leveraging technologies such as dynamic fingerprint database generation and deep reinforcement learning decision-making, the problem of signal acquisition and tracking of UAV spectrum detection in complex electromagnetic environments has been solved, achieving high-precision, real-time UAV signal tracking.

CN120934660BActive Publication Date: 2026-03-27SHENZHEN YANUOXUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing UAV spectrum detection technologies are ill-suited to handle dynamic changes in complex electromagnetic environments, have insufficient signal acquisition capabilities, lag in parameter adjustment, poor environmental adaptability, and centralized processing architectures cannot support swarm intelligence optimization under distributed deployment.

Method used

An adaptive tracking UAV spectrum detection module is adopted, including dynamic fingerprint database generation, deep reinforcement learning decision-making, intent prediction and compensation, extended tracking control and federated learning evolutionary unit. Through polarimetric reconfigurable broadband antenna array, lightweight deep Q network, hidden Markov model and cross-modal data fusion, dynamic acquisition and adaptive tracking of signals are achieved.

Benefits of technology

It achieves high-precision and high-reliability tracking capability for UAV signals, improves signal acquisition rate and tracking real-time performance in complex environments, and possesses self-evolving spectrum detection capability driven by swarm intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120934660B_ABST
    Figure CN120934660B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of unmanned aerial vehicle, especially to a self-adaptive tracking unmanned aerial vehicle spectrum detection module, comprising a dynamic fingerprint library generation unit, a deep reinforcement learning decision unit, an intention prediction and compensation unit, an extended tracking control unit and a federated learning evolution unit, through the intelligent spectrum detection architecture of multi-unit cooperation, the dynamic capture and self-adaptive tracking of unmanned aerial vehicle signals are realized, the dynamic fingerprint library accurately depicts the space-time characteristics of signals, the deep reinforcement learning optimizes the detection parameters in real time, the intention prediction unit actively compensates signal fading, cross-modal data fusion improves tracking robustness, and federated learning realizes group knowledge sharing and continuous evolution, so that the system has high-precision and high-reliable tracking capability in complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of unmanned aerial vehicles, and particularly relates to a self-adaptive tracking unmanned aerial vehicle spectrum detection module. BACKGROUND

[0002] With the rapid development and wide application of unmanned aerial vehicles, the unmanned aerial vehicles play an increasingly important role in the fields of security, logistics, surveying and mapping, etc., but also bring security risks such as illegal intrusion and privacy leakage. Traditional unmanned aerial vehicle spectrum detection technology mainly relies on fixed frequency band scanning and static parameter configuration, and is difficult to cope with the dynamic changes of unmanned aerial vehicle signals in a complex electromagnetic environment. The existing technology has the following limitations: (1) insufficient signal acquisition capability, narrowband antenna and fixed sampling rate cannot adapt to new communication protocols such as frequency hopping and spread spectrum; (2) parameter adjustment lag, rule-based control strategy is difficult to track the rapid frequency offset and Doppler effect caused by unmanned aerial vehicle maneuvering in real time; (3) poor environmental adaptability, single sensor data lacks the ability to cooperatively compensate signal fading and multipath interference. In addition, the existing system mostly adopts a centralized processing architecture and cannot support group intelligence optimization under distributed deployment. Although some research attempts to introduce machine learning algorithms, key problems such as dynamic signal fingerprint modeling, cross-modal data fusion and collaborative learning under privacy protection have not been solved. Therefore, it is urgent to develop a self-adaptive spectrum detection module with environmental perception, intention prediction and group evolution capability to improve the unmanned aerial vehicle monitoring efficiency in complex scenarios. SUMMARY

[0003] The application overcomes the deficiencies of the prior art and provides a self-adaptive tracking unmanned aerial vehicle spectrum detection module.

[0004] To achieve the above-mentioned purpose, the technical scheme adopted by the application is as follows:

[0005] The application discloses a self-adaptive tracking unmanned aerial vehicle spectrum detection module, which comprises a dynamic fingerprint library generation unit, a deep reinforcement learning decision unit, an intention prediction and compensation unit, an extended tracking control unit and a federated learning evolution unit.

[0006] The dynamic fingerprint library generation unit receives unmanned aerial vehicle signals through a polarization reconfigurable wideband antenna array, combines time-frequency-space three-dimensional joint sampling, extracts and fuses the transient frequency offset features, Doppler spread characteristics and spatial arrival angle fluctuation parameters of the signals, and constructs a dynamic signal fingerprint library containing signal time-varying characteristics and space-time correlation.

[0007] The deep reinforcement learning decision unit uses a lightweight deep Q network, takes the dynamic signal fingerprint library as the state input, and realizes real-time game of optimal frequency domain resolution, noise suppression threshold and adaptive filter order to achieve nonlinear dynamic optimization of detection parameters.

[0008] The intention prediction and compensation unit models the unmanned aerial vehicle motion intention by a hidden Markov model, predicts the signal fading interval and frequency offset direction in combination with the space-time correlation of the dynamic signal fingerprint library, and triggers the anti-fading tracking mode and frequency domain compensation strategy in advance;

[0009] The extended tracking control unit fuses the radio frequency signal features optimized by the deep reinforcement learning decision unit and the cross-modal data of the external radar sensor, filters high-confidence features through an attention mechanism, distills an environment-adaptive tracking strategy for the unmanned aerial vehicle signal, and outputs to a distributed edge computing node;

[0010] The federated learning evolution unit aggregates the local adjustment experience of multiple detection modules including the parameter optimization experience of the deep reinforcement learning decision unit, the prediction results of the intention prediction and compensation unit, and the fusion strategy of the extended tracking control unit, iteratively updates the global parameter optimization model through federated learning, and forms a spectrum detection self-evolution ability driven by swarm intelligence.

[0011] The polarized reconfigurable wideband antenna array receives the unmanned aerial vehicle signal, extracts and fuses the transient frequency offset features, Doppler spread characteristics, and spatial angle of arrival fluctuation parameters of the signal in combination with time-frequency-space three-dimensional joint sampling, constructs a dynamic signal fingerprint library containing signal time-varying characteristics and space-time correlation, and specifically includes:

[0012] The polarized reconfigurable wideband antenna array synchronously receives the original radio frequency signals of the unmanned aerial vehicle in horizontal and vertical polarization channels to generate a dual-polarized time-domain signal stream;

[0013] Time-frequency-space three-dimensional synchronous sampling is performed on the dual-polarized time-domain signal stream: the signal stream is divided into millisecond-level time windows to extract transient signal segments; short-time Fourier transform is performed on each transient signal segment to generate a time-frequency matrix; and the spatial angle of arrival is calculated based on the phase difference of the antenna array, and an angle of arrival time sequence is output;

[0014] The carrier frequency hopping points are detected in the time-frequency matrix, the frequency offset and offset rate between the hopping points are calculated, and a transient frequency offset feature vector is generated;

[0015] The Doppler spread characteristics of the transient frequency offset feature vector are analyzed, the Doppler spread coefficient of the frequency offset vector is determined in combination with the signal propagation time delay change rate of the time domain slice, and a Doppler-frequency offset joint feature matrix is output;

[0016] Based on the Doppler-frequency offset joint feature matrix and the angle of arrival time sequence, the covariance relationship between the angle of arrival fluctuation and the Doppler spread coefficient is analyzed through tensor fusion to generate a space-time correlation feature tensor;

[0017] The spatiotemporal correlation feature tensor is subjected to tensor dimension reduction and feature decoupling to extract time-varying characteristic codes and spatial correlation codes, and a signal environment dynamic fingerprint library is constructed by correlating historical states through a rolling time window.

[0018] The dynamic signal fingerprint library is used as a state input to a lightweight deep Q network to realize nonlinear dynamic optimization of detection parameters by real-time gaming of optimal frequency domain resolution, noise suppression threshold and adaptive filter order, specifically as follows:

[0019] The time-varying characteristic codes and spatial correlation codes at the current time are extracted from the dynamic signal fingerprint library, and an environment state feature vector is generated after feature splicing;

[0020] The environment state feature vector is time-series aligned with the parameter adjustment action record of the previous decision period, and a history state transition relationship is fused through a gated recurrent unit to generate an enhanced state tensor;

[0021] The enhanced state tensor is input into the lightweight deep Q network, and an action value function approximation is performed in the network hidden layer: three discrete action vectors of frequency domain resolution adjustment action, noise suppression threshold action and filter order action are generated at the output layer; and a parameter optimization action combination corresponding to the maximum Q value is selected through an ε-greedy strategy;

[0022] The parameter optimization action combination is used to perform real-time configuration of detection parameters: the frequency domain resolution adjustment action is mapped to the FFT point setting of a spectrum analyzer; the noise suppression threshold action is converted into the stopband attenuation depth of a digital filter; and the filter order action is associated with the tap coefficient update of an adaptive filter;

[0023] The signal SNR improvement rate and feature false detection rate changes after parameter adjustment are monitored, and an environment feedback tuple is constructed in combination with the newly generated spatiotemporal correlation feature tensor in the dynamic fingerprint library;

[0024] The environment feedback tuple is stored in a priority experience replay pool, a sampling weight is calculated through a time difference error, the weight parameters of the deep Q network are iteratively updated, and a closed-loop optimization mechanism is formed.

[0025] The dynamic signal fingerprint library is used as a state input to a lightweight deep Q network to realize nonlinear dynamic optimization of detection parameters by real-time gaming of optimal frequency domain resolution, noise suppression threshold and adaptive filter order, specifically as follows:

[0026] The spatiotemporal correlation feature tensor within the current rolling time window is obtained from the dynamic signal fingerprint library, and the time-varying characteristic codes and spatial correlation codes are separated therefrom;

[0027] The time-varying characteristics are encoded into a hidden Markov model state layer, the signal Doppler frequency shift change rule caused by the unmanned aerial vehicle movement is analyzed, and a state transition probability matrix containing acceleration mutation probability is generated;

[0028] The spatial correlation is encoded into a hidden Markov model observation layer, the statistical correlation between the spatial angle of arrival fluctuation and the frequency offset direction is calculated by combining historical signal fading interval data, and an observation probability matrix is output;

[0029] The state transition probability matrix and the observation probability matrix are fused by a Viterbi algorithm to decode and generate an unmanned aerial vehicle intention state sequence in the next three decision periods, including three types of motion intention labels: climbing, diving and sharp turning;

[0030] Based on the unmanned aerial vehicle intention state sequence, the history fading mode of the Doppler-frequency offset combined feature matrix in the dynamic signal fingerprint library is combined to deduce the signal fading starting time point and the duration interval, and a fading prediction vector is generated;

[0031] According to the fading prediction vector and the frequency offset direction label in the intention state sequence, the anti-fading tracking mode is activated in advance: when the climbing intention is predicted, the Doppler fading allowance compensation is started, and when the sharp turning intention is predicted, the spatial beamforming compensation strategy is enabled.

[0032] When the climbing intention is predicted, the Doppler fading allowance compensation is started, and when the sharp turning intention is predicted, the spatial beamforming compensation strategy is enabled, specifically:

[0033] According to the frequency offset direction label in the intention state sequence, the positive Doppler offset corresponding to the climbing intention or the negative Doppler offset corresponding to the sharp turning intention is extracted, and an intention mode identifier is generated;

[0034] When the intention mode identifier is the climbing intention, the maximum fading depth record of the same positive Doppler offset in the historical Doppler-frequency offset combined feature matrix is called to generate a Doppler allowance compensation value; and a frequency compensation operation is triggered before the fading prediction starting time point (such as the previous 200 milliseconds), and the receiver local oscillator frequency is raised by the Doppler allowance compensation value;

[0035] When the intention mode identifier is the sharp turning intention, the azimuth angle mutation rate is calculated based on the angle of arrival time sequence to generate a beam steering angle acceleration; and the phase control unit of the antenna array is driven to generate a spatial beam scanning trajectory according to the beam steering angle acceleration;

[0036] The compensated space-time correlation feature tensor is collected in real time, the Doppler spread coefficient fluctuation rate or the spatial angle of arrival stability is verified, an anti-fading performance index is generated, the anti-fading performance index is associated with the current intention mode identifier and fed back to the anti-fading strategy library to update the compensation parameters of the corresponding scene.

[0037] wherein the optimized radio frequency signal features of the deep reinforcement learning decision unit are fused with the cross-modal data of the external radar sensor, high-confidence features are screened through an attention mechanism, an environment-adaptive tracking strategy for the unmanned aerial vehicle signal is distilled, and the output is distributed to a distributed edge computing node, specifically:

[0038] The radio frequency signal feature vector output by the deep reinforcement learning decision unit is received, and point cloud trajectory data provided by the external radar sensor is synchronously acquired; and the radio frequency feature timestamp is matched with the radar trajectory timestamp to generate a time-space synchronous multi-modal feature tensor; the radio frequency signal includes frequency domain resolution and filter order;

[0039] Cross-domain confidence evaluation is performed on the multi-modal feature tensor: the time-frequency stability index of the radio frequency feature is extracted as a spectral confidence factor; the spatial continuity score of the radar point cloud trajectory is calculated as a trajectory confidence factor; the two confidence factors are input into a gated attention layer to generate a multi-modal confidence attention vector;

[0040] The multi-modal feature tensor is weighted and screened based on the multi-modal confidence attention vector: a first subset of radio frequency features with a spectral confidence factor higher than a first threshold is retained; a second subset of radar trajectory features with a trajectory confidence factor higher than a second threshold is selected; the two feature subsets are fused through feature crossing to generate a high-confidence fusion feature vector;

[0041] The high-confidence fusion feature vector is input into a strategy distillation network: a spatial-spectral joint feature pattern is extracted in a convolutional embedding layer; a tracking strategy parameter group including a beam pointing angle, a scanning rate, and a frequency compensation amount is generated through a fully connected strategy compilation layer;

[0042] The tracking strategy parameter group is binary instruction encoded, and an environment-adaptive tracking strategy instruction set is generated after adding a timestamp and a position label, which is distributed to a distributed edge computing node to perform a beam reconstruction operation through a low-latency communication protocol.

[0043] wherein the aggregation includes multi-probe module local adjustment experience including deep reinforcement learning decision unit parameter optimization experience, intention prediction and compensation unit pre-judgment result, and extended tracking control unit fusion strategy, and the global parameter optimization model is iteratively updated through federated learning to form a group intelligence driven spectral detection self-evolution ability, specifically:

[0044] Local adjustment experience of each probe module is collected, including parameter optimization action record of the deep reinforcement learning decision unit, fading pre-judgment vector of the intention prediction and compensation unit, and high-confidence fusion feature vector of the extended tracking control unit, to generate a multi-source local experience dataset;

[0045] timestamp alignment and feature dimension normalization are performed on the multi-source local experience data set, common feature parameters including frequency domain resolution adjustment amount, Doppler compensation value and beam pointing angle deviation are extracted, and a standardized local experience matrix is generated;

[0046] The standardized local experience matrix is input into a local parameter optimization model of each detection module, a gradient descent calculation model parameter increment is calculated, and an encrypted local gradient update vector is generated;

[0047] The encrypted local gradient update vector from the plurality of detection modules is received, and a global gradient update vector is generated by weighted aggregation through a federated average algorithm, wherein the weights are dynamically allocated according to the local experience data amount of each module;

[0048] Noise disturbance conforming to a Gaussian distribution is added in the global gradient update vector, a privacy-protected global gradient update vector is generated, and data security is ensured;

[0049] The privacy-protected global gradient update vector is distributed to each detection module, the local parameter optimization model is updated, and a globally consistent spectrum detection strategy optimization model is formed;

[0050] The tracking accuracy and response delay of the updated model are monitored, a model performance evaluation index is generated, and the aggregation frequency and learning rate of the federated learning are dynamically adjusted.

[0051] The present application solves the technical defects in the background art, and has the following beneficial effects: through the intelligent spectrum detection architecture of multiple units cooperation, dynamic capture and adaptive tracking of unmanned aerial vehicle signals are realized, the dynamic fingerprint library accurately depicts the space-time characteristics of the signals, the deep reinforcement learning optimizes the detection parameters in real time, the intention prediction unit actively compensates for signal fading, cross-modal data fusion improves the tracking robustness, and the federated learning realizes group knowledge sharing and continuous evolution, so that the system has high-precision and high-reliable tracking capability in complex environments. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings of embodiments according to these drawings without creative labor.

[0053] Figure 1 The system framework diagram of the adaptive tracking unmanned aerial vehicle spectrum detection module is shown in the figure;

[0054] Figure 2 The working flowchart of the adaptive tracking unmanned aerial vehicle spectrum detection module is shown in the figure. DETAILED DESCRIPTION

[0055] In order to enable a more clear understanding of the above-mentioned objects, features and advantages of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0056] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present application, however, the present application can also be implemented in other manners different from those described herein, and therefore, the protection scope of the present application is not limited by the specific embodiments disclosed below.

[0057] As shown in Figure 1 , 2 The present application discloses a kind of adaptive tracking unmanned aerial vehicle spectrum detection module, including dynamic fingerprint library generation unit, deep reinforcement learning decision unit, intention prediction and compensation unit, extension tracking control unit and federal learning evolution unit;

[0058] S1, the dynamic fingerprint library generation unit receives unmanned aerial vehicle signal by polarization reconfigurable wideband antenna array, extracts and fuses the transient frequency offset characteristics, doppler spread characteristics and spatial angle of arrival fluctuation parameters of signal in combination with time-frequency-space three-dimensional joint sampling, constructs dynamic signal fingerprint library including signal time-varying characteristics and space-time correlation;

[0059] S2, the deep reinforcement learning decision unit utilizes lightweight deep Q network, with the dynamic signal fingerprint library as state input, real-time game optimal frequency domain resolution, noise suppression threshold and adaptive filter order, realizes the nonlinear dynamic optimization of detection parameter;

[0060] S3, the intention prediction and compensation unit models unmanned aerial vehicle movement intention by hidden Markov model, in combination with the space-time correlation of dynamic signal fingerprint library, judges signal fading interval and frequency offset direction in advance, triggers anti-fading tracking mode and frequency domain compensation strategy in advance;

[0061] S4, the extension tracking control unit fuses the radio frequency signal characteristics optimized by deep reinforcement learning decision unit and the cross-modal data of external radar sensor, selects high confidence features by attention mechanism, distills environmental adaptive tracking strategy of unmanned aerial vehicle signal, and outputs to distributed edge computing node;

[0062] S5, the federal learning evolution unit aggregates the local adjustment experience of multiple detection modules including deep reinforcement learning decision unit parameter optimization experience, intention prediction and compensation unit prediction result, extension tracking control unit fusion strategy, iteratively updates global parameter optimization model by federal learning, forms spectrum detection self-evolution ability driven by swarm intelligence.

[0063] Wherein, the unmanned aerial vehicle signal is received by the polarization reconfigurable broadband antenna array, combined with time-frequency-space three-dimensional joint sampling, the transient frequency offset characteristics, Doppler spread characteristics and spatial arrival angle fluctuation parameters of the signal are extracted and fused, a dynamic signal fingerprint library containing time-varying characteristics and space-time correlation of the signal is constructed, specifically:

[0064] The original radio frequency signals of the unmanned aerial vehicle in horizontal and vertical polarization channels are synchronously received by the polarization reconfigurable broadband antenna array to generate a dual-polarized time-domain signal stream;

[0065] Time-frequency-space three-dimensional synchronous sampling is performed on the dual-polarized time-domain signal stream: the signal stream is divided into transient signal segments with millisecond-level time windows, and short-time Fourier transform is performed on each transient signal segment to generate a time-frequency matrix; the spatial arrival angle of the signal is calculated based on the phase difference of the antenna array, and an arrival angle time sequence is outputted;

[0066] The carrier frequency hopping points are detected in the time-frequency matrix, the frequency offset and the offset rate between the hopping points are calculated, and a transient frequency offset feature vector is generated;

[0067] The Doppler spread characteristics of the transient frequency offset feature vector are analyzed, the Doppler spread coefficient of the frequency offset vector is determined combined with the signal propagation time delay change rate of the time domain slice, and a Doppler-frequency offset joint feature matrix is outputted;

[0068] Based on the Doppler-frequency offset joint feature matrix and the arrival angle time sequence, the covariance relationship between the arrival angle fluctuation and the Doppler spread coefficient is analyzed by tensor fusion to generate a space-time correlation feature tensor;

[0069] The space-time correlation feature tensor is subjected to tensor dimension reduction and feature decoupling to extract time-varying characteristic codes and spatial correlation codes, and a signal environment dynamic fingerprint library is constructed by associating the historical state through a rolling time window.

[0070] It should be noted that, due to the fact that most of the unmanned aerial vehicles adopt frequency hopping, spread spectrum and other technologies during flight, and are affected by factors such as Doppler effect and environmental shielding, the signal frequency, polarization state and spatial angle of arrival will change rapidly. The application adopts a polarization reconfigurable wideband antenna array (such as a 4x4 UCA array) to synchronously receive the horizontal (H) and vertical (V) polarization channels of the unmanned aerial vehicle radio frequency signals, and generate a dual-polarized time-domain signal stream. The antenna supports dynamic switching of the polarization mode to adapt to the communication polarization characteristics (such as linear / circular polarization) of different unmanned aerial vehicles. The dual-polarized signal stream is segmented in a 10ms time window, and transient signal segments (such as each segment containing 1024 sampling points) are extracted. The short-time Fourier transform is performed on each transient segment to generate a time-frequency matrix (time-frequency-energy distribution), and the carrier frequency hopping points are detected. The spatial angle of arrival is calculated using the phase difference of the antenna array, and the angle of arrival time sequence (such as 100 sampling points per second) is output. In the time-frequency matrix, the carrier frequency hopping points (such as the frequency point mutation caused by frequency hopping communication) are detected, the frequency offset (Δf) and the offset rate (df / dt) of adjacent hopping points are calculated, and a transient frequency offset feature vector is generated. For example, if a certain unmanned aerial vehicle signal jumps from 2.4GHz to 2.42GHz within 50ms, then Δf=20MHz and df / dt=0.4MHz / ms.

[0071] Combined with the signal propagation time delay change rate of the time domain slice (such as based on the correlation peak displacement calculation), the Doppler spread characteristics (such as the frequency spread caused by high-speed motion) in the frequency offset vector are analyzed, the Doppler spread coefficient is determined, and a Doppler-frequency offset joint feature matrix (dimension: time x frequency x Doppler coefficient) is output. The Doppler-frequency offset joint feature matrix is fused with the angle of arrival time sequence (such as using Tucker decomposition), the covariance relationship between the angle of arrival fluctuation and the Doppler spread coefficient is analyzed, and a space-time correlation feature tensor (dimension: time x frequency x space x Doppler) is generated. The space-time correlation feature tensor is subjected to tensor dimension reduction (PCA) and feature decoupling, time-varying characteristic codes (such as LSTM time sequence features) and spatial correlation codes (such as graph convolution features) are extracted, and the historical state is associated through a rolling time window (such as a 30s window) to construct a dynamically updated signal fingerprint library.

[0072] In summary, the application realizes real-time construction of a dynamic signal fingerprint library through a polarization reconfigurable antenna array and a time-frequency-space three-dimensional joint sampling, accurately captures the transient frequency offset, Doppler characteristics and spatial fluctuations of the unmanned aerial vehicle signal, and thus improves the signal capture rate and tracking real-time performance in complex environments.

[0073] Among them, the light-weight deep Q network is used, and the dynamic signal fingerprint library is used as the state input to realize real-time game of the optimal frequency domain resolution, noise suppression threshold and adaptive filter order, and nonlinear dynamic optimization of the detection parameters, specifically:

[0074] extracting time-varying characteristic encoding and spatial correlation encoding of the current moment from the dynamic signal fingerprint library, performing feature splicing to generate an environment state feature vector;

[0075] It should be noted that the time-varying characteristic encoding (such as Doppler shift trend) and the spatial correlation encoding (such as angle of arrival fluctuation) of the current moment are extracted from the dynamic signal fingerprint library, and the two are spliced to generate an environment state feature vector.

[0076] The environment state feature vector is time-series aligned with the parameter adjustment action record of the last decision period, and the historical state transition relationship is fused through a gated recurrent unit to generate an enhanced state tensor;

[0077] It should be noted that the current state feature is time-aligned with the parameter adjustment action (such as the last FFT point setting) of the last decision period, and is input into the gated recurrent unit to fuse the historical state transition relationship and output the enhanced state tensor. The hidden layer of the gated recurrent unit can remember long-term dependencies, such as gradual frequency offset caused by continuous left turn of the unmanned aerial vehicle.

[0078] The enhanced state tensor is input into a lightweight deep Q network, and an action value function approximation is performed in the network hidden layer: the output layer generates three discrete action vectors of frequency domain resolution adjustment action, noise suppression threshold action and filter order action; the parameter optimization action combination corresponding to the maximum Q value is selected through the ε-greedy strategy;

[0079] It should be noted that the enhanced state tensor is input into a lightweight DQN (such as a 4-layer fully connected network, with parameter quantity controlled within 1MB), and the output layer generates three discrete action vectors: frequency domain resolution adjustment action (such as FFT point number selectable 256 / 512 / 1024); noise suppression threshold action (such as stopband attenuation depth selectable 20dB / 30dB / 40dB); filter order action (such as adaptive filter tap number selectable 16 / 32 / 64). Then, the action combination with the maximum Q value is selected through the ε-greedy strategy (ε=0.1) to balance exploration and utilization.

[0080] According to the parameter optimization action combination, the detection parameter real-time configuration is performed: the frequency domain resolution adjustment action is mapped to the FFT point number setting of the spectrum analyzer; the noise suppression threshold action is converted into the stopband attenuation depth of the digital filter; and the filter order action is associated with the tap coefficient update of the adaptive filter;

[0081] The signal-to-noise ratio improvement rate and the feature false detection rate change after parameter adjustment are monitored, and the environment feedback tuple is constructed in combination with the newly generated space-time correlation feature tensor in the dynamic fingerprint library.

[0082] It should be noted that the FFT point number is mapped to the spectrum analyzer, for example, a 512-point FFT is selected to balance the frequency resolution and the calculation delay. The noise suppression threshold is converted into the stopband attenuation depth of the digital filter (such as 30 dB), and the adjacent frequency interference is suppressed. The tap coefficient of the adaptive filter is updated according to the filter order, and the signal separation effect is optimized.

[0083] Then, the adjusted signal-to-noise ratio improvement rate and the feature false detection rate are monitored, and a new space-time correlation feature tensor of the dynamic fingerprint library is combined to construct a feedback tuple (state, action, reward, new state).

[0084] The environment feedback tuple is stored in the priority experience replay pool, the sampling weight is calculated by the time difference error, the weight parameters of the deep Q network are iteratively updated, and a closed-loop optimization mechanism is formed.

[0085] It should be noted that the feedback tuple is stored in the priority experience replay pool, and the sampling weight is allocated according to the time difference error (TD-error). For example, the sample priority of the frequency offset mutation scene is higher. The DQN weight is iteratively updated by minimizing the Bellman error to form an adaptive optimization closed loop.

[0086] In summary, through this step, the unmanned aerial vehicle can meet the real-time tracking demand in the rapid maneuvering scene, effectively improve the signal-to-noise ratio and reduce the bit error rate, and through the closed-loop feedback mechanism, the system can still maintain stable tracking under complex electromagnetic interference.

[0087] Among them, the motion intention of the unmanned aerial vehicle is modeled by a hidden Markov model, the space-time correlation of the dynamic signal fingerprint library is combined, the signal fading interval and the frequency offset direction are predicted, and the anti-fading tracking mode and the frequency domain compensation strategy are triggered in advance. Specifically:

[0088] The space-time correlation feature tensor in the current rolling time window is obtained from the dynamic signal fingerprint library, and the time-varying characteristic code and the spatial correlation code are separated therefrom;

[0089] It should be noted that the space-time correlation feature tensor in the current rolling time window (such as 500 ms) is extracted from the dynamic signal fingerprint library, and the time-varying characteristic code (including Doppler shift trend, signal attenuation rate, etc.) and the spatial correlation code (including angle of arrival fluctuation mode, polarization state change, etc.) are separated by feature decoupling.

[0090] The time-varying characteristic code is input into the state layer of the hidden Markov model, the signal Doppler shift variation law caused by the motion of the unmanned aerial vehicle is analyzed, and a state transition probability matrix containing acceleration mutation probability is generated;

[0091] The spatial correlation coding is input into a hidden Markov model observation layer, historical signal fading interval data is combined, statistical correlation of spatial angle of arrival fluctuation and frequency offset direction is calculated, and an observation probability matrix is output;

[0092] The state transition probability matrix and the observation probability matrix are fused by a Viterbi algorithm to decode a UAV intention state sequence of future three decision periods, including three types of motion intention labels of climbing, diving and sharp turning;

[0093] Based on the UAV intention state sequence, a signal fading starting time point and a duration interval are deduced by combining a historical attenuation mode of a Doppler-frequency offset joint feature matrix in a dynamic signal fingerprint library, and a fading prediction vector is generated;

[0094] It should be noted that the Viterbi algorithm is used to decode the optimal state sequence, and an intention prediction result of future three decision periods (such as 300 ms) is output. Meanwhile, by combining a historical attenuation mode library (recording signal fading characteristics in different motion states), an expected fading starting time and duration are calculated.

[0095] According to the fading prediction vector and the frequency offset direction label in the intention state sequence, an anti-fading tracking mode is activated in advance: when the climbing intention is predicted, the Doppler fading residual compensation is started, and when the sharp turning intention is predicted, the spatial beamforming compensation strategy is enabled.

[0096] It should be noted that when the UAV performs a climbing, diving or sharp turning maneuver, the communication signal will produce a rapid frequency offset due to the Doppler effect, and will be accompanied by a sharp fading of the signal strength. The existing compensation mechanism usually uses a fixed threshold, which cannot predict the motion intention of the UAV, resulting in a "slow half beat" of the compensation action, which seriously affects the continuity and stability of tracking. Especially in a complex electromagnetic environment, multipath effects will further exacerbate signal fading, making it difficult to maintain reliable UAV tracking. In view of this, the present application intelligently predicts the UAV motion intention by a hidden Markov model, analyzes the space-time features of the signal fingerprint library, realizes the early prediction of the signal fading interval and the frequency offset direction, and can actively trigger the targeted anti-fading compensation strategy, thereby improving the continuity and stability of the UAV signal tracking, effectively overcoming the lagging problem of the traditional passive compensation method, and enabling the system to still maintain reliable signal acquisition capability when the UAV rapidly maneuvers.

[0097] When the climbing intention is predicted, the Doppler fading residual compensation is started, and when the sharp turning intention is predicted, the spatial beamforming compensation strategy is enabled, specifically:

[0098] According to the frequency offset direction label in the intention state sequence, a positive Doppler offset corresponding to the climbing intention or a negative Doppler offset corresponding to the sharp turning intention is extracted, and an intention mode identifier is generated;

[0099] When the intention mode identifier is a climb intention, the maximum fading depth record of the same positive Doppler offset in the historical Doppler-frequency offset joint feature matrix is called to generate a Doppler residual compensation value; and a frequency compensation operation is triggered before the fading prediction start time point (such as the previous 200 milliseconds) to up-regulate the receiver local oscillator frequency by the Doppler residual compensation value;

[0100] It should be noted that when the intention identifier is a climb (the frequency offset direction label is “positive”), the maximum fading depth record under the same positive Doppler offset (for example, the maximum fading depth of a certain model is 12 dB at a 20° climb angle) is called from the historical library, and the receiver local oscillator frequency is up-regulated by a compensation value (such as up-regulation +45 kHz) 200 ms before the predicted fading start time point to offset the Doppler frequency shift.

[0101] When the intention mode identifier is a sharp turn intention, the azimuth angle mutation rate is calculated based on the time sequence of the arrival angle to generate a beam steering angle acceleration; and the phase control unit of the antenna array is driven to generate a spatial beam scanning trajectory according to the beam steering angle acceleration;

[0102] It should be noted that the linear fitting of the continuous arrival angle sample values is performed through a sliding time window (such as 100 ms), the ratio of the azimuth angle change amount to the time interval in the window is calculated to obtain the instantaneous angular velocity, and the derivative of the angular velocity difference of adjacent time windows and the window interval is obtained to obtain the angular acceleration. The angular acceleration is used to drive the beam control unit to dynamically adjust the antenna beam steering rate according to the angular acceleration value, realize smooth tracking synchronized with the unmanned aerial vehicle steering action, and avoid signal loss caused by beam jumping.

[0103] The compensated spatiotemporal correlation feature tensor is collected in real time to verify the Doppler spread coefficient fluctuation rate or the spatial arrival angle stability to generate an anti-fading performance index, and the anti-fading performance index is associated with the current intention mode identifier and fed back to the anti-fading strategy library to update the compensation parameters of the corresponding scene.

[0104] It should be noted that the compensated spatiotemporal correlation feature tensor is collected, the Doppler spread coefficient fluctuation rate (greater than or equal to 5% after compensation) or the arrival angle stability (deviation greater than or equal to 2°) is calculated, the anti-fading performance index (such as a stability compliance index of 1, otherwise 0) is generated, and the anti-fading performance index is fed back to the strategy library after being associated with the current intention identifier.

[0105] The radio frequency signal features optimized by the deep reinforcement learning decision unit are fused with the cross-modal data of the external radar sensor, high-confidence features are selected through an attention mechanism, an environment-adaptive tracking strategy for the unmanned aerial vehicle signal is distilled and generated, and the environment-adaptive tracking strategy is output to a distributed edge computing node, specifically:

[0106] receive the radio frequency signal feature vector output by the deep reinforcement learning decision unit, synchronously acquire point cloud trajectory data provided by an external radar sensor, and match the radio frequency feature timestamp with the radar trajectory timestamp to generate a time-space synchronized multi-modal feature tensor;

[0107] It should be noted that the radio frequency feature vector output by the deep reinforcement learning decision unit (including frequency domain resolution, filter order, etc., sampling rate 1 kHz) is received; the point cloud trajectory data of the millimeter wave radar (sampling rate 100 Hz) is synchronously acquired, the radar data is upsampled to 1 kHz through bilinear interpolation, and then a time-space synchronized multi-modal feature tensor (dimension: time x spectral feature x spatial coordinate) is generated by hardware timestamp alignment (PTP protocol).

[0108] Cross-domain confidence evaluation is performed on the multi-modal feature tensor: the time-frequency stability index of the radio frequency feature is extracted as a spectral confidence factor; the spatial continuity score of the radar point cloud trajectory is calculated as a trajectory confidence factor; the two confidence factors are input into a gated attention layer to generate a multi-modal confidence attention vector;

[0109] For the spectral confidence factor: the standard deviation of the carrier frequency of the radio frequency signal in a sliding window (100 ms) is calculated (e.g., 1 point for less than 50 kHz, 0.5 points for 50-100 kHz).

[0110] For the trajectory confidence factor: the continuity and clustering density of the radar point cloud are evaluated (e.g., 1 point for a displacement difference between adjacent frames less than 0.5 m and a point cloud number greater than 50).

[0111] The above two factors are input into a gated attention layer (including a sigmoid activation function) to output an attention vector (e.g., [0.8, 0.2] indicates that the radio frequency feature is preferentially used).

[0112] Based on the multi-modal confidence attention vector, the multi-modal feature tensor is weighted and selected: a first type of radio frequency feature subset with a spectral confidence factor higher than a first threshold is retained; a second type of radar trajectory feature subset with a trajectory confidence factor higher than a second threshold is selected; the two feature subsets are fused through feature cross to generate a high-confidence fusion feature vector;

[0113] It should be noted that the multi-modal feature tensor is masked: the frequency band features with a spectral confidence greater than 0.7 (such as 2.4-2.485GHz sub-band) and the spatial regions with a trajectory confidence greater than 0.6 are retained. For the screened high-confidence radio frequency feature subset and radar trajectory feature subset, a tensor outer product operation is used to construct a cross-modal association matrix, and then an attention weighted pooling is used to generate a fusion feature vector. Specifically, the radio frequency spectrum feature vector (dimension M) and the radar spatial coordinate vector (dimension N) are subjected to a Kronecker product operation to obtain an MxN joint representation matrix, and then a pre-trained attention weight matrix is used to compress the joint representation in a row-column bidirectional manner, and finally a fusion feature vector with a dimension of K is output, wherein each dimension element represents the energy distribution of the spectrum-space coupling feature.

[0114] The high-confidence fusion feature vector is input into a strategy distillation network: a convolutional embedding layer extracts spatial-spectrum joint feature patterns; a fully connected strategy compilation layer generates a tracking strategy parameter group containing beam pointing angle, scanning rate, and frequency compensation;

[0115] The tracking strategy parameter group is binary instruction coded, and an environment adaptive tracking strategy instruction set is generated after adding a timestamp and a position label, which is distributed to distributed edge computing nodes through a low latency communication protocol to perform beam reconstruction operations.

[0116] It should be noted that the strategy parameters are encoded into 16-bit binary instructions (such as 0x5A3F representing an azimuth angle of 35.2°), a GPS timestamp (UTC format) and a base station position label (ECEF coordinates) are added, and the TSN time sensitive network is distributed to the edge nodes to drive the phased array antenna to complete the beam reconstruction.

[0117] As can be seen, the present application realizes intelligent fusion of radio frequency signals and radar trajectories, thereby improving the accuracy and real-time performance of unmanned aerial vehicle tracking in complex environments, ensuring that the system quickly adapts to environmental changes, and effectively solving the problem of insufficient multi-source data fusion and decision lag in traditional methods.

[0118] Among them, the aggregation includes the local adjustment experience of the multi-probe module, including the parameter optimization experience of the deep reinforcement learning decision unit, the pre-judgment result of the intention prediction and compensation unit, and the fusion strategy of the extended tracking control unit. The global parameter optimization model is iteratively updated through federated learning to form a group intelligence driven spectrum detection self-evolution ability, specifically:

[0119] Collect the local adjustment experience of each probe module, including the parameter optimization action record of the deep reinforcement learning decision unit, the fading prediction vector of the intention prediction and compensation unit, and the high-confidence fusion feature vector of the extended tracking control unit, to generate a multi-source local experience dataset;

[0120] timestamp alignment and feature dimension normalization are performed on the multi-source local experience dataset, common feature parameters including frequency domain resolution adjustment amount, Doppler compensation value and beam pointing angle deviation are extracted, and a standardized local experience matrix is generated;

[0121] The standardized local experience matrix is input into the local parameter optimization model of each detection module, and the gradient descent calculation model parameter increment is generated to generate an encrypted local gradient update vector;

[0122] It should be noted that after the standardized local experience matrix is input into the local parameter optimization model of each detection module, the current prediction output is calculated by model forward propagation, and then the update gradient of the model parameter is calculated based on the loss function (such as the mean square loss of policy error) using the back propagation algorithm. Specifically, for each trainable parameter in the model, the partial derivative of the parameter is calculated to form a parameter gradient vector. In order to protect data privacy, differential privacy encryption technology is used to process the gradient vector: first, the gradient value is clipped, and then random noise satisfying Gaussian distribution is added to generate an encrypted local gradient update vector. The vector not only retains the effectiveness of the parameter optimization direction, but also ensures that the original training data cannot be restored.

[0123] The encrypted local gradient update vectors from the plurality of detection modules are received, and a global gradient update vector is generated by weighted aggregation through a federated averaging algorithm, wherein the weights are dynamically allocated according to the local experience data amount of each module;

[0124] It should be noted that the encrypted gradient vectors uploaded by each detection module are received, the aggregation weights are dynamically allocated according to the data amount proportion of the local experience dataset of each module (for example, a module contributes 1000 data, accounting for 20% of the total data amount, and the weight is 0.2), and then the weighted sum of all encrypted gradient vectors (each vector is multiplied by the corresponding weight and then added) is calculated to generate a global gradient update vector. In this process, the weight allocation is proportional to the data amount, ensuring that the nodes with more data contribution have a greater impact on the global model, while the encryption processing ensures the privacy of the original data of each node.

[0125] Gaussian distribution noise disturbance is added to the global gradient update vector to generate a privacy-protected global gradient update vector, ensuring data security;

[0126] The privacy-protected global gradient update vector is distributed to each detection module to update the local parameter optimization model, forming a globally consistent spectrum detection strategy optimization model;

[0127] The tracking accuracy and response delay of the updated model are monitored to generate model performance evaluation indicators, which are used to dynamically adjust the aggregation frequency and learning rate of federated learning.

[0128] In summary, the present application aggregates the local optimization experience of multiple probe modules through the federated learning framework, realizes the cooperative evolution and knowledge sharing of the distributed spectrum probe system, and under the premise of ensuring the data privacy of each node, uses the weighted gradient aggregation to construct a global optimization model, so that the system can continuously adapt to environmental changes, while the dynamic adjustment mechanism ensures the balance between model update efficiency and tracking performance, and improves the accuracy and robustness of group intelligence decision-making.

[0129] The spectrum probe module in the actual application process can further include the following steps:

[0130] Real-time monitoring of the mutation gradient of the Doppler spread coefficient in the spatiotemporal correlation feature tensor generated by the dynamic fingerprint library, when the signal-to-noise ratio drop rate is detected to exceed the environmental adaptive threshold in three consecutive decision cycles, an abnormal fluctuation vector of signal-to-noise ratio is generated;

[0131] The abnormal fluctuation vector of signal-to-noise ratio and the parameter optimization action combination output by the current deep Q network are time-aligned, the correlation convolution kernel is used to analyze the suppression failure features of the abnormal fluctuation of the frequency domain resolution adjustment action and the noise suppression threshold action, and an action failure identifier is generated;

[0132] According to the action failure identifier, the network structure is adjusted: a disturbance suppression special action branch is added to the output layer of the deep Q network, the branch includes three types of discrete actions: trap filter center frequency point offset, adaptive stopband width expansion and polarization anti-interference weight switching, forming an expanded action space;

[0133] The signal-to-noise ratio abnormal fluctuation vector is input into the expanded action space, the Q value of the new action is initialized by transfer learning, the interference suppression action strategy is iteratively optimized by using the time difference error, and the anti-interference action combination instruction is output to the radio frequency front end for real-time reconstruction;

[0134] The spatiotemporal correlation feature tensor after reconstruction is collected, the Doppler spread coefficient recovery rate and the signal-to-noise ratio recovery slope are extracted, an anti-interference performance tensor is generated and fed back to the federated learning evolution unit to drive the global model to update the weight distribution of the interference suppression action strategy.

[0135] It should be noted that the mutation gradient of the Doppler spread coefficient in the dynamic fingerprint library is monitored in real time (such as the gradient value increasing from 0.2 to 1.5 within 1 second); when the signal-to-noise ratio drop rate exceeds the environmental adaptive threshold (such as a drop of more than 5dB per cycle) in three consecutive decision cycles (30ms), an abnormal fluctuation vector of signal-to-noise ratio (including drop rate, frequency point distribution, etc.) is generated; the abnormal vector and the current action combination of the deep Q network are time-aligned, and the action failure features are analyzed by the correlation convolution kernel (for example, if the signal-to-noise ratio continues to drop after the noise suppression threshold is adjusted to 40dB, an action failure identifier "01" is generated).

[0136] It should be noted that in a strong electromagnetic interference environment (such as a sudden signal of a nearby base station), the deep reinforcement learning decision unit of the traditional spectrum detection module may fail. However, regular parameter adjustment (such as frequency domain resolution or filter order) usually cannot suppress the sudden drop in abnormal signal-to-noise ratio, resulting in tracking interruption of the system in unknown interference scenarios. In view of this, in the embodiment, when abnormal fluctuation of signal-to-noise ratio is detected, the system dynamically expands the action space of the decision network, adds special interference suppression strategies, quickly generates the optimal anti-interference scheme through transfer learning and online optimization, enables the radio frequency front end to reconstruct parameters in real time, effectively restores signal quality, and at the same time feeds back the optimization experience to the global model to continuously improve the robustness of the system in complex interference environments.

[0137] The spectrum detection module can further include the following steps in actual application:

[0138] The spatiotemporal correlation feature tensor is subjected to tensor rank decomposition to separate out the angle of arrival fluctuation vector, Doppler spread coefficient sequence and trajectory continuity factor representing different signal sources, and an initial topology graph with signal sources as nodes is constructed;

[0139] The spectral isolation degree weight and spatial conflict probability for similar multi-target scenarios are extracted from the group separation strategy shared by the federated learning evolution unit, and are mapped to the edge weight correction coefficient of the topology graph;

[0140] The spectral isolation degree weight and the angle of arrival fluctuation vector are subjected to Hadamard product operation to generate a frequency domain isolation feature vector; at the same time, the spatial conflict probability and the trajectory continuity factor are subjected to convolution fusion to output a spatial scheduling priority matrix;

[0141] The available frequency band resource blocks are divided according to the frequency domain isolation feature vector, the beam pointing sequence is sorted in combination with the spatial scheduling priority matrix, and a joint resource allocation matrix containing frequency band-beam binding relationship is generated;

[0142] The signal-to-noise ratio balance and feature capture completeness of each signal source after allocation are monitored, and when a trajectory break caused by resource competition is detected, a topology reconstruction instruction is generated and fed back to the federated learning evolution unit to update the group separation strategy.

[0143] It should be noted that in a complex electromagnetic environment, multi-unmanned aerial vehicle target tracking usually faces the problems of frequency domain resource competition and spatial beam pointing conflict. Traditional methods rely on preset fixed allocation strategies, which are difficult to adapt to the dynamic changes of signal sources, resulting in a high trajectory break rate. Therefore, the embodiment realizes dynamic resource scheduling through the following steps:

[0144] Tensor rank decomposition is performed using the spatiotemporal correlation feature tensor (containing features of mixed signal sources) generated by the dynamic fingerprint database. Specifically, it is decomposed into three physical layer feature vectors: (1) angle of arrival fluctuation vector (representing changes in signal orientation); (2) Doppler spread coefficient sequence (reflecting motion velocity characteristics); and (3) trajectory continuity factor (calculating the temporal integrity of the signal). A topological map is constructed with each independent signal source as a node to form a basic framework for resource allocation.

[0145] Next, the historical optimization data shared by the federated learning evolutionary units is used to extract key parameters of the swarm separation strategy for multi-objective scenarios: spectral isolation weight (an empirical value for optimizing frequency band spacing); and spatial conflict probability (a statistic of beam pointing collisions). These two parameters are mapped to weight correction coefficients of the connecting edges between nodes in the topology graph, enabling the graph to inherit the anti-conflict capability of swarm intelligence optimization.

[0146] Subsequently, the spectral isolation weights and the angle-of-arrival fluctuation vector are subjected to a Hadamard product operation to output a frequency domain isolation feature vector (quantifying the required frequency band isolation for each signal source). Simultaneously, the spatial conflict probability and the trajectory continuity factor are convolved and fused to generate a spatial scheduling priority matrix (determining the beam service order). The two are then jointly optimized using a resource conflict resolution operator, ultimately outputting a joint resource allocation matrix representing the frequency band-beam binding relationship.

[0147] Finally, the signal-to-noise ratio (SNR) balance (frequency domain resource competition index) and feature capture completeness (spatial domain tracking continuity index) of each signal source are monitored in real time after resource allocation. When a trajectory break event is detected (completeness is lower than the threshold and accompanied by a sudden change in SNR), a topology reconstruction instruction is generated and fed back to the federated learning evolutionary unit to drive the iterative update of the population separation strategy.

[0148] In summary, this embodiment can effectively eliminate signal trajectory breakage caused by resource contention, improve multi-target parallel tracking capability, and achieve system adaptive evolution.

[0149] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. An adaptive tracking drone spectrum detection module, comprising: The method comprises a dynamic fingerprint library generation unit, a deep reinforcement learning decision unit, an intention prediction and compensation unit, an extended tracking control unit, and a federated learning evolution unit. The dynamic fingerprint library generation unit receives a UAV signal through a polarized reconfigurable broadband antenna array, extracts and fuses the transient frequency offset feature, Doppler spread characteristic, and spatial angle of arrival fluctuation parameter of the signal, and constructs a dynamic signal fingerprint library containing the time-varying characteristics and space-time correlation of the signal. The deep reinforcement learning decision unit uses a lightweight deep Q network, takes the dynamic signal fingerprint library as the state input, and realizes real-time game optimization of the optimal frequency domain resolution, noise suppression threshold, and adaptive filter order to achieve nonlinear dynamic optimization of the detection parameters. The intention prediction and compensation unit models the UAV movement intention through a hidden Markov model, predicts the signal fading interval and frequency offset direction based on the space-time correlation of the dynamic signal fingerprint library, and triggers the anti-fading tracking mode and frequency domain compensation strategy in advance. The extended tracking control unit fuses the radio frequency signal features optimized by the deep reinforcement learning decision unit and the cross-modal data of external radar sensors, selects high-confidence features through an attention mechanism, distills the environment-adaptive tracking strategy of the UAV signal, and outputs to a distributed edge computing node. The federated learning evolution unit aggregates the local adjustment experience of multiple detection modules including the parameter optimization experience of the deep reinforcement learning decision unit, the prediction results of the intention prediction and compensation unit, and the fusion strategy of the extended tracking control unit, iteratively updates the global parameter optimization model through federated learning, and forms a spectrum detection self-evolution ability driven by swarm intelligence.

2. The self-adapting tracking UAV spectrum detection module of claim 1, wherein, The dynamic fingerprint library generation unit receives a UAV signal through a polarized reconfigurable broadband antenna array, extracts and fuses the transient frequency offset feature, Doppler spread characteristic, and spatial angle of arrival fluctuation parameter of the signal, and constructs a dynamic signal fingerprint library containing the time-varying characteristics and space-time correlation of the signal. The polarized reconfigurable broadband antenna array synchronously receives the horizontal and vertical polarization channel original radio frequency signals of the UAV to generate a dual-polarization time domain signal stream. The time-frequency-space three-dimensional synchronous sampling is performed on the dual-polarization time domain signal stream: the signal stream is divided into transient signal segments with a millisecond-level time window; the short-time Fourier transform is performed on each transient signal segment to generate a time-frequency matrix; and the signal spatial angle of arrival is calculated based on the antenna array phase difference to output an angle of arrival time sequence. The carrier frequency hopping points are detected in the time-frequency matrix, the frequency offset and offset rate between the hopping points are calculated, and a transient frequency offset feature vector is generated. The Doppler spread characteristic analysis is performed on the transient frequency offset feature vector, the Doppler spread coefficient of the frequency offset vector is determined based on the signal propagation time delay change rate of the time domain slice, and a Doppler-frequency offset joint feature matrix is output. Based on the Doppler-frequency offset joint feature matrix and the angle of arrival time sequence, the covariance relationship between the angle of arrival fluctuation and the Doppler spread coefficient is analyzed through tensor fusion to generate a space-time correlation feature tensor. The spatiotemporal correlation feature tensor is subjected to tensor dimension reduction and feature decoupling to extract time-varying characteristic codes and spatial correlation codes, and a dynamic signal fingerprint library is constructed by correlating historical states through a rolling time window.

3. The self-adapting tracking UAV spectrum detection module of claim 1, wherein, A lightweight deep Q network is used to take the dynamic signal fingerprint library as state input to perform real-time game optimization of optimal frequency domain resolution, noise suppression threshold and adaptive filter order, thereby realizing nonlinear dynamic optimization of detection parameters, specifically as follows: The time-varying characteristic codes and spatial correlation codes of the current time are extracted from the dynamic signal fingerprint library, and an environment state feature vector is generated after feature splicing; The environment state feature vector is time series aligned with the parameter adjustment action record of the last decision cycle, the historical state transition relationship is fused through a gated recurrent unit to generate an enhanced state tensor; The enhanced state tensor is input into the lightweight deep Q network, an action value function approximation is performed in the network hidden layer, three discrete action vectors of frequency domain resolution adjustment action, noise suppression threshold action and filter order action are generated in the output layer, the parameter optimization action combination corresponding to the maximum Q value is selected through an ε-greedy strategy; Real-time configuration of detection parameters is performed according to the parameter optimization action combination: the frequency domain resolution adjustment action is mapped to the FFT point setting of a spectrum analyzer, the noise suppression threshold action is converted into the stopband attenuation depth of a digital filter, and the filter order action is associated with the tap coefficient update of an adaptive filter; The signal SNR improvement rate and feature false detection rate changes after parameter adjustment are monitored, and an environment feedback tuple is constructed in combination with the newly generated spatiotemporal correlation feature tensor in the dynamic fingerprint library; The environment feedback tuple is stored in a priority experience replay pool, the weight parameters of the deep Q network are iteratively updated through time difference error calculation and sampling weight, and a closed-loop optimization mechanism is formed.

4. The self-adapting tracking UAV spectrum detection module of claim 1, wherein, The motion intention of the UAV is modeled through a hidden Markov model, the spatiotemporal correlation of the dynamic signal fingerprint library is combined, the signal fading interval and frequency offset direction are predicted, and the anti-fading tracking mode and frequency domain compensation strategy are triggered in advance, specifically as follows: The spatiotemporal correlation feature tensor within the current rolling time window is obtained from the dynamic signal fingerprint library, and the time-varying characteristic codes and spatial correlation codes are separated therefrom; The time-varying characteristic codes are input into the state layer of the hidden Markov model to analyze the signal Doppler shift variation law caused by the motion of the UAV and generate a state transition probability matrix containing acceleration mutation probability; The spatial correlation codes are input into the observation layer of the hidden Markov model to calculate the statistical correlation between the spatial angle of arrival fluctuation and the frequency offset direction in combination with historical signal fading interval data, and an observation probability matrix is output; The state transition probability matrix and the observation probability matrix are fused through the Viterbi algorithm to decode and generate a UAV intention state sequence of the next three decision cycles, including three types of motion intention labels of climbing, diving and turning; Based on the UAV intention state sequence, the fading start time point and duration are deduced in combination with the historical fading mode of the Doppler-frequency offset joint feature matrix in the dynamic signal fingerprint library, and a fading prediction vector is generated. According to the fading prediction vector and the frequency offset direction label in the intention state sequence, an anti-fading tracking mode is activated in advance: when the intention is predicted to be climbing, Doppler fading residual compensation is started, and when the intention is predicted to be sharp turning, a spatial beamforming compensation strategy is enabled.

5. The self-adapting tracking UAV spectrum detection module of claim 4, wherein, When the intention is predicted to be climbing, Doppler fading residual compensation is started, and when the intention is predicted to be sharp turning, a spatial beamforming compensation strategy is enabled, specifically: According to the frequency offset direction label in the intention state sequence, a positive Doppler offset corresponding to the climbing intention or a negative Doppler offset corresponding to the sharp turning intention is extracted, and an intention mode identifier is generated; When the intention mode identifier is the climbing intention, the maximum fading depth record of the same positive Doppler offset in the historical Doppler-frequency joint feature matrix is called to generate a Doppler residual compensation value; and a frequency compensation operation is triggered before the fading prediction start time point to raise the receiver local oscillator frequency by the Doppler residual compensation value; When the intention mode identifier is the sharp turning intention, the azimuth angle mutation rate is calculated based on the angle of arrival time sequence to generate a beam steering angle acceleration; And drive the phase control unit of the antenna array to generate a spatial beam scanning trajectory according to the beam steering angle acceleration; Real-time collection of the compensated space-time correlation feature tensor verifies the Doppler spread coefficient fluctuation rate or the spatial angle of arrival stability to generate an anti-fading performance index, and the anti-fading performance index is associated with the current intention mode identifier and fed back to the anti-fading strategy library to update the compensation parameters for the corresponding scene.

6. The self-adapting tracking UAV spectrum detection module of claim 1, wherein, The radio frequency signal features optimized by the deep reinforcement learning decision unit are fused with the cross-modal data of the external radar sensor, high-confidence features are selected through an attention mechanism, an environment-adaptive tracking strategy for the unmanned aerial vehicle signal is distilled, and output to the distributed edge computing node, specifically: Receive the radio frequency signal feature vector output by the deep reinforcement learning decision unit, synchronously acquire the point cloud trajectory data provided by the external radar sensor; and match the radio frequency feature timestamp with the radar trajectory timestamp to generate a space-time synchronous multi-modal feature tensor; Cross-domain confidence evaluation is performed on the multi-modal feature tensor: the time-frequency stability index of the radio frequency feature is extracted as a spectral confidence factor; The spatial continuity score of the radar point cloud trajectory is calculated as a trajectory confidence factor; the two confidence factors are input into a gated attention layer to generate a multi-modal confidence attention vector; Based on the multi-modal confidence attention vector, the multi-modal feature tensor is weighted and selected: a first type of radio frequency feature subset with a spectral confidence factor higher than a first threshold is retained; a second type of radar trajectory feature subset with a trajectory confidence factor higher than a second threshold is selected; Feature fusion is performed on the two feature subsets through feature crossing to generate a high-confidence fusion feature vector; The high-confidence fusion feature vector is input into a strategy distillation network: a space-spectrum joint feature mode is extracted in a convolution embedding layer; a tracking strategy parameter group including a beam pointing angle, a scanning rate, and a frequency compensation amount is generated through a fully connected strategy compilation layer; The tracking strategy parameter set is binary instruction coded, and an environment adaptive tracking strategy instruction set is generated after adding a timestamp and a position label, and is distributed to distributed edge computing nodes through a low latency communication protocol to perform beam reconstruction operations.

7. The self-adapting tracking UAV spectrum detection module of claim 6, wherein: The radio frequency signal includes a frequency domain resolution and a filter order.

8. The self-adapting tracking UAV spectrum detection module of claim 1, wherein, The aggregation includes multi-probe module local adjustment experience, including deep reinforcement learning decision unit parameter optimization experience, intention prediction and compensation unit pre-judgment result, and extended tracking control unit fusion strategy, iteratively updates the global parameter optimization model through federated learning, and forms a group intelligence driven spectrum detection self-evolution ability, specifically: Collecting local adjustment experience of each probe module, including parameter optimization action record of deep reinforcement learning decision unit, fading prediction vector of intention prediction and compensation unit, and high confidence fusion feature vector of extended tracking control unit, generating a multi-source local experience dataset; Timestamp alignment and feature dimension normalization are performed on the multi-source local experience dataset to extract common feature parameters, including frequency domain resolution adjustment amount, Doppler compensation value, and beam pointing angle deviation, to generate a standardized local experience matrix; The standardized local experience matrix is input into the local parameter optimization model of each probe module, and the model parameter increment is calculated through gradient descent to generate an encrypted local gradient update vector; Receiving encrypted local gradient update vectors from multiple probe modules, performing weighted aggregation through federated averaging algorithm to generate a global gradient update vector, wherein the weights are dynamically allocated according to the local experience data amount of each module; Adding noise disturbance conforming to Gaussian distribution in the global gradient update vector to generate a privacy protected global gradient update vector to ensure data security; Distribute the privacy protected global gradient update vector to each probe module to update the local parameter optimization model to form a globally consistent spectrum detection strategy optimization model; Monitoring the tracking accuracy and response delay of the updated model to generate model performance evaluation indicators for dynamically adjusting the aggregation frequency and learning rate of federated learning.

Citation Information

Patent Citations

  • Low-power-consumption star flash control module applied to unmanned aerial vehicle and dummy pilot control method

    CN120183166A

  • Power load prediction method and system based on association rule analysis

    CN120280912A