Low-altitude aircraft identification method and system based on multi-source data fusion

By combining a passive sensor network and a physical model-constrained gated recurrent unit (PI-GRU) prediction model with a reinforcement learning decision model, the problem of long-range, highly concealed identification and resource scheduling of small UAVs in low-altitude security protection is solved, achieving efficient and low-power threat prediction and identification.

CN121561685APending Publication Date: 2026-02-24HONGKE WANGAN (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511408946.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies are insufficient for all-weather, long-range, and highly concealed identification of small, low-speed, and quiet drones in low-altitude security protection. They also lack the ability to predict threats and the flexibility of resource allocation, resulting in low identification efficiency and insufficient concealment.

Method used

By acquiring micro-Doppler feature time series through passive sensor networks, calculating the behavioral anomaly index using a physically constrained gated recurrent unit (PI-GRU) prediction model, updating the target threat level in real time, and optimizing sensor resource scheduling through a reinforcement learning decision model, a balance between dynamism, stealth, and proactivity is achieved.

Benefits of technology

It achieves efficient, low-power, and long-range identification of low-altitude aircraft, enhances the ability to predict threats and the flexibility of resource scheduling, ensures that the system can perform precise operations at the optimal time, and resolves the inherent contradiction between stealth and proactivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121561685A_ABST
    Figure CN121561685A_ABST
Patent Text Reader

Abstract

The invention provides a low-altitude aircraft identification method and system based on multi-source data fusion, relates to the technical field of electrical digital data processing, and aims to solve the technical problems of how to realize pre-judgment identification of potential threats on the premise of keeping long-term strategic concealment of the system and how to realize pre-judgment identification of potential threats according to the cost effectiveness of dynamic change. Sensor resources are optimally scheduled to achieve the highest efficiency of tactical confirmation; specifically, the problem is effectively solved by constructing an intelligent identification method which takes passive behavior pre-judgment as input, takes comprehensive cost optimization as a target and takes reinforcement learning as a decision-making core. The early warning capability of abnormal behaviors is remarkably improved, and dynamic optimal balance of sensor resources among information gain, energy consumption and exposure risk is realized, so that the intelligent level, task cruising ability and overall survivability of the system in a complex confrontation environment are greatly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, specifically to a method and system for identifying low-altitude aircraft based on multi-source data fusion. Background Technology

[0002] In the field of low-altitude security protection for critical infrastructure (such as airports, nuclear power plants, and border areas), there is a significant challenge in the all-weather, long-range, and highly stealthy identification of small, low-speed, and silent drones. Existing technologies either rely on active radar, which easily exposes their own position and consumes a lot of energy; or on photoelectric / acoustic sensors, which have limited range and anti-interference capabilities; or use simple multi-source data fusion in the later stages, resulting in severe information loss and an inability to deal with deceptive, abruptly behaving "gray" targets. Therefore, the market urgently needs an intelligent identification method that can maintain a high degree of stealth, predictively identify potential threats, and confirm them in the most efficient way to solve the technical challenges of long-term, low-power, and high-precision monitoring in unattended scenarios.

[0003] To achieve effective monitoring of low-altitude aircraft, existing technologies have explored various approaches, but still face the following challenges:

[0004] The inherent contradiction between covert detection and active confirmation: Traditional low-altitude defense systems suffer from a disconnect between the detection and identification phases. On the one hand, relying on active radar and other equipment for wide-area searches, while offering long detection ranges, the electromagnetic signals actively emitted by these devices easily reveal their location, limiting their applicability in scenarios requiring radio silence. On the other hand, while passive sensors such as photoelectric and acoustic sensors offer good concealment, their effective range, resistance to weather interference, and nighttime performance are all limited, making it difficult to independently and reliably complete long-range identification tasks.

[0005] The lag in data fusion and the lack of threat prediction capabilities: Existing technologies typically employ post-fusion or decision-level fusion strategies when processing multi-source data. For example, a method disclosed in Chinese patent CN116821823A performs decision fusion of state attributes only after all sensor data collection is completed. Such methods are essentially a "post-confirmation" model, that is, comprehensively interpreting the collected data. However, this model often suffers from time lag in identifying and responding to "gray" targets with deceptive behavioral patterns who only reveal their intentions at critical moments, lacking the ability to capture and predict the "pre-ignition" stage of threat intent.

[0006] The static nature of resource scheduling and the rigidity of decision-making logic: When facing multiple targets, existing systems' sensor resource scheduling strategies largely rely on fixed rules or preset priority thresholds. This static decision-making logic struggles to adapt to dynamically changing battlefield environments, the system's own resource status (such as battery power), and varying mission requirements for stealth. The result is often suboptimal resource allocation, missed optimal detection windows due to rigid adherence to rules, or wasted energy and exposure risks due to unnecessary detection actions. The overall operational efficiency and survivability of the system require further improvement.

[0007] The information disclosed in the background section above is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0008] The purpose of this invention is to provide a method and system for identifying low-altitude aircraft based on multi-source data fusion, so as to solve the problems mentioned in the background art.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] The low-altitude aircraft identification method based on multi-source data fusion includes the following steps:

[0011] S1: Obtain the time series of micro-Doppler features formed by the target aircraft reflecting signals from non-cooperative opportunistic illumination sources through a passive sensor network;

[0012] S2: Input the micro-Doppler feature time series into the preset behavior baseline prediction model, and calculate the behavior anomaly index, which characterizes the degree of current behavior anomaly of the target aircraft, based on the deviation between the prediction result of the preset behavior baseline prediction model and the measured value of the micro-Doppler feature time series.

[0013] S3: Based on the behavioral anomaly index, update the target threat level of the target aircraft in real time;

[0014] S4: Based on the updated target threat level, generate an active sensor resource scheduling strategy for the target aircraft using a preset decision model;

[0015] S5: Execute the active sensor resource scheduling strategy to prioritize the detection and identification of the target aircraft.

[0016] A low-altitude aircraft identification system based on multi-source data fusion, the system being used to execute the aforementioned low-altitude aircraft identification method based on multi-source data fusion, comprising:

[0017] Passive sensing and feature extraction module: used to acquire the time series of micro-Doppler features formed by the target aircraft reflecting signals from non-cooperative opportunistic illumination sources through a passive sensor network;

[0018] Behavior anomaly prediction module: used to input the micro-Doppler feature time series into a preset behavior baseline prediction model, and calculate the behavior anomaly index, which characterizes the degree of current behavior anomaly of the target aircraft, based on the deviation between the prediction result of the preset behavior baseline prediction model and the measured value of the micro-Doppler feature time series.

[0019] Intelligent decision-making and resource scheduling module: used to update the target threat level of the target aircraft in real time based on the behavioral anomaly index;

[0020] Based on the updated target threat level, an active sensor resource scheduling strategy for the target aircraft is generated using a preset decision model.

[0021] Hierarchical active confirmation module: used to execute the active sensor resource scheduling strategy to prioritize the detection and identification of the target aircraft.

[0022] Compared with existing technologies, the beneficial effects of this invention are: by using a completely passive perception network and an innovative behavioral baseline prediction model, it achieves deep prediction of the target's potential intentions, directly addressing the shortcomings of existing technologies in threat prediction capabilities. Furthermore, this invention uses this prediction result (i.e., the target threat level) as the core state, constructing the entire sensor scheduling problem as a reinforcement learning decision-making process aimed at minimizing the "overall task cost." Here, the "overall task cost" innovatively integrates resource scheduling costs and exposure risk costs, thus quantifying and embedding the strategic element of the system's own stealth into a tactical-level decision optimization function for the first time. This design fundamentally resolves the inherent contradiction between stealth and proactivity. Finally, through dynamically calculated response thresholds directly linked to cost-effectiveness, it ensures that every proactive resource allocation is a precise operation at the optimal time and for the highest-value target, thereby solving the rigidity and inefficiency problems of traditional scheduling strategies. Attached Figure Description

[0023] Figure 1 This is a schematic diagram illustrating the application scenario of the low-altitude aircraft of the present invention;

[0024] Figure 2 This is a schematic diagram of the execution logic of S1 and S2 of the present invention;

[0025] Figure 3 This is a schematic diagram of the execution logic of S3 to S5 of the present invention. Detailed Implementation

[0026] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0027] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0028] Example 1:

[0029] Please see Figures 1 to 3 The present invention provides a technical solution:

[0030] The low-altitude aircraft identification method based on multi-source data fusion includes the following steps:

[0031] S1: Obtain the time series of micro-Doppler features formed by the target aircraft reflecting signals from non-cooperative opportunistic illumination sources through a passive sensor network;

[0032] S2: Input the micro-Doppler feature time series into the preset behavior baseline prediction model, and calculate the behavior anomaly index, which characterizes the degree of current behavior anomaly of the target aircraft, based on the deviation between the prediction results of the preset behavior baseline prediction model and the measured values ​​of the micro-Doppler feature time series.

[0033] S3: Based on the behavioral anomaly index, the target threat level of the target aircraft is updated in real time;

[0034] S4: Based on the updated target threat level, generate an active sensor resource scheduling strategy for the target aircraft through a preset decision model;

[0035] S5: Execute an active sensor resource scheduling strategy to prioritize the detection and identification of target aircraft.

[0036] Further explanation: Non-cooperative opportunity exposure source signals include at least one of the following: 5G communication base station signals, digital television broadcast signals, or Wi-Fi signals;

[0037] The micro-Doppler characteristic time series specifically includes "rotor blade scintillation frequency harmonic ratio" and "fuselage vibration modulation spectrum entropy";

[0038] The preset behavior baseline prediction model is a dynamic time-series prediction model based on "gated cyclic units constrained by physical models";

[0039] The "gated loop unit of physical model constraints" is also used to output the "feature dimension importance vector";

[0040] A diagonal weighted matrix is ​​constructed using the "feature dimension importance vector" and participates in quadratic form operations. The results are then normalized using saturated linear unit functions to obtain the "adaptive feature space Mahalanobis distance". The "adaptive feature space Mahalanobis distance" is then represented as the "behavioral anomaly index".

[0041] The following are specific implementation instructions for the above content:

[0042] The core technical feature of this embodiment lies in the following: By utilizing a passive sensor network and non-cooperative opportunistic illumination source signals, multi-dimensional micro-Doppler feature time series that deeply characterize the physical structure and operational state of the aircraft are extracted. A dynamic time-series prediction model based on a "Physics-Informed-Gated-Recurrent-Unit (PI-GRU)" is introduced. This model not only learns the statistical regularities of historical data but also uses prior knowledge of the aircraft's dynamics model as constraints to predict its micro-Doppler "gait" baseline under normal flight conditions. Finally, by calculating the "adaptive feature space Mahalanobis distance" between the measured features and this high-fidelity prediction baseline, a behavioral anomaly index that dynamically adjusts the sensitivity to different feature dimensions is generated. This overcomes the problem of inaccurate predictions by traditional data-driven models when facing small samples or sudden changes in flight attitude, and solves the problem that fixed metrics cannot effectively capture the differences in the contribution of multi-dimensional feature anomalies, thereby improving the sensitivity and accuracy of capturing "threat pre-ignition" signals.

[0043] The core objective of this embodiment, based on the dynamic time-series prediction model of "Physics-Informed-Gated-Recurrent-Unit (PI-GRU)", is to generate a high-fidelity prediction baseline for the micro-Doppler feature time series of the target aircraft by fusing data-driven time-series dependencies with prior physical model constraints.

[0044] The PI-GRU model is based on a standard Gated Recurrent Unit (GRU) network. This network is a variant of a recurrent neural network, and its internal structure includes update and reset gates to adaptively capture and maintain dependencies across different time scales when processing time-series data. The GRU network receives a sequence of historical micro-Doppler feature vectors within a time window as input and encodes the historical information of the sequence through its hidden states. The network's output layer is used to generate predictions of the micro-Doppler feature vectors for the next time step. A key technical feature of this embodiment is the introduction of physical model constraints into the aforementioned basic GRU network, the specific implementation of which is as follows:

[0045] A linear dynamic model is established to describe the dynamic evolution of the micro-Doppler characteristics of a target aircraft under normal flight conditions. This model is based on the fundamental physical assumption that, in the absence of external anomalous disturbances or internal faults, the change of the aircraft's micro-Doppler eigenvectors between adjacent time steps is continuous and inertial. Specifically, the model expresses the eigenvector of the next time step t+1 as a linear combination of the eigenvectors of the current time step t and several previous time steps. The coefficient matrix of this linear combination, i.e., the state transition matrix, is predetermined through analysis of a large number of offline datasets containing only normal flight behavior using a system identification method (least squares method). This model constitutes the physical prior knowledge of the aircraft's normal behavior.

[0046] During model training, for any given time step, there are two independent prediction results: one is the data prediction feature vector obtained by the GRU network itself through data-driven learning; the other is the physical prediction feature vector calculated based on the established linear dynamics model. The difference between these two prediction vectors, either the Euclidean distance or the mean square error, is defined as the physical consistency residual. This residual value quantitatively characterizes the extent to which the GRU network's prediction deviates from the preset physical motion laws.

[0047] The following composite loss function is constructed to simultaneously optimize data fitting accuracy and conformity to physical laws during model training; the composite loss function consists of two weighted components:

[0048] Data fidelity loss: This component measures the difference between the data-predicted feature vectors output by the GRU network and the actual observed feature vectors.

[0049] Physical consistency loss: This component is the physical consistency residual defined above. The final value of the composite loss function is the weighted sum of the two loss components mentioned above.

[0050] The relative importance between the two loss components is adjusted by a pre-defined physical constraint weighting factor. The optimal value of this weighting factor is determined through a systematic hyperparameter optimization method. Specifically, on a set of reserved validation datasets, iterative training and evaluation are performed within the pre-defined weighting factor value range using methods such as grid search or Bayesian optimization. Finally, the weighting factor value that enables the model to achieve the lowest prediction error and the highest anomaly detection discrimination on the validation set is selected as the configuration parameter for the final deployed model.

[0051] This embodiment employs a gradient-based optimization algorithm (Adam optimizer) and backpropagation mechanism to iteratively optimize the network parameters (including weights and biases) of the PI-GRU model to minimize the composite loss function. The training process uses a large-scale micro-Doppler feature time-series dataset containing only normal flight modes until the model's performance converges on the validation set.

[0052] In a specific embodiment of the present invention, the key parameters involved are defined as follows:

[0053] The passive sensor network is deployed within the monitored area and consists of at least three geographically distributed radio frequency (RF) receivers with high-precision time synchronization capabilities. The RF receivers are used to synchronously acquire signals from non-cooperative opportunistic illumination sources; in this embodiment, these signals are preferably downlink signals from a 5G communication base station with a center frequency of 3.5 GHz.

[0054] The harmonic ratio of the rotor blade scintillation frequency, its parameter symbol is: The physical meaning of this parameter is the ratio of the fundamental frequency scintillation peak energy to the second harmonic peak energy generated on the micro-Doppler spectrum of the received signal as the target aircraft's rotor blades periodically block and reflect the signal from the opportunity illumination source during rotation. The value of this parameter is strongly correlated with the number, size, and rotational speed of the aircraft's rotors, and is a key physical fingerprint for distinguishing different types of multi-rotor aircraft. It is obtained by performing time-frequency analysis on the received target reflection signal to obtain the micro-Doppler spectrum, locating the periodic scintillation line spectrum generated by the rotor on the spectrum, and extracting its fundamental frequency peak value. and second harmonic peak value The signal energy at these two locations is denoted as follows: and The calculation model originates from the micro-Doppler theory in the field of radar signal processing; the harmonic ratio of the rotor blade scintillation frequency. The value is obtained by converting the peak energy of the second harmonic. Divide by the peak energy of the fundamental frequency scintillation Then, it is normalized using a preset logarithmic mapping function. The purpose of the logarithmic mapping function is to compress the dynamic range of the parameter and limit its value range to the interval (0,1).

[0055] If the fundamental frequency scintillation peak energy is extracted from the micro-Doppler spectrum of the quadcopter drone at the current moment... -20 dBm, second harmonic peak energy If the value is -26 dBm, then its energy ratio is 0.25. After processing with a logarithmic mapping function (expressed as "1" plus the reciprocal of the natural logarithm of the energy ratio), the normalized rotor blade scintillation frequency harmonic ratio is obtained. The value is 0.8.

[0056] The fuselage vibration modulation spectrum entropy, with parameter symbol as follows: The physical meaning of this parameter is a quantitative representation of the spectral complexity of the sideband modulation caused by the Doppler frequency shift of the target aircraft due to the minute vibrations of the fuselage generated by engine operation or airflow disturbance. The value of this parameter reflects the rigidity, stability, and load state of the aircraft's fuselage structure. The calculation model originates from the Shannon entropy theory in information theory; the calculation logic involves locating the central Doppler spectral line generated by the target aircraft in the micro-Doppler spectrum. Secondly, a preset frequency bandwidth (in this embodiment, the bandwidth) is established on both sides of the central Doppler spectral line. The value was determined by offline experiments and is the rotor fundamental frequency scintillation frequency. The sideband modulation spectrum is extracted from half of the spectrum. Then, this sideband modulation spectrum is normalized so that its total energy is one. Finally, according to the definition of information entropy, the energy probability value of each frequency point in the normalized spectrum is multiplied by its own logarithm, the sum is taken, and the negative value is calculated. The result is then divided by the maximum entropy value within that frequency bandwidth to obtain the fuselage vibration modulation spectrum entropy. Its value range is limited to the interval (0,1). This embodiment has a bandwidth... There are 128 frequency points. If the sideband modulation spectrum energy is highly concentrated at a few frequency points, the calculated normalized fuselage vibration modulation spectrum entropy... A value of 0.2 indicates that the fuselage vibration mode is singular and highly rigid; if the energy is uniformly distributed across all frequency points, then the normalized fuselage vibration modulation spectrum entropy... A value close to 1.0 indicates a complex vibration mode and the presence of abnormal loads or structural damage.

[0057] Behavioral Abnormality Index, whose parameter symbol is The anomaly index represents the statistical deviation between the measured micro-Doppler eigenvector of a target aircraft at the current moment and the baseline eigenvector predicted by a high-fidelity prediction model under normal flight conditions. This index is a comprehensive, dimensionless scalar; the larger the value, the more the aircraft's behavior deviates from its inherent pattern, and the higher the probability of an anomaly. The calculation model for this parameter is an improvement on the Mahalanobis distance in statistics, represented as an "adaptive feature space Mahalanobis distance." Its calculation relies on a dynamic temporal prediction model of a "physical model-constrained gated cyclic unit (PI-GRU)." At each time step, the model outputs, in addition to the predicted eigenvector, a feature dimension importance vector. (Behavioral Anomaly Index) The calculation method is as follows: The difference between the measured micro-Doppler feature vector (hereinafter referred to as the measured feature vector) and the predicted feature vector at the current moment is calculated to obtain the residual vector. Next, each element of the feature dimension importance vector is used as a diagonal element to construct a diagonal weighted matrix. Then, a quadratic operation is performed on the residual vector, the inverse covariance matrix of the historical feature sequence, and the diagonal weighted matrix. The quadratic operation is the transpose of the residual vector multiplied by the inverse covariance matrix, then multiplied by the diagonal weighted matrix, and finally multiplied by the residual vector. Finally, the square root of the result is taken, and its output range is smoothly limited to the interval [0,1] using a saturated linear unit (SaLU) function.

[0058] Furthermore, in this embodiment of the invention, in order to map the calculated adaptive feature space Mahalanobis distance into a behavior anomaly index with a clear physical meaning and convenient for subsequent decision-making modules, the following saturated linear unit function is adopted.

[0059] Saturated linear unit functions are piecewise functions designed to achieve the following technical effects: when the input value is small, the output is linearly related to the input to maintain sensitivity to small anomalies; when the input value exceeds a certain threshold, the output value smoothly approaches and saturates to 1 to avoid numerical overflow caused by extreme outliers and to provide a stable and bounded measure of the degree of anomaly.

[0060] This saturated linear unit function contains two key configurable parameters: the linear scaling factor and the saturation threshold.

[0061] When the calculated result of the Mahalanobis distance in the adaptive feature space is less than or equal to zero, the output of the function is always zero.

[0062] When the input value is greater than zero and less than the saturation threshold, the function's output is the product of that input value and the linear scaling factor. Within this range, the function grows linearly.

[0063] When the input value is greater than or equal to the saturation threshold, the function output is always "1".

[0064] To achieve a smooth transition from the linear growth region to the saturation region and avoid abrupt changes that prevent differentiability at the threshold point, this embodiment preferably employs a smooth approximation function instead of the aforementioned hard-segmentation definition. Specifically, it uses a modified Sigmoid function or a variant of the hyperbolic tangent function, adjusting its parameters to make its behavior approximately linear near the origin, while rapidly saturating in regions far from the origin.

[0065] The determination of the linear scaling factor and saturation threshold is based on statistical analysis of the distribution of historical data.

[0066] To determine the saturation threshold: a large number of data samples containing only normal flight behavior were collected, and the adaptive feature space Mahalanobis distance for each sample was calculated. Then, the distribution of these distance values ​​was statistically analyzed to calculate their mean and standard deviation. The saturation threshold was set as the 99.9th percentile of the normal sample distance distribution, or as the mean plus three or four times the standard deviation. The technical logic behind this setting is that any distance value exceeding this threshold is highly likely to indicate that it does not belong to the normal behavior distribution and should be considered a significant anomaly, assigned a maximum anomaly index close to 1.

[0067] The linear scaling factor is set to the reciprocal of the saturation threshold. This setting ensures that when the input value is exactly equal to the saturation threshold, the function's output in the linear part is exactly one, thus guaranteeing the function's continuity at the breakpoints.

[0068] In one specific implementation of this invention, through statistical analysis, the 99.9th percentile of the Mahalanobis distance distribution in the adaptive feature space under normal flight conditions is determined to be 2.5. Therefore, the saturation threshold is set to 2.5, and the linear scaling factor is set to 1 divided by 2.5, i.e., 0.4.

[0069] At this point, the behavior of the saturated linear unit function is as follows:

[0070] If the input distance is 0.5 (less than 2.5), the output behavior anomaly index is 0.5 multiplied by 0.4, which equals 0.2.

[0071] If the input distance is 2.0 (less than 2.5), the output behavior anomaly index is 2.0 multiplied by 0.4, which equals 0.8.

[0072] If the input distance is 3.0 (greater than or equal to 2.5), the output behavior anomaly index will be 1.0.

[0073] In this embodiment, during the test, the aircraft made a normal turn, and the PI-GRU model accurately predicted the harmonic ratio of the rotor blade scintillation frequency. and fuselage vibration modulation spectrum entropy If the trend of change is small, then the norm of the residual vector is small, and the calculated behavioral anomaly index is... The value is close to 0.05. If the aircraft rotor suddenly malfunctions, causing the rotor blades to flicker at a harmonic ratio... Dramatic changes occur, while the PI-GRU model, based on physical constraints, predicts that it should remain stable. In this case, the residual vector at the rotor blade scintillation frequency harmonic ratio... The weight of each dimension will be large, and the feature dimension importance vector output by the model will also increase the weight of that dimension, ultimately leading to a higher calculated behavioral anomaly index. Value soared.

[0074] The complete calculation process for implementing the core technical features S1 and S2 in this embodiment is specifically broken down into the following steps:

[0075] 1.1) The initial input is the measured micro-Doppler feature vector, which has been synchronized and preprocessed, provided by the front-end signal processing module of the passive sensor network at the current time step t. This vector contains at least the harmonic ratio of the rotor blade scintillation frequency at the current moment. and fuselage vibration modulation spectrum entropy Simultaneously, access the historical micro-Doppler feature time series stored in memory up to the previous time t-1. , and the inverse covariance matrix of the historical feature sequence.

[0076] 1.2) Processing Step 1: Extract historical micro-Doppler feature time series As input, it is fed into a gated recurrent unit (PI-GRU) model that has been pre-trained offline with physical model constraints.

[0077] The PI-GRU model performs one forward propagation computation. A unique feature of this model is that its loss function, during training, includes not only a data prediction error term but also a penalty term. This penalty term is used to punish predictions that violate the predefined constraints of the aircraft dynamics model. The model outputs two results:

[0078] Predicting feature vectors : The predicted value of the micro-Doppler eigenvector at the current time t.

[0079] Feature Dimension Importance Vector A vector with the same dimension as the feature vector, where each element has a value between [0,1], representing the model's assessment of the reliability or importance of the corresponding feature dimension under the current flight condition. In hovering, the fuselage vibration modulation spectrum entropy... The change is compared to the harmonic ratio of the rotor blade scintillation frequency. It can better reflect anomalies, and at this time the feature dimension importance vector The corresponding fuselage vibration modulation spectrum entropy The element value will be higher.

[0080] 1.3) Processing Step Two: Calculate the measured micro-Doppler eigenvectors With predicted feature vectors The residual vector between .

[0081] Construct a diagonal weighted matrix The diagonal elements of this matrix are composed of the feature dimension importance vector. The elements are composed of zero off-diagonal elements.

[0082] 1.4) In step three, perform adaptive Mahalanobis distance calculation in the feature space: calculate the residual vector... The transpose of , the inverse covariance matrix of the historical feature sequence, and the diagonal weighted matrix and residual vector The function performs a series of multiplications; takes the square root of the result to obtain an intermediate distance value; inputs this intermediate distance value into a saturated linear unit function for normalization. This "saturated linear unit function" grows linearly when the input value is small, and its output value smoothly converges and saturates to 1 after the input value exceeds a preset threshold.

[0083] 1.5) The final output is the behavioral abnormality index. Its value is a single floating-point number between [0,1]. Behavioral Anomaly Index It directly quantifies the degree of abnormality in the target aircraft's behavior at time t, serving as a direct input for subsequent threat assessment and decision-making modules;

[0084] The core of this embodiment lies in demonstrating the behavioral abnormality index. The generation is achieved through an adaptive, statistical distance-based deep fusion method; the reason for choosing and improving Mahalanobis distance as the basic fusion framework is as follows:

[0085] Consider feature correlation: Using the covariance matrix to measure distance can naturally handle the correlation between features and avoid distance measurement distortion caused by linear correlation of features.

[0086] Scale independence: It is not sensitive to the scale of the data and does not require scaling the rotor blade scintillation frequency harmonic ratio for different physical dimensions. and fuselage vibration modulation spectrum entropy Perform complex artificial normalization.

[0087] A feature dimension importance vector dynamically generated by the PI-GRU model is introduced. An adaptive Mahalanobis distance for the feature space was constructed as a weighting factor. The improvement is based on the assumption that traditional Mahalanobis distance assumes all feature dimensions are equally important at all times, which contradicts physical reality. During high-speed straight-line flight, the harmonic ratio of the rotor blade scintillation frequency... The minute jitters could be noise, while the fuselage vibration modulation spectrum entropy A sudden increase in these dimensions may indicate a structural problem. The method in this embodiment allows the model to dynamically "focus" on more indicative feature dimensions based on the current flight context, thereby achieving context-aware intelligent fusion.

[0088] The feature dimension importance vector in this mechanism It is obtained through autonomous learning by the PI-GRU model under the dual guidance of data-driven and physical constraints; in the network structure of the PI-GRU model, an additional fully connected layer is added before its output layer, and a sigmoid activation function is used. The output of this layer is the feature dimension importance vector. During model training, this feature dimension importance vector... The feature importance vector is used in the loss function calculation along with the predicted feature vector. Through the backpropagation algorithm, the model learns how to adjust the feature dimension importance vector. This allows the final calculated behavioral anomaly index to most effectively distinguish normal samples from abnormal samples. Essentially, by optimizing the classification task, it learns how to assign appropriate feature weights to different flight scenarios.

[0089] Further detailed implementation instructions for the above content:

[0090] The range of the behavioral abnormality index is limited to the interval [0,1].

[0091] When the abnormal behavior index The closer the calculated value is to 0, the more closely the micro-Doppler "gait" characteristics of the target aircraft match its historical behavior patterns and embedded physical dynamics model, the lower the threat level, and the lower the need to activate high-cost active sensing resources.

[0092] Technically, this means that the system determines that the aircraft is in a predictable and stable "normal" flight state, whether hovering, flying in a straight line at a constant speed, or performing routine maneuvers, its behavior is within the expectations of the baseline model.

[0093] When the abnormal behavior index The closer the calculated value is to 1, the more significantly the micro-Doppler characteristics of the target aircraft deviate from its high-fidelity prediction baseline, the higher the threat level, and the higher the need to activate high-cost active sensing resources.

[0094] Technically, this means that if the aircraft experiences an unexpected physical state change that cannot be explained by the model, the system is more likely to determine that its behavior is abnormal and that there is a potential "threat pre-ignition" signal.

[0095] Abnormality Index Affecting the Final Output Key intermediate parameters include the residual vector Norm and feature dimension importance vector ;

[0096] When other parameters remain constant, the residual vector The norm (i.e., the Euclidean distance between the measured features and the predicted features) and the final behavioral anomaly index It exhibits a strictly positively correlated nonlinear relationship. The norm of the residual vector is the fundamental input for the quadratic form operation that forms the core of the adaptive feature space Mahalanobis distance calculation. According to the mathematical properties of this quadratic form operation, an increase in the norm of the residual vector will lead to an increase in the result of the quadratic form operation, which in turn increases the median distance value after taking the square root. Ultimately, this is mapped to a larger behavioral anomaly index through the saturated linear unit function. The positive correlation design directly maps the "magnitude of prediction bias" to the "degree of anomaly," which aligns with the inherent logic of anomaly detection.

[0097] Feature Dimension Importance Vector The values ​​of each element, and the behavioral abnormality index The relationship between them is dynamic and modulated positively correlated. Feature dimension importance vector. The elements constitute a diagonal weighted matrix. In quadratic form operations, this weighting matrix acts as an "amplifier" or "reducer" for the residual components of different feature dimensions. If the feature dimension importance vector... The middle corresponds to the harmonic ratio of the rotor blade scintillation frequency. A higher element value means that, under the current flight scenario, the model considers the rotor blade scintillation frequency harmonic ratio to be high. Dimensional deviations are more indicative. Therefore, even the rotor blade scintillation frequency harmonic ratio... The residual component is small, and after being amplified by this weight, it has a significant impact on the final behavioral anomaly index. The contribution will also increase significantly. Conversely, the contribution will also increase significantly. This design is one of the core innovations of this invention. It makes the calculation of the behavioral anomaly index no longer isotropic, but has a "context-aware" capability, which can intelligently focus on the most important abnormal signal source at present.

[0098] To verify the superiority of the method of this invention over traditional methods, a series of comparative experiments were designed. The experimental subject was a quadcopter drone, and different types of flight events were injected. The table below shows the behavioral anomaly indices calculated by the method of this invention (using the PI-GRU model and adaptive Mahalanobis distance) and two comparative methods (Method 1: using the standard GRU model and Euclidean distance; Method 2: using the standard GRU model and standard Mahalanobis distance) in three typical scenarios. All indices have been normalized to the [0,1] interval.

[0099] Table 1. Comparison of behavioral anomaly index calculation results using different methods in typical flight scenarios:

[0100]

[0101] The experimental data above clearly reveal the outstanding substantive features and significant progress of the method of the present invention in improving anomaly detection performance.

[0102] Comparing scenario one (stable hovering) and scenario two (normal turning maneuvers), the anomaly indices calculated by the three methods are all at low levels, indicating that they can all identify this as normal behavior. However, the output values ​​of the method of this invention (mean values ​​of 0.03 and 0.08) are significantly lower than the two comparative methods (mean values ​​of 0.055 and 0.145 for method one; and 0.04 and 0.115 for method two). This strongly demonstrates that the introduction of the Physically Constrained Gated Cyclic Unit (PI-GRU) enables the baseline prediction model of this invention to have stronger generalization ability and more accurate prediction ability for normal maneuvers, thereby producing smaller prediction residuals. Compared with comparative method two, the anomaly index of the method of this invention is reduced by 29% in the normal turning scenario, which reduces the false alarm rate of the system under normal flight conditions, proving the superiority of its design.

[0103] Scenario 3 simulates a sudden rotor failure, causing the rotor blades to flicker frequency harmonic ratio. A sharp decline. In this scenario, the behavioral anomaly index calculated by the method of this invention (mean 0.955) is significantly higher than that of comparative method one (mean 0.56) and comparative method two (mean 0.69). The core reason is that the PI-GRU model determines the current state based on the flight context. It should remain stable, therefore the corresponding values ​​in the feature dimension importance vector should be dynamically adjusted. The weight was increased to 0.9. This "intelligent focusing" mechanism greatly amplified the effect. Anomaly signals in the dimension. Compared to the better-performing comparative method two, the anomaly index of the method in this invention is improved by 38%. This result irrefutably demonstrates the immense value of the core innovation of adaptive feature space Mahalanobis distance, which enables this invention to have unparalleled detection sensitivity for fault signals of specific key components.

[0104] Scenario 4 simulates a more subtle anomaly—a drone carrying an undeclared heavy load, primarily manifested in the fuselage vibration modulation spectrum entropy. The significant increase. Similar to scenario three, the method of this invention again demonstrates its "context-aware" capability, which will... The importance weight of the dimension was increased to 0.8, resulting in an anomaly index as high as 0.90. In contrast, the two comparative methods, unable to dynamically adjust feature weights, produced output values ​​(0.495 and 0.625, respectively) far lower than that of this invention. This demonstrates that the method of this invention can not only detect severe hardware failures but also effectively identify more subtle and hidden behavioral anomalies caused by load changes, etc. These data strongly support the feasibility and advancement of this invention in achieving the technical goal of "intent prediction."

[0105] Further explanation: The specific steps for updating the target threat level are as follows: the abnormal behavior index is used as a dynamic risk weight, and it is weighted and fused with the target aircraft's historical threat assessment value and prior identity credibility to generate the updated target threat level;

[0106] The decision-making model is a reinforcement learning model; the input state of the reinforcement learning model is a threat situation map composed of the target threat levels of all target aircraft.

[0107] The reward function of a reinforcement learning model is calculated by including expected information gain, resource scheduling cost, and exposure risk cost.

[0108] The steps for implementing an active sensor resource scheduling strategy include: tiered allocation based on the target threat level;

[0109] The specific logic of the hierarchical invocation is that, based on the first preset threat level threshold and the second preset threat level threshold, operations such as maintaining passive tracking, invoking non-contact sensors, or invoking high-precision sensors are executed respectively.

[0110] The following are specific implementation instructions for the above content:

[0111] The core technical feature of this embodiment lies in the following: First, through a multi-source information fusion mechanism, a comprehensive target threat level is generated by dynamically weighting the behavioral anomaly index (representing the target's immediate behavioral anomalies), the historical threat assessment value (representing its historical behavioral patterns), and the prior identity credibility (representing its background identity). Second, the sensor scheduling problem across the entire airspace is constructed as a Markov decision process based on reinforcement learning, aiming to minimize the "comprehensive task cost." Here, the "comprehensive task cost" not only includes traditional resource scheduling costs but also innovatively introduces the exposure risk cost—the risk of revealing one's own position due to quantified active detection actions. Finally, a pre-trained reinforcement learning model, based on a real-time threat situation map composed of all target threat levels, outputs the optimal active sensor action sequence that maximizes the expected information gain while minimizing the aforementioned comprehensive task cost. Based on this sequence and the target threat level, sensors are then used in a tiered manner. This approach incorporates the strategic cost of "exposure" into the optimization function of tactical-level sensor scheduling for the first time, fundamentally resolving the inherent contradiction of traditional methods in pursuing identification efficiency while failing to consider the long-term stealth of the system.

[0112] In a specific embodiment of the present invention, the key parameters for realizing intelligent decision-making and resource scheduling are defined as follows: the active sensor network includes at least one high-definition photoelectric camera capable of azimuth, pitch, and zoom control and one millimeter-wave radar capable of sector scanning.

[0113] The updated target threat level has the following parameter symbols: This represents the comprehensive quantitative assessment of the potential threat posed by a single target aircraft at the current time step; it is a dimensionless scalar normalized to the interval [0,1], with higher values ​​indicating a more severe threat; the calculation model for this parameter originates from the Bayesian update or recursive filtering ideas in the field of information fusion, iteratively updating the state estimate by continuously incorporating new observational evidence; the updated target threat level It is based on the historical threat assessment values ​​from the previous time step. This is obtained by weighted summation with the immediate risk assessment item at the current time step. The immediate risk assessment item is calculated by taking the behavioral anomaly index at the current moment. After being processed by the Sigmoid function, a non-linear mapping function, it is then multiplied by a factor determined by prior identity confidence. The calculated identity risk coefficient is "1" minus the prior identity credibility. The weighting coefficients in the weighted summation are respectively set as follows: and These are preset system parameters, determined by: inputting a large amount of sample data containing the threat evolution process through offline simulation experiments, with the optimization objective of maximizing the correlation between the final threat level curve and the real threat evolution label, and calibrating the optimal combination of weight coefficients.

[0114] Historical threat assessment values, with parameter symbols as follows: Its physical meaning is a dynamic variable that is recursively updated within each decision-making cycle, characterizing the cumulative threat level of a target aircraft. Its update mechanism ensures the continuity of threat assessment and the retention of historical behavior.

[0115] When a new target aircraft is first detected and its track is established by a passive sensor network, its initial historical threat assessment value It is assigned a preset, low, non-zero initial value. This initial value reflects the inherent, fundamental uncertainty risk of any unknown objective.

[0116] At the end of each decision cycle, after calculating the updated target threat level for the current cycle, this target threat level value will be used as the historical threat assessment value at the start of the next decision cycle. .

[0117] The recursive update process is described as follows: In the next decision cycle, the historical threat assessment value of target aircraft i is equal to the updated target threat level of the target aircraft calculated in the current decision cycle.

[0118] This gives target threat level assessment a "memory effect." If a target exhibits slight anomalies for several consecutive periods, its threat level will accumulate and slowly increase; conversely, if a high-threat target's behavior returns to normal in subsequent periods, its threat level will gradually decrease. This makes threat assessment not only reflect the instantaneous state but also the trend of behavioral evolution, enhancing the robustness of decision-making.

[0119] This embodiment An unknown target ( Historical threat assessment value The value is 0.3. At the current moment, its behavioral abnormality index is... The value is 0.9, which becomes 0.95 after nonlinear mapping. Therefore, its updated target threat level is... The result is 0.43, calculated by multiplying 0.8 by 0.3, adding 0.2 by 0.95, and then multiplying by (1-0).

[0120] Prior identity credibility, its parameter symbol is This represents the confidence level that a target aircraft belongs to the "whitelist" or is a known cooperative target during the passive identification phase; it is a dimensionless scalar normalized to the [0,1] interval. This parameter value is obtained by matching the target's passive identification features (e.g., the device MAC address demodulated from its reflected Wi-Fi signal) against a pre-stored "whitelist" database containing the identities of cooperative aircraft; in this embodiment, "passive identification features" is represented as "the device MAC address demodulated from its reflected Wi-Fi signal."

[0121] If there is a perfect match, then the prior identity credibility is... Assign a value of 1.0; if there is no match at all or the identity identifier cannot be obtained, then the prior identity credibility is reset. If a partial match or fuzzy match exists, a value between 0 and 1 is assigned based on the similarity of the match.

[0122] Expected information gain, with parameter sign as The expected information entropy is the degree to which the uncertainty of the system regarding the true category of the target aircraft decreases after performing a specific active sensor action. In this embodiment, the "true category" is defined as "hostile, neutral, friendly, etc." The calculation model of this parameter is derived from the information entropy theory in information theory. Specifically, the system maintains a probability distribution vector for the target category internally. Before performing the action, the current information entropy is calculated based on this probability distribution. Next, the system uses a pre-trained classifier model to predict how the target's category probability distribution will be updated after performing a sensor action ("sensor action" includes acquiring a high-resolution image) and obtaining the observation results. The expected posterior information entropy after performing the action is calculated by weighted averaging of all possible observation results and their probabilities. The expected information gain is the difference between the current information entropy and the expected posterior information entropy, and its value range is normalized to the [0,1] interval. Specifically, it is obtained by dividing it by the maximum possible value of the current information entropy, i.e., the entropy value under uniform distribution.

[0123] Resource scheduling cost, its parameter symbol is This represents the total quantified value of physical resources consumed to perform active sensor actions, directly related to the system's endurance. This parameter is obtained by weighted summation of the resource consumption involved in performing the action. These resource items include at least: the energy consumption for the sensor to switch from standby to operating state, the power consumption for maintaining operating state per unit time, the energy consumption of the servo mechanism required to drive the sensor to point at the target, and the computational resource consumption required to process the sensor data. The weight coefficients of each resource are determined offline based on their criticality to the system's overall endurance. The final cost value is normalized by dividing by the maximum allowable cost for a single action set by the system.

[0124] The cost of exposing risk, its parameter symbol is: This represents a quantitative assessment of the risk that the enemy's electronic reconnaissance system might intercept information such as the system's location and operating mode due to sensor actions that actively emit electromagnetic signals. The exposure risk cost is a discrete value based on strategy assignment. It is determined as follows: for any sensor action that does not generate active electromagnetic radiation, its exposure risk cost... A constant value of 0 is assigned; for sensor actions that generate active electromagnetic radiation, the exposure risk cost is... A preset penalty value is assigned; the specific size of this penalty value is adjusted according to the concealment requirements of the task scenario. The higher the concealment requirements of the task scenario, the larger the preset penalty value corresponding to the risk cost of exposure.

[0125] Exposure of risk costs The penalty value is a function directly related to the current system stealth level. Specifically, the system internally stores a penalty value lookup table, which predefines several discrete system stealth levels representing the macro-level mission situation. This lookup table establishes and stores a deterministic mapping relationship between each system stealth level and a specific, normalized exposure risk cost penalty value, specifically obtained in advance based on expert knowledge and numerous offline simulation adversarial experiments. These levels, from low to high, represent a decreasing tolerance for system exposure risk; the following three levels are defined:

[0126] Level 3: Proactive Defense: At this level, the mission objective is to prioritize the rapid and accurate identification of threats, with a higher risk of exposure being acceptable.

[0127] Level 2: Routine Alertness: At this level, a balance needs to be struck between identification efficiency and personal concealment.

[0128] Level 1: Silent Surveillance: At this level, the system's primary task is to maintain absolute electromagnetic silence and avoid any behavior that could expose itself.

[0129] The currently effective system stealth level, as an external configuration parameter, is issued and set in real time by the system operator or the superior command and control system according to changes in the mission phase or battlefield environment.

[0130] Expose risk costs in each decision-making cycle of the system operation. The process for determining the penalty value is as follows:

[0131] The system queries the currently externally set system concealment level → using this level as an index, the system retrieves the corresponding penalty value from the penalty value lookup table → the system assigns this retrieved penalty value to the actions of all sensors that generate active electromagnetic radiation, as their exposure risk cost within this decision-making cycle. .

[0132] The penalty value lookup table is defined as follows in this embodiment:

[0133] Level 3 -> Penalty value 0.1; Level 2 -> Penalty value 0.5; Level 1 -> Penalty value 0.9;

[0134] If the system receives an instruction to set the current system stealth level to "Level 1: Silent Surveillance," then in subsequent decision-making cycles, any action that invokes millimeter-wave radar will incur exposure risk costs. It will be automatically assigned a value of 0.9. This will make the reinforcement learning model have a low probability of selecting this action when calculating the reward function, unless the expected information gain is close to 1.0 and there are no other means.

[0135] Conversely, if the system is set to "Level 3: Active Defense", then the risk and cost of radar actions being exposed are... With a value of only 0.1, the model is more inclined to use radar when necessary in exchange for higher information gain, thereby achieving a more proactive detection strategy. Through the above unified and detailed description, it is clear that the determination of exposure risk cost is a two-step process: first, a binary classification (zero / non-zero) is performed based on the physical attributes of the action, and then the non-zero part is accurately assigned a value through an adaptive mechanism that is linked to the requirements of the macro task.

[0136] The "sensor action that does not generate active electromagnetic radiation" in this embodiment includes the case of "using a high-definition photoelectric camera for passive optical imaging";

[0137] The following definition is made public for pre-trained reinforcement learning models:

[0138] The pre-defined decision model is preferably a reinforcement learning model based on a Deep-Q Network (DQN). Its key components are described in detail below.

[0139] The first step is to construct the state space: representing the "state" input as " ";state The system is constructed as a fixed-size two-dimensional tensor. Each row of this tensor corresponds to one target channel that the system can track simultaneously. In this embodiment, the system's maximum tracking capacity is 50 targets, so the tensor has 50 rows. Each row contains key state features describing the target corresponding to that channel, including at least: the updated target threat level. Normalized three-dimensional spatial coordinates of the target And a binary flag indicating whether the channel is active.

[0140] All input features must undergo "min-max normalization" before being fed into the network to ensure that features of different dimensions contribute equally to the network training.

[0141] Step 2: The set of atomic actions in the action space is denoted as AS1 and is defined as a discrete, finite set of actions. Each atomic action... It is a precise tuple that can be directly converted into hardware control instructions. For continuous parameters, they are discretized using a gridding method to accommodate the discrete action space requirements of the DQN model.

[0142] The third step is to construct the neural network architecture:

[0143] At the core of the DQN model is a multi-layered, fully connected feedforward neural network. The number of neurons in its input layer matches the total number of elements in the state tensor. The network contains several hidden layers, each using a rectified linear unit (ReLU) as the activation function. The number of neurons in the output layer is equal to the total number of atomic actions in the action space.

[0144] For a given input state The network outputs a vector, where each element corresponds to the estimated long-term cumulative reward, or Q-value, that can be obtained after performing the corresponding atomic action. ;in This represents the r-th atomic action;

[0145] The model is trained in a high-fidelity digital twin simulation environment. This environment can simulate multiple batches of different types of target aircraft intrusion processes and accurately calculate the expected information gain, resource scheduling cost, and exposure risk cost of each sensor action.

[0146] During training, every experience generated by the interaction between the model and the simulation environment (i.e., the quadruple of state, action, reward, and next state) is stored in a large cache called the "experience replay pool".

[0147] At each training step, a mini-batch of experience is randomly drawn from the experience replay pool. Using this experience, the weights of the DQN network are updated by minimizing the Bellman error. Specifically, gradient descent is employed to make the network... The predicted value continuously approaches the sum of the immediate reward and the maximum estimated Q value of the next state. To improve the stability of training, this embodiment uses an independent "target network" that periodically replicates weights from the main network to calculate the Q value of the next state.

[0148] During actual deployment and runtime, the system will construct the state tensor in real time. The input is fed into a pre-trained DQN model, and a forward propagation is performed to obtain the Q-value vectors of all actions. The system then selects the action with the highest Q-value as the optimal action to be executed in the current decision cycle.

[0149] Furthermore, the complete calculation process for realizing the core technical features of this embodiment can be broken down into the following steps:

[0150] 2.1) The initial input is the set of all N target aircraft tracked by the system in the current decision cycle. For each target aircraft i, i=1,...,N, the input information includes: its latest behavioral anomaly index. Historical threat assessment value and the credibility of prior identity .

[0151] 2.2) Processing Step 1: For each target aircraft i in the system, execute the "Updated Target Threat Level" calculation process to obtain its updated target threat level at the current moment. .

[0152] Identifiers of all target aircraft and their corresponding target threat levels These threats are combined into a vector or matrix to form the threat situation map at the current moment. This threat situation map constitutes the "state" input of the reinforcement learning model.

[0153] 2.3) Processing Step Two: Input the threat situation map into a pre-trained reinforcement learning decision model based on Deep Q-Network (DQN);

[0154] Within this DQN model, for each active sensor action, an estimated long-term cumulative reward, i.e., the Q value, is output based on the current state. In this embodiment, the active sensor actions include "pointing the camera at target aircraft 1 and zooming" and "radar scanning the airspace where target aircraft 3 is located."

[0155] The training process of this DQN model employs a composite reward function comprising three components for optimization. The reward for performing an action given the "state" input is: the expected information gain. Subtract resource scheduling costs Subtract the cost of exposure risk .

[0156] The model outputs the optimal active sensor action for the current decision period by selecting the action that maximizes its Q-value. If multiple target aircraft need to be processed, a sequence of actions ordered by priority is output.

[0157] 2.4) Processing step three: The system receives the optimal action or action sequence output by the reinforcement learning model.

[0158] For each action in the sequence, the system performs a logical judgment on the target aircraft:

[0159] If the target aircraft's updated target threat level Below the first preset threat level threshold If the target is not detected, the system will reject all active detection actions against it and only maintain passive tracking. This ensures zero exposure to low-threat targets, reflecting its "stealth".

[0160] If the target aircraft's updated target threat level Between the first preset threat level threshold Second preset threat level threshold In between, the system only allows the execution of risk exposure costs. Zero-cost, reinforcement learning model-recommended actions (including camera access). This enables low-cost, incremental identification of medium-threat targets;

[0161] If the target aircraft's updated target threat level Higher than the second preset threat level threshold If the system fails to execute the optimal action recommended by the reinforcement learning model, it will unconditionally perform the action, even if the action carries a high risk of exposure (including activating the radar). This ensures the highest priority and accuracy in identifying the most threatening targets, demonstrating "proactiveness" at critical moments.

[0162] 2.5) The final output of the process is a set of specific, executable control instructions, which are sent to the corresponding sensor hardware driver module to complete the detection and identification of the target.

[0163] The design of the reward function defines the optimization objective of the model. This is achieved by considering resource scheduling costs. and the cost of exposure risks As a negative reward.

[0164] Further explanation: The first and second preset threat level thresholds are dynamically calculated and determined based on the current system resource status and task situation; specifically:

[0165] The first dynamic threat level threshold is associated with a cost-benefit balance point for using the lowest-cost active sensors for detection; this balance point is determined by analyzing the expected information gain of the lowest-cost active sensors against the resource scheduling costs.

[0166] The second dynamic threat level threshold is related to the marginal cost-benefit balance point of using high-cost, high-risk active sensors for detection; this balance point is determined by analyzing the marginal expected information gain that high-cost, high-risk active sensors can bring compared to low-cost active sensors, and the marginal comprehensive cost required to use the sensor.

[0167] Marginal comprehensive costs include increased resource allocation costs and exposure risk costs.

[0168] Detailed implementation of the mechanism for determining dynamic threat level thresholds:

[0169] The first dynamic threat level threshold, its parameter symbol is: The physical meaning of this parameter is the "break-even point" at which, under the current system state, the expected benefits of first using a low-cost active sensor (such as a high-definition photoelectric camera) to detect a target, just offset the resource costs, are exactly offset by the threat level of the target. It is a dynamically calculated, dimensionless scalar normalized to the [0,1] interval.

[0170] The second dynamic threat level threshold, its parameter symbol is: The physical meaning of this parameter is the "marginal break-even point" at which, under the current system state, the additional expected benefits of upgrading the detection method from a low-cost sensor to a high-cost, high-risk sensor (such as millimeter-wave radar) just offset the additional overall cost, given the threat level of a target. It is a dynamically calculated, dimensionless scalar normalized to the [0,1] interval, and its value is always greater than or equal to... .

[0171] At the beginning of each decision cycle, before executing the hierarchical strategy judgment, the system first performs the following dynamic threshold calculation process:

[0172] The initial input is the current system status information, which includes at least:

[0173] Performance and cost parameters of all available active sensors, including the average expected information gain and resource scheduling cost of each sensor.

[0174] The current system stealth level is set externally, and the resulting exposure risk cost of high-risk sensors is determined accordingly. .

[0175] Preset baseline threat response values This value represents the minimum threat level that the system must respond to under any circumstances. Preset cost-benefit sensitivity coefficient. This coefficient is used to adjust the sensitivity of the threshold to cost changes.

[0176] The system identifies the sensor with the lowest resource scheduling cost per unit of expected information gain from the sensor list, which serves as the benchmark for the "lowest cost active sensor". In this embodiment, it is a high-definition photoelectric camera.

[0177] Obtain the average expected information gain of the sensor. and resource scheduling costs Calculate its cost-benefit ratio, i.e. and The ratio of .

[0178] First dynamic threat level threshold The calculation method is to multiply the calculated cost-benefit ratio by the cost-benefit sensitivity coefficient. Then add the basic threat response baseline value. To ensure that its range is within [0,1], a saturation function is applied to the calculation results to trim them to the [0,1] interval;

[0179] The system identifies sensors that can provide significantly higher information gain, but also have high resource scheduling costs and non-zero exposure risk costs. As a benchmark for "high-cost, high-risk active sensors", this embodiment uses millimeter-wave radar.

[0180] Calculate the marginal expected information gain Its value represents the average expected information gain of high-cost sensors. Subtract the average expected information gain of the lowest cost sensor .

[0181] Calculate marginal comprehensive cost Its value represents the resource scheduling cost of high-cost sensors. Compared to its current exposure risk costs The sum, minus the resource scheduling cost of the lowest-cost sensor. Calculate the marginal cost-benefit ratio, i.e. and The ratio of .

[0182] Second dynamic threat level threshold This involves multiplying the calculated marginal cost-benefit ratio by a cost-benefit sensitivity coefficient. Then add the calculated first dynamic threat level threshold. The calculation results are then processed using a saturation function.

[0183] The final output is the one that takes effect within the current decision-making cycle. and Two values.

[0184] This embodiment sets a basic threat response baseline value. 0.2: Set by the system designer based on the minimum acceptable safety margin. Cost-benefit sensitivity coefficient. It is a dimensionless adjustment factor, obtained through system-level simulation optimization, used to adjust the magnitude of the threshold change with the cost-benefit ratio. It should be noted that the saturation function is a nonlinear mapping operation whose function is to confine an arbitrary real number input value within a preset closed interval. In this embodiment, the preset interval is [0,1].

[0185] Lower Bound Determination: Compare the input value with the lower bound threshold of the interval, 0. If the input value is less than the lower bound threshold, the function outputs the lower bound threshold, 0.

[0186] Upper limit determination: If the input value is not less than the lower threshold, then it is compared with the upper threshold of the interval (1). If the input value is greater than the upper threshold, then the function outputs the upper threshold (1).

[0187] Values ​​within the range: If the input value is neither less than the lower threshold nor greater than the upper threshold (i.e., within the range [0,1]), then the function's output value is the input value itself. In this embodiment, this function is implemented using simple conditional statements or by calling the `clamp` or `clip` functions from the standard math library.

[0188] The higher the updated target threat level value, the more the comprehensive assessment result of the target aircraft's current and historical behavior by the characterization system tends to be "high-risk"; and the more perception resources and response measures are allocated to it.

[0189] When the target threat level is updated The closer the value is to 1, the more the comprehensive assessment result of the characterization system on the current and historical behavior of the target aircraft tends to be "high-risk". This directly indicates that the target has a higher priority for detection and identification, and the system should allocate more and higher quality perception resources to it and prepare to take more proactive response measures.

[0190] When the target threat level is updated The closer a value is to 0, the more the system judges the target's behavior patterns, historical records, and identity background to be consistent with expectations, and the lower the likelihood that the target poses an actual threat. This indicates that the system should monitor the target with the lowest possible resource cost to maximize the system's stealth and endurance.

[0191] Affects the updated target threat level Key input parameters include: behavioral anomaly index Historical threat assessment value and prior identity credibility .

[0192] Behavioral Abnormality Index The increase will lead to an updated target threat level. Monotonically increasing. This is a positive correlation. In the threat assessment model of this invention, the behavioral anomaly index... As a direct measure of immediate risk, the higher the value, the more the target's current behavior deviates from the normal pattern. This positive correlation design accurately maps the fundamental security logic that "abnormal behavior foreshadows potential threats," ensuring the system's ability to respond quickly to sudden risks.

[0193] Historical threat assessment value The increase will lead to an updated target threat level. Monotonically increasing; this is a positive correlation. As a cumulative record of a target's historical behavior, a higher value indicates that the target has exhibited or continues to exhibit suspicious behavior. Through a recursive update mechanism, this parameter introduces a "memory effect" into threat assessment.

[0194] When other parameters remain unchanged, prior identity credibility The increase will lead to an updated target threat level. Monotonically decreasing; this indicates a negative correlation. In the fusion formula, prior identity credibility... pass The influence of the modulatory behavior anomaly index on the item's performance. When prior identity credibility... When the number approaches 1 (for drones already registered on the whitelist), Even if the abnormal behavior index approaches 0, When the prior identity credibility is high (due to equipment failure), its contribution to the final threat level is also greatly suppressed. Conversely, when the prior identity credibility is low... When the value is 0, the impact of anomalous behavior is completely released. This negative correlation design accurately realizes the intelligent judgment that "background identity determines risk weight," avoiding misjudgment of known friendly targets and waste of resources, enabling the system to focus on truly unknown threats.

[0195] To verify the effectiveness of the core technical features of this invention, a series of experimental scenarios were designed, and key parameter data were collected through high-fidelity environment simulation. The experiment aims to compare this invention (using reinforcement learning decision-making and dynamic threshold) with two baseline methods: baseline method A2 (using rule-based decision-making, i.e., the highest level of detection is initiated when the threat level exceeds a fixed threshold) and baseline method B2 (using a simple greedy algorithm decision-making, i.e., always prioritizing the detection of the target with the highest threat level, without considering cost).

[0196]

[0197] Total cost = resource scheduling cost + exposure risk cost. Camera cost is 0.1, radar cost is 1.6 (resource cost 0.6 + exposure cost 1.0). The fixed threshold for baseline method A2 is 0.5.

[0198] Comparing Experiments 1 and 2: In Experiment 1, although the abnormal behavior index of the "whitelisted drone" was as high as 0.80, due to its prior identity credibility of 1.0, the updated target threat level calculated by the fusion algorithm of this invention was only 0.16, far below the first dynamic threat level threshold of 0.33. Therefore, this invention decided to maintain passive tracking, with a comprehensive cost of 0. In Experiment 2, the abnormal behavior index of the "unknown target" was only 0.60, but due to its prior identity credibility of 0, the calculated threat level was 0.52, higher than the first threshold. The system decided to call the camera, with a comprehensive cost of 0.10. This comparison strongly demonstrates the effectiveness of the multi-source information fusion mechanism of this invention. Compared with baseline methods A2 and B2, which use the highest-cost radar detection (comprehensive cost 1.60) for all abnormal behaviors, this invention can accurately distinguish between "faulty friendly drones" and "suspicious targets," reducing resource waste caused by misjudgment by 100% without sacrificing the ability to respond to real threats. This directly verifies the significant progress of this invention in improving system resilience and resource efficiency.

[0199] Comparing Experiments 2 and 3: In Experiment 2, the threat level was 0.52, falling between two dynamic thresholds, and the system selected a low-cost camera. In Experiment 3, an "unknown target" penetrated the defenses at high speed, causing the threat level to surge to 0.80, exceeding the first threshold, and the system decided to activate the radar. It is noteworthy that the system's stealth level in Experiment 3 was 1 (silent surveillance), resulting in a very high weighting of the radar exposure risk cost in the reinforcement learning model's reward function. Despite this, the model still chose to activate the radar. This comparison demonstrates the collaborative intelligence of the reinforcement learning decision-making model and the dynamic threshold mechanism of this invention. The model does not simply avoid costs, but rather, in the face of high threats, it can weigh the pros and cons and make the correct tactical decision of "firmly confirming even if exposed." The dynamic threshold mechanism provides a clear and reasonable boundary for this tiered response. This invention achieves a dynamic optimal balance between cost and benefit, proving its superior decision-making ability in complex adversarial environments.

[0200] In the table, despite changes in the scenario, the calculated dynamic threshold remains relatively stable. This is because the cost-effectiveness ratio of the sensors was fixed in this experiment. However, in practical applications, if some sensors in the system are damaged or energy is scarce, the dynamic threshold will automatically increase, making the system more "cautious" and prioritizing resource conservation. This cost-effective dynamic threshold design is another major innovation of this invention compared to existing fixed threshold methods. It endows the system with the ability to perceive and adapt to its own state, ensuring that it can make the most reasonable response decision under any resource conditions, greatly improving the system's robustness and task continuity.

[0201] The explanation for the "Second Dynamic Threat Level Threshold" being 1 is as follows: "Under the current situation (routine alert, high exposure risk), upgrading from cameras to radar is a costly undertaking. Radar will not be used lightly unless a target's threat level has reached the theoretical maximum of 1.0." This sets the most stringent threshold for using high-risk assets (radar). Based on the updated quantitative output of target threat levels and the above experimental analysis, the following classification criteria and corresponding operations are established for its practical application.

[0202] The criteria for classifying threats are established based on the principle of "expert experience + data analysis". Expert experience defines the basic response logic (low threat - observation, medium threat - confirmation, high threat - interception preparation), while data analysis (such as ROC curve analysis) is used to accurately locate the optimal operating point between the first dynamic threat level threshold and the second dynamic threat level threshold, in order to achieve the best balance between false positive rate and false negative rate.

[0203]

[0204] This invention deeply integrates passive micro-Doppler gait perception with reinforcement learning-based active sensor scheduling, and uses a newly created "behavioral anomaly index" as the core technological bridge to construct a complete cognitive closed loop of "passive vigilance - intent prediction - intelligent decision-making - hierarchical confirmation," generating a significant synergistic gain effect.

[0205] The synergy between stealth and proactivity: The system utilizes entirely passive "gait" perception to detect and assess target intentions. This assessment directly drives when, where, and how to conduct decisive proactive detection. This synergy enables the system to maintain long-term strategic stealth while also possessing the ability to launch precise tactical confirmation at critical moments.

[0206] Linking predictability with resource efficiency: By calculating the "abnormal behavior index" of targets, the system gains the ability to predict the "pre-ignition" of threats. This predictive ability enables the reinforcement learning decision model to free up valuable, high-power active sensor resources from ineffective monitoring of a massive number of normal targets, and to precisely focus on a very small number of targets with real potential threats, thereby achieving an order-of-magnitude improvement in the overall energy efficiency of the system.

[0207] Example 2:

[0208] Please see Figure 1 A low-altitude aircraft identification system based on multi-source data fusion, wherein the system executes a low-altitude aircraft identification method based on multi-source data fusion, comprising:

[0209] Passive sensing and feature extraction module: used to acquire the time series of micro-Doppler features formed by the target aircraft reflecting signals from non-cooperative opportunistic illumination sources through a passive sensor network;

[0210] Behavior Anomaly Prediction Module: This module is used to input the micro-Doppler feature time series into the preset behavior baseline prediction model, and calculate the behavior anomaly index, which characterizes the degree of current behavior anomaly of the target aircraft, based on the deviation between the prediction results of the preset behavior baseline prediction model and the measured values ​​of the micro-Doppler feature time series.

[0211] Intelligent decision-making and resource scheduling module: used to update the target threat level of target aircraft in real time based on the behavioral anomaly index;

[0212] Based on the updated target threat level, an active sensor resource scheduling strategy for the target aircraft is generated through a preset decision model.

[0213] Hierarchical active confirmation module: used to execute active sensor resource scheduling strategies to prioritize the detection and identification of target aircraft.

[0214] It should be noted that all calculation formulas in this application employ regression analysis, including but not limited to machine learning algorithms, to deeply analyze the collected parameters and identify their natural trends and interrelationships. Specialized software, such as Python's Scikit-learn library or the R language, is used to automatically generate mathematical models that match the data. Then, cross-validation and other methods are used to objectively evaluate the model performance, and continuous feedback and optimization are combined to ensure that the created formulas truly reflect the inherent laws of the data, thereby guaranteeing their effectiveness and accuracy. In all calculation formulas in this application, the parameters in each formula undergo dimensionless processing within a consistent range to ensure that different physical quantities are compared on the same scale; dimensionless processing techniques include, but are not limited to, min-max-normalization and Z-score standardization.

[0215] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A low-altitude aircraft identification method based on multi-source data fusion, characterized in that, The specific steps include: S1: Obtain the time series of micro-Doppler features formed by the target aircraft reflecting signals from non-cooperative opportunistic illumination sources through a passive sensor network; S2: Input the micro-Doppler feature time series into the preset behavior baseline prediction model, and calculate the behavior anomaly index, which characterizes the degree of current behavior anomaly of the target aircraft, based on the deviation between the prediction result of the preset behavior baseline prediction model and the measured value of the micro-Doppler feature time series. S3: Based on the behavioral anomaly index, update the target threat level of the target aircraft in real time; S4: Based on the updated target threat level, generate an active sensor resource scheduling strategy for the target aircraft using a preset decision model; S5: Execute the active sensor resource scheduling strategy to prioritize the detection and identification of the target aircraft.

2. The low-altitude aircraft identification method based on multi-source data fusion according to claim 1, characterized in that: The non-cooperative opportunity irradiation source signal includes at least one of the following: 5G communication base station signal, digital television broadcast signal, or Wi-Fi signal; The micro-Doppler characteristic time series specifically includes "rotor blade scintillation frequency harmonic ratio" and "fuselage vibration modulation spectrum entropy"; The preset behavior baseline prediction model is a dynamic time-series prediction model based on a "gated loop unit constrained by a physical model". The "gated loop unit of physical model constraint" is also used to output the "feature dimension importance vector"; A diagonal weighted matrix is ​​constructed using the "feature dimension importance vector" and participates in quadratic form operations. The operation result is normalized using a saturated linear unit function to obtain the "adaptive feature space Mahalanobis distance". The "adaptive feature space Mahalanobis distance" is represented as the "behavioral anomaly index".

3. The low-altitude aircraft identification method based on multi-source data fusion according to claim 2, characterized in that: The range of the behavioral abnormality index is limited to the interval [0,1]. When the calculated value of the behavioral anomaly index is closer to 0, the more closely the micro-Doppler "gait" characteristics of the target aircraft match its historical behavior patterns and embedded physical dynamics model, the lower the threat level and the lower the need to activate high-cost active sensing resources. The closer the calculated value of the behavioral anomaly index is to 1, the more significantly the micro-Doppler characteristics of the target aircraft deviate from its high-fidelity prediction baseline, the higher the threat level, and the greater the need to activate high-cost active sensing resources.

4. The low-altitude aircraft identification method based on multi-source data fusion according to claim 3, characterized in that: The specific steps for updating the target threat level are as follows: the abnormal behavior index is used as a dynamic risk weight, and it is weighted and fused with the historical threat assessment value and prior identity credibility of the target aircraft to generate the updated target threat level. The decision-making model is a reinforcement learning model; The input state of the reinforcement learning model is a threat situation map consisting of the target threat levels of all target aircraft; The reward function of the reinforcement learning model is calculated by including expected information gain, resource scheduling cost, and exposure risk cost.

5. The low-altitude aircraft identification method based on multi-source data fusion according to claim 4, characterized in that: The steps for implementing the active sensor resource scheduling strategy include: performing tiered allocation based on the target threat level; The specific logic of the hierarchical invocation is that, based on the first preset threat level threshold and the second preset threat level threshold, operations such as maintaining passive tracking, invoking non-contact sensors, or invoking high-precision sensors are executed respectively.

6. The low-altitude aircraft identification method based on multi-source data fusion according to claim 5, characterized in that: Exposure risk cost is a quantitative assessment of the risk that the enemy's electronic reconnaissance system might intercept information such as the location and operating mode of the defense system due to the sensor action of actively emitting electromagnetic signals; for any sensor action that does not generate active electromagnetic radiation, its exposure risk cost is always assigned a value of 0. For sensor actions that generate active electromagnetic radiation, the exposure risk cost is assigned a preset penalty value. The higher the concealment requirement of the task scenario, the greater the pre-set penalty value corresponding to the risk cost of exposure.

7. The low-altitude aircraft identification method based on multi-source data fusion according to claim 6, characterized in that: The first preset threat level threshold and the second preset threat level threshold are dynamically calculated and determined based on the current system resource status and task situation; specifically: The first dynamic threat level threshold is associated with a cost-benefit balance point for calling the lowest-cost active sensor for detection; this balance point is determined by analyzing the expected information gain and resource scheduling cost of the lowest-cost active sensor. The second dynamic threat level threshold is associated with the marginal cost-benefit balance point of calling high-cost, high-risk active sensors for detection; this balance point is determined by analyzing the marginal expected information gain that the high-cost, high-risk active sensors can bring compared to low-cost active sensors, and the marginal comprehensive cost required to call the sensor. Marginal comprehensive costs include increased resource allocation costs and exposure risk costs.

8. The low-altitude aircraft identification method based on multi-source data fusion according to claim 7, characterized in that: The higher the updated target threat level value, the more the system's comprehensive assessment of the target aircraft's current and historical behavior tends to be "high-risk"; and the more perception resources and response measures are allocated to it.

9. A low-altitude aircraft identification system based on multi-source data fusion, characterized in that: The system is used to execute the low-altitude aircraft identification method based on multi-source data fusion as described in any one of claims 1-8, including: Passive sensing and feature extraction module: used to acquire the time series of micro-Doppler features formed by the target aircraft reflecting signals from non-cooperative opportunistic illumination sources through a passive sensor network; Behavior anomaly prediction module: used to input the micro-Doppler feature time series into a preset behavior baseline prediction model, and calculate the behavior anomaly index, which characterizes the degree of current behavior anomaly of the target aircraft, based on the deviation between the prediction result of the preset behavior baseline prediction model and the measured value of the micro-Doppler feature time series. Intelligent decision-making and resource scheduling module: used to update the target threat level of the target aircraft in real time based on the behavioral anomaly index; Based on the updated target threat level, an active sensor resource scheduling strategy for the target aircraft is generated using a preset decision model. Hierarchical active confirmation module: used to execute the active sensor resource scheduling strategy to prioritize the detection and identification of the target aircraft.

Citation Information

Patent Citations

  • Online aircraft multi-source electric signal monitoring and fusion decision recognition method

    CN116821823A