Edge unmanned aerial vehicle identification method and system based on bi-pass fusion-time domain self-attention pulse neural network

By using a dual-channel fusion-temporal self-attention pulse neural network, the problems of limited receptive field and insufficient utilization of temporal dimension in UAV radio frequency signal recognition models are solved, achieving efficient and low-power UAV target recognition.

CN121834447APending Publication Date: 2026-04-10杭州智元研究院有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing UAV radio frequency signal recognition models rely on a single convolution operator, resulting in a limited receptive field and difficulty in capturing important discrimination information across regions or long time spans. Furthermore, the traditional Transformer structure does not fully utilize the temporal dimension information of SNNs, leading to low recognition accuracy and a high risk of misjudgment.

Method used

A dual-channel fusion-temporal self-attention pulse neural network is adopted. Through dual-band signal preprocessing and time-frequency transformation, combined with a pulse sparse dual-channel fusion module and a time-dimensional dual self-attention mechanism, the accurate capture of multi-scale feature information in the time-frequency domain and the mining of temporal-spatial features are achieved.

Benefits of technology

It improves the accuracy and energy efficiency of drone identification, adapts to resource-constrained edge recognition platforms, and achieves efficient and low-power drone target identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834447A_ABST
    Figure CN121834447A_ABST
Patent Text Reader

Abstract

The invention discloses an edge unmanned aerial vehicle identification method and system based on a bi-pass fusion-time domain self-attention pulse neural network. The method comprises the steps of dual-band signal preprocessing and time-frequency transformation, normalization and size adjustment, and target identification by a global feature extraction network. The system comprises a radio frequency signal preprocessing module and a global feature extraction network. According to the invention, high-efficiency and low-power-consumption identification of the unmanned aerial vehicle target is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) target recognition technology, and in particular relates to an edge UAV recognition method and system based on dual-channel fusion-temporal self-attention spiking neural network. Background Technology

[0002] Low-altitude airspace faces significant security risks due to frequent unauthorized drone incursions. Among various monitoring methods, the approach of analyzing the radio frequency (RF) communication characteristics between drones and their ground control devices to achieve target detection and classification has become an important direction in research and engineering practice due to its relatively low deployment cost, wide coverage, and real-time monitoring capabilities.

[0003] In recent years, data-driven methods based on Artificial Neural Networks (ANNs) have made significant progress in the field of UAV RF signal classification, significantly improving classification accuracy and system robustness. However, artificial neural networks typically rely on substantial computing and storage resources, which limits their application in mobile platforms or energy-constrained deployment scenarios.

[0004] In comparison, spiking neural networks (SNNs), with their sparse pulse-driven computation and representation methods closer to biological neural mechanisms, exhibit higher energy efficiency and interpretability, making them an attractive low-power alternative for radio frequency (RF) signal classification. However, current SNN-based RF signal classification models still have the following shortcomings:

[0005] (i) Existing SNN-type UAV radio frequency signal recognition models mainly rely on a single convolution operator, which has a limited receptive field and can only perform local feature extraction. It is easy to lose important discrimination information across regions or long time spans, thus limiting the recognition accuracy. Moreover, existing models mostly collect radio frequency signals from a single frequency band, while UAVs often operate in multiple frequency bands. Inputting only a single frequency band can easily miss the radio frequency signals released by the UAV during operation, leading to misjudgment of the UAV type by the model.

[0006] (II) With the development and maturation of time-frequency representation technology, researchers have begun to explore integrating Transformer-type architectures into specific tasks such as UAV radio frequency (RF) classification. The core objective is to utilize attention mechanisms to capture cross-scale and long-distance dependencies in the time-frequency graph. However, traditional Transformer structures used for UAV RF signal classification only introduce self-attention mechanisms in the spatial and channel dimensions, failing to fully utilize the inherent temporal dimension information of SNNs, thus limiting model performance improvement. Summary of the Invention

[0007] The purpose of this invention is to provide an edge drone recognition method and system based on dual-channel fusion-temporal self-attention spiking neural network, so as to achieve efficient and low-power recognition of drone targets.

[0008] To achieve the objectives of this invention, in one respect, this invention provides an edge drone recognition method based on a dual-channel fusion-temporal self-attention spiking neural network, comprising the following steps:

[0009] S1. Dual-band signal preprocessing and time-frequency transformation: The original baseband IQ signal of the UAV is segmented by IQ sequence to obtain each short-time frame signal. The high-frequency time-frequency feature map and the low-frequency time-frequency feature map of the UAV are obtained by short-time Fourier transform analysis of each short-time frame signal.

[0010] S2. Normalization and size adjustment: The results of the UAV high-frequency time-frequency feature map and UAV low-frequency time-frequency feature map are normalized and the size is adjusted to obtain the standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map.

[0011] S3. Target recognition using a global feature extraction network: The standardized high-frequency time-frequency feature map and the standardized low-frequency time-frequency feature map of the UAV are input into the global feature extraction network. After dual-band feature fusion and time-space feature mining, the UAV model classification result is obtained.

[0012] On the other hand, the present invention also provides a system for edge drone recognition based on dual-channel fusion-temporal self-attention spiking neural network, comprising the following modules:

[0013] The radio frequency signal preprocessing module is used to complete the acquisition, down-conversion, and short-time frequency feature extraction of UAV radio frequency signals. Specifically, it includes an IQ sequence segmentation unit and a short-time Fourier transform unit. The IQ sequence segmentation unit is used to divide the long-time baseband signal into short-time frames of fixed length. The short-time Fourier transform unit is used to perform time-frequency transformation independently on each short-time frame to generate a time-frequency feature map.

[0014] The global feature extraction module includes a pulse sparse dual-channel fusion (SSDP-Fusion) module, a time-dimension dual self-attention (T-DSSA) module, and a grouped wide convolutional sparse feedforward network (GWSFFN) module. These modules are used to fuse dual-band time-frequency feature maps, mine temporal-spatial features, and output UAV model classification results. The module adopts a sparse pulse computing architecture, which is adapted to resource-constrained edge target recognition platforms.

[0015] The significant advancement of this invention compared to existing technologies lies in:

[0016] (1) In view of the diverse information sources required for UAV radio frequency identification application scenarios, this invention proposes a pulse sparse dual-channel fusion module to fuse the time-frequency feature maps of the two input frequency bands; at the same time, an overlapping multi-scale convolution module is designed in this module to achieve accurate capture of multi-scale feature information in the time-frequency domain, effectively make up for the limited receptive field of a single convolution operator, and reduce the loss of important discrimination information.

[0017] (2) This invention introduces the Dual Self-Guided Attention (DSSA) mechanism in the pulse sparse dual-channel fusion module into the UAV radio frequency signal recognition task, and extends the self-attention mechanism of the DSSA model in the time step dimension to design and construct a time-dimensional dual self-attention model; this module enables the network to effectively identify important information in different time dimensions while analyzing spatial features, thereby enabling the SNN model to achieve higher classification accuracy in a shorter time step, taking into account both low energy consumption and high performance requirements.

[0018] To more clearly illustrate the functional characteristics and structural parameters of the present invention, further explanation is provided below in conjunction with the accompanying drawings and specific embodiments. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0020] Figure 1 This is a flowchart of the steps of the present invention;

[0021] Figure 2 This is a diagram of the SSDP-Fusion model architecture of the present invention;

[0022] Figure 3 This is a diagram of the T-DSSA model architecture of the present invention;

[0023] Figure 4 This is a diagram of the GWSFFN architecture of the present invention;

[0024] Figure 5 This is a diagram of the global feature extraction network architecture of the present invention. Detailed Implementation

[0025] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] This invention provides an edge drone recognition method based on a dual-channel fusion-temporal self-attention spiking neural network, combining... Figure 1 This includes the following steps:

[0027] S1. Dual-band signal preprocessing and time-frequency transformation: The original baseband IQ signal of the UAV is segmented by IQ sequence to obtain each short-time frame signal. The short-time Fourier transform is used to analyze each short-time frame signal to obtain the intuitive distribution characteristics of the UAV radio frequency signal in the time and frequency domains at high and low frequencies, namely the UAV high-frequency time-frequency feature map and the UAV low-frequency time-frequency feature map.

[0028] S2. Normalization and size adjustment: The results of the UAV high-frequency time-frequency feature map and UAV low-frequency time-frequency feature map are normalized and the size is adjusted to obtain the standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map.

[0029] Combination Figure 5 S3. Target recognition using a global feature extraction network: The standardized high-frequency time-frequency feature map and the standardized low-frequency time-frequency feature map of the UAV are input into the global feature extraction network. After dual-band feature fusion and time-space feature mining, the UAV model classification result is obtained.

[0030] The short-time Fourier transform analysis of S1 is specifically as follows: a window function of finite length is slid over each short-time subframe signal, a conventional Fourier transform is performed within each time window, and finally, the modulo operation is performed on the results of all time windows to obtain the time-frequency characteristic map of the signal within that time period.

[0031] In this embodiment, the total number of points for each time-frequency feature map sample is set to 10M, the sampling rate is 100M / s, the window function used is the Hamming window, and the number of points for the conventional Fourier transform is set to 4096.

[0032] S2 includes the following steps:

[0033] Step 2-1: Normalize the results of the high-frequency time-frequency feature map and the low-frequency time-frequency feature map of the UAV using the MATLAB normalize function, where the target range is selected as [0, 255].

[0034] Step 2-2: Obtain the standardized high-frequency time-frequency feature map and the standardized low-frequency time-frequency feature map of the UAV by using the normalization result with MATLAB resize(256, 960).

[0035] S3 includes the following steps:

[0036] Combination Figure 2Step 3-1: The standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map are used for feature extraction and fusion through the Spiking Sparse Dual-Path Fusion (SSDP-Fusion) module to obtain the dual-band fused feature map XF. usion_out ;

[0037] Combination Figure 3 and Figure 4 Step 3-2: Combine the dual-band fused feature map X Fusion_out The first sparse feature map X is obtained through the first round of residual block processing. res1 The residual block is composed of a Temporal-Dimensional Dual Self-Guided Attention (T-DSSA) model and a Group-Sparse Wide Convolution Feed-Forward Network (GWSFFN) module connected via membrane shortcut residual connections.

[0038] Step 3-3: Transfer the first sparse feature map X res1 By performing a downsampling operation with a sampling step size of s = 2, the first feature map X after downsampling is obtained. ds1 ;

[0039] Steps 3-4: The downsampled first feature map X... ds1 Perform the residual block processing in step 3-2 of the second round, and extract deep features by inputting two residual blocks to obtain the second sparse feature map X. res2 ;

[0040] Steps 3-5: Transfer the second sparse feature map X res2 By downsampling with a sampling step size of s = 2, the second feature map X after downsampling is obtained. ds2 ;

[0041] Steps 3-6: The downsampled second feature map X ds2 Perform the residual block processing in step 3-2 of the third round, refine the features by inputting 3 residual blocks, and obtain the third sparse feature map X. res3 ;

[0042] Steps 3-7: The third sparse feature map X... res3 Perform spatial average pooling to compress the feature dimension while retaining key information, and obtain the third feature map X after pooling. avg ;

[0043] Steps 3-8: The pooled third feature map X avgInput the classification head to obtain the final target recognition and classification result X. cls .

[0044] Combination Figure 2 The pulse sparse dual-channel fusion module of S3-1 specifically involves: processing the standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map separately through a multi-size convolution module to generate multi-scale features; then accumulating synaptic currents via a membrane shortcut and completing current integration on the neuron membrane to obtain a dual-band fused feature map X. Fusion_out The multi-size convolutional module includes convolution, feature concatenation, layer normalization, batch normalization, pooling layers, and neuron layers.

[0045] S3-1 specifically includes the following steps:

[0046] Step 3-1-1: The standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map are respectively subjected to multi-scale convolution in parallel to extract features; the first channel is sequentially passed through a 3x3 convolutional layer and a 7x7 convolutional layer, and the convolution results are concatenated to obtain the first-scale concatenated feature map X. conv1 The 3x3 convolutional kernel is used to capture narrowband transient burst signals, and the 7x7 convolutional kernel is used to identify grouping and texture structures. The second channel passes through a 3x15 convolutional layer and a 15x3 convolutional layer in sequence, and the convolutional results are stitched together to obtain the second-scale stitched feature map X. conv2 The vertical convolution kernel (3x15) detects broadband transient signals, while the horizontal convolution kernel (15x3) detects the temporal periodic patterns of UAV broadband image transmission and telemetry frequency hopping sequences.

[0047] Step 3-1-2: Stitch the first scale feature map X conv1 The execution layer performs normalization operations, followed by neuron computation and pooling operations to obtain the first pooling feature map X. pool1 ; For the second-scale spliced ​​feature map X conv2 Perform the same layer normalization, neuron computation, and pooling operations to obtain the second pooling feature map X. pool2 ;

[0048] Step 3-1-3: Transfer the first pooling feature map X pool1 Second pooling feature map X pool2 After being processed by 3x3 convolutional layers and then concatenated, the layers are normalized, processed by neurons, pooled, and then processed by another 3x3 convolutional layer before batch normalization to obtain the fused pooled feature map X. branch1 ;

[0049] Step 3-1-4: Combine the fused pooling feature map X branch1The current obtained after neuron computation and pooling, 3x3 convolutional layer processing, and batch normalization calculation is applied to the main path neurons, causing a change in their membrane potential; simultaneously, the fused pooling feature map X... branch1 As a membrane shortcut, the membrane potential is directly added to the main pathway neuron to obtain the final output membrane potential V(t). out ;

[0050] Step 3-1-5: Calculate the output membrane potential V(t). out After neural network computation, the final fusion feature pulse sequence is obtained, which serves as the output dual-band fusion feature map X of the pulse sparse dual-channel fusion module. Fusion_out .

[0051] Combination Figure 3 and Figure 4 S3-2 includes the following steps:

[0052] Step 3-2-1: Combine the dual-band fused feature map X Fusion_out The temporal enhanced feature sequence T-DSSA(X[t]) is obtained by performing a three-level operation of feature dimension splitting, temporal attention calculation, and sparse feature fusion through a time-dimensional dual self-attention model.

[0053] Step 3-2-2: The temporal enhanced feature sequence T-DSSA(X[t]) is processed by channel dimension expansion, spatial local correlation capture and sparse nonlinear activation through a grouped wide convolutional sparse feedforward network module to obtain a sparse feature map with stronger expressive power.

[0054] Step 3-2-1 specifically involves:

[0055] The specific time-dimensional dual self-attention model is shown in the following equation:

[0056] X[t]=X[i,C,B,H,W]=X[i,heads,C / heads,B×H×W],i∈{0,1,...,T / B-1};

[0057] T-Attn(X[t])=SN(DST T (X[t],X[t];f(·))*c1);

[0058] T-DSSA(X[t])=SN(DST(Attn(X[t]),X[t]; f(·))*c2);

[0059]

[0060] Where X[t] is the binary impulse feature tensor of the t-th time block, T is the total time step dimension, B is the time block scale parameter, heads is the number of heads in the multi-head attention, C is the number of feature channels, (H, W, H′, W′) are the height and width dimensions of the original feature map and the feature map after convolution transformation, respectively, and DST, DST T These are dynamic sparse transform and transpose dynamic sparse transform, respectively. SN represents a spiking neuron, and f(·) is a linear transform function. x is the average firing rate of the spiking neurons, and c1 and c2 are correction scaling factors;

[0061] The three-level operation of feature dimension splitting, temporal attention calculation, and sparse feature fusion using the aforementioned time-dimensional dual self-attention module is as follows:

[0062] The dimensionality splitting of the feature tensor: The dual-band fused feature map X... Fusion_out Based on the rules of time blocks and multi-head attention, the split features X[t] with dimensions i, heads, C / heads, B×H×W are obtained, where i is the time block index;

[0063] The time-dimensional self-attention computation involves transforming the split feature X[t] through a transpose dynamic sparse transformation DST. T To achieve feature interaction within the time block, the distribution of the transformation result is corrected by adjusting the scaling factor c1, and then the result is input into the spiking neuron SN to obtain the time attention tensor T-Attn(X[t]) in binary pulse form;

[0064] The sparse feature fusion is as follows: the time attention tensor T-Attn(X[t]) and the split feature X[t] are fused with attention weights and original feature information through dynamic sparse transformation DST, the distribution of the transformation result is corrected by the scaling factor c2, and finally the temporal enhanced feature sequence T-DSSA(X[t]) is obtained through spiking neurons SN.

[0065] Step 3-2-2 specifically involves:

[0066] The specific steps of the grouped wide convolutional sparse feedforward network module for channel dimension expansion, spatial local correlation capture, and sparse nonlinear activation are as follows:

[0067] The temporal augmentation feature sequence T-DSSA(X[t]) is subjected to pointwise convolution (scaling factor R = 4), normalized, and then calculated using neuron spiking to obtain the pointwise convolutional feature map X. n1 ;

[0068] The feature map X after pointwise convolution n1The feature map X after hierarchical convolution is obtained by using 3x3 grouped convolutional layers (number of groups G=64), followed by normalization and neuron spiking calculation. n2 ;

[0069] The feature map X after pointwise convolution n1 With the feature map X after hierarchical convolution n2 A shortcut fusion is performed, and the fused feature map is then subjected to pointwise convolution (scaling factor R = 4), normalized, and then computed through neuronal spiking to obtain a first sparse feature map X with stronger expressive power. res1 .

[0070] A system for implementing the above-described UAV target recognition method based on a spiking neural network, according to the present invention, includes the following modules:

[0071] The radio frequency signal preprocessing module is used to complete the acquisition, down-conversion, and short-time frequency feature extraction of UAV radio frequency signals. Specifically, it includes an IQ sequence segmentation unit and a short-time Fourier transform unit. The IQ sequence segmentation unit is used to divide the long-time baseband signal into short-time frames of fixed length. The short-time Fourier transform unit is used to perform time-frequency transformation independently on each short-time frame to generate a time-frequency feature map.

[0072] The global feature extraction module includes a pulse sparse dual-channel fusion (SSDP-Fusion) module, a time-dimension dual self-attention (T-DSSA) module, and a grouped wide convolutional sparse feedforward network (GWSFFN) module, which are used to fuse dual-band time-frequency feature maps, mine temporal-spatial features, and output UAV model classification results. The module adopts a sparse pulse computing architecture, which is adapted to resource-constrained edge target recognition platforms.

[0073] Example

[0074] Simulation experiments were conducted based on the above scheme:

[0075] (1) Dataset

[0076] The experimental verification in this application is based on the publicly available drone signal dataset (DroneRFa dataset, which contains 25 typical drone signals, covering indoor and outdoor scenarios, with a total of 50,000 samples, and the ratio of training set, validation set and test set is 7:1:2), to ensure that the experimental scenario is highly matched with the technical goal of "drone signal classification" in this application.

[0077] (2) Experimental setup

[0078] In the experiment, the Spiking Neural Network (SNN) model was deployed using the Spiking jelly framework, and a surrogate gradient training strategy was adopted for the SNN, with the Atan function chosen as the surrogate gradient function.

[0079] The core parameters for the training process were set as follows: the learning rate used a cosine annealing scheduling strategy, the optimizer used the Adam optimizer, and the loss function used the cross-entropy loss function. The experimental hardware environment consisted of an NVIDIA GeForce RTX 4090 graphics processor and an Intel(R) Xeon(R) Gold 6338 CPU@2.00GHz central processing unit, providing computational support for model training and inference.

[0080] To ensure the reliability of the experimental results, each experiment was independently repeated 3 times, and the final result was the average of the 3 experiments.

[0081] (3) Model performance comparison

[0082] Table 1 compares the performance of the proposed method (global feature extraction network) with representative baseline models in the field (ResNet-18, SEWnet-18) in three core metrics: classification accuracy, model size, and energy consumption per inference run. All results are the average of three repeated experiments. SSFP-Fusion and T-DSSA are the core innovations of this application. This architecture is specifically adapted to the multi-scale fusion requirements of sparse time-frequency features of UAV pulses, and is the key support for the performance improvement of the proposed method.

[0083] The specific rules for energy consumption statistics are as follows: convolution operations, fully connected operations, and matrix multiplication operations in artificial neural networks are all quantified as floating-point operations (FLOPs), and energy consumption is calculated based on the number of multiply-accumulate operations; the spiking interaction and activation processes in spiking neural networks (SNNs) are defined as synaptic operations (SOPs), and energy consumption is calculated based on the cumulative number of operations. This application uses the above rules to quantitatively analyze the energy consumption of the models, ensuring the objectivity and comparability of energy consumption comparisons between different models.

[0084] To ensure the comparability of energy consumption comparison results with existing studies, this application adopts the Horowitz energy consumption estimation benchmark as a unified evaluation standard. This benchmark clearly stipulates that the energy consumption for each 32-bit floating-point (FP) multiply-accumulate operation (MAC) is 4.6 pJ; and the energy consumption for each accumulation operation (AC) is 0.9 pJ, as shown in the table below:

[0085] Model Accuracy Size Energy Resnet-18 0.9773 42.7M 8.372mJ SEWnet-18 0.9804 42.7M 2.815mJ Global Feature Extraction Network 0.9916 22.3M 5.531mJ

[0086] Experimental results show that, in the test scenario of the DroneRFa dataset, the method proposed in this application achieves the highest classification accuracy of 99.16%, which is 1.43 percentage points higher than ResNet-18 and 1.12 percentage points higher than SEWnet-18, fully verifying the superiority of the network structure proposed in this application in UAV signal feature extraction and classification tasks.

[0087] In terms of model size, the model used in this application has a parameter size of approximately 22.3MB, which is 47.8% smaller than ResNet-18 and SEWnet-18 (both with a parameter size of approximately 42.7MB). This verifies that the network structure proposed in this application can improve classification accuracy while having better parameter efficiency and is more suitable for lightweight deployment requirements.

[0088] In terms of energy consumption, the proposed method consumes 5.531 millijoules (mJ) per inference, which is 33.9% lower than ResNet-18 (8.372 mJ per inference). This fully demonstrates the significant advantage of the proposed solution in low energy consumption performance and can meet the low latency and low power consumption application requirements of UAV signal real-time processing.

[0089] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0090] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An edge drone recognition method based on dual-channel fusion-temporal self-attention spiking neural network, characterized in that, Includes the following steps: S1. Dual-band signal preprocessing and time-frequency transformation: The original baseband IQ signal of the UAV is segmented by IQ sequence to obtain each short-time frame signal. The high-frequency time-frequency feature map and the low-frequency time-frequency feature map of the UAV are obtained by short-time Fourier transform analysis of each short-time frame signal. S2. Normalization and size adjustment: The results of the UAV high-frequency time-frequency feature map and UAV low-frequency time-frequency feature map are normalized and the size is adjusted to obtain the standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map. S3. Target recognition using a global feature extraction network: The standardized high-frequency time-frequency feature map and the standardized low-frequency time-frequency feature map of the UAV are input into the global feature extraction network. After dual-band feature fusion and time-space feature mining, the UAV model classification result is obtained.

2. The edge drone recognition method based on dual-channel fusion-temporal self-attention spiking neural network according to claim 1, characterized in that, The short-time Fourier transform analysis of S1 is specifically as follows: a window function of finite length is slid over each short-time subframe signal, a conventional Fourier transform is performed within each time window, and finally, the modulo operation is performed on the results of all time windows to obtain the time-frequency characteristic map of the signal within that time period.

3. The edge drone recognition method based on dual-channel fusion-temporal self-attention spiking neural network according to claim 1, characterized in that, S2 includes the following steps: Step 2-1: Normalize the results of the high-frequency time-frequency feature map and the low-frequency time-frequency feature map of the UAV using the matlab normalize function; Step 2-2: Obtain the standardized high-frequency time-frequency feature map and the standardized low-frequency time-frequency feature map of the UAV by using the normalization result through MATLAB resize.

4. The edge drone recognition method based on dual-channel fusion-temporal self-attention spiking neural network according to claim 1, characterized in that, S3 includes the following steps: Step 3-1: Extract and fuse the standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map through the pulse sparse dual-channel fusion module to obtain a dual-band fused feature map; Step 3-2: The dual-band fused feature map is processed through the first round of residual block processing to obtain the first sparse feature map; the residual block is composed of a time-dimensional dual self-attention module and a grouped wide convolutional sparse feedforward network module connected by a membrane shortcut residual connection; Step 3-3: The first sparse feature map is downsampled to obtain the downsampled first feature map; Step 3-4: Perform the second round of residual block processing in step 3-2 on the downsampled first feature map, and extract deep features by inputting two residual blocks to obtain the second sparse feature map; Steps 3-5: Downsample the second sparse feature map to obtain the downsampled second feature map; Step 3-6: Perform the third round of residual block processing in step 3-2 on the downsampled second feature map, refine the features by inputting 3 residual blocks, and obtain the third sparse feature map; Steps 3-7: Perform spatial average pooling on the third sparse feature map to compress the feature dimension and retain key information to obtain the pooled third feature map; Steps 3-8: Input the pooled third feature map into the classification head to obtain the final target recognition classification result.

5. The edge drone recognition method based on dual-channel fusion-temporal self-attention spiking neural network according to claim 4, characterized in that, The pulse sparse dual-channel fusion module of S3-1 specifically involves processing the standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map separately through a multi-size convolution module to generate multi-scale features. Then, synaptic currents are accumulated by a membrane shortcut, and current integration is completed on the neuron membrane to obtain a dual-band fused feature map. The multi-size convolution module includes convolution, feature splicing, layer normalization, batch normalization, pooling layer, and neuron layer.

6. The edge drone recognition method based on dual-channel fusion-temporal self-attention spiking neural network according to claim 5, characterized in that, S3-1 specifically includes the following steps: Step 3-1-1: The standardized UAV high-frequency time-frequency feature map and the standardized UAV low-frequency time-frequency feature map are respectively processed by multi-scale convolution in parallel to extract features; the first channel is sequentially processed by a 3x3 convolutional layer and a 7x7 convolutional layer, and the convolution results are concatenated to obtain a first-scale concatenated feature map; the second channel is sequentially processed by a 3x15 convolutional layer and a 15x3 convolutional layer, and the convolution results are concatenated to obtain a second-scale concatenated feature map; Step 3-1-2: Perform layer normalization on the first scale spliced ​​feature map, and then perform neuron calculation and pooling operations to obtain the first pooled feature map; perform the same layer normalization, neuron calculation and pooling operations on the second scale spliced ​​feature map to obtain the second pooled feature map. Step 3-1-3: After processing the first pooling feature map and the second pooling feature map through 3x3 convolutional layers respectively, they are concatenated, and after layer normalization, neuron calculation and pooling, they are processed through 3x3 convolutional layers again and then batch normalized to obtain the fused pooling feature map. Step 3-1-4: The current obtained after the fusion pooling feature map is processed by neuron calculation and pooling, 3x3 convolutional layer processing, and batch normalization calculation is applied to the main path neuron to change its membrane potential; at the same time, the fusion pooling feature map is added to the membrane potential of the main path neuron as a membrane shortcut to obtain the final output membrane potential. Step 3-1-5: The output membrane potential is processed by neurons to obtain the fusion feature pulse sequence, which is used as the output dual-band fusion feature map of the pulse sparse dual-channel fusion module.

7. The edge drone recognition method based on dual-channel fusion-temporal self-attention spiking neural network according to claim 4, characterized in that, S3-2 includes the following steps: Step 3-2-1: The dual-band fused feature map is subjected to a three-level operation of feature dimension splitting, time attention calculation, and sparse feature fusion through a time-dimensional dual self-attention module to obtain a time-series enhanced feature sequence; Step 3-2-2: The temporal enhancement feature sequence is processed by channel dimension expansion, spatial local correlation capture and sparse nonlinear activation through a grouped wide convolutional sparse feedforward network module to obtain a sparse feature map with stronger expressive power.

8. The edge UAV recognition method based on a dual-channel fusion time-domain self-attention spiking neural network according to claim 7, characterized in that, Step 3-2-1 specifically involves: The specific time-dimensional dual self-attention model is shown in the following equation: X[t]=X[i,C,B,H,W]=X[i,heads,C / heads,B×H×W],i∈{0,1,...,T / B-1}; T-Attn(X[t])=SN(DST T (X[t],X[t];f(·))*c1); T-DSSA(X[t])=SN(DST(Attn(X[t]),X[t]; f(·))*c2); Where X[t] is the binary impulse feature tensor of the t-th time block, T is the total time step dimension, B is the time block scale parameter, heads is the number of heads in the multi-head attention, C is the number of feature channels, (H, W, H′, W′) are the height and width dimensions of the original feature map and the feature map after convolution transformation, respectively, and DST, DST T These are dynamic sparse transform and transpose dynamic sparse transform, respectively. SN represents a spiking neuron, and f(·) is a linear transform function. x is the average firing rate of the spiking neurons, and c1 and c2 are correction scaling factors; The three-level operation of feature dimension splitting, temporal attention calculation, and sparse feature fusion using the aforementioned time-dimensional dual self-attention model is as follows: The feature tensor is split in dimension: the dual-band fused feature map is split according to the rules of time blocks and multi-head attention to obtain the split feature X[t] with dimension i, heads, C / heads, B×H×W, where i is the time block index; The time-dimensional self-attention computation involves transforming the split feature X[t] through a transpose dynamic sparse transformation DST. T To achieve feature interaction within the time block, the distribution of the transformation result is corrected by adjusting the scaling factor c1, and the time attention tensor T-Attn(X[t]) in binary pulse form is obtained from the input spiking neuron SN. The sparse feature fusion is described as follows: the temporal attention tensor T-Attn(X[t]) and the split feature X[t] are fused with attention weights and original feature information through dynamic sparse transformation DST. The distribution of the transformation result is corrected by adjusting the scaling factor c2. Finally, the temporal enhanced feature sequence T-DSSA(X[t]) is obtained through spiking neurons SN.

9. The edge UAV recognition method based on dual-channel fusion-temporal self-attention spiking neural network according to claim 7, characterized in that, Step 3-2-2 specifically involves: The specific steps of the grouped wide convolutional sparse feedforward network module for channel dimension expansion, spatial local correlation capture, and sparse nonlinear activation are as follows: The temporal augmentation feature sequence T-DSSA(X[t]) is subjected to pointwise convolution and normalization, followed by neuronal spiking to obtain the pointwise convolutional feature map X. n1 ; The feature map X after pointwise convolution n1 The feature map X after hierarchical convolution is obtained by using 3x3 grouped convolutional layers, followed by normalization and neuron spiking calculation. n2 ; The feature map X after pointwise convolution n1 With the feature map X after hierarchical convolution n2 By performing shortcut fusion, the fused feature map is processed through pointwise convolution and normalization, and then calculated through neuronal spiking to obtain a sparse feature map with stronger expressive power.

10. A system for edge drone recognition based on a dual-channel fusion-temporal self-attention spiking neural network according to any one of claims 1-9, characterized in that, Includes the following modules: The radio frequency signal preprocessing module is used to complete the acquisition, down-conversion, and short-time frequency feature extraction of UAV radio frequency signals. Specifically, it includes an IQ sequence segmentation unit and a short-time Fourier transform unit. The IQ sequence segmentation unit is used to divide the long-time baseband signal into short-time frames of fixed length. The short-time Fourier transform unit is used to perform time-frequency transformation independently on each short-time frame to generate a time-frequency feature map. The global feature extraction module includes a pulse sparse dual-channel fusion module, a time-dimensional dual self-attention module, and a grouped wide convolutional sparse feedforward network module. It is used to fuse dual-band time-frequency feature maps, mine time-series-spatial features, and output UAV model classification results. The module adopts a sparse pulse computing architecture, which is adapted to resource-constrained edge target recognition platforms.

Citation Information

Cited By

  • A GIS defect identification method based on holographic spectrum resonance network

    CN122262656A