Aerial biological recognition method and device based on echo fluctuation and multi-scale characteristics

By extracting the amplitude sequence depth characteristics and multi-scale time-frequency depth characteristics of aerial organisms and performing feature fusion, the problem of low accuracy of aerial biometric identification in the prior art is solved, and a higher accuracy of birds and insects is achieved.

CN119986643AActive Publication Date: 2025-05-13BEIJING INST OF TECH +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510459246.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively distinguish between birds and insects in aerial biometrics, especially due to the similarity of their radar echo signals and the complex wing-flapping patterns, resulting in low recognition accuracy.

Method used

The recognition method based on echo fluctuations and multi-scale features is adopted, and the recognition accuracy of the aerial biological wing pattern is improved by extracting the depth features of amplitude sequences and multi-scale time-frequency depth features, and fusion of features.

Benefits of technology

Through multi-scale feature fusion technology, the micro-moving characteristics and wing-flapping patterns of aerial organisms can be more effectively captured, significantly improving the recognition accuracy of birds and insects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119986643A_ABST
    Figure CN119986643A_ABST
Patent Text Reader

Abstract

The invention provides an aerial biological recognition method and device based on echo fluctuation and multi-scale features, and belongs to the field of low-altitude target radar detection. The method comprises the following steps: extracting amplitude sequence depth features S1 according to track sequence data of aerial creatures; carrying out multi-scale Fourier transform on the track sequence data, and then slicing and recombining to obtain m spliced time-frequency diagrams; s2, performing time-frequency feature extraction on each spliced time-frequency graph to obtain a multi-scale time-frequency feature, capturing time dependence through a recursive network, completing time-frequency time feature extraction, and obtaining a multi-scale time-frequency depth feature; and carrying out feature fusion on the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2, and carrying out aerial biological flapping mode recognition based on the fused features. By using the method, the identification accuracy of the aerial biological flapping mode can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of low-altitude target radar detection, and in particular to an aerial biometric identification method and device based on echo fluctuation and multi-scale features. Background Art

[0002] Migration behavior is an adaptive strategy evolved by organisms to adapt to climate and food changes, which has a wide and profound impact on the ecosystem. Migratory species are an important part of biodiversity, and they play a key role in maintaining the balance and function of ecosystems. For example, migratory birds and insects not only have important ecological value themselves, but also affect other species and ecological processes through their migration behavior, and have a profound impact on human life. However, due to the influence of their flight altitude and body size, there is a lack of effective means of observing aerial organisms. Radar is a powerful tool for the migration of aerial biological targets due to its advantages of all-day, all-weather and large detection range.

[0003] Different types of radars have been developed to study aerial biomigration, including scanning insect radar, vertical insect radar, harmonic radar, bird radar, weather radar network and other systems. There are a lot of studies ranging from large-scale range monitoring to small-scale individual tracking measurements.

[0004] An important progress in insect radar is the development of vertical viewing radar (VLR), which is able to detect the behavioral and biological parameters of migratory insects, although for small insects weighing less than 10 mg, the detection probability is low due to the difficulty of simultaneously performing high-speed continuous sampling and long-term integration. However, by developing insect radars with high range resolution, combined with long-term integration and detection methods, the detection performance of small targets can be effectively improved and the detection range can be increased.

[0005] In the field of bird-detecting radar, although early research focused on radar cross-section measurements, recent research has begun to focus on how to use radar technology to distinguish between echo signals from birds and insects. For example, by analyzing the differential reflectivity of radar echoes, echo signals caused by atmospheric turbulence and insects can be distinguished. In addition, machine learning methods have been proposed to distinguish between echo signals from birds and insects from NEXRAD radar echoes, which shows the great potential of using modern computing technology to improve radar data processing capabilities.

[0006] In order to better count the number ratio of insects and birds, many scholars have conducted in-depth research on the classification of insects and birds. Some studies use the intensity of radar echo signals to identify birds and other targets. However, for a large number of small birds such as passerines, the radar echo intensity and the echo intensity range of large insects are basically the same, and they are basically indistinguishable. Since the flying speed of birds is generally greater than that of insects, some studies use the speed of the target relative to the ground to distinguish between insects and birds. However, the ground speed of birds and insects is affected by wind speed and wind direction, especially the latter. When the wind direction and wind speed information are uncertain, the ground speed of birds and insects has a large overlapping range, which cannot be used as a good means of distinction. The change pattern of echo intensity over time is also the focus of the study of bird and insect flight behavior, reflecting the shape change of the target in the radar beam. The typical pattern of bird echo signals is the periodic fluctuation of the signal, and its fluctuation frequency is generally the frequency of wing flapping. This fluctuation is caused by the changes in the overall body during the wing flapping process. Some studies have pointed out the types of birds corresponding to the three main wing flapping patterns of birds: continuous wing flapping, intermittent flapping and irregular flapping. Insect wings are generally very small relative to their bodies and do not cause significant changes in radar echoes. Their echo signals show a typical sine-like pattern. This distinction between different motion patterns under radar observation is widely used in the identification of birds and insects. Various machine learning methods, including random forests, support vector machines, and neural networks, have been applied to the identification of birds and insects. The input signal mainly includes the original echo and the time-frequency diagram after Fourier transform or wavelet transform.

[0007] In summary, the main challenges in identifying the wing flapping patterns of aerial organisms are as follows: 1. There are many types of aerial biological targets with various wing-flapping patterns. It is difficult for a single feature to effectively distinguish biological targets with different wing-flapping patterns, resulting in low accuracy in wing-flapping pattern recognition.

[0008] 2. The frequency scale of wing flapping of aerial organisms spans a large range. Traditional methods use transformation algorithms with fixed parameters, which cannot effectively describe the typical characteristics of aerial targets and thus affect the accuracy of wing flapping pattern recognition. Summary of the invention

[0009] In view of this, the present invention provides an aerial creature identification method and device based on echo fluctuations and multi-scale features, which can improve the accuracy of aerial creature wing flapping pattern identification.

[0010] In order to solve the above technical problems, the present invention is implemented as follows.

[0011] An aerial biometric recognition method based on echo fluctuation and multi-scale features, comprising: Step 1: Extract the amplitude sequence depth feature S based on the track sequence data of aerial organisms 1, characterizes the fluctuation of radar cross section RCS over time during target movement; Step 2: According to the vibration frequency range of the wing-flapping action of the aerial organisms, the track sequence data is Fourier transformed using a short-time Fourier transform window length of n scales, and n time-frequency graphs are generated for each target; the n time-frequency graphs are sliced, and each time-frequency graph obtains m slices; the slices at the same position in different time-frequency graphs are taken and re-spliced ​​to obtain m spliced ​​time-frequency graphs; each spliced ​​time-frequency graph represents the target micro-motion characteristics with different frequency resolutions in the same time window; n and m are positive integers greater than 2; The time-frequency features of each spliced ​​time-frequency graph are extracted to obtain multi-scale time-frequency features. Then, the time dependency is captured through a recursive network to complete the time-frequency time feature extraction and obtain the multi-scale time-frequency depth feature S 2 ; Step 3: Deep feature S of the amplitude sequence 1 and multi-scale time-frequency depth features S 2 Perform feature fusion; Step 4: Use the fusion features of step 3 to identify the wing flapping patterns of aerial organisms.

[0012] Preferably, in step 2, the range of the n-scale short-time Fourier transform window lengths is selected to be 4.8 ms-48 ms.

[0013] Preferably, in step 2, extracting time-frequency features from each spliced ​​time-frequency graph to obtain multi-scale time-frequency features comprises: performing operations from step 21 to step 23 for each spliced ​​time-frequency graph, and the time-frequency feature extraction results of each spliced ​​time-frequency graph constitute the multi-scale time-frequency features; Step 21: After the spliced ​​time-frequency graph undergoes a layer of convolution processing, it enters four cascaded residual blocks for processing; the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block; the outputs of the first three residual blocks are adjusted to have the same number of channels as the number of channels of the feature extraction result of the fourth residual block through a 1×1 convolution block, and enter the adaptive sliding window pooling layer as the feature extraction result of the corresponding residual block, and on the other hand, serve as the input of the next level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling layer; Step 22: The adaptive sliding window pooling layer adjusts the feature map sizes of the feature extraction results of the four residual blocks to be consistent, and stretches them into one-dimensional vectors using full connections; Step 23: Add and fuse the four one-dimensional vectors to obtain the time-frequency feature extraction result.

[0014] Preferably, the residual block includes two convolution blocks connected in series, namely a first convolution block and a second convolution block; the output of the first convolution block is used as the input of the second convolution block; Each convolution block consists of two convolution layers, namely the first convolution layer and the second convolution layer. The first convolution layer convolves the input x1 of the convolution block to obtain the feature ; Use nonlinear activation function to get the features from the first convolutional layer Perform nonlinear mapping to obtain features , and then the obtained features Input to the second convolutional layer, use the second convolutional layer to feature Perform convolution to obtain features ; The characteristics Add it to the input x1 of the convolution block to get the output of the convolution block.

[0015] Preferably, in step 2, capturing time dependency through a recursive network to complete time-frequency-time feature extraction includes: A two-layer stacked bidirectional long short-term memory Bi-LSTM neural network is used as the recurrent network; Each layer of Bi-LSTM neural network includes m LSTM hidden layers, which process the time-frequency feature extraction results of m spliced ​​time-frequency graphs; each LSTM hidden layer in the same layer of Bi-LSTM neural network learns the temporal correlation of time-frequency features from the positive and reverse directions in turn; The second layer of Bi-LSTM neural network outputs multi-scale time-frequency depth features S 2 .

[0016] Preferably, in step 3, the depth feature S of the amplitude sequence 1 and multi-scale time-frequency depth features S 2 The feature fusion is: Step S31: Depth feature S of amplitude sequence 1 Perform enhanced pooling processing; Step S32: multi-scale time-frequency depth feature S 2 Perform enhanced pooling processing; Step S33: transform the amplitude sequence depth feature S 1 and multi-scale time-frequency depth features S 2 Splicing: The spliced ​​features are processed by channel attention mechanism and spatial attention mechanism, and then pooled and nonlinearly mapped; Step S34: concatenate the results of step 31, step 32 and step 33 to obtain fusion features.

[0017] Preferably, in step 1, the extracted amplitude sequence depth feature S 1 The operation is implemented using 4 layers of stacked 1-D multi-channel convolution operations.

[0018] Preferably, the step 2 further includes performing intensity normalization and scale normalization on the time-frequency diagram; and the step 1 further includes normalizing the amplitude sequence of the track sequence data.

[0019] The present invention also provides an aerial biometric identification device based on echo fluctuation and multi-scale features, comprising: an amplitude sequence depth feature extraction unit, a multi-scale time-frequency depth feature extraction unit, a fusion unit and an identification unit; The amplitude sequence depth feature extraction unit is used to extract the amplitude sequence depth feature S according to the track sequence data of the aerial organisms. 1 , characterizes the fluctuation of RCS over time during target motion; The multi-scale time-frequency depth feature extraction unit includes a splicing module and an extraction module; The splicing module is used to perform Fourier transform on the track sequence data according to the vibration frequency range of the wing flapping action of the aerial creature using a short-time Fourier transform window length of n scales, and generate n time-frequency graphs for each target; slice the n time-frequency graphs, and obtain m slices for each time-frequency graph; take the slices at the same position in different time-frequency graphs and re-splice them to obtain m spliced ​​time-frequency graphs; each spliced ​​time-frequency graph represents the target micro-motion characteristics with different frequency resolutions in the same time window; n and m are positive integers greater than 2; The extraction module is used to extract the time-frequency features of each spliced ​​time-frequency graph to obtain multi-scale time-frequency features, and then capture the time dependency through a recursive network to complete the time-frequency time feature extraction and obtain the multi-scale time-frequency depth feature S 2 ; The fusion unit is used to fusion the amplitude sequence depth feature S 1 and multi-scale time-frequency depth features S 2 Perform feature fusion; The recognition unit is used to recognize the wing flapping pattern of aerial creatures using the fusion features obtained by the fusion unit.

[0020] Preferably, the extraction module includes a multi-scale time-frequency feature extraction module and a time-frequency feature extraction module; The multi-scale time-frequency feature extraction module includes m extraction channels, each of which is used to process one of the m spliced ​​time-frequency images; each extraction channel includes an initial convolution module, four cascaded residual blocks, three 1×1 convolution blocks, an adaptive sliding window pooling module and a feature splicing module; the four residual blocks are respectively the first residual block, the second residual block, the third residual block and the fourth residual block; The spliced ​​time-frequency graph of the extracted channel enters the initial convolution module, and the output of the initial convolution module enters the first residual block; the outputs of the first three residual blocks are adjusted to the same number of channels as the number of channels of the feature extraction result of the fourth residual block through the corresponding 1×1 convolution block, and enter the adaptive sliding window pooling module as the feature extraction result of the corresponding residual block, and on the other hand, serve as the input of the next level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling module; The adaptive sliding window pooling module adjusts the feature map sizes of the feature extraction results of the four residual blocks to be consistent, and stretches them into one-dimensional vectors using full connections; The feature concatenation module adds and fuses the four one-dimensional vectors to obtain the time-frequency feature extraction result of this extraction channel; The time-frequency time feature extraction module includes two stacked Bi-LSTM neural networks; each layer of the Bi-LSTM neural network includes m LSTM hidden layers, and each LSTM hidden layer learns the temporal correlation of the time-frequency features from the positive and reverse directions in sequence.

[0021] Beneficial effects: (1) This invention abandons the traditional method of directly inputting the time-frequency graph into the network, and innovatively intercepts and splices different time-frequency graphs and re-splices them into a new frequency feature graph, thereby obtaining the target micro-motion features with different frequency resolutions in the same time window as the input features of the time-frequency feature extraction module. Compared with the traditional method of directly using the time-frequency graph as input processing, this method introduces multi-frequency resolution information features and splices the frequency features of the same time window, which can enable the algorithm to better capture the key frequency features of the target and improve the accuracy of the recognition of the wingbeat pattern of aerial organisms.

[0022] (2) In a preferred embodiment, the feature fusion of the present invention consists of two parts. The first part fully exploits the different scale features of each type of features through different pooling strategies, and the second part captures the spatial similarity of the two types of features and the feature importance of different channels through the space-channel attention mechanism. The dual-channel feature fusion can further explore the correlation between the two types of features and perform feature fusion to improve the accuracy of biological classification.

[0023] (3) Since the wing flapping patterns of aerial biological targets are complex and their vibration frequencies cover a wide range, in order to better extract the micro-motion characteristics of aerial biological targets, in a preferred embodiment, a preferred value range of the multi-scale short-time Fourier transform window length is designed, the shortest of which is 4.8 ms, corresponding to a frequency of 208 Hz, which can fully extract the micro-motion characteristics of rapidly vibrating insect targets.

[0024] (4) When extracting multi-scale time-frequency depth features, the present invention takes into account that the data set of the present invention is relatively small. Therefore, in the process of selecting the network structure, a basic architecture with relatively few structural layers, better network performance and less prone to training crashes is selected. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 (a) shows the amplitude series and time-frequency diagram of typical insects; Figure 1(b) shows the amplitude series and time-frequency diagram of typical birds; Figure 1(c) shows the amplitude series and time-frequency diagram of other targets; Figure 2 It is a flow chart of the wing flapping pattern recognition method based on echo fluctuation and multi-scale feature fusion of the present invention; Figure 3 It is the overall framework of the algorithm of the present invention; Figure 4 This is an illustration of the splicing of multi-scale time-frequency diagrams; Figure 5 It is the structural diagram of the multi-scale time-frequency feature extraction module; Figure 6 It is the structural diagram of the time-frequency feature extraction module; Figure 7 It is a fusion unit architecture; Figure 8 (a) is a structural diagram of the enhanced pooling module in the fusion unit; Figure 8 (b) shows the structure of the channel space attention module in the fusion unit; Figure 9 (a) shows the amplitude sequence before normalization; Figure 9 (b) shows the normalized amplitude sequence; Figure 10 (a) is the time-frequency diagram before normalization; Figure 10 (b) is the normalized time-frequency diagram; Fig.11 It is a schematic diagram of the wing flapping pattern recognition device based on echo fluctuation and multi-scale feature fusion of the present invention. DETAILED DESCRIPTION

[0026] The present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0027] The present invention provides an aerial biometric recognition scheme based on echo fluctuation and multi-scale features. The basic idea is to extract two depth features, one is the amplitude sequence depth feature S 1 , characterizing the fluctuation of radar cross-section (RCS) over time during target motion; second, multi-scale time-frequency depth feature S 2 , characterize the dynamic changes of biological Doppler at different time-resolutions; fuse the two features to identify the target type, thereby improving the accuracy of aerial biological classification.

[0028] Figure 1 (a) and Figure 1 (b) are the amplitude sequences and time-frequency diagrams of typical insects, birds and other targets, respectively. It can be seen that different types of aerial creatures have different wing flapping patterns, and the amplitude sequences and time-frequency diagrams of the echoes are quite different. When extracting features, especially time-frequency features, the present invention intercepts and splices different time-frequency diagrams to obtain target micro-motion features with different frequency resolutions under the same time window, which are used as input features of the time-frequency feature extraction operation. This makes the time-frequency feature extraction operation not fixed at a certain scale, which can prompt the algorithm to better capture the key frequency features of the target and effectively describe the characteristics of the aerial target. Based on this, a new network architecture is designed to realize the recognition of aerial biological targets, thereby improving the accuracy of aerial biological recognition.

[0029] Figure 2 is a flow chart of the aerial biometric identification method based on echo fluctuation and multi-scale features of the present invention, Figure 3 The algorithm framework of this method is as follows: Figure 2 and Figure 3 As shown, the method comprises the following steps: Step 1: Extract the amplitude sequence depth feature S based on the track sequence data of aerial organisms 1 , characterizes the fluctuation of RCS over time during target motion; Step 2: Select n-scale short-time Fourier transform window lengths according to the window length range determined by the frequency range of the wing vibration of aerial organisms. Use these n-scale short-time Fourier transform window lengths to perform Fourier transform on the track sequence data, and generate n time-frequency graphs for each target; slice the n time-frequency graphs, and obtain m slices for each time-frequency graph; take the slices at the same position in different time-frequency graphs and re-join them to obtain m spliced ​​time-frequency graphs. The splicing diagram is as follows: Figure 4 Each spliced ​​time-frequency diagram represents the target micro-motion characteristics with different frequency resolutions in the same time window; n and m are positive integers greater than 2.

[0030] The time-frequency features of each spliced ​​time-frequency graph are extracted to obtain multi-scale time-frequency features. Then, the time dependency is captured through a recursive network to complete the time-frequency time feature extraction and obtain the multi-scale time-frequency depth feature S 2 .

[0031] Step 3: Deep feature S of the amplitude sequence 1 and multi-scale time-frequency depth features S 2 Perform feature fusion.

[0032] Step 4: Use the fusion features of step 3 to identify the wing flapping patterns of aerial organisms.

[0033] Each step is described in detail below.

[0034] In a preferred embodiment, see Figure 3 In the upper part of the step 1, the implementation method of extracting the depth feature of the amplitude sequence is preferably: using a 4-layer stacked 1-D multi-channel convolution operation to extract the depth feature S of the amplitude sequence 1 Specifically, CNN is used to extract the deep features of the target amplitude sequence. For a typical 1-D multi-channel convolution operation, a residual convolution block is constructed based on the 1-D convolution. Each residual block consists of two convolution blocks. The convolution block consists of the first convolution layer, the activation function, and the second convolution layer. The output of a convolution block is the result of the concatenation of the output of the second convolution layer inside it and the input of this convolution block:

[0035] in: x is the convolution block input, is the output of the second convolutional layer, is the output of the convolutional block.

[0036] In a preferred embodiment, the specific implementation process of step 2 specifically includes steps 20 to 24. Step 20 is slicing and re-joining the time-frequency graph, steps 21 to 23 are multi-scale time-frequency feature extraction, and step 24 is time-frequency temporal feature extraction, and finally a multi-scale time-frequency depth feature S is obtained. 2 Specifically: Step 20: Reassembly of the time-frequency diagram.

[0037] Since the motion patterns of aerial biological targets are complex and their vibration frequencies are widely distributed, in order to better extract the micro-motion features of aerial biological targets, the present invention designs short-time Fourier transform window lengths of different scales, the shortest of which is 4.8ms, corresponding to a frequency of 208Hz, which can fully extract the micro-motion features of fast-vibrating insect targets. In a preferred implementation scheme, the window length is 3.2ms, and a total of n=14 window lengths are selected from 4.8ms-48ms to perform short-time Fourier transform on the signal, and 14 time-frequency graphs are generated for each target. The present invention abandons the traditional method of directly inputting the time-frequency graph into the network, and innovatively intercepts and splices different time-frequency graphs, each of which intercepts a time window of 8 points in turn, and re-splices them into new, for example, m=14 frequency feature graphs, thereby obtaining target micro-motion features with different frequency resolutions under the same time window, as input features of the multi-scale time-frequency deep feature extraction step, and uses this to design a new network architecture to realize the recognition of aerial biological wing flapping patterns.

[0038] Next, the operations of step 21 to step 23 are performed for each spliced ​​time-frequency graph, and the time-frequency feature extraction results of each spliced ​​time-frequency graph constitute the multi-scale time-frequency feature. Figure 5 .

[0039] Step 21: After the spliced ​​time-frequency graph undergoes a layer of convolution processing, it enters four cascaded residual blocks for processing; the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block; the outputs of the first three residual blocks are adjusted through a 1×1 convolution block to have the same number of channels as the number of channels of the feature extraction results of the fourth residual block, and enter the adaptive sliding window pooling layer as the feature extraction results of the corresponding residual blocks, and on the other hand, serve as the input of the next level residual block; the feature extraction results of the fourth residual block directly enter the adaptive sliding window pooling layer.

[0040] Specifically, the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block (corresponding to residual blocks 1, 2, 3, and 4 in the figure). The feature extraction result of the first residual block is F1, the feature extraction result of the second residual block is F2, the feature extraction result of the third residual block is F3, and the feature extraction result of the fourth residual block is F4.

[0041] Among them, the two convolution blocks included in the first residual block are the first convolution block and the second convolution block, and the feature extraction result of the first residual block is F1, and the acquisition method is: (1) Use the first convolutional layer of the first convolutional block to convolve the input x1 of the first convolutional block to obtain the feature , where L represents the number of feature channels, Indicates the first i channels, i =1,2,…,L, It is expressed by the following formula:

[0042] in, represents the convolution operation, Representative i convolution kernels, Representatives i The offset of the convolution kernel; (2) Use a nonlinear activation function to perform nonlinear mapping on the feature F11 obtained by the first convolutional layer to obtain feature F12, then input the obtained feature F12 into the second convolutional layer, and use the second convolutional layer to convolve the feature F12 to obtain feature F13; (3) Add the feature F13 obtained in step (2) to the input x1 of the first convolutional block to obtain the output of the first convolutional block; Similarly, the method for obtaining the output of the second convolutional block in the first residual block is: (4) Use the first convolution layer of the second convolution block to convolve the output of the first convolution block to obtain feature F14; (5) Use a nonlinear activation function to perform nonlinear mapping on the feature F14 obtained by the first convolution layer of the second convolution block to obtain feature F15, then input the obtained feature F15 into the second convolution layer of the second convolution block, and use the second convolution layer of the second convolution block to convolve the feature F15 to obtain feature F16; (6) Add the feature F16 obtained in step (5) to the output of the first convolutional block to obtain the feature extraction result of the first residual block as F1.

[0043] The feature extraction result F1 of the first residual block is used as the input of the first convolution block in the second residual block, and steps (1) to (3) are performed to obtain the output of the first convolution block in the second residual block. The output of the first convolution block in the second residual block is used as the input of the second convolution block in the second residual block, and steps (4) to (6) are performed to obtain the feature extraction result F2 of the second residual block.

[0044] The feature extraction result F2 of the second residual block is used as the input of the first convolution block in the third residual block, and steps (1) to (3) are taken to obtain the output of the first convolution block in the third residual block. The output of the first convolution block in the third residual block is then used as the input of the second convolution block in the third residual block, and steps (4) to (6) are taken to obtain the feature extraction result F3 of the third residual block.

[0045] The feature extraction result F3 of the third residual block is used as the input of the first convolution block in the fourth residual block. Steps (1) to (3) are adopted to obtain the output of the first convolution block in the fourth residual block. The output of the first convolution block in the fourth residual block is used as the input of the second convolution block in the fourth residual block. Steps (4) to (6) are adopted to obtain the feature extraction result F4 of the fourth residual block.

[0046] Next, the feature extraction result F1 of the first residual block, the feature extraction result F2 of the second residual block, and the feature extraction result F3 of the third residual block are each adjusted through a 1×1 convolution block to have the same number of channels as the feature extraction result F4 of the fourth residual block, and then the adjusted feature extraction result F'1 of the first residual block, the feature extraction result F'2 of the second residual block, the feature extraction result F'3 of the third residual block, and the feature extraction result F4 of the fourth residual block enter the adaptive sliding window pooling layer respectively.

[0047] Step 22: The adaptive sliding window pooling layer adjusts the feature map sizes of the feature extraction results of the four residual blocks to be consistent, and stretches them into one-dimensional vectors using full connections.

[0048] Specifically, the feature map sizes of the feature extraction result F'1 of the first residual block, the feature extraction result F'2 of the second residual block, the feature extraction result F'3 of the third residual block, and the feature extraction result F4 of the fourth residual block are adjusted to be consistent; The multi-scale features of the feature extraction result F'1 of the first residual block after the feature map size is adjusted to be consistent in the second step are stretched into a one-dimensional vector using full connection , the multi-scale features of the feature extraction result F'2 of the second residual block after the feature map size is adjusted to be consistent are stretched into a one-dimensional vector using full connection , the multi-scale features of the feature extraction result F'3 of the third residual block after the feature map size is adjusted to be consistent are stretched into a one-dimensional vector using full connection , the multi-scale features of the feature extraction result F'4 of the fourth residual block after the feature map size is adjusted to be consistent are stretched into a one-dimensional vector using full connection .

[0049] Step 23: Add and fuse the four one-dimensional vectors to obtain the time-frequency feature extraction result.

[0050] In this step, the one-dimensional vector , one-dimensional vector , one-dimensional vector and a one-dimensional vector Perform addition fusion to achieve feature splicing and obtain the fused features It is expressed by the following formula:

[0051] in, are weights, which are parameters learned during network training.

[0052] Step 24: Extract time-frequency and time features through a recursive network.

[0053] This step uses a recursive network to integrate m time-frequency features, and inputs the one-dimensional time-frequency features at different times into the time-frequency time feature extraction step.

[0054] In a preferred recursive network, a 2-layer stacked Bi-LSTM neural network is formed. Figure 6 As shown in the figure, each layer of Bi-LSTM neural network includes m LSTM hidden layers, which process the time-frequency feature extraction results of m spliced ​​time-frequency graphs; each LSTM hidden layer in the same layer of Bi-LSTM neural network learns the temporal correlation of one-dimensional time-frequency features from the positive and reverse directions in turn. The recursive network has captured the time dependency, and the output of the second layer of Bi-LSTM neural network is the multi-scale time-frequency depth feature S of the time-frequency graph. 2 .

[0055] In a preferred embodiment, the specific implementation of step 3 is shown in Figure 7 , Figure 8 (a) and Figure 8 (b). The two types of deep features generated in step 2 need to be fused to support the final target classification task. The present invention introduces two feature fusion modules, namely, enhanced pooling module and channel space attention module.

[0056] (1) Enhanced pooling module The enhanced pooling module uses average pooling and global pooling to obtain deep features that are more sensitive to target texture and background respectively, and superimposes the two types of deep features to obtain a more comprehensive feature representation.

[0057] (2) Channel Space Attention Module The channel-spatial attention module extracts information from each feature channel separately, without considering the potential of correlation between different types of features to improve the accuracy of target recognition. In order to break this limitation, the channel attention mechanism and the spatial attention mechanism are introduced respectively, and they are connected in series to further enhance the ability of feature fusion.

[0058] Given an input C×W feature map , the channel space attention module calculates a 1-D channel attention matrix and a 1-D spatial attention matrix ,The whole process operation of the channel space attention module is as follows:

[0059] in, represents element-wise multiplication, represents the output of channel attention processing, represents the final output of the channel space attention module; represents a 1-D channel attention matrix operation, Represents a 1-D spatial attention matrix operation. Channel attention module The output can be expressed as follows:

[0060] in, represents the Sigmoid function, represents the fully connected layer, represents the compression operation, represents the activation function, Represents the adaptive average pooling operation. Spatial Attention Module The output can be expressed as follows:

[0061] in Represents a convolutional module.

[0062] (3) Feature fusion strategy The single feature output by the enhanced pooling module and the interactive feature output by the channel space attention module are fused. The entire feature fusion strategy can be expressed as follows:

[0063] in, Represents the feature map of the enhanced pooling module, represents the depth feature of the amplitude sequence, Represents multi-scale time-frequency depth features, Represents the feature map of the channel space attention module, Represents the connection operation, the final fusion feature output It is obtained by adding the output attention features.

[0064] Therefore, this step 3 adopts Figure 7 The preferred architecture shown implements feature fusion, including the following steps: Step S31: Depth feature S of amplitude sequence 1 Perform enhanced pooling processing; Step S32: multi-scale time-frequency depth feature S 2 Perform enhanced pooling processing; Step S33: transform the amplitude sequence depth feature S 1 and multi-scale time-frequency depth features S 2 Splicing: The spliced ​​features are processed by channel attention mechanism and spatial attention mechanism, and then pooled and nonlinearly mapped; Step S34: concatenate the three results of step 31, step 32 and step 33 to obtain fusion features.

[0065] In step 4, the fused features are linearly mapped through the fully connected layer to obtain a vector of the total number of categories , enter the Softmax classification layer, and get the recognition result. i The probability that a target belongs to each category is as follows:

[0066] Among them, c represents the category number, C represents the total number of categories, and the target is calculated in sequence i Probability of belonging to each category Then, the maximum value is the category to which the target belongs.

[0067] In a preferred embodiment, when processing the track sequence data in step 1 and step 2, data normalization processing is also required.

[0068] In step 1, the amplitude sequence of the track sequence data is further normalized to achieve data normalization. Examples of the amplitude sequence before and after normalization are shown in FIG9 (a) and FIG9 (b). The specific implementation may be:

[0069] in, is the amplitude sequence vector at time t, is the normalized amplitude sequence result. It means to take the maximum value in the sequence.

[0070] In step 2, the time-frequency graph is further subjected to intensity normalization and scale normalization, thereby achieving data normalization. Examples before and after normalization are shown in FIG. 10 (a) and FIG. 10 (b). The specific implementation may be: The track sequence data is subjected to short-time Fourier transform using the implementation method of “keeping the time domain interception window fixed and shifting the signal left” to obtain the time-frequency spectrum:

[0071] in, is the track sequence data to be analyzed, t is the signal moment; and are the center position and time width of the time domain interception window respectively; rect is a rectangular function, exp is an exponential function, f is the Doppler frequency.

[0072] The first step in normalizing the time-frequency spectrum is intensity normalization, and the result is:

[0073] in, is the Doppler vector at time t, To standardize the time-frequency analysis results, max Indicates taking the maximum value.

[0074] The second step of normalizing the time-frequency spectrum is scale normalization. The average value of the instantaneous frequency sequence is extracted as the frequency center of the normalized image, and 400 Hz is taken above and below to form the effective image range.

[0075] When obtaining training samples, the normalized time-frequency graph and amplitude sequence are screened to obtain three wing-beating pattern data sets of typical insects, birds and other targets. When actually performing recognition, the above scheme is also needed to normalize the amplitude sequence and time-frequency graph before performing subsequent operations.

[0076] In order to verify the target wing-flapping pattern recognition method described above, 4500 sets of target echo data sets based on experimental measurements, with 1500 targets in each category, were used to complete their pattern recognition using the aerial biological target wing-flapping pattern recognition method of the present invention.

[0077] Table 1 Confusion matrix of insect and bird classification

[0078] In Table 1, the first row represents the true value label of the target, and the first column represents the target category predicted by the network. The diagonal represents the correct target prediction. The third value "113" in the first column represents that 113 insect targets were incorrectly predicted as other types of targets, and so on for other positions. The four parameter calculation indicators are accuracy, precision, recall, and F1 score, as shown in the following formula, where TP is a positive sample predicted as a positive class, TN is a negative sample predicted as a negative class, FN is a positive sample predicted as a negative class, and FP is a negative sample predicted as a positive class.

[0079]

[0080]

[0081]

[0082]

[0083] Based on the above method, the present invention also provides an aerial biometric identification device based on echo fluctuation and multi-scale features, such as Fig.11 As shown, it includes: an amplitude sequence deep feature extraction unit, a multi-scale time-frequency deep feature extraction unit, a fusion unit and a recognition unit.

[0084] The amplitude sequence depth feature extraction unit is used to extract the amplitude sequence depth feature S according to the track sequence data of the aerial organisms. 1 , characterizing the fluctuation of RCS over time during target motion.

[0085] The multi-scale time-frequency depth feature extraction unit includes a splicing module and an extraction module. The splicing module is used to perform Fourier transform on the track sequence data according to the frequency range of wing vibration of aerial organisms using short-time Fourier transform window lengths of n scales, and generate n time-frequency diagrams for each target; slice the n time-frequency diagrams, and obtain m slices for each time-frequency diagram; take slices at the same position in different time-frequency diagrams and re-splice them to obtain m spliced ​​time-frequency diagrams; each spliced ​​time-frequency diagram represents the target micro-motion characteristics with different frequency resolutions in the same time window; n and m are positive integers greater than 2.

[0086] The extraction module is used to extract the time-frequency features of each spliced ​​time-frequency graph to obtain multi-scale time-frequency features, and then capture the time dependency through the recursive network to complete the time-frequency time feature extraction and obtain the multi-scale time-frequency depth feature S 2 .

[0087] Fusion unit, used to fusion the depth feature S of the amplitude sequence 1 and multi-scale time-frequency depth features S 2 Perform feature fusion; the specific structure is as follows Figure 7 shown.

[0088] The recognition unit is used to identify the type of aerial organisms using the fusion features obtained by the fusion unit.

[0089] Wherein, the extraction module includes a multi-scale time-frequency feature extraction module and a time-frequency time feature extraction module; The multi-scale time-frequency feature extraction module includes m extraction channels, each of which is used to process one of the m spliced ​​time-frequency images; each extraction channel includes an initial convolution module, four cascaded residual blocks, three 1×1 convolution blocks, an adaptive sliding window pooling module and a feature splicing module; the four residual blocks are the first residual block, the second residual block, the third residual block and the fourth residual block; Figure 5 shown.

[0090] The spliced ​​time-frequency map of the extracted channel enters the initial convolution module, and the output of the initial convolution module enters the first residual block; the outputs of the first three residual blocks are adjusted to the same number of channels as the feature extraction result of the fourth residual block through the corresponding 1×1 convolution block, and enter the adaptive sliding window pooling module as the feature extraction result of the corresponding residual block, and on the other hand, serve as the input of the next level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling module.

[0091] The adaptive sliding window pooling module adjusts the feature map sizes of the feature extraction results of the four residual blocks to be consistent, and stretches them into one-dimensional vectors using full connections.

[0092] The feature concatenation module adds and fuses the four one-dimensional vectors to obtain the time-frequency feature extraction result of this extraction channel.

[0093] The time-frequency feature extraction module includes two layers of stacked Bi-LSTM neural networks; Figure 6 As shown in FIG. 1 , each layer of the Bi-LSTM neural network includes m LSTM hidden layers, and each LSTM hidden layer sequentially learns the temporal correlation of the one-dimensional time-frequency features from the positive and reverse directions.

[0094] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in the description may be different and are not limited. Therefore, those skilled in the art in the field of the present invention may modify or replace the technical solutions recorded in the above embodiments; and these modifications and replacements do not deviate from the creative purpose and technical solutions of the present invention and should all fall within the protection scope of the present invention.

Claims

1. A method for aerial biometric identification based on echo fluctuation and multi-scale features, characterized in that: include: Step 1: According to the track sequence data of the aerial organisms, the amplitude sequence depth feature S1 is extracted to characterize the fluctuation of the radar cross-sectional area RCS over time during the target movement; Step 2: According to the vibration frequency range of the wing flapping action of the aerial creature, the track sequence data is Fourier transformed using a short-time Fourier transform window length of n scales, and n time-frequency graphs are generated for each target; the n time-frequency graphs are sliced, and m slices are obtained for each time-frequency graph; Take slices at the same position in different time-frequency graphs and re-join them to obtain m spliced ​​time-frequency graphs; each spliced ​​time-frequency graph represents the target micro-motion characteristics with different frequency resolutions in the same time window; Both n and m are positive integers greater than 2; The time-frequency features of each spliced ​​time-frequency graph are extracted to obtain multi-scale time-frequency features. Then, the time dependency is captured through a recursive network to complete the time-frequency time feature extraction and obtain the multi-scale time-frequency depth feature S2. Step 3: Fusing the amplitude sequence deep feature S1 and the multi-scale time-frequency deep feature S2; Step 4: Use the fusion features obtained in step 3 to recognize the wing flapping patterns of aerial organisms.

2. The method for aerial biometric identification based on echo fluctuation and multi-scale features as claimed in claim 1, characterized in that: In step 2, the selection range of the n-scale short-time Fourier transform window lengths is 4.8ms-48ms.

3. The method for aerial biometric identification based on echo fluctuation and multi-scale features as claimed in claim 1, characterized in that: In step 2, extracting time-frequency features from each spliced ​​time-frequency graph to obtain multi-scale time-frequency features includes: performing operations from step 21 to step 23 for each spliced ​​time-frequency graph, and the time-frequency feature extraction results of each spliced ​​time-frequency graph constitute the multi-scale time-frequency features; Step 21: After the spliced ​​time-frequency graph undergoes a layer of convolution processing, it enters four cascaded residual blocks for processing; the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block; the outputs of the first three residual blocks are adjusted to have the same number of channels as the number of channels of the feature extraction result of the fourth residual block through a 1×1 convolution block, and enter the adaptive sliding window pooling layer as the feature extraction result of the corresponding residual block, and on the other hand, serve as the input of the next level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling layer; Step 22: The adaptive sliding window pooling layer adjusts the feature map sizes of the feature extraction results of the four residual blocks to be consistent, and stretches them into one-dimensional vectors using full connections; Step 23: Add and fuse the four one-dimensional vectors to obtain the time-frequency feature extraction result.

4. The method for aerial biometric identification based on echo fluctuation and multi-scale features as claimed in claim 3, characterized in that: The residual block includes two convolution blocks connected in series, namely a first convolution block and a second convolution block; the output of the first convolution block is used as the input of the second convolution block; Each convolution block consists of two convolution layers, namely the first convolution layer and the second convolution layer. The first convolution layer convolves the input x1 of the convolution block to obtain the feature ; Use nonlinear activation function to get the features from the first convolutional layer Perform nonlinear mapping to obtain features , and then the obtained features Input to the second convolutional layer, use the second convolutional layer to feature Perform convolution to obtain features ; The characteristics Add it to the input x1 of the convolution block to get the output of the convolution block.

5. The method for aerial biometric identification based on echo fluctuation and multi-scale features as claimed in claim 3, characterized in that: In step 2, capturing time dependency through a recursive network to extract time-frequency and time features includes: A two-layer stacked bidirectional long short-term memory Bi-LSTM neural network is used as the recurrent network; Each layer of Bi-LSTM neural network includes m LSTM hidden layers, which process the time-frequency feature extraction results of m spliced ​​time-frequency graphs; each LSTM hidden layer in the same layer of Bi-LSTM neural network learns the temporal correlation of time-frequency features from the positive and reverse directions in turn; The second layer of Bi-LSTM neural network outputs multi-scale time-frequency depth features S2.

6. The method for aerial biometric identification based on echo fluctuation and multi-scale features as claimed in claim 1, characterized in that: In step 3, the feature fusion of the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2 is as follows: Step S31: Perform enhanced pooling processing on the amplitude sequence depth feature S1; Step S32: performing enhanced pooling processing on the multi-scale time-frequency depth feature S2; Step S33: concatenate the amplitude sequence deep feature S1 and the multi-scale time-frequency deep feature S2, perform channel attention mechanism processing and spatial attention mechanism processing on the concatenated features, and then perform pooling and nonlinear mapping; Step S34: concatenate the results of step 31, step 32 and step 33 to obtain fusion features.

7. The method for aerial biometric identification based on echo fluctuation and multi-scale features as claimed in claim 1, characterized in that: In step 1, the operation of extracting the amplitude sequence deep feature S1 is implemented by using a 4-layer stacked 1-D multi-channel convolution operation.

8. The method for aerial biometric identification based on echo fluctuation and multi-scale features as claimed in claim 1, characterized in that: The step 2 further includes performing intensity normalization and scale normalization on the time-frequency diagram; the step 1 further includes normalizing the amplitude sequence of the track sequence data.

9. An aerial biometric identification device based on echo fluctuation and multi-scale features, characterized in that: include: Amplitude sequence deep feature extraction unit, multi-scale time-frequency deep feature extraction unit, fusion unit and recognition unit; The amplitude sequence depth feature extraction unit is used to extract the amplitude sequence depth feature S1 according to the track sequence data of the aerial organism, and characterize the fluctuation of RCS over time during the target movement; The multi-scale time-frequency depth feature extraction unit includes a splicing module and an extraction module; The splicing module is used to perform Fourier transform on the track sequence data according to the vibration frequency range of the wing flapping action of the aerial creatures, using n-scale short-time Fourier transform window lengths, and generating n time-frequency graphs for each target; slicing the n time-frequency graphs, and obtaining m slices for each time-frequency graph; Take slices at the same position in different time-frequency graphs and re-join them to obtain m spliced ​​time-frequency graphs; each spliced ​​time-frequency graph represents the target micro-motion characteristics with different frequency resolutions in the same time window; Both n and m are positive integers greater than 2; The extraction module is used to extract time-frequency features from each spliced ​​time-frequency graph to obtain multi-scale time-frequency features, and then capture time dependencies through a recursive network to complete time-frequency time feature extraction and obtain multi-scale time-frequency depth features S2; The fusion unit is used to perform feature fusion on the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2; The recognition unit is used to recognize the wing flapping pattern of aerial creatures using the fusion features obtained by the fusion unit.

10. The aerial biometric identification device based on echo fluctuation and multi-scale features as claimed in claim 9, characterized in that: The extraction module includes a multi-scale time-frequency feature extraction module and a time-frequency time feature extraction module; The multi-scale time-frequency feature extraction module includes m extraction channels, each of which is used to process one of the m spliced ​​time-frequency images; each extraction channel includes an initial convolution module, four cascaded residual blocks, three 1×1 convolution blocks, an adaptive sliding window pooling module and a feature splicing module; The four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block; The spliced ​​time-frequency graph of the extracted channels enters the initial convolution module, and the output of the initial convolution module enters the first residual block; the outputs of the first three residual blocks are adjusted to the same number of channels as the number of channels of the feature extraction result of the fourth residual block through the corresponding 1×1 convolution block, and enter the adaptive sliding window pooling module as the feature extraction result of the corresponding residual block, and on the other hand, serve as the input of the next level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling module; The adaptive sliding window pooling module adjusts the feature map sizes of the feature extraction results of the four residual blocks to be consistent, and stretches them into one-dimensional vectors using full connections; The feature concatenation module adds and fuses the four one-dimensional vectors to obtain the time-frequency feature extraction result of this extraction channel; The time-frequency time feature extraction module includes two stacked Bi-LSTM neural networks; each layer of the Bi-LSTM neural network includes m LSTM hidden layers, and each LSTM hidden layer learns the temporal correlation of the time-frequency features from the positive and reverse directions in sequence.

Citation Information

Patent Citations

  • Complex flapping mode recognition method based on multi-scale feature fusion

    CN115565041A

  • LPI radar signal spectrogram fusion identification method

    CN117331031A

  • Unmanned aerial vehicle and bird intelligent classification method based on time-frequency micro-motion characteristics

    CN117636023A

  • Signal type identification method and system based on fusion feature and group convolution ViT network

    CN117743946A

  • Classification method for electroencephalogram emotion recognition through multi-scale spatial-temporal feature extraction based on CNN and Transform

    CN118797496A