Air biometric recognition method and device based on echo fluctuation and multi-scale features
By extracting the amplitude sequence and multi-scale time-frequency characteristics of aerial organisms, combined with feature fusion technology, the problem of low accuracy of wing pattern recognition in aerial organisms is solved, and higher biological classification accuracy is achieved.
Patent Information
- Application Number
- CN202510459246.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The prior art is difficult to effectively distinguish the wing-flapping patterns of aerial organisms, resulting in low recognition accuracy. Especially for small birds and insects, traditional methods are difficult to describe the wing-flapping frequency characteristics of their large spans.
By extracting the amplitude sequence depth characteristics and multi-scale time-frequency depth characteristics of aerial biotrack sequence data, short-time Fourier transform and recursive network are used, combined with feature fusion technology, including convolutional neural networks and bidirectional long and short-term memory networks, the time dependence is captured and the accurate identification of aerial bio-flapping wing patterns is achieved.
It improves the accuracy of the recognition of the wing pattern of aerial biological inflight, can better describe the target micro-movement characteristics of different frequency resolutions, and improves the accuracy of biological classification.
Smart Images

Figure CN119986643B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of low-altitude target radar detection, and particularly to an air biological recognition method and device based on echo fluctuation and multi-scale features. Background Art
[0002] Migration behavior is an adaptation strategy evolved by organisms to adapt to climate and food changes, which has had a wide and profound impact on the ecosystem. Migratory species are an important part of biodiversity and play a key role in maintaining the balance and function of the ecosystem. For example, migratory birds and insects not only have important ecological value themselves, but also affect other species and ecological processes through their migration behavior and have a profound impact on human life. However, due to their flight altitude and body size, there is a lack of effective observation means for air organisms. Radar, due to its advantages of all-weather operation, all-day operation, and large detection range, is a powerful tool for the migration of air biological targets.
[0003] Currently, different types of radars have been developed to study the migration of air organisms, including systems such as scanning insect radar, vertical insect radar, harmonic radar, bird detection radar, weather radar network, etc. A large number of studies have been conducted on monitoring at a large scale and individual tracking measurements at a small scale.
[0004] An important advancement in insect radar is the development of the vertical observation radar (VLR), which can detect the behavior and biological parameters of migratory insects. However, for small insects weighing less than 10 milligrams, due to the difficulty of simultaneously performing high-speed continuous sampling and long-time integration, the detection probability is low. However, by developing high-range-resolution insect radar and combining long-time integration and detection methods, the detection performance for small targets can be effectively improved and the detection range can be increased.
[0005] In terms of bird detection radar, although early research mainly focused on radar cross-section measurement, recent research has begun to focus on how to use radar technology to distinguish the echo signals of birds and insects. For example, by analyzing the differential reflectivity of radar echoes, the echo signals caused by atmospheric turbulence and insects can be distinguished. In addition, machine learning methods have been proposed to distinguish the echo signals of birds and insects from NEXRAD radar echoes, which demonstrates the great potential of using modern computing technology to improve radar data processing capabilities.
[0006] To better count the quantity ratio of insect and bird targets, many scholars have conducted in-depth research on the classification of insects and birds. Some studies use the magnitude of radar echo signal intensity to identify birds and other targets. However, for a large number of small birds such as passerines, the radar echo intensity interval is basically the same as that of large insects, making it basically impossible to distinguish. Since the flight speed of birds is generally greater than that of insects, some studies use the speed of the target relative to the ground to distinguish insects and birds. However, the ground speeds of birds and insects are affected by wind speed and wind direction, especially the latter. In the case of uncertain wind speed and wind direction information, the ground speeds of birds and insects have a large overlapping interval and cannot be used as a good distinguishing means either. The change pattern of echo intensity over time is also the focus of the research on the flight behavior of birds and insects, reflecting the shape change of the target in the radar beam. The typical pattern of the bird echo signal is the periodic fluctuation of the signal, and its fluctuation frequency is generally the frequency of wing flapping. This fluctuation is caused by the change of the overall body during the wing flapping process. Some studies have pointed out the main three types of birds corresponding to the wing flapping modes of continuous wing flapping, intermittent flapping, and irregular flapping of birds. Insect wings are generally very small relative to the body and do not cause obvious changes in the radar echo. Their echo signals show a typical pattern similar to a sine wave. This difference in different motion modes under radar observation is widely used in the identification of birds and insects. A variety of machine learning methods including random forests, support vector machines, and neural networks have been applied to the identification of birds and insects. The input signals mainly include the original echo and the time-frequency diagram after Fourier transform or wavelet transform.
[0007] In summary, the main problems in the recognition of the wing flapping mode of aerial organisms are as follows:
[0008] 1. The types of aerial biological targets are complex and the wing flapping modes are diverse. It is difficult for a single feature to effectively distinguish biological targets with different wing flapping modes, resulting in a low accuracy rate of wing flapping mode recognition.
[0009] 2. The wing flapping frequency scale of aerial organisms has a large span. Traditional methods use transformation algorithms with fixed parameters, which are difficult to effectively describe the typical characteristics of aerial targets and affect the accuracy rate of wing flapping mode recognition. Summary of the Invention
[0010] In view of this, the present invention provides a method and device for identifying aerial organisms based on echo fluctuation and multi-scale features, which can improve the accuracy rate of wing flapping mode recognition of aerial organisms.
[0011] To solve the above technical problems, the present invention is implemented as follows.
[0012] A method for identifying aerial organisms based on echo fluctuation and multi-scale features includes:
[0013] Step 1: According to the track sequence data of airborne organisms, extract the amplitude sequence depth feature S1, which characterizes the fluctuation of the radar cross section (RCS) over time during the target movement;
[0014] Step 2: According to the vibration frequency range of the wing flapping action of airborne organisms, perform Fourier transform on the track sequence data using short-time Fourier transform window lengths of n scales. Each target generates n time-frequency diagrams; slice the n time-frequency diagrams, and each time-frequency diagram obtains m slices; take the slices at the same position in different time-frequency diagrams and splice them again to obtain m spliced time-frequency diagrams; each spliced time-frequency diagram characterizes the micro-motion features of the target with different frequency resolutions under the same time window; n and m are positive integers greater than 2;
[0015] Extract time-frequency features from each spliced time-frequency diagram to obtain multi-scale time-frequency features, and then capture time dependence through a recursive network to complete time-frequency time feature extraction and obtain multi-scale time-frequency depth feature S2;
[0016] Step 3: Perform feature fusion on the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2;
[0017] Step 4: Use the fusion features in Step 3 to identify the wing flapping patterns of airborne organisms.
[0018] Preferably, in Step 2, the range of the short-time Fourier transform window lengths of n scales is 4.8 ms - 48 ms.
[0019] Preferably, in Step 2, the extraction of time-frequency features from each spliced time-frequency diagram to obtain multi-scale time-frequency features includes: performing operations in Steps 21 - 23 for each spliced time-frequency diagram, and the time-frequency feature extraction results of each spliced time-frequency diagram constitute the multi-scale time-frequency features;
[0020] Step 21: After performing one layer of convolution processing on the spliced time-frequency diagram, enter four cascaded residual blocks for processing; the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block; the outputs of the first three residual blocks are adjusted to have the same number of channels as the feature extraction result of the fourth residual block through a 1×1 convolution block and serve as the feature extraction results of the corresponding residual blocks to enter the adaptive sliding window pooling layer, and on the other hand, serve as the input of the next-level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling layer;
[0021] Step 22: The adaptive sliding window pooling layer adjusts the feature map sizes of the feature extraction results of the 4 residual blocks to be the same and stretches them into one-dimensional vectors respectively using fully connected layers;
[0022] Step 23: Add and fuse the 4 one-dimensional vectors to obtain the time-frequency feature extraction result.
[0023] Preferably, the residual block includes two cascaded convolutional blocks, namely a first convolutional block and a second convolutional block; the output of the first convolutional block serves as the input of the second convolutional block;
[0024] Each convolutional block includes two convolutional layers, namely a first convolutional layer and a second convolutional layer; the first convolutional layer convolves the input x1 of the convolutional block where it is located to obtain features ; a non-linear activation function is used to perform non-linear mapping on the features obtained by the first convolutional layer to obtain features , and then the obtained features are input into the second convolutional layer, and the second convolutional layer is used to convolve the features to obtain features ;
[0025] The features are added to the input x1 of the convolutional block where they are located to obtain the output of the convolutional block where they are located.
[0026] Preferably, in step 2, the capturing of time dependence through the recurrent network to complete time-frequency time feature extraction includes:
[0027] A two-layer stacked bidirectional long short-term memory Bi-LSTM neural network is used as the recurrent network;
[0028] Each layer of the Bi-LSTM neural network includes m LSTM hidden layers, which correspondingly process the time-frequency feature extraction results of m spliced time-frequency diagrams; in the same layer of the Bi-LSTM neural network, each LSTM hidden layer learns the temporal correlation of time-frequency features from the forward and reverse directions in sequence;
[0029] The second layer of the Bi-LSTM neural network outputs multi-scale time-frequency depth features S2.
[0030] Preferably, in step 3, the feature fusion of the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2 is as follows:
[0031] Step S31: Perform enhanced pooling processing on the amplitude sequence depth feature S1;
[0032] Step S32: Perform enhanced pooling processing on the multi-scale time-frequency depth feature S2;
[0033] Step S33: Concatenate the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2, and perform channel attention mechanism processing and spatial attention mechanism processing on the concatenated features, and then perform pooling and non-linear mapping;
[0034] Step S34: Concatenate the results of step 31, step 32, and step 33 to obtain the fused features.
[0035] Preferably, in step 1, the operation of extracting the amplitude sequence depth feature S1 is implemented by using a 4-layer stacked 1-D multi-channel convolution operation.
[0036] Preferably, step 2 further includes intensity normalization and scale normalization of the time-frequency map; step 1 further includes normalization of the amplitude sequence of the track sequence data.
[0037] The present invention also provides an airborne biometric recognition device based on echo fluctuation and multi-scale features, including: an amplitude sequence depth feature extraction unit, a multi-scale time-frequency depth feature extraction unit, a fusion unit, and an identification unit;
[0038] The amplitude sequence depth feature extraction unit is configured to extract the amplitude sequence depth feature S1 according to the track sequence data of the airborne organism, characterizing the fluctuation change of the RCS over time during the target movement process;
[0039] The multi-scale time-frequency depth feature extraction unit includes a splicing module and an extraction module;
[0040] The splicing module is configured to perform Fourier transform on the track sequence data by using short-time Fourier transform window lengths of n scales according to the vibration frequency range of the wing flapping action of the airborne organism, generating n time-frequency maps for each target; slicing the n time-frequency maps, obtaining m slices for each time-frequency map; taking the slices at the same positions in different time-frequency maps and splicing them again to obtain m spliced time-frequency maps; each spliced time-frequency map characterizes the micro-motion features of the target with different frequency resolutions under the same time window; n and m are positive integers greater than 2;
[0041] The extraction module is configured to perform time-frequency feature extraction on each spliced time-frequency map to obtain multi-scale time-frequency features, and then capture the time dependence through a recursive network to complete the time-frequency time feature extraction, obtaining the multi-scale time-frequency depth feature S2;
[0042] The fusion unit is configured to perform feature fusion on the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2;
[0043] The identification unit is configured to perform wing flapping pattern recognition of the airborne organism by using the fusion features obtained by the fusion unit.
[0044] Preferably, the extraction module includes a multi-scale time-frequency feature extraction module and a time-frequency time feature extraction module;
[0045] The multi-scale time-frequency feature extraction module includes m extraction channels, which are respectively used to process one of the m spliced time-frequency diagrams; each extraction channel includes an initial convolution module, four cascaded residual blocks, three 1×1 convolution blocks, an adaptive sliding window pooling module, and a feature splicing module; the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block respectively;
[0046] The spliced time-frequency diagram of this extraction channel enters the initial convolution module, and the output of the initial convolution module enters the first residual block; the outputs of the first three residual blocks are adjusted to the same number of channels as the feature extraction result of the fourth residual block through the corresponding 1×1 convolution blocks on the one hand, and are used as the feature extraction results of the corresponding residual blocks to enter the adaptive sliding window pooling module, and on the other hand, are used as the input of the next-level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling module;
[0047] The adaptive sliding window pooling module adjusts the feature map sizes of the feature extraction results of the 4 residual blocks to be the same, and stretches them into one-dimensional vectors respectively using fully connected layers;
[0048] The feature splicing module adds and fuses the 4 one-dimensional vectors to obtain the time-frequency feature extraction result of this extraction channel;
[0049] The time-frequency time feature extraction module includes two stacked Bi-LSTM neural networks; each layer of the Bi-LSTM neural network includes m LSTM hidden layers, and each LSTM hidden layer learns the temporal correlation of the time-frequency features from the forward and reverse directions in turn.
[0050] Beneficial effects:
[0051] (1) The present invention abandons the traditional method of directly inputting the time-frequency diagram into the network, and innovatively intercepts and splices different time-frequency diagrams and re-splices them into a new frequency feature diagram, so as to obtain the target micro-motion features with different frequency resolutions under the same time window, as the input features of the time-frequency feature extraction module. Compared with the traditional method of directly using the time-frequency diagram as the input for processing, this method introduces multi-frequency resolution information features, and splices the frequency features of the same time window, which can prompt the algorithm to better capture the key frequency features of the target and improve the accuracy of aerial biological wing flapping pattern recognition.
[0052] (2) In a preferred embodiment, the feature fusion of the present invention consists of two parts. The first part fully excavates the different scale features of various features through different pooling strategies, and the second part captures the spatial similarity of the two types of features and the feature importance of different channels through the spatial-channel attention mechanism. The two-channel feature fusion can further excavate the correlation between the two types of features and perform feature fusion to improve the accuracy of biological classification.
[0053] (3) Since the wing flapping patterns of aerial biological targets are complex and their vibration frequencies cover a wide range, in order to better extract the micro-motion features of aerial biological targets, in a preferred embodiment, an optimal value range of the window length of the multi-scale short-time Fourier transform is designed. The shortest is 4.8 ms, corresponding to a frequency of 208 Hz, which can fully extract the micro-motion features of fast-vibrating insect targets.
[0054] (4) When extracting the multi-scale time-frequency depth features of the present invention, considering that the data volume of the present invention is relatively small, in the process of selecting the network structure, a basic architecture with relatively fewer structural layers, better network performance and not easily collapsing during training is selected. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Fig. 1(a) is the amplitude sequence and time-frequency diagram of typical insects;
[0056] Fig. 1(b) is the amplitude sequence and time-frequency diagram of typical birds;
[0057] Fig. 1(c) is the amplitude sequence and time-frequency diagram of other targets;
[0058] Figure 2 is the flow chart of the wing flapping pattern recognition method based on echo fluctuation and multi-scale feature fusion of the present invention;
[0059] Figure 3 is the overall framework of the algorithm of the present invention;
[0060] Figure 4 is the schematic diagram of multi-scale time-frequency map splicing;
[0061] Figure 5 is the structural diagram of the multi-scale time-frequency feature extraction module;
[0062] Figure 6 is the structural diagram of the time-frequency time feature extraction module;
[0063] Figure 7 is the fusion unit architecture;
[0064] Fig. 8(a) is the structural diagram of the enhanced pooling module in the fusion unit;
[0065] Fig. 8(b) is the structural diagram of the channel spatial attention module in the fusion unit;
[0066] Fig. 9(a) is the amplitude sequence before normalization;
[0067] Fig. 9(b) is the amplitude sequence after normalization;
[0068] Fig. 10(a) is the time-frequency diagram before normalization;
[0069] Fig. 10(b) is the time-frequency diagram after normalization;
[0070] Figure 11 This is a schematic diagram of the flapping wing pattern recognition device based on echo fluctuation and multi-scale feature fusion of the present invention. Specific implementation manners
[0071] The present invention will be described in detail below with reference to the accompanying drawings and by way of examples.
[0072] The present invention provides an aerial biological recognition solution based on echo fluctuation and multi-scale features. Its basic idea is to extract two types of deep features. One is the amplitude sequence deep feature S1, which characterizes the fluctuation change of the radar cross-sectional area (RCS) over time during the target movement. The other is the multi-scale time-frequency deep feature S2, which characterizes the biological Doppler dynamic change at different time-resolution. The two features are fused to identify the target type, thereby improving the accuracy of aerial biological classification.
[0073] Figures 1(a), 1(b), and 1(c) are respectively the amplitude sequences and time-frequency diagrams of typical insects, birds, and other targets. It can be seen that for different types of aerial organisms, due to their different flapping wing patterns, there are obvious differences in the amplitude sequences and time-frequency diagrams of the echoes. When extracting features in the present invention, especially time-frequency features, different time-frequency diagrams are intercepted and spliced to obtain the micro-motion features of the target at different frequency resolutions under the same time window, as the input features for the time-frequency feature extraction operation. This makes the time-frequency feature extraction operation not fixed at a certain scale, which can prompt the algorithm to better capture the key frequency features of the target, effectively describe the features of the aerial target, and design a new network architecture to achieve the recognition of aerial biological targets, thereby improving the accuracy of aerial biological recognition.
[0074] Figure 2 This is a flowchart of the aerial biological recognition method based on echo fluctuation and multi-scale features of the present invention. Figure 3 This is the algorithm framework of this method, as Figure 2 and Figure 3 shown. The method includes the following steps:
[0075] Step 1: According to the track sequence data of the aerial organism, extract the amplitude sequence deep feature S1, which characterizes the fluctuation change of the RCS over time during the target movement.
[0076] Step 2: According to the window length range determined by the flapping wing vibration frequency range of the aerial organism, select the short-time Fourier transform window lengths of n scales. Use the short-time Fourier transform window lengths of these n scales to perform Fourier transform on the track sequence data. Each target generates n time-frequency diagrams; slice the n time-frequency diagrams, and each time-frequency diagram obtains m slices; take the slices at the same position in different time-frequency diagrams and splice them again to obtain m spliced time-frequency diagrams. The splicing schematic is as Figure 4As shown. Each spliced time-frequency diagram represents the target micro-motion features with different frequency resolutions under the same time window; n and m are positive integers greater than 2.
[0077] Perform time-frequency feature extraction on each spliced time-frequency diagram to obtain multi-scale time-frequency features; then capture the time dependence through a recurrent network to complete the time-frequency time feature extraction and obtain the multi-scale time-frequency depth feature S2.
[0078] Step 3: Perform feature fusion on the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2.
[0079] Step 4: Use the fusion features in Step 3 for aerial biological wing flapping pattern recognition.
[0080] The following describes each step in detail.
[0081] In a preferred embodiment, referring to Figure 3 the upper part of, the implementation manner of Step 1 amplitude sequence depth feature extraction is preferably: adopt a 4-layer stacked 1-D multi-channel convolution operation to extract the amplitude sequence depth feature S1. Specifically, use a CNN to extract the depth feature of the target amplitude sequence. For a typical 1-D multi-channel convolution operation, build a residual convolution block on the basis of 1-D convolution. Each residual block consists of two convolution blocks, and each convolution block is composed of a first convolution layer, an activation function, and a second convolution layer. The output of a convolution block is the result of concatenating the output of the second convolution layer inside it and the input of this convolution block:
[0082]
[0083] Among them: x is the input of the convolution block, is the output of the second convolution layer, is the output of the convolution block.
[0084] In a preferred embodiment, the specific implementation process of Step 2 specifically includes Step 20 - Step 24. Step 20 is the slicing and re-splicing of the time-frequency diagram, Steps 21 - 23 are multi-scale time-frequency feature extraction, and Step 24 is time-frequency time feature extraction, and finally the multi-scale time-frequency depth feature S2 is obtained. Specifically:
[0085] Step 20: Re-splicing of the time-frequency diagram.
[0086] Due to the complex motion patterns of aerial biological targets and the wide range of their vibration frequencies, in order to better extract the micro-motion features of aerial biological targets, the present invention designs short-time Fourier transform window lengths of different scales. The shortest is 4.8 ms, corresponding to a frequency of 208 Hz, which can fully extract the micro-motion features of fast-vibrating insect targets. In a preferred implementation, the window lengths are selected at intervals of 3.2 ms from 4.8 ms to 48 ms, with a total of n = 14 window lengths for performing short-time Fourier transform on the signal, and 14 time-frequency diagrams are generated for each target. The present invention abandons the traditional method of directly inputting the time-frequency diagrams into the network and innovatively intercepts and stitches different time-frequency diagrams. Each time-frequency diagram is sequentially intercepted with a time window of 8 points and re-stitched into, for example, m = 14 frequency feature diagrams, so as to obtain the micro-motion features of the target with different frequency resolutions under the same time window, which are used as the input features for the multi-scale time-frequency depth feature extraction step, and a new network architecture is designed based on this to achieve the recognition of the wing-beating patterns of aerial organisms.
[0087] The following operations of Step 21 to Step 23 are performed on each stitched time-frequency diagram, and the time-frequency feature extraction results of each stitched time-frequency diagram constitute the multi-scale time-frequency features. See Figure 5 。
[0088] Step 21: After performing one-layer convolution processing on the stitched time-frequency diagram, it enters four cascaded residual blocks for processing; the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block; the outputs of the first three residual blocks are adjusted to have the same number of channels as the feature extraction result of the fourth residual block through a 1×1 convolution block, and are used as the feature extraction results of the corresponding residual blocks to enter the adaptive sliding window pooling layer, and on the other hand, are used as the inputs of the next-level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling layer.
[0089] Specifically, the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block (corresponding to Residual Block 1, 2, 3, and 4 in the figure), the feature extraction result of the first residual block is F1, the feature extraction result of the second residual block is F2, the feature extraction result of the third residual block is F3, and the feature extraction result of the fourth residual block is F4.
[0090] Among them, the first residual block includes two convolution blocks, namely the first convolution block and the second convolution block. The method for obtaining the feature extraction result F1 of the first residual block is as follows:
[0091] (1) Convolve the input x1 of the first convolution block using the first convolution layer of the first convolution block to obtain the feature , where L represents the number of channels of the feature, represents the i rd channel in the feature F11i = 1, 2, …, L, is represented by the following formula:
[0092]
[0093] where, represents a convolution operation, represents the i th convolution kernel, represents the offset of the i th convolution kernel;
[0094] (2) Use a non - linear activation function to perform a non - linear mapping on the feature F11 obtained from the first convolutional layer to obtain the feature F12, and then input the obtained feature F12 into the second convolutional layer. Use the second convolutional layer to perform convolution on the feature F12 to obtain the feature F13;
[0095] (3) Add the feature F13 obtained in step (2) to the first convolutional block input x1 to obtain the output of the first convolutional block;
[0096] Similarly, the method for obtaining the output of the second convolutional block in the first residual block is as follows:
[0097] (4) Use the first convolutional layer of the second convolutional block to perform convolution on the output of the first convolutional block to obtain the feature F14;
[0098] (5) Use a non - linear activation function to perform a non - linear mapping on the feature F14 obtained from the first convolutional layer of the second convolutional block to obtain the feature F15, and then input the obtained feature F15 into the second convolutional layer of the second convolutional block. Use the second convolutional layer of the second convolutional block to perform convolution on the feature F15 to obtain the feature F16;
[0099] (6) Add the feature F16 obtained in step (5) to the output of the first convolutional block to obtain the feature extraction result of the first residual block as F1.
[0100] Taking the feature extraction result of the first residual block as F1 as the input of the first convolutional block in the second residual block, adopt the methods of steps (1) - (3) to obtain the output of the first convolutional block in the second residual block, and then take the output of the first convolutional block in the second residual block as the input of the second convolutional block in the second residual block, adopt the methods of steps (4) - (6) to obtain the feature extraction result of the second residual block as F2.
[0101] Take the feature extraction result F2 of the second residual block as the input of the first convolutional block in the third residual block, and adopt the methods in steps (1)-(3) to obtain the output of the first convolutional block in the third residual block. Then, take the output of the first convolutional block in the third residual block as the input of the second convolutional block in the third residual block, and adopt the methods in steps (4)-(6) to obtain the feature extraction result F3 of the third residual block.
[0102] Take the feature extraction result F3 of the third residual block as the input of the first convolutional block in the fourth residual block, and adopt the methods in steps (1)-(3) to obtain the output of the first convolutional block in the fourth residual block. Then, take the output of the first convolutional block in the fourth residual block as the input of the second convolutional block in the fourth residual block, and adopt the methods in steps (4)-(6) to obtain the feature extraction result F4 of the fourth residual block.
[0103] Next, pass the feature extraction result F1 of the first residual block, the feature extraction result F2 of the second residual block, and the feature extraction result F3 of the third residual block respectively through a 1×1 convolutional block to adjust the number of channels to be the same as that of the feature extraction result F4 of the fourth residual block. Then, enter the adjusted feature extraction result F’1 of the first residual block, the feature extraction result F’2 of the second residual block, the feature extraction result F’3 of the third residual block, and the feature extraction result F4 of the fourth residual block into the adaptive sliding window pooling layer respectively.
[0104] Step 22: The adaptive sliding window pooling layer adjusts the feature map sizes of the feature extraction results of the 4 residual blocks to be the same, and stretches them into one-dimensional vectors respectively using fully connected layers.
[0105] Specifically, the feature map sizes of the feature extraction result F’1 of the first residual block, the feature extraction result F’2 of the second residual block, the feature extraction result F’3 of the third residual block, and the feature extraction result F4 of the fourth residual block are adjusted to be the same;
[0106] Stretch the multi-scale features of the feature extraction result F’1 of the first residual block after adjusting the feature map size in the second step into a one-dimensional vector using a fully connected layer and stretch the multi-scale features of the feature extraction result F’2 of the second residual block after adjusting the feature map size into a one-dimensional vector using a fully connected layer and stretch the multi-scale features of the feature extraction result F’3 of the third residual block after adjusting the feature map size into a one-dimensional vector using a fully connected layer and stretch the multi-scale features of the feature extraction result F’4 of the fourth residual block after adjusting the feature map size into a one-dimensional vector using a fully connected layer .
[0107] Step 23: Add and fuse the four one-dimensional vectors to obtain the time-frequency feature extraction result.
[0108] In this step, the one-dimensional vectors , the one-dimensional vector , the one-dimensional vector and the one-dimensional vector are added and fused to achieve feature splicing, and the fused feature is represented by the following formula:
[0109]
[0110] where is the weight, which is a parameter learned during network training.
[0111] Step 24: Extract time-frequency time features through a recurrent network.
[0112] In this step, a recurrent network is used to integrate m time-frequency features, and the one-dimensional time-frequency features at different times are input into the time-frequency time feature extraction step.
[0113] In a preferred embodiment, the recurrent network consists of two stacked Bi-LSTM neural networks. As Figure 6 shown, each layer of the Bi-LSTM neural network includes m LSTM hidden layers, corresponding to processing the time-frequency feature extraction results of m spliced time-frequency diagrams; in the same layer of the Bi-LSTM neural network, each LSTM hidden layer learns the temporal correlation of the one-dimensional time-frequency features in the forward and reverse directions in turn. The recurrent network has captured the time dependence, and the output of the second layer of the Bi-LSTM neural network is the multi-scale time-frequency depth feature S2 of the time-frequency diagram.
[0114] In a preferred embodiment, for the specific implementation of step 3, refer to Figure 7 , Figure 8(a) and Figure 8(b). The two types of depth features generated in step 2 need to be fused to support the final target classification task. The present invention introduces two feature fusion modules, namely an enhanced pooling module and a channel spatial attention module.
[0115] (1) Enhanced pooling module
[0116] The enhanced pooling module uses average pooling and global pooling to obtain depth features that are more sensitive to the target texture and background respectively, and superimposes the two types of depth features to obtain a more comprehensive feature representation.
[0117] (2) Channel spatial attention module
[0118] The channel spatial attention module extracts the information of each feature channel separately, without considering the potential of the correlation of different types of features to improve the target recognition accuracy. To break this limitation, the channel attention mechanism and the spatial attention mechanism are introduced respectively and they are cascaded to further enhance the feature fusion ability.
[0119] Given an input feature map of C×W , the channel spatial attention module calculates a 1-D channel attention matrix and a 1-D spatial attention matrix . The entire process of the channel spatial attention module is as follows:
[0120]
[0121] where, represents element-wise multiplication, represents the output of the channel attention process, represents the final output of the channel spatial attention module; represents the 1-D channel attention matrix operation, represents the 1-D spatial attention matrix operation. The output of the channel attention module can be expressed as follows:
[0122]
[0123] where, represents the Sigmoid function, represents the fully connected layer, represents the compression operation, represents the activation function, represents the adaptive average pooling operation. The output of the spatial attention module can be expressed as follows:
[0124]
[0125] where represents the convolutional module.
[0126] (3)Feature Fusion Strategy
[0127] Fuse the single feature output by the enhanced pooling module and the interactive feature output by the channel spatial attention module. The entire feature fusion strategy can be expressed as follows:
[0128]
[0129] where, represents the feature map of the enhanced pooling module, represents the amplitude sequence depth feature, Represents multi-scale time-frequency depth features, Represents the feature map of the channel spatial attention module, Represents the connection operation and the final fused feature output Is obtained by adding the output attention features.
[0130] Therefore, in this step 3, Figure 7 The preferred architecture shown is used to achieve feature fusion, including the following steps:
[0131] Step S31: Perform enhanced pooling processing on the amplitude sequence depth feature S1;
[0132] Step S32: Perform enhanced pooling processing on the multi-scale time-frequency depth feature S2;
[0133] Step S33: Concatenate the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2, perform channel attention mechanism processing and spatial attention mechanism processing on the concatenated features, and then perform pooling and non-linear mapping;
[0134] Step S34: Concatenate the three results of step 31, step 32 and step 33 to obtain the fused feature.
[0135] In step 4, the fused feature is linearly mapped through a fully connected layer to obtain a vector of the size of the total number of categories , enter the Softmax classification layer to obtain the identification result, and the probabilities that the i th target in the sample belongs to each category are as follows:
[0136]
[0137] Among them, c represents the category serial number, C represents the total number of categories, and the probabilities that the target i belongs to each category are calculated in turn , and the maximum value among them is taken as the category to which the target belongs.
[0138] In a preferred embodiment, when processing the track sequence data in step 1 and step 2, data normalization processing is also required.
[0139] In step 1, it further includes normalizing the amplitude sequence of the track sequence data to achieve data normalization. Examples of the amplitude sequence before and after normalization are shown in FIGS. 9(a) and 9(b). The specific implementation can be:
[0140]
[0141] Among them, is the amplitude sequence vector at time t, is the result of the normalized amplitude sequence. Indicates taking the maximum value in the sequence.
[0142] In step 2, it further includes intensity normalization and scale normalization of the time-frequency diagram to achieve data normalization. Examples before and after normalization are shown in Figures 10(a) and 10(b). The specific implementation can be as follows:
[0143] Perform short-time Fourier transform on the track sequence data using the implementation method of "keeping the time-domain intercept window stationary and shifting the signal to the left" to obtain the time-frequency spectrum diagram:
[0144]
[0145] Among them, is the track sequence data to be analyzed, t is the signal time; and are the center position and time width of the time-domain intercept window respectively; rect is the rectangular function, exp is the exponential function, f is the Doppler frequency.
[0146] The first step of normalizing the time-frequency spectrum diagram is intensity normalization, and the processing result is:
[0147]
[0148] Among them, is the Doppler vector at time t, is the result of the standardized time-frequency analysis diagram, max Indicates taking the maximum value.
[0149] The second step of normalizing the time-frequency spectrum diagram is scale normalization. Extract the average value of the instantaneous frequency sequence as the frequency center of the normalized image, and take 400 Hz above and below to form the effective image range.
[0150] When obtaining the training samples, screen the obtained normalized time-frequency diagram and amplitude sequence to obtain three datasets of wing-beating patterns for typical insects, birds, and other targets. When actually performing recognition, it is also necessary to use the above scheme to first normalize the amplitude sequence and time-frequency diagram, and then perform subsequent operations.
[0151] To verify the above-mentioned target wing-beating pattern identification method, for 4500 groups of target echo datasets based on experimental measurements, 1500 for each category, use the wing-beating pattern identification method of aerial biological targets of the present invention to complete its pattern identification.
[0152] Table 1 Insect-bird classification confusion matrix
[0153]
[0154] In Table 1, the first row represents the true labels of the targets, and the first column represents the target categories predicted by the network. The diagonal represents correct target predictions. The third value in the first column, "113", means that 113 insect targets were mispredicted as other types of targets, and so on for other positions. The four parameter calculation metrics are Precision, Accuracy, Recall, and F1-score, as shown in the following formulas, where TP is the positive sample predicted as the positive class, TN is the negative sample predicted as the negative class, FN is the positive sample predicted as the negative class, and FP is the negative sample predicted as the positive class.
[0155]
[0156]
[0157]
[0158]
[0159] Based on the above method, the present invention also provides an aerial biometric recognition device based on echo fluctuation and multi-scale features, as Figure 11 shown, including: an amplitude sequence depth feature extraction unit, a multi-scale time-frequency depth feature extraction unit, a fusion unit, and an identification unit.
[0160] The amplitude sequence depth feature extraction unit is used to extract the amplitude sequence depth feature S1 according to the track sequence data of the aerial organism, which characterizes the fluctuation of the RCS over time during the target movement process.
[0161] The multi-scale time-frequency depth feature extraction unit includes a splicing module and an extraction module. Among them,
[0162] The splicing module is used to perform Fourier transform on the track sequence data using short-time Fourier transform window lengths of n scales according to the wing flapping vibration frequency range of the aerial organism. Each target generates n time-frequency diagrams; slice the n time-frequency diagrams, and each time-frequency diagram obtains m slices; take the slices at the same positions in different time-frequency diagrams and splice them again to obtain m spliced time-frequency diagrams; each spliced time-frequency diagram characterizes the micro-motion features of the target with different frequency resolutions under the same time window; n and m are positive integers greater than 2.
[0163] The extraction module is used to extract time-frequency features from each spliced time-frequency diagram to obtain multi-scale time-frequency features, and then capture time dependence through a recurrent network to complete time-frequency time feature extraction and obtain multi-scale time-frequency depth feature S2.
[0164] The fusion unit is used to perform feature fusion on the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2; the specific structure is as Figure 7as shown
[0165] An identification unit for identifying the type of airborne organism by using the fusion features obtained by the fusion unit.
[0166] Wherein, the extraction module includes a multi-scale time-frequency feature extraction module and a time-frequency time feature extraction module;
[0167] The multi-scale time-frequency feature extraction module includes m extraction channels, which are respectively used to process one of the m spliced time-frequency diagrams; each extraction channel includes an initial convolution module, four cascaded residual blocks, three 1×1 convolution blocks, an adaptive sliding window pooling module, and a feature splicing module; the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block; as Figure 5 as shown
[0168] The spliced time-frequency diagram of this extraction channel enters the initial convolution module, and the output of the initial convolution module enters the first residual block; the outputs of the first three residual blocks are adjusted to the same number of channels as the feature extraction result of the fourth residual block through the corresponding 1×1 convolution blocks and serve as the feature extraction results of the corresponding residual blocks to enter the adaptive sliding window pooling module, and on the other hand, serve as the input of the next-level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling module.
[0169] The adaptive sliding window pooling module adjusts the feature map sizes of the feature extraction results of the 4 residual blocks to be the same, and stretches them into one-dimensional vectors respectively by using fully connected layers.
[0170] The feature splicing module adds and fuses the 4 one-dimensional vectors to obtain the time-frequency feature extraction result of this extraction channel.
[0171] The time-frequency time feature extraction module includes two stacked Bi-LSTM neural networks; as Figure 6 as shown, each layer of the Bi-LSTM neural network includes m LSTM hidden layers, and each LSTM hidden layer learns the temporal correlation of one-dimensional time-frequency features in the forward and reverse directions in turn.
[0172] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in this description can be different and are not limited. Therefore, those skilled in the art of the present invention can modify or equivalently replace the technical solutions recorded in the foregoing embodiments; and these modifications and replacements do not depart from the spirit and technical solutions of the present invention, and should all fall within the protection scope of the present invention.
Claims
1. An airborne biometric recognition method based on echo fluctuation and multi-scale features, characterized in that, Including: Step 1: According to the track sequence data of airborne organisms, extract the amplitude sequence depth feature S1, which characterizes the fluctuation of the radar cross section (RCS) with time during the target movement. Step 2: According to the vibration frequency range of the wing flapping action of airborne organisms, perform Fourier transform on the track sequence data using short-time Fourier transform window lengths of n scales. Each target generates n time-frequency diagrams; slice the n time-frequency diagrams, and each time-frequency diagram obtains m slices. Take the slices at the same position in different time-frequency diagrams and splice them again to obtain m spliced time-frequency diagrams; each spliced time-frequency diagram characterizes the micro-motion features of the target with different frequency resolutions under the same time window. Both n and m are positive integers greater than 2. Perform time-frequency feature extraction on each spliced time-frequency diagram to obtain multi-scale time-frequency features, and then capture the time dependence through a recursive network to complete the time-frequency time feature extraction and obtain the multi-scale time-frequency depth feature S2. Step 3: Perform feature fusion on the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2. Step 4: Use the fusion features obtained in Step 3 to perform wing flapping pattern recognition of airborne organisms.
2. The airborne biometric recognition method based on echo fluctuation and multi-scale features according to claim 1, characterized in that In Step 2, the selection range of the short-time Fourier transform window lengths of the n scales is 4.8 ms - 48 ms.
3. The airborne biometric recognition method based on echo fluctuation and multi-scale features according to claim 1, characterized in that In Step 2, the performing time-frequency feature extraction on each spliced time-frequency diagram to obtain multi-scale time-frequency features includes: performing operations of Step 21 - Step 23 for each spliced time-frequency diagram, and the time-frequency feature extraction results of each spliced time-frequency diagram form the multi-scale time-frequency features. Step 21: After the spliced time-frequency diagram is subjected to one layer of convolution processing, it enters four cascaded residual blocks for processing; the four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block respectively; the outputs of the first three residual blocks are adjusted to have the same number of channels as the feature extraction result of the fourth residual block through a 1×1 convolution block and serve as the feature extraction results of the corresponding residual blocks to enter the adaptive sliding window pooling layer, and on the other hand, serve as the input of the next-level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling layer. Step 22: The adaptive sliding window pooling layer adjusts the feature map sizes of the feature extraction results of the 4 residual blocks to be the same and stretches them into one-dimensional vectors respectively using fully connected layers. Step 23: Add and fuse the 4 one-dimensional vectors to obtain the time-frequency feature extraction result.
4. The airborne biometric recognition method based on echo fluctuation and multi-scale features according to claim 3, wherein Each of the four residual blocks includes two cascaded convolution blocks, namely the first convolution block and the second convolution block; the output of the first convolution block serves as the input of the second convolution block. Each convolutional block includes two convolutional layers, namely the first convolutional layer and the second convolutional layer; the first convolutional layer convolves the input x1 of the convolutional block where it is located to obtain features ; a non-linear activation function is used to perform a non-linear mapping on the features obtained by the first convolutional layer to obtain features , and then the obtained features are input into the second convolutional layer, and the second convolutional layer is used to convolve the features to obtain features ; Add the feature to the input x1 of the convolutional block where it is located to obtain the output of the convolutional block where it is located.
5. The airborne biometric recognition method based on echo fluctuation and multi-scale features according to claim 3, characterized in that In Step 2, the capturing the time dependence through a recursive network to complete the time-frequency time feature extraction includes: Adopting a two-layer stacked bidirectional long short-term memory (Bi-LSTM) neural network as the recursive network. Each layer of the Bi-LSTM neural network includes m LSTM hidden layers, which correspondingly process the time-frequency feature extraction results of m spliced time-frequency diagrams; in the same layer of the Bi-LSTM neural network, each LSTM hidden layer learns the temporal correlation of the time-frequency features from the forward and reverse directions in sequence. The second layer of the Bi-LSTM neural network outputs the multi-scale time-frequency depth feature S2.
6. The airborne biometric recognition method based on echo fluctuation and multi-scale features according to claim 1, wherein In step 3, the feature fusion of the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2 is as follows: Step S31: Perform enhanced pooling processing on the amplitude sequence depth feature S1; Step S32: Perform enhanced pooling processing on the multi-scale time-frequency depth feature S2; Step S33: Concatenate the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2, perform channel attention mechanism processing and spatial attention mechanism processing on the concatenated features, and then perform pooling and non-linear mapping; Step S34: Concatenate the results of step 31, step 32, and step 33 to obtain the fusion feature.
7. The airborne biometric recognition method based on echo fluctuation and multi-scale features according to claim 1, characterized in that, In step 1, the operation of extracting the amplitude sequence depth feature S1 is implemented by using a 4-layer stacked 1-D multi-channel convolution operation.
8. The airborne biometric recognition method based on echo fluctuation and multi-scale features according to claim 1, characterized in that Step 2 further includes intensity normalization and scale normalization of the time-frequency diagram; step 1 further includes normalization of the amplitude sequence of the track sequence data.
9. An airborne biometric recognition device based on echo fluctuation and multi-scale features, characterized in that It includes: An amplitude sequence depth feature extraction unit, a multi-scale time-frequency depth feature extraction unit, a fusion unit, and an identification unit; The amplitude sequence depth feature extraction unit is used to extract the amplitude sequence depth feature S1 according to the track sequence data of the airborne creature, which characterizes the fluctuation of the RCS over time during the target movement; The multi-scale time-frequency depth feature extraction unit includes a splicing module and an extraction module; The splicing module is used to perform Fourier transform on the track sequence data by using short-time Fourier transform window lengths of n scales according to the vibration frequency range of the wing flapping action of the airborne creature, and generate n time-frequency diagrams for each target; slice the n time-frequency diagrams, and obtain m slices for each time-frequency diagram; Take the slices at the same position in different time-frequency diagrams and splice them again to obtain m spliced time-frequency diagrams; each spliced time-frequency diagram characterizes the micro-motion features of the target with different frequency resolutions under the same time window; Both n and m are positive integers greater than 2; The extraction module is used to extract time-frequency features from each spliced time-frequency diagram to obtain multi-scale time-frequency features, and then capture time dependence through a recurrent network to complete time-frequency time feature extraction and obtain the multi-scale time-frequency depth feature S2; The fusion unit is used to perform feature fusion on the amplitude sequence depth feature S1 and the multi-scale time-frequency depth feature S2; The identification unit is used to perform airborne creature wing flapping pattern recognition by using the fusion feature obtained by the fusion unit.
10. The airborne biometric recognition device based on echo fluctuation and multi-scale features according to claim 9, characterized in that, The extraction module includes a multi-scale time-frequency feature extraction module and a time-frequency time feature extraction module; The multi-scale time-frequency feature extraction module includes m extraction channels, which are respectively used to process one of the m spliced time-frequency diagrams; each extraction channel includes an initial convolution module, four cascaded residual blocks, three 1×1 convolution blocks, an adaptive sliding window pooling module, and a feature splicing module; The four residual blocks are the first residual block, the second residual block, the third residual block, and the fourth residual block; The concatenated time-frequency map of the extraction channel enters the initial convolution module, and the output of the initial convolution module enters the first residual block; the outputs of the first three residual blocks are adjusted to have the same number of channels as the feature extraction result of the fourth residual block through the corresponding 1×1 convolution blocks on the one hand, and enter the adaptive sliding window pooling module as the feature extraction results of the corresponding residual blocks, and on the other hand, serve as the input of the next-level residual block; the feature extraction result of the fourth residual block directly enters the adaptive sliding window pooling module. The adaptive sliding window pooling module adjusts the feature map sizes of the feature extraction results of the 4 residual blocks to be the same, and stretches them into one-dimensional vectors respectively using fully connected layers. The feature concatenation module adds and fuses the 4 one-dimensional vectors to obtain the time-frequency feature extraction result of this extraction channel. The time-frequency time feature extraction module includes two stacked Bi-LSTM neural networks; each Bi-LSTM neural network includes m LSTM hidden layers, and each LSTM hidden layer learns the temporal correlation of the time-frequency features from the forward and reverse directions in turn.
Citation Information
Patent Citations
Complex flapping mode recognition method based on multi-scale feature fusion
CN115565041A
LPI radar signal spectrogram fusion identification method
CN117331031A