A Few-Shot Optimization Bird Sound Recognition Method Based on Bridged Transformer

Through the bridge Transformer structure and sample loss optimization module SLOBlock in the BTNN model, the problems of overfitting bird sound recognition and insufficient feature interaction under the small sample data set are solved, and bird sound recognition with high accuracy is achieved.

CN115762536BActive Publication Date: 2025-07-18NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211512964.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2025-07-18
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

The existing bird sound recognition methods are prone to overfitting problems when sample data is scarce, and the deep learning network does not fully extract and utilize bird sound features, and there is a lack of interaction between different feature information.

Method used

The BTNN model is designed, and the short-time Fourier transform of bird singing signals is extracted by bridging the Transformer structure to generate a spectral graph, combined with the sample loss optimization module SLOBlock, and the feature relationship model is modeled using the cross-attention mechanism of the single-layer Transformer encoder to realize the information interaction of local and global features and the optimization of internal network losses.

Benefits of technology

The accuracy of bird sound recognition on the small sample data set is improved, the utilization of input features is improved, the sample is expanded without expanding external data, and the test accuracy of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762536B_ABST
    Figure CN115762536B_ABST
Patent Text Reader

Abstract

The present invention discloses a few-shot optimized bird sound recognition method based on a bridged Transformer, which includes obtaining a BTNN bird sound recognition network model, extracting the short-time Fourier transform of the bird sound signal to generate a spectrogram as the input feature of the overall network model; using the bridged Transformer structure to extract and complementarily fuse the information of the local and global features of the STFT spectrogram to obtain bird sound feature parameters; introducing a sample loss optimization module SLOBlock, using the cross-attention mechanism of a single-layer Transformer encoder to perform relationship modeling on the output feature map from the backbone network, and internally optimizing the training and testing of the network itself on the few-shot dataset; conducting experiments on the Birdsdata dataset and the xeno-canto dataset, and inputting the optimized features into a Softmax classifier to obtain the recognition result. The present invention improves the accuracy of bird sound recognition tests in the case of scarce sample data by designing the BTNN model, and at the same time strengthens the information interaction of the input spectrogram at the global and local levels, and improves the extraction and utilization of the input features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bird sound recognition, and specifically to a few-shot optimized bird sound recognition method based on a bridging Transformer. Background Art

[0002] As an important part of the ecosystem, birds are widely distributed and sensitive to environmental changes. Most scholars regard birds as indicator species for monitoring environmental changes. Therefore, the monitoring, recognition, and classification of bird species are of great significance. Nowadays, with the development of acoustic technology and digital processing technology, it has gradually become the mainstream to understand the survival status of bird species through the analysis and research of bird calls.

[0003] With the development of deep learning, some domestic and foreign studies have shown that deep neural networks such as convolutional neural networks (CNNs), convolutional recurrent neural networks (CRNNs), long short-term memory networks (LSTMs), etc. can extract more valuable and richer feature information in bird sound recognition. For example, Sprengelt et al. used the short-time Fourier transform to convert bird sound signals into spectrograms and first used a convolutional neural network to train the spectrograms, achieving an accuracy rate of 58% in the public dataset of the BirdCLEF 2016 challenge and winning the championship. Liu H et al. used log-Mel spectrograms as inputs and self-built a neural network model that cascades bidirectional LSTMs and DenseNets, achieving an accuracy rate of 92.2% in a 20-class bird sound database. Jean P fragmented and then spliced the spectrograms generated by the short-time Fourier transform and input them into a classic Transformer neural network. Through the test of 397-class bird sound recognition in the CLEF competition database of that year, the accuracy rate reached 77.55% and won the gold medal. Qiu Zhibin et al. input Mel spectrograms into a self-built 24-layer CNN model and could achieve an accuracy rate of 96.1% in a dataset containing 40 types of bird calls by fine-tuning network parameters. Feng Yuqian proposed a bimodal feature fusion bird species recognition method, which optimized the bird sound recognition algorithm by fusing the time-frequency domain features of bird sounds through a cascaded structure of CNN and LSTM, achieving an average accuracy rate of 93.9% in the recognition of 6 types of birds.

[0004] From the introduction of the research methods of the above domestic and foreign scholars, it can be seen that in the current bird sound recognition method based on deep learning, converting the bird sound audio signal into a spectrogram and inputting it into the neural network to learn the time-frequency domain feature information in the spectrogram to obtain the recognition result has become the mainstream. However, there are still several problems at present. For example, the sample size of the existing bird sound dataset is scarce, which will lead to overfitting problems in the recognition process and the test effect is not good. There are also problems such as the insufficient extraction and utilization of bird sound features by the deep learning network and the lack of interaction between different feature information. Therefore, a small-sample optimized bird sound recognition method based on Bridging Transformer is proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide a small-sample optimized bird sound recognition method based on Bridging Transformer, which designs a BTNN model to improve the accuracy of bird sound recognition test in the case of scarce sample data, and at the same time strengthens the information interaction at the global and local levels of the input spectrogram to improve the extraction and utilization of the input features.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] A small-sample optimized bird sound recognition method based on Bridging Transformer, the recognition method includes the following steps:

[0008] S1. Obtain the BTNN bird sound recognition network model, and extract the short-time Fourier transform of the bird sound signal to generate a spectrogram as the input feature of the overall network model.

[0009] S2. Use the Bridging Transformer structure to extract, complement and fuse the information of the local features and global features of the STFT spectrogram to obtain bird sound feature parameters.

[0010] S3. Introduce the sample loss optimization module SLOBlock, and use the cross-attention mechanism of the single-layer Transformer encoder to perform relationship modeling on the output feature map from the backbone network, and optimize the training and testing of the network itself on the small-sample dataset from the inside.

[0011] S4. Conduct experiments on the Birdsdata dataset and the xeno-canto dataset, and input the optimized features into the Softmax classifier to obtain the recognition result.

[0012] Further, the specific operation of S1 is as follows:

[0013] Perform preprocessing operations such as pre-emphasis, framing and windowing on the original audio signal.

[0014] Obtain the STFT spectrogram through the short-time Fourier transform as the input feature of the overall network.

[0015] Input it into a single ordinary convolutional layer to obtain an operable feature map.

[0016] Furthermore, the specific operations of the STFT spectrogram include the following steps:

[0017] (1) When the energy of the bird sound signal is mostly concentrated in the low-frequency band, a high-pass filter is used to pre-emphasize the bird sound audio signal. The expression of the high-pass filter is as follows:

[0018] H(Z) = 1 - αZ -1 (1)

[0019] Where α ranges from (0.9, 1).

[0020] (2) When the signal is complex and variable during bird calls, it is necessary to perform frame segmentation operations on it. The pre-emphasized bird sound signal is subjected to frame segmentation and windowing operations. The window function is selected as the Hamming window, the frame length is set to 23 ms, and the frame shift is set to 11 ms.

[0021] (3) Each frame of the bird sound signal with the window function added after preprocessing is processed separately. The discrete Fourier transform is used to replace the original Fourier transform to implement the discrete STFT. The expression is as follows:

[0022]

[0023] Where x(n) is the input signal, l represents the frame shift amount, k represents the current spectral line number, N = 44100, and n represents the number of sampling points and the current nth frame respectively. The STFT spectrogram is obtained by using the relationship between the amplitude change with respect to time and frequency and the relationship between the energy magnitude with respect to time and frequency.

[0024] Furthermore, the bridging Transformer module includes a ConvBlock, a FormerBlock, a Conv to Former structure, and a Former to Conv structure:

[0025] The ConvBlock adopts a three-layer inverted bottleneck convolution structure, using the characteristics of depthwise separable convolution to reduce the computational complexity during channel convolution operations while retaining the high efficiency of convolution itself.

[0026] The FormerBlock includes a multi-head attention MHA module and a multi-layer perceptron module MLP of the Transformer encoder structure. The MHA calculates the incoming feature maps in parallel, and the MLP is used to integrate and screen the obtained information. The calculation formulas for the attention mechanism and multi-head attention in the MHA are as follows:

[0027]

[0028] MHA(Q, K, V) = Concat(head1, head2, …, head h )W 0 (4)

[0029] Among them, in formula (3), Q, K, and V are all weight matrices, and d k is the dimension corresponding to the K matrix. For the calculation of multi-head attention in formula (4), the above group of weight matrices is extended to use multiple groups, and each group is calculated synchronously to speed up. Finally, multiple output matrices are linearly transformed and concatenated using the W 0 matrix and sent into the subsequent MLP. For each head i there is: Set the number of attention heads h = 4, and each attention head is set with d k = d v = d / h = 32. Since each attention head performs dimensionality reduction, the output matrix finally obtained by concatenation is basically the same size as that calculated by a single attention mechanism. The specific expression of the MLP is as follows:

[0030] MLP(x) = [ReLU(xW1 + b1)]·W2 + b2 (5)

[0031] Among them, W1 and W2 are weight matrices, and b1 and b2 are bias vectors. The overall MLP model is composed of two linear layers and a ReLU activation function nested together.

[0032] The Conv to Former structure fuses the local feature information x i input by the ConvBlock with the learnable token z i of the FormerBlock. The operation is completed by attention mapping, but only the input tokens are query-mapped, and the feature maps are not mapped, which can reduce the internal computational complexity during the attention operation. Finally, the residual structure is used to integrate the local information into the global information of the tokens. The specific operation formula is as follows:

[0033]

[0034] Among them, H is the number of multi-head attention heads. According to the size of H, the input features x and z are evenly divided into x h and z h , is the query projection matrix of the h-th head, and W O is used to finally combine multiple heads together.

[0035] The Former to Conv structure realizes the global information token z output by the FormerBlock i+1 and the feature x' of the ConvBlock for extracting local information i In the process of fusion, similar to the Conv to Former structure, the ConvBlock feature does not obtain the global feature information after passing through the ConvBlock module. Therefore, the second bridging structure still uses the attention mapping operation to map the global information token z of the FormerBlock i+1 to obtain the key and value through two mappings, and the output x' of the ConvBlock module i is directly used as the query to guide the integration of global information, and finally the final output feature x is obtained through the residual structure for feature fusion i+1 , and the specific operation formula is as follows:

[0036]

[0037] Among them, and are the projection matrices of the key and value, the local feature provides the query, and other operations are the same as formula (6).

[0038] Furthermore, the specific operation of the sample loss optimization module SLOBlock is as follows:

[0039] (1) Input the feature map into the sample loss optimization module. Each output feature map is based on a batch, and the relationship between different category samples of each batch of spectrograms is modeled. The cross-attention mechanism in the single-layer Transformer encoder is used to assign the corresponding weight attention to the spectrograms of different types of bird sounds in the same batch

[0040] (2) When the cross-attention mechanism is not added, all losses only propagate gradients on the corresponding samples and categories, that is, one-to-one sample optimization. After using the SLOBlock, the losses on other samples also start to provide gradient-optimized loss feedback, and the gradient optimization formula is as follows:

[0041]

[0042] Furthermore, the evaluation metrics of the BTNN bird sound recognition network model in the S4 experiment include accuracy and F1-score. The F1-score is weighted by two metrics, precision and recall. The evaluation formula is as follows:

[0043]

[0044]

[0045] Among them, TP represents the number of correctly classified samples, and both FP and FN represent the number of misclassified samples.

[0046] Advantages of the present invention:

[0047] 1. The recognition method of the present invention generates a spectrogram by extracting the short-time Fourier transform of the bird call signal as the input feature of the overall network model;

[0048] 2. The recognition method of the present invention extracts local features based on the convolutional module, extracts global features based on the Transformer module, and uses the attention mechanism to build a bridging structure to realize the interactive sharing of information between the two modules;

[0049] 3. The recognition method of the present invention uses a single-layer Transformer encoder to share losses for each batch of feature maps obtained by training the bridging structure, and optimizes by superimposing losses between different categories inside the network, so as to expand samples without using data augmentation and improve the accuracy of testing small sample datasets in the model;

[0050] 4. The recognition method of the present invention obtained recognition accuracies of 91.34% and 82.63% on the Birdsdata dataset and the xeno-canto dataset processed with small samples, proving the effectiveness of the bird sound recognition model of the present invention and its applicability to real-time bird sound monitoring and recognition. Description of the Drawings

[0051] The present invention will be further described below with reference to the accompanying drawings.

[0052] Figure 1 is the network diagram of the BTNN model of the recognition method of the present invention;

[0053] Figure 2 is the structural diagram of the bridging Transformer module of the recognition method of the present invention;

[0054] Figure 3 is the schematic diagram of the SLOBlock of the recognition method of the present invention. Detailed Embodiments

[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0056] A small sample optimization method for bird sound recognition based on bridged Transformer, such as Figure 1 As shown, the identification method comprises the following steps:

[0057] S1. Obtain the BTNN bird sound recognition network model, extract the short-time Fourier transform of the bird sound signal to generate a spectrogram as the input feature of the overall network model

[0058] The original audio signal is first pre-processed by pre-emphasis, frame splitting and windowing, and then the STFT spectrogram is obtained by short-time Fourier transform as the input feature of the overall network; it is then input into a single ordinary convolutional layer to obtain an operational feature map.

[0059] For the original bird sound audio signal obtained, the bird sound spectrum within each frame time is considered to be unchanged, but this short-time spectrum that is considered to be unchanged can only be used to reflect the static characteristics of the bird sound when it is singing. In order to reflect the dynamic frequency characteristics of the bird sound signal and realize the analysis of non-stationary time-varying signals, the short-time Fourier transform is used to generate the STFT spectrogram to obtain the time-frequency characteristics of the bird signal. The STFT spectrogram is specifically operated as follows:

[0060] (1) When the energy of the bird sound signal is mostly concentrated in the low frequency band, while the energy of the high frequency part is relatively small, in order to compensate for the influence of the damaged high frequency signal transmission, a high-pass filter is used to pre-emphasize the bird sound audio signal. The expression of the high-pass filter is as follows:

[0061] H(Z)=1-αZ -1 (1)

[0062] Among them, the value range of α is (0.9,1), and 0.935 is taken here.

[0063] (2) When birds are singing, the signal is complex and changeable, so it needs to be framed. The pre-emphasized bird sound signal is framed and windowed. The window function is selected as the Hamming window, the frame length is set to 23 ms, and the frame shift is set to 11 ms.

[0064] (3) After preprocessing, the bird sound signal with the window function added to each frame is processed separately, and the discrete Fourier transform is used to replace the original Fourier transform to realize the discrete STFT. The expression is as follows:

[0065]

[0066] Among them, x(n) is the input signal, l represents the frame shift amount, k represents the current spectral line number, N = 44100, and n represents the number of sampling points and the current nth frame respectively. By using the relationship between the amplitude change with respect to time and frequency and the relationship between the energy magnitude with respect to time and frequency, the STFT spectrogram can be obtained, with the size set to 224×224 pixels. In the spectrogram, the abscissa is time, the ordinate is frequency, and the depth of the color in the figure represents the magnitude of the bird sound intensity.

[0067] S2. Use the bridging Transformer structure to achieve the information extraction, complementation and fusion of the local and global features of the STFT spectrogram, and obtain the bird sound feature parameters.

[0068] Based on the convolutional module (ConvBlock), the extraction of local features is completed. The feature map obtained for the first time is input into the bridging Transformer module (FormerBlock) composed of multiple cascaded Conv-Former Blocks to learn local and global features. The Transformer module completes the extraction of global features. The bridging structure is built using the attention mechanism to achieve the interactive sharing of information between the two modules, and the output feature map with complemented feature information is obtained through the concatenation operation.

[0069] As Figure 2 shown, the bridging Transformer module includes ConvBlock, FormerBlock, Conv to Former structure and Former to Conv structure. Figure 2 For the first bridging Transformer module, the input z i is an initialized two-dimensional vector. Among them, M represents the number of tokens, and d represents the dimension of the tokens. Here, M = 3 and d = 128. Compared with Transformer networks such as ViT that take images as input, only tokens are used to conduct the local features provided by ConvBlock, and the number of tokens used is much smaller than the number of input tokens of Transformer in conventional image processing tasks. FormerBlock only needs to consider the tokens mapped to the intermediate MHA attention calculation part and hardly needs to consider the high computational complexity caused by the too large input size, and can completely obtain the local feature information to achieve the subsequent interactive complementation with the global features. The z obtained after passing through FormerBlock i+1The output token is the global feature after finally completing the local information. For the FormerBlock in the subsequent module, the initial value of the input token is completely filled by the output of the previous module, and the token dimension parameters of all modules are the same.

[0070] The ConvBlock refers to the MobileNetV3 network and adopts a three-layer inverted bottleneck convolution structure. It uses the characteristics of depthwise separable convolution to reduce the computational complexity during channel convolution operations while retaining the efficiency of convolution itself. Here, x i The input is the two-dimensional matrix after the spectrogram is deformed. The local feature x' obtained after passing through the ConvBlock i is used as the input for the subsequent bridging module.

[0071] The FormerBlock consists of a multi-head attention (MHA) module and a multi-layer perceptron (MLP) module of the Transformer encoder structure. The MHA can grasp the key regional information of the global input by parallel computing the input feature map; the MLP functions similar to a fully connected layer to integrate and screen the obtained information. The calculation formulas for the attention mechanism and multi-head attention in the MHA are as follows:

[0072]

[0073] MHA(Q, K, V) = Concat(head1, head2, …, head h )W 0 (4)

[0074] Among them, in formula (3), Q, K, and V are all weight matrices, and d k is the dimension corresponding to the K matrix. For the calculation of multi-head attention in formula (4), it is to expand the above group of weight matrices into multiple groups and calculate each group synchronously to speed up. Finally, the multiple output matrices are linearly transformed and concatenated using the W 0 matrix and then sent into the subsequent MLP. For each head i there is: Set the number of attention heads h = 4, and each attention head is set with d k = d v = d / h = 32. Since each attention head has been dimension-reduced, the size of the finally concatenated output matrix is basically the same as that calculated by a single attention mechanism. The specific expression of the MLP is as follows:

[0075] MLP(x) = [ReLU[xW1 + b1)]·W2 + b2 (5)

[0076] Among them, W1 and W2 are weight matrices, and b1 and b2 are bias vectors. The overall MLP model consists of two linear layers and a ReLU activation function nested together.

[0077] The Conv to Former structure fuses the local feature information x i input by the ConvBlock with the learnable token z i of the FormerBlock. The operation is completed using attention mapping. However, only the input tokens are mapped for query, and the feature maps are not mapped, which can reduce the internal computational complexity during the attention operation. Finally, the residual structure is used to integrate the local information into the global information of the tokens. The specific operation formula is as follows:

[0078]

[0079] Among them, H is the number of multi-head attention heads. According to the size of H, the input features x h and z h are evenly divided into x is the query projection matrix of the h-th head, and W O is then used to combine the multiple heads together finally.

[0080] The Former to Conv structure realizes the process of fusing the global information token z i+1 output by the FormerBlock with the feature x' i extracting local information of the ConvBlock. Similar to the Conv to Former structure, the ConvBlock features do not obtain complementary global feature information after passing through the ConvBlock module. Therefore, the second bridging structure still uses the attention mapping operation to map the global information token z i+1 of the FormerBlock twice to obtain the key and value. The output x' i of the ConvBlock module directly serves as the query to guide the integration of the global information. Finally, the feature fusion is performed through the residual structure to obtain the final output feature x i+1 , and the specific operation formula is as follows:

[0081]

[0082] Among them, and are the projection matrices of the key and value. The local features provide the query, and other operations are the same as formula (6).

[0083] The x obtained from the bridging Transformer module composed of the above four parts i+1 and z i+1 are then passed into the next module, and through multiple interactions, a feature fusion process that complements global information and local information is achieved. During this process, since the feature map of the ConvBlock itself has never undergone any mapping, it ensures its original local information for the learning of the FormerBlock token. At the same time, the information representing global features in the FormerBlock token is continuously updated and fused into the feature map of the ConvBlock in multiple modules, better realizing the fusion and utilization of the provided spectrogram features. Secondly, the optimization of the parameters used is considered multiple times in each module, making the number of parameters of this module much smaller than that of traditional deep learning networks.

[0084] S3. Introduce a Sample Loss Optimization Block (SLOBlock). Using the cross-attention mechanism of a single-layer Transformer encoder, it realizes relationship modeling for the output feature map from the backbone network, internally optimizes the training and testing of the network itself on small sample datasets, and can achieve the implicit expansion of its own internal samples without expanding external sample data, thereby quickly optimizing the sample loss. The mechanism principle of the SLOBlock is as Figure 3 shown.

[0085] Input the feature map into the sample loss optimization module. Use the cross-attention mechanism in a single-layer Transformer encoder to assign weight attention to the spectrograms of different types of bird sounds in the same batch. The single-layer Transformer encoder shares the loss for each batch of feature maps obtained from the training of the bridging structure. By modeling the relationships between samples, the purpose of superimposing the sample losses of different types of bird sounds is achieved, enabling the implicit expansion of the data of corresponding types of bird sounds without using external data augmentation methods such as data enhancement, thereby realizing the gradient optimization of the loss function and improving the accuracy of testing small sample datasets in the model.

[0086] Each output feature map is based on a batch (batch size) unit, and models the relationships between different category samples for the spectrograms of each batch. When the cross-attention mechanism is not added, all losses only propagate gradients on the corresponding samples and categories, that is, one-to-one sample optimization; while after using the SLOBlock, other samples also start to provide loss feedback for gradient optimization (dashed line). The gradient optimization formula is as follows:

[0087]

[0088] Specifically, given N samples (from 0 to N-1) in a batch, the SLOBlock brings new gradient terms, that is, for sample X i in terms of L i also optimizes the network according to samples X of other classes j (i≠j), and implicitly expands N-1 virtual samples for each label from within the network by modeling the relationships between the mini-batch samples.

[0089] After tuning the internal loss of the output feature map of each batch through the sample optimization module and sharing the weight parameters on the shared Softmax classifier, it is possible to speed up the recognition speed during the final test without additionally changing the inference structure of the network when testing the recognition results.

[0090] S4. Experiments were conducted on the Birdsdata dataset and the xeno-canto dataset. The optimized features were input into the Softmax classifier to obtain the recognition results. Recognition accuracies of 91.34% and 82.63% were obtained on the Birdsdata dataset and the xeno-canto dataset processed with few samples, respectively, demonstrating the effectiveness of the bird sound recognition model proposed in this paper.

[0091] (1) Experimental data collection:

[0092] Birdsdata is a manually annotated standard dataset of natural sounds released by Beijing Birds Data Technology Co., Ltd. This dataset publicly collects a total of 20 types of common bird songs in China, with a total of 14,311 wav audio files. All the provided data has been standardized and segmented for 2s and noise reduction processing. The xeno-canto bird sound data comes from the global field bird sound database, which contains 44 common bird audio recordings in Eurasia and are all recorded in natural environments, with a total of 7,032 mp3 audio files, and the duration ranges from 30s to 5min. There are blanks in the middle of the audio and it comes with environmental noise. The sampling frequency of the above datasets is 44.1kHz

[0093] (2) Experimental settings

[0094] The operating system of the experimental hardware is Ubuntu20.04, the GPU model is GTX2080Ti, the CUDA version is 10.1, and the entire network model is built using the Pytorch1.8.0 deep learning framework. During the overall training process, the number of iterations (epoch) is set to 100, the single training step size (batch_size) of the input data is set to 32, the Adam algorithm is used as the optimizer to update the weight parameters, the momentum is set to 0.9, and the learning rate (learning_rate) uses the step decay method, with the initial learning rate set to 10 -4, and it decays to 0.1 times the previous learning rate at 56% and 78% of the total number of iterations. The Dropout layer is set to 0.2.

[0095] (3) Experimental evaluation

[0096] Accuracy and F1-score are used as evaluation metrics to assess the performance of the self-model and compare with other models. The F1-score is obtained by weighting two metrics, Precision and Recall. The evaluation formula is as follows:

[0097]

[0098]

[0099] Among them, TP represents the number of correctly classified samples, and both FP and FN represent the number of misclassified samples. In the specific experiment, the overall dataset is divided into a training set and a test set in a ratio of 8:2, and five-fold cross-validation is used to conduct five experiments respectively. The test results after each training and the final average value are recorded. Finally, to verify the effectiveness of the BTNN model, the BTNN model of the present invention is compared with other deep learning-based methods (BiLSTM-DenseNet, CNN-LSTM, VGGNet, CRNN, and Transformer-CNN).

[0100] (4) Verifying the effectiveness of the BTNN model

[0101] 1) Comparative experiment on the original sample dataset: Without any processing on the original large sample database, five experiments are conducted through five-fold cross-validation to calculate the average value and standard deviation. The specific comparison test results are shown in Table 1 below:

[0102] Table 1

[0103]

[0104] As can be seen from Table 1, compared with the above methods, the BTNN model proposed in the present invention has corresponding improvements in the accuracy in the original sample dataset; on the Birdsdata dataset, the accuracy of the BTNN network model can reach 96.89%, and the F1-score can reach 96.30%, second only to the Transformer-CNN network model. However, the number of model parameters of Transformer-CNN is much larger than that of BTNN, and there will be data comparison and explanation in the subsequent experiments. On the xeno-canto dataset, the accuracy of the BTNN network model can reach 91.64%, and the F1-score can reach 90.55%, both higher than the existing deep learning methods. By comparing the sample styles and the number of samples in the two datasets, it can be found that when the generated STFT spectrogram information in the xeno-canto dataset is affected by factors such as insufficient samples, noise interference, and blank spaces, it is very important to extract more critical time-frequency domain information. The high accuracy indicates that BTNN has better robustness when trained in a dataset with interference. In addition, a model comparison experiment after removing the SLOBlock was added. It can be found that on the Birdsdata dataset with sufficient sample data volume, whether the SLOBlock is added or not does not have much impact on the improvement of accuracy. On the xeno-canto dataset, even without conducting small-sample experiments, the accuracy has increased by 1.7%, which provides an experimental basis for subsequent small-sample dataset experiments.

[0105] 2) Comparison experiment on small-sample datasets: The original bird sound dataset is randomly divided. That is, the Birdsdata dataset is randomly divided into 10 parts (each part has 1431 samples), and the xeno-canto dataset is randomly divided into 5 parts (each part has 1406 samples). Five-fold cross-validation is performed on each small-sample set of each dataset, and the results are averaged and the standard deviation is calculated. At the same time, experiments are conducted on other deep learning methods under the same conditions. The specific comparison experiment results are shown in Table 2 below.

[0106] Table 2

[0107]

[0108]

[0109] As can be seen from Table 2, when the sample data volume is reduced to 10% - 20% of the original, the overall accuracy of the bird sound recognition method based on deep learning decreases. The sample quantity is an important factor affecting the training and testing accuracy of the neural network. By comparing the above methods, it is found that methods with deeper depths and larger network models such as VGGNet and Transformer-CNN are more significantly affected by the lack of sample data. Although the Transformer-CNN network model performs well in the original sample dataset, its accuracy in testing on the small sample dataset is not as good as other methods. In contrast, the BTNN model proposed in the present invention is least affected on the two datasets, obtaining the highest average accuracies of 91.34% and 82.63% on the Birdsdata dataset and the xeno-canto dataset respectively, which are only reduced by 5.55% and 9.01% compared to the testing accuracy during the original sample training. Thus, it can be proven that the BTNN model has a certain optimization in the training of small sample datasets.

[0110] 3) Ablation experiment and parameter comparison of SLOBlock:

[0111] To prove that SLOBlock plays an optimization role in the training of small sample datasets in the BTNN model and at the same time prove the versatility of this module in deep neural networks, the present invention conducts an ablation experiment on SLOBlock alone. The dataset for the experiment is still the small sample dataset randomly divided as mentioned above. The experimental results are shown in Table 3 below. It can be seen from the table that on the xeno-canto dataset, the accuracy of each method has improved, and on the Birdsdata dataset, except for the MobileNetV3 convolutional neural network, the others have also improved slightly, especially on the Transformer type neural network with a relatively large demand for sample data volume, the improvement is more obvious. Through data experiments, it is proven that SLOBlock has good versatility in different deep neural networks.

[0112] Table 3

[0113]

[0114]

[0115] 4) The BTNN optimizes the computational complexity and the number of parameters of the model in the following aspects:

[0116] (1) Only use the depthwise separable convolution mentioned in MobileNetV3 to construct the ConvBlock, reducing the number of parameters for channel convolution calculation;

[0117] (2) FormerBlock only uses a very small number (6 - 8) of tokens for calculating the attention parameters output by the conduction bridging structure;

[0118] (3) The bridging Transformer module itself only performs attention mapping on tokens, and no operations are performed on the feature map itself. Therefore, the overall computational complexity of the two bridging structures is only equivalent to that of one attention mechanism. To demonstrate that the present invention plays a certain role in optimizing the network of the BTNN model, the internal parameters and operation time of BTNN and each comparison method are shown in Table 4 below:

[0119] Table 4 Comparison table of the number of parameters of each method

[0120]

[0121] As can be seen from the above table, some detailed optimizations mentioned in the present invention can help the BTNN model reduce the overall number of parameters, and at the same time, the model itself processes spectrogram features relatively fast, second only to the lightweight MobileNetV3 network. This provides support for subsequent experiments to use the model of the present invention for real-time bird sound monitoring and recognition.

[0122] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0123] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A few-shot optimized bird sound recognition method based on a bridging Transformer, characterized in that, The recognition method includes the following steps: S1. Obtain the BTNN bird sound recognition network model, and extract the short-time Fourier transform of the bird sound signal to generate a spectrogram as the input feature of the overall network model; S2. Use the bridging Transformer structure to extract, complement, and fuse the information of the local and global features of the STFT spectrogram to obtain bird sound feature parameters. The bridging Transformer structure includes a ConvBlock, a FormerBlock, a Conv to Former structure, and a Former to Conv structure; S3. Introduce the sample loss optimization module SLOBlock, and use the cross-attention mechanism of the single-layer Transformer encoder to perform relationship modeling on the output feature map from the backbone network, and optimize the training and testing of the small sample data set by the network itself internally; S4. Conduct experiments on the Birdsdata data set and the xeno-canto data set, and input the optimized features into the Softmax classifier to obtain the recognition result.

2. The small-sample optimized bird sound recognition method based on a bridged Transformer according to claim 1, wherein The specific operation of S1 is as follows: Perform pre-emphasis, framing, and windowing preprocessing operations on the original audio signal; Obtain the STFT spectrogram through short-time Fourier transform as the input feature of the overall network; Input it into a single ordinary convolutional layer to obtain an operable feature map.

3. The small-sample optimized bird sound recognition method based on a bridging Transformer according to claim 2, wherein The specific operation of the STFT spectrogram includes the following steps: (1) When the energy of the bird sound signal is mostly concentrated in the low-frequency band, a high-pass filter is used to perform pre-emphasis processing on the bird sound audio signal. The expression of the high-pass filter is as follows: H(Z) = 1 - αZ -1 (1) where α takes values in the range of (0.9, 1); (2) When the signal is complex and variable during bird calls, it is necessary to perform framing operations on it. Perform framing and windowing operations on the pre-emphasized bird sound signal. The window function is selected as the Hamming window, the frame length is set to 23 ms, and the frame shift is set to 11 ms; (3) Process each frame of the bird sound signal with the window function added separately, and use the discrete Fourier transform to replace the original Fourier transform to implement the discrete STFT. The expression is as follows: where x(n) is the input signal, l represents the frame translation amount, k represents the current spectral line number, N = 44100, n represents the number of sampling points and the current nth frame respectively, and the STFT spectrogram is obtained by using the relationship between the amplitude change with respect to time and frequency and the relationship between the energy magnitude with respect to time and frequency.

4. The small-sample optimized bird sound recognition method based on a bridging Transformer according to claim 1, wherein The ConvBlock adopts a three-layer inverted bottleneck convolutional structure, and uses the characteristics of depthwise separable convolution to reduce the computational complexity during channel convolution operations while retaining the high efficiency of convolution itself; The FormerBlock includes a multi-head attention MHA module and a multi-layer perceptron module MLP of the Transformer encoder structure. MHA calculates the incoming feature maps in parallel, and MLP is used to integrate and screen the obtained information. The calculation formulas of the attention mechanism and multi-head attention in MHA are as follows: MHA(Q, K, V) = Concat(head1, head2, …, head h )W 0 (4) Among them, in formula (3), Q, K, and V are all weight matrices, and d k is the dimension corresponding to the K matrix. For the calculation of multi-head attention in formula (4), the above set of weight matrices is extended to use multiple sets, and each set is calculated synchronously to speed up. Finally, multiple output matrices are linearly transformed and spliced using the W 0 matrix and then sent to the subsequent MLP. For each head i there is: Set the number of attention heads h = 4, and each attention head is set to d k = d v = d / h = 32. Since each attention head performs dimensionality reduction, the size of the output matrix finally obtained by splicing is basically the same as that calculated by a single attention mechanism. The specific expression of the MLP is as follows: MLP(x) = [ReLU(xW1 + b1)]·W2 + b2 (5) Among them, W1 and W2 are weight matrices, and b1 and b2 are bias vectors. The overall MLP model consists of two linear layers and a ReLU activation function nested together; The described Conv to Former structure takes the local feature information x input by the ConvBlock i and the learnable token z of the FormerBlock i for fusion. The operation is completed using attention mapping, but only query mapping is performed on the input tokens, not on the feature maps, which can reduce the internal computational complexity during the attention operation. Finally, the residual structure is used to integrate the local information into the global information of the tokens. The specific operation formula is as follows: where H is the number of multi-head attention heads, and the input features x and z are evenly divided into x h and z h , is the query projection matrix of the h-th head, and W O is then used to finally combine multiple heads together; The Former to Conv structure realizes the global information token z output by the FormerBlock i+1 and the feature x' of the ConvBlock that extracts local information i The fusion process is similar to the Conv to Former structure. After passing through the ConvBlock module, the ConvBlock feature does not obtain the global feature information. Therefore, the second bridging structure still uses the attention mapping operation to map the global information token z of the FormerBlock i+1 to obtain the key and value through two mappings, and the output x' of the ConvBlock module i is directly used as the query to guide the integration of global information. Finally, the final output feature x is obtained through feature fusion using the residual structure i+1 , and the specific operation formula is as follows: Among them, and are the projection matrices of the key and value. The local features provide the query, and other operations are the same as those in formula (6).

5. The small-sample optimized bird sound recognition method based on a bridging Transformer according to claim 1, characterized in that The specific operation of the sample loss optimization module SLOBlock is as follows: (1) Input the feature map into the sample loss optimization module. Each time the output feature map is in units of a batch. The relationship modeling between different category samples of the spectrogram of each batch is carried out, and the cross-attention mechanism in the single-layer Transformer encoder is used to endow the weight attention corresponding to the spectrogram of different types of bird sounds in the same batch; (2) When the cross-attention mechanism is not added, all losses only propagate gradients on the corresponding samples and categories, that is, one-to-one sample optimization; after using SLOBlock, the losses that provide gradient optimization feedback also start on other samples. The gradient optimization formula is as follows:

6. The small-sample optimized bird sound recognition method based on a bridged Transformer according to claim 1, wherein The evaluation metrics of the BTNN bird sound recognition network model in the S4 experiment include the accuracy Accuracy and F1-score. The F1-score is obtained by weighting two metrics, precision Precision and recall Recall. The evaluation formula is as follows: Among them, TP represents the number of correctly classified samples, and both FP and FN represent the number of misclassified samples.

Citation Information

Patent Citations

  • Small sample rare bird identification method based on comparative learning

    CN114548256A

  • Multi-modal emotion recognition method and system based on improved Transform

    CN115272908A