Bird sound recognition method based on multi-scale fusion
Patent Information
- Application Number
- CN202510163252.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-09
AI Technical Summary
Existing bird sound recognition technologies are difficult to fully perform long-term correlation modeling on the time and frequency axes of the spectrum, resulting in the possibility of missing important timing and frequency characteristics when processing complex and changeable bird sound data, while insufficient exploration of the impact of different scale features, and face the challenges of computing power and storage space when deploying on resource-constrained devices.
A bird sound recognition method based on multi-scale fusion is proposed. The CLDNN framework is adopted to establish long-term frequency dependence by adding jump connection layer and LSTM paths, and a multi-feature fusion module is introduced to fusion features of different scales. At the same time, lightweight student models are generated through knowledge distillation to reduce resource consumption.
It improves the sound recognition accuracy, reduces the possibility of feature omissions during long-term monitoring, reduces the computing power resource consumption required for computing power and storage, and realizes the rapid identification of bird sounds in a wild environment without network coverage.
Smart Images

Figure CN119964581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bird sound recognition, and in particular to a bird sound recognition method based on multi-scale fusion. Background Art
[0002] With the continuous leap forward in sound recognition technology, the application of sound technology in the fields of wildlife population assessment, protection, and biodiversity research has become more and more extensive, and has received deep attention from academia and practice. Birds, as a sensitive indicator of ecological health in nature, have vocal characteristics that vary from species to species. This unique distinction provides a stable and reliable basis for species identification. For this reason, an innovative method of bird monitoring - monitoring using acoustic detection equipment - is gradually becoming a research focus. This method cleverly uses acoustic monitoring equipment to automatically collect bird sounds, and accurately identifies bird species based on vocal characteristics, thereby efficiently monitoring birds.
[0003] However, although bird sound recognition technology based on deep learning has made some progress, existing methods still have limitations. It is difficult for them to fully perform long-term correlation modeling on the time and frequency axes of the spectrogram, which means that important timing and frequency features may be missed when processing complex and changing bird sound data. In addition, these technologies have not fully explored how features at different scales affect the final recognition results. At the same time, due to the constraints of computing power and storage space, these advanced recognition methods still face many challenges in achieving real-time bird sound recognition, especially when deployed on resource-constrained devices. Summary of the invention
[0004] In order to solve the problem in the prior art that features are easily missed during long-term monitoring of bird sound recognition models and computing resources are consumed greatly, the present invention proposes a bird sound recognition method based on multi-scale fusion.
[0005] The specific technical solution is as follows: A bird sound recognition method based on multi-scale fusion, the steps include:
[0006] S1: Acquire bird sound sample signals and pre-process the bird sound sample signals;
[0007] S2: construct a bird sound recognition model, wherein the bird sound recognition model adopts a hybrid architecture CLDNN with a frequency capture path, wherein the hybrid architecture CLDNN is composed of CNN, LSTM and DNN, and the time dependency and frequency dependency are obtained through the bird sound recognition model, and the obtained features are subjected to feature fusion to obtain training classification results;
[0008] S3: Use the trained bird sound recognition model as the teacher model and generate a lightweight student model through knowledge distillation.
[0009] Further, the preprocessing includes:
[0010] The attenuation of the bird sound sample signal in the sound production mechanism is compensated by the filter;
[0011] The compensated bird sound sample signal is divided into several frames, and all the framed sound signals are windowed respectively;
[0012] The windowed sound signal is subjected to fast Fourier transform, and the transform result is input into the Mel filter bank to obtain the Mel spectrum of bird sound.
[0013] Furthermore, the construction of the bird sound recognition model includes:
[0014] Add a skip connection layer to the feature extraction module of the CNN, and pass the output feature map of each combination module of the convolutional layer and the skip connection layer in the model to the next combination module and the multi-scale fusion module along different paths;
[0015] Set the time-dependent path and frequency-dependent path of the bird sounds for LSTM, rearrange the spectrogram dimensions, swap the time quantity and frequency quantity of the spectrogram feature tensor, and use LSTM to process the rearranged spectrogram to capture the time dependency and frequency dependency;
[0016] The output feature map with time dependency and frequency dependency is stacked according to the channel dimension as input, and feature fusion is performed in the channel and spatial directions respectively.
[0017] Furthermore, the feature extraction process of the CNN includes: setting a BN layer and a ReLU activation function for each convolutional layer of the CNN, and using the formula Extract detailed features, where k represents the size of the convolution kernel.
[0018] Furthermore, the feature fusion includes: stacking the features obtained by combining the convolution layer and the jump connection layer of the CNN with the time dependency and the frequency dependency according to the channel dimension, taking the stacking result as input, and performing feature fusion in the channel and spatial directions.
[0019] Furthermore, the combination acquisition of the features specifically includes:
[0020] Perform global maximum pooling and average pooling operations on the features to obtain two feature vectors;
[0021] The two feature vectors are passed through a fully connected layer with shared parameters, and the two feature vectors output by the fully connected layer are summed element by element and activated to generate a channel attention mask.
[0022] Furthermore, the combination acquisition of the features also includes:
[0023] Multiply the channel attention mask by the original feature tensor element-wise to obtain a feature map with channel attention;
[0024] Perform spatial pooling on the feature map with channel attention, stack the pooling results along the channel dimension, reduce the number of channels through 2D convolution, and add Sigmoid activation function to output the spatial attention mask;
[0025] The feature map is weighted by the channel attention mask and the spatial attention mask to obtain a fused feature map.
[0026] Furthermore, the construction of the bird sound recognition model also includes: generating a feature vector through a global pooling operation on the fusion result, inputting the feature vector into a fully connected layer, and obtaining the final classification result through a softmax activation function.
[0027] Furthermore, the generation of the student model includes:
[0028] Train the teacher model to fully learn the feature distribution of the data and obtain a teacher model with feature recognition ability and generalization performance;
[0029] The knowledge of the teacher model is transferred to the student model through joint training.
[0030] Furthermore, the knowledge transfer includes: setting a temperature coefficient T, softening the category probability distribution of the teacher model by the temperature coefficient T, and generating a soft target containing more information.
[0031] The above technical solution has the following advantages or technical effects:
[0032] 1. The present invention proposes a bird recognition model based on multi-scale fusion, which is improved on the CLDNN framework with high sound recognition accuracy, increases the long-term dependence of the LSTM path establishment frequency, and introduces a multi-feature fusion module to fuse features of different scales, thereby improving the sound recognition accuracy.
[0033] 2. The present invention performs multi-scale fusion transformation on CLDNN, extracts features from the bird sound spectrum, and then models the time and frequency correlation through two independent paths to capture time dependency and frequency dependency. Finally, the extracted features are input into the multi-scale fusion module for feature fusion in the channel and spatial directions to obtain the classification results, thereby reducing the possible feature omissions in the long-term monitoring process.
[0034] 3. The present invention reduces resource consumption by obtaining a lightweight bird recognition model through knowledge distillation of a bird recognition model based on multi-scale fusion. The lightweight bird recognition model can realize rapid recognition of bird sounds in a wild environment without network coverage, thereby reducing the consumption of computing power and computing resources required for storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flow chart of the method of the present invention;
[0036] Figure 2 This is a schematic diagram of the traditional CLDNN model;
[0037] Figure 3 It is a schematic diagram of the CLDNN model after the modification of the present invention;
[0038] Figure 4 is a schematic diagram of a multi-feature fusion module of the present invention;
[0039] Figure 5 It is a schematic diagram of a student model of the present invention. DETAILED DESCRIPTION
[0040] In order to make the technical solution of the present invention clearer, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] like Figure 1 As shown, a bird sound recognition method based on multi-scale fusion comprises the following steps:
[0042] S1: Acquire bird sound sample signals and pre-process the bird sound sample signals;
[0043] S2: construct a bird sound recognition model, wherein the bird sound recognition model adopts a hybrid architecture CLDNN with a frequency capture path, wherein the hybrid architecture CLDNN is composed of CNN, LSTM and DNN, and the time dependency and frequency dependency are obtained through the bird sound recognition model, and the obtained features are subjected to feature fusion to obtain training classification results;
[0044] S3: Use the trained bird sound recognition model as the teacher model and generate a lightweight student model through knowledge distillation.
[0045] In order to enable the model to extract high-quality feature parameters and improve the model's recognition performance of bird sounds, after obtaining the bird sound sample signal, the bird sound sample signal data is preprocessed to solve possible environmental noise interference, sound instability and other problems, including:
[0046] S11: The filter is used to compensate for the attenuation of the bird sound sample signal in the sound generation mechanism, mainly to compensate and emphasize the high-frequency part of the audio signal. The filter formula is as follows:
[0047]
[0048] Where b is the pre-emphasis factor of the filter, ranging from 0.9 to 1.
[0049] S12: Divide the compensated bird sound sample signal into several frames. Since the bird sound signal is non-stationary, it is impossible to directly use the fast Fourier transform (FFT) to obtain the relationship between the frequency and time. The original signal is divided into several frames, each frame is 20ms to 40ms long, and the frame shift is 10ms to 20ms. After the framing, the signal in each frame can be approximately regarded as a stationary signal, and each frame is subjected to the fast Fourier transform.
[0050] S13: Window all framed sound signals separately to reduce spectrum leakage. Windowing can smooth the signal boundary, reduce the sudden change between the signal and its part outside the window, and avoid spectrum leakage in fast Fourier transform calculation. The Hamming window formula is:
[0051]
[0052] Where N is the length of the window, n is the sampling point in the window, and w[n] is the weight after windowing.
[0053] S14: Perform fast Fourier transform on the windowed sound signal. The fast Fourier transform formula is as follows:
[0054]
[0055] in, is the nth sampling point of the signal in the time domain, is the corresponding frequency domain representation, N is the length of the signal, and f is the frequency.
[0056] S15: Input the transformation result into the Mel filter bank to obtain the Mel spectrum of bird sounds. The formula is as follows:
[0057] .
[0058] In order to fully focus on the local characteristics of different bird vocalizations and make full use of the long-term temporal dependencies of bird vocalizations, a hybrid architecture CLDNN is selected. Figure 2As shown in the figure, the hybrid architecture CLDNN method can effectively capture local features and long-term temporal correlation features. This method fully utilizes the complementarity between CNN, long short-term memory (LSTM) network and fully connected neural network (DNN). This method can fully utilize the ability of CNN to reduce frequency changes, the potential of LSTM in modeling long-term dependencies, and the advantages of DNN in mapping features to more separable spaces. However, when the number of classification categories increases, this method is prone to model non-convergence.
[0059] There are long-range correlations in the bird sound spectrogram along both the time axis and the frequency axis. CLDNN only models the long-term correlation of the bird sound spectrogram along the time axis, but not the long-term correlation of the frequency axis. Figure 3 As shown, a bird sound recognition model is constructed. The bird sound recognition model includes a hybrid architecture CLDNN and a multi-scale fusion module. The hybrid architecture CLDNN is composed of CNN, LSTM and DNN. The bird sound recognition model is trained, and the hybrid architecture CLDNN is transformed into a multi-scale fusion. The features of the bird sound spectrum are extracted, and the time and frequency correlation are modeled through two independent paths to capture the time dependency and frequency dependency. The extracted features are input into the multi-scale fusion module, and the features are fused in the channel and spatial directions to obtain the model training classification results.
[0060] In the feature extraction module of CNN, a skip connection layer is added, and the feature extraction formula is:
[0061]
[0062] Among them, k represents the size of the convolution kernel, It represents a 2D convolution with k=3 and stride S=1, and sets a BN layer and a ReLU activation function on each convolution layer. The skip connection layer includes an AvgPool layer and a convolution layer. The detailed feature extraction formula is , represents a 2D convolution with k=3 and stride S=1. The output feature map of the combination module of each convolution layer and the jump connection layer is transmitted to the next combination module and the multi-scale fusion module along different paths. The number of output channels of the convolution layer is , the S of the AvgPool layer in the skip connection layer is set to 2 to ensure the same size of the output feature map.
[0063] The bird sound spectrogram has long-range correlation along both the time axis and the frequency axis. Along the time axis, it contains global correlation, and along the frequency axis, it has harmonic correlation. The LSTM in the traditional CLDNN mainly performs time series modeling on the features extracted by CNN, capturing the long-term dependencies in the data while ignoring the frequency correlation of the sound.
[0064] An independent path is added to the hybrid architecture CLDNN in addition to the original path to capture the frequency dependency of bird sounds. In the frequency capture path, since LSTM can capture temporal dependency, the spectrogram dimensions are rearranged, the frequency dimension is treated as a time series, the time and frequency quantities of the spectrogram feature tensor are interchanged, and then the rearranged spectrogram is processed using LSTM to capture the long-term dependency between each frequency component and other components. The original time capture path is to perform LSTM operation on the feature tensor obtained after CNN feature extraction to obtain the long-term dependency of time. Finally, the time dependency and frequency dependency are added to capture both time dependency and frequency dependency at the same time. Capturing time dependency and frequency dependency at the same time can comprehensively perform long-term correlation modeling on the time and frequency axes of the spectrogram, reduce the omission of important temporal and frequency features when processing complex and changeable bird sound data, and increase the accuracy of bird sound recognition.
[0065] like Figure 4 As shown, in the feature fusion module, the features obtained by combining the convolution layer and the jump connection layer module and the time dependency and frequency dependency obtained by time and frequency modeling are stacked according to the channel dimension as input, and then the features are fused in the channel and spatial directions respectively.
[0066] The obtained features are subjected to global maximum pooling and average pooling operations to obtain two feature vectors, which are then passed through fully connected layers (MLPs) with shared parameters. The two feature vectors output by the fully connected layer are then element-wise summed and activated to generate a channel attention mask.
[0067] The channel attention mask is multiplied element-by-element with the original feature tensor to obtain a feature map with channel attention. The specific formula is:
[0068]
[0069] This will enhance the information of important channels and suppress the features of unimportant channels. The obtained feature map with channel attention is pooled in the spatial direction. The pooling operation includes global maximum pooling and global average pooling. The two pooling results are stacked along the channel dimension, and the number of channels is reduced by 2D convolution. Then, a spatial attention mask is obtained using the Sigmoid activation function. The spatial attention mask can be used to adjust the importance of the spatial region. The specific formula is: .
[0070] The feature map is weighted by channel attention and spatial attention mask to obtain the fused feature map, and then the feature map is generated into a feature vector through global pooling operation. The feature vector is input into a fully connected layer and the final classification result is obtained through the softmax activation function, completing the transformation of the entire multi-scale feature fusion of CLDNN.
[0071] In order to reduce resource consumption and realize rapid recognition of bird sounds in the wild environment without network coverage, the trained bird sound recognition model is defined as a teacher model, and a lightweight learning model is obtained through knowledge distillation. The specific process is as follows:
[0072] Student model design: Only one layer is retained for each module, reducing the number of parameters while maintaining feature extraction capabilities.
[0073] Teacher model training: First, the R-FCN model is fully trained to ensure that it can fully learn the feature distribution of the data, thereby improving the model's recognition ability and generalization performance.
[0074] Knowledge transfer and joint training: Joint training is used to transfer the knowledge of the teacher model to the student model. The generated student model is as follows: Figure 5 As shown in Figure 2. To achieve this goal, a temperature coefficient T is introduced to soften the category probability distribution of the teacher model, thereby generating a soft target that contains more information. By softening the category probability distribution, the student model can better learn the intrinsic representation of the teacher model, thereby improving its own performance.
[0075] The method of the present invention constructs a lightweight bird sound recognition model with few parameters and high performance. By developing monitoring equipment and deploying it in the wild, real-time bird sound detection can be achieved, and accurate recognition of bird sounds in the wild can be realized.
[0076] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A bird sound recognition method based on multi-scale fusion, characterized in that the steps include: S1: Acquire bird sound sample signals and pre-process the bird sound sample signals; S2: construct a bird sound recognition model, wherein the bird sound recognition model adopts a hybrid architecture CLDNN with a frequency capture path, wherein the hybrid architecture CLDNN is composed of CNN, LSTM and DNN, and the time dependency and frequency dependency are obtained through the bird sound recognition model, and the obtained features are subjected to feature fusion to obtain training classification results; S3: Use the trained bird sound recognition model as the teacher model and generate a lightweight student model through knowledge distillation.
2. The bird sound recognition method based on multi-scale fusion according to claim 1, characterized in that: The pre-processing comprises: The attenuation of the bird sound sample signal in the sound production mechanism is compensated by the filter; The compensated bird sound sample signal is divided into several frames, and all the framed sound signals are windowed respectively; The windowed sound signal is subjected to fast Fourier transform, and the transform result is input into the Mel filter bank to obtain the Mel spectrum of bird sound.
3. The bird sound recognition method based on multi-scale fusion according to claim 1, characterized in that: The construction of the bird sound recognition model includes: Add a skip connection layer to the feature extraction module of the CNN, and pass the output feature map of each combination module of the convolutional layer and the skip connection layer in the model to the next combination module and the multi-scale fusion module along different paths; Set the time-dependent path and frequency-dependent path of the bird sounds for LSTM, rearrange the spectrogram dimensions, swap the time quantity and frequency quantity of the spectrogram feature tensor, and use LSTM to process the rearranged spectrogram to capture the time dependency and frequency dependency; The output feature map with time dependency and frequency dependency is stacked according to the channel dimension as input, and feature fusion is performed in the channel and spatial directions respectively.
4. The bird sound recognition method based on multi-scale fusion according to claim 3 is characterized in that: The feature extraction process of the CNN includes: setting a BN layer and a ReLU activation function for each convolutional layer of the CNN, and using the formula Extract detailed features, where k represents the size of the convolution kernel.
5. The bird sound recognition method based on multi-scale fusion according to claim 3, characterized in that: The feature fusion includes: stacking the features obtained by combining the convolution layer and the jump connection layer of the CNN with the time dependency and the frequency dependency according to the channel dimension, taking the stacking result as input, and performing feature fusion in the channel and spatial directions.
6. The bird sound recognition method based on multi-scale fusion according to claim 5, characterized in that: The combination acquisition of the features specifically includes: Perform global maximum pooling and average pooling operations on the features to obtain two feature vectors; The two feature vectors are passed through a fully connected layer with shared parameters, and the two feature vectors output by the fully connected layer are summed element by element and activated to generate a channel attention mask.
7. The bird sound recognition method based on multi-scale fusion according to claim 6, characterized in that: The combined acquisition of the features further includes: Multiply the channel attention mask by the original feature tensor element-wise to obtain a feature map with channel attention; Perform spatial pooling on the feature map with channel attention, stack the pooling results along the channel dimension, reduce the number of channels through 2D convolution, and add Sigmoid activation function to output the spatial attention mask; The feature map is weighted by the channel attention mask and the spatial attention mask to obtain a fused feature map.
8. The bird sound recognition method based on multi-scale fusion according to claim 3 is characterized in that: The construction of the bird sound recognition model also includes: generating a feature vector through a global pooling operation of the fusion result, inputting the feature vector into a fully connected layer, and obtaining the final classification result through a softmax activation function.
9. The bird sound recognition method based on multi-scale fusion according to claim 1, characterized in that: The generation of the student model includes: Train the teacher model to fully learn the feature distribution of the data and obtain a teacher model with feature recognition ability and generalization performance; The knowledge of the teacher model is transferred to the student model through joint training.
10. The bird sound recognition method based on multi-scale fusion according to claim 9, characterized in that: The knowledge transfer includes: setting a temperature coefficient T, softening the category probability distribution of the teacher model through the temperature coefficient T, and generating a soft target containing more information.