ECAPA-TDNN bird sound identification method based on dynamic frequency band division

Through the ECAPA-TDNN model of dynamic frequency division band, the problem of insufficient identification of traditional bird sound recognition models under frequency distribution inhomogeneity and environmental noise interference is solved, and more efficient bird sound feature extraction and recognition is achieved.

CN120496542APending Publication Date: 2025-08-15GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717370.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional bird sound recognition models lack generalization ability when dealing with bird sound signals distributed in different frequency, insufficient key feature extraction, and face problems such as mixed frequency domain features, environmental noise interference and insufficient cross-band correlation modeling in complex wild environments.

Method used

The ECAPA-TDNN model based on dynamic frequency division band is constructed. By integrating the dynamic frequency band segmentation module, a multi-band parallel processing architecture and a cross-band gated fusion unit, the learned parameterized Sigmoid function is used to dynamically adjust the band boundaries, and combined with the multi-band exclusive convolution kernel and channel attention mechanism, feature extraction and fusion are performed.

Benefits of technology

It improves the recognition accuracy and robustness of the model in complex environments, effectively extracts the characteristics of different frequency bands of bird sounds, and enhances the recognition ability of different bird sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496542A_ABST
    Figure CN120496542A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bird sound recognition, and provides an ECAPA-TDNN bird sound recognition method based on dynamic frequency band division, and the method is characterized in that the method comprises the steps: recognition data set construction, which comprises the steps: downloading a target bird species audio file from an original bird species audio through a bird sound event detection model, constructing an identification data set containing the target bird species audio file; the construction of a dynamic sub-band ECAPA-TDNN model comprises the step of integrating a dynamic band segmentation module, a multi-band parallel processing architecture and a cross-band gating fusion unit on the basis of an existing ECAPA-TDNN model. And carrying out verification and post-processing on the dynamic sub-band ECAPA-TDNN model. According to the method, the dynamic sub-band ECAPA-TDNN model is constructed, so that the features of different frequency bands of the bird sound are effectively extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bird sound recognition, and in particular relates to an ECAPA-TDNN bird sound recognition method based on dynamic frequency band division. Background Art

[0002] In bird sound recognition tasks, audio signals typically contain rich time-frequency information, and the frequency distribution of different bird calls varies significantly. Therefore, how to properly divide frequency bands to extract key features directly affects the model's recognition performance. Due to the complexity of bird sounds and the interference of environmental noise, the acoustic features of current bird sound recognition networks are mostly trained on human voice features. However, human voice features are mainly concentrated in low frequencies. Therefore, traditional bird sound recognition models suffer from insufficient generalization ability and inadequate key feature extraction when processing bird sound signals with different frequency distributions.

[0003] In addition, although the traditional equally spaced frequency band division method is simple and easy to implement, it ignores the non-uniform frequency distribution of bird sounds, resulting in the loss of some key information or the increase of redundant calculations. Therefore, the existing bird sound recognition methods face the challenges of frequency domain feature mixing, environmental noise interference and insufficient cross-band correlation modeling in complex wild environments. Summary of the Invention

[0004] In view of the above-mentioned shortcomings of the existing technology, the present invention proposes an ECAPA-TDNN bird sound recognition method based on dynamic frequency band division. To address the problems of traditional bird sound recognition models in processing bird sound signals with different frequency distributions, such as insufficient generalization ability and inadequate key feature extraction, the technical solution designed by the present invention includes the following steps: Step 1: Construction of the recognition dataset, which includes downloading the target bird species audio files from the original bird species audio using the bird sound event detection model and constructing a recognition dataset containing the target bird species audio files; Step 2: Dynamic band-splitting ECAPA-TDNN model construction, which includes integrating the dynamic band-splitting module, multi-band parallel processing architecture, and cross-band gating fusion unit on the existing ECAPA-TDNN model; Step 3: Verification and post-processing of the dynamic frequency-band ECAPA-TDNN model, including the use of a double verification strategy. Strategy 1 is that the target bird species audio files with the dynamic frequency-band ECAPA-TDNN model output probability ≥ 0.7 are output as correctly identified samples. Strategy 2 is that the target bird species audio files corresponding to Strategy 1 with the dynamic frequency-band ECAPA-TDNN model output probability ≥ 0.5 within 30 consecutive seconds are also output as correctly identified samples.

[0005] Preferably, the step 1 further includes: The audio files of target bird species were manually verified, labeled, and cut. The cutting rule was to split long audio files every 30 seconds, retain the terminal segments with a length of ≥15 seconds, and delete the segments with a length of <2 seconds.

[0006] Preferably, the dynamic frequency band segmentation module in step 2 includes: The frequency band boundaries are dynamically determined through a learnable parameterized Sigmoid function, so that the model can adaptively adjust the frequency band boundaries according to the input target bird species audio file. The frequency band boundary parameters and steepness are adaptively optimized during training.

[0007] Preferably, the multi-band processing architecture in step 2 includes: After the dynamic frequency band segmentation module, dedicated convolution channels for low frequency, medium frequency and high frequency are constructed respectively. The low frequency channel uses a large convolution kernel with a kernel size of 15, the medium frequency channel uses a hollow convolution kernel with a kernel size of 3, and the high frequency channel uses a small convolution kernel with a kernel size of 3.

[0008] Preferably, the cross-band gating fusion unit in step 2 includes: The features of the low-frequency, medium-frequency, and high-frequency branches are globally averaged and pooled to obtain a global description at the channel level. Then, a lightweight fully connected network is used to calculate the attention weights of each frequency band and weightedly fuse the features. After inputting them into the backbone ECAPA-TDNN network, the final bird sound prediction result is obtained through Softmax.

[0009] Preferably, the core calculation formula of the dynamic frequency band segmentation module is as follows:

[0010] Where, is the learnable band boundary parameter, To control the steepness of the boundary.

[0011] Preferably, the multi-band processing architecture further includes: A channel attention mechanism is added to each frequency band processing branch.

[0012] Preferably, the attention weight of each frequency band is calculated using the following formula:

[0013] Where, and is a trainable parameter.

[0014] Preferably, the weighted fusion feature is formulated as follows:

[0015] Where, is the weight of each frequency band.

[0016] Preferably, the method further comprises: Preprocessing and data enhancement of audio files of target bird species; The data preprocessing includes uniformly sampling the target bird species audio files to 32,000 Hz, converting them to WAV format, randomly slicing them into 10-second segments, and performing amplitude normalization during the dynamic frequency band ECAPA-TDNN model training process; The data enhancement includes spectrum enhancement on the Mel-spectrogram, random masking of 0 to 5 frames in the time domain, random masking of 0 to 10 channels in the frequency domain, and mixing in background noise fragments with a probability of 0.5.

[0017] Beneficial effects: The present application provides an ECAPA-TDNN bird sound recognition method based on dynamic frequency division bands. A dynamic frequency division band ECAPA-TDNN model is constructed through a differentiable dynamic frequency segmentation module, a multi-band parallel processing architecture and a cross-band gated fusion unit to effectively extract the characteristics of different frequency bands of bird sounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of a preferred embodiment of the present invention; Figure 2 1 is a schematic structural diagram of a dynamic frequency division band ECAPA-TDNN model according to a preferred embodiment of the present invention.

[0019] Figure 3 This is a schematic diagram of the network structure of SE-Res2Block in a preferred embodiment of the present invention; Figure 4 Schematic diagram of the network structure of ECAPA-TDNN in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0020] The embodiments of the present invention are described in detail below. The following embodiments are implemented based on the technical solutions of the present invention, and provide detailed implementation methods and specific operating procedures. However, the protection scope of the present invention is not limited to the following embodiments.

[0021] The present invention designs an ECAPA-TDNN bird sound recognition method based on dynamic frequency band division. The technical solution includes the following steps: Figure 1-2 As shown, specifically including: Step 1: Construction of the recognition dataset, which includes downloading the target bird species audio files from the original bird species audio using the bird sound event detection model and constructing a recognition dataset containing the target bird species audio files; Step 2: Dynamic band-splitting ECAPA-TDNN model construction, which includes integrating the dynamic band-splitting module, multi-band parallel processing architecture, and cross-band gating fusion unit on the existing ECAPA-TDNN model; Step 3: Verification and post-processing of the dynamic frequency-band ECAPA-TDNN model, including the use of a double verification strategy. Strategy 1 is that the target bird species audio files with the dynamic frequency-band ECAPA-TDNN model output probability ≥ 0.7 are output as correctly identified samples. Strategy 2 is that the target bird species audio files corresponding to Strategy 1 with the dynamic frequency-band ECAPA-TDNN model output probability ≥ 0.5 within 30 consecutive seconds are also output as correctly identified samples.

[0022] Preferably, step 1 further includes: The audio files of target bird species were manually verified, labeled, and cut. The cutting rule was to split long audio files every 30 seconds, retain the terminal segments with a length of ≥15 seconds, and delete the segments with a length of <2 seconds.

[0023] Specifically, for step 1, we constructed the recognition dataset. We first trained a bird sound event detection model using three datasets obtained from the Bird Audio Detection Challenge 2018 - DCASE Task 3. We then obtained a catalog of bird species recorded in the project area from the China Birdwatching Recording Center. We then downloaded audio files for the corresponding bird species from Xeno-Canto. We found that most audio files were between 2 seconds and 5 minutes in length, and faced the following challenges: (a) Because each audio file is annotated by the uploader, it may not be accurate, so verification was required to minimize errors caused by mislabeling. (b) For longer audio files, other noise and silence accounted for the majority of the entire audio file, while bird calls accounted for a smaller portion. Therefore, we manually trimmed the long audio files to eliminate long, useless segments. To address these two issues, we manually inspected the spectrograms and waveforms of all the data, filtered out low-quality and mislabeled audio, and segmented the long audio files every 30 seconds. If the last segment of the audio file was less than 30 seconds and was longer than 15 seconds, it was retained; if it was less than 15 seconds, it was discarded. We also deleted any original audio segments shorter than 2 seconds. Finally, the bird sound recognition dataset is constructed. This dataset construction method is highly versatile and extensible. Bird sound datasets in various complex areas can be constructed according to this method, and it has broad application prospects.

[0024] Preferably, the dynamic frequency band segmentation module in step 2 includes: The frequency band boundaries are dynamically determined through a learnable parameterized Sigmoid function, so that the model can adaptively adjust the frequency band boundaries according to the input target bird species audio file. The frequency band boundary parameters and steepness are adaptively optimized during training.

[0025] Preferably, the multi-band processing architecture in step 2 includes: After the dynamic frequency band segmentation module, dedicated convolution channels for low frequency, medium frequency and high frequency are constructed respectively. The low frequency channel uses a large convolution kernel with a kernel size of 15, the medium frequency channel uses a hollow convolution kernel with a kernel size of 3, and the high frequency channel uses a small convolution kernel with a kernel size of 3.

[0026] Preferably, the cross-band gating fusion unit in step 2 includes: The features of the low-frequency, medium-frequency, and high-frequency branches are globally averaged and pooled to obtain a global description at the channel level. Then, a lightweight fully connected network is used to calculate the attention weights of each frequency band and weightedly fuse the features. After inputting them into the backbone ECAPA-TDNN network, the final bird sound prediction result is obtained through Softmax.

[0027] Preferably, the core calculation formula of the dynamic frequency band segmentation module is as follows:

[0028] Where, is the learnable band boundary parameter, To control the steepness of the boundary.

[0029] Preferably, the multi-band processing architecture further includes: A channel attention mechanism is added to each frequency band processing branch.

[0030] Preferably, the attention weight of each frequency band is calculated as follows:

[0031] Where, and is a trainable parameter.

[0032] Preferably, the weighted fusion features are expressed as follows:

[0033] Where, is the weight of each frequency band.

[0034] Specifically, such as Figure 2-4As shown, the dynamic frequency-band ECAPA-TDNN model in step 2 is based on the Time Delay Neural Network (TDNN) and is currently one of the best single-unit models for speaker recognition. Building on TDNN, it further utilizes dilated convolutions and introduces Res2Net to obtain multi-scale contextual information. The SENet module also allows the neural network to independently learn the weights for each feature channel. Therefore, the ECAPA-TDNN network can extract more subtle speech features than networks like ResNet and MobileNet, which is very helpful for identifying subtle differences between multiple bird calls.

[0035] Separately, the SE-Res2Block combines Res2Net with SENet, where the dilated convolutions contain dense layers before and after each with a 1-frame context. The first dense layer reduces the feature dimensionality, while the second restores the number of features to the original dimensionality. This is followed by a SE block to rescale each channel. The entire setup is overlaid with skip connections. The Res2Net module also enhances the central convolutional layer, enabling it to handle multi-scale features by constructing hierarchical residual connections internally. This integration improves performance while significantly reducing the number of model parameters.

[0036] In addition, based on the ECAPA-TDNN model, a dynamic frequency band division ECAPA-TDNN model is proposed. In this model, the differentiable dynamic frequency segmentation module realizes adaptive learning of frequency band boundaries through parameterized sigmoid function. The model can automatically optimize the frequency band division during training, so that the characteristics of different bird species can be extracted within the appropriate frequency band; a multi-band parallel processing architecture is constructed, and frequency band-specific convolution kernels and attention mechanisms are used to realize differentiated feature extraction, which is used to capture the long-term dependency patterns of low-frequency bands, expand the receptive field to enhance the resolution of mid-frequency regions, and retain the detailed structure of high-frequency information, and use Dropout to suppress noise.

[0037] Preferably, the method further comprises: Preprocessing and data enhancement of audio files of target bird species; Data preprocessing, including uniformly sampling the target bird species audio files to 32,000 Hz, converting them to WAV format, randomly slicing them into 10-second segments, and performing amplitude normalization during the dynamic frequency-splitting ECAPA-TDNN model training process; The data enhancement includes spectral enhancement on the Mel-spectrogram, random masking of frames 0 to 5 in the time domain, random masking of channels 0 to 10 in the frequency domain, and mixing in background noise fragments with a probability of 0.5.

[0038] Specifically, by preprocessing the audio files of target bird species, all audio samples are uniformly converted to a WAV format without high-frequency loss, eliminating the impact of amplitude differences in bird audio samples on model training. Data enhancement of the target bird species' audio files is also performed to reduce the impact of spatial mismatches caused by differences in the signal-to-noise ratio between training and field recordings, as well as differences in the sound collection environment (equipment, sampling rate, temperature, weather, etc.). Furthermore, speech rate perturbation, volume enhancement, and the addition of Gaussian noise are used in training to improve model robustness.

[0039] In addition, this application provides the following evaluation metrics to compare and evaluate the performance of each model in the application:

[0040]

[0041]

[0042]

[0043] in TP, TN, FP and FN Represent true positive, true negative, false positive and false negative samples respectively, and use accuracy (Accuracy), precision (Precision), recall (Recall) and F1 score evaluation indicators. Accuracy: Indicates the proportion of samples correctly classified by the model to the total samples. Precision: Indicates the proportion of all samples classified as positive that truly belong to the positive class. It is suitable for evaluating false alarms (false positives). Recall: Indicates the proportion of all true positive samples that are correctly classified as positive. It is suitable for evaluating omissions (false negatives). F1 score: It is the harmonic mean of precision and recall, which comprehensively considers the impact of false positives and false negatives, and is suitable for tasks with unbalanced categories.

[0044] The above describes in detail the preferred embodiments of the present invention. It should be understood that numerous modifications and variations based on the concepts of the present invention are possible by those skilled in the art without inventive effort. Therefore, any technical solution that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. An ECAPA-TDNN bird sound recognition method based on dynamic frequency band division, characterized in that: include: Step 1: Construction of the recognition dataset, which includes downloading the target bird species audio files from the original bird species audio using the bird sound event detection model and constructing a recognition dataset containing the target bird species audio files; Step 2: Dynamic band-splitting ECAPA-TDNN model construction, which includes integrating the dynamic band-splitting module, multi-band parallel processing architecture, and cross-band gating fusion unit on the existing ECAPA-TDNN model; Step 3: Verification and post-processing of the dynamic frequency-band ECAPA-TDNN model, including the use of a double verification strategy. Strategy 1 is that the target bird species audio files with the dynamic frequency-band ECAPA-TDNN model output probability ≥ 0.7 are output as correctly identified samples. Strategy 2 is that the target bird species audio files corresponding to Strategy 1 with the dynamic frequency-band ECAPA-TDNN model output probability ≥ 0.5 within 30 consecutive seconds are also output as correctly identified samples.

2. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 1, characterized in that: The step 1 further includes: The audio files of target bird species were manually verified, labeled, and cut. The cutting rule was to split long audio files every 30 seconds, retain the terminal segments with a length of ≥15 seconds, and delete the segments with a length of <2 seconds.

3. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 1, characterized in that: The dynamic frequency band segmentation module in step 2 includes: The frequency band boundaries are dynamically determined through a learnable parameterized Sigmoid function, so that the model can adaptively adjust the frequency band boundaries according to the input target bird species audio file. The frequency band boundary parameters and steepness are adaptively optimized during training.

4. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 3, characterized in that: The multi-band processing architecture in step 2 includes: After the dynamic frequency band segmentation module, dedicated convolution channels for low frequency, medium frequency and high frequency are constructed respectively. The low frequency channel uses a large convolution kernel with a kernel size of 15, the medium frequency channel uses a hollow convolution kernel with a kernel size of 3, and the high frequency channel uses a small convolution kernel with a kernel size of 3.

5. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 4, characterized in that: The cross-band gating fusion unit in step 2 includes: The features of the low-frequency, medium-frequency, and high-frequency branches are globally averaged and pooled to obtain a global description at the channel level. Then, a lightweight fully connected network is used to calculate the attention weights of each frequency band and weightedly fuse the features. After inputting them into the backbone ECAPA-TDNN network, the final bird sound prediction result is obtained through Softmax.

6. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 3, characterized in that: The core calculation formula of the dynamic frequency band segmentation module is as follows: Where, is the learnable band boundary parameter, To control the steepness of the boundary.

7. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 4, characterized in that: The multi-band processing architecture further includes: A channel attention mechanism is added to each frequency band processing branch.

8. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 5, characterized in that: The calculation formula for the attention weight of each frequency band is as follows: Where, and is a trainable parameter.

9. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 5, characterized in that: The weighted fusion feature is formulated as follows: Where, is the weight of each frequency band.

10. The ECAPA-TDNN bird sound recognition method based on dynamic frequency band division according to claim 1, characterized in that: The method further comprises: Preprocessing and data enhancement of audio files of target bird species; The data preprocessing includes uniformly sampling the target bird species audio files to 32,000 Hz, converting them to WAV format, randomly slicing them into 10-second segments, and performing amplitude normalization during the dynamic frequency band ECAPA-TDNN model training process; The data enhancement includes spectrum enhancement on the Mel-spectrogram, random masking of 0 to 5 frames in the time domain, random masking of 0 to 10 channels in the frequency domain, and mixing in background noise fragments with a probability of 0.5.