A Bird Sound Detection Method and System Based on JDC-CRNN
Through the bird sound detection method based on JDC-CRNN, using the Mel spectrogram to train and optimize the model, the problem of insufficient detection accuracy and speed in the existing technology is solved, and efficient bird sound detection is achieved.
Patent Information
- Application Number
- CN202310084987.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-01-17
AI Technical Summary
The existing bird sound detection methods have shortcomings in taking into account both detection accuracy and detection speed, especially the method based on JDC-CNN and multi-feature fusion model has high computational complexity and cannot respond quickly on lightweight devices.
The bird sound detection method based on JDC-CRNN is adopted, and the audio data with bird sound annotation information is obtained for preprocessing, converted into a Meer spectrum diagram, and the optimization model is trained using the training set and verification set, the loss function and early stop mechanism are set, and the detector and classifier are built to reduce the computational complexity.
It realizes accurate detection of bird sounds, while reducing the complexity of calculation time, improving the detection speed, and is suitable for fast response of lightweight equipment.
Smart Images

Figure CN116246640B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bird sound detection, and more specifically, to a bird sound detection method and system based on JDC-CRNN. Background Art
[0002] As a widely distributed animal, birds play an important role in nature and are excellent natural indicators of environmental quality. When observing birds outdoors, birdwatchers often only hear their voices but not see their shadows. Therefore, it is of great significance to extract information from bird sounds. Through audio data analysis, the computer can achieve automatic processing of bird sounds and detect whether there are birds through audio. The emergence of machine learning and convolutional neural networks has made it possible to automatically detect bird sounds. Automatic bird sound detection has helped scholars to a certain extent in monitoring the changing trends of bird community populations and species diversity in biological systems.
[0003] The bird sound datasets used in bird sound detection tasks are usually weakly labeled because the audio sometimes contains other redundant information. The difficulty in bird sound audio detection lies in how to accurately and sensitively perceive relatively weak bird calls when there are too many and too strong interference signals in the audio; another potential aspect is to reduce the computational complexity of the algorithm model while ensuring accuracy, so that the model still has good response capabilities on lightweight devices, enabling researchers to conveniently detect bird sounds when using mobile devices outdoors and reducing the inconvenience caused by poor signals or the increased computational load on the cloud.
[0004] The existing solution, JDC-CNN (Joint Detection and Classification ConvolutionNeural Network), was proposed by Kong et al. in 2017. JDC-CNN first defines a baseline CNN model based on VGG (VisualGeometry Group Network) for comparison, which is divided into global max-pooling CNN and global average-pooling CNN according to the different global pooling layers; after introducing a detector based on the former as a classifier, JDC-CNN is obtained. JDC-CNN uses a VGG-CNN-like as a classifier and a single-layer CNN as a detector. In JDC-CNN, the detector determines whether a piece of audio needs further analysis or is directly skipped, and the classifier outputs the probability of the presence of bird sounds in the audio after further analyzing the audio. Compared with the baseline CNN model, the classification performance of JDC-CNN has a slight improvement; however, it is based on a large CNN network. To accurately detect sound events in the audio, it requires a large amount of time for calculation and training, and has a high computational complexity, which is not conducive to rapid response; especially the classifier network used as the classifier has a slow calculation speed, affecting the flexibility and response capabilities of the model.
[0005] The prior art discloses a bird sound recognition method based on multi-feature fusion and a combined model, including: preprocessing the read original bird sound audio, including pre-emphasis and framing and windowing; extracting four features of the mel cepstral coefficients, energy coefficients after mel filtering, short-time zero-crossing rate, and short-time spectral centroid of the bird sound, normalizing them respectively and then vertically splicing them to form a fusion feature; drawing an STFT spectrogram; respectively inputting the fusion feature and the drawn STFT spectrogram into two CNN models based on the Inception module for training. After training is completed, the probability arrays output by the two models are spliced to form a feature array, and this feature array is used as the input of the ANN model for training. After training is completed, the optimal parameters of the above three models are loaded; any bird sound audio to be measured is input into the three models after loading the optimal parameters to obtain the bird sound recognition classification result; the recognition model of this application is constructed based on a large CNN network. Although it can detect the bird sound in the audio, it requires a large amount of time for calculation and training, and the computational complexity is relatively high, which is not conducive to rapid response and cannot balance detection accuracy and detection speed. Summary of the Invention
[0006] The present invention aims to overcome the defect of the prior art that it cannot balance detection accuracy and detection speed when detecting bird sounds, and provides a bird sound detection method and system based on JDC-CRNN, which can achieve accurate detection of bird sounds, while reducing the computational time complexity and improving the detection speed.
[0007] To solve the above technical problems, the technical solution of the present invention is as follows:
[0008] The present invention provides a bird sound detection method based on JDC-CRNN, including:
[0009] S1: Obtain audio data with bird sound annotation information, and preprocess the audio data to obtain preprocessed audio data;
[0010] S2: Convert the preprocessed audio data into a mel spectrogram, and divide the mel spectrogram into a training set and a validation set according to a preset ratio;
[0011] S3: Use the training set to train the constructed bird sound detection model based on JDC-CRNN for a preset number of rounds, set a loss function, and adjust the network parameters of the bird sound detection model to obtain a trained bird sound detection model;
[0012] S4: Set an early stopping mechanism, and use the validation set to test the trained bird sound detection model to obtain an optimized bird sound detection model;
[0013] S5: Obtain the audio data to be detected and convert it into a Mel spectrogram; use the optimized bird sound detection model to detect the Mel spectrogram of the audio data to be detected to obtain the bird sound detection result.
[0014] Preferably, the preprocessing performed on the audio data includes a segmentation operation and a filtering operation.
[0015] Preferably, the constructed bird sound detection model based on JDC-CRNN includes a detector, a classifier, and an output layer connected in sequence; the output end of the detector is also connected to the input end of the output layer.
[0016] The detector is used to first perform a preliminary detection on the Mel spectrogram in the training set, and transmit the detector result to the output layer; the output layer processes the detector result and outputs a binary initial detection result, indicating whether the current audio data should be ignored or passed to the classifier for further detection; when the preliminary detection result is 0, it means there is no bird sound, ignore the current audio data, and backpropagate to adjust the network parameters of the detector and the output layer; when the preliminary detection result is 1, it means there is a bird sound, then return to the detector to start training the classifier; the classifier then determines in detail whether there is a bird sound in the audio, transmits the classifier result to the output layer, and outputs a binary final detection result; the role of the detector is to ignore the audio data without bird sound and save computing resources.
[0017] Preferably, the detector includes a first max pooling layer, and the output end of the first max pooling layer serves as the output end of the detector;
[0018] The classifier includes a first convolutional layer, a second max pooling layer, a second convolutional layer, a third max pooling layer, a third convolutional layer, a fourth max pooling layer, a stacking processing layer, a first recurrent layer, a second recurrent layer, and a fifth max pooling layer connected in sequence;
[0019] The output layer includes a forward propagation unit and an activation function connected in sequence; the input end of the forward propagation unit serves as the input end of the output layer.
[0020] Preferably, the specific method for training the constructed bird sound detection model based on JDC-CRNN with the training set for a preset number of rounds is as follows:
[0021] S3.1: Input the Mel spectrogram in the training set into the detector, and the detector compresses the Mel spectrogram in the training set along the time axis direction to detect whether there is a bird sound, obtain the detector result, and transmit it to the output layer;
[0022] S3.2: The output layer processes the detector results to obtain the initial detection results. If the initial detection results are 0, it is determined that there is no bird sound in the audio data corresponding to the current Mel spectrogram, and the initial detection is used as the final detection result. If the initial binarization result is 1, the detector is returned.
[0023] S3.3: The detector inputs the current Mel spectrogram into the classifier, restores it to the preprocessed audio data, and extracts local time-frequency features. Then, after compressing and stacking the local time-frequency features along the frequency axis direction, sequential global features in different directions are extracted, and the classifier results are output.
[0024] S3.4: The output layer processes the classifier results to obtain the final detection results. If the final detection results are 0, it is determined that there is no bird sound in the audio data corresponding to the current Mel spectrogram. If the final detection results are 1, it is determined that there is bird sound in the audio data corresponding to the current Mel spectrogram.
[0025] S3.5: Repeat steps S3.1 - 3.4 until the training of the preset number of rounds is completed.
[0026] Preferably, the specific method of step S3.3 is as follows:
[0027] S3.3.1: The detector inputs the current Mel spectrogram into the classifier and restores it to the corresponding preprocessed audio data.
[0028] S3.3.2: The preprocessed audio data is input into the first convolutional layer for feature extraction to obtain the first local time-frequency features.
[0029] S3.3.3: The second max pooling layer compresses and filters the first local sequential features along the frequency axis direction to obtain the first compressed local time-frequency features.
[0030] S3.3.4: The second convolutional layer extracts features from the first compressed local time-frequency features to obtain the second local time-frequency features.
[0031] S3.3.5: The third max pooling layer compresses and filters the second local time-frequency features along the frequency axis direction to obtain the second compressed local time-frequency features.
[0032] S3.3.6: The third convolutional layer extracts features from the second compressed local time-frequency features to obtain the third local time-frequency features.
[0033] S3.3.7: The fourth max pooling layer compresses and filters the third local time-frequency features along the frequency axis direction to obtain the third compressed local time-frequency features.
[0034] S3.3.8: The stacking layer stacks the third compressed local time-frequency features along the frequency axis direction to obtain the stacked local time-frequency features.
[0035] S3.3.9: The first recursive layer extracts features from the stacked local time-frequency features in the positive direction along the time axis to obtain the forward-order time series global features;
[0036] S3.3.10: The second recursive layer extracts features from the forward-order time series global features in the negative direction along the time axis to obtain the reverse-order time series global features;
[0037] S3.3.11: The fifth max pooling layer compresses the reverse-order time series global features along the time axis direction to obtain the classifier result.
[0038] The first max pooling layer is a single-layer time series max pooling layer, which compresses the Mel spectrogram along the time axis, so that the frequency dimension of the Mel spectrogram remains unchanged while the time dimension is 1, and it is sent to the output layer as the detector result. If the processed result by the output layer is 0, it is determined that there is no bird sound; if the processed result by the output layer is 1, it returns to the detector link to continue the further calculation and determination of the classifier; before entering the classifier, the Mel spectrogram will be restored to the preprocessed audio data; when entering the convolutional layer, the convolutional kernel will extract local time-frequency features, and after being extracted by multiple convolutional layers, it plays the role of a filter; the second max pooling layer, the third max pooling layer, and the fourth max pooling layer are frequency-axis max pooling layers, and the local time-frequency features will be compressed along the frequency axis; when entering the recursive layer, the compressed local time-frequency features are further extracted along the positive and negative directions of the time axis to obtain the time series global features; finally, it enters the fifth max pooling layer, and the time dimension is compressed to 1 along the time axis and sent to the output layer as the classifier result.
[0039] Preferably, a loss function is set, and the specific method for adjusting the network parameters of the bird sound detection model is as follows:
[0040] According to the final detection result of the Mel spectrogram of the audio data and the bird sound annotation information of the audio data, a binary cross-entropy loss function is constructed; the network parameters of the detector and the classifier are adjusted according to the binary cross-entropy loss value.
[0041] Preferably, the specific method of step S4 is as follows:
[0042] Input the Mel spectrogram in the validation set into the trained bird sound detection model. According to the final detection result corresponding to the Mel spectrogram in the validation set output by the trained bird sound detection model, combined with the bird sound annotation information corresponding to the Mel spectrogram in the validation set, draw an ROC curve; use the area under the ROC curve value AUC as the validation index. If the area under the ROC curve value AUC does not improve in the preset number of rounds of training, stop training, save the network parameters of the current detector and classifier, and obtain the optimized bird sound detection model; otherwise, adjust the network parameters of the detector and the classifier, and repeat the training process of step S3.
[0043] Preferably, the activation function in the output layer is the sigmoid activation function.
[0044] The present invention also provides a bird sound detection system based on JDC-CRNN for implementing the above-mentioned bird sound detection method based on JDC-CRNN, including:
[0045] A data acquisition and preprocessing module, configured to acquire audio data with bird sound annotation information, preprocess the audio data, and obtain preprocessed audio data;
[0046] A data conversion and division module, configured to convert the preprocessed audio data into a Mel spectrogram, and divide the Mel spectrogram into a training set and a validation set at a preset ratio;
[0047] A model training module, which uses the training set to train the constructed bird sound detection model based on JDC-CRNN for a preset number of rounds, sets a loss function, adjusts the network parameters of the bird sound detection model, and obtains a trained bird sound detection model;
[0048] A model testing module, configured to set an early stopping mechanism, use the validation set to test the trained bird sound detection model, and obtain an optimized bird sound detection model;
[0049] A bird sound detection module, configured to acquire audio data to be detected and convert it into a Mel spectrogram; use the optimized bird sound detection model to perform bird sound detection on the Mel spectrogram of the audio data to be detected, and obtain a bird sound detection result.
[0050] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0051] In this application, audio data with bird sound annotation information is acquired, preprocessed and then converted into a Mel spectrogram; the Mel spectrogram is used to train and test the constructed bird sound detection model based on JDC-CRNN to obtain an optimized bird sound detection model, which is used to perform bird sound detection on the Mel spectrogram of the audio data to be detected to obtain a bird sound detection result. The bird sound detection model of this application is based on the CRNN network, which realizes accurate detection of bird sounds, reduces the computational time complexity, and improves the detection speed. Description of the Drawings
[0052] Figure 1 It is a flowchart of a bird sound detection method based on JDC-CRNN described in Embodiment 1.
[0053] Figure 2 It is a schematic structural diagram of a bird sound detection model based on JDC-CRNN described in Embodiment 2.
[0054] Figure 3Schematic diagram of a bird sound detection system based on JDC-CRNN described in Embodiment 3. Detailed implementation manners
[0055] The accompanying drawings are only for illustrative purposes and should not be construed as limitations on this patent.
[0056] To better illustrate this embodiment, some components in the accompanying drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product.
[0057] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0058] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0059] Embodiment 1
[0060] This embodiment provides a bird sound detection method based on JDC-CRNN, as Figure 1 shown, including:
[0061] S1: Obtain audio data with bird sound annotation information, preprocess the audio data to obtain preprocessed audio data;
[0062] S2: Convert the preprocessed audio data into a Mel spectrogram, and divide the Mel spectrogram into a training set and a validation set at a preset ratio;
[0063] S3: Use the training set to train the constructed bird sound detection model based on JDC-CRNN for a preset number of rounds, set a loss function, and adjust the network parameters of the bird sound detection model to obtain a trained bird sound detection model;
[0064] S4: Set an early stopping mechanism, use the validation set to test the trained bird sound detection model to obtain an optimized bird sound detection model;
[0065] S5: Obtain the audio data to be detected, convert it into a Mel spectrogram; use the optimized bird sound detection model to perform bird sound detection on the Mel spectrogram of the audio data to be detected to obtain a bird sound detection result.
[0066] In the specific implementation process, this embodiment obtains audio data with bird sound annotation information, converts it into a Mel spectrogram after preprocessing; uses the Mel spectrogram to train and test the constructed bird sound detection model based on JDC-CRNN to obtain an optimized bird sound detection model, which is used to perform bird sound detection on the Mel spectrogram of the audio data to be detected to obtain a bird sound detection result. The bird sound detection model of this application is based on the CRNN network, which realizes accurate detection of bird sounds, reduces the computational time complexity at the same time, and improves the detection speed.
[0067] Example 2
[0068] This embodiment provides a bird sound detection method based on JDC-CRNN, including:
[0069] S1: Obtain audio data with bird sound annotation information, preprocess the audio data to obtain preprocessed audio data;
[0070] The preprocessing includes a segmentation operation and a filtering operation; in this embodiment, the audio data is divided into sub-audio data of 10s, and after filtering, it is used as the preprocessed audio data;
[0071] S2: Use the Mel filter bank method to convert the preprocessed audio data into a Mel spectrogram, and divide the Mel spectrogram into a training set and a validation set at a preset ratio;
[0072] S3: Use the training set to train the constructed bird sound detection model based on JDC-CRNN for a preset number of rounds, set a loss function, adjust the network parameters of the bird sound detection model, and obtain a trained bird sound detection model;
[0073] As Figure 2 shown, the bird sound detection model based on JDC-CRNN includes a detector, a classifier, and an output layer connected in sequence; the output end of the detector is also connected to the input end of the output layer;
[0074] The detector includes a first max pooling layer;
[0075] The classifier includes a first convolutional layer, a second max pooling layer, a second convolutional layer, a third max pooling layer, a third convolutional layer, a fourth max pooling layer, a stacking processing layer, a first recurrent layer, a second recurrent layer, and a fifth max pooling layer connected in sequence;
[0076] The output layer includes a forward propagation unit and an activation function connected in sequence; the activation function is a sigmoid activation function;
[0077] The detector is used to initially detect the Mel spectrogram in the training set and pass the detector results to the output layer; the output layer processes the detector results and outputs a binary initial detection result, indicating whether the current audio data should be ignored or passed to the classifier for further detection; when the initial detection result is 0, it means there is no bird sound, and the current audio data is ignored, and the network parameters of the detector and the output layer are adjusted through backpropagation; when the initial detection result is 1, it means there is a bird sound, and then it returns to the detector to start training the classifier; the classifier then determines in detail whether there is a bird sound in the audio, passes the classifier results to the output layer, and outputs a binary final detection result; the role of the detector is to ignore the audio data without bird sound and save computing resources;
[0078] The specific method for training the constructed bird sound detection model based on JDC-CRNN with the training set for a preset number of rounds is as follows:
[0079] S3.1: Input the Mel spectrogram in the training set into the detector. The detector compresses the Mel spectrogram in the training set along the time axis direction to detect whether there is a bird sound, obtains the detector results, and transmits them to the output layer;
[0080] S3.2: The output layer processes the detector results to obtain the initial detection result; if the initial detection result is 0, it is determined that there is no bird sound in the audio data corresponding to the current Mel spectrogram, and the initial detection is used as the final detection result; if the initial binary result is 1, it returns to the detector;
[0081] S3.3: The detector inputs the current Mel spectrogram into the classifier, restores it to the preprocessed audio data, and extracts local time-frequency features; then compresses and stacks the local time-frequency features along the frequency axis direction to extract the time-series global features in different directions, and outputs the classifier results; specifically:
[0082] S3.3.1: The detector inputs the current Mel spectrogram into the classifier and restores it to the corresponding preprocessed audio data;
[0083] S3.3.2: The preprocessed audio data is input into the first convolutional layer for feature extraction to obtain the first local time-frequency features;
[0084] S3.3.3: The second max-pooling layer compresses and screens the first local time-series features along the frequency axis direction to obtain the first compressed local time-frequency features;
[0085] S3.3.4: The second convolutional layer extracts features from the first compressed local time-frequency features to obtain the second local time-frequency features;
[0086] S3.3.5: The third max pooling layer compresses and filters the second local time-frequency feature along the frequency axis to obtain the second compressed local time-frequency feature;
[0087] S3.3.6: The third convolutional layer extracts features from the second compressed local time-frequency feature to obtain the third local time-frequency feature;
[0088] S3.3.7: The fourth max pooling layer compresses and filters the third local time-frequency feature along the frequency axis to obtain the third compressed local time-frequency feature;
[0089] S3.3.8: The stacking layer stacks the third compressed local time-frequency feature along the frequency axis to obtain the stacked local time-frequency feature;
[0090] S3.3.9: The first recurrent layer extracts features from the stacked local time-frequency feature along the positive direction of the time axis to obtain the forward sequential global feature;
[0091] S3.3.10: The second recurrent layer extracts features from the forward sequential global feature along the negative direction of the time axis to obtain the reverse sequential global feature;
[0092] S3.3.11: The fifth max pooling layer compresses the reverse sequential global feature along the time axis to obtain the classifier result;
[0093] S3.4: The output layer processes the classifier result to obtain the final detection result; if the final detection result is 0, it is determined that there is no bird sound in the audio data corresponding to the current mel spectrogram; if the final detection result is 1, it is determined that there is bird sound in the audio data corresponding to the current mel spectrogram;
[0094] S3.5: Repeat steps S3.1 - 3.4 until the training of the preset number of rounds is completed;
[0095] The first max pooling layer is a single-layer sequential max pooling layer that compresses the mel spectrogram along the time axis, keeping the frequency dimension of the mel spectrogram unchanged while the time dimension is 1, and sending it to the output layer as the detector result. If it is processed by the output layer to be 0, it is considered that there is no bird sound; if it is processed by the output layer to be 1, it returns to the detector link to continue the further calculation and determination of the classifier; before entering the classifier, the mel spectrogram will be restored to the preprocessed audio data; when entering the convolutional layer, the convolutional kernel will extract local time-frequency features, and after being extracted by multiple convolutional layers, it acts as a filter; the second max pooling layer, the third max pooling layer, and the fourth max pooling layer are frequency-axis max pooling layers, and the local time-frequency features will be compressed along the frequency axis; when entering the recurrent layer, the compressed local time-frequency features are further extracted along the positive and negative directions of the time axis to obtain sequential global features; finally, it enters the fifth max pooling layer, which compresses the time dimension to 1 along the time axis and sends it to the output layer as the classifier result;
[0096] Construct a binary cross-entropy loss function based on the final detection result of the Mel spectrogram of the audio data and the bird sound annotation information of the audio data; adjust the network parameters of the detector and classifier according to the binary cross-entropy loss value;
[0097] In this embodiment, the training for a preset number of rounds is 50 rounds.
[0098] S4: Set an early stopping mechanism, use the validation set to test the trained bird sound detection model, and obtain an optimized bird sound detection model; specifically:
[0099] Input the Mel spectrogram in the validation set into the trained bird sound detection model. According to the final detection result corresponding to the Mel spectrogram in the validation set output by the trained bird sound detection model, combine the bird sound annotation information corresponding to the Mel spectrogram in the validation set, and draw an ROC curve; use the area under the ROC curve value AUC as the validation metric. If the area under the ROC curve value AUC does not improve during the training of the preset number of rounds, stop the training, save the network parameters of the current detector and classifier, and obtain an optimized bird sound detection model; otherwise, adjust the network parameters of the detector and classifier, and repeat the training process in step S3.
[0100] S5: Obtain the audio data to be detected, convert it into a Mel spectrogram; use the optimized bird sound detection model to perform bird sound detection on the Mel spectrogram of the audio data to be detected, and obtain a bird sound detection result.
[0101] In the specific implementation process, an example is used to explain the change process after the preprocessed audio data enters the classifier; the tensor shape before entering the classifier is (500, 40), and after reduction, it enters the first convolutional layer (500, 40, 96) → the second max pooling layer (500, 8, 96) → the second convolutional layer (500, 8, 96) → the third max pooling layer (500, 2, 96) → the third convolutional layer (500, 2, 96) → the fourth max pooling layer (500, 1, 96), which is mainly composed of a convolutional layer with a rectified linear unit ReLU activation function and a non-overlapping pooling layer on the frequency axis; enter the stacking processing layer (500, 96) → the first recurrent layer (500, 96) → the second recurrent layer (500, 96); the main structure of the recurrent layer is a gated recurrent unit GRU. The purpose of using GRU twice is to avoid missing the temporal global information of the forward and reverse orders; finally, the fifth max pooling layer outputs a tensor (1, 96);
[0102] The method of this application is compared with traditional baseline CNN models and JDC-CNN. The baseline CNN models include max-pooling CNN and average-pooling CNN. The experimental data for comparison are the bird audio datasets Warblrb10k, Freefield1010, and BirdVox-DCASE-20k provided by the Detection and Classification of Acoustic Scenes and Events Challenge (DCASE). Each dataset provides audio files and annotation information on the presence or absence of bird sounds. Using AUC as the classification evaluation metric, the comparison results are shown in the following table:
[0103] Max-pooling CNN Average-pooling CNN JDC-CNN The method of this application Test results (%) 77.78 81.05 81.23 79.84
[0104] The comparison results of the time used by the method of this application are shown in the following table:
[0105] Max-pooling CNN Average-pooling CNN JDC-CNN The method of this application Training time / min 32.4 32.7 43.1 31.3 Average prediction time / s 2.7 2.7 2.9 2.5
[0106] To further detect the optimization of the method of this application, the number of parameters and the amount of computation of each model are detected. The comparison results are shown in the following table:
[0107] Max-pooling CNN Average-pooling CNN JDC-CNN The method of this application Computational complexity / MB 2932.8 2932.9 3736.0 2326.5 Number of parameters / piece 167,907 167,907 169,908 192,793
[0108] It can be seen from the above three tables that when detecting the presence or absence of bird sounds in 10s of audio, the method of this application has a classification performance that does not exceed 5% compared with other traditional methods, and the computing time is less than that of other traditional methods. The number of trainable parameters and the amount of computation correspond to the space complexity and time complexity of the algorithm. The larger the number of parameters, the larger the hard disk space occupied, the greater the space complexity, and the higher the requirement for the video memory or memory space of the hardware. The larger the amount of computation, the longer the running time of the algorithm network, the higher the time complexity, and the higher the requirement for the computing power of the hardware chip. The method of this application has a decrease in the amount of computation accompanied by an increase in the number of parameters, which means an increase in space complexity and a relatively large occupation of the hard disk space of the hardware. However, it can be solved by deep learning model compression, such as network structure optimization, model distillation, and pruning.
[0109] Example 3
[0110] This embodiment provides a bird sound detection system based on JDC-CRNN for implementing the bird sound detection method based on JDC-CRNN described in Embodiment 1 or 2, as Figure 3 shown, including:
[0111] A data acquisition and preprocessing module for acquiring audio data with bird sound annotation information, preprocessing the audio data, and obtaining preprocessed audio data;
[0112] A data conversion and division module, which is used to convert the preprocessed audio data into a Mel spectrogram, and divide the Mel spectrogram into a training set and a validation set according to a preset ratio;
[0113] A model training module, which uses the training set to train the constructed bird sound detection model based on JDC-CRNN for a preset number of rounds, sets a loss function, adjusts the network parameters of the bird sound detection model, and obtains a trained bird sound detection model;
[0114] A model testing module, which is used to set an early stopping mechanism, and uses the validation set to test the trained bird sound detection model to obtain an optimized bird sound detection model;
[0115] A bird sound detection module, which is used to obtain the audio data to be detected and convert it into a Mel spectrogram; uses the optimized bird sound detection model to perform bird sound detection on the Mel spectrogram of the audio data to be detected, and obtains a bird sound detection result
[0116] The same or similar reference numerals correspond to the same or similar components;
[0117] The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0118] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A bird sound detection method based on JDC-CRNN, characterized in that, Including: S1: Obtain audio data with bird sound annotation information, preprocess the audio data, and obtain the preprocessed audio data; S2: Convert the preprocessed audio data into a Mel spectrogram, and divide the Mel spectrogram into a training set and a validation set at a preset ratio; S3: Use the training set to train the constructed bird sound detection model based on JDC-CRNN for a preset number of rounds, set a loss function, adjust the network parameters of the bird sound detection model, and obtain the trained bird sound detection model; The constructed bird sound detection model based on JDC-CRNN includes a detector, a classifier, and an output layer connected in sequence; the output end of the detector is also connected to the input end of the output layer; The specific method for using the training set to train the constructed bird sound detection model based on JDC-CRNN for a preset number of rounds is: S3.1: Input the Mel spectrogram in the training set into the detector. The detector compresses the Mel spectrogram in the training set along the time axis direction, detects whether there is a bird sound, obtains the detector result, and transmits it to the output layer; S3.2: The output layer processes the detector result to obtain an initial detection result; If the initial detection result is 0, it is determined that there is no bird sound in the audio data corresponding to the current Mel spectrogram, and the initial detection is used as the final detection result; If the initial binary result is 1, return to the detector; S3.3: The detector inputs the current Mel spectrogram into the classifier, restores it to the preprocessed audio data, and extracts local time-frequency features; then compresses and stacks the local time-frequency features along the frequency axis direction, extracts sequential global features in different directions, and outputs the classifier result; S3.4: The output layer processes the classifier result to obtain the final detection result; If the final detection result is 0, it is determined that there is no bird sound in the audio data corresponding to the current Mel spectrogram; If the final detection result is 1, it is determined that there is a bird sound in the audio data corresponding to the current Mel spectrogram; S3.5: Repeat steps S3.1 - 3.4 until the preset number of rounds of training is completed; S4: Set an early stopping mechanism, use the validation set to test the trained bird sound detection model, and obtain the optimized bird sound detection model; S5: Obtain the audio data to be detected and convert it into a Mel spectrogram; Use the optimized bird sound detection model to detect the bird sound in the Mel spectrogram of the audio data to be detected, and obtain the bird sound detection result.
2. The bird sound detection method based on JDC-CRNN according to claim 1, wherein The preprocessing performed on the audio data includes a segmentation operation and a filtering operation.
3. The bird sound detection method based on JDC-CRNN according to claim 1, characterized in that The detector includes a first max pooling layer, and the output end of the first max pooling layer is used as the output end of the detector; The classifier includes a first convolutional layer, a second max pooling layer, a second convolutional layer, a third max pooling layer, a third convolutional layer, a fourth max pooling layer, a stacking processing layer, a first recurrent layer, a second recurrent layer, and a fifth max pooling layer connected in sequence; The output layer includes a forward propagation unit and an activation function connected in sequence; the input end of the forward propagation unit is used as the input end of the output layer.
4. The bird sound detection method based on JDC-CRNN according to claim 3, wherein The specific method of step S3.3 is: S3.3.1: The detector inputs the current Mel spectrogram into the classifier and restores it to the corresponding preprocessed audio data; S3.3.2: The preprocessed audio data is input into the first convolutional layer for feature extraction to obtain the first local time-frequency feature; S3.3.3: The second max pooling layer compresses and filters the first local time-series feature along the frequency axis direction to obtain the first compressed local time-frequency feature; S3.3.4: The second convolutional layer performs feature extraction on the first compressed local time-frequency feature to obtain the second local time-frequency feature; S3.3.5: The third max pooling layer compresses and filters the second local time-frequency feature along the frequency axis direction to obtain the second compressed local time-frequency feature; S3.3.6: The third convolutional layer performs feature extraction on the second compressed local time-frequency feature to obtain the third local time-frequency feature; S3.3.7: The fourth max pooling layer compresses and filters the third local time-frequency feature along the frequency axis direction to obtain the third compressed local time-frequency feature; S3.3.8: The stacking layer stacks the third compressed local time-frequency feature along the frequency axis direction to obtain the stacked local time-frequency feature; S3.3.9: The first recurrent layer performs feature extraction on the stacked local time-frequency feature along the positive direction of the time axis to obtain the forward-order time-series global feature; S3.3.10: The second recurrent layer performs feature extraction on the forward-order time-series global feature along the negative direction of the time axis to obtain the reverse-order time-series global feature; S3.3.11: The fifth max pooling layer compresses the reverse-order time-series global feature along the time axis direction to obtain the classifier result.
5. The bird sound detection method based on JDC-CRNN according to claim 1, characterized in that The specific method for setting the loss function and adjusting the network parameters of the bird sound detection model is: According to the final detection result of the Mel spectrogram of the audio data and the bird sound annotation information of the audio data, construct a binary cross-entropy loss function; adjust the network parameters of the detector and the classifier according to the binary cross-entropy loss value.
6. The bird sound detection method based on JDC-CRNN according to claim 5, wherein The specific method of step S4 is: Input the Mel spectrogram in the validation set into the trained bird sound detection model. According to the final detection result corresponding to the Mel spectrogram in the validation set output by the trained bird sound detection model, combined with the bird sound annotation information corresponding to the Mel spectrogram in the validation set, draw an ROC curve; use the area under the ROC curve value AUC as the validation index. If the area under the ROC curve value AUC does not improve in the preset number of rounds of training, stop training, save the network parameters of the current detector and classifier, and obtain the optimized bird sound detection model; otherwise, adjust the network parameters of the detector and the classifier, and repeat the training process of step S3.
7. The bird sound detection method based on JDC-CRNN according to claim 3, wherein The activation function in the output layer is the sigmoid activation function.
8. A bird sound detection system based on JDC-CRNN, which is used to implement the bird sound detection method based on JDC-CRNN according to any one of claims 1-7, characterized in that, It includes: A data acquisition and preprocessing module, which is used to acquire audio data with bird sound annotation information, preprocess the audio data, and obtain the preprocessed audio data; A data conversion and partitioning module, which is used to convert the preprocessed audio data into a Mel spectrogram, and partition the Mel spectrogram into a training set and a validation set at a preset ratio; A model training module, which uses the training set to train the constructed bird sound detection model based on JDC-CRNN for a preset number of rounds, sets the loss function, adjusts the network parameters of the bird sound detection model, and obtains the trained bird sound detection model; A model testing module, which is used to set an early stopping mechanism, test the trained bird sound detection model using a validation set, and obtain an optimized bird sound detection model; A bird sound detection module, which is used to obtain audio data to be detected and convert it into a Mel spectrogram; Use the optimized bird sound detection model to perform bird sound detection on the Mel spectrogram of the audio data to be detected, and obtain a bird sound detection result.
Citation Information
Patent Citations
Bird sound recognition method based on multi-feature fusion and combination model
CN113724712A
Method for generating acoustic model
US20210183375A1