A bird sound classification method and device based on single-step progressive representation transfer learning

By adopting a single-step progressive representation transfer learning method in bird sound classification, using self-supervised and supervised learning branches, combining data augmentation and weighted loss functions, the problems of strong dependence on labeled data and overfitting models in the existing technology are solved, and more efficient bird sound classification and generalization capabilities are achieved.

CN115294971BActive Publication Date: 2025-05-13GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210852000.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-05-13
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

When processing and utilizing bird monitoring data, especially bird call data, the prior art has problems such as strong dependence on labeled data and overfitting of models, making it difficult to effectively utilize labeled data.

Method used

The single-step progressive representation transfer learning method is adopted, and the bird sound signal self-supervised representation learning branch and supervised classification learning branch are constructed, and the data is enhanced by using the Mel spectrogram, and the model is updated through the weighted loss function, reducing the number of training times and improving training efficiency.

Benefits of technology

The classification and generalization ability of bird sound classification has been improved, the dependence on labeled data has been reduced, and the utilization rate of labeled bird sound data has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294971B_ABST
    Figure CN115294971B_ABST
Patent Text Reader

Abstract

The present invention discloses a bird sound classification method and device of single-step progressive representation transfer learning, the method comprising: constructing a bird sound data set; extracting a mel spectrogram of the bird sound data; performing different data enhancement processing on the mel spectrogram to obtain an enhanced spectrogram; constructing a self-supervised representation learning branch for bird sound signals, taking the enhanced spectrogram as the input of the self-supervised representation learning branch for bird sound signals, and calculating a first loss; constructing a supervised classification learning branch for bird sound signals, taking the enhanced spectrogram as the input of the supervised classification learning branch for bird sound signals, and calculating a second loss; weighting the first loss and the second loss to obtain a final loss; predicting the bird sound to be predicted through a supervised classification learning branch model to determine the type of the bird sound. The present invention can improve classification ability and generalization ability, reduce the number of training times, and improve training efficiency, and can be widely used in the field of machine learning technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a bird sound classification method and device based on single-step progressive representation transfer learning. Background Art

[0002] The number of bird species is a sensitive indicator for evaluating environmental quality, and they play an important role in ecological balance and biodiversity. In recent years, the Global Red List of Birds shows that a series of bird species are gradually decreasing. Monitoring birds and collecting bird data have become an important part of bird species protection. In the past decade, passive acoustic monitoring (PAM) has become the main method of bird monitoring. By deploying acoustic transducers in wildlife activity areas to record animal calls, monitoring data can be obtained in a non-invasive way for a long time. The data collected by PAM helps us better analyze bird activities in order to develop more effective protection measures for them. However, how to process and use this data is indeed a huge challenge.

[0003] With the rapid development of deep learning, deep neural networks have been successfully applied to bird species monitoring and classification. Under supervised classification learning, if there is enough labeled data, the neural network will perform well, otherwise it will cause the model to overfit. However, labeling bird calls is very labor-intensive and time-consuming, and can only be done by professional ornithologists or experienced bird lovers. With the emergence of self-supervised learning, supervised information can be mined from a large amount of unlabeled data, so that neural networks can first perform general representation learning on unlabeled data, and then be fine-tuned and applied to downstream classification tasks. Among them, contrastive learning is a very important part of self-supervised learning and is widely used in computer vision, natural language processing and other fields. The goal of contrastive learning is to use the contrastive loss function to map the representation distance of positive samples in the feature space together, and to pull the representation distance between positive and negative samples in the feature space apart. Among them, positive samples are enhanced copies or views from the same input, while negative samples are enhanced copies or views from different inputs. Migrating the self-supervised learning method to the field of bird sound signal recognition can improve people's utilization of unlabeled bird sound data, reduce dependence on labeled data, and maintain feature generalization. Although self-supervised learning has been successfully applied to various downstream tasks, if the model is only fine-tuned using the labeled information provided by the downstream task, this may not be able to take advantage of the generality and applicability of self-supervised learning. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a bird sound classification method and device based on single-step progressive representation transfer learning to improve classification and generalization capabilities, reduce the number of training times, and improve training efficiency.

[0005] An aspect of an embodiment of the present invention provides a bird sound classification method using single-step progressive representation transfer learning, comprising:

[0006] Build a bird sound dataset;

[0007] Extracting a mel spectrogram of the bird sound data in the bird sound dataset;

[0008] Performing different data enhancement processing on the Mel-size spectrogram to obtain enhanced spectrograms of different enhanced versions;

[0009] Constructing a self-supervised characterization learning branch for bird sound signals, taking the enhanced spectrogram as an input of the self-supervised characterization learning branch for bird sound signals, and calculating a first loss of the self-supervised classification learning branch for bird sound signals;

[0010] Constructing a supervised classification learning branch for bird sound signals, taking the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculating a second loss of the supervised classification learning branch for bird sound signals;

[0011] The first loss and the second loss are weighted to obtain a final loss; the final loss is used to update the self-supervised representation learning branch model and the supervised classification learning branch model;

[0012] The bird sounds to be predicted are predicted by the supervised classification learning branch model to determine the type of the bird sounds.

[0013] Optionally, constructing a bird sound dataset comprises:

[0014] According to the preset number of bird sound categories, the total duration of bird sound data of each category is preset, and the bird sound data is collected; wherein the total duration of the data of each category is greater than or equal to 1200 seconds; the data of each category includes at least one audio file, the duration of each audio file is greater than or equal to 10 seconds, the duration of the bird sound in each audio file is greater than or equal to 50% of the total duration of the audio; the duration of the continuous non-bird sound segment in the audio file is less than 25% of the total duration of the audio file;

[0015] Unify the data format of the collected bird sound dataset;

[0016] A stratified sampling strategy is used to divide the bird sound dataset into a training set, a validation set and a test set;

[0017] All data are one-hot encoded, and the encoding length of the one-hot encoding is equal to the number of bird categories in the bird sound dataset.

[0018] Optionally, extracting the mel spectrogram of the bird sound data in the bird sound dataset includes:

[0019] Configure the data sampling points, frame length information, and frame shift information of the bird sound signal, and calculate the total number of frames in the recording;

[0020] Perform frame stacking processing on any two adjacent frames to obtain the target frame;

[0021] Performing windowing processing on the target frame to obtain a windowed frame;

[0022] Perform DFT operation on each windowed frame to determine the amplitude spectrum of each frame signal to analyze the frequency points of the spectrum;

[0023] According to the upper and lower limits of the actual frequency of bird sounds obtained by analysis, the upper and lower limits of the Mel frequency are determined according to the conversion relationship between the Mel frequency and the actual frequency;

[0024] Construct a Mel filter and set multiple bandpass filters in the frequency range of the bird sound;

[0025] The amplitude spectrum of the bird sound signal is passed through the Mel filter function to obtain the corresponding Mel spectrogram.

[0026] Optionally, performing different data enhancement processing on the Mel-level spectrogram to obtain different enhanced versions of enhanced spectrograms includes at least one of the following:

[0027] After slicing random noise data according to a preset signal-to-noise ratio range, the data is superimposed on the bird sound signal; the types of the noise data include white noise, pink noise and brown noise;

[0028] Alternatively, the sliced ​​data is divided into multiple pieces of data at equal distances on the time axis, and then the multiple pieces of data are spliced ​​in random order to obtain a new bird sound signal;

[0029] Alternatively, the amplitude values ​​of all sampling points of the bird sound signal are multiplied by a set amplitude gain factor to adjust the volume of the bird sound signal within a random amplitude range;

[0030] Alternatively, perform time-frequency masking on the bird sound signal.

[0031] Optionally, the step of constructing a self-supervised characterization learning branch for bird sound signals, taking the enhanced spectrogram as an input of the self-supervised characterization learning branch for bird sound signals, and calculating a first loss of a self-supervised classification learning branch for bird sound signals comprises:

[0032] Constructing a self-supervised representation learning branch model for bird sound signals; the self-supervised representation learning branch model for bird sound signals includes an encoder, a projection layer, and a prediction layer;

[0033] The enhanced spectrogram is used as the input of the self-supervised representation learning branch of the bird sound signal, and the first loss of the self-supervised classification learning branch of the bird sound signal is calculated by a cosine similarity loss function with a gradient stop feedback module.

[0034] Optionally, the step of constructing a supervised classification learning branch for bird sound signals, taking the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculating a second loss of the supervised classification learning branch for bird sound signals comprises:

[0035] Constructing a supervised classification learning branch model for bird sound signals; the supervised classification learning branch model for bird sound signals includes an encoder and a classification layer;

[0036] The enhanced spectrogram is used as the input of the supervised classification learning branch of the bird sound signal, and the second loss of the supervised classification learning branch of the bird sound signal is calculated by the cross entropy loss function.

[0037] Optionally, the method further includes: updating and optimizing the model of each branch by configuring an optimizer and a learning rate; this step specifically includes:

[0038] Set the number of epochs and the number of batchsizes;

[0039] Use the stochastic batch gradient descent optimizer and configure the initial learning rate;

[0040] Cosine annealing learning rate strategy, select the learning rate strategy;

[0041] During the training process, the model training is completed when the loss value of the validation set is observed to decrease and stabilize.

[0042] Another aspect of the embodiment of the present invention further provides a bird sound classification device for single-step progressive representation transfer learning, comprising:

[0043] The first module is used to build a bird sound dataset;

[0044] The second module is used to extract the mel spectrogram of the bird sound data in the bird sound dataset;

[0045] The third module is used to perform different data enhancement processing on the Mel spectrogram to obtain enhanced spectrograms of different enhanced versions;

[0046] The fourth module is used to construct a self-supervised characterization learning branch for bird sound signals, use the enhanced spectrogram as an input of the self-supervised characterization learning branch for bird sound signals, and calculate a first loss of the self-supervised classification learning branch for bird sound signals;

[0047] A fifth module is used to construct a supervised classification learning branch for bird sound signals, use the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculate a second loss of the supervised classification learning branch for bird sound signals;

[0048] A sixth module is used to weight the first loss and the second loss to obtain a final loss; the final loss is used to update the self-supervised representation learning branch model and the supervised classification learning branch model;

[0049] The seventh module is used to predict the bird sounds to be predicted through the supervised classification learning branch model to determine the type of bird sounds.

[0050] Another aspect of an embodiment of the present invention further provides an electronic device, including a processor and a memory;

[0051] The memory is used to store programs;

[0052] The processor executes the program to implement the method described above.

[0053] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the method described above.

[0054] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.

[0055] The embodiment of the present invention constructs a bird sound data set; extracts the Mel spectrogram of the bird sound data in the bird sound data set; performs different data enhancement processing on the Mel spectrogram to obtain enhanced spectrograms of different enhanced versions; constructs a self-supervised characterization learning branch for bird sound signals, uses the enhanced spectrogram as the input of the self-supervised characterization learning branch for bird sound signals, and calculates the first loss of the self-supervised classification learning branch for bird sound signals; constructs a supervised classification learning branch for bird sound signals, uses the enhanced spectrogram as the input of the supervised classification learning branch for bird sound signals, and calculates the second loss of the supervised classification learning branch for bird sound signals; weights the first loss and the second loss to obtain the final loss; the final loss is used to update the self-supervised characterization learning branch model and the supervised classification learning branch model; predicts the bird sounds to be predicted through the supervised classification learning branch model to determine the type of bird sounds. The present invention can improve classification and generalization capabilities, reduce the number of training times, and improve training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0057] Figure 1 An overall step flow chart provided for an embodiment of the present invention;

[0058] Figure 2 A schematic diagram of test results provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0060] In view of the problems existing in the prior art, an embodiment of the present invention provides a bird sound classification method based on single-step progressive representation transfer learning, comprising:

[0061] Build a bird sound dataset;

[0062] Extracting a mel spectrogram of the bird sound data in the bird sound dataset;

[0063] Performing different data enhancement processing on the Mel-size spectrogram to obtain enhanced spectrograms of different enhanced versions;

[0064] Constructing a self-supervised characterization learning branch for bird sound signals, taking the enhanced spectrogram as an input of the self-supervised characterization learning branch for bird sound signals, and calculating a first loss of the self-supervised classification learning branch for bird sound signals;

[0065] Constructing a supervised classification learning branch for bird sound signals, taking the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculating a second loss of the supervised classification learning branch for bird sound signals;

[0066] The first loss and the second loss are weighted to obtain a final loss; the final loss is used to update the self-supervised representation learning branch model and the supervised classification learning branch model;

[0067] The bird sounds to be predicted are predicted by the supervised classification learning branch model to determine the type of the bird sounds.

[0068] Optionally, constructing a bird sound dataset comprises:

[0069] According to the preset number of bird sound categories, the total duration of bird sound data of each category is preset, and the bird sound data is collected; wherein the total duration of the data of each category is greater than or equal to 1200 seconds; the data of each category includes at least one audio file, the duration of each audio file is greater than or equal to 10 seconds, the duration of the bird sound in each audio file is greater than or equal to 50% of the total duration of the audio; the duration of the continuous non-bird sound segment in the audio file is less than 25% of the total duration of the audio file;

[0070] Unify the data format of the collected bird sound dataset;

[0071] A stratified sampling strategy is used to divide the bird sound dataset into a training set, a validation set and a test set;

[0072] All data are one-hot encoded, and the encoding length of the one-hot encoding is equal to the number of bird categories in the bird sound dataset.

[0073] Optionally, extracting the mel spectrogram of the bird sound data in the bird sound dataset includes:

[0074] Configure the data sampling points, frame length information, and frame shift information of the bird sound signal, and calculate the total number of frames in the recording;

[0075] Perform frame stacking processing on any two adjacent frames to obtain the target frame;

[0076] Performing windowing processing on the target frame to obtain a windowed frame;

[0077] Perform DFT operation on each windowed frame to determine the amplitude spectrum of each frame signal to analyze the frequency points of the spectrum;

[0078] According to the upper and lower limits of the actual frequency of bird sounds obtained by analysis, the upper and lower limits of the Mel frequency are determined according to the conversion relationship between the Mel frequency and the actual frequency;

[0079] Construct a Mel filter and set multiple bandpass filters in the frequency range of the bird sound;

[0080] The amplitude spectrum of the bird sound signal is passed through the Mel filter function to obtain the corresponding Mel spectrogram.

[0081] Optionally, performing different data enhancement processing on the Mel-level spectrogram to obtain different enhanced versions of enhanced spectrograms includes at least one of the following:

[0082] After slicing random noise data according to a preset signal-to-noise ratio range, the data is superimposed on the bird sound signal; the types of the noise data include white noise, pink noise and brown noise;

[0083] Alternatively, the sliced ​​data is divided into multiple pieces of data at equal distances on the time axis, and then the multiple pieces of data are spliced ​​in random order to obtain a new bird sound signal;

[0084] Alternatively, the amplitude values ​​of all sampling points of the bird sound signal are multiplied by a set amplitude gain factor to adjust the volume of the bird sound signal within a random amplitude range;

[0085] Alternatively, perform time-frequency masking on the bird sound signal.

[0086] Optionally, the step of constructing a self-supervised characterization learning branch for bird sound signals, taking the enhanced spectrogram as an input of the self-supervised characterization learning branch for bird sound signals, and calculating a first loss of a self-supervised classification learning branch for bird sound signals comprises:

[0087] Constructing a self-supervised representation learning branch model for bird sound signals; the self-supervised representation learning branch model for bird sound signals includes an encoder, a projection layer, and a prediction layer;

[0088] The enhanced spectrogram is used as the input of the self-supervised representation learning branch of the bird sound signal, and the first loss of the self-supervised classification learning branch of the bird sound signal is calculated by a cosine similarity loss function with a gradient stop feedback module.

[0089] Optionally, the step of constructing a supervised classification learning branch for bird sound signals, taking the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculating a second loss of the supervised classification learning branch for bird sound signals comprises:

[0090] Constructing a supervised classification learning branch model for bird sound signals; the supervised classification learning branch model for bird sound signals includes an encoder and a classification layer;

[0091] The enhanced spectrogram is used as the input of the supervised classification learning branch of the bird sound signal, and the second loss of the supervised classification learning branch of the bird sound signal is calculated by the cross entropy loss function.

[0092] Optionally, the method further includes: updating and optimizing the model of each branch by configuring an optimizer and a learning rate; this step specifically includes:

[0093] Set the number of epochs and the number of batchsizes;

[0094] Use the stochastic batch gradient descent optimizer and configure the initial learning rate;

[0095] Cosine annealing learning rate strategy, select the learning rate strategy;

[0096] During the training process, the model training is completed when the loss value of the validation set is observed to decrease and stabilize.

[0097] Another aspect of the embodiment of the present invention further provides a bird sound classification device for single-step progressive representation transfer learning, comprising:

[0098] The first module is used to build a bird sound dataset;

[0099] The second module is used to extract the mel spectrogram of the bird sound data in the bird sound dataset;

[0100] The third module is used to perform different data enhancement processing on the Mel spectrogram to obtain enhanced spectrograms of different enhanced versions;

[0101] The fourth module is used to construct a self-supervised characterization learning branch for bird sound signals, use the enhanced spectrogram as an input of the self-supervised characterization learning branch for bird sound signals, and calculate a first loss of the self-supervised classification learning branch for bird sound signals;

[0102] A fifth module is used to construct a supervised classification learning branch for bird sound signals, use the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculate a second loss of the supervised classification learning branch for bird sound signals;

[0103] A sixth module is used to weight the first loss and the second loss to obtain a final loss; the final loss is used to update the self-supervised representation learning branch model and the supervised classification learning branch model;

[0104] The seventh module is used to predict the bird sounds to be predicted through the supervised classification learning branch model to determine the type of bird sounds.

[0105] Another aspect of an embodiment of the present invention further provides an electronic device, including a processor and a memory;

[0106] The memory is used to store programs;

[0107] The processor executes the program to implement the method described above.

[0108] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the method described above.

[0109] The embodiment of the present invention also discloses a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device can read the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above method.

[0110] The specific implementation process of the present invention is described in detail below in conjunction with the accompanying drawings:

[0111] like Figure 1 As shown, the overall steps of the present invention include:

[0112] S1. Build a bird sound dataset, unify the format of the dataset, and divide it into training set and validation set;

[0113] S2, extracting the Mel spectrogram of bird sound data;

[0114] S3, performing two different data enhancements on the spectrogram to obtain two different enhanced versions of the enhanced spectrogram;

[0115] S4, constructing an encoder, a projection layer and a prediction layer, and constructing a self-supervised representation learning branch for bird sound signals; using the enhanced spectrogram of step S3 as the input of the self-supervised representation learning branch for bird sound signals, and calculating the loss of the self-supervised classification learning branch for bird sound signals;

[0116] S5. Construct an encoder and a classification layer, and construct a supervised classification learning branch for bird sound signals. Use the enhanced spectrogram of step S3 as the input of the supervised classification learning branch for bird sound signals, and calculate the loss of the supervised classification learning branch for bird sound signals;

[0117] S6. Construct a cosine annealing strategy to weight the losses obtained in steps S5 and S6 to obtain a final loss for updating the self-supervised representation learning branch model and the supervised classification learning branch model;

[0118] S7. Set the optimizer, learning rate, and learning rate strategy to update each branch model. After training, use the supervised classification learning branch model to predict bird sounds;

[0119] Among them, the specific details of S1 are:

[0120] 1) Requirements for constructing the bird sound dataset: The number of bird sound categories is N S , the total duration of bird sound data of each category K si , i=1,2,…N S , requiring the total duration K of each category of data in the bird sound dataset si The duration of the audio files in each category shall not be less than 10 seconds. The duration of the bird sounds in each audio file shall not be less than 50% of the total duration of the audio file. The continuous non-bird sound segments in the audio file shall not be greater than 25% of the total duration of the entire audio file. The requirements for data set construction are shown in Table 1.

[0121] 2) The above data sets are processed in a unified format: audio format: wav, sampling frequency: 32000Hz, number of audio channels: single channel.

[0122] Table 1 Dataset construction requirements

[0123]

[0124] 3) A stratified sampling strategy is used to divide the data set into training set, validation set, and test set with a ratio of 7:3.

[0125] 4) Perform one-hot encoding on all data labels, and the encoding length is equal to the number of bird categories Ns in the bird sound dataset.

[0126] The specific details of step S2 are:

[0127] 1) Record the number of data sampling points N of a bird sound signal, set the frame length wlen to 1024, the frame shift inc to 512, and get the total number of recording frames nf. Each frame after the frame is x in (n,λ), where n is the sampling point number and λ is the frame number.

[0128]

[0129] 2) Set the current frame x in (n,λ) and the previous frame x in (n,λ-1) is overlapped with a length of 512 to obtain x on (n,λ), the total length of the frame after stacking is still wlen, n=0,1,…,wlen-1,.

[0130]

[0131] 3) For x on (n,λ) is windowed, and the window type is Hamming window w(n,α), where α is 0.46 and the window length is equal to the frame length wlen=1024. Thus, all windowed frames x are obtained. w (λ,n).

[0132]

[0133] x w (n,λ)=x on (n,λ)*w(n,α)0≤n≤wlen-1 (4)

[0134] 4) For each frame x w (n,λ) performs a DFT operation of N = 1024 points, and takes the modulus of the DFT operation result to obtain the amplitude spectrum X(λ,k) of each frame signal, where k represents the frequency point and λ is the frame number. Due to the symmetry of Fourier transform, only the first N frames of the spectrum are processed. f frequency points for analysis, where N f =N / 2+1.

[0135]

[0136] 5) According to the actual upper and lower limits of the frequency of most bird sounds, l With f h (Unit: Hz), here we set f l =300 Hz, f h =14000hz, according to the conversion relationship between Mel frequency and actual frequency:

[0137]

[0138] The upper and lower limits of the Mel frequency (in mel) of the Mel filter bank to be determined are F mel (f l ) and F mel (f h ).

[0139] 6) Then construct a Mel filter bank in the frequency range of bird sounds [F mel (f l ),F mel (f h )], 1≤m≤M, where k represents the frequency point, m is the index of the filter, and M is the total number of filters to be set, where M=64. The expression of the H(k,m) filter is shown in equation (9), and each filter has a triangular filtering characteristic, and its center frequency is f(m), as shown in equation (10).

[0140]

[0141]

[0142] Among them, f l is the lowest frequency in the filter frequency range (in Hz), f h is the highest frequency of the filter frequency range, N is the length of FFT, f s is the sampling frequency, F mel The inverse function of (f).

[0143] 7) The amplitude spectrum X(k,λ) of the bird sound signal is passed through the Mel filter function H(k,m) to obtain the Mel sound spectrogram X of the bird sound signal H (m,λ)See formula (11).

[0144]

[0145] The specific details of step S3 are:

[0146] 1) Add noise data to the bird sound signal. The noise types are white noise, pink noise, and brown noise. Each added noise is a random one of the above. According to the set signal-to-noise ratio range (min_dB, max_dB), the noise slice data is superimposed on the bird sound signal. The signal-to-noise ratio range threshold needs to be set in advance, for example, min_dB = 3, max_dB = 15.

[0147] 2) Perform time interval shift transformation on the bird sound signal, that is, divide the slice data into n equal parts (n is less than or equal to 3) at equal distances on the time axis, and splice the n equal parts of data in random order to form a new bird sound signal.

[0148] 3) Change the volume of the bird sound signal. Multiply the amplitude value of all sampling points of the bird sound signal by the set amplitude gain factor a=10 (b / 20) The volume is adjusted within a random amplitude range, where b = (min_dB, max_dB), and the maximum and minimum decibel thresholds need to be set in advance, for example, min_dB = -12, max_dB = 12.

[0149] 4) Perform time-frequency masking on the bird sound signal. H (λ,m) for frequency masking. The frequency masking strategy is set according to the range of the sound frequency of different bird species. Set the masking frequency channel [m, m+f * ), where m∈[f l , f h -f * ], that is, f from the upper and lower limits of the Mel bandpass filter bank frequency f l ,f h Randomly select from * is the adjustable masking width, which is generally set to 20% of the species' vocalization bandwidth. Then set the number of masked frequency channels N F , usually set to 2. H (λ,m) performs time masking and sets the masking time channel [λ,λ+T * ), where λ∈[1, total number of frames - T * ], T * It is an adjustable time frame mask width, which is generally set to 20% of the total number of slice frames (rounded up). Then set the number of masked time channels N T , usually set to 2.

[0150] 5) The enhancement methods 1, 2, 3, and 4 in the above step S3 are randomly combined with probability p (for example, p=0.5) as a data enhancement method for bird sound signals.

[0151] 6) The bird sound signal that has completed the time domain data enhancement method is named as enhanced spectrogram.

[0152] The specific details of step S4 are:

[0153] 1) Construct a self-supervised representation learning branch model for bird sound signals. The model consists of three parts in order, namely encoder, projection layer, and prediction layer. The structures of the encoder, projection layer, and prediction layer are shown in Table 2 below:

[0154] Table 2 Self-supervised representation learning branch model

[0155]

[0156]

[0157] 2) The loss function of the self-supervised representation learning branch model is the cosine similarity loss function with a gradient stop feedback module, as follows:

[0158] p 1 =h(g(f(x 1 )))z 2 =g(f(x 1 )) (12)

[0159]

[0160]

[0161] Among them, x1 and x2 are enhanced spectrograms of two different enhanced versions of the bird sound signal x, f(·) represents the encoder, g(·) represents the projection layer, h(·) represents the prediction layer, D(p1,z2) represents the cosine similarity, stopgrad(·) indicates that the gradient on this branch will no longer be used for the subsequent back propagation process, p1, p2 are the output vectors of the image after the encoder, projection layer and prediction layer, and z1, z2 are the output vectors of the image after the encoder and projection layer.

[0162] The specific details of step S5 are:

[0163] 1) Construct a supervised classification learning branch model for bird sound signals. The model consists of two parts in order, namely the encoder and the classification layer. The structures of the encoder and the classification layer are shown in Table 3 below:

[0164] Table 3. Supervised classification learning branch model

[0165]

[0166]

[0167] 2) The loss function of the supervised classification learning model is the cross entropy loss function, as follows:

[0168] pred=c(f(x 1 )) (15)

[0169] L ce =CrossEntropyLoss(pred,ture) (16)

[0170] Where x1 is an enhanced spectrogram of an enhanced version of the bird sound signal x, f(·) represents the encoder, c(·) represents the classification layer, and CrossEntroLoss(·) is the cross entropy loss function. pred is the output vector after the encoder and classification layer, and true is the label vector after one-hot encoding.

[0171] The specific details of step S6 are:

[0172] The cosine annealing strategy is used to weight the loss of the self-supervised representation learning branch model and the loss of the supervised classification learning model to obtain the final loss function, as follows:

[0173] L final =λ(t)L ssl +(1-λ(t))L sl (17)

[0174] L ssl =L ssl +1 (18)

[0175] L sl =L ce (19)

[0176]

[0177] where t∈[1,epoch]].

[0178] The specific details of step S7 are:

[0179] 1) Set the number of epochs and batch sizes to 100 and 32 respectively.

[0180] 2) Optimizer selection. Use the random batch gradient descent optimizer (Batch_SGD) and set the initial learning rate lnitial_lr to 0.001

[0181] 3) Learning rate strategy selection. The cosine annealing learning rate strategy is adopted. The specific strategy is shown in the following formula:

[0182]

[0183] Among them, new_lr is the new learning rate obtained at the beginning of each epoch training, lnitial_lr is the initial learning rate, eta_min is the parameter eta_min represents the minimum learning rate, and T_max represents 1 / 4 of the cosine period. For example, lnitial_lr = 1e-3, eta_min = 1e-5, and T_max = epoch can be set.

[0184] During the training process, the model training is completed when the loss value of the validation set is observed to decrease and stabilize.

[0185] like Figure 2 As shown, the present invention achieved an accuracy of 97.9% on the large-scale bird song dataset Birdsdata released by Zhiyuan Research Institute and Bainiao Data, which has a significant improvement effect.

[0186] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.

[0187] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified, one or more of the functions and / or features described may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0188] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.

[0189] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0190] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0191] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0192] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0193] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0194] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A bird sound classification method based on single-step progressive representation transfer learning, characterized in that: include: Build a bird sound dataset; Extracting a mel spectrogram of the bird sound data in the bird sound dataset; The Mel-size spectrogram is subjected to different data enhancement processes to obtain different enhanced versions of enhanced spectrograms, including at least one of the following: After slicing random noise data according to a preset signal-to-noise ratio range, the data is superimposed on the bird sound signal; the types of the noise data include white noise, pink noise and brown noise; Alternatively, the sliced ​​data is divided into multiple pieces of data at equal distances on the time axis, and then the multiple pieces of data are spliced ​​in random order to obtain a new bird sound signal; Alternatively, the amplitude values ​​of all sampling points of the bird sound signal are multiplied by a set amplitude gain factor to adjust the volume of the bird sound signal within a random amplitude range; Alternatively, perform time-frequency masking on the bird sound signal; Constructing a self-supervised representation learning branch model for bird sound signals; the self-supervised representation learning branch model for bird sound signals comprises an encoder, a projection layer and a prediction layer, wherein the output dimensions of the encoder, the projection layer and the prediction layer are all 2048; The enhanced spectrogram is used as the input of the self-supervised representation learning branch of the bird sound signal, and the first loss of the self-supervised classification learning branch of the bird sound signal is calculated by a cosine similarity loss function with a gradient stop backpropagation module; Constructing a supervised classification learning branch for bird sound signals, taking the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculating a second loss of the supervised classification learning branch for bird sound signals; The first loss and the second loss are weighted to obtain a final loss; the final loss is used to update the self-supervised representation learning branch model and the supervised classification learning branch model; The bird sounds to be predicted are predicted by the supervised classification learning branch model to determine the type of the bird sounds.

2. The bird sound classification method of single-step progressive representation transfer learning according to claim 1 is characterized in that: The step of constructing a bird sound dataset comprises: According to the preset number of bird sound categories, the total duration of bird sound data of each category is preset, and the bird sound data is collected; wherein the total duration of the data of each category is greater than or equal to 1200 seconds; the data of each category includes at least one audio file, the duration of each audio file is greater than or equal to 10 seconds, the duration of the bird sound in each audio file is greater than or equal to 50% of the total duration of the audio; the duration of the continuous non-bird sound segment in the audio file is less than 25% of the total duration of the audio file; Unify the data format of the collected bird sound dataset; A stratified sampling strategy is used to divide the bird sound dataset into a training set, a validation set and a test set; All data are one-hot encoded, and the encoding length of the one-hot encoding is equal to the number of bird categories in the bird sound dataset.

3. The bird sound classification method of single-step progressive representation transfer learning according to claim 1 is characterized in that: The step of extracting the mel spectrogram of the bird sound data in the bird sound dataset comprises: Configure the data sampling points, frame length information, and frame shift information of the bird sound signal, and calculate the total number of frames in the recording; Perform frame stacking processing on any two adjacent frames to obtain the target frame; Performing windowing processing on the target frame to obtain a windowed frame; Perform DFT operation on each windowed frame to determine the amplitude spectrum of each frame signal to analyze the frequency points of the spectrum; According to the upper and lower limits of the actual frequency of bird sounds obtained by analysis, the upper and lower limits of the Mel frequency are determined according to the conversion relationship between the Mel frequency and the actual frequency; Construct a Mel filter and set multiple bandpass filters in the frequency range of the bird sound; The amplitude spectrum of the bird sound signal is passed through the Mel filter function to obtain the corresponding Mel spectrogram.

4. The bird sound classification method of single-step progressive representation transfer learning according to claim 1 is characterized in that: The step of constructing a supervised classification learning branch for bird sound signals, taking the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculating a second loss of the supervised classification learning branch for bird sound signals includes: Constructing a supervised classification learning branch model for bird sound signals; the supervised classification learning branch model for bird sound signals includes an encoder and a classification layer; The enhanced spectrogram is used as the input of the supervised classification learning branch of the bird sound signal, and the second loss of the supervised classification learning branch of the bird sound signal is calculated by the cross entropy loss function.

5. The bird sound classification method of single-step progressive representation transfer learning according to claim 1 is characterized in that: The method further includes: updating and optimizing the models of each branch by configuring an optimizer and a learning rate; this step specifically includes: Set the number of epochs and the number of batchsizes; Use the stochastic batch gradient descent optimizer and configure the initial learning rate; Cosine annealing learning rate strategy, select the learning rate strategy; During the training process, the model training is completed when the loss value of the validation set is observed to decrease and stabilize.

6. A bird sound classification device using single-step progressive representation transfer learning, characterized in that: include: The first module is used to build a bird sound dataset; The second module is used to extract the mel spectrogram of the bird sound data in the bird sound dataset; The third module is used to perform different data enhancement processing on the Mel-size spectrogram to obtain different enhanced versions of enhanced spectrograms, including at least one of the following: After slicing random noise data according to a preset signal-to-noise ratio range, the data is superimposed on the bird sound signal; the types of the noise data include white noise, pink noise and brown noise; Alternatively, the sliced ​​data is divided into multiple pieces of data at equal distances on the time axis, and then the multiple pieces of data are spliced ​​in random order to obtain a new bird sound signal; Alternatively, the amplitude values ​​of all sampling points of the bird sound signal are multiplied by a set amplitude gain factor to adjust the volume of the bird sound signal within a random amplitude range; Alternatively, perform time-frequency masking on the bird sound signal; The fourth module is used to construct a self-supervised representation learning branch model for bird sound signals; the self-supervised representation learning branch model for bird sound signals includes an encoder, a projection layer, and a prediction layer, wherein the output dimensions of the encoder, the projection layer, and the prediction layer are all 2048; the enhanced spectrogram is used as the input of the self-supervised representation learning branch for bird sound signals, and the first loss of the self-supervised classification learning branch for bird sound signals is calculated by a cosine similarity loss function with a gradient stop return module; A fifth module is used to construct a supervised classification learning branch for bird sound signals, use the enhanced spectrogram as an input of the supervised classification learning branch for bird sound signals, and calculate a second loss of the supervised classification learning branch for bird sound signals; A sixth module is used to weight the first loss and the second loss to obtain a final loss; the final loss is used to update the self-supervised representation learning branch model and the supervised classification learning branch model; The seventh module is used to predict the bird sounds to be predicted through the supervised classification learning branch model to determine the type of bird sounds.

7. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text classification method and device

    CN111522958A

  • Birdsound recognition model training method, recognition method and storage medium

    CN113936667A