Single-channel time-domain bird sound separation method, device, and computer-readable storage medium
By optimizing the structure of the bird sound separation network, including feature segmentation module and DPTTNet block, the existing time domain separation methods have solved the problem of large computing volume and high memory demand, and efficient bird sound separation has been achieved and training time has been shortened.
Patent Information
- Application Number
- CN202211354718.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-11-01
AI Technical Summary
When the existing time domain separation methods separate the characteristics of different sound sources, the computing volume and memory requirements are large, resulting in high requirements for running equipment and a longer time to train the separation model.
A single-channel time domain bird sound separation method is proposed. By constructing bird sound data sets and separation networks, including encoder, separator and decoder, the structure of the separator is optimized to reduce the computational amount and memory requirements using feature segmentation modules, DPTTNet blocks, dual-path blocks and overlapping addition modules.
The bird sound separation with small calculation volume, high efficiency and low cost is achieved, which reduces the problems of slow computing speed, large calculation volume and high memory requirements of the separation model, and shortens the training time of the separation model.
Smart Images

Figure CN115731924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a single-channel time-domain bird sound separation method, device, and computer-readable storage medium. Background Art
[0002] In recent years, acoustic monitoring has been widely applied in bird monitoring, research, and conservation. This method does not invade or damage the natural environment and can reduce the impact of human interference on birds. The audio files collected in acoustic monitoring can be used as important data for tracking the distribution of bird communities over time. Through deep learning technology, the recorded bird sound data can be used to automatically classify birds and quickly understand the species composition and quantity of birds in the current environment. Many studies have proposed various methods to improve the accuracy of sound-based bird classification. However, the field environment is very complex, and there are many factors affecting the accuracy of sound-based bird species classification, such as noise interference, too small recorded bird sounds, bird sound overlap, etc. When collecting bird sounds in the wild, bird sound overlap is a common problem because birds are social animals and usually chirp together. Bird sound overlap is one of the important factors affecting the accuracy of bird species classification. In the face of bird sound overlap, people have made a lot of efforts in identifying bird species, such as multi-label methods. When multiple bird sounds appear simultaneously in a piece of audio, the recognition model is trained to assign multiple labels to identify multiple bird species. Compared with the traditional single-label method, this method improves the recognition accuracy, but when bird sounds overlap in the time domain and frequency domain, the recognition accuracy is not high. In recent years, due to the rapid development of deep learning, significant progress has been made in source separation. However, there are few studies on bird sound separation. Whether the source separation method can be directly used for bird sound separation remains to be studied, but this provides a reference for the separation of bird sound overlap.
[0003] In related technologies, deep learning-based monaural source separation can be described by an "encoder-separator-decoder" framework. The encoder converts the input audio into high-dimensional features; the separator learns the masks of different sound sources and multiplies them with the input high-dimensional features to achieve the separation of different sound source features; the decoder converts the separated high-dimensional features into one-dimensional time-domain signals. This framework can be applied to source separation in both the frequency domain and the time domain. Most encoders convert the time-domain mixed audio into another feature representation through the short-time Fourier transform. In deep learning-based source separation, people prefer to learn the coefficients of the encoder in a data-driven manner, and one-dimensional convolution is usually used as the encoder.
[0004] In the current time-domain separation method, the encoder obtained through network learning converts the input mixed signal into high-dimensional features, and the separator then separates the high-dimensional features of each sound source from this mixed high-dimensional feature. The decoder restores the high-dimensional features of each sound source into the sound signals corresponding to the sound sources.
[0005] In the current time-domain separation method, the features encoded by the encoder have a high dimensionality and a large length, and need to pass through multiple transformer blocks in the separator, resulting in large requirements for the amount of computation, memory, etc. when the separator separates the features of different sound sources, high requirements for the operating device, and a long time for training the separation model. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a single-channel time-domain bird sound separation method, device, and computer-readable storage medium with small computation amount, high efficiency, and low cost.
[0007] One aspect of the embodiments of the present invention provides a single-channel time-domain bird sound separation method, including:
[0008] Construct a bird sound dataset, and perform data partitioning on the bird sound dataset to obtain a training set and a validation set; wherein, the bird sound dataset is single-channel bird sound data;
[0009] Construct a bird sound separation network, where the bird sound separation network includes an encoder, a separator, and a decoder;
[0010] Construct a separation model loss function, and configure an optimizer, a learning rate, and a learning rate strategy;
[0011] According to the separation model loss function, the optimizer, the learning rate, and the learning rate strategy, perform model training on the training set through the bird sound separation network to obtain a bird sound separation model;
[0012] According to the bird sound separation model, perform separation processing on the mixed bird sound data to obtain a bird sound separation result.
[0013] Optionally, the constructing a bird sound dataset, and performing data partitioning on the bird sound dataset to obtain a training set and a validation set includes:
[0014] Obtain bird sound data of different categories; wherein, the audio playback duration of the bird sound data of each category is not less than 1200 seconds; the existence duration of the bird sound in each audio file is not less than 50% of the total duration of the audio file; the continuous non-bird sound frequency band in the audio file is not greater than 25% of the total duration of the entire audio file;
[0015] Perform normalization processing on the obtained bird sound data to unify the audio format, sampling frequency, and number of audio channels of the bird sound data;
[0016] Adopt a stratified sampling strategy to partition the normalized bird sound data into a training set and a validation set;
[0017] Mix the bird sound data of the training set and the validation set to obtain a mixed training set and a mixed validation set.
[0018] Optionally, the mixing of the bird sound data of the training set and the validation set to obtain a mixed training set and a mixed validation set includes:
[0019] Configure the input bird sound length of the network to be 4 seconds and obtain the sampling points corresponding to the signal length.
[0020] Randomly select two different bird sound signals. If the number of sampling points of the selected bird sound signal is less than 64,000 and the audio time is less than 4 seconds, perform zero-padding on the bird sound signal to make it 64,000 points; if the number of sampling points of the selected bird sound signal is greater than 64,000 and the audio time is greater than 4 seconds, randomly select 64,000 of the sampling points.
[0021] After obtaining two bird sound signals of equal length, perform a mixing process on the two bird sound signals until the mixing process of any two bird sound signals of any two types in the training set and the validation set is completed.
[0022] Among them, the expression of the mixing process is:
[0023] s(t) = s 1 (t) + α · s 2 (t)
[0024] Among them, s(t) is the mixed bird sound signal, s 1 (t) and s 2 (t) are two different types of bird sound signals, and α is the gain coefficient of s 2 (t) during the mixing process.
[0025] Optionally, the method further includes the step of data augmentation for the bird sound data set, and this step specifically includes at least one of the following:
[0026] Overlay the noise slice data on the bird sound data set according to the set signal-to-noise ratio range and add noise data to the bird sound data set.
[0027] Or, divide the slice data into several equal parts at equal intervals on the time axis and splice the data of each equal part in a random order to complete the time interval displacement transformation of the bird sound data set.
[0028] Or, multiply the amplitude values of all sampling points of the bird sound signal in the bird sound data set by the set amplitude gain factor to perform volume adjustment within a random amplitude range on the bird sound signal and complete the volume transformation of the bird sound signal on the bird sound data set.
[0029] Optionally, the construction of the bird sound separation network includes:
[0030] Construct an encoder for the bird sound separation network; wherein, the encoder is composed of a one-dimensional convolutional layer and a ReLU activation function; the number N of convolutional kernels of the one-dimensional convolutional layer is set to 256, the size of the convolutional kernel is set to 16, and the convolutional stride is set to 8;
[0031] Construct a separator for the bird sound separation network; wherein, the separator is composed of four parts: a feature segmentation module, a DPTTNet block, a dual-path block, and an overlap-and-add module;
[0032] Construct a decoder for the bird sound separation network.
[0033] Optionally,
[0034] The separator for constructing the bird sound separation network includes:
[0035] Segment the features in the bird sound signal into several overlapping blocks, and splice all the segmented overlapping blocks into a three-dimensional tensor; wherein, there is a 50% overlap between two adjacent overlapping blocks;
[0036] Halve the feature length through a one-dimensional convolutional layer, then perform multi-head attention calculation, followed by processing through a normalization layer and a ReLU activation layer, and finally restore the feature length through an inverse one-dimensional convolutional layer;
[0037] Complete the sequence modeling process through local transformer processing and global transformer processing to obtain the target features;
[0038] Perform overlap-and-add processing on the target features to obtain masks for different sound source estimations, and complete the construction of the separator of the bird sound separation network;
[0039] The decoder for constructing the bird sound separation network includes:
[0040] Reconstruct the high-dimensional feature vector into a bird sound audio signal;
[0041] Use a transposed convolutional layer as the decoder. After obtaining the masks for different sound source estimations through the separator, then perform a dot product of the masks and the output of the encoder to obtain the estimated features of different sound sources, and then obtain the sound signal through the decoder.
[0042] Another aspect of the embodiments of the present invention further provides a single-channel time-domain bird sound separation device, including:
[0043] A first module for constructing a bird sound dataset and partitioning the bird sound dataset to obtain a training set and a validation set; wherein, the bird sound dataset is single-channel bird sound data;
[0044] The second module is used to construct a bird sound separation network, and the bird sound separation network includes an encoder, a separator, and a decoder;
[0045] The third module is used to construct a separation model loss function, and configure an optimizer, a learning rate, and a learning rate strategy;
[0046] The fourth module is used to perform model training on the training set through the bird sound separation network according to the separation model loss function, the optimizer, the learning rate, and the learning rate strategy, to obtain a bird sound separation model;
[0047] The fifth module is used to perform separation processing on the mixed bird sound data according to the bird sound separation model to obtain a bird sound separation result.
[0048] Another aspect of the embodiments of the present invention also provides an electronic device, including a processor and a memory;
[0049] The memory is used to store a program;
[0050] The processor executes the program to implement the method as described above.
[0051] Another aspect of the embodiments of the present invention also provides a computer-readable storage medium, and the storage medium stores a program, and the program is executed by a processor to implement the method as described above.
[0052] The embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the method as described above.
[0053] The embodiments of the present invention first construct a bird sound data set, and perform data partitioning on the bird sound data set to obtain a training set and a validation set; wherein, the bird sound data set is single-channel bird sound data; construct a bird sound separation network, and the bird sound separation network includes an encoder, a separator, and a decoder; construct a separation model loss function, and configure an optimizer, a learning rate, and a learning rate strategy; perform model training on the training set through the bird sound separation network according to the separation model loss function, the optimizer, the learning rate, and the learning rate strategy, to obtain a bird sound separation model; perform separation processing on the mixed bird sound data according to the bird sound separation model to obtain a bird sound separation result. The present invention has small computational complexity, high efficiency, and low cost. Description of the Drawings
[0054] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0055] Figure 1 It is the overall step flowchart provided by the embodiment of the present invention;
[0056] Figure 2 It is the schematic diagram of feature segmentation provided by the embodiment of the present invention;
[0057] Figure 3 It is the operation structure and process schematic diagram of the DPTTNet block provided by the embodiment of the present invention. Detailed implementation manners
[0058] In order to make the objectives, technical solutions and advantages of the present application more clear, the following further details the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0059] In view of the problems existing in the prior art, one aspect of the embodiment of the present invention provides a single-channel time-domain bird sound separation method, including:
[0060] Construct a bird sound dataset, and perform data division on the bird sound dataset to obtain a training set and a validation set; wherein, the bird sound dataset is single-channel bird sound data;
[0061] Construct a bird sound separation network, and the bird sound separation network includes an encoder, a separator and a decoder;
[0062] Construct a separation model loss function, and configure an optimizer, a learning rate and a learning rate strategy;
[0063] According to the separation model loss function, the optimizer, the learning rate and the learning rate strategy, perform model training on the training set through the bird sound separation network to obtain a bird sound separation model;
[0064] According to the bird sound separation model, perform separation processing on the mixed bird sound data to obtain a bird sound separation result.
[0065] Optionally, the constructing a bird sound dataset, and performing data division on the bird sound dataset to obtain a training set and a validation set includes:
[0066] Obtain bird sound data of different categories; among them, the audio playback duration of the bird sound data of each category is not less than 1200 seconds; the existence duration of the bird sound in each audio file is not less than 50% of the total duration of the audio file; the continuous non-bird sound frequency band in the audio file is not greater than 25% of the total duration of the entire audio file;
[0067] Perform normalization processing on the obtained bird sound data to unify the audio format, sampling frequency, and number of audio channels of the bird sound data;
[0068] Use a stratified sampling strategy to divide the normalized bird sound data into a training set and a validation set;
[0069] Mix the bird sound data of the training set and the validation set to obtain a mixed training set and a validation set.
[0070] Optionally, the mixing of the bird sound data of the training set and the validation set to obtain a mixed training set and a validation set includes:
[0071] Configure the input bird sound length of the network to be 4 seconds and obtain the sampling points corresponding to the signal length;
[0072] Randomly select two different bird sound signals. If the number of sampling points of the selected bird sound signal is less than 64000 and the audio time is less than 4 seconds, perform zero-padding on the bird sound signal to make it 64000 points; if the number of sampling points of the selected bird sound signal is greater than 64000 and the audio time is greater than 4 seconds, randomly select 64000 of the sampling points;
[0073] After obtaining two bird sound signals of equal length, perform mixing processing on the two bird sound signals until the mixing processing of any two bird sound signals of any two categories in the training set and the validation set is completed;
[0074] Among them, the expression of the mixing processing is:
[0075] s(t) = s 1 (t) + α·s 2 (t)
[0076] Among them, s(t) is the mixed bird sound signal, s 1 (t) and s 2 (t) are two different types of bird sound signals, and α is the gain coefficient of s 2 (t) during the mixing process.
[0077] Optionally, the method further includes the step of data augmentation for the bird sound data set, and this step specifically includes at least one of the following:
[0078] Overlay the noise slice data onto the bird sound dataset according to the set signal-to-noise ratio range, and add noise data to the bird sound dataset;
[0079] Alternatively, divide the slice data into several equal parts at equal intervals on the time axis, and splice the data of each equal part in a random order to complete the time interval displacement transformation of the bird sound dataset;
[0080] Alternatively, multiply the amplitude values of all sampling points of the bird sound signal in the bird sound dataset by a set amplitude gain factor to adjust the volume within a random amplitude range, and complete the volume transformation of the bird sound signal on the bird sound dataset.
[0081] Optionally, the construction of the bird sound separation network includes:
[0082] Construct an encoder for the bird sound separation network; wherein, the encoder consists of a one-dimensional convolutional layer and a ReLU activation function; the number of convolutional kernels N of the one-dimensional convolutional layer is set to 256, the convolutional kernel size is set to 16, and the convolutional stride is set to 8;
[0083] Construct a separator for the bird sound separation network; wherein, the separator consists of four parts: a feature segmentation module, a DPTTNet block, a dual-path block, and an overlap-and-add module;
[0084] Construct a decoder for the bird sound separation network.
[0085] Optionally,
[0086] The separator for constructing the bird sound separation network includes:
[0087] Segment the features in the bird sound signal into several overlapping blocks, and splice all the segmented overlapping blocks into a three-dimensional tensor; wherein, there is a 50% overlap between two adjacent overlapping blocks;
[0088] Halve the feature length through a one-dimensional convolutional layer, then perform multi-head attention calculation, followed by processing through a normalization layer and a ReLU activation layer, and finally restore the feature length through an inverse one-dimensional convolutional layer;
[0089] Complete the sequence modeling process through local transformer processing and global transformer processing to obtain the target features;
[0090] Perform overlap-and-add processing on the target features to obtain masks for different sound source estimations, and complete the construction of the separator of the bird sound separation network;
[0091] The decoder for constructing the bird sound separation network includes:
[0092] Reconstruct the high-dimensional feature vector into a bird sound audio signal;
[0093] Use a transposed convolutional layer as the decoder. After obtaining the masks for estimating different sound sources through the separator, multiply the masks and the output of the encoder point - by - point to obtain the estimated features of different sound sources, and then obtain the sound signal through the decoder.
[0094] Another aspect of the embodiments of the present invention also provides a single - channel time - domain bird sound separation device, including:
[0095] A first module for constructing a bird sound data set and partitioning the bird sound data set to obtain a training set and a validation set; wherein, the bird sound data set is single - channel bird sound data;
[0096] A second module for constructing a bird sound separation network, the bird sound separation network including an encoder, a separator and a decoder;
[0097] A third module for constructing a separation model loss function and configuring an optimizer, a learning rate and a learning rate strategy;
[0098] A fourth module for training a model on the training set through the bird sound separation network according to the separation model loss function, the optimizer, the learning rate and the learning rate strategy to obtain a bird sound separation model;
[0099] A fifth module for separating the mixed bird sound data according to the bird sound separation model to obtain a bird sound separation result.
[0100] Another aspect of the embodiments of the present invention also provides an electronic device, including a processor and a memory;
[0101] The memory is used for storing programs;
[0102] The processor executes the program to implement the method as described above.
[0103] Another aspect of the embodiments of the present invention also provides a computer - readable storage medium, the storage medium stores a program, and the program is executed by a processor to implement the method as described above.
[0104] The embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer - readable storage medium. The processor of the computer device can read the computer instructions from the computer - readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the method as described above.
[0105] The following combines the specification drawings to describe the specific implementation process of the present invention in detail:
[0106] Refer to Figure 1Overall method flowchart. A method for training and testing a bird sound recognition model based on deep learning in complex acoustic scenarios provided by an embodiment of the present invention includes the following steps:
[0107] S1. Construct a bird sound dataset, unify the format of the dataset, and divide it into a training set and a validation set;
[0108] S2. Perform data augmentation on the dataset;
[0109] S3. Construct a bird sound separation network, namely an encoder, a separator, and a decoder;
[0110] S4. Construct a separation model loss function;
[0111] S5. Set the optimizer, learning rate, and learning rate strategy. After training is completed, use the mixed bird sound data in the validation set for separation verification.
[0112] Specifically, the specific implementation steps of S1 include:
[0113] 1). Requirements for constructing the bird sound dataset: The number of bird sound categories is N S , and the total duration of the bird sound data for each category is K si , i = 1, 2,... N S , and it is required that the total duration K of each category of data in the bird sound dataset si is not less than 1200 seconds. The data of each category can include several audio files. The duration of the audio files in each category cannot be shorter than 10 seconds. The duration of the bird sound in each audio file shall not be less than 50% of the total duration of the audio file. The continuous non-bird sound segments in the audio file cannot be greater than 25% of the total duration of the entire audio file.
[0114] 2). Perform unified processing on the data format of the above dataset. Audio format: wav, sampling frequency: 16000 Hz, audio channel number: single channel. The requirements for unified data format processing are shown in Table 1 below:
[0115] Table 1 Requirements for dataset construction
[0116]
[0117] 3). Use the stratified sampling strategy to divide the dataset into a training set and a validation set, with a ratio of 7:3.
[0118] 4). Mix the bird sound data in the training set and the validation set respectively to obtain a mixed training set and a validation set. The mixing process is as follows:
[0119] ①. Set the input bird sound length of the network to 4 seconds. Since the sampling rate is 16,000 Hz, the required signal length is 4 * 16,000 = 64,000 sampling points. Then randomly select two different bird sound signals. If the number of sampling points of the selected bird sound signal is less than 64,000, that is, the audio time is less than 4 seconds, then zeros need to be padded to the signal to make the short signal padded to 64,000 points; if the number of sampling points of the selected bird sound signal is greater than 64,000, that is, the audio time is greater than 4 seconds, then randomly select 64,000 of these sampling points.
[0120] ②. After obtaining two bird sound signals of equal length, mix them according to Formula 1
[0121]
[0122] where s(t) is the mixed bird sound signal, s 1 (t) and s 2 (t) are two different types of bird sound signals, α is the gain coefficient of s 2 (t) during the mixing process, q represents the magnitude level of the two bird sound signals, measured in decibels, and the value of q ranges from -5 to +5. When synthesizing the mixed bird sound signal, q is a random number in the range of -5 to +5.
[0123] ③. Mix every two bird sound audios of every two types in the training set and the validation set respectively.
[0124] Specifically, the specific implementation steps of S2 include:
[0125] 1). Add noise data to the bird sound signal. The noise types are white noise, pink noise, and brown noise. Each time the added noise is a random one of the above, and the noise slice data is superimposed on the bird sound signal according to the set signal-to-noise ratio range (min_dB, max_dB). The signal-to-noise ratio range threshold needs to be set in advance. For example, min_dB = 3 and max_dB = 15.
[0126] 2). Perform time interval displacement transformation on the bird sound signal. That is, divide the slice data into n equal parts (n is less than or equal to 3) at equal distances on the time axis, and splice the n equal parts of data in a random order to form a new bird sound signal.
[0127] 3). Perform volume transformation on the bird sound signal. That is, multiply the amplitude value of all sampling points of the bird sound signal by the set amplitude gain factor a = 10 (b / 20) Perform volume adjustment within a random amplitude range, where b = (min_dB, max_dB), and the maximum and minimum decibel thresholds need to be set in advance. For example, min_dB = -12 and max_dB = 12.
[0128] 4) The enhancement methods in 1), 2), and 3) in step S2 above are randomly combined according to the probability p (for example, p = 0.5) as the data enhancement method for bird sound signals.
[0129] Save the mixed bird sound signals that have completed the time-domain data enhancement method as a new dataset for subsequent model training.
[0130] Specifically, for the above S3, the bird sound separation network structure includes three parts: an encoder, a separator, and a decoder, which need to be constructed in sequence. Step S3 includes the following steps 1)-3):
[0131] 1) Construct the encoder of the bird sound separation network. The encoder consists of a one-dimensional convolutional layer and a ReLU activation function. The number of convolutional kernels N of the one-dimensional convolutional layer is set to 256, the convolutional kernel size is set to 16, and the convolutional stride is set to 8. The formula for the encoder process can be expressed as:
[0132] w = ReLU(conv1d(x))
[0133] where w represents the features output by the encoder after encoding.
[0134] 2) Construct the separator of the bird sound separation network. The separator is the core part of the entire bird sound separation network, mainly consisting of four parts: feature segmentation, DPTTNet block, dual-path block processing, and overlap and add. Among them, the DPTTNet block plays an important role in reducing the computational amount and the number of parameters of the entire separation network.
[0135] The above step 2) specifically includes the following (1)-(4):
[0136] (1) Feature segmentation
[0137] As Figure 2 shown, this step divides the feature w into S = 136 overlapping blocks with a length of K = 120. There is a 50% overlap between adjacent blocks to maintain the correlation between different blocks, and then all the divided blocks are concatenated into a three-dimensional tensor D ∈ R N×K×S , which is convenient for overall modeling of the input features from two dimensions later. If the length of the input feature does not meet the segmentation condition, the input feature needs to be padded with 0.
[0138] (2) DPTTNet block
[0139] Refer to Figure 3, To address the issue of excessive computational complexity and high memory requirements when calculating multi-head attention for long input features, a one-dimensional convolutional layer is used in the DPTTNet block to reduce the feature length. The number of convolutional kernels in the one-dimensional convolutional layer is the same as the input feature dimension, the length of the convolutional kernel is 4, and the convolutional stride is 2. After passing through the one-dimensional convolutional layer, the feature length is halved, and then multi-head attention is calculated, followed by layer normalization and ReLU activation. To restore the feature length, an inverse one-dimensional convolutional layer is required. The parameter settings of the inverse one-dimensional convolutional layer are the same as those of the one-dimensional convolutional layer.
[0140] (3) Dual-path processing procedure
[0141] In the dual-path processing stage, each complete sequence modeling includes local transformer processing and global transformer processing, and a total of B complete sequence modelings need to be repeated. The operation process of local transformer processing is as follows:
[0142] D b intra = IntraTransformer b [D b-1 inter
[0143] = [DPTTNet block(D b tnter [:, :, i]), i = 1,..., S]
[0144] The features D output by the segmentation stage are first passed to local transformer processing, and local transformer processing acts on the second dimension of the features D. Next, it goes to global transformer processing. The features after local transformer processing are passed to global transformer processing and then act on the last dimension of the features. The operation process of global transformer processing is as follows:
[0145] D b inter = InterTransformer b [D b-1 intra
[0146] = [transformer(D b intra [:, j, :]), j = 1,..., K]
[0147] Among them, b = 0, 1,..., B - 1 represents the number of times of the dual-path processing procedure. When b = 0, D 0 intra Represents the input feature D after segmentation. It should also be noted that layer normalization for each local transformer processing and global transformer processing is applied to all dimensions.
[0148] A global modeling operation process in the dual-path processing can be summarized as follows:
[0149] D b+1 = f inter (ρ(f intra (D b )))
[0150] where f inter (·) and f intra (·) represent inter-transformer and intra-transformer respectively, and ρ represents swapping the last two dimensions of D b ∈ R N×K×S . After B complete dual-path processes, the output D B is obtained, and then followed by a one-dimensional convolutional layer to change the output channels to C*N, where C represents the number of sound sources. The process is as follows:
[0151] D output = ψ -1 (f output (ψ(D B )))
[0152] where D B ∈ R N×K×S , D output ∈ R C×N×K×S , f output (·) represents one-dimensional convolution. ψ represents reshaping the input feature because one-dimensional convolution cannot directly operate on D B ∈ R N×K×S . The last two dimensions need to be merged into one dimension, and then inverse-transformed after one-dimensional convolution.
[0153] (4) Overlap and add
[0154] The D output obtained after B dual-path processes needs to be converted to the same shape as before segmentation through the overlap and add method. The calculation process is the reverse operation of the previous feature segmentation. After passing through a one-dimensional convolutional layer and an activation function, the mask estimation of C sound sources is obtained.
[0155] m i = max(0, conv1d(OverlapAdd(D output )))
[0156] Among them, OverlapAdd represents the operation of performing overlap and add on D output For the operation of overlap and add, each small segment of features has an overlap of 50%.
[0157] 3), Construct a decoder
[0158] The decoder reconstructs the high-dimensional feature vector into a bird audio signal. We use a transposed convolutional layer as the decoder, and its convolutional kernel size and stride are the same as those of the encoder. After obtaining the masks for different sound source estimations through the separator, the estimated features of different sound sources are obtained by multiplying them pointwise with the output w of the encoder, and then the sound signal can be obtained through the decoder. The transformation of the decoder can be expressed as:
[0159]
[0160] Among them, ⊙ represents pointwise multiplication, represents the estimation of the original sound source.
[0161] Specifically, in the steps of S4 above, the scale-invariant signal-to-noise ratio (SI-SNR) is used as the loss function during the training process because it is usually used as an evaluation metric for source separation. In this embodiment, utterance-level permutation-invariant training (uPIT) is used to train the proposed model to maximize SI-SNR. The calculation process of SI-SNR is as follows:
[0162]
[0163]
[0164]
[0165] Among them, represents the estimated audio signal, and s represents the labeled audio signal.
[0166] Specifically, S5 above specifically includes the following steps:
[0167] 1), Set the number of epochs and the number of batch sizes to 100 and 8 respectively.
[0168] 2), Select an optimizer. The Adam optimizer is used as the stochastic batch gradient descent optimizer, and the initial learning rate is set to Initial_Ir = 0.00015.
[0169] 3), Select a learning rate strategy. The cosine annealing learning rate strategy is adopted, and the strategy is specifically shown in the following formula:
[0170]
[0171] Among them, new_Ir is the new learning rate obtained at the start of training for each epoch, Initial_Ir is the initial learning rate, eta_min is the parameter eta_min representing the minimum learning rate, and T_max represents 1 / 4 of the period of cos. For example, Initial_Ir = 1e-3, eta_min = 1e-5, and T_max = epoch can be set.
[0172] 4) During the training process, when it is observed that the validation set loss value has not decreased after 10 epochs, the model training is completed.
[0173] Next, this embodiment will illustrate the comparison between the single-channel time-domain bird sound separation method provided by the present invention and related methods:
[0174] This embodiment has achieved similar performance to other sound source separation models in a mixed bird sound dataset of 20 different bird species, but the amount of computation, the number of parameters, and the memory requirement have been greatly reduced.
[0175] The separation performance of different separation models is shown in Table 2, and the floating-point operations, running time, and memory occupancy of different separation models are shown in Table 3:
[0176] Table 2
[0177] Network SI-SNRi (dB) SDRi (dB) Params (M) FLOPs (G) DPRNN 19.3 20 2.6 60.016 DPTNet 21.5 22.1 2.639 42.961 Ours 19.3 20.1 0.444 5.893
[0178] Table 3
[0179] Network FLOPs (G) CPU Time (s) GPU Time (ms) F / B GPU Memory (GB) DPRNN 60.016 2.047 54.994 0.890 / 1390 DPTNet 42.961 2.256 52.001 0.844 / 1.996 Ours 5.893 0.316 13.507 0.602 / 1.002
[0180] Among them, both SI-SNRi and SDRi are indicators for measuring the bird sound separation performance, and the larger the better. Params represents the size of the model parameters, FLOPs represents the floating-point operations, CPU Time represents the time spent by the model to separate a 4-second mixed bird sound when running on the CPU, GPU Time represents the time spent by the model to separate a 4-second mixed bird sound when running on the GPU, and F / B GPU Memory represents the size of the video memory consumed by the model during forward inference and backpropagation when running on the GPU.
[0181] In summary, the present invention can reduce the amount of computation and memory requirement of the time-domain sound source separation model, accelerate the separation speed, enable the separation model to be better applied to various devices with low performance, shorten the training time of the separation model, and solve the problems of slow operation speed, large amount of computation, and high memory requirement of the time-domain separation model.
[0182] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks may sometimes be executed in reverse order. Further, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.
[0183] In addition, although the present invention has been described in the context of functional modules, it should be understood that unless otherwise stated to the contrary, one or more of the described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are illustrative only and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0184] If the described functions are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art or a part of this technical solution, may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0185] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
[0186] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0187] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well-known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0188] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0189] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0190] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A single-channel time-domain bird sound separation method, characterized in that, it includes: Construct a bird sound dataset, and perform data partitioning on the bird sound dataset to obtain a training set and a validation set; wherein, the bird sound dataset is single-channel bird sound data; Construct a bird sound separation network, and the bird sound separation network includes an encoder, a separator, and a decoder; Construct a separation model loss function, and configure an optimizer, a learning rate, and a learning rate strategy; According to the separation model loss function, the optimizer, the learning rate, and the learning rate strategy, perform model training on the training set through the bird sound separation network to obtain a bird sound separation model; According to the bird sound separation model, perform separation processing on the mixed bird sound data to obtain a bird sound separation result; The construction of the bird sound separation network includes: Construct an encoder of the bird sound separation network; wherein, the encoder is composed of a one-dimensional convolutional layer and a ReLU activation function; the number of convolutional kernels N of the one-dimensional convolutional layer is set to 256, the convolutional kernel size is set to 16, and the convolutional stride is set to 8; Construct a separator of the bird sound separation network; wherein, the separator is composed of four parts: a feature segmentation module, a DPTTNet block, a dual-path block, and an overlap-and-add module; Construct a decoder of the bird sound separation network; The construction of the separator of the bird sound separation network includes: Segment the features in the bird sound signal into several overlapping blocks, and splice all the segmented overlapping blocks into a three-dimensional tensor; wherein, there is a 50% overlap between two adjacent overlapping blocks; Halve the feature length through a one-dimensional convolutional layer, then perform multi-head attention calculation, followed by normalization layer and ReLU activation layer processing, and finally restore the feature length through an inverse one-dimensional convolutional layer; Complete the sequence modeling process through local transformer processing and global transformer processing to obtain target features; Perform overlap-and-add processing on the target features to obtain masks for different sound source estimations, and complete the construction of the separator of the bird sound separation network; The construction of the decoder of the bird sound separation network includes: Reconstruct the high-dimensional feature vector into a bird sound audio signal; Use a transposed convolutional layer as the decoder. After obtaining the masks for different sound source estimations through the separator, multiply the masks and the output of the encoder point by point to obtain the estimated features of different sound sources, and then obtain the sound signal through the decoder.
2. The single-channel time-domain bird sound separation method according to claim 1, characterized in that, The construction of the bird sound dataset and the data partitioning of the bird sound dataset to obtain a training set and a validation set include: Obtain bird sound data of different categories; wherein, the audio playback duration of the bird sound data of each category is not less than 1200 seconds; the existence duration of the bird sound in each audio file is not less than 50% of the total duration of the audio file; the continuous non-bird sound frequency band in the audio file is not greater than 25% of the total duration of the entire audio file; Perform normalization processing on the obtained bird sound data to unify the audio format, sampling frequency, and audio channel number of the bird sound data; Use a stratified sampling strategy to divide the normalized bird sound data into a training set and a validation set; Mix the bird sound data of the training set and the validation set to obtain a mixed training set and a mixed validation set.
3. A single-channel time-domain bird sound separation method according to claim 2, wherein, the mixing of the bird sound data of the training set and the validation set to obtain a mixed training set and a mixed validation set includes: Configure the input bird sound length of the network to be 4 seconds and obtain the sampling points corresponding to the signal length; Randomly select two different bird sound signals. If the number of sampling points of the selected bird sound signal is less than 64,000 and the audio time is less than 4 seconds, perform zero-padding on the bird sound signal to make the bird sound signal have 64,000 points; if the number of sampling points of the selected bird sound signal is greater than 64,000 and the audio time is greater than 4 seconds, randomly select 64,000 of the sampling points; After obtaining two bird sound signals of equal length, perform a mixing process on the two bird sound signals until the mixing process of any two bird sound signals of any two types in the training set and the validation set is completed; wherein, the expression of the mixing process is: Among them, is the mixed bird sound signal, and are two different types of bird sound signals, is during the mixing process of the gain coefficient.
4. A single-channel time-domain bird sound separation method according to claim 1, wherein, the method further includes the step of data augmentation for the bird sound data set, and this step specifically includes at least one of the following: Overlay the noise slice data on the bird sound data set according to the set signal-to-noise ratio range, and add noise data to the bird sound data set; Or, divide the slice data into several equal parts at equal intervals on the time axis, splice the data of each equal part in a random order, and complete the time interval displacement transformation of the bird sound data set; Or, multiply the amplitude values of all sampling points of the bird sound signal in the bird sound data set by a set amplitude gain factor, perform volume adjustment within a random amplitude range on the bird sound signal, and complete the volume transformation of the bird sound signal on the bird sound data set.
5. A single-channel time-domain bird sound separation device, wherein, it includes: The first module is used to construct a bird sound data set and perform data division on the bird sound data set to obtain a training set and a validation set; wherein, the bird sound data set is single-channel bird sound data; The second module is used to construct a bird sound separation network, and the bird sound separation network includes an encoder, a separator, and a decoder; The third module is used to construct a separation model loss function and configure an optimizer, a learning rate, and a learning rate strategy; The fourth module is used to train the model on the training set through the bird sound separation network according to the separation model loss function, the optimizer, the learning rate, and the learning rate strategy to obtain a bird sound separation model; The fifth module is used to perform separation processing on the mixed bird sound data according to the bird sound separation model to obtain a bird sound separation result; The construction of the bird sound separation network includes: Construct the encoder of the bird sound separation network; wherein, the encoder is composed of a one-dimensional convolutional layer and a ReLU activation function; the number N of convolutional kernels of the one-dimensional convolutional layer is set to 256, the convolutional kernel size is set to 16, and the convolutional stride is set to 8; Construct the separator of the bird sound separation network; wherein, the separator is composed of four parts: a feature segmentation module, a DPTTNet block, a dual-path block, and an overlap-and-add module; Construct a decoder for a bird sound separation network; The separator for constructing the bird sound separation network includes: Divide the features in the bird sound signal into several overlapping blocks, and splice all the divided overlapping blocks into a three-dimensional tensor; among them, there is a 50% overlap between two adjacent overlapping blocks; Halve the feature length through a one-dimensional convolutional layer, then perform multi-head attention calculation, followed by processing of a normalization layer and a ReLU activation layer, and finally restore the feature length through an inverse one-dimensional convolutional layer; Complete the sequence modeling process through local transformer processing and global transformer processing to obtain the target features; Perform an overlap-and-add process on the target features to obtain masks for different sound source estimations, and complete the construction of the separator for the bird sound separation network; The decoder for constructing the bird sound separation network includes: Reconstruct the high-dimensional feature vector into a bird sound audio signal; Use a transposed convolutional layer as the decoder. After obtaining the masks for different sound source estimations through the separator, then perform a dot product of the masks and the output of the encoder to obtain the estimated features of different sound sources, and then obtain the sound signal through the decoder.
6. An electronic device, Characterized in that, It includes a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, Characterized in that, The storage medium stores a program, and the program is executed by the processor to implement the method according to any one of claims 1 to 4.
8. A computer program product, including a computer program, Characterized in that, The computer program, when executed by the processor, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Bird sound recognition model training method, classification method, device and medium
CN114372513A
Birdwatching System
US20200273484A1