Method and apparatus for speech source separation based on convolutional neural network
By aggregating the multi-scale convolutional neural network (CNN) speech source separation method, the robustness problem of speech separation under complex background noise is solved, efficient speech extraction in professional content is achieved, and the separation effect and quality are improved.
Patent Information
- Application Number
- CN202080035468.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-24
- Filing Date
- 2020-05-13
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2040-05-13
AI Technical Summary
Existing speech separation methods are ineffective when dealing with complex background noise, especially in professional content where it is difficult to effectively extract dialogue or speech. Traditional methods assume that the background noise is steady-state and cannot adapt to dynamic backgrounds.
An aggregated multi-scale convolutional neural network (CNN) is used for speech source separation, which extracts speech features through multiple parallel convolution paths, generates output masks, and performs post-processing to improve the separation effect, including cascaded pooling and filtering operations.
The robustness and separation effect of speech separation are improved in both steady-state and non-steady-state backgrounds, the extraction quality of dialogue or speech is enhanced, and noise interference is reduced.
Smart Images

Figure CN114341979B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of International Patent Application No. PCT / CN2019 / 086769, filed on May 14, 2019, U.S. Provisional Patent Application No. 62 / 856,888, filed on June 4, 2019, and European Patent Application No. 19188010.3, filed on July 24, 2019, each of which is hereby incorporated by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to a method and apparatus for speech source separation based on convolutional neural networks (CNNs), and more particularly to improving speech extraction from raw noisy speech signals using an aggregated multi-scale CNN.
[0004] Although some embodiments will be described herein with particular reference to the present disclosure, it will be understood that the disclosure is not limited to such areas of use and is applicable in a much broader context. Background Art
[0005] Any discussion of the background art throughout the disclosure should not be considered as an admission that the background art is widely known or forms part of the common general knowledge in the field.
[0006] Speech source separation aims to recover the target speech from background noise and has numerous applications in speech and / or audio technology. In this context, speech source separation is often referred to as the "cocktail party problem." In this context, extracting dialogue from professional content (such as movies and television) presents challenges due to the complex background.
[0007] Currently, most separation methods only focus on static background or noise. Two traditional monophonic speech separation methods are speech enhancement and computational auditory scene analysis (CASA).
[0008] The simplest and most widely used enhancement method is spectral subtraction [SF Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE Trans. Acoust. Speech Sig. Process., vol. 27, pp. 113-120, 1979], in which the power spectrum of the estimated noise is subtracted from the power spectrum of the noisy speech. Background estimation assumes that the background noise is stationary, meaning that its spectral characteristics do not change abruptly over time, or at least are more stationary than the speech. However, this assumption can be limiting when applied to professional content.
[0009] CASA works by using perceptual principles from auditory scene analysis and exploiting grouping cues such as pitch and onset. For example, the tandem algorithm separates voiced speech by alternating between pitch estimation and pitch-based grouping [G. Hu and D. L. Wang, “A tandem algorithm for pitch estimation and voiced speech segregation,” IEEE Trans. Audio Speech Lang. Proc., vol. 18, pp. 2067–2079, 2010].
[0010] A recent approach treats speech separation as a supervised learning problem that has benefited from the rapid rise of deep learning. The original idea of supervised speech separation was inspired by the concept of time-frequency (TF) masking in CASA.
[0011] Deep neural networks (DNNs) have been shown to significantly improve the performance of supervised speech separation. Types of DNNs include feedforward multilayer perceptrons (MLPs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), and generative adversarial networks (GANs). CNNs are a type of feedforward network.
[0012] However, despite the use of DNNs for speech separation, there is still a need for a robust separation method to extract dialogue / speech in professional content in both steady-state and non-stationary (dynamic) contexts. Summary of the Invention
[0013] According to a first aspect of the present disclosure, a method for speech source separation based on a convolutional neural network (CNN) is provided. The method may include the following steps: (a) providing multiple frames of time-frequency transform of an original noisy speech signal. The method may further include the step of (b) inputting the time-frequency transform of the multiple frames into an aggregated multi-scale CNN having multiple parallel convolution paths, wherein each parallel convolution path includes one or more convolution layers. The method may further include the step of (c) extracting and outputting features from the time-frequency transform of the multiple frames input through each parallel convolution path. The method may further include the step of (d) obtaining an aggregated output of the outputs of the parallel convolution paths. And the method may further include the step of (e) generating an output mask (mask) for extracting speech from the original noisy speech signal based on the aggregated output.
[0014] In some embodiments, the original noisy speech signal may include one or more of high-pitched, cartoon, and other abnormal speech.
[0015] In some embodiments, the time-frequency transform of the multiple frames may pass through a two-dimensional convolutional layer and then a leaky rectified linear unit (LeakyRelu) before being input into the aggregated multi-scale CNN.
[0016] In some embodiments, obtaining the aggregated output in step (d) may further include applying weights to corresponding outputs of the parallel convolutional paths.
[0017] In some embodiments, different weights may be applied to the corresponding outputs of the parallel convolutional paths based on speech and / or audio domain knowledge and one or more of the trainable parameters learned from the training process.
[0018] In some embodiments, obtaining the aggregated output in step (d) may include concatenating the weighted outputs of the parallel convolutional paths.
[0019] In some embodiments, obtaining the aggregate output in step (d) may include summing the weighted outputs of the parallel convolutional paths.
[0020] In some embodiments, in step (c), speech harmonic features may be extracted and outputted by each parallel convolution path.
[0021] In some embodiments, the method may further include step (f): post-processing the output mask.
[0022] In some embodiments, the output mask may be a single-frame spectral amplitude mask, and post-processing the output mask may include at least one of the following steps: (i) limiting the output mask to in is set based on a statistical analysis of the target masks in the training data; (ii) if the average mask of the current frame is less than ε, the output mask is set to 0; (iii) if the input is zero, the output mask is set to zero; or (iv) J*K median filtering.
[0023] In some embodiments, the output mask may be a single-frame spectrum amplitude mask, and the method may further include the following step (g): multiplying the output mask with the amplitude spectrum of the original noisy speech signal, performing ISTFT and obtaining a wav signal.
[0024] In some embodiments, generating the output mask in step (e) may include applying cascaded pooling to the aggregated output.
[0025] In some embodiments, the cascaded pooling may include performing one or more stages of paired convolutional layers and pooling processes, where the one or more stages are followed by a final convolutional layer.
[0026] In some embodiments, a flattening operation may be performed after cascaded pooling.
[0027] In some embodiments, as the pooling process, an average pooling process may be performed.
[0028] In some embodiments, each of the multiple parallel convolution paths of the CNN may include L convolution layers, where L is a natural number ≥ 1, and the lth layer of the L layers has N l filters, where l=1…L.
[0029] In some embodiments, for each parallel convolution path, the number of filters N in layer l is l N l =l*N0, where N0 is a predetermined constant ≥1.
[0030] In some embodiments, within each parallel convolution path, the filter sizes of the filters may be the same.
[0031] In some embodiments, the filter sizes of the filters may differ between different parallel convolution paths.
[0032] In some embodiments, for a given parallel convolution path, the input may be zero-padded before performing the convolution operation in each of the L convolutional layers.
[0033] In some embodiments, for a given parallel convolution path, the filter may have a filter size of n*n, or the filter may have filter sizes of n*1 and 1*n.
[0034] In some embodiments, the filter size may depend on the harmonic length for feature extraction.
[0035] In some embodiments, for a given parallel convolutional path, filters of at least one of the layers of the parallel convolutional path may be dilated two-dimensional convolutional filters.
[0036] In some embodiments, the dilation operation of the filter of at least one of the layers of the parallel convolutional path may be performed only on the frequency axis.
[0037] In some embodiments, for a given parallel convolution path, filters of two or more layers in the layers of the parallel convolution path may be dilated two-dimensional convolution filters, and the dilation factor of the dilated two-dimensional convolution filter may increase exponentially with the increase of the number of layers l.
[0038] In some embodiments, for a given parallel convolution path, the dilation in the first of the L convolutional layers can be (1, 1), the dilation in the second of the L convolutional layers can be (1, 2), the dilation in the lth of the L convolutional layers can be (1, 2^(l-1)), and the dilation in the last of the L convolutional layers can be (1, 2^(L-1)), where (c, d) can represent the dilation factor c along the time axis and the dilation factor d along the frequency axis.
[0039] In some embodiments, for a given parallel convolution path, nonlinear operations may also be performed in each of the L convolutional layers.
[0040] In some embodiments, the nonlinear operation may include one or more of a parametric rectified linear unit (PRelu), a rectified linear unit (Relu), a leaky rectified linear unit (LeakyRelu), an exponential linear unit (Elu), and a scaled exponential linear unit (Selu).
[0041] In some embodiments, as a nonlinear operation, a rectified linear unit (ReLU) may be performed.
[0042] According to a second aspect of the present disclosure, a device for speech source separation based on a convolutional neural network (CNN) is provided, wherein the device includes a processor configured to perform the steps of a method for speech source separation based on a convolutional neural network (CNN).
[0043] According to a third aspect of the present disclosure, there is provided a computer program product comprising a computer-readable storage medium having instructions, the instructions being adapted to, when executed by a device having processing capabilities, cause the device to perform a method for speech source separation based on a convolutional neural network (CNN). BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Example embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which:
[0045] Figure 1 A flow chart illustrating an example of a method for speech source separation based on a convolutional neural network (CNN) is illustrated.
[0046] Figure 2 A flow chart of a further example of a method for speech source separation based on a convolutional neural network (CNN) is illustrated.
[0047] Figure 3 Illustrated is an example of an aggregated multi-scale convolutional neural network (CNN) for speech source separation.
[0048] Figure 4Illustrated is an example of a portion of a processing flow for convolutional neural network (CNN) based speech source separation.
[0049] Figure 5 An example of a cascaded pooling structure is shown.
[0050] Figure 6 An example of complexity reduction is illustrated. DETAILED DESCRIPTION
[0051] Speech Source Separation Based on Aggregated Multi-Scale Convolutional Neural Network (CNN)
[0052] The following describes a method and apparatus for speech source separation based on a convolutional neural network (CNN). This approach is particularly valuable for extracting dialogue from professional content, such as movies or television. CNN-based speech source separation is based on feature extraction from the spectrum of the original noisy signal with different receptive fields. Multi-scale features and multi-frame input enable the model to fully exploit both temporal and frequency information.
[0053] Overview
[0054] refer to Figure 1 An example is provided of a method for speech source separation based on a convolutional neural network (CNN). In step 101, a plurality of frames (e.g., M frames) of a time-frequency transform of an original noisy speech signal are provided. Although normal speech can be used as the original noisy speech signal, in one embodiment, the original noisy speech signal may include one or more of high-pitched, cartoon, and other abnormal speech. Abnormal speech may include, for example, emotional speeches, excited and / or angry voices, children's voices used in cartoons. Abnormal speech, high-pitched speech is characterized by a high dynamic range of audio and / or sparse harmonics of audio components. Although there is no limit to the number of frames, in one embodiment, 8 frames of the time-frequency transform of the original noisy speech signal may be provided.
[0055] Alternatively or additionally, an N-point short-time Fourier transform (STFT) can be used to provide the spectral amplitude of the speech based on multiple frames of the original noisy speech signal. In this case, the selection of N can be based on the expected time length of the sampling rate and the frame according to the following formula:
[0056] N = time length * sampling rate
[0057] There may or may not be data overlap between two adjacent frames. Generally, a longer N may result in better frequency resolution but increase computational complexity. In one embodiment, N may be selected to be 1024 and the sampling rate to be 16 kHz.
[0058] In step 102, the multi-frame time-frequency transform of the original noisy speech signal is input into a converged multi-scale CNN having multiple parallel convolutional paths. Each convolutional path includes one or more convolutional layers, such as a cascade of one or more convolutional layers. The structure of the converged multi-scale CNN is described in more detail below.
[0059] In step 103, features are extracted and output from the input multi-frame time-frequency transforms by each parallel convolution path of the aggregated multi-scale CNN. In one embodiment, speech harmonic features can be extracted and output by each parallel convolution path. Alternatively or additionally, correlations between harmonic features in different receptive fields can be extracted.
[0060] In step 104, an aggregate output of the outputs of the parallel convolutional paths is obtained. In one embodiment, obtaining the aggregate output may further include applying weights to the respective outputs of the parallel convolutional paths. In one embodiment, different weights may be applied based on one or more of speech and / or audio domain knowledge, or may be trainable parameters that can be learned from a training process, as further described below.
[0061] In step 105 , an output mask for extracting speech from the original noisy speech signal is generated based on the aggregated output.
[0062] refer to Figure 2 For example, in one embodiment, the method for speech source separation based on convolutional neural network (CNN) may further include a step 106 of post-processing the output mask. In one embodiment, the output mask may be a single-frame spectrum amplitude mask. The single-frame spectrum amplitude mask may be defined as follows:
[0063]
[0064] Where S(t,f) represents the spectral amplitude of clean speech and Y(t,f) represents the spectral amplitude of noisy speech. Post-processing the output mask in step 106 may include at least one of the following steps: (i) limiting the output mask to in The upper limit of the soft mask estimated from the network is set based on the statistical analysis of the target mask in the training data; the soft mask represents the portion of speech in each time-frequency block and is usually between 0 and 1. However, in some cases where phase cancellation occurs, the soft mask (i.e. ) can be greater than 1. To avoid CNN generating inappropriate masks, the soft mask is limited to a maximum value therefore, It can be equal to 1 or greater than 1, for example, equal to 2, or, for example, equal to any other intermediate real number between 1 and 2; (ii) if the average mask of the current frame is less than ε, the output mask is set to 0; (iii) if the input is zero, the output mask is set to zero; or (iv) J*K median filtering. The J*K median filter is a filter of size J*K, where J is the range of the frequency dimension and K is the range of the time dimension. Using the median filter, the target soft mask is replaced by the median of the mask, for example, in its J*K surrounding neighbors. The median filter is used for smoothing to avoid sudden changes in the frequency and time dimensions. In one embodiment, J=K=3. In other embodiments, J*K can be equal to 3*5, or 7*3, or 5*5. However, J*K can be equal to any other filter size suitable for the specific implementation of CNN. The post-processing step (i) ensures that the separation result does not cause audio clipping. The post-processing step (ii) can remove residual noise for use as sound activation detection. Post-processing step (iii) can avoid edge effects associated with applying the median filter in step (iv). Post-processing step (iv) can smooth the output mask and remove audible artifacts. By performing post-processing, the perceptual quality of the separation result can be improved.
[0065] In step 107, in one embodiment, the output mask may be a single-frame spectrum amplitude mask, and the method may further include multiplying the output mask with the amplitude spectrum of the original noisy speech signal, performing an inverse short-time Fourier transform (ISTFT), and obtaining a wav signal.
[0066] The above-mentioned method for speech source separation based on convolutional neural network (CNN) can be implemented on a corresponding device including a processor configured to perform the method. Alternatively or additionally, the above-mentioned method for speech source separation based on convolutional neural network (CNN) can be implemented as a computer program product including a computer-readable storage medium having instructions suitable for causing a device to perform the method.
[0067] Aggregate multi-scale convolutional neural network structure
[0068] refer to Figure 3 An example of an aggregated multi-scale convolutional neural network (CNN) for speech source separation is shown. In the described method, a pure convolutional network is used for feature extraction. The aggregated multi-scale CNN includes multiple parallel convolution paths. Although the number of parallel convolution paths is not limited, the aggregated multi-scale CNN may include three parallel convolution paths. Using these parallel convolution paths, it is possible to extract different (e.g., local and global) feature information of the time-frequency transform of multiple frames of the original noisy speech signal at different scales.
[0069] refer to Figure 3In the example of , in step 201, the time-frequency transform of multiple frames (e.g., M frames) of the original noisy speech signal can be input into an aggregated multi-scale CNN with multiple parallel convolution paths. Figure 3 In the example shown, three parallel convolution paths are shown. N-point short-time Fourier transforms (STFTs) can be applied to the multiple frames (e.g., M frames). Therefore, the input to the CNN may correspond to a dimension of M*(N / 2+1). N can be 1024.
[0070] refer to Figure 4 For example, in one embodiment, before being input into the aggregated multi-scale CNN, in step 201, the time-frequency transform of multiple frames of the original noisy speech signal may be subjected to a two-dimensional convolutional layer in step 201a, followed by a leaky rectified linear unit (LeakyRelu) in step 201b. The two-dimensional convolutional layer may have N filters (also referred to as N_filters), where N is a natural number greater than 1. The filter size of this layer may be (1, 1). In addition, this layer may not be dilated.
[0071] like Figure 3 As shown in the example, in step 201, the time-frequency transforms of multiple frames are input (in parallel) into multiple parallel convolution paths. In one embodiment, each of the multiple parallel convolution paths of the CNN may include L convolution layers 301, 302, 303, 401, 402, 403, 501, 502, 503, where L is a natural number > 1, and the lth layer of the L layers has Nl filters, where l = 1...L. Although the number of layers L in each parallel convolution path is not limited, each parallel convolution path may include, for example, L = 5 layers. In one embodiment, for each parallel convolution path, the number of filters Nl in the lth layer may be given by Nl = l*N0, where N0 is a predetermined constant > 1.
[0072] In one embodiment, the filter size of the filters in each parallel convolution path can be the same (i.e., uniform). For example, a filter size of (3, 3) (i.e., 3*3) can be used in each layer L in the parallel convolution paths 301-303 in the multiple parallel convolution paths. By using the same filter size in each parallel convolution path, the mixing of features of different scales can be avoided. In this way, the CNN learns to extract features of the same scale in each path, which greatly improves the convergence speed of the CNN.
[0073] In one embodiment, the filter sizes of the filters in different parallel convolution paths 301-303, 401-403, and 501-503 can be different. For example, but not by way of limitation, if the aggregated multi-scale CNN includes three parallel convolution paths, the filter size in the first parallel convolution path 301-303 can be (3,3), the filter size in the second parallel convolution path 401-403 can be (5,5), and the filter size in the third parallel convolution path 501-503 can be (7,7). However, other filter sizes are also feasible, where the same filter size can be used within each parallel convolution path, and different filter sizes can be used between different parallel convolution paths. The different filter sizes of the filters in different parallel convolution paths represent the different scales of the CNN. In other words, by using multiple filter sizes, multi-scale processing is possible. For example, when the filter size is small (e.g., 3*3), a small range of information is processed around the target frequency-time slice, while when the filter size is large (e.g., 7*7), a large range of information is processed. Processing information over a small area is equivalent to extracting so-called "local" features. Processing information over a large area is equivalent to extracting so-called "global" features. The inventors have discovered that features extracted by different filter sizes have different characteristics. Using a large filter size tends to preserve more speech harmonics but also retains more noise, while using a smaller filter size preserves more of the key components of the speech and more aggressively removes noise.
[0074] In one embodiment, the filter size may depend on the harmonic length for feature extraction.
[0075] In one embodiment, for a given convolution path, before performing the convolution operation in each of the L convolutional layers, the input of each layer may be zero-padded. In this way, the same data shape can be maintained from input to output.
[0076] In one embodiment, for a given parallel convolution path, a nonlinear operation may also be performed in each of the L convolutional layers. Although the nonlinear operation is not limited, in one embodiment, the nonlinear operation may include one or more of a parametric rectified linear unit (PRelu), a rectified linear unit (Relu), a leaky rectified linear unit (LeakyRelu), an exponential linear unit (Elu), and a scaled exponential linear unit (Selu). In one embodiment, a rectified linear unit (Relu) may be performed as the nonlinear operation. The nonlinear operation may be used as an activation in each of the L convolutional layers.
[0077] In one embodiment, for a given parallel convolution path, the filter of at least one layer in the layers of the parallel convolution path may be a dilated two-dimensional convolution filter. The use of dilated filters enables the extraction of correlations between harmonic features in different receptive fields. Dilation enables reaching a distant receptive field by jumping (i.e., skipping, jumping over) a series of time-frequency (TF) segments. In one embodiment, the dilation operation of the filter of at least one layer in the layers of the parallel convolution path may be performed only on the frequency axis. For example, in the context of the present disclosure, dilation (1, 2) may indicate no dilation along the time axis (dilation factor 1), while skipping every other segment on the frequency axis (dilation factor 2). In general, dilation (1, d) may indicate skipping (d-1) segments along the frequency axis between segments used for feature extraction by the corresponding filter.
[0078] In one embodiment, for a given convolution path, the filters of two or more layers in the layers of the parallel convolution path can be dilated two-dimensional convolution filters, where the dilation factor of the dilated two-dimensional convolution filter increases exponentially with the increase of the number of layers l. In this way, exponential receptive field growth with depth can be achieved. Figure 3 As shown in the example, in one embodiment, for a given parallel convolution path, the dilation in the first layer of the L convolution layers may be (1, 1), the dilation in the second layer of the L convolution layers may be (1, 2), the dilation in the lth layer of the L convolution layers may be (1, 2^(l-1)), and the dilation in the last layer of the L convolution layers may be (1, 2^(L-1)), where (c, d) represents the dilation factor c along the time axis and the dilation factor d along the frequency axis.
[0079] An aggregated multi-scale CNN can be trained. The training of an aggregated multi-scale CNN may involve the following steps:
[0080] (i) Calculate the frame FFT coefficients of the original noisy speech and the target speech;
[0081] (ii) Obtain the amplitude of the noisy speech and the target speech by ignoring the phase;
[0082] (iii) The target output mask is obtained by calculating the difference between the amplitude of the noisy speech and the target speech as follows:
[0083] Target mask = ||Y(t,f)|| / ||X(t,f)||
[0084] Where Y(t,f) and X(t,f) represent the spectral amplitudes of the target speech and the noisy speech;
[0085] (iv) limiting the target mask to a small range based on the statistical histogram;
[0086] Due to the negative correlation between the target speech and the interference, the initial target mask can have a very wide range of values. According to statistical results, masks in [0, 2] may account for approximately 90%, and masks in [0, 4] may account for approximately 98%. Based on training results, the mask may be restricted to [0, 2] or [0, 4]. These statistical results may be related to the speech and background type, as well as the mask restrictions, but they may be important for training CNNs.
[0087] (v) using multiple frame rate amplitudes of noisy speech as input;
[0088] (vi) Use the corresponding target mask from step (iii) as output.
[0089] To train the aggregated multi-scale CNN, high-pitched, cartoon, and abnormal speech can be covered to increase robustness.
[0090] Path weighting and aggregation
[0091] refer to Figure 3 For example, from steps 303, 403, and 503, features extracted from the time-frequency transform of multiple frames of the original noisy speech signal input in step 201 in each parallel convolution path of the aggregated multi-scale CNN are output. Then, in step 202, the outputs from each parallel convolution path are aggregated to obtain an aggregated output.
[0092] In one embodiment, the step of obtaining the aggregated output may include applying weights 304 (W1), 404 (W2), 504 (W3) to the corresponding outputs 303, 403, 503 of the parallel convolutional paths. In one embodiment, different weights 304 (W1), 404 (W2), 504 (W3) may be applied to the corresponding outputs of the parallel convolutional paths based on speech and / or audio domain knowledge and one or more trainable parameters learned from the training process. The trainable parameters may be obtained during the training process of the aggregated multi-scale CNN, where the trainable parameters may be the weights themselves, which may be directly learned from the entire training process along with other parameters.
[0093] In general, a larger filter size of a CNN may retain more speech components while involving more noise, while a smaller filter size may only retain some key components of speech while removing more noise. For example, if a larger weight is selected for a path with a larger filter size, the model may be relatively more conservative and have relatively better speech preservation at the expense of more residual noise. On the other hand, if a larger weight is selected for a path with a smaller filter size, the model may remove noise more aggressively and may also lose some speech components. Therefore, applying weights to the outputs of parallel convolution paths can be used to control the aggressiveness of the CNN, such as by achieving a preferred trade-off between speech preservation and noise removal in the above example.
[0094] In one embodiment, obtaining the aggregated output in step 202 may include concatenating the weighted outputs of the parallel convolution paths. If the input to the aggregated multi-scale CNN is, for example, M*(N / 2+1), then the dimension of the output in the case of concatenation may be 3*(n_filters*n)*M*(N / 2+1).
[0095] In one embodiment, obtaining the aggregate output in step 202 may include adding the weighted outputs of the parallel convolution paths. If the input to the aggregated multi-scale CNN is, for example, M*(N / 2+1), then the dimension of the output in the case of addition may be (n_filters*n)*M*(N / 2+1). It should be noted that in the present disclosure, for example, for the filters of the L convolution layers of the parallel convolution paths of the CNN, the number of filters may be expressed as N_filters*n, expressed as N0_filters*{l}, or expressed as N0*l=N l , while for other convolutional layers, the number of filters can be expressed as N filters or N_filters.
[0096] Cascade Pooling
[0097] Since multiple frames of the time-frequency transform of the original noisy speech signal are input into the aggregated multi-scale CNN, the feature extraction of the CNN is also performed on multiple frames. Figure 5An example of , showing a cascade pooling structure. In one embodiment, generating the output mask includes applying cascade pooling to the aggregated output in step 601. By applying cascade pooling, multi-frame features can be used to predict a single-frame output mask by finding the most effective features. In one embodiment, cascade pooling may include performing one or more stages of paired convolution layers 602, 604, 606 and pooling processing 603, 605, 607, wherein the one or more stages are followed by a final convolution layer, 608. The pooling process may be performed only on the time axis. In one embodiment, average pooling may be performed as the pooling process. In the convolution layers 602, 604, 606, the number of filters may be reduced. While the number of filters in the first convolution layer 602 is not limited, for example, it may be N_filters*4 or N_filters*2, the number of filters may be gradually reduced, otherwise, the performance may degrade. In addition, in the final convolution layer 608, the number of filters must be 1. In Figure 5 In the example, without limitation, the number of filters ranges from N_filters*4 in the first convolutional layer 602, to N_filters*2 in the second convolutional layer 604, to N_filters in the third convolutional layer 606, to 1 in the final convolutional layer 608. The filter size may depend on the number of frames M. Alternatively or additionally, the filter size of each convolutional layer may depend on the output frame size of the previous pooling layer. If the output frame size of the previous pooling layer is larger than the filter size on the time axis, the filter sizes in each convolutional layer may be the same, for example (3,1). If the output frame size of the previous pooling layer is smaller than the filter size on the time axis of the previous convolutional layer, assuming that the previous pooling layer has M' frame outputs, for example M'<3, the filter size of the current convolutional layer may be (M', 1). In Figure 5 In the example, in the first convolution layer 602, the filter size (3, 1) is used, in the second convolution layer 604, the filter size (3, 1) is used, in the third convolution layer 606, the filter size (2, 1) is used, and in the last convolution layer 608, the filter size (1, 1) is used. Non-linear operations can be performed in each convolution layer. Figure 5 In the example of FIG, , rectified linear units (ReLU) are performed in convolutional layers 602, 604, and 606, and leaky rectified linear units (Leaky ReLU) are performed in the last convolutional layer 608. In one embodiment, a flattening operation 609 may be performed after the cascaded pooling.
[0098] Reduced complexity
[0099] refer to Figure 6For example, in one embodiment, for a given parallel convolution path, the filter can have a filter size of n*n, 701, or the filter can have a filter size of n*1, 701a and a filter size of 1*n, 701b. The filter can be applied in the frequency-time dimension, so the filter size n*n can represent a filter with a filter length of n in the frequency axis and a filter length of n in the time axis. Similarly, the filter size n*1 can represent a filter with a filter length of n in the frequency axis and a filter length of 1 in the time axis, while the filter size 1*n can represent a filter with a filter length of 1 in the frequency axis and a filter length of n in the time axis. The filter size n*n can be replaced by a concatenation of a filter of size n*1 and a filter of size 1*n. Therefore, complexity reduction can be achieved as follows. For example, for an n*n filter, there are n*n parameters. If it is assumed that there are 64 such filters in one layer of the L layer, the number of parameters will be 64*n*n. By replacing the filter size n*n with the concatenation of two filters of size n*1 and 1*n respectively, the parameters will be only 64*n*1*2, thus reducing the complexity of the model.
[0100] explain
[0101] Unless otherwise specifically noted, as will be apparent from the following discussion, it should be understood that throughout this specification, discussions utilizing terms such as "process," "compute," "calculate," "determine," etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that processes and / or transforms data represented as physical quantities (e.g., electronic quantities) into other data similarly represented as physical quantities.
[0102] In a similar manner, the term "processor" may refer to any device or portion of a device that processes electronic data, for example from registers and / or memory, to transform that electronic data into other electronic data, for example, that may be stored in registers and / or memory. A "computer" or "computing machine" or "computing platform" may include one or more processors.
[0103] In one example embodiment, the methods described herein may be executed by one or more digital processors that accept computer-readable (also referred to as machine-readable) code comprising an instruction set that, when executed by the one or more processors, performs at least one of the methods described herein. Any processor capable of executing an instruction set (sequential or otherwise) specifying an action to be taken may be included. Thus, an example is a typical processing system comprising one or more processors. Each processor may comprise one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may also include a memory subsystem comprising main RAM and / or static RAM and / or ROM. A bus subsystem may be included for communication between components. The processing system may also be a distributed processing system having processors coupled via a network. If the processing system requires a display, it may include a display such as a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data entry is required, the processing system may also include an input device, such as one or more of an alphanumeric input unit such as a keyboard, a pointing control device such as a mouse, and the like. The processing system may also include a storage system such as a disk drive. In some configurations, the processing system may include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium carrying computer-readable code (e.g., software), the computer-readable code including a set of instructions to cause one or more of the methods described herein to be performed when executed by one or more processors. It should be noted that when the method includes several elements, e.g., several steps, no ordering of these elements is implied unless otherwise specified. The software may reside on a hard disk, or may reside entirely or at least partially in RAM and / or a processor during execution by the computer system. Thus, the memory and the processor also constitute a computer-readable carrier medium carrying the computer-readable code. Furthermore, the computer-readable carrier medium may form or be included in a computer program product.
[0104] In alternative example embodiments, the one or more processors operate as a standalone device or may be connected to other processors in a networked deployment, for example, a network, one or more processors may operate as a server or user machine in a server-user network environment, or as a peer machine in a peer-to-peer or distributed network environment. The one or more processors may form a personal computer (PC), a tablet computer, a personal digital assistant (PDA), a cellular phone, a network appliance, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify operations to be taken by the machine.
[0105] It should be noted that the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0106] Thus, an example embodiment of each method described herein is in the form of a machine-readable carrier medium carrying an instruction set, such as a computer program, for execution on one or more processors, such as one or more processors that are part of a network server arrangement. Thus, as will be appreciated by those skilled in the art, example embodiments of the present disclosure may be embodied as a method, an apparatus such as a dedicated apparatus, an apparatus such as a data processing system, or a computer-readable carrier medium, such as a computer program product. The computer-readable carrier medium carries computer-readable code comprising an instruction set that, when executed on one or more processors, causes the one or more processors to implement a method. Thus, various aspects of the present disclosure may take the form of a method, a fully hardware example embodiment, a fully software example embodiment, or an example embodiment combining software and hardware aspects. Furthermore, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable carrier medium) carrying computer-readable program code embodied in the medium.
[0107] Software can also be sent or received over a network via a network interface device. Although the carrier medium is a single medium in the example embodiments, the term "carrier medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instruction sets. The term "carrier medium" should also be understood to include any medium that can store, encode, or carry an instruction set for execution by one or more processors and cause one or more processors to perform any one or more of the methods according to the present disclosure. The carrier medium can take a variety of forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory, such as main memory. Transmission media include coaxial cables, copper wires, and optical fibers, including lines comprising bus subsystems. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communications. For example, the term "carrier medium" shall accordingly include, but not be limited to, solid-state memories, computer products embodied in optical and magnetic media; media carrying propagated signals detectable by at least one processor or one or more processors and representing a set of instructions for implementing a method when executed; and transmission media in a network carrying propagated signals detectable by at least one of one or more processors and representing a set of instructions.
[0108] It will be appreciated that in one example embodiment, the steps of the method discussed are performed by one (or more) suitable processors of a processing (e.g., computer) system executing instructions (computer-readable code) stored in a memory. It will also be appreciated that the present disclosure is not limited to any particular implementation or programming technique, and that the present disclosure may be implemented using any suitable technique for implementing the functionality described herein. The present disclosure is not limited to any particular programming language or operating system.
[0109] References throughout this specification to "one example embodiment," "some example embodiments," or "example embodiments" mean that a particular feature, structure, or characteristic described in connection with the example embodiment is included in at least one example embodiment of the present disclosure. Thus, appearances of the phrases "one example embodiment," "some example embodiments," or "example embodiments" throughout this specification are not necessarily all referring to the same example embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more example embodiments, as will be apparent to one of ordinary skill in the art from this disclosure.
[0110] As used herein, unless otherwise indicated, the use of ordinal adjectives "first," "second," "third," etc., to describe common objects merely indicates that different instances of similar objects are being referred to and is not intended to imply that the objects so described must be in a given sequence in time, space, ranking, or in any other manner.
[0111] In the following claims and the description herein, any of the terms comprising, consisting of or including is an open term, which means that at least the elements / features that follow are included, but other elements / features are not excluded. Therefore, when used in the claims, the term comprising should not be interpreted as being limited to the means or elements or steps listed thereafter. For example, the scope of the statement that the device comprises A and B should not be limited to the device consisting of only elements A and B. Any of the terms comprising or including as used herein is also an open term, which also means that at least the elements / features that follow the term are included, but other elements / features are not excluded. Therefore, comprising is synonymous with including and refers to including.
[0112] It should be understood that in the above description of exemplary embodiments of the present invention, various features are sometimes grouped together into a single embodiment, figure, or description thereof in order to simplify the disclosure and aid in understanding one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claims require more features than those expressly recited therein. On the contrary, as reflected in the following claims, inventive aspects lie in less than all the features of a single, previously disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate exemplary embodiment of the present invention.
[0113] In addition, although some example embodiments described herein include some features that are included in other example embodiments but not others, as will be understood by those skilled in the art, combinations of features of different example embodiments are contemplated within the scope of this disclosure and form different embodiments. For example, in the following claims, any of the claimed example embodiments may be used in any combination.
[0114] In the description provided herein, numerous specific details are set forth. However, it should be understood that embodiments of the present invention may be practiced without these specific details. In other cases, well-known methods, structures, and techniques are not shown in detail in order not to obscure understanding of this description.
[0115] Thus, while what is believed to be the best mode of the present invention has been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the present invention, and it is intended that all such changes and modifications be claimed as falling within the scope of the present invention. For example, any formulas given above are merely representative of processes that may be used. Functionality may be added to or deleted from the block diagrams, and operations may be interchanged between functional blocks. Steps may be added to or deleted from the methods described within the scope of the present invention.
[0116] Various aspects of the present invention may be understood from the following enumerated example embodiments (EEE):
[0117] EEE1. A method for speech source separation based on a convolutional neural network (CNN), wherein the method comprises the following steps:
[0118] (a) providing multiple frames of the time-frequency transform of the original noisy speech signal;
[0119] (b) inputting the time-frequency transform of the multiple frames into an aggregated multi-scale CNN with multiple parallel convolution paths;
[0120] (c) extracting and outputting features from the time-frequency transform of the input multiple frames through each parallel convolution path;
[0121] (d) obtaining an aggregate output of the outputs of the parallel convolutional paths; and
[0122] (e) Generate an output mask based on the aggregated output for extracting speech from the original noisy speech signal.
[0123] EEE 2. The method according to EEE 1, wherein the original noisy speech signal includes one or more of high-pitched, cartoon, and other abnormal speech.
[0124] EEE 3. The method according to EEE 1 or EEE 2, wherein the time-frequency transform of the multiple frames passes through a two-dimensional convolution layer and then a leaky rectified linear unit (LeakyRelu) before being input into the aggregated multi-scale CNN.
[0125] EEE 4. The method according to any one of EEEs 1-3, wherein obtaining the aggregate output in step (d) further comprises applying weights to corresponding outputs of the parallel convolution paths.
[0126] EEE 5. The method according to EEE 4, wherein different weights are applied to the corresponding outputs of the parallel convolution paths based on speech and / or audio domain knowledge and one or more trainable parameters learned from a training process.
[0127] EEE 6. The method according to EEE 4 or EEE 5, wherein obtaining the aggregate output in step (d) includes concatenating the weighted outputs of the parallel convolution paths.
[0128] EEE 7. The method according to EEE 4 or EEE 5, wherein obtaining the aggregate output in step (d) comprises adding weighted outputs of the parallel convolution paths.
[0129] EEE 8. The method according to any one of EEE 1-7, wherein, in step (c), speech harmonic features are extracted and output through each parallel convolution path.
[0130] EEE 9. The method according to any one of EEEs 1-8, wherein the method further comprises step (f): post-processing the output mask.
[0131] EEE 10. The method according to EEE 9, wherein the output mask is a single-frame spectral amplitude mask, and wherein post-processing the output mask comprises at least one of the following steps:
[0132] (i) Limit the output mask to in It is set based on the statistical analysis of the target masks in the training data;
[0133] (ii) If the average mask of the current frame is less than ε, set the output mask to 0;
[0134] (iii) if the input is zero, set the output mask to zero; or
[0135] (iv) J*K median filtering.
[0136] EEE 11. A method according to any one of EEE 1-10, wherein the output mask is a single-frame spectrum amplitude mask, and the method further comprises the following step (g): multiplying the output mask with the amplitude spectrum of the original noisy speech signal, performing ISTFT and obtaining a wav signal.
[0137] EEE 12. The method according to any one of EEEs 1-11, wherein generating the output mask in step (e) comprises applying cascaded pooling to the aggregated output. EEE 13.
[0138] EEE 13. The method according to EEE 12, wherein the cascaded pooling comprises performing one or more stages of paired convolutional layers and pooling processes, wherein the one or more stages are followed by a final convolutional layer.
[0139] EEE 14. The method according to EEE 12 or EEE 13, wherein a flattening operation is performed after the cascade pooling.
[0140] EEE 15. The method according to any one of EEE 12 to 14, wherein an average pooling process is performed as the pooling process.
[0141] EEE 16. A method according to any one of EEE 1 to 15, wherein each parallel convolution path of the multiple parallel convolution paths of the CNN includes L convolution layers, where L is a natural number ≥ 1, wherein the lth layer of the L layers has Nl filters, where l = 1…L.
[0142] EEE 17. The method according to EEE 16, wherein for each parallel convolution path, the number of filters N in layer l l By N l =l*N0, where N0 is a predetermined constant ≥1.
[0143] EEE 18. The method according to EEE 16 or EEE 17, wherein within each parallel convolution path, the filter sizes of the filters are the same.
[0144] EEE 19. The method according to EEE 18, wherein the filter sizes of the filters are different between different parallel convolution paths.
[0145] EEE 20. The method according to EEE 19, wherein, for a given parallel convolution path, the filter has a filter size of n*n, or the filter has a filter size of n*1 and 1*n.
[0146] EEE 21. The method according to EEE 19 or EEE 20, wherein the filter size depends on the harmonic length for feature extraction.
[0147] EEE 22. A method according to any one of EEEs 16-21, wherein, for a given parallel convolution path, the input is zero-padded before performing the convolution operation in each of the L convolutional layers.
[0148] EEE 23. A method according to any one of EEEs 16-22, wherein, for a given parallel convolution path, a filter of at least one of the layers of the parallel convolution path is a dilated two-dimensional convolution filter.
[0149] EEE 24. The method according to EEE 23, wherein the dilation operation of the filter of at least one of the layers of the parallel convolution path is performed only on the frequency axis.
[0150] EEE 25. A method according to EEE 23 or EEE 24, wherein, for a given parallel convolution path, filters of two or more layers in the layers of the parallel convolution path are dilated two-dimensional convolution filters, and wherein the dilation factor of the dilated two-dimensional convolution filter increases exponentially with the increase of the number of layers l.
[0151] EEE 26. The method according to EEE 25, wherein, for a given parallel convolution path, the dilation in the first layer of the L convolutional layers is (1, 1), the dilation in the second layer of the L convolutional layers is (1, 2), the dilation in the lth layer of the L convolutional layers is (1, 2^(l-1)), and the dilation in the last layer of the L convolutional layers is (1, 2^(L-1)), where (c, d) represents the dilation factor c along the time axis and the dilation factor d along the frequency axis.
[0152] EEE 27. A method according to any one of EEEs 16-26, wherein, for a given parallel convolution path, a nonlinear operation is additionally performed in each of the L convolutional layers.
[0153] EEE 28. The method according to EEE 27, wherein the nonlinear operation includes one or more of a parametric rectified linear unit (PRelu), a rectified linear unit (Relu), a leaky rectified linear unit (LeakyRelu), an exponential linear unit (Elu), and a scaled exponential linear unit (Selu).
[0154] EEE 29. The method according to EEE 28, wherein a rectified linear unit (ReLU) is performed as the nonlinear operation.
[0155] EEE 30. An apparatus for speech source separation based on convolutional neural network (CNN), wherein the apparatus comprises a processor configured to perform the steps of the method according to any one of EEE 1 to 29.
[0156] EEE 31. A computer program product comprising a computer-readable storage medium having instructions, the instructions being adapted to, when executed by a device having processing capabilities, cause the device to perform the method according to any one of EEE 1-29. EEE 32.
Claims
1. A method for speech source separation based on a convolutional neural network (CNN), the method comprising the following steps: Providing a plurality of frames of time-frequency transform of an original noisy speech signal; The time-frequency transforms of the multiple frames are input into an aggregated multi-scale CNN having multiple parallel convolution paths, each of the multiple parallel convolution paths of the CNN comprising a cascade of L convolutional layers, where L is a natural number > 1, and the lth layer of the L layers has N l filters, where l = 1…L, where the filter sizes of the filters differ between different parallel convolution paths, and where the filter sizes of the filters are the same within each parallel convolution path; Extracting and outputting features from the time-frequency transform of the plurality of input frames through each parallel convolution path; Get the aggregate output of the outputs of the parallel convolution paths; as well as An output mask for extracting speech from the original noisy speech signal is generated based on the aggregated output.
2. The method according to claim 1, wherein The time-frequency transforms of the multiple frames are passed through a 2D convolutional layer followed by a leaky rectified linear unit (LeakyRelu) before being input into the aggregated multi-scale CNN.
3. The method of claim 1 , wherein obtaining the aggregated output further comprises applying weights to corresponding outputs of the parallel convolutional paths.
4. The method according to claim 3, wherein: Based on speech and / or audio domain knowledge and one or more trainable parameters learned from the training process, different weights are applied to corresponding outputs of the parallel convolutional paths.
5. The method of claim 3, wherein obtaining the aggregate output comprises concatenating weighted outputs of parallel convolutional paths.
6. The method of claim 3, wherein obtaining the aggregate output comprises adding weighted outputs of the parallel convolutional paths.
7. The method according to claim 1, wherein In the step of extracting and outputting features, speech harmonic features are extracted and output through each parallel convolution path. The method according to claim 1 , further comprising the step of post-processing the output mask.
9. The method according to claim 8, wherein The output mask is a single-frame spectral magnitude mask, and wherein post-processing the output mask comprises at least one of the following steps: Constrain the output mask to [0,φ], where φ is set based on statistical analysis of target masks in the training data; If the average mask of the current frame is less than ε, the output mask is set to 0; If the input is zero, set the output mask to zero; or Perform a median filter of size J*K, where J is an integer representing the size of the frequency dimension and K is an integer representing the size of the time dimension.
10. The method according to claim 1, wherein The output mask is a single-frame spectrum amplitude mask, and the method further comprises the following steps: multiplying the output mask with the amplitude spectrum of the original noisy speech signal, performing ISTFT and obtaining a wav signal. The method of claim 1 , wherein generating the output mask comprises applying cascaded pooling to the aggregated output.
12. The method of claim 11, wherein the cascaded pooling comprises performing one or more stages of paired convolutional layers and pooling processes, wherein the one or more stages are followed by a final convolutional layer. The method of claim 11 , wherein a flattening operation is performed after the cascade pooling.
14. The method according to claim 12, wherein: As the pooling process, average pooling is performed.
15. The method according to any one of claims 1 to 14, wherein for each parallel convolution path, the number of filters N in the lth layer is l By N l = l * N0, where N0 is a predetermined constant ≥ 1.
16. The method according to any one of claims 1 to 14, wherein For a given parallel convolution path, the filter has a filter size of n*n, or the filter has a filter size of n*1 and 1*n.
17. The method according to any one of claims 1 to 14, wherein For a given parallel convolutional path, the input is zero-padded before performing the convolution operation in each of the L convolutional layers.
18. The method according to any one of claims 1 to 14, wherein: For a given parallel convolutional path, filters of at least one of the layers of the parallel convolutional path are dilated two-dimensional convolutional filters.
19. The method according to claim 18, wherein The dilation operation of the filter of at least one of the layers of the parallel convolution path is performed only on the frequency axis.
20. The method according to claim 18, wherein For a given parallel convolution path, filters of two or more layers in the layers of the parallel convolution path are dilated two-dimensional convolution filters, and wherein a dilation factor of the dilated two-dimensional convolution filters increases exponentially with an increase in the number of layers l.
21. The method according to claim 20, wherein For a given parallel convolutional path, the dilation in the first of the L convolutional layers is (1, 1), the dilation in the second of the L convolutional layers is (1, 2), the dilation in the lth of the L convolutional layers is (1, 2^(l-1)), and the dilation in the last of the L convolutional layers is (1, 2^(L-1)), where (c, d) represents the dilation factor c along the time axis and the dilation factor d along the frequency axis.
22. The method according to any one of claims 1 to 14, wherein For a given parallel convolutional path, additionally, nonlinear operations are performed in each of the L convolutional layers.
23. The method according to claim 22, wherein the nonlinear operation comprises one or more of a parametric rectified linear unit (PRelu), a rectified linear unit (Relu), a leaky rectified linear unit (LeakyRelu), an exponential linear unit (Elu), and a scaled exponential linear unit (Selu).
24. The method according to claim 23, wherein a rectified linear unit (ReLU) is performed as the nonlinear operation.
25. An apparatus for speech source separation based on a convolutional neural network (CNN), wherein the apparatus comprises a processor configured to execute the method according to any one of claims 1-24.
26. A computer program product comprising a computer-readable storage medium having instructions adapted, when executed by a device having processing capabilities, to cause the device to perform the method according to any one of claims 1 to 24.
27. A computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-24.
28. A device for speech source separation based on a convolutional neural network (CNN), comprising: one or more processors, and A computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 24.
29. An apparatus for speech source separation based on a convolutional neural network (CNN), comprising components for performing the method according to any one of claims 1-24.
Citation Information
Patent Citations
Computationally efficient method for filtering noise
CN105849804A
Semantic segmentation method based on multi-scale convolutional neural networks (CNNs)
CN108230329A