Two-step mixed sound source separation and de-reverberation method
Through the cascading method of Sepformer network and time convolution network, the separation and dereverberation of speech signals in the real sound field are solved, and efficient speech separation in noise and reverberation environments are achieved, and speech quality and system robustness are improved.
Patent Information
- Application Number
- CN202410133000.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-07-25
AI Technical Summary
In real indoor sound field environments, voice signals are easily disturbed by multiple speakers, noise and reverbs, and the prior art is difficult to effectively separate and de-reverbs, especially for single microphone systems, which lead to difficulty in understanding speech in people with hearing impairment.
The two-step mixed sound source separation and dereverberation method is adopted. First, the separation and noise reduction are performed through the Sepformer network, and then the dereverberation is performed using the time convolution network (TCN). Combining the scale invariance signal-to-noise ratio and the transposition-residue mechanism, the network parameters are optimized to improve the separation effect.
It significantly improves the separation quality and robustness of the voice signal, can effectively separate out clear voice signals in the presence of noise and reverberation, and improves the performance of the automatic audio recognition and recognition system.
Smart Images

Figure CN120375853A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice signal, and specifically to a two-step hybrid sound source separation and dereverberation method. Background Art
[0002] In a real indoor sound field environment, there may be multiple speakers, and the voice signals recorded by a microphone are often interfered by the noise and reverberation in the room, which poses a huge challenge to the recognition of indoor voice signals. Especially for hearing-impaired listeners, it is more difficult to understand the heard voice under the conditions of multiple speakers, noise and reverberation at the same time.
[0003] A single voice separation system aims to separate each target voice signal by using a single microphone. Compared with the situation where a multi-voice separation system can use multiple microphones, obviously a single microphone has more application scenarios.
[0004] Traditional voice separation methods such as Nonnegative matrix factorization (NMF) have been studied for many years. In recent years, deep learning algorithms have greatly improved the separation effect. Among them, Deep neural network (DNM), Recurrent neural network (RNN) and Long short-term memory network (LSTM) learn pure time-frequency domain signals or pure time-frequency domain signal masks from the mixed signals, including mask ratio methods based on Ideal binary mask (IBM), Ideal ratio mask (IRM), Ideal target mask (ITM), complex Ideal ratio mask (cIRM). Whether it is a direct method or a mask estimation method, the inverse short-time Fourier transform of the estimated amplitude spectrum of each source combined with the original sound source phase is used as a single sound source separation method. However, since the processed time-frequency domain features still need to be reconstructed by the inverse short-time Fourier transform with the help of the phase of the original signal in the end, this will greatly interfere with the upper limit of the accuracy of the reconstructed waveform. Moreover, the time-frequency domain separation method requires the highest possible resolution frequency domain calculation of the original signal, and the window width of the short-time Fourier transform needs to be widened, which will increase the calculation time of the system, which is a very obvious limitation for hearing devices with low latency and low power consumption. And the time-domain separation method that can avoid the problem of unifying the amplitude and phase of the sound source has become a new research hotspot.
[0005] Although in the field of source separation, time-domain end-to-end separation methods have achieved great success, most of them only target anechoic and noise-free speech signals or anechoic speech signals. Traditional speech dereverberation techniques such as the weight prediction method use constructed filters to complete the prediction of late reverberant signals (greater than 50 milliseconds). After multiple iterations and by using multi-channel audio information simultaneously, speech signal dereverberation is achieved. Dereverberation and noise reduction methods often require multi-channel information and the filter prediction ability drops severely in a strong noise environment. When only using single-channel audio information for iterative calculations, due to the reduced amount of information provided, the number of filter iterations needs to increase and it is easy to fall into a local optimum situation. In recent years, the application of deep neural networks to the field of speech dereverberation has emerged. By using a deep learning network to assist in finding the power spectral density (PSD) of the initial dereverberated signal, this provides a more accurate idea. In summary, previous methods often only studied the separation problem of mixed sound sources, or separately solved the problem of sound source dereverberation, or put the problem of sound source noise reduction into both of them at the same time. Therefore, in an actual sound field, it is difficult to simultaneously solve the separation, noise reduction, and dereverberation of mixed sound sources. Summary of the Invention
[0006] To solve the problem that existing speech separation methods may have multiple speakers in a real indoor environment and the speech signals recorded by microphones are often disturbed by room noise and room reverberation, the present invention provides a two-step method for separating and dereverberating mixed sound sources.
[0007] The present invention is implemented by the following technical solutions:
[0008] A two-step method for separating and dereverberating mixed sound sources, comprising the following steps:
[0009] Step S1: First, an encoding module converts the waveform of the original mixed audio signal collected by a single microphone into corresponding data feature blocks; then the data feature blocks are input into a masking module for estimating the masking feature blocks of each channel; then the masking feature blocks of each channel are multiplied pointwise with the mixed audio signal to obtain the data feature blocks of each channel; the data feature blocks of each channel are input into a decoding module to obtain the separated reverberant audio signal while reducing the noise of the mixed audio signal.
[0010] Step S2: Taking the separated audio signal as the input, using the early reverberation enhanced signal as the target to be learned, and using the scale-invariant signal-to-noise ratio as the loss function, after converting the input to the frequency domain, transposing and taking the logarithm, and then inputting it into a temporal convolutional network for dereverberation.
[0011] Step S3: Cascade the separation and noise reduction network in Step S1 and the demixing network in Step S2. The separation and noise reduction network is a separation attention network (res1 / res2 / noise are the targets to be learned during training), and the demixing network is a temporal convolutional - transposed - residual - WPE network (TCN - Inverse - Residual - WPE). The mixed audio signal first enters the separation and noise reduction network to output the reverberant audio signal. Each audio signal has been separated and the noise has been basically removed. Then, the two obtained reverberant audio signals are input into the same dereverberation network to complete the reverberation removal work under strong or weak noise conditions.
[0012] Furthermore, in the above - mentioned Step S1, the specific method for separating and reducing noise from the mixed audio signal is as follows: The mixed audio signal is sequentially input into the encoding module and the masking module, and finally, separation and noise reduction are completed through the decoding module. For the encoding module and the decoding module, first, the mixed audio signal is divided into N segments at a certain overlap rate, and each segment is sampled L times at a fixed frequency and sent into the encoding module to be transformed into features w. The decoding module transforms the estimated multiple features into the separated output, and then the overlapping - add method is used to reconstruct each separated signal.
[0013] (1);
[0014] where is the non - linear unit (relu) of the encoder, is the basic matrix of the encoder, and x refers to the original signal after overlapping sampling;
[0015] (2);
[0016] where is the basic matrix of the decoder, is the feature estimation value of each separated signal;
[0017] For the mask module, the Sepformer network is used as the core part of the separation and noise reduction stage. The latent representation output by the encoding module is subjected to layer regularization and linearization operations, and then sliced with an overlapping factor of 50% to adjust the representation dimension. After that, it is sent to the Sepfomer sub-unit. Each Sepfomer sub-unit contains one inter-Transformer and one intra-Transformer. The structural parameters of the two are the same and are used to extract the correlation of the intra-block and inter-block time series respectively. Each inter-Transformer and intra-Transformer contains layer regularization, multi-head attention mechanism (MHA), feed-forward network (FFW), and several residual connection mechanisms. Finally, the representation after several repetitions of the Sepfomer network architecture is successively subjected to activation layer (PRelu), linearization, padding, and forward propagation operations to obtain the output mask;
[0018] Finally, the mixed audio signal is multiplied by the output mask, and each result obtained by the multiplication is decoded by the decoding module to obtain the enhanced signal corresponding to each speaker after separation and noise reduction.
[0019] Further, in step S2, the specific method for dereverberating the mixed audio signal is as follows:
[0020] The acoustic signal of the indoor scene collected by a single microphone is modeled as follows:
[0021] (3);
[0022] Among them, [t] is the noise-free and reverberation-free original signal of the m-th sound source, n is the noise signal, is the room impulse response (RIR) between the m-th sound source and the microphone, is the reverberation delay time, is the signal received by the microphone; the signals of multiple sound sources are superimposed with the noise signal after reverberation in the indoor environment and are highly overlapped in the time domain. The purpose of speech enhancement is to restore the near-field target sound source without reverberation, noise, and other sound sources;
[0023] After the time-domain separation and noise reduction network, the reverberant speech signal in Equation 3 is stripped out:
[0024] (4);
[0025] Among them is the audio signal received by the m-th microphone; the time-domain model of Equation 4 is transformed into a frequency-domain model using the short-time Fourier transform;
[0026] (5);
[0027] is the frequency-domain form of the audio signal, the bin index is and the frame index is where K and N are the number of bins and the number of frames; is the frequency-domain form of the room impulse response (also known as the acoustic transfer function) between the sound source and the microphone, is the number of frames of the delay time. is the frequency-domain form of the near-field speech signal of the sound source;
[0028] The Weight prediction method (WPE) divides the right side of Equation (5) into a useful early reverberation signal and a late reverberation signal to be removed, and predicts the late reverberation signal by a multi-channel linear filter:
[0029] (6);
[0030] is the frequency-domain form of the audio signal recorded by the selected microphone, is a multi-channel linear filtering convolutional unit. Ignoring k and n, if the estimated value of the filter matrix is obtained then the estimated value of the demixed signal d can be obtained ;
[0031] (7);
[0032] Traditional WPE models the demixed signal d as a zero-mean circular Gaussian random variable distribution. Since each frame is independent of each other, the variance of each frame of can be obtained by calculating and further obtaining , which is also known as the power spectral density (PSD) of the target signal. The probability density function of d can be written as:
[0033] (8);
[0034] The bin corresponding to each frame follows distribution, Taking the maximum log-likelihood function of (8) can be transformed into:
[0035] (9);
[0036] WPE splits the demixing problem into the solution problems of g, d, and solves them step by step and iteratively. Starting from the result of the i-th iteration,
[0037] (10);
[0038] where: ;
[0039] Use Equation (7) to obtain , and simplify to obtain: (11);
[0040] Repeat the above process until the convergence criterion is met or the maximum number of iterations is exceeded;
[0041] To better limit the range of input values, a trained transpose-residual-temporal convolutional network is used to estimate , specifically, the logarithm of the energy of the short-time Fourier transform of the pre-prepared reverberant speech signal is input into the temporal convolutional network (TCN): (B is the batch, T is the time-domain frame, and F is the frequency-domain bin), which is consistent with the data format in the input bidirectional long short-term memory network (BLSTM); before inputting into the TCN, it needs to go through the transposition operation of swapping the time and frequency dimensions to become ; The TCN unit is a deep residual network consisting of a group of Q one-dimensional (1-D) modules for expanding the receptive field, where the expansion factor increases exponentially with a base of 2 to ensure a large enough receptive field. The TCN unit is repeated N times to form the overall TCN architecture; Conv1×1 is a pointwise convolution with a convolution kernel of 1 for changing the dimension, the parametric rectified linear unit (PreLU) is the activation function, layer normalization (LN) is the regularization operation, and the depthwise (DW) Conv1D is a one-dimensional convolution with an exponentially increasing expansion parameter; the network output (B×T×F) is transposed again to obtain the output in the B×T×F format, and then the exponential operation is performed to obtain , if a residual mechanism is added between the TCN input and output, that is, the original B×T×F is added to the transpose (B×T×F) of the B×F×T output of the TCN, and then enters the WPE iteration process to obtain the frequency-domain form of the demixed signal, calculate R and r, calculate g to obtain the demixed signal, and use the demixed signal for the new estimation, and perform iterative calculation again until the iterative condition is met, the number of iterations is exceeded, or the demixed signal is obtained after recursive calculation of g.
[0042] Furthermore, in the step S3, the separation network uses the anechoic reverberant speech signals and noise signals of the two speakers respectively as the training targets (res1, res2, noise) , and the network output is the estimated reverberant signal and noise signal;
[0043] During the training process, the utterance-level permutation invariant training (uPIT) method is used to select the maximum scale-invariant signal-to-noise ratio (sisdr) to accelerate the convergence of the network, and the training objective is to minimize the negative sisdr; the sisdr is defined as follows, is the target signal, is the estimated signal:
[0044] (14);
[0045] The demixing network uses the reverberant audio signals with random noise of the two speakers respectively as the input values, and the training targets are the anechoic early reverberation enhanced signals of the two speakers respectively. The training objective is to minimize the negative scale-invariant signal-to-noise ratio (sisdr) (demixing estimated value, demixing target value), and the estimated value is obtained by the TCN-WPE demixing network; considering that the early reverberation signal can be used to enhance the speech quality, the original pure s1 / s2 in the dataset (WHAMR!) is not used as the training target in the TCN-WPE, but the early reverberation enhanced signal (that is, the superposition value of the direct signal and the early reverberation signal) after convolving the original signal with a duration of at least 4 seconds and the early room impulse response (the cut-off time et of the early room impulse response signal is the sum of the direct signal arrival time dt and the delay time dlt, and the time length of the early room impulse response signal can be changed by adjusting the time length of the room impulse response); the superposition value of the direct signal and the early reverberation signal (also called the early reverberation enhanced signal) can unify the appearance time of the original signal and the reverberant signal, thereby ensuring the convergence of the TCN-WPE.
[0046] Furthermore, in the step S3, the set evaluation criteria include perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI), signal-to-interference ratio (SIR), scale-invariant signal-to-noise ratio (SI-SDR), and scale-invariant improvement value (SI-SDRI). Among them, the main evaluation criteria are SI-SDRI and SI-SDR. Higher values represent better effects. When calculating the metrics on the test set, the separated estimated reverberant signals 1 and 2 and the noise need to be sorted and adjusted first using the unordered permutation invariant training (uPIT) principle.
[0047] The structure of the present invention is reasonably designed and reliable. It provides a two-step hybrid sound source separation and dereverberation method for the actual sound field environment, cascades the separation network based on the Sepformer architecture and the dereverberation network based on the TCN architecture, and simultaneously solves the problems of single sound source signal separation and enhancement under the conditions of other sound sources, noise, and reverberation, enabling the automatic audio recognition and discrimination system to obtain a single sound source signal with higher quality and improving the robustness of the sound source separation and enhancement system.
[0048] By introducing the reverberant target signal and the noise signal as the training targets of the Sepformer separation network, the parameters of the Sepformer separation network are adjusted, and rapid optimization of the parameters of the separation network is achieved.
[0049] By introducing the residual mechanism and the transposed mechanism to adjust the parameter information of the input temporal convolutional network to solve the problem that the deep network has inaccurate signal energy estimation, and using the scale-invariant signal-to-noise ratio as the loss function of the dereverberation network to solve the problem that the deep network has poor convergence ability, with the early reverberation enhanced signal as the training target, the rapid convergence and good dereverberation ability of the temporal convolutional network are achieved.
[0050] By introducing noise signals with different signal-to-noise ratios, the problem that the separated and denoised signal still has actual noise interference is solved on the basis of simulating the actual dereverberation working scenario.
[0051] By introducing appropriate signal delays for the evaluation of separation and denoising and dereverberation metrics, the problem that far-field signals require more path propagation time than near-field signals is alleviated to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a flowchart of two-step hybrid sound source separation and dereverberation for the actual sound field environment.
[0053] Figure 2 It is a schematic diagram of the mask unit of the Sudo-rmrf network.
[0054] Figure 3 It is a schematic diagram of the mask unit of the Sepformer network in the present invention.
[0055] Figure 4 It is a schematic diagram of the TCN dereverberation network in the present invention. Detailed implementation manners
[0056] A two-step hybrid sound source separation and dereverberation method, which is implemented through the following steps:
[0057] Step S1: The sound source separation and noise reduction unit is as shown in the left part of the appendix. First, the encoder converts the waveform of the equal-length mixed audio signal collected by the original single microphone into corresponding data feature blocks; then the data feature blocks are input into the mask module to estimate the mask feature blocks of each channel; then the mask feature blocks of each channel are multiplied point by point with the waveform of the mixed audio signal to obtain the data feature blocks of each channel; finally, the data feature blocks are input into the decoder to obtain the reverberant audio signals of each speaker after separation while reducing the noise of the mixed audio signal. Figure 1 The specific method for separating and reducing the noise of the mixed audio signal is that the mixed audio signal is input into the encoder and the mask module in sequence, and finally the separation and noise reduction are completed through the decoder; for the encoder and the decoder, first the mixed audio signal is divided into N segments at a certain overlap rate, and each segment is sampled L times at a fixed frequency and sent into the encoder to be converted into feature w, and the decoder will estimate multiple features
[0058] and convert them into the separated output, and then use the overlap-add method to reconstruct each separated signal: (1);
[0059] (1);
[0060] wherein, is the non-linear unit (relu) of the encoder, is the basic matrix of the encoder, and x refers to the original signal after overlapping sampling;
[0061] (2);
[0062] wherein, is the basic matrix of the decoder ( ), is the feature estimated value of each separated signal;
[0063] For the mask module, the Separable Attention Network (Sepformer) adopted in the present invention is illustrated by taking the Successive Downsampling and Resampling Multi-Resolution Separable Network (Sudo-rmrf) as a reference.
[0064] Successive Downsampling and Resampling Multi-Resolution Separable Network (Sudo-rmrf): First, the latent representation output by the encoder is subjected to layer regularization and one-dimensional convolution operations to achieve the purpose of dimensionality reduction. ; Subsequently It is fed into M sequentially connected separable module units with consistent parameters. The structure of each unit is similar to that of a U-shaped network (U-Net). The downward part of the U-Net represents successive temporal downsampling operations, and the upward part represents successive complementary resampling operations, enabling the module to extract multi-temporal resolution information and achieve the purpose of expanding the receptive field within the limit of a relatively small number of parameters.
[0065] Specifically for one unit, First, it passes through , Prelu for channel expansion and non-linear processing, and then undergoes multiple downsamplings through , global layer regularization, . After the last sampling, it passes through a multi-head attention mechanism (MHA) unit, and then continuous upsampling is achieved through nearest neighbor interpolation , and finally, the output is completed through and layer regularization , where in the resampling operation, when the input has the same feature dimension as the downsampling output, residual addition is required.
[0066] In the Successive Downsampling and Resampling Multi-Resolution Separable Network (Sudo-rmrf), a multi-head attention mechanism (MHA) is added as a self-attention module at the bottommost part of the separable module; the feature vectors after successive downsampling pass through a linear unit, a position encoding unit, layer regularization, a multi-head attention unit, a global regularization unit successively and are output after using residual connection, serving as the input features for the bottommost resampling unit.
[0067] = +GloN (Multi-HeadAttention(GloN( ))) ;
[0068] represents a linear operation, represents a position encoding operation, GloN represents global regularization, Multi-HeadAttention represents an attention mechanism operation, represents the original input mixed signal, output features.
[0069] The Separation Attention Network (Sepformer) adopted in the present invention is as follows: taking the Sepformer network as the core part of the separation and noise reduction stage, performing layer regularization and linearization operations on the latent representation output by the encoder, and then performing a slicing operation with an overlapping factor of 50% to adjust the representation dimension; then sending it into the Sepfomer subunit, and each Sepfomer subunit contains one inter-Transformer and one intra-Transformer. The structural parameters of the two are the same and are respectively used to extract the correlation of the intra-block and inter-block time series. Each inter-Transformer and intra-Transformer contains layer regularization, multi-head attention mechanism (MHA), feed-forward network (FFW), and several residual connection mechanisms. Finally, the representation after repeating the Sepfomer network architecture several times is successively subjected to operations such as activation layer (PRelu), linearization, padding, and forward propagation to obtain the output mask.
[0070] Finally, the mixed audio signal is multiplied pointwise with the output mask, and each result obtained by the pointwise multiplication passes through the decoder to obtain the enhanced signal corresponding to each speaker after separation and noise reduction.
[0071] Step S2: Taking the separated audio signal as the input quantity, using the early reverberation enhanced signal as the target to be learned, and using the scale-invariant signal-to-noise ratio as the loss function. After converting the input quantity to the frequency domain, transposing and taking the logarithm, and then inputting it into the temporal convolutional network for dereverberation.
[0072] In step S2, the specific method for dereverberating the mixed audio signal is as follows: as shown in the right part of the appendix Figure 1 shown on the right side,
[0073] Model the acoustic signal of the indoor scene collected by a single microphone as follows:
[0074] (3);
[0075] Among them, [t] is the noise-free and reverberation-free original signal of the m-th sound source, n is the noise signal, is the room impulse response (RIR) between the m-th sound source and the microphone, is the reverberation delay time, is the signal received by the microphone; multiple source signals are reverberated in an indoor environment and then superimposed with the noise signal, presenting as highly overlapping in the time domain. The purpose of speech separation and enhancement is to restore the near-field target sound source without reverberation, noise, and other sound sources; after passing through the time-domain separation and noise reduction network in Equation 3, the reverberant speech signal is separated out:
[0076] (4);
[0077] where is the audio signal received by the m-th microphone, and the time-domain model in Equation 4 is transformed into a frequency-domain model using the short-time Fourier transform;
[0078] (5);
[0079] is the frequency-domain form of the audio signal, the bin index is and the frame index is , where K and N are the number of bins and frames. is the frequency-domain form of the room impulse response (also known as the acoustic transfer function) between the sound source and the microphone, is the number of frames of the delay time. is the frequency-domain form of the near-field speech signal of the sound source.
[0080] The Weight prediction method (WPE) divides the right side of Equation (5) into the useful early reverberation signal and the late reverberation signal to be removed, and predicts the late reverberation signal by a multi-channel linear filter:
[0081] (6);
[0082] is the frequency-domain form of the audio signal recorded by the selected microphone, is the multi-channel linear filtering convolutional unit. Ignoring k and n, if the estimated value of the filter matrix is obtained, then the estimated value of the demixed signal d can be obtained.
[0083] (7);
[0084] Traditional WPE models the demixed signal d as a zero-mean circular Gaussian random variable distribution. Since each frame is independent, the variance of each frame of d can be obtained, and further is obtained, Also known as the power spectral density (PSD) of the target signal; the probability density function of d can be written as:
[0085] (8);
[0086] The bins corresponding to each frame follow distribution, Taking the maximized log-likelihood function of (8) can be transformed into:
[0087] (9);
[0088] WPE splits the dereverberation problem into the solution problems of g, d, and solves them step by step and iteratively. Starting from the result of the i-th iteration,
[0089] (10);
[0090] where ;
[0091] Using equation (7) to obtain and simplifying to obtain: (11);
[0092] Repeat the above process until the convergence criterion is met or the maximum number of iterations is exceeded.
[0093] Although WPE and its improved versions have achieved good results in the field of speech dereverberation, at the beginning of the iteration, is unknown. Generally, is initialized as or approximated using and the mean of adjacent frames ( is the observed signal of a selected channel, also known as PSD).
[0094] (12);
[0095] (13);
[0096] The full prediction error method reduces approximation errors by averaging over multiple channels. However, in the case of a small number of channels or even a single channel, the estimation effect of filter g is significantly reduced. Therefore, a deep neural network can be used to learn and predict the PSD of the demixed signal. Convolutional neural networks (Cnn), long short-term memory networks (Lstm), and bidirectional long short-term memory networks (Blstm) can be used to predict the PSD. However, the phase information of the reverberant signal is not utilized. The phase information of the reverberant signal can also be added to the TCN using the Griffin-Lim’s algorithm, and the TCN can directly predict the demixed signal.
[0097] Based on this, the present invention proposes to introduce the temporal convolutional network (TCN) as a deep neural network into the PSD estimation of WPE, realizing the TCN-based WPE filter algorithm (TCN-WPE).
[0098] The present invention uses a trained transposed-residual-temporal convolutional network (TCN-inverse-residual) for estimation , specifically, the logarithm of the energy of the short-time Fourier transform of the pre-prepared reverberant speech signal is input into the temporal convolutional network (Temporal convolutional network, TCN): (B is the batch, T is the time domain frame, and F is the frequency domain bin), which is consistent with the data format in the input bidirectional long short-term memory network (Bidirectional long short-term memory, BLSTM); before inputting into the TCN, it needs to go through a transpose operation of swapping the time and frequency dimensions to become ; the TCN unit is a deep residual network consisting of a group of Q one-dimensional (1-Dimension, 1-D) modules for expanding the receptive field, where the expansion factor increases exponentially with a base of 2 to ensure a large enough receptive field. The TCN unit is repeated N times to form the overall TCN architecture; Conv1×1 is a pointwise convolution with a kernel of 1 for changing the dimension, the parametric rectified linear unit (Parametric rectified linear unit, PreLU) is the activation function, layer normalization (Layer Normalization, LN) is the regularization operation, and the depthwise (DW) Conv1D is a one-dimensional convolution with an exponentially increasing expansion parameter; the network output (B×F×T) is transposed again to obtain an output in the B×T×F format, and then an exponential operation is performed to obtain , as Figure 4 shown, if a residual mechanism is added between the TCN input and output, that is, the original B×T×F is added to the transpose (B×T×F) of the TCN output B×F×T.
[0099] After entering the WPE iteration stage, the frequency-domain form of the demixed signal is obtained, the calculations of R and r are performed, and after calculating g, the demixed signal is obtained and used for new estimation. Iterative calculations are performed again until the iterative conditions are met, the number of iterations is exceeded, or the demixed signal is obtained after recursive calculation of g.
[0100] Step S3: Cascade the separation and noise reduction network in Step S1 and the demixing network in Step S2. The separation and noise reduction network is a separation attention network (res1 / res2 / noise are the targets to be learned during training), and the demixing network is a time convolution - transpose - residual - WPE network (TCN - Inverse - Resisual - WPE). The mixed audio signal first enters the separation and noise reduction network to output the reverberant audio signals. Each audio signal has been separated, and the noise has been basically removed. Then, the two obtained reverberant audio signals are input into the same dereverberation network to complete the reverberation removal work under strong noise or weak noise conditions.
[0101] In Step S3, the separation network uses the noise - free reverberant speech signals and noise signals of the two speakers respectively as the training targets (res1, res2, noise) , and the network output is the estimated reverberant signal and noise signal.
[0102] During the training process, the speech sorting invariance training method (Utterance - level permutation invariant training, uPIT) is used to select the maximum scale - invariant signal - to - noise ratio (scale - invariant source - to - noise ratio, sisdr) to accelerate network convergence. The training objective is to minimize the negative sisdr; the sisdr is defined as follows, is the target signal, is the estimated signal:
[0103] (14);
[0104] The dereverberation network takes the reverberant audio signals with random noise of two speakers as input values, and the training target is the noise-free early reverberation enhanced signals of the two speakers respectively. The training objective is to minimize the negative scale-invariant signal-to-distortion ratio (sisdr) (dereverberation estimate value, dereverberation target value). The estimate value is obtained by the TCN-WPE dereverberation network. Considering that the early reverberation signal can be used to enhance the speech quality, the original clean s1 / s2 in the dataset (WHAMR!) is not used as the training target in TCN-WPE. Instead, the early reverberation enhanced signal (i.e., the superposition value of the direct signal and the early reverberation signal) after convolving the original signal with a duration of greater than or equal to 4 seconds and the early room impulse response (the cut-off time et of the early room impulse response signal is the superposition of the direct signal arrival time dt and the delay time dlt, and the time length of the early room impulse response signal can be changed by adjusting the time length of the room impulse response) is used. Using the superposition value of the direct signal and the early reverberation signal (also known as the early reverberation enhanced signal) can unify the appearance time of the original signal and the reverberation signal, thereby ensuring the convergence of TCN-WPE. The original deep neural network-weight prediction error (DNN-WPE) only calculates the loss function using the amplitude spectrum information of the estimated signal and processes the original reverberation phase information. However, the present invention goes further, restores the time-domain form of the estimated signal, uses sisdr as the loss function, and fully utilizes the phase information during the network backpropagation process.
[0105] In step S3, the set evaluation criteria include perceptual evaluation of speech quality (pesq), short-time objective intelligibility (stoi), signal-to-interference ratio (sir), scale-invariant signal-to-distortion ratio (sisdr), and scale-invariant improvement value (sisdri). Among them, the main evaluation criteria are sisdri and sisdr, and higher values represent better effects. When calculating the metrics on the test set, the separated estimated reverberant signals 1 and 2, and the noise need to be sorted and adjusted first using the unordered permutation invariant training (uPIT) principle.
[0106] Figure 1 It is a flowchart of a two-step hybrid sound source separation and dereverberation method for an actual sound field environment. The left part is the hybrid sound source separation and noise reduction part, which takes the mixed audio signal as input and outputs the reverberant audio signal and the noise signal. The right part is the audio signal dereverberation part, which takes the reverberant audio signal as input and outputs the anechoic audio signal.
[0107] Figure 3It is a schematic diagram of the Sepformer network mask module, which can realize the mask output of each signal to be estimated. To prove the effectiveness of the separation effect, the separation algorithms used as comparison benchmarks are: successive downsampling and resampling multi-resolution separation network (sudo-rmrf), successive downsampling and resampling multi-resolution separation (sudo-rmrf) + attention mechanism network (the network structure is as Figure 2 shown), and the training objectives used for comparison are the original signals s1 / s2, the original signals s1 / s2 / noise, and the reverberant signals res1 / res2. The training objectives used in the present invention are res1 / res2 / noise.
[0108] Figure 4 It is a schematic diagram of the temporal convolutional network (TCN). During the training process, the reverberant signal first undergoes a short-time Fourier transform, then the modulus, square, and logarithm of the signal are taken, and after being fed into the TCN unit, the predicted frequency-domain signal is output. After exponential transformation, weight prediction error, and inverse short-time Fourier transform, the scale-invariant signal-to-noise ratio is calculated with the original anechoic signal. To prove the effectiveness of the dereverberation effect, a bidirectional long short-term network is used as the benchmark dereverberation network unit, and the mean square error maximization method is used as the loss function.
[0109] To verify the effectiveness of the present method, it is illustrated through the following comparative experiments.
[0110] The data used in the experiment is the WHAMR! dataset, which is used for speech signal separation and noise reduction training. The WHAMR! dataset is a mixed speech version with noise and reverberation made according to the wsj0 dataset. Its dataset includes a training set, a validation set, and a test set, each containing 20,000, 5,000, and 3,000 data respectively. Among them, the signal-to-noise ratio of the speech signal and the noise signal randomly appears between -6 dB and +3 dB, and the relevant parameters of the room impulse response, such as the length, width, and height of the room, randomly appear in (5 m - 10 m, 5 m - 10 m, 3 m - 4 m), randomly appears in (0.1 s - 1 s) under different reverberation conditions, the center position of the microphone randomly appears in (7.5 m ± 0.2 m, 7.5 m ± 0.2 m, 0.9 m - 1.8 m), and the source position randomly appears in (height: 0.9 m - 1.8 m, distance: 0.66 m - 2 m, :0 - 2 ) appears.
[0111] The present invention uses a self-made wsj0-RIRS dataset for de-reverberation network training. Data greater than 4 seconds (including 17075) in the Wall Street Journal LDC93S6A (wsj0-2mix) training set are selected and convolved with the REVERB challenge data single-channel room impulse response signal (RIRS) to obtain a noise-free wsj0-RIRS reverberation dataset. RIRS simulates a total of 6 reverberation conditions, with the reverberation time (T60) of small, medium and large rooms being 0.25s, 0.5s, and 0.75s, respectively, and the distance between the microphone and the speaker being 0.5m (near field) and 2m (far field), respectively. The present invention randomly takes 16675 data from the noise-free wsj0-RIRS reverberation dataset and two groups of random noise (signal-to-noise ratios of 15dB / 2dB, respectively) to generate a noisy wsj0-RIRS reverberation dataset, and the remaining 400 data are used to generate a test set.
[0112] By comparing multiple separation and denoising networks, it is proved that Sepformer+res1 / res2 / noise has the best separation effect. By comparing multiple de-reverberation networks, the time convolution+transposition network achieves the best de-reverberation effect under strong noise conditions, and the time convolution+transposition+residual network achieves the best de-reverberation effect under weak noise conditions. In a real sound field environment, due to the presence of other speakers and noise, the spectrogram of the mixed audio signal is difficult to distinguish. After Sepformer+res1 / res2 / noise and the time convolution network+transposition+residual cascade network, the best separation denoising and de-reverberation effects are achieved. The two-step aliasing sound source separation and de-altering method proposed in the present invention has the ability to restore the original pure signal most closely compared to the one-step direct separation and de-altering method and the cascade network of Sepformer+res1 / res2 / noise and weighted prediction error method and bidirectional long short-term memory network. This reflects the advantage of using the improved time convolutional network in estimating the signal energy spectrum. While maintaining a faster calculation speed, it can significantly improve the speech quality of the signals to be separated and enhanced. The above shows that the proposed two-step mixed sound source separation and de-altering method can achieve high robustness enhancement of speech quality.
[0113] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A two-step hybrid sound source separation and dereverberation method, characterized in that: Including the following steps: Step S1: First, the encoding module converts the waveform of the original mixed audio signal collected by the original single microphone into corresponding data feature blocks; then the data feature blocks are input into the masking module to estimate the masking feature blocks of each channel; then the masking feature blocks of each channel are multiplied pointwise with the mixed audio signal to obtain the data feature blocks of each channel; the data feature blocks of each channel are input into the decoding module to obtain the separated reverberant audio signal while reducing the noise of the mixed audio signal; Step S2: Taking the separated audio signal as the input quantity, using the early reverberation enhancement signal as the target to be learned, and using the scale-invariant signal-to-noise ratio as the loss function, after converting the input quantity to the frequency domain, transposing and taking the logarithm, and then inputting it into the temporal convolutional network for dereverberation; Step S3: Cascade the separation and noise reduction network in Step S1 and the dereverberation network in Step S2. The separation and noise reduction network is a separation attention network (res1 / res2 / noise are the targets to be learned during training), and the dereverberation network is a temporal convolutional - transposed - residual - WPE network (TCN - Inverse - Residual - WPE). The mixed audio signal first enters the separation and noise reduction network to output the reverberant audio signal, where each audio signal has been separated and the noise has been basically removed. Then the two obtained reverberant audio signals are input into the same dereverberation network to complete the reverberation removal work under strong or weak noise conditions.
2. The two-step hybrid sound source separation and dereverberation method according to claim 1, characterized in that: In the step S1, the specific method for separating and denoising the mixed audio signal is as follows: the mixed audio signal is sequentially input into an encoding module and a masking module, and finally the separation and denoising are completed through a decoding module; for the encoding module and the decoding module, first, the mixed audio signal is divided into N segments at a certain overlapping rate, and each segment is sampled L times at a fixed frequency and sent into the encoding module to be converted into features w, and the decoding module converts the estimated multiple features into the separated output, and then the overlapping addition method is used to reconstruct each separated signal; (1); Among them, is the non-linear unit (relu) of the encoder, is the basic matrix of the encoder, and x refers to the original signal after overlapping sampling; (2); Among them, is the decoder basis matrix, is the feature estimate value of each separated signal; For the masking module, taking the Sepformer network as the core part of the separation and noise reduction stage, performing layer regularization and linearization operations on the latent representation output by the encoding module, then performing a slicing operation with an overlap factor of 50% to adjust the representation dimension, and then sending it into the Sepfomer subunit. Each Sepfomer subunit contains one inter - Transformer and one intra - Transformer. The structural parameters of the two are the same and are respectively used to extract the correlation of the intra - block and inter - block time series. Each inter - Transformer and intra - Transformer contains layer regularization, multi - head attention mechanism (Multi - headattention, MHA), feed - forward network (Feed - forward network, FFW) and several residual connection mechanisms. Finally, the representation after repeating the Sepfomer network architecture several times is successively subjected to activation layer (PRelu), linearization, padding and forward propagation operations to obtain the output mask; Finally, the mixed audio signal is multiplied pointwise with the output mask, and each result obtained by the pointwise multiplication passes through the decoding module to obtain the enhanced signals of the respective speakers after separation and noise reduction.
3. The two-step hybrid sound source separation and dereverberation method according to claim 2, characterized in that: In the said Step S2, the specific method for dereverberating the mixed audio signal is as follows: Model the acoustic signal of the indoor scene collected by a single microphone as follows: (3); Among them, [t] is the original signal without noise and reverberation of the m-th sound source, and n is the noise signal, is the room impulse response (RIR) between the m-th sound source and the microphone, is the reverberation delay time, is the signal received by the microphone; the signals of multiple sound sources are superimposed with the noise signal after reverberation in the indoor environment, showing a high degree of overlap in the time domain, and the purpose of speech enhancement is to restore the near-field target sound source without reverberation, noise and other sound sources; The reverberant speech signal is separated from Equation 3 by the time-domain separation and noise reduction network: (4); wherein is the audio signal received by the m-th microphone; the time-domain model of Equation 4 is transformed into a frequency-domain model using the short-time Fourier transform; (5); is the frequency domain form of the audio signal, with the bin index being , and the frame index being , where K and N are the number of bins and frames respectively; is the frequency domain form of the room impulse response (also known as the acoustic transfer function) between the sound source and the microphone, is the number of frames of the delay time, is the frequency domain form of the near-field speech signal of the sound source; The Weight prediction method (WPE) divides the right side of Equation (5) into the useful early reverberation signal and the late reverberation signal to be removed, and predicts the late reverberation signal by a multi-channel linear filter: (6); is the frequency-domain form of the audio signal recorded by the selected microphone, is a multi-channel linear filtering convolution unit. Ignoring k and n, if the estimated value of the filter matrix is obtained , then the estimated value of the demixed signal d can be obtained ; (7); Traditional WPE models the demixed signal d as a zero-mean circular Gaussian random variable distribution. Since each frame is independent of each other, by calculating the variance of each frame we can further obtain , which is also known as the power spectral density (PSD) of the target signal. The probability density function of d can be written as: (8); The bins corresponding to each frame follow a distribution , maximizing the log-likelihood function for (8) can be transformed into: (9); WPE splits the demixing problem into the problems of solving for g and d, and solves them step by step and iteratively. Starting from the result of the i-th iteration, (10); Wherein: ; Obtained by using Equation (7) , and simplified to obtain: (11); Repeat the above process until the convergence criterion is met or the maximum number of iterations is exceeded; To better limit the range of input values, a trained transposed-residual-temporal convolutional network is used for estimation , specifically, the logarithm of the energy of the short-time Fourier transform of the pre-prepared reverberant speech signal is input into the Temporal convolutional network (TCN): (B is the batch, T is the time-domain frame, and F is the frequency-domain bin), which is consistent with the data format in the input Bidirectional long short-term memory (BLSTM); before inputting into the TCN, it needs to be transposed in the time and frequency dimensions, that is, the transpose operation is performed to become ; the TCN unit is a deep residual network consisting of a group of Q one-dimensional (1-Dimension, 1-D) modules for expanding the receptive field, where the expansion factor increases in the form of an exponent of 2 to ensure a large enough receptive field. The TCN unit is repeated N times to form the overall TCN architecture; Conv1×1 is a pointwise convolution with a convolution kernel of 1 for changing dimensions, the Parametric rectified linear unit (PreLU) is the activation function, LayerNormalization (LN) is the regularization operation, and the depthwise (DW) Conv1D is a one-dimensional convolution with an exponentially increasing expansion parameter; the network output (B×T×F) is transposed again to obtain an output in the B×T×F format, and then an exponential operation is performed to obtain , if a residual mechanism is added between the TCN input and output, that is, the original B×T×F is added to the transpose (B×T×F) of the TCN output of B×F×T, and then it enters the WPE iteration process to obtain the frequency-domain form of the dereverberated signal, calculate R and r, calculate g to obtain the dereverberated signal, and use the dereverberated signal for the new estimation, and perform iterative calculations again until the iteration condition is met, the iteration times are exceeded, or the dereverberated signal is obtained after recursive calculation of g.
4. A two-step hybrid sound source separation and dereverberation method according to claim 3, characterized in that: In the step S3, the separation network uses the anechoic reverberant speech signals and noise signals of the two speakers respectively as training targets (res1, res2, noise) , and the network outputs to estimate the reverberant signal and the noise signal; During the training process, the Utterance-level permutation invariant training (uPIT) method is used to select the maximum scale-invariant signal-to-noise ratio (sisdr) to accelerate the convergence of the network. The training objective is to minimize the negative sisdr. The sisdr is defined as follows: is the target signal, is the estimated signal: (14); The demixing network takes the reverberant audio signals with random noise of the two speakers as input values, and the training target is the noise-free early reverberation enhanced signals of the two speakers respectively. The training purpose is to minimize the negative scale-invariant signal-to-noise ratio (sisdr) (demixing estimated value, demixing target value), and the estimated value is obtained by the TCN-WPE demixing network; Considering that the early reverberation signal can be used to enhance the speech quality, the original clean s1 / s2 in the dataset (WHAMR!) is not used as the training target in TCN-WPE. Instead, the early reverberation enhanced signal (that is, the superposition value of the direct signal and the early reverberation signal) after convolving the original signal with a duration of at least 4 seconds and the early room impulse response (the cut-off time et of the early room impulse response signal is the sum of the direct signal arrival time dt and the delay time dlt. The time length of the early room impulse response signal can be changed by adjusting the time length of the room impulse response); The superposition value of the direct signal and the early reverberation signal (also called the early reverberation enhanced signal) can unify the appearance time of the original signal and the reverberant signal, thus ensuring the convergence of TCN-WPE.
5. A two-step hybrid sound source separation and dereverberation method according to claim 4, characterized in that: In step S3, the set evaluation criteria include perceptual evaluation of speech quality (pesq), short-time objective intelligibility (stoi), signal-to-interference ratio (sir), scale-invariant signal-to-noise ratio (sisdr), and scale-invariant improvement value (sisdri). Among them, the main evaluation criteria are sisdri and sisdr. Higher values represent better effects. When calculating the metrics on the test set, the separated estimated reverberant signals 1 and 2, and the noise need to be sorted and adjusted first using the unordered permutation invariant training (uPIT) principle.
Citation Information
Cited By
Chinese opera sound source separation method based on domain adaptation and time domain information guidance
CN121148415A
Atrial arrhythmia identification method and system based on space-time diagram neural network and atrial signal source reconstruction, and storage medium
CN121795923A