Filtering and Summing Multi-Channel Speech Separation Method Based on Simplified Attention Encoder-Decoder Network
By simplifying the attention codec network calculation filter parameters, the speech separation problem in low signal-to-noise ratio and high reverberation environments is solved, and the efficient speech separation effect is achieved in different environments.
Patent Information
- Application Number
- CN202211515165.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-29
AI Technical Summary
In a low signal-to-noise ratio and high reverberation environment, the performance of existing multi-channel speech separation algorithms has significantly decreased, making it difficult to effectively separate the target speech.
The simplified attention codec network is used as the timing modeling network for filter summing structure, and the filter parameters are calculated through attention mechanism and position encoding, and combined with the multi-head attention mechanism to pay attention to information at different time steps and resolutions, improving speech separation performance.
The performance of mainstream algorithms is close to that of high signal-to-noise ratio and low reverb environments, and performs better than mainstream algorithms in low signal-to-noise ratio and high reverb environments, improving the effect of speech separation.
Smart Images

Figure CN115910092B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of speech signal separation, and particularly relates to a filter-sum multi-channel speech separation method based on a simplified attention encoder-decoder network. Background Art
[0002] As early as 1953, Cherry proposed the "cocktail party problem", which refers to separating the speech of a target speaker under conditions of noise, reverberation, and multiple speakers. Multi-channel speech separation refers to two or more speech acquisition devices that simultaneously use speech spectrum information and the spatial information of speakers to separate the speech of multiple speakers. In recent years, significant progress has been made in speech separation based on deep learning. Array-based speech separation algorithms use network structures such as the Temporal Convolutional Network (TCN) and the Dual-Path Recurrent Neural Network (DPRNN). However, in an environment with low signal-to-noise ratio and high reverberation, the speech separation performance of the above neural network methods drops significantly. Summary of the Invention
[0003] The purpose of the present invention is to provide a filter-sum multi-channel speech separation method based on a simplified attention encoder-decoder network. By using the simplified attention encoder-decoder network as the temporal modeling network of the filter-sum structure, it can comprehensively integrate information from different time steps through the attention mechanism, and can also use positional encoding and multi-head attention to enable different query vectors obtained by encoding to focus on the comprehensive information of different time steps and different resolutions, so as to better calculate the filter parameters. In an environment with a relatively high signal-to-noise ratio and low reverberation, its performance is very close to that of mainstream algorithms. In an environment with low signal-to-noise ratio and high reverberation, the performance of the algorithm of the present invention is superior to that of mainstream algorithms, so as to solve the problem of designing spatial filters in speech separation algorithms in an environment with low signal-to-noise ratio and high reverberation.
[0004] To solve the above technical problems, the specific technical solution of the present invention is as follows:
[0005] A filter-sum multi-channel speech separation method based on a simplified attention encoder-decoder network, the method comprising the following steps:
[0006] Step 1: Frame a multi-channel speech signal containing multiple sound sources to obtain frame-level speech signals of each channel. Select the speech signal of one channel as the reference channel speech signal, calculate the normalized cross-correlation feature with the speech signals of the remaining channels, calculate the embedding feature of the reference channel speech signal, splice the normalized cross-correlation feature and the embedding feature, and use the spliced parameters as the input feature of the first simplified attention encoder-decoder network to output the filter parameters for the reference channel speech signal;
[0007] Step 2: Filter the reference channel speech signal by using the filter parameters output by the simplified attention encoder-decoder network in Step 1 to obtain the pre-separated speech signals of each sound source;
[0008] Step 3: Calculate the normalized cross-correlation features between the pre-separated speech signals of each sound source in Step 2 and the speech signals of the remaining channels, calculate the embedding features of the speech signals of the remaining channels, splice the normalized cross-correlation features and the embedding features, and use the spliced parameters as the input of the second simplified attention encoder-decoder network to output the filter parameters for the speech signals of the remaining channels;
[0009] Step 4: Filter the speech signals of the corresponding channels by using the filter parameters of the remaining channels obtained in Step 3 to obtain the speech signals of each sound source separated from the speech signals of the remaining channels, and add the pre-separated speech signals of each sound source to the speech signals of each sound source separated from the remaining channels to obtain the final separated speech of each sound source.
[0010] Further, Step 1 specifically includes the following steps: Select a reference channel, set its number as 1, and first calculate the normalized cross-correlation values between the speech signal of the reference channel and the speech signals of the remaining channels:
[0011]
[0012] where is the signal of the t-th frame received by the reference channel, and the data of the previous and subsequent frames is introduced, that is, the signal with a length of 3L composed of samples from tH - L to tH + 2L - 1, and H represents the frame shift during frame division; represents the signal with a length of L taken from the reference channel signal with a length of 3L where j represents the starting serial number of the sample point, represents a real vector with a dimension of 1×L; represents the signal of the t-th frame received by the m-th channel, || || represents the vector norm, and <> represents the inner product calculation; represents the normalized cross-correlation value between the signal of the t-th frame of the reference channel and the signal of the m-th channel, and is connected according to the j serial number to obtain a 2L + 1-dimensional normalized cross-correlation function
[0013] Average the normalized cross-correlation function for different channels to obtain the normalized cross-correlation feature NCC of the t-th frame t :
[0014]
[0015] where is the normalized cross-correlation feature of the t-th frame;
[0016] The calculation formula for the embedding feature is as follows:
[0017]
[0018] where is the signal of the t-th frame received by the reference channel; represents the weight matrix for calculating the embedding, which is a learnable parameter, K u represents the dimension of the embedding vector;
[0019] These two types of features are concatenated as the feature parameters of the speech signal of this frame, and the feature parameters are input into the first simplified attention encoder-decoder network to obtain the filter parameters for the speech signal of the reference channel of this frame;
[0020] The filter parameter corresponding to the t-th frame and the k-th sound source has the following calculation formula:
[0021]
[0022] where represents the first simplified attention encoder-decoder network, and for the feature with a dimension of 2L + 1 + K input at each moment, after the temporal convolution and dimension transformation therein, K vectors P with a dimension of K U are output U of the vector P t 1,1 , P t 1,2 ,..., P t 1 ,k ,..., P t 1,K ; K represents the number of sound sources; ⊙ represents the element-wise multiplication of two vectors or matrices with the same dimension; is the coefficient of the simplified attention encoder-decoder network ; is the bias; tanh() is the hyperbolic tangent operation, and σ() is the Sigmoid function.
[0023] Furthermore, the calculation formula for the pre-separated speech signal of each sound source obtained in step 2 is:
[0024]
[0025] where represents the pre-separated signal of the t-th frame and the k-th sound source filtered from the reference channel signal.
[0026] Furthermore, step 3 specifically includes the following steps: Calculate the normalized cross-correlation value between the speech signals of the remaining channels and the pre-separated speech signals of each sound source:
[0027]
[0028] Wherein is the t-th frame signal received by the m-th channel, and the data of the previous and next frames is introduced, that is, the samples from tH-L to tH+2L-1 form a signal with a length of 3L; denotes a vector with a length of L taken from denotes the pre-separated signal of the t-th frame and the k-th sound source and the m-th channel signal The normalized cross-correlation value, which is connected according to the j sequence number to obtain a 2L+1-dimensional normalized cross-correlation function
[0029] Taking the normalized cross-correlation function directly as the normalized cross-correlation feature of the remaining channels;
[0030] The embedding feature of the remaining channel speech signals is calculated as:
[0031]
[0032] Wherein denotes the embedding feature of the m-th channel and the t-th frame;
[0033] Concatenating the two types of features as the input of the second simplified attention encoder-decoder network to obtain the filter parameters for the remaining channel speech signals, denotes the filter parameter corresponding to the t-th frame and the k-th sound source for the m-th channel:
[0034]
[0035] Wherein denotes the second simplified attention encoder-decoder network.
[0036] Furthermore, the calculation formula for the final separated speech of each sound source obtained in step 4 is:
[0037] Wherein denotes the separated speech of the k-th sound source and the t-th frame.
[0038] Furthermore, the simplified attention encoder-decoder network is composed of a bidirectional long short-term memory network, an encoding layer and a decoding layer containing an attention mechanism. The simplified attention encoder-decoder network uses the bidirectional long short-term memory network to implement position encoding, and the position encoding is used to integrate the information of the entire input feature sequence and provide the position relationship of different feature vectors for the attention structure.
[0039] Furthermore, the encoding layer and decoding layer of the simplified attention encoder-decoder network are composed of layer normalization, multi-head attention structure, fully connected layer, and Gaussian error linear unit.
[0040] Furthermore, the multi-head attention structure of the simplified attention encoder-decoder network adopts a narrow multi-head form.
[0041] Furthermore, the two simplified attention encoder-decoder networks in Step 1 and Step 3 have the same structure but different network parameters.
[0042] The filtering and summing multi-channel speech separation method based on the simplified attention encoder-decoder network of the present invention has the following advantages: The simplified attention encoder-decoder network consists of a bidirectional long short-term network, an encoding layer and a decoding layer containing an attention mechanism. This network can directly synthesize information from different time steps from a global perspective through the attention mechanism, and can also use positional encoding and multi-head attention to make different query vectors obtained by encoding focus on the comprehensive information of different time steps and different resolutions, which can better calculate the filter parameters. The present invention improves the effective extraction of spatial information and spectral information in the array signal through the attention mechanism, improves the calculation performance of the spatial filter parameters, and thus improves the array speech separation performance.
[0043] The algorithm proposed by the present invention is very close to the performance of mainstream algorithms in a high signal-to-noise ratio and low reverberation environment, while in a low signal-to-noise ratio and high reverberation environment, the performance of the algorithm of the present invention is superior to that of mainstream algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is the structural diagram of the speech separation algorithm of the present invention;
[0045] Figure 2 It is the structural diagram of the simplified attention encoder-decoder network of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0046] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a filtering and summing multi-channel speech separation method based on a simplified attention encoder-decoder network of the present invention with reference to the accompanying drawings.
[0047] Step 1: Frame the multi-channel speech signal containing multiple sound sources to obtain the frame-level speech signals of each channel. Select the speech signal of one channel as the reference channel speech signal, calculate the normalized cross-correlation feature with the speech signals of the remaining channels, calculate the embedding feature of the reference channel speech signal, splice the normalized cross-correlation feature with the normalized cross-correlation feature, and use the spliced parameters as the input feature of the first simplified attention encoder-decoder network to output the filter parameters for the reference channel speech signal.
[0048] Select a reference channel, with its number set to 1. First, calculate the normalized cross-correlation values between the speech signal of the reference channel and the speech signals of the other channels:
[0049]
[0050] where is the signal of the t-th frame received by the reference channel, and the data of the previous and next frames is introduced, that is, the signal with a length of 3L composed of samples from tH-L to tH+2L-1, where H represents the frame shift during frame segmentation; represents the signal with a length of L taken from the reference channel signal of length 3L where j represents the starting serial number of the sample point, represents a real vector with a dimension of 1×L; represents the signal of the t-th frame received by the m-th channel, || || represents the vector norm, and 〈 〉 represents the inner product calculation; represents the normalized cross-correlation value between the signal of the t-th frame, the reference channel signal and the signal of the m-th channel, and is connected according to the j serial number to obtain a 2L+1-dimensional normalized cross-correlation function
[0051] Average the normalized cross-correlation function for different channels to obtain the normalized cross-correlation feature NCC of the t-th frame t :
[0052]
[0053] where is the normalized cross-correlation feature of the t-th frame.
[0054] The embedding feature mainly reflects the short-term stationarity of the speech signal and reduces the dimension of the original one-frame reference channel signal. Its calculation formula is:
[0055]
[0056] where is the signal of the t-th frame received by the reference channel; represents the weight matrix for calculating the embedding, which is a learnable parameter, and K u represents the dimension of the embedding vector.
[0057] Concatenate these two types of features as the feature parameters of this frame of speech signal, and input the feature parameters into the first simplified attention encoder-decoder network to obtain the filter parameters for this frame of reference channel speech signal.
[0058] The structure of the simplified attention encoder-decoder network is as Figure 2As shown in the figure. The simplified attention encoder-decoder network consists of a bidirectional long short-term memory network, an encoding layer and a decoding layer that contain an attention mechanism. The bidirectional long short-term memory network is used for position encoding, so that the position encoding at each time step can integrate the information of the entire time series, and can enable the attention module to recognize the position relationship of different feature vectors. The encoding layer and the decoding layer are composed of layer normalization, multi-head attention structure, fully connected layer and Gaussian error linear unit.
[0059] The multi-head adopts the form of narrow multi-head, that is: the input feature matrix is split into multiple vectors of equal length along the feature dimension, and then respectively input into the attention modules with different network parameters. Finally, when outputting, the outputs of different modules are concatenated along the feature dimension. In this way, each attention module will focus on different dimensions of the feature vector, which can avoid some dimensions being too prominent and ignoring the information of other dimensions. Moreover, the narrow multi-head can also reduce the calculation scale of matrix multiplication. In addition, there is a fully connected layer and a dropout operation in the output part of the multi-head attention module for feature mapping, adjusting the output vector, and residual connection can be performed.
[0060] The filter parameters corresponding to the k-th sound source at the t-th frame The calculation formula is:
[0061]
[0062] where represents the first simplified attention encoder-decoder network. For the features with a dimension of 2L + 1 + K input at each moment, after the temporal convolution and dimension transformation therein, K vectors P with a dimension of K U are output U ; K represents the number of sound sources; ⊙ represents element-wise multiplication of two vectors or matrices with the same dimension; t 1,1 , P t 1,2 ,..., P t 1 ,k ,..., P t 1,K ; is the coefficient of the simplified attention encoder-decoder network ; is the bias; tanh() is the hyperbolic tangent operation, and σ() is the Sigmoid function. ;
[0063] Step 2: Use the filter output by the simplified attention encoder-decoder network in Step 1 to filter the reference channel speech signal to obtain the pre-separated speech signals of each sound source. The calculation formula is:
[0064]
[0065] wherein represents the pre-separated signal of the k-th sound source in the t-th frame filtered from the reference channel signal.
[0066] Step 3: Calculate the normalized cross-correlation features of the pre-separated speech signals of each sound source in Step 2 and the speech signals of the remaining channels, calculate the embedding features of the speech signals of the remaining channels, splice the normalized cross-correlation features with the normalized cross-correlation features, and use the spliced parameters as the input of the second simplified attention encoder-decoder network to output the filter parameters for the speech signals of the remaining channels.
[0067] The remaining channel speech signals calculate the normalized cross-correlation values with the pre-separated speech signals of each sound source:
[0068]
[0069] wherein is the signal of the t-th frame received by the m-th channel, and the data of the previous and subsequent frames is introduced, that is, the samples from tH-L to tH+2L-1 form a signal with a length of 3L; represents from a vector with a length of L taken out, represents the pre-separated signal of the k-th sound source in the t-th frame and the m-th channel signal the normalized cross-correlation value, which is connected in sequence according to the j sequence number to obtain a 2L+1-dimensional normalized cross-correlation function
[0070] The normalized cross-correlation function is directly used as the normalized cross-correlation feature of the remaining channels.
[0071] The embedding features of the remaining channel speech signals are calculated as:
[0072]
[0073] wherein represents the embedding feature of the m-th channel and the t-th frame.
[0074] Splice the two types of features as the input of the second simplified attention encoder-decoder network to obtain the filter parameters for the speech signals of the remaining channels, represents the filter parameter corresponding to the k-th sound source in the t-th frame for the m-th channel:
[0075]
[0076] wherein It represents the second simplified attention encoder-decoder network, which is different from H1 in that for the features with a dimension of 2L + 1 + K input at each moment u after temporal convolution and dimensional transformation, a vector with a dimension of K u is output; is the coefficient of the simplified attention encoder-decoder network and is the bias.
[0077] In terms of network structure K vectors are output, that is, for K sound sources, there are K different sets of filter parameters for separating K target sound sources, and different sets of filter parameters for different channels and different sound sources are obtained according to different input feature parameters.
[0078] Step 4: Filter the speech signals of the corresponding channels by using the filters of the remaining channels obtained in Step 3 to obtain the speech signals of each sound source separated from the speech signals of the remaining channels, and add the pre-separated speech signals of each sound source to the speech signals of each sound source separated from the remaining channels to obtain the final separated speech of each sound source.
[0079]
[0080] where represents the separated speech of the k-th sound source and the t-th frame.
[0081] During the training process, the scale-invariant signal-to-noise ratio SI-SNRi (Scale Invariant Signal to Noise Ratio improvement) between the separated speech and the clean speech is used as the training objective. To solve the problem of the permutation of the sound source numbers of the separated speech output by the network and the sound source numbers of the clean speech, the permutation invariant training PIT (Permutation Invariant Training) structure is used here to calculate the loss function.
[0082] The present invention uses SI-SNRi and short-time objective intelligibility STOI (Short-Time Objective Intelligibility) as evaluation indicators. The multi-channel speech separation method based on the simplified attention encoder-decoder SABED (Simplified Attention Between Encoder-Decoder) network of the present invention is compared with the multi-channel speech separation algorithms that use the temporal convolutional network TCN and the dual-path recurrent neural network DPRNN as the temporal modeling network for filtering and summing. The performance evaluation is shown in Tables 1 and 2.
[0083] Table 1 Comparison of SI-SNRi values of different algorithms in multiple environments
[0084]
[0085] Table 2 STOI Comparison of Different Algorithms in Multiple Environments
[0086]
[0087] The signal-to-noise ratio SNR (Signal Noise Ratio) in the table is the signal-to-noise ratio of additive noise. T0, T 200 , T 500 respectively represent that the reverberation time RT60 is 0 s, 200 ms, and 500 ms.
[0088] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A filtering and summing multi-channel speech separation method based on a simplified attention encoder-decoder network, characterized in that The method comprises the following steps: Step 1: Frame the multi-channel speech signal containing multiple sound sources to obtain the speech signals of each channel at the frame level. Select the speech signal of one channel as the reference channel speech signal, calculate the normalized cross-correlation features with the speech signals of the remaining channels, calculate the embedding features of the reference channel speech signal, splice the normalized cross-correlation features and the embedding features, and use the spliced parameters as the input features of the first simplified attention encoder-decoder network to output the filter parameters for the reference channel speech signal. Step 2: Filter the reference channel speech signal using the filter parameters output by the simplified attention encoder-decoder network in Step 1 to obtain the pre-separated speech signals of each sound source. Step 3: Calculate the normalized cross-correlation features between the pre-separated speech signals of each sound source in Step 2 and the speech signals of the remaining channels, calculate the embedding features of the speech signals of the remaining channels, splice the normalized cross-correlation features and the embedding features, and use the spliced parameters as the input of the second simplified attention encoder-decoder network to output the filter parameters for the speech signals of the remaining channels. Step 4: Filter the speech signals of the corresponding channels using the filter parameters of the remaining channels obtained in Step 3 to obtain the speech signals of each sound source separated from the speech signals of the remaining channels. Add the pre-separated speech signals of each sound source to the speech signals of each sound source separated from the remaining channels to obtain the final separated speech of each sound source.
2. The filtering and summing multi-channel speech separation method based on the simplified attention encoder-decoder network according to claim 1, wherein Step 1 specifically includes the following steps: Select a reference channel, set its number to 1, and first calculate the normalized cross-correlation values between the reference channel speech signal and the speech signals of the remaining channels: Among them is the t-th frame signal received by the reference channel, and the data of the previous and subsequent frames is introduced, that is, the signal with a length of 3L composed of samples from tH-L to tH+2L-1, where H represents the frame shift during frame division; represents the signal with a length of L extracted from the reference channel signal with a length of 3L where j represents the starting serial number of the sample points represents a real vector with a dimension of 1×L; represents the t-th frame signal received by the m-th channel, || || represents the vector norm, and < > represents the inner product calculation; represents the normalized cross-correlation value between the t-th frame, the reference channel signal and the m-th channel signal, and the 2L+1-dimensional normalized cross-correlation function is obtained by connecting them according to the j serial number The normalized cross-correlation function is averaged over different channels to obtain the normalized cross-correlation feature NCC of the t-th frame t : Among them is the normalized cross-correlation feature of the t-th frame; The calculation formula for the embedding features is: wherein is the t-th frame signal received by the reference channel; represents the weight matrix for calculating the embedding, which is a learnable parameter, K u represents the dimension of the embedding vector; Splice the normalized cross-correlation features and the embedding features as the feature parameters of this frame of speech signal, and input the feature parameters into the first simplified attention encoder-decoder network to obtain the filter parameters for this frame of reference channel speech signal. The filter parameters corresponding to the k-th sound source in the t-th frame The calculation formula is as follows: Among them represents the first simplified attention encoder-decoder network, and for the features with a dimension of 2L+1+K input at each moment U after the temporal convolution and dimension transformation therein, outputs K vectors with a dimension of K U ; K represents the number of sound sources; ⊙ represents the element-wise multiplication of two vectors or matrices with the same dimension; is the coefficient of the simplified attention encoder-decoder network ; is the bias; tanh() is the hyperbolic tangent operation, and σ() is the Sigmoid function.
3. The filtering and summing multi-channel speech separation method based on the simplified attention encoder-decoder network according to claim 2, wherein The calculation formula for obtaining the pre-separated speech signals of each sound source in Step 2 is: Among them represents the pre-separated signal of the k-th sound source in the t-th frame filtered from the reference channel signal.
4. The filtering and summing multi-channel speech separation method based on the simplified attention encoder-decoder network according to claim 3, wherein Step 3 specifically includes the following steps: Calculate the normalized cross-correlation values between the speech signals of the remaining channels and the pre-separated speech signals of each sound source: wherein is the t-th frame signal received by the m-th channel, and the data of the previous and next frames is introduced, that is, the samples from tH-L to tH+2L-1 form a signal with a length of 3L; denotes from a vector with a length of L taken out, denotes the pre-separated signal of the k-th sound source in the t-th frame and the m-th channel signal the normalized cross-correlation value of, which is connected according to the j sequence number to obtain a 2L+1-dimensional normalized cross-correlation function Take the normalized cross-correlation function directly as the normalized cross-correlation features of the remaining channels; The embedding features of the speech signals of the remaining channels are calculated as: wherein represents the embedding feature of the m-th channel and the t-th frame; Concatenate the two types of features as the input of the second simplified attention encoder-decoder network to obtain the filter parameters for the speech signals of the remaining channels. Denote the filter parameter corresponding to the k-th sound source at the t-th frame of the m-th channel: Among them represents the second simplified attention encoder-decoder network.
5. The filtering and summing multi-channel speech separation method based on the simplified attention encoder-decoder network according to claim 4, wherein, The calculation formula for obtaining the final separated speech of each sound source in Step 4 is: wherein represents the separated speech of the k-th sound source at the t-th frame.
6. The filtering and summing multi-channel speech separation method based on the simplified attention encoder-decoder network according to claim 1, wherein The simplified attention encoder-decoder network consists of a bidirectional long short-term memory network, an encoding layer and a decoding layer containing an attention mechanism. The simplified attention encoder-decoder network uses the bidirectional long short-term memory network to implement position encoding, and the position encoding is used to integrate the information of the entire input feature sequence and provide the position relationship of different feature vectors for the attention structure.
7. The filtering and summing multi-channel speech separation method based on the simplified attention encoder-decoder network according to claim 6, wherein The encoding layer and the decoding layer of the simplified attention encoder-decoder network consist of layer normalization, a multi-head attention structure, a fully connected layer and a Gaussian error linear unit.
8. The filtering and summing multi-channel speech separation method based on the simplified attention encoder-decoder network according to claim 7, characterized in that, The multi-head attention structure of the simplified attention encoder-decoder network adopts a narrow multi-head form.
9. The filtering and summing multi-channel speech separation method based on a simplified attention encoder-decoder network according to claim 1, wherein The two simplified attention encoder-decoder networks in Step 1 and Step 3 have the same structure but different network parameters.
Citation Information
Patent Citations
Microphone array voice separation method based on TC-ResNet network
CN112201276A
Blind source separation method and acoustic signal processing system for improving interference estimation in binaural wiener filtering
US20100183178A1