Co-attention multichannel TF-GridNet voice separation method based on ad-hoc microphone array

By introducing a common attention module in the multi-channel TF-GridNet network, the problem of insufficient exploration of speaker embedding correlation in the self-organized microphone array is solved, and more efficient channel fusion and speech separation performance is achieved, especially in complex acoustic environments.

CN120048281APending Publication Date: 2025-05-27RES & DEV INST OF NORTHWESTERN POLYTECHNICAL UNIV IN SHENZHEN +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510035447.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing multi-channel speech separation method lacks in-depth exploration of the correlation between embeddings of different speakers when dealing with self-organized microphone arrays, resulting in limited representation of extracted features and poor performance in complex acoustic environments.

Method used

The multi-channel TF-GridNet speech separation method based on the ad-hoc microphone array is adopted. By integrating the common attention module (Co-Attention) into the TF-GridNet network, the correlation and differences between different speakers between different channels are captured, and the feature representation ability is enhanced.

Benefits of technology

It significantly improves the speech separation performance, especially in environments with long reverberation time, channel fusion can be more efficiently performed to improve the speech separation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048281A_ABST
    Figure CN120048281A_ABST
Patent Text Reader

Abstract

The invention discloses a common attention multichannel TF-GridNet voice separation method based on an ad-hoc microphone array, based on a TF-GridNet network, the method integrates a common attention module (Co-Attention) to capture the correlation and difference of different speakers among different channels, so as to enhance the representation of extracted features, and improve the speech separation efficiency. Therefore, high-efficiency channel fusion and voice separation performance are realized. And the common attention module can be used for respectively selecting two attention mechanisms, namely, a self-attention (Co-SA) mechanism and a graph attention (Co-Graph) mechanism, and the common attention module can be used for respectively selecting two attention mechanisms, namely, the self-attention (Co-SA) mechanism and the graph attention (Co-Graph) mechanism. According to the invention, a single-channel TF-GridNet model is expanded, and a common attention module is integrated into a multi-channel TF-GridNet, so that the multi-channel condition of the self-organizing microphone array is processed. Experimental results show that under the conditions of different microphone channel numbers and different reverberation time, the proposed method is remarkably superior to a baseline method, excellent performance is shown, and the voice separation quality in a complex acoustic environment can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech recognition, and particularly relates to a multi-channel TF-GridNet speech separation method based on co-attention of an ad-hoc microphone array. Background Art

[0002] The goal of speech separation is to separate overlapping signals from multiple speakers, which is widely applied in reality in fields such as intelligent assistants (e.g., Siri, Alexa), speech input methods, real-time translation, customer service systems, and smart homes, significantly improving the efficiency and convenience of human-computer interaction. According to the number of microphone array channels, speech separation methods can be divided into single-channel methods and multi-channel methods. Multi-channels can utilize rich spatial information, which is also the focus of the present invention.

[0003] Traditional multi-channel speech separation is achieved through beamforming methods. In recent years, the widespread application of deep neural networks (DNNs) has brought new solutions for multi-channel speech separation. These include: (1) neural time-frequency (TF) masking methods for predicting the spectral mask of a reference channel. (2) Neural time-frequency mapping methods for directly predicting the time-frequency spectrum of speech, such as TF-GridNet, which is a state-of-the-art separation model. (3) Neural beamforming methods for estimating filter parameters in the time domain or frequency domain using neural networks to output the separated speech signals. (4) Explicitly utilizing features containing spatial information, such as spatial information like the inter-channel phase difference (IPD) between input channels to learn the position features of different speakers and improve separation performance.

[0004] However, the above methods are designed for fixed microphone arrays. In contrast, some recent research has focused on self-organizing microphone arrays. The key feature of the latter is that microphone nodes can be randomly and flexibly placed. Nevertheless, this topic seems to be not yet mature, especially in the context of speech separation. Existing work mainly focuses on speech separation under conditions of limited reverberation time, usually not exceeding 0.8 seconds. At the same time, the embeddings from different speakers are usually closely related, but most separation models only use basic convolutional or linear layers. These methods have two main drawbacks: (1) lack of in-depth exploration of the correlation between embeddings of different speakers, which may limit the representation ability of the extracted features, and (2) often need to fuse multiple speakers of multi-channel inputs through simple concatenation and convolution, generating suboptimal fusion features. Summary of the Invention

[0005] To overcome the deficiencies of the prior art, the present invention provides a multi-channel TF-GridNet voice separation method based on co-attention of an ad-hoc microphone array. Based on the TF-GridNet network, this method integrates a co-attention module (Coordinate Graph Attention, Co-Attention) to capture the correlations and differences of different speakers among different channels, which is used to enhance the representation of the extracted features, thereby achieving efficient channel fusion and voice separation performance. The co-attention module can respectively select two attention mechanisms: self-attention (Coordinate Self Attention, Co-SA) and graph attention (Coordinate Graph Attention, Co-Graph). The present invention extends the single-channel TF-GridNet model and integrates the co-attention module into the multi-channel TF-GridNet to handle the multi-channel situation of the self-organizing microphone array. Experimental results show that under the conditions of different numbers of microphone channels and different reverberation times, the proposed method is significantly superior to the baseline method, showing excellent performance and can significantly improve the voice separation quality in complex acoustic environments.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0007] Step 1: Signal model;

[0008] In the time-frequency (T-F) domain, the multi-channel signal captured by an M-channel microphone array is expressed as:

[0009]

[0010] where m ∈ [1, M], f ∈ [1, F], t ∈ [1, T], and c ∈ [1, C] respectively represent the indices of the microphone, frequency, time range, and speaker; and respectively represent the direct sound and reverberation of the speaker; F represents the frequency domain, T represents the time domain, C represents the number of speakers, represents the complex number domain;

[0011] Step 2: Model architecture;

[0012] Using the complex spectrogram of the short-time Fourier transform (STFT) as the acoustic feature, the input is a tensor with a shape of M × 2 × T × F, where 2 represents the stacked real and imaginary parts; after applying a two-dimensional convolution (Conv2D) with a convolution kernel of 3 × 3, global layer normalization (gLN) is used to obtain an embedding tensor with a shape of M × D × T × F;

[0013] Slice the second dimension of the multi-channel tensor according to the number of speakers, and then input the slice into the co-attention module to enhance the feature representation ability by using the feature correlation between different speakers; after co-attention processing, the tensor is re-concatenated and average pooling is performed in the channel dimension to form a single-channel tensor feature with the shape of D×T×F; then, a residual connection is established, and the vector corresponding to the unprocessed reference channel is connected to the output to alleviate the problem of vanishing gradients; after fusion using convolution, the obtained tensor is sent to the time-frequency grid network module TF-GridNet blocks for processing, and a 2D deconvolution Deconv2D with an output channel number of 2C and a convolution kernel size of 3×3 is used to output the spectrum predictions of different speakers, and the inverse short-time Fourier transform iSTFT is used to convert the mapped time-frequency domain signal into a time-domain signal;

[0014] Step 3: Co-attention module;

[0015] Given the input of the self-attention module, denoted as Split along the second dimension according to the number of speakers to obtain the embeddings corresponding to different speakers;

[0016] Assume the number of speakers C = 2, and obtain the embeddings Z of the two speakers 1 and After applying co-attention, obtain F 1 and Then fuse through convolution with a kernel size of 3 to produce the output of this module, denoted as

[0017] Step 4: Co-attention module based on self-attention;

[0018] For the h-th attention head, Use the convolution of Z 1 and Z 2 to calculate the query vectors of the two speakers respectively ( and ), key vectors ( and ), and value vectors ( and ), where the size of each corresponding subspace is E; all these subspaces where d h = E / H, and H represents the total number of attention heads; subsequently, the output is calculated as follows:

[0019]

[0020]

[0021] where and The outputs of all attention heads are as follows: Obtain F 2,t in the same way, where W 1 and are the weight matrices of the linear projection layer, and

[0022] Step 5: Co-attention module based on graph attention;

[0023] Use graph attention to further aggregate information, which applies the self-attention mechanism based on the graph convolutional network (GCN) layer to calculate the attention weights on the graph channels; assuming that all channels are interconnected, the adjacency matrix A is defined as:

[0024] Given the input of co-graph attention, denoted as Z 1t and Z 2t , for the h-th attention head, this process is achieved by projecting into a d and -dimensional space using learnable parameters; the query matrices and key matrices of different speakers are denoted as h and and

[0025] For speaker 1, the score corresponding to the query tensor-key matrix pairing from channel m to channel j is calculated using the following formula:

[0026]

[0027]

[0028] where is a learnable vector; the Softmax function is used to normalize the attention scores of all adjacent channels;

[0029] For two speakers, the aggregated output of channel m is represented by the and of the h-th head:

[0030]

[0031]

[0032] The aggregated features of all nodes are concatenated, and then the aggregated output is obtained as follows:

[0033]

[0034]

[0035] F 2,t Obtained by a similar method.

[0036] A computer program that causes a computer to execute the above-mentioned multi-channel TF-GridNet voice separation method with co-attention.

[0037] An electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned multi-channel TF-GridNet voice separation method with co-attention.

[0038] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned multi-channel TF-GridNet voice separation method with co-attention is implemented.

[0039] A chip, comprising: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the above-mentioned multi-channel TF-GridNet voice separation method with co-attention.

[0040] A computer program product, the computer program product includes a computer storage medium, the computer storage medium stores a computer program, the computer program includes instructions that can be executed by at least one processor, and when the instructions are executed by the at least one processor, the above-mentioned multi-channel TF-GridNet voice separation method with co-attention is implemented.

[0041] The beneficial effects of the present invention are as follows:

[0042] The present invention proposes a multi-channel model based on co-attention for voice separation using a self-organizing microphone array. The core lies in integrating the co-attention mechanism into the model to enhance the interaction between different speakers across multiple channels, thereby promoting more effective channel fusion. Experimental results show that the proposed model has superior performance compared with existing methods, especially in challenging environments with a long reverberation time. These results emphasize the importance of the co-attention mechanism in improving the performance of multi-channel voice separation tasks based on channel fusion. Description of the Drawings

[0043] Figure 1 It is a block diagram of the multi-channel separation model of the present invention.

[0044] Figure 2 It is a block diagram of the co-attention module of the present invention.

[0045] Figure 3 It is a block diagram of the co-attention module based on self-attention of the present invention.

[0046] Figure 4 This is the structural block diagram of the co-attention module based on graph attention for the present invention. Detailed implementation manners

[0047] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0048] The present invention applies co-attention to multi-channel speech separation, focuses on the correlations and differences between different speakers, and uses the attention mechanism to learn the global dependencies of spatio-temporal features, so that the embedding information learned from one speaker can be used to enhance the separation performance of another speaker.

[0049] A multi-channel TF-GridNet speech separation method based on co-attention of an ad-hoc microphone array specifically includes the following steps:

[0050] (1) Signal model:

[0051] In the time-frequency (T-F) domain, the multi-channel signal captured by an M-channel microphone array is expressed as:

[0052]

[0053] where m ∈ [1, M], f ∈ [1, F], t ∈ [1, T], c ∈ [1, C], representing the indices of the microphone, frequency, time range, and speaker respectively. and represent the direct sound and reverberation of the speaker. The focus of this work is to separate the direct signal Y from the mixture X.

[0054] (2) Model architecture;

[0055] The structure of this model is as shown in the attached Figure 1 figure. The main innovations of this model include: (i) it extends the single-channel TF-GridNet model to the multi-channel case applicable to ad-hoc microphone arrays, and (ii) it integrates the co-attention module into the multi-channel TF-GridNet in an end-to-end manner. The architecture of the proposed model is shown in Figure 1 this figure. This method uses the complex spectrogram of the short-time Fourier transform (STFT) as the acoustic feature, and the input is a tensor of shape M × 2 × T × F, where 2 represents the stacked real and imaginary parts. After applying a two-dimensional convolution (Conv2D) with a convolution kernel of 3 × 3, global layer normalization (gLN) is used to obtain an embedding tensor of shape M × D × T × F.

[0056] The subsequent processing process is as shown in the appendix Figure 2 As shown. According to the number of speakers, slice the second dimension of the multi-channel tensor, and then input it into the co-attention module to enhance the feature representation ability by using the feature correlation between different speakers. After co-attention processing, the tensor is re-concatenated and average-pooled in the channel dimension to form a single-channel tensor feature with the shape of D×T×F. Then, a residual connection is established, and the vector corresponding to the unprocessed reference channel is connected to the output to alleviate the vanishing gradient problem. After completing the fusion using convolution, the obtained tensor is sent to the time-frequency grid network module (TF-GridNet blocks) for processing, and a 2D deconvolution (Deconv2D) with an output channel number of 2C and a convolution kernel size of 3×3 is used to output the spectral predictions of different speakers, and the inverse short-time Fourier transform (iSTFT) is used to convert the mapped time-frequency domain signal into a time-domain signal. Co-attention module

[0057] Given the input to the self-attention module, denoted as According to the number of speakers, split along the second dimension to obtain the embeddings corresponding to different speakers. Here, assuming the number of speakers C = 2, the embeddings Z 1 and After applying co-attention, obtain F 1 and Then fuse through convolution with a kernel size of 3 to produce the output of this module, denoted as Appendix Figure 3 and appendix Figure 4 respectively show the co-attention based on self-attention and graph attention, which are alternately applied between the time and frequency dimensions. Below, taking the time dimension as an example, the use of co-attention is explained. Consider the input tensor as T independent sequences, each with a length of T, and then apply co-attention to model and infer the information within each sub-band. Of course, its use in the frequency domain dimension is similar.

[0058] (3) Co-attention module based on self-attention;

[0059] For the h-th attention head, Using the convolution of Z 1 and Z 2 can calculate the query vectors of the two speakers ( and ), key vectors ( and ), and value vectors ( and ), where the size of each corresponding subspace is E. All these subspaces where d h = E / H, and H represents the total number of attention heads. Subsequently, the output is calculated as follows:

[0060]

[0061]

[0062] where and The outputs of all attention heads are: Obtaining F 2,t is done in the same way, where W 1 and are the weight matrices of the linear projection layer, and

[0063] (4) Co-attention module based on graph attention;

[0064] As shown in the appendix Figure 4 , graph attention is used to further aggregate information. It applies a self-attention mechanism based on the graph convolutional network (GCN) layer to calculate the attention weights on the graph channels. The adjacency matrix is derived from the graph by using each channel itself as the query vector to attend to other channels. In this case, assuming that all channels are interconnected, the adjacency matrix A is defined as:

[0065] Given the input of co-graph attention, denoted as Z 1t and Z 2t , for the h-th attention head, this process is achieved by projecting into a d and dimensional space using the learnable parameters h In this case, the query matrix and key matrix for different speakers are denoted as and

[0066] For speaker 1, the score corresponding to the query tensor-key matrix pairing from channel m to channel j is calculated using the following formula:

[0067]

[0068]

[0069] where is a learnable vector. The Softmax function is used to normalize the attention scores of all adjacent channels. The processing for speaker 2 is similar and will not be elaborated here. For the two speakers, the aggregated output of channel m by the h-th head is and denoted as:

[0070]

[0071]

[0072] The aggregated features of all nodes are concatenated. Then, the aggregated output is obtained as follows:

[0073]

[0074]

[0075] F 2,t can be obtained by a similar method.

[0076] Example:

[0077] (1) Experimental setup:

[0078] The self-attention-based method is denoted as Co-SA (Coordinate Self Attention), and the graph-attention-based method is denoted as Co-Graph (Coordinate Graph Attention). The co-attention module contains two spatio-temporal blocks, and each spatio-temporal block has four attention heads. Scale-Invariant Signal-to-Noise Ratio improvement (SI-SNRi), Signal-to-Distortion Ratio improvement (SDRi), Short-Time Objective Intelligibility (STOI), and Perceptual Evaluation of Speech Quality (PESQ) are used to evaluate the separation performance.

[0079] (2) Data preparation:

[0080] A method for evaluating on the multi-channel two-speaker reverberant speech separation task using a self-organizing microphone array is presented. A multi-channel reverberant dataset is created, which contains 20,000, 5,000, and 3,000 4-second utterances from the WSJ0-2mix dataset. For each utterance, a room with a random size is simulated, where the length is randomly selected from [3, 10] meters, the width is randomly selected from [3, 10] meters, and the height is randomly selected from [2.5, 4] meters. The reverberation time T 60 is randomly set in the range of [0.2, 1.2] seconds. The room impulse response function is generated using the gpuRIR method. The positions of the speakers and microphones are randomly placed in the room, with the distance between different speakers restricted to be greater than 0.5 meters and the speaker height between 1 and 2 meters. Additionally, the training data uses a self-organizing microphone array with 8 channels, while the test data includes microphone arrays with 4, 8, 12, and 16 channels.

[0081] (3) Comparison methods:

[0082] The baseline comparison methods mainly include three categories: ① Conventional signal processing methods. Weighted Prediction Error (WPE): The NARA-WPE algorithm is used to generate a single-channel input for TF-GridNet, and the default settings are directly applied here. ② Channel selection. (i) Oracle one-best: The output of the microphone closest to each speaker is selected as the separation prediction for that speaker, assuming the microphone positions are known. (ii) Beamforming: After inputting the multi-channel audio into a single-channel model respectively, traditional acoustic beamforming techniques are used to synthesize the multi-channel output into a single channel. ③ Multi-channel deep learning methods for self-organizing microphones. (i) Mean pooling: Equal weights are assigned to all channels for fusion and then fed into a single-channel TF-GridNet network. (ii) FaSNet-TAC. (iii) Channel attention: In this paper, the co-attention module is replaced by channel attention. (iv) Flow attention: Its usage method is similar to (iii), replacing channel attention with flow attention.

[0083] (4) Experimental results:

[0084] Effect of the number of microphone channels on performance: Table 1 shows the performance of various comparison methods on the WSJ0-2mix-reverb test set, where the number of microphones in the "8Mic" test scenario is the same as that in the training data. The results show that the proposed method Co-SA is significantly better than the baseline methods. Its performance is even better than the scenario where the microphone positions are known in advance, indicating that co-attention enhances the representation ability of the extracted features through the mutual correlation between different speakers, demonstrating its superior performance in the matching scenario. Generally speaking, Co-SA is better than Co-Graph, and Co-Graph is better than other methods. It is also observed that the proposed method has strong generalization ability in the mismatched test scenarios. Specifically, the results of the mismatched scenarios with 4, 12, and 16 microphones all show good performance. In the case of 4 microphones, although the number of channels is reduced, the SDRi of Co-SA only drops by about 4%, while STOI and PESQ remain almost unchanged, indicating its strong robustness to the change of the number of channels. In addition, as the number of microphones increases, the separation performance of Co-SA will be further improved. Generally speaking, Co-SA shows excellent performance and robustness in both matching and mismatching scenarios, especially in terms of SI-SNRi and PESQ.

[0085] Table 1 Results of comparison methods with different numbers of microphones

[0086]

[0087] Effect of reverberation time on performance: Table 2 lists the effects of different reverberation times on the performance of the comparison methods. The proposed method, especially Co-SA, is better than all other methods in all evaluation metrics, especially in terms of SDRi, SI-SNRi, and PESQ. Although the Co-Graph method maintains stable performance in SDRi and SI-SNRi, it is still lower than Co-SA but still better than most other methods. Under high reverberation conditions, the performance of previous methods such as beamforming and FaSNet-TAC drops significantly, especially in SDRi and PESQ. In contrast, advanced methods such as Co-SA and Stream Attention show strong robustness and can significantly improve the speech separation quality in complex acoustic environments.

[0088] Table 2 Results of comparison methods in different reverberation scenarios

[0089]

Claims

1. A multi-channel TF-GridNet speech separation method based on joint attention of ad-hoc microphone array, characterized in that The steps include: Step 1: Signal model; In the time-frequency (TF) domain, a multi-channel signal captured by an M-channel microphone array It is expressed as: Where m∈[1,M], f∈[1,F], t∈[1,T], c∈[1,C] represent the indices of microphone, frequency, time range and speaker respectively; and Represent the direct sound and reverberation of the speaker respectively; F represents the frequency domain, T represents the time domain, and C represents the number of speakers. represents a complex domain; Step 2: Model architecture; The complex spectrum of the short-time Fourier transform STFT is used as the acoustic feature, and the input is a tensor of shape M×2×T×F, where 2 represents the stacked real and imaginary parts; after applying a two-dimensional convolution Conv2D with a convolution kernel of 3×3, a global normalization gLN is used to obtain an embedding tensor of shape M×D×T×F; According to the number of speakers, the second dimension of the multi-channel tensor is sliced ​​and then input into the joint attention module to improve the feature representation capability by using the feature correlation between different speakers; After the joint attention processing, the tensor is reconstructed and average-pooled in the channel dimension to form a single-channel tensor feature with a shape of D×T×F. Then, a residual connection is established, and the vector corresponding to the unprocessed reference channel is connected to the output to alleviate the gradient vanishing problem. After the fusion is completed using convolution, the obtained tensor is sent to the time-frequency grid network module TF-GridNet blocks for processing, and the two-dimensional deconvolution Deconv2D with an output channel number of 2C and a convolution kernel size of 3×3 is used to output the spectrum prediction of different speakers. The mapped time-frequency domain signal is converted into a time domain signal using the inverse short-time Fourier transform iSTFT. Step 3: Joint attention module; Given the input of the self-attention module, denoted as Split along the second dimension according to the number of speakers to obtain embeddings corresponding to different speakers; Assume the number of speakers C = 2, and get the embeddings Z1 and Z2 of the two speakers. After applying joint attention, we get F1 and Then the convolution with kernel size 3 is used to fuse the output of this module, which is recorded as Step 4: Co-attention module based on self-attention; For the h-th attention head, The query vectors for the two speakers are calculated using the convolution of Z1 and Z2 respectively ( and ), key vector( and ) and the value vector ( and ), where each corresponding subspace is of size E; all these subspaces where d h =E / H, where H represents the total number of attention heads; then, the output is calculated as follows: in and The output of all attention heads is: Get F 2,t The method is the same, where W1 and is the weight matrix of the linear projection layer, and Step 5: Joint attention module based on graph attention; Use graph attention to further aggregate information, which applies the self-attention mechanism based on the graph convolutional network GCN layer to calculate the attention weights on the graph channels; Assuming that all channels are correlated, the adjacency matrix A is defined as: Given the input of the joint graph attention, denoted as Z 1t and Z 2t , for the hth attention head, by using the learnable parameters and Projection to d h dimensional space to realize this process; the query matrix and key matrix of different speakers are represented as and For speaker 1, the score for the query tensor-key matrix pairing from channel m to channel j is calculated using the following formula: in is a learnable vector; the Softmax function is used to normalize the attention scores of all adjacent channels; For two speakers, the aggregate output of channel m is composed of the hth head and express: The aggregated features of all nodes are concatenated and then the aggregated output is obtained as follows: F 2,t Obtained by a similar method.

2. A computer program, characterized in that The computer program enables a computer to execute the method as claimed in claim 1.

3. An electronic device, characterized in that: include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method as claimed in claim 1.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method as claimed in claim 1 is implemented.

5. A chip, characterized in that: include: A processor, used to call and run a computer program from a memory, so that a device equipped with the chip executes the method as claimed in claim 1.

6. A computer program product, characterized in that The computer program product comprises a computer storage medium storing a computer program, wherein the computer program comprises instructions executable by at least one processor, and when the instructions are executed by the at least one processor, the method according to claim 1 is implemented.