Multi-channel speech enhancement method based on multi-scale spatial information and spectral feature fusion

This multi-channel speech enhancement method, which integrates multi-scale spatial information and spectral features, solves the problem of lack of effective fusion in multi-channel speech enhancement, achieves better speech signal separation, and improves the model's performance in complex environments.

CN119889338BActive Publication Date: 2025-10-31SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411912274.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-10-31
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing multi-channel speech enhancement methods lack effective fusion of spatial information and spectral features, resulting in poor performance in separating target speech from background noise in complex environments.

Method used

A method of fusing multi-scale spatial information and spectral features is adopted. Spatial features within, between, and across all channels are extracted through grouped convolution. Then, a hierarchical feature fusion strategy and a multi-scale encoder-decoder are used in conjunction with a recurrent neural network for noise suppression and speech enhancement.

Benefits of technology

It improves the speech enhancement effect of the model in complex environments, enhances the ability to process speech signals, adapts to the directionality and coherence of speech, and improves robustness and overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889338B_ABST
    Figure CN119889338B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features. It recombines different spectral components according to spectral characteristics to extract intra-channel, inter-channel, and full-channel feature patterns. These features are then fused to create a unified deep feature set. A local feature extraction module is introduced to enhance the feature weights of the current frame, and a feature attention mechanism is used to fuse features at different scales. A decomposition attention mechanism is introduced to fuse encoder and decoder outputs in multiple decomposition spaces, allowing detailed features to be used by the deep module. This invention combines spatial and spectral features, using feature fusion methods to create a unified feature representation. By adaptively learning and utilizing the patterns contained in spatial features through the attention module, rather than fitting physically meaningful directional features, it can more flexibly adapt to different scenarios and has good application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features, and belongs to the field of speech enhancement technology. Background Technology

[0002] Speech enhancement aims to improve the accuracy and quality of noise-contaminated speech and is an indispensable component in many applications such as hearing aids, smart terminal devices, and video conferencing systems. In these applications, speech is inevitably affected by environmental noise, room reverberation, and other factors. Furthermore, with the development of automatic speech recognition and large-scale language models, the importance of speech enhancement as a fundamental front-end module is steadily increasing, as high-quality speech signals are a prerequisite for accurate prediction in these applications. Based on the number of recording microphones, speech enhancement can be divided into two categories: single-channel speech enhancement and multi-channel speech enhancement. Single-channel speech enhancement primarily extracts clear speech from a single noisy speech channel to remove background noise. Limited by the number of microphones, single-channel speech enhancement can only utilize spectral information, which to some extent limits its performance. Due to the lack of spatial information, single-channel models often struggle to effectively distinguish target speech from background noise, especially in complex acoustic environments. Multi-channel speech enhancement, on the other hand, can utilize spatial information captured from recordings from multiple microphones, resulting in better speech clarity and quality compared to single-channel speech enhancement. Multichannel speech enhancement utilizes spatial filtering to amplify speech signals from specific directions while suppressing noise from other directions, resulting in better separation of speech from noise, especially in complex environments. However, classical spatial filtering also has many limitations. It is typically based on linear operations and cannot effectively handle nonlinear features or complex feature patterns. In some cases, linear filtering may lead to the loss of important information. Furthermore, different filter and parameter settings can produce drastically different results, and selecting appropriate parameters often requires experience and trial and error, which can increase the complexity of the process.

[0003] Recently, with the significant progress made by deep learning technology in single-channel speech enhancement, some researchers have proposed methods combining deep neural networks with traditional beamforming, known as neural beamforming. The core idea driving these methods is to replace traditional algorithms with deep neural network models to estimate components used in beamformers or subsequent filtering techniques, such as the posterior probability of speech presence, interaural phase difference, and amplitude spectrum of interfering microphone signals. Neural beamformers generally have lower complexity and represent an improvement over traditional methods; however, challenges such as linearization and the ability to handle non-stationary noise remain to be overcome. Furthermore, matrix operations such as inverting matrices can sometimes become numerically unstable when training with deep neural networks. Therefore, more researchers are turning to data-driven deep learning methods. Data-driven deep learning methods can fully utilize the nonlinear fitting and time-series modeling capabilities of models, making them more suitable for multi-channel speech enhancement. The channel-attention-intensive CA Dense U-net model implements beamforming processing by introducing a channel attention mechanism, enabling the model to perform nonlinear beamforming. The Dense Frequency-Time Attention Network DeFT-AN integrates three different types of modules to handle feature patterns in the spatial, frequency, and temporal domains, respectively. SpatialNet models, on the other hand, process time-frequency features independently. These methods demonstrate that how spatial information is handled is crucial for distinguishing speech from noise, as speech is typically directional and spatially coherent, while noise is often diffuse or has low spatial correlation. These methods use separate modules to process spatial and spectral information without fusing them. While this design effectively extracts features from each, it may limit the model's flexibility and overall performance when processing complex signals. Without effective integration, the model may not fully utilize the interrelationship between spatial and spectral information, thus affecting the effectiveness of multi-channel speech enhancement. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features. By extracting and fusing spatial information and spectral features of different channel combinations, the advanced nature of single-channel speech enhancement models is extended to the multi-channel domain, thereby better separating target speech from background noise.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0006] A multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features includes the following steps:

[0007] Step 1: For the noisy spectrum of a multi-channel microphone, the spatial features within the channel, between channels, and across all channels are extracted using grouped convolution. The spatial features include spatial information and spectral features.

[0008] Step 2: The spatial features extracted in Step 1, including intra-channel, inter-channel, and full-channel features, are fused using a hierarchical feature fusion strategy to obtain a unified deep feature representation.

[0009] Step 3: Construct a multi-scale spatial aggregation encoder and a multi-scale decoder. A recurrent neural network connects the multi-scale spatial aggregation encoder and the multi-scale decoder. The unified deep feature representation obtained in Step 2 is fed into the multi-scale spatial aggregation encoder for encoding and feature mapping. The output of the multi-scale spatial aggregation encoder is fed into the recurrent neural network to extract feature patterns at frame-by-frame time steps, and to perform noise suppression and speech enhancement. The output of the recurrent neural network is fed into the multi-scale decoder for signal recovery and reconstruction.

[0010] Step 4: Fuse the outputs of the multi-scale spatial aggregation encoder and the multi-scale decoder, and enlarge the depth features obtained by fusion and map them to the spectral domain to obtain the enhanced speech signal spectrum.

[0011] As a preferred embodiment of the present invention, the specific process of step 1 is as follows:

[0012] Step 1.1: For the microphone noisy spectrum of M channels, extract the spatial features within each channel. Each channel corresponds to a set of convolutions, then the spatial features X of the M channels are obtained. inner Represented as:

[0013]

[0014] Where, δ and X represents the PReLU activation function and group normalization operation, respectively; (m) (t,k) represents the spatial characteristics of the m-th channel. Represents a real number, where T and F represent the number of frames and the number of frequency points, respectively; Φ m The parameters of the m-th channel convolution kernel are represented, where m = 1, ..., M, and * indicates the convolution operation; (t, k) represents the frame and frequency index, where t = 1, ..., T, and k = 1, ..., F.

[0015] Step 1.2: For the microphone noisy spectrum of the M channels, combine the real parts of the spectrum of all channels together, and combine the imaginary parts of the spectrum of all channels together. Extract spatial features from the combinations of real and imaginary parts respectively. Each combination corresponds to a set of convolutions, represented as:

[0016]

[0017] Among them, X outer (r / i) represents the spatial characteristics of the combination of the real and imaginary parts of the M-channel.

[0018]

[0019] Θ r / i This represents the kernel parameters corresponding to the convolution of the real and imaginary parts. Let these represent the real parts of the spectrum of the 1st and Mth channels, respectively. Let represent the imaginary parts of the spectrum of the 1st and Mth channels, respectively;

[0020] Step 1.3: For the microphone noisy spectrum of the M channel, a set of convolutions is used to extract spatial features. The receptive field of the convolution kernel includes all microphone channels, resulting in the full-channel spatial feature X. full .

[0021] As a preferred embodiment of the present invention, the specific process of step 2 is as follows:

[0022] Step 2.1: The feature attention fusion method is used to fuse the spatial features within and between channels to obtain the fused features; specifically:

[0023] Add the spatial features within the channel to the spatial features between channels point-to-point:

[0024]

[0025] Where Z represents the feature after point-to-point addition, and C represents the number of channels of the intermediate feature;

[0026] Calculate the soft selection fusion weight W:

[0027]

[0028] Where σ represents the Sigmoid operator, σ(x) = 1 / 1 + e -x Operator This indicates that the global feature representation is calculated along the sub-band. PWC represents point-to-point addition, and PWC represents point-to-point convolution with a kernel size of (1,1). The representation layer normalization operation, θ1 and θ2 respectively represent The corresponding parameters for the first and second layers of PwC, φ1 and φ2 respectively represent The corresponding parameters for PwC first and second layers;

[0029] Based on the soft-selection fusion weight W, the spatial characteristics within and between channels are fused:

[0030] X1=W⊙X inner +(IW)⊙X outer

[0031] Where X1 represents the fused features, ⊙ represents point-to-point multiplication, and I represents a one-dimensional vector;

[0032] Step 2.2: Based on the temporal attention mechanism, the fused features obtained in Step 2.1 are fused with the full-channel spatial features to obtain a unified deep feature representation; specifically:

[0033] Using each frequency component in X1 as a key value, X full The frequency components in all channels are used as query and action values ​​to generate a spatial aggregation attention map that spans frames and channels.

[0034] For frequency point k, calculate the key value of the attention weight. The query value is pass Calculate the weight matrix A for frequency point k. k ,in, The masking matrix is ​​formed by combining the weight matrices of all frequency points. And act on By aggregating spatial information and spectral features, a unified deep feature representation is obtained, namely...

[0035]

[0036] Here, @ represents matrix multiplication.

[0037] As a preferred embodiment of the present invention, the specific process of step 3 is as follows:

[0038] Step 3.1: Construct a multi-scale spatial aggregation encoder, comprising several layers of multi-scale spatial aggregation modules. The input X1 of the first-layer multi-scale spatial aggregation module is the unified depth feature representation obtained in Step 2, and the input X of the i-th-layer multi-scale spatial aggregation module is... i It is the cumulative combination of the regression outputs of all previous layers, expressed as in, These represent the output results of the (i-1)th and 1st layer multi-scale spatial aggregation modules, respectively;

[0039] The i-th layer multi-scale spatial aggregation module processes the input X i Perform dilated convolution to obtain the global feature X. global Simultaneously, the i-th layer multi-scale spatial aggregation module processes the input X. i Perform convolution operations to obtain local features X local ;

[0040]

[0041] Using a feature attention fusion method for X global and X local By fusing, we obtain:

[0042]

[0043] By using two cascaded conformer modules, spatial information and spectral features are aggregated from the time dimension and the spectral dimension, respectively. Specifically, the time-domain conformer module is used to... Remodeling The format is defined, and the attention weight of each frequency component in the time dimension is calculated. Using the frequency domain dimension conformer module to Remodeling The format is defined, and attention weights are calculated based on the frequency component correlation matrix. Both the time-domain and frequency-domain Conformer modules apply the calculated attention weights to the time-domain and frequency-domain dimensions. Thus, the output of the i-th layer multi-scale spatial aggregation module is obtained.

[0044] Step 3.2: Construct a multi-scale decoder, including several layers of multi-scale decoding modules. The input of the first layer of the multi-scale decoding module is the output Y1 of the recurrent neural network, and the input Y of the i-th layer of the multi-scale decoding module is... i It is the cumulative combination of the regression outputs of all previous layers, expressed as in, These represent the output results of the (i-1)th and 1st layer multi-scale decoding modules, respectively;

[0045] The i-th layer multi-scale decoding module processes the input Y i Perform dilated convolution to obtain the global feature Y. global Meanwhile, the i-th layer multi-scale decoding module decodes the input Y. i Perform convolution operations to obtain local features Y local ;

[0046]

[0047] The feature attention fusion method is used to analyze Y global and Y local The output of the i-th layer multi-scale decoding module is obtained by fusion.

[0048]

[0049] As a preferred embodiment of the present invention, the specific process of step 4 is as follows:

[0050] Step 4.1: Use parallel linear transformation to map the output features of the multi-scale spatial aggregation encoder to H factorized subspaces, where the h-th subspace features... Represented as:

[0051]

[0052] Where h∈[1,2,…,H], C is the dimension of each subspace, i.e., the dimension of the channel, and X e W represents the output features of a multi-scale spatial aggregation encoder. e This represents the feature transformation matrix corresponding to the multi-scale spatial aggregation encoder;

[0053] Step 4.2: Map the output features of the multi-scale decoder to the same subspace as the multi-scale spatial aggregation encoder, and use the normalized exponential function operator softmax to calculate the weights of each subspace. Represented as:

[0054]

[0055] Among them, X d W represents the output features of the multi-scale decoder. d This represents the feature transformation matrix corresponding to the multi-scale decoder;

[0056] Step 4.3: Apply the weights obtained by the multi-scale decoder in different subspaces to the spatial features corresponding to those subspaces, and sum them to obtain the enhanced speech signal features, expressed as:

[0057]

[0058] Among them, X ffa This indicates the features of the enhanced speech signal.

[0059] A computer device includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the computer program to implement the steps of the multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features.

[0060] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features.

[0061] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:

[0062] 1. The method designed in this invention aims to simultaneously capture spatial information and spectral features at different scales, and enhances the processing capability of multi-channel speech signals through an effective fusion mechanism. This design not only improves the model's performance in complex environments, but also better adapts to the directionality and coherence of speech, thereby achieving superior speech enhancement effects.

[0063] 2. This invention fully extracts and aggregates implicit spatial information across channels, frequencies, and frames, as well as captures the cross-frame dynamic characteristics of spectral features, and aggregates spectral and spatial information based on dynamic features to create a unified deep feature representation.

[0064] 3. This invention constructs a multi-scale spatial aggregation encoder and a multi-scale decoder. It enhances the weight of the current frame spectrum through local processing and feature attention fusion methods, thereby alleviating the problem of weakening local features in the receptive field over a long period of time. It further integrates the deep feature representations of different time scales by utilizing the inherent relationships within the spectrum.

[0065] 4. This invention connects the outputs of the encoder and decoder based on the skip connection module of decomposed attention, which solves the problem of continuous loss of details of features during the forward feed of the model, and further improves the overall performance and robustness. Attached Figure Description

[0066] Figure 1 This is a flowchart of the multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features of the present invention;

[0067] Figure 2 This is a network structure diagram of the multi-channel speech enhancement model of the present invention;

[0068] Figure 3 The graph shows a comparison between the spectrum of noisy speech and the spectrum of speech enhanced by the method of the present invention; wherein, (a) is the spectrum of noisy speech, (b) is the spectrum of clean speech, and (c) is the spectrum of speech enhanced by the method of the present invention. Detailed Implementation

[0069] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0070] like Figure 1 As shown, the multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features of the present invention includes the following steps:

[0071] (1) Input the noisy speech spectrum of the multi-channel input into the multi-channel spatial spectrum feature extractor to extract spatial information from the intra-channel, inter-channel and full-channel receptive fields respectively;

[0072] (2) The three outputs extracted in step (1) are sent to the multi-layer spatial fusion module, which is divided into two stages of feature fusion, namely feature attention fusion and T-conformer multi-head attention fusion:

[0073] 1) The feature attention module receives spatial information depth features from within and between channels. Based on its global and local information, it calculates soft selection weights and fuses spatial information from two different receptive fields while ensuring the dynamic characteristics of the features.

[0074] 2) Using the output of the feature attention module and the full-channel features as input, the conformer module is used to take the output of the feature attention module as the key and the full-channel features as the query and value. A causal mask is introduced to aggregate spatial information and spectral features to generate a unified deep feature representation.

[0075] (3) The unified deep feature representation generated by the multi-layer spatial fusion module is fed into the multi-scale spatial aggregation encoder for encoding and feature mapping.

[0076] (4) The encoder output is fed into a recurrent neural network to extract feature patterns at frame time steps, and noise suppression and extraction are performed.

[0077] (5) The output of the recurrent neural network is fed into the multi-scale decoder for signal recovery and reconstruction;

[0078] (6) By using decompositional attention to fuse the outputs of the encoder and decoder, lost details are introduced, and the depth features are enlarged and mapped to the spectral domain to reconstruct the signal features and output the enhanced speech signal spectrum.

[0079] Figure 2 This is a network structure diagram of the multi-channel speech model of the present invention, which is mainly described as follows:

[0080] Step (A) involves constructing a multi-channel spatial spectrum feature extractor to extract spatial information from different channel combinations, thereby capturing cross-channel spatial patterns.

[0081] In multi-channel speech spectra, spatial information is implicit between channels and between spectra. This invention employs grouped convolution to extract spatial patterns within channels, between channels, and across all channels. The specific steps are as follows:

[0082] (A1) For spatial information within a channel, assuming there are M microphone input channels, the spectral characteristics of the m-th channel can be expressed as: The extraction process can then be represented as:

[0083]

[0084] Where, δ and These represent the PReLU activation function and the group normalization operation, respectively; Φ m This represents the parameters of the m-th channel convolution kernel. * indicates the convolution operation. t and k represent the frame and frequency index, respectively.

[0085] (A2) For inter-channel spatial information, first combine the real and imaginary parts of all microphone spectra to form two sets, namely... Then, spatial features are extracted using two different sets of convolutional modules, which can be represented as:

[0086]

[0087] Where, Θ r / i This represents the kernel parameters corresponding to the two sets of convolutions.

[0088] (A3) For full-channel spatial information X full It is calculated in a similar way to the extraction of spatial information within the channel, except that in this stage, there is only one set of convolution operations, and the receptive field of the convolution kernel includes all microphone channels.

[0089] Step (B) involves constructing a spatial feature fusion strategy. Based on the multi-path spatial spectral feature output, a hierarchical fusion mechanism is adopted to gradually fuse spatial information and spectral features to create a unified deep feature representation.

[0090] This invention employs a hierarchical feature fusion strategy to reduce the complexity of learning. The fusion process consists of two steps: first, a feature attention method is used to merge inter-channel and intra-channel features; then, a narrowband processing module based on a conformer structure is proposed to integrate it with the full-channel output. The specific steps are as follows:

[0091] (B1) This invention fuses intra-channel and inter-channel spatial information using a feature attention fusion method. These features contain spatial patterns at different scales. To preserve the original numerical dynamic range and distribution characteristics of the fused features, a soft selection approach is employed. Given the point-to-point summed features are... The soft selection fusion weights can then be expressed as:

[0092]

[0093] Where σ represents the Sigmoid operator, σ(x) = 1 / 1 + e -x Operator This indicates that the global feature representation is calculated along the sub-band; while the module with θ / φ as a parameter... This can be represented as two cascaded convolutions:

[0094]

[0095] PWC is a point-to-point convolution with a kernel size of (1,1). Presentation layer standardization. Parameter subscripts indicate the layer number, such as one of the modules. The parameters are denoted by θ, where θ1 and θ2 represent the parameters of the first and second layer PWCs, respectively. After calculating the soft selection weights, the overall fusion process can be represented as:

[0096] X1=W⊙X inner +(IW)⊙X outer

[0097] Where ⊙ represents point-to-point multiplication, and I represents a completely uniform vector.

[0098] (B2) Use the Conformer attention module to further combine the output X1 of step (B1) with the full-channel feature X full Fusion. A narrowband approach is used to capture the dynamic characteristics of each frequency component, using each frequency component in X1 as a key value, X... full The frequency components across all channels are used as query and action values ​​to generate a spatially aggregated attention map that spans both frames and channels, and this map is used to aggregate the two inputs. For frequency point k, the key value for calculating the attention weights can be represented as... The query value is pass Calculate the weight matrix for frequency point k, where It is a masking matrix used to control the causal properties of the model, and C represents the number of channels for the intermediate features. The feature matrices of all frequency points are combined. And act on By aggregating spatial information and spectral features, a unified feature representation is obtained, namely:

[0099]

[0100] Here, @ represents matrix multiplication.

[0101] Step (C) involves constructing a multi-scale spatial aggregation encoder and a multi-scale decoder. Local convolutions are introduced on top of a densely connected convolutional network to enhance the resolution of the current frame's spectral features, thus constructing the multi-scale encoder and decoder. To introduce a spatial aggregation strategy into the multi-scale encoder, this invention first utilizes a feature attention mechanism to fuse feature representations at different time scales, and then fuses spatial features along the channel dimension using narrowband and full-band methods. The specific steps are as follows:

[0102] (C1) A PWC convolution is added to each dilated convolution step of the dense convolutional encoder as a local feature processing step, and deep feature representations at different scales are fused through a feature attention mechanism, namely cross-frame global features and intra-frame cross-band spectral mode features. Given input features The input to the i-th layer is the cumulative combination of the outputs of all previous regressions, denoted as: in This represents the output of the feature attention fusion module at layer i-1, which fuses the global feature X. global and local features X local X global and X local The extraction process can be represented as:

[0103]

[0104] (C2) Fuse global features X using the feature attention fusion method proposed in step (B1). global and local features X local , can be represented as in This is the output of the feature attention fusion module at the i-th layer of the multi-scale encoder-decoder. If there are subsequent iterations, it will be combined with the input. It is fed into the (i+1)th layer, otherwise it is used as the output of the multi-scale codec.

[0105] (C3) In order to aggregate spatial information and spectral features during the encoding stage, the output of each layer of the encoder is processed. A post-processing module, the instantaneous frequency domain conformer, is added. This module utilizes two cascaded conformer modules to aggregate spatial information and spectral features from the time and spectral dimensions, respectively. The time-domain conformer block reshapes the input representation into... And calculate the attention weight of each frequency component in the time dimension. In contrast, the frequency domain dimension conformer reshapes the input into a format. Attention weights are calculated based on the frequency component correlation matrix. Both T-conformer and F-conformer apply the calculated attention weights to the input features to obtain a unified deep feature representation after fusion.

[0106] Step (D) constructs a decomposed attention fusion connection module to connect the multi-scale spatial aggregation encoder and the multi-scale decoder, solving the problem of continuous loss of details in feature transmission in the model.

[0107] (D1) Introduces a factorial attention method to integrate the deep representations of the encoder and decoder. Given the encoder output X... e and decoder output X d This invention uses parallel linear transformation to map the encoder's output features to H factorized subspaces, where the h-th subspace feature can be represented as:

[0108]

[0109] Where h∈[1,2,…,H], and C is the dimension of each subspace, that is, the dimension of the channel. W e This represents the feature transformation matrix corresponding to the encoder, which is a variable learned by the model.

[0110] (D2) will convert the decoder's output feature X d Mapped to the same subspace as the encoder, and the weights for each subspace are computed using the normalized exponential function operator softmax, expressed as:

[0111]

[0112] Among them, W d This represents the feature transformation matrix corresponding to the decoder, which is a variable learned by the model.

[0113] (D3) Apply the weights obtained by the decoder in different subspaces to the decoder features corresponding to those subspaces, and sum them to obtain the fused features. This process can be represented as:

[0114]

[0115] Table 1 demonstrates the state-of-the-art performance of the proposed method on the 4-channel L3DAS dataset. In terms of performance metrics, the proposed multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features significantly improves multiple speech quality perception metrics and outperforms currently popular multi-channel speech enhancement methods.

[0116] Table 1 Algorithm Performance Comparison

[0117] Model Parameter order of M Complexity: FLOPS / G PESQ STOI WER (%) Unproc - - 1.21 0.62 34.10 Baseline 5.52 32.15 1.63 0.86 19.18 SEUSpeech 2.18 16.57 2.62 0.92 5.44 DeFT-AN 2.67 37.99 3.09 0.95 4.95 McNet 1.85 27.96 2.74 0.94 4.01 This invention model 3.10 29.75 3.37 0.96 3.63

[0118] Figure 3 The graph shows a comparison between the spectrum of noisy speech and the spectrum of speech enhanced by the method of the present invention; wherein, (a) is the spectrum of noisy speech, (b) is the spectrum of clean speech, and (c) is the spectrum of speech enhanced by the method of the present invention.

[0119] Based on the same inventive concept, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the aforementioned multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features.

[0120] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features.

[0121] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0122] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0123] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0124] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0125] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.

Claims

1. A multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features, characterized in that, Includes the following steps: Step 1: For the noisy spectrum of a multi-channel microphone, the spatial features within the channel, between channels, and across all channels are extracted using grouped convolution. The spatial features include spatial information and spectral features. The specific process is as follows: Step 1.1, for The microphone noise spectrum of each channel is used to extract spatial features within each channel. Each channel corresponds to a set of convolutions. Spatial characteristics of the passage Represented as: , in, and These represent the PReLU activation function and the group normalization operation, respectively. Indicates the first Spatial characteristics of each channel , Represent real numbers, These represent the number of frames and the number of frequency points, respectively. Indicates the first The parameters of the convolution kernel for each channel. , Indicates the convolution operation; Indices representing frames and frequency points. , ; Step 1.2, for The microphone noisy spectrum of each channel is combined by concatenating the real parts of all channel spectra and the imaginary parts of all channel spectra. Spatial features are extracted from the combinations of the real and imaginary parts respectively. Each combination corresponds to a set of convolutions, represented as follows: , in, express Spatial characteristics of the combination of real and imaginary parts of a channel. , , This represents the kernel parameters corresponding to the convolution of the real and imaginary parts. They represent the first The real part of the spectrum of each channel. They represent the first The imaginary part of the spectrum of each channel; Step 1.3, for The microphone noisy spectrum of each channel is extracted using a set of convolutions to obtain spatial features. The receptive field of the convolution kernels includes all microphone channels, resulting in full-channel spatial features. ; Step 2: The spatial features extracted in Step 1, including intra-channel, inter-channel, and full-channel features, are fused using a hierarchical feature fusion strategy to obtain a unified deep feature representation. Step 3: Construct a multi-scale spatial aggregation encoder and a multi-scale decoder. A recurrent neural network connects the multi-scale spatial aggregation encoder and the multi-scale decoder. The unified deep feature representation obtained in Step 2 is fed into the multi-scale spatial aggregation encoder for encoding and feature mapping. The output of the multi-scale spatial aggregation encoder is fed into the recurrent neural network to extract feature patterns at frame-by-frame time steps, and to perform noise suppression and speech enhancement. The output of the recurrent neural network is fed into the multi-scale decoder for signal recovery and reconstruction. Step 4: Fuse the outputs of the multi-scale spatial aggregation encoder and the multi-scale decoder, and enlarge the depth features obtained by fusion and map them to the spectral domain to obtain the enhanced speech signal spectrum.

2. The multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features according to claim 1, characterized in that, The specific process of step 2 is as follows: Step 2.1: The feature attention fusion method is used to fuse the spatial features within and between channels to obtain the fused features; specifically: Add the spatial features within the channel to the spatial features between channels point-to-point: , in, The feature is the result of point-to-point addition, where C represents the number of channels in the intermediate feature; Calculate the soft selection fusion weights : , , , in, This represents the Sigmoid operator. Operator This indicates that the global feature representation is calculated along the sub-band. PWC represents point-to-point addition, and PWC represents point-to-point convolution with a kernel size of (1,1). Presentation layer standardized operations, and They represent The corresponding parameters for PwC's first and second layers, and They represent The corresponding parameters for PwC first and second layers; Based on soft selection fusion weights Integrating the spatial features within and between passageways: , in, Indicates the characteristics after fusion. This represents point-to-point multiplication. Represents a vector that is all one; Step 2.2: Based on the temporal attention mechanism, the fused features obtained in Step 2.1 are fused with the full-channel spatial features to obtain a unified deep feature representation; specifically: by Each frequency component is used as a key value. The frequency components in all channels are used as query and action values ​​to generate a spatial aggregation attention map that spans frames and channels. For frequency points Calculate the key value of attention weights The query value is ,pass Calculate frequency points weight matrix ,in, The masking matrix is ​​formed by combining the weight matrices of all frequency points. and act on By aggregating spatial information and spectral features, a unified deep feature representation is obtained, namely... , in, This represents matrix multiplication.

3. The multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features according to claim 2, characterized in that, The specific process of step 3 is as follows: Step 3.1: Construct a multi-scale spatial aggregation encoder, including several layers of multi-scale spatial aggregation modules. The input of the first layer of multi-scale spatial aggregation module... The unified deep feature representation obtained in step 2, the first Input of the multi-scale spatial aggregation module It is the cumulative combination of the regression outputs of all previous layers, expressed as ,in, They represent the first Output results of the 1st layer multi-scale spatial aggregation module; No. Layered multi-scale spatial aggregation module for input Perform dilated convolution to obtain global features. At the same time, the first Layered multi-scale spatial aggregation module for input Perform convolution operations to obtain local features. ; , Employing a feature attention fusion method to and By fusing, we obtain: , By using two cascaded conformer modules, spatial information and spectral features are aggregated from the time dimension and the spectral dimension, respectively. Specifically, the time-domain conformer module is used to... Remodeling The format is defined, and the attention weight of each frequency component in the time dimension is calculated. The frequency domain dimension conformer module is used to... Remodeling The format is defined, and attention weights are calculated based on the frequency component correlation matrix. The Conformer module applies the calculated attention weights to both the time-domain and frequency-domain dimensions. Thus, the first Output of multi-scale spatial aggregation module ; Step 3.2: Construct a multi-scale decoder, which includes several layers of multi-scale decoding modules. The input of the first layer of the multi-scale decoding module is the output of the recurrent neural network. , No. Input of the multi-scale decoding module It is the cumulative combination of the regression outputs of all previous layers, expressed as ,in, They represent the first Output results of the 1st layer multi-scale decoding module; No. The multi-scale decoding module of the layer performs input... Perform dilated convolution to obtain global features. At the same time, the first The multi-scale decoding module of the layer performs input... Perform convolution operations to obtain local features. ; , Employing a feature attention fusion method to and By merging, we obtain the first... Output of multi-scale decoding module : 。 4. The multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features according to claim 3, characterized in that, The specific process of step 4 is as follows: Step 4.1: Use parallel linear transformation to map the output features of the multi-scale spatial aggregation encoder to... The factorized subspace, the first factorization subspace Subspace features Represented as: , in, , The dimension of each subspace, i.e., the dimension of the channel, This represents the output features of a multi-scale spatial aggregation encoder. This represents the feature transformation matrix corresponding to the multi-scale spatial aggregation encoder; Step 4.2: Map the output features of the multi-scale decoder to the same subspace as the multi-scale spatial aggregation encoder, and use the normalized exponential function operator. Calculate the weight of each subspace , represented as: , in, This represents the output features of the multi-scale decoder. This represents the feature transformation matrix corresponding to the multi-scale decoder; Step 4.3: Apply the weights obtained by the multi-scale decoder in different subspaces to the spatial features corresponding to those subspaces, and sum them to obtain the enhanced speech signal features, expressed as: , in, This indicates the features of the enhanced speech signal.

5. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features as described in any one of claims 1 to 4.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-channel speech enhancement method based on the fusion of multi-scale spatial information and spectral features as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Microphone array-oriented channel attention weighted speech enhancement method

    CN112151059A

  • Multi-channel target voice extraction method based on FFC-LSTM and electronic equipment

    CN116863940A