Cascade mask network speech enhancement method fusing auditory process and related device

Through the cascading mask network combining auditory process and speaker embedding technology, the problems of speech clarity and hearing quality in noisy environments are solved, and effective noise suppression and speech enhancement under low signal-to-noise ratio conditions are achieved.

CN120412604AActive Publication Date: 2025-08-01XIANGJIANG LAB

Patent Information

Application Number
CN202510912373.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Existing speech enhancement methods are difficult to effectively suppress non-stationary noise and human voice noise in noisy environments, resulting in poor speech clarity and hearing quality, especially under low signal-to-noise ratio conditions, and excessive voice suppression of target speakers.

Method used

The cascade mask network that integrates the auditory process is adopted to initially enhance the voice waveform through the time domain coarse-grained mask network, and combine the frequency domain fine-grained mask network and the target speaker embed information to generate the ultimate enhanced voice, and use the dual-path sparse attention module and the speaker embedding fusion module to optimize the enhancement effect.

Benefits of technology

Improve the listening quality and clarity of the desired voice in noisy environments, effectively suppress background noise and vocal interference, reduce calculation complexity, and be suitable for resource-constrained environments, and improve the robustness and generalization of voice enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412604A_ABST
    Figure CN120412604A_ABST
Patent Text Reader

Abstract

The invention provides a cascade mask network speech enhancement method fusing an auditory process and a related device, and relates to the technical field of speech signal processing. The method comprises the following steps: acquiring noisy voice, and extracting coarse-grained features from the noisy voice through a time domain coarse-grained mask network to generate a preliminarily enhanced voice waveform; and performing short-time Fourier transform on the preliminarily enhanced voice waveform, and generating a final enhanced voice by combining a frequency domain fine-grained mask network with embedded information of a target speaker. The effect of improving the hearing quality and definition of the expected voice in the noisy environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech signal processing, and in particular, to a cascaded mask network speech enhancement method and related device that integrates the auditory process. Background Art

[0002] In recent years, with the rapid development of deep learning technology, speech enhancement has gradually evolved into a supervised learning task based on deep learning. By inputting noisy speech into the network and learning the non-linear mapping relationship with clean speech, compared with traditional statistical signal processing methods, deep learning-based speech enhancement methods have a stronger ability to suppress non-stationary noise. To improve the speech enhancement performance of the network, deep learning-based speech enhancement methods are developing in the direction of using more powerful neural network basic modules and more effective structural designs. Among them, according to different training objectives, they can be divided into two categories: mask estimation-based methods and mapping-based methods. Compared with mask estimation, the mapping method has unstable regulation in detail values, while the mask method has a limited dynamic range to learn and a faster convergence speed. Therefore, the mask estimation-based enhancement method is superior to the mapping-based enhancement method and has received more attention from researchers. In terms of network structure design, the decoupled speech enhancement framework decouples the original complex spectrum estimation into two subtasks of optimizing amplitude and phase, alleviating the two-way compensation effect between amplitude and phase in single-stage complex domain enhancement methods. Therefore, many time-frequency domain methods adopt the decoupled dual-path framework and its similar architectures. These existing decoupled or single-stage speech enhancement frameworks are designed based on various characteristics of speech signals, such as time-domain characteristics and frequency-domain characteristics, and usually perform better under relatively high signal-to-noise ratio conditions. When the signal-to-noise ratio is extremely low, there are problems such as incomplete noise removal and excessive suppression of the target speaker's speech, which makes the final speech enhancement effect still not good. In terms of the design of basic modules, existing methods usually improve speech enhancement performance by increasing the number of learnable parameters or deepening the hierarchical structure of the neural network, which makes the model parameter quantity too large and the computational complexity too high. At the same time, general speech enhancement models are limited to eliminating environmental noise and cannot effectively eliminate human voice noise in the environment. Therefore, there is an urgent need to propose a speech enhancement method.

[0003] Various types of interference signals will seriously reduce the performance of speech processing-related applications such as automatic speech recognition, online audio and video calls, and hearing aids in real scenarios, such as environmental sounds, human voices, music, indoor reverberation, etc. Therefore, how to improve the listening quality and clarity of the desired speech in a noisy environment has become an urgent technical problem to be solved. Summary of the Invention

[0004] In order to improve the listening quality and clarity of the desired speech in a noisy environment, the present application provides a cascaded mask network speech enhancement method and related device that integrates the auditory process.

[0005] In a first aspect, a cascaded mask network speech enhancement method for integrating the auditory process provided by the present application adopts the following technical solutions: A cascaded mask network speech enhancement method for integrating the auditory process, comprising: Obtain noisy speech, and extract coarse-grained features from the noisy speech through a time-domain coarse-grained mask network to generate a preliminarily enhanced speech waveform; Perform a short-time Fourier transform on the preliminarily enhanced speech waveform, and generate a final enhanced speech by combining target speaker embedding information through a frequency-domain fine-grained mask network; Wherein, the time-domain coarse-grained mask network includes an encoder, an enhancement layer, a mask network and a decoder, and the frequency-domain fine-grained mask network includes an encoder, a dual-path sparse attention module, a speaker embedding fusion module and a decoder; The dual-path sparse attention module models the time-frequency distribution through sparse multi-head self-attention, temporal convolutional attention and frequency convolutional attention, and the speaker embedding fusion module performs cross-attention calculation on the target speaker embedding and intermediate features to optimize the enhancement effect.

[0006] Optionally, the encoder of the time-domain coarse-grained mask network includes a two-dimensional convolutional block and a dense dilated convolutional block for extracting high-level features of the input speech; the enhancement layer uses a grouped bidirectional GRU and a unidirectional GRU to model the intra-frame and inter-frame dependencies respectively; the mask network generates a coarse-grained mask through a parallel two-dimensional grouped convolutional path, multiplies it element-wise with the encoder output, and then reconstructs the preliminarily enhanced speech waveform through the decoder, and the decoder uses a dense dilated convolutional block and a sub-pixel convolution to achieve waveform reconstruction.

[0007] Optionally, the grouped bidirectional GRU of the enhancement layer groups the input features into two parts along the channel dimension, processes them in parallel through a bidirectional GRU respectively, and after the output is concatenated, performs feature enhancement through a one-dimensional grouped convolution and group normalization, and performs a residual connection with the features output by the encoder to alleviate gradient disappearance.

[0008] Optionally, the dual-path sparse attention module of the frequency-domain fine-grained mask network includes: A sparse multi-head self-attention mechanism that respectively adopts local windows and dilated attention along the frequency axis and the time axis to reduce the computational complexity; A temporal convolutional attention module and a frequency convolutional attention module that respectively capture the long-range correlations of the time axis and the frequency axis through a compressed temporal convolutional network, and generate an attention map to multiply with the input features; Wherein, the temporal convolutional attention module performs global average pooling along the frequency axis, and the frequency convolutional attention module performs global average pooling along the time axis.

[0009] Optionally, in the sparse multi-head self-attention mechanism, causal masking and dual masking are used for attention calculation along the time axis. Among them, the look-ahead mask in the dual mask allows the model to fuse limited future context information, and the look-behind mask ensures that the attention calculation only focuses on the information of local adjacent frames.

[0010] Optionally, the speaker embedding fusion module fuses the target speaker embedding vector with the intermediate features through cross-attention calculation, including: Performing cross-attention calculation between the fixed speaker embedding vector and the intermediate features to generate an adaptive speaker embedding; After concatenating the adaptive speaker embedding and the intermediate features, integrating them through two-dimensional pointwise convolution to form fused features; Among them, the target speaker embedding vector is generated by a speaker embedding extraction network and is processed in the compressed ERB domain. The speaker embedding extraction network is stacked by recalibrated encoders. The recalibrated encoder includes a two-dimensional gated linear unit and a U-Net block, and optimizes the distinguishability of the embedding through AAMsoftmax additive angular margin loss.

[0011] Optionally, the frame division parameters of the time-domain coarse-grained masking network are a frame length of 640 samples and a frame shift of 320 samples; The frame division parameters of the frequency-domain fine-grained masking network are a frame length of 512 samples and a frame shift of 256 samples, and a square root Hanning window is used for short-time Fourier transform.

[0012] In a second aspect, the present application proposes a cascaded masking network speech enhancement device for a fusion auditory process, which executes the method described above, including: A speech acquisition module, configured to acquire noisy speech, and extract coarse-grained features from the noisy speech through a time-domain coarse-grained masking network to generate a preliminarily enhanced speech waveform; A speech enhancement module, configured to perform short-time Fourier transform on the preliminarily enhanced speech waveform, and generate a final enhanced speech by combining target speaker embedding information through a frequency-domain fine-grained masking network.

[0013] In a third aspect, the present application provides a computer device, which includes: a memory, a processor, and when the processor runs computer instructions stored in the memory, it executes the method described above.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, including instructions, and when the instructions run on a computer, the computer is made to execute the method described above.

[0015] In summary, the present application obtains noisy speech, extracts coarse-grained features from the noisy speech through a time-domain coarse-grained mask network to generate a preliminarily enhanced speech waveform; performs a short-time Fourier transform on the preliminarily enhanced speech waveform, and combines target speaker embedding information through a frequency-domain fine-grained mask network to generate a final enhanced speech. The effect of improving the listening quality and clarity of the desired speech in a noisy environment is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic structural diagram of a computer device for the hardware operating environment involved in the solution of the embodiment of the present application; Figure 2 is a schematic flowchart of the first embodiment of the cascade mask network speech enhancement method that integrates the auditory process of the present application; Figure 3 is an overall network structure diagram of the cascade mask network speech enhancement method that integrates the auditory process of the present application; Figure 4 is a schematic internal structure diagram of the dual-path sparse attention module of the present application; Figure 5 is a schematic internal structure diagram of the time convolution attention and frequency convolution attention modules of the present application; Figure 6 is a structural block diagram of the first embodiment of the cascade mask network speech enhancement device that integrates the auditory process of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0018] Refer to Figure 1 , Figure 1 is a schematic structural diagram of a computer device for the hardware operating environment involved in the solution of the embodiment of the present application.

[0019] As Figure 1As shown in the figure, a computer device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0020] Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or have a different component layout.

[0021] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and a cascaded mask network voice enhancement program for fusing the auditory process.

[0022] In Figure 1 the computer device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in this application may be set in the computer device. The computer device calls the cascaded mask network voice enhancement program stored in the memory 1005 through the processor 1001 and executes the cascaded mask network voice enhancement method provided in the embodiments of this application.

[0023] The embodiments of this application provide a cascaded mask network voice enhancement method for fusing the auditory process. Referring to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the cascaded mask network voice enhancement method for fusing the auditory process in this application.

[0024] In this embodiment, the cascaded mask network voice enhancement method for fusing the auditory process includes the following steps: Step S10: Obtain noisy speech, and extract coarse-grained features from the noisy speech through a time-domain coarse-grained mask network to generate a preliminarily enhanced speech waveform.

[0025] The encoder of the time-domain coarse-grained mask network includes a two-dimensional convolutional block and a dense dilated convolutional block for extracting high-level features of the input speech; the enhancement layer uses grouped bidirectional GRU and unidirectional GRU to model the intra-frame and inter-frame dependencies respectively; the mask network generates a coarse-grained mask through a parallel two-dimensional grouped convolutional path, multiplies it element-wise with the encoder output, and then reconstructs the preliminarily enhanced speech waveform through a decoder. The decoder uses a dense dilated convolutional block and sub-pixel convolution to achieve waveform reconstruction.

[0026] It can be understood that the grouped bidirectional GRU of the enhancement layer groups the input features into two parts along the channel dimension, processes them in parallel through bidirectional GRU respectively, performs feature enhancement on the output concatenation through one-dimensional grouped convolution and group normalization, and performs a residual connection with the features output by the encoder to alleviate gradient disappearance.

[0027] In a specific implementation, in this embodiment, a time-domain coarse-grained mask network is constructed. In order to simulate the process of the human ear "roughly listening" in a noisy environment in the first stage, the time-domain frame length is longer than the frequency-domain frame length in the second stage. Specifically, the noisy speech undergoes coarse-grained sequence framing preprocessing before entering the encoder. Each frame covers more time samples, and each mask value is calculated based on longer-time signal information, thereby achieving a coarser-grained analysis. Although the time resolution is reduced, the model can capture the signal features over a longer time, such as the overall energy change of the speech signal, the speech envelope, etc., reflecting the overall trend and dynamic changes of the speech signal. The time-domain coarse-grained mask network in the first stage includes three modules: an encoder, an enhancement layer, a mask network, and a decoder. The encoder uses a two-dimensional convolutional block and a dense dilated convolutional block to learn the encoded representation of the input speech signal. The enhancement layer uses two grouped DPRNNs to model the intra-frame and inter-frame dependencies. The mask network estimates the optimal mask based on the features learned by the enhancement layer to enhance the noisy speech, including two two-dimensional convolutions and two parallel two-dimensional grouped convolutional paths. The decoder applies a dense dilated convolutional block, sub-pixel convolution, and two-dimensional pointwise convolution to reconstruct the speech features, and reconstructs the preliminarily enhanced speech waveform through overlap-and-add.

[0028] In a specific implementation, constructing a time-domain coarse-grained mask network includes: Perform coarse-grained framing on the input long speech sequence, stack the sequence frames to form a 3D tensor and input it into the encoder; Use the encoder to obtain the learned signal representation S of the input sequence frames; Use the enhancement layer to model the intra-frame dependencies.

[0029] Model the dependencies between frames in the enhancement layer. The process is similar to modeling the intra-frame dependencies in the enhancement layer, except for the dimension transformation and using a unidirectional GRU as the basic modeling module. That is, when flattening the tensor, merge dimension B and dimension F to prepare for the inter-frame RNN processing; The mask network uses two-dimensional convolution to double the output of the enhancement layer along the channel dimension to obtain features , and combines the dual-path two-dimensional grouped convolution with different activation functions to obtain features by integrating the advantages of the two activation functions ; Features Are processed through two-dimensional grouped convolution and the ReLU activation function, and use non-linear transformation to extract and learn the feature representation of the target speaker, and then generate a coarse-grained mask of the target speaker; Multiply the obtained mask element-wise with the encoder output , and obtain the preliminary enhanced features of the target speaker through the decoder, and use overlap and add to reconstruct the preliminarily enhanced speech waveform.

[0030] Among them, the steps of using the enhancement layer to model the intra-frame dependencies specifically also include: Rearrange and reshape the learning signal representation S, that is, rearrange the last two dimensions, the number of frames N and the frame length F, in front of the channel dimension C in the 3D tensor, and then flatten the tensor into (B×N, F, C); Use grouped bidirectional GRU for intra-frame modeling. The reshaped encoded representation Is grouped along the last dimension C to obtain , , and the segmented And Correspond to different subsets of the input features respectively. This helps to reduce the computational burden of a single GRU layer and allows the model to process different feature groups in parallel. At the same time, after the hidden state Is initialized as a all-zero tensor, the hidden state Is segmented into And , corresponding to the initial hidden states of rnn1 and rnn2 respectively. The segmentation method is the same as that of the input , that is, divide Into two parts along the last dimension; Use rnn1 to process the input And , and obtain the output And the new hidden state ; Use rnn2 to process the input And , and obtain the output And the new hidden state . Two GRU layers can process different feature groups in parallel, thus improving the computational efficiency; Concatenate and along the last dimension. The new hidden state and are concatenated to obtain the final output and the final hidden state ; Rearrange the output after intra-frame modeling. Use one-dimensional grouped convolution to perform linear transformation and feature extraction on the rearranged output to enhance the feature expression ability, and then transpose the features after the linear transformation to restore the original feature dimension; Perform group normalization on the features after convolution. Stabilize the training process through group normalization and accelerate the model convergence; After group normalization and the rearranged input encoding representation perform residual connection to promote information flow and alleviate the problem of gradient disappearance.

[0031] The time-domain coarse-grained mask network described in this embodiment is set as follows: The frame length and frame shift for splitting the long speech sequence are 640 (40 ms) and 320 (20 ms) respectively. The sampling rate of the noisy speech is 16000 Hz. The encoder uses two-dimensional pointwise convolution to double the number of channels to 64, and uses dense dilated convolution blocks with dilation sizes of 1, 2, 4, and 8 to gradually expand the receptive field to extract high-level features. The dense dilated convolution block introduces depthwise separable convolution to replace the standard convolution for pruning. The decoder uses depthwise separable convolution to downsample the frame length F, with a convolution kernel size of (1, 3) and a stride of (1, 2). Sub-pixel convolution in the decoder is used to perform upsampling to restore the dimension to F. The input feature of the enhancement layer is , and the grouped bidirectional GRU models the local dependencies. Before input, it undergoes rearrangement and reshaping to obtain the feature , and the specific formula is: where is the output of the grouped bidirectional GRU, and are the mapping functions defined by the bidirectional GRU respectively, and are features grouped along the last dimension, and are the features of the hidden state after grouping. The hidden state Initialized as a tensor of all zeros, the one-dimensional grouped convolution pair Performs linear transformation and feature extraction. Group normalization divides the channels into 4 groups and calculates the mean and variance within each group for normalization. The specific formula is: Where And Are the mean and variance of all elements within the group; Is a small constant used to prevent division by zero errors; Is the scaling factor; Is the offset; Is the feature input to group normalization. The inter-frame global dependency modeling is based on a unidirectional GRU implementation. The whole process is the same as the local dependency modeling. Among them, when reordering and reshaping the input features, the obtained feature is The coarse-grained mask of the target speaker generated by the mask network is multiplied element-wise with the encoder output.

[0032] Step S20: Perform short-time Fourier transform on the preliminarily enhanced speech waveform, and generate the final enhanced speech by combining the target speaker embedding information through the frequency-domain fine-grained mask network; Among them, the time-domain coarse-grained mask network includes an encoder, an enhancement layer, a mask network, and a decoder. The frequency-domain fine-grained mask network includes an encoder, a dual-path sparse attention module, a speaker embedding fusion module, and a decoder.

[0033] In a specific implementation, this embodiment constructs a frequency-domain fine-grained mask network, uses a shorter frame length for short-time Fourier transform in the second stage, and reduces the redundancy of the input features using an equivalent rectangular bandwidth filter bank before entering the encoder. The encoder and decoder are composed of two two-dimensional convolutions and three grouped temporal convolutional blocks. The grouped temporal convolutional block is composed of two two-dimensional pointwise convolutions and two-dimensional depthwise dilated convolutions. At the same time, skip connections are made between the encoder and the decoder to reduce information loss. The enhancement layer uses a dual-path sparse attention module including sparse multi-head self-attention, temporal convolutional attention, and frequency convolutional attention to model the time-frequency distribution. The fine-grained mask generated by the decoder further refines the signal on the basis of the preliminary enhancement. In the second stage, in addition to the three parts of the encoder, the enhancement layer, and the decoder, a speaker fusion module is introduced after the enhancement layer to fuse the target speaker embedding with the intermediate features of the fine-grained mask network. The target speaker embedding is output after the speaker embedding extraction network processes the target speaker's registered speech.

[0034] In a specific implementation, the steps of constructing the frequency-domain fine-grained mask network include: The short-time Fourier transform is used to obtain the spectrum Z of the preliminarily enhanced speech. The frame length is shorter than that in the first stage. Z is decoupled to obtain the original complex spectrum Y and the original magnitude spectrum Y0. The magnitude, real part, and imaginary part are stacked into a new feature tensor and input into the frequency-domain fine-grained mask network; ERB filters are used to merge features in the high-frequency band above 2 kHz, effectively reducing the computational burden of the model while ensuring performance; The encoder extracts features from the feature map, doubling the number of channels from 16 to 32 and downsampling along the frequency direction; The dual-path sparse attention module models the intra-frame spectral patterns and long-range inter-frame dependencies in the output of the encoder along the frequency axis and the time axis respectively, enhancing the model's ability to model the entire time-frequency representation.

[0035] The speaker embedding fusion module fuses the output of the dual-path sparse attention module with the target speaker embedding. The fixed speaker embedding vector and perform cross-attention calculation. The attention weights reflect the correlation between each time-frequency unit and the fixed speaker embedding. The output of the cross-attention is connected to the original speaker embedding through a residual connection to obtain an adaptive speaker embedding. The residual connection is used to aggregate the adaptive embedding, and the adaptive speaker embedding is directly connected to the intermediate features for fusion. The speaker embedding extraction network consists of stacked recalibrated encoders, reducing information loss while expanding the feature scale. The entire network is executed in the compressed ERB domain, not only occupying less computational resources but also generating speaker embeddings that can be optimized for the speech enhancement task.

[0036] The decoder upsamples the frequency dimension layer by layer, reducing the number of channels from 32 back to 16, corresponding to the downsampling operation of the encoder in the step "The encoder extracts features from the feature map, doubling the number of channels from 16 to 32 and downsampling along the frequency direction", completing the reconstruction of the features.

[0037] ERB filters are used to split the frequency bands of the decoder output to restore the original resolution, corresponding to the downsampling operation in the step "ERB filters are used to merge features in the high-frequency band above 2 kHz, effectively reducing the computational burden of the model while ensuring performance".

[0038] The dual-path sparse attention module models the time-frequency distribution through sparse multi-head self-attention, temporal convolutional attention, and frequency convolutional attention. The speaker embedding fusion module performs cross-attention calculation between the target speaker embedding and the intermediate features to optimize the enhancement effect.

[0039] Among them, the dual-path sparse attention module models the in-frame spectral patterns and long-range inter-frame dependencies of the encoder output along the frequency axis and the time axis respectively, enhancing the model's ability to model the entire time-frequency representation. The steps specifically include: First, model the spectral patterns within the frame along the frequency axis. The sparse multi-head self-attention only focuses on adjacent frequency bands at specific frequencies within the local window size, and dilated attention is used outside the window. The dilated attention step size is , and for frequency bands with an interval of and multiples outside the window are attended to, enabling the model to effectively filter or emphasize specific frequency components, reducing the redundant calculations in the full multi-head attention mechanism and lowering the model complexity; The features output by the sparse multi-head self-attention mechanism are linearly transformed through a linear layer to enhance the feature expression ability. The output of the reshaping and rearrangement linear layer is normalized by calculating the mean and variance independently along the channel dimension through an instance normalization layer to stabilize the training process; Add the normalized features to the encoder output to form a residual connection, which helps the model better transmit gradients and alleviate the vanishing gradient problem in deep networks; Rearrange the output after the residual connection and input it into the frequency convolutional attention module. The frequency convolutional attention module uses a compressed temporal convolutional network (S-TCM) to capture long-range correlations. Global average pooling is performed along the time axis before inputting to the S-TCM. The initial one-dimensional convolution in the S-TCM is used for channel mixing, and the dilated convolution is used to explore long-range sequence relationships. The final one-dimensional convolution projects the number of output channels back to the original size. A gated branch is added in parallel to the main dilated convolution branch. This gated branch uses the sigmoid function to adjust its value, which helps the flow of gradients during backpropagation and alleviates the vanishing gradient problem. After generating the attention map using the sigmoid activation function, the attention map is replicated along the time axis and multiplied by the output after the residual connection to obtain the features ; Model the long-range inter-frame dependencies along the time axis using causal sparse multi-head self-attention. The look-ahead mask allows the first layer of attention to fuse limited future context information, and the look-back mask ensures that the attention calculation only focuses on the information of local adjacent frames. The look-ahead mask is removed from the remaining layers, and a causal mask is set to prevent further expansion of the future receptive field. Compared with the full multi-head self-attention mechanism, it ensures that the model runs in a streaming manner and can focus on local time features, improving efficiency and relevance; The output after residual connection is rearranged and input into the temporal convolutional attention module. The temporal convolutional attention module has the same structure as the frequency convolutional attention module, but global average pooling is performed along the frequency axis before inputting to the S-TCM. Finally, after applying the sigmoid activation function to generate the attention map, the attention map is replicated along the frequency axis and multiplied by the output after residual connection to obtain the features , generating an attention map to guide the model to focus on "which time frame" and effectively modeling the energy distribution on the time axis.

[0040] In this embodiment, the auditory process of the human ear is incorporated into the speech enhancement framework. The time-domain coarse-grained mask network analyzes the overall waveform shape, temporal structure, and energy change of the speech signal to obtain a preliminary enhanced signal dominated by the target speaker. In the frequency-domain enhancement stage, speaker embeddings are introduced as prior information, and the generated fine-grained mask performs more refined enhancement. The overall framework design follows the "coarse-to-fine" auditory process, improving the problems of incomplete noise removal and over-suppression of the target speaker's speech in existing speech enhancement methods at low signal-to-noise ratios. It can not only eliminate background noise but also effectively suppress other interfering speakers. A new dual-path sparse attention module is proposed in the frequency-domain fine-grained mask network. The temporal convolutional attention module and the frequency convolutional attention module can effectively model the energy distribution on the time axis and the frequency axis. The sparse multi-self-attention mechanism significantly reduces the number of parameters and computational overhead while enhancing the model's expressive ability compared to the full attention mechanism. The dual-mask design along the time axis also ensures that the model can be executed in a streaming manner. The overall setting of the module makes the model more efficient, lightweight, and suitable for resource-constrained environments. The present invention provides a new perspective for the construction of the speech enhancement framework, and the overall model improves the speech enhancement effect while maintaining a small number of parameters and computational complexity.

[0041] It should be noted that the dual-path sparse attention module of the frequency-domain fine-grained mask network includes: a sparse multi-head self-attention mechanism that respectively adopts local windows and dilated attention along the frequency axis and the time axis to reduce the computational complexity; a temporal convolutional attention module and a frequency convolutional attention module that respectively capture the long-range correlations of the time axis and the frequency axis through compressed temporal convolutional networks and generate attention maps to multiply with the input features; among them, the temporal convolutional attention module performs global average pooling along the frequency axis, and the frequency convolutional attention module performs global average pooling along the time axis.

[0042] In a specific implementation, in the sparse multi-head self-attention mechanism, causal masks and dual masks are set for the attention calculation along the time axis. Among them, the look-ahead mask in the dual mask allows the model to fuse limited future context information, and the look-behind mask ensures that the attention calculation only focuses on the information of local adjacent frames.

[0043] It should be noted that the speaker embedding fusion module fuses the target speaker embedding vector with the intermediate features through cross-attention calculation, including: cross-attention calculation between the fixed speaker embedding vector and the intermediate features to generate an adaptive speaker embedding; after the adaptive speaker embedding and the intermediate features are concatenated, they are integrated through two-dimensional pointwise convolution to form fused features; among them, the target speaker embedding vector is generated by the speaker embedding extraction network and processed in the compressed ERB domain. The speaker embedding extraction network is stacked by recalibrated encoders. The recalibrated encoder includes a two-dimensional gated linear unit and a U-Net block, and optimizes the distinguishability of the embedding through the AAMsoftmax additive angular margin loss.

[0044] In a specific implementation, the frame division parameters of the time-domain coarse-grained masking network are a frame length of 640 samples and a frame shift of 320 samples; the frame division parameters of the frequency-domain fine-grained masking network are a frame length of 512 samples and a frame shift of 256 samples, and a square root Hanning window is used for short-time Fourier transform.

[0045] The overall structure diagram of this embodiment is as Figure 3 shown. The model includes three parts: a time-domain coarse-grained masking network, a frequency-domain fine-grained masking network, and a speaker embedding extraction network. The time-domain coarse-grained masking network includes four modules: an encoder, an enhancement layer, a masking network, and a decoder. The encoder uses two-dimensional convolutional blocks and dense dilated convolutional blocks to learn the encoded representation of the input speech signal. The enhancement layer uses two grouped DPRNNs to model the intra-frame and inter-frame dependency relationships. The masking network estimates the optimal mask based on the features learned by the enhancement layer to enhance the noisy speech, including two two-dimensional convolutions and two parallel two-dimensional grouped convolution paths. The decoder applies dense dilated convolutional blocks, sub-pixel convolution, and two-dimensional pointwise convolution to reconstruct the speech features and reconstruct the preliminarily enhanced speech waveform through overlap and add; the frequency-domain fine-grained masking network includes an encoder, an enhancement layer, a speaker embedding fusion module, and a decoder. The encoder and the decoder are composed of two two-dimensional convolutional blocks and three grouped temporal convolutional blocks. The grouped temporal convolutional block is composed of two two-dimensional pointwise convolutions and two-dimensional depth dilated convolution. Skip connections are made between the encoder and the decoder to reduce information loss. The enhancement layer uses a dual-path sparse attention module including sparse multi-head self-attention, temporal convolutional attention, and frequency convolutional attention to model the time-frequency distribution. The speaker embedding fusion module calculates the cross-attention between the fixed speaker embedding vector and the intermediate features, and fuses the generated target speaker adaptive embedding with the intermediate features; the speaker embedding extraction network is composed of stacked recalibrated encoders. Through the combination of continuous downsampling and upsampling operations, different scales of information can be effectively utilized without generating excessive computational load.

[0046] In a specific implementation, the specific enhancement process of the frequency-domain fine-grained masking network in this embodiment is as follows: The time-frequency spectrum of the first-stage time-domain preliminary enhanced signal is obtained through short-time Fourier transform. The frame length and frame shift are 512 (32 ms) and 256 (16 ms) respectively, and the window function is the square root of the Hann window. The real part and imaginary part are extracted from the input spectral data, and the amplitude, real part, and imaginary part are stacked together to form a new feature tensor and input it into the network. The second-stage fine-grained enhancement is performed in a compressed domain, and a filter is used to compress the frequency band above 2 kHz. The specific formula for converting Hertz to scale is: where is the center frequency (kHz) of the filter, and is the bandwidth (Hz) of the filter.

[0047] The first two two-dimensional convolutional blocks in the encoder include two-dimensional convolution, batch normalization, and the PReLU activation function. The two-dimensional convolutional kernel has a size of (1, 5), and downsampling is performed in the width direction while increasing the number of channels from 9 to 32. The last three grouped temporal convolutional blocks gradually expand the receptive field through pointwise convolution and depthwise separable convolution, thereby capturing longer-range temporal dependencies and enhancing the feature representation ability.

[0048] The dual-path sparse attention module models the spectral patterns within a single frame and the long-range dependencies between frames along the frequency axis and time axis respectively. The internal structure of the dual-path sparse attention module is as Figure 4 shown. The output of the encoder has dimensions B×C×T×F, and after rearrangement and reshaping to , it is input into the sparse multi-head self-attention module with dimensions BT×F×C. Three matrices , , are defined to perform three linear transformations on respectively, obtaining new vectors , , . The obtained vectors are respectively concatenated into a large matrix, denoted as , , , which correspond to the query matrix, key matrix, and value matrix respectively. The query vector is successively multiplied by the key matrix, and the relevant attention weights are obtained. The sparse attention mask is applied to the attention weights, and the obtained value is then multiplied by the corresponding value vector respectively. The specific formula is: where is the adjustment factor. The sparse multi-head self-attention mask value is 0 or 1. When the mask value is 0, the frame focus frame and calculate the corresponding attention score. When the mask value is 1, the attention score is masked or discarded. The specific formula is: In sparse multi-head self-attention, local attention is used to control each frame to focus on the frames within the context window, and dilated attention calculates the distance between the frames outside the local context window for the positions of several time frames. The mask is applied to the similarity matrix. Assume is the index distance. For the local attention mask, when the index distance is exceeded, by setting it to 0, the weight of this part becomes extremely small and hardly affects the result. The specific formula is: For the dilated attention mask, the attention calculation is effective when the index distance is a multiple of, and the specific formula is: A linear layer is used to perform a linear transformation on the output of the sparse multi-head attention module. The output of the linear layer is normalized by calculating the mean and variance independently along the channel dimension by the instance normalization layer. The normalized features are connected to the encoder output with a residual connection.

[0049] The frequency convolution attention module models the energy distribution in the frequency dimension. The internal structures of the frequency convolution attention module and the time convolution attention module are as Figure 5 shown. Frequency convolution attention will perform global average pooling along the time axis, that is, after the model "scans" the entire time axis, the overall statistical features of the frequency dimension are obtained. The specific formula is: where is the frequency energy representation, represents the number of channels, represents the time step, is the frequency dimension. The obtained frequency energy representation is used by the frequency convolution attention module to capture long-range correlations. The frequency convolution attention module sets the dilation rate of the dilated convolution in S-TCM to 5, where the type of the normalization layer is the immediate layer normalization, which normalizes independently in the channel dimension for each time step. After applying the sigmoid activation function to generate the attention map in the frequency dimension, the attention map is replicated along the time axis and multiplied by the input of the frequency convolution attention module.

[0050] In the first layer of attention calculation, the sparse multi-head self-attention adds a double-mask setting. That is, the look-ahead mask allows the model to fuse limited future context information, and the look-back mask ensures that the attention calculation only focuses on the information of local adjacent frames. In the remaining layers, a causal mask is set and the look-ahead mask is cancelled to prevent the number of future frames relied on by the model from growing linearly due to the fixed look-ahead as the number of attention calculation layers increases. Additionally, the frequency convolution attention module is replaced with a temporal convolution attention module. The specific formula for global average pooling along the frequency axis is: where is the temporal energy representation. In the temporal convolution attention module, the dilation rate of the dilated convolution in S-TCM is set to 2. After applying the sigmoid activation function to generate the attention map in the temporal dimension, the attention map is replicated along the frequency axis and multiplied by the input of the temporal convolution attention module.

[0051] The speaker embedding fusion module performs cross-attention calculation on the fixed speaker embedding vector and the intermediate features generated by the frequency-domain fine-grained mask network, where and are the intermediate features, is the speaker embedding. The single speaker embedding calculates the dot product with the intermediate features at each time step to obtain an attention score vector with a length equal to the sequence length, representing the similarity between the fixed speaker embedding and the intermediate features at each time step. According to the matching degree between the intermediate features at each time step and the speaker embedding, the model dynamically adjusts the feature representation at that time step. The time step features highly correlated with the speaker embedding will be assigned higher weights. The output features after cross-attention calculation are connected with the original speaker embedding through a residual connection to obtain an adaptive speaker embedding. The specific formula is: In the formula, are the intermediate features, is the speaker embedding vector, are learnable parameters used to control the influence intensity of the cross-attention correction term on the speaker embedding vector . The generated adaptive embedding is concatenated and fused with the intermediate features of the original backbone network. The specific formula is: In the formula, represents two-dimensional pointwise convolution, which effectively integrates the intermediate features and the adaptive embedding.

[0052] The decoder adds the feature map of the encoder to the output of the current decoder layer through skip connections corresponding to the layers of the encoder. In the first three grouped temporal convolutional blocks, pointwise convolution and depthwise separable convolution with different dilation rates (5, 2, and 1 respectively) are used to gradually expand the receptive field for feature reconstruction. Transposed convolution is used to upsample the feature map. In the last two 2D convolutional blocks, 2D convolution with a kernel size of (1, 5) is used to reduce the number of channels from 32 to 16, and then from 16 to 2, and upsampling is performed with a stride of 2 in the frequency dimension.

[0053] In the embodiment, the speaker embedding extraction network consists of four stacked recalibrated encoders. The speaker's registered speech undergoes short-time Fourier transform to obtain the time-frequency spectrum. The real part, imaginary part, and magnitude are stacked as a new feature tensor and input into the speaker embedding extraction network. The input spectral features are downsampled by ERB filters and then passed through the recalibrated encoder to obtain the speaker embedding. The recalibrated encoder includes a 2D gated linear unit and a U-Net block. The speaker embedding encodes the key information of a specific speaker. Therefore, in addition to calculating the loss between the clean speech and the enhanced speech, a classification task loss function, AAMsoftmax additive angular margin loss, is set to enhance the discriminability of the embedding. This loss function solves the obvious ambiguity of the softmax loss at the decision boundary by maximizing the inter-class distance and minimizing the intra-class distance.

[0054] In the specific implementation, the beneficial effects that can be finally achieved by this embodiment through the above description are: 1. Improve the speech enhancement effect in a low signal-to-noise ratio environment Technical means: The time-domain coarse-grained mask network captures the overall energy change and envelope characteristics of the speech signal through long-frame segmentation preprocessing (40 ms frame length), and the frequency-domain fine-grained mask network refines the modeling of the time-frequency distribution by short-frame segmentation (32 ms frame length) combined with a dual-path sparse attention module.

[0055] Logical deduction: In the coarse-grained stage, noise is initially suppressed and the dominant signal of the target speech is enhanced. In the fine-grained stage, residual noise is further eliminated and the speech details are optimized. The cascading of the two alleviates the problems of noise residue and over-suppression of the target speech under low signal-to-noise ratio.

[0056] 2. Effectively suppress background noise and human voice interference Technical means: The speaker embedding fusion module introduces the target speaker embedding information and dynamically adjusts the weights of the time-frequency units through cross-attention to enhance the ability to focus on a specific speaker.

[0057] Logical deduction: Traditional methods are difficult to distinguish the target speech from interfering human voices. The speaker embedding, as prior information, guides the model to focus on the target sound source. Combined with the refined processing of the frequency-domain fine-grained mask, it significantly suppresses the cross-interference in a multi-human voice environment.

[0058] 3. Reduce computational complexity and support streaming processing Technical means: The dual-path sparse attention module combines local windows with dilated attention, and sparse multi-head self-attention reduces redundant calculations; the temporal convolutional attention module introduces a causal mask to limit future frame dependencies.

[0059] Logical deduction: The sparse attention mechanism reduces the computational complexity from O(n²) to O(n). The causal mask ensures that the model only relies on historical information, meeting the requirements of real-time streaming processing and being applicable to scenarios such as online audio and video calls.

[0060] 4. Lightweight model design for resource-constrained environments Technical means: Technologies such as grouped bidirectional GRU, depthwise separable convolution, and ERB filters are used to compress high-frequency redundant features to optimize the number of parameters; the speaker embedding extraction network uses compressed ERB domain processing to reduce the computational load.

[0061] Logical deduction: The modular lightweight design reduces memory occupancy and computational resource consumption while ensuring performance, enabling the model to be deployed on mobile devices or embedded systems (such as hearing aids).

[0062] 5. Improve the robustness and generalization of speech enhancement Technical means: Technologies such as residual connections, group normalization, and instance normalization stabilize the training process; the AAMsoftmax loss function optimizes the inter-class discriminability of speaker embeddings.

[0063] Logical deduction: Normalization and residual connections alleviate the vanishing gradient problem and improve the convergence stability of the model; the high discriminability of speaker embeddings enhances the model's adaptability to diverse noises and speaker variations.

[0064] In this embodiment, noisy speech is obtained, and coarse-grained features are extracted from the noisy speech through a time-domain coarse-grained mask network to generate a preliminarily enhanced speech waveform; the preliminarily enhanced speech waveform is subjected to a short-time Fourier transform, and the final enhanced speech is generated through a frequency-domain fine-grained mask network combined with target speaker embedding information. The effect of improving the listening quality and clarity of the desired speech in a noisy environment is achieved.

[0065] In addition, an embodiment of the present application also proposes a computer-readable storage medium, on which a program for speech enhancement of a cascaded mask network integrating the auditory process is stored. When the program for speech enhancement of the cascaded mask network integrating the auditory process is executed by a processor, the steps of the method for speech enhancement of the cascaded mask network integrating the auditory process as described above are implemented.

[0066] Refer to Figure 6 , Figure 6This is the structural block diagram of the first embodiment of the cascaded mask network speech enhancement device that integrates the auditory process in the present application.

[0067] As Figure 6 shown, the cascaded mask network speech enhancement device that integrates the auditory process proposed in the embodiment of the present application includes: A speech acquisition module 10, configured to acquire noisy speech, and extract coarse-grained features from the noisy speech through a time-domain coarse-grained mask network to generate a preliminarily enhanced speech waveform; A speech enhancement module 20, configured to perform a short-time Fourier transform on the preliminarily enhanced speech waveform, and generate a final enhanced speech by combining the target speaker embedding information through a frequency-domain fine-grained mask network.

[0068] It should be understood that the above is only an example for illustration, and does not constitute any limitation to the technical solution of the present application. In specific applications, those skilled in the art can set according to needs, and the present application does not make any restrictions.

[0069] In this embodiment, by acquiring noisy speech, extracting coarse-grained features from the noisy speech through a time-domain coarse-grained mask network to generate a preliminarily enhanced speech waveform; performing a short-time Fourier transform on the preliminarily enhanced speech waveform, and generating a final enhanced speech by combining the target speaker embedding information through a frequency-domain fine-grained mask network. The effect of improving the listening quality and clarity of the desired speech in a noisy environment is achieved.

[0070] It should be noted that the above-described work process is only illustrative and does not constitute a limitation to the protection scope of the present application. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no restrictions are made here.

[0071] In addition, for the technical details not described in detail in this embodiment, reference can be made to the method for cascaded mask network speech enhancement that integrates the auditory process provided in any embodiment of the present application, which will not be elaborated here.

[0072] In addition, it should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.

[0073] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0074] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as Read-Only Memory (ROM) / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application. The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A cascaded mask network speech enhancement method integrating the auditory process, characterized in that, Including: Obtain noisy speech, and extract coarse-grained features from the noisy speech through a time-domain coarse-grained mask network to generate a preliminarily enhanced speech waveform; Perform short-time Fourier transform on the preliminarily enhanced speech waveform, and generate a final enhanced speech by combining target speaker embedding information through a frequency-domain fine-grained mask network; Wherein, the time-domain coarse-grained mask network includes an encoder, an enhancement layer, a mask network and a decoder, and the frequency-domain fine-grained mask network includes an encoder, a dual-path sparse attention module, a speaker embedding fusion module and a decoder; The dual-path sparse attention module models the time-frequency distribution through sparse multi-head self-attention, temporal convolutional attention and frequency convolutional attention, and the speaker embedding fusion module performs cross-attention calculation on the target speaker embedding and intermediate features to optimize the enhancement effect.

2. The method according to claim 1, characterized in that, The encoder of the time-domain coarse-grained mask network includes a two-dimensional convolutional block and a dense dilated convolutional block for extracting high-level features of the input speech; the enhancement layer uses grouped bidirectional GRU and unidirectional GRU to model the intra-frame and inter-frame dependencies respectively; the mask network generates a coarse-grained mask through a parallel two-dimensional grouped convolutional path, multiplies it element-wise with the encoder output, and then reconstructs the preliminarily enhanced speech waveform through the decoder. The decoder uses a dense dilated convolutional block and a sub-pixel convolution to achieve waveform reconstruction.

3. The method according to claim 2, characterized in that The grouped bidirectional GRU of the enhancement layer groups the input features into two parts along the channel dimension, processes them in parallel through bidirectional GRU respectively, and after the output is concatenated, performs feature enhancement through one-dimensional grouped convolution and group normalization, and performs residual connection with the features output by the encoder to alleviate gradient disappearance.

4. The method according to claim 1, wherein The dual-path sparse attention module of the frequency-domain fine-grained mask network includes: A sparse multi-head self-attention mechanism that respectively uses local windows and dilated attention along the frequency axis and the time axis to reduce the computational complexity; A temporal convolutional attention module and a frequency convolutional attention module that respectively capture long-range correlations of the time axis and the frequency axis through a compressed temporal convolutional network, and generate an attention map to multiply with the input features; Wherein, the temporal convolutional attention module performs global average pooling along the frequency axis, and the frequency convolutional attention module performs global average pooling along the time axis.

5. The method according to claim 4, wherein In the sparse multi-head self-attention mechanism, causal mask and double mask settings are adopted for the attention calculation along the time axis. Among them, the look-ahead mask in the double mask allows the model to fuse limited future context information, and the look-back mask ensures that the attention calculation only focuses on the information of local adjacent frames.

6. The method according to claim 1, characterized in that, The speaker embedding fusion module fuses the target speaker embedding vector and intermediate features through cross-attention calculation, including: Performing cross-attention calculation on the fixed speaker embedding vector and intermediate features to generate an adaptive speaker embedding; After the adaptive speaker embedding and intermediate features are concatenated, they are integrated through two-dimensional pointwise convolution to form a fused feature; Among them, the target speaker embedding vector is generated by a speaker embedding extraction network and processed in the compressed ERB domain. The speaker embedding extraction network is stacked by recalibrated encoders. The recalibrated encoder includes a two-dimensional gated linear unit and a U-Net block, and optimizes the distinguishability of the embedding through the AAMsoftmax additive angular margin loss.

7. The method according to claim 1, characterized in that, The frame division parameters of the time-domain coarse-grained mask network are a frame length of 640 samples and a frame shift of 320 samples; The frame division parameters of the frequency-domain fine-grained mask network are a frame length of 512 samples and a frame shift of 256 samples, and a square root Hanning window is used for short-time Fourier transform.

8. A cascaded mask network speech enhancement device integrating an auditory process, characterized in that, Performing the method according to claim 1, comprising: A voice acquisition module, configured to acquire noisy speech, and extract coarse-grained features from the noisy speech through a time-domain coarse-grained mask network to generate a preliminarily enhanced speech waveform; A voice enhancement module, configured to perform short-time Fourier transform on the preliminarily enhanced speech waveform, and generate a final enhanced speech by combining the target speaker embedding information through a frequency-domain fine-grained mask network.

9. A computer device, characterized in that, The device includes: a memory and a processor. When the processor runs computer instructions stored in the memory, it executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Including instructions, when the instructions run on a computer, the computer is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Single-channel speech enhancement method based on waveform spectrum fusion network

    CN116682444A

  • Audio enhancement method and apparatus, and electronic device and readable storage medium

    WO2023226839A1

Cited By

  • Voice signal denoising method and device, electronic equipment and program product

    CN121306167A

  • Sound signal periodic feature extraction method, network model training method, storage medium and equipment

    CN121617417A

  • Target voice extraction method and system and medium

    CN121811905A