Lightweight single-channel speech enhancement method based on improved convolutional recurrent network

By improving the convolutional recurrent network, combining encoder, decoder, aggregated packet dual-path recurrent network and convolutional hybrid packet dual-path recurrent neural network, the convolutional hybrid packet dual-path recurrent neural network solves the confinement of traditional models in time-frequency dynamic modeling and feature space integration, achieving better speech enhancement performance and model lightweighting.

CN119993175AActive Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510157170.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The traditional G-DPRNN and CRN combined models show limited capabilities in time-frequency dynamic modeling, feature space integration and detail capture, and the recursive neural network structure of G-DPRNN results in a lack of direct interaction between channels, limiting the diversity of the network.

Method used

An improved convolutional recurrent network is adopted, including an encoder, decoder, aggregated packet dual-path recurrent network and a convolutional hybrid packet dual-path recurrent neural network. The deep features of speech are extracted through these components, the complex ratio mask of the signal is enhanced, and the convolutional recurrent network is obtained for enhancing the speech signal by training network weights and biases.

Benefits of technology

Improve speech enhancement performance, improve enhanced speech clarity and intelligibility, and maintain the model's lightweight, suitable for lightweight voice enhancement applications of edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993175A_ABST
    Figure CN119993175A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech enhancement, in particular to a lightweight single-channel speech enhancement method based on an improved convolutional recurrent network. The outstanding ability of the improved convolutional recurrent network during feature extraction is fully utilized; an aggregation packet dual-path loop network and a convolutional hybrid packet dual-path loop network are used to improve the depth time-frequency features of multiple channels and fuse the features of the channels, so that the voice information contained in the depth features is richer, and then the depth features are used to train a separation model, so that the voice performance is further enhanced, and the voice processing efficiency is improved. An aggregation packet dual-path loop network and a convolutional hybrid packet dual-path loop network are provided, and the packet dual-path loop network architecture is improved, so that the speech enhancement performance of the convolutional loop network is improved, the lightweight of the model is maintained, the effectiveness of the enhancement model is improved, and the speech enhancement efficiency is improved. And the sharpness and the intelligibility of the voice are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech enhancement, and in particular to a lightweight single-channel speech enhancement method based on an improved convolutional recurrent network. Background Art

[0002] In daily life, speech is a medium for human communication and its importance cannot be ignored. Speech enhancement is a technology that restores clean speech from noisy speech as much as possible to improve speech intelligibility and perceived quality. This technology has been widely used in voice communication, digital hearing aids and other fields, and is used as the front end of many speech systems, such as automatic speech recognition, speech coding, human-computer interaction, etc.

[0003] With the rapid development of AI and deep neural networks, artificial neural networks have become the darling of the current computer technology and communication technology fields with their excellent modeling capabilities, highly abstract prediction capabilities, and excellent relationship mapping capabilities. Single-channel speech enhancement algorithms based on deep learning have been widely used and studied in the field of speech enhancement.

[0004] Compared with traditional speech enhancement algorithms, methods based on deep neural networks (DNNs) have overwhelming performance, but are often accompanied by greater model complexity. DNN-based speech enhancement can be mainly divided into time-frequency and time-domain methods. Time-frequency domain methods aim to extract noise characteristics of acoustic features (e.g., complex spectrum or log power spectrum). Common training targets include ideal ratio mask (IRI) and target magnitude spectrum (TMS). Phase spectrum is also considered to be beneficial to speech quality. Time domain methods directly estimate clean speech waveforms through end-to-end training, avoiding the trouble of estimating phase information in the time-frequency domain. Time domain methods have difficulty modeling extremely long sequences, and traditional recurrent neural networks (RNNs) cannot effectively model such long sequences. Therefore, a dual-path recurrent neural network (DPRNN) is proposed to solve this problem. Among them, long sequence features are divided into smaller blocks and iteratively processed by intra-block and inter-block RNNs, thereby reducing the length of the sequence to be processed by each RNN. The intra-block operation in DPRNN is designed to model the signal features within a frame. This method is also applicable to the frequency domain and has the potential advantage of making full use of the harmonic spectrum structure of speech. In lightweight scenarios, in order to reduce computational overhead, a grouped DPRNN (Grouped Dual-path RNN, G-DPRNN) is proposed, which replaces the RNN in DPRNN with a grouped RNN, and uses a group recurrent layer to effectively reduce the number of model parameters and computational complexity. Recently, a network structure called a convolutional recurrent network (CRN) has been proposed. CRN takes advantage of the advantages of CNN and RNN, and can not only capture the local patterns of the spectrogram, but also model the dependencies between consecutive time frames. Therefore, combining the characteristics of G-DPRNN and CRN in the time-frequency domain, it can achieve performance comparable to or better than the traditional RNN model with a significantly reduced number of parameters and computational cost.

[0005] The traditional G-DPRNN and CRN combined model faces the limitation of lightweight network and shows limited ability in time-frequency dynamic modeling, feature space integration and detail capture. In addition, the recurrent neural network (RNN) structure of G-DPRNN leads to the lack of direct interaction between channels, which limits the diversity of the network. Summary of the invention

[0006] The purpose of the present invention is to provide a lightweight single-channel speech enhancement method based on an improved convolutional recurrent network, so as to solve the problem that the traditional G-DPRNN and CRN combined model faces the limitations of lightweight networks, exhibits limited capabilities in time-frequency dynamic modeling, feature space integration and detail capture, and the recursive neural network (RNN) structure of G-DPRNN leads to a lack of direct interaction between channels, which limits the diversity of the network.

[0007] To achieve the above object, the present invention provides a lightweight single-channel speech enhancement method based on an improved convolutional recurrent network, and the lightweight single-channel speech enhancement method based on the improved convolutional recurrent network comprises the following steps:

[0008] Step 1: preprocess the input single-pass speech signal, extract the real spectrum, imaginary spectrum and amplitude spectrum features of the speech signal respectively, form a three-channel real-valued tensor input, and use an equivalent rectangular bandwidth band processing module to compress the high frequency band;

[0009] Step 2: Expand and reorganize the compressed frequency band features using a spectrum feature enhancement module;

[0010] Step 3: Input the expanded and reorganized frequency band features into the improved convolutional recurrent network, extract the deep features of the speech, and enhance the complex ratio mask of the signal as the output. By training the network weights and biases, a trained convolutional recurrent network for enhancing the speech signal is obtained.

[0011] Step 4: Input the audio features of the noisy speech signal used for testing into the trained convolutional recurrent network for enhancing the speech signal to enhance the noisy speech signal.

[0012] As a further improvement of the present invention, the preprocessing in step 1 includes framing and windowing; the equivalent rectangular bandwidth band processing in step 1 is to merge the high frequency bands and use linear transformation to map the amplitude characteristics of multiple high frequency bands into one feature dimension.

[0013] As a further improvement of the present invention, the spectrum feature enhancement module in step 2 remaps the sub-band information of the frequency dimension to the channel dimension through spectrum expansion, conversion of the spectrum dimension to the channel dimension and convolution operation.

[0014] As a further improvement of the present invention, the improved convolutional recurrent network in step 3 consists of four parts: an encoder, a decoder, an aggregated grouping dual-path recurrent network, and a convolutional hybrid grouping dual-path recurrent neural network;

[0015] Among them, the encoder is used to generate a high-dimensional embedding representation, the decoder is used to gradually restore the original dimension of the feature, the aggregated grouped dual-path recurrent network further enhances the high-dimensional time-frequency feature input, the convolutional mixed grouped dual-path recurrent neural network is used for local feature extraction and feature fusion between channels, and the output of the decoder deconvolution block is a complex ratio mask of the enhanced signal. After multiplying the mask with the input complex spectrum matrix, the enhanced speech signal is reconstructed through an inverse short-time Fourier transform, and parameter optimization is performed through a constrained loss function to train the convolutional recurrent network.

[0016] As a further improvement of the present invention, the encoder in step 3 is composed of two layers of convolution and three layers of grouped convolution. Specifically: the encoder input feature size is 9×T×129, the convolution kernel size of the first two layers is 1×5, the number of input and output channels is 16, the stride in the frequency dimension is 2, and the stride in the time dimension is 1, so the number of time frames T remains unchanged, and the output feature size of the first two layers is 16×T×33; the convolution kernel size of the last three layers of grouped convolution is 3×3, and the expansion coefficients are gradually increased to 1, 2, and 5 respectively for expanding the receptive field. The number of input and output channels is 16, and the output size of the last three layers is 16×T×33. Deep features;

[0017] The decoder is jump-connected with the encoder through a series of deconvolution layers to gradually restore the time-frequency resolution of the features and reconstruct the enhanced speech signal;

[0018] The aggregated grouped dual-path recurrent network combines the grouped dual-path recurrent network with a time-frequency attention mechanism, wherein the time-frequency attention mechanism includes depthwise separable convolution, adaptive mean pooling addition, intra-band convolution, and inter-band convolution to generate attention weights and assign weights to each time-frequency position.

[0019] As a further improvement of the present invention, the adaptive mean pooling is divided into average pooling for the time dimension and average pooling for the frequency dimension, respectively generating two different axis vectors T of the time axis and the frequency axis. P and F P , and then perform broadcast addition processing; the intra-band convolution and inter-band convolution use 1×k and k×1 convolution kernels to perform strip convolution modeling in the intra-band and inter-band directions respectively, generate attention weights in the time-frequency domain, and then use 3×3 deep convolution to further extract local details of the input features, use Hadamard product to assign attention features to each time-frequency position, and finally use a multi-layer perceptron module to perform nonlinear mapping on the time-frequency features.

[0020] As a further improvement of the present invention, the intra-band convolution and inter-band convolution expressions are:

[0021]

[0022] Among them, γ represents the large kernel strip convolution, k represents the kernel size of the strip convolution, Represents the time axis vector T P and the frequency axis vector F P The result after broadcast addition, φ represents batch normalization after ReLU function, and δ represents Sigmoid activation function;

[0023] The expression of attention feature allocation is:

[0024]

[0025] Among them, γ 3*3 represents a 3×3 depth convolution with kernel, y is the attention feature obtained in the previous step, stands for Hadamard product.

[0026] As a further improvement of the present invention, the convolution hybrid grouped dual-path recurrent neural network in step 3 is used to combine an improved depth-separable convolution module with a grouped dual-path recurrent network, wherein the improved depth-separable convolution module comprises depth-wise convolution, point-wise convolution and residual connection, wherein the depth-wise convolution uses a large convolution kernel to extract global information of each channel, and the point-wise convolution adopts an inverted bottleneck design, wherein the hidden dimension between two point-wise convolution layers is set to be four times as wide as the input dimension, so as to fully fuse the global information between the channels, and finally, the GELU activation function and BatchNorm regularization are used after each convolution layer;

[0027] The expression of the improved depth-wise separable convolution module is:

[0028] X′ out =BN(σ l {DepthwiseConv(X in )})+X in

[0029] X″ out =BN(σ l {PointwiseConv(X′ out )})

[0030] X out =BN(σ l {PointwiseConv(X″ out )})

[0031] Among them, X in denotes the output of the grouped dual-path recurrent network as the input of the depthwise separable convolutional module, σ lrepresents the GELU activation function, and BN represents BatchNorm regularization.

[0032] As a further improvement of the present invention, step 4 is specifically: preprocessing the noisy speech signal used for testing, extracting the real spectrum, imaginary spectrum and amplitude spectrum features of the noisy speech signal, inputting them into the improved convolutional recurrent network after optimization training, obtaining a predicted complex ratio mask, multiplying the predicted complex ratio mask with the real spectrum and imaginary spectrum features to obtain a target complex spectrum, reconstructing the speech signal through an inverse short-time Fourier transform, and completing the enhancement of the noisy speech signal.

[0033] As a further improvement of the present invention, the parameter amount of the improved convolutional recurrent network is only 30.2K, and the computational complexity is only 48.2MMACs, which is suitable for lightweight speech enhancement applications in edge devices.

[0034] As a further improvement of the present invention, the lightweight single-channel speech enhancement method based on the improved convolutional recurrent network also includes combining perceptual contrast stretching technology to further optimize the perceptual quality of the enhanced speech, significantly improving speech clarity and intelligibility; the perceptual contrast stretching technology utilizes the sensitivity of the human ear to specific frequency bands to reallocate weights of specific frequency bands, and uses the PCS weight array to multiply the logarithmic amplitude spectrum band by band to amplify the specific frequency band.

[0035] The present invention discloses a lightweight single-channel speech enhancement method based on an improved convolutional recurrent network, which makes full use of the excellent ability of the improved convolutional recurrent network in extracting features, uses an aggregated grouped dual-path recurrent network and a convolutional mixed grouped dual-path recurrent network to enhance the deep time-frequency features of multiple channels and fuse the features between channels, so that the speech information contained in the deep features is richer, and then uses the deep features to train the separation model, so that the performance of the enhanced speech is further improved. In addition, the technical solution aims at the problems that the time-frequency feature extraction ability, feature space integration ability and feature interaction ability between channels of the convolutional recurrent network are limited under the requirements of the lightweight model, and the quality improvement of the enhanced speech is limited. On the basis of the traditional convolutional recurrent network model, an aggregated grouped dual-path recurrent network and a convolutional mixed grouped dual-path recurrent network are proposed, and the grouped dual-path recurrent network architecture is improved, which not only improves the speech enhancement performance of the convolutional recurrent network, but also maintains the lightweight of the model, thereby improving the effectiveness of the enhancement model, and improving the clarity and intelligibility of the enhanced speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0037] Figure 1 This is a structural diagram of a lightweight single-channel speech enhancement model based on an improved convolutional recurrent network provided by the present invention.

[0038] Figure 2 The present invention provides Figure 1 Model structure diagram of the aggregated grouped dual-path recurrent network AG-DPRNN.

[0039] Figure 3 The present invention provides Figure 1 Model structure diagram of the medium convolutional hybrid grouped dual-path recurrent network CG-DPRNN. DETAILED DESCRIPTION

[0040] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0041] See also Figures 1 to 3 The present invention provides a lightweight single-channel speech enhancement method based on an improved convolutional recurrent network, and the lightweight single-channel speech enhancement method based on the improved convolutional recurrent network comprises the following steps:

[0042] Step 1: preprocess the input single-pass speech signal, extract the real spectrum, imaginary spectrum and amplitude spectrum features of the speech signal respectively, form a three-channel real-valued tensor input, and use an equivalent rectangular bandwidth band processing module to compress the high frequency band;

[0043] Step 2: Expand and reorganize the compressed frequency band features using a spectrum feature enhancement module;

[0044] Step 3: Input the expanded and reorganized frequency band features into the improved convolutional recurrent network, extract the deep features of the speech, and enhance the complex ratio mask of the signal as the output. By training the network weights and biases, a trained convolutional recurrent network for enhancing the speech signal is obtained.

[0045] Step 4: Input the audio features of the noisy speech signal used for testing into the trained convolutional recurrent network for enhancing the speech signal to enhance the noisy speech signal.

[0046] Furthermore, the improved convolutional recurrent network in step 3 consists of four parts: an encoder, a decoder, an aggregated grouping dual-path recurrent network, and a convolutional hybrid grouping dual-path recurrent neural network;

[0047] Among them, the encoder is used to generate a high-dimensional embedding representation, the decoder is used to gradually restore the original dimension of the feature, the aggregated grouped dual-path recurrent network further enhances the high-dimensional time-frequency feature input, the convolutional mixed grouped dual-path recurrent neural network is used for local feature extraction and feature fusion between channels, and the output of the decoder deconvolution block is a complex ratio mask of the enhanced signal. After multiplying the mask with the input complex spectrum matrix, the enhanced speech signal is reconstructed through an inverse short-time Fourier transform, and parameter optimization is performed through a constrained loss function to train the convolutional recurrent network.

[0048] In this embodiment, the excellent ability of the improved convolutional recurrent network in extracting features is fully utilized, and the aggregated grouped dual-path recurrent network and the convolutional mixed grouped dual-path recurrent network are used to enhance the deep time-frequency features of multiple channels and fuse the features between channels, so that the speech information contained in the deep features is richer, and the deep features are then used to train the separation model, so that the performance of the enhanced speech is further improved. In addition, this technical solution aims to address the problems of limited time-frequency feature extraction capabilities, feature space integration capabilities, and feature interaction capabilities between channels of the convolutional recurrent network under the requirements of a lightweight model, and limited improvement in the quality of enhanced speech. Based on the traditional convolutional recurrent network model, an aggregated grouped dual-path recurrent network and a convolutional mixed grouped dual-path recurrent network are proposed, and the grouped dual-path recurrent network architecture is improved, which not only improves the speech enhancement performance of the convolutional recurrent network, but also maintains the lightweight of the model, thereby improving the effectiveness of the enhancement model and improving the clarity and intelligibility of the enhanced speech.

[0049] The steps of the specific implementation mode of the present invention are as follows:

[0050] Step 1: Preprocess multiple speech signals to extract the real spectrum, imaginary spectrum and amplitude spectrum features of the noisy speech signals respectively, and use the equivalent rectangular bandwidth (ERB) band processing module to compress the high frequency band.

[0051] Step 1.1: Preprocess multiple speech signals.

[0052] Preprocessing technology is of certain importance in the field of digital speech processing. Since speech signals have the characteristics of short-term stability, speech signals must be preprocessed before Fast Fourier Transform (FFT) is performed on them. Common preprocessing methods include framing, windowing, pre-emphasis and other operations. In speech signal enhancement technology, preprocessing technology can improve the effect of speech enhancement to a certain extent.

[0053] In this embodiment, the noisy speech signal is preprocessed, the sampling rate of each speech signal is 48kHz, and the downsampling rate is 16kHz. Of course, in other embodiments, the number of speech signals can be set to other values, and the sampling rate of each speech signal can be set to other values, as long as the preprocessing of the speech signal can be achieved, there is no limitation here.

[0054] Step 1.2: Extract multi-channel features of the preprocessed speech signal.

[0055] The pre-processed noisy speech signal is subjected to a short-time Fourier transform (STFT), and the real spectrum, imaginary spectrum, and amplitude spectrum features of the noisy speech signal are extracted. For each time frame, the real and imaginary parts of the STFT each contain 257 frequency bands, and the amplitude spectrum also has 257 frequency bands. Therefore, in the time dimension, the spectrum of each frame contains 257 frequency band features of the real part, imaginary part, and amplitude. These features are spliced ​​in the channel dimension to obtain 3 channels.

[0056] Step 1.3: Use the Equivalent Rectangular Bandwidth (ERB) band processing module to compress the high frequency band.

[0057] The spectrum of each channel is divided into multiple ERB subbands, where the high-frequency part is compressed using the ERB fc function to reduce high-frequency details while retaining the low-frequency characteristics of the speech signal. The high-frequency and low-frequency features are merged to return the compressed time-frequency features.

[0058] Step 2: Expand and reorganize the compressed frequency band features using the spectral feature enhancement module (Subband Feature Extraction, SFE).

[0059] Specifically, multiple local sub-band feature blocks are separated in a sliding window manner on the frequency axis, and the features extracted from each window are rearranged according to a new shape, so that the output feature map has a new channel dimension, that is, the frequency dimension of each feature map will become larger and the number of channels will increase. By extracting local frequency band features, the model's local modeling ability for each frequency band is enhanced, and complex time-frequency characteristics can be better modeled.

[0060] Step 3: Input the above features into the improved convolutional recurrent network, which consists of four parts: an encoder, a decoder, an aggregated grouped dual-path recurrent network, and a convolutional mixed grouped dual-path recurrent neural network.

[0061] Step 3.1: The encoder extracts features and the decoder restores features.

[0062] The time-frequency features extracted by SFE are processed by the encoder to obtain deep features, where the encoder input feature size is 9×T×129, the convolution kernel size of the first two layers is 1×5, the number of input and output channels is 16, the stride in the frequency dimension is 2, and the stride in the time dimension is 1, so the number of time frames T remains unchanged, and the output feature size of the first two layers is 16×T×33; the convolution kernel size of the last three layers of grouped convolution is 3×3, and the expansion coefficients are gradually increased to 1, 2, and 5 respectively to expand the receptive field. The number of input and output channels is 16, and the output size of the last three layers is 16×T×33. Deep features.

[0063] See also Figure 1 As shown in FIG. 1 , the decoder is jump-connected with the encoder through a series of deconvolution layers to gradually restore the time-frequency resolution of the features and reconstruct the enhanced speech signal, which will not be described in detail here.

[0064] Step 3.2: Aggregated Grouped Dual-path RNN (AG-DPRNN) is used to extract deep time-frequency features.

[0065] Faced with the limitations of lightweight networks in terms of the number of channels, the time-frequency dynamic modeling capability of the Grouped Dual-path RNN (G-DPRNN) is also limited. To solve this problem, the Grouped Dual-path RNN is combined with the time-frequency attention mechanism to enhance the time-frequency dynamic modeling capability. First, the G-DPRNN divides the input features and hidden states into two disjoint groups, each of which is fed into a recurrent layer with 2 times fewer parameters than the original one. Then a representation rearrangement layer is applied to obtain the grouped output, which is then fed into the Dual-path Recurrent Neural Network (DPRNN). Its intra-frame RNN can model the spectral pattern in a single frame, while the inter-frame RNN models the temporal dependency of specific frequency points. The intra-frame modeling uses a grouped bidirectional GRU, and the inter-frame modeling uses a grouped unidirectional GRU, which ensures the causality of the model. Finally, the output features of size 16×T×33 are processed using the time-frequency attention mechanism.

[0066] See also Figure 2 As shown in the figure, the time-frequency attention mechanism includes deep separable convolution, adaptive mean pooling addition, intra-band convolution and inter-band convolution to generate attention weights and assign weights to each time-frequency position. Among them, deep separable convolution is used to capture the local time-frequency relationship of features; adaptive mean pooling is divided into average pooling of the time dimension and average pooling of the frequency dimension to capture the axial global context information in two directions, and generate two different axis vectors T of the time axis and frequency axis respectively. P and F P , and then perform broadcast addition processing; the intra-band convolution and inter-band convolution use 1×k and k×1 convolution kernels to perform strip convolution modeling in the intra-band and inter-band directions respectively to generate attention weights in the time-frequency domain. The intra-band convolution and inter-band convolution expressions are:

[0067]

[0068] Among them, γ represents the large kernel strip convolution, k represents the kernel size of the strip convolution, Represents the time axis vector T P and the frequency axis vector F P The result after broadcast addition. φ represents batch normalization after ReLU function, and δ represents Sigmoid activation function.

[0069] Next, a 3×3 deep convolution is used to further extract the local details of the input features, and the Hadamard product is used to assign the attention features to each time-frequency position. The expression of attention feature assignment is:

[0070]

[0071] Among them, γ 3*3 represents a 3×3 depth convolution with kernel, y is the attention feature obtained in the previous step, stands for Hadamard product.

[0072] Finally, a multi-layer perceptron (MLP) module is used to perform nonlinear mapping on the time-frequency features. It contains two layers of 1×1 convolution instead of the traditional fully connected layer, which performs nonlinear mapping and enhances features with lower computational overhead, and the output feature dimension is still 16×T×33.

[0073] Step 3.3: Convolutional mixed grouped dual-path recurrent neural network (CG-DPRNN) is used to further extract deep time-frequency features.

[0074] Under the limitation of lightweight network, the Grouped Dual-path RNN (G-DPRNN) is insufficient in terms of feature space integration and detail capture. In addition, the recurrent neural network (RNN) structure of G-DPRNN leads to indirect interactions between channels, which in turn limits the diversity of features. Therefore, CGDPRNN combines an improved depthwise separable convolution module with G-DPRNN, providing G-DPRNN with an additional feature enhancement and fusion module that can effectively fuse and associate features between channels at the same time.

[0075] See also Figure 2 As shown in Figure 1, specifically, the output features of AG-DPRNN are first processed by an improved depthwise separable convolution module. This module consists of depthwise convolution (i.e., the number of groups is equal to the number of channels) and pointwise convolution (i.e., the convolution kernel size is 1×1), where the depthwise convolution is used to extract fine-grained features within the local frequency domain within each time frame, and then residually connected with the time-frequency features of the previous step. The connected features use pointwise convolution operations to mix the time-frequency information between channels.

[0076] In order to fully mix the time-frequency dimension and the information between channels, we applied two point-by-point convolutions after the depthwise convolution and performed an inverted bottleneck design on them. Specifically, this design involves setting the hidden dimension between the two point-by-point convolution layers to four times the input dimension. The enlarged hidden dimension can fully and completely mix the fine-grained features extracted by the depthwise convolution. In addition, we used GELU activation and post-activation BatchNorm layers after each convolution. At the same time, from the perspective of model lightweighting, depthwise separable convolution can effectively reduce network parameters and computational costs compared to ordinary convolution. The expression of the improved depthwise separable convolution module is:

[0077] X′ out =BN(σ l {DepthwiseConv(X in )})+X in

[0078] X″ out =BN(σ l {PointwiseConv(X′ out )})

[0079] X out =BN(σ l {PointwiseConv(X″ out )})

[0080] Among them, X in denotes the output of the grouped dual-path recurrent network as the input of the depthwise separable convolutional module, σ l Denotes GELU activation function, and BN denotes BatchNorm regularization. The output of the improved depthwise separable convolution module is used as the input of G-DPRNN to further extract deep time-frequency features.

[0081] Step 3.4: Train and improve the weights and biases of the convolutional recurrent network.

[0082] The parameter optimization process of the improved convolutional recurrent network is divided into the forward propagation stage and the back propagation stage. The forward propagation (FP) stage is used to randomly initialize the weights and biases of each layer of neurons; the back propagation (BP) stage is used to minimize the joint constraint loss function through integrated optimization, and iteratively update and adjust the weights and biases of each layer of neurons. The main body of the BP and FP stages is the loss function, which is optimized by the gradient descent algorithm to make it as close to the minimum value as possible. The expression of the loss function is:

[0083]

[0084] in, and s represent the enhanced speech and pure speech respectively. and S represent the enhanced speech spectrogram and the pure speech spectrogram respectively. The values ​​of parameters γ and δ are set to 0.01 and 0.3 respectively. The rest of the above formula is expressed as:

[0085]

[0086]

[0087] The amplitude loss is calculated with a scale of 0.3 to reduce the sensitivity to the amplitude difference, while the real part loss and imaginary part loss are calculated with a scale of 0.7 to balance the errors under different amplitudes.

[0088] Finally, a trained improved convolutional recurrent network for enhancing speech signals is obtained.

[0089] Step 4: Input the time-frequency features of the noisy speech signal used for testing into the improved convolutional recurrent network after optimization training to enhance the noisy speech signal.

[0090] 4.1. Testing phase.

[0091] First, the noisy speech signal is preprocessed, and then STFT is performed to extract the real spectrum, imaginary spectrum and amplitude spectrum features of the noisy speech signal.

[0092] 4.2. Reconstruct speech time domain signal.

[0093] In the stage of reconstructing the speech time domain signal, it is input into the improved convolutional recurrent network after optimization training to obtain the predicted complex ratio mask, and the predicted complex ratio mask is multiplied with the real spectrum and imaginary spectrum features to obtain the target complex spectrum. The speech signal is reconstructed through inverse short-time Fourier transform to complete the enhancement of the noisy speech signal.

[0094] 4.3 Performance evaluation

[0095] To evaluate the performance of our proposed lightweight model, we chose to use the VoiceBank+DEMAND dataset, a recognized benchmark in the field of speech enhancement. This dataset combines the VoiceBank corpus and the noisy samples of the DEMAND dataset, the former consisting of clean speech recordings and the latter creating a set of realistic noisy speech samples. The VoiceBank+DEMAND dataset contains a total of 11,572 pairs of noisy and clean speech samples for training and 824 pairs of noisy and clean speech samples for testing. To ensure the uniformity of all data, each audio sample is resampled to a consistent 16kHz frequency. The experimental results are the average of the 824 speech results.

[0096] In the process of speech preprocessing, the calculation of short-time Fourier transform (STFT) uses a Hanning window with square root weighting, the length is set to 32 milliseconds, the jump length is 16 milliseconds, and the length of Fourier transform is 512. The input features are the real spectrum, imaginary spectrum and amplitude spectrum features of the speech signal, and are combined by channel.

[0097] The present invention adopts multiple speech indicators to measure the accuracy and effectiveness of the proposed algorithm, including Perceptual Evaluation of Speech Quality (PESQ), Scale-Invariant Signal-to-Noise Ratio (SISNR) and Short-Time Objective Intelligibility (STOI). The values ​​of these indicators are positively correlated with the speech enhancement performance.

[0098] Additionally, during the evaluation phase, we used Perceptual Contrast Stretching (PCS), a spectral enhancement technique that uses the human ear’s sensitivity to specific frequency bands to improve the auditory quality of speech. This method adjusts the amplitude spectrum of speech based on the perceptual significance of each frequency band, thereby improving clarity in the most important areas. In our study, PCS was used as an auxiliary step in the evaluation phase after the initial enhancement phase to further improve speech quality, with a focus on optimizing perceptual auditory properties.

[0099] In order to verify that the two improvement strategies for G-DPRNN are effective, we conducted ablation experiments on the improved models. The experimental results are shown in Table 1. Both AG-DPRNN, which uses the time-frequency attention mechanism to further enhance the high-dimensional time-frequency feature input, and CG-DPRNN, which uses the improved deep separable convolution module for local feature extraction and feature fusion between channels, outperform G-DPRNN under very limited increments of computing resources. The best performance indicators can be achieved by integrating AG-DPRNN and CG-DPRNN.

[0100] In order to verify the lightweight degree and performance advantages of the improved convolutional recurrent network, the method of the present invention is compared with other algorithms. The experimental results are shown in Table 2. The improved overall model ACCRN only uses 30k model parameters and 48M MACs of calculation to achieve PESQ 2.95.

[0101] Although the performance of DeepFilterNet2, CCFNet+, and Dense-TSNet models are slightly better than ACCRN, the model size of DeepFilterNet2 is 2.3M, ACCRN is only 3.4% of its size, and the computational complexity of CCFNet+ is 1.47G, ACCRN is only 3% of its size.

[0102] At the same time, compared with the recently proposed FSPEN (2024) model, ACCRN significantly reduces parameters and computation while maintaining comparable performance. As for Dense-TSNet (2024), although its model is more lightweight than our model, its huge amount of computation will inevitably limit its real-time performance, making it less versatile than our model in the deployment of edge devices. The results show that ACCRN achieves highly competitive performance with minimal resource usage, making it more suitable for wearable devices and IoT devices.

[0103] Table 1 Ablation experiment of improved convolutional recurrent network ACCRN

[0104]

[0105] Table 2 Comparison of improved convolutional recurrent network ACCRN and other classic lightweight models

[0106]

[0107] In summary, the present invention provides a lightweight single-channel speech enhancement method based on an improved convolutional recurrent network. The backbone network is a convolutional recurrent network based on an encoder and a decoder, and two improvement strategies are used for the grouped dual-path recurrent network. One is to combine the time-frequency attention mechanism with the grouped dual-path recurrent network, and the other is to use an improved depth-separable convolution for local feature extraction and feature fusion between channels, thereby improving the speech enhancement performance and improving the clarity and intelligibility of the enhanced speech. These two strategies not only improve the performance but also maintain the lightweight of the model, and have a wider range of real-world application scenarios. In the performance evaluation stage, the perceptual contrast stretching technology is combined to further optimize the perceptual quality of the enhanced speech, significantly improving the speech clarity and intelligibility.

[0108] What is disclosed above is only a preferred embodiment of the present invention, and it certainly cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made according to the claims of the present invention still fall within the scope of the invention.

Claims

1. A lightweight single-channel speech enhancement method based on an improved convolutional recurrent network, characterized in that: The steps include: Step 1: preprocess the input single-pass speech signal, extract the real spectrum, imaginary spectrum and amplitude spectrum features of the speech signal respectively, form a three-channel real-valued tensor input, and use an equivalent rectangular bandwidth band processing module to compress the high frequency band; Step 2: Expand and reorganize the compressed frequency band features using a spectrum feature enhancement module; Step 3: Input the expanded and reorganized frequency band features into the improved convolutional recurrent network, extract the deep features of the speech, and enhance the complex ratio mask of the signal as the output. By training the network weights and biases, a trained convolutional recurrent network for enhancing the speech signal is obtained. Step 4: Input the audio features of the noisy speech signal used for testing into the trained convolutional recurrent network for enhancing the speech signal to enhance the noisy speech signal.

2. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 1, characterized in that: The preprocessing in step 1 includes framing and windowing; the equivalent rectangular bandwidth band processing in step 1 is to merge the high frequency bands and use linear transformation to map the amplitude features of multiple high frequency bands into one feature dimension.

3. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 1, characterized in that: The spectrum feature enhancement module in step 2 remaps the sub-band information of the frequency dimension to the channel dimension through spectrum expansion, conversion of the spectrum dimension to the channel dimension and convolution operation.

4. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 1, characterized in that: The improved convolutional recurrent network in step 3 consists of four parts: encoder, decoder, aggregation grouping dual-path recurrent network and convolutional hybrid grouping dual-path recurrent neural network; Among them, the encoder is used to generate a high-dimensional embedding representation, the decoder is used to gradually restore the original dimension of the feature, the aggregated grouped dual-path recurrent network further enhances the high-dimensional time-frequency feature input, the convolutional mixed grouped dual-path recurrent neural network is used for local feature extraction and feature fusion between channels, and the output of the decoder deconvolution block is a complex ratio mask of the enhanced signal. After multiplying the mask with the input complex spectrum matrix, the enhanced speech signal is reconstructed through an inverse short-time Fourier transform, and parameter optimization is performed through a constrained loss function to train the convolutional recurrent network.

5. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 4, characterized in that: The encoder in step 3 is composed of two layers of convolution and three layers of grouped convolution. Specifically: the encoder input feature size is 9×T×129, the convolution kernel size of the first two layers is 1×5, the number of input and output channels is 16, the stride in the frequency dimension is 2, and the stride in the time dimension is 1, so the number of time frames T remains unchanged, and the output feature size of the first two layers is 16×T×33; the convolution kernel size of the last three layers of grouped convolution is 3×3, and the expansion coefficients are gradually increased to 1, 2, and 5 respectively for expanding the receptive field. The number of input and output channels is 16, and the output size of the last three layers is 16×T×33. Deep features; The decoder is jump-connected with the encoder through a series of deconvolution layers to gradually restore the time-frequency resolution of the features and reconstruct the enhanced speech signal; The aggregated grouped dual-path recurrent network combines the grouped dual-path recurrent network with a time-frequency attention mechanism, wherein the time-frequency attention mechanism includes depthwise separable convolution, adaptive mean pooling addition, intra-band convolution, and inter-band convolution to generate attention weights and assign weights to each time-frequency position.

6. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 5, characterized in that: The adaptive mean pooling is divided into average pooling for the time dimension and average pooling for the frequency dimension, generating two different axis vectors T for the time axis and the frequency axis respectively. P and F P , and then perform broadcast addition processing; the intra-band convolution and inter-band convolution use 1×k and k×1 convolution kernels to perform strip convolution modeling in the intra-band and inter-band directions respectively, generate attention weights in the time-frequency domain, and then use 3×3 deep convolution to further extract local details of the input features, use Hadamard product to assign attention features to each time-frequency position, and finally use a multi-layer perceptron module to perform nonlinear mapping on the time-frequency features.

7. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 6, characterized in that: The intra-band convolution and inter-band convolution expressions are: Among them, γ represents the large kernel strip convolution, k represents the kernel size of the strip convolution, Represents the time axis vector T P and the frequency axis vector F P The result after broadcast addition, φ represents batch normalization after ReLU function, and δ represents Sigmoid activation function; The expression of attention feature allocation is: Among them, γ 3*3 represents a 3×3 depth convolution with kernel, y is the attention feature obtained in the previous step, stands for Hadamard product.

8. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 4, characterized in that: The convolution hybrid grouped dual-path recurrent neural network in step 3 is used to combine an improved depth-wise separable convolution module with a grouped dual-path recurrent network, wherein the improved depth-wise separable convolution module includes depth-wise convolution, point-wise convolution, and residual connection, wherein the depth-wise convolution uses a large convolution kernel to extract the global information of each channel, and the point-wise convolution adopts an inverted bottleneck design, wherein the hidden dimension between two point-wise convolution layers is set to be four times as wide as the input dimension, so as to fully fuse the global information between the channels, and finally, the GELU activation function and BatchNorm regularization are used after each convolution layer; The expression of the improved depth-wise separable convolution module is: X' out =BN(σ l {DepthwiseConv(X in )})+X in X” out =BN(σ l {PointwiseConv(X’ out )}) X out =BN(σ l {PointwiseConv(X” out )}) Among them, X in denotes the output of the grouped dual-path recurrent network as the input of the depthwise separable convolutional module, σ l represents the GELU activation function, and BN represents BatchNorm regularization.

9. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 1, characterized in that: Step 4 is specifically as follows: preprocessing the noisy speech signal used for testing, extracting the real spectrum, imaginary spectrum and amplitude spectrum features of the noisy speech signal, inputting them into the improved convolutional recurrent network after optimization training, obtaining a predicted complex ratio mask, multiplying the predicted complex ratio mask with the real spectrum and imaginary spectrum features to obtain a target complex spectrum, reconstructing the speech signal through inverse short-time Fourier transform, and completing the enhancement of the noisy speech signal.

10. The lightweight single-channel speech enhancement method based on an improved convolutional recurrent network as claimed in claim 1, characterized in that: The lightweight single-channel speech enhancement method based on the improved convolutional recurrent network also includes combining perceptual contrast stretching technology to further optimize the perceptual quality of the enhanced speech, significantly improving speech clarity and intelligibility; the perceptual contrast stretching technology uses the sensitivity of the human ear to specific frequency bands to reallocate weights for specific frequency bands, and uses the PCS weight array to multiply the logarithmic amplitude spectrum band by band to amplify the specific frequency band.

Citation Information

Patent Citations

  • Speech enhancement model calculation amount compression method based on recurrent neural network

    CN115273874A

  • Convolutional recurrent neural network and speech enhancement method and device

    CN115273883A

  • Single-channel end-to-end speech extraction method based on DPRNN-Ext

    CN116189704A

  • Single-channel speech enhancement method based on waveform spectrum fusion network

    CN116682444A

  • Speech enhancement method and device based on attention-enhanced double-path convolutional recurrent network

    CN118887967A