A method and device for speech denoising under low signal-to-noise ratio

By introducing time-frequency converters and improving dense blocks into the TFDense-Net model, and using multi-spectrum discriminators for training, the problem of poor voice denoising effect in low signal-to-noise ratio environments is solved, and more efficient voice denoising and audio reconstruction effects are achieved.

CN119229889BActive Publication Date: 2025-05-13BEIJING LANGUAGE AND CULTURE UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411778837.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-13
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

The prior art has poor voice denoising effect in low signal-to-noise ratio environments, high computing resource overhead, and difficult to effectively remove background noise, resulting in unclear voice.

Method used

The TFDense-Net model is adopted, combined with the U-net network structure and the Transformer model structure, a time-frequency converter module is introduced and dense blocks are improved. The model is conducted adversarial iterative training through a multi-spectral discriminator to improve speech denoising performance.

Benefits of technology

Effectively enhance audio noise reduction performance in complex noise environments, improve speech clarity, reduce feature loss caused by noise interference, improve audio reconstruction quality, and maintain speech naturalness and clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229889B_ABST
    Figure CN119229889B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for speech denoising under low signal-to-noise ratio, and relates to the technical field of speech denoising. The method comprises: recording audio through a microphone to obtain pure speech data; preprocessing the pure speech data to obtain training speech data; constructing a TFDense-Net speech denoising model to be trained according to the U-net network structure and the Transformer model structure; based on a multi-spectral discriminator, according to the training speech data, using the Adam optimizer to perform adversarial iterative training on the TFDense-Net speech denoising model to be trained to obtain the TFDense-Net speech denoising model; in a low signal-to-noise ratio environment, the speech data to be denoised is collected through a microphone; the speech data to be denoised is input into the TFDense-Net speech denoising model to obtain denoised speech data. The present invention is an efficient and clear speech denoising method under low signal-to-noise ratio that combines an improved dense block and a video transformer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech denoising, and in particular to a speech denoising method and device for low signal-to-noise ratio. Background Art

[0002] Speech denoising technology is usually divided into time domain methods and time-frequency domain methods. The time domain method is simple to implement, can retain complete audio information, and can perform clear and definite audio processing. Due to the high sampling rate and inconspicuous audio features of audio sequences, speech denoising models do not show significant noise reduction performance under large parameter amounts. Based on signal theory, time-frequency domain methods are widely used in audio signal processing tasks, and short-time Fourier transform (STFT) is usually used to convert one-dimensional audio into a two-dimensional spectrum. By performing Fourier transform through a sliding window, STFT reduces the speech time series and extracts neighboring features of the audio sequence, enhancing the model's ability to recognize audio signals. The time-frequency domain method is compact and shows strong robustness in complex environments and signal distortion scenarios. Due to its rich audio characteristics, it has been widely used in speech denoising research.

[0003] In existing speech denoising methods, the dual-branch full-band-sub-band fusion network performs downsampling before using the self-attention mechanism to participate in the calculation, which does not take advantage of the U-shaped convolutional network in feature fusion and greatly increases the computing resource overhead. The dual-path self-attention mechanism needs to calculate the attention of the full band and sub-band, which poses challenges when extracting features from the overall spectrum.

[0004] In the prior art, there is a lack of an efficient and clear speech denoising method under low signal-to-noise ratio that combines improved dense blocks and video transformers. Summary of the invention

[0005] In order to solve the technical problems of large computing resource overhead and unclear denoised speech at low signal-to-noise ratio in the prior art, an embodiment of the present invention provides a method and device for speech denoising at low signal-to-noise ratio. The technical solution is as follows:

[0006] On the one hand, a method for speech denoising under low signal-to-noise ratio is provided, the method being implemented by a speech denoising device, the method comprising:

[0007] Recording audio through a microphone to obtain clean voice data; preprocessing the clean voice data to obtain training voice data;

[0008] Construct the TFDense-Net speech denoising model to be trained based on the U-net network structure and the Transformer model structure;

[0009] Based on the multi-spectral discriminator, according to the training speech data, using the Adam optimizer to perform adversarial iterative training on the TFDense-Net speech denoising model to be trained, to obtain the TFDense-Net speech denoising model;

[0010] In a low signal-to-noise ratio environment, speech data to be denoised is collected by a microphone; the speech data to be denoised is input into the TFDense-Net speech denoising model to obtain denoised speech data.

[0011] On the other hand, a speech denoising device for low signal-to-noise ratio is provided, which is applied to a speech denoising method for low signal-to-noise ratio, and the device comprises:

[0012] A training voice acquisition module is used to record audio through a microphone to obtain clean voice data; pre-process the clean voice data to obtain training voice data;

[0013] The model building module is used to build the TFDense-Net speech denoising model to be trained based on the U-net network structure and the Transformer model structure;

[0014] A model training module is used to perform adversarial iterative training on the TFDense-Net speech denoising model to be trained based on the multi-spectral discriminator and the training speech data using an Adam optimizer to obtain the TFDense-Net speech denoising model;

[0015] The speech denoising module is used to collect the speech data to be denoised through a microphone in a low signal-to-noise ratio environment; the speech data to be denoised is input into the TFDense-Net speech denoising model to obtain denoised speech data.

[0016] On the other hand, a speech denoising device is provided, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned speech denoising methods for low signal-to-noise ratio is implemented.

[0017] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned speech denoising methods under low signal-to-noise ratio.

[0018] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0019] The present invention proposes a speech denoising method for low signal-to-noise ratio. By introducing a time-frequency converter module between the encoder and decoder of TFDense-Net, the time domain and frequency domain features in the audio signal can be captured simultaneously. In a complex noise environment, the audio denoising performance can be effectively enhanced and the speech clarity can be improved. This module can also reduce the feature loss caused by noise interference, improve the quality of audio reconstruction, effectively remove background noise while maintaining the naturalness and clarity of speech, and improve the audio denoising effect.

[0020] For the optimization of the dense block module, a combination of deep convolution and point convolution is used to improve the dense block, so that it not only retains the advantages of dense connections in the dense network structure, but also expands the receptive field through dilated convolution to better capture multi-scale information in the time and frequency domain. This module is used in both the encoder and decoder of TFDense-Net, and can effectively retain and transmit detailed information of speech. The extraction and fusion capabilities of audio signal features are optimized, so that the network can better maintain the key information in the speech, improve the noise reduction effect and reduce the loss of details.

[0021] The proposed TFDense-Net adopts a U-shaped network convolutional architecture, combined with a time-frequency converter and an improved dense block, for feature extraction and reconstruction of audio signals. The encoder gradually compresses the time-frequency features by downsampling, the bottleneck layer uses a self-attention mechanism for global feature fusion, and the decoder restores the signal by upsampling. This architecture ensures the model's ability to capture time-frequency features while maintaining a low computational complexity, and is suitable for audio denoising tasks. While ensuring the denoising effect, the computational complexity and parameter amount of the model are greatly reduced, and the training efficiency and inference speed of the model are improved. The present invention is an efficient and clear speech denoising method under low signal-to-noise ratio that combines an improved dense block and a video converter. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 is a flow chart of a method for speech denoising under low signal-to-noise ratio provided by an embodiment of the present invention;

[0024] Figure 2 It is a block diagram of a speech denoising device for low signal-to-noise ratio provided by an embodiment of the present invention;

[0025] Figure 3It is a structural schematic diagram of a speech denoising device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0027] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0028] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.

[0029] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0030] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0031] The embodiment of the present invention provides a method for speech denoising under low signal-to-noise ratio, which can be implemented by a speech denoising device, which can be a terminal or a server. Figure 1 The flowchart of the speech denoising method for low signal-to-noise ratio is shown in the figure. The processing flow of the method may include the following steps:

[0032] S1. Record audio through a microphone to obtain pure voice data; pre-process the pure voice data to obtain training voice data.

[0033] The clean voice data is preprocessed to obtain training voice data, including:

[0034] Unifying the audio length of the pure voice data to obtain first processed voice data;

[0035] Based on a sampling rate of 16 kHz, resampling the first processed voice data to obtain second processed voice data;

[0036] Performing data filtering on the second processed voice data to obtain third processed voice data;

[0037] According to the third processed speech data, noise is added by adopting a distortion modeling method to obtain training speech data.

[0038] In a feasible implementation manner, the present invention uses the voice data in WAV format recorded by a microphone, and a total of 20k single-channel pure voice data is used.

[0039] The original data is preprocessed, and the voice data is uniformly adjusted to a length of 2 seconds. If the length of the original audio signal is less than 2 seconds, zeros are appended to the end of the data to extend it. If the signal length exceeds 2 seconds, 2 seconds of the signal are randomly cropped as the training sample of the current period to ensure the consistency of the signal length within the same batch.

[0040] The speech data is resampled at a sampling rate of 16kHz to ensure that each speech data contains 32k sampling points. The audio data with low information entropy is filtered out from the second processed speech data to improve the purity of the speech data for the training process. The third processed speech data is preprocessed into speech with complex noise using distortion modeling as training data. Among them, 19k data is taken as the training set and 1k data is taken as the verification set.

[0041] S2. Construct the TFDense-Net speech denoising model to be trained based on the U-net network structure and the Transformer model structure.

[0042] Among them, the TFDense-Net speech denoising model includes an encoder, a decoder, and a bottleneck layer;

[0043] The encoder is composed of multiple identical modules stacked together; each module is composed of an improved dense block, a 3×3 convolutional layer, and a parameterized rectified linear unit activation function;

[0044] The bottleneck layer includes a transformer module and a global self-attention module; the transformer module includes a time-frequency transformer, a frequency transformer, and a global transformer;

[0045] The decoder includes a number of identical modules; the structure of the decoder is symmetrical to that of the encoder.

[0046] In a feasible implementation manner, the structure of the Time-Frequency Dense Network (TFDense-Net) proposed in the present invention, which is called Time-Frequency Dense Network in Chinese, mainly consists of three modules: an encoder, a decoder and a bottleneck layer.

[0047] The encoder takes the original Mel spectrum as input and is composed of multiple identical layers stacked together, each of which includes a multi-head self-attention mechanism and a feedforward neural network. The bottleneck layer takes the output of the encoder as input, uses a time-frequency transformer to extract features, and provides the denoised features as output to the decoder for spectrum reconstruction. The decoder then restores the denoised features to the denoised Mel spectrum, and finally restores the denoised Mel spectrum to the denoised speech data through an inverse short-time Fourier transform.

[0048] In the encoder part based on the self-attention mechanism, in order to be more suitable for the dual-path network, the classic linear layer is replaced by a gated recurrent unit based on a position-aware feedforward network, so as to better learn the position information and perform feature fusion.

[0049] In the encoder structure, the input value vector, key vector, and query vector are passed through linear layers to complete their respective linear transformations. These linearly transformed vectors enter the scaled dot product attention module, where the attention weights are calculated through dot products and the attention output is generated. In the multi-head attention mechanism, the outputs of these different heads are concatenated and passed through a linear layer to further process the merged features. The output of the attention module is then added to the input through a residual connection and layer normalized to enhance the training stability of the model.

[0050] The output after layer normalization enters the gated recurrent unit layer for further extraction of sequence information. The output is processed by the rectified linear unit activation function and then passes through a linear layer to generate the output of this layer. This output is added to the output of the previous layer through a residual connection and layer normalization is performed again. This complete process extracts the contextual information of the input through the multi-head attention module, and further learns the sequence features through the gated recurrent unit layer, so that the model can effectively capture complex sequence dependencies, thereby enhancing the understanding and expression of input data.

[0051] A time-frequency transformer is introduced in the bottleneck layer of TFDense-Net, and a global self-attention mechanism is added to further integrate features. A 1×1 convolution layer is used between the self-attention mechanisms to learn more two-dimensional features.

[0052] In the bottleneck layer structure of TFDense-Net, the input passes through a 1×1 convolutional layer for preliminary feature extraction and compression. The features enter a module consisting of a time-frequency transformer, a frequency transformer, and a global transformer mechanism, and the overall structure is repeated B times.

[0053] In each cycle, the time dimension information is processed by the time-frequency transformer module to capture the temporal dependencies. The features are group normalized to normalize the data distribution, and the original input and the output after the time-frequency transformer are added through the residual connection to retain the original information and avoid gradient disappearance.

[0054] The audio features are transformed into dimensions through permutation operations to adapt them to the input format of the frequency transformer module. The bottleneck layer structure focuses on feature extraction in the frequency dimension and further explores the correlation in frequency. After another group normalization and residual connection, it ensures that the features in the frequency dimension are effectively captured while maintaining a stable gradient flow.

[0055] After processing the frequency information, the features are arranged again and enter the global self-attention mechanism module. This module focuses on the integration of global features to generate more representative global information. After group normalization, residual connection is performed with the input to fuse the global features with the original information, further enhancing the expression ability.

[0056] After a layer of 1×1 convolution and a 1×1 group convolution, the features are compressed and the channels are combined to generate the final output. This series of processes captures features in time, frequency, and global dimensions through different self-attention mechanism modules, combined with convolution and normalization operations, to ensure that the model fully extracts and expresses the multi-dimensional information of the input.

[0057] The features processed by the encoder are passed to the time-frequency transformer module, which further optimizes the feature representation by capturing the dependencies in time and frequency to make it contain more time-frequency information. The features processed by the self-attention mechanism module are passed to the decoder for restoration.

[0058] The decoder structure is symmetrical to the encoder and is also composed of multiple modules, each of which contains a parameterized rectified linear unit activation function, a 3×3 convolutional layer, and an improved dense block. These modules gradually restore the features processed by the self-attention mechanism to a form similar to the input signal. The output of the decoder is transformed through the inverse short-time Fourier transform to restore the frequency domain features back to the time domain signal, and the reconstructed time domain signal is output. This whole process extracts features through the encoder, enhances features through the self-attention mechanism module, and then restores the signal through the decoder, thereby achieving effective processing and reconstruction of time-frequency information. Speech enhancement is achieved.

[0059] Among them, the improved dense block refers to the dense block optimized by using the dilated convolution layer and the point-by-point convolution layer; the improved dense block is used for the feature integration of the input TFDense-Net speech denoising model data.

[0060] In a feasible implementation, in the structure of TFDense-Net, the input signal is converted from the time domain signal to the frequency domain through short-time Fourier transform, which facilitates the subsequent feature extraction. The converted frequency domain features enter multiple modules in the encoder, each of which consists of an improved dense block, a 3×3 convolutional layer, and a parameterized rectified linear unit activation function. The improved dense block is used to extract and enrich the frequency domain feature information, the convolutional layer further processes the features, and the parameterized rectified linear unit activation function introduces nonlinearity to enhance the expressiveness of the model.

[0061] Based on the dense block, the present invention uses the dilated convolution layer and the point-by-point convolution layer to optimize the dense block. This optimized module is implemented in the encoder and decoder of the U-shaped convolutional network for feature integration.

[0062] In the improved dense block structure, the processed input is passed through a 2×3 convolution layer to extract features, which are then concatenated and processed through a convolutional network. The features are sequentially convolved with multiple dilation rates, including convolutional layers with dilation rates of 1, 3, and 5. After each convolution, the features are concatenated and passed through the convolutional network again, gradually enriching the multi-scale information of the features.

[0063] The features are then further processed in multi-scale convolutional network modules, which correspond to different channel dimensions (2B, 3B, 4B) but are similar in structure. Each module first performs layer normalization on the input and performs point-by-point feature extraction through point convolution. The features are combined with gated linear units through depthwise convolution to enhance feature selection between channels. After another pointwise convolution, the output is added to the previous layer through residual connections to form the first output layer. This process is repeated multiple times, each time including depthwise convolution, pointwise convolution, and residual connections, so that the features are enriched and enhanced layer by layer.

[0064] Through such a stacked structure, the model can effectively capture multi-scale features, and optimize information flow and gradient propagation through gating mechanisms and residual connections, thereby achieving deep processing and expression of input features.

[0065] S3. Based on the multi-spectral discriminator and the training speech data, the Adam optimizer is used to perform adversarial iterative training on the TFDense-Net speech denoising model to be trained to obtain the TFDense-Net speech denoising model.

[0066] Optionally, based on the multi-spectral discriminator and the training speech data, an Adam optimizer is used to perform adversarial iterative training on the TFDense-Net speech denoising model to be trained, so as to obtain the TFDense-Net speech denoising model, including:

[0067] Perform short-time Fourier transform on the training speech data to obtain the training speech Mel spectrum;

[0068] Construct a generative adversarial network based on the multi-spectral discriminator and the TFDense-Net speech denoising model to be trained;

[0069] According to the training speech data and the training speech Mel spectrum, the generative adversarial network is trained adversarially to obtain an adversarial loss function; the adversarial loss function includes a time domain loss function and a time-frequency domain loss function;

[0070] According to the adversarial loss function, the Adam optimizer is used to adjust the parameters of the TFDense-Net speech denoising model to be trained to obtain the TFDense-Net speech denoising model.

[0071] In a feasible implementation, after the data is normalized, the data is converted into a three-dimensional tensor representing a Mel spectrum using a short-time Fourier transform, and is input into a speech denoising model TFDense-Net to obtain a denoised Mel spectrum.

[0072] The training speech data and its corresponding spectral features and the denoised speech signal and its spectral features are fed into the discriminator network as input. The network aims to learn to distinguish between the original and denoised audio and use the adversarial loss function to evaluate the effect of the denoising algorithm.

[0073] The present invention uses the Adam optimizer, the initialization learning rate is 0.001, and 300 batches are trained. The window length of the short-time Fourier transform and the inverse short-time Fourier transform and the fast Fourier transform (Fast Fourier Transform, FFT) point size are set to 512, and the frame shift is 256. The number of feature maps of the speech time-frequency map in the time-frequency domain is set to 64.

[0074] As the training progresses, the model parameters gradually stabilize and eventually reach an optimized state. In this state, the model can accurately identify the noise component in the speech signal and effectively separate it from the signal. After the training is completed, the model parameters are fixed and no further adjustments are made.

[0075] Optionally, based on the multi-spectral discriminator, adversarial training is performed on the generative adversarial network according to the training speech data and the training speech Mel spectrum to obtain an adversarial loss function, including:

[0076] Use the training speech Mel spectrum to perform frequency domain speech denoising on the training TFDense-Net speech denoising model to obtain the denoised training speech Mel spectrum;

[0077] Perform inverse short-time Fourier transform on the mel spectrum of the denoised training speech to obtain denoised training speech data;

[0078] Copying the denoised training speech data and the training speech data to obtain a plurality of training speech branch data and a plurality of denoised training speech branch data;

[0079] Performing multi-frequency domain short-time Fourier transform on multiple training speech branch data and multiple denoising training speech branch data to obtain multi-frequency domain training speech frequency domain features and multi-frequency domain denoising training speech frequency domain features;

[0080] According to the frequency domain characteristics of the training speech and the frequency domain characteristics of the denoised training speech, the loss function is summarized and calculated through a multi-spectral discriminator to obtain an adversarial loss function.

[0081] In a feasible implementation, the discriminator uses different window sizes and frame shifts to perform short-time Fourier transform on the speech data. For each processed result, a convolutional neural network is used to further fuse and extract the features of the speech signal, and gradually downsample it to map the high-dimensional features to a low-dimensional space, thereby providing a compact and informative representation for quality assessment and ultimately evaluating the speech quality.

[0082] During the adversarial training process, the goal of the discriminator network is to minimize the difference in quality assessment between the real speech and the denoised speech, that is, to minimize the adversarial loss function. This process involves adversarial training in a generative adversarial network, where the discriminator and the denoising model compete with each other to improve the performance of the denoising model. In this way, a denoising model is trained that can effectively remove noise and maintain speech quality.

[0083] In the multi-spectral discriminator structure, the input audio signal is copied into multiple branches and short-time Fourier transform is performed on each branch to convert the time domain audio signal into a frequency domain representation so that the discriminator can analyze the frequency information.

[0084] The multi-spectral discriminator receives audio input, evaluates the quality of the generated audio, outputs adversarial loss, and strengthens the model's discriminative ability. The whole process combines the reconstruction loss of TFDense-Net and the adversarial loss of the multi-spectral discriminator, making the generated audio both high-fidelity and closer to the characteristics of real audio.

[0085] Among them, the multi-spectral discriminator includes multiple processing branches and a summarization layer;

[0086] The processing branch includes a low-resolution processing layer, a feature normalization layer, and a 3×3 convolutional layer;

[0087] The aggregation layer is used to aggregate and calculate the frequency domain features of the outputs of multiple processing branches.

[0088] In a feasible implementation, the present invention proposes a multi-spectral discriminator to score the denoised speech, thereby better improving the denoising quality of the speech denoising model. In the frequency discriminator, low-resolution processing is performed in the low-resolution processing layer to reduce the detail level of the feature so as to more efficiently focus on the main mode of the frequency information. Based on the feature normalization layer, the feature passes through the weight normalization module to ensure the stability of the data distribution, which helps to optimize the learning effect of the model. After a 3×3 convolution layer, the local frequency features are further extracted.

[0089] The output of the frequency discriminator is passed to a summary layer, which combines the results of all branches and finally calculates the adversarial loss. This loss is used to measure the authenticity of the audio, allowing the entire network to effectively identify and evaluate the frequency characteristics of the audio and optimize the performance of the denoising model.

[0090] S4. In a low signal-to-noise ratio environment, the speech data to be denoised is collected through a microphone; the speech data to be denoised is input into the TFDense-Net speech denoising model to obtain denoised speech data.

[0091] In a feasible implementation, the audio signal to be denoised is transformed into a frequency domain representation through a short-time Fourier transform and then enters the TFDense-Net speech denoising model. In this model, frequency domain features are extracted by an encoder; the interaction of time-frequency information is further enhanced by a time-frequency converter module. These features are gradually restored by a decoder and converted back to the time domain through an inverse short-time Fourier transform to generate a reconstructed audio signal.

[0092] The trained model can automatically activate the noise feature extraction mechanism it has learned. Through the forward propagation process, the model quickly identifies and locates the noise areas, and uses the denoising network to process these areas to achieve clarity of speech signals and complete the denoising task.

[0093] In a feasible implementation, the present invention uses a self-made data set to evaluate the effect of the TFDense-Net denoising model. The evaluation indicators used are short-time objective intelligibility and objective speech quality assessment, which are important indicators for measuring the quality of speech signals.

[0094] Short-term objective intelligibility is an indicator used to evaluate the intelligibility of speech signals. The larger the short-term objective intelligibility, the higher the clarity and comprehensibility of the speech signal. Objective speech quality assessment is a standard indicator used to evaluate the quality of speech signals. The larger the objective speech quality assessment, the higher the subjective assessment of speech quality.

[0095] Whether the discriminator can improve the model effect is verified by using the discriminator on the same speech denoising model TFDense-Net. The verification results are shown in Table 1 (the denoising effect table of TFDense-Net on the dataset with and without the discriminator).

[0096] Table 1

[0097]

[0098] The denoising effect of TFDense-Net under different signal-to-noise ratios is shown in Table 2 (Table of denoising effect of TFDense-Net under different signal-to-noise ratios).

[0099] Table 2

[0100]

[0101] An objective comparison was conducted with other state-of-the-art denoising models on this dataset. It turns out that the TFDense-Net denoising model has achieved better performance in both denoising performance and sound quality, further verifying its competitiveness in speech enhancement tasks. The denoising effect is shown in Table 3 (denoising effect table of different models on this dataset).

[0102] Table 3

[0103]

[0104] The present invention proposes a speech denoising method for low signal-to-noise ratio. By introducing a time-frequency converter module between the encoder and decoder of TFDense-Net, the time domain and frequency domain features in the audio signal can be captured simultaneously. In a complex noise environment, the audio denoising performance can be effectively enhanced and the speech clarity can be improved. This module can also reduce the feature loss caused by noise interference, improve the quality of audio reconstruction, effectively remove background noise while maintaining the naturalness and clarity of speech, and improve the audio denoising effect.

[0105] For the optimization of the dense block module, a combination of deep convolution and point convolution is used to improve the dense block, so that it not only retains the advantages of dense connections in the dense network structure, but also expands the receptive field through dilated convolution to better capture multi-scale information in the time and frequency domain. This module is used in both the encoder and decoder of TFDense-Net, and can effectively retain and transmit detailed information of speech. The extraction and fusion capabilities of audio signal features are optimized, so that the network can better maintain the key information in the speech, improve the noise reduction effect and reduce the loss of details.

[0106] The proposed TFDense-Net adopts a U-shaped network convolutional architecture, combined with a time-frequency converter and an improved dense block, for feature extraction and reconstruction of audio signals. The encoder gradually compresses the time-frequency features by downsampling, the bottleneck layer uses a self-attention mechanism for global feature fusion, and the decoder restores the signal by upsampling. This architecture ensures the model's ability to capture time-frequency features while maintaining a low computational complexity, and is suitable for audio denoising tasks. While ensuring the denoising effect, the computational complexity and parameter amount of the model are greatly reduced, and the training efficiency and inference speed of the model are improved. The present invention is an efficient and clear speech denoising method under low signal-to-noise ratio that combines an improved dense block and a video converter.

[0107] Figure 2 1 is a block diagram of a speech denoising device for low signal-to-noise ratio according to an exemplary embodiment, wherein the device is used in a speech denoising method for low signal-to-noise ratio. Figure 2 The device includes a training speech acquisition module 210, a model construction module 220, a model training module 230 and a speech denoising module 240. Among them:

[0108] The training speech acquisition module 210 is used to record audio through a microphone to obtain pure speech data; pre-process the denoised speech pure speech data to obtain training speech data;

[0109] A model building module 220 is used to build a TFDense-Net speech denoising model to be trained according to the denoised speech U-net network structure and the Transformer model structure;

[0110] The model training module 230 is used to perform adversarial iterative training on the TFDense-Net speech denoising model to be trained for the denoised speech based on the multi-spectral discriminator and the denoised speech training speech data using the Adam optimizer to obtain the TFDense-Net speech denoising model;

[0111] The speech denoising module 240 is used to collect speech data to be denoised through a microphone in a low signal-to-noise ratio environment; the denoised speech data to be denoised is input into the denoised speech TFDense-Net speech denoising model to obtain denoised speech data.

[0112] Optionally, the training speech acquisition module 210 is further used to:

[0113] Unifying the audio length of the pure voice data to obtain first processed voice data;

[0114] Based on a sampling rate of 16 kHz, resampling the first processed voice data to obtain second processed voice data;

[0115] Performing data filtering on the second processed voice data to obtain third processed voice data;

[0116] According to the third processed speech data, noise is added by adopting a distortion modeling method to obtain training speech data.

[0117] Among them, the TFDense-Net speech denoising model includes an encoder, a decoder, and a bottleneck layer;

[0118] The encoder is composed of multiple identical modules stacked together; each module is composed of an improved dense block, a 3×3 convolutional layer, and a parameterized rectified linear unit activation function;

[0119] The bottleneck layer includes a transformer module and a global self-attention module; the transformer module includes a time-frequency transformer, a frequency transformer, and a global transformer;

[0120] The decoder includes a number of identical modules; the structure of the decoder is symmetrical to that of the encoder.

[0121] Among them, the improved dense block refers to the dense block optimized by using the dilated convolution layer and the point-by-point convolution layer; the improved dense block is used for the feature integration of the input TFDense-Net speech denoising model data.

[0122] Optionally, the model training module 230 is further used to:

[0123] Perform short-time Fourier transform on the training speech data to obtain the training speech Mel spectrum;

[0124] Construct a generative adversarial network based on the multi-spectral discriminator and the TFDense-Net speech denoising model to be trained;

[0125] According to the training speech data and the training speech Mel spectrum, the generative adversarial network is trained adversarially to obtain an adversarial loss function; the adversarial loss function includes a time domain loss function and a time-frequency domain loss function;

[0126] According to the adversarial loss function, the Adam optimizer is used to adjust the parameters of the TFDense-Net speech denoising model to be trained to obtain the TFDense-Net speech denoising model.

[0127] Optionally, the model training module 230 is further used to:

[0128] Use the training speech Mel spectrum to perform frequency domain speech denoising on the training TFDense-Net speech denoising model to obtain the denoised training speech Mel spectrum;

[0129] Perform inverse short-time Fourier transform on the mel spectrum of the denoised training speech to obtain denoised training speech data;

[0130] Copying the denoised training speech data and the training speech data to obtain a plurality of training speech branch data and a plurality of denoised training speech branch data;

[0131] Performing multi-frequency domain short-time Fourier transform on multiple training speech branch data and multiple denoising training speech branch data to obtain multi-frequency domain training speech frequency domain features and multi-frequency domain denoising training speech frequency domain features;

[0132] According to the frequency domain characteristics of the training speech and the frequency domain characteristics of the denoised training speech, the loss function is summarized and calculated through a multi-spectral discriminator to obtain an adversarial loss function.

[0133] Among them, the multi-spectral discriminator includes multiple processing branches and a summarization layer;

[0134] The processing branch includes a low-resolution processing layer, a feature normalization layer, and a 3×3 convolutional layer;

[0135] The aggregation layer is used to aggregate and calculate the frequency domain features of the outputs of multiple processing branches.

[0136] The present invention proposes a speech denoising method for low signal-to-noise ratio. By introducing a time-frequency converter module between the encoder and decoder of TFDense-Net, the time domain and frequency domain features in the audio signal can be captured simultaneously. In a complex noise environment, the audio denoising performance can be effectively enhanced and the speech clarity can be improved. This module can also reduce the feature loss caused by noise interference, improve the quality of audio reconstruction, effectively remove background noise while maintaining the naturalness and clarity of speech, and improve the audio denoising effect.

[0137] For the optimization of the dense block module, a combination of deep convolution and point convolution is used to improve the dense block, so that it not only retains the advantages of dense connections in the dense network structure, but also expands the receptive field through dilated convolution to better capture multi-scale information in the time and frequency domain. This module is used in both the encoder and decoder of TFDense-Net, and can effectively retain and transmit detailed information of speech. The extraction and fusion capabilities of audio signal features are optimized, so that the network can better maintain the key information in the speech, improve the noise reduction effect and reduce the loss of details.

[0138] The proposed TFDense-Net adopts a U-shaped network convolutional architecture, combined with a time-frequency converter and an improved dense block, for feature extraction and reconstruction of audio signals. The encoder gradually compresses the time-frequency features by downsampling, the bottleneck layer uses a self-attention mechanism for global feature fusion, and the decoder restores the signal by upsampling. This architecture ensures the model's ability to capture time-frequency features while maintaining a low computational complexity, and is suitable for audio denoising tasks. While ensuring the denoising effect, the computational complexity and parameter amount of the model are greatly reduced, and the training efficiency and inference speed of the model are improved. The present invention is an efficient and clear speech denoising method under low signal-to-noise ratio that combines an improved dense block and a video converter.

[0139] Figure 3 is a structural diagram of a speech denoising device provided by an embodiment of the present invention, such as Figure 3 As shown, the speech denoising device may include the above Figure 2 The speech denoising device shown is used for speech denoising under low signal-to-noise ratio. Optionally, the speech denoising device 310 may include a first processor 2001 .

[0140] Optionally, the speech denoising device 310 may further include a memory 2002 and a transceiver 2003 .

[0141] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0142] Combine the following Figure 3 The components of the speech denoising device 310 are introduced in detail:

[0143] The first processor 2001 is the control center of the speech denoising device 310, and may be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be application specific integrated circuits (ASICs), or may be configured to implement one or more integrated circuits of the embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs).

[0144] Optionally, the first processor 2001 can perform various functions of the speech denoising device 310 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.

[0145] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 3 CPU0 and CPU1 are shown in FIG.

[0146] In a specific implementation, as an embodiment, the speech denoising device 310 may also include multiple processors, such as Figure 3 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0147] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0148] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently, and access the first processor 2001 through the interface circuit ( Figure 3 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0149] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0150] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0151] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently and communicate with the first processor 2001 through the interface circuit ( Figure 3 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0152] It should be noted that Figure 3 The structure of the speech denoising device 310 shown in the figure does not constitute a limitation on the router, and the actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0153] In addition, the technical effects of the speech denoising device 310 can refer to the technical effects of the speech denoising method for low signal-to-noise ratio described in the above method embodiment, which will not be repeated here.

[0154] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0155] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0156] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0157] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0158] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0159] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0160] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0161] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0162] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0163] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0164] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0165] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0166] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A speech denoising method for low signal-to-noise ratio, characterized in that: The method comprises: Recording audio through a microphone to obtain clean voice data; preprocessing the clean voice data to obtain training voice data; Construct the TFDense-Net speech denoising model to be trained based on the U-net network structure and the Transformer model structure; Based on the multi-spectral discriminator, according to the training speech data, using the Adam optimizer to perform adversarial iterative training on the TFDense-Net speech denoising model to be trained, to obtain the TFDense-Net speech denoising model; The TFDense-Net speech denoising model includes an encoder, a decoder and a bottleneck layer; The encoder is composed of a plurality of identical modules stacked together; each module is composed of an improved dense block, a 3×3 convolutional layer, and a parameterized rectified linear unit activation function; The bottleneck layer includes a transformer module and a global self-attention module; the transformer module includes a time-frequency transformer, a frequency transformer and a global transformer; The decoder comprises a plurality of identical modules; the structure of the decoder is symmetrical to that of the encoder; The method of performing adversarial iterative training on the TFDense-Net speech denoising model to be trained based on the training speech data using the Adam optimizer to obtain the TFDense-Net speech denoising model includes: Performing short-time Fourier transform on the training speech data to obtain a training speech Mel spectrum; Constructing a generative adversarial network based on the multi-spectral discriminator and the TFDense-Net speech denoising model to be trained; According to the training speech data and the training speech Mel spectrum, the generative adversarial network is subjected to adversarial training to obtain an adversarial loss function; the adversarial loss function includes a time domain loss function and a time-frequency domain loss function; According to the adversarial loss function, using the Adam optimizer to adjust the parameters of the TFDense-Net speech denoising model to be trained to obtain the TFDense-Net speech denoising model; In a low signal-to-noise ratio environment, speech data to be denoised is collected by a microphone; the speech data to be denoised is input into the TFDense-Net speech denoising model to obtain denoised speech data.

2. The method for speech denoising under low signal-to-noise ratio according to claim 1, characterized in that: The preprocessing of the clean voice data to obtain training voice data includes: Unifying the audio length of the clean voice data to obtain first processed voice data; Based on a sampling rate of 16 kHz, resampling the first processed voice data to obtain second processed voice data; performing data filtering on the second processed voice data to obtain third processed voice data; According to the third processed speech data, noise is added by adopting a distortion modeling method to obtain training speech data.

3. The speech denoising method for low signal-to-noise ratio according to claim 1, characterized in that: The improved dense block refers to a dense block optimized by using an expanded convolutional layer and a point-by-point convolutional layer; the improved dense block is used for feature integration of input TFDense-Net speech denoising model data.

4. The method for speech denoising under low signal-to-noise ratio according to claim 1, characterized in that: The method of performing adversarial training on the generative adversarial network based on the multi-spectral discriminator and the training speech data and the training speech Mel spectrum to obtain an adversarial loss function includes: Using the training speech Mel spectrum, performing frequency domain speech denoising on the TFDense-Net speech denoising model to be trained to obtain a denoised training speech Mel spectrum; Performing an inverse short-time Fourier transform on the mel spectrum of the denoised training speech to obtain denoised training speech data; Copying the denoised training speech data and the training speech data to obtain a plurality of training speech branch data and a plurality of denoised training speech branch data; Performing multi-frequency domain short-time Fourier transform on the multiple training speech branch data and the multiple denoised training speech branch data to obtain multi-frequency domain training speech frequency domain features and multi-frequency domain denoised training speech frequency domain features; According to the training speech frequency domain features and the denoised training speech frequency domain features, a loss function is summarized and calculated through a multi-spectral discriminator to obtain an adversarial loss function.

5. The method for speech denoising under low signal-to-noise ratio according to claim 1, characterized in that: The multi-spectral discriminator includes a plurality of processing branches and a summarization layer; The processing branch includes a low-resolution processing layer, a feature normalization layer and a 3×3 convolution layer; The aggregation layer is used to aggregate and calculate the frequency domain features output by the multiple processing branches.

6. A speech denoising device for low signal-to-noise ratio, the speech denoising device for low signal-to-noise ratio being used to implement the speech denoising method for low signal-to-noise ratio as claimed in any one of claims 1 to 5, characterized in that: The device comprises: A training voice acquisition module is used to record audio through a microphone to obtain clean voice data; pre-process the clean voice data to obtain training voice data; The model building module is used to build the TFDense-Net speech denoising model to be trained based on the U-net network structure and the Transformer model structure; A model training module is used to perform adversarial iterative training on the TFDense-Net speech denoising model to be trained based on the multi-spectral discriminator and the training speech data using an Adam optimizer to obtain the TFDense-Net speech denoising model; The speech denoising module is used to collect the speech data to be denoised through a microphone in a low signal-to-noise ratio environment; the speech data to be denoised is input into the TFDense-Net speech denoising model to obtain denoised speech data.

7. A speech denoising device, characterized in that: The speech denoising device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Single-channel speech enhancement method based on deep reconvolution network

    CN114360567A

  • Attention generation confrontation speech enhancement method based on joint perception loss

    CN115410589A