Real-time full-band voice noise reduction enhancement method based on hybrid DSP deep learning method
By combining mixed DSP and deep learning methods in speech noise reduction technology, real-time noise reduction and enhancement of full-band voice signals is achieved, and the problems of poor adaptability, insufficient real-time and high computational complexity in complex noise environments in the prior art are solved, which significantly improves the quality and processing efficiency of voice signals.
Patent Information
- Application Number
- CN202510306980.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-15
- Publication Date
- 2025-06-06
AI Technical Summary
The existing speech noise reduction technology has poor adaptability, insufficient real-time performance and high computational complexity in complex noise environments.
Real-time full-band speech noise reduction enhancement method based on hybrid DSP deep learning method is adopted to achieve real-time noise reduction and enhancement of voice signals through steps such as signal preprocessing, spectrum division, feature extraction, gain estimation, gain smoothing and gain application.
It significantly improves the clarity and intelligibility of the voice signal, solves the problems of speech distortion or unclear caused by traditional methods in complex noisy environments, ensures the naturalness and reality of the speech, and achieves low latency and efficient processing performance on low-power devices.
Smart Images

Figure CN120108412A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of speech signal processing and deep learning, and in particular to a real-time full-band speech noise reduction and enhancement method based on a hybrid DSP deep learning method. Background Art
[0002] With the widespread application of intelligent communication devices and speech recognition technology, speech noise reduction and enhancement technology has become one of the core technologies to improve the quality of voice communication and speech recognition accuracy. Especially in noisy environments, the quality of speech signals will be seriously affected. Traditional speech enhancement technology faces many challenges in such environments. In order to meet these challenges, researchers have been seeking more efficient and adaptable speech noise reduction methods.
[0003] At present, speech noise reduction and enhancement technology mainly relies on two categories of methods: traditional digital signal processing (DSP) technology and deep learning-based technology. Traditional DSP methods remove noise by performing frequency domain analysis, filtering, and gain adjustment on the signal. These methods usually rely on preset noise models. However, traditional DSP methods have great limitations, especially under complex background noise. The fixedness of the noise model makes its performance unstable in different noise environments. In addition, traditional DSP methods also have certain problems in computational complexity and processing delay, especially in real-time applications, it is difficult to meet the requirements of low latency and high efficiency.
[0004] On the other hand, in recent years, deep learning technology has been widely used in the field of speech noise reduction, especially the successful application of convolutional neural networks (CNN) and recurrent neural networks (RNN) in feature extraction and gain estimation. Deep learning methods can automatically learn the difference between noise and speech according to the input signal characteristics, thereby achieving better noise reduction effects. Compared with traditional DSP methods, deep learning methods can have better adaptability and robustness in dynamic noise environments. However, deep learning technology still has some problems, especially in terms of real-time and computing resources. Since the training of deep learning models usually requires a lot of computing resources and data, real-time processing is still a big challenge.
[0005] At present, most methods in the field of speech noise reduction either rely on fixed noise models and lack adaptability, or have difficulty meeting real-time requirements due to the high computational overhead of deep learning models. In addition, when processing full-band speech signals, most existing technologies use full-band processing, ignoring the different contributions of different frequency bands to speech signals, resulting in low processing efficiency and unsatisfactory results. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention provides a real-time full-band speech noise reduction enhancement method based on a hybrid DSP deep learning method, which solves the problems of poor adaptability, insufficient real-time performance and high computational complexity in the existing speech noise reduction technology in complex noise environments.
[0007] To achieve the above objectives, the present invention is implemented by the following technical scheme: a real-time full-band speech noise reduction and enhancement method based on a hybrid DSP deep learning method, comprising the following steps: Signal preprocessing: The input time-domain speech signal is framed to obtain a series of speech frames, and a window function is applied to each frame of speech signal; Spectrum division: performing fast Fourier transform on each frame signal to obtain spectrum data, and dividing the spectrum data into multiple frequency bands to form a multi-band spectrum; Feature extraction: extract the features of each frequency band and process them through a convolutional neural network to generate high-level features for gain estimation; Gain estimation: Processing the high-level features through a deep learning model to output gain estimates for each frequency band; Gain smoothing: performing gain smoothing on the gain estimation value to ensure smooth transition of the spectrum gain in time and avoid distortion of the speech signal caused by sudden changes; Gain application: Apply the gain to the corresponding frequency band spectrum to obtain enhanced spectrum data; Inverse transform restoration: Perform inverse fast Fourier transform on the enhanced spectrum data to convert the frequency domain signal into a time domain signal to obtain a de-noised speech signal.
[0008] Preferably, the signal preprocessing step comprises the following steps: Performing frame segmentation on the input time domain speech signal, dividing the speech signal into fixed-length frames with a frame length of 15 ms to 25 ms and a frame overlap rate of 30% to 50%; Applying a window function to each frame of speech signal, wherein the window function is a Hamming window or a Hanning window, for reducing spectrum leakage; Perform fast Fourier transform on each frame of windowed signal to convert the time domain signal into frequency domain signal to obtain spectrum data X b (t).
[0009] Preferably, the spectrum division step comprises the following steps: The obtained spectrum data X b (t) Applying the Bark scale, the spectrum is divided into 22 frequency bands to reduce computational complexity and preserve speech characteristics; The spectral energy of each frequency band is extracted and used as the input feature of the subsequent deep learning model.
[0010] Preferably, the feature extraction step comprises the following steps: Extract spectrum energy, pitch, and spectrum non-stationarity features from the spectrum data of each frequency band. These features can reflect the difference between speech and noise. The extracted features are input into the convolutional neural network for adaptive feature learning to generate high-level feature representations.
[0011] Preferably, the gain estimation step comprises the following steps: The frequency band features are input into multiple convolutional layers through a convolutional neural network for local feature extraction; Perform pooling operations on the extracted features to reduce feature dimensions and retain key information; The pooled features are processed through a fully connected layer to output the gain estimate g for each frequency band b (t).
[0012] Preferably, the gain smoothing step comprises the following steps: The gain estimate g for each frequency band is b (t) Perform Kalman filtering to smooth the gain and avoid sudden changes; Calculate the Kalman gain K(t); Update the gain estimate g b (t), ensuring a smooth transition of the gain in time.
[0013] Preferably, the Kalman gain formula is as follows: Among them, P(t-1) is the gain estimation error covariance at the previous moment, and R(t) is the noise covariance at the current moment.
[0014] Preferably, the gain estimation step includes the following model compression and optimization techniques: Use L1 regularization to prune the neural network and remove unimportant neurons and connections; Quantize the network weights and convert floating point numbers to integers to reduce computation and memory usage.
[0015] The present invention provides a real-time full-band speech noise reduction and enhancement method based on a hybrid DSP deep learning method. It has the following beneficial effects: 1. The present invention adopts speech noise reduction and enhancement technology based on hybrid DSP deep learning method, which can significantly improve the clarity and intelligibility of speech signals through precise frequency band gain adjustment and gain smoothing. Compared with the traditional noise suppression method in the prior art, the present invention optimizes gain estimation through deep learning model, solves the problem that traditional methods cause speech distortion or unclearness when dealing with complex noise environments, and ensures the naturalness and realism of speech.
[0016] 2. The present invention enhances the robustness of the model through deep learning model training and adversarial training, so that it can adapt to different types of noise environments. Compared with the method of using a fixed noise suppression model in the prior art, the present invention can dynamically adjust the gain value and spectrum adjustment strategy, thereby effectively improving the quality of speech signals in a wider range of noise environments. Especially in dynamic noise environments, the present invention can effectively remove noise and enhance speech signals, significantly improving the intelligibility of speech signals.
[0017] 3. The present invention effectively reduces the computational complexity through the optimized design of multi-band processing and gain application of speech signals. During processing, the spectrum division technology is adopted so that each frequency band can be processed independently, thereby reducing unnecessary computational burden. Compared with the traditional full-band processing method, the present invention significantly improves the computational efficiency and processing speed by optimizing the algorithm and model compression while ensuring the speech enhancement effect, ensuring that the system can run with low latency in real-time speech processing, and is particularly suitable for low-power devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] Please refer to the attached Figure 1 The embodiment of the present invention provides a real-time full-band speech noise reduction and enhancement method based on a hybrid DSP deep learning method, comprising the following steps: S1. Signal preprocessing: Frame the input time-domain speech signal to obtain a series of speech frames, and apply a window function to each frame of speech signal; The purpose of the signal preprocessing in this embodiment is to provide high-quality input data for subsequent spectrum analysis and feature extraction by performing operations such as framing and windowing on the input time-domain speech signal.
[0021] In this embodiment, the signal preprocessing step includes the following key operations: Generally speaking, the input time domain speech signal is a continuous signal, and directly performing spectrum analysis on the entire signal may lead to inaccurate results. In order to analyze the local characteristics of the speech signal, the signal is usually divided into several short signal frames, each frame signal is relatively independent, so that the signal can be better presented in both time and frequency domains. Frame segmentation is to reduce the computational complexity without losing key information.
[0022] In one implementation, assuming that the length of each frame is set to 20ms and the frame overlap rate is set to 50%, the number of samples contained in each frame is: frame length=sampling rate×frame duration=160samples; Assuming the sampling rate is 8kHz, each frame of the signal contains 160 samples.
[0023] Each frame of signal is extracted from the entire signal according to its starting position, and there is a 30% to 50% overlap between adjacent frames. Frame overlap helps improve the continuity of speech signals, especially in spectrum analysis, which can effectively avoid breaks between time windows and ensure full utilization of information.
[0024] After the framing process is completed, in order to reduce the sudden change of the signal at the frame boundary and avoid spectrum leakage, a window function needs to be applied to each frame of the signal. Common window functions include Hamming window, Hanning window, Blackman window, etc. In this embodiment, the Hamming window is selected to perform windowing on each frame of the signal.
[0025] The expression of the Hamming window is as follows: Wherein: w(n) is the window function value of the nth sample; N is the length of the window function, which is usually set to be consistent with the length of the signal frame; j is the index of the window function, 0≤n≤N-1.
[0026] Specifically, assuming that the length of each frame signal is N = 256 points, the window function value w(n) will be applied to the samples of each frame signal. Through the windowing operation, the sample values at both ends of the signal will gradually decrease, thereby reducing the impact of the boundary effect on the spectrum estimation.
[0027] As an option, other types of window functions can be considered, such as the Hanning window, which is expressed as: Compared with the Hamming window, the spectrum leakage effect of this window function is slightly larger, but it can still effectively reduce the discontinuity of the signal endpoints and is suitable for certain specific speech signal processing requirements.
[0028] Through the above frame segmentation and windowing operations, each frame of the signal is processed into a smooth short-time signal in the time domain. Next, we perform spectrum analysis. To this end, we need to perform a fast Fourier transform (FFT) on each frame of the signal to convert the time domain signal into a frequency domain signal.
[0029] Specifically, for each frame signal X(k), its frequency domain representation X(k) can be obtained by applying FFT, and its calculation formula is: Where: X(k) is the amplitude information of the kth frequency point of the frequency domain signal; x(n) is the nth sample of the time domain signal; N is the number of sample points per frame signal; k is the index of the output frequency point, 0≤k≤N-1.
[0030] Through this FFT operation, the spectrum information of each frame signal is obtained, which can reveal the energy distribution of the frame signal in the frequency domain. The frequency domain signal will be used as a feature input for gain estimation and noise suppression in subsequent steps.
[0031] S2, spectrum division: performing fast Fourier transform on each frame signal to obtain spectrum data, and dividing the spectrum data into multiple frequency bands to form a multi-band spectrum; The spectrum division step of this embodiment is to process the spectrum data of each frame and divide it into multiple frequency bands, thereby forming a multi-band spectrum, which is convenient for subsequent feature extraction and gain estimation. Through this step, the spectrum information of the signal can be effectively decomposed into multiple sub-bands, which is convenient for processing each frequency band separately, achieving better speech noise reduction and enhancement effects.
[0032] In the aforementioned step S1, after the frame processing and windowing by the window function, the spectrum data of each frame of the speech signal has been generated, and the spectrum of each frame has been converted in the frequency domain and becomes suitable for further processing. In order to perform further gain estimation in a multi-band range, the spectrum data of each frame will be spectrum divided next.
[0033] In this embodiment, in order to reduce the computational complexity and effectively maintain the characteristics of the speech signal, the Bark scale is used for spectrum division. The Bark scale is a perceptual frequency scale that is more consistent with the auditory characteristics of the human ear, and thus can better characterize the frequency characteristics of the speech signal. Specifically, the Bark scale captures speech information in different frequency bands by dividing the spectrum data into multiple frequency bands, thereby providing a more detailed and accurate frequency response for subsequent gain estimation.
[0034] Generally, the goal of spectrum division is to divide the spectrum data into multiple frequency bands, each of which contains a certain range of frequency information. In this embodiment, the spectrum data is divided into 22 frequency bands, each of which corresponds to a specific frequency range. In this division method, the spectrum information of the speech signal can more finely reflect the energy distribution of the speech signal in different frequency segments.
[0035] Specifically, assuming that the FFT length of each frame signal is N, the spectrum data obtained after FFT conversion will be divided into several frequency bands. The spectrum energy of each frequency band will be used as the input feature of the subsequent deep learning model. This division method can effectively reduce the computational complexity while retaining the main frequency characteristics of the speech signal.
[0036] For example, in one possible implementation, assume that the frequency range of the spectrum of each frame signal after FFT conversion is 0Hz to 4kHz, and within this range, the Bark scale divides the frequency range into 22 bands, so that the frequency range covered by each band will gradually increase according to the design of the Bark scale. In this way, the spectrum information of different frequency bands can effectively capture the details of the speech signal, especially the energy changes in the low and high frequency bands.
[0037] In order to achieve the division of Bark scale, the following conversion formula is usually used: Where: f b is the center frequency of the bth frequency band; f is the frequency in Hertz (Hz); and b is the index of the frequency band.
[0038] Through this formula, the frequency range of each frequency band can be calculated, and the spectrum data can be divided into corresponding frequency bands to form a multi-band spectrum.
[0039] As an option, in some application scenarios, the number of frequency bands or the width of the frequency bands may also be adjusted as needed. For example, some applications may require more frequency bands to capture high-frequency features more finely, or the number of frequency bands may be adjusted according to actual computing power to reduce computing overhead.
[0040] In this embodiment, not only the Bark scale but also other perceptual frequency scales (such as the Mel scale, the ERB scale, etc.) can be used to divide the spectrum data. Different frequency scales have different effects on different types of speech signal processing. Therefore, in different application scenarios, the most appropriate spectrum division method can be selected according to the needs.
[0041] The introduction of the spectrum division step divides the original spectrum data into multiple frequency bands, so that the signal of each frequency band can be processed independently, optimizing the accuracy of subsequent gain estimation. Through the Bark scale division method, the spectrum information of the signal can be more consistent with the auditory characteristics of the human ear and can accurately reflect the energy distribution of the speech signal in each frequency band. This can not only improve the noise reduction effect, but also enhance the intelligibility of the signal, especially in a noisy environment.
[0042] In addition, spectrum division also effectively reduces computational complexity. Compared with directly processing the entire spectrum, by dividing it into multiple frequency bands for independent processing, the amount of spectrum analysis calculation can be reduced and the processing of each frequency band can be made more efficient.
[0043] S3, feature extraction: extracting features of each frequency band and processing the features through a convolutional neural network to generate high-level features for gain estimation; The feature extraction step of this embodiment is intended to extract high-level information that can effectively represent the characteristics of the speech signal from the spectrum data of each frequency band. Through this process, the system can identify the key signal characteristics of the speech signal and provide accurate input data for subsequent gain estimation. Feature extraction is a key step in speech signal processing, especially in noisy environments, where accurate feature extraction can significantly improve speech enhancement effects.
[0044] In the aforementioned step S2, a fast Fourier transform (FFT) is performed on each frame signal and the spectrum data is divided into multiple frequency bands, thereby obtaining a multi-band spectrum of each frame signal. This spectrum data contains the energy distribution of the speech signal in each frequency band and is the basis for subsequent feature extraction and gain estimation.
[0045] In this embodiment, these frequency band features are processed by a convolutional neural network (CNN) to extract more abstract and discriminative high-level features. These high-level features can fully express the speech characteristics of the speech signal and provide strong support for the gain estimation step.
[0046] In this embodiment, the basic features of the spectrum data of each frequency band are first extracted. Generally, the spectrum data of the frequency band may include information such as spectrum energy, frequency distribution, pitch, etc. These basic features can effectively reflect the energy distribution and corresponding frequency features of the speech signal of each frequency band.
[0047] Specifically, spectral energy is the sum of the frequency domain information of each frequency band, representing the "intensity" of the frequency band. Pitch reflects the frequency component of the speech signal, and spectral non-stationarity can reflect the time-varying characteristics of the signal. In speech signals, non-stationary characteristics are particularly important.
[0048] In the formula, it is assumed that the spectrum data of the bth frequency band is X b (t), the spectrum energy of the frequency band can be calculated by the following formula: Where: E b (t) is the energy of the bth frequency band at time t; X b (t,k) is the spectrum value of the b-th frequency band at time t and output frequency point k; N b is the frequency resolution of the bth frequency band.
[0049] This formula provides an effective measure of the energy of each frequency band in the time domain, which can provide important information for subsequent feature extraction.
[0050] As an option, the pitch feature can be obtained through the main frequency component in the spectrum. For each frequency band, the pitch can be calculated based on the weighted average of the frequency or obtained through peak detection and other methods. This can further capture the pitch changes in the speech signal and provide more detailed signal features for subsequent gain estimation.
[0051] After the basic features are extracted, the extracted features are processed using a convolutional neural network (CNN) to generate higher-level, recognizable features. The CNN structure consists of multiple convolutional layers and pooling layers, which can automatically learn the deep relationship between features of different frequency bands and extract more abstract feature representations.
[0052] Specifically, the convolution layer of CNN filters the input frequency band features and extracts local features. After each convolution operation, the feature dimension gradually decreases, but the information contained is richer. The convolution operation can be expressed by the following formula: Among them: F b (t,k) is the feature map output by the convolutional layer; X b (t,n) is the spectrum value of the b-th frequency band at the input frequency point y at time t; W b (k, y) is the weight parameter of the convolution kernel; b b is the bias of the convolutional layer.
[0053] Through the convolution operation, the convolution kernel W b (k,n) can effectively extract local patterns in frequency band features and help identify local features in speech signals, especially information such as pitch and timbre in speech signals.
[0054] In one possible implementation, the output feature map of the convolutional layer is reduced in dimension by the pooling layer. The pooling operation can effectively compress the dimension of the feature while retaining key information and avoiding redundant data. In the pooling operation, maximum pooling or average pooling is generally used, where maximum pooling retains the most important features by taking the maximum value in the feature map.
[0055] Finally, the features processed by the convolutional layer and the pooling layer will be further combined and optimized by the fully connected layer to generate high-level features for gain estimation. The fully connected layer is processed by the following formula: Where: H b (t) is a high-level feature; W fc is the weight of the fully connected layer; b fc is the bias; F b (t,k) is the feature map output by the convolutional layer.
[0056] In this embodiment, the final high-level feature H b (t) will be provided as input to the subsequent gain estimation module to ensure accurate gain estimation based on the characteristics of each frequency band.
[0057] Through the feature extraction step, the original frequency band features are converted into high-level feature representations, thereby improving the recognition of important information in the speech signal. These high-level features can more accurately reflect the content of the speech signal and contribute to the accuracy of subsequent gain estimation.
[0058] Using convolutional neural networks (CNN) for feature processing can automatically learn the complex nonlinear relationship between frequency bands, further improving the system's processing capabilities, especially in noisy environments, to better extract the key features of the signal. This step can effectively improve the speech noise reduction and enhancement effects, and enhance the system's adaptability and robustness in various environments.
[0059] S4, gain estimation: processing the high-level features through a deep learning model and outputting a gain estimation value for each frequency band; In the aforementioned step S3, the features of each frequency band are processed by a convolutional neural network (CNN) to generate high-level features for gain estimation. These high-level features already contain the main speech components of the speech signal, noise information, and the relationship between frequency bands, and are the basis for subsequent gain estimation. The purpose of gain estimation is to output the gain value of each frequency band based on these high-level features, thereby optimizing the signal and improving the quality of the speech signal.
[0060] In this embodiment, gain estimation is achieved by processing high-level features with a deep learning model. The deep learning model can predict the gain value of each frequency band based on the input high-level features (i.e., the features extracted by CNN above). This gain value indicates how to adjust the amplitude of each frequency band to achieve the best speech enhancement effect.
[0061] In general, gain estimation can be inferred through a trained deep neural network model, which includes multiple hidden layers and is trained through a back-propagation algorithm, so that the network can effectively map input features to output gain values.
[0062] Specifically, the gain estimation model is usually a regression model whose goal is to output a gain coefficient for each frequency band. The training process of the regression model is optimized by minimizing the loss function to minimize the difference between the output gain value and the true gain value. The goal of gain estimation is to estimate the gain coefficient g of each frequency band based on the characteristics of the frequency band. b (t), this gain coefficient represents the gain value of this frequency band.
[0063] For example, suppose we have a set of high-level features H b (t), these features are the features extracted by CNN in the previous step. The gain estimation model accepts these features as input and outputs the gain g of the corresponding frequency band b (t), the gain is used to adjust the spectrum value of each frequency band. Gain g b (t) can be estimated by the following regression formula: g b (t) = f(H b (t),W g ,b g ); Where: g b (t) is the gain estimate of the b-th frequency band at time t, which represents the gain coefficient of the frequency band; H b (t) is the high-level feature of the b-th frequency band at time t, which comes from the features extracted by CNN. g is the weight of the gain estimation model, representing the parameters of each layer in the network, b g is the bias term of the gain estimation model, which is used to adjust the output result. f(·) is the mapping function of the gain estimation, which is usually a deep neural network or linear regression model, performing nonlinear transformation on the input features.
[0064] As an option, the gain estimation model can adopt a variety of neural network structures. For example, a simple fully connected neural network (MLP) is used to implement the regression estimation of the gain. In this case, each layer of the network processes the output of the previous layer and gradually adjusts the gain value of the frequency band. It is also possible to use a convolutional neural network (CNN) to further optimize the feature extraction and gain estimation process, especially in the local pattern recognition of the frequency band.
[0065] In one possible implementation, the gain estimation model estimates the gain not only based on a single feature, but also combines multiple different features (such as spectral energy, pitch, spectral non-stationarity, etc.). By inputting multiple features into the gain estimation network, the network can learn a more accurate gain value based on these diverse features.
[0066] The goal of gain estimation is to optimize the quality of the speech signal to the greatest extent possible. Usually, a loss function such as mean square error (MSE) is used to optimize the output of the model. The optimization process of the loss function can ensure that the error between the gain value output by the model and the ideal gain value is minimized.
[0067] The gain estimation step converts the high-level features obtained from the feature extraction into gain values for each frequency band. This gain value will be used to adjust the spectral data of the frequency band in the subsequent gain application step to improve the quality of the speech signal. In a noisy environment, gain estimation can effectively suppress noise components and preserve the clarity and intelligibility of the speech signal by accurately calculating the gain value of each frequency band.
[0068] The deep learning model of gain estimation can automatically learn the nonlinear relationship between frequency bands based on the input features, which enables it to effectively enhance speech under complex noise conditions. Through gain estimation, the system can provide personalized gain adjustment to achieve the best speech enhancement effect.
[0069] In the gain estimation step, in order to further improve computational efficiency and reduce resource consumption, the gain estimation model also includes the following model compression and optimization techniques, which help to reduce the amount of computation and memory usage without significantly losing performance, thereby ensuring efficient operation of the system in real-time speech processing.
[0070] In this embodiment, L1 regularization is used to prune the neural network to remove unimportant neurons and connections. L1 regularization imposes L1 norm constraints on the weights of the network, which causes the network to generate sparse weights during training, that is, some weights become zero. In this way, the network can remove neurons and connections that do not contribute significantly to the gain estimation results, thereby reducing the amount of computation and memory usage.
[0071] Specifically, the goal of L1 regularization is to minimize the loss function of the network while making the weight vector as sparse as possible. The pruned network is more compact and computationally efficient. This not only reduces the size of the model, but also speeds up the inference process, which is particularly important in application scenarios that require real-time processing. Through pruning, the structure of the neural network is simplified and computing resources are used more efficiently.
[0072] To further optimize network performance, the gain estimation model uses network weight quantization technology. Specifically, the floating-point weights in the network are converted to integer weights. The quantization process reduces the memory and computational precision required in storage and calculation by mapping floating-precision weight values to integer values in a limited range.
[0073] The quantized network requires less memory at runtime and can be computed with lower power consumption. This approach is particularly suitable for low-power devices such as mobile and embedded devices, where memory and computing power are limited. Through quantization, the inference speed can be significantly improved while reducing computing latency, ensuring that the system can run efficiently during real-time speech processing.
[0074] The computational and storage efficiency of the gain estimation model has been significantly improved through L1 regularization pruning and network weight quantization. These optimization techniques reduce the number of parameters required for calculation, reduce memory usage, and speed up the inference process, allowing the model to run efficiently on resource-constrained devices. For real-time speech processing tasks, especially in low-power environments, model compression and optimization techniques are critical to ensure that the system achieves low-latency and efficient processing performance while ensuring the quality of speech enhancement.
[0075] Through these optimization measures, the gain estimation step of the present invention can not only provide high-quality speech enhancement effects in various noise environments, but also run smoothly on devices with low power consumption and limited computing resources, providing wider applicability and flexibility for practical applications.
[0076] S5, gain smoothing: performing gain smoothing on the gain estimation value to ensure smooth transition of the spectrum gain in time and avoid distortion of the speech signal caused by sudden changes; In the aforementioned step S4 (gain estimation), high-level features are processed by a deep learning model and gain estimates for each frequency band are output. The gain estimates provide a suitable gain coefficient for each frequency band, aiming to enhance the speech signal. However, in practical applications, the gain estimates may change suddenly or dramatically over time, resulting in distortion of the speech signal. In order to ensure a smooth transition of the spectral gain over time and avoid the instability of the gain estimates affecting the signal quality, a gain smoothing step is introduced.
[0077] In this embodiment, the goal of gain smoothing is to adjust the gain estimate through smoothing processing so that the gain transition is smooth in the time dimension, avoiding audio distortion or unnatural effects caused by too fast changes in the gain estimate. Gain smoothing introduces a smoothing factor in time so that the gain values at the previous and next moments remain relatively consistent, thereby effectively reducing mutations.
[0078] Generally, gain smoothing is often achieved through filtering methods. Kalman filtering is a common method in gain smoothing, which can generate a smoothly transitioned gain value based on a weighted average of the gain estimate at the current moment and the gain estimate at the previous moment.
[0079] Specifically, in this embodiment, gain smoothing uses Kalman filtering to process the gain estimation value of each frequency band. Kalman filtering is a recursive filtering method, which calculates the smoothed gain value by weighted averaging the gain estimation value of the previous moment and the gain estimation value of the current moment at each moment. This method can flexibly control the degree of smoothing according to different weighting coefficients to ensure a smooth change of the gain value.
[0080] The core formula of Kalman filtering can be expressed as: in: is the smoothed gain estimate of the b-th frequency band at time t, indicating the gain after smoothing; is the gain estimate of the b-th frequency band at time t, indicating the gain value estimated by the deep learning model; K b (t) is the Kalman gain, which represents the weighting coefficient of the current moment to the gain of the previous moment; is the smoothing gain value of the b-th frequency band at the previous time t-1; is the gain estimate of the b-th frequency band at the previous time t-1.
[0081] Kalman gain K b The calculation method of (t) usually depends on the covariance of the gain estimation error and the noise, and the formula is as follows: Where: P(t-1) is the gain estimation error covariance at the previous moment, indicating the error range of the gain estimation at the previous moment; R(t) is the noise covariance at the current moment, indicating the influence of noise in the current gain estimation process.
[0082] The value of the Kalman gain determines the trade-off between the current gain estimate and the smoothing gain value. A smaller Kalman gain makes the gain more dependent on the smoothing gain at the previous moment, while a larger Kalman gain makes the current gain estimate have a greater impact on the smoothing result.
[0083] As an option, the Kalman filter can be replaced with other types of filters, such as weighted average filters, low-pass filters, etc., which can also achieve gain smoothing to a certain extent. However, the Kalman filter has better performance, especially when dealing with noise and dynamic changes, and can effectively balance the signal and noise.
[0084] In one possible implementation, the gain smoothing process is not limited to performing Kalman filtering on the gain estimate, but can also incorporate more time series information, such as using a long short-term memory network (LSTM) to process the gain estimate sequence. This method can take into account gain changes over a longer time span, which helps to further improve the effect of gain smoothing.
[0085] Through the gain smoothing step, the problem of sudden gain changes that may occur is effectively alleviated, thereby avoiding distortion in the speech signal. Gain smoothing makes the transition of spectral gain in time smooth and natural, ensuring the stability of the gain estimation process.
[0086] In a noisy environment, gain smoothing can not only reduce the fluctuation of the gain estimation value, but also improve the system's adaptability to dynamic noise. Through the smooth gain value, the speech enhancement effect is significantly improved, the speech signal is clearer and more natural, and the artifacts or unnatural audio effects caused by improper noise suppression are reduced.
[0087] S6, gain application: applying the gain to the corresponding frequency band spectrum to obtain enhanced spectrum data; In the aforementioned step S5 (gain smoothing), the gain estimation value of each frequency band has been smoothed by Kalman filtering to ensure smooth transition of gain in time and avoid distortion of speech signal caused by mutation. Next, the gain application step will apply the smoothed gain value of each frequency band to optimize the spectrum data and finally obtain the enhanced speech signal.
[0088] In this embodiment, the process of gain application is to smooth the gain value Spectral data applied to each frequency band X b (t,k). The spectrum value of each frequency band will be adjusted according to the gain value, so that the target frequency band of the signal is enhanced, while the irrelevant noise part is relatively suppressed. The gain application step essentially adjusts the amplitude of the frequency band according to the gain value, thereby enhancing the speech signal and suppressing the noise.
[0089] Gain application is performed by multiplying the smoothed gain value with the spectrum data. Assuming the gain value is the smoothed gain value of each frequency band obtained through the gain smoothing step, and the spectrum data X b(t, k) is the original spectrum value of each frequency band at time t and frequency point k, and the gain application formula is: in: is the enhanced spectrum data, which represents the spectrum value of the bth frequency band at time t and output frequency point k after the gain is applied. This data will be used to generate the enhanced speech signal in the subsequent steps. is the smoothing gain value of the bth frequency band at time t, is the smoothing gain coefficient output by the gain smoothing step, and is used to adjust the amplitude of the frequency band. b (t, k) is the original spectrum value of the b-th frequency band at the output frequency point k at time t, indicating the spectrum data without enhancement processing.
[0090] In general, the gain value during gain application is It is possible to make adjustments that are too large, causing certain frequency bands to have excessive amplitudes, which can cause distortion or unacceptable effects. To avoid this, gain values are often limited to ensure that the spectral amplitudes are not overly amplified or suppressed. To this end, the gain application process can include a limit or constraint that restricts the range of gain values.
[0091] Specifically, the formula for gain application can be restricted as follows when gain is applied: Where: clip(·) means limiting the spectrum value after the gain is applied so that it is within g min and g max This function ensures that the spectrum value does not exceed the set minimum and maximum ranges; g min Indicates the minimum limit of the gain value to prevent the gain from being too small, resulting in the signal amplitude being too low; g max Indicates the maximum limit of the gain value to prevent excessive gain from causing excessive signal amplitude and distortion; The enhanced spectrum data represents the spectrum value of the bth frequency band at time t and output frequency point k after the gain is applied. is the smoothing gain value of the bth frequency band at time t, is the smoothing gain coefficient output by the gain smoothing step, and is used to adjust the amplitude of the frequency band. b (t, k) is the original spectrum value of the b-th frequency band at the output frequency point k at time t, indicating the spectrum data that has not been enhanced.
[0092] For example, in practical applications, g min It may be set to 0.1, which means that if the gain is too small, the spectrum amplitude cannot be lower than 10%; g maxIt may be set to 3.0, which means that the maximum gain is 3 times, that is, the maximum gain does not exceed three times the original signal. By setting an appropriate gain limit, you can avoid over-enhancement or over-suppression of the spectrum after the gain is applied, thereby maintaining the naturalness and clarity of the signal.
[0093] As an option, during the gain application process, the gain value can be dynamically adjusted based on different noise environments, speech signal characteristics, etc. For example, in a quieter environment, the gain value can remain relatively stable without too much change; while in a noisy environment, the gain value can be appropriately increased to enhance the intelligibility of the speech signal.
[0094] In this case, the gain value can be adjusted based on dynamic factors such as background noise intensity, voice pitch, and voice energy. The formula for dynamic gain adjustment can be expressed as: in: is the smoothed gain estimate of the b-th frequency band at time t, indicating the gain after smoothing; Noise level is the noise level at the current moment, which can be estimated by the signal-to-noise ratio (SNR); Speech energy is the energy of the current speech signal; pitch is the pitch feature of the current speech.
[0095] This method of dynamically adjusting the gain value can optimize the effect of gain application in real time according to the actual voice environment, ensuring that good voice quality can be maintained in various environments.
[0096] Through the gain application step, the original spectrum data is effectively adjusted, and the performance of the enhanced spectrum data in the time domain and frequency domain is more in line with the expected speech enhancement goal. The gain application ensures that the amplitude of the frequency band is reasonably amplified or suppressed, thereby improving the clarity of the speech signal, enhancing intelligibility, and reducing the impact of noise components.
[0097] Gain limiting and dynamic adjustment mechanisms further ensure the stability and flexibility of gain application, prevent over-amplification or suppression, and optimize the gain application effect. In a noisy environment, precise control of gain application enables the speech signal to be effectively denoised and retain key signal features during the enhancement process, improving speech quality.
[0098] S7, inverse transformation recovery: perform inverse fast Fourier transform on the enhanced spectrum data, convert the frequency domain signal into a time domain signal, and obtain a noise-reduced speech signal.
[0099] In the aforementioned step S6, the spectrum data of each frequency band has been adjusted by the gain application process to generate enhanced spectrum data. Next, the inverse transform recovery is responsible for converting these enhanced spectrum data back to the time domain signal and restoring it to the denoised speech signal. The inverse transform recovery reversely converts the frequency domain signal into the time domain signal through the inverse fast Fourier transform (IFFT), providing a clear signal for the final speech output.
[0100] In this embodiment, the core of the inverse transform recovery step is to perform an inverse fast Fourier transform (IFFT) to convert the enhanced spectrum data of each frequency band back to the time domain. This step is opposite to the FFT operation in step S2 (spectrum division) and aims to restore the time domain waveform of the original signal based on the gain-adjusted spectrum data.
[0101] Spectral data after gain application is the enhanced spectrum of each frequency band in the time domain t and the output frequency point k. In the recovery process, the IFFT operation is used to convert it back to the time domain signal to obtain the enhanced time domain signal That is, the recovered signal for each frequency band.
[0102] Specifically, the recovered signal of each frequency band can be achieved by IFFT: in: is the restored signal of the bth frequency band in the time domain, which represents the denoised signal obtained by inverse transformation; is the enhanced spectrum data of the bth frequency band at time t and output frequency point k after gain application; IFFT is the inverse fast Fourier transform, which means that the frequency domain signal Convert to time domain signal
[0103] As an option, the inverse FFT operation can also be combined with some optimization strategies, such as increasing the FFT length to improve the spectrum resolution, or using a window function to reduce spectrum leakage, thereby improving the quality of the recovered signal.
[0104] When performing frame processing, each frame of the signal will overlap with the adjacent frames. This overlap helps to improve the accuracy of spectrum estimation, but windowing processing is required during inverse transformation to ensure smooth transition and seamless connection of the time domain signal.
[0105] Specifically, for each frame signal The application window function w(n) is overlapped and windowed, and the final time domain signal can be expressed as: in: is the final restored time domain signal, the noise-reduced speech signal after overlapping and windowing; w(n) is the window function value of the nth sample; is the restored signal of each frame, and c represents the time position of the current frame.
[0106] The purpose of overlapping windowing is to ensure that the signal obtained by frame segmentation is continuous after restoration without any mutations or unnatural seams, thereby ensuring the smoothness of the restored signal.
[0107] In one possible implementation, the signal recovery process is not limited to inverse transforming and overlapping and windowing each frequency band individually. In order to improve processing efficiency, time domain synthesis between frequency bands can also be performed after the inverse transform operation of all frequency bands is completed. The synthesis operation combines the time domain signals of each frequency band into an overall time domain signal to form the final noise-reduced speech signal. Specifically, this synthesis step is usually completed by a simple addition operation: Where: x final (t) is the final noise-reduced speech signal, which represents the time domain signal after all frequency bands are synthesized. is the restored signal after overlapping and windowing for each frequency band b, and B is the total number of frequency bands.
[0108] Through the inverse transform restoration step, the gain-adjusted spectral data is successfully converted back to the time domain signal and the denoised speech signal is generated. The successful implementation of inverse transform restoration ensures that the spectrally enhanced signal can be restored to a time domain waveform close to the original speech while maintaining the enhanced speech quality.
[0109] The overlapping windowing operation effectively avoids the boundary effect caused by framing and inverse transformation, ensures the smooth transition of the time domain signal, and makes the restored signal continuous and natural. The final restored signal minimizes the noise component in the speech signal and improves the clarity and intelligibility of the speech signal.
[0110] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A real-time full-band speech noise reduction and enhancement method based on a hybrid DSP deep learning method, characterized in that: The following steps are involved: Signal preprocessing: The input time-domain speech signal is framed to obtain a series of speech frames, and a window function is applied to each frame of speech signal; Spectrum division: performing fast Fourier transform on each frame signal to obtain spectrum data, and dividing the spectrum data into multiple frequency bands to form a multi-band spectrum; Feature extraction: extract the features of each frequency band and process them through a convolutional neural network to generate high-level features for gain estimation; Gain estimation: Processing the high-level features through a deep learning model to output gain estimates for each frequency band; Gain smoothing: performing gain smoothing on the gain estimation value to ensure smooth transition of the spectrum gain in time and avoid distortion of the speech signal caused by sudden changes; Gain application: Apply the gain to the corresponding frequency band spectrum to obtain enhanced spectrum data; Inverse transform restoration: Perform inverse fast Fourier transform on the enhanced spectrum data to convert the frequency domain signal into a time domain signal to obtain a de-noised speech signal.
2. The real-time full-band speech noise reduction and enhancement method based on the hybrid DSP deep learning method according to claim 1 is characterized in that: The signal preprocessing step comprises the following steps: Performing frame segmentation on the input time domain speech signal, dividing the speech signal into fixed-length frames with a frame length of 15 ms to 25 ms and a frame overlap rate of 30% to 50%; Applying a window function to each frame of speech signal, wherein the window function is a Hamming window or a Hanning window, for reducing spectrum leakage; Perform fast Fourier transform on each frame of windowed signal to convert the time domain signal into frequency domain signal to obtain spectrum data X b (t).
3. The real-time full-band speech noise reduction and enhancement method based on the hybrid DSP deep learning method according to claim 1 is characterized in that: The spectrum division step comprises the following steps: The obtained spectrum data X b (t) Applying the Bark scale, the spectrum is divided into 22 frequency bands to reduce computational complexity and preserve speech characteristics; The spectral energy of each frequency band is extracted and used as the input feature of the subsequent deep learning model.
4. The real-time full-band speech noise reduction and enhancement method based on the hybrid DSP deep learning method according to claim 1 is characterized in that: The feature extraction step comprises the following steps: Extract spectrum energy, pitch, and spectrum non-stationarity features from the spectrum data of each frequency band. These features can reflect the difference between speech and noise. The extracted features are input into the convolutional neural network for adaptive feature learning to generate high-level feature representations.
5. The real-time full-band speech noise reduction and enhancement method based on the hybrid DSP deep learning method according to claim 1 is characterized in that: The gain estimation step comprises the following steps: The frequency band features are input into multiple convolutional layers through a convolutional neural network for local feature extraction; Perform pooling operations on the extracted features to reduce feature dimensions and retain key information; The pooled features are processed through a fully connected layer to output the gain estimate g for each frequency band b (t).
6. The real-time full-band speech noise reduction and enhancement method based on the hybrid DSP deep learning method according to claim 1 is characterized in that: The gain smoothing step comprises the following steps: The gain estimate g for each frequency band is b (t) Perform Kalman filtering to smooth the gain and avoid sudden changes; Calculate the Kalman gain K(t); Update the gain estimate g b (t), ensuring a smooth transition of the gain in time.
7. The real-time full-band speech noise reduction and enhancement method based on the hybrid DSP deep learning method according to claim 6 is characterized in that: The Kalman gain formula is as follows: Among them, P(t-1) is the gain estimation error covariance at the previous moment, and R(t) is the noise covariance at the current moment.
8. The real-time full-band speech noise reduction and enhancement method based on the hybrid DSP deep learning method according to claim 1 is characterized in that: The gain estimation step includes the following model compression and optimization techniques: Use L1 regularization to prune the neural network and remove unimportant neurons and connections; Quantize the network weights and convert floating point numbers to integers to reduce computation and memory usage.
Citation Information
Patent Citations
Voice noise reduction method for conference terminal based on neural network model
CN109065067A
Noise suppression method and device and mobile terminal
CN110335620A
Noise suppression method and device and mobile terminal
CN113113039A
Training method of frequency band gain model and voice noise reduction method for vehicle-mounted scene
CN113782011A
Speech enhancement model training method and apparatus, device, medium and program product
WO2025035943A1