Real-time single-microphone voice noise reduction algorithm based on voice enhancement residual error and continuous spectrum estimation
By using the minimum control averaging method based on speech enhancement residuals and continuous spectrum estimation, the problem of incomplete noise suppression in walkie-talkies in complex noise environments is solved, and real-time, low-distortion speech noise reduction effect is achieved in walkie-talkies.
Patent Information
- Application Number
- CN202510460329.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2026-01-16
AI Technical Summary
Existing speech noise reduction algorithms suffer from incomplete noise suppression, large noise suppression delays or the generation of musical noise in complex noise environments, poor ability to track and suppress transient noise, and high computational complexity, making it difficult to meet the real-time processing requirements of devices such as walkie-talkies.
A minimum control averaging method based on speech enhancement residuals and continuous spectrum estimation is adopted. Through frame segmentation, windowing and FFT preprocessing, combined with an improved minimum control recursive averaging method and continuous minimum value tracking algorithm, noise power spectrum estimation and gain function calculation are used to perform real-time single-microphone speech noise reduction.
It effectively suppresses various noises in noisy environments, ensures voice clarity, reduces distortion, and meets the real-time processing needs of walkie-talkies.
Smart Images

Figure CN121354580A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a real-time single-microphone speech noise reduction algorithm, belonging to the field of speech signal processing technology. Background Technology
[0002] In the information age, the importance of voice communication is increasingly prominent. Whether it's communication between people or interaction between people and machines, voice is a crucial medium for information transmission. However, voice signals are often interfered with by various noises during transmission, which can originate from environmental factors, transmission media, electrical equipment, and more. These interferences can distort the voice, affecting the listener's ability to understand it. For example, when using walkie-talkies, users frequently encounter various noises, such as engine noise from vehicles, operating noise from machinery, construction noise, and noise generated by social activities. These noises can severely impact the communication quality of walkie-talkies, making it difficult for users to clearly hear the other party's voice in noisy environments. To address this problem, voice noise reduction technology has emerged. The goal of voice noise reduction technology is to separate the purest possible voice from a noisy voice signal. This technology can effectively reduce voice distortion, improve voice quality, reduce auditory fatigue, and enhance auditory perception. Therefore, voice communication devices must utilize voice noise reduction technology when used in noisy environments.
[0003] Existing speech denoising algorithms have significant shortcomings in complex noisy environments. Traditional methods are incomplete in suppressing stationary noise, easily leaving background noise or producing "musical noise." Furthermore, their ability to track and suppress transient noise (such as keyboard clicks and sudden mechanical noise) is poor, resulting in large noise suppression delays or severe speech distortion. In addition, while AI-based algorithms offer excellent performance, they rely on large amounts of training data and have high computational complexity, making them unsuitable for the real-time processing requirements of devices such as walkie-talkies. Existing noise power spectrum estimation methods have insufficient response speed under non-stationary noise conditions and do not fully consider the recursive update problem when speech and noise coexist, thus limiting the accuracy of noise estimation. Therefore, there is an urgent need for a single-channel speech denoising scheme that balances real-time performance, low distortion, and robustness to non-stationary noise to improve the communication quality of digital walkie-talkies in noisy environments. Summary of the Invention
[0004] To address the problem of walkie-talkies being buried by noise in dedicated communication scenarios due to noisy and diverse background noise, this invention proposes a real-time single-microphone speech denoising algorithm based on speech enhancement residuals and continuous spectrum estimation.
[0005] The technical solution adopted by the present invention to solve the above problems is as follows: The steps of the present invention include: Step 1: Preprocess the input signal by framing, windowing, and using FFT to obtain the spectral amplitude of the current frame. ; Step 2: Using the minimum control averaging method based on speech enhancement residuals and continuous spectrum estimation, the noise power spectrum of the input signal is estimated to obtain the estimated value. ; Step 3: Estimate based on noise power spectrum The gain function is calculated using a log-MMSE estimator based on optimal correction. Considering the processing performance of the walkie-talkie, we can approximate the exponential integral using a closed expression; Step 4: Measure the spectral amplitude of the input signal in the current frame. With gain function Multiply to get Then, IFFT and inter-frame merging are performed on it to output noise-reduced speech.
[0006] Furthermore, step 1 specifically includes: Step 101: Set the number of sampling points per frame to 256 and the number of frame shift points to 96. Extract the sampled audio signal and store it in the dynamic buffer. Step 102: Design a hybrid window function. The part that overlaps with other frames is designed as a Hanning window, and the rest is designed as a rectangular window to ensure that the amplitude remains unchanged when the signals of adjacent frames are superimposed, and the sound is continuous. Step 103: After the time-domain signal is processed by STFT, 256 frequency points are obtained. The power spectrum of the signal is calculated and saved for noise power spectrum estimation.
[0007] Furthermore, step 2 specifically includes: Step 201: Make a coarse estimate of the probability of speech presence, and obtain... : (1); Step 202: Regarding the presence of speech By default, the noise estimation result of the previous frame is used instead; for silent mode... The noise spectrum estimation is performed using a decision-guided method, utilizing the noise power spectrum estimate from the previous frame. Smooth the power spectrum of the input signal in the current frame: (2); Step 203: Calculate the speech enhancement residual factor according to formula (3). : (3); Step 204, for The state, through the previous frame and Weighted summation for the current frame Update: (4).
[0008] Furthermore, step 3 specifically includes: Step 301: Solve the prior signal-to-noise ratio using the decision-guided method according to formula (5). This calculation method can prevent the generation of musical noise. (5); Step 302: Calculate the posterior signal-to-noise ratio : (6); Step 303: Calculate from both local and global perspectives. At the same time, a new variable is introduced. right The values are corrected to prevent the spectral amplitude from gradually decreasing at the end of the speech. (7); Step 304: According to formula (1), , as well as Substituting the values, we can obtain an accurate estimate of the probability of speech presence. ; Step 305: Calculate according to formula (8) ,in The threshold value is fixed and set based on the spectral amplitude of the input signal. (8).
[0009] The beneficial effects of this invention are as follows: Based on the characteristics of dedicated communication applications of walkie-talkies, and considering the diverse and varied background noise, including both stationary and non-stationary noise, this invention proposes a speech denoising algorithm based on speech enhancement residuals and continuous spectrum estimation. This aims to improve the performance of existing noise estimation algorithms in non-stationary noise. First, the input signal needs to be preprocessed through framing, windowing, and FFT. The spectral amplitude information of each frame is then processed sequentially. An improved minimum controlled reverberation averaging (IMCRA) method is used to estimate the noise power spectrum. Since this method is slow in tracking non-stationary noise, the idea of a continuous minimum tracking algorithm is borrowed to correct the minimum power spectrum in IMCRA. Furthermore, because the original time-varying reverberation averaging method only considers the case where speech and noise segments are separate, but speech and noise often coexist, this invention introduces the speech enhancement residual as an approximation of the true noise and applies it to the reverberation averaging process. The obtained noise power spectrum estimate is then input into the OM-LSA estimator to obtain the speech presence probability estimate, and then the gain function is calculated. Finally, with the spectral amplitude Multiplication yields an estimate of the speech spectrum amplitude. The final denoised speech signal is obtained through IFFT and inter-frame merging. This invention can effectively suppress background noise in a variety of noise types, including white noise, pink noise, aircraft noise, and wind noise, while ensuring speech clarity. Attached Figure Description
[0010] Figure 1 This is a flowchart of the present invention; Figure 2 This is a flowchart of the input signal preprocessing in step 1; Figure 3 This is a flowchart of the noise power estimation process for the input signal in step 2; Figure 4 This is a flowchart of the speech amplitude spectrum estimation process for the input signal in step 3. Detailed Implementation
[0011] Specific implementation method one: as follows Figures 1 to 4 As shown, a real-time single-microphone speech denoising algorithm based on speech enhancement residuals and continuous spectrum estimation includes the following steps: Step 1: The input speech signal to the terminal first needs to be preprocessed to meet the terminal's requirements for speech sampling rate, frame length, and processing time. After reading the input speech signal, the speech is divided into frames, and each frame is windowed. The windowed signal is then subjected to a short-time Fourier transform to obtain the speech signal's spectrum. and the power spectrum of the speech signal The specific process includes: Step 101: Set the number of sampling points per frame to 256 and the number of frame shift points to 96. Extract the sampled audio signal and store it in the dynamic buffer. Step 102: Design a hybrid window function. The part that overlaps with other frames is designed as a Hanning window, and the rest is designed as a rectangular window to ensure that the amplitude remains unchanged when the signals of adjacent frames are superimposed, and the sound is continuous. Step 103: After the time-domain signal is processed by STFT, 256 frequency points are obtained. The power spectrum of the signal is calculated and saved for noise power spectrum estimation. Step 2: The probability of speech presence is obtained by using minimum control estimation, and then time-varying recursive averaging is performed using this probability to obtain the noise estimation result. To address the low accuracy of noise estimation under non-stationary noise conditions, the speech enhancement residual is introduced as an approximation of the true noise and added to the recursive averaging process. The specific process includes: Step 201: Make a coarse estimate of the probability of speech presence, and obtain... : (1); Step 202: Regarding the presence of speech By default, the noise estimation result of the previous frame is used instead; for silent mode... The noise spectrum estimation is performed using a decision-guided method, utilizing the noise power spectrum estimate from the previous frame. Smooth the power spectrum of the input signal in the current frame: (2); Step 203: Calculate the speech enhancement residual factor according to formula (3). : ; Step 204, for The state, through the previous frame and Weighted summation for the current frame Update: (4); Step 3: Estimate based on noise power spectrum Calculate the prior signal-to-noise ratio, the posterior signal-to-noise ratio, and the prior probability of the silent state. Then, the gain function of the speech amplitude spectrum estimator is obtained. The specific process includes: Step 301: Solve the prior signal-to-noise ratio using the decision-guided method according to formula (5). This calculation method can prevent the generation of musical noise: (5); Step 302: Calculate the posterior signal-to-noise ratio : (6); Step 303: Calculate from both local and global perspectives. At the same time, a new variable is introduced. right The values are corrected to prevent the frequency points corresponding to the gradually decreasing spectral amplitude at the end of the speech signal from being misjudged as not having speech: (7); Step 304: According to formula (1), , as well as Substituting the values, we can obtain an accurate estimate of the probability of speech presence. ; Step 305: Calculate according to formula (8) ,in The threshold value is fixed and set based on the spectral amplitude of the input signal. (8); Step 4: Measure the spectral amplitude of the input signal in the current frame. With gain function Multiply to get Then, perform IFFT on it. When merging frames, the frame shifts of the two frames need to be aligned to ensure that the output speech does not lose energy due to frame splitting, and finally output noise-reduced speech.
[0012] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent substitutions, and improvements made to the above embodiments without departing from the scope of the present invention, based on the technical essence of the present invention and within the spirit and principles of the present invention, shall still fall within the protection scope of the present invention.
Claims
1. A real-time single-microphone speech denoising algorithm based on speech enhancement residual and continuous spectrum estimation, characterized in that, The specific steps include: Step 1, pre-process the input signal, get the spectrum amplitude of the current frame by framing, windowing and FFT ; Step 2, estimate the noise power spectrum of the input signal using a minimum control mean method based on speech enhancement residual and continuous spectrum estimation to obtain an estimated value ; Step 3, estimating the gain function from the noise power spectrum estimate , computing the gain function using a log-MMSE estimator based on optimal corrections , approximating the exponential integral using a closed-form expression considering the processing capabilities of the intercom Step 4: Measure the spectral amplitude of the input signal in the current frame. With gain function Multiply to get Then, IFFT and inter-frame merging are performed on it to output noise-reduced speech.
2. The real-time single-microphone speech denoising algorithm based on speech enhancement residual and continuous spectrum estimation according to claim 1, characterized in that, Step 1 specifically includes: Step 101, set the number of sampling points of each frame to 256, the number of frame shift points to 96, cut the sampled voice signal, and store it in a dynamic buffer; Step 102, design a mixed window function, the overlapping part with other frames is designed as a Hanning window, and the remaining part is designed as a rectangular window, to ensure that the amplitude does not change when the adjacent frame signals are superimposed, and the sound is continuous; Step 103, after the time domain signal is processed by STFT, 256 frequency point values are obtained, the power spectrum of the signal is calculated, and they are saved to be used for noise power spectrum estimation.
3. The real-time single-microphone speech denoising algorithm based on speech enhancement residual and continuous spectrum estimation according to claim 1, characterized in that, Step 2 specifically includes: Step 201, a rough estimation is made on the voice existence probability, obtaining : (1); Step 202, for voice presence state , the noise estimation result of the previous frame is used by default to replace it; for the silence state , the noise spectrum estimation is performed using the decision-directed method, and the noise power spectrum estimation value of the previous frame is smoothed with the input signal power spectrum of the current frame: (2); Step 203, calculate the speech enhancement residual factor according to formula (3) : (3); Step 204, for each pixel in the current frame state, update the state by weighted sum of the state of the previous frame and the state of the current frame perform update: (4)。 4. The real-time single-microphone speech denoising algorithm based on speech enhancement residual and continuous spectrum estimation according to claim 1, characterized in that, Step 3 specifically includes: Step 301, solve the prior SNR according to formula (5) by using decision-directed method This calculation can prevent the generation of musical noise. (5); Step 302, compute posterior signal-to-noise ratio : (6); Step 303, calculating from the local and global two angles while introducing a new variable to the value of the correction to prevent the gradual decrease in the spectral amplitude of the tail of the voice, (7); Step 304: According to formula (1), , as well as Substituting the values, we can obtain an accurate estimate of the probability of speech presence. ; Step 305, calculating according to formula (8) wherein is a fixed threshold value, set according to the spectral amplitude values of the input signal: (8)。