A single-channel speech enhancement method balancing noise reduction and speech quality
By using PEFAC to estimate the fundamental frequency and an adaptive prior method to optimize the noise power spectral density in single-channel speech enhancement, the balance between noise reduction and speech intelligibility in single-channel speech enhancement technology is solved, achieving efficient noise suppression and speech recovery with low complexity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV
- Filing Date
- 2023-06-15
- Publication Date
- 2026-04-28
AI Technical Summary
Single-channel speech enhancement technology struggles to achieve both high noise reduction and high speech clarity and intelligibility while maintaining low computational complexity. Existing algorithms, such as OMLSA, damage speech components and leave noise residue in low signal-to-noise ratio and non-stationary noise scenarios.
The fundamental frequency is estimated using the PEFAC method. Combined with unbiased minimum mean square error and adaptive prior methods, speech enhancement is achieved by optimizing the estimation of speech presence probability and noise power spectral density through cepstral smoothing and the logarithmic spectral amplitude gain of the generalized gamma prior χ.
With low computational complexity, it achieves good noise suppression and preservation of weak speech components, restores damaged harmonics, and achieves a good balance between noise reduction and speech quality.
Smart Images

Figure CN116913308B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech enhancement algorithms, specifically relating to a single-channel speech enhancement method that balances noise reduction and speech quality. Background Technology
[0002] Speech enhancement is one of the core issues in audio signal processing. Its main purpose is to recover speech from signals contaminated by environmental noise, interfering human voices, reverberation, and echoes, thereby improving speech quality and intelligibility. It has been widely used in mobile communications, teleconferencing, Bluetooth headsets, hearing aids, and speech recognition front-ends. Single-channel speech enhancement is an important technique, primarily used in scenarios where only a single microphone collects the signal, but it can also serve as a post-processing step in multi-channel techniques.
[0003] Single-channel speech enhancement technology based on traditional signal processing algorithms has advantages such as low computational complexity, high stability, and strong parameter interpretability. It has developed into various types of algorithms, including spectral subtraction, Wiener filtering, statistical model-based algorithms, wavelet transform-based algorithms, and subspace-based algorithms. Among these, the minimum mean square error-based algorithm in the discrete Fourier transform domain has gained wider application due to its low computational complexity, excellent processing performance, and ease of integration with other systems.
[0004] Single-channel speech enhancement technology lacks information from the spatial dimension, thus it can only process and enhance based on the diversity of the spectrum and the characteristics of the source. However, in practical applications, due to differences in environmental complexity, noise type and intensity, and speech signal quality, single-channel speech enhancement always struggles to simultaneously achieve high noise reduction and high clarity and intelligibility. This is one of the current challenges in single-channel speech enhancement technology research.
[0005] The Optimally-Modified Log-Spectral Amplitude (OMLSA) algorithm, which considers the uncertainty of speech presence (SPU), is an excellent single-channel speech enhancement algorithm (Cohen I, Berdugo B. Speech enhancement for non-stationary noise environments[J]. Signal processing, 2001, 81(11):2403-2418.). Subsequent studies have further improved upon it, and it remains one of the commonly used algorithms in this field. However, its noise tracking ability is poor, and it damages speech components in low signal-to-noise ratio and non-stationary noise scenarios, with noticeable musical noise residue.
[0006] Temporal Cepstrum Smoothing (TCS) is a priori signal-to-noise ratio estimation method that can effectively preserve speech harmonic components while reducing estimation variance and suppressing musical noise. Combined with an unbiased minimum mean square error noise power spectral density estimator (MMS) with strong tracking capabilities and low computational complexity, it can achieve good speech enhancement performance. However, while preserving speech harmonic components, TCS also retains many noise components, and it can also degrade speech quality in complex environments. A major reason for this problem is that the cepstrum fundamental frequency estimation method it uses has low accuracy in both fundamental frequency estimation and voiced frame determination. Many studies have developed more accurate fundamental frequency estimation methods, such as the PEFAC (Pitch Estimation Filter with Amplitude Compression) method (Gonzalez S, Brookes MA pitch estimation filter robust to high levels of noise (PEFAC) [C] / / 2011 19th European Signal Processing Conference.IEEE,2011:451-455.). Summary of the Invention
[0007] Due to variations in environmental complexity, noise type and intensity, and speech signal quality, single-channel speech enhancement often struggles to simultaneously achieve high noise reduction and high clarity and intelligibility. Therefore, achieving a good balance between noise reduction and speech quality while maintaining low computational complexity is crucial in practical applications. In this regard, this invention proposes a single-channel speech enhancement method that balances noise reduction and speech quality.
[0008] The technical solution adopted in this invention is as follows:
[0009] A single-channel speech enhancement method that balances noise reduction and speech quality includes the following steps:
[0010] Step 1: Transform the noisy signal into the time-frequency domain and estimate the fundamental frequency using the PEFAC method;
[0011] Step 2: Calculate the posterior signal-to-noise ratio. Based on the fundamental frequency estimate obtained in Step 1, smooth the posterior signal-to-noise ratio in the cepstral domain, and then use the fixed prior method to estimate the posterior probability of speech existence.
[0012] Step 3: Using the unbiased minimum mean square error method, estimate the noise power spectral density based on the posterior speech existence probability obtained in Step 2.
[0013] Step 4: Based on the noise power spectral density estimate obtained in Step 3, calculate the posterior signal-to-noise ratio estimate and the maximum likelihood estimate of the speech power spectral density.
[0014] Step 5: Based on the fundamental frequency estimate obtained in Step 1, smooth the maximum likelihood estimate of the speech power spectral density obtained in Step 4 in the cepstral domain, and simultaneously perform cepstral fundamental frequency enhancement to obtain the prior signal-to-noise ratio estimate.
[0015] Step 6: Using the adaptive prior method, based on the posterior signal-to-noise ratio estimate obtained in Step 4 and the prior signal-to-noise ratio estimate obtained in Step 5, estimate the posterior speech existence probability again.
[0016] Step 7: Based on the posterior signal-to-noise ratio estimate obtained in Step 4 and the prior signal-to-noise ratio estimate obtained in Step 5, calculate the logarithmic spectral amplitude gain based on the generalized gamma prior χ, and then combine it with the posterior speech existence probability estimate obtained in Step 6 to derive the gain estimate based on the uncertainty of speech existence.
[0017] Step 8: Use the gain estimate obtained in Step 7 to enhance the speech, and transform the enhanced spectrum back to the time domain to obtain the enhanced signal.
[0018] Compared with existing technologies, the method of the present invention has low computational complexity, good noise suppression capability, and can preserve weak speech components, restore damaged harmonics, and achieve a good balance between noise reduction and speech quality. Attached Figure Description
[0019] Figure 1 This is a flowchart of the method of the present invention.
[0020] Figure 2 The figure shows the average objective score results of the test set of the present invention and the comparison method under different signal-to-noise ratios. Among them, (a)-(e) are the average score results of broadband PESQ, STOI, DNSMOS-OVRL, DNSMOS-SIG and DNSMOS-BAK, respectively.
[0021] Figure 3 Examples of spectrograms for speech enhancement by the present invention and comparison methods are shown. (a)-(d) are spectrograms of noisy signal, clean speech, OMLSA algorithm enhancement result, and the enhancement result of the present invention method, respectively. Detailed Implementation
[0022] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and embodiments.
[0023] This invention includes the following steps:
[0024] Step 1: Transform the noisy signal into the time-frequency domain and estimate the fundamental frequency using the PEFAC method;
[0025] Step 2: Calculate the posterior signal-to-noise ratio. Based on the fundamental frequency estimate described in Step 1, smooth the posterior signal-to-noise ratio in the cepstral domain, and then use the fixed prior (FP) method to estimate the posterior probability of speech presence.
[0026] Step 3: Using the unbiased minimum mean square error method, estimate the noise power spectral density based on the posterior speech existence probability obtained in Step 2.
[0027] Step 4: Based on the noise power spectral density estimate obtained in Step 3, calculate the posterior signal-to-noise ratio estimate and the maximum likelihood estimate of the speech power spectral density.
[0028] Step 5: Based on the fundamental frequency estimate obtained in Step 1, smooth the maximum likelihood estimate of the speech power spectral density obtained in Step 4 in the cepstral domain, and simultaneously perform cepstral fundamental frequency enhancement to obtain the prior signal-to-noise ratio estimate.
[0029] Step 6: Using the Adaptive Prior (AP) method, based on the posterior signal-to-noise ratio estimate obtained in Step 4 and the prior signal-to-noise ratio estimate obtained in Step 5, estimate the posterior speech existence probability again.
[0030] Step 7: Based on the posterior signal-to-noise ratio estimate obtained in Step 4 and the prior signal-to-noise ratio estimate obtained in Step 5, calculate the logarithmic spectral amplitude gain based on the generalized gamma prior χ, and then combine it with the posterior speech existence probability estimate obtained in Step 6 to derive the gain estimate based on the uncertainty of speech existence.
[0031] Step 8: Use the gain estimate obtained in Step 7 to enhance the speech, and transform the enhanced spectrum back to the time domain to obtain the enhanced signal.
[0032] Furthermore, the cepstral smoothing method in step 2 is as follows:
[0033] set up This represents the estimated posterior signal-to-noise ratio, where k and l represent the frequency band index and frame index, respectively. Using the noise power spectral density estimate from the previous frame Calculation of the power spectral density of the noisy signal Y(k,l) in the current frame:
[0034]
[0035] Will Transformed to the cepstral domain, denoted as γ ceps (q,l), that is, taking the logarithm and performing the inverse discrete Fourier transform:
[0036] P ceps (q,l)=IDFT{log(P(k,l)| k=0,1,...,N-1 )},q=0,1,...,N-1
[0037] Where IDFT{·} denotes the inverse discrete Fourier transform, q denotes the cepstral frequency index, and N denotes the length of the discrete Fourier transform used in step 1. Due to symmetry, the following operations only apply to... conduct.
[0038] Transform the fundamental frequency estimate f0(l) obtained in step 1 into the cepstral frequency q0(l):
[0039]
[0040] Among them, f s Indicates the sampling rate; This indicates rounding down. Furthermore, the fundamental frequency is extended to a range of size 2Δq0+1, centered at q0(l).
[0041]
[0042] Where v(l) is the voiced frame discrimination result, given by the fundamental frequency estimator in step 1; v(l) = 1 indicates that the current frame is a voiced frame, and v(l) = 0 indicates that the current frame is an unvoiced frame. Then, the smoothing factor α is determined. ceps (q,l):
[0043]
[0044] Where, β ceps Used to smooth α ceps (q,l); α const (q) is a preset cepstral frequency-dependent smoothing factor, characterized by smaller values at lower cepstral frequencies and larger values at others; α0 is a smaller fundamental frequency smoothing factor. This results in weaker smoothing of the cepstral coefficients corresponding to the fundamental harmonic components and spectral envelope of speech, and stronger smoothing of the noise-dominated cepstral coefficients, thus preserving speech components as much as possible while smoothing.
[0045] Smoothing γ ceps (q,l), we get
[0046]
[0047] The inverse transform returns the signal to the frequency domain, yielding a biased smoothed result γ. b (k,l):
[0048]
[0049] Where DFT{·} denotes the Discrete Fourier Transform.
[0050] By performing bias compensation, an unbiased smoothing result is obtained.
[0051]
[0052] Where B(l) is the deviation compensation factor, based on χ 2 Calculated. Assuming before smoothing The shape parameter is μ γ χ 2 The distribution, after smoothing, also conforms to χ². 2 Distribution, with shape parameters as For the preset shape parameter μ γ It can calculate the cepstral variance var offline. q {γ ceps}:
[0053]
[0054] Where ζ(·,·) denotes the Riemann zeta function; κ m Denotes the logarithmic covariance; M represents the non-zero κ. m The corresponding maximum subscript; κ m The calculation formula is:
[0055]
[0056]
[0057]
[0058] Where Γ(·) represents the gamma function; ψ(·) represents the psi function; and the correlation coefficient ρ m The value of depends on the window function used in the short-time Fourier transform in step 1. Then, the smoothed cepstral variance is obtained based on the smoothing factor. Approximately
[0059]
[0060] Where, ν q Represents the weighting coefficients, ν0,ν N / 2 =1 / 2, the remainder is 2. Solving the following equation yields the smoothed shape parameters.
[0061]
[0062] This leads to the deviation compensation factor:
[0063]
[0064] Where ψ(·) represents the psi function.
[0065] Furthermore, the specific steps for estimating the probability of speech presence using the fixed prior method in step 2 are as follows:
[0066] Using the posterior signal-to-noise ratio after cepstral smoothing in step 2 and shape parameters Calculate the generalized likelihood ratio Λ(k,l):
[0067]
[0068] Let and represent the fixed prior probability of speech existence and the prior signal-to-noise ratio, respectively. This represents the state of speech presence. Then, the posterior probability of speech presence is calculated:
[0069]
[0070] Recursive smoothing It also checks the average value; if the average value is too large, then a constraint is imposed. The upper limit is set so that it does not remain close to 1 for a long time, in order to avoid the noise power spectral density estimation in step 3 getting stuck.
[0071] Furthermore, in step 3, the noise power spectral density is estimated using the unbiased minimum mean square error algorithm. The calculation formula is:
[0072]
[0073] Where, α N This represents the noise power spectral density smoothing factor.
[0074] Furthermore, the cepstral smoothing method in step 5 is as follows:
[0075] set up This represents the estimated value of the speech power spectral density, first using the estimated value of the noise power spectral density of the current frame. Update the estimated posterior signal-to-noise ratio:
[0076]
[0077] Next, calculate the maximum likelihood estimate of the speech power spectral density:
[0078]
[0079] Where, ξ min This represents the lower limit of the prior signal-to-noise ratio;
[0080] Will Transformed to the cepstral domain, denoted as That is, take the logarithm and perform the inverse discrete Fourier transform:
[0081]
[0082] Using the method described in step 2, the smoothing factor α is determined based on the fundamental frequency estimate f0(l). ceps (k,l); smoothing get
[0083]
[0084] Perform fundamental frequency enhancement:
[0085]
[0086] Where ρ is the fundamental frequency amplification factor, which is a constant greater than 1;
[0087] The inverse transform is performed back to the frequency domain to obtain a biased smoothed result.
[0088]
[0089] Where DFT{·} denotes the Discrete Fourier Transform.
[0090] By performing bias compensation, an unbiased smoothing result is obtained.
[0091]
[0092] Where B is a fixed deviation compensation factor:
[0093] B = exp(0.5 * 0.5772)
[0094] Based on the unbiased smoothing results, the formula for calculating the prior signal-to-noise ratio is:
[0095]
[0096] Furthermore, in step 6, the specific steps for estimating the posterior speech existence probability using the adaptive prior method are as follows:
[0097] Calculate the time-domain recursive average of the prior signal-to-noise ratio obtained in step 5:
[0098]
[0099] Where, α S This is a smoothing factor. A size of 2W is used. L +1 smaller smooth window wL Time-domain recursive average Perform local frequency domain smoothing
[0100]
[0101] Then calculate the probability of the existence of local prior speech.
[0102]
[0103] Among them, P min It is the lower bound of probability; and Let represent the upper and lower bounds of the local prior signal-to-noise ratio, respectively. The global prior probability of speech existence is calculated in the same way. Simply replace the superscript or subscript L in the above formula with G.
[0104] Calculate the frame average prior signal-to-noise ratio Frame-averaged prior signal-to-noise ratio with upper and lower constraints
[0105]
[0106]
[0107] Among them, K min and K max Constrain the frequency band range for calculating the average prior signal-to-noise ratio of the frame; and They represent The upper and lower bounds are then used to calculate the probability of the existence of prior speech in the frame.
[0108]
[0109] in, and Let represent the upper and lower limits of the frame prior signal-to-noise ratio, respectively; the variable p is given by the following formula:
[0110]
[0111] Calculate the adaptive prior speech existence probability
[0112]
[0113] in, This represents the lower bound of the probability of prior speech existence. Then, the probability of posterior speech existence based on the adaptive method is calculated.
[0114]
[0115]
[0116] Furthermore, in step 7, log-spectral amplitude gain based on speech presence uncertainty (SPU) is used. The calculation formula is:
[0117]
[0118] Among them, G min Indicates the lower limit of gain; This represents the log-spectral amplitude gain based on a prior of χ with shape parameter μ, where the prior χ is a class of generalized gamma distributions, calculated as follows:
[0119]
[0120] in,
[0121]
[0122]
[0123] Where 1F1(a,b;x) represents the merging hypergeometry function.
[0124] The advantages of the method proposed in this invention are more readily apparent in experimental evaluations:
[0125] 1. Evaluation settings
[0126] The clean speech used for evaluation came from the TIMIT dataset, consisting of 128 randomly selected clean speech samples, with male and female voices each comprising 50%. The noise used for evaluation included background noise and pink noise from the NOISEX-92 dataset, and railway, subway, car, and road noise, as well as modulated white Gaussian noise from the ITU-T P.501 dataset. Signal-to-noise ratios were set to -10, -5, 0, 5, 10, and 15 dB.
[0127] The comparison method used was the OMLSA algorithm. Evaluation metrics included the Deep Noise Suppression Mean Opinion Score (DNSMOS), the broadband PESQ score mapped to the Mean Opinion Score – Listening Quality Objective (MOS-LQO), and the STOI. DNSMOS simulates subjective MOS scoring through a deep neural network and includes three components: OVRL, SIG, and BAK, representing overall speech quality, speech tone quality, and noise reduction effectiveness, respectively.
[0128] 2. Parameter settings
[0129] Sampling rate f s The frame rate is 16kHz, the frame length and the number of discrete Fourier transform points N are 512 points, with 50% overlap, and the short-time Fourier transform window function is the square root Hanning window.
[0130] Fixed prior speech existence probability Set to 0.5; fix the prior signal-to-noise ratio Set to 15dB; smooth the shape parameter μ of the forward and backward arithmetic signal-to-noise ratio. γ Set to 0.7; the recursive smoothing factor for the posterior speech probability in the avoidance step is set to 0.9; the noise power spectral density smoothing factor α N Set to 0.8; prior signal-to-noise ratio lower limit ξ min Set to -25dB; Δq0 is set to 2; β ceps In steps 2 and 5, the values are set to 0.9 and 0.96 respectively; for the square root Hanning window, ρ1 = 0.5, M = 1; the fundamental frequency amplification factor ρ is set to 2.5; the upper and lower limits for estimating the posterior speech existence probability using the adaptive method are set as follows: P min =0.005,
[0131] K min =3,K max =257; the time-domain recursive averaging factor α of the prior signal-to-noise ratio. S Set to 0.7; both the local and global smoothing windows are set to normalized Hanning windows, with the parameter W determining their size. L and W G Set to 1 and 15 respectively; set the prior shape parameter μ of χ to 0.6; and set the lower limit of gain G. min Set to -18dB; the fundamental frequency smoothing factor α0 in steps 2 and 5 is set to 0.2 and 0.1 respectively, and the fixed smoothing factor is set to:
[0132]
[0133] In step 2 and step 5, α1, α2, and α3 are set to 0.2, 0.4, 0.997 and 0.1, 0.6, 0.95, respectively.
[0134] 3. Evaluation Results
[0135] Figure 2The average objective scores of the proposed method and the OMLSA algorithm under different signal-to-noise ratios are shown. (a)-(e) are the average scores of wideband PESQ, STOI, DNSMOS-OVRL, DNSMOS-SIG, and DNSMOS-BAK, respectively. It can be seen that the proposed method has significant advantages.
[0136] Figure 3 The spectrogram of a sample is shown. The signal-to-noise ratio of the noisy signal is 10 dB, and the background noise type is background noise. (a)-(d) are the spectrograms of the noisy signal, clean speech, OMLSA algorithm enhancement result, and the enhancement result of the method of this invention, respectively. The spectrograms show that the method of this invention has better performance in both noise reduction and speech recovery.
Claims
1. A single-channel speech enhancement method that balances noise reduction and speech quality, characterized in that, The method includes the following steps: Step 1: Transform the noisy signal into the time-frequency domain and estimate the fundamental frequency using the PEFAC method; Step 2: Calculate the posterior signal-to-noise ratio. Based on the fundamental frequency estimate obtained in Step 1, smooth the posterior signal-to-noise ratio in the cepstral domain, and then use the fixed prior method to estimate the posterior probability of speech existence. Step 3: Using the unbiased minimum mean square error method, estimate the noise power spectral density based on the posterior speech existence probability obtained in Step 2. Step 4: Based on the noise power spectral density estimate obtained in Step 3, calculate the posterior signal-to-noise ratio estimate and the maximum likelihood estimate of the speech power spectral density. Step 5: Based on the fundamental frequency estimate obtained in Step 1, smooth the maximum likelihood estimate of the speech power spectral density obtained in Step 4 in the cepstral domain, and simultaneously perform cepstral fundamental frequency enhancement to obtain the prior signal-to-noise ratio estimate. Step 6: Using the adaptive prior method, based on the posterior signal-to-noise ratio estimate obtained in Step 4 and the prior signal-to-noise ratio estimate obtained in Step 5, estimate the posterior speech existence probability again. Step 7: Based on the posterior signal-to-noise ratio estimate obtained in Step 4 and the prior signal-to-noise ratio estimate obtained in Step 5, calculate the logarithmic spectral amplitude gain based on the generalized gamma prior χ, and then combine it with the posterior speech existence probability estimate obtained in Step 6 to derive the gain estimate based on the uncertainty of speech existence. Step 8: Use the gain estimate obtained in Step 7 to enhance the speech, and transform the enhanced spectrum back to the time domain to obtain the enhanced signal.
2. The single-channel speech enhancement method balancing noise reduction and speech quality according to claim 1, characterized in that, The cepstral smoothing method in step 2 is as follows: set up This represents the estimated posterior signal-to-noise ratio, where k and l represent the frequency band index and frame index, respectively; Using the noise power spectral density estimate from the previous frame Calculation of the power spectral density of the noisy signal Y(k,l) in the current frame: Will Transformed to the cepstral domain, denoted as γ ceps (q,l), that is, taking the logarithm and performing the inverse discrete Fourier transform: Where IDFT{·} denotes the inverse discrete Fourier transform, q denotes the cepstral frequency index, and N denotes the length of the discrete Fourier transform used in step 1; due to symmetry, the following operations only apply to... conduct; Transform the fundamental frequency estimate f0(l) obtained in step 1 into the cepstral frequency q0(l): Among them, f s Indicates the sampling rate; This indicates rounding down; furthermore, the fundamental frequency is extended to a range of size 2Δq0+1 with q0(l) as the center: Where v(l) is the voiced frame discrimination result, v(l) = 1 indicates that the current frame is a voiced frame, and v(l) = 0 indicates that the current frame is an unvoiced frame; then the smoothing factor α is determined. ceps (k,l): Where, β ceps Used to smooth α ceps (q,l); α const (q) is a preset cepstral frequency-related smoothing factor, which is smaller at low cepstral frequencies and larger at others; α0 is a smaller fundamental frequency smoothing factor. Smoothing γ ceps (q,l), we get The inverse transform returns the signal to the frequency domain, yielding a biased smoothed result γ. b (k,l): Where DFT{·} denotes the Discrete Fourier Transform; By performing bias compensation, an unbiased smoothing result is obtained. Where B(l) is the deviation compensation factor, based on χ 2 The distribution was calculated; it is assumed that before smoothing... The shape parameter is μ γ χ 2 The distribution, after smoothing, also conforms to χ². 2 Distribution, with shape parameters as For the preset shape parameter μ γ It can calculate the cepstral variance var offline. q {γ ceps }: Where ζ(·,·) denotes the Riemann zeta function; κ m Denotes the logarithmic covariance; M represents the non-zero κ. m The corresponding maximum subscript; κ m The calculation formula is: Where ·(·) represents the gamma function; ψ(·) represents the psi function; and the correlation coefficient ρ m The value depends on the window function used in the short-time Fourier transform in step 1; then, the smoothed cepstral variance is obtained based on the smoothing factor. Approximately Where, ν q Represents the weighting coefficients, ν0,ν N / 2 =1 / 2, the remainder is 2; solve the following equation to obtain the smoothed shape parameters. This leads to the deviation compensation factor:
3. The single-channel speech enhancement method balancing noise reduction and speech quality according to claim 1, characterized in that, The cepstral smoothing method in step 5 is as follows: set up This represents the estimated speech power spectral density, where k and l represent the frequency band index and frame index, respectively; first, the estimated noise power spectral density of the current frame is used. Update the estimated posterior signal-to-noise ratio: Next, calculate the maximum likelihood estimate of the speech power spectral density: Where, ξ min This represents the lower limit of the prior signal-to-noise ratio; Will Transformed to the cepstral domain, denoted as That is, take the logarithm and perform the inverse discrete Fourier transform: Using the method described in step 2, the smoothing factor α is determined based on the fundamental frequency estimate f0(l). ceps (k,l); smoothing get Perform fundamental frequency enhancement: Where ρ is the fundamental frequency amplification factor, which is a constant greater than 1; The inverse transform is performed back to the frequency domain to obtain a biased smoothed result. Where DFT{·} denotes the Discrete Fourier Transform; By performing bias compensation, an unbiased smoothing result is obtained. Where B is a fixed deviation compensation factor: B = exp(0.5 * 0.5772).
4. The single-channel speech enhancement method balancing noise reduction and speech quality according to claim 1, characterized in that, In step 7, the logarithmic spectrum amplitude gain based on the uncertainty of speech is used. The calculation formula is: in, This indicates that the posterior speech obtained in step 6 has a probability estimate, where Indicates the state of the presence of speech; G min Indicates the lower limit of gain; The log-spectral amplitude gain based on the prior χ with shape parameter μ is expressed by the following formula: in, Where 1F1(a,b;x) represents the merging hypergeometry function; and These represent the posterior signal-to-noise ratio (SNR) estimate obtained in step 4 and the prior SNR estimate obtained in step 5, respectively.