A speech recognition method based on wavelet threshold denoising and ICEEMDAN comprehensive denoising
Patent Information
- Application Number
- CN202311317556.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-11
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-10-11
AI Technical Summary
[0004]本发明的目的是提供一种基于小波阈值去噪和ICEEMDAN综合去噪的语音识别方法,用于解决现有技术的小波阈值去噪法中难以保证阈值函数在阈值点处的连续性,或者存在小波系数的偏差问题的技术问题
[0032]本发明具有以下优点:本发明提出了自适应的门限阈值和一种改进的阈值函数,能够通过多个自适应参数灵活调节阈值函数,既能保证函数在阈值点处的连续性,也能解决小波系数的偏差问题。在此基础上,综合了ICEEMDAN算法进一步改进小波阈值去噪算法,能够极大提高语音识别中连续小波变换的精确性,为语音识别技术进一步改进解决了识别模糊,远场识别困难和应用场景局限等问题。
Smart Images

Figure CN117392974B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology and relates to a speech recognition method based on wavelet threshold denoising and ICEEMDAN integrated denoising. Background Technology
[0002] Speech recognition technology is used for human-computer interaction with digital devices. Essentially, speech recognition is a pattern recognition based on speech feature parameters. Through learning, the system can classify input speech according to certain patterns and then find the best matching result based on judgment criteria. The development of wavelet transform technology has provided new processing methods and techniques for speech signals, leading to rapid advancements in speech processing technology. However, the wavelet transform process is susceptible to noise, causing distortion of the recovered signal; therefore, the accuracy of speech recognition is easily affected by noise. Real-world traffic scenarios involve significant environmental noise, thus requiring the voice control module in intelligent pedestrian walkway assistance systems to have noise reduction capabilities to improve speech recognition accuracy.
[0003] Wavelet thresholding is a commonly used speech recognition denoising algorithm. However, the hard thresholding function sets wavelet coefficients smaller than the threshold to zero, which causes the hard thresholding function to jump at the threshold point. The soft thresholding function, on the other hand, sets wavelet coefficients smaller than the threshold to zero while subtracting the threshold from the portion greater than or equal to the threshold. Although this ensures the continuity of the thresholding function, some energy of the effective signal is lost in the portion of the wavelet coefficients greater than the threshold, resulting in wavelet coefficient deviation. Both thresholding functions have their drawbacks. Summary of the Invention
[0004] The purpose of this invention is to provide a speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising, which solves the technical problem that existing wavelet threshold denoising methods have difficulty in ensuring the continuity of the threshold function at the threshold point, or have the problem of wavelet coefficient deviation.
[0005] The aforementioned speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising includes decomposing the interfered signal EMD into a set of modal components (IMFs) from high frequency to low frequency. Wavelet threshold denoising is responsible for processing the high-frequency modal components, while ICEEMDAN denoising is responsible for processing the low-frequency modal components. The wavelet threshold denoising method for the high-frequency modal components obtains an adaptive threshold based on the feature information of the modal components. In the processing of general white noise, the adaptive parameter of the threshold is set to a fixed value. In the processing of some special noises, the threshold has a linear relationship with the feature value of the noise. The threshold function of wavelet threshold denoising is set to accurately process the noise.
[0006] Preferably, the speech recognition method includes the following steps:
[0007] (1) Signal initialization processing;
[0008] (2) Using EMD decomposition, the discrete digital signal is decomposed into a set of intrinsic mode components (IMFs) distributed from high frequency to low frequency, and the wavelet power spectral density of the corresponding noisy speech signal is calculated.
[0009] (3) Based on the signal characteristics obtained in step (1), a correlation threshold is obtained. Using the correlation threshold, the intrinsic mode components (IMF) are divided into two groups: high-frequency IMF components with a correlation coefficient greater than the correlation threshold and low-frequency IMF components with a correlation coefficient less than the correlation threshold. The high-frequency IMF components should be processed in step (4), and the low-frequency IMF components should be processed in step (5).
[0010] (4) Wavelet thresholding for noise reduction;
[0011] (5) ICEEMDAN noise reduction;
[0012] (6) The denoised high-frequency IMF component and the denoised low-frequency IMF component are superimposed to form a clean signal. The clean signal is then decomposed by wavelet and proceeded to step (7).
[0013] (7) Perform wavelet transform, wavelet packet transform and wavelet multiresolution analysis on the wavelet components of the pure signal to extract signal features and perform speech recognition.
[0014] Preferably, in step (2), the intrinsic mode components (IMFs) are [ω1, ω2, ..., ω]. n ];
[0015] In step (3), the relevant threshold is λ. k The formula is Among them, c i Let be the correlation coefficient of the i-th signal; the high-frequency IMF components are [ω1, ω2, ..., ω]. k The low-frequency IMF component is [ω]. k+1 ω k+2 ,…,ω n ].
[0016] Preferably, step (4) specifically includes the following sub-steps:
[0017] 4-1) First, the high-frequency IMF components [ω1, ω2, ..., ω] are analyzed. k Perform wavelet transform and decompose to extract coefficients, and obtain low-frequency approximation coefficients and high-frequency detail coefficients of each layer;
[0018] 4-2) The adaptive threshold λ is obtained, and the adaptive threshold function is:
[0019]
[0020] In the formula, σ represents the high-frequency IMF components [ω1, ω2, ..., ω]. k The variance of ] is given by N, where N is the observation length; γ is the spectral flatness, and g(γ) is the spectral flatness function. The formula includes... The adaptability of the spectral flattening function to different power spectra can be adjusted by customizing the parameter α. The parameter β∈(0,1) takes the value of the reciprocal of the signal-to-noise ratio of the original signal, which is used to dilute high-density noise, and j is the decomposition scale;
[0021] 4-3) Use spectral flatness γ to adjust the rising function segment |ω j The range of |∈[λ,γ) is used to obtain the denoised high-frequency IMF components using the improved threshold function;
[0022] 4-4) Using the wavelet reconstruction function, the j-layer wavelet components are combined and reconstructed into denoised high-frequency IMF components [ε1, ε2, ..., ε]. k Then proceed to step (6).
[0023] Preferably, in sub-step 4-2), the feature information of the modal component is the wavelet power spectral density of the noisy speech signal, and the formula for spectral flatness γ is: γ = 10lg(u g -u a ), u g The geometric mean of the wavelet power spectral density of the noisy speech signal is calculated as follows: u a The arithmetic mean of the wavelet power spectral density of the noisy speech signal is calculated as follows: The spectral flatness function is:
[0024] Preferably, in sub-step 4-2), the following is adopted: A formula for replacing the arithmetic mean of wavelet power spectral density of noisy speech signals.
[0025] Preferably, in sub-step 4-3), the segment |ω j A hard threshold function is used for |∈[γ,+∞) For segment | ω j |∈(-∞,λ) uses the function The denoised high-frequency IMF components are obtained using the improved threshold function.
[0026]
[0027] Where, ω jγ is the value of the j-th layer in the high-frequency IMF component, μ is an improved adaptive parameter that can adjust the sharpening degree of denoising according to the influence of noise on the original signal, and the specific value of γ varies according to the actual waveform.
[0028] Preferably, step (5) specifically includes the following sub-steps:
[0029] 5-1) Take the low-frequency IMF component [ω] k+1 ω k+2 ,…,ω n The mean is used as the first IMF component obtained from the ICEEMDAN decomposition;
[0030] 5-2) After adding specific noise to the residual signals of each stage obtained after EMD decomposition, EMD decomposition is performed again. The residual signals after decomposition will be freed from the noise interference and will present the original spectral characteristics of the signal. If the obtained residual signal is sufficient to represent the characteristics of the original signal, proceed to sub-step (5-3); otherwise, return to sub-step (5-1).
[0031] 5-3) Reconstruct the residual signal into a denoised low-frequency IMF component [ε] k+1 ,…ε t-1 , ε t+b+1 ,…ε n Then proceed to step (6).
[0032] This invention offers the following advantages: It proposes an adaptive threshold and an improved threshold function, which can flexibly adjust the threshold function through multiple adaptive parameters. This ensures the continuity of the function at the threshold point and also addresses the deviation problem of wavelet coefficients. Furthermore, it integrates the ICEEMDAN algorithm to further improve the wavelet thresholding denoising algorithm, significantly enhancing the accuracy of continuous wavelet transform in speech recognition. This solves problems such as fuzzy recognition, difficulties in far-field recognition, and limitations in application scenarios, thus further improving speech recognition technology. Attached Figure Description
[0033] Figure 1 This is a flowchart illustrating a speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising in this invention.
[0034] Figure 2 The figure shows the simulation results of the improved wavelet threshold denoising process for low-frequency signals according to the present invention.
[0035] Figure 3 The figure shows the simulation results of the improved wavelet threshold denoising process for high-frequency signals according to the present invention.
[0036] Figure 4 The image shows the simulation results of ICEEMDAN denoising in this invention.
[0037] Figure 5 The simulation results of the improved wavelet threshold denoising and ICEEMDAN integrated denoising method of this invention are shown in the figure.
[0038] Figure 6 This is a comparison of the denoising effects of the present invention and VisuShrink threshold wavelet threshold denoising. (VisuShrink threshold wavelet threshold denoising, reference: Liu Chong, Ma Lixiu, Pan Jinfeng, et al. Denoising of partial discharge signals by combining VMD and improved wavelet threshold [J]. Modern Electronics Technology, 2021, 44(21): 6.)
[0039] Figure 7 This is a schematic diagram of the threshold function in the improved wavelet thresholding denoising method of this invention. Detailed Implementation
[0040] The following detailed description of the embodiments, with reference to the accompanying drawings, will further illustrate the specific implementation of the present invention, in order to help those skilled in the art to have a more complete, accurate, and in-depth understanding of the inventive concept and technical solution of the present invention.
[0041] like Figure 1-7 As shown, this invention discloses a speech recognition method based on a combination of wavelet threshold denoising and ICEEMDAN denoising. The denoising method used in this invention combines wavelet threshold denoising and an improved fully adaptive noise set empirical mode decomposition (EMD) algorithm to decompose the interfered signal into a set of modal components (IMFs) from high frequency to low frequency. Wavelet threshold denoising handles the high-frequency modal components, while ICEEMDAN handles the low-frequency modal components. Wavelet threshold denoising for the high-frequency modal components requires obtaining an adaptive threshold based on the characteristic information of the modal components. For general white noise processing, the adaptive parameter of the threshold can be considered a constant value; for some special noise, it has a linear relationship with the characteristic value of the noise. The threshold function of wavelet threshold denoising, for accurate noise processing, can be approximated as multiple convergent curve functions that infinitely approach the threshold.
[0042] Specifically, the speech recognition method disclosed in this invention includes the following steps.
[0043] (1) Signal initialization processing: use the device to acquire the voice signal and analyze the characteristics of the signal.
[0044] (2) Using EMD decomposition, the discrete digital signal is decomposed into a set of intrinsic mode components (IMFs) distributed from high frequency to low frequency: [ω1, ω2, ..., ω n | and further calculate the wavelet power spectral density of the corresponding noisy speech signal.
[0045] (3) Based on the signal features obtained in step (1), a correlation threshold λ is obtained. k The formula is Among them, c i Let be the correlation coefficient of the i-th signal; using this correlation threshold, signals with correlation coefficients greater than λ are classified. k High-frequency IMF components [ω1, ω2, ..., ω k The correlation coefficient is less than λ. k Low-frequency IMF components [ω] k+1 ω k+2 ,…,ω n Two groups, high-frequency IMF components [ω1, ω2, ..., ω k The process should proceed to step (4), where the low-frequency IMF component [ω] is processed. k+1 ω k+2 ,…,ω n The process should proceed to step (5).
[0046] (4) Wavelet thresholding for noise reduction. This step specifically includes the following sub-steps.
[0047] 4-1) First, the high-frequency IMF components [ω1, ω2, ..., ω] are analyzed. k Wavelet transform is performed and coefficients are extracted to obtain low-frequency approximation coefficients and high-frequency detail coefficients of each layer.
[0048] 4-2) The fixed threshold λ is obtained by the following formula: In the formula, σ represents the high-frequency IMF components [ω1, ω2, ..., ω]. k The variance of λ is given by N, where N is the observation length. To adapt to different actual noise conditions, this invention modifies the formula for the fixed threshold λ as follows: Where j is the decomposition scale.
[0049] The formula for spectral flatness γ is: γ = 10lg(u g -u a The characteristic information of the modal components is the wavelet power spectral density of the noisy speech signal, u g The geometric mean of the wavelet power spectral density of the noisy speech signal is calculated as follows: u a The arithmetic mean of the wavelet power spectral density of the noisy speech signal is calculated as follows: In the two formulas above, f i Let (i = 1, 2, ..., n) be the wavelet power spectral density of the noisy speech signal. Due to the limitations of the arithmetic mean in terms of adaptability, this invention uses... Replace this formula with a spectral flatness function. Improve the adaptive threshold λ, and add The adaptability of the spectral flattening function to different power spectra can be adjusted by customizing the parameter α.
[0050] The improved adaptive threshold function is: The parameter β∈(0,1) is the reciprocal of the signal-to-noise ratio of the original signal, used to dilute high-density noise. The adaptive threshold takes into account the characteristic that the value of λ gradually decreases as the scale j increases, so as to make it consistent with the propagation characteristics of noise at each scale of wavelet transform. It also takes into account the noise and speech characteristics of noisy speech, making the threshold estimation more accurate and the denoising effect better.
[0051] 4-3) Use spectral flatness γ to adjust the range of the rising function segment, |ω j |∈[λ,γ), for the segment|ω j A hard threshold function is used for |∈[γ,+∞) For segment | ω j |∈(-∞,λ) uses the function The denoised high-frequency IMF components are obtained using the improved threshold function.
[0052]
[0053] The added high-frequency IMF components [ω1, ω2, ..., ω] k The variance σ is used to adapt to different noise spectra, increase the denoising range of the signal, and enable the denoising algorithm to extract the features of the original signal more efficiently. j γ is the value of the j-th layer in the high-frequency IMF component, μ is an improved adaptive parameter that can adjust the sharpening degree of denoising according to the influence of noise on the original signal, and the specific value of γ varies according to the actual waveform to adapt to the diversity of the original signal.
[0054] 4-4) Using the wavelet reconstruction function, the j-layer wavelet components are combined and reconstructed into denoised high-frequency IMF components [ε1, ε2, ..., ε]. k Then proceed to step (6).
[0055] (5) ICEEMDAN noise reduction. This step specifically includes the following sub-steps.
[0056] 5-1) Take the low-frequency IMF component [ω] k+1 ω k+2 ,…,ω n The mean was used as the first IMF component obtained from the ICEEMDAN decomposition.
[0057] 5-2) After adding specific noise to the residual signals of each stage obtained after EMD decomposition, EMD decomposition is performed again. The residual signals after decomposition will be freed from the noise interference and will present the original spectral characteristics of the signal. If the obtained residual signal is sufficient to represent the characteristics of the original signal, proceed to sub-step (5-3); otherwise, return to sub-step (5-1).
[0058] 5-3) Reconstruct the residual signal into a denoised low-frequency IMF component [ε] k+1 ,…ε t-1 , ε t+b+1 ,…ε n Then proceed to step (6).
[0059] (6) Denoise the high-frequency IMF components [ε1, ε2, ..., ε k [ε] and the low-frequency IMF component for noise reduction k+1 ,…ε t-1 , ε t+b+1 ,…ε n The signals are superimposed and recombined to form a pure signal. The pure signal is then decomposed by wavelet decomposition, and the process proceeds to step (7).
[0060] (7) Perform wavelet transform, wavelet packet transform and wavelet multiresolution analysis on the wavelet components of the pure signal to extract signal features and perform speech recognition.
[0061] This invention presents a comprehensive denoising method applicable to various complex scenarios. It addresses the issue that the enhancement effect is unsatisfactory when the same threshold is used for different decomposition scales *j*, and that the threshold should decrease as the noise modulus maxima decreases with increasing scale. The threshold design method ensures that useful speech signal information is largely preserved, adapting to most noisy environments, and its denoising effect is not significantly affected even in non-stationary noise environments. Simulation results show that our self-developed improved wavelet threshold denoising method far surpasses other improved methods in all evaluation metrics. This denoising algorithm, with its significant denoising effect and ability to greatly preserve and amplify signal features, can be applied to speech recognition to address the problem of a sharp drop in signal-to-noise ratio in far-field recognition systems.
[0062] The present invention has been described above by way of example with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited to the above-described manner. Any non-substantial improvements made using the inventive concept and technical solution of the present invention, or the direct application of the inventive concept and technical solution of the present invention to other occasions without modification, are all within the protection scope of the present invention.
Claims
1. A speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising, characterized in that: This includes decomposing the interfered signal EMD into a set of modal components IMF from high frequency to low frequency; wavelet threshold denoising is responsible for processing the high frequency modal components; and ICEEMDAN denoising is responsible for processing the low frequency modal components. The wavelet thresholding denoising responsible for high-frequency modal components obtains an adaptive threshold based on the feature information of the modal components. In the processing of general white noise, the adaptive parameter of the threshold is set to a fixed value. The speech recognition method includes the following steps: (1) Signal initialization processing; (2) Using EMD decomposition, the discrete digital signal is decomposed into a set of intrinsic mode components (IMFs) distributed from high frequency to low frequency, and the wavelet power spectral density of the corresponding noisy speech signal is calculated. (3) Based on the signal characteristics obtained in step (1), a correlation threshold is obtained. Using the correlation threshold, the intrinsic mode components (IMFs) are divided into two groups: high-frequency IMFs with a correlation coefficient greater than the correlation threshold and low-frequency IMFs with a correlation coefficient less than the correlation threshold. The high-frequency IMFs should be processed in step (4), and the low-frequency IMFs should be processed in step (5). (4) Wavelet thresholding for noise reduction; (5) ICEEMDAN noise reduction; (6) The denoised high-frequency IMF components and the denoised low-frequency IMF components are superimposed to form a clean signal. The clean signal is then decomposed by wavelet and proceeded to step (7). (7) Perform wavelet transform, wavelet packet transform and wavelet multiresolution analysis on the wavelet components of the pure signal to extract signal features and perform speech recognition; In step (2), the Intrinsic Modal Components (IMFs) are: ; In step (3), the relevant threshold is The formula is ,in, For the first The correlation coefficient of each signal; the high-frequency IMF component is The low-frequency IMF component is .
2. The speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising as described in claim 1, characterized in that: Step (4) specifically includes the following sub-steps: 4-1) First, analyze the high-frequency IMF components. Perform wavelet transform and decompose to extract coefficients, and obtain low-frequency approximation coefficients and high-frequency detail coefficients of each layer; 4-2) Obtain the adaptive threshold The adaptive threshold function is: In the formula, For high-frequency IMF components variance The length of the observation; For spectral flatness, For the spectral flatness function, add the following to the equation: To use custom parameters Adjusting the adaptability of the spectral flattening function to different power spectra, i.e. ;parameter , The value is the reciprocal of the signal-to-noise ratio of the original signal, used to dilute high-density noise. For decomposition scale; 4-3) Using spectral flatness To adjust the ascending function segment The range of noise-reduced high-frequency IMF components is obtained by using the improved threshold function; 4-4) Use wavelet reconstruction functions to combine and reconstruct the j-level wavelet components into denoised high-frequency IMF components. Then proceed to step (6).
3. The speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising according to claim 2, characterized in that: In sub-step 4-2), the characteristic information of the modal components is the wavelet power spectral density and spectral flatness of the noisy speech signal. The formula is: , The geometric mean of the wavelet power spectral density of the noisy speech signal is calculated as follows: ; The arithmetic mean of the wavelet power spectral density of the noisy speech signal is calculated as follows: The spectral flatness function is: .
4. The speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising as described in claim 3, characterized in that: In sub-step 4-2), the following is adopted: A formula for replacing the arithmetic mean of wavelet power spectral density of noisy speech signals.
5. The speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising as described in claim 4, characterized in that: In sub-step 4-3), the segment Using a hard threshold function For the section Use function The denoised high-frequency IMF components are obtained using the improved threshold function. : in, The value of the j-th layer in the high-frequency IMF components. The improved adaptive parameters can adjust the degree of denoising sharpening based on the impact of noise on the original signal. The specific value varies depending on the actual waveform.
6. The speech recognition method based on wavelet threshold denoising and ICEEMDAN combined denoising as described in claim 5, characterized in that: Step (5) specifically includes the following sub-steps: 5-1) Take the low-frequency IMF component The mean is used as the first IMF component obtained from the ICEEMDAN decomposition; 5-2) After adding specific noise to the residual signals of each stage obtained after EMD decomposition, EMD decomposition is performed again. The residual signals after decomposition will be freed from the noise interference and will show the original spectral characteristics of the signal. If the obtained residual signal is sufficient to represent the characteristics of the original signal, proceed to sub-step (5-3); otherwise, return to sub-step (5-1). 5-3) Reconstruct the residual signal into a denoised low-frequency IMF component. Then proceed to step (6).
Citation Information
Patent Citations
Combined denoising method and system based on improved complementary set mode decomposition
CN114886378A
Equipment operation trend prediction method based on ICEEMDAN secondary decomposition coupling informer model
CN115438301A