Low signal to noise ratio voice endpoint detection method based on time-frequency instaneous energy spectrum
An instantaneous energy spectrum, endpoint detection technology, applied in speech analysis, speech recognition, instruments, etc., can solve the problems of modal aliasing, unsatisfactory voice endpoint accuracy, and weak anti-noise ability, to improve stability, The effect of reducing program running time and improving accuracy
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Publication Date
- 2013-05-22
Smart Images
Figure 1 Figure 2 Figure 3
Abstract
Description
technical field
[0001] The invention belongs to the field of voice processing, and relates to a low signal-to-noise ratio voice endpoint detection method based on time-frequency instantaneous energy spectrum. Background technique
[0002] Various noises will inevitably be introduced in the process of voice acquisition, transmission and communication, and the existence of noise will directly affect the clarity and intelligibility of voice. Endpoint detection of noisy speech signals to obtain the start and end points of effective speech segments plays a very important role in subsequent speech enhancement, coding and recognition. At present, the traditional endpoint detection methods mainly include average energy, average zero-crossing rate, cepstral coefficient, short-time frequency band variance, short-time energy-frequency value, cepstrum distance, autocorrelation similarity distance, information entropy and spectral entropy. But they are all based on the assumption that t...
Examples
Embodiment Construction
[0021] Below in conjunction with accompanying drawing, the present invention will be further described, and the concrete steps of the inventive method are:
[0022] Step (1) For the noisy speech signal under the background of strong noise (Such as figure 1 shown) plus Hamming window processing. Use the db3 wavelet basis function in Daubechies to perform three-layer wavelet packet decomposition on the windowed noisy speech signal, and the schematic diagram of the wavelet packet decomposition binary tree is as follows figure 2 shown. Reconstruct the decomposed result to obtain the reconstruction signal, denoted as , and the corresponding frequency bands are ,in for the minimum frequency resolution, , is the sampling frequency.
[0023] Step (2) will reconstruct the obtained low frequency component signal Perform adaptive EMD decomposition (the first 7 IMF components such as image 3 shown), thus obtaining a finite number of IMF components and residual signal ...