Pitch Estimation Using Frequency Domain Phase Differences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional pitch estimation techniques in audio communications systems, such as in-car-communications (ICC) systems, face challenges with short window lengths that are too short to capture full pitch periods, particularly for low male voices, leading to difficulties in resolving low pitch frequencies and high latency in speech processing.
Innovation Solution
A method that estimates pitch frequency directly in the frequency domain by computing phase differences between multiple short windows, employing weighted sums and smoothing constants, and using normalized cross-spectra to determine linearity, allowing for efficient detection of voiced speech and estimation of pitch periods even with very low pitch frequencies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If short window length is used for pitch estimation, then latency is reduced and processing speed is improved, but pitch estimation accuracy deteriorates because the window is too short to capture full pitch periods
Solution Approach 1:
The patent divides the pitch estimation problem into two independent components: (1) voiced speech detection using autocorrelation on short windows to detect periodicity, and (2) pitch frequency estimation using phase differences from frequency domain representations of multiple short windows. This segmentation allows each component to use optimal window lengths for its specific function, resolving the contradiction between short windows for low latency and long windows for accurate pitch period capture.
Solution Approach 2:
The patent transitions from time-domain pitch estimation (which requires long windows to capture full pitch periods) to frequency-domain pitch estimation. By computing phase differences between frequency domain representations of multiple short windows and analyzing their linearity over frequency, the system can estimate pitch frequency without requiring the window length to exceed the pitch period, thus achieving low latency while maintaining accuracy.
2Productivity
If conventional pitch estimation techniques are used with short windows, then processing speed is improved, but the ability to resolve low pitch frequencies deteriorates
Solution Approach 1:
The patent performs preliminary voiced speech detection using autocorrelation on short windows before proceeding to pitch frequency estimation. This preliminary action identifies frames containing periodic signals, allowing the subsequent phase difference analysis to focus only on relevant data. This two-stage approach maintains processing speed while improving low pitch frequency resolution by avoiding computations on non-periodic signals.
Solution Approach 2:
The patent replaces the mechanical time-domain autocorrelation method (which requires shifting and comparing entire long windows) with a frequency-domain phase difference method. By transforming short windows to frequency domain representations and analyzing phase differences, the system achieves better low pitch frequency resolution with shorter windows, as the phase information across multiple frequency bins provides more sensitive pitch period estimation.
Data Source
Figure 1A
Figure 1B
Figure 2
AI summary
A low-complexity method and apparatus for detection of voiced speech and pitch estimation is disclosed that is capable of dealing with special constraints given by applications where low latency is required, such as in-car communication (ICC) systems. An example embodiment employs very short frames that may capture only a single excitation impulse of voiced speech in an audio signal. A distance between multiple such impulses, corresponding to a pitch period, may be determined by evaluating phase differences between low-resolution spectra of the very short frames. An example embodiment may perform pitch estimation directly in a frequency domain based on the phase differences and reduce computational complexity by obviating transformation to a time domain to perform the pitch estimation. In an event the phase differences are determined to be substantially linear, an example embodiment enhances voice quality of the voiced speech by applying speech enhancement to the audio signal.