Multi-Dimensional Pitch Tracking Using Extremum Strength Values
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech detection processes fail to accurately estimate pitch periods in noisy environments and struggle with signals containing two voices with similar pitches, as they rely on one-dimensional models susceptible to noise and harmonic interactions.
Innovation Solution
A multi-dimensional model is used to compare input signals with delayed versions, calculating the strength of extrema to identify pitches, and a filter module extracts the pitch of the input signal based on these strengths, effectively distinguishing between voices with similar pitches.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If one-dimensional functions are used for speech detection, then the process is simple, but pitch period estimation accuracy deteriorates in noisy environments
Solution Approach 1:
The patent transitions from one-dimensional pitch detection to two-dimensional pitch detection by analyzing both time lag and frequency shift dimensions simultaneously. The two-dimensional function evaluates pitch periods by comparing signals across multiple time delays and frequency offsets, creating a dimensional space where true pitch periods form distinct peaks that are resistant to noise interference.
2Use of energy by moving object
If one-dimensional functions are used for speech detection, then computational requirements are low, but reliability deteriorates when voices have similar pitches
Solution Approach 1:
By introducing a second dimension (frequency shift) to the pitch detection process, the patent creates a two-dimensional search space where pitch periods of different voices manifest as separate peaks. This dimensional expansion enables reliable distinction between voices with similar pitches, as each voice occupies a unique position in the two-dimensional lag-frequency space.
3Adaptability or versatility
If traditional speech detection processes are used, then they work for single voice signals, but they fail to distinguish between multiple voices with similar pitches
Solution Approach 1:
The two-dimensional function creates a search space with lag and frequency shift dimensions, allowing multiple pitch periods to be identified simultaneously. Each voice in the mixture produces peaks at its specific lag and frequency shift values, enabling the system to distinguish and track multiple voices even when their pitches are similar or when one pitch is a multiple of another.
Data Source
AI summary
An apparatus includes a function module, a strength module, and a filter module. The function module compares an input signal, which has a component, to a first delayed version of the input signal and a second delayed version of the input signal to produce a multi-dimensional model. The strength module calculates a strength of each extremum from a plurality of extrema of the multi-dimensional model based on a value of at least one opposite extremum of the multi-dimensional model. The strength module then identifies a first extremum from the plurality of extrema, which is associated with a pitch of the component of the input signal, that has the strength greater than the strength of the remaining extrema. The filter module extracts the pitch of the component from the input signal based on the strength of the first extremum.


