Speech Data Classification Using Formant and Pitch Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speaker normalization techniques face challenges in accurately classifying speakers due to the variability of vocal tract length and shape, leading to suboptimal recognition rates and high computational complexity.
Innovation Solution
The method employs formant frequencies and pitch, dependent on vocal tract length and shape, as feature parameters for clustering to train a classifier, which selects a warping factor for normalizing speech, thereby improving speaker classification and reducing computational load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If linear search method is used to obtain warping factor, then speaker classification accuracy is improved, but computational complexity increases significantly
Solution Approach 1:
The patent changes the approach from searching for optimal warping factors to directly estimating them using formant frequencies and pitch as features. This parameter transformation converts a computationally intensive search problem into a more efficient estimation problem, reducing computational complexity while maintaining classification accuracy.
Solution Approach 2:
The patent replaces the mechanical linear search process with a classifier-based estimation system that uses acoustic features (formant frequencies and pitch). This substitution eliminates the need for iterative searching and replaces it with a direct classification approach, significantly reducing computational burden.
2Productivity
If parametric approach is used to obtain warping factor, then computational load is reduced, but recognition rate becomes unstable
Solution Approach 1:
The patent incorporates feedback mechanisms through the classifier training process, where the system learns from training data to establish reliable relationships between acoustic features and speaker classes. This feedback loop ensures that the estimation process becomes stable and reliable, overcoming the instability of pure parametric approaches.
Solution Approach 2:
The patent performs preliminary classification and feature extraction before warping factor estimation. By pre-processing the speech data to extract formant frequencies and pitch, and training a classifier on these features, the system establishes a stable foundation that improves recognition rate consistency while maintaining computational efficiency.
3Measurement precision
If more speaker classes are used for normalization, then recognition accuracy is improved, but computation load increases heavily
Solution Approach 1:
The patent changes the approach from using multiple speaker classes requiring extensive computation to using acoustic features (formant frequencies and pitch) that directly capture speaker characteristics. This parameter transformation allows the system to handle speaker variability efficiently without requiring numerous discrete speaker classes, thus reducing computation load while maintaining recognition accuracy.
Data Source
AI summary
A method for processing speech data includes obtaining a pitch and at least one formant frequency for each of a plurality of first speech data; constructing a first feature space with the obtained fundamental frequencies and formant frequencies as features; and classifying the plurality pieces of first speech data using the first feature space, and thus a plurality of speech data classes and the corresponding description are obtained.


