Speech Data Classification Using Formant and Pitch Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speaker normalization techniques face challenges in accurately classifying speakers due to the variability of vocal tract length and shape, leading to suboptimal recognition rates and high computational complexity.

Innovation Solution

The method employs formant frequencies and pitch, dependent on vocal tract length and shape, as feature parameters for clustering to train a classifier, which selects a warping factor for normalizing speech, thereby improving speaker classification and reducing computational load.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If linear search method is used to obtain warping factor, then speaker classification accuracy is improved, but computational complexity increases significantly

Engineering Contradiction:
Improvespeaker classification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the approach from searching for optimal warping factors to directly estimating them using formant frequencies and pitch as features. This parameter transformation converts a computationally intensive search problem into a more efficient estimation problem, reducing computational complexity while maintaining classification accuracy.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical linear search process with a classifier-based estimation system that uses acoustic features (formant frequencies and pitch). This substitution eliminates the need for iterative searching and replaces it with a direct classification approach, significantly reducing computational burden.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If parametric approach is used to obtain warping factor, then computational load is reduced, but recognition rate becomes unstable

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidrecognition rate stability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent incorporates feedback mechanisms through the classifier training process, where the system learns from training data to establish reliable relationships between acoustic features and speaker classes. This feedback loop ensures that the estimation process becomes stable and reliable, overcoming the instability of pure parametric approaches.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary classification and feature extraction before warping factor estimation. By pre-processing the speech data to extract formant frequencies and pitch, and training a classifier on these features, the system establishes a stable foundation that improves recognition rate consistency while maintaining computational efficiency.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If more speaker classes are used for normalization, then recognition accuracy is improved, but computation load increases heavily

Engineering Contradiction:
Improverecognition accuracyVSAvoidcomputation load
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the approach from using multiple speaker classes requiring extensive computation to using acoustic features (formant frequencies and pitch) that directly capture speaker characteristics. This parameter transformation allows the system to handle speaker variability efficiently without requiring numerous discrete speaker classes, thus reducing computation load while maintaining recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7957959B2Method and apparatus for processing speech data with classification models
Publication Date: 2011.06.07 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7957959B2 patent drawing
  • US7957959B2 patent drawing
  • US7957959B2 patent drawing

AI summary

A method for processing speech data includes obtaining a pitch and at least one formant frequency for each of a plurality of first speech data; constructing a first feature space with the obtained fundamental frequencies and formant frequencies as features; and classifying the plurality pieces of first speech data using the first feature space, and thus a plurality of speech data classes and the corresponding description are obtained.