Advanced Feature Discrimination Vectors for Speaker-Invariant Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in accurately recognizing speech due to variations in speech characteristics among individual speakers, including fundamental frequency, time duration, and accent, which complicates the discrimination of phonemes and sub-phonemes, especially in continuous speech recognition.
Innovation Solution
The method involves generating advanced feature discrimination vectors (AFDVs) by renormalizing high-resolution oscillator peaks extracted from audio signals, eliminating variations in fundamental frequency and time duration, allowing these vectors to be aligned in a common coordinate space for accurate comparison across speakers and utterances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional speech recognition systems use pre-trained statistical models to account for speaker variations, then they can handle diverse speakers, but the system complexity and computational requirements increase significantly
Solution Approach 1:
The patent extracts and removes the harmful variations in fundamental frequency and time duration from the speech signal through renormalization. By separating these problematic variations from the core phonemic information, the system eliminates the need for complex statistical models to compensate for speaker differences, thereby reducing system complexity while maintaining adaptability.
Solution Approach 2:
The patent applies parameter changes by renormalizing the speech signal to eliminate variations in fundamental frequency and time duration. This transformation changes the parameter space of the speech signal, making it invariant to speaker-specific variations and enabling simpler comparison across different speakers without requiring extensive statistical modeling.
2Measurement precision
If the system uses extensive statistical modeling across vast populations to account for speech variations, then recognition accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary renormalization of the speech signal to eliminate fundamental frequency and time duration variations before the recognition process. This preliminary action pre-processes the signal to remove sources of variation, so that subsequent recognition can proceed with simpler, faster comparisons rather than requiring extensive statistical modeling during processing.
Solution Approach 2:
By extracting and removing the problematic variations in fundamental frequency and time duration through renormalization, the system eliminates the need for time-consuming statistical modeling to account for these variations, thereby maintaining recognition accuracy while reducing processing time.
3Loss of information
If the system processes high-resolution oscillator peaks without renormalization, then detailed spectral information is preserved, but variations in fundamental frequency and time duration prevent accurate comparison across speakers
Solution Approach 1:
The patent applies parameter changes by renormalizing the high-resolution oscillator peaks to eliminate variations in fundamental frequency and time duration. This transformation maintains the detailed spectral information while changing the parameter space to make it comparable across different speakers, thereby achieving both information preservation and accurate phoneme discrimination.
Solution Approach 2:
The patent applies local quality by selectively addressing only the specific parameters (fundamental frequency and time duration) that cause comparison problems, while preserving all other spectral information. The renormalization process locally transforms the problematic dimensions without affecting the overall spectral structure, enabling accurate phoneme discrimination while maintaining detailed spectral characteristics.
Data Source
Figure 1
Figure 2
Figure 3A~3B
AI summary
A method of renormalizing high-resolution oscillator peaks, extracted from windowed samples of an audio signal, is disclosed. Feature vectors are generated for which variations in both fundamental frequency and time duration of speech are substantially mitigated. The feature vectors may be aligned within a common coordinate space, free of those variations in frequency and time duration that occurs between speakers, and even over speech by a single speaker, to facilitate a simple and accurate determination of matches between those AFDVs generated from a sample of the audio signal and corpus AFDVs generated for known speech at the phoneme and sub-phoneme level. The renormalized feature vectors can be combined with traditional feature vectors such as MFCCs, or they can be used exclusively to identify voiced, semi-voiced and unvoiced sounds.