Advanced Feature Discrimination Vectors for Speaker-Invariant Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in accurately recognizing speech due to variations in speech characteristics among individual speakers, including fundamental frequency, time duration, and accent, which complicates the discrimination of phonemes and sub-phonemes, especially in continuous speech recognition.

Innovation Solution

The method involves generating advanced feature discrimination vectors (AFDVs) by renormalizing high-resolution oscillator peaks extracted from audio signals, eliminating variations in fundamental frequency and time duration, allowing these vectors to be aligned in a common coordinate space for accurate comparison across speakers and utterances.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional speech recognition systems use pre-trained statistical models to account for speaker variations, then they can handle diverse speakers, but the system complexity and computational requirements increase significantly

Engineering Contradiction:
Improveability to handle speaker variationsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes the harmful variations in fundamental frequency and time duration from the speech signal through renormalization. By separating these problematic variations from the core phonemic information, the system eliminates the need for complex statistical models to compensate for speaker differences, thereby reducing system complexity while maintaining adaptability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies parameter changes by renormalizing the speech signal to eliminate variations in fundamental frequency and time duration. This transformation changes the parameter space of the speech signal, making it invariant to speaker-specific variations and enabling simpler comparison across different speakers without requiring extensive statistical modeling.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If the system uses extensive statistical modeling across vast populations to account for speech variations, then recognition accuracy improves, but the processing time and computational resources increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary renormalization of the speech signal to eliminate fundamental frequency and time duration variations before the recognition process. This preliminary action pre-processes the signal to remove sources of variation, so that subsequent recognition can proceed with simpler, faster comparisons rather than requiring extensive statistical modeling during processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By extracting and removing the problematic variations in fundamental frequency and time duration through renormalization, the system eliminates the need for time-consuming statistical modeling to account for these variations, thereby maintaining recognition accuracy while reducing processing time.

Inventive Principle:
Principle #2Taking out (Extraction)

3Loss of information

If the system processes high-resolution oscillator peaks without renormalization, then detailed spectral information is preserved, but variations in fundamental frequency and time duration prevent accurate comparison across speakers

Engineering Contradiction:
Improvespectral information preservationVSAvoidphoneme discrimination accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by renormalizing the high-resolution oscillator peaks to eliminate variations in fundamental frequency and time duration. This transformation maintains the detailed spectral information while changing the parameter space to make it comparable across different speakers, thereby achieving both information preservation and accurate phoneme discrimination.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality by selectively addressing only the specific parameters (fundamental frequency and time duration) that cause comparison problems, while preserving all other spectral information. The renormalization process locally transforms the problematic dimensions without affecting the overall spectral structure, enabling accurate phoneme discrimination while maintaining detailed spectral characteristics.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3042377B1Method and system for generating advanced feature discrimination vectors for use in speech recognition
Publication Date: 2023.01.11 XMOS INC
  • EP3042377B1 patent drawingFigure 1
  • EP3042377B1 patent drawingFigure 2
  • EP3042377B1 patent drawingFigure 3A~3B

AI summary

A method of renormalizing high-resolution oscillator peaks, extracted from windowed samples of an audio signal, is disclosed. Feature vectors are generated for which variations in both fundamental frequency and time duration of speech are substantially mitigated. The feature vectors may be aligned within a common coordinate space, free of those variations in frequency and time duration that occurs between speakers, and even over speech by a single speaker, to facilitate a simple and accurate determination of matches between those AFDVs generated from a sample of the audio signal and corpus AFDVs generated for known speech at the phoneme and sub-phoneme level. The renormalized feature vectors can be combined with traditional feature vectors such as MFCCs, or they can be used exclusively to identify voiced, semi-voiced and unvoiced sounds.