Accent Compensative Speech Recognition via Phoneme Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems face significant challenges in accurately recognizing speech from non-native speakers due to accent-related pronunciation variations, leading to reduced accuracy and potential system inoperability, especially when the user interface is limited to a specific language.

Innovation Solution

An accent compensative speech recognition system that uses a first-language acoustic module to determine a phoneme sequence from a voice-induced electrical signal, paired with a second-language lexicon module to recognize speech segments in a desired language, without the need for new acoustic models or extensive memory allocation, allowing for improved recognition accuracy across languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech recognition systems use language-specific acoustic models, then recognition accuracy is improved for native speakers, but recognition accuracy deteriorates for non-native speakers with accents

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidlanguage adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the speech recognition task into two independent components: accent-independent phoneme recognition (acoustic model) and language-specific word identification (lexicon model). This segmentation allows the acoustic model to focus on universal phoneme features that work across all languages and accents, while the lexicon model handles language-specific variations, thereby resolving the contradiction between native speaker accuracy and non-native speaker adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal acoustic model that can process speech from any language or accent by focusing on phoneme-level features rather than language-specific features. This universal model serves multiple languages simultaneously, eliminating the need for separate acoustic models for each language while maintaining high recognition accuracy for both native and non-native speakers.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If speech recognition systems create separate acoustic models for different languages and accents, then recognition accuracy for non-native speakers is improved, but device complexity and memory requirements increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the speech recognition system into two functional segments: a universal accent-independent acoustic model for phoneme recognition and a language-specific lexicon model for word identification. This segmentation eliminates the need to create and maintain multiple complete acoustic models for different languages and accents, thereby reducing device complexity while maintaining recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of creating separate acoustic models for each language and accent combination, the patent uses a single copied universal acoustic model that works across all languages. The language-specific information is stored separately in lexicon models, avoiding the exponential growth of system complexity that would result from creating dedicated acoustic models for every language-accent pair.

Inventive Principle:
Principle #26Copying

3Measurement precision

If speech recognition systems use language-specific acoustic models, then recognition accuracy is improved, but memory allocation requirements increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidmemory allocation
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments memory requirements into two parts: a shared universal acoustic model that is loaded once and reused for all languages, and separate language-specific lexicon models that are loaded as needed. This segmentation dramatically reduces total memory allocation compared to loading complete language-specific acoustic models for each language.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The universal acoustic model serves as a multi-functional component that processes speech from any language or accent without requiring separate model instances. This universality eliminates redundant memory storage across multiple acoustic models, reducing overall memory allocation while maintaining recognition accuracy through the language-specific lexicon models.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7640159B2System and method of speech recognition for non-native speakers of a language
Publication Date: 2009.12.29 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7640159B2 patent drawing
  • US7640159B2 patent drawing
  • US7640159B2 patent drawing

AI summary

An accent compensative speech recognition system and related methods for use with a signal processor generating one or more feature vectors based upon a voice-induced electrical signal are provided. The system includes a first-language acoustic module that determines a first-language phoneme sequence based upon one or more feature vectors, and a second-language lexicon module that determines a second-language speech segment based upon the first-language phoneme sequence. A method aspect includes the steps of generating a first-language phoneme sequence from at least one feature vector based upon a first-language phoneme model, and determining a second-language speech segment from the first-language phoneme sequence based upon a second-language lexicon model.