Voice Recognition Phoneme Competition Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reading software often focuses on reading skills other than fluency, and existing speech recognition systems struggle to accurately assess reading fluency by failing to differentiate between correct and incorrect pronunciations, leading to false positives and negatives, which hinders the development of decoding skills, comprehension, and vocabulary.

Innovation Solution

A method that segments words into consecutive phonemes, stores sequences including complete, truncated, and mispronunciation phoneme sequences, and compares utterances to determine accuracy using predefined pronunciation levels (loose, medium, and strict) to assess reading fluency, reducing false positives and negatives by using competition models that represent potential mispronunciations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech recognition systems are used to assess reading fluency, then the system is simple to operate, but the measurement precision is poor due to inability to differentiate between correct and incorrect pronunciations

Engineering Contradiction:
Improvepronunciation assessment accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments words into phoneme sequences and creates multiple competition models for each word, where each model represents a different pronunciation variant (correct or incorrect). This segmentation allows the system to precisely differentiate between correct and incorrect pronunciations by comparing the spoken input against multiple phoneme-level competition models, thereby improving measurement precision without requiring complex external systems

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of pronunciation representation from single-word recognition to multi-variant phoneme sequence recognition. By representing each word with multiple phoneme sequences (including correct and incorrect pronunciations) and assigning different confidence thresholds, the system achieves higher measurement precision while maintaining operational simplicity through parameter-based differentiation

Inventive Principle:
Principle #35Parameter changes

2Reliability

If speech recognition systems use single-word recognition models, then the device complexity is low, but the reliability is poor due to false positives and negatives in pronunciation assessment

Engineering Contradiction:
Improvepronunciation assessment reliabilityVSAvoidrecognition model complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-generating multiple competition models for each word before the actual speech recognition task. These competition models include both correct and incorrect pronunciation variants, allowing the system to anticipate and differentiate between various pronunciation outcomes. This preliminary preparation improves reliability by ensuring that all possible pronunciation variants are accounted for in advance, reducing false positives and negatives

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates multiple copies of phoneme sequences representing different pronunciation variants (correct and incorrect) for each word. These copied phoneme models serve as competition models that the speech recognition system can compare against the spoken input. By having multiple copied representations of pronunciation variants, the system improves reliability through comparative analysis while keeping the base phoneme inventory manageable

Inventive Principle:
Principle #26Copying

Data Source

PatentUS7624013B2Word competition models in voice recognition
Publication Date: 2009.11.24 SCIENTIFIC LEARNING CORP
  • US7624013B2 patent drawing
  • US7624013B2 patent drawing
  • US7624013B2 patent drawing

AI summary

A system and method relate to voice recognition software, and more particularly to voice recognition tutoring software to assist in reading development.