Voice Recognition Phoneme Competition Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Reading software often focuses on reading skills other than fluency, and existing speech recognition systems struggle to accurately assess reading fluency by failing to differentiate between correct and incorrect pronunciations, leading to false positives and negatives, which hinders the development of decoding skills, comprehension, and vocabulary.
Innovation Solution
A method that segments words into consecutive phonemes, stores sequences including complete, truncated, and mispronunciation phoneme sequences, and compares utterances to determine accuracy using predefined pronunciation levels (loose, medium, and strict) to assess reading fluency, reducing false positives and negatives by using competition models that represent potential mispronunciations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speech recognition systems are used to assess reading fluency, then the system is simple to operate, but the measurement precision is poor due to inability to differentiate between correct and incorrect pronunciations
Solution Approach 1:
The patent segments words into phoneme sequences and creates multiple competition models for each word, where each model represents a different pronunciation variant (correct or incorrect). This segmentation allows the system to precisely differentiate between correct and incorrect pronunciations by comparing the spoken input against multiple phoneme-level competition models, thereby improving measurement precision without requiring complex external systems
Solution Approach 2:
The patent changes the parameter of pronunciation representation from single-word recognition to multi-variant phoneme sequence recognition. By representing each word with multiple phoneme sequences (including correct and incorrect pronunciations) and assigning different confidence thresholds, the system achieves higher measurement precision while maintaining operational simplicity through parameter-based differentiation
2Reliability
If speech recognition systems use single-word recognition models, then the device complexity is low, but the reliability is poor due to false positives and negatives in pronunciation assessment
Solution Approach 1:
The patent performs preliminary action by pre-generating multiple competition models for each word before the actual speech recognition task. These competition models include both correct and incorrect pronunciation variants, allowing the system to anticipate and differentiate between various pronunciation outcomes. This preliminary preparation improves reliability by ensuring that all possible pronunciation variants are accounted for in advance, reducing false positives and negatives
Solution Approach 2:
The patent creates multiple copies of phoneme sequences representing different pronunciation variants (correct and incorrect) for each word. These copied phoneme models serve as competition models that the speech recognition system can compare against the spoken input. By having multiple copied representations of pronunciation variants, the system improves reliability through comparative analysis while keeping the base phoneme inventory manageable
Data Source
AI summary
A system and method relate to voice recognition software, and more particularly to voice recognition tutoring software to assist in reading development.


