Word Alignment Speech Processing for Real-Time Fluency Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computerized language assessment systems struggle to accurately determine a user's proficiency and/or fluency in a foreign language, leading to delays and disparities in scoring due to reliance on human examiners.
Innovation Solution
A speech processing system that aligns acoustic speech models with user utterances to identify skipped, repeated, or inserted speech sounds, using networks to account for various pronunciations and silences, and employs speech scoring features to generate fluency and proficiency scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human examiners manually assess each user's speech, then the accuracy and reliability of fluency and proficiency scoring is maintained, but the time required for assessment increases significantly and productivity decreases
Solution Approach 1:
The patent creates a computerized system that copies and replicates human examiner assessment capabilities through acoustic models, language models, and scoring algorithms. The system captures human assessment criteria in digital form, allowing multiple users to be assessed simultaneously without requiring actual human examiners for each case, thus maintaining reliability while dramatically improving productivity
Solution Approach 2:
The patent replaces the mechanical system of human examiners physically listening and evaluating speech with an automated acoustic processing system. The mechanical action of human hearing and cognitive evaluation is substituted with electronic signal processing, acoustic model matching, and automated scoring algorithms, enabling high-speed assessment without sacrificing accuracy
2Reliability
If multiple human examiners assess each user to avoid score disparity, then the reliability and consistency of scoring improves, but the time delay for obtaining results increases further
Solution Approach 1:
The patent merges multiple assessment perspectives into a single integrated computerized system. Instead of requiring multiple sequential human examiners, the system combines acoustic model analysis, language model evaluation, and scoring algorithms into one unified process that can simultaneously evaluate speech from multiple dimensions and produce a single consistent score, eliminating the time delay of sequential human assessment
Solution Approach 2:
The patent enables continuous, uninterrupted assessment processing where the computerized system can evaluate multiple users simultaneously without the breaks, consultations, and sequential processing required by human examiners. The system maintains continuous operation, processing speech assessments in real-time or near-real-time, thereby eliminating the delays inherent in human-based multi-examiner protocols
3Productivity
If a computerized marking system is implemented to reduce reliance on human examiners, then productivity and assessment speed improve, but the accuracy and reliability of proficiency determination deteriorates
Solution Approach 1:
The patent employs parameter changes by adjusting and optimizing acoustic model parameters, language model parameters, and scoring algorithm parameters based on training data from human-assessed speech samples. The system learns and adapts parameter values that correlate human speech characteristics with human examiner scores, enabling the computerized system to achieve measurement precision comparable to human examiners while maintaining high productivity
Solution Approach 2:
The patent incorporates feedback mechanisms where the computerized assessment system is trained and calibrated using feedback from human examiner assessments. The system compares its automated scores with human examiner scores, identifies discrepancies, and adjusts its acoustic models and scoring algorithms accordingly. This feedback loop enables the system to continuously improve its accuracy and reliability in determining language proficiency
4Adaptability or versatility
If the system accounts for various pronunciations, skips, repetitions, and inserted sounds through multiple network paths, then the adaptability and coverage of speech assessment improves, but the device complexity increases
Solution Approach 1:
The patent segments the complex task of handling speech variations into separate, modular components: acoustic models handle phonetic recognition, language models handle grammatical structure, and specific modules handle skips, repetitions, and insertions. By dividing the overall assessment function into discrete segments, the system can manage complexity through modular design while maintaining high adaptability to various speech patterns and variations
Data Source
AI summary
A speech processing system includes an input for receiving an input utterance spoken by a user and a word alignment unit configured to align different sequences of acoustic speech models with the input utterance spoken by the user. Each different sequence of acoustic speech models corresponds to a different possible utterance that a user might make. The system identifies any parts of a read prompt text that the user skipped; any parts of the read prompt text that the user repeated; and any speech sounds that the user inserted between words of the read prompt text. The information from the word alignment unit can be used to assess the proficiency and/or fluency of the user's speech.


