Word Alignment Speech Processing for Real-Time Fluency Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computerized language assessment systems struggle to accurately determine a user's proficiency and/or fluency in a foreign language, leading to delays and disparities in scoring due to reliance on human examiners.

Innovation Solution

A speech processing system that aligns acoustic speech models with user utterances to identify skipped, repeated, or inserted speech sounds, using networks to account for various pronunciations and silences, and employs speech scoring features to generate fluency and proficiency scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human examiners manually assess each user's speech, then the accuracy and reliability of fluency and proficiency scoring is maintained, but the time required for assessment increases significantly and productivity decreases

Engineering Contradiction:
Improvescoring accuracyVSAvoidassessment speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent creates a computerized system that copies and replicates human examiner assessment capabilities through acoustic models, language models, and scoring algorithms. The system captures human assessment criteria in digital form, allowing multiple users to be assessed simultaneously without requiring actual human examiners for each case, thus maintaining reliability while dramatically improving productivity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical system of human examiners physically listening and evaluating speech with an automated acoustic processing system. The mechanical action of human hearing and cognitive evaluation is substituted with electronic signal processing, acoustic model matching, and automated scoring algorithms, enabling high-speed assessment without sacrificing accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If multiple human examiners assess each user to avoid score disparity, then the reliability and consistency of scoring improves, but the time delay for obtaining results increases further

Engineering Contradiction:
Improvescoring consistencyVSAvoidassessment delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent merges multiple assessment perspectives into a single integrated computerized system. Instead of requiring multiple sequential human examiners, the system combines acoustic model analysis, language model evaluation, and scoring algorithms into one unified process that can simultaneously evaluate speech from multiple dimensions and produce a single consistent score, eliminating the time delay of sequential human assessment

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent enables continuous, uninterrupted assessment processing where the computerized system can evaluate multiple users simultaneously without the breaks, consultations, and sequential processing required by human examiners. The system maintains continuous operation, processing speech assessments in real-time or near-real-time, thereby eliminating the delays inherent in human-based multi-examiner protocols

Inventive Principle:
Principle #20Continuity of useful action

3Productivity

If a computerized marking system is implemented to reduce reliance on human examiners, then productivity and assessment speed improve, but the accuracy and reliability of proficiency determination deteriorates

Engineering Contradiction:
Improveassessment efficiencyVSAvoidproficiency assessment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent employs parameter changes by adjusting and optimizing acoustic model parameters, language model parameters, and scoring algorithm parameters based on training data from human-assessed speech samples. The system learns and adapts parameter values that correlate human speech characteristics with human examiner scores, enabling the computerized system to achieve measurement precision comparable to human examiners while maintaining high productivity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent incorporates feedback mechanisms where the computerized assessment system is trained and calibrated using feedback from human examiner assessments. The system compares its automated scores with human examiner scores, identifies discrepancies, and adjusts its acoustic models and scoring algorithms accordingly. This feedback loop enables the system to continuously improve its accuracy and reliability in determining language proficiency

Inventive Principle:
Principle #23Feedback

4Adaptability or versatility

If the system accounts for various pronunciations, skips, repetitions, and inserted sounds through multiple network paths, then the adaptability and coverage of speech assessment improves, but the device complexity increases

Engineering Contradiction:
Improvespeech variation handlingVSAvoidnetwork structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the complex task of handling speech variations into separate, modular components: acoustic models handle phonetic recognition, language models handle grammatical structure, and specific modules handle skips, repetitions, and insertions. By dividing the overall assessment function into discrete segments, the system can manage complexity through modular design while maintaining high adaptability to various speech patterns and variations

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12380882B2Speech processing system and method
Publication Date: 2025.08.05 THE CHANCELLOR MASTERS & SCHOLARS OF THE UNIV OF CAMBRIDGE
  • US12380882B2 patent drawing
  • US12380882B2 patent drawing
  • US12380882B2 patent drawing

AI summary

A speech processing system includes an input for receiving an input utterance spoken by a user and a word alignment unit configured to align different sequences of acoustic speech models with the input utterance spoken by the user. Each different sequence of acoustic speech models corresponds to a different possible utterance that a user might make. The system identifies any parts of a read prompt text that the user skipped; any parts of the read prompt text that the user repeated; and any speech sounds that the user inserted between words of the read prompt text. The information from the word alignment unit can be used to assess the proficiency and/or fluency of the user's speech.