Spontaneous Speech Pronunciation Assessment via Acoustic Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated systems face challenges in evaluating spontaneous speech pronunciation of non-native speakers due to difficulties in recognizing and assessing spontaneous speech, which limits the assessment of communicative competence beyond scripted texts.
Innovation Solution
A computer-implemented system using non-native and native acoustic models for speech recognition, time alignment, and feature calculation to assess spontaneous speech pronunciation, generating an assessment score based on phonetic and phonemic statistics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated speech recognition systems are used to evaluate spontaneous speech, then assessment efficiency is improved, but measurement precision deteriorates due to difficulties in recognizing and assessing spontaneous speech
Solution Approach 1:
The patent introduces an intermediary alignment process that maps recognized words to phonemes using acoustic models. This intermediary step bridges the gap between automated speech recognition and pronunciation assessment, allowing the system to evaluate phoneme-level pronunciation accuracy even when complete word recognition is challenging in spontaneous non-native speech
Solution Approach 2:
The system changes the assessment parameters from word-level recognition metrics to phoneme-level pronunciation metrics. By calculating pronunciation statistics based on aligned phonemes rather than relying solely on word recognition accuracy, the system maintains measurement precision while preserving automated assessment efficiency
2Adaptability or versatility
If non-native acoustic models are used for speech recognition, then adaptability to non-native speakers is improved, but device complexity increases due to multiple acoustic models
Solution Approach 1:
The patent segments the acoustic modeling into distinct components: a non-native acoustic model for initial speech recognition and a separate reference acoustic model for alignment and pronunciation assessment. This segmentation allows each model to be optimized for its specific function without requiring a single complex model to handle all aspects
Solution Approach 2:
The reference acoustic model serves multiple functions: it provides the basis for time alignment between recognized speech and reference speech, enables phoneme-level pronunciation assessment, and establishes ground truth for calculating pronunciation statistics. This multi-functionality reduces the need for additional specialized models
Data Source
AI summary
Computer-implemented systems and methods are provided for assessing non-native spontaneous speech pronunciation. Speech recognition on digitized speech is performed using a non-native acoustic model trained with non-native speech to generate word hypotheses for the digitized speech. Time alignment is performed between the digitized speech and the word hypotheses using a reference acoustic model trained with native-quality speech. Statistics are calculated regarding individual words and phonemes in the word hypotheses based on the alignment. A plurality of features for use in assessing pronunciation of the speech are calculated based on the statistics, an assessment score is calculated based on one or more of the calculated features, and the assessment score is stored in a computer-readable memory.


