Pronunciation Error Detection via Dual Model Reliability Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing pronunciation evaluation systems require correct sentences for GOP score calculation, making it difficult to handle misrecognition errors and achieve learning effects in real-world scenarios.
Innovation Solution
A pronunciation error detection apparatus that includes a non-native speaker speech recognition model and a native speaker speech recognition model under weakly constraining grammar, allowing for the detection of pronunciation errors without relying on correct sentences, by comparing the reliability of phoneme recognition results between the two models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a non-native speaker speech recognition model is used, then the system can recognize speech from non-native speakers, but it cannot detect pronunciation errors without correct sentences
Solution Approach 1:
The patent introduces a native speaker speech recognition model as an intermediary to mediate between the non-native speaker speech and the pronunciation error detection task. This intermediary model provides reliable phoneme-level probability information that serves as a reference for detecting pronunciation errors in non-native speaker speech, enabling error detection without requiring correct sentences.
2Reliability
If correct sentences are required for GOP score calculation, then pronunciation evaluation can be performed, but learners cannot achieve learning effects in real-world scenarios with erroneous readings
Solution Approach 1:
The patent inverts the traditional approach by not requiring correct sentences as input. Instead, it uses the native speaker model's phoneme probabilities as a reference to identify deviations in non-native speaker speech, thereby detecting pronunciation errors directly from erroneous readings without needing correct sentences for comparison.
Solution Approach 2:
The system provides feedback by comparing the non-native speaker's phoneme probabilities against the native speaker model's probabilities. This feedback mechanism identifies pronunciation errors by highlighting significant deviations, enabling learners to improve their pronunciation based on the detected errors without requiring correct sentences.
3Measurement precision
If a native speaker speech recognition model under weakly constraining grammar is used, then phoneme reliability can be accurately assessed, but the system complexity increases
Solution Approach 1:
The patent segments the speech recognition task into two specialized components: a non-native speaker model for initial speech recognition and a native speaker model for pronunciation error detection. Each model is optimized for its specific function, with the native speaker model operating under weakly constraining grammar to provide accurate phoneme-level probabilities without the complexity of full sentence grammatical constraints.
Data Source
AI summary
The present invention provides a pronunciation error detection apparatus capable of following a text without the need for a correct sentence even when erroneous recognition such as a reading error occurs. The pronunciation error detection apparatus comprises: a speech recognition part that recognizes the speech in speech data based on a speech recognition model for a non-native speaker, and outputs speech recognition results, reliability and time information; a reliability determination part that outputs the speech recognition results with higher reliability than a predetermined threshold and the corresponding time information as the determined speech recognition results and the determined time information; and a pronunciation error detection part that outputs a phoneme as a pronunciation error when reliability for each phoneme in the speech recognition results using the native speaker speech recognition model under a weakly constraining grammar is greater than the reliability of the corresponding phoneme in the speech recognition results using the native speaker acoustic model under a constraining grammar in which the determined speech recognition results are correct for the speech data in a segment specified by the determined time information.


