Pronunciation Error Detection via Dual Model Reliability Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing pronunciation evaluation systems require correct sentences for GOP score calculation, making it difficult to handle misrecognition errors and achieve learning effects in real-world scenarios.

Innovation Solution

A pronunciation error detection apparatus that includes a non-native speaker speech recognition model and a native speaker speech recognition model under weakly constraining grammar, allowing for the detection of pronunciation errors without relying on correct sentences, by comparing the reliability of phoneme recognition results between the two models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a non-native speaker speech recognition model is used, then the system can recognize speech from non-native speakers, but it cannot detect pronunciation errors without correct sentences

Engineering Contradiction:
Improveability to handle non-native speaker speechVSAvoidpronunciation error detection accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a native speaker speech recognition model as an intermediary to mediate between the non-native speaker speech and the pronunciation error detection task. This intermediary model provides reliable phoneme-level probability information that serves as a reference for detecting pronunciation errors in non-native speaker speech, enabling error detection without requiring correct sentences.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If correct sentences are required for GOP score calculation, then pronunciation evaluation can be performed, but learners cannot achieve learning effects in real-world scenarios with erroneous readings

Engineering Contradiction:
Improvepronunciation evaluation reliabilityVSAvoidapplicability to real-world learning scenarios
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent inverts the traditional approach by not requiring correct sentences as input. Instead, it uses the native speaker model's phoneme probabilities as a reference to identify deviations in non-native speaker speech, thereby detecting pronunciation errors directly from erroneous readings without needing correct sentences for comparison.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The system provides feedback by comparing the non-native speaker's phoneme probabilities against the native speaker model's probabilities. This feedback mechanism identifies pronunciation errors by highlighting significant deviations, enabling learners to improve their pronunciation based on the detected errors without requiring correct sentences.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If a native speaker speech recognition model under weakly constraining grammar is used, then phoneme reliability can be accurately assessed, but the system complexity increases

Engineering Contradiction:
Improvephoneme reliability assessment accuracyVSAvoiddual recognition model system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech recognition task into two specialized components: a non-native speaker model for initial speech recognition and a native speaker model for pronunciation error detection. Each model is optimized for its specific function, with the native speaker model operating under weakly constraining grammar to provide accurate phoneme-level probabilities without the complexity of full sentence grammatical constraints.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11568761B2Pronunciation error detection apparatus, pronunciation error detection method and program
Publication Date: 2023.01.31 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11568761B2 patent drawing
  • US11568761B2 patent drawing
  • US11568761B2 patent drawing

AI summary

The present invention provides a pronunciation error detection apparatus capable of following a text without the need for a correct sentence even when erroneous recognition such as a reading error occurs. The pronunciation error detection apparatus comprises: a speech recognition part that recognizes the speech in speech data based on a speech recognition model for a non-native speaker, and outputs speech recognition results, reliability and time information; a reliability determination part that outputs the speech recognition results with higher reliability than a predetermined threshold and the corresponding time information as the determined speech recognition results and the determined time information; and a pronunciation error detection part that outputs a phoneme as a pronunciation error when reliability for each phoneme in the speech recognition results using the native speaker speech recognition model under a weakly constraining grammar is greater than the reliability of the corresponding phoneme in the speech recognition results using the native speaker acoustic model under a constraining grammar in which the determined speech recognition results are correct for the speech data in a segment specified by the determined time information.