Automatic Pronunciation Scoring via Acoustic Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional human evaluation of non-native speech pronunciation in foreign language learning is inconsistent and time-consuming, and existing automatic scoring systems face challenges due to immature speech recognition technology and difficulty in extracting effective acoustic features.

Innovation Solution

An automatic evaluation system that employs linguistically verified acoustic features to detect prosody and fluency errors, using speech recognition, text-to-phone conversion, feature extraction, and score integration, with specific formulas for rhythm and fluency features to provide a consistent and reliable pronunciation accuracy score.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human raters are used for pronunciation evaluation, then evaluation quality can be maintained through linguistic knowledge and intuition, but inter- and intra-rater inconsistencies occur and evaluation is time-consuming

Engineering Contradiction:
Improvepronunciation evaluation qualityVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates a copy of human rating behavior by training an automatic scoring system on human-rated speech data. The system learns to replicate human raters' evaluation patterns and linguistic intuition through machine learning, enabling automated pronunciation assessment that mimics human judgment while eliminating time consumption and inconsistency issues.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical human rating process with an automated computational system. Instead of relying on human raters' linguistic knowledge and intuition, the system uses acoustic feature extraction and machine learning algorithms to automatically evaluate pronunciation, substituting human cognitive processes with computational mechanisms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If human raters are hired for large-scale evaluation, then evaluation quality can be maintained, but costs increase significantly

Engineering Contradiction:
Improveevaluation qualityVSAvoidevaluation cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent enables the evaluation system to serve itself by automatically processing and scoring speech samples without requiring human rater intervention. The automated system performs all evaluation tasks independently, eliminating the need to hire and train human raters while maintaining consistent quality across large numbers of test-takers.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the fundamental parameters of the evaluation system from human-based to machine-based. By transitioning from human cognitive evaluation to automated acoustic analysis, the system dramatically reduces costs while maintaining evaluation quality through consistent application of learned scoring criteria.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If speech recognition technology is used for automatic scoring, then time and cost are reduced, but incorrect pronunciation of non-native speakers causes many recognition errors

Engineering Contradiction:
Improvescoring efficiencyVSAvoidspeech recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent extracts and focuses specifically on pronunciation-related acoustic features rather than attempting full speech recognition. By isolating and analyzing only the relevant phonetic and prosodic characteristics, the system avoids the pitfalls of general speech recognition while maintaining high accuracy in pronunciation assessment.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces acoustic feature extraction as an intermediary between speech input and scoring output. Instead of directly using speech recognition technology that fails on non-native pronunciation, the system first transforms speech into acoustic features that capture pronunciation characteristics, then uses these features for accurate scoring.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Extent of automation

If features simulating human rating rubric are created, then automatic scoring can be implemented, but effective acoustic feature extraction remains difficult due to qualitative vs quantitative mismatch

Engineering Contradiction:
Improveautomatic scoring capabilityVSAvoidacoustic feature extraction difficulty
Core Design Contradiction:
Extent of automationVSDifficulty of detecting and measuring

Solution Approach 1:

The patent transforms qualitative human rating criteria into quantitative acoustic measurements. By identifying specific acoustic parameters that correspond to pronunciation quality dimensions, the system converts subjective human evaluation concepts into objective, measurable features that can be automatically processed and scored.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent discards the limitations of direct speech recognition and recovers effective scoring capability through acoustic feature extraction. By abandoning reliance on text-based recognition and instead focusing on acoustic properties, the system recovers the ability to accurately assess pronunciation while enabling full automation.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS10354645B2Method for automatic evaluation of non-native pronunciation
Publication Date: 2019.07.16 HANKUK UNIV OF FOREIGN STUDIES RES & BUSINESS FOUND
  • US10354645B2 patent drawing
  • US10354645B2 patent drawing

AI summary

Provided herein is a method of automatically evaluating non-native speakers' pronunciation proficiency, wherein an input speech signal is received, the speech recognition module transcribes it into a text; the text is converted into a segmental string with temporal information for each phone; in the key module of feature extraction, nine features of speech rhythm and fluency are computed; and then all feature values are integrated into an overall pronunciation proficiency score. This method can be applied to both types of non-native speech input: read-aloud speech and spontaneous speech.