Automatic Pronunciation Scoring via Acoustic Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional human evaluation of non-native speech pronunciation in foreign language learning is inconsistent and time-consuming, and existing automatic scoring systems face challenges due to immature speech recognition technology and difficulty in extracting effective acoustic features.
Innovation Solution
An automatic evaluation system that employs linguistically verified acoustic features to detect prosody and fluency errors, using speech recognition, text-to-phone conversion, feature extraction, and score integration, with specific formulas for rhythm and fluency features to provide a consistent and reliable pronunciation accuracy score.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human raters are used for pronunciation evaluation, then evaluation quality can be maintained through linguistic knowledge and intuition, but inter- and intra-rater inconsistencies occur and evaluation is time-consuming
Solution Approach 1:
The patent creates a copy of human rating behavior by training an automatic scoring system on human-rated speech data. The system learns to replicate human raters' evaluation patterns and linguistic intuition through machine learning, enabling automated pronunciation assessment that mimics human judgment while eliminating time consumption and inconsistency issues.
Solution Approach 2:
The patent replaces the mechanical human rating process with an automated computational system. Instead of relying on human raters' linguistic knowledge and intuition, the system uses acoustic feature extraction and machine learning algorithms to automatically evaluate pronunciation, substituting human cognitive processes with computational mechanisms.
2Measurement precision
If human raters are hired for large-scale evaluation, then evaluation quality can be maintained, but costs increase significantly
Solution Approach 1:
The patent enables the evaluation system to serve itself by automatically processing and scoring speech samples without requiring human rater intervention. The automated system performs all evaluation tasks independently, eliminating the need to hire and train human raters while maintaining consistent quality across large numbers of test-takers.
Solution Approach 2:
The patent changes the fundamental parameters of the evaluation system from human-based to machine-based. By transitioning from human cognitive evaluation to automated acoustic analysis, the system dramatically reduces costs while maintaining evaluation quality through consistent application of learned scoring criteria.
3Productivity
If speech recognition technology is used for automatic scoring, then time and cost are reduced, but incorrect pronunciation of non-native speakers causes many recognition errors
Solution Approach 1:
The patent extracts and focuses specifically on pronunciation-related acoustic features rather than attempting full speech recognition. By isolating and analyzing only the relevant phonetic and prosodic characteristics, the system avoids the pitfalls of general speech recognition while maintaining high accuracy in pronunciation assessment.
Solution Approach 2:
The patent introduces acoustic feature extraction as an intermediary between speech input and scoring output. Instead of directly using speech recognition technology that fails on non-native pronunciation, the system first transforms speech into acoustic features that capture pronunciation characteristics, then uses these features for accurate scoring.
4Extent of automation
If features simulating human rating rubric are created, then automatic scoring can be implemented, but effective acoustic feature extraction remains difficult due to qualitative vs quantitative mismatch
Solution Approach 1:
The patent transforms qualitative human rating criteria into quantitative acoustic measurements. By identifying specific acoustic parameters that correspond to pronunciation quality dimensions, the system converts subjective human evaluation concepts into objective, measurable features that can be automatically processed and scored.
Solution Approach 2:
The patent discards the limitations of direct speech recognition and recovers effective scoring capability through acoustic feature extraction. By abandoning reliance on text-based recognition and instead focusing on acoustic properties, the system recovers the ability to accurately assess pronunciation while enabling full automation.
Data Source
AI summary
Provided herein is a method of automatically evaluating non-native speakers' pronunciation proficiency, wherein an input speech signal is received, the speech recognition module transcribes it into a text; the text is converted into a segmental string with temporal information for each phone; in the key module of feature extraction, nine features of speech rhythm and fluency are computed; and then all feature values are integrated into an overall pronunciation proficiency score. This method can be applied to both types of non-native speech input: read-aloud speech and spontaneous speech.

