Speech Similarity Analysis Through Pronunciation Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing language learning software struggles with analyzing speech similarity across multiple languages due to increased module volume and hardware requirements when adding similarity analysis functions for different languages.
Innovation Solution
Extract only the evaluation pronunciation feature corresponding to a standard pronunciation feature from the evaluation audio, using a trained encoder to reduce processing volume and enable similarity analysis across various languages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If similarity analysis function is added for multiple languages, then language versatility is improved, but module volume and hardware requirements increase
Solution Approach 1:
The patent segments the speech analysis process into distinct feature extraction components that can be independently processed. By dividing the complex speech analysis into separate pronunciation feature extraction steps, the system can handle multiple languages without proportionally increasing overall module volume, as each language can use the same segmented processing pipeline with language-specific feature parameters.
Solution Approach 2:
The patent implements a universal speech analysis module that can process multiple languages through a single integrated system. The analysis module uses language-agnostic speech recognition technology combined with language-specific pronunciation feature databases, allowing one module to serve multiple language analysis functions without requiring separate dedicated modules for each language.
2Measurement precision
If full speech analysis is performed, then measurement precision is improved, but computational complexity increases
Solution Approach 1:
The patent extracts only the essential pronunciation features from speech signals that are necessary for similarity determination, rather than performing complete speech analysis. By taking out and focusing on specific pronunciation characteristics (such as phoneme sequences, stress patterns, and intonation features), the system achieves adequate measurement precision while significantly reducing computational complexity compared to full speech analysis.
Solution Approach 2:
The patent applies partial action by performing speech analysis to the extent necessary for pronunciation similarity determination, without executing complete speech processing pipelines. The system extracts and compares only the pronunciation-relevant features from speech signals, achieving sufficient accuracy for language learning purposes while avoiding the excessive computational burden of comprehensive speech analysis including semantics, syntax, and contextual interpretation.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Embodiments provide a method and an apparatus for determining speech similarity, and a program product, which relate to speech technology. The method includes: playing exemplary audio, and acquiring evaluation audio of a user, where the exemplary audio is audio of specified content that is read by using a specified language; acquiring a standard pronunciation feature corresponding to the exemplary audio, and extracting, from the evaluation audio, an evaluation pronunciation feature corresponding to the standard pronunciation feature, where the standard pronunciation feature is used to reflect a specific pronunciation of the specified content in the specified language; and determining a feature difference between the standard pronunciation feature and the evaluation pronunciation feature, and determining similarity between the evaluation audio and the exemplary audio according to the feature difference. In the scheme of the present application, the evaluation pronunciation feature corresponding to the standard pronunciation feature corresponding to the exemplary audio can be extracted from the evaluation audio, thereby achieving relatively small volume of a module functioned with similarity analysis of follow-up reading.