Speech Similarity Analysis Through Pronunciation Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing language learning software struggles with analyzing speech similarity across multiple languages due to increased module volume and hardware requirements when adding similarity analysis functions for different languages.

Innovation Solution

Extract only the evaluation pronunciation feature corresponding to a standard pronunciation feature from the evaluation audio, using a trained encoder to reduce processing volume and enable similarity analysis across various languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If similarity analysis function is added for multiple languages, then language versatility is improved, but module volume and hardware requirements increase

Engineering Contradiction:
Improvelanguage versatilityVSAvoidmodule volume
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the speech analysis process into distinct feature extraction components that can be independently processed. By dividing the complex speech analysis into separate pronunciation feature extraction steps, the system can handle multiple languages without proportionally increasing overall module volume, as each language can use the same segmented processing pipeline with language-specific feature parameters.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a universal speech analysis module that can process multiple languages through a single integrated system. The analysis module uses language-agnostic speech recognition technology combined with language-specific pronunciation feature databases, allowing one module to serve multiple language analysis functions without requiring separate dedicated modules for each language.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If full speech analysis is performed, then measurement precision is improved, but computational complexity increases

Engineering Contradiction:
Improvespeech similarity accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential pronunciation features from speech signals that are necessary for similarity determination, rather than performing complete speech analysis. By taking out and focusing on specific pronunciation characteristics (such as phoneme sequences, stress patterns, and intonation features), the system achieves adequate measurement precision while significantly reducing computational complexity compared to full speech analysis.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by performing speech analysis to the extent necessary for pronunciation similarity determination, without executing complete speech processing pipelines. The system extracts and compares only the pronunciation-relevant features from speech signals, achieving sufficient accuracy for language learning purposes while avoiding the excessive computational burden of comprehensive speech analysis including semantics, syntax, and contextual interpretation.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4276827B1Speech similarity determination method, device and program product
Publication Date: 2025.10.22 LEMON INC(GB)
  • EP4276827B1 patent drawingFigure 1
  • EP4276827B1 patent drawingFigure 2
  • EP4276827B1 patent drawingFigure 3

AI summary

Embodiments provide a method and an apparatus for determining speech similarity, and a program product, which relate to speech technology. The method includes: playing exemplary audio, and acquiring evaluation audio of a user, where the exemplary audio is audio of specified content that is read by using a specified language; acquiring a standard pronunciation feature corresponding to the exemplary audio, and extracting, from the evaluation audio, an evaluation pronunciation feature corresponding to the standard pronunciation feature, where the standard pronunciation feature is used to reflect a specific pronunciation of the specified content in the specified language; and determining a feature difference between the standard pronunciation feature and the evaluation pronunciation feature, and determining similarity between the evaluation audio and the exemplary audio according to the feature difference. In the scheme of the present application, the evaluation pronunciation feature corresponding to the standard pronunciation feature corresponding to the exemplary audio can be extracted from the evaluation audio, thereby achieving relatively small volume of a module functioned with similarity analysis of follow-up reading.