Automatic Speech Syllable Duration Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
When individuals speak a language other than their native tongue, mispronunciations due to differences in syllable duration can lead to communication difficulties, as existing speech processing technologies struggle to accurately correct these issues in real-time.
Innovation Solution
An automated telecommunication system that encodes speech using techniques like LPC and CELP, detects language and accent, and adjusts syllable duration and amplitude to correct mispronunciations by matching the pronunciation patterns of the native language, thereby improving communication understandability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech is processed digitally to correct pronunciation issues, then speech intelligibility is improved, but processing complexity increases
Solution Approach 1:
The patent segments speech into distinct syllables and identifies specific syllable duration parameters that cause mispronunciation. By dividing the speech signal into manageable units (syllables) and targeting only the problematic duration parameters rather than processing the entire speech signal, the system achieves effective pronunciation correction with reduced processing complexity
Solution Approach 2:
The patent modifies specific speech production parameters (syllable duration and amplitude) to correct mispronunciations. By changing only the problematic duration parameters of affected syllables rather than reprocessing the entire speech signal, the system achieves intelligibility improvement with minimal processing complexity
2Measurement precision
If real-time speech correction is implemented, then communication understandability is improved, but processing time increases
Solution Approach 1:
The patent extracts and modifies only the specific syllable duration parameters that cause mispronunciation, rather than processing the entire speech signal in real-time. This selective extraction and modification approach enables effective correction with minimal processing time
Solution Approach 2:
The patent uses pre-stored duration values for correctly pronounced syllables in the target language. When a mispronunciation is detected, the system retrieves the correct duration value from storage and applies it immediately, avoiding the need for complex real-time calculation and reducing processing time
Data Source
AI summary
A very common problem is when people speak a language other than the language which they are accustomed, syllables can be spoken for longer or shorter than the listener would regard as appropriate. An example of this can be observed when people who have a heavy Japanese accent speak English. Since Japanese words end with vowels, there is a tendency for native Japanese to add a vowel sound to the end of English words that should end with a consonant. Illustratively, native Japanese speakers often pronounce “orange” as “orenji.” An aspect provides an automatic speech-correcting process that would not necessarily need to know that fruit is being discussed; the system would only need to know that the speaker is accustomed to Japanese, that the listener is accustomed to English, that “orenji” is not a word in English, and that “orenji” is a typical Japanese mispronunciation of the English word “orange.”


