Speech Recognition Phoneme Synthesis for Accented Conference Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current conference call systems face challenges in accurately understanding speakers with heavily accented speech, particularly when participants are geographically diverse and lack visual aids, leading to potential miscommunication due to limitations in speech recognition and synthesis technologies.
Innovation Solution
A method and system that captures and identifies phonemes in spoken words, accesses corresponding text, synthesizes a pronunciation, and substitutes it into the speech string, using a phoneme dictionary and text-to-speech synthesizer, allowing for improved understanding by all participants without interrupting the communication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition systems use personalized phonemes to recognize accented speech, then recognition accuracy for individual speakers improves, but processing intensity and error rates increase due to variations in phoneme sounds when pronounced together
Solution Approach 1:
The patent segments the speech recognition process into distinct phases: phoneme identification, text access, pronunciation synthesis, and substitution. By breaking down the complex task of recognizing accented speech into smaller, manageable components, the system reduces processing intensity at each stage while maintaining overall recognition accuracy.
Solution Approach 2:
The system performs preliminary actions by pre-storing pronunciations in a phoneme dictionary and pre-synthesizing pronunciations before substitution. This allows the speech recognition system to avoid real-time synthesis calculations, significantly reducing processing intensity during actual speech recognition while maintaining accuracy.
2Measurement precision
If conference call participants use visual aids such as drawings or slides, then understanding of spoken language improves, but availability and accessibility to all participants decreases due to hardware requirements
Solution Approach 1:
The patent introduces an intermediary system that translates accented speech into synthesized pronunciations based on phoneme dictionaries. This intermediary layer acts as a mediator between the speaker and listeners, converting difficult-to-understand accented speech into clear, standardized pronunciations that all participants can understand, regardless of their hardware capabilities.
Solution Approach 2:
The system replaces the need for visual aids (mechanical/optical system) with a speech processing system that uses phoneme identification and synthesis. Instead of requiring participants to view slides or drawings, the system substitutes these visual requirements with audio-based phoneme processing that works across all telephone devices.
3Adaptability or versatility
If speech synthesizers emulate accented speech, then vocabulary coverage improves, but naturalness and acceptance by listeners decreases
Solution Approach 1:
The patent applies local quality by selectively synthesizing only those pronunciations that are needed based on the phoneme dictionary, rather than emulating the entire accented speech pattern. This localized approach maintains naturalness for recognized words while providing accurate pronunciations for unrecognized words, balancing vocabulary coverage with listener acceptance.
Data Source
AI summary
Speech recognition processing captures phonemes of words in a spoken speech string and retrieves text of words corresponding to particular combinations of phonemes from a phoneme dictionary. A text-to-speech synthesizer then can produce and substitute a synthesized pronunciation of that word in the speech string. If the speech recognition processing fails to recognize a particular combination of phonemes of a word, as spoken, as may occur when a word is spoken with an accent or when the speaker has a speech impediment, the speaker is prompted to clarify the word by entry, as text, from a keyboard or the like for storage in the phoneme dictionary such that a synthesized pronunciation of the word can be played out when the initially unrecognized spoken word is again encountered in a speech string to improve intelligibility, particularly for conference calls.


