Speech Recognition Phoneme Synthesis for Accented Conference Calls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current conference call systems face challenges in accurately understanding speakers with heavily accented speech, particularly when participants are geographically diverse and lack visual aids, leading to potential miscommunication due to limitations in speech recognition and synthesis technologies.

Innovation Solution

A method and system that captures and identifies phonemes in spoken words, accesses corresponding text, synthesizes a pronunciation, and substitutes it into the speech string, using a phoneme dictionary and text-to-speech synthesizer, allowing for improved understanding by all participants without interrupting the communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems use personalized phonemes to recognize accented speech, then recognition accuracy for individual speakers improves, but processing intensity and error rates increase due to variations in phoneme sounds when pronounced together

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing intensity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech recognition process into distinct phases: phoneme identification, text access, pronunciation synthesis, and substitution. By breaking down the complex task of recognizing accented speech into smaller, manageable components, the system reduces processing intensity at each stage while maintaining overall recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-storing pronunciations in a phoneme dictionary and pre-synthesizing pronunciations before substitution. This allows the speech recognition system to avoid real-time synthesis calculations, significantly reducing processing intensity during actual speech recognition while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If conference call participants use visual aids such as drawings or slides, then understanding of spoken language improves, but availability and accessibility to all participants decreases due to hardware requirements

Engineering Contradiction:
Improvelanguage understanding accuracyVSAvoidparticipant accessibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary system that translates accented speech into synthesized pronunciations based on phoneme dictionaries. This intermediary layer acts as a mediator between the speaker and listeners, converting difficult-to-understand accented speech into clear, standardized pronunciations that all participants can understand, regardless of their hardware capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces the need for visual aids (mechanical/optical system) with a speech processing system that uses phoneme identification and synthesis. Instead of requiring participants to view slides or drawings, the system substitutes these visual requirements with audio-based phoneme processing that works across all telephone devices.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If speech synthesizers emulate accented speech, then vocabulary coverage improves, but naturalness and acceptance by listeners decreases

Engineering Contradiction:
Improvevocabulary coverageVSAvoidlistener acceptance
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by selectively synthesizing only those pronunciations that are needed based on the phoneme dictionary, rather than emulating the entire accented speech pattern. This localized approach maintains naturalness for recognized words while providing accurate pronunciations for unrecognized words, balancing vocabulary coverage with listener acceptance.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8849666B2Conference call service with speech processing for heavily accented speakers
Publication Date: 2014.09.30 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8849666B2 patent drawing
  • US8849666B2 patent drawing
  • US8849666B2 patent drawing

AI summary

Speech recognition processing captures phonemes of words in a spoken speech string and retrieves text of words corresponding to particular combinations of phonemes from a phoneme dictionary. A text-to-speech synthesizer then can produce and substitute a synthesized pronunciation of that word in the speech string. If the speech recognition processing fails to recognize a particular combination of phonemes of a word, as spoken, as may occur when a word is spoken with an accent or when the speaker has a speech impediment, the speaker is prompted to clarify the word by entry, as text, from a keyboard or the like for storage in the phoneme dictionary such that a synthesized pronunciation of the word can be played out when the initially unrecognized spoken word is again encountered in a speech string to improve intelligibility, particularly for conference calls.