Text to Speech Accent Adaptation via Region-Specific Phoneme Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Speech to Text (STT) and Text to Speech (TTS) systems require separate, lengthy training processes and struggle to capture user pronunciations of domain terminology effectively, leading to a one-size-fits-all approach in TTS systems, which can result in unfamiliar pronunciation patterns for users with varying dialects and accents.

Innovation Solution

The system determines region-specific phoneme sequences for domain terminology from evaluated STT data, adapting TTS outputs to match the user's accent by using a region-specific pronunciation dictionary, allowing for customized text-to-speech responses that reflect the user's local pronunciation differences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate training processes are used for STT and TTS systems, then each system can be optimized independently, but the training time and complexity increase significantly

Engineering Contradiction:
Improvesystem optimizationVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent combines STT and TTS training processes into a unified framework where phoneme sequence data from STT training directly informs TTS training. This integration allows both systems to be optimized together rather than separately, reducing total training time while maintaining or improving optimization quality through shared phoneme representations.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of manufacture

If default phoneme sequences are used in TTS systems, then the system is simpler to implement, but it cannot adapt to users with varying dialects and accents

Engineering Contradiction:
Improvesystem implementationVSAvoidaccent adaptation
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic phoneme sequence selection in TTS systems. Instead of using fixed default phoneme sequences, the system dynamically selects phoneme sequences based on the user's detected accent or dialect, allowing the TTS output to adapt to different users while maintaining a unified system architecture.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality by using different phoneme sequences for different regions or accent groups. The system maintains multiple phoneme sequence variants and selects the appropriate local variant based on user characteristics, allowing customization for specific dialects without redesigning the entire system.

Inventive Principle:
Principle #3Local quality

3Ease of operation

If region-specific phoneme sequences are implemented, then user familiarity and comfort improve, but the system complexity and data requirements increase

Engineering Contradiction:
Improveuser familiarityVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-processing and organizing phoneme sequence data according to different regions and accents during the training phase. This pre-organization allows the TTS system to quickly select appropriate phoneme sequences at runtime without complex real-time analysis, reducing operational complexity while maintaining user familiarity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11699430B2Using speech to text data in training text to speech models
Publication Date: 2023.07.11 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11699430B2 patent drawing
  • US11699430B2 patent drawing
  • US11699430B2 patent drawing

AI summary

A system and method for providing a text to speech output by receiving user audio data, determining a user region-specific-pronunciation classification according to the audio data, determining text for a response to the user according to the audio data, identifying a portion from the text, where a region specific-pronunciation dictionary includes the portion, and using a phoneme string, from the dictionary selected according to the user region-specific pronunciation classification, for the word in a text to speech output to the user.