Text-To-Speech Phonetic Transcription Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Text-To-Speech systems based on concatenative technology face challenges in generating natural-sounding synthetic speech due to mismatches between speaker-independent Front-End phonetic transcriptions and speaker-specific pronunciation styles, leading to degraded signal quality and increased processing requirements.

Innovation Solution

A Text-To-Speech system that generates multiple phonetic transcriptions for each word and uses a cost function to select the most appropriate ones based on dynamic programming, ensuring better matching with the speaker's pronunciation style, thereby reducing signal processing and improving quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speaker-independent statistical models or rules are used for phonetic transcription, then the Front-End can generate phonetic forms, but the phonetic forms do not match the speaker's pronunciation style, degrading output signal quality

Engineering Contradiction:
ImproveFront-End phonetic transcription capabilityVSAvoidphonetic form matching accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent changes the parameter of phonetic transcription from a single fixed form to multiple alternative forms with associated costs. The TTS engine evaluates different phonetic transcriptions using a cost function that considers pronunciation style matching, and selects the transcription with the lowest cost, thereby improving phonetic form matching accuracy while maintaining Front-End operation capability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system introduces dynamic selection of phonetic transcriptions based on speaker-specific characteristics. Instead of using a static speaker-independent model, the system dynamically adapts to the speaker's pronunciation style by evaluating multiple phonetic forms and selecting the most appropriate one, resolving the contradiction between operational ease and matching precision

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If manual adaptation of rules by expert linguists is performed for each new voice, then speaker-specific pronunciation styles can be accommodated, but the process becomes very time consuming

Engineering Contradiction:
Improvespeaker-specific pronunciation matchingVSAvoidtime for rule adaptation
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system enables automatic adaptation to speaker-specific pronunciation styles without requiring manual intervention by expert linguists. The TTS engine uses a cost function to automatically evaluate and select phonetic transcriptions that best match the speaker's style, making the system self-adapting and eliminating time-consuming manual rule adaptation while maintaining high phonetic matching accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual rule adaptation by linguists with an automated computational system. The cost function and dynamic programming algorithm automatically perform the adaptation task, substituting human expert labor with an efficient computational mechanism that achieves the same goal without time loss

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Manufacturing precision

If a statistical Front-End dedicated to the speaker is trained, then speaker-specific phonetic forms can be generated, but the training process is also time consuming

Engineering Contradiction:
Improvespeaker-specific phonetic transcription accuracyVSAvoidtraining time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

Instead of performing full statistical training of a speaker-specific Front-End model, the system uses a partial approach by generating multiple phonetic transcriptions and selecting the best one using a cost function. This partial action achieves speaker-specific phonetic accuracy without the excessive time investment required for complete statistical training

Inventive Principle:
Principle #16Partial or excessive action

4Productivity

If speaker-independent Front-End systems force pronunciations, then phonetic transcriptions can be generated efficiently, but the pronunciations are not natural for the recorded speakers, negatively impacting final signal quality

Engineering Contradiction:
Improvephonetic transcription generation speedVSAvoidnaturalness of pronunciation
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system maintains the efficiency of speaker-independent Front-End by allowing it to generate phonetic transcriptions quickly, then adds a dynamic selection layer that evaluates multiple transcriptions using a cost function. This dynamic approach selects pronunciations that are natural for the recorded speakers, resolving the contradiction between productivity and pronunciation naturalness

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS7869999B2Systems and methods for selecting from multiple phonectic transcriptions for text-to-speech synthesis
Publication Date: 2011.01.11 CERENCE OPERATING CO
  • US7869999B2 patent drawing
  • US7869999B2 patent drawing
  • US7869999B2 patent drawing

AI summary

A system and method for generating synthetic speech, which operates in a computer implemented Text-To-Speech system. The system comprises at least a speaker database that has been previously created from user recordings, a Front-End system to receive an input text and a Text-To-Speech engine. The Front-End system generates multiple phonetic transcriptions for each word of the input text, and the TTS engine uses a cost function to select which phonetic transcription is the more appropriate for searching the speech segments within the speaker database to be concatenated and synthesized.