Text-To-Speech Phonetic Transcription Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Text-To-Speech systems based on concatenative technology face challenges in generating natural-sounding synthetic speech due to mismatches between speaker-independent Front-End phonetic transcriptions and speaker-specific pronunciation styles, leading to degraded signal quality and increased processing requirements.
Innovation Solution
A Text-To-Speech system that generates multiple phonetic transcriptions for each word and uses a cost function to select the most appropriate ones based on dynamic programming, ensuring better matching with the speaker's pronunciation style, thereby reducing signal processing and improving quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speaker-independent statistical models or rules are used for phonetic transcription, then the Front-End can generate phonetic forms, but the phonetic forms do not match the speaker's pronunciation style, degrading output signal quality
Solution Approach 1:
The patent changes the parameter of phonetic transcription from a single fixed form to multiple alternative forms with associated costs. The TTS engine evaluates different phonetic transcriptions using a cost function that considers pronunciation style matching, and selects the transcription with the lowest cost, thereby improving phonetic form matching accuracy while maintaining Front-End operation capability
Solution Approach 2:
The system introduces dynamic selection of phonetic transcriptions based on speaker-specific characteristics. Instead of using a static speaker-independent model, the system dynamically adapts to the speaker's pronunciation style by evaluating multiple phonetic forms and selecting the most appropriate one, resolving the contradiction between operational ease and matching precision
2Manufacturing precision
If manual adaptation of rules by expert linguists is performed for each new voice, then speaker-specific pronunciation styles can be accommodated, but the process becomes very time consuming
Solution Approach 1:
The system enables automatic adaptation to speaker-specific pronunciation styles without requiring manual intervention by expert linguists. The TTS engine uses a cost function to automatically evaluate and select phonetic transcriptions that best match the speaker's style, making the system self-adapting and eliminating time-consuming manual rule adaptation while maintaining high phonetic matching accuracy
Solution Approach 2:
The patent replaces the mechanical process of manual rule adaptation by linguists with an automated computational system. The cost function and dynamic programming algorithm automatically perform the adaptation task, substituting human expert labor with an efficient computational mechanism that achieves the same goal without time loss
3Manufacturing precision
If a statistical Front-End dedicated to the speaker is trained, then speaker-specific phonetic forms can be generated, but the training process is also time consuming
Solution Approach 1:
Instead of performing full statistical training of a speaker-specific Front-End model, the system uses a partial approach by generating multiple phonetic transcriptions and selecting the best one using a cost function. This partial action achieves speaker-specific phonetic accuracy without the excessive time investment required for complete statistical training
4Productivity
If speaker-independent Front-End systems force pronunciations, then phonetic transcriptions can be generated efficiently, but the pronunciations are not natural for the recorded speakers, negatively impacting final signal quality
Solution Approach 1:
The system maintains the efficiency of speaker-independent Front-End by allowing it to generate phonetic transcriptions quickly, then adds a dynamic selection layer that evaluates multiple transcriptions using a cost function. This dynamic approach selects pronunciations that are natural for the recorded speakers, resolving the contradiction between productivity and pronunciation naturalness
Data Source
AI summary
A system and method for generating synthetic speech, which operates in a computer implemented Text-To-Speech system. The system comprises at least a speaker database that has been previously created from user recordings, a Front-End system to receive an input text and a Text-To-Speech engine. The Front-End system generates multiple phonetic transcriptions for each word of the input text, and the TTS engine uses a cost function to select which phonetic transcription is the more appropriate for searching the speech segments within the speaker database to be concatenated and synthesized.


