Text-to-Speech Synthesis Using Domain-Specific Prosody Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Text-to-Speech (TTS) synthesis systems in call centers and automated voice services fail to provide high-quality, natural-sounding speech, especially in complex transactions, due to limitations in voice talent recordings and lack of application-specific training, resulting in unsatisfactory voice quality and user experience.
Innovation Solution
A method and system that records and stores audio files of live voices speaking in specific domains using various prosodies, and trains a TTS system to select audio segments based on dialog states and speech acts, creating a domain-specific speech database for enhanced voice synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If conventional TTS systems use limited voice talent recordings, then the system complexity is reduced, but the voice quality and naturalness deteriorate
Solution Approach 1:
The patent segments the training data into domain-specific subsets (e.g., customer service, technical support) and records multiple prosodies for different dialog states and speech acts. This segmentation allows the system to achieve high voice quality for specific domains without requiring all possible voice talent recordings, thus managing system complexity while improving manufacturing precision.
Solution Approach 2:
The patent changes the parameter of prosody by recording the same text in multiple prosodic variations (different tones, speeds, emotions) for different dialog states and speech acts. This parameter change enables the TTS system to select the most appropriate prosody for each context, significantly improving voice quality and naturalness without proportionally increasing system complexity.
2Manufacturing precision
If TTS systems are trained with domain-specific speech data, then the voice quality and expressiveness improve, but the data collection and processing complexity increase
Solution Approach 1:
The patent creates a universal domain-specific speech database that can serve multiple functions: training the TTS system, providing prosodic variations for different dialog states, and enabling selective audio segment retrieval. This multi-functionality justifies the increased data collection complexity by delivering multiple benefits from a single comprehensive database.
Solution Approach 2:
The patent performs preliminary action by pre-recording and pre-processing domain-specific speech data with multiple prosodies before the TTS system needs to generate responses. This preliminary data preparation simplifies the online processing requirements and allows the system to focus on selecting and synthesizing appropriate segments rather than processing raw data in real-time.
3Productivity
If the TTS system selects audio segments based on multiple dialog states and speech acts, then the responsiveness and user satisfaction improve, but the selection process complexity increases
Solution Approach 1:
The patent implements feedback mechanisms where the TTS system receives information about the current dialog state and speech act from the dialog management system, uses this feedback to select appropriate audio segments from the trained database, and generates responses accordingly. This feedback loop enables high responsiveness by ensuring the selected segments match the current context, while the systematic approach to selection manages complexity.
Data Source
AI summary
A system, method and computer readable medium that trains a text-to-speech synthesis system for use in speech synthesis is disclosed. The method may include recording audio files of one or more live voices speaking language used in a specific domain, the audio files being recorded using various prosodies, storing the recorded audio files in a speech database; and training a text-to-speech synthesis system using the speech database, wherein the text-to-speech synthesis system selects audio selects audio segments having a prosody based on at least one dialog state and one speech act.


