Blending Recorded Speech with TTS Output for Domain-Specific Prosody
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Combining high-quality recorded speech with text-to-speech (TTS) synthesizer output in various applications, such as navigation and voice-activated systems, is time-consuming and difficult due to the need for manual selection and alignment of specific prompts.
Innovation Solution
A text-to-speech engine that identifies the domain of input text and selects domain-specific recorded speech to blend with TTS output, using a domain detector and blending unit to match and refine the prosody of static phrases with synthesized speech, thereby smoothing the acoustic trajectory.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual selection and alignment of recorded speech prompts is performed, then high-quality domain-specific speech output is achieved, but the process becomes time-consuming and difficult
Solution Approach 1:
The system automatically identifies the domain from input text and selects appropriate recorded speech prompts without manual intervention. The domain detector and blending unit work autonomously to match recorded speech with TTS output, eliminating the need for manual selection and alignment while maintaining high speech quality.
Solution Approach 2:
The system changes the parameter of domain identification from manual to automatic based on text analysis. By detecting the domain from input text characteristics and using this information to select recorded speech prompts, the system achieves both high quality and time efficiency.
2Manufacturing precision
If manual selection and alignment of recorded speech prompts is performed, then accurate domain-specific speech is achieved, but the operation becomes difficult and complex
Solution Approach 1:
The system performs self-service by automatically detecting the domain from input text and selecting the appropriate recorded speech prompts. The blending unit automatically aligns and combines recorded speech with TTS output, eliminating complex manual operations while maintaining accurate domain-specific speech generation.
Solution Approach 2:
The system segments the speech generation process into distinct components: domain detection, recorded speech selection, and blending with TTS output. This segmentation allows each component to operate autonomously and simplifies the overall operation while maintaining precision.
3Manufacturing precision
If recorded speech is combined with TTS synthesizer output, then high-quality domain-specific speech is produced, but the process becomes time-consuming
Solution Approach 1:
The system achieves self-service automation where the domain detector automatically identifies the domain from input text, selects appropriate recorded speech prompts, and the blending unit automatically combines them with TTS output. This eliminates manual processing steps while maintaining high speech quality, significantly improving productivity.
Solution Approach 2:
The system performs preliminary action by pre-identifying the domain from input text before selecting and blending recorded speech. This advance domain detection enables efficient selection of appropriate recorded prompts, reducing overall processing time while maintaining quality.
Data Source
AI summary
A text-to-speech (TTS) engine combines recorded speech with synthesized speech from a TTS synthesizer based on text input. The TTS engine receives the text input and identifies the domain for the speech (e.g. navigation, dialing, . . . ). The identified domain is used in selecting domain specific speech recordings (e.g. pre-recorded static phrases such as “turn left”, “turn right” . . . ) from the input text. The speech recordings are obtained based on the static phrases for the domain that are identified from the input text. The TTS engine blends the static phrases with the TTS output to smooth the acoustic trajectory of the input text. The prosody of the static phrases is used to create similar prosody in the TTS output.


