Text-to-Speech Synthesis Using Spoken Example Prosodic Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional text-to-speech systems struggle to produce natural-sounding speech due to the lack of effective prosodic control, requiring manual and tedious processes for users to achieve desired pitch and duration, leading to unnatural and monotone speech outputs.
Innovation Solution
A system and method that automatically determines prosodic parameters from a spoken utterance, generates marked-up text using these parameters, and synthesizes a waveform to mimic the spoken input's style and pronunciation, including pitch contour, duration, and energy information, allowing for direct specification of prosodic attributes in markup elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional TTS systems use flat pitch contours with constant pitch values, then the system operation is simple, but the synthesized speech sounds unnatural and monotone
Solution Approach 1:
The system performs preliminary action by automatically analyzing the reference spoken utterance to extract prosodic parameters (pitch contours, duration, energy) before synthesis. This pre-extraction of prosodic information enables the TTS system to generate natural-sounding speech without requiring manual intervention, thus resolving the contradiction between operational simplicity and speech naturalness
Solution Approach 2:
The system copies the prosodic characteristics from the reference spoken utterance by extracting pitch contours, duration patterns, and energy information from the reference audio. These extracted parameters are then applied to the synthesized speech output, allowing the system to replicate natural speech patterns automatically while maintaining simple operation
2Manufacturing precision
If users manually generate marked-up text to control pitch and duration, then fine control of prosodic attributes is achieved, but the process becomes very burdensome and tedious
Solution Approach 1:
The system performs self-service by automatically extracting prosodic parameters from the reference spoken utterance without requiring user intervention. The automated extraction process eliminates the need for users to manually create marked-up text, thereby achieving precise prosodic control while eliminating the time-consuming manual effort
Solution Approach 2:
The system replaces the mechanical process of manual text markup with an automated computational extraction process. By using speech processing algorithms to automatically extract pitch, duration, and energy parameters from audio, the system substitutes manual user actions with automated mechanical processing, achieving both precision and efficiency
3Reliability
If conventional TTS systems process annotated text inputs with marked-up text, then more fluent and human-like speech is produced, but the system complexity increases
Solution Approach 1:
The system takes out the essential prosodic information from the reference spoken utterance by extracting pitch contours, duration patterns, and energy parameters. This extraction of key features simplifies the input to the synthesis process, achieving fluent and human-like speech without requiring complex annotated text processing systems
Solution Approach 2:
The system changes the input format from complex annotated text to extracted prosodic parameters (pitch, duration, energy contours). This parameter transformation simplifies the system architecture by directly using acoustic parameters from the reference audio, thereby achieving human-like speech with reduced system complexity
Data Source
AI summary
Systems and methods for speech synthesis and, in particular, text-to-speech systems and methods for converting a text input to a synthetic waveform by processing prosodic and phonetic content of a spoken example of the text input to accurately mimic the input speech style and pronunciation. Systems and methods provide an interface to a TTS system to allow a user to input a text string and a spoken utterance of the text string, extract prosodic parameters from the spoken input, and process the prosodic parameters to derive corresponding markup for the text input to enable a more natural sounding synthesized speech.


