Text-to-Speech Synthesis Using Spoken Example Prosodic Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-speech systems struggle to produce natural-sounding speech due to the lack of effective prosodic control, requiring manual and tedious processes for users to achieve desired pitch and duration, leading to unnatural and monotone speech outputs.

Innovation Solution

A system and method that automatically determines prosodic parameters from a spoken utterance, generates marked-up text using these parameters, and synthesizes a waveform to mimic the spoken input's style and pronunciation, including pitch contour, duration, and energy information, allowing for direct specification of prosodic attributes in markup elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional TTS systems use flat pitch contours with constant pitch values, then the system operation is simple, but the synthesized speech sounds unnatural and monotone

Engineering Contradiction:
Improvesystem operation simplicityVSAvoidnaturalness of synthesized speech
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system performs preliminary action by automatically analyzing the reference spoken utterance to extract prosodic parameters (pitch contours, duration, energy) before synthesis. This pre-extraction of prosodic information enables the TTS system to generate natural-sounding speech without requiring manual intervention, thus resolving the contradiction between operational simplicity and speech naturalness

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system copies the prosodic characteristics from the reference spoken utterance by extracting pitch contours, duration patterns, and energy information from the reference audio. These extracted parameters are then applied to the synthesized speech output, allowing the system to replicate natural speech patterns automatically while maintaining simple operation

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If users manually generate marked-up text to control pitch and duration, then fine control of prosodic attributes is achieved, but the process becomes very burdensome and tedious

Engineering Contradiction:
Improvecontrol precision of pitch and durationVSAvoiduser effort and time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically extracting prosodic parameters from the reference spoken utterance without requiring user intervention. The automated extraction process eliminates the need for users to manually create marked-up text, thereby achieving precise prosodic control while eliminating the time-consuming manual effort

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces the mechanical process of manual text markup with an automated computational extraction process. By using speech processing algorithms to automatically extract pitch, duration, and energy parameters from audio, the system substitutes manual user actions with automated mechanical processing, achieving both precision and efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If conventional TTS systems process annotated text inputs with marked-up text, then more fluent and human-like speech is produced, but the system complexity increases

Engineering Contradiction:
Improvefluency and human-like pronunciationVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system takes out the essential prosodic information from the reference spoken utterance by extracting pitch contours, duration patterns, and energy parameters. This extraction of key features simplifies the input to the synthesis process, achieving fluent and human-like speech without requiring complex annotated text processing systems

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the input format from complex annotated text to extracted prosodic parameters (pitch, duration, energy contours). This parameter transformation simplifies the system architecture by directly using acoustic parameters from the reference audio, thereby achieving human-like speech with reduced system complexity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8886538B2Systems and methods for text-to-speech synthesis using spoken example
Publication Date: 2014.11.11 CERENCE OPERATING CO
  • US8886538B2 patent drawing
  • US8886538B2 patent drawing
  • US8886538B2 patent drawing

AI summary

Systems and methods for speech synthesis and, in particular, text-to-speech systems and methods for converting a text input to a synthetic waveform by processing prosodic and phonetic content of a spoken example of the text input to accurately mimic the input speech style and pronunciation. Systems and methods provide an interface to a TTS system to allow a user to input a text string and a spoken utterance of the text string, extract prosodic parameters from the spoken input, and process the prosodic parameters to derive corresponding markup for the text input to enable a more natural sounding synthesized speech.