Prosodic Contour Selection for Natural Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis systems fail to produce natural-sounding speech due to the lack of human-like prosody, which affects the intelligibility and naturalness of synthesized speech.
Innovation Solution
A computer-implemented method and system that extracts prosodic contours from human speech utterances, generates a model to estimate distances between prosodic contours and text attributes, and selects the most appropriate prosodic contour for synthesized speech based on these distances, ensuring improved prosody and naturalness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional speech synthesis methods are used, then the system complexity is low, but the naturalness and intelligibility of synthesized speech deteriorates
Solution Approach 1:
The system pre-extracts prosodic contours from audio data and stores them in a database before synthesis operations. This preliminary preparation allows the synthesis system to quickly retrieve and select appropriate prosodic contours without performing complex real-time extraction, thereby improving naturalness while maintaining computational efficiency.
Solution Approach 2:
The patent introduces prosodic contours as an intermediary representation between raw audio and synthesized speech. These contours capture essential prosodic information (pitch, energy, duration) and serve as reusable templates that bridge the gap between text input and natural-sounding speech output, resolving the contradiction between simplicity and naturalness.
2Reliability
If multiple prosodic contours are stored and selected from a database, then the naturalness of synthesized speech is improved, but the processor usage and selection complexity increases
Solution Approach 1:
The system extracts and stores only the essential prosodic contour parameters (pitch, energy, duration) from complete audio signals. By separating and storing only these critical features in the database, the system reduces the amount of data that needs to be processed during synthesis while maintaining the ability to generate natural-sounding speech.
Solution Approach 2:
The patent transforms complex audio signals into simplified parametric representations (prosodic contours with discrete parameters). This parameter transformation enables efficient storage and quick comparison during selection, reducing processor usage while preserving the naturalness benefits of having multiple prosodic options.
3Adaptability or versatility
If prosodic contours are extracted and stored for every possible utterance, then the adaptability of the synthesis system is improved, but the storage requirements and data volume increases
Solution Approach 1:
The system creates prosodic contours that serve multiple functions: they can be used for exact matching when text is identical, for similarity-based selection when text is related, and as templates for generating prosody in novel utterances. This multi-functionality allows a smaller set of contours to support greater system adaptability, reducing the total data volume required.
Data Source
AI summary
The subject matter of this specification can be implemented in a computer-implemented method that includes receiving utterances and transcripts thereof. The method includes analyzing the utterances and transcripts to determine certain attributes, such as distances between prosodic contours for pairs of utterances. A model can be generated that can be used to estimate a distance between a determined prosodic contour for a received utterance and an unknown prosodic contour for a synthesized utterance when given a distance between attributes for text associated with the received utterance and the synthesized utterance.


