Prosodic Contour Selection for Natural Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis systems fail to produce natural-sounding speech due to the lack of human-like prosody, which affects the intelligibility and naturalness of synthesized speech.

Innovation Solution

A computer-implemented method and system that extracts prosodic contours from human speech utterances, generates a model to estimate distances between prosodic contours and text attributes, and selects the most appropriate prosodic contour for synthesized speech based on these distances, ensuring improved prosody and naturalness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speech synthesis methods are used, then the system complexity is low, but the naturalness and intelligibility of synthesized speech deteriorates

Engineering Contradiction:
Improvenaturalness of synthesized speechVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system pre-extracts prosodic contours from audio data and stores them in a database before synthesis operations. This preliminary preparation allows the synthesis system to quickly retrieve and select appropriate prosodic contours without performing complex real-time extraction, thereby improving naturalness while maintaining computational efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces prosodic contours as an intermediary representation between raw audio and synthesized speech. These contours capture essential prosodic information (pitch, energy, duration) and serve as reusable templates that bridge the gap between text input and natural-sounding speech output, resolving the contradiction between simplicity and naturalness.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If multiple prosodic contours are stored and selected from a database, then the naturalness of synthesized speech is improved, but the processor usage and selection complexity increases

Engineering Contradiction:
Improvenaturalness of synthesized speechVSAvoidprocessor usage
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system extracts and stores only the essential prosodic contour parameters (pitch, energy, duration) from complete audio signals. By separating and storing only these critical features in the database, the system reduces the amount of data that needs to be processed during synthesis while maintaining the ability to generate natural-sounding speech.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms complex audio signals into simplified parametric representations (prosodic contours with discrete parameters). This parameter transformation enables efficient storage and quick comparison during selection, reducing processor usage while preserving the naturalness benefits of having multiple prosodic options.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If prosodic contours are extracted and stored for every possible utterance, then the adaptability of the synthesis system is improved, but the storage requirements and data volume increases

Engineering Contradiction:
Improveadaptability of synthesis systemVSAvoiddata volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system creates prosodic contours that serve multiple functions: they can be used for exact matching when text is identical, for similarity-based selection when text is related, and as templates for generating prosody in novel utterances. This multi-functionality allows a smaller set of contours to support greater system adaptability, reducing the total data volume required.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9093067B1Generating prosodic contours for synthesized speech
Publication Date: 2015.07.28 GOOGLE LLC
  • US9093067B1 patent drawing
  • US9093067B1 patent drawing
  • US9093067B1 patent drawing

AI summary

The subject matter of this specification can be implemented in a computer-implemented method that includes receiving utterances and transcripts thereof. The method includes analyzing the utterances and transcripts to determine certain attributes, such as distances between prosodic contours for pairs of utterances. A model can be generated that can be used to estimate a distance between a determined prosodic contour for a received utterance and an unknown prosodic contour for a synthesized utterance when given a distance between attributes for text associated with the received utterance and the synthesized utterance.