Multi-Style Speech Synthesis Using Universal Database

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Text-to-Speech (TTS) systems face challenges in rendering text as speech in multiple styles efficiently, as creating databases for each style is expensive and time-consuming, and requires significant storage space, while also struggling with acoustic and prosodic characteristics overlap between styles.

Innovation Solution

The method involves using speech segments from different styles to generate speech in a specified style by identifying and matching acoustic and prosodic characteristics, allowing for multi-style synthesis by training the TTS system to estimate similarity between speech segments and apply transformations to match desired styles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If separate speech databases are created for each speaking style, then speech quality and style accuracy are improved, but storage space requirements and system complexity increase significantly

Engineering Contradiction:
Improvespeech style accuracyVSAvoidstorage space
Core Design Contradiction:
Manufacturing precisionVSVolume of stationary object

Solution Approach 1:

The patent implements a universal speech database that stores speech segments with multiple style labels, allowing a single database to serve multiple speaking styles. The system can select and adapt segments from this unified database based on the target style, eliminating the need for separate databases for each style while maintaining speech quality through style-specific segment selection and acoustic parameter adjustment.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes acoustic and prosodic parameters of speech segments to match the target speaking style. By adjusting parameters such as pitch, duration, and energy, the system can adapt segments from one style to another, reducing the need for style-specific databases while preserving style accuracy through parameter transformation.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If separate speech databases are created for each speaking style, then style-specific speech quality is improved, but the time and cost required to create and maintain these databases increase

Engineering Contradiction:
Improvespeech style qualityVSAvoiddatabase creation and maintenance time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

A single universal database serves all speaking styles, eliminating the need to create and maintain separate databases for each style. The database stores speech segments annotated with multiple style characteristics, allowing the system to efficiently select and adapt segments for any target style without additional database creation or maintenance overhead.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

Speech segments are pre-annotated with multiple style labels and acoustic characteristics during a single database creation process. This preliminary annotation enables rapid style adaptation during synthesis by simply selecting and adjusting pre-labeled segments, rather than creating style-specific databases whenever new styles are needed.

Inventive Principle:
Principle #10Preliminary action

3Volume of stationary object

If speech segments from different styles are used for synthesis, then storage requirements are reduced, but the difficulty of matching acoustic and prosodic characteristics increases

Engineering Contradiction:
Improvestorage spaceVSAvoidacoustic and prosodic matching complexity
Core Design Contradiction:
Volume of stationary objectVSDifficulty of detecting and measuring

Solution Approach 1:

The system employs feedback mechanisms where acoustic and prosodic characteristics of selected speech segments are compared against the target style profile, and adjustments are made accordingly. This iterative matching process ensures that segments from different styles can be effectively adapted to the target style by measuring and adjusting acoustic parameters such as pitch, duration, and spectral characteristics.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system systematically adjusts acoustic parameters (pitch, duration, energy, spectral characteristics) of speech segments to match the target style. By applying parameter transformations based on the difference between the source and target style characteristics, the system can effectively adapt segments from any style while reducing storage requirements through the universal database approach.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9990915B2Systems and methods for multi-style speech synthesis
Publication Date: 2018.06.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9990915B2 patent drawing
  • US9990915B2 patent drawing
  • US9990915B2 patent drawing

AI summary

Techniques for performing multi-style speech synthesis. The techniques include using at least one computer hardware processor to perform: obtaining input comprising text and an identification of a desired speaking style to use in rendering the text as speech; identifying a plurality of speech segments for use in synthesizing the text as speech, the identifying comprising identifying a first speech segment recorded and/or synthesized in a first speaking style that is different from the desired speaking style based at least in part on a measure of similarity between the desired speaking style and the first speaking style; synthesizing speech from the text in the desired speaking style at least in part by using the first speech segment; and outputting the synthesized speech.