Multi-Style Speech Synthesis Using Universal Database
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Text-to-Speech (TTS) systems face challenges in rendering text as speech in multiple styles efficiently, as creating databases for each style is expensive and time-consuming, and requires significant storage space, while also struggling with acoustic and prosodic characteristics overlap between styles.
Innovation Solution
The method involves using speech segments from different styles to generate speech in a specified style by identifying and matching acoustic and prosodic characteristics, allowing for multi-style synthesis by training the TTS system to estimate similarity between speech segments and apply transformations to match desired styles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If separate speech databases are created for each speaking style, then speech quality and style accuracy are improved, but storage space requirements and system complexity increase significantly
Solution Approach 1:
The patent implements a universal speech database that stores speech segments with multiple style labels, allowing a single database to serve multiple speaking styles. The system can select and adapt segments from this unified database based on the target style, eliminating the need for separate databases for each style while maintaining speech quality through style-specific segment selection and acoustic parameter adjustment.
Solution Approach 2:
The system changes acoustic and prosodic parameters of speech segments to match the target speaking style. By adjusting parameters such as pitch, duration, and energy, the system can adapt segments from one style to another, reducing the need for style-specific databases while preserving style accuracy through parameter transformation.
2Manufacturing precision
If separate speech databases are created for each speaking style, then style-specific speech quality is improved, but the time and cost required to create and maintain these databases increase
Solution Approach 1:
A single universal database serves all speaking styles, eliminating the need to create and maintain separate databases for each style. The database stores speech segments annotated with multiple style characteristics, allowing the system to efficiently select and adapt segments for any target style without additional database creation or maintenance overhead.
Solution Approach 2:
Speech segments are pre-annotated with multiple style labels and acoustic characteristics during a single database creation process. This preliminary annotation enables rapid style adaptation during synthesis by simply selecting and adjusting pre-labeled segments, rather than creating style-specific databases whenever new styles are needed.
3Volume of stationary object
If speech segments from different styles are used for synthesis, then storage requirements are reduced, but the difficulty of matching acoustic and prosodic characteristics increases
Solution Approach 1:
The system employs feedback mechanisms where acoustic and prosodic characteristics of selected speech segments are compared against the target style profile, and adjustments are made accordingly. This iterative matching process ensures that segments from different styles can be effectively adapted to the target style by measuring and adjusting acoustic parameters such as pitch, duration, and spectral characteristics.
Solution Approach 2:
The system systematically adjusts acoustic parameters (pitch, duration, energy, spectral characteristics) of speech segments to match the target style. By applying parameter transformations based on the difference between the source and target style characteristics, the system can effectively adapt segments from any style while reducing storage requirements through the universal database approach.
Data Source
AI summary
Techniques for performing multi-style speech synthesis. The techniques include using at least one computer hardware processor to perform: obtaining input comprising text and an identification of a desired speaking style to use in rendering the text as speech; identifying a plurality of speech segments for use in synthesizing the text as speech, the identifying comprising identifying a first speech segment recorded and/or synthesized in a first speaking style that is different from the desired speaking style based at least in part on a measure of similarity between the desired speaking style and the first speaking style; synthesizing speech from the text in the desired speaking style at least in part by using the first speech segment; and outputting the synthesized speech.


