Speech Synthesis Decoder for User-Controlled Speech Style Variation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional deep learning-based speech synthesis models fail to output synthetic voices with the desired speech style of the user, as they are trained on specific voice data and lack personal control over voice features.
Innovation Solution
A speech synthesis device that includes a processor to acquire voice feature information through a text and user input, using a decoder supervised-trained to minimize the difference between learning text and voice characteristics, allowing for the generation of synthetic voices in various styles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional deep learning-based speech synthesis models are trained on specific voice data, then the model learns the voice characteristics, but the model cannot output synthetic voices with diverse speech styles desired by users
Solution Approach 1:
The speech synthesis model is segmented into multiple independent style components (first style component, second style component, etc.), each responsible for a specific speech style. This allows the model to generate diverse speech styles by combining different components without retraining the entire model, resolving the contradiction between style diversity and training complexity.
Solution Approach 2:
The style components are designed to be universal and interchangeable, where each component can be applied to different text inputs to produce different speech styles. This multi-functional design enables the system to achieve diverse speech styles without requiring separate models for each style, reducing overall complexity.
2Manufacturing precision
If Concatenative synthesis is used to connect phonemes, then the sound quality is natural, but the rhythm becomes unstable
Solution Approach 1:
A style component acts as an intermediary between the text input and the phoneme connection process. It introduces style-specific parameters that guide the synthesis of phonemes while maintaining stable rhythm patterns, thus preserving natural sound quality while ensuring rhythm stability.
Solution Approach 2:
The style components modify specific parameters (such as pitch, duration, and intensity) of the synthesized phonemes to achieve different speech styles while maintaining stable rhythmic structures. This allows rhythm stability to be preserved while enabling diverse speech styles.
3Stability of the object's composition
If Statistical parametric speech synthesis is used, then the rhythm is stable, but noise (buzzing) is caused in the vocoding process
Solution Approach 1:
The patent replaces the traditional vocoding mechanism with a neural network-based synthesis approach. The style components generate acoustic features directly through learned patterns, substituting the mechanical vocoding process that causes buzzing noise, while maintaining stable rhythm through the learned temporal patterns.
Data Source
AI summary
Provided is a speech synthetic device capable of outputting a synthetic voice having various speech styles. The speech synthesis device includes a speaker, and a processor to acquire voice feature information through a text and a user input; generate a synthetic voice, by receiving the text and the voice feature information inputs into a decoder supervised-trained to minimize a difference between feature information of a learning text and characteristic information of a learning voice, and output the generated synthetic voice through the speaker.


