Acoustic Space Model for Natural Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-speech synthesis systems fail to accurately replicate the natural human voice, lacking emotional expressiveness and speaker identity, resulting in synthesized speech that sounds uniform and unhuman-like.
Innovation Solution
A method utilizing a deep neural network to generate a continuous acoustic space model by correlating phonetic, linguistic, and vocoder features with defined speech attributes, allowing for the synthesis of speech with selected attributes such as emotions, genders, and accents, enabling the creation of more natural and varied human-like voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional text-to-speech synthesis is used, then the system is simple and easy to implement, but the synthesized speech lacks emotional expressiveness and speaker identity
Solution Approach 1:
The patent applies parameter changes by mapping multiple discrete speech attributes (emotions, genders, accents, ages) to continuous acoustic space parameters. The acoustic space model transforms these discrete attribute selections into continuous acoustic parameter adjustments, enabling smooth transitions and natural variation in synthesized speech while maintaining a unified model structure.
Solution Approach 2:
The patent introduces a new dimension by creating an acoustic space model that represents speech attributes in a continuous multi-dimensional space. Instead of using separate discrete models for each attribute, the system maps speech attributes to points in a continuous acoustic space, allowing for infinite variations along each dimension while maintaining coherence across all attributes.
2Reliability
If multiple discrete speech attributes are modeled separately, then each attribute can be optimized independently, but the overall system becomes complex and difficult to manage
Solution Approach 1:
The patent merges multiple discrete attribute models into a single unified acoustic space model. This model simultaneously represents and can generate speech with any combination of attributes (emotions, genders, accents, ages) by navigating the continuous acoustic space, eliminating the need for separate discrete models while maintaining accurate representation of each attribute.
Solution Approach 2:
The acoustic space model serves as a universal representation that can handle any combination of speech attributes. A single model structure performs multiple functions by mapping different attribute combinations to corresponding points in acoustic space, making the system more versatile and easier to manage compared to multiple specialized models.
3Manufacturing precision
If a continuous acoustic space model is used, then the synthesized speech sounds more natural and varied, but the training data requirements and processing complexity increase
Solution Approach 1:
The patent uses parameter changes to transform the training process. Instead of training on raw audio waveforms, the system extracts acoustic parameters and maps them to the acoustic space model. This parameter transformation enables more efficient training by working with compressed acoustic representations rather than full audio data, reducing processing requirements while maintaining naturalness.
Data Source
AI summary
There are disclosed methods and systems for text-to-speech synthesis for outputting a synthetic speech having a selected speech attribute. First, an acoustic space model is trained based on a set of training data of speech attributes, using a deep neural network to determine interdependency factors between the speech attributes in the training data, the dnn generating a single, continuous acoustic space model based on the interdependency factors, the acoustic space model thereby taking into account a plurality of interdependent speech attributes and allowing for modelling of a continuous spectrum of the interdependent speech attributes. Next, a text is received; a selection of one or more speech attribute is received, each speech attribute having a selected attribute weight; the text is converted into synthetic speech using the acoustic space model, the synthetic speech having the selected speech attribute; and the synthetic speech is outputted as audio having the selected speech attribute.


