Acoustic Space Model for Natural Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech synthesis systems fail to accurately replicate the natural human voice, lacking emotional expressiveness and speaker identity, resulting in synthesized speech that sounds uniform and unhuman-like.

Innovation Solution

A method utilizing a deep neural network to generate a continuous acoustic space model by correlating phonetic, linguistic, and vocoder features with defined speech attributes, allowing for the synthesis of speech with selected attributes such as emotions, genders, and accents, enabling the creation of more natural and varied human-like voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional text-to-speech synthesis is used, then the system is simple and easy to implement, but the synthesized speech lacks emotional expressiveness and speaker identity

Engineering Contradiction:
Improvespeech attribute variationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by mapping multiple discrete speech attributes (emotions, genders, accents, ages) to continuous acoustic space parameters. The acoustic space model transforms these discrete attribute selections into continuous acoustic parameter adjustments, enabling smooth transitions and natural variation in synthesized speech while maintaining a unified model structure.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a new dimension by creating an acoustic space model that represents speech attributes in a continuous multi-dimensional space. Instead of using separate discrete models for each attribute, the system maps speech attributes to points in a continuous acoustic space, allowing for infinite variations along each dimension while maintaining coherence across all attributes.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If multiple discrete speech attributes are modeled separately, then each attribute can be optimized independently, but the overall system becomes complex and difficult to manage

Engineering Contradiction:
Improveattribute representation accuracyVSAvoidmodel structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple discrete attribute models into a single unified acoustic space model. This model simultaneously represents and can generate speech with any combination of attributes (emotions, genders, accents, ages) by navigating the continuous acoustic space, eliminating the need for separate discrete models while maintaining accurate representation of each attribute.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The acoustic space model serves as a universal representation that can handle any combination of speech attributes. A single model structure performs multiple functions by mapping different attribute combinations to corresponding points in acoustic space, making the system more versatile and easier to manage compared to multiple specialized models.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If a continuous acoustic space model is used, then the synthesized speech sounds more natural and varied, but the training data requirements and processing complexity increase

Engineering Contradiction:
Improvespeech naturalnessVSAvoidtraining efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent uses parameter changes to transform the training process. Instead of training on raw audio waveforms, the system extracts acoustic parameters and maps them to the acoustic space model. This parameter transformation enables more efficient training by working with compressed acoustic representations rather than full audio data, reducing processing requirements while maintaining naturalness.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9916825B2Method and system for text-to-speech synthesis
Publication Date: 2018.03.13 Y E HUB ARMENIA LLC
  • US9916825B2 patent drawing
  • US9916825B2 patent drawing
  • US9916825B2 patent drawing

AI summary

There are disclosed methods and systems for text-to-speech synthesis for outputting a synthetic speech having a selected speech attribute. First, an acoustic space model is trained based on a set of training data of speech attributes, using a deep neural network to determine interdependency factors between the speech attributes in the training data, the dnn generating a single, continuous acoustic space model based on the interdependency factors, the acoustic space model thereby taking into account a plurality of interdependent speech attributes and allowing for modelling of a continuous spectrum of the interdependent speech attributes. Next, a text is received; a selection of one or more speech attribute is received, each speech attribute having a selected attribute weight; the text is converted into synthetic speech using the acoustic space model, the synthetic speech having the selected speech attribute; and the synthetic speech is outputted as audio having the selected speech attribute.