Text to Speech Voice and Attribute Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech systems fail to accurately replicate the human voice, lacking the ability to independently control speaker voice and attributes such as emotion and accent, which limits their naturalness and versatility in applications like electronic games and E-books.

Innovation Solution

A text-to-speech method that separates speaker voice and attribute parameters, using a Gaussian probability distribution model to generate speech with selected voice characteristics, allowing for continuous control of voice and attribute combinations, and enabling the transplantation of voice attributes between speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If current text-to-speech systems are used, then basic speech output is achieved, but the ability to accurately replicate human voice and independently control speaker voice and attributes is lacking

Engineering Contradiction:
Improveability to control speaker voice and attributesVSAvoidaccuracy of human voice replication
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the speaker characteristics into two independent parts: speaker voice (identity) and speaker attributes (emotion, accent, style). This is achieved by factorizing the acoustic model into a speaker-independent base model and speaker-specific components, allowing independent control of voice identity and attributes without compromising either aspect.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses parameter changes by introducing continuous control parameters for speaker attributes (emotion, accent, speaking style) that can be independently adjusted. The acoustic model parameters are modified to allow continuous variation in attributes while maintaining speaker voice identity, enabling smooth transitions and precise control over speech characteristics.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If speaker voice and attributes are controlled independently, then versatility and naturalness are improved, but system complexity increases

Engineering Contradiction:
Improveindependent control of voice and attributesVSAvoidacoustic model structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal acoustic model that serves multiple functions: it can generate speech for different speakers, express different emotions, convey different accents, and combine these attributes in various ways. The base model handles commonalities across all speakers and attributes, while modular components handle speaker-specific and attribute-specific variations, reducing overall complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent adds new dimensions to the acoustic model by introducing separate parameter spaces for speaker voice and speaker attributes. Instead of a single complex parameter space, the model uses multiple lower-dimensional parameter spaces that can be independently controlled, making the system more manageable and easier to implement despite the increased versatility.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9269347B2Text to speech system
Publication Date: 2016.02.23 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9269347B2 patent drawing
  • US9269347B2 patent drawing
  • US9269347B2 patent drawing

AI summary

A text-to-speech method configured to output speech having a selected speaker voice and a selected speaker attribute, including: inputting text; dividing the inputted text into a sequence of acoustic units; selecting a speaker for the inputted text; selecting a speaker attribute for the inputted text; converting the sequence of acoustic units to a sequence of speech vectors using an acoustic model; and outputting the sequence of speech vectors as audio with the selected speaker voice and a selected speaker attribute. The acoustic model includes a first set of parameters relating to speaker voice and a second set of parameters relating to speaker attributes, which parameters do not overlap. The selecting a speaker voice includes selecting parameters from the first set of parameters and the selecting the speaker attribute includes selecting the parameters from the second set of parameters.