Speech Synthesis Device Target Voice Data Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech synthesis devices using voice conversion methods often fail to achieve sufficient similarity to a target uttered voice due to the lack of utilization of target voice data during synthesis, relying solely on converted voice data.

Innovation Solution

A speech synthesis device that incorporates both target voice data and converted voice data, with the target voice data being prioritized to enhance similarity, by generating a voice data set that includes target voice data obtained from a target uttered voice and converted voice data obtained from an arbitrary voice, to produce synthesized speech with improved resemblance to the target voice.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If only converted voice data is used in speech synthesis, then the synthesis process can be simplified, but the similarity to the target uttered voice becomes insufficient

Engineering Contradiction:
Improvesynthesis process complexityVSAvoidsimilarity to target voice
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges converted voice data and target voice data into a unified voice data set, allowing the system to leverage both the conversion capability and the authentic target voice characteristics. This combination resolves the contradiction by maintaining synthesis simplicity while improving voice similarity through the integrated data source.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent applies local quality by using target voice data specifically for portions of the speech where high similarity to the target uttered voice is most important, while relying on converted voice data for other portions. This selective approach improves overall similarity without requiring complete restructuring of the synthesis process.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If both target voice data and converted voice data are incorporated, then similarity to target voice is improved, but the complexity of data management increases

Engineering Contradiction:
Improvesimilarity to target voiceVSAvoiddata management complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a voice data set as an intermediary structure that organizes both target voice data and converted voice data. This mediator simplifies data management by providing a unified interface and structure, allowing the synthesis system to access both data types without dealing with their individual complexities separately.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9135910B2Speech synthesis device, speech synthesis method, and computer program product
Publication Date: 2015.09.15 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9135910B2 patent drawing
  • US9135910B2 patent drawing
  • US9135910B2 patent drawing

AI summary

According to an embodiment, a speech synthesis device includes a first storage, a second storage, a first generator, a second generator, a third generator, and a fourth generator. The first storage is configured to store therein first information obtained from a target uttered voice. The second storage is configured to store therein second information obtained from an arbitrary uttered voice. The first generator is configured to generate third information by converting the second information so as to be close to a target voice quality or prosody. The second generator is configured to generate an information set including the first information and the third information. The third generator is configured to generate fourth information used to generate a synthesized speech, based on the information set. The fourth generator configured to generate the synthesized speech corresponding to input text using the fourth information.