Speech Synthesis Device Target Voice Data Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech synthesis devices using voice conversion methods often fail to achieve sufficient similarity to a target uttered voice due to the lack of utilization of target voice data during synthesis, relying solely on converted voice data.
Innovation Solution
A speech synthesis device that incorporates both target voice data and converted voice data, with the target voice data being prioritized to enhance similarity, by generating a voice data set that includes target voice data obtained from a target uttered voice and converted voice data obtained from an arbitrary voice, to produce synthesized speech with improved resemblance to the target voice.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If only converted voice data is used in speech synthesis, then the synthesis process can be simplified, but the similarity to the target uttered voice becomes insufficient
Solution Approach 1:
The patent merges converted voice data and target voice data into a unified voice data set, allowing the system to leverage both the conversion capability and the authentic target voice characteristics. This combination resolves the contradiction by maintaining synthesis simplicity while improving voice similarity through the integrated data source.
Solution Approach 2:
The patent applies local quality by using target voice data specifically for portions of the speech where high similarity to the target uttered voice is most important, while relying on converted voice data for other portions. This selective approach improves overall similarity without requiring complete restructuring of the synthesis process.
2Measurement precision
If both target voice data and converted voice data are incorporated, then similarity to target voice is improved, but the complexity of data management increases
Solution Approach 1:
The patent introduces a voice data set as an intermediary structure that organizes both target voice data and converted voice data. This mediator simplifies data management by providing a unified interface and structure, allowing the synthesis system to access both data types without dealing with their individual complexities separately.
Data Source
AI summary
According to an embodiment, a speech synthesis device includes a first storage, a second storage, a first generator, a second generator, a third generator, and a fourth generator. The first storage is configured to store therein first information obtained from a target uttered voice. The second storage is configured to store therein second information obtained from an arbitrary uttered voice. The first generator is configured to generate third information by converting the second information so as to be close to a target voice quality or prosody. The second generator is configured to generate an information set including the first information and the third information. The third generator is configured to generate fourth information used to generate a synthesized speech, based on the information set. The fourth generator configured to generate the synthesized speech corresponding to input text using the fourth information.


