Voice Synthesis Model Using Separated Sound Source and Style Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice synthesis techniques require extensive preparation of voice units for each combination of speaking persons and performance styles, making them burdensome to implement.
Innovation Solution
A computer-based information processing method using a machine learning-generated synthesis model that inputs sound source, style, and synthesis data to generate target sounds without pre-prepared voice units, utilizing deep neural networks to produce acoustic features and audio signals for various combinations of speakers and styles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If unit-concatenating-type voice synthesis is used to synthesize target sounds with different combinations of speaking persons and performance styles, then the synthesis can accommodate variety in speakers and styles, but the preparation burden of voice units becomes excessively large
Solution Approach 1:
The patent extracts the essential characteristics of voice units and separates them into independent sound source data and style data components. Instead of preparing complete voice units for each speaker-style combination, the system extracts and stores only the essential sound source characteristics and style characteristics separately, which can then be combined during synthesis to generate target sounds for any speaker-style combination without requiring pre-prepared voice units for each combination.
Solution Approach 2:
The patent segments the voice unit preparation process into two independent parts: sound source data preparation (representing speaker characteristics) and style data preparation (representing performance style characteristics). This segmentation allows the system to prepare a manageable set of sound source data and style data independently, then combine them during synthesis to create target sounds for any combination, avoiding the need to prepare complete voice units for every possible speaker-style combination.
2Manufacturing precision
If voice units are prepared for each combination of speaking persons and performance styles, then accurate synthesis for each combination is achieved, but the preparation work and system complexity increase significantly
Solution Approach 1:
The patent performs preliminary action by pre-extracting and pre-storing sound source data and style data in a prepared state, but not pre-computing complete voice units for each combination. The sound source data and style data are prepared in advance and stored in databases, allowing the synthesis system to quickly retrieve and combine these pre-prepared components for any speaker-style combination without requiring time-consuming pre-preparation of complete voice units for each combination.
Solution Approach 2:
The patent introduces sound source data and style data as intermediary representations that mediate between the raw audio data and the final synthesized output. Instead of directly manipulating complete voice units for each combination, the system uses these intermediary data structures to represent speaker characteristics and style characteristics separately, which can then be efficiently combined during synthesis to generate target sounds for any combination without requiring pre-prepared voice units for each combination.
3Adaptability or versatility
If extensive voice unit preparation is performed for each speaker-style combination, then comprehensive synthesis capability is achieved, but the ease of operation and system maintenance deteriorates
Solution Approach 1:
The patent creates a universal synthesis system that can handle any speaker-style combination by using a single set of sound source data and style data that are prepared independently. The sound source data serves as a universal representation of speaker characteristics that can be combined with any style data, and the style data serves as a universal representation of performance styles that can be applied to any sound source. This multi-functional approach allows the system to synthesize target sounds for any combination without requiring separate voice unit preparations for each combination, thereby improving ease of operation and maintenance while maintaining comprehensive synthesis capability.
Data Source
AI summary
An information processing system includes at least one memory storing a program and at least one processor. The at least one processor implements the program to input a piece of sound source data obtained by encoding a first identification data representative of a sound source, a piece of style data obtained by encoding a second identification data representative of a performance style, and synthesis data representative of sounding conditions into a synthesis model generated by machine learning, and to generate, using the synthesis model, feature data representative of acoustic features of a target sound of the sound source to be generated in the performance style and according to the sounding conditions, and to generate an audio signal corresponding to the target sound using the generated feature data.


