Voice Synthesis Model Using Separated Sound Source and Style Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis techniques require extensive preparation of voice units for each combination of speaking persons and performance styles, making them burdensome to implement.

Innovation Solution

A computer-based information processing method using a machine learning-generated synthesis model that inputs sound source, style, and synthesis data to generate target sounds without pre-prepared voice units, utilizing deep neural networks to produce acoustic features and audio signals for various combinations of speakers and styles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If unit-concatenating-type voice synthesis is used to synthesize target sounds with different combinations of speaking persons and performance styles, then the synthesis can accommodate variety in speakers and styles, but the preparation burden of voice units becomes excessively large

Engineering Contradiction:
Improveability to synthesize different combinations of speakers and performance stylesVSAvoidpreparation burden of voice units
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts the essential characteristics of voice units and separates them into independent sound source data and style data components. Instead of preparing complete voice units for each speaker-style combination, the system extracts and stores only the essential sound source characteristics and style characteristics separately, which can then be combined during synthesis to generate target sounds for any speaker-style combination without requiring pre-prepared voice units for each combination.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the voice unit preparation process into two independent parts: sound source data preparation (representing speaker characteristics) and style data preparation (representing performance style characteristics). This segmentation allows the system to prepare a manageable set of sound source data and style data independently, then combine them during synthesis to create target sounds for any combination, avoiding the need to prepare complete voice units for every possible speaker-style combination.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If voice units are prepared for each combination of speaking persons and performance styles, then accurate synthesis for each combination is achieved, but the preparation work and system complexity increase significantly

Engineering Contradiction:
Improvesynthesis accuracy for each speaker-style combinationVSAvoidpreparation time for voice units
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-extracting and pre-storing sound source data and style data in a prepared state, but not pre-computing complete voice units for each combination. The sound source data and style data are prepared in advance and stored in databases, allowing the synthesis system to quickly retrieve and combine these pre-prepared components for any speaker-style combination without requiring time-consuming pre-preparation of complete voice units for each combination.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces sound source data and style data as intermediary representations that mediate between the raw audio data and the final synthesized output. Instead of directly manipulating complete voice units for each combination, the system uses these intermediary data structures to represent speaker characteristics and style characteristics separately, which can then be efficiently combined during synthesis to generate target sounds for any combination without requiring pre-prepared voice units for each combination.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If extensive voice unit preparation is performed for each speaker-style combination, then comprehensive synthesis capability is achieved, but the ease of operation and system maintenance deteriorates

Engineering Contradiction:
Improvecomprehensive synthesis capability for various speakers and stylesVSAvoidease of system operation and maintenance
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent creates a universal synthesis system that can handle any speaker-style combination by using a single set of sound source data and style data that are prepared independently. The sound source data serves as a universal representation of speaker characteristics that can be combined with any style data, and the style data serves as a universal representation of performance styles that can be applied to any sound source. This multi-functional approach allows the system to synthesize target sounds for any combination without requiring separate voice unit preparations for each combination, thereby improving ease of operation and maintenance while maintaining comprehensive synthesis capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11942071B2Information processing method and information processing system for sound synthesis utilizing identification data associated with sound source and performance styles
Publication Date: 2024.03.26 YAMAHA CORP
  • US11942071B2 patent drawing
  • US11942071B2 patent drawing
  • US11942071B2 patent drawing

AI summary

An information processing system includes at least one memory storing a program and at least one processor. The at least one processor implements the program to input a piece of sound source data obtained by encoding a first identification data representative of a sound source, a piece of style data obtained by encoding a second identification data representative of a performance style, and synthesis data representative of sounding conditions into a synthesis model generated by machine learning, and to generate, using the synthesis model, feature data representative of acoustic features of a target sound of the sound source to be generated in the performance style and according to the sounding conditions, and to generate an audio signal corresponding to the target sound using the generated feature data.