Sound Synthesis Model Segmentation for Natural Fluctuations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound synthesis techniques struggle to generate high-quality sounds that include significant fluctuations, such as vibrato and stochastic fluctuations, due to the suppression of dynamic components in estimation models constructed using machine learning.

Innovation Solution

The method involves using two separate models: a first model to generate series of fluctuations based on first control data, and a second model to generate series of features, including pitches and spectral features, using both control data and the generated fluctuations, allowing for a more natural and dynamic sound synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a single estimation model is constructed using machine learning with training data including series of pitches, then the model can generate synthesis sounds, but the series of fluctuations is suppressed and sound quality deteriorates

Engineering Contradiction:
Improvesound qualityVSAvoidmodel structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the single estimation model into two separate models: a first model that generates the series of fluctuations from control data, and a second model that generates the series of pitches using both control data and the generated fluctuations. This segmentation allows each model to specialize in one aspect, preventing the suppression of fluctuations while maintaining manageable complexity through modular design.

Inventive Principle:
Principle #1Segmentation

2Extent of automation

If machine learning is used to construct an estimation model with training data, then the model can estimate series of pitches, but dynamic components (series of fluctuations) are suppressed

Engineering Contradiction:
Improveautomatic pitch estimationVSAvoidloss of dynamic fluctuations
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The patent applies preliminary action by first generating the series of fluctuations using the first model before using them as input for the second model to estimate pitches. This preliminary generation of fluctuations ensures that dynamic components are preserved and fed into the pitch estimation process, preventing information loss while maintaining automation.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If a single model estimates both control data and pitches directly, then the process is simple, but the generated sounds lack natural fluctuations

Engineering Contradiction:
Improveprocessing process simplicityVSAvoidnaturalness of sound
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the processing into two distinct stages: first generating fluctuations from control data, then using those fluctuations to estimate pitches. This segmentation, while slightly increasing structural complexity, ensures that natural fluctuations are preserved and incorporated into the pitch estimation, significantly improving sound naturalness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The series of fluctuations acts as an intermediary between the control data and the pitch estimation. The first model transforms control data into fluctuations, which then serve as essential input for the second model's pitch estimation. This intermediary role ensures natural fluctuations are transmitted through the processing pipeline without being suppressed.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11875777B2Information processing method, estimation model construction method, information processing device, and estimation model constructing device
Publication Date: 2024.01.16 YAMAHA CORP
  • US11875777B2 patent drawing
  • US11875777B2 patent drawing
  • US11875777B2 patent drawing

AI summary

An information processing device includes a memory storing instructions, and a processor configured to implement the stored instructions to execute a plurality of tasks. The tasks includes: a first generating task that generates a series of fluctuations of a target sound based on first control data of the target sound to be synthesized, using a first model trained to have an ability to estimate a series of fluctuations of the target sound based on first control data of the target sound, and a second generating task that generates a series of features of the target sound based on second control data of the target sound and the generated series of fluctuations of the target sound, using a second model trained to estimate a series of features of the target sound based on second control data of the target sound and a series of fluctuations of the target sound.