Sound Synthesis Model Segmentation for Natural Fluctuations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound synthesis techniques struggle to generate high-quality sounds that include significant fluctuations, such as vibrato and stochastic fluctuations, due to the suppression of dynamic components in estimation models constructed using machine learning.
Innovation Solution
The method involves using two separate models: a first model to generate series of fluctuations based on first control data, and a second model to generate series of features, including pitches and spectral features, using both control data and the generated fluctuations, allowing for a more natural and dynamic sound synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a single estimation model is constructed using machine learning with training data including series of pitches, then the model can generate synthesis sounds, but the series of fluctuations is suppressed and sound quality deteriorates
Solution Approach 1:
The patent divides the single estimation model into two separate models: a first model that generates the series of fluctuations from control data, and a second model that generates the series of pitches using both control data and the generated fluctuations. This segmentation allows each model to specialize in one aspect, preventing the suppression of fluctuations while maintaining manageable complexity through modular design.
2Extent of automation
If machine learning is used to construct an estimation model with training data, then the model can estimate series of pitches, but dynamic components (series of fluctuations) are suppressed
Solution Approach 1:
The patent applies preliminary action by first generating the series of fluctuations using the first model before using them as input for the second model to estimate pitches. This preliminary generation of fluctuations ensures that dynamic components are preserved and fed into the pitch estimation process, preventing information loss while maintaining automation.
3Device complexity
If a single model estimates both control data and pitches directly, then the process is simple, but the generated sounds lack natural fluctuations
Solution Approach 1:
The patent segments the processing into two distinct stages: first generating fluctuations from control data, then using those fluctuations to estimate pitches. This segmentation, while slightly increasing structural complexity, ensures that natural fluctuations are preserved and incorporated into the pitch estimation, significantly improving sound naturalness.
Solution Approach 2:
The series of fluctuations acts as an intermediary between the control data and the pitch estimation. The first model transforms control data into fluctuations, which then serve as essential input for the second model's pitch estimation. This intermediary role ensures natural fluctuations are transmitted through the processing pipeline without being suppressed.
Data Source
AI summary
An information processing device includes a memory storing instructions, and a processor configured to implement the stored instructions to execute a plurality of tasks. The tasks includes: a first generating task that generates a series of fluctuations of a target sound based on first control data of the target sound to be synthesized, using a first model trained to have an ability to estimate a series of fluctuations of the target sound based on first control data of the target sound, and a second generating task that generates a series of features of the target sound based on second control data of the target sound and the generated series of fluctuations of the target sound, using a second model trained to estimate a series of features of the target sound based on second control data of the target sound and a series of fluctuations of the target sound.


