Sound Generation Using Machine Learning Model for Natural Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound generation methods require detailed specification of time series musical feature amounts like amplitude, volume, and pitch to produce naturally changing sounds, making it difficult for users to generate natural sounds such as singing or instrumental performances.
Innovation Solution
A sound generation method using a trained model that processes representative values of musical feature amounts for each section of a musical note, generating a sound data sequence with continuously changing musical feature amounts, where the model is trained using machine learning to learn the input-output relationship between input and output feature amount sequences extracted from reference sound data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If detailed time series of musical feature amounts are specified to generate natural sounds, then sound quality is improved, but ease of operation deteriorates
Solution Approach 1:
The patent segments the continuous time series of musical feature amounts into discrete sections (e.g., attack, body, release phases of notes). Instead of requiring users to specify detailed values for every time point, the system only requires specification of representative values for each section, which are then interpolated to generate the complete time series. This segmentation approach maintains sound quality while significantly reducing the complexity of user input.
Solution Approach 2:
The patent applies preliminary action by pre-defining the structure and boundaries of sections (attack, body, release) based on musical theory and analysis of reference data. The system prepares the framework of sections in advance, and users only need to fill in representative values for each pre-established section rather than creating the entire time series from scratch. This preliminary structuring simplifies the user's task while preserving the ability to generate high-quality natural sounds.
2Ease of operation
If representative values for each section are used instead of detailed time series, then ease of operation is improved, but manufacturing precision deteriorates
Solution Approach 1:
The patent introduces an intermediary process that automatically interpolates and expands the discrete representative values into a continuous detailed time series. The system uses algorithms to generate intermediate values between the user-specified representative points, effectively acting as a mediator that transforms simple user input into the detailed data required for high-quality sound generation. This intermediary step bridges the gap between ease of operation and manufacturing precision.
Solution Approach 2:
The patent applies parameter changes by transforming the density and distribution of feature amount specifications. Instead of requiring high-density specifications throughout the entire time series, the system changes the parameter distribution to concentrate representative values at critical section boundaries and phases, then uses mathematical interpolation to recover the continuous detailed time series. This parameter transformation maintains sound quality while reducing input complexity.
Data Source
AI summary
A sound generation method that is realized by a computer includes receiving a representative value of a musical feature amount for each of a plurality of sections of a musical note, and using a trained model to process a first feature amount sequence in accordance with the representative value for each section, thereby generating a sound data sequence corresponding to a second feature amount sequence in which the musical feature amount changes continuously.


