Sound Generation Model Coarse-to-Fine Feature Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound generation methods require users to specify detailed time series of musical feature amounts like amplitude, volume, and pitch, making it difficult to generate naturally changing sounds such as those from singing or performing.
Innovation Solution
A sound generation method using a trained model that processes a first feature amount sequence with a lower fineness to generate a sound data sequence with a higher fineness, allowing for the easy acquisition of natural sounds by learning the input-output relationship between the sequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If detailed time series of musical feature amounts are specified, then sound generation accuracy is improved, but ease of operation deteriorates
Solution Approach 1:
The system performs preliminary extraction of musical feature amounts from reference sound data before the actual sound generation process. By pre-processing and storing feature amount sequences at high fineness, the system eliminates the need for users to manually specify detailed time series data during operation, thus improving ease of operation while maintaining sound generation accuracy
Solution Approach 2:
The system creates copies of musical feature amount sequences from reference sound data at different fineness levels. These copied sequences serve as templates that can be directly used for sound generation without requiring users to manually create detailed specifications, resolving the contradiction between accuracy and ease of operation
2Reliability
If fineness of feature amount sequence is increased, then sound naturalness is improved, but data processing complexity increases
Solution Approach 1:
The system segments the feature amount sequence into multiple representations at different fineness levels. By dividing the data into coarse and fine granularity versions, the system can process lower fineness data for basic operations while maintaining high fineness data for natural sound generation, thus reducing overall data processing complexity while preserving sound naturalness
Solution Approach 2:
The system adds a fineness dimension to the feature amount data structure. By organizing data in multiple dimensional layers (different fineness levels), the system can efficiently process data at appropriate granularity for each operation, reducing complexity while maintaining the capability to generate natural sounds when needed
Data Source
AI summary
A sound generation method that is realized by a computer includes receiving a first feature amount sequence in which a musical feature amount changes over time, and using a trained model that has learned an input-output relationship between an input feature amount sequence in which the musical feature amount changes over time at a first fineness and a reference sound data sequence corresponding to an output feature amount sequence in which the musical feature amount changes over time at a second fineness that is higher than the first fineness, to process the first feature amount sequence, thereby generating a sound data sequence corresponding to a second feature amount sequence in which the musical feature amount changes at the second fineness.


