Machine Learning Signal Processing for Flexible Sound Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI singer technologies require users to provide detailed instructions for pitch and volume to generate high-quality sound signals, which is burdensome for users.
Innovation Solution
A signal processing method that uses a trained model to generate acoustic feature sequences based on user input, allowing users to select from different degrees of enforcement, reducing the need for detailed control value specification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If detailed instructions pertaining to pitch or volume are provided to the synthesis model, then high-quality sound signals can be generated, but the user experience becomes burdensome
Solution Approach 1:
The system changes the parameter of control granularity by introducing degree of enforcement levels. Instead of requiring detailed pitch and volume instructions, users can provide simpler control values and select an enforcement degree (first, second, or third), which transforms the complexity from the user side while maintaining high-quality sound generation through the trained model's interpretation of these simplified parameters
2Extent of automation
If the synthesis model uses a trained model to generate sound signals, then automation increases, but the model complexity increases
Solution Approach 1:
The system performs preliminary action by pre-training the synthesis model with multiple degrees of enforcement (first, second, and third) during the training phase. This preliminary training enables the model to automatically handle various levels of control enforcement without requiring complex runtime decision-making logic, thus achieving high automation while keeping the operational complexity manageable through pre-computed model configurations
Data Source
AI summary
A signal processing method, which is realized by a computer, includes receiving a control value representing a musical feature, receiving a selection signal for selecting either a first degree of enforcement or a second degree of enforcement that is lower than the first degree of enforcement, and generating, by using a trained model, in accordance with the selection signal, either an acoustic feature amount sequence that reflects the control value in accordance with the first degree of enforcement, or an acoustic feature amount sequence that reflects the control value in accordance with the second degree of enforcement.


