Generative Model Acoustic Feedback for Natural Sound Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound processing techniques, such as the neural parametric singing synthesizer, face limitations in generating perceptually natural target sounds due to the lack of reflection of fluctuations in acoustic features, leading to unnatural audio properties.
Innovation Solution
A sound processing method that uses a trained generative model to sequentially generate and process acoustic features, where the input data at each time point includes the previously generated acoustic features, allowing for the reflection of fluctuations in waveform signal generation, resulting in a more natural target sound.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If acoustic features are generated by probabilistic processing using random numbers, then the audio property of the waveform signal fluctuates in accordance with the random number, but the acoustic features returned to the generative model do not reflect these fluctuations, limiting the generation of perceptually natural target sound
Solution Approach 1:
The patent introduces a feedback mechanism where the second acoustic feature amount (generated from the waveform signal) is returned to the input side of the generative model as input data for subsequent time points. This feedback loop enables the generative model to receive and process fluctuation information that was present in the original acoustic features, thereby improving the perceptual naturalness of the generated target sound while preserving the fluctuation characteristics introduced by probabilistic processing
Data Source
AI summary
A sound processing method includes: generating with a trained generative model, for each of a plurality of time points including a first time point, a first acoustic feature amount of a target sound to be generated, by sequentially processing input data including condition data representing conditions of the target sound; generating, for each of the plurality of time points, a time-domain waveform signal representing a waveform of the target sound based on the first acoustic feature amount; and generating, for each of the plurality of time points, a second acoustic feature amount based on the time-domain waveform signal. The input data at the first time point includes the second acoustic feature amount generated before the first time point.


