Sound Source Separation via Hidden Variable Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for separating mixed sound signals, particularly in pop music, into human voice and accompaniment are inefficient, as they rely on end-to-end deterministic models that struggle to accurately isolate sound sources, impacting music editing and retrieval processes.
Innovation Solution
A method and apparatus that utilize a coding model to extract hidden variables representing human voice and accompaniment features from mixed sound signals, followed by decoding models to separate these components, employing neural networks trained with loss functions to minimize errors and improve sound source separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If an end-to-end deterministic model is used to separate voice and accompaniment, then the separation process is simplified, but the accuracy of sound source isolation deteriorates
Solution Approach 1:
The patent segments the separation process into two distinct stages: first extracting hidden variables from the mixed sound signal using a coding model, then separating the actual sound sources using decoding models. This segmentation allows each stage to focus on specific tasks, improving overall separation accuracy while maintaining manageable complexity
Solution Approach 2:
The patent introduces hidden variables as an intermediary representation between the mixed sound signal and the final separated sources. These hidden variables capture essential features of the sound sources without directly revealing them, enabling more accurate separation through the decoding models while keeping the process structured
2Adaptability or versatility
If model training is performed online, then real-time adaptation is improved, but resource consumption and processing time increase
Solution Approach 1:
The patent performs model training in advance (offline) before actual separation tasks. The coding and decoding models are pre-trained on training datasets to learn the characteristics of different sound sources. This preliminary action allows the models to be ready for deployment without requiring continuous online training, reducing real-time computational burden while maintaining adaptability
Data Source
AI summary
This application can provide a method and electronic device for separating mixed sound signal. The method includes: obtaining a first hidden variable representing a human voice feature and a second hidden variable representing an accompaniment sound feature by inputting feature data of a mixed sound extracted from a mixed sound signal into a coding model for the mixed sound; obtaining first feature data of a human voice and second feature data of an accompaniment sound by inputting the first hidden variable and the second hidden variable into a first decoding model for the human voice and a second decoding model for the accompaniment sound respectively; and obtaining, based on the first feature data and the second feature data, the human voice and the accompaniment sound.


