Semi-supervised Source Separation Using Non-negative Matrix Factorization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current signal processing technologies face challenges in effectively separating mixed signals from multiple sources, such as audio signals containing overlapping sounds, where computers struggle to differentiate between constituent sound sources like the human auditory system can.
Innovation Solution
The implementation of semi-supervised source separation using non-negative techniques, specifically employing non-negative hidden Markov (N-HMM) and non-negative factorial hidden Markov (N-FHMM) models to model and separate signals of interest from noise or other sources, allowing for independent processing of signal components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional signal processing methods are used to separate mixed signals, then the separation process becomes computationally complex and requires extensive training data, but the separation effectiveness and signal-to-interference ratio remain insufficient
Solution Approach 1:
The patent segments the mixed signal into distinct source components using non-negative matrix factorization, decomposing the spectrogram into source signals and spectral profiles. This segmentation approach enables effective separation without requiring complex processing frameworks, directly resolving the contradiction between separation effectiveness and processing complexity.
Solution Approach 2:
The patent transforms the signal processing problem into the spectrogram domain, changing the parameter representation from time-domain to frequency-time domain. This parameter transformation simplifies the separation task by exploiting the non-negativity property in the spectrogram representation, improving separation effectiveness while reducing computational complexity.
2Reliability
If supervised learning methods with extensive training data are employed, then model accuracy improves, but the requirement for large amounts of training data increases processing complexity and time
Solution Approach 1:
The patent implements a semi-supervised approach where the system performs self-service by automatically learning source characteristics during the separation process itself. The non-negative matrix factorization algorithm adapts to the specific mixture being processed without requiring external training data, achieving high reliability while eliminating the time loss associated with extensive training.
Solution Approach 2:
The patent performs preliminary spectral analysis and non-negativity constraint application before the actual separation process. By pre-processing the signal in the spectrogram domain and establishing non-negativity constraints upfront, the system prepares the data structure to enable accurate separation without requiring subsequent extensive training, thus reducing training time while maintaining accuracy.
3Manufacturing precision
If conventional source separation algorithms are applied to mixed audio signals, then signal components can be separated, but artifacts and interference in the separated signals increase
Solution Approach 1:
The patent converts the harmful interference and artifacts into beneficial separation information by applying non-negativity constraints. The algorithm exploits the fact that source signals and spectral profiles are non-negative in the spectrogram domain, transforming what would be interference into structured components that can be cleanly separated, thus reducing artifacts while achieving precise component separation.
Solution Approach 2:
The patent introduces the spectrogram domain as an intermediary representation between the original mixed signal and the separated sources. By transforming the signal to the frequency-time domain and performing separation in this intermediate space using non-negative matrix factorization, the system achieves clean separation with minimal artifacts, avoiding the interference problems of direct time-domain separation.
Data Source
AI summary
Systems and methods for semi-supervised source separation using non-negative techniques are described. In some embodiments, various techniques disclosed herein may enable the separation of signals present within a mixture, where one or more of the signals may be emitted by one or more different sources. In audio-related applications, for instance, a signal mixture may include speech (e.g., from a human speaker) and noise (e.g., background noise). In some cases, speech may be separated from noise using a speech model developed from training data. A noise model may be created, for example, during the separation process (e.g., “on-the-fly”) and in the absence of corresponding training data.


