NMF Source Separation with Automatic Base Count via SSVI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech denoising and source separation techniques fail to effectively handle non-stationary background noise due to the requirement of choosing an appropriate number of dictionary elements a priori in nonnegative matrix factorization (NMF) models, leading to inefficiencies and reduced accuracy when using Bayesian nonparametric models with variational inference algorithms.
Innovation Solution
The structured stochastic variational inference (SSVI) algorithm is applied to determine the number of bases for a nonnegative matrix factorization (NMF) model, restoring dependencies between model parameters and latent variables, allowing for automatic selection of the appropriate number of dictionary elements and improving the efficiency and accuracy of source separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If Bayesian nonparametric models with variational inference algorithms are used to automatically determine the number of dictionary elements, then the requirement of choosing an appropriate number a priori is addressed, but dependencies between model parameters and variables are broken introducing local optima and reducing accuracy
Solution Approach 1:
The patent introduces an intermediary mechanism (structured stochastic variational inference framework) that restores the broken dependencies between model parameters and variables. This intermediary allows the system to automatically determine dictionary size while maintaining accurate posterior distributions, resolving the contradiction between automation and precision.
Solution Approach 2:
The patent changes the parameter representation by using a structured stochastic variational approach that maintains dependency relationships. By transforming how parameters are inferred (from independent variational inference to structured joint inference), the system achieves both automatic dictionary size determination and maintained accuracy.
2Productivity
If the number of dictionary elements is chosen a priori in NMF models, then computational efficiency is maintained, but the model cannot adapt to non-stationary noise and reduces source separation accuracy
Solution Approach 1:
The patent makes the dictionary size dynamic by using Bayesian nonparametric models with structured stochastic variational inference. Instead of fixing the number of dictionary elements a priori, the system dynamically determines the appropriate size based on the data, enabling adaptation to non-stationary noise while maintaining computational efficiency through the variational framework.
3Adaptability or versatility
If MCMC models are used for Bayesian nonparametric inference, then automatic discovery of latent components is enabled, but computational expense and inefficiency increase significantly
Solution Approach 1:
The patent substitutes the MCMC mechanical system with a structured stochastic variational inference system. Instead of using sampling-based MCMC methods that are computationally expensive, the patent employs deterministic variational inference with structured dependencies, achieving automatic latent component discovery with significantly improved computational efficiency.
Data Source
AI summary
Methods and systems for source separation based on determining a number of bases for a nonnegative matrix factorization (NMF) model are disclosed. A method includes receiving, at a computing device, a mixed signal including a combination of first signal data and second signal data. The method also includes generating, by the computing device, a time-frequency representation of the mixed signal. The method further includes determining, by applying a structured stochastic variational inference (SSVI) algorithm to the NMF model, a number of bases for a dictionary of signal-related components of the mixed signal. The method uses the number of bases and the time-frequency representation to construct the dictionary and an activation matrix of weights, the weights indicating how active each one of the signal-related components is at a given time. The method then uses the dictionary and the activation matrix to separate the first signal data from the second signal data.


