RNN-Regularized NMF for Audio Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional sound decomposition techniques fail to accurately separate sound sources due to their inability to capture long-term temporal dependencies, leading to inaccuracies and resource-intensive processes.
Innovation Solution
The use of nonnegative matrix factorization techniques combined with recurrent neural networks (RNNs) to capture temporal dependencies and incorporate global temporal information, employing a Cosine distance as a cost function to improve sound source separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional sound decomposition techniques are used, then the process is simpler and less resource-intensive, but the accuracy of sound source separation deteriorates due to inability to capture temporal dependencies
Solution Approach 1:
The patent combines nonnegative matrix factorization (NMF) with recurrent neural networks (RNNs) to merge the strengths of both techniques. NMF provides efficient feature extraction while RNNs capture long-term temporal dependencies, achieving accurate sound source separation without requiring overly complex architectures
Solution Approach 2:
The patent introduces an intermediary model that bridges NMF and RNNs, where NMF extracts initial features and RNNs process temporal patterns. This intermediary approach allows the system to leverage temporal information for improved separation accuracy while maintaining computational efficiency
2Reliability
If conventional sound decomposition techniques are used, then the computational resources required are reduced, but the reliability of sound source identification deteriorates due to mislabeling of sound portions
Solution Approach 1:
The patent performs preliminary feature extraction using NMF before applying RNNs for temporal analysis. This preliminary action reduces the dimensionality and complexity of data that the RNN must process, thereby lowering computational resource requirements while maintaining reliable sound source identification
Solution Approach 2:
The patent segments the sound decomposition process into distinct stages: NMF-based feature extraction followed by RNN-based temporal modeling. This segmentation allows each component to specialize in specific tasks, improving overall reliability while optimizing resource usage at each stage
Data Source
AI summary
Sound processing techniques using recurrent neural networks are described. In one or more implementations, temporal dependencies are captured in sound data that are modeled through use of a recurrent neural network (RNN). The captured temporal dependencies are employed as part of feature extraction performed using nonnegative matrix factorization (NMF). One or more sound processing techniques are performed on the sound data based at least in part on the feature extraction.


