RNN-Regularized NMF for Audio Source Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sound decomposition techniques fail to accurately separate sound sources due to their inability to capture long-term temporal dependencies, leading to inaccuracies and resource-intensive processes.

Innovation Solution

The use of nonnegative matrix factorization techniques combined with recurrent neural networks (RNNs) to capture temporal dependencies and incorporate global temporal information, employing a Cosine distance as a cost function to improve sound source separation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional sound decomposition techniques are used, then the process is simpler and less resource-intensive, but the accuracy of sound source separation deteriorates due to inability to capture temporal dependencies

Engineering Contradiction:
Improveaccuracy of sound source separationVSAvoidcomplexity of decomposition technique
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines nonnegative matrix factorization (NMF) with recurrent neural networks (RNNs) to merge the strengths of both techniques. NMF provides efficient feature extraction while RNNs capture long-term temporal dependencies, achieving accurate sound source separation without requiring overly complex architectures

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary model that bridges NMF and RNNs, where NMF extracts initial features and RNNs process temporal patterns. This intermediary approach allows the system to leverage temporal information for improved separation accuracy while maintaining computational efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If conventional sound decomposition techniques are used, then the computational resources required are reduced, but the reliability of sound source identification deteriorates due to mislabeling of sound portions

Engineering Contradiction:
Improvereliability of sound source identificationVSAvoidcomputational resources required
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary feature extraction using NMF before applying RNNs for temporal analysis. This preliminary action reduces the dimensionality and complexity of data that the RNN must process, thereby lowering computational resource requirements while maintaining reliable sound source identification

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the sound decomposition process into distinct stages: NMF-based feature extraction followed by RNN-based temporal modeling. This segmentation allows each component to specialize in specific tasks, improving overall reliability while optimizing resource usage at each stage

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9721202B2Non-negative matrix factorization regularized by recurrent neural networks for audio processing
Publication Date: 2017.08.01 ADOBE INC
  • US9721202B2 patent drawing
  • US9721202B2 patent drawing
  • US9721202B2 patent drawing

AI summary

Sound processing techniques using recurrent neural networks are described. In one or more implementations, temporal dependencies are captured in sound data that are modeled through use of a recurrent neural network (RNN). The captured temporal dependencies are employed as part of feature extraction performed using nonnegative matrix factorization (NMF). One or more sound processing techniques are performed on the sound data based at least in part on the feature extraction.