Multichannel Sound Source Estimation via Nonnegative Tensor Factorization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional signal processing techniques for microphone arrays face challenges in accurately estimating sound sources due to variations in acoustic characteristics and microphone arrangements, leading to decreased quality of sound source separation, especially when multiple sound sources are present.

Innovation Solution

A signal processing system that applies nonnegative matrix factorization (NMF) and nonnegative tensor factorization (NTF) to multichannel signal processing outputs, allowing for adaptive estimation of model parameters and noise suppression filters, without prior assumptions about the environment, to separate sound sources and improve estimation accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional signal processing techniques are used for microphone arrays, then the system structure is simple, but the sound source estimation accuracy decreases due to acoustic variations and microphone arrangement deviations

Engineering Contradiction:
Improvesound source estimation accuracyVSAvoidrobustness to acoustic variations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the fundamental parameters of signal processing by transitioning from conventional linear filtering methods to nonnegative matrix factorization (NMF) and nonnegative tensor factorization (NTF). This mathematical transformation allows the system to adapt to acoustic variations by learning spatial and spectral patterns directly from data, rather than relying on fixed microphone arrangement assumptions. The parameter change enables the system to maintain high estimation accuracy despite variations in acoustic characteristics and microphone positioning.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamic adaptability through iterative optimization algorithms that continuously adjust model parameters based on input signals. The NMF/NTF framework allows the system to dynamically learn spatial basis vectors and spectral characteristics, making it adaptable to changing acoustic environments and microphone arrangements. This dynamic approach replaces static conventional methods with learning-based adaptive processing.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If nonnegative matrix factorization and nonnegative tensor factorization are applied, then sound source estimation accuracy improves independently of acoustic variations, but the computational complexity increases

Engineering Contradiction:
Improvesound source estimation accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex sound source separation problem into manageable components through NMF/NTF factorization. By decomposing the multichannel signal into spatial basis vectors, spectral basis vectors, and activity vectors, the system handles each component separately. This segmentation reduces computational complexity compared to attempting to solve the entire problem simultaneously, while maintaining high estimation accuracy through the coordinated processing of decomposed elements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by focusing computational resources on extracting the most significant spatial and spectral patterns through NMF/NTF. Rather than processing all signal characteristics equally, the method identifies and processes dominant components that contribute most to sound source estimation accuracy. This selective approach reduces overall computational complexity while maintaining effectiveness.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10373628B2Signal processing system, signal processing method, and computer program product
Publication Date: 2019.08.06 KK TOSHIBA
  • US10373628B2 patent drawing
  • US10373628B2 patent drawing
  • US10373628B2 patent drawing

AI summary

A signal processing system includes a filter unit, a conversion unit, a decomposition unit, and an estimation unit. The filter unit applies, to a plurality of time series input signals, N filters estimated by independent component analysis of the input signals to output N output signals. The conversion unit converts the output signals into nonnegative signals each taking on a nonnegative value. The decomposition unit decomposes the nonnegative signals into a spatial basis that includes nonnegative three-dimensional elements, that is, K first elements, N second elements, and I third elements, a spectral basis matrix of I rows and L columns that includes L nonnegative spectral basis vectors expressed by I-dimensional column vectors, and a nonnegative L-dimensional activity vector. The estimation unit estimates sound source signals representing signals of the signal sources based on the output signals using the spatial basis, the spectral basis matrix, and the activity vector.