Sound Source Separation Filter Estimation with Correlated Covariance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound source separation models, such as ILRMA and ICA, assume no correlation between time frequency bins of sound source spectra, making them unsuitable for modeling unsteady signals like vocal sounds, leading to inaccurate separation.

Innovation Solution

An estimation device and method that calculates a covariance matrix incorporating correlations between sound source spectra and channels to improve sound source separation performance, using models like ILRMA-F, ILRMA-T, and ILRMA-FT that consider frequency, time, or both correlations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If models of the related art (ILRMA, ICA, NMF) are used to perform sound source separation, then the separation process can be implemented with a simple model assuming no correlation between time frequency bins, but the separation accuracy deteriorates when applied to unsteady signals like vocal sounds that have correlation between time frequency bins

Engineering Contradiction:
Improvemodel simplicityVSAvoidsound source separation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent changes the modeling parameter from assuming independence to modeling correlation between time frequency bins. Specifically, it introduces a first-order autoregressive model where the spectral component at time frequency bin (f, t) is modeled as a linear combination of the spectral component at (f, t-1) and Gaussian noise, thereby capturing the temporal correlation in unsteady signals while maintaining computational feasibility

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent makes the model dynamic by allowing the statistical properties of sound sources to vary over time through the autoregressive parameter. This enables the model to adapt to unsteady signals like vocal sounds where the correlation structure changes over time, rather than assuming static independence as in traditional models

Inventive Principle:
Principle #15Dynamics

2Device complexity

If a model assuming no correlation between time frequency bins is used, then the computational complexity remains low, but the model becomes unsuitable for modeling unsteady signals such as vocal sounds

Engineering Contradiction:
Improvecomputational complexityVSAvoidmodel suitability for unsteady signals
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent introduces an autoregressive parameter to model the correlation between adjacent time frequency bins. This parameter change allows the model to capture temporal dynamics in unsteady signals while maintaining a computationally tractable formulation through efficient estimation algorithms

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs preliminary modeling of the correlation structure through the autoregressive framework before performing sound source separation. This preliminary action of capturing temporal correlation enables subsequent separation algorithms to work more effectively on unsteady signals without requiring complex real-time adjustments

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11967328B2Estimation device, estimation method, and estimation program
Publication Date: 2024.04.23 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11967328B2 patent drawing
  • US11967328B2 patent drawing
  • US11967328B2 patent drawing

AI summary

A sound source separation filter information estimation device (10) estimates a covariance matrix having information on a correlation between sound source spectra and information on a correlation between channels as information on sound source separation filter information for separating an individual sound source signal from a mixed acoustic signal.