Sound Source Separation via Spatial Frequency Masking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound source separation technologies face significant calculation costs as the number of microphones in a microphone array increases, particularly when using methods like multichannel NMF, which requires enormous computational resources for optimization calculations.

Innovation Solution

The proposed solution involves transforming multichannel sound signals into a spatial frequency domain using an orthonormal base like the Fourier base, diagonalizing the microphone correlation matrix, and employing a spatial frequency mask for sound source separation, reducing the computational complexity by approximating the inverse matrix as a diagonal matrix.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the number of microphones in the microphone array is increased, then the sound collection capability and spatial resolution are improved, but the calculation cost of the inverse matrix of the microphone correlation matrix increases significantly

Engineering Contradiction:
Improvespatial resolutionVSAvoidcalculation cost
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the sound separation problem from the time domain to the spatial frequency domain by applying spatial Fourier transform. This parameter transformation changes the mathematical properties of the problem, allowing the use of diagonalization techniques that reduce calculation complexity from O(N^3) to O(N^2) per frequency bin, making large-scale microphone arrays computationally feasible

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the traditional time-domain optimization approach with a frequency-domain approach based on spatial spectral analysis. By substituting the mechanical optimization process with spectral decomposition and diagonalization operations, the system achieves lower computational complexity while maintaining separation performance

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If traditional sound source separation methods like multichannel NMF are used, then accurate sound source separation is achieved, but enormous calculation cost is required for optimization calculations

Engineering Contradiction:
Improveseparation accuracyVSAvoidcalculation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent changes the domain parameter from time to spatial frequency, and applies diagonalization to the spatial spectral matrix. This parameter transformation converts the complex optimization problem into a series of simpler operations that can be efficiently computed using Fast Fourier Transform and diagonal matrix operations, significantly improving calculation efficiency

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the sound separation problem into independent frequency bin operations. By processing each frequency bin separately in the spatial frequency domain, the system avoids the need for global optimization across all frequencies, reducing computational complexity and enabling parallel processing

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10650841B2Sound source separation apparatus and method
Publication Date: 2020.05.12 SONY GROUP CORP
  • US10650841B2 patent drawing
  • US10650841B2 patent drawing
  • US10650841B2 patent drawing

AI summary

The present technology relates to a sound source separation apparatus and a method which make it possible to separate a sound source at lower calculation cost. A communication unit receives a spatial frequency spectrum of a sound collection signal which is obtained by a microphone array collecting a plane wave of sound from a sound source, and a spatial frequency mask generating unit generates a spatial frequency mask for masking a component of a predetermined region in a spatial frequency domain on the basis of the spatial frequency spectrum. A sound source separating unit extracts a component of a desired sound source from the spatial frequency spectrum as an estimated sound source spectrum on the basis of the spatial frequency mask. The present technology can be applied to a spatial frequency sound source separator.