Sound Source Separation via Spatial Frequency Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source separation technologies face significant calculation costs as the number of microphones in a microphone array increases, particularly when using methods like multichannel NMF, which requires enormous computational resources for optimization calculations.
Innovation Solution
The proposed solution involves transforming multichannel sound signals into a spatial frequency domain using an orthonormal base like the Fourier base, diagonalizing the microphone correlation matrix, and employing a spatial frequency mask for sound source separation, reducing the computational complexity by approximating the inverse matrix as a diagonal matrix.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the number of microphones in the microphone array is increased, then the sound collection capability and spatial resolution are improved, but the calculation cost of the inverse matrix of the microphone correlation matrix increases significantly
Solution Approach 1:
The patent transforms the sound separation problem from the time domain to the spatial frequency domain by applying spatial Fourier transform. This parameter transformation changes the mathematical properties of the problem, allowing the use of diagonalization techniques that reduce calculation complexity from O(N^3) to O(N^2) per frequency bin, making large-scale microphone arrays computationally feasible
Solution Approach 2:
The patent replaces the traditional time-domain optimization approach with a frequency-domain approach based on spatial spectral analysis. By substituting the mechanical optimization process with spectral decomposition and diagonalization operations, the system achieves lower computational complexity while maintaining separation performance
2Measurement precision
If traditional sound source separation methods like multichannel NMF are used, then accurate sound source separation is achieved, but enormous calculation cost is required for optimization calculations
Solution Approach 1:
The patent changes the domain parameter from time to spatial frequency, and applies diagonalization to the spatial spectral matrix. This parameter transformation converts the complex optimization problem into a series of simpler operations that can be efficiently computed using Fast Fourier Transform and diagonal matrix operations, significantly improving calculation efficiency
Solution Approach 2:
The patent segments the sound separation problem into independent frequency bin operations. By processing each frequency bin separately in the spatial frequency domain, the system avoids the need for global optimization across all frequencies, reducing computational complexity and enabling parallel processing
Data Source
AI summary
The present technology relates to a sound source separation apparatus and a method which make it possible to separate a sound source at lower calculation cost. A communication unit receives a spatial frequency spectrum of a sound collection signal which is obtained by a microphone array collecting a plane wave of sound from a sound source, and a spatial frequency mask generating unit generates a spatial frequency mask for masking a component of a predetermined region in a spatial frequency domain on the basis of the spatial frequency spectrum. A sound source separating unit extracts a component of a desired sound source from the spatial frequency spectrum as an estimated sound source spectrum on the basis of the spatial frequency mask. The present technology can be applied to a spatial frequency sound source separator.


