Mask Estimation Using Gaussian Mixture Model for Sound Source Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The accuracy of sound source separation using a single microphone is compromised because different criteria are used for mask estimation during learning and operation, leading to suboptimal parameter learning and decreased separation accuracy.

Innovation Solution

A neural network model is trained to convert input audio signals into embedded vectors, which are then fitted to a mixed Gaussian model to calculate mask information for extracting specific sound sources, with parameter updates ensuring consistency between learning and operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If different criteria are used for mask estimation during learning and operation, then the system can be simpler to implement, but the accuracy of sound source separation decreases

Engineering Contradiction:
ImproveEase of implementationVSAvoidAccuracy of sound source separation
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies homogeneity by using the same mask estimation method (Gaussian mixture model) during both learning and operation phases. This ensures consistency in the criteria for mask estimation, eliminating the discrepancy between training and deployment while maintaining implementation feasibility through a unified approach.

Inventive Principle:
Principle #33Homogeneity

2Measurement precision

If the same method is used for mask estimation during learning and operation, then the accuracy of sound source separation improves, but the system complexity increases

Engineering Contradiction:
ImproveAccuracy of sound source separationVSAvoidSystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements universality by designing a single mask estimation system based on Gaussian mixture model that serves both learning and operation functions. This multi-functional approach uses the same algorithmic framework for both phases, improving accuracy while avoiding the need for separate complex systems for each phase.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If parameters are updated to minimize distance between teacher mask and estimated mask, then learning convergence is faster, but the mask estimation criteria become inconsistent with operation phase

Engineering Contradiction:
ImproveLearning convergence speedVSAvoidConsistency of mask estimation criteria
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent uses feedback mechanisms where the system continuously compares estimated masks with target masks during learning, and the same Gaussian mixture model criteria are applied during operation. This feedback loop ensures both rapid convergence during learning and consistency during operation by maintaining the same estimation criteria throughout the system lifecycle.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11562765B2Mask estimation apparatus, model learning apparatus, sound source separation apparatus, mask estimation method, model learning method, sound source separation method, and program
Publication Date: 2023.01.24 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11562765B2 patent drawing
  • US11562765B2 patent drawing
  • US11562765B2 patent drawing

AI summary

A mask estimation apparatus for estimating mask information for specifying a mask used to extract a signal of a specific sound source from an input audio signal includes a converter which converts the input audio signal into embedded vectors of a predetermined dimension using a trained neural network model and a mask calculator which calculates the mask information by fitting the embedded vectors to a mixed Gaussian model.