Mask Estimation Using Gaussian Mixture Model for Sound Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy of sound source separation using a single microphone is compromised because different criteria are used for mask estimation during learning and operation, leading to suboptimal parameter learning and decreased separation accuracy.
Innovation Solution
A neural network model is trained to convert input audio signals into embedded vectors, which are then fitted to a mixed Gaussian model to calculate mask information for extracting specific sound sources, with parameter updates ensuring consistency between learning and operation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If different criteria are used for mask estimation during learning and operation, then the system can be simpler to implement, but the accuracy of sound source separation decreases
Solution Approach 1:
The patent applies homogeneity by using the same mask estimation method (Gaussian mixture model) during both learning and operation phases. This ensures consistency in the criteria for mask estimation, eliminating the discrepancy between training and deployment while maintaining implementation feasibility through a unified approach.
2Measurement precision
If the same method is used for mask estimation during learning and operation, then the accuracy of sound source separation improves, but the system complexity increases
Solution Approach 1:
The patent implements universality by designing a single mask estimation system based on Gaussian mixture model that serves both learning and operation functions. This multi-functional approach uses the same algorithmic framework for both phases, improving accuracy while avoiding the need for separate complex systems for each phase.
3Productivity
If parameters are updated to minimize distance between teacher mask and estimated mask, then learning convergence is faster, but the mask estimation criteria become inconsistent with operation phase
Solution Approach 1:
The patent uses feedback mechanisms where the system continuously compares estimated masks with target masks during learning, and the same Gaussian mixture model criteria are applied during operation. This feedback loop ensures both rapid convergence during learning and consistency during operation by maintaining the same estimation criteria throughout the system lifecycle.
Data Source
AI summary
A mask estimation apparatus for estimating mask information for specifying a mask used to extract a signal of a specific sound source from an input audio signal includes a converter which converts the input audio signal into embedded vectors of a predetermined dimension using a trained neural network model and a mask calculator which calculates the mask information by fitting the embedded vectors to a mixed Gaussian model.


