Binary Mask Error Correction via Statistical Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for calculating binary masks from noisy speech introduce errors due to the lack of clean target speech, making it difficult to improve speech intelligibility in noisy environments.
Innovation Solution
A statistical model, such as a Hidden Markov Model, is used to identify and correct errors in noisy binary masks by training on both clean and noisy signals, allowing for the estimation of a more accurate binary mask representing the target signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If binary masks are calculated using noisy speech instead of clean speech, then the method can operate without requiring clean target speech, but errors are introduced in the binary masks
Solution Approach 1:
The patent applies feedback by using the noisy binary mask to generate a statistical model that is then used to correct errors in the mask. The corrected mask is fed back to improve the statistical model, creating an iterative refinement process that progressively improves accuracy without requiring clean speech input.
Solution Approach 2:
The patent introduces a statistical model as an intermediary between the noisy speech input and the final binary mask output. This intermediary component processes the noisy information and produces corrected binary masks, effectively mediating the transformation from noisy input to accurate output.
2Loss of information
If traditional methods working on waveforms or time-frequency representation are used, then more information is available for processing, but the processing algorithm becomes more complex
Solution Approach 1:
The patent extracts only the essential information needed for binary mask generation by working directly in the binary domain rather than processing full waveforms or time-frequency representations. This extraction approach retains sufficient information for accurate speech processing while dramatically reducing algorithmic complexity.
Solution Approach 2:
The patent segments the processing into distinct stages: generating the initial binary mask from noisy speech, creating a statistical model from this mask, and then using the model to correct errors. This segmentation allows each stage to be optimized independently, reducing overall system complexity while maintaining information integrity.
3Device complexity
If the binary domain is used for processing, then the complexity of the processing algorithm is reduced, but the upper limit on what can be achieved is reduced
Solution Approach 1:
The patent changes the parameters of the binary mask by using a statistical model to adjust individual bit values based on probability calculations. This parameter modification allows the system to achieve higher reliability in the binary domain by refining the mask values through statistical inference rather than direct waveform processing.
Solution Approach 2:
The patent replaces direct mechanical processing of waveforms with a statistical probabilistic system. Instead of processing signals through complex waveform transformations, the system uses statistical models to infer and correct binary mask values, substituting mechanical signal processing with statistical reasoning to achieve better performance.
Data Source
AI summary
The invention relates to a method of identifying and correcting errors in a noisy binary mask. An object of the present invention is to provide a scheme for improving a binary mask representing speech. The problem is solved in that the method comprises a) providing a noisy binary mask comprising a binary representation of the power density of an acoustic signal comprising a target signal and a noise signal at a predefined number of discrete frequencies and a number of discrete time instances; b) providing a statistical model of a clean binary mask representing the power density of the target signal; and c) using the statistical model to detect and correct errors in the noisy binary mask. This has the advantage of providing an alternative and relatively simple way of improving an estimate of a binary mask representing a speech signal. The invention may e.g. be used for speech processing, e.g. in a hearing instrument.


