Binary Mask Sound Recognition for Hearing Aids
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sound recognition technologies face challenges in effectively separating speech from noise in mixed auditory environments, particularly in improving speech intelligibility and quality, especially in low-power devices like hearing aids.
Innovation Solution
A method and system for automatic sound recognition using binary masks that represent sound elements as binary time-frequency units, allowing for the estimation and modification of sound patterns to enhance speech intelligibility and quality, utilizing a training database of models and statistical methods like Hidden Markov Models for pattern recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional speech recognition methods are used to separate speech from noise in mixed auditory environments, then speech recognition capability is provided, but speech intelligibility and quality deteriorate in noisy conditions
Solution Approach 1:
The patent segments the speech signal into time-frequency units (TFUs) and creates binary masks that selectively preserve or suppress specific TFUs based on their likelihood of containing speech versus noise. This segmentation in the time-frequency domain allows precise control over which spectral components are retained, thereby maintaining speech intelligibility while filtering noise.
Solution Approach 2:
The patent applies different processing qualities to different regions of the time-frequency spectrum. By estimating binary masks that identify speech-dominated TFUs versus noise-dominated TFUs, the system applies local quality enhancement to speech regions while suppressing noise regions, thereby improving overall speech quality without uniformly processing the entire signal.
2Measurement precision
If complex speech processing algorithms are used to improve speech separation, then speech quality is improved, but device complexity and power consumption increase
Solution Approach 1:
The patent extracts only the essential binary mask information from the time-frequency representation of the signal, rather than processing the entire complex signal. By taking out and processing only the binary mask data that indicates speech versus noise regions, the system achieves effective speech separation with reduced computational complexity suitable for low-power devices.
Solution Approach 2:
The patent changes the parameter representation from continuous amplitude values to binary values (0 or 1) indicating noise or speech dominance. This parameter change simplifies the data structure and reduces computational requirements while maintaining the essential information needed for speech quality enhancement.
3Measurement precision
If binary masks are estimated and modified to match training patterns, then speech intelligibility is enhanced, but processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary action by pre-computing and storing binary mask templates for various phonemes and speech elements during an offline training phase. During real-time processing, the system only needs to compare the estimated binary mask against these pre-computed templates, significantly reducing online processing time while maintaining high speech intelligibility.
Solution Approach 2:
The patent creates simplified binary mask representations (copies) of speech elements that capture the essential time-frequency energy distribution patterns. These binary mask copies are computationally lightweight compared to full spectral representations, enabling fast comparison and matching operations in real-time speech processing.
Data Source
Figure 1
Figure 2
Figure 3a~3b
AI summary
The invention relates to a method of automatic sound recognition. The object of the present invention is to provide an alternative scheme for automatically recognizing sounds, e.g. human speech. The problem is solved by providing a training database comprising a number of models, each model representing a sound element in the form of a binary mask comprising binary time frequency (TF) units which indicate the energetic areas in time and frequency of the sound element in question, or of characteristic features or statistics extracted from the binary mask; providing an input signal comprising an input sound element; estimating the input sound element based on the models of the training database to provide an output sound element. The method has the advantage of being relatively simple and adaptable to the application in question. The invention may e.g. be used in devices comprising automatic sound recognition, e.g. for sound, e.g. voice control of a device, or in listening devices, e.g. hearing aids, for improving speech perception.