Binary Mask Sound Recognition for Hearing Aids

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound recognition technologies face challenges in effectively separating speech from noise in mixed auditory environments, particularly in improving speech intelligibility and quality, especially in low-power devices like hearing aids.

Innovation Solution

A method and system for automatic sound recognition using binary masks that represent sound elements as binary time-frequency units, allowing for the estimation and modification of sound patterns to enhance speech intelligibility and quality, utilizing a training database of models and statistical methods like Hidden Markov Models for pattern recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speech recognition methods are used to separate speech from noise in mixed auditory environments, then speech recognition capability is provided, but speech intelligibility and quality deteriorate in noisy conditions

Engineering Contradiction:
Improvespeech recognition capabilityVSAvoidspeech intelligibility and quality
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the speech signal into time-frequency units (TFUs) and creates binary masks that selectively preserve or suppress specific TFUs based on their likelihood of containing speech versus noise. This segmentation in the time-frequency domain allows precise control over which spectral components are retained, thereby maintaining speech intelligibility while filtering noise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities to different regions of the time-frequency spectrum. By estimating binary masks that identify speech-dominated TFUs versus noise-dominated TFUs, the system applies local quality enhancement to speech regions while suppressing noise regions, thereby improving overall speech quality without uniformly processing the entire signal.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If complex speech processing algorithms are used to improve speech separation, then speech quality is improved, but device complexity and power consumption increase

Engineering Contradiction:
Improvespeech qualityVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential binary mask information from the time-frequency representation of the signal, rather than processing the entire complex signal. By taking out and processing only the binary mask data that indicates speech versus noise regions, the system achieves effective speech separation with reduced computational complexity suitable for low-power devices.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter representation from continuous amplitude values to binary values (0 or 1) indicating noise or speech dominance. This parameter change simplifies the data structure and reduces computational requirements while maintaining the essential information needed for speech quality enhancement.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If binary masks are estimated and modified to match training patterns, then speech intelligibility is enhanced, but processing time and computational resources increase

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-computing and storing binary mask templates for various phonemes and speech elements during an offline training phase. During real-time processing, the system only needs to compare the estimated binary mask against these pre-computed templates, significantly reducing online processing time while maintaining high speech intelligibility.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates simplified binary mask representations (copies) of speech elements that capture the essential time-frequency energy distribution patterns. These binary mask copies are computationally lightweight compared to full spectral representations, enabling fast comparison and matching operations in real-time speech processing.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP2306457B1Automatic sound recognition based on binary time frequency units
Publication Date: 2016.10.12 OTICON
  • EP2306457B1 patent drawingFigure 1
  • EP2306457B1 patent drawingFigure 2
  • EP2306457B1 patent drawingFigure 3a~3b

AI summary

The invention relates to a method of automatic sound recognition. The object of the present invention is to provide an alternative scheme for automatically recognizing sounds, e.g. human speech. The problem is solved by providing a training database comprising a number of models, each model representing a sound element in the form of a binary mask comprising binary time frequency (TF) units which indicate the energetic areas in time and frequency of the sound element in question, or of characteristic features or statistics extracted from the binary mask; providing an input signal comprising an input sound element; estimating the input sound element based on the models of the training database to provide an output sound element. The method has the advantage of being relatively simple and adaptable to the application in question. The invention may e.g. be used in devices comprising automatic sound recognition, e.g. for sound, e.g. voice control of a device, or in listening devices, e.g. hearing aids, for improving speech perception.