Acoustic Training Data Generation for Hearing Aid Speech Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern hearing support devices struggle to differentiate between desired target sounds and noise, particularly in complex environments like restaurants, where background speech can be indistinguishable from desired speech, known as the 'cocktail party problem.

Innovation Solution

A method involving the creation of a virtual three-dimensional room for simulating sound sources and receivers, generating synthetic acoustic data samples, and training a machine learning model to distinguish between desired and interfering sounds, using proximity-based enhancement, speech enhancement, and dereverberation algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If hearing support devices amplify all sounds reaching the device, then the overall sound level is improved, but the ability to distinguish target sound from noise deteriorates

Engineering Contradiction:
Improvesound levelVSAvoidsound discrimination
Core Design Contradiction:
Illumination intensityVSLoss of information

Solution Approach 1:

The patent segments the acoustic signal into multiple components by using multiple microphones positioned at different locations. Each microphone captures a portion of the acoustic scene, and the system processes these segmented signals separately to identify and enhance target speech while suppressing noise, thereby resolving the contradiction between amplifying all sounds and distinguishing target sound.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by enhancing specific spatial regions rather than uniformly processing all sounds. The system identifies the spatial location of target speech sources and applies selective amplification to those regions while applying suppression to other regions containing noise, thus improving sound discrimination while maintaining overall sound level.

Inventive Principle:
Principle #3Local quality

2Device complexity

If hearing support devices amplify all sounds without differentiation, then the device complexity is reduced, but the sound clarity in noisy environments deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidsound clarity
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by pre-processing the acoustic signals from multiple microphones to estimate the acoustic environment and identify target speech sources before final sound enhancement. This preliminary analysis of spatial and spectral characteristics enables the system to distinguish target sound from noise efficiently, achieving good sound clarity without requiring overly complex real-time processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating virtual representations of the acoustic scene through signal processing. The system generates estimated acoustic signals that replicate the characteristics of target speech sources while filtering out noise, allowing the device to enhance clarity without directly manipulating all raw acoustic signals with high complexity.

Inventive Principle:
Principle #26Copying

3Area of stationary object

If hearing support devices amplify speech from all directions equally, then the coverage area is improved, but the ability to suppress distant noise deteriorates

Engineering Contradiction:
Improvecoverage areaVSAvoiddistant noise
Core Design Contradiction:
Area of stationary objectVSObject-affected harmful factors

Solution Approach 1:

The patent applies asymmetry by treating sounds from different directions differently. The system uses the spatial arrangement of multiple microphones to create asymmetric processing: sounds from certain directions (where target speech is identified) are enhanced, while sounds from other directions (containing distant noise) are suppressed. This directional asymmetry allows the device to maintain wide coverage while selectively suppressing harmful distant noise.

Inventive Principle:
Principle #4Asymmetry

Data Source

PatentUS11937073B1Systems and methods for curating a corpus of synthetic acoustic training data samples and training a machine learning model for proximity-based acoustic enhancement
Publication Date: 2024.03.19 AUDIOFOCUS INC
  • US11937073B1 patent drawing
  • US11937073B1 patent drawing
  • US11937073B1 patent drawing

AI summary

A system and method includes generating a virtual n-dimensional space that includes one or more positions of one or more source nodes and a position of a receiver node; executing a plurality of simulations including simulating acoustic signals emanating from the one or more source nodes within the virtual n-dimensional room; estimating a measure of the acoustic signals received at the receiver node; computing a plurality of acoustic signal data samples based on the estimation for each of the plurality of simulations; and creating a training data corpus for training an artificial neural network, the training data corpus including at least a sampling of the plurality of acoustic data samples, and the artificial neural network, once trained, is configured to generate an inference indicating a likely intended sound to a target receiver of a mixture of acoustic signals.