Hearing Device Throat-Vibration Sensing for Noisy Speech Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Hearing aid users face challenges in understanding speech in noisy environments, particularly at low signal-to-noise ratios, where traditional beamforming and noise reduction algorithms fail to provide effective enhancement.

Innovation Solution

A hearing device that incorporates a high-speed video camera focused on the throat region of a target talker to detect vibrations of the vocal cords, combined with microphone signals, to enhance noisy speech by using visual information to improve noise reduction algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional beamforming and noise reduction algorithms are used, then speech processing is effective at high signal-to-noise ratios, but speech intelligibility deteriorates at low signal-to-noise ratios

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidnoise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent transitions from purely acoustic signal processing to audio-visual processing by incorporating optical information from a video camera. The system captures visual data of the talker's throat region and processes it alongside audio signals to extract speech information, effectively adding a spatial dimension (visual domain) to complement the acoustic domain and improve speech intelligibility in noisy environments where traditional acoustic methods fail

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces visual information from a video camera as an intermediary to bridge the gap when acoustic signals are too noisy for reliable speech processing. The video camera captures optical signals of the talker's vocal region, which are then processed to extract speech features that supplement or replace degraded acoustic signals, enabling speech understanding at low signal-to-noise ratios

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a high-speed video camera is added to detect vocal cord vibrations, then speech enhancement capability is improved, but device complexity increases

Engineering Contradiction:
Improvevocal cord vibration detectionVSAvoidsystem structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the video camera serve multiple functions: it captures visual information for speech enhancement, tracks the talker's throat region, and provides optical signals that can be processed to extract speech features. By making the visual processing system multi-functional, the patent justifies the added device complexity through enhanced capabilities that go beyond simple speech detection

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces or supplements the mechanical/acoustic signal capture method (microphones) with an optical measurement system (video camera). Instead of relying solely on acoustic waves that are easily masked by noise, the system uses optical fields to detect vocal cord vibrations visually, substituting a more robust measurement modality that is less susceptible to acoustic interference

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If visual information from a video camera is used to enhance speech, then noise reduction performance is improved, but loss of time for processing additional data increases

Engineering Contradiction:
Improvenoise reduction effectivenessVSAvoidsignal processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of the video signal to extract only the relevant throat region information before combining it with audio processing. By pre-identifying and isolating the vocal cord region in the video feed, the system reduces the amount of data that needs to be processed in real-time, minimizing the time penalty of adding visual processing while maintaining noise reduction effectiveness

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances speech intelligibility in noisy conditions by effectively extending the signal-to-noise ratio range, allowing for improved speech clarity even in challenging acoustic environments.

Implementation Method 1

an auxiliary input unit for receiving an auxiliary electric signal representing a current vibration of the vocal cords of a target talker, wherein the auxiliary electric signal is derived from visual information, e.g. provided by light sensitive sensor

Methodology Applied
Scientific EffectLight reflection: Reflection

Data Source

PatentEP3618457B1A hearing device configured to utilize non-audio information to process audio signals
Publication Date: 2025.07.16 OTICON
  • EP3618457B1 patent drawingFigure 1A
  • EP3618457B1 patent drawingFigure 1B
  • EP3618457B1 patent drawingFigure 2A

AI summary

A hearing device, e.g. a hearing aid, is configured to be worn by a user, e.g. fully or partially on the head of the user, comprises a) an input transducer for converting a sound comprising a target sound from a target talker and possible additional sound in an environment of the user, when the user wears the hearing device, to an electric sound signal representative of said sound, b) an auxiliary input unit configured to provide an auxiliary electric signal representative of said target signal or properties thereof, c) a processor connected to said input transducer and to said auxiliary input unit, and wherein said processor is configured to apply a processing algorithm to said electric sound signal, or a signal derived therefrom, to provide an enhanced signal by attenuating components of said additional sound relative to components of said target sound in said electric sound signal, or said signal derived therefrom. The auxiliary electric signal is derived from visual information, e.g. from a camera, containing information of current vibrations of a facial or throat region of said target talker, and the processing algorithm is configured to use the auxiliary electric signal or the signal derived therefrom to provide the enhanced signal.