Hearing Device Throat-Vibration Sensing for Noisy Speech Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hearing aid users face challenges in understanding speech in noisy environments, particularly at low signal-to-noise ratios, where traditional beamforming and noise reduction algorithms fail to provide effective enhancement.
Innovation Solution
A hearing device that incorporates a high-speed video camera focused on the throat region of a target talker to detect vibrations of the vocal cords, combined with microphone signals, to enhance noisy speech by using visual information to improve noise reduction algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional beamforming and noise reduction algorithms are used, then speech processing is effective at high signal-to-noise ratios, but speech intelligibility deteriorates at low signal-to-noise ratios
Solution Approach 1:
The patent transitions from purely acoustic signal processing to audio-visual processing by incorporating optical information from a video camera. The system captures visual data of the talker's throat region and processes it alongside audio signals to extract speech information, effectively adding a spatial dimension (visual domain) to complement the acoustic domain and improve speech intelligibility in noisy environments where traditional acoustic methods fail
Solution Approach 2:
The patent introduces visual information from a video camera as an intermediary to bridge the gap when acoustic signals are too noisy for reliable speech processing. The video camera captures optical signals of the talker's vocal region, which are then processed to extract speech features that supplement or replace degraded acoustic signals, enabling speech understanding at low signal-to-noise ratios
2Measurement precision
If a high-speed video camera is added to detect vocal cord vibrations, then speech enhancement capability is improved, but device complexity increases
Solution Approach 1:
The patent makes the video camera serve multiple functions: it captures visual information for speech enhancement, tracks the talker's throat region, and provides optical signals that can be processed to extract speech features. By making the visual processing system multi-functional, the patent justifies the added device complexity through enhanced capabilities that go beyond simple speech detection
Solution Approach 2:
The patent replaces or supplements the mechanical/acoustic signal capture method (microphones) with an optical measurement system (video camera). Instead of relying solely on acoustic waves that are easily masked by noise, the system uses optical fields to detect vocal cord vibrations visually, substituting a more robust measurement modality that is less susceptible to acoustic interference
3Reliability
If visual information from a video camera is used to enhance speech, then noise reduction performance is improved, but loss of time for processing additional data increases
Solution Approach 1:
The patent performs preliminary processing of the video signal to extract only the relevant throat region information before combining it with audio processing. By pre-identifying and isolating the vocal cord region in the video feed, the system reduces the amount of data that needs to be processed in real-time, minimizing the time penalty of adding visual processing while maintaining noise reduction effectiveness
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances speech intelligibility in noisy conditions by effectively extending the signal-to-noise ratio range, allowing for improved speech clarity even in challenging acoustic environments.
Implementation Method 1
an auxiliary input unit for receiving an auxiliary electric signal representing a current vibration of the vocal cords of a target talker, wherein the auxiliary electric signal is derived from visual information, e.g. provided by light sensitive sensor
Data Source
Figure 1A
Figure 1B
Figure 2A
AI summary
A hearing device, e.g. a hearing aid, is configured to be worn by a user, e.g. fully or partially on the head of the user, comprises a) an input transducer for converting a sound comprising a target sound from a target talker and possible additional sound in an environment of the user, when the user wears the hearing device, to an electric sound signal representative of said sound, b) an auxiliary input unit configured to provide an auxiliary electric signal representative of said target signal or properties thereof, c) a processor connected to said input transducer and to said auxiliary input unit, and wherein said processor is configured to apply a processing algorithm to said electric sound signal, or a signal derived therefrom, to provide an enhanced signal by attenuating components of said additional sound relative to components of said target sound in said electric sound signal, or said signal derived therefrom. The auxiliary electric signal is derived from visual information, e.g. from a camera, containing information of current vibrations of a facial or throat region of said target talker, and the processing algorithm is configured to use the auxiliary electric signal or the signal derived therefrom to provide the enhanced signal.