Wearable Hearing Aid Lip-Visual Audio Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current hearing enhancement devices fail to effectively filter ambient noise and enhance audio frequencies in environments with multiple speakers, particularly in crowded settings where background noise interferes with conversations.
Innovation Solution
A wearable device comprising a camera, microphone, and speaker system connected to a processor that associates visual lip movements with audio inputs, filters irrelevant sounds, builds a database of faces and their associated frequencies, and amplifies voices by monitoring specific frequency ranges based on gender and age, allowing for effective noise cancellation and voice enhancement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If ambient noise filtering is implemented using traditional microphones, then noise reduction is achieved, but audio clarity in crowded environments with multiple speakers deteriorates
Solution Approach 1:
The system segments the audio spectrum into multiple frequency bands and processes each band separately. By dividing the complex acoustic environment into manageable frequency segments, the system can identify and enhance speech frequencies while filtering out ambient noise in each segment, thereby improving voice detection accuracy without sacrificing noise reduction capability.
Solution Approach 2:
The patent introduces visual information from the camera as an intermediary to assist audio processing. By capturing lip movements and facial expressions, the system creates a visual channel that complements the audio channel, enabling more accurate speech detection and differentiation between multiple speakers even in noisy environments.
2Measurement precision
If frequency-based voice enhancement is applied to all detected sounds, then audio clarity improves, but ability to distinguish between multiple speakers deteriorates
Solution Approach 1:
The system applies different processing qualities to different spatial locations and frequency bands. By analyzing the directional information and applying localized enhancement only to relevant speech sources while maintaining different processing characteristics for different speakers, the system preserves speaker differentiation while enhancing audio clarity for each individual speaker.
Solution Approach 2:
The patent adds the visual dimension to the audio processing system. By incorporating camera data that captures facial movements and lip patterns, the system creates an additional dimension for speaker identification and differentiation, allowing it to distinguish between multiple speakers even when their audio frequencies overlap.
3Measurement precision
If a database of faces and frequencies is built, then voice association accuracy improves, but device complexity increases
Solution Approach 1:
The system performs preliminary actions by pre-processing and storing characteristic features of faces and their associated audio frequencies in a database. By preparing and organizing this data in advance, the system reduces the computational burden during real-time operation, as it only needs to query and compare against pre-processed data rather than performing full analysis on every input.
Solution Approach 2:
The patent creates simplified copies or representations of face-audio associations in the database. Instead of storing complete raw data, the system extracts and stores key characteristic features that can be quickly matched and compared, reducing storage requirements and processing complexity while maintaining association accuracy.
Data Source
AI summary
An apparatus comprising a camera, a microphone, a speaker and a processor. The camera may be mounted to a first portion of a pair of eyeglasses. The microphone may be mounted to a second portion of the pair of eyeglasses. The speaker may be mounted to the pair of eyeglasses. The processor may be electronically connected to the camera, the microphone and the speaker. The processor may be configured to (i) associate a visual display of the movement of the lips of a target received from the camera with an audio portion of the target received from the microphone, (ii) filter sounds not related to the target, and (iii) play sounds not filtered through the speaker.


