AR Eyeglass Captioning With Dual Microphones for Voice Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional augmented reality glasses and smartphone speech-to-text apps struggle with background noise suppression, inadequate processing for severe hearing loss, and fail to distinguish between the wearer's voice and others', leading to inefficient communication assistance for individuals with hearing loss.
Innovation Solution
An augmented reality device with dual microphone systems, one inwardly targeting the wearer and one outwardly targeting the talker, processes signals to distinguish voices and provide real-time text captions, optionally translating languages and capturing emotional valence, with a display integrated into eyeglasses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional augmented reality glasses use single microphone systems to capture speech, then device complexity is reduced, but the ability to distinguish between wearer's voice and talker's voice deteriorates
Solution Approach 1:
The patent divides the microphone system into two separate systems: a first microphone system positioned to capture the wearer's voice and a second microphone system positioned to capture the talker's voice. This segmentation allows the device to distinguish between different voice sources by analyzing signals from spatially separated microphones, resolving the contradiction between simple device design and accurate voice distinction.
2Object-affected harmful factors
If hearing aid devices use beamforming microphone arrays to target prominent sounds, then background noise suppression is improved, but the ability to capture desired speech accurately deteriorates when the most prominent sound is not the desired speech
Solution Approach 1:
The patent introduces an intermediary processing step that analyzes signals from both microphone systems before determining which sound to capture. Rather than directly targeting the most prominent sound, the system uses intermediate signal processing to identify and isolate the desired speech source, allowing background noise suppression while accurately capturing the intended speech even when it's not the most prominent sound.
3Productivity
If smartphone speech-to-text apps capture all audio input, then real-time captioning is provided, but the user experience becomes unnatural by captioning the user's own voice
Solution Approach 1:
The patent extracts and separates the wearer's voice signal from the overall audio input using the first microphone system. By taking out the wearer's own voice from the mixed audio signal, the system can provide real-time captioning of only the talker's speech, maintaining productivity while improving user experience naturalness by eliminating self-captioning.
4Adaptability or versatility
If augmented reality devices integrate multiple sensors and features, then functionality is enhanced, but device complexity and cost increase making them difficult for older people with disabilities to use
Solution Approach 1:
The patent implements a universal processing architecture where a single processor handles multiple functions: analyzing signals from both microphone systems, distinguishing between wearer and talker voices, suppressing background noise, and generating real-time captions. This multi-functionality approach enhances adaptability while managing device complexity by consolidating processing tasks rather than requiring separate dedicated systems for each function.
Data Source
AI summary
A method and apparatus to assist people with hearing loss. An augmented reality device with microphones and a display captured speech of a person talking to the wearer of the device and displays real-time captions in the wearer's field of view, while optionally not captioning the wearer's own speech. The microphone system in this apparatus inverts the use of microphones in an augmented reality device by analyzing and processing environmental sounds while ignoring the wearer's own voice.


