Wearable AR Captioning for Multi-Speaker Audio Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hearing impaired individuals face difficulties in understanding conversations involving multiple speakers, as existing visual confirmation techniques like lip reading become ineffective when multiple speakers are present or when speakers are outside the person's line of sight.
Innovation Solution
A wearable device equipped with a multi-directional microphone and augmented reality display that captures and analyzes speech from multiple speakers, providing captioned text within the wearer's field of view, even when speakers are outside their line of sight, using a combination of on-board and offsite computing resources for speech recognition and natural language processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If lip reading is used to understand speakers, then direct visual confirmation is possible, but it becomes ineffective when multiple speakers are present or when speakers are outside line of sight
Solution Approach 1:
The system segments the audio signal into multiple directional channels using an array of microphones, allowing separate identification and tracking of multiple speakers in different spatial locations. Each speaker's voice is processed independently through beamforming algorithms, enabling accurate identification even in multi-speaker environments.
Solution Approach 2:
The wearable device integrates multiple functions: directional audio capture, speech-to-text conversion, and augmented reality display. This multi-functional system can handle both single and multiple speakers, speakers within and outside the visual field, making it universally applicable to various conversation scenarios.
2Loss of information
If visual confirmation techniques are used, then speaker understanding is possible, but they are not possible when the speaker is outside a person's line of sight
Solution Approach 1:
The system replaces the mechanical requirement of visual line-of-sight with an acoustic field-based detection system. The array of microphones captures sound waves from all directions in three-dimensional space, allowing speaker identification and speech capture without requiring the speaker to be within the wearer's visual field.
Solution Approach 2:
The system introduces audio signals as an intermediary medium between the speaker and the wearer. Instead of requiring direct visual contact, the audio captured by directional microphones serves as the mediator that conveys speaker information to the wearable device for processing and display.
3Adaptability or versatility
If multiple speakers are involved in a conversation, then social interaction complexity increases, but hearing impaired persons struggle to keep up with the conversation
Solution Approach 1:
The system performs preliminary separation and identification of multiple speaker voices through directional audio processing before the wearer needs to process the information. By pre-organizing the audio signals into distinct speaker channels with spatial information, the system reduces the cognitive load and processing time required for the wearer to follow complex multi-speaker conversations.
Data Source
AI summary
A wearable device providing an augmented reality experience for the benefit of hearing impaired persons is disclosed. The augmented reality experience displays a virtual text caption box that includes text that has been translated from speech detected from surrounding speakers.


