Wearable AR Captioning for Multi-Speaker Audio Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Hearing impaired individuals face difficulties in understanding conversations involving multiple speakers, as existing visual confirmation techniques like lip reading become ineffective when multiple speakers are present or when speakers are outside the person's line of sight.

Innovation Solution

A wearable device equipped with a multi-directional microphone and augmented reality display that captures and analyzes speech from multiple speakers, providing captioned text within the wearer's field of view, even when speakers are outside their line of sight, using a combination of on-board and offsite computing resources for speech recognition and natural language processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If lip reading is used to understand speakers, then direct visual confirmation is possible, but it becomes ineffective when multiple speakers are present or when speakers are outside line of sight

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidmulti-speaker environment adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments the audio signal into multiple directional channels using an array of microphones, allowing separate identification and tracking of multiple speakers in different spatial locations. Each speaker's voice is processed independently through beamforming algorithms, enabling accurate identification even in multi-speaker environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The wearable device integrates multiple functions: directional audio capture, speech-to-text conversion, and augmented reality display. This multi-functional system can handle both single and multiple speakers, speakers within and outside the visual field, making it universally applicable to various conversation scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If visual confirmation techniques are used, then speaker understanding is possible, but they are not possible when the speaker is outside a person's line of sight

Engineering Contradiction:
Improvespeech information accessibilityVSAvoidline of sight requirement
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system replaces the mechanical requirement of visual line-of-sight with an acoustic field-based detection system. The array of microphones captures sound waves from all directions in three-dimensional space, allowing speaker identification and speech capture without requiring the speaker to be within the wearer's visual field.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces audio signals as an intermediary medium between the speaker and the wearer. Instead of requiring direct visual contact, the audio captured by directional microphones serves as the mediator that conveys speaker information to the wearable device for processing and display.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If multiple speakers are involved in a conversation, then social interaction complexity increases, but hearing impaired persons struggle to keep up with the conversation

Engineering Contradiction:
Improvecomplex conversation handling capabilityVSAvoidinformation processing speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system performs preliminary separation and identification of multiple speaker voices through directional audio processing before the wearer needs to process the information. By pre-organizing the audio signals into distinct speaker channels with spatial information, the system reduces the cognitive load and processing time required for the wearer to follow complex multi-speaker conversations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10878819B1System and method for enabling real-time captioning for the hearing impaired via augmented reality
Publication Date: 2020.12.29 UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)
  • US10878819B1 patent drawing
  • US10878819B1 patent drawing
  • US10878819B1 patent drawing

AI summary

A wearable device providing an augmented reality experience for the benefit of hearing impaired persons is disclosed. The augmented reality experience displays a virtual text caption box that includes text that has been translated from speech detected from surrounding speakers.