AR Eyeglass Captioning With Dual Microphones for Voice Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional augmented reality glasses and smartphone speech-to-text apps struggle with background noise suppression, inadequate processing for severe hearing loss, and fail to distinguish between the wearer's voice and others', leading to inefficient communication assistance for individuals with hearing loss.

Innovation Solution

An augmented reality device with dual microphone systems, one inwardly targeting the wearer and one outwardly targeting the talker, processes signals to distinguish voices and provide real-time text captions, optionally translating languages and capturing emotional valence, with a display integrated into eyeglasses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If conventional augmented reality glasses use single microphone systems to capture speech, then device complexity is reduced, but the ability to distinguish between wearer's voice and talker's voice deteriorates

Engineering Contradiction:
Improvemicrophone system complexityVSAvoidvoice distinction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent divides the microphone system into two separate systems: a first microphone system positioned to capture the wearer's voice and a second microphone system positioned to capture the talker's voice. This segmentation allows the device to distinguish between different voice sources by analyzing signals from spatially separated microphones, resolving the contradiction between simple device design and accurate voice distinction.

Inventive Principle:
Principle #1Segmentation

2Object-affected harmful factors

If hearing aid devices use beamforming microphone arrays to target prominent sounds, then background noise suppression is improved, but the ability to capture desired speech accurately deteriorates when the most prominent sound is not the desired speech

Engineering Contradiction:
Improvebackground noise suppressionVSAvoiddesired speech capture accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary processing step that analyzes signals from both microphone systems before determining which sound to capture. Rather than directly targeting the most prominent sound, the system uses intermediate signal processing to identify and isolate the desired speech source, allowing background noise suppression while accurately capturing the intended speech even when it's not the most prominent sound.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If smartphone speech-to-text apps capture all audio input, then real-time captioning is provided, but the user experience becomes unnatural by captioning the user's own voice

Engineering Contradiction:
Improvereal-time captioning capabilityVSAvoiduser experience naturalness
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent extracts and separates the wearer's voice signal from the overall audio input using the first microphone system. By taking out the wearer's own voice from the mixed audio signal, the system can provide real-time captioning of only the talker's speech, maintaining productivity while improving user experience naturalness by eliminating self-captioning.

Inventive Principle:
Principle #2Taking out (Extraction)

4Adaptability or versatility

If augmented reality devices integrate multiple sensors and features, then functionality is enhanced, but device complexity and cost increase making them difficult for older people with disabilities to use

Engineering Contradiction:
Improvefunctional capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal processing architecture where a single processor handles multiple functions: analyzing signals from both microphone systems, distinguishing between wearer and talker voices, suppressing background noise, and generating real-time captions. This multi-functionality approach enhances adaptability while managing device complexity by consolidating processing tasks rather than requiring separate dedicated systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250363993A1Eyeglass augmented reality speech to text device and method
Publication Date: 2025.11.27 XANDERGLASSES INC
  • US20250363993A1 patent drawing
  • US20250363993A1 patent drawing
  • US20250363993A1 patent drawing

AI summary

A method and apparatus to assist people with hearing loss. An augmented reality device with microphones and a display captured speech of a person talking to the wearer of the device and displays real-time captions in the wearer's field of view, while optionally not captioning the wearer's own speech. The microphone system in this apparatus inverts the use of microphones in an augmented reality device by analyzing and processing environmental sounds while ignoring the wearer's own voice.