Eyewear Diarization via Stereo Camera and Spatial Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current portable eyewear devices, such as smart glasses, lack effective solutions for providing immersive three-dimensional vision and efficient speech recognition, particularly for users with partial or total blindness, and struggle to distinguish between multiple speakers in real-time environments.
Innovation Solution
The eyewear device employs stereo cameras with visible light and infrared capabilities, combined with advanced image processing and machine learning algorithms to generate three-dimensional images and perform diarization, allowing for real-time speech recognition and text-to-speech conversion, while using eye and head movement tracking to adjust the field of view and display information in a non-obstructive manner.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speech recognition is performed in real-time environments with multiple speakers, then speech recognition capability is improved, but ability to distinguish between multiple speakers deteriorates
Solution Approach 1:
The system segments the mixed audio signal into individual speaker components using source separation algorithms. This allows the speech recognition system to process each speaker's speech independently, maintaining both high speech recognition accuracy and the ability to distinguish between multiple speakers by analyzing their separate audio streams.
2Loss of information
If text is displayed on eyewear display to provide speech recognition feedback, then information provision is improved, but vision obstruction increases
Solution Approach 1:
The system transitions the feedback modality from purely visual (2D display) to spatial audio (3D soundscape). By rendering speech recognition results and speaker distinctions as spatial audio cues that match the physical speaker positions, the system provides comprehensive information feedback without obstructing the user's visual field.
Solution Approach 2:
The system introduces audio feedback as an intermediary medium between the speech recognition system and the user. Instead of directly displaying text on the eyewear, the system uses spatial audio to convey speech information and speaker identification, eliminating the need for visual displays that would obstruct vision.
3Adaptability or versatility
If stereo cameras with visible light and infrared capabilities are used, then three-dimensional vision capability is improved, but device complexity increases
Solution Approach 1:
The system merges the visible light and infrared camera systems into a unified imaging pipeline. By combining the data streams from both camera types and processing them through a single three-dimensional reconstruction algorithm, the system achieves enhanced adaptability for different lighting conditions and depths while managing device complexity through integrated processing.
Data Source
AI summary
An eyewear device that performs diarization by segmenting spoken language into different speakers and remembering each speaker over the course of a session. The speech of each speaker is translated to text and the text of each speaker is displayed on an eyewear display. The text of each user has a different attribute such that the eyewear user can distinguish the text of different speakers. Examples of the text attribute can be a text color, font, and font size. The text is displayed on the eyewear display such that it does not substantially obstruct the user's vision.


