Computational Eyewear Case Offloads Speech Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing hearing aid technologies struggle to improve speech comprehension for individuals with severe to profound hearing loss, especially in noisy or multi-speaker environments, as they primarily focus on amplifying sound rather than enhancing comprehension.
Innovation Solution
An integrated system of extended reality (XR) eyewear coupled with a dedicated computational eyewear case, which includes a multi-sensor microphone array and a rechargeable battery, to provide real-time speech-to-text captioning by offloading processing and communication tasks from smartphones to a single-function computational case.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If processing and communication tasks are offloaded to a smartphone, then the eyewear device complexity is reduced, but the smartphone battery life is depleted quickly and processing reliability varies due to competing tasks
Solution Approach 1:
The system divides processing tasks between the eyewear device and a dedicated computational case. The eyewear contains only essential components (microphones, display, basic processor) while the computational case handles speech-to-text processing and communication tasks. This segmentation reduces eyewear complexity while maintaining reliability through dedicated processing resources in the case.
Solution Approach 2:
The computational case acts as an intermediary device between the eyewear and the cloud/server. It receives audio data from the eyewear, performs local processing, and communicates with external services. This intermediary role ensures reliable processing without depleting smartphone battery or competing for resources.
2Power
If a tethered smartphone is used for real-time processing, then the eyewear can leverage smartphone computational power, but the smartphone battery life is quickly depleted due to continuous processing tasks
Solution Approach 1:
The system separates the processing function from the power source by using a dedicated computational case with its own battery. The case handles all intensive processing tasks (speech-to-text conversion, noise reduction) independently, eliminating the need for the smartphone to consume battery power for these operations. The eyewear and case form a self-contained processing unit.
Solution Approach 2:
Instead of using the smartphone's processing power, the system creates a dedicated computational case that copies the necessary processing functions (speech-to-text engine, noise reduction algorithms) into a separate device. This copying allows the same processing capabilities to exist without the side effect of depleting smartphone battery.
3Measurement precision
If cloud-based speech-to-text systems are used, then the Word Error Rate is minimized, but the reliability of cellular or Wi-Fi connections varies greatly depending on environmental factors
Solution Approach 1:
The computational case performs speech-to-text processing locally before needing to communicate with the cloud. By preparing and processing audio data in advance using local resources, the system reduces dependency on real-time network connectivity. This preliminary local processing ensures accurate transcription even when cloud connection is unreliable.
Solution Approach 2:
The system uses a hybrid approach where local processing in the computational case provides immediate feedback and transcription, while cloud-based services provide supplementary processing when available. The local processor continues operating independently, ensuring continuous reliable service regardless of network conditions.
4Adaptability or versatility
If the eyewear contains all processing capabilities, then the system operates independently, but the production costs increase and the device becomes more complex
Solution Approach 1:
The system segments functionality between the eyewear device and the computational case. The eyewear contains only essential components for capturing audio and displaying results (microphones, simple processor, display). The computational case contains the intensive processing capabilities (speech-to-text engine, noise reduction, communication modules). This segmentation maintains system independence while keeping the eyewear itself simple and cost-effective to manufacture.
Data Source
AI summary
An integrated system provides real-time speech-to-text captioning. The system includes an eyewear component comprising an eyewear frame, one or more microphones, a sensor, a display system, a wireless transceiver, and a processor. The eyewear component captures audio using the microphones and detects when the wearer is speaking using the sensor. The processor receives audio from the microphones, determines if the audio is from the wearer speaking, and transmits non-wearer audio to an eyewear case component for speech-to-text conversion. The eyewear component receives the speech-to-text conversion from the eyewear case component and displays it in the wearer's field of view using the display system. The eyewear case component includes a case housing, a wireless transceiver, at least one microphone, and a processor. The eyewear case component receives audio from the microphone, performs speech-to-text conversion on the received audio data, and transmits the text data to the eyewear component.


