Electronic Eyewear Visual-Voice Search for Contextual AR Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing wearable electronic devices, such as electronic eyewear, struggle to provide contextually relevant augmented reality features based on user intent and the environment, often delivering irrelevant search results due to a lack of integration of visual and voice data processing.

Innovation Solution

The system captures images and voice commands using integrated cameras and microphones, processes them to extract contextual signals and keywords, and selects augmented reality features that match the user's intent and the viewed scene, refining search results for enhanced relevance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If visual and voice data processing are integrated to provide contextually relevant augmented reality features, then search result relevance is improved, but device complexity increases

Engineering Contradiction:
Improvesearch result relevanceVSAvoiddata processing integration
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines visual data from cameras and voice data from microphones into a unified processing system. The visual processor and voice processor both feed into a common augmented reality feature selection system that integrates multiple data sources to determine contextually relevant features, resolving the contradiction by merging separate processing streams into a coordinated system that improves relevance while managing complexity through structured integration.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an image processor and voice processor as intermediary components that pre-process visual and audio data before feeding them to the augmented reality feature selection system. These intermediaries extract relevant features and convert raw data into structured information, reducing the complexity burden on the main processing system while maintaining high search result relevance through multi-stage processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple sensors and processors are integrated for contextual analysis, then user intent accuracy is improved, but manufacturing complexity increases

Engineering Contradiction:
Improveuser intent accuracyVSAvoiddevice assembly
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent divides the electronic eyewear device into distinct functional modules: cameras for visual input, microphones for voice input, an image processor for visual data analysis, a voice processor for audio data analysis, and an augmented reality feature selection system. This segmentation allows each component to be manufactured and tested independently, then assembled into the final device, improving manufacturability while maintaining high user intent accuracy through specialized processing in each module.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12450837B2Contextual visual and voice search from electronic eyewear device
Publication Date: 2025.10.21 SNAP INC
  • US12450837B2 patent drawing
  • US12450837B2 patent drawing
  • US12450837B2 patent drawing

AI summary

Augmented reality features are selected for presentation to a display of an electronic eyewear device by using a camera of the electronic eyewear device to capture a scan image and processing the scan image to extract contextual signals. Simultaneously, voice data from the user is captured by a microphone of the electronic eyewear device and voice-to-text conversion of the captured voice data is performed to identify keywords in the voice data. The extracted contextual signals and the identified keywords are then used to select at least one augmented reality feature that matches the extracted contextual signals and the identified keywords, and the selected augmented reality feature is presented to the display for user selection. The contextual information thus refines the search results to provide the augmented reality feature best suited for the context of the scan image captured by the electronic eyewear device.