Head-Mounted Device Gesture-Triggered OCR for Low-Power Content Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional head-mounted devices (HMDs) face limitations in providing real-time, low-latency, and low-power content access due to power constraints and processing limitations, leading to inefficient content interaction, especially for users with visual impairments or reading challenges.

Innovation Solution

A head-mounted device configured to capture content within its field of view, perform optical character recognition, and enable navigation between content items using touch, hold, or look gestures, with minimal image capture and OCR occurrence, allowing users to access content independently of their physical reach and language translation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional HMDs perform continuous image capture and OCR processing to enable content access, then content recognition accuracy is improved, but power consumption increases and latency increases

Engineering Contradiction:
Improvecontent recognition accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs image capture and OCR processing periodically or on-demand based on user gestures rather than continuously. The camera captures images only when triggered by specific gestures (touch, hold, look), and OCR processing is performed selectively on captured images, reducing overall power consumption while maintaining content recognition accuracy when needed.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system uses the user's own gestures and natural interactions to trigger content capture and processing. The HMD detects gestures such as touching the device, holding it in a specific position, or looking at content for a duration, and automatically initiates the content recognition process without requiring separate commands, thereby reducing unnecessary processing and power consumption.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If conventional HMDs perform continuous image capture and OCR processing to enable content access, then content recognition accuracy is improved, but processing latency increases

Engineering Contradiction:
Improvecontent recognition accuracyVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by capturing images and performing OCR processing only when triggered by user gestures, rather than waiting for continuous processing cycles. This on-demand approach reduces latency by initiating processing immediately when content access is needed, rather than through continuous periodic processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system skips unnecessary processing steps by only capturing images and performing OCR when triggered by specific gestures. Rather than continuously processing all visible content, the system rushes through the essential steps (gesture detection → image capture → OCR → result delivery) only when needed, reducing overall processing latency.

Inventive Principle:
Principle #21Skipping (Rushing through)

3Measurement precision

If conventional HMDs require repetitive pointing actions for content navigation, then content access precision is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvecontent access precisionVSAvoidnavigation convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system extracts the navigation function from repetitive pointing actions by using alternative gesture-based controls. Users can navigate content by touching the device, holding it in specific positions, or using gaze direction, which eliminates the need for continuous precise pointing while maintaining content access precision through gesture recognition.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system introduces gestures as an intermediary between the user and content navigation. Instead of directly pointing at content on the display, users perform gestures (touch, hold, look) that serve as mediators to select and navigate content, improving ease of operation while maintaining precision through the mapping of gestures to specific content items.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Speed

If conventional HMDs perform frequent OCR processing to enable real-time content access, then content access speed is improved, but power consumption increases

Engineering Contradiction:
Improvecontent access speedVSAvoidpower consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system performs OCR processing periodically or on-demand based on gesture triggers rather than continuously. Images are captured and processed only when the user performs a relevant gesture, reducing the frequency of OCR operations and associated power consumption while maintaining real-time content access speed when needed.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system uses user gestures to self-trigger OCR processing, performing the computationally intensive task only when the user indicates a need for content access. This eliminates unnecessary OCR processing that would consume power without providing immediate benefit, while maintaining fast response times when content is actually needed.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11435857B1Content access and navigation using a head-mounted device
Publication Date: 2022.09.06 GOOGLE LLC
  • US11435857B1 patent drawing
  • US11435857B1 patent drawing
  • US11435857B1 patent drawing

AI summary

In a general aspect, a head-mounted device (HMD) can be configured to receive a selection of a mode of operation of a content reader of the HMD from a plurality of modes of operation, and initiate, in response to the selection, a content capture process to capture content within a field of view (FOV) of a camera of the HMD. The HMD can be further configured to identify, in the captured content, a plurality of content items, receive a navigation command to select a content item from the plurality of content items, and provide, in response to the navigation command, readback of the selected content item.