Head Mounted Device Text Capture and Action Triggering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing smart glasses and augmented/virtual reality devices lack mechanisms to efficiently capture and utilize text from the real world environment, such as billboards or product labels, to trigger follow-up actions.

Innovation Solution

The implementation of a method and apparatus that capture images or videos of the environment, analyze text items within these images or videos to determine interesting text, and use this text to trigger various actions, such as copying, pasting, extracting phone numbers or websites, adding contacts, or translating text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If smart glasses capture and process text from the real world environment, then text processing capability and user interaction are improved, but device complexity and computational requirements increase

Engineering Contradiction:
Improvetext processing capabilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The text processing system is divided into separate modules: text detection module, text recognition module, and action triggering module. This segmentation allows each component to be optimized independently and reduces overall system complexity while maintaining comprehensive text processing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An intermediary processing layer is introduced between text capture and action execution. This layer analyzes detected text, determines user intent, and selectively triggers actions, thereby managing complexity by filtering and organizing information flow rather than directly mapping all text to actions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If smart glasses analyze all text items in captured images, then text detection accuracy is improved, but processing time and energy consumption increase

Engineering Contradiction:
Improvetext detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts only relevant text items from captured images based on predefined criteria such as text size, contrast, and spatial distribution. By filtering out irrelevant text elements, the system maintains high detection accuracy for meaningful text while significantly reducing processing time and computational load.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of analyzing every text item in the field of view, the system performs partial analysis on selected text regions that meet specific importance thresholds. This selective approach ensures accurate detection of critical text while avoiding unnecessary processing of insignificant text elements.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If smart glasses trigger multiple actions based on captured text, then functionality and user productivity are improved, but system reliability and error risk increase

Engineering Contradiction:
Improveuser productivityVSAvoidsystem reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary validation and confirmation steps before triggering actions. Detected text is verified against known formats (e.g., validating phone number patterns, checking URL structures), and users are presented with confirmation prompts for critical actions. This preliminary processing reduces errors and increases reliability while maintaining high productivity through automated workflow initiation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms where action results are monitored and fed back to the text processing module. If an action fails or produces unexpected results, the system adjusts its text interpretation and retriggers appropriate actions. This closed-loop feedback system enhances reliability by continuously validating the correctness of text-based action triggering.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250182410A1Methods, apparatuses and computer program products for facilitating actions based on text captured by head mounted devices
Publication Date: 2025.06.05 META PLATFORMS INC
  • US20250182410A1 patent drawing
  • US20250182410A1 patent drawing
  • US20250182410A1 patent drawing

AI summary

A system and method for determining interesting text to trigger actions of devices are provided. The system may include one or more head-mounted devices associated with a network. A head mounted device(s) may capture an image(s) and/or a video(s) corresponding to an environment detected in a field of view of a camera(s). The image(s) and/or the video(s) may include one or more text items associated with the environment. The head mounted device may determine whether a text item(s) of the one or more text items is interesting. The head mounted device may extract the text item(s) determined as being interesting and may superimpose the text item(s) at a position in the image(s) and/or the video(s). The head mounted device may trigger, based on the text item(s) determined as being interesting, one or more actions capable of being performed by the head mounted device.