Camera View Text Overlay for AR Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for providing input to electronic devices are tedious and time-consuming, such as manually inputting phone numbers or web addresses, and lack efficient ways to utilize detected text in real-time camera views for quick actions.

Innovation Solution

A portable computing device processes images to recognize and locate text, identifies text entity types, and renders overlays on the camera view that allow users to perform associated functions by selecting the overlays, such as calling a number or opening a web browser, using optical character recognition and machine learning algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If manual input methods are used for entering phone numbers or web addresses, then input accuracy can be maintained, but user time and effort are significantly increased

Engineering Contradiction:
Improvetime for manual inputVSAvoidmanual input effort
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The patent replaces manual mechanical input (typing phone numbers or web addresses) with an optical recognition system. The camera captures text in the real world, and optical character recognition (OCR) technology automatically converts the visual text into actionable data, substituting the mechanical typing process with an automated visual recognition system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by allowing the device to automatically extract and process text information from the environment without requiring user intervention. The camera and OCR system work autonomously to identify text, determine its type (phone number, URL, etc.), and present relevant actions, making the device serve itself in gathering and processing information.

Inventive Principle:
Principle #25Self-service

2Productivity

If text recognition and overlay rendering are added to the camera system, then user interaction speed is improved, but device complexity increases

Engineering Contradiction:
Improvetext interaction speedVSAvoidcamera processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the text processing task into distinct functional modules: text detection in the camera view, optical character recognition to convert text to actionable data, entity type determination to classify the text (phone number, URL, etc.), and overlay rendering to display interactive elements. This segmentation allows each module to be optimized independently and processed in a streamlined pipeline.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-processing the camera feed to detect and recognize text before the user needs to interact with it. The OCR and entity type determination occur in the background, preparing the data and presenting ready-to-execute actions, so when the user views the overlay, the work is already done and they can immediately select an action.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If real-time text detection and overlay rendering are implemented, then user convenience is enhanced, but processing time and computational resources are increased

Engineering Contradiction:
Improveuser convenienceVSAvoidcamera processing energy
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by focusing computational resources only on detecting and processing text elements within the camera view, rather than analyzing the entire scene. The system selectively identifies text regions, applies OCR only to those regions, and renders overlays only where text is detected, thereby reducing overall computational load and energy consumption while maintaining user convenience.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9922431B2Providing overlays based on text in a live camera view
Publication Date: 2018.03.20 AMAZON TECH INC
  • US9922431B2 patent drawing
  • US9922431B2 patent drawing
  • US9922431B2 patent drawing

AI summary

Approaches are described for rendering augmented reality overlays on an interface displaying the active field of view of a camera. The interface can display to a user an image or video, for example, and the overlay can be rendered over, near, or otherwise positioned with respect to any text or other such elements represented in the image. The overlay can have associated therewith at least one function or information, and when an input associated with the overlay is selected, the function can be performed (or caused to be performed) by the portable computing device.