Optical Character Recognition Text Extraction for Mobile Input

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for interacting with computing devices, such as entering phone numbers or web addresses, are tedious and time-consuming, and there is a need for more efficient ways to provide textual information to applications or services on portable devices.

Innovation Solution

A portable computing device is equipped with a camera that captures images of text, processes them using optical character recognition (OCR) algorithms, and identifies patterns like phone numbers, email addresses, or URLs, enabling the device to perform associated actions automatically or with user confirmation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual input methods are used for entering phone numbers, email addresses, or web addresses, then users have direct control over input accuracy, but the process becomes tedious and time-consuming

Engineering Contradiction:
Improvespeed of text inputVSAvoidtime required for manual typing
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the mechanical manual typing process with an optical recognition system. The camera captures images of text, and OCR algorithms automatically convert the visual text into digital input, eliminating the need for manual keyboard entry and dramatically reducing input time

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service text input by automatically capturing, recognizing, and processing text from images. The device performs the entire text extraction workflow autonomously without requiring user intervention for typing, allowing users to simply point and capture

Inventive Principle:
Principle #25Self-service

2Ease of operation

If optical character recognition is used to automatically recognize text from images, then manual input time is reduced, but the device processing complexity increases

Engineering Contradiction:
Improveconvenience of text inputVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent divides the text processing task into distinct modular components: image capture module, OCR processing module, pattern recognition module, and application execution module. This segmentation allows each component to be optimized independently and simplifies the overall system architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs a universal OCR engine that can recognize multiple text formats (phone numbers, email addresses, URLs, physical addresses) and integrate with various applications (messaging, browsing, mapping). This multi-functional approach consolidates what would otherwise require separate specialized systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9916514B2Text recognition driven functionality
Publication Date: 2018.03.13 AMAZON TECH INC
  • US9916514B2 patent drawing
  • US9916514B2 patent drawing
  • US9916514B2 patent drawing

AI summary

Various approaches for providing textual information to an application, system, or service are disclosed. In particular, various embodiments enable a user to capture an image with a camera of a portable computing device. The computing device is capable of taking the image and processing it to recognize, identify, and/or isolate the text in order to forward the text to an application or function. The application or function can then utilize the text to perform an action in substantially real-time. The text may include an email, phone number, URL, an address, and the like and the application or function may be dialing the phone number, navigating to the URL, opening an address book to save contact information, displaying a map to show the address, and so on.