Text Recognition for Search Results Using OCR and NLP Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for providing users with information through computing devices are inefficient, as they require manual input and struggle to filter relevant information from images, leading to tedious and time-consuming processes, especially when dealing with large volumes of text containing irrelevant data.

Innovation Solution

The system employs optical character recognition (OCR) and Neuro-linguistic programming (NLP) techniques to analyze images, filter out irrelevant text, and prioritize relevant words for product searches, using image information such as position and size to determine the importance of recognized text, and correct OCR errors, ultimately forming a coherent query for product search engines.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual input methods are used for search queries, then input accuracy is maintained, but time consumption increases and user convenience deteriorates

Engineering Contradiction:
Improveuser convenienceVSAvoidtime consumption
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical input (typing, clicking) with automated optical recognition. The camera captures images of products, and the system automatically extracts text and identifies products through image processing and database matching, eliminating the need for users to manually input search queries.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically analyzing captured images, extracting relevant information, and generating search queries without user intervention. The automated text extraction and product identification processes occur independently, allowing the system to serve itself rather than requiring continuous manual input.

Inventive Principle:
Principle #25Self-service

2Loss of information

If all text in captured images is processed, then information completeness is improved, but processing time increases and relevance filtering becomes difficult

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts only the relevant text information from captured images using optical character recognition. Instead of processing all text, the system selectively extracts text from product labels, packaging, and visible product information, filtering out irrelevant text such as background elements or non-product-related content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different processing quality to different regions of the captured image. Text regions are processed with high accuracy using OCR, while non-text regions are processed differently or not at all. The system identifies and prioritizes text elements based on their location and characteristics within the image.

Inventive Principle:
Principle #3Local quality

3Device complexity

If basic text recognition is used, then implementation simplicity is maintained, but text accuracy deteriorates especially for similar characters

Engineering Contradiction:
Improveimplementation simplicityVSAvoidtext recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent incorporates feedback mechanisms where the system uses contextual information from the image to verify and correct text recognition results. The captured image context is used to validate extracted text, and the system can iteratively refine its recognition accuracy by comparing against known product characteristics and database information.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system combines multiple text recognition approaches and validation methods to achieve high accuracy. It integrates optical character recognition with image processing techniques, contextual analysis, and database verification to create a composite system that overcomes the limitations of any single method.

Inventive Principle:
Principle #40Composite materials

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables users to quickly and accurately identify products from images, providing relevant search results by isolating key terms and phrases, thus enhancing the efficiency of product searches on portable computing devices.

Implementation Method 1

an optical character recognition (OCR) engine and methods, as set forth herein, can automatically attempt to select the most relevant words associated with products available for purchase from an electronic marketplace

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS9934526B1Text recognition for search results
Publication Date: 2018.04.03 AMAZON TECH INC
  • US9934526B1 patent drawing
  • US9934526B1 patent drawing
  • US9934526B1 patent drawing

AI summary

Various embodiments enable a process to automatically attempt to select the most relevant words associated with products available for purchase from an electronic marketplace from an image frame. For example, an image frame containing text can be obtained and analyzed with an optical character recognition. The recognized words can then be preprocessed using various filtering and scoring techniques to narrow down a volume of text to a few relevant query terms. These query terms can then be sent to a search engine associated with the electronic marketplace to return relevant products to a user.