Text Recognition for Search Results Using OCR and NLP Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for providing users with information through computing devices are inefficient, as they require manual input and struggle to filter relevant information from images, leading to tedious and time-consuming processes, especially when dealing with large volumes of text containing irrelevant data.
Innovation Solution
The system employs optical character recognition (OCR) and Neuro-linguistic programming (NLP) techniques to analyze images, filter out irrelevant text, and prioritize relevant words for product searches, using image information such as position and size to determine the importance of recognized text, and correct OCR errors, ultimately forming a coherent query for product search engines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual input methods are used for search queries, then input accuracy is maintained, but time consumption increases and user convenience deteriorates
Solution Approach 1:
The patent replaces manual mechanical input (typing, clicking) with automated optical recognition. The camera captures images of products, and the system automatically extracts text and identifies products through image processing and database matching, eliminating the need for users to manually input search queries.
Solution Approach 2:
The system performs self-service by automatically analyzing captured images, extracting relevant information, and generating search queries without user intervention. The automated text extraction and product identification processes occur independently, allowing the system to serve itself rather than requiring continuous manual input.
2Loss of information
If all text in captured images is processed, then information completeness is improved, but processing time increases and relevance filtering becomes difficult
Solution Approach 1:
The patent extracts only the relevant text information from captured images using optical character recognition. Instead of processing all text, the system selectively extracts text from product labels, packaging, and visible product information, filtering out irrelevant text such as background elements or non-product-related content.
Solution Approach 2:
The system applies different processing quality to different regions of the captured image. Text regions are processed with high accuracy using OCR, while non-text regions are processed differently or not at all. The system identifies and prioritizes text elements based on their location and characteristics within the image.
3Device complexity
If basic text recognition is used, then implementation simplicity is maintained, but text accuracy deteriorates especially for similar characters
Solution Approach 1:
The patent incorporates feedback mechanisms where the system uses contextual information from the image to verify and correct text recognition results. The captured image context is used to validate extracted text, and the system can iteratively refine its recognition accuracy by comparing against known product characteristics and database information.
Solution Approach 2:
The system combines multiple text recognition approaches and validation methods to achieve high accuracy. It integrates optical character recognition with image processing techniques, contextual analysis, and database verification to create a composite system that overcomes the limitations of any single method.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables users to quickly and accurately identify products from images, providing relevant search results by isolating key terms and phrases, thus enhancing the efficiency of product searches on portable computing devices.
Implementation Method 1
an optical character recognition (OCR) engine and methods, as set forth herein, can automatically attempt to select the most relevant words associated with products available for purchase from an electronic marketplace
Data Source
AI summary
Various embodiments enable a process to automatically attempt to select the most relevant words associated with products available for purchase from an electronic marketplace from an image frame. For example, an image frame containing text can be obtained and analyzed with an optical character recognition. The recognized words can then be preprocessed using various filtering and scoring techniques to narrow down a volume of text to a few relevant query terms. These query terms can then be sent to a search engine associated with the electronic marketplace to return relevant products to a user.


