Gesture-Based Text Region Isolation for OCR Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing methods using Optical Character Recognition (OCR) are inefficient and error-prone when trying to recognize specific text in images, as they process large images with abundant textual and graphical information, hiding the desired text.
Innovation Solution
A method and system that uses a user's underline gesture on a touch-sensitive device to selectively recognize text, involving image display, gesture detection, text region identification, and OCR processing, allowing the user to highlight the text region of interest for accurate extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional OCR technology is used to process the entire image, then all text in the image can be recognized, but the processing time increases and accuracy decreases due to the large amount of data
Solution Approach 1:
The patent divides the large image into multiple smaller sub-images or regions of interest. By segmenting the image, the system processes only relevant portions containing text, significantly reducing processing time while maintaining recognition accuracy. This is achieved through gesture-based selection of specific areas followed by OCR processing of those segmented regions.
Solution Approach 2:
The patent extracts only the necessary text-containing regions from the full image using gesture input. Instead of processing the entire large image, the system identifies and extracts specific areas of interest where text is located, reducing the data volume for OCR processing while ensuring accurate recognition of the desired text.
2Loss of information
If traditional OCR technology processes the entire image, then comprehensive text recognition is achieved, but the desired text is hidden among other recognized text
Solution Approach 1:
The patent extracts only the specific text regions of interest from the full image through gesture-based selection. This extraction principle ensures that only the desired text is processed and output, eliminating irrelevant text from the results. The user can precisely select which text to extract, improving both selection accuracy and extraction efficiency.
Solution Approach 2:
The patent applies different processing qualities to different regions of the image. Areas containing text of interest receive full OCR processing attention, while other areas are either excluded or processed with lower priority. This local quality approach ensures that the desired text is accurately recognized and extracted without being lost among other text, while optimizing overall processing efficiency.
3Productivity
If the user manually selects text regions, then processing efficiency improves, but the operation becomes more complex
Solution Approach 1:
The patent replaces complex manual text region selection with intuitive gesture-based control. Instead of requiring users to manually draw boundaries or use multiple steps to select text areas, simple gestures (such as tapping or swiping) automatically identify and select text regions. This substitution maintains high processing efficiency while dramatically improving ease of operation.
Solution Approach 2:
The system automatically identifies and processes text regions based on simple user gestures, performing the complex task of text region selection without requiring detailed user input. The gesture-based interface allows users to initiate text selection with minimal effort, and the system handles the complex work of identifying and processing the appropriate text areas, improving both efficiency and ease of use.
Data Source
AI summary
An image is displayed on a touch screen. A user's underline gesture on the displayed image is detected. The area of the image touched by the underline gesture and a surrounding region approximate to the touched area are identified. Skew for text in the surrounding region is determined and compensated. A text region including the text is identified in the surrounding region and cropped from the image. The cropped image is transmitted to an optical character recognition (OCR) engine, which processes the cropped image and returns OCR'ed text. The OCR'ed text is outputted.


