Automated Highlighted Text Extraction in Scanned Documents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting and extracting highlighted regions in documents are inefficient, requiring manual transcription and are time-consuming, especially in the legal industry where differentiation between highlighted and non-highlighted regions is necessary.
Innovation Solution
A method and system using color masking and multi-mask compression technology, combined with Optical Character Recognition (OCR), to automatically extract and recognize highlighted regions in scanned text and image documents, reducing manual effort by converting images into binary form and applying morphological operations to enhance extraction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual transcription of highlighted portions is used, then extraction accuracy can be maintained, but time consumption increases significantly
Solution Approach 1:
The patent replaces the mechanical manual transcription process with an automated image processing system that uses color space conversion, thresholding, and morphological operations to detect and extract highlighted regions, thereby eliminating time-consuming manual work while maintaining extraction accuracy
Solution Approach 2:
The system enables self-service extraction by allowing the document processing system to automatically identify, segment, and extract highlighted portions without human intervention, using automated color-based detection and image processing algorithms
2Difficulty of detecting and measuring
If color scanner is used to detect highlighted regions, then detection capability is improved, but processing complexity increases
Solution Approach 1:
The patent transforms the color image into different color spaces (HSV, LAB, YCbCr) and applies thresholding operations to convert continuous color values into discrete binary masks, simplifying the detection process by changing the parameter representation from continuous RGB values to segmented binary regions
Solution Approach 2:
The patent segments the document image into different regions (highlighted vs. non-highlighted) by applying color-based thresholding and morphological operations, separating the highlighted portions from the rest of the document for independent processing
3Productivity
If mid-tone portion is screened with low frequency screen to convert to binary form, then extraction speed is improved, but precision is reduced
Solution Approach 1:
The patent applies adaptive morphological operations that dynamically adjust to the local characteristics of the image, using operations like closing and opening to refine the binary masks while preserving the boundaries of highlighted regions, thereby maintaining precision during the binary conversion process
Data Source
AI summary
An automated method for extracting highlighted regions in a scanned text documents includes color masking of highlight regions, extracting text from highlighted regions, recognizing the characters in extracted text optically and inserting the recognized characters to new document in order to easily identify highlighted text in scanned images. Using a two-layer multi-mask compression technology configured in a scanned export image path, edges and text regions can be extracted and together with the use of mask coordinates and associated mask colors, all highlighted texts can be easily identified and extracted. Optical Character Recognition (OCR) can then be utilized to appropriate summarization of different extracted highlighted texts.


