Automated Highlighted Text Extraction in Scanned Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for detecting and extracting highlighted regions in documents are inefficient, requiring manual transcription and are time-consuming, especially in the legal industry where differentiation between highlighted and non-highlighted regions is necessary.

Innovation Solution

A method and system using color masking and multi-mask compression technology, combined with Optical Character Recognition (OCR), to automatically extract and recognize highlighted regions in scanned text and image documents, reducing manual effort by converting images into binary form and applying morphological operations to enhance extraction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual transcription of highlighted portions is used, then extraction accuracy can be maintained, but time consumption increases significantly

Engineering Contradiction:
Improveextraction accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical manual transcription process with an automated image processing system that uses color space conversion, thresholding, and morphological operations to detect and extract highlighted regions, thereby eliminating time-consuming manual work while maintaining extraction accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service extraction by allowing the document processing system to automatically identify, segment, and extract highlighted portions without human intervention, using automated color-based detection and image processing algorithms

Inventive Principle:
Principle #25Self-service

2Difficulty of detecting and measuring

If color scanner is used to detect highlighted regions, then detection capability is improved, but processing complexity increases

Engineering Contradiction:
Improvedetection capabilityVSAvoidprocessing complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The patent transforms the color image into different color spaces (HSV, LAB, YCbCr) and applies thresholding operations to convert continuous color values into discrete binary masks, simplifying the detection process by changing the parameter representation from continuous RGB values to segmented binary regions

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the document image into different regions (highlighted vs. non-highlighted) by applying color-based thresholding and morphological operations, separating the highlighted portions from the rest of the document for independent processing

Inventive Principle:
Principle #1Segmentation

3Productivity

If mid-tone portion is screened with low frequency screen to convert to binary form, then extraction speed is improved, but precision is reduced

Engineering Contradiction:
Improveextraction speedVSAvoidextraction precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies adaptive morphological operations that dynamically adjust to the local characteristics of the image, using operations like closing and opening to refine the binary masks while preserving the boundaries of highlighted regions, thereby maintaining precision during the binary conversion process

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8494280B2Automated method for extracting highlighted regions in scanned source
Publication Date: 2013.07.23 GENESEE VALLEY INNOVATIONS LLC
  • US8494280B2 patent drawing
  • US8494280B2 patent drawing
  • US8494280B2 patent drawing

AI summary

An automated method for extracting highlighted regions in a scanned text documents includes color masking of highlight regions, extracting text from highlighted regions, recognizing the characters in extracted text optically and inserting the recognized characters to new document in order to easily identify highlighted text in scanned images. Using a two-layer multi-mask compression technology configured in a scanned export image path, edges and text regions can be extracted and together with the use of mask coordinates and associated mask colors, all highlighted texts can be easily identified and extracted. Optical Character Recognition (OCR) can then be utilized to appropriate summarization of different extracted highlighted texts.