Document Image Segmentation for Efficient Content Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for capturing and processing paper documents into electronic format are inefficient, leading to time-consuming error correction, delays, and unnecessary processing of non-relevant content, especially when outsourcing to third-party service bureaus, and do not allow for online indexing and content filtering.

Innovation Solution

A method and system that enables users to segment document images, define metadata fields, and apply filtration criteria to exclude non-relevant objects, allowing for online indexing and efficient content retrieval, with interactive user input for error correction and content management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If full-text indexing is performed on all document objects, then complete search capability is achieved, but processing time and resources are excessively consumed

Engineering Contradiction:
Improvesearch capabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the document image into multiple regions of interest (ROIs) based on user-defined criteria such as text density, object types, or semantic categories. Only these segmented regions are subjected to full-text indexing and detailed processing, while other regions are excluded or processed at a lower level. This segmentation approach maintains search capability for relevant content while dramatically reducing the overall processing scope and time consumption.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If all image objects are processed for metadata extraction, then comprehensive information is obtained, but processing efficiency decreases due to inclusion of non-relevant content

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent extracts and processes only the relevant portions of the document image that contain meaningful information for the user's specific needs. By applying filtration criteria to identify and exclude non-relevant objects and regions, the system obtains sufficient information for effective search and retrieval without the overhead of processing the entire document. This extraction approach balances information completeness with processing efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If user interaction is enabled for content filtering and error correction, then processing accuracy improves, but system complexity increases

Engineering Contradiction:
Improveprocessing accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent incorporates feedback mechanisms where users can interact with the system to provide corrections, confirmations, or adjustments to the automatically extracted metadata and identified regions of interest. This feedback loop allows the system to learn from user interactions and improve its accuracy over time. The feedback is integrated into the existing processing pipeline in a modular manner, allowing accuracy improvement without proportionally increasing overall system complexity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8532384B2Method of retrieving information from a digital image
Publication Date: 2013.09.10 HOWIE CAMERON TELFER
  • US8532384B2 patent drawing
  • US8532384B2 patent drawing
  • US8532384B2 patent drawing

AI summary

The method includes segmenting a document image to identify image objects within the document image and applying an automatic algorithm to the image objects so as to assign initial metadata to the image objects. The method further includes selecting image objects from the set of image objects whose metadata satisfy filtration criteria so as to exclude those selected image objects from processing and processing a rest of the image objects in the set of image objects. The method further includes presenting the document image, image objects, and metadata to the user and enabling input from the user to manage the image objects, subsets, and metadata.