Document Image Binarization via Content-Type Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Document images containing multiple types of content, such as text, graphics, and photos, are difficult to binarize effectively using existing methods, which often result in unsatisfactory results and high computation costs.

Innovation Solution

A method that divides the document image into sub-images, determines the type of each sub-image based on horizontal projection profiles and density calculations, and applies specific binarization processes tailored to each content type, allowing for separate and efficient binarization of text, graphics, and photos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If global or local thresholding approaches are used for binarization, then the binarization process is simple, but the results are unsatisfactory for composite images containing photos and the computation cost is high

Engineering Contradiction:
Improvebinarization qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent divides the document image into multiple sub-images and categorizes each sub-image according to its content type (text-only, photo-only, or mixed). This segmentation allows different binarization strategies to be applied to different regions, improving overall binarization quality while managing computational complexity through targeted processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different binarization methods to different content types within the document image. Text regions use one binarization approach optimized for character recognition, while photo regions use another approach optimized for image quality, achieving local optimization of binarization quality for each content type.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If content-type-specific binarization processes are applied to different sub-images, then binarization quality is improved, but the processing complexity increases

Engineering Contradiction:
Improvebinarization qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the document image into content-type-specific sub-images and applies dedicated binarization processes to each type. This segmentation strategy improves binarization quality by optimizing processing for each content type while managing complexity through systematic categorization and targeted processing of homogeneous regions.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If advanced binarization methods are used for composite images, then binarization quality improves, but the computation cost increases

Engineering Contradiction:
Improvebinarization qualityVSAvoidcomputation cost
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the document image into homogeneous content regions and applies appropriate binarization methods only where needed. This reduces computation cost by avoiding unnecessary complex processing in uniform regions while maintaining high binarization quality in challenging mixed-content areas through targeted advanced methods.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies computationally intensive advanced binarization methods only to specific content types that require them (such as mixed text-photo regions), while using simpler methods for homogeneous regions. This local optimization strategy improves binarization quality where needed while minimizing overall computation cost.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9965695B1Document image binarization method based on content type separation
Publication Date: 2018.05.08 KONICA MINOLTA SYSTEMS LABORATORY INC
  • US9965695B1 patent drawing
  • US9965695B1 patent drawing
  • US9965695B1 patent drawing

AI summary

A method for binarizing a grayscale document image, which first divides the document image into a plurality of sub-images and determining a type of each sub-image based on a horizontal projection profile and a density of each sub-image, the type being 1: text only, 2: graphics only, 3: photo only, 4: text and graphics, 5: text and photo, 6: graphics and photo, or 7: text and graphics and photo. Then a selected one of first to seventh binarization processes is applied to binarize each sub-image based on its type to generate a binary sub-image. All binary sub-images are then combine to generate a binary image of the grayscale document image. Of the first to seventh binarization processes respectively applied to the first to seventh types of sub-images, at least those for the first, second, third, fifth, sixth and seventh type are different from each other.