Document Image Binarization via Content-Type Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Document images containing multiple types of content, such as text, graphics, and photos, are difficult to binarize effectively using existing methods, which often result in unsatisfactory results and high computation costs.
Innovation Solution
A method that divides the document image into sub-images, determines the type of each sub-image based on horizontal projection profiles and density calculations, and applies specific binarization processes tailored to each content type, allowing for separate and efficient binarization of text, graphics, and photos.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If global or local thresholding approaches are used for binarization, then the binarization process is simple, but the results are unsatisfactory for composite images containing photos and the computation cost is high
Solution Approach 1:
The patent divides the document image into multiple sub-images and categorizes each sub-image according to its content type (text-only, photo-only, or mixed). This segmentation allows different binarization strategies to be applied to different regions, improving overall binarization quality while managing computational complexity through targeted processing.
Solution Approach 2:
The patent applies different binarization methods to different content types within the document image. Text regions use one binarization approach optimized for character recognition, while photo regions use another approach optimized for image quality, achieving local optimization of binarization quality for each content type.
2Manufacturing precision
If content-type-specific binarization processes are applied to different sub-images, then binarization quality is improved, but the processing complexity increases
Solution Approach 1:
The patent segments the document image into content-type-specific sub-images and applies dedicated binarization processes to each type. This segmentation strategy improves binarization quality by optimizing processing for each content type while managing complexity through systematic categorization and targeted processing of homogeneous regions.
3Manufacturing precision
If advanced binarization methods are used for composite images, then binarization quality improves, but the computation cost increases
Solution Approach 1:
The patent segments the document image into homogeneous content regions and applies appropriate binarization methods only where needed. This reduces computation cost by avoiding unnecessary complex processing in uniform regions while maintaining high binarization quality in challenging mixed-content areas through targeted advanced methods.
Solution Approach 2:
The patent applies computationally intensive advanced binarization methods only to specific content types that require them (such as mixed text-photo regions), while using simpler methods for homogeneous regions. This local optimization strategy improves binarization quality where needed while minimizing overall computation cost.
Data Source
AI summary
A method for binarizing a grayscale document image, which first divides the document image into a plurality of sub-images and determining a type of each sub-image based on a horizontal projection profile and a density of each sub-image, the type being 1: text only, 2: graphics only, 3: photo only, 4: text and graphics, 5: text and photo, 6: graphics and photo, or 7: text and graphics and photo. Then a selected one of first to seventh binarization processes is applied to binarize each sub-image based on its type to generate a binary sub-image. All binary sub-images are then combine to generate a binary image of the grayscale document image. Of the first to seventh binarization processes respectively applied to the first to seventh types of sub-images, at least those for the first, second, third, fifth, sixth and seventh type are different from each other.


