PDF Clarity Improvement via Rectangular Object Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The clarity of PDF documents converted from paper documents is poor due to retained impurity data from the scanning process, affecting reading experiences, especially on small screens.
Innovation Solution
A method involving scanning, margin data removal, dividing into rectangular objects, and filtering out invalid data to create a tailored PDF document that retains only effective regions, improving clarity by reducing impurity data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If the paper document is scanned to create a PDF document using conventional methods, then the conversion process is simple and fast, but the resulting PDF document contains impurity data that degrades clarity and reading experience
Solution Approach 1:
The patent segments the PDF document into multiple rectangular objects (text blocks, images, tables, etc.) and processes each object individually to identify and remove impurity data. This segmentation allows precise control over which data to retain and which to remove, thereby improving clarity while managing complexity through systematic processing of discrete elements.
Solution Approach 2:
The patent extracts and removes impurity data from specific regions (margins, headers, footers, and other non-essential areas) of the PDF document. By identifying and taking out only the harmful impurity data while preserving the essential content, the method improves document clarity without requiring complete reconstruction of the document.
2Ease of operation
If impurity data is retained during format conversion, then the conversion process is simple, but the reading experience deteriorates especially on small screens
Solution Approach 1:
The patent applies different processing quality standards to different regions of the PDF document. Essential content areas are preserved with high fidelity, while margin areas and non-essential regions are cleaned of impurity data. This local quality approach ensures that reading experience is improved in critical areas without unnecessarily affecting other parts of the document.
Solution Approach 2:
The patent identifies impurity data that originally degraded reading experience and transforms the situation by systematically removing these harmful elements. The process converts the harmful presence of impurity data into a benefit by using automated detection and removal mechanisms that improve overall document quality and readability, especially on small screens where impurity data had the most negative impact.
Data Source
AI summary
The present invention relates to a method for improving clarity of a PDF file converted from a paper file. The method comprises: step 1: scanning a paper file to obtain an electronic image file; step 2: determining a top margin, a bottom margin, a left margin, and a right margin of the electronic image file, and deleting data within the top margin, the bottom margin, the left margin, and the right margin of the electronic image file, and converting a first cropped file into a first PDF file; step 3: dividing the first PDF file into some rectangular objects, determining effective areas of the rectangular objects, and deleting data outside effective areas of the rectangular objects, so as to obtain cropped rectangular objects corresponding to the rectangular objects in a one-to-one manner; and step 4: combining the cropped rectangular objects according to same position distribution of the rectangular objects corresponding to the cropped rectangular objects on the first PDF file, so as to obtain and output a second PDF file. By means of the present invention, clarity of a PDF file converted from a paper file can be improved.

