Paragraph Detection Across Page and Column Boundaries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document processing technologies struggle to accurately identify and tag paragraphs that span across columns, pages, or other reading units, leading to incomplete or incorrect grouping of text.
Innovation Solution
A method and apparatus that involve obtaining a set of candidate paragraphs, identifying pairs of candidate paragraphs that span columns, and outputting tagged paragraphs corresponding to these identified pairs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional document processing methods are used to identify paragraphs, then the process is simple, but the accuracy of identifying paragraphs spanning columns or pages is poor
Solution Approach 1:
The patent segments the paragraph identification process into multiple distinct stages: obtaining candidate paragraphs, determining their spatial relationships, identifying spanning pairs, and tagging them. This segmentation allows each stage to be optimized independently, improving overall accuracy while managing complexity through modular processing steps.
Solution Approach 2:
The patent performs preliminary actions by first obtaining a set of candidate paragraphs and analyzing their spatial relationships before final identification. This preliminary analysis of candidate positions and overlaps enables more accurate identification of spanning paragraphs, as the system prepares and filters candidates before making final determinations.
2Reliability
If advanced methods are used to accurately identify spanning paragraphs, then the identification accuracy improves, but the processing complexity increases
Solution Approach 1:
The patent incorporates feedback mechanisms where the system evaluates candidate paragraph pairs, determines their spatial relationships, and uses this information to refine identifications of spanning paragraphs. The feedback loop allows the system to adjust its identification criteria based on the analyzed spatial patterns, improving reliability while managing complexity through iterative refinement.
Solution Approach 2:
The patent introduces intermediary concepts such as candidate paragraph positions, spatial relationship descriptors, and tagging mechanisms that mediate between raw document data and final paragraph identifications. These intermediaries simplify the complex task of detecting spanning paragraphs by breaking it down into manageable computational steps involving position comparison and relationship classification.
3Manufacturing precision
If basic paragraph detection is used, then the processing is fast, but paragraphs spanning columns or pages are incorrectly grouped
Solution Approach 1:
The patent applies partial action by focusing computational resources only on candidate paragraph pairs that exhibit spatial characteristics of spanning paragraphs. Rather than analyzing all possible paragraph combinations, the system identifies and processes only those candidates with relevant spatial relationships, achieving high precision without proportionally increasing processing time across the entire document.
Data Source
AI summary
Systems, methods, apparatuses, and computer program products for detecting and tagging of paragraphs that span columns, pages, or other reading units are provided. For example, a method can include obtaining a set of candidate paragraphs for a document. The method can also include identifying a pair of candidate paragraphs spanning columns from among the set of candidate paragraphs. The method can further include outputting a tagged paragraph corresponding to the pair of candidate paragraphs spanning columns.


