Paragraph Detection Across Page and Column Boundaries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document processing technologies struggle to accurately identify and tag paragraphs that span across columns, pages, or other reading units, leading to incomplete or incorrect grouping of text.

Innovation Solution

A method and apparatus that involve obtaining a set of candidate paragraphs, identifying pairs of candidate paragraphs that span columns, and outputting tagged paragraphs corresponding to these identified pairs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional document processing methods are used to identify paragraphs, then the process is simple, but the accuracy of identifying paragraphs spanning columns or pages is poor

Engineering Contradiction:
Improveparagraph identification accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the paragraph identification process into multiple distinct stages: obtaining candidate paragraphs, determining their spatial relationships, identifying spanning pairs, and tagging them. This segmentation allows each stage to be optimized independently, improving overall accuracy while managing complexity through modular processing steps.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by first obtaining a set of candidate paragraphs and analyzing their spatial relationships before final identification. This preliminary analysis of candidate positions and overlaps enables more accurate identification of spanning paragraphs, as the system prepares and filters candidates before making final determinations.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If advanced methods are used to accurately identify spanning paragraphs, then the identification accuracy improves, but the processing complexity increases

Engineering Contradiction:
Improveparagraph grouping reliabilityVSAvoiddetection algorithm complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent incorporates feedback mechanisms where the system evaluates candidate paragraph pairs, determines their spatial relationships, and uses this information to refine identifications of spanning paragraphs. The feedback loop allows the system to adjust its identification criteria based on the analyzed spatial patterns, improving reliability while managing complexity through iterative refinement.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces intermediary concepts such as candidate paragraph positions, spatial relationship descriptors, and tagging mechanisms that mediate between raw document data and final paragraph identifications. These intermediaries simplify the complex task of detecting spanning paragraphs by breaking it down into manageable computational steps involving position comparison and relationship classification.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If basic paragraph detection is used, then the processing is fast, but paragraphs spanning columns or pages are incorrectly grouped

Engineering Contradiction:
Improvetext grouping precisionVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent applies partial action by focusing computational resources only on candidate paragraph pairs that exhibit spatial characteristics of spanning paragraphs. Rather than analyzing all possible paragraph combinations, the system identifies and processes only those candidates with relevant spatial relationships, achieving high precision without proportionally increasing processing time across the entire document.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12293143B2Detection and tagging of paragraphs spanning columns, pages, or other reading units
Publication Date: 2025.05.06 KONICA MINOLTA BUSINESS SOLUTIONS USA INC
  • US12293143B2 patent drawing
  • US12293143B2 patent drawing
  • US12293143B2 patent drawing

AI summary

Systems, methods, apparatuses, and computer program products for detecting and tagging of paragraphs that span columns, pages, or other reading units are provided. For example, a method can include obtaining a set of candidate paragraphs for a document. The method can also include identifying a pair of candidate paragraphs spanning columns from among the set of candidate paragraphs. The method can further include outputting a tagged paragraph corresponding to the pair of candidate paragraphs spanning columns.