Document Structure Extraction via Layout Position Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for extracting reference information from digital layout files are inefficient, relying on metadata extraction which is not effective.

Innovation Solution

A method and apparatus that determine the position of reference information in a layout file and directly extract items related to it, bypassing metadata extraction by using predefined keywords and analyzing layout structures to identify start and end pages, and classify text blocks into rows and columns for accurate extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If metadata extraction methods are used to obtain reference information from digital layout files, then the extraction process can be automated, but the extraction efficiency is low and processing time is excessive

Engineering Contradiction:
Improveautomation of reference information extractionVSAvoidextraction efficiency
Core Design Contradiction:
Extent of automationVSProductivity

Solution Approach 1:

The patent extracts reference information directly from the layout file structure by identifying specific positional relationships between text blocks, rather than extracting through metadata. This involves detecting reference text blocks based on their spatial relationships with main text blocks and directly parsing the content from these identified blocks, thereby improving extraction efficiency while maintaining automation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces text block position information and spatial relationship analysis as an intermediary mechanism between the layout file and reference information extraction. By using positional coordinates and structural relationships of text blocks as intermediaries, the system can efficiently locate and extract reference information without relying on slow metadata extraction processes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If metadata extraction is used to obtain reference information, then the process can handle various document formats, but the accuracy in identifying reference items is insufficient

Engineering Contradiction:
Improvehandling of various document formatsVSAvoidaccuracy of reference item identification
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies local quality analysis by examining specific local structural characteristics of text blocks in the layout file. It identifies reference information based on local positional relationships (such as text blocks positioned below or beside main text blocks) and local formatting patterns, thereby achieving high accuracy in reference item identification while maintaining adaptability to different document formats through pattern recognition.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs preliminary analysis of the layout file structure to identify text blocks that contain reference information before actual extraction. This preliminary action involves detecting positional relationships and structural patterns in advance, which improves the accuracy of reference item identification by pre-filtering and pre-classifying potential reference text blocks based on their spatial and structural characteristics.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If template methods are used for extraction, then the process follows a standardized approach, but the processing time is excessive and efficiency is low

Engineering Contradiction:
Improvestandardization of extraction processVSAvoidprocessing time
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The patent segments the extraction process into distinct operational stages: identifying main text blocks, detecting reference text blocks based on positional relationships, classifying reference items, and extracting content. This segmentation allows each stage to be optimized independently and executed efficiently, reducing overall processing time while maintaining a standardized approach through consistent application of positional relationship rules across all segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces traditional metadata-based extraction mechanisms with a direct structural analysis approach that uses text block position information and spatial relationships. This substitution eliminates the intermediate metadata extraction step, directly parsing reference information from the layout file structure, thereby significantly reducing processing time while maintaining standardization through algorithmic positional analysis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9418051B2Methods and devices for extracting document structure
Publication Date: 2016.08.16 NEW FOUNDER HLDG DEV LLC
  • US9418051B2 patent drawing
  • US9418051B2 patent drawing
  • US9418051B2 patent drawing

AI summary

A method for extracting a document structure is disclosed. The method may include determining a position of reference information in a layout file, and extracting items related to the reference information from the determined position of the layout file. An apparatus for extracting a document structure is also disclosed. The apparatus may include a processor configured to determine a position of reference information in a layout file; and to extract items related to the reference information from the determined position of the layout file. The apparatus may further include a storage device configured to store the extracted items.