Document Structure Extraction via Layout Position Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting reference information from digital layout files are inefficient, relying on metadata extraction which is not effective.
Innovation Solution
A method and apparatus that determine the position of reference information in a layout file and directly extract items related to it, bypassing metadata extraction by using predefined keywords and analyzing layout structures to identify start and end pages, and classify text blocks into rows and columns for accurate extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If metadata extraction methods are used to obtain reference information from digital layout files, then the extraction process can be automated, but the extraction efficiency is low and processing time is excessive
Solution Approach 1:
The patent extracts reference information directly from the layout file structure by identifying specific positional relationships between text blocks, rather than extracting through metadata. This involves detecting reference text blocks based on their spatial relationships with main text blocks and directly parsing the content from these identified blocks, thereby improving extraction efficiency while maintaining automation.
Solution Approach 2:
The patent introduces text block position information and spatial relationship analysis as an intermediary mechanism between the layout file and reference information extraction. By using positional coordinates and structural relationships of text blocks as intermediaries, the system can efficiently locate and extract reference information without relying on slow metadata extraction processes.
2Adaptability or versatility
If metadata extraction is used to obtain reference information, then the process can handle various document formats, but the accuracy in identifying reference items is insufficient
Solution Approach 1:
The patent applies local quality analysis by examining specific local structural characteristics of text blocks in the layout file. It identifies reference information based on local positional relationships (such as text blocks positioned below or beside main text blocks) and local formatting patterns, thereby achieving high accuracy in reference item identification while maintaining adaptability to different document formats through pattern recognition.
Solution Approach 2:
The patent performs preliminary analysis of the layout file structure to identify text blocks that contain reference information before actual extraction. This preliminary action involves detecting positional relationships and structural patterns in advance, which improves the accuracy of reference item identification by pre-filtering and pre-classifying potential reference text blocks based on their spatial and structural characteristics.
3Ease of manufacture
If template methods are used for extraction, then the process follows a standardized approach, but the processing time is excessive and efficiency is low
Solution Approach 1:
The patent segments the extraction process into distinct operational stages: identifying main text blocks, detecting reference text blocks based on positional relationships, classifying reference items, and extracting content. This segmentation allows each stage to be optimized independently and executed efficiently, reducing overall processing time while maintaining a standardized approach through consistent application of positional relationship rules across all segments.
Solution Approach 2:
The patent replaces traditional metadata-based extraction mechanisms with a direct structural analysis approach that uses text block position information and spatial relationships. This substitution eliminates the intermediate metadata extraction step, directly parsing reference information from the layout file structure, thereby significantly reducing processing time while maintaining standardization through algorithmic positional analysis.
Data Source
AI summary
A method for extracting a document structure is disclosed. The method may include determining a position of reference information in a layout file, and extracting items related to the reference information from the determined position of the layout file. An apparatus for extracting a document structure is also disclosed. The apparatus may include a processor configured to determine a position of reference information in a layout file; and to extract items related to the reference information from the determined position of the layout file. The apparatus may further include a storage device configured to store the extracted items.


