Layout Text Block Extraction Using Spatial Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional information extraction methods for text with layout suffer from low accuracy due to the conversion of layout text into plain text, losing metadata and spatial information, which limits the effectiveness of extraction algorithms.
Innovation Solution
An information extraction method that utilizes feature information at a text block granularity, including metadata and spatial location, to directly process text with layout, enhancing accuracy and applicability across various layouts and line crossings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If text with layout is converted into text without layout for information extraction, then the extraction process becomes simpler, but accuracy of information extraction decreases due to loss of metadata and spatial information
Solution Approach 1:
The patent segments the text with layout into multiple text blocks while preserving their spatial relationships and metadata. Each text block is treated as an independent unit with associated feature information, allowing the extraction process to work on structured segments rather than converting to unstructured plain text. This maintains both operational simplicity and extraction accuracy.
Solution Approach 2:
The patent introduces a new dimensional approach by adding spatial location information and metadata as additional dimensions to the text data. Instead of reducing text with layout to one-dimensional plain text, the system maintains multi-dimensional information including position, size, and visual characteristics, enabling accurate extraction while preserving structural context.
2Ease of manufacture
If conventional text extraction methods are used on text with layout, then processing is easier, but feature information richness is insufficient leading to low extraction accuracy
Solution Approach 1:
The patent performs preliminary processing to extract and organize feature information from text blocks before the main extraction task. By pre-processing the text with layout to identify text blocks, their positions, sizes, and visual characteristics, the system prepares rich feature data in advance, making subsequent extraction more accurate without complicating the overall process.
Solution Approach 2:
The patent introduces an intermediary processing layer that bridges text with layout and extraction algorithms. This intermediary layer extracts feature information including spatial locations, metadata, and visual characteristics, transforming the input into a format that retains richness while being suitable for processing, thus preventing information loss.
3Productivity
If template-based methods are used for text with layout extraction, then extraction can be performed, but the method is limited to specific templates and lacks versatility
Solution Approach 1:
The patent creates a universal extraction method that works across different text with layout formats without requiring template-specific configurations. By using general feature extraction techniques that identify text blocks and their properties, the system can handle diverse layouts, formats, and structures, making the extraction capability both productive and highly adaptable.
Solution Approach 2:
The patent enables adaptability by dynamically adjusting extraction parameters based on the actual text with layout being processed. Instead of being constrained by fixed templates, the system can modify its extraction strategy according to the detected text block characteristics, spatial relationships, and content types, providing both productivity and versatility.
Data Source
AI summary
An information extraction method includes: determining that a text block that belongs to a target category and that is in text with layout is to be extracted; recognizing, based on feature information at a text block granularity, the text block that belongs to the target category and that is in the text with layout; and outputting an identifier of the text block that belongs to the target category and that is in the text with layout.


