Document Semantic Representation Through Text-Layout Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document AI models struggle with accurately capturing semantic information from document images due to focusing solely on text-level manipulation, leading to inefficiencies and errors in tasks like document image classification and form understanding.
Innovation Solution
A solution that jointly processes both textual and layout information from document images using a deep learning model to generate semantic feature representations, incorporating spatial arrangements of text elements to enhance understanding and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing document AI models focus solely on text-level manipulation, then the processing speed is maintained, but the accuracy of capturing semantic information deteriorates
Solution Approach 1:
The patent segments the document processing into two distinct streams: a text processing stream that handles textual information and a layout processing stream that handles spatial arrangement information. These segmented streams are processed separately and then integrated, allowing the model to capture both semantic meaning and structural context without excessive complexity
Solution Approach 2:
The patent transitions from traditional one-dimensional text-only processing to two-dimensional processing by incorporating layout information (spatial positions, coordinates, and visual structure). This dimensional expansion enables the model to capture semantic information more accurately by considering both what the text says and where it is positioned in the document
2Reliability
If existing models use only textual information, then the computational resources are conserved, but the understanding of document structure deteriorates
Solution Approach 1:
The patent merges text information and layout information into a unified processing framework. The text processing stream and layout processing stream are combined through integration mechanisms that allow the model to leverage both textual content and spatial structure, improving document structure understanding while managing computational resources efficiently through shared components
3Measurement precision
If traditional text recognition approaches are used, then the processing time is reduced, but the semantic feature extraction accuracy deteriorates
Solution Approach 1:
The patent performs preliminary processing of layout information to extract spatial features and structural relationships before the main semantic analysis. By pre-processing and organizing layout data (such as determining hierarchical relationships and spatial groupings), the model prepares the structural context in advance, which accelerates the subsequent semantic feature extraction while maintaining high accuracy
Data Source
AI summary
There is provided a solution for semantic representation of text in a document. In this solution, textual information comprising a sequence of text elements (220) and layout information (230) of the text element are determined from a document. The layout information (230) indicates a spatial arrangement of the plurality of text elements (220) presented within the document. Based at least in part on the plurality of text elements (220) and the layout information (230), respective semantic feature representations (180) of the plurality of text elements (220) are generated. By jointly using both the textual information and the layout information (230), rich semantics of the text elements (220) in the document can be effectively captured in the feature representations.


