Document Semantic Representation Through Text-Layout Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document AI models struggle with accurately capturing semantic information from document images due to focusing solely on text-level manipulation, leading to inefficiencies and errors in tasks like document image classification and form understanding.

Innovation Solution

A solution that jointly processes both textual and layout information from document images using a deep learning model to generate semantic feature representations, incorporating spatial arrangements of text elements to enhance understanding and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing document AI models focus solely on text-level manipulation, then the processing speed is maintained, but the accuracy of capturing semantic information deteriorates

Engineering Contradiction:
Improveaccuracy of capturing semantic informationVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the document processing into two distinct streams: a text processing stream that handles textual information and a layout processing stream that handles spatial arrangement information. These segmented streams are processed separately and then integrated, allowing the model to capture both semantic meaning and structural context without excessive complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from traditional one-dimensional text-only processing to two-dimensional processing by incorporating layout information (spatial positions, coordinates, and visual structure). This dimensional expansion enables the model to capture semantic information more accurately by considering both what the text says and where it is positioned in the document

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If existing models use only textual information, then the computational resources are conserved, but the understanding of document structure deteriorates

Engineering Contradiction:
Improvedocument structure understandingVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent merges text information and layout information into a unified processing framework. The text processing stream and layout processing stream are combined through integration mechanisms that allow the model to leverage both textual content and spatial structure, improving document structure understanding while managing computational resources efficiently through shared components

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If traditional text recognition approaches are used, then the processing time is reduced, but the semantic feature extraction accuracy deteriorates

Engineering Contradiction:
Improvesemantic feature extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of layout information to extract spatial features and structural relationships before the main semantic analysis. By pre-processing and organizing layout data (such as determining hierarchical relationships and spatial groupings), the model prepares the structural context in advance, which accelerates the subsequent semantic feature extraction while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12374141B2Semantic representation of text in document
Publication Date: 2025.07.29 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12374141B2 patent drawing
  • US12374141B2 patent drawing
  • US12374141B2 patent drawing

AI summary

There is provided a solution for semantic representation of text in a document. In this solution, textual information comprising a sequence of text elements (220) and layout information (230) of the text element are determined from a document. The layout information (230) indicates a spatial arrangement of the plurality of text elements (220) presented within the document. Based at least in part on the plurality of text elements (220) and the layout information (230), respective semantic feature representations (180) of the plurality of text elements (220) are generated. By jointly using both the textual information and the layout information (230), rich semantics of the text elements (220) in the document can be effectively captured in the feature representations.