Intermediate Text Representations for Table Structure Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automated computer analysis of electronic documents, particularly tables in contracts, is error-prone and inefficient, leading to inaccurate extraction of relevant content, which is crucial for software asset management and licensing information.
Innovation Solution
A system that converts table structures into intermediate text representations and uses machine learning models to extract information of interest, improving the accuracy and efficiency of content extraction from electronic documents by analyzing text and structures, including tables, and storing the extracted information in a structured format.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated computer analysis is used to extract information from electronic documents, then productivity is improved, but measurement precision deteriorates due to error-prone analysis
Solution Approach 1:
The patent introduces intermediate representations as a mediator between the original table structure and the final extracted information. These intermediate representations serve as a bridge that captures both structural context and content meaning, allowing the system to maintain high extraction accuracy while achieving automated processing. The intermediate representations include encoded structural information (row, column, cell relationships) combined with semantic content, enabling the machine learning model to accurately interpret tabular data without manual analysis.
2Measurement precision
If manual analysis is used to extract information from documents, then measurement precision is maintained, but productivity deteriorates due to time-consuming processes
Solution Approach 1:
The patent replaces the mechanical manual analysis process with an automated system that uses machine learning models trained on intermediate representations. The system automatically processes electronic documents, converts tables to intermediate text representations, and extracts information without human intervention. This substitution maintains high extraction accuracy through sophisticated algorithms while dramatically improving productivity by eliminating time-consuming manual review processes.
3Productivity
If table structures are directly processed by machine learning models, then productivity is improved, but measurement precision deteriorates due to difficulty in interpreting tabular structures
Solution Approach 1:
The patent segments the table structure into discrete, meaningful components within the intermediate representations. Each representation captures specific structural elements (row identifiers, column headers, cell boundaries, relationships) separately from the content, allowing the machine learning model to process both structure and meaning independently and accurately. This segmentation enables the model to understand tabular relationships without confusion, maintaining high precision while achieving automated processing.
Data Source
AI summary
A document is received. The document is analyzed to discover text and structures of content included in the document. A result of the analysis is used to determine intermediate text representations of segments of the content included in the document, wherein at least one of the intermediate text representations includes an added text encoding the discovered structure of the corresponding content segment within a structural layout of the document. The intermediate text representations are used as an input to a machine learning model to extract information of interest in the document. One or more structured records of the extracted information of interest are created.


