Intermediate Text Representations for Table Structure Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated computer analysis of electronic documents, particularly tables in contracts, is error-prone and inefficient, leading to inaccurate extraction of relevant content, which is crucial for software asset management and licensing information.

Innovation Solution

A system that converts table structures into intermediate text representations and uses machine learning models to extract information of interest, improving the accuracy and efficiency of content extraction from electronic documents by analyzing text and structures, including tables, and storing the extracted information in a structured format.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated computer analysis is used to extract information from electronic documents, then productivity is improved, but measurement precision deteriorates due to error-prone analysis

Engineering Contradiction:
Improvecontent extraction efficiencyVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces intermediate representations as a mediator between the original table structure and the final extracted information. These intermediate representations serve as a bridge that captures both structural context and content meaning, allowing the system to maintain high extraction accuracy while achieving automated processing. The intermediate representations include encoded structural information (row, column, cell relationships) combined with semantic content, enabling the machine learning model to accurately interpret tabular data without manual analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual analysis is used to extract information from documents, then measurement precision is maintained, but productivity deteriorates due to time-consuming processes

Engineering Contradiction:
Improveextraction accuracyVSAvoidcontent extraction efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical manual analysis process with an automated system that uses machine learning models trained on intermediate representations. The system automatically processes electronic documents, converts tables to intermediate text representations, and extracts information without human intervention. This substitution maintains high extraction accuracy through sophisticated algorithms while dramatically improving productivity by eliminating time-consuming manual review processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If table structures are directly processed by machine learning models, then productivity is improved, but measurement precision deteriorates due to difficulty in interpreting tabular structures

Engineering Contradiction:
Improveprocessing speedVSAvoidstructure interpretation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the table structure into discrete, meaningful components within the intermediate representations. Each representation captures specific structural elements (row identifiers, column headers, cell boundaries, relationships) separately from the content, allowing the machine learning model to process both structure and meaning independently and accurately. This segmentation enables the model to understand tabular relationships without confusion, maintaining high precision while achieving automated processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11880435B2Determination of intermediate representations of discovered document structures
Publication Date: 2024.01.23 SERVICENOW INC
  • US11880435B2 patent drawing
  • US11880435B2 patent drawing
  • US11880435B2 patent drawing

AI summary

A document is received. The document is analyzed to discover text and structures of content included in the document. A result of the analysis is used to determine intermediate text representations of segments of the content included in the document, wherein at least one of the intermediate text representations includes an added text encoding the discovered structure of the corresponding content segment within a structural layout of the document. The intermediate text representations are used as an input to a machine learning model to extract information of interest in the document. One or more structured records of the extracted information of interest are created.