Automated PDF Table Consolidation and Trend Visualization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting and consolidating data from tables in PDF documents are cumbersome due to the lack of structural and semantic information, requiring manual effort and interrupting the reading flow, and making it difficult to identify trends across multiple documents.

Innovation Solution

A method and system that automatically identify and consolidate similar tables from multiple documents by determining their schemas using machine learning, extracting data, and visualizing the consolidated data to facilitate trend identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual extraction and consolidation of table data from PDF documents is performed, then data can be extracted, but the process is tedious and time-consuming

Engineering Contradiction:
Improvedata extraction speedVSAvoidtime for manual extraction
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically extracting table data from PDF documents without requiring manual user intervention. The extraction process is automated through computational algorithms that identify and parse table structures, populate data, and consolidate results across multiple documents, eliminating the need for tedious manual copying and pasting operations

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual process of table extraction with an automated computational system. Instead of manually selecting, copying, and pasting table data, the system uses algorithmic processing to automatically extract, structure, and consolidate table information from multiple PDF documents, substituting human manual labor with automated mechanical processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If tables in PDF documents are viewed without structural information, then the documents can be read, but understanding table data requires cumbersome interactions that interrupt reading flow

Engineering Contradiction:
Improveease of table understandingVSAvoidcomplexity of interactions
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary consolidation table that serves as a mediator between the original PDF tables and the user. This consolidation table aggregates data from multiple source tables and presents it in a unified format with clear structure and context, eliminating the need for users to navigate multiple separate tables and their associated complex interactions across different documents

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system merges multiple separate table structures from different PDF documents into a single consolidated table. By combining the data and standardizing the structure, the system creates a unified view that simplifies understanding and eliminates the need for users to switch between multiple tables and documents, thereby reducing interaction complexity

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If tables are extracted without semantic information, then extraction is simpler, but the extracted data lacks context for understanding trends across documents

Engineering Contradiction:
Improvesemantic information retentionVSAvoidtrend identification efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system performs preliminary action by pre-processing and consolidating table data from multiple documents before the user needs to analyze trends. The consolidation process preserves semantic information including table captions, row labels, and contextual metadata, organizing this information in advance so that when users view the consolidated table, the contextual information is already available and properly structured for immediate trend analysis

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240202435A1Automatic cross document consolidation and visualization of data tables
Publication Date: 2024.06.20 OHIO STATE INNOVATION FOUND
  • US20240202435A1 patent drawing
  • US20240202435A1 patent drawing
  • US20240202435A1 patent drawing

AI summary

In an embodiment, a set of related documents (105) is selected (305). Each document may include at least one table of data (107). The tables may not include semantic or structural data that can be used to understand the data in the tables. Each table is processed to determine a schema for the table that includes a name and type for each column of the table (320). A consolidated schema is received for a consolidated table (320). The consolidated schema includes a name and type for each column of the consolidated table. The data from each table is extracted from the table and added to the consolidated table based on the schema associated with the table and the schema associated with the consolidated table (325). Later, the data in the consolidated table can be visualized to help identify one or more trends (400).