Physics Model for Semi-Structured Document Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Variations in the layouts and designs of semi-structured electronic documents hinder efficient data extraction and transfer, requiring manual template creation or data entry, which is time-consuming and error-prone.

Innovation Solution

A system that generates a physics model of semi-structured documents using relationships among data elements, represented as mechanical objects, to automatically extract data by identifying locations and extracting data without manual input or custom template creation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual template creation or manual data entry is used to extract data from semi-structured documents, then data extraction accuracy can be maintained, but time consumption and operational complexity increase significantly

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical operations (template creation and data entry) with an automated physics-based model that uses mechanical object relationships to automatically locate and extract data elements from semi-structured documents, thereby eliminating time-consuming manual processes while maintaining extraction accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service data extraction by allowing the physics model to automatically identify data element locations and extract information without requiring manual intervention or template creation, making the extraction process autonomous and efficient

Inventive Principle:
Principle #25Self-service

2Extent of automation

If custom templates are created to match exact document layouts, then data extraction can be automated, but system complexity and development effort increase

Engineering Contradiction:
Improvedata extraction automationVSAvoidtemplate creation complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The physics-based model serves as a universal extraction mechanism that can handle multiple document layouts and formats without requiring separate custom templates for each, thereby reducing system complexity while maintaining automation capability across diverse document types

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the approach from fixed template parameters to dynamic physics-based relationships that adapt to different document layouts, allowing automation without requiring complex custom template creation for each document variation

Inventive Principle:
Principle #35Parameter changes

3Reliability

If manual data entry is required for semi-structured documents, then data accuracy can be ensured, but productivity decreases

Engineering Contradiction:
Improvedata accuracyVSAvoiddata extraction efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent substitutes manual data entry operations with automated physics-based extraction that maintains data accuracy through structured relationship modeling while dramatically improving productivity by eliminating repetitive manual input tasks

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10614125B1Modeling and extracting elements in semi-structured documents
Publication Date: 2020.04.07 INTUIT INC
  • US10614125B1 patent drawing
  • US10614125B1 patent drawing
  • US10614125B1 patent drawing

AI summary

The disclosed embodiments provide a system that describes a semi-structured document for the purpose of acquiring a set of data elements from the semi-structured document. During operation, the system obtains a physics model of a semi-structured document, wherein the physics model includes a set of relationships represented by physical objects that describe relative positions of a set of data elements in the semi-structured document. Next, the system applies the physics model to a representation of the semi-structured document to automatically extract a set of data from the representation. The system then provides the extracted set of data for use with one or more applications without requiring manual input of the data into the one or more applications.