Electronic Document Data Extraction via Hash Map Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional document format converters fail to extract and index relevant data from electronic documents, limiting their usability for analysis and reporting purposes.

Innovation Solution

A computer-implemented system and method that accesses electronic files, extracts data using a mapping structure, organizes it into a hash map, and stores it in a database, allowing for customizable document viewing and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional document format converters are used to convert between document formats, then the document can be accessed and edited in its current form, but the data cannot be extracted, indexed, or manipulated for analysis and reporting purposes

Engineering Contradiction:
Improvedocument accessibilityVSAvoiddata usability for analysis
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent extracts data from electronic documents by identifying structural elements (tables, text blocks, images) and pulling out relevant information using mapping structures. This extraction process separates the data from its original document format, enabling it to be stored in a database and used for various analytical purposes while the original document remains accessible.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary system that acts as a bridge between conventional document formats and analytical programs. This intermediary extracts and transforms data into a usable format, allowing both document accessibility and data usability for analysis to coexist.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If conventional converters simply recreate text from converted documents, then the document format can be changed, but no categorization or indexing is performed to make the data more useful

Engineering Contradiction:
Improveconversion simplicityVSAvoiddata organization and indexing
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent performs preliminary actions by pre-defining mapping structures that identify how data should be extracted and organized. Before the actual conversion process, the system establishes the framework for categorization and indexing, ensuring that data is properly organized from the outset rather than as an afterthought.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the document into distinct structural elements (tables, text blocks, images) and processes each segment separately. This segmentation allows for targeted extraction and indexing of specific data types, preventing loss of organizational information during conversion.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If data is extracted and organized into a database with mapping structures and hash maps, then the data can be indexed and manipulated for analysis, but the process becomes more complex than simple format conversion

Engineering Contradiction:
Improvedata manipulabilityVSAvoidextraction system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal extraction system that can handle multiple document formats and extract various types of data (tables, text, images) through a single unified approach. This multi-functionality reduces the need for separate complex processes for each document type, managing the overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes parameters by transforming data from its original document-specific format into a standardized database format using mapping structures. This parameter transformation allows the system to handle diverse document types through a consistent process, balancing complexity with versatility.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If the system reads parameters, forms rectangles, and iteratively analyzes parent and child rectangles to detect data tables, then accurate data extraction can be achieved, but the processing time and computational resources increase

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-reading parameters and establishing the coordinate system and rectangle structures before detailed analysis. This preliminary setup organizes the data in a way that facilitates faster subsequent processing while maintaining extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the document analysis into distinct phases: reading parameters, forming rectangles, iterative parent-child rectangle analysis, and final data extraction. This segmentation allows for optimized processing at each stage, balancing accuracy with processing efficiency.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9092417B2Systems and methods for extracting data from a document in an electronic format
Publication Date: 2015.07.28 WEB ACCESS
  • US9092417B2 patent drawing
  • US9092417B2 patent drawing
  • US9092417B2 patent drawing

AI summary

A computer-implemented method of extracting data from a document in an electronic format. The method includes the steps of accessing a file in an electronic format from a memory module; extracting data from the file corresponding to a plurality of keys contained within a mapping structure stored in the memory module; organizing the extracted data into values, wherein each value maps to one of the plurality of keys to form a hash map; storing the hash map in a database; and providing a user access to the database via an output device. The output device allows the user to view a customizable document whose content is derived from the values and keys stored in the database.