Electronic Document Data Extraction via Hash Map Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional document format converters fail to extract and index relevant data from electronic documents, limiting their usability for analysis and reporting purposes.
Innovation Solution
A computer-implemented system and method that accesses electronic files, extracts data using a mapping structure, organizes it into a hash map, and stores it in a database, allowing for customizable document viewing and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional document format converters are used to convert between document formats, then the document can be accessed and edited in its current form, but the data cannot be extracted, indexed, or manipulated for analysis and reporting purposes
Solution Approach 1:
The patent extracts data from electronic documents by identifying structural elements (tables, text blocks, images) and pulling out relevant information using mapping structures. This extraction process separates the data from its original document format, enabling it to be stored in a database and used for various analytical purposes while the original document remains accessible.
Solution Approach 2:
The patent introduces an intermediary system that acts as a bridge between conventional document formats and analytical programs. This intermediary extracts and transforms data into a usable format, allowing both document accessibility and data usability for analysis to coexist.
2Ease of manufacture
If conventional converters simply recreate text from converted documents, then the document format can be changed, but no categorization or indexing is performed to make the data more useful
Solution Approach 1:
The patent performs preliminary actions by pre-defining mapping structures that identify how data should be extracted and organized. Before the actual conversion process, the system establishes the framework for categorization and indexing, ensuring that data is properly organized from the outset rather than as an afterthought.
Solution Approach 2:
The patent segments the document into distinct structural elements (tables, text blocks, images) and processes each segment separately. This segmentation allows for targeted extraction and indexing of specific data types, preventing loss of organizational information during conversion.
3Adaptability or versatility
If data is extracted and organized into a database with mapping structures and hash maps, then the data can be indexed and manipulated for analysis, but the process becomes more complex than simple format conversion
Solution Approach 1:
The patent creates a universal extraction system that can handle multiple document formats and extract various types of data (tables, text, images) through a single unified approach. This multi-functionality reduces the need for separate complex processes for each document type, managing the overall system complexity.
Solution Approach 2:
The patent changes parameters by transforming data from its original document-specific format into a standardized database format using mapping structures. This parameter transformation allows the system to handle diverse document types through a consistent process, balancing complexity with versatility.
4Measurement precision
If the system reads parameters, forms rectangles, and iteratively analyzes parent and child rectangles to detect data tables, then accurate data extraction can be achieved, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary actions by pre-reading parameters and establishing the coordinate system and rectangle structures before detailed analysis. This preliminary setup organizes the data in a way that facilitates faster subsequent processing while maintaining extraction accuracy.
Solution Approach 2:
The patent segments the document analysis into distinct phases: reading parameters, forming rectangles, iterative parent-child rectangle analysis, and final data extraction. This segmentation allows for optimized processing at each stage, balancing accuracy with processing efficiency.
Data Source
AI summary
A computer-implemented method of extracting data from a document in an electronic format. The method includes the steps of accessing a file in an electronic format from a memory module; extracting data from the file corresponding to a plurality of keys contained within a mapping structure stored in the memory module; organizing the extracted data into values, wherein each value maps to one of the plurality of keys to form a hash map; storing the hash map in a database; and providing a user access to the database via an output device. The output device allows the user to view a customizable document whose content is derived from the values and keys stored in the database.


