Auto-populating Database Fields from Heterogeneous Document Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computers are unable to automatically interpret and transfer text data from optical character recognition (OCR) into specific data fields for entry into a unified database, requiring manual human intervention, which is time-consuming and inefficient, especially when dealing with heterogeneous documents from various sources.
Innovation Solution
A computer-implemented method that receives an image file, performs OCR, identifies and compares text parameters to stored parameters, sorts the text into categories, and auto-populates data entry fields, allowing for automatic entry into a unified database, with the option for user verification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data entry is used for heterogeneous documents, then data accuracy can be maintained through human verification, but time consumption and labor effort increase significantly
Solution Approach 1:
The patent segments the data entry process into distinct functional modules: OCR text extraction module, parameter identification module, text sorting module, and field auto-population module. Each module handles a specific aspect of document processing, enabling automated handling of heterogeneous documents while maintaining data accuracy through structured verification workflows.
Solution Approach 2:
The patent introduces an intermediary processing layer between OCR extraction and final data entry. This intermediary layer includes parameter identification and text sorting mechanisms that bridge the gap between raw extracted text and structured database fields, enabling automated accurate data entry without direct manual intervention.
2Productivity
If automated OCR processing is implemented without interpretation capability, then processing speed increases, but the system cannot automatically transfer information to specific data fields
Solution Approach 1:
The patent performs preliminary actions by pre-defining parameter comparison criteria and field mapping rules before actual document processing. The system pre-processes the OCR output by identifying text parameters and sorting them according to predetermined categories, enabling seamless automatic transfer to specific data fields without requiring real-time interpretation decisions.
Solution Approach 2:
The patent replaces the mechanical interpretation process (human reading and understanding) with an automated parameter identification and comparison system. The system uses programmed logic to identify text parameters, compare them against stored parameters, and automatically determine field mappings, substituting human cognitive functions with automated computational processes.
3Reliability
If heterogeneous documents from various sources are processed manually, then data quality can be controlled, but the complexity and time required for data entry increase
Solution Approach 1:
The patent creates a universal processing framework that handles multiple types of heterogeneous documents through a single integrated system. The parameter identification and text sorting mechanisms are designed to work across different document formats and sources, providing consistent data quality control without requiring separate manual processes for each document type.
Data Source
AI summary
Enabling a computer to automatically enter information into a unified database from heterogenous documents. An image file is received. The image file is displayed in a first area of a window rendered on a tangible display device. The fields for data entry are displayed in a second area of the window. Optical character recognition is performed on the image file. At least one parameter of text is identified in the image file. The at least one parameter of the text is compared to at least one of a plurality of stored parameters. The text is sorted according to the at least one of the plurality of stored parameters into a plurality of categories, wherein sorted text is formed. The fields are auto-populated and displayed in the second area of the window based on the sorted text.


