Data Extraction from Electronic Documents Using Cosine Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing software tools fail to extract data from electronic documents with high accuracy, especially when forms have varying data patterns such as multiple choice options, underlined values, and default values, due to differences in document structure caused by different vendors or record keepers.
Innovation Solution
A system and method using a server computing device to process electronic documents with a text extraction and data processing application, determining data extraction formulas, and identifying data elements based on cosine similarity scores to extract data values accurately from documents with diverse patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing software tools are used to extract data from electronic documents, then data extraction can be performed automatically, but extraction accuracy is low especially when forms have varying data patterns
Solution Approach 1:
The patent implements dynamic data extraction by training machine learning models on diverse document patterns and using cosine similarity to adaptively match query documents against training data. The system dynamically adjusts its extraction approach based on the specific pattern recognized, rather than using a fixed extraction rule set.
Solution Approach 2:
The system changes parameters by using cosine similarity scores to determine the best match between document patterns. It transforms the extraction problem from rule-based to similarity-based, allowing flexible adaptation to different data patterns through parameter adjustment rather than structural changes.
2Ease of operation
If software tools use fixed extraction rules, then extraction process is simple, but they fail to extract proper information from documents with different structures
Solution Approach 1:
The patent applies preliminary action by pre-training machine learning models on diverse document patterns before actual extraction. The system performs advance preparation by creating a library of trained models that can be quickly applied to new documents, maintaining simplicity while improving reliability.
Solution Approach 2:
The system introduces cosine similarity as an intermediary mechanism between the document structure and extraction rules. This intermediary allows the system to bridge fixed extraction processes with variable document structures, maintaining operational simplicity while achieving reliable extraction across different patterns.
3Adaptability or versatility
If a tool is designed to handle multiple document patterns, then versatility increases, but system complexity increases
Solution Approach 1:
The patent implements universality by creating a single machine learning-based extraction system that can handle multiple document patterns through cosine similarity matching. Instead of separate tools for each pattern type, one universal system adapts to various patterns by comparing against trained templates.
Solution Approach 2:
The system uses copying by creating trained model copies from sample documents. These trained models serve as reusable templates that can be applied to similar documents, reducing the need for complex real-time analysis while maintaining high extraction accuracy across different patterns.
4Measurement precision
If manual data extraction is performed, then extraction accuracy can be high, but extraction time increases significantly
Solution Approach 1:
The patent implements self-service by enabling the system to automatically train and improve its own extraction capabilities. The machine learning models self-adjust based on training data and cosine similarity measurements, eliminating the need for manual rule configuration while maintaining high extraction accuracy at automated speeds.
Solution Approach 2:
The system substitutes mechanical manual extraction with automated machine learning-based extraction. By replacing human manual processes with algorithmic cosine similarity matching and trained model application, the system achieves both high accuracy and automated processing speed.
Data Source
AI summary
Systems and methods for extracting data from electronic documents based on data patterns. The method includes receiving electronic template documents. Each template document corresponds to a type of electronic document. The method further includes, for each template document, processing the template document using a text extraction and data processing application. The method also includes, for each template document, determining a data extraction formula corresponding to the type of electronic document. The method further includes, storing the data extraction formula in a first database. The method also includes, receiving an electronic document including user data and a Unicode corresponding to the type of document. The method also includes, processing and classifying the electronic document into the type of document corresponding to the Unicode. The method also includes identifying data elements in the electronic document based on the data extraction formula and extracting data values for each of the identified data elements.


