Document Data Segment Identification Using OCR and NLP
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer vision and process automation technologies are unable to effectively review and extract data from complex, data-rich electronic documents, such as those containing a mixture of text, tables, images, and other content, leading to the need for manual review which is time-consuming and prone to human error.
Innovation Solution
A computer-implemented method using a trained natural language processing model and optical character recognition processor to extract text data, determine candidate entity data, access n-gram words from a knowledge base, and calculate similarity scores to identify and select the most accurate entity data, enabling the review of complex data sets 25 times faster than manual methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing computer vision and process automation technology is used, then simple electronic documents can be reviewed and extracted, but complex data-rich electronic documents cannot be effectively processed
Solution Approach 1:
The patent segments complex data-rich documents into multiple data segments including text segments, table segments, and image segments. Each segment type is processed by specialized extraction modules, enabling the system to handle diverse content types that existing automation technology cannot process effectively.
Solution Approach 2:
The patent creates a universal document processing system that can handle multiple document types and content formats (text, tables, images) through a single integrated platform. The system uses multiple extraction modules that work together to process various data segments, achieving multi-functionality that replaces manual review across different document complexities.
2Measurement precision
If manual review of complex documents is performed, then accurate extraction is achieved, but processing is time-consuming and expensive
Solution Approach 1:
The system performs self-service by automatically identifying, segmenting, and extracting data from complex documents without requiring manual intervention. The multiple extraction modules work autonomously to process different data segments, achieving both high accuracy through specialized processing and high productivity through automation.
Solution Approach 2:
The patent replaces the mechanical manual review process with an automated computer-based system that uses multiple extraction modules to process document segments. This substitution maintains or improves accuracy while dramatically increasing processing speed and reducing costs associated with manual review.
3Productivity
If existing automation technology is used, then processing speed is maintained, but accuracy and reliability of extraction deteriorate for complex documents
Solution Approach 1:
By segmenting complex documents into distinct data segments (text, tables, images) and applying specialized extraction methods to each segment type, the system maintains high accuracy for complex documents while preserving fast automated processing speeds. Each segment is processed by the most appropriate extraction module.
Solution Approach 2:
The patent applies local quality by using different extraction approaches for different data segments based on their specific characteristics. Text segments use natural language processing, table segments use structured data extraction, and image segments use optical character recognition, ensuring each segment is processed with the appropriate method for maximum accuracy.
Data Source
AI summary
A data processing system receives a plurality of electronic documents in image format. For each signature segment, the system determines associated surrounding text using an optical character recognition processor. The system accesses first data stored in a signature knowledge base. The system determines first similarity scores based on the associated surrounding text and the first data using a statistical measure technique. The system selects an optimum signature segment based on a distance metric between each signature segment and each associated surrounding text. For each stamp segment, the system extracts text from the stamp segment using an optical character recognition processor. The system accesses second data stored in a stamp knowledge base. The system determines second similarity scores based on the extracted text and the second data using the statistical measure technique. The system selects an optimum stamp segment based on the second similarity scores.


