Implicit Coordinates and Local Neighborhood for Document Data Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document processing systems rely on human-maintained templates for OCR, which are inflexible and require significant resources to maintain, especially for documents with variability in structure and format.
Innovation Solution
A system and method that use machine vision and natural language processing to compare visual and textual elements between new documents and historical documents, creating match metrics to extract relevant data from regions of interest without requiring human interaction or rigid templates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If template-based OCR is used with human-maintained templates, then data extraction accuracy is improved, but device complexity and maintenance resources increase significantly
Solution Approach 1:
The system automatically learns document structures and templates through machine learning from historical documents, eliminating the need for manual template creation and maintenance by humans. The AI model self-updates and adapts to new document formats autonomously.
Solution Approach 2:
The patent replaces manual template-based OCR mechanisms with an AI-powered machine learning system that automatically processes documents, substituting human operators and rigid template systems with adaptive algorithms.
2Productivity
If rigid rule-based methods are used, then processing speed is improved, but adaptability to document variability decreases
Solution Approach 1:
The system transitions from static rigid rules to dynamic machine learning models that can adapt to varying document formats. The AI model learns from historical documents and automatically adjusts to new structures, maintaining both speed and flexibility.
Solution Approach 2:
The patent changes the fundamental parameters of document processing from fixed rule-based parameters to adaptive learned parameters through machine learning, enabling the system to handle document variability while maintaining efficiency.
3Measurement precision
If templates are created for every document type, then extraction accuracy is improved, but loss of time and resources increases
Solution Approach 1:
The system performs preliminary learning actions by training on historical documents to pre-establish patterns and structures. This preliminary machine learning process eliminates the need for time-consuming manual template creation for each new document type.
Solution Approach 2:
The AI system learns from copies of historical documents, analyzing patterns and structures from existing document examples to automatically create extraction capabilities without requiring manual template development for each document type.
4Measurement precision
If X-Y coordinates are used for target location, then positioning precision is improved, but reliability decreases when document layout changes
Solution Approach 1:
The patent changes the coordinate system from absolute fixed X-Y positions to relative positioning based on document structure and content relationships. The AI model understands contextual relationships between elements, making positioning reliable even when layouts vary.
Data Source
AI summary
A system and method are disclosed for using a local neighborhood for determining similar targets in different documents or using implicit coordinates for obtaining a coordinate location of a target. The local neighborhood method may include identifying a first target in a first document; identifying one or more first elements within a first distance range from the first target; creating a first local neighborhood based on the identifying; determining that that first local neighborhood is similar to a third local neighborhood in a second document; and determining a second target in the second document that corresponds to the first target in the first document, based on the determining the similarity. The implicit coordinates method may include performing OCR on the first document to find the first target; and obtaining a first location of the first target by using at least one of OCR or element recognition.


