Unstructured Data Extraction via Virtual Cursor Grouping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting and organizing data from unstructured documents, such as images, often fail when documents contain internal headers, as they rely on whitespace detection, which is not always reliable.
Innovation Solution
A system and method that group data elements based on matching criteria and horizontal overlaps, allowing for the assignment of data elements to rows and columns in a structured output document, even in the presence of internal headers, by using a computer configured with instructions to perform grouping and column assignment processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If whitespace detection is used to extract and organize data from unstructured documents, then the process is simple and fast, but it fails when documents contain internal headers
Solution Approach 1:
The patent introduces an intermediary layer of processing between raw document scanning and final data organization. A virtual cursor system acts as a mediator that tracks logical positions through the document, allowing the system to distinguish between whitespace that separates data and whitespace that is part of internal header structures. This intermediary tracking mechanism enables reliable data extraction even in documents with complex formatting.
Solution Approach 2:
The patent adds a dimensional layer by introducing virtual cursor positions that operate independently from physical whitespace detection. Instead of relying solely on horizontal/vertical whitespace distances, the system creates a logical dimension through cursor-based tracking that can navigate through internal headers and other complex structures. This dimensional transformation allows the system to maintain accurate data organization despite the presence of internal headers.
2Device complexity
If traditional whitespace-based grouping is used, then the method is simple to implement, but it cannot handle documents with internal headers
Solution Approach 1:
The patent creates a universal data extraction framework that can handle multiple document types and formatting scenarios through a single cohesive system. The virtual cursor mechanism serves multiple functions: it tracks data element positions, navigates through internal headers, determines grouping boundaries, and assigns logical positions. This multi-functional approach eliminates the need for separate specialized algorithms for different document scenarios.
Solution Approach 2:
The patent performs preliminary actions by pre-establishing virtual cursor positions and grouping structures before actual data extraction begins. The system pre-processes the document structure to identify potential grouping boundaries and establishes a framework for data organization in advance. This preliminary structuring enables the system to handle internal headers and complex formatting without requiring complex real-time decision-making during extraction.
Data Source
AI summary
Data elements from an input document can be automatically organized into rows and columns in a structured output document using a grouping process that automatically applies matching criteria based on horizontal position, data content and horizontal extent and tests for horizontal overlaps between data elements and neighbors of data elements in existing groups, assigns columns to those groups based on horizontal positions of data elements from groups that have already been assigned to columns. Rows may be assigned to data elements based on those data elements' vertical positions.


