Unstructured Document Navigation via Positional Linking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Unstructured domain-specific documents, such as summary plan documents, are difficult to interpret and extract information from due to their complex and lengthy nature, often containing significant domain-specific information that is challenging to navigate and convert into structured formats without losing positional information.
Innovation Solution
A method and system that convert unstructured documents into structured documents using domain-specific natural language processing engines and ontologies, maintaining contextual information and linking extracted information to its original position, allowing for the generation of navigable structures and graphical user interfaces for efficient information extraction and verification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If unstructured documents are converted to structured formats using general NLP techniques, then information extraction becomes easier, but domain-specific contextual information and positional accuracy are lost
Solution Approach 1:
The patent changes the parameters of the NLP engine by making it domain-specific rather than general-purpose. The system tailors the NLP engine to specific domains (e.g., healthcare, finance) by training it on domain-specific corpora and configuring it with domain-specific ontologies, thereby improving both information extraction capability and contextual information preservation simultaneously
Solution Approach 2:
The patent introduces an intermediary layer between the unstructured document and the structured output. This intermediary consists of domain-specific ontologies and knowledge graphs that mediate the transformation process, ensuring that domain-specific contextual information is preserved and mapped appropriately to the structured format while maintaining positional accuracy
2Loss of information
If complex unstructured documents are processed to extract detailed information, then information completeness improves, but processing time and system complexity increase
Solution Approach 1:
The patent segments the processing task by dividing the document into logical sections and processing them independently using the structured document format. This segmentation allows parallel processing and reduces the overall processing time while maintaining complete information extraction from each segment
Solution Approach 2:
The patent performs preliminary actions by pre-processing the document to identify and mark key information elements before the main extraction process. This preliminary structuring reduces the complexity of the subsequent extraction phase and accelerates the overall processing time
3Productivity
If general NLP techniques are used on unstructured documents, then processing speed is maintained, but information extraction accuracy deteriorates
Solution Approach 1:
The patent changes the parameters of the NLP engine from general-purpose to domain-specific configurations. By adjusting parameters such as vocabulary, entity recognition rules, and relationship extraction patterns to match the specific domain, the system achieves high extraction accuracy without sacrificing processing speed
4Ease of operation
If unstructured documents are transformed into structured formats without position tracking, then navigation becomes easier, but verification of extracted information becomes difficult
Solution Approach 1:
The patent adds another dimension to the structured document format by incorporating positional metadata that tracks the location of extracted information in the original unstructured document. This additional dimensional information enables both easy navigation through the structured format and reliable verification by mapping back to the source document
Data Source
AI summary
Aspects of the present disclosure describe techniques for generating a machine learning model for extracting information from textual content. The method generally includes receiving an unstructured document and a structured document including information extracted from the unstructured document and position information associated with the extracted information. The unstructured document is rendered in a first pane, and a graphical rendering of the structured document is rendered in a second pane. The graphical rendering generally may be a structure in which content from the structured document is displayed in a hierarchical format. Each element in the structured document is linked to the rendered unstructured document based on position information included in the structured document.


