Unstructured Document Navigation via Positional Linking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Unstructured domain-specific documents, such as summary plan documents, are difficult to interpret and extract information from due to their complex and lengthy nature, often containing significant domain-specific information that is challenging to navigate and convert into structured formats without losing positional information.

Innovation Solution

A method and system that convert unstructured documents into structured documents using domain-specific natural language processing engines and ontologies, maintaining contextual information and linking extracted information to its original position, allowing for the generation of navigable structures and graphical user interfaces for efficient information extraction and verification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If unstructured documents are converted to structured formats using general NLP techniques, then information extraction becomes easier, but domain-specific contextual information and positional accuracy are lost

Engineering Contradiction:
Improveinformation extractionVSAvoiddomain-specific contextual information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent changes the parameters of the NLP engine by making it domain-specific rather than general-purpose. The system tailors the NLP engine to specific domains (e.g., healthcare, finance) by training it on domain-specific corpora and configuring it with domain-specific ontologies, thereby improving both information extraction capability and contextual information preservation simultaneously

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary layer between the unstructured document and the structured output. This intermediary consists of domain-specific ontologies and knowledge graphs that mediate the transformation process, ensuring that domain-specific contextual information is preserved and mapped appropriately to the structured format while maintaining positional accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If complex unstructured documents are processed to extract detailed information, then information completeness improves, but processing time and system complexity increase

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent segments the processing task by dividing the document into logical sections and processing them independently using the structured document format. This segmentation allows parallel processing and reduces the overall processing time while maintaining complete information extraction from each segment

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-processing the document to identify and mark key information elements before the main extraction process. This preliminary structuring reduces the complexity of the subsequent extraction phase and accelerates the overall processing time

Inventive Principle:
Principle #10Preliminary action

3Productivity

If general NLP techniques are used on unstructured documents, then processing speed is maintained, but information extraction accuracy deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidinformation extraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters of the NLP engine from general-purpose to domain-specific configurations. By adjusting parameters such as vocabulary, entity recognition rules, and relationship extraction patterns to match the specific domain, the system achieves high extraction accuracy without sacrificing processing speed

Inventive Principle:
Principle #35Parameter changes

4Ease of operation

If unstructured documents are transformed into structured formats without position tracking, then navigation becomes easier, but verification of extracted information becomes difficult

Engineering Contradiction:
ImprovenavigationVSAvoidverification accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent adds another dimension to the structured document format by incorporating positional metadata that tracks the location of extracted information in the original unstructured document. This additional dimensional information enables both easy navigation through the structured format and reliable verification by mapping back to the source document

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11392753B2Navigating unstructured documents using structured documents including information extracted from unstructured documents
Publication Date: 2022.07.19 ANTHROPIC PBC
  • US11392753B2 patent drawing
  • US11392753B2 patent drawing
  • US11392753B2 patent drawing

AI summary

Aspects of the present disclosure describe techniques for generating a machine learning model for extracting information from textual content. The method generally includes receiving an unstructured document and a structured document including information extracted from the unstructured document and position information associated with the extracted information. The unstructured document is rendered in a first pane, and a graphical rendering of the structured document is rendered in a second pane. The graphical rendering generally may be a structure in which content from the structured document is displayed in a hierarchical format. Each element in the structured document is linked to the rendered unstructured document based on position information included in the structured document.