Anchor-Based Data Extraction from Unstructured Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data extraction methods from semi-structured or unstructured documents face challenges in accurately identifying and extracting information due to variations in document format, terminology, and structural inconsistencies, leading to inefficiencies and inaccuracies.

Innovation Solution

The implementation of a data extraction system that uses flexible anchor elements, identified through machine learning models, to recognize relationships between anchor elements and information elements based on multiple metrics such as position, style, structure, and semantics, allowing for consistent data extraction across varying document formats and changes over time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional data extraction methods are used on semi-structured or unstructured documents, then the extraction process can be performed, but the accuracy and reliability of extracted information deteriorates due to variations in document format, terminology, and structural inconsistencies

Engineering Contradiction:
Improveextraction accuracyVSAvoiddocument format variability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic anchor element identification where the system adapts to different document formats by learning from multiple examples. The anchor elements are not fixed but are dynamically identified based on their relationship to information elements across various document structures, allowing the extraction system to maintain accuracy despite format variations.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes parameters by using multiple metrics (position, style, structure, semantics) to identify anchor elements instead of relying on a single fixed parameter. This multi-parameter approach allows the system to adapt to different document formats while maintaining extraction accuracy.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If flexible anchor elements with multiple metrics are used to improve extraction accuracy across varying formats, then adaptability improves, but the complexity of the extraction system increases

Engineering Contradiction:
Improveformat flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the complex extraction task into identifying specific anchor elements that have relationships to information elements. By focusing on these key anchor points rather than processing the entire document structure, the system achieves format flexibility while managing complexity through targeted analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The anchor elements serve as intermediaries between the document structure and the information extraction process. These anchor elements mediate the relationship between varying document formats and the extraction system, simplifying the overall complexity by providing stable reference points.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual identification of anchor elements is performed to ensure accuracy, then extraction precision improves, but the productivity and efficiency of the extraction process decreases

Engineering Contradiction:
Improveanchor identification accuracyVSAvoidextraction efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically identifying anchor elements through machine learning algorithms that analyze multiple metrics. The system learns from training data to autonomously identify anchor elements with their relationships to information elements, eliminating the need for manual identification while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses feedback from multiple metrics (position, style, structure, semantics) to continuously improve anchor element identification. By incorporating feedback from various document features and relationship types, the system achieves high accuracy automatically without manual intervention.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240311581A1Data identification and extraction from unstructured documents
Publication Date: 2024.09.19 ADOBE INC
  • US20240311581A1 patent drawing
  • US20240311581A1 patent drawing
  • US20240311581A1 patent drawing

AI summary

Aspects of the method, apparatus, non-transitory computer readable medium, and system include obtaining a document and an information element. The aspects further include identifying, from the document, an anchor element that has an anchor type and a relationship type, wherein the anchor type describes a structure of a set of anchor elements, and the relationship type describes a relationship between the anchor element and the information element. The aspects further include extracting information corresponding to the information element based on the anchor element, the anchor type, and the relationship type, and displaying the extracted information to a user.