Unstructured Document Tagging via NLP Highlighting and User Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches for analyzing unstructured data using object-based data modeling platforms are hindered by noisy automated metadata tagging and laborious manual tagging, which requires significant manual review and often results in errors due to complex object ontologies and user interface struggles.

Innovation Solution

A computing system that facilitates the tagging of unstructured documents by using natural language processing to highlight matching terms and prompting users to select appropriate terms for structured data object creation, transforming unstructured documents into structured data objects suitable for analysis via an object-based data modeling framework.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated metadata tagging is used, then tagging speed is improved, but tagging accuracy deteriorates due to noisy results requiring significant manual review

Engineering Contradiction:
Improvetagging speedVSAvoidtagging accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary semi-structured template format that bridges automated tagging and final structured data objects. The template acts as a mediator that organizes automated tagging results into a controlled format, making it easier to identify and correct errors while preserving automation benefits.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The tagging process is segmented into multiple stages: automated metadata extraction, template-based organization, and structured object creation. This segmentation allows each stage to be optimized independently, with automated tools handling volume and human reviewers focusing on quality control at critical transition points.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If manual tagging is used, then tagging accuracy is improved, but productivity deteriorates due to laborious and error-filled processes

Engineering Contradiction:
Improvetagging accuracyVSAvoidtagging speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

Automated preprocessing and template preparation are performed before manual tagging begins. The system pre-structures the document, identifies potential data elements, and prepares templates in advance, so human taggers work with pre-organized content rather than raw unstructured data, significantly reducing their cognitive load and time requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs self-service automated tagging for routine, high-confidence extractions, allowing human reviewers to focus only on ambiguous or complex cases. This self-service capability handles the volume of straightforward tagging automatically while human intelligence is allocated to challenging instances.

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If complex object ontologies are enforced, then data structure quality is improved, but ease of operation deteriorates as users struggle with the interface mechanisms

Engineering Contradiction:
Improvedata structure qualityVSAvoiduser interface usability
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The semi-structured template serves as an intermediary layer between the user and the complex object ontology. Users interact with simplified template fields rather than directly manipulating complex ontology structures, while the system automatically handles the mapping to the underlying data model, shielding users from ontology complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Instead of requiring users to understand and navigate complex object ontologies to create structured data, the system inverts the approach by presenting simplified templates that automatically map to the ontology. The complexity is hidden in the reverse direction, from structured data back to unstructured source.

Inventive Principle:
Principle #13The other way round (Inversion)

4Measurement precision

If full manual review of automated tagging is performed, then tagging accuracy is improved, but loss of time increases due to significant manual review requirements

Engineering Contradiction:
Improvetagging accuracyVSAvoidmanual review time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of performing full manual review of all tagged content, the system applies partial review focused on high-value or high-risk fields identified through the template structure. The template guides reviewers to prioritize specific sections that require human verification, performing excessive action only where necessary rather than uniformly across all content.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11244102B2Systems and methods for facilitating data object extraction from unstructured documents
Publication Date: 2022.02.08 PALANTIR TECHNOLOGIES INC
  • US11244102B2 patent drawing
  • US11244102B2 patent drawing
  • US11244102B2 patent drawing

AI summary

Systems and methods are provided for facilitating data object extraction from unstructured documents. Unstructured documents may include data in an unorganized format, such as raw text. The system may use natural language processing to determine characteristics of the terms used in the unstructured document. The system may prompt a user to select terms from the document corresponding in characteristics to properties of a data object being generated. The user may select terms from the document and the system may generate a data object according to the selected terms.