Metadata Augmentation for NLP Accuracy and Resource Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural language processing systems face inaccuracies and inefficiencies when dealing with documents lacking metadata values or having predicted accuracy below a threshold, leading to suboptimal processing outcomes.

Innovation Solution

The method involves accessing and augmenting metadata values for documents by selecting appropriate value extraction processes, using artificial intelligence models, and normalizing phrases across different entities and document types, thereby improving accuracy and reducing computational resource usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If metadata values are augmented using value extraction processes, then the accuracy of natural language processing is improved, but the computational resource usage increases

Engineering Contradiction:
Improveaccuracy of natural language processingVSAvoidcomputational resource usage
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary actions by extracting and normalizing metadata values before the natural language processing task. Value extraction processes are applied in advance to documents that lack complete or accurate metadata, preparing the data structures needed for subsequent NLP operations. This preliminary processing ensures that when NLP is executed, the metadata is already optimized for accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The metadata extraction system operates autonomously by automatically identifying documents with incomplete metadata and applying appropriate value extraction processes without human intervention. The system self-regulates by determining which documents need metadata augmentation based on predefined criteria, and automatically selects and applies the most suitable extraction processes to minimize computational overhead while maximizing accuracy improvement.

Inventive Principle:
Principle #25Self-service

2Device complexity

If a common ontology is used across entities and document types, then downstream processing is simplified, but the difficulty of detecting and measuring accurate metadata values increases

Engineering Contradiction:
Improvedownstream processing complexityVSAvoiddifficulty of detecting accurate metadata values
Core Design Contradiction:
Device complexityVSDifficulty of detecting and measuring

Solution Approach 1:

The system implements a universal ontology framework that serves multiple functions across different entities and document types. This common ontology structure provides standardized metadata fields and value formats that can be applied consistently throughout the system, simplifying downstream processing by eliminating the need for entity-specific or document-type-specific processing logic.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system manages the complexity of detecting accurate metadata values by dynamically adjusting extraction parameters and thresholds based on the specific document type and entity being processed. The value extraction processes modify their detection parameters adaptively, allowing the common ontology to maintain simplicity while still achieving high accuracy in metadata value detection across diverse content types.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240220736A1Metadata Processing
Publication Date: 2024.07.04 ASTRATA INC
  • US20240220736A1 patent drawing
  • US20240220736A1 patent drawing
  • US20240220736A1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for processing metadata. One of the methods includes accessing a document a) that has, in metadata that can be used to perform natural language processing, a first set of values for a first subset of metadata types from a plurality of metadata types and b) for which a second set of values for a second subset of metadata types from the plurality can be augmented; selecting, for a metadata type from the second subset of metadata types, one or more value extraction processes from a plurality of value extraction processes each of which is for a corresponding metadata type from the plurality; determining, using the one or more value extraction processes and the document, a value for the metadata type; and storing the value for the metadata type as part of the metadata for the document.