Cross-Document Entity Metadata Linkage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural language processing (NLP) systems struggle to effectively propagate metadata across multiple documents related to the same subject, limiting the ability to provide a comprehensive view of concepts mentioned in different documents.

Innovation Solution

A computer-implemented method and system that identifies related documents, uses NLP to create annotations for concepts, and establishes metadata linkages between matching concepts across documents, allowing for the aggregation of mentions and providing a unified view of the concept when viewing any document.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If NLP is used to identify concepts and create metadata in individual documents, then concept identification accuracy is improved, but metadata propagation across related documents is lost

Engineering Contradiction:
Improveconcept identification accuracyVSAvoidmetadata propagation
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent merges metadata from multiple related documents by identifying concept matches across documents and creating aggregated metadata that combines information from all source documents. This resolves the contradiction by maintaining individual document concept identification accuracy while adding cross-document metadata propagation through the merging process.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary metadata aggregation system that sits between individual document processing and final information delivery. This intermediary layer evaluates annotations across documents, identifies concept matches, and creates unified metadata linkages, thereby preventing metadata loss while preserving individual concept identification accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If metadata is only associated with source documents containing identified concepts, then processing complexity is reduced, but comprehensive concept understanding across documents is limited

Engineering Contradiction:
Improvemetadata processing complexityVSAvoidcomprehensive concept understanding
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the metadata processing into distinct stages: individual document concept identification, cross-document annotation evaluation, concept match identification, and metadata linkage creation. This segmentation maintains manageable processing complexity at each stage while achieving comprehensive concept understanding through the cumulative effect of all stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a cross-document dimension to metadata processing by evaluating annotations against all documents in a set and creating linkages across documents. This dimensional expansion transforms single-document metadata into multi-document metadata networks, enhancing concept understanding versatility without proportionally increasing processing complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of information

If cross-document metadata linkage is implemented, then comprehensive concept information is provided, but system complexity increases

Engineering Contradiction:
Improveconcept information completenessVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by first identifying related documents and creating individual document annotations before attempting cross-document matching. This preliminary processing organizes data in advance, making the subsequent metadata linkage creation more efficient and manageable, thereby reducing overall system complexity while achieving complete concept information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses self-service mechanisms where annotations from one document automatically serve as evaluation criteria for matching concepts in other documents. This self-referential approach reduces the need for external complex processing logic, as the metadata system uses its own generated annotations to drive cross-document linkage creation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11461540B2Cross-document propagation of entity metadata
Publication Date: 2022.10.04 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11461540B2 patent drawing
  • US11461540B2 patent drawing
  • US11461540B2 patent drawing

AI summary

Embodiments include cross-document propagation of entity metadata. Aspects include identifying a set of documents from a plurality of documents, the set of documents being related to one another and identifying a concept in a first document of the set of documents and creating an annotation corresponding to the concept. Aspects also include evaluating the annotation from the first document against all of the documents in the set of documents and identifying a concept match between the annotation and a mention discovered in a second document in the set of documents. Aspects further include creating a metadata linkage between the concept in the first document to the mention in the second document.