Cross-Document Entity Metadata Linkage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Natural language processing (NLP) systems struggle to effectively propagate metadata across multiple documents related to the same subject, limiting the ability to provide a comprehensive view of concepts mentioned in different documents.
Innovation Solution
A computer-implemented method and system that identifies related documents, uses NLP to create annotations for concepts, and establishes metadata linkages between matching concepts across documents, allowing for the aggregation of mentions and providing a unified view of the concept when viewing any document.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If NLP is used to identify concepts and create metadata in individual documents, then concept identification accuracy is improved, but metadata propagation across related documents is lost
Solution Approach 1:
The patent merges metadata from multiple related documents by identifying concept matches across documents and creating aggregated metadata that combines information from all source documents. This resolves the contradiction by maintaining individual document concept identification accuracy while adding cross-document metadata propagation through the merging process.
Solution Approach 2:
The patent introduces an intermediary metadata aggregation system that sits between individual document processing and final information delivery. This intermediary layer evaluates annotations across documents, identifies concept matches, and creates unified metadata linkages, thereby preventing metadata loss while preserving individual concept identification accuracy.
2Device complexity
If metadata is only associated with source documents containing identified concepts, then processing complexity is reduced, but comprehensive concept understanding across documents is limited
Solution Approach 1:
The patent segments the metadata processing into distinct stages: individual document concept identification, cross-document annotation evaluation, concept match identification, and metadata linkage creation. This segmentation maintains manageable processing complexity at each stage while achieving comprehensive concept understanding through the cumulative effect of all stages.
Solution Approach 2:
The patent adds a cross-document dimension to metadata processing by evaluating annotations against all documents in a set and creating linkages across documents. This dimensional expansion transforms single-document metadata into multi-document metadata networks, enhancing concept understanding versatility without proportionally increasing processing complexity.
3Loss of information
If cross-document metadata linkage is implemented, then comprehensive concept information is provided, but system complexity increases
Solution Approach 1:
The patent performs preliminary actions by first identifying related documents and creating individual document annotations before attempting cross-document matching. This preliminary processing organizes data in advance, making the subsequent metadata linkage creation more efficient and manageable, thereby reducing overall system complexity while achieving complete concept information.
Solution Approach 2:
The system uses self-service mechanisms where annotations from one document automatically serve as evaluation criteria for matching concepts in other documents. This self-referential approach reduces the need for external complex processing logic, as the metadata system uses its own generated annotations to drive cross-document linkage creation.
Data Source
AI summary
Embodiments include cross-document propagation of entity metadata. Aspects include identifying a set of documents from a plurality of documents, the set of documents being related to one another and identifying a concept in a first document of the set of documents and creating an annotation corresponding to the concept. Aspects also include evaluating the annotation from the first document against all of the documents in the set of documents and identifying a concept match between the annotation and a mention discovered in a second document in the set of documents. Aspects further include creating a metadata linkage between the concept in the first document to the mention in the second document.


