Document Provenance Analysis Using Content Similarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Determining the influence of documents on the creation of other documents, known as the provenance question, is challenging due to the lack of associated tracking information, especially for legacy documents, and existing methods require manual association of tracking data.

Innovation Solution

An algorithmic approach using similarity measurements between document contents, without requiring tracking information, to determine the subset of documents from which a particular document was derived, employing forward-in-time and reversed-in-time graphs, Markov chains, and matrix stabilization to calculate derivation probabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual association of tracking information is used to determine document provenance, then the accuracy of provenance determination is improved, but the productivity and automation level deteriorate

Engineering Contradiction:
Improveprovenance determination accuracyVSAvoiddocument provenance analysis efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables documents to self-identify their provenance by automatically comparing their content against the document universe using similarity algorithms. Documents independently determine their own derivation relationships without requiring manual tracking information association, thus achieving both high accuracy and full automation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of tracking information association with an automated computational system using similarity algorithms (e.g., cosine similarity, Jaccard similarity). This substitution eliminates manual labor while maintaining or improving provenance determination accuracy through algorithmic content analysis

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If tracking information is required to be associated with documents, then the reliability of provenance data is improved, but the adaptability to legacy documents deteriorates

Engineering Contradiction:
Improveprovenance data reliabilityVSAvoidlegacy document compatibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces similarity algorithms as an intermediary mechanism that bridges the gap between documents with and without tracking information. By using content-based similarity comparison, the system can determine provenance relationships for legacy documents that lack tracking information, while still providing reliable results comparable to manual tracking association

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter used for provenance determination from relying on tracking information metadata to using content-based similarity metrics. This parameter change enables universal applicability across all document types including legacy documents, while maintaining reliability through sophisticated similarity measurement algorithms

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If content-based similarity analysis is used without tracking information, then the ease of operation and automation are improved, but the measurement precision of provenance determination may deteriorate

Engineering Contradiction:
Improveautomated provenance analysis easeVSAvoidprovenance determination accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system implements feedback mechanisms through iterative similarity comparison and probability calculation. The Markov chain model uses feedback from multiple comparison rounds to refine provenance probability estimates, ensuring high measurement precision while maintaining full automation and ease of operation

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs excessive similarity comparisons by analyzing documents against the entire document universe rather than just a selected subset. This exhaustive approach ensures that no potential source document is missed, thereby maintaining high measurement precision while the automated process keeps ease of operation intact

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8793264B2Determining a subset of documents from which a particular document was derived
Publication Date: 2014.07.29 HEWLETT PACKARD ENTERPRISE DEV LP
  • US8793264B2 patent drawing
  • US8793264B2 patent drawing
  • US8793264B2 patent drawing

AI summary

Embodiments of the present invention pertain to determining a subset of documents from which a particular document was derived. According to one embodiment, similarity measurements indicating similarities between contents of documents are received. A subset of the documents that the particular document was derived from is determined based on dates the documents were created and the similarity measurements without requiring document tracking information to be associated with the documents to determine the subset.