Document Provenance Analysis Using Content Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Determining the influence of documents on the creation of other documents, known as the provenance question, is challenging due to the lack of associated tracking information, especially for legacy documents, and existing methods require manual association of tracking data.
Innovation Solution
An algorithmic approach using similarity measurements between document contents, without requiring tracking information, to determine the subset of documents from which a particular document was derived, employing forward-in-time and reversed-in-time graphs, Markov chains, and matrix stabilization to calculate derivation probabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual association of tracking information is used to determine document provenance, then the accuracy of provenance determination is improved, but the productivity and automation level deteriorate
Solution Approach 1:
The system enables documents to self-identify their provenance by automatically comparing their content against the document universe using similarity algorithms. Documents independently determine their own derivation relationships without requiring manual tracking information association, thus achieving both high accuracy and full automation
Solution Approach 2:
The patent replaces the manual mechanical process of tracking information association with an automated computational system using similarity algorithms (e.g., cosine similarity, Jaccard similarity). This substitution eliminates manual labor while maintaining or improving provenance determination accuracy through algorithmic content analysis
2Reliability
If tracking information is required to be associated with documents, then the reliability of provenance data is improved, but the adaptability to legacy documents deteriorates
Solution Approach 1:
The patent introduces similarity algorithms as an intermediary mechanism that bridges the gap between documents with and without tracking information. By using content-based similarity comparison, the system can determine provenance relationships for legacy documents that lack tracking information, while still providing reliable results comparable to manual tracking association
Solution Approach 2:
The system changes the parameter used for provenance determination from relying on tracking information metadata to using content-based similarity metrics. This parameter change enables universal applicability across all document types including legacy documents, while maintaining reliability through sophisticated similarity measurement algorithms
3Ease of operation
If content-based similarity analysis is used without tracking information, then the ease of operation and automation are improved, but the measurement precision of provenance determination may deteriorate
Solution Approach 1:
The system implements feedback mechanisms through iterative similarity comparison and probability calculation. The Markov chain model uses feedback from multiple comparison rounds to refine provenance probability estimates, ensuring high measurement precision while maintaining full automation and ease of operation
Solution Approach 2:
The patent performs excessive similarity comparisons by analyzing documents against the entire document universe rather than just a selected subset. This exhaustive approach ensures that no potential source document is missed, thereby maintaining high measurement precision while the automated process keeps ease of operation intact
Data Source
AI summary
Embodiments of the present invention pertain to determining a subset of documents from which a particular document was derived. According to one embodiment, similarity measurements indicating similarities between contents of documents are received. A subset of the documents that the particular document was derived from is determined based on dates the documents were created and the similarity measurements without requiring document tracking information to be associated with the documents to determine the subset.


