Partial Document Similarity Analysis for Precise Relationship Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies for associating document data lack precision in identifying relationships between partial documents due to similarity calculations based on entire document data.
Innovation Solution
An information processing apparatus and method that divides document data into partial documents, derives partial document relationship information based on similarity and date information, and then generates overall document relationship information using partial document relationships.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If similarity calculation is performed on entire document data, then the calculation process is simple, but the precision of association between document data deteriorates
Solution Approach 1:
The patent divides document data into multiple partial documents (e.g., paragraphs, sections, or sentences) and performs similarity calculations on these segmented units rather than on entire documents. This segmentation enables more precise matching by comparing specific corresponding portions of documents, thereby improving association precision while maintaining manageable calculation complexity through focused local comparisons.
2Measurement precision
If document data is divided into partial documents for similarity calculation, then the precision of association improves, but the device complexity increases
Solution Approach 1:
The patent segments document data into partial documents and processes them in a structured manner: extracting partial documents from multiple source documents, calculating similarity matrices for these partial documents, and then integrating results to determine overall document associations. This systematic segmentation approach improves precision while controlling complexity through organized processing steps.
Solution Approach 2:
The patent combines multiple partial document similarity calculations into an overall document association result by integrating the partial similarity matrices. This merging process synthesizes the precise local comparisons into a comprehensive view of document relationships, achieving high precision association without requiring entirely new processing frameworks.
3Loss of time
If entire document data is used for similarity calculation, then processing time is reduced, but the accuracy of identifying document relationships deteriorates
Solution Approach 1:
The patent divides documents into partial documents and calculates similarity for these smaller units, which can be processed more efficiently than entire documents. The segmented approach allows for faster computation of individual partial document similarities while maintaining or improving relationship identification accuracy through focused comparisons of relevant content portions.
Solution Approach 2:
The patent performs similarity calculations on selected partial documents rather than processing entire documents in full. By focusing computational resources on specific partial document comparisons that are most relevant to identifying relationships, the system achieves accurate relationship identification with reduced processing time compared to exhaustive entire-document analysis.
Data Source
AI summary
An information processing apparatus acquires a plurality of pieces of document data, each including a plurality of partial documents, and date information associated with each of the plurality of pieces of document data, derives partial document relationship information representing a relationship between the partial documents, based on a similarity between the partial documents and the date information, among the plurality of pieces of document data, and derives overall document relationship information representing an overall relationship among the plurality of pieces of document data, based on the partial document relationship information.


