Universal Document Similarity Measure via Semantic Structures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current document processing systems are limited in comparing documents across different languages and formats, failing to accurately detect similarity or difference, especially in cross-language plagiarism detection and handling diverse types of information within documents.
Innovation Solution
A method that estimates and visualizes similarity/difference between documents using exhaustive syntactic and semantic analyses, converting textual parts into language-independent semantic structures, and comparing various types of information such as text, images, and video, employing weights for different content types, and providing visualization techniques like highlighting and underlining.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If lexical-based similarity measures are used for document comparison, then the computation is simple and fast, but the accuracy is insufficient for cross-language and cross-format document similarity detection
Solution Approach 1:
The patent segments document comparison into multiple independent analysis dimensions: lexical analysis, syntactic analysis, semantic analysis, and visual analysis. Each dimension processes specific features separately before integrating results, allowing the system to achieve high accuracy through comprehensive multi-dimensional comparison while managing complexity through modular architecture
Solution Approach 2:
The patent creates a universal comparison framework that handles multiple document types (text, images, video) and multiple languages through a single integrated system. The system uses language-independent semantic structures and universal feature extraction mechanisms to compare diverse document formats, eliminating the need for separate specialized systems for each document type
2Measurement precision
If comprehensive syntactic and semantic analysis is performed on documents, then the similarity estimation accuracy is improved, but the processing time increases
Solution Approach 1:
The patent performs preliminary processing steps including document segmentation, feature extraction, and semantic structure construction before the actual similarity comparison. By pre-processing and indexing documents in advance, the system can quickly retrieve and compare only the relevant features during similarity estimation, significantly reducing real-time processing time while maintaining high accuracy
Solution Approach 2:
The patent applies partial analysis by focusing computational resources on the most discriminative features for each document pair. The system dynamically adjusts the depth of syntactic and semantic analysis based on document characteristics and comparison requirements, performing exhaustive analysis only when necessary and using lighter analysis for preliminary filtering
3Adaptability or versatility
If language-independent semantic structures are used for cross-language comparison, then cross-language plagiarism detection capability is improved, but the system complexity increases
Solution Approach 1:
The patent introduces language-independent semantic structures as an intermediary layer between source language documents and the comparison system. This intermediary converts documents from various languages into a universal semantic representation space, enabling cross-language comparison without requiring direct translation or language-specific processing rules, thus simplifying the overall system architecture
Solution Approach 2:
The patent transforms language-specific textual parameters into language-independent semantic parameters. By changing the representation parameters from surface-level lexical forms to deep-level semantic meanings, the system achieves cross-language comparability. The semantic analysis module extracts invariant semantic features that remain consistent across language boundaries, eliminating language-specific complexity
4Measurement precision
If multiple types of information (text, images, video) are compared within documents, then the comprehensive similarity assessment is improved, but the system complexity and computational burden increase
Solution Approach 1:
The patent merges multiple comparison tasks (text similarity, image similarity, video similarity) into a single unified framework. The system combines different types of feature vectors and similarity scores through weighted integration, producing a comprehensive document similarity assessment. This merging approach handles diverse information types through common processing mechanisms, reducing overall system complexity compared to separate specialized systems
Data Source
AI summary
Described herein are methods for finding substantially similar/different sources (files and documents), and estimating similarity or difference between given sources. Similarity and difference may be found across a variety of formats. Sources may be in one or more languages such that similarity and difference may be found across any number and types of languages. A variety of characteristics may be used to arrive at an overall measure of similarity or difference including determining or identifying syntactic roles, semantic roles and semantic classes in reference to sources.


