Universal Document Similarity Measure via Semantic Structures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current document processing systems are limited in comparing documents across different languages and formats, failing to accurately detect similarity or difference, especially in cross-language plagiarism detection and handling diverse types of information within documents.

Innovation Solution

A method that estimates and visualizes similarity/difference between documents using exhaustive syntactic and semantic analyses, converting textual parts into language-independent semantic structures, and comparing various types of information such as text, images, and video, employing weights for different content types, and providing visualization techniques like highlighting and underlining.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If lexical-based similarity measures are used for document comparison, then the computation is simple and fast, but the accuracy is insufficient for cross-language and cross-format document similarity detection

Engineering Contradiction:
Improvedocument similarity detection accuracyVSAvoidcomparison system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments document comparison into multiple independent analysis dimensions: lexical analysis, syntactic analysis, semantic analysis, and visual analysis. Each dimension processes specific features separately before integrating results, allowing the system to achieve high accuracy through comprehensive multi-dimensional comparison while managing complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal comparison framework that handles multiple document types (text, images, video) and multiple languages through a single integrated system. The system uses language-independent semantic structures and universal feature extraction mechanisms to compare diverse document formats, eliminating the need for separate specialized systems for each document type

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If comprehensive syntactic and semantic analysis is performed on documents, then the similarity estimation accuracy is improved, but the processing time increases

Engineering Contradiction:
Improvesimilarity estimation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing steps including document segmentation, feature extraction, and semantic structure construction before the actual similarity comparison. By pre-processing and indexing documents in advance, the system can quickly retrieve and compare only the relevant features during similarity estimation, significantly reducing real-time processing time while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial analysis by focusing computational resources on the most discriminative features for each document pair. The system dynamically adjusts the depth of syntactic and semantic analysis based on document characteristics and comparison requirements, performing exhaustive analysis only when necessary and using lighter analysis for preliminary filtering

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If language-independent semantic structures are used for cross-language comparison, then cross-language plagiarism detection capability is improved, but the system complexity increases

Engineering Contradiction:
Improvecross-language comparison capabilityVSAvoidsemantic analysis system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces language-independent semantic structures as an intermediary layer between source language documents and the comparison system. This intermediary converts documents from various languages into a universal semantic representation space, enabling cross-language comparison without requiring direct translation or language-specific processing rules, thus simplifying the overall system architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms language-specific textual parameters into language-independent semantic parameters. By changing the representation parameters from surface-level lexical forms to deep-level semantic meanings, the system achieves cross-language comparability. The semantic analysis module extracts invariant semantic features that remain consistent across language boundaries, eliminating language-specific complexity

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If multiple types of information (text, images, video) are compared within documents, then the comprehensive similarity assessment is improved, but the system complexity and computational burden increase

Engineering Contradiction:
Improvecomprehensive similarity assessment accuracyVSAvoidmulti-format comparison system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple comparison tasks (text similarity, image similarity, video similarity) into a single unified framework. The system combines different types of feature vectors and similarity scores through weighted integration, producing a comprehensive document similarity assessment. This merging approach handles diverse information types through common processing mechanisms, reducing overall system complexity compared to separate specialized systems

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9235573B2Universal difference measure
Publication Date: 2016.01.12 ABBYY DEVELOPMENT INC
  • US9235573B2 patent drawing
  • US9235573B2 patent drawing
  • US9235573B2 patent drawing

AI summary

Described herein are methods for finding substantially similar/different sources (files and documents), and estimating similarity or difference between given sources. Similarity and difference may be found across a variety of formats. Sources may be in one or more languages such that similarity and difference may be found across any number and types of languages. A variety of characteristics may be used to arrive at an overall measure of similarity or difference including determining or identifying syntactic roles, semantic roles and semantic classes in reference to sources.